《電子技術應用》
您所在的位置:首頁 > 其他 > 设计应用 > 基于增强语义信息理解的场景图生成
基于增强语义信息理解的场景图生成
2023年电子技术应用第5期
曾军英,陈运雄,秦传波,陈宇聪,王迎波,田慧明,顾亚谨
(五邑大学 智能制造学部,广东 江门 529020)
摘要: 场景图生成(SGG)任务旨在检测图像中的视觉关系三元组,即主语、谓语、宾语,为场景理解提供结构视觉布局。然而,现有的场景图生成方法忽略了预测的谓词频率高但却无信息性的问题,从而阻碍了该领域进步。为了解决上述问题,提出一种基于增强语义信息理解的场景图生成算法。整个模型由特征提取模块、图像裁剪模块、语义转化模块、拓展信息谓词模块四部分组成。特征提取模块和图像裁剪模块负责提取视觉特征并使其具有全局性和多样性。语义转化模块负责将谓词之间的语义关系从常见的预测中恢复信息预测。拓展信息谓词模块负责扩展信息谓词的采样空间。在数据集VG和VG-MSDN上与其他方法进行比较,平均召回率分别达到59.5%和40.9%。该算法可改善预测出来的谓词信息性不足问题,进而提升场景图生成算法的性能。
中圖分類號:TP391
文獻標志碼:A
DOI: 10.16157/j.issn.0258-7998.223276
中文引用格式: 曾軍英,陳運雄,秦傳波,等. 基于增強語義信息理解的場景圖生成[J]. 電子技術應用,2023,49(5):52-56.
英文引用格式: Zeng Junying,Chen Yunxiong,Qin Chuanbo,et al. Scene graph generation based on enhanced semantic information understanding[J]. Application of Electronic Technique,2023,49(5):52-56.
Scene graph generation based on enhanced semantic information understanding
Zeng Junying,Chen Yunxiong,Qin Chuanbo,Chen Yucong,Wang Yingbo,Tian Huiming,Gu Yajin
(Department of Intelligent Manufacturing, Wuyi University, Jiangmen 529020,China)
Abstract: The Scene Graph Generation (SGG) task aims to detect visual relation triples in images, i.e. subject, predicate and object, to provide a structural visual layout for scene understanding. However, existing approaches to scene graph generation ignore the high frequency but uninformative problem of predicted predicates, hindering progress in this field. In order to solve the above problems, this paper proposes a scene graph generation algorithm based on enhanced semantic information understanding. The whole model consists of four parts: feature extraction module, image cropping module, semantic transformation module and extended information predicate module. Feature extraction module and image cropping module are responsible for extracting visual features and making them global and diverse. The semantic transformation module is responsible for restoring the semantic relationship between predicates from common predictions to informative predictions. The extended information predicate module is responsible for extending the sampling space of the information predicate. Comparing with other methods on datasets VG and VG-MSDN, the average recall reaches 59.5% and 40.9%, respectively. The algorithm in this paper can improve the problem of insufficient information of the predicted predicate, and then improve the performance of the scene graph generation algorithm.
Key words : scene graph generation;image cropping;semantic transformation;extended information

0 引言

場景圖生成 (SGG) 任務的目標是從給定圖像生成圖結構表示,以抽象出對象(以邊界框為基礎)及其成對關系。場景圖旨在促進對圖像中復雜場景的理解,并具有廣泛的下游應用潛力,例如圖像檢索、視覺推理、視覺問答(VQA)、圖像字幕、結構化圖像生成和外繪和機器人技術。好的場景圖可以在感興趣的實例之間提供信息豐富的關系?,F有的場景圖生成大多遵循通用的范式,即從圖像中檢測目標,提取區域特征,然后在標準分類目標函數的指導下識別謂詞類別。但是,這種范式有幾方面的缺點。


本文詳細內容請下載:http://www.tom3567.com/resource/share/2000005313




作者信息:

曾軍英,陳運雄,秦傳波,陳宇聰,王迎波,田慧明,顧亞謹

(五邑大學 智能制造學部,廣東 江門 529020)


微信圖片_20210517164139.jpg

此內容為AET網站原創,未經授權禁止轉載。
主站蜘蛛池模板: 国产精品久久久久久久久久久久| 中文字幕av日韩精品| 精品激情国产视频| 欧美在线中文字幕| 欧美日韩免费观看一区| 日韩中文字幕网站| 亚洲欧美日韩精品久久久| 国产福利精品视频| 日韩资源av在线| 久久手机精品视频| 日韩手机在线观看视频| 久久久国产在线视频| 久久99精品久久久久久久青青日本| 国产精品久久久久av| 欧美亚洲激情在线| 国产av不卡一区二区| 国产www精品| 97精品久久久| 人妻少妇精品无码专区二区| 国产日韩欧美黄色| 国产欧美日韩精品专区| 99精品国产高清在线观看| 亚洲91精品在线亚洲91精品在线| 日韩人妻一区二区三区蜜桃视频| 欧美成在线观看| 国产不卡一区二区在线播放| 久久九九国产精品怡红院| 亚洲一区二区免费| 久久久久亚洲av无码专区喷水| 久久精品国产精品亚洲色婷婷| 欧美亚洲精品日韩| 少妇人妻无码专区视频| 欧美视频在线第一页| 久久6免费高清热精品| 一区二区三区四区欧美日韩| 久久亚洲成人精品| 大波视频国产精品久久| 日韩中文字幕网址| 视频一区二区三区在线观看| 国产男女激情视频| 久久久久免费精品|