Grounded Semantic Quality and Diversity Diagnostics for Image Caption Evaluation
编号:106
访问权限:仅限参会人
更新:2026-07-22 16:10:01 浏览:15次
Online
摘要
Image caption generation has advanced significantly with vision-language models which generate fluent, meaningful description for the Image. The conventional image captioning metrics like BLEU, METEOR, ROUGE and CIDEr mainly evaluate based on lexical or consensus-based similarity with the ground truth captions and provide limited explanation about the semantic grounding of the generated caption. To address this problem, Grounded Semantic Quality and Diversity diagnostic framework is proposed which not only evaluates the semantic grounding but also decompose them in categories to understand the reason of weak semantic grounding. To evaluate candidate set caption, Grounded Semantic Diversity metric is introduced. In addition to evaluation, the experiments to use the metrics for reranking and optimization are demonstrated to show its effectiveness for the purpose. Experimental results show that when GSQ is incorporated as a reward component in self-critical sequence training, the fine-tuned model improves CIDEr by 19.3% and object-F1 by 10.6%, while reducing the unsupported semantic rate by 19.6% compared with the BLIP baseline.
关键词
Image captioning Metrics,Knowledge Graph,Reranking,,Entity awareness,Hallucination mitigation,Context awareness,Diversity in image captioning
稿件作者
Sharmila Kharat
MIT Academy of Engineering Alandi Pune
Sunita Barve
MIT Academy of Engineering Alandi Pune
发表评论