Quality Over Quantity: Image Curation for Multimodal Fake News Detection
编号:37
访问权限:仅限参会人
更新:2026-07-27 20:05:24
浏览:16次
Online
摘要
Most multimodal fake news detection research assumes that more training images lead to better performance. We test this assumption directly. Using a Light Cross-Modal Attention model that fuses frozen DeBERTa-v3-base and CLIP ViT-B/32 embeddings through a 247K-parameter attention module, we compare detection accuracy across three image collection strategies: generic web scraping via Bing (679 images), curated article-specific images from FakeNewsNet (60 images), and manually collected images from 15 fact-checking organizations (150 images). In a size-controlled comparison (150 vs. 150), curated images outperform generic ones by 14.7 percentage points (93.3% vs. 78.6%, 5-fold CV). The 150 curated images also beat all 679 generic images by 4.5 points. A second experiment on mixed-source data reveals that multimodal fusion degrades by only 1.4% when data sources are combined, while text-only and image-only models drop by 19.6% and 13.5% respectively. These results point to image-text relevance as a more important factor than dataset size in multimodal fake news detection.
关键词
fake news detection;multimodal learning;cross-modal attention;image quality;image curation;CLIP;DeBERTa
稿件作者
manish rai
Manipal university jaipur
Amrit Raj
Manipal university jaipur
发表评论