Vision-language modeling has significantly advanced radiology by enabling models that jointly learn from medical images and radiology reports for tasks such as disease classification, report generation, and visual question answering. However, most existing approaches treat an entire medical image as a single entity during image-text alignment, overlooking the fine-grained anatomical reasoning process employed by radiologists. In clinical practice, radiologists systematically examine individual anatomical regions, associate findings with specific structures, and integrate these region-specific observations before arriving at a conclusion about the image as a whole. To address this limitation, we propose an anatomy-aware vision-language framework that learns anatomy-specific representations using dedicated anatomy tokens and anatomical segmentation masks. The framework further incorporates context-aware anatomical representations and jointly learns anatomical localization, aligns anatomical regions with their corresponding findings, and aligns global image representations with image-level disease categories within a unified vision-language framework. Extensive experiments on out-of-distribution datasets demonstrate the effectiveness of the proposed framework. The model achieves strong performance in zero-shot disease classification and anatomical segmentation, demonstrating robust generalization to unseen data and accurate localization of anatomical structures. Comprehensive ablation studies further validate the contribution of each component and the effectiveness of the proposed design choices.
07月30日
2026
08月01日
2026
注册截止日期
初稿截稿日期
2026年07月30日 印度 Trichy
2026 International Conference on Networks Computers and Communications
发表评论