Adaptive Confidence-Guided Multi-Scale Vision Transformer Framework for Explainable Microscopic Fungal Species Classification
编号:96
访问权限:仅限参会人
更新:2026-07-25 18:00:37
浏览:15次
Online
摘要
Pathogenic fungi are difficult to identify in the field. The task of reading microscopic slides is still primarily one that requires training from human eyes — a point of delay in clinical
decisions and an introduction of inter-observer error. This paper proposes a new confidence-guided multi-scale fusion framework called ACMF. The framework runs two Vision Transformer (ViT-B/16) encoders at different input resolutions — 224 and
384 pixels — and merges the outputs in four different ways. These range from simple averaging (S1) and accuracy-weighted blending (S2) to a per-image entropy-based weighting scheme (S3) and an expert MLP fusion head (S4). Experiments on the
DeFungi benchmark — five clinically relevant fungal species, 6,801 microscopic images — demonstrate competitive accuracy. The best-performing variant is S2, with a macro F1 of 91.63% on the held-out test set — a 3.16-point gain over ViT-B/16-224 used alone. Attention Rollout maps are generated at inference time, highlighting hyphae and spore structures without additional annotation. Comparisons with ResNet-50, EfficientNet-B3, and Swin-Tiny confirm that ACMF delivers the most consistent value.
关键词
fungal classification,vision transformer,deep learning,explainable AI,multi-scale fusion,DeFungi,microscopy
稿件作者
Rishi Jayanath A
Karunya Institute of Technology and Sciences
G Naveen Sundar
Karunya Institute of Technology and Sciences
发表评论