Label anchored virtual consultant: Literature guided multimodal prototype learning for retinal disease classification from unpaired data

Document Type

Article

Source of Publication

Journal of Computational Science

Publication Date

11-1-2026

Abstract

Medical image classification models often struggle when labelled data are limited, disease classes are imbalanced, and paired image-text clinical data are unavailable. Medical literature contains useful disease-level knowledge, but most multimodal learning methods require matched image-text pairs, such as fundus images linked with clinical reports. These paired datasets are rarely available in routine medical settings. In this study, we propose a Virtual Consultant (VC) framework for retinal disease classification that integrates retinal fundus images with disease-specific medical literature without requiring paired image-text data. The framework introduces a label-anchored multimodal prototype architecture centred on a global class embedding table. An image encoder learns visual representations from fundus images, while a text encoder processes disease-labelled medical literature. Instead of enforcing sample-level image-text alignment, both modalities interact through class-level prototypes, and a cross-modal alignment objective encourages consistency between visual and textual disease representations. This design distinguishes VC from conventional image-only and paired multimodal approaches by enabling disease-labelled textual representations to regularize class-level prototype learning while retaining image-only inference. We evaluate VC on an eight-class retinal disease classification task comprising 2757 fundus images and 697 curated disease-labelled literature abstracts. Under the multi-seed evaluation protocol, VC achieved competitive full-data performance, with 91.22 ± 1.01% accuracy and 0.883 ± 0.018 macro-F1, when compared against ResNet50 image-only, ResNet50 prototype-based, DenseNet121, and EfficientNet-B0 baselines. In this study, VC provides a different methodological advantage by integrating unpaired literature through label-anchored prototype learning. In reduced-data experiments, VC showed its clearest advantage at the most data-limited setting, improving accuracy from 81.40% to 83.49% at 20% labelled training data compared with the image-only ResNet50 baseline. The prototype analysis further showed that same-class image and text prototypes converged during training, supporting the role of literature-guided class-level alignment as a semantic regularization mechanism. A single-seed alignment-weight sensitivity analysis under the 20% labelled-data setting showed that removing alignment reduced performance, while the highest tested alignment-weight produced the best numerical result in this exploratory run. Class-wise and confusion-pair analyses revealed clinically interpretable error patterns, although their clinical significance was not independently validated. In the exploratory single-seed literature-control experiment, disease-name-only text produced the highest numerical accuracy and macro-F1, but the conditions were not statistically compared; therefore, the added value of full abstract-level content remains unresolved. VC offers a practical approach to label-level semantic prototype regularization using independently collected disease-labelled text when paired image-text data are unavailable. However, external and prospective validation across independent datasets, imaging devices, institutions, and patient populations is required before broader clinical claims can be made.

ISSN

1877-7503

Publisher

Elsevier BV

Volume

101

Disciplines

Medicine and Health Sciences

Keywords

Label-Anchored Learning, Medical Image Analysis, Multimodal Learning, Prototype-Based Classification, Retinal Disease Classification, Semantic Regularization, Unpaired Multimodal Learning

Scopus ID

105048068853

Indexed in Scopus

yes

Open Access

no

Share

COinS