An autonomous multimodal AI agent for evidence-grounded ophthalmic diagnosis

Cell Reports Medicine · Published 2026-08-01 · DOI 10.1016/j.xcrm.2026.102969

Free full text

Authors (10)

Kaikai Zhao, Qixuan Sun, Daohuan Kang, Tao Yu, Wenzheng Han, Rui Yao, Rupesh Agrawal, Gui-shuang Ying, Andrzej Grzybowski, Kai Jin

Abstract

Summary: Multimodal ophthalmic diagnosis requires integrating fundus photography, B-scan ultrasonography, and medical evidence, yet most artificial intelligence (AI) systems remain single-task or weakly grounded. AgentEYE is an auditable multimodal agent that routes ocular images to specialized fundus and B-scan tools, retrieves guideline/web evidence, and synthesizes evidence-grounded reports. In a 302-case internal benchmark, AgentEYE shows higher diagnostic correctness and completeness than large language model (LLM)-only baselines and an ablation without specialized imaging tools; performance remains similar to the no-retrieval ablation, indicating that retrieval mainly supports evidence grounding and citation auditability. Blinded evaluation of 200 cases by three ophthalmologists confirms improved diagnostic correctness, completeness, safety, and citation grounding versus an LLM-only self-citation baseline. External analyses show distribution-dependent performance. These findings support AgentEYE as a traceable decision-support prototype requiring prospective multicenter validation.

Abstract from DOAJ. Public domain (CC0 1.0).

Read the article at the publisher →

Publication details

Year
2026

Citation

Zhao, K., Sun, Q., Kang, D., et al. (2026). An autonomous multimodal AI agent for evidence-grounded ophthalmic diagnosis. Cell Reports Medicine. https://doi.org/10.1016/j.xcrm.2026.102969

Related articles