
Researchers at IIIT-Hyderabad evaluated four vision-language models against thousands of chest X-rays and found discrepancies between AI-highlighted regions and radiologist assessments.
The study authors questioned whether visual heatmaps accurately reflect a model's ability to locate disease independently rather than relying on pre-established diagnostic conclusions.
The findings highlight the importance of clinical evaluations and independent medical assessments rather than complete reliance on AI-generated outputs.