- Article
Evaluating the Performance of Large Language Models in Evidence-Scarce Scenario: The Diabetic Foot Ulcer Transition Phase
- Kamran Shakir,
- Hadi Sarlak and
- Paolo Caravaggi
- + 3 authors
Diabetic foot ulcers (DFUs) impose a substantial burden on people with diabetes and healthcare systems. The post-healing “transition phase” remains clinically challenging with limited guideline support. While large language models (LLMs) are increasingly proposed as clinical decision-support tools, their reliability in evidence-scarce scenarios is largely untested. This exploratory study benchmarked leading LLMs against European clinician consensus for the evidence-scarce scenario of DFU transition-phase management. Six LLMs (ChatGPT-4o, ChatGPT-5.0, Gemini 2.5 Flash, Gemini 2.5 Pro, Claude Sonnet 4.0, and Perplexity) were evaluated for accuracy and hallucination using a two-stage framework. Benchmarks were derived from an online survey of European DFU experts reflecting European-level and national practices (Denmark, Netherlands, UK). Binary outcomes were summarized as proportions with Wilson 95% confidence intervals. Paired within-item comparisons across models were assessed using Cochran’s Q, followed by exact McNemar tests with Holm correction (α = 0.05); inferential results were considered supportive due to the limited number of paired items (n = 8). Across 192 accuracy assessments and 384 reference checks, LLM accuracy ranged from 50–75% and declined when simulating national practices. Hallucination rates exceeded 50% in several models. LLMs rely on generic recommendations, which contrast with clinicians’ contextual, patient-centered reasoning, suggesting current limitations in their suitability for clinical decision support in DFU transition phase clinical care.
Informatics
20 July 2026








