Condition: Emergency Medicine · Diagnostic Errors · Artificial Intelligence (AI) in Diagnosis · Sponsor: Marmara University Pendik Training and Research Hospital
This retrospective diagnostic accuracy study evaluates the ability of two large language models (LLMs) - GPT-4o (gpt-4o-2024-11-20; OpenAI) and Claude 4.6 Sonnet (claude-sonnet-4-6; Anthropic) - to generate correct diagnoses from anonymized Turkish-language emergency department (ED) anamnesis notes, and compares their performance with the diagnosis entered by the treating emergency physician. A consensus gold standard is established by three independent board-certified emergency medicine specialists who blindly review each note and vote on the primary diagnosis using ICD-10 three-character codes; the majority vote (at least 2 of 3 specialists agreeing) constitutes the reference standard. Both LLMs are evaluated using a standardized zero-shot direct prompting strategy (temperature=0, stateless API sessions). The primary outcome is diagnostic accuracy (proportion of ICD-10 chapter-level matches) and Cohen's kappa for each LLM against the gold standard. Secondary outcomes include top-3 accuracy, treating physician accuracy, inter-model agreement, and subgroup analyses by ESI triage level and ICD-10 chapter. Inter-rater reliability among the three specialists is quantified using Fleiss' kappa. Analyses are performed in Jamovi. This study represents the first evaluation of LLM diagnostic accuracy using Turkish-language clinical notes and the first to benchmark LLM performance against an independent three-specialist majority-vote gold standard rather than against the treating physi…
This description comes directly from the study's public registry record.
Emir Ünal, Assistant Professor · +905327766010 · emirunal@gmail.com
Emir Unal, Assistant Professor · emirunal@gmail.com
Always discuss trial participation with your own doctor first.
| Marmara University Pendik Training and Research Hospital | Istanbul, Istanbul, Turkey (Türkiye) | Recruiting |
Get one email when the public record changes — results posted, or the study's status changes. Nothing else, ever.
We email about this public record only. Unsubscribe anytime with one click. Never medical advice.
This page is independently generated by Eichor from the public ClinicalTrials.gov record and re-synced daily. It is not the sponsor's official website unless claimed. Nothing here is medical advice; eligibility is always determined by the study team — talk to your own doctor first.
Source record: clinicaltrials.gov/study/NCT07632859