2 min readfrom Machine Learning

What Sante's 83.83 on DiagnosisArena-MCQ actually measures [D]

What Sante's 83.83 on DiagnosisArena-MCQ actually measures [D]
What Sante's 83.83 on DiagnosisArena-MCQ actually measures [D]

https://preview.redd.it/xv2epabu6ioh1.png?width=1171&format=png&auto=webp&s=4e22c69855ec509bd63a038a24f765caeaafa59c

Ant Ling reports 83.83 on DiagnosisArena-MCQ for Ling-3.0-flash-Sante, its new medical reasoning model. The suffix matters: the task provides case information, examinations and tests, then asks the model to choose from four diagnoses.

That result tells us about selecting an answer when the candidate set and case evidence are supplied. It does not establish how the same model would generate an unrestricted differential, decide what history is missing, or choose which investigation to request next. Those would require different evaluations.

The release also reports two other medical results:

Evaluation Sante result What the task adds
MedXpertQA-Text 53.88 Challenging medical questions in a text subset.
HealthBench Professional 45.73 Open-ended professional clinical chat, assessed with physician-written rubrics.

The published HealthBench Professional definition includes care consultation, writing/documentation and medical research. Its score is not percentage accuracy. The Sante chart does not provide enough scoring detail to identify the reported value as length-adjusted or unadjusted, so a comparison with another published HBP result would need that checked first.

This is why the three results are useful together. They give Sante a broader medical-text evaluation profile than an exam score alone, while leaving specific questions open. For a case-answering application, the first decision is whether users supply the alternatives or expect the model to construct them. The release supports including Sante in that evaluation; the 83.83 figure applies to the supplied-options version.

submitted by /u/Expert_Coffee_203
[link] [comments]

Want to read more?

Check out the full article on the original site

View original article

Tagged with

#Medical Reasoning
#DiagnosisArena-MCQ
#Sante
#Ling-3.0-flash
#Medical Model
#Diagnosis
#Case Information
#Case Evidence
#Differential Diagnosis
#Medical Investigations
#MedXpertQA-Text
#HealthBench Professional
#Clinical Chat
#Physician Rubrics
#Care Consultation
#Medical Research
#Evaluation Metrics
#Machine Learning
#Medical Text
#Candidate Set