AI health advice, graded
We asked AI to answer 2,800 user health questions and graded them with real doctors.
Overall
| Model | Score (% info correct) | Graded |
|---|---|---|
| 57.7 | 2800 | |
| 52.0 | 2783 | |
| 39.6 | 2799 |
By specialty
Score = mean HealthBench rubric score (0–100). Bold = best model for that specialty. Category rows use the 2782 items graded across all models. Each model answered at medium reasoning; every answer is graded against physician-written rubrics by gpt-5.4-mini.
Questions and rubrics are from the HealthBench benchmark (arxiv.org/pdf/2505.08775). Some questions were based on real Google search queries, some were physician-written, and some were AI-generated based on physician-provided guidelines. We selected only English, consumer-posed questions from the original 5,000 HealthBench questions. In total, 262 doctors wrote rubrics to grade the questions.
* Length-adjusted scores reward conciseness comparable to human doctors, preventing an AI from gaming the score by just giving longer and longer answers (to have higher chance of saying something our doctor-written rubric likes).
→ See scores broken down by task type (diagnosis, treatment, triage…)
Search a symptom or issue
Relevant AI answers will show here.