Blog

What medical AI leaves out

/about says we map four kinds of alternative. This is the measurement behind that choice: when a medical answer fails, it usually fails by omission — and an omission is invisible unless something goes looking for it.

What the models actually miss

We ran GPT-5.2 and Claude Opus 4.8 over 951 clinical conversations from OpenAI’s HealthBench — filtered to the ones where evidence could change a high-stakes decision — and graded every answer against its doctor-written rubric.

When these models get something wrong, roughly 9 times in 10 it is something they left out — not something they got wrong.

The direct answer is usually strong. What goes missing is the clinical context around it:

What the model left out GPT-5.2 Opus 4.8
Didn’t surface an alternative diagnosis 48% 55%
Didn’t flag a drug’s side effects or monitoring 15% 18%

Each figure is the share of the 951 conversations containing at least one instance of that pattern. Most are recall misses — the ideal answer would also have said X — not whole answers that are wrong or harmful. Graded by GPT-4.1, which is HealthBench’s own rubric grader — the same grading apparatus we audited separately.

See the full failure-pattern study →

We are not the first to find this. NOHARM (2025) reports that 80%+ of severe medical-AI errors were omissions (arXiv:2512.01241), npj Digital Medicine reports the same shape, and missed and delayed diagnoses are a long-standing patient-safety problem in human medicine (systematic review).


One case, three tools

A NoBSmed user wanted a GLP-1 weight-loss plan. His doctor generated one with OpenEvidence. He pasted that into ChatGPT, which added what OpenEvidence had left out. Then he ran the whole thing through NoBSmed.

A causal graph in three columns. The doctor's OpenEvidence plan covers semaglutide, appetite and weight. ChatGPT adds titration, resistance training to protect muscle, stopping the drug, and weight regain. NoBSmed adds possible sleep apnea never screened, glycemic load as a possible driver, and insulin resistance never measured — each linked to weight regain and to the focus and energy the patient actually wanted.
One de-identified case, shared with permission (n = 1). The first two columns were drawn by us from what OpenEvidence and ChatGPT reported; the third is the NoBSmed evidence graph. See the full evidence graph for this case →

We audit the benchmarks too

OpenAI grades its flagship models on HealthBench, its own medical benchmark. We audited the benchmark itself — 1,298 clinical claims across its gold answers and rubrics — and found 29 that would change a decision, including fabricated citations in the reference answers. Errors in an answer key propagate into every model graded against it. Read the audit →

← Back to what we do about it