Our mission is to flag potential BS in your health stack. Scientifically, that means four things: missing upstream causes, overlooked harms, weak evidence, and unexplored alternatives.
For $5 you get a report on how your health stack lines up with science: a more complete bird’s-eye view of the clinical evidence than a rushed doctor visit or hours of ChatGPT research can provide.
We flag 4 ways your health stack may not be lining up with science:
| Blind spot | Why it matters | What’s yours to decide | Example |
|---|---|---|---|
| Upstream causes(false negative) nobody investigated what may be driving the symptom | You can treat the symptom for years while the thing causing it goes unmeasured. | Whether to trade short-term convenience for a longer-term fix. | Low mood treated with an SSRI — bipolar never screened for, exercise never tried. |
| Downstream side effects(false negative) trade-offs and harms the treatment itself creates | The fix quietly becomes the next problem you have to treat. | Which side-effects you can live with, and which you can't. | Semaglutide started for weight loss, with no resistance training to hold back muscle loss. |
| Iffy evidence(false positive) a contested or weakly supported claim treated as settled | You are betting on a claim that is weaker than it sounded. | Whether a weak-evidence bet is one you want to take. | Daily aspirin after 70 framed as a balance of risks, when guidelines recommend against starting it. |
| Alternative treatments(false negative) narrowed to one path before the others were weighed | An option that fits your situation better was never on the table. | How you weigh efficacy, durability, convenience and side-effects against each other. | CPAP prescribed for sleep apnea, with an oral appliance never weighed against it. |
In evaluation terms the health stack is what is under test, so those labels name ITS error class and not ours: three of the four are false negatives — something real that nobody looked at — and one is a false positive, asserted more firmly than the evidence carries. Applicability is a qualifier rather than a fifth flag: a finding can be true in general and still unknown for you until a test or your own particulars settle it, which is why every flag hands a judgement back instead of making it.
Examples are illustrations of each pattern, not real cases.
The report comes with the underlying data as a machine-readable file, not just a PDF. Upload it into ChatGPT, Claude or whatever you already use and keep digging — the studies, the concept ids and every relationship we drew are all in there.
NoBSmed is for people with complex, nuanced, non-emergency health problems who want to do deeper research before an appointment.
NoBSmed takes your current health story — what’s going on, what you’re taking or considering, relevant labs, what you’ve been told, and what you want to improve — and builds the causal story behind your particular problem. Then it compares that picture against findings from clinical studies.
More technically, our product is a causal clinical-evidence diff engine for evaluating the completeness of a health plan or regimen in non-urgent care.
A NoBSmed user wanted a GLP-1 weight-loss plan. His doctor generated one with OpenEvidence. He pasted that into ChatGPT, which added what OpenEvidence had left out. Then he ran the whole thing through NoBSmed.
We ran GPT-5.2 and Claude Opus 4.8 over 951 clinical conversations from OpenAI’s HealthBench — filtered to the ones where evidence could change a high-stakes decision — and graded every answer against its doctor-written rubric.
When these models get something wrong, roughly 9 times in 10 it is something they left out — not something they got wrong.
The direct answer is usually strong. What goes missing is the clinical context around it:
| What the model left out | GPT-5.2 | Opus 4.8 |
|---|---|---|
| Didn’t surface an alternative diagnosis | 48% | 55% |
| Didn’t flag a drug’s side effects or monitoring | 15% | 18% |
Each figure is the share of the 951 conversations containing at least one instance of that pattern. Most are recall misses — the ideal answer would also have said X — not whole answers that are wrong or harmful. Graded by GPT-4.1, which is HealthBench’s own rubric grader — the same grading apparatus we audited separately.
That is why the four things above are the four things we look for. Three of them are omissions. See the full failure-pattern study →
OpenAI grades its flagship models on HealthBench, its own medical benchmark. We audited the benchmark itself — 1,298 clinical claims across its gold answers and rubrics — and found 29 that would change a decision, including fabricated citations in the reference answers.
Errors in an answer key propagate into every model graded against it. Read the audit →
We don’t tell you what to do. We don’t diagnose, prescribe, or tell you to start, stop, or change a treatment. We surface what clinical studies actually said about people in situations like yours — and where the evidence runs out — so you and your clinician can decide. This is not a substitute for clinical care.
Science writer and breast cancer survivor. Keeps us creative and grounded.
LinkedInNoBSmed retrieves and structures clinical-study evidence. It does not diagnose, prescribe, or replace professional medical judgment. Users should consult a qualified healthcare professional before making medical decisions.