THE SIGNAL IN ONE SENTENCE

OpenEvidence has organized its new medical AI models around how quickly and deeply they search, while restricting its most advanced biological reasoning system.

01

WHAT ACTUALLY CHANGED

OpenEvidence released three production models for verified clinicians. Osler is the default and aims to answer in roughly five seconds. Sackett spends about thirty seconds on questions that require deeper evidence review. Snow, the successor to Deep Consult, can spend roughly five minutes investigating medical literature before producing a report.

The company says all three are held to the same clinical-accuracy standard. Their intended difference is the amount of reasoning and evidence search applied to the question. They are rolling out on the web and mobile apps with unlimited use at no charge for verified clinicians.

A fourth model, Darwin, is available only through a research application. OpenEvidence says frontier reasoning in virology, immunology, and human genetics creates dual-use concerns, including work relevant to biological weapons or unsafe germline editing. Institutional partners and accredited researchers will help evaluate its safeguards before access expands.

OpenEvidence reports that Darwin answered all 660 questions correctly in its final physician-reviewed MedQA set. The original test split contained 1,273 questions. The company re-annotated the set, excluded questions for ambiguity or missing information, and conducted another physician review of questions missed by any evaluated model.

02

WHY THIS MATTERS

Most model menus make users decode brand names that reveal little about the practical choice. OpenEvidence is presenting something a clinician can understand immediately: how much time and search depth does this question deserve?

That distinction mirrors real clinical work. Looking up a routine reference during rounds is different from weighing contradictory evidence before a specialist decision. A five-minute answer is not automatically better, but the product now exposes the cost of deeper investigation instead of pretending every question belongs in the same queue.

Darwin creates a second boundary. The production models are separated by deliberation time, while the research model is separated by permission. Medical AI is becoming a stack of capability, evidence depth, identity checks, and access rules rather than one universal doctor-shaped chatbot.

FIG. 047MATCH THE SEARCH DEPTH TO THE QUESTION
1CLINICAL QUESTION→
2FAST REFERENCE→
3DEEPER SEARCH→
4LITERATURE REVIEW→
5CLINICIAN JUDGMENT
The model choice becomes a decision about evidence depth and time, with the final clinical interpretation remaining in human hands.

03

WHERE IT COULD HELP

  • Check routine medical evidence during clinical work
  • Compare conflicting studies before a treatment discussion
  • Prepare deeper case reviews for specialist consultation
  • Route sensitive biological research through controlled access

KEEP A HAND ON THE WHEEL

The reported scores come from OpenEvidence, including a filtered and physician-reviewed benchmark set. They do not demonstrate patient outcomes or clinical safety. Clinicians remain responsible for interpretation, and restricted research access does not eliminate dual-use risk.

04

TERMS WORTH KEEPING

SOURCES AND VERIFICATION STATUS

This article was written from the materials below. Product claims and dates were checked against those sources on September 5, 2026.

PUBLICATION RECEIPT: Revision 1. Approved by Zak and published September 5, 2026.

THE PUBLICATION ENGINE

WANT A SIGNAL OF YOUR OWN?

We build source-grounded publications, private briefings, and editorial systems for organizations with something useful to say.

WORK WITH US