Medical AI Valuations Surge as Benchmark Reveals Critical Omission Errors in Clinical Software
Independent research highlights safety gaps in AI platforms like OpenEvidence and Doximity amidst complex legal and regulatory shifts.
A rigorous study evaluating artificial intelligence in clinical settings reveals that even top-performing medical AI tools frequently make critical errors of omission, exposing a stark divide between soaring Silicon Valley valuations and real-world medical safety.
The independent evaluation, known as the NOHARM benchmark, examined clinical AI platforms across 1,100 actual patient cases using nearly 13,000 physician annotations to quantify potential patient harm. Jointly conducted by researchers at Stanford University, Harvard University, and the ARISE network, the study benchmarked specialized tools alongside general frontier models, including OpenAI’s GPT-5.6 Sol and Anthropic’s Claude Fable 5.
Among the platforms tested, Doximity’s enterprise tool Ask achieved the highest score, ahead of specialized competitors such as OpenEvidence. However, researchers discovered a pervasive vulnerability across every system: 76.6% of all harmful errors were omissions—instances where the AI left out vital medical facts or treatment options rather than stating incorrect information.
Eric Topol, a scientist at Scripps Research and co-chair of Doximity’s PeerCheck verification program, noted that this pattern reveals a persistent “illusion of readiness” in medical AI. Topol emphasized that while doctors using AI delivered better care than those working without it, errors of omission must approach zero before health systems can rely heavily on these platforms.
The benchmark results arrive amid a wave of intense venture capital investment into point-of-care tools. OpenEvidence, founded in 2021 as a free search engine querying peer-reviewed journals, saw its valuation rise from $1 billion in February to $12 billion in January following an investment round co-led by Thrive Capital and DST Global. The company raised roughly $700 million over 12 months with backing from Sequoia, Kleiner Perkins, and GV.
OpenEvidence Chief Executive Daniel Nadler challenged the NOHARM findings, arguing that the study allowed model re-testing and noting that the benchmark itself had not completed peer review.
Doximity has taken a different commercial route, integrating its Ask AI tool into enterprise contracts with over 150 health systems to assist physicians with note summarization, administrative documentation, and drug interaction analysis. Outputs are filtered through its PeerCheck mechanism, where physicians verify references against source material. Doximity reported $145.4 million in quarterly revenue, reflecting 5% year-over-year growth.
The findings coincide with shifting regulatory and legal frameworks governing clinical decision support software. Federal guidelines relaxed rules to allow broader operation provided clinicians can independently verify the AI’s reasoning, while state laws passed in 2026 mandate direct human physician sign-off before AI-driven decisions reach patients. Meanwhile, malpractice law remains unsettled over whether liability for clinical errors falls on physicians, hospital networks, or AI vendors.









