{"id":10768,"date":"2026-07-29T12:41:26","date_gmt":"2026-07-29T12:41:26","guid":{"rendered":"https:\/\/nile1.com\/en\/?p=10768"},"modified":"2026-07-29T12:41:43","modified_gmt":"2026-07-29T12:41:43","slug":"medical-ai-valuations-surge-as-benchmark-reveals-critical-omission-errors-in-clinical-software","status":"publish","type":"post","link":"https:\/\/nile1.com\/en\/2026\/07\/29\/medical-ai-valuations-surge-as-benchmark-reveals-critical-omission-errors-in-clinical-software\/","title":{"rendered":"Medical AI Valuations Surge as Benchmark Reveals Critical Omission Errors in Clinical Software"},"content":{"rendered":"<p>A rigorous study evaluating artificial intelligence in clinical settings reveals that even top-performing medical AI tools frequently make critical errors of omission, exposing a stark divide between soaring Silicon Valley valuations and real-world medical safety.<\/p>\n<p>The independent evaluation, known as the NOHARM benchmark, examined clinical AI platforms across 1,100 actual patient cases using nearly 13,000 physician annotations to quantify potential patient harm. Jointly conducted by researchers at <a href=\"https:\/\/www.stanford.edu\" target=\"_blank\" rel=\"noopener\">Stanford University<\/a>, Harvard University, and the ARISE network, the study benchmarked specialized tools alongside general frontier models, including OpenAI\u2019s GPT-5.6 Sol and Anthropic\u2019s Claude Fable 5.<\/p>\n<p>Among the platforms tested, Doximity\u2019s enterprise tool Ask achieved the highest score, ahead of specialized competitors such as OpenEvidence. However, researchers discovered a pervasive vulnerability across every system: 76.6% of all harmful errors were omissions\u2014instances where the AI left out vital medical facts or treatment options rather than stating incorrect information.<\/p>\n<p>Eric Topol, a scientist at Scripps Research and co-chair of Doximity\u2019s PeerCheck verification program, noted that this pattern reveals a persistent &#8220;illusion of readiness&#8221; in medical AI. Topol emphasized that while doctors using AI delivered better care than those working without it, errors of omission must approach zero before health systems can rely heavily on these platforms.<\/p>\n<p>The benchmark results arrive amid a wave of intense venture capital investment into point-of-care tools. OpenEvidence, founded in 2021 as a free search engine querying peer-reviewed journals, saw its valuation rise from $1 billion in February to $12 billion in January following an investment round co-led by <a href=\"https:\/\/nile1.com\/en\/2026\/07\/28\/thrive-capital-leads-4-2-billion-bid-for-fifa-commercial-unit-as-european-football-reacts\/\" class=\"auto-internal-link\" title=\"Thrive Capital Leads $4.2 Billion Bid for FIFA Commercial Unit as European Football Reacts\">Thrive Capital<\/a> and DST Global. The company raised roughly $700 million over 12 months with backing from <a href=\"https:\/\/nile1.com\/en\/2026\/07\/16\/bunkerhill-health-secures-25-million-to-scale-ai-platform-tackling-healthcares-dual-crises-of-burnout-and-missed-diagnoses\/\" class=\"auto-internal-link\" title=\"Bunkerhill Health Secures $25 Million to Scale AI Platform Tackling Healthcare\u2019s Dual Crises of Burnout and Missed Diagnoses\">Sequoia<\/a>, Kleiner Perkins, and GV.<\/p>\n<p>OpenEvidence Chief Executive Daniel Nadler challenged the NOHARM findings, arguing that the study allowed model re-testing and noting that the benchmark itself had not completed peer review.<\/p>\n<p>Doximity has taken a different commercial route, integrating its Ask AI tool into enterprise contracts with over 150 health systems to assist physicians with note summarization, administrative documentation, and drug interaction analysis. Outputs are filtered through its PeerCheck mechanism, where physicians verify references against source material. Doximity reported $145.4 million in quarterly revenue, reflecting 5% year-over-year growth.<\/p>\n<p>The findings coincide with shifting regulatory and legal frameworks governing <a href=\"https:\/\/www.fda.gov\/medical-devices\/digital-health-center-excellence\" target=\"_blank\" rel=\"noopener\">clinical decision support software<\/a>. Federal guidelines relaxed rules to allow broader operation provided clinicians can independently verify the AI&#8217;s reasoning, while state laws passed in 2026 mandate direct human physician sign-off before AI-driven decisions reach patients. Meanwhile, malpractice law remains unsettled over whether liability for clinical errors falls on physicians, hospital networks, or AI vendors.<\/p>\n<div class=\"related-news-box\">\n<h3 class=\"related-news-title\">Read also:<\/h3>\n<ul class=\"related_news_list\">\n<li><a href=\"https:\/\/nile1.com\/en\/2026\/07\/29\/enterprise-ai-deployments-face-financial-reckoning-as-token-costs-explode\/\">Enterprise AI Deployments Face Financial Reckoning as Token Costs Explode<\/a><\/li>\n<li><a href=\"https:\/\/nile1.com\/en\/2026\/07\/29\/six-flags-great-adventure-unveils-382-foot-bakunawa-coaster-for-2027-debut\/\">Six Flags Great Adventure Unveils 382-Foot &#8216;Bakunawa&#8217; Coaster for 2027 Debut<\/a><\/li>\n<li><a href=\"https:\/\/nile1.com\/en\/2026\/07\/29\/orla-nixon-leads-claim-operations-at-new-york-life-group-benefit-solutions\/\">Orla Nixon Leads Claim Operations at New York Life Group Benefit Solutions<\/a><\/li>\n<\/ul>\n<\/div>\n","protected":false},"excerpt":{"rendered":"<p>A rigorous study evaluating artificial intelligence in clinical settings reveals that even top-performing medical AI tools frequently make critical errors of omission, exposing a stark divide between soaring Silicon Valley valuations and real-world medical safety. The independent evaluation, known as the NOHARM benchmark, examined clinical AI platforms across 1,100 actual patient cases using nearly 13,000 &hellip;<\/p>\n","protected":false},"author":1,"featured_media":10770,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_sitemap_exclude":false,"_sitemap_priority":"","_sitemap_frequency":"","footnotes":""},"categories":[3],"tags":[13539,13533,13537,13535,13536,13534,13532,13538,5653,4626,13290],"class_list":["post-10768","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-business","tag-clinical-decision-support-software","tag-doximity","tag-dst-global","tag-eric-topol","tag-kleiner-perkins","tag-noharm","tag-openevidence","tag-peercheck","tag-sequoia","tag-stanford-university","tag-thrive-capital"],"_links":{"self":[{"href":"https:\/\/nile1.com\/en\/wp-json\/wp\/v2\/posts\/10768","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/nile1.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/nile1.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/nile1.com\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/nile1.com\/en\/wp-json\/wp\/v2\/comments?post=10768"}],"version-history":[{"count":3,"href":"https:\/\/nile1.com\/en\/wp-json\/wp\/v2\/posts\/10768\/revisions"}],"predecessor-version":[{"id":10772,"href":"https:\/\/nile1.com\/en\/wp-json\/wp\/v2\/posts\/10768\/revisions\/10772"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/nile1.com\/en\/wp-json\/wp\/v2\/media\/10770"}],"wp:attachment":[{"href":"https:\/\/nile1.com\/en\/wp-json\/wp\/v2\/media?parent=10768"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/nile1.com\/en\/wp-json\/wp\/v2\/categories?post=10768"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/nile1.com\/en\/wp-json\/wp\/v2\/tags?post=10768"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}