Your doctor is probably looking up your symptoms on their phone.
OpenEvidence, founded in 2021, is a free, ad-supported AI search engine for doctors that pulls answers straight from peer-reviewed medical journals at the point of care and labels how strong the evidence is. It’s become somewhat of the poster child of the medical AI boom.
In February, a Sequoia-led round valued the company at $1 billion. In the following months, that number grew to $3.5 billion by July with GV and Kleiner Perkins co-leading, $6 billion in October to $12 billion this January, in a round co-led by Thrive Capital and DST Global. That’s roughly $700 million raised in about a year.
Meanwhile, competitor Doximity has taken an enterprise approach. Best known as a professional networking platform for physicians, it now sells an AI assistant called Ask that helps doctors summarize patient notes, check drug interactions, and draft documentation. Ask is folded into paid, enterprise contracts with more than 150 health systems. Every answer runs through a human-review layer called PeerCheck, where physicians check AI output against the original sources it’s citing. Doximity posted $145.4 million in quarterly revenue this spring, up 5% year over year.
In mid-July, a new independent benchmark called NOHARM put both Doximity and OpenEvidence’s AI tools to the test, alongside OpenAI’s GPT-5.6 Sol and Anthropic’s Claude Fable 5. Built by researchers at Stanford, Harvard, and the ARISE network, it ran 1,100 real clinical cases through each model and collected roughly 13,000 physician annotations to score for patient harm. Doximity Ask came out on top. OpenEvidence contested the accuracy of Doximity’s score, however.
“We believe that a rigorous study methodology does not allow AIs with perfect memories to ask for ‘re-tests.’” CEO Daniel Nadler told me via email. “We suspect this would not pass peer review. To be clear, the NOHARM study is not even peer-reviewed—the basic table stakes requirement in medicine for even the flimsiest medical conclusions.”
But the finding isn’t really about who “won.” It’s that even the best-performing models still miss things. Across every AI system tested, 76.6% of harmful errors were omissions, meaning the AI left something out, not that it stated something factually wrong.
Eric Topol, cardiologist, Scripps Research scientist, and co-chair of Doximity’s PeerCheck program, has spent his career studying diagnostic error. He said that distinction is significant: “Errors of omission need to be brought as close to zero as possible,” he told me, adding that today’s models maintain an “illusion of readiness” that has followed medical AI even as it improves. (NOHARM still noted that doctors equipped with AI give better care compared to those without)
The regulatory backdrop makes NOHARM’s timing pointed. The FDA loosened its stance on AI-powered clinical decision-support tools this January, giving them more room to operate as long as doctors can independently check the AI’s reasoning. States have moved in the opposite direction, passing more than a dozen new laws in 2026 governing AI use in healthcare (most requiring a human to sign off before any AI-assisted decision reaches a patient). Malpractice law hasn’t caught up to either trend: courts are still untangling who’s liable, the doctor, the hospital, or the AI vendor, when a model’s suggestion turns out to be wrong.
That legal gray zone stands to be an early signal of what regulators, hospital systems, and the next wave of investors will start asking.
See you tomorrow,
Lily Mae Lazarus
X: @LilyMaeLazarus
Email: lily.lazarus@fortune.com
Submit a deal for the Term Sheet newsletter here.
Joey Abrams curated the deals section of today’s newsletter. Subscribe here.