m
Recent Posts
HomeProviderNew Study Says OpenEvidence Beats General AI

New Study Says OpenEvidence Beats General AI

OpenEvidence clinical AI study

OpenEvidence outperformed Claude Opus 4.8, Gemini 3.1 Pro and GPT-5.5 on real-world clinical questions in a new physician-graded OpenEvidence clinical AI study, contradicting a June Nature Medicine paper that found general-purpose AI models beat specialized clinical tools. The conflicting results leave health system leaders navigating competing evidence on how these tools stack up.

How the OpenEvidence Clinical AI Study Was Conducted

The study, posted June 27 as a preprint on arXiv, has not yet been peer-reviewed. It had 149 practicing physicians across 36 states rate AI answers to 620 real point-of-care questions drawn from OpenEvidence’s platform, plus 187 questions from the HealthBench benchmark, judging responses on accuracy, clinical utility, source quality, verifiability and completeness.

Clear Win Margins Across All Dimensions

OpenEvidence posted positive win margins on all five dimensions in this OpenEvidence clinical AI study, with differences ranging from 25 to 39 percentage points over the three general-purpose models. Claude Opus 4.8 and Gemini 3.1 Pro scored near parity with each other, while GPT-5.5 recorded the lowest win rates on every axis.

How This Contradicts the Earlier Nature Medicine Study

The results stood in contrast to a Nature Medicine study published June 12, which found GPT-5.2, Gemini 3.1 Pro and Claude Opus 4.6 outperformed OpenEvidence and Wolters Kluwer’s UpToDate Expert AI. This direct contradiction between the two studies has intensified scrutiny of both papers’ methodologies within the broader OpenEvidence clinical AI study debate.

OpenEvidence Seeks a Retraction

OpenEvidence has since asked Nature Medicine to retract that study, alleging flawed methods; the journal has pointed the company to its formal rebuttal process instead. This escalation reflects how high the stakes have become for companies competing in the clinical AI decision-support space.

Why the Two Studies Reached Different Conclusions

The new paper’s authors attribute the divergence partly to design differences. Their evaluation used 149 physicians in specialty-matched, head-to-head comparisons, versus 12 clinicians at one institution using isolated rubric scoring in the earlier study, a methodological distinction central to understanding why this OpenEvidence clinical AI study reached such different results.

Questions About Independence and Involvement

OpenEvidence codesigned the data collection plan, administered the survey and paid participating physicians, though the paper’s authors, based at University of California San Francisco, Harvard Medical School and Stanford University, among others, report no affiliation with the company. This level of company involvement in study design is worth noting for readers evaluating the strength of the findings independently.

What This Means for Health System AI Decisions

The competing results leave hospital and health system leaders with conflicting independent evidence on how general-purpose AI models compare with specialized clinical decision support tools as adoption of both accelerates. This OpenEvidence clinical AI study debate underscores a broader challenge facing health IT leaders: evaluating AI performance claims when even peer-reviewed and preprint studies produce contradictory conclusions.

What to Watch Going Forward

As both studies continue to draw scrutiny, and as OpenEvidence pursues its retraction request with Nature Medicine, health system leaders may want to look beyond any single study’s conclusions and instead track how these tools perform within their own clinical environments before making adoption decisions based on this OpenEvidence clinical AI study or its Nature Medicine counterpart alone.

For more healthcare industry updates, insights and news, visit DistilINFOClick here to subscribe to stay informed.

Share

No comments

Sorry, the comment form is closed at this time.