Artificial intelligence is rapidly reshaping healthcare. For years, healthcare organizations invested heavily in specialized clinical AI tools designed specifically for medical professionals. However, a new study suggests that general-purpose AI models may now outperform these healthcare-focused systems.
Researchers compared leading AI chatbots such as ChatGPT, Gemini, and Claude with specialized clinical tools including OpenEvidence and UpToDate Expert AI. Surprisingly, the general AI models consistently delivered better results across several medical benchmarks. These findings could influence how hospitals and physicians adopt AI in the coming years.
General AI Surpasses Clinical AI
The study evaluated two clinical AI assistants against three frontier large language models (LLMs). Researchers tested the systems using medical knowledge exams, clinician-alignment assessments, and real-world clinical questions.
The results were striking. General-purpose AI models outperformed specialized healthcare AI tools in every category. In particular, ChatGPT, Gemini, and Claude demonstrated stronger reasoning, clearer communication, and better alignment with physician expectations.
MedQA Performance
The MedQA benchmark measures medical knowledge using standardized healthcare questions.
Gemini achieved the highest accuracy with 97.4%. ChatGPT followed closely with 94.2%, while Claude scored 90.2%. In comparison, OpenEvidence achieved 89.6%, and UpToDate Expert AI scored 88.4%.
Although the specialized tools performed well, they failed to surpass the general AI models. This result challenges the assumption that healthcare-specific AI automatically offers superior expertise.
HealthBench Scores
Researchers also used HealthBench, a benchmark designed to measure how closely AI responses align with clinicians.
ChatGPT recorded the highest score at 88 out of 100. Gemini earned 79.3, while Claude scored 77. Meanwhile, OpenEvidence and UpToDate Expert AI lagged behind with scores of 62.6 and 61.3 respectively.
Therefore, general AI models not only possessed stronger medical knowledge but also communicated more effectively with clinicians.
Real Clinical Queries Test
The most important assessment involved real clinical questions submitted by physicians.
Researchers collected 100 de-identified clinical queries and asked twelve clinicians to evaluate the AI-generated responses through a randomized and blinded review process. Once again, ChatGPT, Gemini, and Claude formed the top-performing tier.
Interestingly, the clinical AI tools performed similarly to Google Search AI Overview rather than surpassing it. Consequently, healthcare leaders may need to reconsider the value proposition of specialized AI systems.
Why General AI Models Excel
Several factors explain why general AI models are gaining an advantage.
First, companies behind ChatGPT, Gemini, and Claude invest billions of dollars into model training and infrastructure. Their models learn from enormous datasets and receive frequent updates.
Second, these frontier models possess stronger reasoning and communication capabilities. They can synthesize information from multiple sources, explain complex concepts, and adapt their responses to different contexts.
Finally, general AI systems continue to evolve at a rapid pace. As a result, specialized healthcare tools may struggle to keep up unless they integrate the latest advances in large language models.
Implications for Healthcare Providers
The findings do not mean hospitals should abandon specialized clinical AI.
Instead, researchers suggest a hybrid approach. Healthcare organizations could use general AI models for tasks such as medical education, drafting summaries, and clinical research. At the same time, hospitals can develop institution-specific AI systems that incorporate local guidelines and patient data.
Moreover, healthcare leaders must remain cautious. AI models still produce hallucinations and may provide incorrect medical advice. Therefore, clinicians should continue reviewing AI-generated recommendations before making patient-care decisions.
Balancing Innovation and Safety
Patient safety remains the top priority.
Although AI performance has improved dramatically, no model should replace physicians entirely. Instead, AI should function as a decision-support tool that enhances human expertise.
Researchers believe the future will involve collaboration between doctors and AI systems. This partnership could improve diagnostic accuracy, streamline workflows, and reduce administrative burdens across healthcare organizations.
Future of AI in Clinical Practice
Healthcare AI is entering a new phase.
General-purpose models are no longer simple chatbots. They increasingly demonstrate advanced clinical reasoning and can outperform specialized healthcare systems on established benchmarks.
Nevertheless, hospitals must carefully evaluate accuracy, transparency, privacy, and regulatory compliance before adopting these technologies at scale. The organizations that successfully balance innovation with safety will likely lead the next generation of healthcare transformation.
Conclusion
The latest research highlights a significant shift in healthcare AI. ChatGPT, Gemini, and Claude outperformed specialized clinical AI tools across multiple medical benchmarks, including medical knowledge, clinician alignment, and real-world clinical queries.
As AI continues to evolve, healthcare organizations may increasingly rely on a combination of frontier AI models and institution-specific solutions. Ultimately, the future of healthcare AI will depend on how effectively these technologies support clinicians while maintaining patient safety and trust.
