General-purpose large language models outperformed specialized clinical AI tools

For the past few years, the healthcare industry has assumed that AI built specifically for medicine would naturally outperform general-purpose AI. That assumption now deserves a closer look. A new study published in Nature Medicine found that leading general-purpose large language models consistently performed better than specialized clinical AI systems across several important medical evaluations.

This matters because healthcare organizations have invested heavily in purpose-built AI products. The expectation has been that domain-specific training would produce better clinical reasoning, more accurate medical knowledge, and stronger support for physicians. In several independent evaluations, the latest frontier models demonstrated stronger performance despite being designed for a much broader range of tasks.

The evaluation was comprehensive rather than relying on a single benchmark. Researchers measured medical knowledge through US Medical Licensing Examination-style questions, assessed how closely AI responses aligned with expert clinicians, and evaluated answers to real clinical questions submitted by physicians. Strong performance across all three areas indicates that the newest generation of general-purpose models is becoming increasingly capable of handling complex medical information.

For executives, this changes how AI strategy should be evaluated. Instead of assuming that industry-specific branding equals better outcomes, organizations should focus on independently measured performance. The pace of improvement in general-purpose AI is exceptionally fast because these models receive continuous investment, extensive testing, and large-scale deployment across many industries. That momentum can translate into rapid improvements for healthcare applications as well.

The study published in Nature Medicine evaluated GPT-5.2, Gemini 3.1 Pro, and Claude Opus 4.6 against the specialized clinical AI systems OpenEvidence and UpToDate Expert AI. Researchers tested the models using 500 MedQA questions based on the US Medical Licensing Examination, 500 HealthBench evaluation items, and 100 real clinical queries reviewed through a randomized, blinded assessment by 12 clinicians.

Benchmark results consistently favored the general-purpose AI models

The most important finding was not that a general-purpose model won one benchmark. It was that they performed better across every major evaluation. Consistency matters much more than isolated success, particularly in healthcare, where reliability directly affects clinical decision-making.

In the medical knowledge assessment, Gemini 3.1 Pro achieved 97.4% accuracy. GPT-5.2 followed with 94.2%, while Claude Opus 4.6 reached 90.2%. The specialized systems, OpenEvidence and UpToDate Expert AI, scored 89.6% and 88.4%, respectively. Although these differences may appear modest at first glance, they become significant when applied across thousands or millions of clinical interactions.

The pattern continued in the HealthBench evaluation, which measures how closely AI responses match expert clinical judgment rather than simply identifying the correct answer. GPT-5.2 received a score of 88 out of 100. Gemini 3.1 Pro scored 79.3, and Claude Opus 4.6 scored 77. OpenEvidence scored 62.6, while UpToDate Expert AI scored 61.3. This gap suggests that the leading frontier models currently communicate clinical recommendations in ways that more closely reflect expert medical reasoning.

Researchers also evaluated the models using 100 real clinical queries collected from physicians. Here, two clear performance groups emerged. The three general-purpose models consistently formed the higher-performing group, while the specialized clinical AI systems formed the lower-performing group. OpenEvidence and UpToDate Expert AI performed at roughly the same level as Google’s AI Overview in this real-world evaluation.

For business leaders, this reinforces an important principle. Procurement decisions should be driven by measurable outcomes rather than product positioning. Specialized healthcare AI may still offer advantages in workflow integration, regulatory features, documentation, or enterprise support. However, if the objective is achieving the highest-quality reasoning and clinical knowledge, independent benchmark performance deserves significant weight in investment decisions.

The numbers also highlight how quickly frontier AI is advancing. Organizations evaluating AI platforms today should expect performance leadership to change rapidly as new foundation models are released. Procurement strategies that remain flexible and benchmark-driven will be better positioned than those tied to assumptions about specialization alone.

Okoone experts
LET'S TALK!

A project in mind?
Schedule a 30-minute meeting with us.

Senior experts helping you move faster across product, engineering, cloud & AI.

Please enter a valid business email address.

Limited transparency surrounding proprietary clinical AI models makes independent validation more difficult

Performance is only part of the decision. Transparency is equally important, especially in healthcare. The study highlights a major challenge with many proprietary clinical AI products: healthcare providers often have very little visibility into how these systems are built or validated.

According to the researchers, companies behind specialized clinical AI tools generally do not disclose their model architecture, the foundation models they rely on, or the details of their training pipelines. As a result, hospitals and clinicians must evaluate these products without fully understanding how they were developed or why they perform the way they do.

This creates practical challenges for organizations responsible for patient safety. Healthcare leaders must assess whether an AI system is reliable, whether it produces consistent results across different clinical situations, and whether its claims are supported by independent evidence. When technical information is unavailable, those assessments become more dependent on vendor-provided information instead of external validation.

The issue extends beyond technical performance. Regulators, compliance teams, and clinical governance committees increasingly expect organizations to understand the systems they deploy. As AI becomes part of routine clinical workflows, transparency supports risk management, procurement decisions, and long-term accountability. Products that cannot be independently evaluated may face greater scrutiny as governance standards continue to evolve.

For executives, this shifts the conversation from simply asking whether an AI tool works to asking how confidently its performance can be verified. Independent benchmarking, published research, and reproducible evaluations become increasingly valuable when comparing competing AI platforms.

The Nature Medicine study specifically notes that the proprietary clinical AI systems evaluated do not publicly disclose their architectures, base models, or training pipelines. The researchers argue that this limits independent assessment of both their claimed advantages and their safety.

The findings support combining general-purpose AI with hospital-specific models rather than relying on a single approach

The researchers are not arguing that healthcare organizations should abandon specialized clinical AI. Their recommendation is more balanced. They suggest combining the strengths of frontier general-purpose models with AI systems built using each hospital’s own institutional knowledge and data.

This approach recognizes that different AI systems serve different purposes. General-purpose models continue to improve rapidly and perform well across broad medical reasoning tasks. At the same time, hospital-specific models can incorporate local clinical guidelines, treatment pathways, operational procedures, and organizational policies that general models cannot fully capture on their own.

For healthcare providers, this offers a practical path forward. General-purpose models can support lower-risk activities such as drafting documentation, summarizing information, or assisting with administrative work. More sensitive clinical decisions can benefit from models that have been adapted to the institution’s own standards and governance processes.

This strategy also provides greater flexibility. Healthcare organizations are not locked into a single vendor or technology stack. As frontier models continue to improve, institutions can update the underlying foundation models while preserving the local knowledge, safeguards, and workflows that make their AI systems clinically relevant.

For business leaders, this is an important strategic consideration. AI should be viewed as an evolving capability rather than a fixed product purchase. Building an architecture that allows multiple models to work together can reduce dependence on individual vendors, accelerate adoption of future advances, and create stronger long-term resilience as both technology and regulation continue to change.

The researchers conclude that specialized clinical AI tools may still carry institutional legitimacy and be appropriate for routine clinical use. However, based on their findings, they recommend developing hospital-specific large language models that leverage institutional data while using leading general-purpose models for less-sensitive tasks.

Investment in specialized clinical AI continues to grow

The market for clinical AI remains highly active. Investors continue to commit substantial capital to companies developing healthcare-specific AI products, even as independent research raises questions about whether these systems consistently outperform the latest general-purpose models. This reflects confidence that specialized AI can deliver value beyond benchmark performance alone.

OpenEvidence illustrates this trend. Earlier this year, the company raised $250 million in a closed Series D funding round, bringing its valuation to $12 billion. Since then, it has expanded its platform with new capabilities, including audio telehealth, AI-assisted medical coding, prescription support, and patient prioritization features. These additions show that healthcare AI companies are increasingly competing by offering integrated clinical workflows rather than standalone language models.

For healthcare executives, this distinction is important. A language model is only one component of an enterprise healthcare solution. Hospitals also evaluate how well a platform integrates with electronic health records, protects patient data, supports regulatory compliance, manages user permissions, and fits into existing clinical processes. These operational capabilities can influence purchasing decisions even if another model achieves higher benchmark scores.

At the same time, the study suggests organizations should separate platform capabilities from underlying model performance during procurement. A product may provide excellent workflow integration while relying on an AI model that no longer represents the state of the art. Evaluating these dimensions independently allows organizations to make more informed investment decisions and avoid assuming that a specialized platform automatically delivers superior clinical reasoning.

The broader implication is that competition in healthcare AI is entering a new phase. Success is likely to depend on combining high-performing foundation models with strong enterprise software, trusted governance, security, and seamless clinical integration. Organizations that continuously reassess both components will be better positioned as AI capabilities continue to evolve.

Earlier this year, OpenEvidence raised $250 million in a closed Series D funding round, reaching a valuation of $12 billion. Following the funding, the company expanded its product portfolio with audio telehealth, AI coding, prescription support, and patient prioritization features. These developments highlight continued investor confidence in specialized clinical AI platforms, even as independent benchmarking shows that leading general-purpose models currently outperform them on several medical evaluations.

Key takeaways for decision-makers

  • Prioritize proven performance over specialization: Independent testing found that leading general-purpose AI models outperformed specialized clinical AI across medical knowledge, clinical reasoning, and real physician queries. Leaders should base AI adoption decisions on validated outcomes rather than assuming healthcare-specific tools are inherently superior.
  • Benchmark AI before making long-term commitments: The performance gap was consistent across multiple evaluations, suggesting that regular, independent benchmarking should become part of AI procurement and governance. This helps ensure organizations invest in models that continue to deliver the strongest results as the technology evolves.
  • Make transparency part of vendor selection: Proprietary clinical AI systems often provide limited visibility into their architectures and training methods, making independent validation more difficult. Executives should favor vendors that support rigorous evaluation with published evidence and clear documentation to strengthen governance and reduce risk.
  • Build a flexible hybrid AI strategy: The study recommends combining high-performing general-purpose models with hospital-specific AI built on institutional data. This approach allows organizations to benefit from rapid advances in foundation models while maintaining local clinical standards, governance, and patient safety.
  • Separate platform value from model performance: Strong workflow integration, security, and compliance remain important differentiators for enterprise healthcare AI, but they should not be confused with model quality. Leaders should evaluate the underlying AI model and the surrounding platform independently to make better long-term technology investments.

Alexander Procter

July 30, 2026

9 Min

Okoone experts
LET'S TALK!

A project in mind?
Schedule a 30-minute meeting with us.

Senior experts helping you move faster across product, engineering, cloud & AI.

Please enter a valid business email address.