Why AI Can't Just Be Accurate: The New Framework Forcing Healthcare to Rethink Trust
Artificial intelligence systems can generate answers, but many cannot clearly explain how they arrived at those answers, creating a critical trust gap in healthcare and other regulated industries. A new peer-reviewed study published in Frontiers in Pharmacology argues that the next generation of AI innovation must prioritize not just performance, but credibility, introducing a framework designed to evaluate AI systems beyond raw accuracy.
The research, published by Unisys, explores one of the most significant barriers to AI adoption in healthcare: the inability of many AI systems to show their work. When a doctor or pharmacist receives a recommendation from an AI system, they need to understand not just what the system recommends, but why it made that recommendation. Without that visibility, organizations struggle to validate results and build confidence in the technology.
What Makes AI Trustworthy in Healthcare?
To address this challenge, the research introduces a neurosymbolic AI approach, which combines advanced artificial intelligence techniques with human knowledge and reasoning to provide greater transparency. Think of it as teaching AI systems to think more like humans do, combining raw computational power with logical reasoning that can be traced and understood.
The study also introduces a new evaluation framework called MURP, which stands for Mechanistic Coherence, Uncertainty, Robustness and Provenance. This framework extends how organizations evaluate AI systems beyond predictive accuracy to include criteria needed for use in highly regulated environments like healthcare and pharmaceuticals.
"The next phase of AI innovation will be shaped not only by performance, but by credibility. Organizations are increasingly exploring AI in complex and regulated settings, making the ability to evaluate, govern and stand behind AI-driven outcomes just as important as the outcomes themselves," said Salvatore Sinno, vice president of innovation, Enterprise Computing Solutions, Unisys.
Salvatore Sinno, Vice President of Innovation, Enterprise Computing Solutions, Unisys
How to Evaluate AI Systems for Regulated Industries
- Mechanistic Coherence: The AI system must be able to explain its reasoning in a way that makes logical sense to human experts, not just produce a number or recommendation.
- Uncertainty Quantification: The system should acknowledge when it is less confident in its conclusions, rather than presenting all outputs with equal certainty.
- Robustness Testing: The AI must be tested to ensure it produces reliable results across different scenarios and edge cases, not just on ideal datasets.
- Provenance Tracking: Every recommendation should be traceable back to the data and reasoning that produced it, creating an audit trail for regulatory compliance.
The paper specifically focuses on drug repurposing, a process where researchers identify new medical uses for existing medications. This is an area where explainability becomes critical, because pharmaceutical decisions affect patient safety and regulatory approval.
The research reflects a broader shift in how organizations are approaching AI adoption. Rather than simply deploying the most accurate model available, companies in regulated industries are asking harder questions about governance, accountability, and the ability to defend AI-driven decisions to regulators and stakeholders.
This publication is part of Unisys's broader AI-First strategy, which focuses on moving organizations from experimental AI projects to scalable, real-world adoption. The company's recent AI and Cloud Insights Report 2026 reinforces this theme, emphasizing that credibility and trust are now as important as raw performance metrics in determining whether AI tools will actually be adopted at scale.
The implications extend beyond healthcare. Any industry with strict regulatory requirements, compliance obligations, or high-stakes decision-making faces similar challenges. Financial services, insurance, and government agencies all need AI systems that can justify their recommendations to auditors, regulators, and the public. The MURP framework provides a structured way to evaluate whether an AI system meets those demands.