Every technology vendor claims their product works. The question that matters for life sciences is: how do you know, and can someone else verify it?
When Medeloop’s Chief Epidemiologist John Ayers set out to evaluate HealthVerity eXOs, he brought the skepticism of a scientist, not the enthusiasm of a vendor. And what he was most skeptical of was not the AI, but the status quo.
"We evaluated studies from the New England Journal of Medicine, JAMA, the Lancet, the British Medical Journal. It turned out about 90% of those studies misinterpreted their primary finding." — John Ayers, Chief Epidemiologist, Medeloop
That finding reframes the question. Compared with how RWE is done today, how reliable, transparent and verifiable is this platform? And by that measure, the relevant standard is not perfection, but whether agentic AI can make real-world evidence generation more transparent, reproducible and verifiable than the status quo.
One of the most frequently cited concerns about AI systems in high-stakes settings is hallucination: the generation of confident-sounding outputs that have no basis in the underlying data. In a real-world evidence context, hallucinated numbers could find their way into regulatory submissions, payer dossiers, or clinical guidelines.
Across 250 independent runs, HealthVerity eXOs recorded zero hallucinations.1 The platform did not invent numbers not present in the data or generate results from language model inference. The platform wrote and executed code against real data, then reported what the code returned. The outputs were grounded in the data by construction.
That distinction between a system that generates text about data and a system that analyzes data and then explains the results, is fundamental to understanding why agentic AI can be more trustworthy than generative AI for this class of problems. Read our blog to understand the differences between agentic and generative AI.
The HealthVerity eXOs validation study reported a reproducibility consistency score of 0.758 across repeated runs of the same research questions. For non-technical audiences, that number requires some context to interpret correctly.1
First, the benchmark: in published epidemiological literature, fewer than 2% of studies are independently judged reproducible meaning a reader could understand the methods well enough to replicate the study and expect to arrive at a comparable result. eXOs's score is not measured against a theoretical ideal; it is measured against the noisy, variable reality of how real-world research is actually conducted.
Second, the gap between the score and perfect reproducibility reflects that different methodological choices can produce different results, which is why visibility into study design, cohort logic, code sets and analysis steps is essential.
The gold standard for accuracy in observational research is concordance with previously published studies. If an agentic AI system produces results that are systematically inconsistent with the existing evidence base, that is a signal worth taking seriously.
Across the 250 eXOs runs, more than 90% achieved high or medium accuracy alignment with published literature. That figure is not derived from cherry-picked questions, rather it reflects performance across 50 diverse research questions designed to represent the range of analyses RWE teams actually conduct.
The code prompt adherence score, 99.2 out of 100, captures something related but distinct: whether the system executed the analysis it said it would execute. The protocol, the code, and the results are consistent. What the system documented doing is what it actually did.
One consequence of these validation results that deserves more attention is the equity dimension. Not every organization has the same access to experienced biostatisticians, seasoned epidemiologists and rigorous QA processes. Smaller teams, earlier-stage companies, and functions like Medical Affairs that have historically received a "lesser version" of RWE — they can now access the same analytical rigor as the best-resourced teams in the industry.
What the validation evidence demonstrates is that HealthVerity eXOs works consistently, transparently and in a way that can be independently verified. For the field, that is the proof of concept the industry has been waiting for.
Curious how eXOs performs when tested against published literature? Read the full validation study: Validating HealthVerity eXOs
HealthVerity. Validating HealthVerity eXOs: An Agentic AI Platform for Real-World Evidence. Published May 4, 2026. https://info.healthverity.com/validating-healthverity-exos-an-agentic-ai-platform-for-real-world-evidence