Healthcare leaders do not need another list of theoretical AI use cases. They need evidence that a system worked with real clinicians, patients, EHRs, contact centers, privacy controls, and operating constraints.
This index organizes eight named deployments by evidence strength. It includes peer-reviewed studies, health-system reports, and vendor-supported customer cases. Those sources are not interchangeable. Randomized evidence can establish a stronger causal signal. Large observational deployments show whether a system can survive routine practice. Vendor cases can reveal useful operating metrics but require independent verification.
Disclosure: Conscious Engines did not deliver the implementations described. This is an independent synthesis of public evidence intended to help enterprise buyers design and evaluate their own systems.
1. The Results in One View
| Organization | Workflow | Scale | Headline result | Evidence type |
|---|---|---|---|---|
| UCLA Health | Ambient documentation | 238 physicians randomized | Nabla reduced time in notes 9.5% versus control; DAX difference was not significant | Randomized trial, NEJM AI |
| Providence | Ambient documentation | 1,547 active clinicians, 16,149 clinician-months | Median time in notes fell from 7.1 to 6.1 minutes; no appointment-per-day association | Peer-reviewed observational study, JAMA Network Open |
| McLeod Health | Ambient documentation | 23 clinicians in core pilot analysis, more than 150,000 later notes | Time in notes down 28.3%, 81% system adoption | Peer-reviewed implementation study |
| TPMG | Ambient documentation | 7,260 physicians, 2.58 million encounters | 15,791 estimated documentation hours saved | Large observational evaluation |
| Cleveland Clinic | Ambient documentation | More than 4,800 users, 3.5 million encounters | 70% encounter-level use among established users | Peer-reviewed implementation report |
| SolutionHealth | Ambient and active medical speech | More than 1,000 clinicians, nearly 60,000 encounters | 56% improvement in average documentation time | Vendor-supported named customer case |
| Inova Health | Patient-access voice agents | About 338,000 calls per month | 50% appointment calls resolved, 4,272 hours released monthly, 8.8x reported ROI | Vendor-supported named customer case |
| Houston Methodist | Vaccine information and scheduling voice agent | More than 200,000 calls in one month | 91% automation, every call answered on first ring | Health-system case study |
The results establish three points.
First, healthcare AI is already operating at enterprise scale. TPMG and Cleveland Clinic together report more than six million ambient-supported encounters. Second, average effects vary sharply by system, clinician, workflow, and measurement method. Providence and the UCLA randomized trial found modest documentation-time changes, while SolutionHealth and McLeod reported larger improvements. Third, voice AI can create value outside the clinical note by completing patient-access work at call-center scale.
2. What the Evidence Actually Supports
Ambient documentation can reduce work, but not equally for everyone
The strongest causal evidence in this set comes from UCLA's randomized trial of 238 physicians, assigned to DAX Copilot, Nabla, or a control group. Nabla users showed a statistically significant 9.5% reduction in time in notes versus control. DAX users did not show a significant difference on the primary time measure. Both products improved several physician-experience measures. Clinicians reported occasional clinically significant inaccuracies, and one mild patient-safety event was identified.
Providence provides scale and objective EHR telemetry. Its study found a small immediate reduction in time per note, followed by a gradual decrease in after-hours documentation. It found no significant association with daily appointment count.
These studies set a credible base case: ambient AI can help, but the value may appear in cognitive load, patient attention, or time returned after work rather than additional scheduled volume.
Adoption and utilization determine realized value
McLeod reported 81% adoption. Cleveland Clinic reported 70% encounter-level utilization among established users, after onboarding more than 4,000 clinicians in 16 weeks. TPMG's highest-use group activated ambient AI in 89% of eligible encounters.
These measures are not identical. Adoption can mean a user tried the system. Encounter-level utilization measures how often it is used when it could be used. Accepted-note rate measures how often the output survived review. Buyers should demand all three.
Voice agents create a separate healthcare business case
Inova's reported 8.8x return came from patient-access voice agents integrated with Epic, CRM, and telephony. Houston Methodist used a voice assistant to absorb a public-health demand spike, including 14,583 calls in one day.
The clinical-risk boundary is different from ambient notes. Appointment scheduling, modification, provider search, FAQs, and routing can often be constrained through deterministic workflows. Emergency symptoms, clinical advice, complex identity cases, and uncertain transactions should move to trained staff.
3. A Buyer's Evidence Ladder
Healthcare buyers should grade claims before using them in a board paper.
| Evidence level | What it can establish | Examples in this index | Main limitation |
|---|---|---|---|
| Randomized controlled trial | Stronger causal comparison | UCLA | Short period, single institution, selected physicians |
| Peer-reviewed EHR observational study | Objective real-world association at scale | Providence | Voluntary use and confounding |
| Peer-reviewed implementation study | Selection, workflow, adoption, and rollout evidence | McLeod, Cleveland Clinic | Often no randomized counterfactual |
| Large health-system evaluation | Feasibility and experience at production scale | TPMG | Selection and survey bias |
| First-party health-system case | Operating performance in a named workflow | Houston Methodist | Limited external validation |
| Vendor-supported named customer case | Commercial implementation and metrics | SolutionHealth, Inova | Strong incentive to emphasize favorable results |
No level should be ignored. A randomized two-month trial cannot prove five-year maintainability. A vendor case cannot prove causality. Together they define a more useful range of expected outcomes and failure modes.
Every claim should retain its qualifier:
- McLeod's $2,629 per provider per month was a projected revenue gain.
- Providence found no increase in appointments per day.
- Inova's 8.8x ROI formula was not publicly disclosed.
- Houston Methodist's results came from the exceptional 2021 vaccine context.
- SolutionHealth's 2.5 additional appointments was a capacity equivalent, not confirmed booked volume.
- TPMG's time saving was estimated from observed documentation behavior.
4. How to Evaluate a Bespoke Healthcare Model
The evidence suggests a 90-day evaluation should compare the bespoke model against the current workflow and at least one credible alternative.
For medical speech and ambient documentation, measure:
- clinical term and medication accuracy;
- speaker-attribution and negation errors;
- unsupported and omitted clinical facts;
- time in notes and after-hours EHR activity;
- edit distance and time to signed note;
- clinician adoption, encounter-level use, and accepted-note rate;
- performance by specialty, accent, language, and encounter complexity;
- cost per accepted note.
For patient-access voice agents, measure:
- intent and critical-field accuracy;
- correctly completed appointment or administrative task;
- transfer, abandonment, and rapid repeat contact;
- identity and authorization failures;
- emergency and clinical escalation performance;
- patient satisfaction by language and access cohort;
- cost per durable resolution.
The model stack should match the task. Medical ASR converts speech. A note model structures the encounter. Retrieval supplies current institutional knowledge. Deterministic checks validate critical fields. A voice agent manages dialogue and invokes approved actions. A routing layer sends only difficult cases to a larger model.
The enterprise should own the evaluation set even when it buys model APIs. That corpus contains the specialties, language, failure cases, and business rules that make the system trustworthy locally.
5. What This Means for Healthcare Leaders
The public record no longer supports asking whether healthcare AI can reach production. It has. The better questions are:
- Which workflow has a measurable denominator and a tolerable automation boundary?
- Which evidence applies to our specialties, languages, systems, and baseline burden?
- What outcome will finance and clinical governance accept as real value?
- Which parts of the model stack should we own to protect data, control cost, and improve locally?
- How will we detect regression when a model, prompt, or source document changes?
Conscious Engines builds bespoke healthcare AI systems around the enterprise's own language and operating environment. That can include medical speech-to-text, text-to-speech, clinical voice agents, private enterprise RAG, task-specific small language models, EHR integration, and continuous evaluation.
The commercial thesis is precise: use the smallest capable model for each step, keep clinical authority with qualified people, and measure cost per accepted outcome. The evidence above provides external benchmarks. The deployment still has to earn trust on the client's own data.
Read the full case studies
- TPMG: 2.58 million ambient-supported encounters
- SolutionHealth: ambient and active medical speech infrastructure
- McLeod Health: evidence-first vendor selection and rollout
- Cleveland Clinic: 4,000 clinicians onboarded in 16 weeks
- Providence: what objective EHR data showed
- Inova Health: patient-access voice AI and reported 8.8x ROI
- Houston Methodist: voice AI under extreme call volume
- The broader enterprise AI model stack for healthcare
Primary sources
- NEJM AI randomized trial via PubMed Central
- JAMA Network Open Providence study
- JMIR McLeod Health implementation study
- npj Health Systems Cleveland Clinic implementation report
- American Medical Association summary of the TPMG evaluation
- Microsoft SolutionHealth customer report
- Hyro Inova Health customer report
- Houston Methodist voice automation case study