Industry: Healthcare
Organization: McLeod Health
Use case: Ambient clinical documentation and vendor selection
Evidence basis: 2026 peer-reviewed JMIR Medical Informatics study
Disclosure: This is an independent analysis by Conscious Engines based on public research. Financial gains were projected from observed volume and coding changes, not reported as audited revenue.
1. Outcome at a Glance
McLeod Health used a three-phase selection and rollout process to compare four ambient AI products before committing to one. A 90-day pilot across five ambulatory specialties included 23 clinicians in the core EHR-metric analysis. The system later reached 81% adoption and generated more than 150,000 notes.
Key Outcomes
28.3% reduction
Time in notes
Statistically significant, P<.001, n=23 reported in the cited case.
35.4% reduction
After-hours “pajama time”
Trend, P=.054, not conventionally significant reported in the cited case.
3.8% increase
Level 4 established-patient coding
P=.05 reported in the cited case.
8.5% increase
Established-patient volume
Observed during pilot period.
$2,629 per provider
Projected monthly revenue
Projection tied to volume and coding patterns reported in the cited case.
| Measure | Reported result | Evidence interpretation |
|---|---|---|
| Time in notes | 28.3% reduction | Statistically significant, P<.001, n=23 |
| After-hours “pajama time” | 35.4% reduction | Trend, P=.054, not conventionally significant |
| Level 4 established-patient coding | 3.8% increase | P=.05 |
| Established-patient volume | 8.5% increase | Observed during pilot period |
| Projected monthly revenue | $2,629 per provider | Projection tied to volume and coding patterns |
| System-wide adoption | 81% | Reported after rollout |
| Notes generated | More than 150,000 | Reported operating volume |
This is one of the more useful public ambient-AI cases because it reports statistical significance, shows the vendor-selection process, and separates observed changes from projected revenue. It also represents a nonprofit, nonacademic regional system with seven hospitals and more than 1,200 providers, rather than only a major academic medical center.
2. The Operational Problem
McLeod Health faced the same buying problem that many mid-sized health systems face: more than 100 ambient documentation products may be available, but vendor demonstrations are difficult to compare. A polished demo can hide weak specialty performance, poor EHR integration, excessive note editing, or coding problems.
The organization needed to determine four things before scaling:
- whether the product captured complex clinical facts accurately;
- whether the note was readable and usable by physicians;
- whether documentation supported appropriate coding;
- whether the tool worked naturally inside Epic.
This is a procurement and evaluation problem as much as a model problem. A health system that skips structured comparison can lock itself into the wrong workflow, then mistake low adoption for resistance to AI.
The financial problem also requires care. Time released after hours improves wellbeing but may not create cash savings. Higher patient volume and more complete coding may affect revenue, but the organization must confirm that changes reflect clinically appropriate documentation rather than model-driven inflation.
3. What Was Built
McLeod created an evaluation funnel before creating an enterprise rollout.
System at a Glance
Simulated evaluation
Four vendors processed 15 complex outpatient scripts with leaders acting as standardized patients.
Multidisciplinary scoring
Physicians, revenue-cycle specialists, and nonclinical reviewers assessed accuracy, billing quality, and readability.
Workflow demonstration
The top two vendors demonstrated Epic integration and usability.
Clinical pilot
The selected product ran for 90 days across five ambulatory specialties.
System rollout
Adoption, documentation, coding, volume, satisfaction, and financial indicators were monitored.
| Phase | Design |
|---|---|
| Simulated evaluation | Four vendors processed 15 complex outpatient scripts with leaders acting as standardized patients |
| Multidisciplinary scoring | Physicians, revenue-cycle specialists, and nonclinical reviewers assessed accuracy, billing quality, and readability |
| Workflow demonstration | The top two vendors demonstrated Epic integration and usability |
| Clinical pilot | The selected product ran for 90 days across five ambulatory specialties |
| System rollout | Adoption, documentation, coding, volume, satisfaction, and financial indicators were monitored |
This architecture makes the health system's evaluation set a strategic asset. The 15 scripted encounters create controlled comparisons across products. Live pilot data then show how the chosen system behaves under real audio, specialty language, clinician styles, and workflow pressure.
A bespoke alternative would preserve the same structure while allowing the enterprise to choose each model layer independently: domain-adapted speech-to-text, diarization, note generation, specialty templates, private RAG, structured validation, and Epic integration. The health system can route specialties or languages to the model that performs best instead of accepting a single vendor's model for every encounter.
4. How It Reached Production
The rollout is valuable as a repeatable buyer method.
Use the same cases for every vendor. Comparable audio and requirements reduce the influence of sales presentation quality.
Score different stakeholder outcomes. Physicians care about clinical accuracy and edit burden. Revenue-cycle teams care about evidence and coding integrity. Information technology teams care about integration, identity, latency, and support.
Pilot across heterogeneous specialties. A model that works in primary care may fail in oncology, surgery, pediatrics, or complex multi-speaker encounters.
Use objective EHR telemetry. Time in notes, after-hours activity, same-day closure, and encounter use are stronger than satisfaction surveys alone.
Keep statistical and commercial meaning separate. The 35.4% pajama-time decline narrowly missed conventional statistical significance. The $2,629 figure is projected monthly revenue, not an audited realization.
Before deployment, a health system should define a release threshold for critical clinical errors, unsupported facts, missing facts, edit distance, time to signed note, clinician use, patient consent, and cost per accepted note.
5. What Healthcare Leaders Should Take Away
McLeod Health shows that a regional provider can conduct serious AI diligence without the research budget of a large academic institution. The evaluation started with four vendors and 15 controlled cases, moved through a 23-clinician measurement cohort, and reached 81% adoption with more than 150,000 notes.
A production-ready clinical model program follows this evidence-first pattern. The provider owns the evaluation corpus, clinical glossary, note policies, integrations, security boundary, and acceptance thresholds. Models can change without surrendering the operating system around them.
The right commercial metric is cost per clinically accepted note at the required safety level, with separate accounting for time released, capacity used, coding effects, and revenue actually realized.
Related Conscious Engines research
- The enterprise AI model stack for healthcare
- When generic speech models enter the clinic
- Building the private clinical knowledge layer
- TPMG's 2.58-million-encounter ambient AI rollout
- Why your evaluation set is your AI moat
Sources
- JMIR Medical Informatics, Selecting, Scaling, and Measuring the Value of Ambient AI in a Nonacademic Health System
- The study reported P values and a limited core EHR-metric sample. The projected revenue figure should remain labeled as projected.