We’re putting $5.5M behind research and development of smaller, specialized AI models for real-world deployment — Read our manifesto →
    All posts

    [Case Study] From a 23-Clinician Pilot to 150,000 Notes: McLeod Health's Evidence-First Ambient AI Rollout

    A community health system tested four vendors, measured documentation and coding outcomes, and reached 81% adoption before scaling.

    Conscious Engines

    Industry: Healthcare
    Organization: McLeod Health
    Use case: Ambient clinical documentation and vendor selection
    Evidence basis: 2026 peer-reviewed JMIR Medical Informatics study
    Disclosure: This is an independent analysis by Conscious Engines based on public research. Financial gains were projected from observed volume and coding changes, not reported as audited revenue.

    1. Outcome at a Glance

    McLeod Health used a three-phase selection and rollout process to compare four ambient AI products before committing to one. A 90-day pilot across five ambulatory specialties included 23 clinicians in the core EHR-metric analysis. The system later reached 81% adoption and generated more than 150,000 notes.

    Key Outcomes

    28.3% reduction

    Time in notes

    Statistically significant, P<.001, n=23 reported in the cited case.

    35.4% reduction

    After-hours “pajama time”

    Trend, P=.054, not conventionally significant reported in the cited case.

    3.8% increase

    Level 4 established-patient coding

    P=.05 reported in the cited case.

    8.5% increase

    Established-patient volume

    Observed during pilot period.

    $2,629 per provider

    Projected monthly revenue

    Projection tied to volume and coding patterns reported in the cited case.

    MeasureReported resultEvidence interpretation
    Time in notes28.3% reductionStatistically significant, P<.001, n=23
    After-hours “pajama time”35.4% reductionTrend, P=.054, not conventionally significant
    Level 4 established-patient coding3.8% increaseP=.05
    Established-patient volume8.5% increaseObserved during pilot period
    Projected monthly revenue$2,629 per providerProjection tied to volume and coding patterns
    System-wide adoption81%Reported after rollout
    Notes generatedMore than 150,000Reported operating volume

    This is one of the more useful public ambient-AI cases because it reports statistical significance, shows the vendor-selection process, and separates observed changes from projected revenue. It also represents a nonprofit, nonacademic regional system with seven hospitals and more than 1,200 providers, rather than only a major academic medical center.

    2. The Operational Problem

    McLeod Health faced the same buying problem that many mid-sized health systems face: more than 100 ambient documentation products may be available, but vendor demonstrations are difficult to compare. A polished demo can hide weak specialty performance, poor EHR integration, excessive note editing, or coding problems.

    The organization needed to determine four things before scaling:

    • whether the product captured complex clinical facts accurately;
    • whether the note was readable and usable by physicians;
    • whether documentation supported appropriate coding;
    • whether the tool worked naturally inside Epic.

    This is a procurement and evaluation problem as much as a model problem. A health system that skips structured comparison can lock itself into the wrong workflow, then mistake low adoption for resistance to AI.

    The financial problem also requires care. Time released after hours improves wellbeing but may not create cash savings. Higher patient volume and more complete coding may affect revenue, but the organization must confirm that changes reflect clinically appropriate documentation rather than model-driven inflation.

    3. What Was Built

    McLeod created an evaluation funnel before creating an enterprise rollout.

    System at a Glance

    Simulated evaluation

    Four vendors processed 15 complex outpatient scripts with leaders acting as standardized patients.

    Multidisciplinary scoring

    Physicians, revenue-cycle specialists, and nonclinical reviewers assessed accuracy, billing quality, and readability.

    Workflow demonstration

    The top two vendors demonstrated Epic integration and usability.

    Clinical pilot

    The selected product ran for 90 days across five ambulatory specialties.

    System rollout

    Adoption, documentation, coding, volume, satisfaction, and financial indicators were monitored.

    PhaseDesign
    Simulated evaluationFour vendors processed 15 complex outpatient scripts with leaders acting as standardized patients
    Multidisciplinary scoringPhysicians, revenue-cycle specialists, and nonclinical reviewers assessed accuracy, billing quality, and readability
    Workflow demonstrationThe top two vendors demonstrated Epic integration and usability
    Clinical pilotThe selected product ran for 90 days across five ambulatory specialties
    System rolloutAdoption, documentation, coding, volume, satisfaction, and financial indicators were monitored

    This architecture makes the health system's evaluation set a strategic asset. The 15 scripted encounters create controlled comparisons across products. Live pilot data then show how the chosen system behaves under real audio, specialty language, clinician styles, and workflow pressure.

    A bespoke alternative would preserve the same structure while allowing the enterprise to choose each model layer independently: domain-adapted speech-to-text, diarization, note generation, specialty templates, private RAG, structured validation, and Epic integration. The health system can route specialties or languages to the model that performs best instead of accepting a single vendor's model for every encounter.

    4. How It Reached Production

    The rollout is valuable as a repeatable buyer method.

    Use the same cases for every vendor. Comparable audio and requirements reduce the influence of sales presentation quality.

    Score different stakeholder outcomes. Physicians care about clinical accuracy and edit burden. Revenue-cycle teams care about evidence and coding integrity. Information technology teams care about integration, identity, latency, and support.

    Pilot across heterogeneous specialties. A model that works in primary care may fail in oncology, surgery, pediatrics, or complex multi-speaker encounters.

    Use objective EHR telemetry. Time in notes, after-hours activity, same-day closure, and encounter use are stronger than satisfaction surveys alone.

    Keep statistical and commercial meaning separate. The 35.4% pajama-time decline narrowly missed conventional statistical significance. The $2,629 figure is projected monthly revenue, not an audited realization.

    Before deployment, a health system should define a release threshold for critical clinical errors, unsupported facts, missing facts, edit distance, time to signed note, clinician use, patient consent, and cost per accepted note.

    5. What Healthcare Leaders Should Take Away

    McLeod Health shows that a regional provider can conduct serious AI diligence without the research budget of a large academic institution. The evaluation started with four vendors and 15 controlled cases, moved through a 23-clinician measurement cohort, and reached 81% adoption with more than 150,000 notes.

    A production-ready clinical model program follows this evidence-first pattern. The provider owns the evaluation corpus, clinical glossary, note policies, integrations, security boundary, and acceptance thresholds. Models can change without surrendering the operating system around them.

    The right commercial metric is cost per clinically accepted note at the required safety level, with separate accounting for time released, capacity used, coding effects, and revenue actually realized.

    Sources