We’re putting $5.5M behind research and development of smaller, specialized AI models for real-world deployment — Read our manifesto →
    All posts

    Healthcare AI in Production: Eight Named Deployments With Measured Results [Evidence Index]

    A source-ranked guide to ambient documentation, medical speech AI, and patient-access voice agents across leading health systems.

    Conscious Engines

    Healthcare leaders do not need another list of theoretical AI use cases. They need evidence that a system worked with real clinicians, patients, EHRs, contact centers, privacy controls, and operating constraints.

    This index organizes eight named deployments by evidence strength. It includes peer-reviewed studies, health-system reports, and vendor-supported customer cases. Those sources are not interchangeable. Randomized evidence can establish a stronger causal signal. Large observational deployments show whether a system can survive routine practice. Vendor cases can reveal useful operating metrics but require independent verification.

    Disclosure: Conscious Engines did not deliver the implementations described. This is an independent synthesis of public evidence intended to help enterprise buyers design and evaluate their own systems.

    1. The Results in One View

    OrganizationWorkflowScaleHeadline resultEvidence type
    UCLA HealthAmbient documentation238 physicians randomizedNabla reduced time in notes 9.5% versus control; DAX difference was not significantRandomized trial, NEJM AI
    ProvidenceAmbient documentation1,547 active clinicians, 16,149 clinician-monthsMedian time in notes fell from 7.1 to 6.1 minutes; no appointment-per-day associationPeer-reviewed observational study, JAMA Network Open
    McLeod HealthAmbient documentation23 clinicians in core pilot analysis, more than 150,000 later notesTime in notes down 28.3%, 81% system adoptionPeer-reviewed implementation study
    TPMGAmbient documentation7,260 physicians, 2.58 million encounters15,791 estimated documentation hours savedLarge observational evaluation
    Cleveland ClinicAmbient documentationMore than 4,800 users, 3.5 million encounters70% encounter-level use among established usersPeer-reviewed implementation report
    SolutionHealthAmbient and active medical speechMore than 1,000 clinicians, nearly 60,000 encounters56% improvement in average documentation timeVendor-supported named customer case
    Inova HealthPatient-access voice agentsAbout 338,000 calls per month50% appointment calls resolved, 4,272 hours released monthly, 8.8x reported ROIVendor-supported named customer case
    Houston MethodistVaccine information and scheduling voice agentMore than 200,000 calls in one month91% automation, every call answered on first ringHealth-system case study

    The results establish three points.

    First, healthcare AI is already operating at enterprise scale. TPMG and Cleveland Clinic together report more than six million ambient-supported encounters. Second, average effects vary sharply by system, clinician, workflow, and measurement method. Providence and the UCLA randomized trial found modest documentation-time changes, while SolutionHealth and McLeod reported larger improvements. Third, voice AI can create value outside the clinical note by completing patient-access work at call-center scale.

    2. What the Evidence Actually Supports

    Ambient documentation can reduce work, but not equally for everyone

    The strongest causal evidence in this set comes from UCLA's randomized trial of 238 physicians, assigned to DAX Copilot, Nabla, or a control group. Nabla users showed a statistically significant 9.5% reduction in time in notes versus control. DAX users did not show a significant difference on the primary time measure. Both products improved several physician-experience measures. Clinicians reported occasional clinically significant inaccuracies, and one mild patient-safety event was identified.

    Providence provides scale and objective EHR telemetry. Its study found a small immediate reduction in time per note, followed by a gradual decrease in after-hours documentation. It found no significant association with daily appointment count.

    These studies set a credible base case: ambient AI can help, but the value may appear in cognitive load, patient attention, or time returned after work rather than additional scheduled volume.

    Adoption and utilization determine realized value

    McLeod reported 81% adoption. Cleveland Clinic reported 70% encounter-level utilization among established users, after onboarding more than 4,000 clinicians in 16 weeks. TPMG's highest-use group activated ambient AI in 89% of eligible encounters.

    These measures are not identical. Adoption can mean a user tried the system. Encounter-level utilization measures how often it is used when it could be used. Accepted-note rate measures how often the output survived review. Buyers should demand all three.

    Voice agents create a separate healthcare business case

    Inova's reported 8.8x return came from patient-access voice agents integrated with Epic, CRM, and telephony. Houston Methodist used a voice assistant to absorb a public-health demand spike, including 14,583 calls in one day.

    The clinical-risk boundary is different from ambient notes. Appointment scheduling, modification, provider search, FAQs, and routing can often be constrained through deterministic workflows. Emergency symptoms, clinical advice, complex identity cases, and uncertain transactions should move to trained staff.

    3. A Buyer's Evidence Ladder

    Healthcare buyers should grade claims before using them in a board paper.

    Evidence levelWhat it can establishExamples in this indexMain limitation
    Randomized controlled trialStronger causal comparisonUCLAShort period, single institution, selected physicians
    Peer-reviewed EHR observational studyObjective real-world association at scaleProvidenceVoluntary use and confounding
    Peer-reviewed implementation studySelection, workflow, adoption, and rollout evidenceMcLeod, Cleveland ClinicOften no randomized counterfactual
    Large health-system evaluationFeasibility and experience at production scaleTPMGSelection and survey bias
    First-party health-system caseOperating performance in a named workflowHouston MethodistLimited external validation
    Vendor-supported named customer caseCommercial implementation and metricsSolutionHealth, InovaStrong incentive to emphasize favorable results

    No level should be ignored. A randomized two-month trial cannot prove five-year maintainability. A vendor case cannot prove causality. Together they define a more useful range of expected outcomes and failure modes.

    Every claim should retain its qualifier:

    • McLeod's $2,629 per provider per month was a projected revenue gain.
    • Providence found no increase in appointments per day.
    • Inova's 8.8x ROI formula was not publicly disclosed.
    • Houston Methodist's results came from the exceptional 2021 vaccine context.
    • SolutionHealth's 2.5 additional appointments was a capacity equivalent, not confirmed booked volume.
    • TPMG's time saving was estimated from observed documentation behavior.

    4. How to Evaluate a Bespoke Healthcare Model

    The evidence suggests a 90-day evaluation should compare the bespoke model against the current workflow and at least one credible alternative.

    For medical speech and ambient documentation, measure:

    • clinical term and medication accuracy;
    • speaker-attribution and negation errors;
    • unsupported and omitted clinical facts;
    • time in notes and after-hours EHR activity;
    • edit distance and time to signed note;
    • clinician adoption, encounter-level use, and accepted-note rate;
    • performance by specialty, accent, language, and encounter complexity;
    • cost per accepted note.

    For patient-access voice agents, measure:

    • intent and critical-field accuracy;
    • correctly completed appointment or administrative task;
    • transfer, abandonment, and rapid repeat contact;
    • identity and authorization failures;
    • emergency and clinical escalation performance;
    • patient satisfaction by language and access cohort;
    • cost per durable resolution.

    The model stack should match the task. Medical ASR converts speech. A note model structures the encounter. Retrieval supplies current institutional knowledge. Deterministic checks validate critical fields. A voice agent manages dialogue and invokes approved actions. A routing layer sends only difficult cases to a larger model.

    The enterprise should own the evaluation set even when it buys model APIs. That corpus contains the specialties, language, failure cases, and business rules that make the system trustworthy locally.

    5. What This Means for Healthcare Leaders

    The public record no longer supports asking whether healthcare AI can reach production. It has. The better questions are:

    1. Which workflow has a measurable denominator and a tolerable automation boundary?
    2. Which evidence applies to our specialties, languages, systems, and baseline burden?
    3. What outcome will finance and clinical governance accept as real value?
    4. Which parts of the model stack should we own to protect data, control cost, and improve locally?
    5. How will we detect regression when a model, prompt, or source document changes?

    Conscious Engines builds bespoke healthcare AI systems around the enterprise's own language and operating environment. That can include medical speech-to-text, text-to-speech, clinical voice agents, private enterprise RAG, task-specific small language models, EHR integration, and continuous evaluation.

    The commercial thesis is precise: use the smallest capable model for each step, keep clinical authority with qualified people, and measure cost per accepted outcome. The evidence above provides external benchmarks. The deployment still has to earn trust on the client's own data.

    Read the full case studies

    Primary sources