We’re putting $5.5M behind research and development of smaller, specialized AI models for real-world deployment — Read our manifesto →
    All posts

    [Case Study] What the EHR Data Actually Showed: Providence's 1,547-Clinician Ambient AI Study

    A large real-world analysis found modest documentation gains, no increase in daily appointments, and a gradual reduction in after-hours work.

    Conscious Engines

    Industry: Healthcare
    Organization: Providence St Joseph Health
    Use case: Ambient clinical documentation
    Evidence basis: 2026 JAMA Network Open study using EHR metadata
    Disclosure: This is an independent analysis by Conscious Engines. The study was observational, adoption was voluntary, and it evaluated one ambient AI product.

    1. Outcome at a Glance

    Providence analyzed 16,149 clinician-month observations from 1,547 active users of Microsoft DAX Copilot. Active use required at least 25 ambient-AI encounters in a month. The study used objective EHR metadata from July 2023 through March 2025.

    Key Outcomes

    7.1 to 6.1 minutes

    Median time in notes per appointment

    Immediate reduction of 0.26 minutes per note, P<.001 reported in the cited case.

    22.0 to 20.6 minutes

    Median after-hours documentation

    No immediate drop, followed by a sustained monthly decline, P=.02.

    12.4 to 12.7

    Median appointments per day

    No significant immediate or sustained association reported in the cited case.

    382.3 to 399.1

    Median monthly RVUs

    Immediate increase of 7.40 RVUs, P=.03, not sustained reported in the cited case.

    5.0 to 5.1

    Clinician efficiency profile

    No immediate association reported in the cited case.

    MeasureBefore and after summaryInterrupted time-series finding
    Median time in notes per appointment7.1 to 6.1 minutesImmediate reduction of 0.26 minutes per note, P<.001
    Median after-hours documentation22.0 to 20.6 minutesNo immediate drop, followed by a sustained monthly decline, P=.02
    Median appointments per day12.4 to 12.7No significant immediate or sustained association
    Median monthly RVUs382.3 to 399.1Immediate increase of 7.40 RVUs, P=.03, not sustained
    Clinician efficiency profile5.0 to 5.1No immediate association

    This is a particularly valuable evidence piece because the effects were modest, not spectacular. Providence did not find that ambient AI caused clinicians to see more patients each day. It found a small immediate reduction in time spent on notes, a gradual reduction in after-hours documentation, and an immediate but not sustained RVU association.

    2. The Operational Problem

    Ambient AI is often sold using self-reported hours saved. Those reports are useful for understanding experience, but memory, enthusiasm, and selection can affect them. EHR event logs provide a more objective view of what clinicians actually did.

    Even event-log metrics require interpretation. Time in notes is not total cognitive burden. A physician may spend less time typing but more time reviewing. After-hours activity can reflect schedule, specialty, personal preference, or other workflow changes. Appointment and RVU measures can change for reasons unrelated to the model.

    Providence needed to examine whether active use was associated with both immediate and sustained changes, rather than compare two simple snapshots.

    The study also shows the difference between availability and adoption. Licences were offered broadly in a health system serving more than two million patients, but active use rose from 0.1% of clinicians in January 2024 to a maximum of 28.8% in March 2025. A technically available tool is not automatically part of clinical practice.

    3. What Was Built

    Providence integrated DAX Copilot into its clinical environment and connected use data with EHR productivity measures.

    System at a Glance

    Ambient encounter logs

    Identify active clinicians and adoption timing.

    EHR event metadata

    Measure time in notes and after-hours work.

    Appointment data

    Test whether capacity changed.

    RVU data

    Assess a productivity indicator.

    Clinician attributes

    Segment provider type, specialty, and region.

    Interrupted time series

    Separate immediate change from trend after adoption.

    Data or systemEvaluation role
    Ambient encounter logsIdentify active clinicians and adoption timing
    EHR event metadataMeasure time in notes and after-hours work
    Appointment dataTest whether capacity changed
    RVU dataAssess a productivity indicator
    Clinician attributesSegment provider type, specialty, and region
    Interrupted time seriesSeparate immediate change from trend after adoption

    Most active users were physicians, and almost two-thirds practiced in primary care. That composition matters when applying findings to a specialty-heavy organization.

    For a bespoke model program, the evaluation layer should go further. It should join EHR activity with model latency, note acceptance, edit distance, critical clinical errors, encounter complexity, specialty, and model version. This helps identify whether time changes reflect better output or more aggressive copy-forward behavior.

    4. How It Reached Production

    Providence offers a template for honest post-deployment measurement.

    Define active use. Requiring 25 supported encounters in a month is more meaningful than counting licence activation.

    Anchor analysis around each clinician's adoption date. Staggered rollout makes a single calendar comparison misleading.

    Separate immediate and sustained effects. After-hours improvement emerged gradually, suggesting that learning and workflow integration matter.

    Report null results. No appointment-per-day association and no immediate efficiency-profile change constrain the business case.

    Use causal language carefully. The design identifies associations. It cannot rule out selection effects or concurrent operational changes.

    The study authors explicitly advised measured expectations. That makes this article useful on a commercial site: credible buyers are more likely to trust evidence that includes boundaries.

    5. What Healthcare Leaders Should Take Away

    Providence's data suggest that ambient AI can reduce documentation burden, but the average effect may be measured in seconds per note rather than hours per day. At enterprise volume, seconds still compound. Value may also emerge through lower cognitive load and after-hours work rather than added appointments.

    Objective telemetry belongs in the product contract. A health system should know time to draft, time to accepted note, edit burden, error, use, after-hours work, and cost by specialty. A bespoke medical speech and note system can then be improved where local evidence shows weak performance.

    The primary metric remains cost per accepted note at a defined clinical quality threshold. Providence's measured results help set a realistic base case, while TPMG, McLeod, Cleveland Clinic, and SolutionHealth show the range possible under different workflows and adoption levels.

    Sources