Industry: Healthcare
Organization: Providence St Joseph Health
Use case: Ambient clinical documentation
Evidence basis: 2026 JAMA Network Open study using EHR metadata
Disclosure: This is an independent analysis by Conscious Engines. The study was observational, adoption was voluntary, and it evaluated one ambient AI product.
1. Outcome at a Glance
Providence analyzed 16,149 clinician-month observations from 1,547 active users of Microsoft DAX Copilot. Active use required at least 25 ambient-AI encounters in a month. The study used objective EHR metadata from July 2023 through March 2025.
Key Outcomes
7.1 to 6.1 minutes
Median time in notes per appointment
Immediate reduction of 0.26 minutes per note, P<.001 reported in the cited case.
22.0 to 20.6 minutes
Median after-hours documentation
No immediate drop, followed by a sustained monthly decline, P=.02.
12.4 to 12.7
Median appointments per day
No significant immediate or sustained association reported in the cited case.
382.3 to 399.1
Median monthly RVUs
Immediate increase of 7.40 RVUs, P=.03, not sustained reported in the cited case.
5.0 to 5.1
Clinician efficiency profile
No immediate association reported in the cited case.
| Measure | Before and after summary | Interrupted time-series finding |
|---|---|---|
| Median time in notes per appointment | 7.1 to 6.1 minutes | Immediate reduction of 0.26 minutes per note, P<.001 |
| Median after-hours documentation | 22.0 to 20.6 minutes | No immediate drop, followed by a sustained monthly decline, P=.02 |
| Median appointments per day | 12.4 to 12.7 | No significant immediate or sustained association |
| Median monthly RVUs | 382.3 to 399.1 | Immediate increase of 7.40 RVUs, P=.03, not sustained |
| Clinician efficiency profile | 5.0 to 5.1 | No immediate association |
This is a particularly valuable evidence piece because the effects were modest, not spectacular. Providence did not find that ambient AI caused clinicians to see more patients each day. It found a small immediate reduction in time spent on notes, a gradual reduction in after-hours documentation, and an immediate but not sustained RVU association.
2. The Operational Problem
Ambient AI is often sold using self-reported hours saved. Those reports are useful for understanding experience, but memory, enthusiasm, and selection can affect them. EHR event logs provide a more objective view of what clinicians actually did.
Even event-log metrics require interpretation. Time in notes is not total cognitive burden. A physician may spend less time typing but more time reviewing. After-hours activity can reflect schedule, specialty, personal preference, or other workflow changes. Appointment and RVU measures can change for reasons unrelated to the model.
Providence needed to examine whether active use was associated with both immediate and sustained changes, rather than compare two simple snapshots.
The study also shows the difference between availability and adoption. Licences were offered broadly in a health system serving more than two million patients, but active use rose from 0.1% of clinicians in January 2024 to a maximum of 28.8% in March 2025. A technically available tool is not automatically part of clinical practice.
3. What Was Built
Providence integrated DAX Copilot into its clinical environment and connected use data with EHR productivity measures.
System at a Glance
Ambient encounter logs
Identify active clinicians and adoption timing.
EHR event metadata
Measure time in notes and after-hours work.
Appointment data
Test whether capacity changed.
RVU data
Assess a productivity indicator.
Clinician attributes
Segment provider type, specialty, and region.
Interrupted time series
Separate immediate change from trend after adoption.
| Data or system | Evaluation role |
|---|---|
| Ambient encounter logs | Identify active clinicians and adoption timing |
| EHR event metadata | Measure time in notes and after-hours work |
| Appointment data | Test whether capacity changed |
| RVU data | Assess a productivity indicator |
| Clinician attributes | Segment provider type, specialty, and region |
| Interrupted time series | Separate immediate change from trend after adoption |
Most active users were physicians, and almost two-thirds practiced in primary care. That composition matters when applying findings to a specialty-heavy organization.
For a bespoke model program, the evaluation layer should go further. It should join EHR activity with model latency, note acceptance, edit distance, critical clinical errors, encounter complexity, specialty, and model version. This helps identify whether time changes reflect better output or more aggressive copy-forward behavior.
4. How It Reached Production
Providence offers a template for honest post-deployment measurement.
Define active use. Requiring 25 supported encounters in a month is more meaningful than counting licence activation.
Anchor analysis around each clinician's adoption date. Staggered rollout makes a single calendar comparison misleading.
Separate immediate and sustained effects. After-hours improvement emerged gradually, suggesting that learning and workflow integration matter.
Report null results. No appointment-per-day association and no immediate efficiency-profile change constrain the business case.
Use causal language carefully. The design identifies associations. It cannot rule out selection effects or concurrent operational changes.
The study authors explicitly advised measured expectations. That makes this article useful on a commercial site: credible buyers are more likely to trust evidence that includes boundaries.
5. What Healthcare Leaders Should Take Away
Providence's data suggest that ambient AI can reduce documentation burden, but the average effect may be measured in seconds per note rather than hours per day. At enterprise volume, seconds still compound. Value may also emerge through lower cognitive load and after-hours work rather than added appointments.
Objective telemetry belongs in the product contract. A health system should know time to draft, time to accepted note, edit burden, error, use, after-hours work, and cost by specialty. A bespoke medical speech and note system can then be improved where local evidence shows weak performance.
The primary metric remains cost per accepted note at a defined clinical quality threshold. Providence's measured results help set a realistic base case, while TPMG, McLeod, Cleveland Clinic, and SolutionHealth show the range possible under different workflows and adoption levels.
Related Conscious Engines research
- TPMG's ambient AI rollout across 2.58 million encounters
- McLeod Health's evidence-first rollout
- Cleveland Clinic's scaling playbook
- Medical speech recognition and clinical ASR
- Why your evaluation set is your AI moat