AI
September 9, 2026

Someone Still Has to Run the Workflow

What Epic’s AI integration changes for operations – and how I evaluate speed to impact

Share this post

A capability is not a workflow. That distinction is about to cost some organizations a year.

OpenAI has connected Epic to ChatGPT for Healthcare. Authorized clinicians can pull clinical notes, lab results, medications, and specialist documentation from the record into ChatGPT to review histories and prepare for appointments, or they can work with ChatGPT inside supported Epic workflows without leaving the chart. The connection is read-only; nothing is written back. A companion plugin reaches nine authoritative public sources – ClinicalTrials.gov, CMS Coverage, RxNorm, DailyMed, and PubMed among them – at the level of specific records, fields, identifiers, and versions. Role-based access, single sign-on, audit logs, and a business associate agreement carry the enterprise controls. UCSF Health is a pilot partner.

The evaluation work behind it is substantial. Physicians rated responses across 27 connected-record use cases – pre-visit review, clinical timelines, medication review, handoff summaries – and 99.1% of 4,363 ratings came back safe. Across five connected data sources, more than 93% of responses were rated good or better by accuracy. Those are vendor-reported figures and they are strong ones.

What the announcement does not contain is a workflow. Every Epic client in the country receives the same capability, at the same time, on the same terms. AdventHealth’s chief AI officer made the point in the same release: the value of AI starts with people, and the goal is to move practical innovation into action. Capability arrives from the vendor. Performance is built on site, by an operations team, against a baseline someone had the discipline to capture first.

What follows is the framework I use with Epic clients adopting these capabilities: the operating areas, what fails in each, how I sequence the work, and how I decide when to expect a number.

The framework: six operating areas & a measurement lag

AI capability inside the chart reaches a dollar only by passing through the record and everything downstream of it. There are six places where that passage either holds or breaks. A defect at any one of them is invisible at the other five, which is why an organization can report enthusiastic adoption and unchanged financial performance in the same quarter.

# Operating area The question it answers Primary failure mode
1 Documentation standard Did the clinician’s better understanding reach the note? Informed physician, unchanged documentation
2 Code-level specificity Does the record support the code, not just the diagnosis? Summary reads well; approach, levels, and laterality still absent
3 Access and permitted use Can the roles that do revenue cycle work reach it? Clinician-only deployment; coding, auth, and denials excluded
4 Authorization handoff Does the clinical timeline become an auth packet? Necessity history assembled for the visit and lost before submission
5 Denial and appeal support Are we citing the policy in force on the date of service? Public-source citation without the payer’s own policy version
6 Verification threshold Which outputs require a human, and who decided? Undefined review line discovered during a payer audit
Measurement lag When should we expect a number, and from where? Thirty-day ROI claims that measure adoption, not impact

The measurement lag is not a seventh area. It sits underneath all six, and it is the difference between a program that survives its first quarterly review and one that does not.

Where the operating work sits

Operating area 1: The read-only boundary and documentation quality

The medical record is a clinical document, a legal document, and a financial document at once. Every code, every appeal, every audit response traces to what a physician wrote and attested to. Read-only preserves that chain – the clinician remains the author of record, and nothing enters the chart that a physician did not write. Treat that boundary as a design principle to protect, not a first-version gap waiting to be closed. It is the reason the tool is usable at all in a billing environment.

The boundary also caps the benefit: the integration cannot improve documentation. It can only improve the clinician’s understanding of what is already there. Whether that understanding reaches the note is a workflow question, answered by documentation standards, physician education, and CDI review – not by the integration.

What we do: set the documentation standard the specialty requires, run coding and CDI review against it, and measure note specificity before and after deployment so the organization can tell whether anything changed.

Operating area 2: Summarization is not documentation

This is the distinction most organizations will get wrong, and it fails quietly. A model that reads a chart well produces something useful for a clinician to read. It does not produce something a payer will pay on.

Specificity is authored, not summarized: the surgical approach rather than the diagnosis, levels and laterality, decompression distinguished from fusion, the elements that establish medical decision-making complexity, and the necessity thread connecting a symptom to a conservative care history to an imaging finding to a procedure. None of that appears because the summary was good. It appears because a physician stated it.

The failure mode looks like success: the physician enters the visit better informed, the note comes out unchanged, adoption metrics are excellent, and the coding queue is identical three months later.

What we do: translate model output into a documentation standard – what the pre-visit summary should prompt the physician to state, in the terms the code set, and the payer policy require – and keep a coder in the loop on the specialties where the money moves.

Operating area 3: The access model decides who benefits

The integration inherits role-based access, single sign-on, and chart permissions already configured. That is the correct security posture, and it also quietly determines who in the organization gets anything out of this.

In most groups this is clinician-facing: coders, CDI specialists, prior authorization staff, and AR analysts sit in different security classes with narrower chart access, by design and for good reason. Left alone, the revenue cycle receives a second-order benefit at best.

The question is not whether you have AI in Epic: it is which roles can see it, whether that maps to where the work sits, and what the access review and permitted-use policy have to say before you extend it. Decide that deliberately rather than letting a license bundle decide it.

What we do: run the role mapping against real revenue cycle workflows, build the access recommendation with compliance and IT, and write the permitted-use policy – what it may be used for, what it may never be used for, who approves exceptions, and how use is monitored.

Operating area 4: Pre-visit review is an authorization asset first

For referral-heavy specialties, the highest-value output of a clinical timeline is not the visit. It is the authorization. Prior imaging, prior conservative care, prior injections and the response to them, prior surgical history – that is exactly the material determining whether an authorization is approved on first submission and whether the claim survives a medical necessity review. Today it lives across scanned packets, faxed summaries, and outside-record repositories, assembled by hand and often incomplete by the time the decision has to be made.

One caution, and it is a real one: outside-record retrieval operates under its own governance – network participation agreements, framework rules of the road, and purpose-of-use limitations that generally restrict retrieval to treatment. Those rules govern the AI layer exactly as they govern any other access path. A summarization tool does not expand the permitted purpose of the underlying data.

What we do: design the handoff – how a clinical timeline becomes a structured authorization packet, who touches it, where it enters the queue – and get the purpose-of-use question answered by counsel before the workflow is built rather than after.

Operating area 5: The public data plugin is a denials asset

PubMed, ClinicalTrials.gov, RxNorm, DailyMed, and CMS Coverage are being read as clinical reference tools. Four of the five are also denial management resources, and the fifth changes the calculus entirely.

Version-level coverage policy is the material point: the ability to work with the specific policy in force on the date of service is the difference between an appeal that argues and an appeal that cites. Drug labeling supports coverage arguments on infusion and injectable therapy. RxNorm supports clean medication identification across systems. Trial registry status bears directly on whether a payer treats a service as investigational. Literature is frequently what a peer-to-peer turns on when a policy is silent, or the patient sits off pathway.

The limit is equally clear: commercial payer medical policy is not in those datasets. National coverage material is not a substitute for the policy a commercial plan published, and neither is a substitute for reading it.

Note who the natural user is: this is a denial and appeals capability sitting inside a clinician-facing tool, which returns directly to the access question above.

What we do: build the appeal architecture that pairs a public-source citation with the payer’s own policy version and make sure the denials team – not only the clinician – can reach it.

Operating area 6: The accuracy numbers tell you where the human belongs

Read the evaluation figures carefully, because they are more useful than when they first appear. A 99.1% safety rating across 4,363 physician ratings is a strong result for clinical synthesis. More than 93% rated good or better on accuracy across five connected data sources is also strong – for clinical synthesis.

It is not a rate you would accept on a determination: framed as a revenue cycle threshold, better than 93% means a meaningful minority of outputs are not good, and in this domain those outputs carry a dollar value and an audit trail. That is not a criticism of technology. It is a specification.

Numbers like these are design inputs, not guarantees: they tell an operator exactly where verification is required and where it is not. The organizations that do well here are the ones that write that line down deliberately instead of discovering it during a payer audit.

What we do: define which outputs require human verification and which do not, stand up the sampling and QA program behind it, and make sure the audit trail can answer the question a payer will eventually ask about how a determination was reached.

How I evaluate speed to impact

When a capability like this arrives, the question I am asked is whether it works. The question I need answered is how long before it shows up in an indicator I already report, and what must be true along the way. In a clinical revenue cycle, nothing reaches the dollar without passing through the record, and every step between the tool and the claim adds lag and dilutes attribution. I evaluate four questions, in this order:

How many hands sit between the output and the dollar?

Count the hops. A public-source citation that strengthens appeal is two hops from a payer decision. A pre-visit summary that improves a note is five: model to clinician, clinician to note, note to coder, coder to claim, claim to adjudication. Both are worth doing. They are not the same investment on the same reporting horizon and treating them as though they are how an AI program loses its sponsor.

Does it change a decision, or only inform one?

Informing is real and unmeasurable. I look for the decision that comes out different – an authorization submitted with the conservative care history attached, a code selected on approach rather than diagnosis, an appeal filed against the policy version in force on the date of service. If I cannot name the decision, I cannot measure the impact, and I should not be promising one.

Does an indicator I already report move?

If seeing the improvement requires inventing a new metric, impact is further away than it looks and the attribution argument will be weak when finance asks. First-pass authorization approval, medical necessity denial rate, coding queue aging, over-90 AR, appeal overturn rate. If one of those moves, the case makes itself.

What is the natural lag?

Use case Hops to the dollar Indicator that moves Signal window Ceiling
Appeal and peer-to-peer citation support 2 Appeal overturn rate; medical necessity write-offs 30–60 days Moderate
Medication identification and labeling 2–3 Infusion and injectable billing accuracy; drug denial rate 45–75 days Moderate
Pre-visit review into the authorization packet 3 First-pass authorization approval; peer-to-peer volume 60–90 days High
Clinical timeline for necessity defense 3–4 Medical necessity denial rate; appeal cycle time One to two quarters High
Chart summarization improving documentation 5 Coding queue aging; code distribution; documentation-driven denials Two quarters or more Highest
Handoff summaries Clinical only No direct revenue cycle indicator Not applicable Indirect

Signal windows are directional, drawn from typical adjudication and appeal cycles. Tune them to your own payer mix and charge lag before committing them in front of a board.

The practical consequence is a sequencing rule, and it runs against intuition. The use cases with the highest ceiling carry the longest signal delay. Documentation quality is the largest prize in this entire announcement, and it is also the one that will take two quarters to prove. The use cases with the fastest signal – appeals support, medication identification – have modest ceilings by comparison.
So, start where the hops are fewest, not where the prize is biggest. You need a measured win inside one cycle to hold the sponsorship that funds the two-quarter work. Programs that lead with the largest prize tend to get defunded about a month before their evidence would have arrived.

How this creates value

We are not implementing Epic, and we are not selling the model. We do the operating work between them – the part that determines whether a capability produces a measurable revenue cycle result or a well-received pilot that changes nothing. Concretely, that means five things:

  1. Baseline before anything goes live: first-pass authorization approval, medical necessity denial rate, coding queue aging, and time from encounter to closed note. Without a pre-period, no improvement can be attributed to anything, and leadership will ask.
  2. Design the access and the policy: role mapping across clinical and revenue cycle functions, the access recommendation, and the written permitted-use policy. One page is enough. Zero pages is not.
  3. Build the handoffs, not the enthusiasm: summary to documentation standard, timeline to authorization packet, public-source citation to appeal. Named owners, defined queues, real turnaround expectations.
  4. Draw the verification line: which outputs require human review, and which do not, with sampling behind it and an audit trail that holds up when a payer asks how a determination was reached.
  5. Report on the cadence the lag supports: post-period against the same baseline, at the KPI level, with an honest read on what moved and what did not. That cadence is slower than the one the organization wants and considerably more defensible.

A model that reads the chart well is worth having, and read-only is the right first constraint – it protects the one thing the revenue cycle cannot function without. But an organization that mistakes reading the record for improving it will find out at the denial.
If you are a healthcare executive, practice administrator or physician leader at an Epic organization adopting these capabilities, start with three questions this week:

  1. Which of our revenue cycle roles can see this capability today?
  2. What was our first-pass authorization approval rate the week before go-live?
  3. Which outputs require a human to verify, and who wrote that down?

The value of AI in the chart will not be measured in clinician adoption. It will be measured in whether the documentation got more specific, the authorization was approved the first time, and the claim went out correct. The technology is new. The test is not.

Source: Fierce Healthcare, “ChatGPT for Healthcare unveils new integrations with Epic EHR, public health data.” Evaluation figures are as reported by OpenAI. PracticeCore is not affiliated with, and does not speak for, Epic Systems or OpenAI.

Moses B. Landon, MBA, is Head of Revenue Cycle Management at PracticeCore, where he leads our organization and operations for specialty practices. He has more than two decades of senior executive experience in healthcare administration and margin improvement, AI RCM strategy, and enterprise transformation.