Healthcare

AI medical coding is only as accurate as the note underneath it

AI medical coding now auto-generates ICD-10 and CPT codes from the clinical note. But the code inherits the documentation — and a confident code on a thin note books a denial faster.

7 min read

On March 5, AWS introduced Amazon Connect Health, and one capability inside it is the quiet tell for where healthcare AI is heading: a medical-coding agent that generates ICD-10 and CPT codes straight from the clinical note, with confidence scores and source traceability. It’s a real advance. It’s also the clearest example of why AI medical coding will disappoint the organizations that treat it as a coding-accuracy upgrade. The coding step was rarely the weak link. The note it codes from is.

TL;DR: AI medical coding can now read a clinical note and produce ICD-10 and CPT codes with a confidence score attached. The mapping from documented facts to codes is close to solved. But the code is only as right as the note, and notes are the part that’s thin — missing medical-necessity language, an undocumented complexity level, a modifier the record doesn’t support. Point a confident coding agent at an incomplete note and it generates a clean-looking code that gets denied downstream, faster than a human coder would have flagged the gap. Coding is already ~13% of revenue-cycle leakage and denials are near 11.8% and climbing. Fix the documentation-completeness contract first, then run the agent.

This is the same shape I wrote about when an AI customer-service agent turns an internal record into a public promise. Coding is the revenue-cycle version: the agent turns an incomplete note into a submitted claim, and the payer holds the receipt.

The regulator just named the priority, and it isn’t the model

Watch what the people setting the agenda are actually worried about. On June 9, CMS stood up a new Office of Health Technology Products, led by Deputy Administrator and Chief Product Officer Amy Gleason, to run AI, interoperability, and digital-health strategy across its programs — with a stated agenda that includes replacing paper intake forms with digital check-in. Read that against the industry read-out: MarketScale’s July 2026 signal piece put governance, data quality, and interoperability at the top of the healthcare-AI agenda for mid-2026, and framed the limiting factor bluntly — the constraint on healthcare AI is not the sophistication of the model, it’s the quality and consistency of the underlying data.

That’s the operator thesis with a federal office attached to it. The bottleneck is the data the AI runs on. And nowhere is that more literal than coding, where the “underlying data” is a clinician’s note written in the last ninety seconds of a visit.

The coding agent isn’t the hard part. The note is.

Adoption is already moving. A 2025 HFMA and AKASA survey of 519 CFOs and revenue-cycle leaders found 80% of health systems exploring, piloting, or implementing generative AI for the revenue cycle — up from 58% merely considering it in 2023, a 38-point jump in under two years. The demand is real because the pain is real. The question is what the agent inherits when it gets there.

Here’s what a coding agent actually reads: a note. Sometimes an ambient-documentation model wrote it from the visit audio; sometimes the clinician typed three lines between patients. The agent’s job is to map what’s in that note to the right ICD-10 and CPT codes. It does that well. What it cannot do is code the encounter that happened — only the encounter that got documented. If the visit justified a level-four evaluation-and-management code but the note only supports a level three, the agent codes the three, confidently, and the practice quietly undercharges. If the note is missing the medical-necessity language a payer requires, the agent can’t invent it. The clinical reality and the documented reality diverged before the agent ever saw the chart.

MGMA’s own members put numbers on where the money leaks. In an MGMA Stat poll from January 6, 2026 (288 responses), practices named denials and appeals the single biggest revenue-cycle leak at 48%, with coding itself at 13% and front-end issues at 23%. The coding leakage MGMA describes isn’t exotic model failure — it’s undercoding, especially in E/M, plus modifier issues like a Modifier 25 the documentation doesn’t support and gaps such as insufficient medical-necessity language. Those are documentation defects. An agent that codes faster doesn’t fix them. It commits to them at machine speed.

Speed amplifies whatever the note already is

The denials aren’t hypothetical, and they’re getting worse. Kodiak Solutions’ benchmark — drawn from 2,300-plus hospitals and 375,000 physicians — put the initial claim denial rate at 11.8% in 2024, up from 10.2% a few years earlier, and running near that through 2025. Experian Health’s 2025 State of Claims report found 41% of providers now say more than 10% of their claims are denied, up from 30% in 2022 and 38% in 2024. The trend line is the wrong direction, and payers are tightening the edits that trigger denials, not loosening them.

Now drop a coding agent into that environment. Amazon Connect Health’s own framing is that it gets “medical codes ready to review by the time each patient visit ends” and completes billing “in minutes instead of days.” Take that at face value. The minutes you save are the minutes a coder would have spent looking at the note and thinking this doesn’t support that code. Remove the review and you don’t remove the defect — you remove the catch. A confident code on a thin note becomes a submitted claim becomes a denial becomes the appeals work that was already 48% of your leakage. The optimizer got faster. The note it optimizes over didn’t get more complete.

The confidence score is the trap, not the safeguard

The reassuring feature is the dangerous one. A confidence score tells you how well the code matches the note. It tells you nothing about whether the note matches the visit. Those are different questions, and the second one is the one that gets you denied. An agent can be 98% confident it coded the documentation correctly and be wrong about the encounter, because the documentation was incomplete — and the score will read green the whole way to adjudication. Source traceability has the same limit: it lets a coder trace a code back to the sentence that justifies it, which is genuinely useful, but it can’t surface the sentence that was never written. You cannot audit your way to a complete note from an incomplete one.

What the agent does, and what the note still needs

The coding-agent stepWhat it automatesWhat it inherits from the noteWhat owning the documentation adds
Read the encounterParsing the clinical note into structured findingsWhatever the clinician did — or didn’t — write downA completeness standard: what a note must contain per specialty before it’s codeable
Assign ICD-10 / CPTMapping documented facts to codesUndocumented complexity, missing modifiers, thin medical-necessity languageThe justifying detail captured at the point of care, not assumed after
Score confidenceRating code-to-note matchA green score on a note that doesn’t match the visitA separate check on note-to-visit completeness, before coding runs
Submit the claimSending codes to the payerA clean-looking code that may violate a payer-specific editA current payer-rules feed that fails the claim loud before submission

The working version

The fix is unglamorous and it comes before you turn anything on. Define what a complete note looks like — per specialty, because a dermatology visit and a cardiology visit don’t share a bar — so the agent isn’t coding from a summary that’s missing the element that justifies the charge. Structure the documentation capture, including whatever ambient-scribe output feeds it, so the required detail is recorded rather than inferred. Keep the eligibility and payer-edit rules current, because a code can be clinically right and still denied on a specific payer’s rule. And make the coding step fail loud: a note too thin to support the code it would generate should stop and route back to the clinician, not sail through with a confident score.

That’s the same layer under every one of these deployments — the integration and data-contract work that decides whether AI does anything useful. In coding it happens to sit directly on top of your cash flow. Start the agent narrow: high-volume, low-complexity encounters with coder sign-off, and widen only into the visit types where the documentation is reliably complete. Making the note codeable is the project; the coding agent is the last step.

The operator read

AWS built a strong tool, and “generate the code from the note with a confidence score” is an honest description of what it does. The catch is the preposition. It codes from the note — and the note is the thing your organization controls and most hasn’t standardized. A coding agent doesn’t give you a better clinician’s documentation; it gives you faster codes from whatever documentation you already have, denials and all. If you’re about to point one at your claims and you can’t say what “complete” means for a note in each specialty you bill, that’s the conversation to have before the confidence scores start looking reassuring.

FAQ

Is AI medical coding accurate?
An AI medical coding tool is only as accurate as the clinical note it reads. The model can map documented findings to ICD-10 and CPT codes very reliably — that part is close to solved. What it can't do is code what the clinician never wrote down. If the note omits the medical-necessity language, understates the complexity of the visit, or is missing the detail that justifies a higher E/M level, the agent produces a clean, confident code for the thin note it was given. The accuracy problem was never the coding step; it's the documentation underneath it. Coding is already the source of about 13% of revenue-cycle leakage on files humans code today (MGMA Stat, January 2026), and an agent inherits that gap rather than closing it.
Will AI medical coding reduce or increase claim denials?
It does whichever the documentation already points at, faster. Initial claim denial rates are already near 11.8% and rising (Kodiak Solutions, 2025), and 41% of providers now report that more than 10% of their claims are denied (Experian Health, 2025). A coding agent pointed at complete, consistent notes will cut denials by coding them right the first time. Pointed at notes with documentation gaps — insufficient medical-necessity language, a Modifier 25 the record doesn't support, an undercoded E/M level — it generates codes that look clean and get denied downstream, at machine speed. Speed is neutral. It amplifies whatever the note already is.
What is Amazon Connect Health's medical coding feature?
Amazon Connect Health is an agentic AI service AWS announced on March 5, 2026. Its medical-coding capability, in preview, generates ICD-10 and CPT codes directly from clinical notes with confidence scores and full source traceability, so a coder can validate each prediction against the underlying documentation. Its ambient-documentation feature, generally available, writes the clinical note from the patient-clinician conversation across 22+ specialties. It's a capable tool. What it can't ship you is a documentation standard that guarantees the note it codes from is complete — that's still the provider organization's job.
What data does an AI coding agent need to work reliably?
A complete clinical note and a clean payer-rules feed. The note has to actually contain what the encounter justified — the complexity, the medical necessity, the modifiers' supporting detail — not just a shorthand summary. And the eligibility and payer-specific edit rules have to be current, because a code can be clinically valid and still violate a specific payer's edit and get denied. Both are data-contract problems, not model problems. Define what a 'complete' note looks like per specialty, keep the payer rules fresh, and make the pipeline flag an incomplete note before it codes — rather than coding it confidently and finding out at adjudication.
How should a health system prepare before deploying an AI medical coding agent?
Fix the documentation layer before you point the agent at the claim. Define what 'complete' means for a note in each specialty, so the agent isn't coding from a summary that's missing the justifying detail. Structure intake and ambient-documentation output so the required elements are captured, not assumed. Keep the eligibility and payer-edit rules current and make the coding step fail loud when a note is too thin to support the code it would generate — so a gap stops the claim instead of flowing into a denial. Then start narrow: high-volume, low-complexity encounters with coder sign-off, and widen only where the documentation is reliably complete.