CASE STUDY

AI Codes Alongside ED Experts: Results From
1,000 Emergency Department Charts

Emergency department coding puts autonomous medical coding under one of its toughest tests.

To evaluate whether Amy AI could perform at expert-human standards, CombineHealth and Edelberg & Associates' expert coders independently coded 1,000 ED charts, with Edelberg's audit team evaluating their performance.

98%+
Coding accuracy
~85%
Charts coded autonomously
~5×
Documentation gaps identified
Get the Full Case Study
OUTCOMES

CombineHealth Matched or Outperformed Expert Human Coders

Higher accuracy
Outperformed expert coders on E/M and ICD-10 accuracy
~85%
Charts coded autonomously by CombineHealth's medical coding platform
50%
Faster coding turnaround as compared to human coders
~5x
More CDI-triggering documentation gaps identified by CombineHealth
01 · THE TEST

A Deliberately Difficult Test: High-Complexity ED Environment

Emergency department coding compresses significant complexity into a high-volume workflow. Coders must interpret complete encounters, determine E/M levels, assign procedures and modifiers, select defensible ICD-10 diagnoses, evaluate medical necessity, and work through documentation that can support more than one compliant coding interpretation.

That complexity made the ED a rigorous environment for evaluating whether CombineHealth's autonomous medical coding platform, Amy AI, could perform at expert-human standards.

For 15 days, approximately 100 charts per day were assigned separately to CombineHealth and Edelberg's expert coding team, covering 1,000 emergency department charts in total.

The outputs were not judged simply by comparing CombineHealth's codes with the human coders' codes. Instead, Edelberg's audit team independently reviewed coding decisions against defined criteria, including documentation support, coding guidelines, specificity, sequencing, medical necessity, and applicable payer requirements.

This distinction was particularly important for ICD-10, where two different coding decisions can both be defensible when supported by the clinical documentation and official coding guidelines. The evaluation therefore measured whether each coding decision was supportable—not whether AI and human coders always produced identical outputs.

02 · THE PERFORMANCE

CombineHealth Matched Expert-Level Coding Performance

Across the evaluation, CombineHealth achieved 95.9% E/M accuracy, 99% CPT accuracy, 100% modifier accuracy, and 93.2% ICD-10 accuracy. Edelberg's expert human coding team recorded 94.7%, 99%, 100%, and 92.2%, respectively.

At the same time, approximately 85% of charts were coded autonomously, while approximately 15% were escalated when the platform detected ambiguity, conflicting documentation, or insufficient clinical support.

03 · THE BIGGER FINDING

~5× More CDI-Triggering Documentation Gaps Identified

CombineHealth identified approximately 5× more CDI-triggering documentation issues than the traditional workflow.

These included missing critical care time, incomplete procedure details, missing independent interpretation of tests, missing documentation of external physician discussions, and insufficient medical necessity support.

This reflects how CombineHealth's automation platform approaches coding: analyzing the complete encounter, applying the medical coding methodology, evaluating documentation sufficiency and medical necessity, and incorporating payer-aware intelligence into billing-ready coding decisions.

04 · THE IMPACT

Operational Impact From the Evaluation

CombineHealth Matched or Outperformed Expert Human Coders

Across every measured coding dimension in the 1,000-chart evaluation

~85%
Charts coded autonomously by CombineHealth's medical coding platform
50%
Faster coding turnaround

~12 hours with the platform vs. ~24 hours with human coders

~5×
More CDI-triggering documentation gaps identified by CombineHealth

The evaluation was designed to determine whether autonomous medical coding could operate at expert-human standards in a complex, high-volume environment—not to prove that AI was better than human coders.

The results showed that CombineHealth could deliver expert-level coding performance while increasing autonomous execution, reducing turnaround time, and identifying substantially more documentation gaps.

Scroll to Explore the Case Study
OUR PLATFORM. REAL IMPACT.

Go Behind the Results

The full case study shows what happened inside the 1,000-chart evaluation: where CombineHealth and expert coders reached different decisions, how Edelberg's audit team determined what was defensible, and where the platform uncovered documentation gaps that traditional coding workflows missed.

See the detailed methodology, coding-level findings, documentation patterns, and operational takeaways behind the headline results.

Get the Full Case Study