Why Published Accuracy Figures Are Hard to Compare

Every vendor publishes an accuracy number and almost none define it identically. Some measure word-level transcription accuracy, which is nearly meaningless for a system composing structured notes. Some measure against a reference note written by a clinician, which depends heavily on who wrote the reference. Some report internal QA pass rates. These are three different quantities wearing the same label, which is why the useful comparison is one you run yourself on your own encounters.

Not All Errors Are Equal

Weight errors by consequence rather than counting them evenly. A misspelled ordinary word is cosmetic. A wrong laterality, incorrect medication, wrong dose, or inverted finding — hyper for hypo — is clinically significant. A hallucinated finding that was never discussed is the most serious category, because it is fluent and therefore easy to sign past. A sensible scoring approach counts significant errors per note separately from cosmetic ones and reports them separately.

Omission Is the Error People Forget to Measure

Systems are usually tested for what they got wrong and rarely for what they left out. Omission is harder to spot precisely because nothing on the page looks incorrect — the counselling that was delivered, the option that was considered and declined, the instruction that was given. Build omission into your test explicitly: after reading the generated note, list what you remember discussing that is not there. Our AI clinical documentation is structured to your template partly so that expected sections make omissions visible.

A Practical Test Protocol

Twenty consecutive real encounters, not curated ones. For each, read the generated note and record: significant errors, cosmetic errors, omissions of material content, structural placement errors, and review time in seconds. Twenty encounters is enough to see a pattern and small enough that a single clinician can complete it in a week. Include your messiest visit types deliberately — interrupted visits, multi-problem encounters, difficult audio.

Interpreting the Results

What matters is the shape of the errors rather than a single percentage. Consistent structural placement errors mean template configuration is wrong and is fixable. Scattered terminology errors in one specialty suggest the system is not specialty-aware for your field. Hallucinated content, even rarely, is the finding that should stop a purchase pending explanation. And review time is the practical bottom line: if it exceeds a couple of minutes, the tool has relocated work rather than removed it.

Re-Test After Configuration

First-week results reflect an unconfigured system as much as the underlying technology. Run the protocol once early to establish a baseline, adjust templates and preferences based on what you find, then re-run it. The delta tells you how much of the initial error rate was configuration — usually a substantial share — and gives you a defensible number for the decision.

Keep Measuring After Go-Live

Accuracy is not a one-time acceptance test. Spot-check five notes monthly against the same criteria, and watch review time as the leading indicator. Drift in either is worth raising early rather than discovering during an audit. If you want a baseline on your own encounters at no cost, our free trial is designed for exactly this kind of evaluation.