The job moved. It did not vanish.
Physicians still talk. Someone still has to turn that talk into a chart that another clinician can trust at 2 a.m. What changed is where the first draft comes from. For twenty years a transcriptionist listened and typed. Now a speech model produces a rough text in seconds, and a person spends their time on the sentences the model got wrong. If your vendor's slide says 'no humans,' ask who fixes 'left' when you said 'right.'
Medical transcription was never just typing speed. It was knowing that 'LAD' in a cardiology dictation is not the same problem as 'lad' in a casual sentence, that a mumbled dose is more dangerous than a mumbled adjective, and that your EHR template has a place for the assessment that a raw paragraph will miss. AI is good at the common sentence. The chart is full of uncommon sentences.
Clinics that threw out transcription and kept only a raw speech-to-text box usually discovered the hidden work: the physician became the transcriptionist. That is cheaper on an invoice and expensive on a Tuesday evening. The useful version of AI in this workflow is a first pass that a trained editor can finish faster than they could have typed it, not a product that pretends the edit is optional.
A transcript is not a note
Speech recognition writes down what it heard, in order, with the pauses removed. A clinical note has a shape: history, exam, data, assessment, plan. A physician who dictates in that shape already is easy to help. A physician who dictates a story, then a lab, then goes back to correct the story, produces audio that a model will linearize into something that looks fluent and is factually scrambled. Fluency is the new risk. The old risk was a blank.
Watch for three failure modes that look like success. First, the model 'cleans up' a number. You said 15 of metformin; it prints 50 because 50 is common. Second, it resolves an unclear pronoun to the wrong problem. Third, it drops a negation. 'No chest pain' becomes 'chest pain' because the acoustic energy was on the noun. A human editor listening at 1.5x will catch the negation. A physician skimming a pretty paragraph at the end of clinic will not.
This is why turnaround time alone is a bad scoreboard. A note that returns in four minutes and contains one inverted laterality is worse than a note that returns in two hours and matches the dictation. Ask the service for a correction log on medications, allergies, laterality, and negations for a sample week. If they cannot show it, they are selling speed, not transcription.
Where the model actually saves time
The time savings are real on high-volume, relatively clean work. A family physician who dictates into a quiet phone app, uses the same headings every day, and speaks doses the way they are written will see the editor's job shrink to punctuation, headings, and the occasional drug name. That is a legitimate product. It is also a narrow product. It describes a minority of the audio that arrives from hospitals, procedure rooms, and accented, interrupted, multi-speaker dictation.
AI helps the queue in three concrete places. It starts the document so the editor is not facing a blank screen. It proposes a split into your template headings when the dictator used those words. It flags low-confidence spans — if the vendor exposes them — so the editor slows down exactly there instead of re-listening to the whole file. Those three features are the whole economic case. Everything else is marketing.
Background noise, speaker overlap, and telephone dictation from a car still break models that demo well on a headset in a booth. If your group has hospitalists dictating from a corridor, do not pilot on the medical director's office recordings. Pilot on the worst hour of audio you actually send. If the edit time on that hour does not fall, you have not bought a faster service. You have bought a demo.
What the human editor is still paid to do
The editor's remaining work is judgment, not keystrokes. They decide that a sentence belongs under assessment and not under history. They hear 'Keflex' through a mask and do not accept the model's guess of a similar-sounding word. They notice that the dictator corrected themselves thirty seconds later and that the model kept both versions. They leave a blank and a question rather than invent a dose. That last habit is the difference between a transcription service and a language model.
Specialty knowledge still matters, and it matters more now because the draft looks finished. An orthopedic editor knows the difference between a tendon and a ligament in the way that surgeon talks. A psychiatry editor knows not to 'smooth' a patient's words into clinical jargon the physician did not use. A radiology editor knows that a measurement repeated wrong is not a style issue. General AI writing tools do not carry that ear. A service that routes your audio to whoever is free, with no specialty bench, will show it in the bounce-back rate.
You should be able to name the accountability. Who edited this note? Where are they? What may they change without calling you? Sunrise's long practice has been that a person stands behind the document before it reaches you to sign. AI can sit in front of that person. It should not sit instead of them when the document will live in the legal record.
Templates, files, and the handoff into the EHR
The unglamorous half of transcription is routing. Audio arrives from a phone, an app, or a recorder. It has to be tied to the right patient and encounter, land in the right template, and return to the right inbox. Models do not do that. Workflow does. If your AI transcription rollout did not include a written map of 'file in, signed note out,' you will spend the savings on staff hunting for documents.
Keep dictation and ambient scribing as different pipes. Dictation is you speaking the note, often after the patient has left, often for operative reports and addenda. Ambient scribing is software listening to the visit. Mixing them in one queue without labels produces notes that look alike and were produced under different consent and different risk. Your staff should see which path a document took before they file it.
EHR paste is where good transcripts go to die. If the returned document is a single block and your chart expects discrete fields, someone is still doing data entry. Ask whether the service can return headings that match your note type, and whether a human checked that the assessment did not get pasted into the exam. Integration is not a logo on a slide. It is the number of clicks between 'transcription ready' and 'signed.'
How to buy this without getting a toy
Run a two-week sample on your real audio, not the vendor's sample files. Include at least one noisy file, one heavy accent, one procedure or operative dictation if you do those, and one file where you correct yourself mid-sentence. Score medications, laterality, negations, and whether the headings match your template. Ignore the font.
Contract for the edit, not the model. You want a named quality standard: critical errors per hundred lines, turnaround by note type, and a way to send a note back that is tracked. You want a business associate agreement that covers the audio, the draft, and the editor, including any subprocessors who touch the audio to run the model. If the model vendor is unnamed, you do not have a BAA. You have a brochure.
Price the physician time you are trying to avoid. If the 'AI transcript' returns so rough that you spend six minutes a note repairing it, multiply that by your daily volume and compare it to a human-edited service. Many groups discover that the cheap tier is the expensive tier. The honest offer is a faster edit, with a person still on the hook, at a price you can compare to the evening you currently spend charting.
What to watch after go-live
For the first month, sample ten notes a week yourself. Not the easy ones. Look only at meds, allergies, side, and 'no.' Keep a tally. If critical corrections rise after the vendor 'updates the model,' freeze the update for your account until you re-sample. Models change underneath you. Your standard should not.
Tell dictators what helps. Speak the heading. Spell the rare drug once. Say the correction out loud ('correction: right knee, not left') instead of hoping the model noticed the tone. These are old transcription habits, and they matter more when a fluent draft hides the mistake. A five-minute huddle beats a twenty-page tip sheet.
Keep a path that is not AI. Operative notes, disability narratives, and anything you would not want summarized should still be able to go to a human who types from the audio. A clinic with only one pipe will eventually force the wrong encounter through it. Transcription's future is a choice of pipes — edited dictation, ambient draft, or both — with the same rule at the end: a clinician signs, and a human who can hear the audio was in the chain when the words were clinical.
A practical standard you can post in the dictation room
Use this standard until you have your own numbers. The first draft may be produced by software. The document you are asked to sign must have been compared to the audio by a person for medications, doses, laterality, and negations. Headings must match the note type. Unclear audio must be marked, not guessed. Turnaround is whatever you contracted, but a fast wrong note is a defect, not a win.
If a vendor cannot live with that paragraph, they are not selling medical transcription. They are selling a text generator that happens to accept audio. Physicians have always been willing to pay for the former, because the alternative is doing it themselves after the last patient. AI belongs in that paid workflow as the start of the edit. The signature is still yours, and the ear that protects the signature should still be a person's.
Post the standard where people dictate, not in a policy binder. The doctors who will use the service are tired at the end of the list. They will follow a four-line card. They will not reread a contract. Your quality program is that card plus the weekly sample of ten notes, done by someone who still listens to audio when the text looks too clean.
Accents, code-switching, and audio the demo never used
Most public demos use a calm speaker of the variety of English the model saw most in training. Your dictators include surgeons who trained abroad, hospitalists who switch into another language for a drug name, and clinicians who talk fast because the next patient is waiting. Error rates are not evenly distributed. If you only average them, you will bless a system that is excellent for half your group and unsafe for the rest.
Ask for a breakdown, even a rough one, by speaker, not just by day. A model that is '98 percent accurate' can still miss every other medication for one physician. That physician will quietly go back to typing, and your adoption numbers will lie. The fix is not to tell them to 'speak more clearly' as if clarity were a moral trait. The fix is a human editor who has heard that physician before, plus a feedback loop that sends their corrections back as examples rather than as complaints.
Code-switching is normal in real clinics and rare in benchmarks. A dictator may say the diagnosis in English and the patient's words in the language of the visit, or spell a name and then resume at full speed. Models often latinize, drop, or 'correct' the non-English span into something plausible and wrong. Your instruction to editors should be explicit: preserve the patient's words, mark what you could not hear, and do not translate unless the dictator translated. A polished English sentence the patient did not say is a clinical invention.
Operative notes, addenda, and audio you should not summarize
Some documents are the procedure. An operative note, a procedure note, a discharge summary built from dictation, a disability narrative, and a medical-legal addendum are not candidates for a 'helpful summary.' The sentences are the work. An AI pass may still produce the first typing so an editor can move faster, but the editor must work from the audio with the same attention they used when there was no model. If your contract gives those note types the same automated confidence threshold as a sore-throat visit, change the contract.
Addenda are where fluent models cause quiet harm. The physician says 'add to the plan, start the statin, and ignore my earlier comment about aspirin.' A summary-style model may fold that into a clean plan that still contains the aspirin line because it was 'in the note.' Transcription of an addendum should show the change, not dissolve it. Train editors to keep the correction visible. Train physicians to say 'delete the aspirin sentence' rather than 'scratch that' and hope.
Keep a non-AI route on the order form. It can be a checkbox: human type-up from audio, no generative rewrite. You will use it less than you fear and you will be glad on the day a case looks like it may be reviewed. The existence of the checkbox also disciplines the default path. Staff learn that AI-assisted is a method, not a law of nature. That is the whole cultural change worth making: software drafts, a person who can hear the file edits, and you sign only what you would have said.
Use the checkbox the first week even on ordinary notes, once per dictator, so staff remember it exists. A control that nobody has clicked is not a control. On that one file, compare the human type-up to the AI-assisted version side by side. You are not looking for style. You are looking for any fact that appeared in only one of them. Those facts are your real error list, and they are worth more than a quarterly business review slide.
What a sane week looks like after the switch
Picture a group of eight clinicians sending about forty dictated files a day. Before AI, editors typed from scratch and the median note returned in a few hours, longer for operative reports. After a careful switch, the same editors open a draft, correct it against the audio, and the median on clean clinic dictation drops. Operative reports move less, because they were never limited by typing speed. Physicians notice the clinic notes first. They should not be told that everything got faster. They should be told which note types got faster and which still take the old window.
The medical director's job that week is boring and sufficient. Ten notes, headphones, a tally of meds, side, and the word 'no.' One conversation with the dictator whose tally is an outlier. One written decision: keep the model version, or roll it back for that speaker. No all-hands. No celebration of minutes saved until the tally is dull. Dull is the goal. Transcription is supposed to disappear into the day. AI belongs there only if it makes the dullness more reliable, not if it makes the draft more eloquent than the dictator was.