How to Pick a Medical Transcription Service That Works on Your Audio

Four tabs open, four accuracy claims. Ninety-nine percent. Ninety-nine point five. “Clinical-grade.” One that just says “industry-leading” and gives no number at all. None of them tells you which audio that was measured on, who produced the reference transcript it was scored against, or whether a second of it came from a room that sounds like yours.

The short answer. You cannot choose a medical transcription service from published accuracy figures, because word error rate only means something relative to a specific test set, and no vendor publishes theirs. The buyer’s job is to build a test set from their own recordings first, then score every candidate on it identically. Twenty real appointments, a checked human reference transcript and an afternoon of scoring will separate a shortlist more reliably than every comparison page on the internet.

Do it in that order and you usually find something else: the gap between the top few products on your audio is smaller than the gap a better microphone makes.

General information for practices and clinicians, not legal or clinical advice. Your compliance officer and attorney get the final word, especially on recording consent.

What “accuracy” means, and why the percentage misleads

The industry metric is word error rate, and Microsoft’s speech documentation states it plainly: WER = (I + D + S) / N × 100, where an insertion is a word “incorrectly added in the hypothesis transcript,” a deletion is a word “undetected,” a substitution is a word “substituted between reference and hypothesis,” and N is “the total number of words provided in the human-labeled transcript.” A vendor’s 99% is almost always 1 minus WER on a test set they chose. NIST has distributed the standard scoring machinery for decades in its SCTK toolkit, built around sclite, so the arithmetic is not in dispute. The inputs are.

Two properties of that formula should change how you read any headline number. The first is that it depends entirely on the audio. WER on clean read speech, one speaker, close mic, scripted sentences, no crosstalk, is a different quantity from WER in an exam room with a fan running, a patient who trails off, a chair scraping and two people finishing each other’s sentences. Both are legitimately word error rate. They are not comparable, and nothing in the marketing tells you which one you are looking at.

The second is that it weights every word the same, which is the part that matters clinically. “The” costs exactly what “metoprolol” costs. Count it on one of your own transcripts: a fifteen-minute consultation might run three or four thousand words, of which the drug names, doses, units, frequencies, lateralities, dates, numbers and negations come to a few dozen. That slice carries almost all the clinical risk and is statistically invisible inside a whole-transcript percentage. A product can post an excellent WER while reliably mangling exactly those words, because there are so few of them that getting every one wrong barely moves the number.

Microsoft documents a third measure worth knowing about, token error rate, same formula but scored on “the final end-to-end display format” including “punctuation, capitalization, and ITN,” the normalisation that turns spoken numbers into written ones. “Fifteen milligrams” rendered as “50 mg” is not a word error in the lexical sense. It is a dosing error in your record.

Build your own test set before you shortlist a medical transcription service

This is the whole method, and almost nobody does it, which is why almost everybody buys on a demo.

Collect around twenty recordings that look like your actual work. Microsoft’s guidance for evaluating a speech model is to “provide 30 minutes to 5 hours of representative audio,” and twenty appointments sits inside that. Representative is doing the work in that sentence. Spread the set across every clinician who will use it, including the fast talker and the one with the accent the software will find hardest; across your visit types, since a medication review and a post-operative check produce different vocabulary; across your rooms, because the one with the hard floor behaves nothing like the quiet one. Include at least three encounters with a carer or family member present, keep telehealth as its own category rather than mixing it in, and include one deliberately terrible recording. The interesting question is not how a product performs on a good day.

Two ways to get that audio without creating a compliance problem before you have a contract. Settle the consent and BAA questions below first, which you need anyway. Or stage the encounters: staff reading realistic scripted dialogue in the real rooms on the real microphones, with drug names, doses, lateralities and negations planted deliberately. Staged audio is kinder than the real thing, and it has one large advantage. You know the ground truth exactly, because you wrote it, and no patient information leaves the building during a bake-off with four vendors you have not signed with.

Produce a human reference transcript. No shortcut exists. One person transcribes verbatim, a second checks against the audio, disagreements get resolved by listening again rather than guessing. Budget a working day. Every score you produce is measured against this file, so an error here becomes permanent.

Then write a scoring rule that reflects clinical risk rather than word count. Three buckets, scored and reported separately:

  • Bucket A, clinical content. Drug names, doses, units, frequencies and routes, laterality, anatomical sites, every number, every date, named diagnoses, allergies, device names. Counted individually, never averaged into anything else.
  • Bucket B, negation and uncertainty. No, not, denies, without, ruled out, unlikely, possibly, query, suspected. A dropped “no” inverts the clinical meaning of a sentence while costing one word of WER. This bucket catches the failure a percentage hides best.
  • Bucket C, everything else. Filler, false starts, backchannels, social talk. Track it, weight it near zero, stop worrying about it.

A product with a worse overall WER and zero bucket A errors is the better buy, and no comparison table will tell you that. Run the same files through every candidate on the same day, without telling any vendor which recordings they are getting or what you are scoring. If a service is non-deterministic, run each file twice and check the output matches.

Speaker labels: what diarisation does well, and where it quietly fails

Diarisation, spelled diarization in most vendor documentation, answers a different question from transcription. Transcription asks what was said. Diarisation asks who spoke when, and it is scored by its own metric rather than by word error rate, which is why an accuracy claim tells you nothing at all about speaker labels. NIST treated it as a separate task for that reason, running “Who Spoke When” speaker diarisation as its own track in the Rich Transcription evaluations.

It works well in the case everyone demos: two people, distinct voices, polite turn-taking, one good microphone, nobody interrupting. Consultations are not reliably that. Five failure modes, in roughly the order you will meet them.

  • Overlapping speech. Two people talking at once is the hardest case in the field, and the segment usually gets assigned wholesale to one of them.
  • A third person in the room. Carer, interpreter, student, chaperone. Published ceilings are generous, and Amazon Transcribe for instance “can differentiate between a maximum of 30 unique speakers.” The ceiling is not the problem. Accuracy at three voices is.
  • Short utterances. “Mm-hm,” “right,” “okay.” Too little audio to identify, so they attach to whoever spoke around them, which is how a patient ends up appearing to agree to something.
  • Telehealth. Compressed audio, two capture chains, and the far end coming out of your speakers back into your own microphone. Test it separately or you will average a bad result into a good one.
  • Label drift. Labels are anonymous by design, Speaker 1 and Speaker 2, and someone has to map them to names. Check whether the mapping holds across the whole recording or whether the speakers swap identities halfway through.

Test it rather than trust it. Mark every speaker change in your reference transcripts, then count two things separately: turns given to the wrong speaker, and clinically meaningful content attributed to the wrong person. Those are different severities. A misattributed “okay” is noise. A daughter’s own medication list landing in the patient’s history is a chart error with a long tail.

One practical move beats all of it. If you can capture two channels, one microphone per speaker, or a telehealth platform that records each participant to its own track, use channel separation instead of diarisation. Speaker attribution stops being a prediction and becomes a fact about which track the audio arrived on.

The audio chain decides more than the vendor does

Buyers argue about models and ignore the twelve inches between a mouth and a microphone. Google’s Speech-to-Text best practices are blunt: “Position the microphone as close as possible to the person that is speaking, particularly when background noise is present,” and “Capture audio with a sampling rate of 16,000 Hz or higher.” Echo and background noise “may reduce accuracy, especially if a lossy codec is also used.”

One item in the same guidance catches people out. Pre-processing audio to make it sound nicer usually makes recognition worse: Google advises against automatic gain control and states that “applying noise-reduction signal processing to the audio before sending it to the service typically reduces recognition accuracy.” A conference speakerphone with aggressive noise suppression can lose to a cheap lapel mic feeding raw audio.

  • Get the microphone off the laptop. A built-in mic at the far end of the desk, behind a monitor, pointed at the ceiling, is the worst common setup in healthcare and it is nearly universal. A boundary mic on the desk between you, or a lapel mic, costs less than a month of most subscriptions.
  • Fix the room before you blame the software. The vent, the corridor door, the hard surfaces, the printer. Stand where the mic is and listen for thirty seconds. You will hear things you stopped noticing years ago.
  • Check what the app uploads. Some capture at a low sample rate or compress hard to save bandwidth. If the chain resamples, information was lost before any model saw it.

The test that settles the argument takes an afternoon. Take the two worst-scoring files from your bake-off, re-record the same content with a better microphone in the same room, and rescore them on the product that already lost. If a fifty-dollar microphone moves your bucket A count further than changing vendors would, you have your answer and you can buy the cheaper service.

Transcription, dictation and ambient documentation are three different purchases

People arrive at this decision having conflated the three, then buy the wrong category.

Transcription turns a recorded conversation into text of what was said, verbatim or lightly cleaned, with speaker labels and usually timestamps. It does not write your note. You get a record of the encounter.

Dictation is one known speaker narrating a document into a microphone, usually straight into an EHR field. Most accurate of the three by a distance: one voice, controlled phrasing, close mic, a human correcting in real time.

Ambient documentation listens to the consultation and produces a structured clinical note. Transcription is an input to it, not a competitor.

Choose by what needs to exist afterwards. If you need the words, because this is a therapy session, a multidisciplinary meeting, an independent medical examination or anything that may be read back later, buy transcription. If you need the note written, that is ambient documentation, and its economics and governance are separate subjects: we have covered whether an AI medical scribe pays for itself in a small practice and what a public hospital must settle before deploying one, and neither argument is repeated here. If you already compose the note in your head and just want it typed, dictation is cheaper and more accurate than either.

The compliance layer, compressed

A transcription vendor handling recordings of your patients is a business associate, and the agreement comes before the first recording, not after the pilot. Check it against the elements in HHS’s sample business associate agreement provisions: permitted uses and disclosures specified rather than gestured at, subcontractors required to “agree to the same restrictions, conditions, and requirements that apply to the business associate,” an obligation to “report to covered entity any use or disclosure of protected health information not provided for by the Agreement” including breaches of unsecured PHI, and return or destruction at termination with the clause people skip, that the “business associate shall retain no copies of the protected health information.”

One vendor answer is worth pre-empting. “We encrypt everything and cannot see your data” is not a reason to skip a BAA. HHS’s cloud guidance is explicit: “Lacking an encryption key for the encrypted data it receives and maintains does not exempt a CSP from business associate status,” and an entity that maintains ePHI on your behalf “is a business associate, even if the entity cannot actually view the ePHI.”

The same guidance hands you a checklist, noting that service level agreements can address HIPAA concerns including “system availability and reliability,” “back-up and data recovery,” the “manner in which data will be returned to the customer after service use termination,” security responsibility and “use, retention and disclosure limitations.” Turn each into a question whose answer contains a number.

  • The audio. Stored where, in which country, for how long after delivery, playable back by whom inside the vendor, deleted automatically or on request. Days, not the word “temporarily.”
  • The transcripts, separately. Retention for text is often different from retention for audio, and people ask about only one.
  • Training. Are recordings, transcripts or your corrections used to train, fine-tune or evaluate models, or reviewed by humans for quality? If there is an opt-out, is it a contract term or a console toggle a product update can reset? Put it in the contract.
  • Who listens. Most human-assisted services have people in the loop, which is fine and often the point. Ask whether they are workforce members of the business associate or subcontractors, where they work, and whether the downstream BAAs exist.
  • Consent to recording. Not answered by HIPAA. Recording-consent law is state law and it varies, including states that require all parties to consent. HHS notes the Privacy Rule preserves “contrary State laws that… provide greater privacy protections or privacy rights.” Ask your attorney which rule applies, then build the plainest practice available: a line in the check-in paperwork, a spoken sentence from the clinician, and a refusal path with no friction attached.

Where the words land, and the review step that cannot be optional

The question that decides whether you save any time is dull and rarely asked in a demo. Does the output arrive in the record, or as a file somebody re-types? Settle these before signing.

  • Format. What you actually receive: plain text, Word, JSON, a timestamped caption format. Whether speaker labels survive export or exist only in the vendor’s web viewer. Whether timestamps are per segment or per word, which matters the first time someone checks a disputed passage against the audio.
  • Delivery. API, SFTP, portal download, or write-back into the chart. Email is not a delivery mechanism for this. If you want write-back, call your EHR vendor before the transcription vendor and ask what an interface costs and how long it takes. Weeks, not days, and it is the step that wrecks timelines.
  • Turnaround. Streaming and near-real-time are a different product from batch. Ask the committed turnaround, then ask what the actual distribution looked like in their busiest month and what happens over a holiday weekend. Get a number into the contract instead of “typically.”
  • Custom vocabulary. Can you maintain a term list: your formulary, local clinician and facility names, device names, your practice’s abbreviations? On specialist vocabulary this is usually the biggest accuracy lever you control inside the product, and it is free.
  • The correction loop. When someone fixes an error, does anything improve, or does the same drug name come back wrong next week?

Then the part that does not move. A transcript entering the clinical record is reviewed and approved by the clinician who signs it, because the signature attests to the content regardless of what produced the draft. Design the review rather than hoping for it. Name what gets read in full every time and never skimmed: medications, doses, units and frequencies, allergies, every number and date, laterality, and the negations and hedges. Do it while you still remember the encounter. Working through eleven transcripts at nine at night is not review, it is signing.

The human option, priced properly

Treating a trained human transcriptionist as what you did before the software got good is a mistake, and it costs practices money in both directions. Any medical transcription service worth shortlisting will quote you a human tier, and humans still win clearly in several places: strong or unfamiliar accents, genuinely poor audio you cannot fix, three or more speakers in unstructured conversation, dense specialty vocabulary, psychiatry and therapy where verbatim phrasing carries clinical meaning, anything heading into a legal, disability or IME file, and any recording where the cost of an undetected error is high enough that “the clinician will spot it” is not a control you want to lean on.

Most human services now run edit-from-draft: speech recognition produces the first pass, a trained healthcare documentation specialist corrects it against the audio. Much faster than typing from nothing, much more accurate than raw output, and the reason the price gap between “AI” and “human” services is narrower than the marketing implies.

Compare them in three steps. Convert every quote to cost per hour of audio, since vendors price per audio minute, per line, per word or per hour, and a per-line price means nothing until you ask how many characters count as a line. Add your own correction time to the cheaper option. If a service saves a few dollars per encounter but costs a clinician four extra minutes of editing, twenty encounters a day is eighty minutes of clinical time. Price that yourself. It is almost always the larger number. Then ask for a paid sample on your three ugliest files and score it with the same buckets. A provider who will not transcribe your difficult audio for money before you commit has told you how they expect it to go.

Two parts of this are work AB7 Solutions does directly, either side of the decision. The first is the evaluation: human reference transcripts built and double-checked against your own recordings, scored with clinical terms and negations counted separately from filler, so your shortlist is ranked on your audio rather than somebody’s benchmark. The second is whatever the test says you need next, which is medical scribing and healthcare documentation support, edit-from-draft review of speech recognition output by trained people, and remote healthcare administrative staffing, under a signed business associate agreement with the retention and subcontractor questions above answered in writing rather than on a slide. The protocol in this article is the one we would run for you, which is the only credential worth offering on a question like this. Call +1 321 341 7733, email ab@ab7solutions.com or director@ab7solutions.com, or start at www.ab7solutions.com.

Sources: Microsoft Learn, Test accuracy of a custom speech model (word error rate and token error rate formulas, representative audio guidance); NIST, Rich Transcription Evaluation and the Speech Recognition Scoring Toolkit (SCTK/sclite); Google Cloud, Speech-to-Text best practices; Amazon Web Services, Amazon Transcribe speaker partitioning; HHS Office for Civil Rights, Sample Business Associate Agreement Provisions, Guidance on HIPAA and Cloud Computing and HIPAA laws and regulations, on State law preemption. Accessed September 2026.

Leave a Comment

Your email address will not be published. Required fields are marked *