APIs, integration & security — in depth

Word Error Rate as a Meeting Transcription Quality Metric

Word Error Rate alone misses what matters most in real meeting transcriptions.

Senior Writer · · 7 min read
Cover illustration for “Word Error Rate as a Meeting Transcription Quality Metric”
Speech Recognition · September 22, 2026 · 7 min read · 1,611 words

Word Error Rate is the number most people reach for when they want to know if a transcription tool works, and for good reason: it's the closest thing this field has to a standard unit of measure. WER counts the minimum number of word-level edits, substitutions, deletions, and insertions, needed to turn a machine transcript into the reference version a human would produce. The formula is simple: add up substitutions, deletions, and insertions, then divide by the total word count of the reference. Zero percent means a perfect match, 100% means the output is unusable, and everything between tells you roughly how much cleanup a transcript needs before anyone can trust it. Word-level scoring, rather than character-level, exists because a wrong word costs a reader real meaning, while a typo inside a correctly identified word usually doesn't.

That's the theory. What follows is where the theory holds up, where it starts to crack, and what a buyer actually needs to check before trusting a transcription vendor's number.

The gap between clean-audio benchmarks and real meeting conditions

Vendors advertise word accuracy in the mid-90s to high-90s for 2026-era models, and on the audio they're testing against, that claim holds. Leading systems now run in the 1 to 4% WER range on clean benchmarks, assessments built on controlled, high-quality audio that is far removed from everyday meeting conditions.

Meetings don't sound like audiobooks. The same research that produces those low single-digit numbers on clean audio also reports WER climbing as high as 28% once the input shifts to accented, conversational speech. That's not a small variance: a transcript at one end can be handed to a client, while one at the other end needs a full read-through before anyone trusts it. The marketing number and the meeting-room number are measuring two different worlds, and buyers need to plan for that gap.

The four conditions that degrade WER in meeting audio

Four recurring conditions explain most of that gap, and any meeting that includes even one of them should make a buyer discount the vendor's clean-audio number accordingly.

Multiple speakers and overlapping speech. Crosstalk and interruptions are structurally different from the sequential, one-voice-at-a-time audio most benchmarks are built on. When more than one person is talking, WER alone stops being enough. Researchers use cpWER instead, a version that penalizes both transcription mistakes and speaker-attribution errors in a single combined score. A 2025 systematic review found WER climbing past 50% in conversational, multi-speaker scenarios, a number that should give pause to anyone assuming a vendor's clean-audio score will hold up in a five-person standup.

Accented and non-native English speech. The same 28% WER figure cited above applies to accented dialogue, and global teams compound this problem: a call with speakers from three regions compounds accent variation with the ordinary variation in audio quality, microphone type, and background noise that every meeting already has. Models systematically underperform on accented speech, a pattern that research on accented dialogue documents repeatedly.

Technical and domain-specific vocabulary. A standard speech model has no built-in expectation that someone will say "Kubernetes," a client's product name, or a proprietary internal acronym, so these terms turn into substitutions or outright deletions. That matters more than the raw error count suggests: a single missed technical term inside an action item can corrupt the entire downstream task, turning a meeting-summary error into a real operational one. This is why keyword WER and missed-entity rate turn out to be more diagnostic than overall WER when the meeting in question is technical.

When WER stops measuring what matters for meeting intelligence

WER measures words. Meetings run on meaning, and that mismatch produces three specific blind spots.

First, semantic equivalence. "Cannot" versus "can't" registers as an error under strict WER scoring, and so does "yep" for "yes," even though neither substitution changes what was said or how a summary would read. Second, error asymmetry: WER treats a substitution (a plausible wrong word swapped in) and a deletion (a word dropped entirely) as equally costly, each counted as one error, but in practice they aren't equivalent. A deletion can strip a step out of an action item or remove context that a substitution of similar meaning would preserve, which makes deletions structurally more damaging in many meeting contexts. Third, and maybe most consequential: WER measures words, not who said them. A transcript where every word is transcribed correctly but every speaker label is wrong is functionally useless for meeting minutes, decision logs, or CRM attribution, and WER alone would still call that transcript a success.

The downstream use case is what decides which of these errors a buyer can tolerate and which ones sink the whole exercise. A compliance recording of a regulated conversation demands very high accuracy on substantive content; a dropped filler word matters far less there. A sales call transcript feeding a CRM depends heavily on correct speaker attribution and correct product or deal names, where errors carry real operational weight. A research interview may require fine-grained verbatim detail, while a summary meant for a busy executive needs exactly the opposite: the noise stripped out, the substance kept.

WER also says nothing about latency, and that omission hides the speed-versus-accuracy tradeoff real-time meeting tools face but post-meeting transcription tools never have to deal with. Real-time meeting tools face a speed-versus-accuracy tradeoff that a post-meeting transcription tool never has to deal with, since it can take its time. A tool boasting a technically lower WER but meaningful latency may perform worse in practice, for a live use case, than a competitor with a slightly higher WER that updates almost instantly.

The complementary metrics that fill the gaps WER leaves

None of this means WER should get thrown out. It means WER needs company, and three metrics in particular are gaining ground for good reason.

Semantic WER uses a large language model as a judge, scoring whether meaning survived the transcription rather than whether every character lines up with the reference. Open-source efforts like Pipecat's STT benchmark are working to standardize Semantic WER using reasoning models as judges, specifically to cut down on scoring bias between systems. This metric matters most when a transcript is feeding something downstream, a summary, an action-item extractor, a CRM update, where semantic fidelity counts for more than lexical exactness.

cpWER, the concatenated minimum-permutation WER, folds transcription accuracy and speaker attribution into a single score, penalizing both a wrong word and a wrong speaker label. It's a tougher and more useful signal than diarization error rate alone for any multi-speaker meeting, because it captures the failure mode that matters most in practice: correct words attached to the wrong person.

Keyword and entity accuracy tracks whether the names, product terms, and acronyms that actually carry business weight survive the transcription process. A tool posting a strong overall WER can still fail the meeting that matters most if it consistently drops or garbles product names, and this gap is often measured separately as keyword WER or missed-entity rate. Several vendors now let users supply a custom vocabulary list specifically to push this number up.

Leading tools compared against the full framework

Published vendor numbers and independent testing don't always agree, and where they diverge, the independent figures deserve more weight, since vendors are grading their own homework.

On clean audio, the numbers cluster tightly. AssemblyAI's Universal-3.5 Pro Realtime model posts a pooled WER of 6.99% among real-time systems. Independent conversational-English testing from 2025 put AssemblyAI around 4.5%, Deepgram around 5.26% in vendor-reported batch conditions (climbing to roughly 7 to 10% in independent real-world testing), OpenAI's Whisper base model around 5.1%, and Google Speech-to-Text around 4.8%, a tight, competitive field on clean input. In multi-speaker audio, the same tests show Deepgram around 5.8%, AssemblyAI around 6%, Google around 6.5%, and Whisper base around 7%, all higher, all consistent with the pattern this framework predicts.

Among consumer-facing meeting note-taker products, independent head-to-head testing reports accuracy in the low-to-mid 90s across the board, with variance tied directly to audio clarity and speaker count rather than to any single product being fundamentally ahead. One 2026 review of the eight leading meeting note-taker tools found all eight clearing 90 to 95%+ accuracy on English audio: raw English transcription accuracy has become table stakes, not a differentiator. The tools that separate themselves now do it on speaker attribution, domain vocabulary handling, and latency, not on the headline WER number.

A practical evaluation method for buyers who can't run their own benchmarks

Most buyers don't have a research team on hand to run a controlled WER study, and they don't need one. Two steps get most of the way there.

Start by stress-testing with real audio. Record an actual meeting, or something close to it, matching the speaker count, accent mix, and technical vocabulary of the meetings the tool will actually handle. Run that same file through every tool under consideration, then check the output against a transcript a human has verified line by line. That's a homemade WER test, and it will surface more truth than any spec sheet.

Then narrow the check to the words that matter most. Build a list of 20 to 30 terms the business depends on: product names, competitor names, technical acronyms, the names of key people who show up in these meetings repeatedly. Running that list against each tool's output shows what survives. Overall accuracy is a useful headline number, but a tool that transcribes the vast majority of words correctly while consistently mangling the one product name driving the deal is failing at the exact moment it needs to succeed. That five-minute check tells a buyer more than any benchmark table ever will.

Sources

  1. How accurate is speech-to-text in 2026?
  2. Evaluating the performance of artificial intelligence-based speech recognition for clinical documentation: a systematic review
  3. The Instrumental Dissolution of Typing: Why AI Challenges the Keyboard Era in Knowledge Work

More in Speech Recognition