PHILOSOPHY

What is Zero-Edit Rate? A Voice Dictation Metric Defined

Accuracy measures words. Users care whether the message is finished.

By Rubil TeamJuly 14, 20266 min read

Zero-Edit Rate measures the percentage of voice dictations that produce output ready to send with no edits. It exists because transcription accuracy does not describe whether a message is finished. ZER is counted with four rule categories (Content, Voice, Format, Mechanics), and one edit in any category fails the whole message. It is a send-readiness benchmark for voice dictation tools.

Accuracy measures words. Users care whether the message is finished.

Why accuracy is the wrong yardstick

The dictation category has spent a decade competing on Word Error Rate. WER is the percentage of transcribed words that match a reference transcript. It is well-defined, easy to compare, and beside the point for the person using the tool.

Users do not read WER. Users read the message on the screen and decide whether they can hit send. A 97% accurate transcript that needs two minutes of reformatting is a failure. A message with one word slightly off but perfect structure is a pass. The metric should match the job.

What Zero-Edit Rate measures

A dictation passes when the output needs zero edits before sending. An edit is a content error, a voice alteration, a format violation, or a mechanics mistake. One comma kills a pass. The bar is strict because strict is what makes the number mean something.

A softer scoring rule (say, “fewer than three edits counts”) would blur exactly the thing that matters. The binary pass or fail is the point. Either you sent it, or you opened your keyboard.

The four edit categories

Every output is graded across four categories. If any one records an edit, the message fails. Examples below are drawn from the current baseline suite.

  • Content. A stated fact is missing, invented, or changed. Reordering that alters meaning counts here too. If the speaker said “push the launch back a week” and the output says “push the launch,” that is a Content edit.
  • Voice. The speaker's phrasing or intensity was altered. The baseline flagged “kind of satisfying” shortened to “satisfying” as a Voice edit. The hedge was doing work; removing it changed the message.
  • Format. The output shape does not fit the platform. The baseline failed a multi-step AI Chat prompt that came back as prose instead of a numbered list. The words were right; the shape was wrong.
  • Mechanics. Punctuation, capitalization, greeting comma, sentence boundaries. An unwanted comma in “Hi, Kim” instead of “Hi Kim” fails a message. Small, deterministic, binary.

What ZER does not measure

ZER is a text-pipeline metric. It does not score transcription from audio, microphone quality, or how a system handles a strong accent. Those upstream factors affect what lands in the pipeline, but they are not what the pipeline is being graded on.

It also does not score subjective writing quality. “This is good prose” is not a rule the benchmark can check. The rules are things a checker can verify: greeting present, list rendered, capitalization matches convention, comma where a reasonable user would expect one.

How Rubil measures it

The current suite is 24 dictation scenarios spread across 12 platforms (email, email reply, Slack, Slack DM, Teams, LinkedIn, LinkedIn comment, Twitter/X, document, tech ticket, tech comment, AI chat). Every output runs through the deterministic checker against published rules.

An independent AI judge from a different vendor than the formatting model reads each output alongside the rules and flags anything the checker might have missed. It is advisory only. It cannot change a published number. The published pass or fail is settled by the deterministic checker.

Rubil's current Zero-Edit Rate is 95.3% across 12 platforms. The full breakdown, methodology, and monthly changelog live at /quality.

We format what you said, we never say it for you.

Why publish this

Not because transparency is virtuous. Because the category is full of unverifiable claims. “99% accurate.” “90% zero-edit.” “Human-level output.” None of them come with a rulebook, a scenario list, or a version stamp. A number with rules behind it is worth more than an adjective.

Rubil publishes a definition and a reference implementation for zero-edit measurement. Not the industry standard, which is earned, not declared. A starting point that anyone can copy, criticize, or beat. See the live results and rules at rubil.io/quality, and if you want a picture of how a formatting-first tool compares to a rewriting one, read Rubil vs Wispr Flow.

Frequently asked

What is a zero-edit rate?

Zero-Edit Rate (ZER) is the share of dictated messages that need no edits before the user sends them. A message either passes with zero edits, or it fails. Partial credit is not counted. It measures send-readiness, not transcription accuracy.

How is ZER calculated?

A fixed test suite of realistic dictation scenarios is run through the formatting pipeline. Each output is graded by a deterministic checker against published rules covering four categories: Content, Voice, Format, and Mechanics. If any category records an edit, the message fails. ZER is the pass count divided by the total scenario count.

Why not measure transcription accuracy?

Transcription accuracy (Word Error Rate) captures whether the words came through, not whether the message is finished. A 97% accurate transcript that still needs a greeting, a bullet list, and two comma fixes has failed the user. ZER measures the thing users care about: can I hit send.

Who runs this benchmark?

Rubil publishes and runs the benchmark. The grading is deterministic (rules, not opinion), and an independent AI judge from a different vendor than the formatting model reads every output as an advisory flagger. The judge cannot change any published number.

Can I rerun it?

Yes. The rules, the four counting categories, the reasonable-user definition, and the meta-instruction rule are published on the /quality page. The scenario suite version is stamped on every result. Anyone can implement the same grading against their own tool.

Zero-Edit RateVoice dictationQuality metricsBenchmarkMethodology

Try Rubil free

1,000 words/day. No credit card. No setup.

← Back to all postsrubil.io/blog · © 2026 Rubil