QUALITY

The Zero-Edit Standard

Accuracy measures words. Users care whether the message is finished.

Zero-Edit Rate measures the percentage of voice dictations that produce output ready to send with no edits. It scores send-readiness, not transcription accuracy. We run 24 realistic dictation scenarios across 12 platforms, grade every output against published rules, and report the results here monthly.

We format what you said, we never say it for you.

95.3%
Zero-Edit Rate
July 2026 baseline, v2.2.0 · 64 evaluations across 12 platforms

Results by platform

Sorted worst-first. Small samples show their sample size. Baseline v2.2.0 · last updated July 14, 2026.

PlatformZEREdits/msgnNote
AI Chat67%0.333List structure not preserved in multi-step prompts
Slack88%0.1217Intensity softened beyond profanity scope
Document100%01
Email100%07
Email Reply100%06
LinkedIn100%02
LinkedIn Comment100%02
Slack DM100%013
Teams100%07
Tech Comment100%02
Tech Ticket100%02
Twitter/X100%02

Vocabulary metrics (Glossary terms, @handles, acronym rules) are tracked internally and excluded from the public ZER. They will be promoted when Glossary adoption warrants it.

How this is measured

We format what you said, we never say it for you.

This is a text-pipeline benchmark. A test suite of 24 dictation scenarios is run through the formatting pipeline, and every output is graded by a deterministic checker against published rules. A message passes only if it needs zero edits before sending.

The four counting categories

  • Content: a stated fact is missing, invented, or changed. Reordering that alters meaning counts here too.
  • Voice: the speaker’s phrasing, hedging, or intensity was altered. Softening “kind of satisfying” to “satisfying” is a Voice edit.
  • Format: the output shape does not fit the platform. A multi-step prompt returned as prose instead of a list is a Format edit.
  • Mechanics: punctuation, capitalization, greeting comma, sentence boundaries. One missing comma is enough to fail.

The reasonable-user definition

An output passes if a reasonable user, on the platform the message is intended for, would send it without making any edit. If the reasonable user would fix even a single character before sending, it fails.

The meta-instruction rule

If the speaker dictates instructions about a message rather than the message itself, Rubil formats the instructions as spoken. Composing the implied message would be generation, which Rubil does not do.

The role of the AI judge

An independent AI judge (a different vendor from the formatting model) reads each output alongside the rules and flags anything that looks off. Its role is advisory only. It cannot change any number on this page. Every published pass or fail is settled by the deterministic checker.

Scope

These are text-pipeline metrics. We do not measure microphones, accents, or end-to-end speech accuracy. Transcription quality upstream affects results but is not the variable being scored.

For the full definition and why the category needs it, read What is a Zero-Edit Rate?

Changelog

  • v2.2.0 · July 2026
    Initial public baseline. 24 test scenarios, 12 platforms. Vocabulary tracked internally, reported separately from ZER.
Last updated: July 14, 2026← Back to Rubil