The Zero-Edit Standard
Accuracy measures words. Users care whether the message is finished.
Zero-Edit Rate measures the percentage of voice dictations that produce output ready to send with no edits. It scores send-readiness, not transcription accuracy. We run 24 realistic dictation scenarios across 12 platforms, grade every output against published rules, and report the results here monthly.
We format what you said, we never say it for you.
Results by platform
Sorted worst-first. Small samples show their sample size. Baseline v2.2.0 · last updated July 14, 2026.
| Platform | ZER | Edits/msg | n | Note |
|---|---|---|---|---|
| AI Chat | 67% | 0.33 | 3 | List structure not preserved in multi-step prompts |
| Slack | 88% | 0.12 | 17 | Intensity softened beyond profanity scope |
| Document | 100% | 0 | 1 | |
| 100% | 0 | 7 | ||
| Email Reply | 100% | 0 | 6 | |
| 100% | 0 | 2 | ||
| LinkedIn Comment | 100% | 0 | 2 | |
| Slack DM | 100% | 0 | 13 | |
| Teams | 100% | 0 | 7 | |
| Tech Comment | 100% | 0 | 2 | |
| Tech Ticket | 100% | 0 | 2 | |
| Twitter/X | 100% | 0 | 2 |
Vocabulary metrics (Glossary terms, @handles, acronym rules) are tracked internally and excluded from the public ZER. They will be promoted when Glossary adoption warrants it.
How this is measured
We format what you said, we never say it for you.
This is a text-pipeline benchmark. A test suite of 24 dictation scenarios is run through the formatting pipeline, and every output is graded by a deterministic checker against published rules. A message passes only if it needs zero edits before sending.
The four counting categories
- Content: a stated fact is missing, invented, or changed. Reordering that alters meaning counts here too.
- Voice: the speaker’s phrasing, hedging, or intensity was altered. Softening “kind of satisfying” to “satisfying” is a Voice edit.
- Format: the output shape does not fit the platform. A multi-step prompt returned as prose instead of a list is a Format edit.
- Mechanics: punctuation, capitalization, greeting comma, sentence boundaries. One missing comma is enough to fail.
The reasonable-user definition
An output passes if a reasonable user, on the platform the message is intended for, would send it without making any edit. If the reasonable user would fix even a single character before sending, it fails.
The meta-instruction rule
If the speaker dictates instructions about a message rather than the message itself, Rubil formats the instructions as spoken. Composing the implied message would be generation, which Rubil does not do.
The role of the AI judge
An independent AI judge (a different vendor from the formatting model) reads each output alongside the rules and flags anything that looks off. Its role is advisory only. It cannot change any number on this page. Every published pass or fail is settled by the deterministic checker.
Scope
These are text-pipeline metrics. We do not measure microphones, accents, or end-to-end speech accuracy. Transcription quality upstream affects results but is not the variable being scored.
For the full definition and why the category needs it, read What is a Zero-Edit Rate?
Changelog
- v2.2.0 · July 2026Initial public baseline. 24 test scenarios, 12 platforms. Vocabulary tracked internally, reported separately from ZER.