Quality statusLast run August 15, 2026 · Suite v3.1.0

Post-Edit Rate 0.58%

All four writing surfaces operational. 99.42% of dictations arrived ready to send with no edits at all.

Post-Edit Rate is the share of voice dictations that need at least one edit before they are fit to send. It scores send-readiness rather than transcription accuracy. Lower is better and zero is the goal.

Rubil is a voice dictation tool for Chrome and Mac. You speak, and formatted text appears in whatever you are writing in: email structure in Gmail, a short message in Slack, clean prose in Notion. It formats your words and never writes new ones.

Overall0.96% across the window
3 runs agoLatest
General writing · Notes, documents, long-form0.77%
Email · Gmail, Outlook, Superhuman0%
Direct message · Slack, Teams, iMessage0.46%
AI chat prompt · Claude, ChatGPT, Gemini1.39%
Operationalunder 1%Degraded1–5%Outage5% and above

Run history

August 15, 2026Suite v3.1.0

Operational · Post-Edit Rate 0.58% · 99.42% sent as written

August 14, 2026Value-preservation guard

Operational · Post-Edit Rate 0.43% · 99.57% sent as written

August 14, 2026Formatting model upgrade

Degraded · Post-Edit Rate 1.88% · 98.12% sent as written

What this measures

Post-Edit Rate is the percentage of voice dictations that need at least one edit before they are fit to send.

The name comes from post-editing, the term translators use for the cleanup a human does on machine output. That cleanup is the real cost of dictation software, and this puts a number on how often you have to pay it.

Every output is graded on four things: content (a fact missing, invented or changed), voice (phrasing, hedging or intensity altered), format (the shape does not fit the surface), and mechanics (punctuation, capitalisation, sentence boundaries). One edit in any category fails the whole message. Grading is deterministic; an independent AI judge reviews every output but cannot change a score.

Each run sends the full scenario set through the live production pipeline several times over, because the failures that matter are intermittent and a single clean run hides them. Scenarios are written to be hard, so this is a stress test rather than a picture of an average day.

The bar is deliberately unforgiving in one direction. Rubil formats your words and never rewrites, generates or summarises them, so an output that reads more smoothly than what you said is a failure, not a bonus. Deleting a hedge like “I think”, softening “pissed off”, or tidying a dictated number all count against the score, because each one puts words in your mouth that you did not say.

Why each surface is scored separately

A dictation that is perfect for Slack is wrong for email. Averaging those into one figure hides the surface you write on every day, so each is graded against its own rules and reported on its own row.

Email

Greeting on its own line, paragraph breaks at topic shifts, the closing the speaker dictated and no closing they did not. Names stay as spoken, so “Hi Sarah” never becomes “Hi, Sarah”.

Direct message

Chat register: one or two compact blocks, greeting inline, no email scaffolding. Shorter in shape, never shorter in substance: if the email version makes four points, the message makes the same four.

AI chat prompt

Structure is preserved rather than prettified. A dictated sequence renders as a numbered list, and an instruction stays an instruction: Rubil formats what you asked for, it does not answer it.

General writing

Documents, notes and posts. Long dictation is broken into paragraphs at natural topic shifts, and values stay exactly as dictated, so a build number or a time is never quietly reformatted.

Why not Word Error Rate?

Word Error Rate is the standard for speech recognition and it answers a real question: how many words did the recognizer get wrong. It is the right metric for judging a recognizer and the wrong one for judging whether you can send the message. A dictation can score a perfect 0% Word Error Rate and still cost you a minute of cleanup.

Word Error RatePost-Edit Rate
CountsIndividual wordsWhole messages
AsksDid it hear me right?Can I send this as is?
JudgesThe recognizerThe finished output
Punctuation & structureUsually ignoredCounted as a failure

Common questions

What is Rubil?

Rubil is a voice dictation tool for Chrome and Mac. You speak, and formatted text appears in whatever application you are writing in: email structure in Gmail, a short message in Slack, clean prose in Notion, across 20+ surfaces. It is speech-to-text, not text-to-speech, and it formats your own words rather than generating new ones. Rubil costs $5 a month, with 1,000 words a day free.

What is Post-Edit Rate?

Post-Edit Rate (PER) is the percentage of voice dictations that need at least one edit before they are fit to send. It scores send-readiness rather than transcription accuracy. A dictation counts against the rate if a reasonable user would fix even a single character before sending. Lower is better and zero is the goal.

What is Rubil’s Post-Edit Rate?

On the most recent run, Rubil’s Post-Edit Rate was 0.58%, meaning better than 99 in 100 dictations arrived ready to send with no edits at all. Direct messages and AI chat prompts routinely record zero edits. The suite runs daily and every result is posted here.

How is Post-Edit Rate different from Word Error Rate?

Word Error Rate counts how many individual words a speech recognizer got wrong against a reference transcript. Post-Edit Rate counts whole messages that needed human intervention before sending. A dictation can score a perfect 0% Word Error Rate and still fail on Post-Edit Rate: every word transcribed correctly, but no punctuation, no paragraph breaks, and filler left in, so the user still has to fix it. Word Error Rate measures the recognizer. Post-Edit Rate measures whether the user got to skip the cleanup.

What counts as an edit?

Four things. Content: a fact is missing, invented, or changed. Voice: the speaker’s phrasing, hedging, or intensity was altered, so softening “kind of satisfying” to “satisfying” counts. Format: the shape does not fit the surface, such as a dictated list returned as a paragraph. Mechanics: punctuation, capitalisation, or sentence boundaries. One edit in any category fails the whole dictation, which is why the bar is high.

How often is this measured?

Daily. Each run sends the full scenario set through the live production pipeline several times over, not a staging copy, so the number reflects what users receive. Results are posted here whether they improve or not.

Why does formatting differ between Slack, email and AI chat?

Because a good message looks different in each. Email needs a greeting, paragraph breaks and a closing. A Slack message reads as one or two compact blocks with the greeting inline. An AI chat prompt keeps its structure and lists intact. Rubil detects the surface and applies the right shape, which is why the board reports each one separately rather than averaging them into a single figure.

Does Rubil rewrite what you say?

No. Rubil formats the user’s own words and never rewrites, generates, or summarizes them. Hedges such as “I think” and “kind of” are preserved deliberately, because hedging is part of how a person speaks. Numbers dictated as digits are copied character-for-character, so an order number is never reformatted, and intensity is left alone.

Measured daily · Last run August 15, 2026 · Suite v3.1.0 · Full history as JSON← Back to Rubil