Last updated July 13, 2026

How accurate is Zonlo's feedback?

Feedback you can't trust is worse than no feedback. This page explains how Zonlo grades what you say, how we measure that grading, what the numbers mean, and what we haven't validated yet. It is an engineering writeup, not a sales page.

What Zonlo is trying to do

Zonlo is a private space to rehearse speaking out loud before a real conversation happens. It is built for people who know the words but freeze in the moment, so every design decision follows from two rules: the feedback has to deserve trust, and it can never punish a correct answer. A grader that wrongly fails someone who already freezes when speaking would recreate the exact harm the app exists to remove.

That's why this page exists. If we're asking you to speak into an app at your most self-conscious, you deserve to know exactly what happens to your voice, what judges your reply, and how often that judge is right.

How grading works

Every reply you speak goes through the same pipeline, and the whole thing is built so the most damaging mistake, wrongly failing a correct answer, is the one we guard against hardest.

  1. Your iPhone transcribes your voice on-device with Apple's speech recognition. The audio never leaves your phone and is never written to disk. Only your words, as text, go any further.
  2. The transcript travels over an encrypted connection to our server, along with which scenario and which exchange of the conversation you're on. Nothing else about you is needed to grade a sentence.
  3. Japanese replies are parsed by a morphological analyzer. Kagome, an open-source tool, segments the sentence and produces the readings, the furigana and romaji, shown in your feedback. This is deterministic grammar analysis, not model guesswork.
  4. A large language model grades the transcript against that scenario's rubric: did your reply work in that moment, is there a more natural way to say it, and did the politeness level fit the scene. We tested feeding the deterministic parse into the register call and found it made the model over-cautious, wrongly flagging casual-but-fine replies, so today the model makes that judgment directly. Improving it is active work.
  5. The verdict is validated before you see it. If the grader errors, times out, or returns something malformed, Zonlo falls back to gentle, generic feedback instead of blocking you or showing a wrong red mark. You are never stuck waiting on a model.

What each number means

Cases in the corpus
How many real learner phrasings we test the grader against. Each case is a transcript with a human-decided correct verdict, including clusters of valid alternative phrasings: the correct-but-different answers a strict grader would wrongly mark down.
Pass and correctness
Of all corpus cases, how often the grader's pass/fail verdict and its correctness call match the human label. This is the headline number: does the grader agree with a careful human.
Register judgments
How often the grader correctly judges politeness level, for example casual form where the scene called for です/ます. This is our softest number and the one we are actively working on: the model makes this call directly, and its misses split between being too strict on short polite replies and too lenient on blunt ones.
False-fails
Correct replies the grader wrongly failed. We track this as its own number because it is the most damaging mistake Zonlo can make: telling someone who already freezes when speaking that a correct reply was wrong is exactly the harm this app exists to avoid. Our bar for this number is zero, and we are not there yet: the current run has 6, mostly on newer follow-up exchanges whose grading we are still tuning. Closing them is the top thing we are working on.

How we measure it

Every change to the grader, whether it's the prompt, the model, or the validation logic, runs against the full corpus before it ships. If it grades worse than what's live, it doesn't ship. That's the whole rule, and it has already vetoed model upgrades that looked better on paper.

The corpus is deliberately weighted toward the hard cases: valid alternative phrasings, politeness borderlines, and replies that are correct but unusual. A grader can look great on easy cases; these are the ones that decide whether the feedback is trustworthy.

An independent check on register

There's a trap in measuring a language model against labels you wrote yourself: if the model just agrees with you, you've learned that you agree with yourself, not that either of you is right. Politeness level is where this bites hardest, because it's the judgment we're least sure of and, for Japanese, the one we can't yet fully check with a native speaker.

So for register we added a second opinion that comes from neither us nor any AI: three public datasets where each sentence's politeness or formality was labeled by researchers, professional translators, or crowd raters. We wrote a small deterministic checker for each language, from grammar rules alone, and scored how often it agrees with those human labels:

These are agreement rates between our rule checker and outside human labels, not the app's own grading accuracy. The point is that our softest number now has an anchor that didn't come from us or from an AI.

Current results

500 cases in the corpus
94.6% pass and correctness
94.5% register judgments
6 false-fails

Measured on Gemini 2.5 Flash-Lite, last run July 13, 2026. Percentages on a corpus this size should be read as a direction, not a guarantee: one new hard case can move them, which is exactly why we publish the corpus size next to them. These are our own measurements against our own labeled corpus. There is no third-party or industry benchmark for spoken-conversation feedback, so we cannot compare Zonlo to other apps here, and about half the corpus, the Japanese cases, has not yet been reviewed by a native speaker. Read these as our honest internal number, not a verified or comparative claim.

What we haven't validated yet

Why we publish this

Language apps love the word "AI" and hate showing their work. Zonlo is built for people who freeze when speaking, and that only works if the feedback deserves trust. So we publish the corpus size, the misses, and the gaps, and we update this page as they change. The date at the top is the date of the numbers, not of the prose. For plain-language detail on how AI is used and its limits, see our AI Disclosure.

Curious how this compares to other apps? Read Zonlo vs Duolingo and Zonlo vs Speak, or see how a conversation works.

Rehearse it here first.

Zonlo is coming soon to the App Store. Join the waitlist and be first in when it lands.