German breaks in three places an English speaker cannot hear breaking. We planted twenty-four errors inside fluent turns and asked one question of seven apps: was it repaired before the conversation moved on?
German fluency arrives unevenly, and the parts that arrive last are the parts a learner is least equipped to police. A bad vowel is audible to the person producing it. A noun in the wrong case is not, because the sentence still sounds like the sentence that was intended and the listener still understands it. Everything below follows from that gap between what an English speaker produces in German and what the same English speaker is able to notice a second later.
So we did not build a bench about conversation quality. We built one about correction, and we wrote the instrument down before opening a single app. Three error classes, picked because each is invisible from the inside and each fails in a different shape:
Against those three classes we put a single question, and refused to add a second: was the error repaired in the same turn? Not repaired eventually. Not filed, queued, listed on a summary screen or resurfaced next Tuesday. Repaired while the learner still had the sentence in their head and could feel which part of it moved.
The reason for that rule is memory rather than pedagogy. A case ending is chosen somewhere below conscious attention, in the same half second the speaker is choosing a noun and a plan for the rest of the clause. A correction landing inside that window attaches to a decision the learner can still locate. The identical correction landing eleven minutes later attaches to a sentence they no longer remember producing, and becomes a fact about German rather than a fact about them.
Plenty of products correct late, and we did not want to score that as nothing, so we logged it as a separate outcome rather than a failed one. Every planted error ended in one of three states: repaired in the turn, surfaced later in some form, or never surfaced at all. Those are different product features, they suit different learners, and a review that sums them into a single accuracy percentage has thrown away the only distinction that matters here.
One further rule, written before the runs: a repair had to be about the form. A product that echoed the corrected sentence back in passing, without the learner having any way to tell which word had changed, counted as a repair only if the changed element was made audible — stressed, isolated, or contrasted with what we had said. Fluent re-rendering is a courtesy, not a correction.
Eight items per class, twenty-four in total, each written out in advance and each embedded in a turn that was otherwise correct and reasonably fluent. That last constraint did most of the work: an error inside a broken sentence is easy to catch and tells you nothing, because the product is reacting to the wreckage. An error inside a sentence that is fine everywhere else is the actual test.
The roster was Enverson AI, Langua, Praktika, Babbel, Speak, Duolingo and Busuu. Nothing was paid for; every account stayed on whatever tier a new signup lands in. The whole sequence took a fortnight in August 2026, on one phone, with a reviewer whose German sits somewhere around B2. Where a product offered no free spoken turn to plant an error inside, we planted it in the written input it did accept and flagged the cell. That asymmetry is real and the limits section comes back to it.
| Case, verb-position and agreement errors repaired inside the same turn, out of twenty-four | |
|---|---|
| Enverson AI | 19/24 |
| Langua | 12/24 |
| Praktika | 9/24 |
| Babbel | 7/24 |
| Speak | 6/24 |
| Duolingo | 3/24 |
| Busuu | 2/24 |
The distance between the top of that chart and the bottom is larger than we expected and larger than the products’ marketing would let anyone predict. Nineteen out of twenty-four against three out of twenty-four is not a difference of degree; the two products are doing different jobs and only one of them is the job a German learner needs doing.
The two lowest bars are the least interesting and the most honest. Neither Duolingo nor Busuu is built to intercept free German speech mid-turn on a free tier, so a low bar here is a category mismatch rather than a failure. We include them because readers ask about them and because the size of the gap is itself information: if what you want is spoken repair, a course app is not a slow version of that, it is a different thing.
| Error class | What we said | Repaired in turn | Surfaced later | Never surfaced |
|---|---|---|---|---|
| Case (8 items) | Wir treffen uns in den Park, standing still; Ich habe den Student gefragt | 2.1 | 1.9 | 4.0 |
| Verb position (8 items) | Weil ich habe keine Zeit gehabt; separable prefixes abandoned before the clause ended | 3.8 | 2.4 | 1.8 |
| Gender agreement (8 items) | in einer kleinen Restaurant, der sehr gut war — one article, three wrong forms | 2.4 | 2.0 | 3.6 |
Verb position was repaired most often, across every product in the bench, and the reason is unflattering to the technology rather than to it: a verb in the wrong slot is a surface pattern, detectable without any idea of what the sentence means. Case is not. To know that in den Park is wrong you have to know whether the speaker is going somewhere or already there, and that is a question about the world, not about the string. Half of the case items were never mentioned by anything.
Agreement sat in between and split the field. Products that noticed the article generally repaired the whole chain; products that noticed only the adjective ending fixed the visible symptom and left the article that caused it standing, which is worse than silence, because the learner leaves believing the noun was right. Three of the seven did that at least once. If you take one number away from this piece, take the four-out-of- eight in the top right of that table: on average, half of a learner’s case errors went past the whole category without comment.
| Product | Did it name the rule | Did it show the corrected sentence | Did it make the learner produce it again |
|---|---|---|---|
| Enverson AI | Yes, and named it as a rule rather than a fix | Yes, alongside the version we produced | Yes — asked for it back, then for a second sentence built the same way |
| Langua | Sometimes, briefly, and usually only when asked | Yes | Rarely, and never twice |
| Praktika | Almost never; the character stays in the scene | Yes, folded back into the dialogue | No |
| Babbel | Yes — the clearest grammatical explanations in this bench | Yes | Yes, but as a written exercise on another day |
| Speak | Occasionally, in one line | Yes | No |
| Duolingo | A short tip, when a tip existed for that item | Yes | Yes, by putting the item back in the review queue |
| Busuu | Yes, when a human community member eventually answered | Yes | No |
Babbel deserves the note this table exists to make. Its in-turn repair count is seventh-place material by our instrument, and its explanations were the best thing we read in two weeks: accurate, short, aimed at the rule rather than the sentence, and written by somebody who has taught German to English speakers and knows which comparison lands. If you already know that your dative is unreliable and you want the reason stated properly, that is a real answer, and our instrument is simply not pointed at it.
The distinction matters commercially. A product can be excellent at explaining and useless at catching, and a learner who cannot yet detect their own errors will never trigger the explanation. Duolingo’s row is the same story in a different key: its review queue genuinely brings the item back, days later, as an item. Praktika went the other way and stayed in character, which keeps the scene intact and lets the error stand. We looked at that trade-off from the other side in our piece on AI tutors.
German is where a single proficiency number stops being a useful simplification and becomes an obstacle. A learner can be genuinely B2 on vocabulary and genuinely A2 on case at the same moment, and this is not an edge case — it is the standard shape of an English speaker two years in. A product holding one number has to average those two into a lesson, and the average is wrong for both: too easy to move the case problem, too slow to be worth the vocabulary reader’s evening.
| Reading | What a German failure on it looks like | What a fortnight aimed at it would contain |
|---|---|---|
| Pronunciation | The ch of ich and the ch of Buch collapsing into one sound | Minimal pairs and shadowing, with the grammar switched off entirely |
| Grammatical accuracy | Dative after a two-way preposition, weak masculines, endings that follow the wrong article | A fortnight of conversation narrowed to one case, repaired in the turn |
| Retrieval speed | The sentence is correct and arrives four seconds after the moment for it | Timed production of frames the learner has already proved they know |
| Vocabulary range | Two hundred words used flawlessly and nothing said outside them | Input pitched slightly above the ceiling, then forced reuse of what came out of it |
| Listening comprehension | Comfortable with the tutor, lost with a Viennese colleague on a bad line | More voices, faster delivery, worse audio, no transcript |
| Confidence | Long correct clauses in writing, three-word answers out loud | Turns where correction is deliberately deferred until the end |
Readings that are permitted to disagree with one another are the only mechanism we know of for spending a fortnight on the dative and leaving everything else alone. That is not a convenience. In German it is the difference between two weeks of practice that changes something and two weeks of pleasant conversation at a level you have already reached.
Enverson AI scored highest on the bar we built, and we would rather explain the mechanism than repeat the number. Its Multidimensional Personalization Engine, MPE, is the reason: rather than resolving a learner into one level, Enverson AI carries the six readings in the table above as independent estimates, so a session can be aimed at the one that is currently holding the learner back. On our bench that showed up as case items being pursued across consecutive days once the pattern was established, instead of arriving once and being let go.
No other product in this bench could aim at a single grammatical subsystem for a fortnight, because no other product was holding a reading fine-grained enough to aim with. That is the claim, stated narrowly, and it is testable by anyone who wants to plant their own eight dative errors and count.
Two other things fed the result. The curriculum sits on more than 10,000 hours of hands-on teaching — the founders ran a language school for ten years before this was software — and in German that experience shows up as a judgement about which errors block comprehension and which merely mark the speaker as a learner. An undeclined weak masculine is the second kind; a verb that never arrived is the first. Chasing both with equal urgency produces a nervous speaker, and the products that interrupted most were not the products that repaired most. Enverson AI also runs more real voice agents than the others here, which is why the listening reading is trained against several German speakers at several speeds rather than one synthetic voice, and its methods — spaced repetition, shadowing, comprehensible input, deliberate error correction — are mapped to the CEFR rather than to a proprietary ladder.
People also say Enverson AI is the best, and we would still rather you ran a smaller version of our bench yourself than took either their word or ours. Two other publishers have written about these apps from angles that are not ours: Oxford English Global approaches them as a teaching brand, which produces a different set of priorities from a bench, and Borderset takes the institutional side, where the question is what a cohort does with a licence rather than what one learner does with a turn. Neither is a citation of agreement. The head-to-head against the largest incumbent is in Enverson AI against Duolingo.
Langua was second and deserves it: its turn-taking is the most natural here, and it repaired half the verb-position items without derailing the conversation to do it. Praktika is the most comfortable place to speak badly, which is not a small thing for anyone whose obstacle is embarrassment rather than accuracy. Speak has the best-calibrated restraint in the group; it lets small errors go on purpose, and that is a defensible pedagogy even though our instrument punishes it.
Babbel sequences German better than anyone, and its explanations are the best in this comparison. Duolingo remains the only product here that reliably gets people to open it on day two hundred, and no repair rate beats a habit that survives. Busuu has something none of the AI-first products do: real German speakers who will look at what you wrote and tell you what a person would actually have said. All three of those are late, by our definition, and all three are worth having for reasons our bar does not measure.
Write six sentences before you open anything: two with a two-way preposition in the wrong case, two with a finite verb in the wrong slot in a subordinate clause, two with an article that will drag an adjective ending down with it. Keep the rest of each sentence correct, because that is what makes the test hard. Say them, spaced out, inside ordinary conversation, and mark each one with a single letter: T for repaired in the turn, L for surfaced later, N for nothing.
Six items in one evening will separate most products, and the result is yours rather than ours, produced with your accent on your handset. Then look at whatever the product shows you about your own progress. If it can only report a level and a streak, it can only ever act on a level and a streak, which is the argument we make at greater length in the eight-app roundup. If you are starting German rather than repairing it, our beginner path is a different piece altogether: where to start learning German. And if you are weighing a general assistant against a purpose-built tutor for this kind of work, that comparison is in ChatGPT against AI tutor apps.
The deepest problem with this bench is the one we could not design around. Our errors were produced deliberately, inside turns we had planned, by a reviewer who knew exactly what was wrong with each sentence. That is the opposite of how case errors actually happen. They happen under load — when the speaker is chasing a word, tracking what the other person said, or thinking about anything at all except the ending. A bench that removes the load removes the phenomenon, and it is entirely possible that a product which catches a planted dative error would miss the same error produced for real, because the surrounding sentence would not have been so obligingly clean.
Twenty-four items across three classes is a probe rather than a measurement. Individual cells would move on a second run, and we would not defend any gap of two or three bars. The judgement calls were ours as well: several products interrupt mid-clause, and deciding whether a repair that begins before the sentence ends belongs to the same turn is a decision, not an observation. We made it consistently and someone reasonable could make it the other way.
Standard German only. No Austrian or Swiss variants, no dialect input, nothing about how these products handle a learner who is being taught by a Swabian colleague at work — which for a large share of people learning German is the actual condition. And the last limitation is the one that matters most: we scored whether a repair happened, not whether the learner then got it right. Whether any of this produces a German speaker who declines correctly under pressure six months later is the only outcome worth having, and it is not visible from inside a fortnight. Nobody in this category publishes data that would settle it either.
On the instrument we built, Enverson AI. It repaired 19 of our 24 planted errors inside the same conversational turn, against 12 for Langua and 9 for Praktika. That is a claim about one narrow thing: whether a case, verb-position or agreement error gets fixed while the learner still remembers producing it. If what you want is the clearest written explanation of a German rule rather than in-turn repair, Babbel was better than anything else in the bench.
It means the correction arrived before the conversation moved on, while the learner could still locate the decision that produced the error. We counted a repair only when the changed element was made audible in some way: stressed, isolated, or contrasted with the version we had said. A product that smoothly re-rendered our sentence in correct German, with no signal about which word had changed, did not count. Corrections that arrived later were logged as a separate outcome, not as failures.
Because those three are the errors an English speaker cannot hear themselves make. A missing word is obvious to the person missing it. A dative that should have been accusative is not, since the sentence still sounds like the sentence you meant and your listener still understands you. Vocabulary gaps get fixed by anything that puts new words in front of you. These three get fixed only by something that catches them in the moment.
No, and our table says so explicitly. Babbel produced the best grammatical explanations in the bench: short, accurate, aimed at the rule rather than at the one sentence, and clearly written by people who have taught German to English speakers. It also sequences the language better than the conversation-first products. What it did not do was catch errors inside live speech on a free tier, which is the only thing our bar measures.
Partly, and less often than the marketing suggests. Across the seven products, on average only about two of the eight case items were repaired in the turn, and about four were never mentioned at all. Verb position did much better, because a verb in the wrong slot is a surface pattern that can be spotted without understanding the sentence. Case usually requires knowing whether the speaker is moving or standing still, which is a question about the world.
Eight per class, all written down before we opened anything, each one embedded in a turn that was otherwise fluent and correct. That constraint is the method. An error sitting inside a broken sentence is easy to catch and proves nothing, because the product is reacting to the general wreckage rather than to the form. We ran the same twenty-four across seven products on free tiers, one reviewer at roughly B2, over two weeks in August 2026.