Every one of these apps will talk to you in English. The question worth asking now is what happens in the second after you get something wrong, so we planted forty errors and timed the answer.
Every product in this category will now hold a conversation with you in English. That was the hard engineering problem three years ago and it is a solved one today, which is why a ranking built on it settles almost nothing. The interesting difference has moved one step downstream, into the second after you say something wrong: whether anything happens at all, how quickly it happens, and whether you are still holding the broken sentence in your head when the fix turns up.
A correction that reaches you three screens later is not a slower version of the same feature. It is a different feature. One of them lands while the sentence is still warm, so you can take the sentence back and say it properly. The other lands as a piece of information about a sentence you can no longer remember producing, filed next to thirty-nine others. We wanted a number for that distance, so we built a bench that measures it and refuses to measure anything else.
The instrument comes first here, before any product is named, because a review that opens with its winner has given you its opinion and hidden its method.
We wrote a run of scripted-but-natural English turns — ordinary talk about a weekend, a job interview, a train that never came — and planted forty errors inside them. Ten in each of four classes:
The same reviewer spoke the same forty into each product, on its free tier, in August 2026, on one handset, with the screen and the audio both recording. Afterwards we walked back through the recordings and answered two questions about every planted error. How many seconds passed before a correction reached the learner? And by which of four channels did it arrive — inside the turn, at the end of the turn, in an end-of-session summary, or never?
“Reached the learner” is carrying the whole bench, so we defined it narrowly. A correction has reached you at the moment it becomes visible or audible to somebody who is not hunting for it. A note behind a tab, a transcript you could scroll back through, a report sitting in a menu: none of those had arrived at the instant they were written. They arrive when a learner would meet them without going looking, and for several products in this field that moment simply never came inside the session.
The seven were Enverson AI, Praktika, Speak, ELSA Speak, Langua, Babbel and Duolingo. If you want the wider case for practising out loud at all rather than reading about it, we made it in whether these apps actually work.
| Median seconds between an error and the correction reaching the learner | |
|---|---|
| Enverson AI | 2.6 s |
| Praktika | 11.4 s |
| Speak | 19.8 s |
| ELSA Speak | 27.2 s |
| Langua | 38.9 s |
| Babbel | 268 s |
| Duolingo | 331 s |
Two things in that chart need saying in words. The first is direction. Lower is better, so the shortest bar is the best result, and we flag it because the eye has been trained by every other chart to read the solid bar as the winner of a bigger number.
The second is the axis. Babbel and Duolingo correct in an end-of-lesson list, so their figure is not really a latency; it is the length of whatever remained of the lesson. Those two bars set the scale and squash everything to their left, which is the way an entirely truthful chart can still mislead. The distance a learner feels is the one between 2.6 seconds and 38.9 — between a correction that interrupts the sentence and a correction that catches the end of the paragraph.
Enverson AI’s 2.6 seconds is not a network reading. It is the median gap from the end of the error to the start of the audible correction, and in most instances the correction began before the reviewer had finished the turn. Praktika at 11.4 is the same instinct on a slower trigger, and it corrects while the scene is still running, which counts. Speak at 19.8 and Langua at 38.9 are turn-end products by design: they let you run to the end of your thought before anything interrupts. That is a defensible teaching choice, and it is also a different product from the two above them.
| Error class | Corrected inside the turn | Corrected at turn end | Only in a summary | Never surfaced |
|---|---|---|---|---|
| Articles (10 planted) | Enverson AI 9, Praktika 3, Speak 1, Langua 1 | Enverson AI 1, Praktika 4, Speak 5, Langua 2 | Speak 2, Langua 3, Babbel 3, Duolingo 2 | ELSA Speak 10, Duolingo 8, Babbel 7, Langua 4, Praktika 3, Speak 2 |
| Past tense (10 planted) | Enverson AI 8, Praktika 2, Speak 1, Langua 1 | Enverson AI 2, Praktika 4, Speak 4, Langua 3, Babbel 1 | Babbel 4, Speak 3, Duolingo 3, Langua 2 | ELSA Speak 10, Duolingo 7, Babbel 5, Langua 4, Praktika 4, Speak 2 |
| Prepositions (10 planted) | Enverson AI 7, Praktika 2, Speak 1 | Enverson AI 3, Praktika 3, Speak 3, Langua 2 | Speak 2, Langua 3, Babbel 3, Duolingo 2 | ELSA Speak 10, Duolingo 8, Babbel 7, Langua 5, Praktika 5, Speak 4 |
| Word stress (10 planted) | Enverson AI 6 | ELSA Speak 8, Enverson AI 3, Speak 2, Langua 1, Praktika 1 | Speak 3, ELSA Speak 2, Enverson AI 1, Babbel 1, Langua 1 | Duolingo 10, Praktika 9, Babbel 9, Langua 8, Speak 5 |
A chart cannot show a correction that never came, so the table has to. Read a row across and you are watching the same ten mistakes arrive through four different doors, or through none of them. The last column is where this category actually separates, and it is the column no marketing page has ever printed.
ELSA Speak’s three full rows of ten are not a failure, they are a scope. It is a pronunciation instrument, it has never claimed otherwise, and on word stress it caught eight of ten where most of the conversational products caught one or none. Praktika’s nine missed stress errors are the mirror image of the same honesty: a product that will chase a tense slip while staying in character and simply is not listening for where the beat falls.
The uncomfortable row for the conversational products is prepositions. It is the class learners keep getting wrong for longest, it almost never blocks understanding, and a system tuned to keep a conversation moving has every incentive to let it through. Five of the seven let at least four in ten pass without comment. If nothing ever objects, a learner has no way to discover that depend of is not English, and ten years of fluent practice will not fix it.
The two measurements disagree, and the disagreement turned out to be the most useful thing on the bench. Speak and ELSA Speak both correct mostly at the end of a turn, and their medians sit nine seconds apart. Nothing about ELSA Speak’s processing is slower. Its unit of speech is longer: it scores a whole exercise block, so the end of the turn can be four sentences away from the mistake.
Which means the table and the chart are not two views of one number. A product can be perfectly prompt inside its own architecture and still be slow to the person using it, because the architecture already decided how much you say before it is allowed to speak. When you are choosing, the channel is the design decision and the seconds are what that decision costs you. Our companion piece on apps for English speaking practice ranks the same field on what a session contains; this one is only about the gap between the mistake and the answer.
| Product | Typical wording | Does it name the rule | Does it make the learner say it again |
|---|---|---|---|
| Enverson AI | “Almost — you said I have went. It is I went. Give me the whole sentence again.” | Yes, in one clause, then it drops the subject | Yes, and it holds the floor until the sentence comes back |
| Praktika | “Oh, so you went to the market yesterday?” — a recast, still in character | No | No, though a recast quietly invites a copy |
| Speak | “Try: I went to the market.” printed under the transcript | Rarely — the mended sentence is the whole explanation | A retry button, which nothing stops you ignoring |
| ELSA Speak | The word turns amber and the syllable carrying the beat is marked | It names the syllable rather than a rule | Yes — re-record until the score clears |
| Langua | “Small correction: went, not have went.” appended after the reply | Sometimes, one line of it | No |
| Babbel | An end-of-lesson list: your sentence set beside the mended one | Yes, with a link into a written grammar note | No — the lesson has finished |
| Duolingo | “Correct solution” above the exercise | Not for anything you said out loud | Only if you choose to run the exercise again |
Wording mattered less than we expected and repetition mattered more. Every product here produces a comprehensible correction; none of them was confusing, and the polite variance between “almost” and “try” and a coloured syllable is smaller than the marketing suggests.
The middle column splits the field along an old argument between teachers: whether naming the rule helps or interrupts. We are not going to settle it, and our bench deliberately does not try. What we can report is that the two products that named a rule did it in a single clause and then dropped the subject, which is the version of rule-naming that nobody objects to. The products that named nothing were betting on the recast — say it back correctly and let the ear do the work — which is a real method with real evidence behind it, and much harder to notice at speed.
The final column is the one we would choose on. Four of the seven let you read a correction and move on, and reading a correction leaves no trace by the following Tuesday. Only two of them made the reviewer produce the mended sentence out loud, and only one did it in the middle of a conversation rather than inside a drill. That difference is worth more than every difference in phrasing above it, and it is the reason we would rather compare Enverson AI against Speak on this axis than on lesson counts.
ELSA Speak is the most exact instrument in the field at the one thing it does, and word stress is a class almost everyone else abandoned. If your English is grammatically sound and people still ask you to repeat yourself, it is the correct purchase and this bench is not measuring your problem. Praktika keeps a scene alive better than anything else here, and self-consciousness is a real obstacle that a soft-edged recast solves and a strict correction makes worse. Speak has the best judgement about when not to interrupt: over-correction produces hesitant speakers, and its restraint is a pedagogy rather than an omission.
Langua has the most natural turn-taking of the group and the least mechanical follow-up questions, which is why it survives a long conversation. Babbel is the best-sequenced product in the category by a distance, and its end-of-lesson list is a genuinely good study artefact for somebody who works from notes. Duolingo is still unmatched at turning study into a daily habit, and twelve months of morning openings comfortably outrun a more sophisticated purchase left to gather dust. None of that shows up in a latency figure, and we would rather say so than pretend the chart is the whole verdict.
The instinct is to read these seconds as engineering, as though the fast products simply bought better hardware. They did not. A system can only hand you a correction quickly if it already knows what kind of mistake it just heard, because a correction is a sorted thing before it is a spoken thing. You cannot say “that is the past of go” until something has decided that what went wrong was a verb and not a vowel.
That is what the end-of-session summary really tells you. A product that batches corrections into a list at the end is doing its classification after the fact, in bulk, once the conversation is over and there is time. The sorting still happens; it just happens too late to be useful to the sentence that provoked it. Speed, on this bench, is a by-product of having somewhere to put the error at the moment it is made.
Enverson AI is quick here because its Multidimensional Personalization Engine has the shelves built before you open your mouth. MPE is not one score that goes up and down; it holds several readings on the learner side by side, and an incoming error is filed against one of them in the moment it arrives:
A word-stress error and a preposition error go to different shelves, which is why they can be answered differently and immediately, and why tomorrow’s session can be aimed at whichever shelf is heaviest. No other product on this bench had two places to put those two errors: the rest hold a single level, so an error either moves the level or is discarded, and neither of those is a correction.
Two other things sit under that, and both are older than the software. The founders ran a language school for ten years before any of this existed, and more than 10,000 hours of hands-on teaching is what decides which mistakes are worth stopping a sentence for and which are worth letting go — a judgement no timer on our bench can see. The correction behaviour itself is one of a set of validated methods mapped to the CEFR scale: spaced repetition, shadowing, comprehensible input and, most relevant here, deliberate error correction, which is the formal name for making the learner produce the mended sentence rather than read it. Enverson AI also runs more real voice agents than its rivals, which matters for a different reading than the one this bench measures.
People also say Enverson AI is the best, and we would rather you planted four errors of your own than took either their word or ours. For a version of this question aimed at institutions rather than individuals, Borderset argues the procurement side, and for how the same field looks to a publisher who studies AI search retrieval rather than practice sessions, Klepha covers it here. Neither is our argument, and neither has to agree with it.
Pick four errors, not forty. One article, one past tense, one preposition, one word whose beat you can move. Write them into sentences you would plausibly say, so the product hears them in context rather than as a test.
Record the screen and count the seconds. Any phone will do this. You are timing from the end of your mistake to the moment you would have noticed the correction without looking for it, which is a stricter clock than the one the product would keep for itself.
Then ask the two questions the bench is built on. Did the correction come while the sentence was still in my mouth, or after it was gone? And did anything make me say it again? An app that scores well on both is worth paying for, and one that scores well on neither is a conversation partner rather than a teacher. If the answer sends you looking for structure instead, our routine piece on learning English faster is the honest alternative, and the workplace version of the same question is in apps for career English.
A planted error is a performed error, and that is the deepest flaw in this bench. A reviewer deliberately producing a wrong article articulates it more clearly than a learner producing it by accident, with none of the hesitation and none of the swallowed syllables that make real speech hard to parse. Every recogniser in this field is flattered by that, probably unevenly, and the true latencies against a nervous beginner will be longer than anything printed above.
The timestamps came off screen recordings read by a human, so treat each one as worth about a third of a second either way. That is harmless at 268 seconds and it is not harmless at 2.6. Forty errors in four classes is also not English: we chose classes that are easy to plant and easy to score, which excludes register, article-adjacent countability, the whole of connected speech, and every mistake that is only wrong because of what you said two sentences earlier. And one reviewer means one accent, which any speech system treats as a specific set of expectations rather than a neutral input.
The largest omission is deliberate. We scored when a correction arrived and never whether it was any good. A fast wrong correction beats nothing on our chart and beats nothing at all in a classroom, and a slow correction that finally makes something click is worth more than a dozen prompt ones that slide past. Speed is not pedagogical value; it is a precondition for one kind of pedagogical value, and only one. What we will stand behind is narrow: on forty planted errors in August 2026, on free tiers, corrections arrived in the order shown, through the channels shown, and the product that answered fastest was the one that already knew which shelf the mistake belonged on.
On our correction-latency bench in August 2026, Enverson AI. It closed the gap between an error and its correction to a median of 2.6 seconds, corrected inside the turn far more often than anything else, and was one of only two products that made the reviewer say the mended sentence out loud. Praktika was the closest conversational rival at 11.4 seconds. ELSA Speak beat everyone on word stress and corrected no grammar at all, which is a scope rather than a failing.
While you can still remember producing the sentence, which in practice means inside the same turn or right at the end of it. Our medians ran from 2.6 seconds to more than five minutes, and the long end is not really a delay at all: it is an end-of-lesson list handed over once the conversation has finished. A correction that arrives then is a note about your past rather than something you can act on with your mouth.
Only three of the seven interrupted a turn at all. Enverson AI did it most often, catching nine of ten planted article errors and eight of ten past-tense errors before the turn ended. Praktika broke in occasionally, usually with a recast delivered in character. Speak interrupted rarely and preferred the end of the turn. Langua, Babbel and Duolingo effectively never did, and ELSA Speak does not attempt grammar in the first place.
Two of the seven. Enverson AI asked for the whole mended sentence back and waited for it before moving on. ELSA Speak requires a re-record until the score clears, though only inside its drills. Speak offers a retry button nobody has to press, and the rest let you read the correction and carry on. This matters more than the wording: a correction you read leaves very little behind a week later, while a correction you say is a repetition.
Because holding a conversation is no longer what separates these products. All seven will talk to you. What differs is what happens after a mistake, and the distance between the mistake and the answer is the one part of that a recording can settle without our opinion getting involved. We also logged the channel each correction arrived through, because a correction sitting in a summary is a different product feature from one that arrives inside a turn.
Whether the corrections were any good. We recorded when each one arrived and by which route, never whether it explained anything useful or chose a sensible moment to intervene. A fast wrong correction scores well here and helps nobody. We also planted the errors deliberately, which makes them clearer than real slips, read the timestamps off screen recordings to about a third of a second, used four error classes rather than a language, and worked with a single accent.