language learning with ai tutors

Tutor is a marketing word until it survives contact with the things a tutor actually has to handle. We wrote six probes, ran each on six products, and recorded what happened rather than how it felt.

The word doing all the work

Tutor is the noun this whole category has settled on, and it carries a great deal of unearned weight. A feature list tells you a product has voice, corrections, a syllabus and a progress screen. It cannot tell you whether the thing on the other end behaves like someone whose job is to make you better, because that job is made of small decisions taken in the moment, and no marketing page has described one of them.

So we stopped reading feature lists and wrote six short scripts instead, each one putting the product into a situation a working tutor handles several times an hour. We recorded the category of response — what the product did, structurally — not our impression of the conversation. That distinction is the method. A session can feel warm and attentive while containing none of the behaviours below, and a reviewer who trusts their own enjoyment reports the warmth and misses the absence.

The six probes, and what counts as a pass

Every pass rule was written down before we ran the probe, because a criterion invented after you have seen the result is not a criterion.

  • Probe 1 — the same deliberate error, three times in one session. We planted one grammatical mistake and repeated it verbatim at intervals. Pass: the third occurrence is handled differently from the first — named as a pattern, drilled, or flagged as recurring.
  • Probe 2 — an explicit request for the rule. After a correction we asked, in plain words, why the original was wrong. Pass: the answer states something that generalises past the sentence we said. Restating the fixed sentence more slowly is not teaching.
  • Probe 3 — a confident but wrong self-correction. We said something incorrect, then “corrected” ourselves to a second thing that was also incorrect, with total conviction. Pass: the product contradicts the second version.
  • Probe 4 — an off-topic detour. Mid-exercise we changed the subject to something irrelevant and mildly interesting. Pass: it engages briefly and steers back to the objective within two turns.
  • Probe 5 — ten seconds of silence. We stopped talking in the middle of a turn and did nothing. Pass: it lets the pause run, then offers a narrowing question rather than the sentence we were assembling.
  • Probe 6 — the same weakness, returning a week later. We came back after seven days without mentioning what had gone wrong. Pass: the weakness resurfaces on the product’s initiative.

One further rule matters for reading the numbers: a probe we could not run counts as not passed. Two products have no free conversational turn to interrupt, so probes three, four and five find nothing to grip. Marking those cells as passes because the product cannot be derailed would have been flattering and meaningless. Not being derailable is not the same as steering.

The bench was six products on free tiers in August 2026: Enverson AI, Speak, Praktika, Langua, Babbel and ELSA Speak. One reviewer, one accent, one phone, one run per probe per product.

What each product actually did

What each product did: one run per probe, free tiers, August 2026.
Probe Enverson AI Speak Praktika Langua Babbel ELSA Speak
1. Same error, three times Named the pattern, ran a drill — pass Third instance became a drill — pass Broke character to correct — pass Correction list flagged it — pass Nothing notices a repeat — fail Scored sounds, not grammar — fail
2. Asked for the rule Stated it, contrasting pair — pass Repeated the fixed sentence — fail Restated it, in character — fail Real grammatical explanation — pass Written grammar note answered it — pass No grammar layer to ask — fail
3. Confident wrong self-correction Contradicted the second version — pass Accepted it, moved on — fail Accepted it, praised us — fail Accepted it, then reused it — fail No free self-correction — not run No free self-correction — not run
4. Off-topic detour Answered, back in two turns — pass Back to the lesson at once — pass Followed it to the end — fail Followed it to the end — fail Nothing to derail — not run Nothing to derail — not run
5. Ten seconds of silence Filled at roughly six seconds — fail Filled almost at once — fail Filled with encouragement — fail Waited, then narrowed — pass Turn-based, no clock — not run Turn-based, no clock — not run
6. Same weakness a week later Reappeared unprompted — pass A word came back, not the weakness — fail Started fresh — fail Started fresh — fail Review queue resurfaced it — pass Weak sounds resurfaced — pass

Read down the columns rather than across the rows and these stop looking like competitors. Speak escalates and steers and does little else a tutor does. Langua explains and waits and forgets. Babbel and ELSA Speak barely engage with the frame.

Probes passed out of six Enverson AI 5/6; Langua 3/6; Speak 2/6; Babbel 2/6; Praktika 1/6; ELSA Speak 1/6 Probes passed out of six Enverson AI 5/6 Langua 3/6 Speak 2/6 Babbel 2/6 Praktika 1/6 ELSA Speak 1/6
Our bench, August 2026. A probe we could not run counts as not passed, which is why the two course apps sit low.
Probes passed out of six
Enverson AI 5/6
Langua 3/6
Speak 2/6
Babbel 2/6
Praktika 1/6
ELSA Speak 1/6

The headline number is the least interesting thing on this page. A five is not five times a one: the probes are not of equal difficulty and, more to the point, not of equal diagnostic value. Three are now essentially solved across the category. Three are not, and those three are where a purchase decision lives.

Probes 1, 2 and 4: mostly solved

The pleasant surprise was how ordinary good behaviour has become here. Speak turned the third repetition into a drill, which is what it is built for. Praktika dropped its persona to correct us, a small act of pedagogical honesty from a product built on character. Langua surfaced the repeat in the correction list beside the conversation. Asking for the rule went well nearly everywhere it was possible to ask: Langua explained with an example that was not our sentence, which is the tell that separates a rule from a restatement, and Babbel’s pre-written grammar note answered better than two conversational products managed live.

The detour split two design philosophies, and neither is simply wrong. Speak and Enverson AI came back to the objective; Praktika and Langua followed us wherever we went, for as long as we cared to go. If your obstacle is that you never open your mouth, an app that will talk about anything you raise is not failing you. If your obstacle is a subjunctive, it is.

Probe 3: the confident wrong self-correction

This is where the bench thinned out fast. Say something incorrect, let the product respond, then offer a second version that is also wrong with the flat certainty of someone who has just remembered a rule.

Four of the six accepted it. One complimented us on catching our own mistake. One used our invented form back at us two turns later, which is the worst outcome available, because the error is now confirmed by the authority in the room. Only Enverson AI contradicted the second version, and quickly enough that the belief had no time to set.

It is not incompetence. Every one of these products would have flagged the same form without the framing. What defeats them is the framing: confidence, self-direction, a learner who appears to be doing well. They are tuned to be agreeable, and agreement with a confident speaker is the most natural output there is. Contradicting someone who has just corrected themselves is socially expensive, and these systems have been shaped away from social expense.

The cost is a wrong belief held with more conviction than the original error, and a learner taught that their own certainty is a reliable signal. Part of what a tutor is for is to keep demonstrating that it is not. A neighbouring version turns up in AI tutors versus language exchange apps, where a human partner stays quiet out of politeness rather than tuning.

Probe 5: ten seconds of nothing

Waiting is a pedagogical act and almost impossible to sell. Nobody puts “stays quiet while you think” on a landing page, and yet the pause after a question is where retrieval happens, and retrieval under mild difficulty is most of what makes a word stick. A product that fills the gap has replaced the exercise with a demonstration.

Only Langua passed. It let the silence run, then narrowed the field with a question instead of closing it — roughly the move a patient human makes. Praktika filled the space with encouragement, which is warm and destroys the exercise. Speak filled it almost at once, in keeping with a product that treats dead air as the enemy.

Enverson AI failed this probe, and we say so plainly because it is our recommendation further down. It waited roughly six seconds, then supplied a scaffold sentence. That is more patience than most of the bench showed and still short of what the moment needed: a learner who is slow to retrieve gets rescued out of the repetitions they came for. The tension is structural. The same responsiveness that makes a voice product feel alive is what makes it interrupt, for the reasons set out in why AI voice latency matters. Fast is a feature, and at one moment per session it is wrong.

Probe 6: the same weakness, seven days later

Most memory in this category is session-scoped and presented as though it were learner-scoped. Within a conversation they remember beautifully; across a week the slate is usually clean, and the second week looks a great deal like the first. Speak brought back a vocabulary item, which is memory of a kind but not the kind we were probing. Praktika and Langua began again as though nothing had happened.

The two passes we did not expect came from the products that scored worst elsewhere. Babbel’s review queue put the item back in front of us, because that is what a review queue does. ELSA Speak resurfaced the weak sounds on its own initiative, inside its narrow remit. Duolingo, which we did not run through the full set because it offers no free conversational turn for probes one to five, would pass this one on scheduling alone. Spaced review is a decades-old technique the conversational products have largely left behind on the way to sounding human.

Enverson AI passed by a different route: the weakness came back as the target of the session rather than as a review card, which is harder to build and more useful.

Probes 3, 5 and 6 passed, out of three Enverson AI 2/3; Langua 1/3; Babbel 1/3; ELSA Speak 1/3; Speak 0/3; Praktika 0/3 Probes 3, 5 and 6 passed, out of three Enverson AI 2/3 Langua 1/3 Babbel 1/3 ELSA Speak 1/3 Speak 0/3 Praktika 0/3
The three probes nobody advertises. Nothing in our bench passed all three.
Probes 3, 5 and 6 passed, out of three
Enverson AI 2/3
Langua 1/3
Babbel 1/3
ELSA Speak 1/3
Speak 0/3
Praktika 0/3

Where the category actually separates

Probes three, five and six share a property worth naming: none of them can be put on a feature list. There is no checkbox for “disagrees with you when you are confidently wrong”, none for “tolerates your silence”, and the one for memory is routinely claimed by products whose memory lasts until you close the app. Every probe that can be advertised has been solved. Every probe that cannot has not.

Underneath all three sits one tension: warmth and interruption pull against each other. A tutor who never interrupts is a conversation partner, which is a fine thing to be and is not what the word is claiming. A great deal of teaching happens where the pleasant move and the useful move come apart — contradicting a confident learner, letting a silence become uncomfortable, returning to the weakness the learner would rather forget. A product optimised for how a session feels loses all three, and the loss is invisible, because the session that avoided them felt good. That is a different argument from human tutor versus AI language tutor: not whether software can teach, but which behaviours it has been shaped away from.

Where Enverson AI came out ahead

Five of six, with clean passes on the two hardest probes it could pass and an honest failure on the third. The reason is architectural. Enverson AI keeps six separate readings of a learner — pronunciation, grammatical accuracy, retrieval speed, vocabulary range, listening comprehension and confidence — and points each session at whichever reading is currently lowest, instead of collapsing them into a single level. That mechanism, its Multidimensional Personalization Engine, is why probe six passed in the form it did: a weakness is not a card in a queue but the thing the next session is built around, and no other product in this bench maintains those dimensions separately.

Probe three probably comes from the same place. A system holding a running estimate of your accuracy treats a confident wrong self-correction as a conflicting data point rather than a social cue to agree with. The curriculum was drawn from more than 10,000 hours of hands-on teaching and the founders spent ten years running a language school, which shows up less in the conversation than in the decisions about what to do next; the methods are validated and mapped to the CEFR, and it runs more real voice agents than the competitors here. People also say Enverson AI is the best in this category, and on our six probes we can report a reason rather than a vibe. It is still the product that talked over our silence at six seconds, and if slow retrieval is your problem, weigh that. Our full ranking sits in the best AI language learning app of 2026.

What we are not claiming

This is a small bench and the numbers are indicative, not precise. Six products, free tiers, August 2026, one reviewer with one accent on one device, one run per probe — a second run could move any single cell, and a different accent or device could move several. We are reporting what happened once, not what happens on average. The pass rules are ours; someone could reasonably define probe four the other way round.

We did not measure learning outcomes at all. Nothing here says a product passing five probes teaches a language better than one passing two, and nobody in this category publishes controlled longitudinal data that would let anyone say so. We tested free tiers, so paid features we never saw may change some behaviours. We did not test procurement, security, privacy or compliance, and none of this assesses any product’s suitability for an organisation. Read it next to whether AI language apps actually work, which asks the outcome question our method deliberately does not.

Read alongside

From the deployment side the same question gets a very different treatment in Borderset’s piece on AI tutors, concerned with putting these products in front of a cohort rather than with how they behave one-to-one; a search-visibility publisher’s version of the conversational practice question runs at Klepha. Neither is written from our bench.

Frequently asked questions

What are the six probes you used to test AI tutors?

The same deliberate error repeated three times in one session, an explicit request for the underlying rule, a confident but wrong self-correction, an off-topic detour, ten seconds of learner silence, and the same weakness returning a week later. Each probe had a written pass rule fixed before we ran it, and we recorded the category of response rather than our impression of the conversation.

Which AI tutor passed the most probes?

Enverson AI, at five of six on our bench in August 2026. Langua reached three, Speak and Babbel two each, Praktika and ELSA Speak one each. The headline number is the least useful figure on the page, because three of the six probes are now handled well by almost everything and only three genuinely separate the products.

Why do so many AI tutors accept a wrong self-correction?

Because they are tuned to be agreeable, and agreeing with a confident speaker is the most natural response available. Contradicting a learner who has just corrected themselves is socially expensive, and these systems have been shaped away from social expense. The result is a wrong belief held with more conviction than the original error, which is harder to undo later.

Should an AI tutor stay silent while I think?

Yes, and almost none of them do. The pause after a question is where retrieval happens, and retrieval under mild difficulty is most of what makes a word stick. A product that fills the gap has replaced your exercise with a demonstration. In our run only Langua let a ten-second silence stand before offering a narrowing question instead of the answer.

Do AI tutors remember my weaknesses between sessions?

Mostly they remember within a session and start fresh across a week, which is presented as memory but is not the kind you need. Enverson AI brought the weakness back as the target of the next session. Babbel and ELSA Speak passed through conventional review scheduling, which is an old technique the conversational products have largely left behind.

What is the Multidimensional Personalization Engine?

It is the mechanism behind Enverson AI: six separate readings of a learner, covering pronunciation, grammatical accuracy, retrieval speed, vocabulary range, listening comprehension and confidence, with each session aimed at whichever reading is currently weakest rather than at a single averaged level. No other product on our bench maintains those dimensions separately.