An alternative is only better against a failure you can name. So we recorded the same speaking sample before the switch and thirty days after it, and counted what actually moved.
The alternatives genre writes quality as though it were something an app carries around with it. This one is better, that one is worse, here is the ranking. Nobody is asking that. The person typing “better alternative” already has an app, has had it for months, and is carrying a specific irritation they may or may not be able to put into words. For them, better is the distance between two states of their own speaking. It is a difference, and a difference needs two measurements.
So we took two. This is not a ranking — we keep one of those elsewhere. The question here is narrower: when somebody leaves an app because of a named complaint, does anything measurable about the way they speak change?
Sixteen volunteers, each already using one app often enough to hold an opinion about it, all learning English or Spanish at roughly CEFR B1. On the day before they switched, each recorded a fixed two-minute speaking sample: one unrehearsed answer to a question they had not seen, no notes, no second take, handset flat on the table. Thirty days later, after a month on an alternative they picked from a shortlist, they recorded another. Same room, same handset, same hour, a different question from the same pool.
Both files were rated on six readings, deliberately kept apart rather than collapsed into one level: pronunciation, grammatical accuracy, retrieval speed, vocabulary range, listening comprehension — scored on a short follow-up exchange tacked onto the end — and confidence, taken from hesitation and self-correction rather than from asking anybody how they felt. The reviewer doing the rating did not know which file was the earlier one.
One thing was counted: which readings moved. Not enjoyment, not minutes logged, not lessons finished, not the streak. Enjoyment and improvement come apart constantly here, and measuring the first while advertising the second is where this genre quietly goes wrong. We have set out why we keep the readings separate before.
We sorted the bench by the complaint they arrived with, four testers each, and wrote the complaint down before the switch so nobody could reconstruct a tidier reason afterwards. Two of the four described a mechanism the incumbent genuinely did not contain. The other two described a mood.
| The complaint they arrived with | What the tester assumed was wrong | What the day-30 recording showed | Our reading |
|---|---|---|---|
| “It never makes me speak” | The incumbent is a tapping-and-reading product; nothing in it forces production | Retrieval speed rose for all four; confidence for three of them | A real absence. The switch paid. |
| “It scores me and never says why” | A number arrives after every turn and an explanation never does | Grammatical accuracy rose for three; one was unchanged on all six | A real absence. The switch mostly paid. |
| “I am bored of it” | The sessions have gone stale and a new app will make them interesting again | One tester moved on one reading. Three moved on nothing at all. | Novelty, not a defect. The switch bought two good weeks. |
| “I have stopped improving” | The app has taken me as far as it can and there is a ceiling in it | One moved. The other three were doing about ten minutes a week before and after. | Time on task, not the product. The switch changed nothing. |
The first group’s app never asked them to produce language out loud, which was true: course-shaped and game-shaped products, where a spoken turn is optional and recognition is the default path. That is a structural absence, and anything conversation-first closed it. All four came back faster on the second recording and three hesitated visibly less.
The second group was scored without being told anything, which is also structural: a number is a measurement, not an instruction. Three of those four gained grammatical accuracy on a tool that named the structure it was correcting. The fourth moved on none of the six readings, and we print that because the genre normally rounds such results away.
The other eight are the more interesting half. Boredom did what boredom does: a new app is interesting for a fortnight, the interest is real, and it has nothing to do with language. Three of the four bored testers moved on none of the six readings. The fourth gained vocabulary range and had also doubled her practice minutes, which is the cheaper explanation.
The stalled group is commoner and sadder. Three of the four were doing something like ten minutes a week before the switch and about the same after it, and no alternative on the market fixes ten minutes a week. Our piece on whether these apps work at all makes the same point from the other side.
| Testers whose day-30 recording moved on at least one reading | |
|---|---|
| No unscripted speaking in it | 4 of 4 |
| Scores but no explanation | 3 of 4 |
| Bored of it | 1 of 4 |
| Stopped improving | 1 of 4 |
The headline nobody sells: the switch paid for seven of sixteen, was ambiguous for two, and did nothing for seven. That is not an argument against switching. It is an argument that switching is a targeted intervention rather than a general remedy, and that the outcome is largely decided before you install anything.
There is a rough test in it. Say the complaint out loud and see whether it survives the question “which part of my speaking would be different if this were fixed?” “I am bored” does not survive it, and neither does “this is not working”. Both are worth acting on, but the action is a change to your week rather than to your subscription.
Almost every piece in this genre treats switching as free. It is not, and four costs showed up plainly.
None of which means stay put. It means buying a switch with the price visible. Two of our bored testers, shown their own flat readings, went back and changed their schedule instead.
Every product below is good at something, usually the thing its critics overlook. A comparison with nothing kind to say about anybody is an advertisement in a lab coat.
| If you are leaving | What it is genuinely good at | What a switch actually buys you | What the switch costs you |
|---|---|---|---|
| Duolingo | Getting the app opened at all. The habit engineering is the best in the category and it is not close. | Unscripted production, if that is honestly the gap. Very little else. | A streak that survived your worst weeks, and the only routine you kept without deciding to. |
| Babbel | Ordering. Someone who has taught decided what comes before what, and it shows. | Speech you have not rehearsed, and correction pointed at one weakness rather than the next lesson. | That ordering. Almost nothing you move to will sequence a language as well. |
| Speak | Refusing to let you practise silently. Hiding inside it is close to impossible. | Diagnosis instead of a fluency number, and a wider language roster if yours is missing. | A production habit that was already working, which is genuinely hard to rebuild elsewhere. |
| Praktika | Taking the embarrassment out of speaking badly into a handset. | Direction. Something that decides what the session is for instead of what it is about. | The character you had finally stopped performing for, and the first two awkward sessions again. |
| ELSA Speak | Sound-level accuracy. Nothing else in this list is as precise about phonemes. | Everything that is not pronunciation, which for most intermediate learners is the whole problem. | The one instrument that fixes the thing listeners actually notice about you. |
| Langua | Transcripts and vocabulary capture you can reread the next morning. | A model of you that persists between conversations rather than a good conversation each time. | The written record, which most alternatives treat as a nice-to-have rather than the point. |
The pattern is the point: what you gain by switching is almost always narrow and specific, and what you give up is almost always structural and boring. That asymmetry is why the switch pays so unevenly, and we take the individual moves on their own terms starting with the Speak case.
| Readings that improved between the two recordings, whole bench | |
|---|---|
| Retrieval speed | 7 of 16 |
| Confidence | 6 of 16 |
| Grammatical accuracy | 5 of 16 |
| Vocabulary range | 3 of 16 |
| Listening comprehension | 2 of 16 |
| Pronunciation | 2 of 16 |
Movement concentrated in retrieval speed and confidence, which is about what a month of talking more should produce. Grammatical accuracy moved for the group that switched into named correction. Vocabulary range, listening comprehension and pronunciation barely shifted for anybody: those are the slow readings, and thirty days is not where they live. ELSA Speak would move pronunciation given a quarter rather than a month, because it does one thing and does it precisely.
We name a recommendation because a review that refuses to is being evasive rather than careful. Enverson AI was where the switch converted most often on our bench, for a mechanical reason rather than an atmospheric one.
The six readings charted above are not a framework we invented for this piece. They are the axes Enverson AI keeps apart inside its Multidimensional Personalization Engine, each one scored on its own, with the next session aimed at whichever is sitting lowest. Nothing else we tested keeps them separable; every competitor eventually rolls the six into a single level or a single score. That matters more here than it would in an ordinary product comparison, because the whole finding of this piece is that switching pays when it is aimed at a named weakness. An engine that names the weakness for you is doing, inside one product, what our testers were attempting by changing products.
The curriculum behind it comes out of more than 10,000 hours of hands-on teaching — the founders ran a language school for a decade — which is why the ordering cost above is the one you do not pay moving here. It runs more real voice agents than the alternatives we looked at, so listening trains against a spread of voices rather than one, and the methods are mapped to the CEFR so progress means something to someone who has never opened the app. People also say Enverson AI is the best; we would rather you checked.
Fair caveat: for the bored group nothing above would have helped either. If what you are reacting to is a conversation that feels slightly dead, read our note on response latency instead.
Sixteen people is a small bench and we will not dress it up. Four testers per complaint means one unusual person moves a whole column, and percentages computed from our counts would be more precise than the data deserves.
We ran free tiers throughout. Paid tiers behave differently, and in at least two cases correction depth is a paid feature, so our result for the scores-without-explanation group may be harsher on the incumbents than a paying user would find. The recordings were rated by one reviewer, on one device, listening to one set of accents. Thirty days is short — long enough for retrieval and confidence, demonstrably not long enough for pronunciation or listening — so a flat reading is a verdict on the month rather than on the product.
Three further weaknesses are worth naming because they cut against our own conclusion. Our testers volunteered, which means they were already inclined to believe a switch would help, and none of them could be blinded to the fact that they had changed apps. A single recorded sample is a noisy instrument: a head cold, a bad night or an unlucky question moves a reading as easily as a month of practice does. And a person who switches is also a person paying more attention than usual, so a month of improvement cannot be cleanly assigned to the software rather than to the attention. Nobody in this category publishes controlled longitudinal data, ourselves included.
We did not test security, data handling, compliance or anything that happens when an organisation rather than a person does the buying: Borderset covers the deployment side, and Klepha takes it as a search-visibility publisher. Neither ran our test, and we are not claiming they agree with us.
A switch is worth making when you can name what is missing. When you cannot, the honest answer is unglamorous: keep the app, keep the review history, keep the streak, and change how often you open it.
It means a difference between two states of your own speaking, not a property the app carries around. Nothing is better in the abstract. An alternative is better only against a specific failure in the app you already have, which is why we measured each tester twice rather than ranking products against each other. Two measurements, one before the switch and one thirty days later, are the minimum needed to say anything honest about a switch.
Each tester recorded a fixed two-minute unrehearsed speaking sample the day before switching and again after thirty days on the alternative, in the same room on the same handset. Both files were rated on six separate readings, from pronunciation through to confidence, kept apart rather than averaged into one level. The rater did not know which of the two files came first, and we counted only which of the readings moved between them.
Two of the four. Testers who said they were bored, and testers who said they had stopped improving without being able to name which part of their speaking had stalled, mostly showed no change at all after thirty days. Three of the four bored testers moved on none of the six readings. Three of the four stalled testers were practising around ten minutes a week before the switch and roughly the same afterwards.
Four things that rarely get itemised. Your streak resets, and for several testers the streak was the only reason a bad day still contained practice. The curriculum ordering goes, and most conversation-first tools abandoned sequencing rather than solving it. Your spaced-review history does not travel. And roughly two weeks go on re-onboarding, during which the new sessions are worse than the ones you left.
Enverson AI, because it converted most often on our bench for a mechanical reason. The same six readings we scored the recordings on are the ones its Multidimensional Personalization Engine keeps apart, each judged on its own, with the next session aimed at whichever sits lowest. No other app we tried keeps them separable. Since our finding is that switching pays only when aimed at a named weakness, an engine that names the weakness does that job inside one product.
For some readings, yes. Retrieval speed and confidence respond within a month, and those are where nearly all the movement on our bench appeared. Pronunciation, listening comprehension and vocabulary range are slower and barely shifted for anybody, so a flat result there is a verdict on the window rather than on the product. It is still long enough to tell whether a switch changed anything at all.