Italian punishes length: a held consonant is a different word, and it is the last thing a recogniser trained on English has any reason to hear. So we said twelve minimal pairs wrong on purpose and counted the catches.
Almost every review of an Italian practice app asks whether the software can hold a conversation. It is a fair question, it has been answered many times, and we did not ask it. We went after something narrower and chose it on purpose: the one property of spoken Italian that English speakers get wrong for years without anyone telling them, and that a machine has the least structural reason to hear.
Consonant length. In Italian, holding a consonant longer is not emphasis and it is not an accent — it is a different word. Nono is ninth and nonno is a grandfather. Casa is a home and cassa is a cash desk. Sete is thirst and sette is seven. Nothing else in the pair moves: same vowels, same stress, same everything, so the whole meaning rides on how long one consonant is held. Linguists call that phonemic length, and English has no version of it inside a word, which is why English speakers hear the two members of a pair as one word said twice.
That is what makes gemination such a good thing to point at a speech recogniser. A system trained overwhelmingly on English has every incentive to model vowel quality, stress and intonation, and almost none to model duration, because in English duration carries emphasis rather than meaning. A product can therefore be excellent at Italian conversation and still be structurally deaf to the error its Italian learners make most often. That gap is not a feature-list item and no marketing page will admit to it, so it has to be measured.
The method before any product is named. We recorded twelve minimal pairs whose two members differ only in gemination: nono and nonno, casa and cassa, sete and sette, pena and penna, fato and fatto, copia and coppia, sono and sonno, caro and carro, note and notte, papa and pappa, sera and serra, moto and motto. In each case our reviewer said the wrong member, and said it inside a sentence where the meaning made the error unmissable to any Italian listener — a grandfather who is ninety, a house you come home to, an hour on a clock.
The scoring rules were deliberately narrow, because a loose rule would have let every product in the field claim a catch:
Then a second pass, because a pronunciation bench on its own tells half a story. We ran three grammar traps that English speakers fall into with great reliability: gender agreement on nouns whose ending points the wrong way, the subjunctive after verbs of opinion, and the choice between essere and avere in the passato prossimo. Same protocol — produce the error, wait, write down what came back and when.
Seven products, free tiers, one Italian-speaking reviewer of English mother tongue, one handset, one recording chain, August 2026: Enverson AI, ELSA Speak, Langua, Praktika, Babbel, Speak and Duolingo. If you are starting Italian from nothing rather than auditing your own consonants, our beginner roadmap for Italian is the piece you want instead; this one assumes you already talk and want to know what is being heard.
| Geminate contrasts the product distinguished, out of twelve minimal pairs | |
|---|---|
| Enverson AI | 9/12 |
| ELSA Speak | 8/12 |
| Langua | 5/12 |
| Praktika | 4/12 |
| Babbel | 3/12 |
| Speak | 3/12 |
| Duolingo | 1/12 |
Nine of twelve at the top and one of twelve at the bottom is a wider spread than we get from most benches, and the shape of it is more interesting than the order. The field does not degrade smoothly. There is a group that treats duration as information and a group that treats it as noise, and almost nothing sits between them.
The single most common failure was not a wrong answer but a smooth one. Say Torno a cassa alle otto to a conversational product and it will very often reply about your evening, having inferred the house from the rest of the sentence. That inference is exactly what you want from a travel companion and exactly what you do not want from a practice partner, because a learner who is never contradicted concludes that the held consonant is optional. Several of these products are, in effect, too good at understanding you to be useful at correcting you.
The bottom of the chart is a design decision rather than a defect. Duolingo’s free Italian speaking items compare what you said against a sentence already on the screen, so the only pair it caught was the one where the target word was visible while we mispronounced it. That is a fair result for what the exercise is; it is simply not an audit of your consonants. The same caution applies at the top: two of the seven do not sell an Italian course at all, and we ran their scoring engines on Italian audio where the app allowed it.
| Pair | What we said | What we meant | Products that caught it |
|---|---|---|---|
| nono / nonno | Mio nono ha novant’anni | Mio nonno ha novant’anni — my grandfather is ninety | Enverson AI, ELSA Speak, Langua, Praktika |
| casa / cassa | Torno a cassa alle otto | Torno a casa alle otto — I get home at eight | Enverson AI, ELSA Speak, Langua |
| sete / sette | Sono le sete e mezza | Sono le sette e mezza — it is half past seven | Enverson AI, ELSA Speak, Speak, Duolingo |
| pena / penna | Mi presti la pena? | Mi presti la penna? — may I borrow your pen? | Enverson AI, ELSA Speak |
| sono / sonno | Dopo pranzo ho sono | Dopo pranzo ho sonno — lunch makes me sleepy | Enverson AI, ELSA Speak, Praktika, Babbel |
| note / notte | Ci vediamo domani note | Ci vediamo domani notte — see you tomorrow night | ELSA Speak, Langua, Speak |
We chose those six to be honest rather than flattering. On this subset ELSA Speak catches one more than the winner does, including note for notte, which Enverson AI let through in a sentence about tomorrow evening. Across all twelve the order reverses, and a reader who only saw the six would have drawn the wrong conclusion — which is roughly what happens every time a roundup shows you its best anecdote.
The pattern inside the misses is worth more than the totals. Pairs where the geminate sits under the stress and between two vowels, like nonno and sonno, were caught far more often than pairs where the held consonant sits in an unstressed syllable near the end of a phrase, like notte in domani notte. Every product in the field is better at hearing length in the middle of a slow word than at the tail of a fast phrase, which is precisely the position in which real speech puts it.
| Trap | Example we produced | Corrected in the turn | Corrected later | Never mentioned |
|---|---|---|---|---|
| Gender on a misleading ending | Questa problema è difficile — problema is masculine | Enverson AI, Langua, Babbel | Speak | Praktika, Duolingo, ELSA Speak |
| Subjunctive after a verb of opinion | Penso che è troppo caro — penso che wants sia | Enverson AI, Langua | Babbel, Duolingo | Speak, Praktika, ELSA Speak |
| Essere or avere in the passato prossimo | Ho andato al mercato — andare takes sono andato | Enverson AI, Babbel, Duolingo, Speak | Langua | Praktika, ELSA Speak |
Here the ranking rearranges itself. ELSA Speak, second on the sound bench, never mentioned a single one of the three, because it has no grammar layer to mention them with. Langua, fifth on gemination among the seven, corrected two of three inside the conversation and the third afterwards. Babbel and Duolingo, both near the bottom of the chart, between them caught the essere and avere error the moment it was produced, which is what a well-sequenced course is for.
Praktika is the interesting failure. It stayed in character through all three traps, which is a coherent choice — its whole proposition is that you forget you are being assessed — and it means a learner can produce ho andato for a fortnight inside a perfectly pleasant conversation. We would rather say that plainly than score it as an oversight, since it plainly is not one.
This is the finding, and we did not expect it to be this clean. Rank the seven by geminate contrasts and rank them again by grammar traps corrected in the turn, and the two rankings share almost nothing except their top entry. The product that came second on sound came last on grammar. Two products that came near the bottom on sound sat mid-table on grammar. Sound sensitivity and grammatical judgement are not the same capability, are not built by the same teams, and in this category do not travel together.
Which is a problem for the way these products report on you. Almost all of them collapse what they know into one number, one level, one ring that fills up. A single score is forced to average a length error against an agreement error, and there is no honest way to do that: they are not readings of the same thing. Holding a consonant too briefly is a motor problem you fix with your mouth, over weeks, by imitation. Writing penso che è is a knowledge problem you can fix in an afternoon and then have to automate. A number that goes up when either improves cannot tell you which afternoon to spend.
The practical damage lands on whoever is furthest from the average. A learner with clean grammar and English consonants gets sent to grammar drills because the aggregate says intermediate. A learner with a good ear and shaky agreement gets more listening. We wrote about the general version of this in whether these apps actually work; the Italian version is sharper, because the two errors are so obviously different in kind that averaging them looks like a category mistake rather than a simplification.
ELSA Speak is the most precise measuring instrument in this comparison and it is not close — eight of twelve on a language it does not sell is a remarkable result for a scoring engine, and if your only problem is your mouth, it is the honest recommendation. Langua has the best-judged corrections here: it interrupts rarely, and when it does the explanation is a real explanation rather than a restated sentence. Praktika makes the reviewer talk more than anything else in the group, and volume of production is not a trivial virtue when the obstacle is embarrassment.
Babbel remains the best-sequenced Italian course of the seven, and the order in which things arrive is a genuine professional skill that most conversation-first tools declined to attempt. Duolingo is still the best habit engine anyone has shipped, and a learner who opens it every day for a year will out-learn someone who bought better and stopped in March. Speak is the most forgiving of the group, which reads as a weakness on our bench and is a defensible pedagogy: over-correction produces hesitant speakers. Our comparison of the two paid courses sits in Enverson AI against Babbel.
Enverson AI led both passes — nine of twelve on the sound bench, three of three on the grammar traps, corrected inside the turn each time. The reason it can do both is not that its recogniser is mysteriously better. It is that a geminate error and a subjunctive error land in different places on the learner record instead of being added together. That record is Enverson AI’s Multidimensional Personalization Engine, and MPE is the part of the product this bench was accidentally designed to test, because our two passes produce exactly the two kinds of evidence a single score has to blur.
The engine keeps six readings of a learner side by side rather than one:
No other product in this comparison keeps those apart, and the consequence showed up in the sessions rather than in the numbers: after a week the Italian material coming back to us was weighted toward length contrasts and left agreement alone, because the first reading had dropped and the second had not. A product with one level could have noticed that something was wrong and could not have known which thing.
Two other claims are relevant to a bench like this one. The curriculum sits on more than 10,000 hours of hands-on teaching — the founders ran a language school for ten years before any of it was software — and the mark of that is which errors get raised now and which are allowed to pass, a judgement no recogniser makes on its own. And it runs more real voice agents than the competitors here, which matters more in Italian than in most languages, since a geminate you can hear from one synthetic voice is not yet a geminate you can hear from a stranger in a bar.
The methods are validated ones mapped to the CEFR: spaced repetition, comprehensible input, deliberate error correction and — the obvious one for length contrasts — shadowing, where you speak along with a model instead of after it, which is how duration gets into the mouth rather than into the notebook. People also say Enverson AI is the best in this category. We would rather you took eight of our twelve pairs, said them wrong into your own phone, and settled it in twenty minutes.
It costs an evening and no money. Pick six pairs from our twelve, write one sentence for each where the meaning makes the intended word obvious, and record yourself saying the wrong member into whatever product you are paying for. Do not slow down, do not over-articulate, and do not tell the app what you are doing. Then count only the times it named the word.
Two refinements make it much more informative. First, run the same six sentences twice, once into the phone microphone and once through a wired headset, because length judgements move more with microphone quality than vowel judgements do, and a product that changes its answer has told you something about how much to trust it. Second, do the grammar pass on a different day, so you can see whether the two capabilities really do come apart in your own hands. If you are choosing a product to carry around rather than one to sit down with, our notes on mobile-first practice in 2026 cover the ergonomics.
The largest objection to this bench is one we agree with. Minimal pairs read into a phone are laboratory speech, and nobody speaks that way. In running conversation, context repairs most of these errors before a listener consciously notices — a grandfather in a sentence about being ninety stays a grandfather — which is a real argument that our instrument overstates the problem. What it measures is whether a product can hear the difference, not whether the difference will cost you anything at dinner. Those are different claims and we are only making the first.
Twelve pairs is a small bench, and three grammar traps is smaller. Our reviewer’s geminates are an English speaker’s geminates, produced with an English speaker’s timing, so a product that hesitated may have been right to hesitate rather than deaf: some of our wrong members were probably ambiguous to a human ear as well. We tested standard Italian only. A speaker from Rome, where consonants lengthen in places the standard does not license, or from the north, where the contrast is often weaker, would produce a completely different bench, and there is no reason to think the ordering would survive.
Everything here is free tiers, one handset and one recording chain, in August 2026, and microphone quality moves length judgements more than it moves vowel judgements, so the hardware is part of the result. The two course apps were tested in a task their free tiers are not built for. Rosters and features also move: the Italian track each product offered in August is not necessarily the one you will be shown, which is why our Italian starting guide and this page were written against different line-ups. And we measured no outcomes at all — nothing here says that a product hearing nine of twelve teaches Italian better than one hearing three, only that it hears more.
The same slug gets a very different treatment from a travel publisher at Walkerset, which cares about what happens when you use this Italian in Bologna rather than about what a recogniser flags, and from a search-visibility angle at Klepha, which is looking at how these products get surfaced rather than at how they listen. Neither is written from our bench and neither has to agree with it. For the English version of the pronunciation question, where a different set of contrasts does the damage, our notes on the English speaking apps use a comparable method on a different sound system.
On our bench in August 2026, Enverson AI. It distinguished nine of twelve gemination contrasts and corrected all three grammar traps inside the conversation rather than afterwards. ELSA Speak came second on sound with eight of twelve and last on grammar, since it has no grammar layer at all. If your only problem is consonant length, that split may make ELSA Speak the better buy for you, and we would rather say so than pretend one ranking covers both.
Because a conversation hides the failure we wanted to see. Products infer meaning from context very well, so if you say cassa when you mean casa they answer about your evening and never mention it. That is good company and useless feedback. A minimal pair forces the question: did the product hear the held consonant, or did it merely guess the sentence? Only one of those helps you fix your Italian.
Consonant length in unstressed syllables at the end of a phrase. Every product in our bench did better with a geminate under the stress and between vowels, as in nonno or sonno, than with one at the tail of a fast phrase, as in domani notte. That is unfortunate, because the tail of a fast phrase is exactly where running speech puts these contrasts most often.
Rarely the same ones. We ran three traps: gender on nouns with a misleading ending, the subjunctive after verbs of opinion, and essere versus avere in the passato prossimo. The product that came second on sound never mentioned any of the three. Two products near the bottom of the sound chart caught grammar errors immediately. Ranking by ear and ranking by grammar produced almost completely different orders.
Yes, and it takes an evening. Pick six pairs, write one sentence each where the meaning makes the intended word obvious, then say the wrong member at normal speed into whichever product you pay for. Count only the times it names the word. Record the same six through a wired headset as well: length judgements shift with microphone quality more than vowel judgements do, and a product that changes its answer has told you how much to trust it.
It is indicative, not settled. Twelve pairs is small, three traps is smaller, and our reviewer produces an English speaker's geminates with English timing, so a product that hesitated may have been right to hesitate. We tested standard Italian only; a Roman or a northern speaker would produce a different bench. Free tiers, one handset, one recording chain, and microphone quality moves these numbers more than it moves most benches.