The apps look interchangeable. The thinking behind them isn't
Open five language apps and they look like variations on a theme: colourful tiles, a progress bar, a voice reading a sentence. They are not variations on a theme. Each one is the descendant of a specific school of language teaching, and those schools have spent the last century and a half approaching the same question from genuinely different starting points. They don't all agree with each other — but each of them noticed something real, which is exactly why all seven are still in use today.
This matters to you for a practical reason: an app is shaped by what its tradition believes learning is. A tradition that treats language as a system of rules will teach you rules and test you on them, and you may end up able to write a neat sentence and still freeze in a conversation. A tradition that treats language as something absorbed from meaningful input will give you hours of listening and never once explain a rule, and you may end up understanding a great deal while making mistakes you cannot hear yourself making. Neither app is broken. They are optimised for different things.
Which brings us to the finding this article builds toward, and the reason it is worth knowing the map: no single school is complete, and the people who actually get fluent almost always end up combining several of them — usually without ever learning their names. The right mix is personal. It depends on what you want the language for, how you like to be taught, how much time you have and in what shape, and frankly on what you'll enjoy enough to keep doing. What follows is the map, so you can make that combination on purpose rather than by accident.
The map
Almost no app is pure. The mainstream course apps — Duolingo, Babbel, Mondly — sit squarely in the grammar–translation column with a communicative dialogue bolted on top, which is roughly what school textbooks have done for forty years. The apps that look least like school, on the other hand, are the committed ones: Anki does retrieval practice and literally nothing else, Lingopie is pure comprehensible input, italki is pure communicative practice with a human being.
1. Grammar–translation: learn the rule, translate the sentence
The oldest school in the room, and the one most people were taught in. It comes from nineteenth-century classrooms where the model subject was Latin: you learned the declensions, you translated passages, you were graded on accuracy. Speaking was not really the point — reading great literature and mental discipline were.
On your phone it looks like this: a sentence in one language, an instruction to render it in the other, and an explicit note explaining the rule you just used. In our hands-on testing, Duolingo's core exercise was exactly this — a Vietnamese sentence to be turned into English from a bank of word tiles — and Babbel's lessons pause to tell you outright that Swedish 'hej' works in both formal and informal settings. That explanation is the grammar–translation tradition, alive and well.
Its strength is that it makes the invisible visible. Adults especially benefit from being told the rule rather than being left to infer it; research on explicit instruction has been fairly kind to it. Its weakness is the one every schoolchild discovers: translation ability is not speaking ability. You are practising the act of converting between languages, which is a real skill, but it is not the skill of thinking in one.
There's a specific trap in the app version. Because word tiles only fit together a few ways, you can often assemble the right answer without understanding the sentence — recognition dressed up as production. If your app offers a setting to type the answer instead of tapping tiles, that single change moves you from recognising to producing, and it is probably the highest-value setting in this entire category.
2. Audiolingual drilling: repeat the pattern until it is automatic
Born in the Second World War, when the American army needed people who could function in Japanese and German quickly, and formalised in the 1950s as the audiolingual method. Its theory came from behaviourist psychology: language is a set of habits, and habits are built by drilling correct responses until they are automatic. Mimicry, memorisation, and endless pattern practice — no explanations, no translation, no negotiating.
Chomsky demolished the underlying theory in 1959, and no serious linguist today believes language is just a stack of conditioned responses. But the method's engineering survives because the drills genuinely work for what they cover. Pimsleur is its most direct descendant: audio lessons built around prompting you for a phrase just before you would have forgotten it. Rosetta Stone's repetition-heavy exercises come from the same lineage.
What it's good for: pronunciation, fixed phrases, and the automaticity that lets you say a common sentence without assembling it consciously. What it can't do: prepare you for anything unscripted, because a drill has one right answer and a conversation does not. Its other virtue is entirely practical — audio drills work with your eyes and hands busy, which is why this tradition owns the commute.
3. Immersion / the direct method: no translation, ever
A late-nineteenth-century revolt against grammar–translation, commercialised into the Berlitz schools and still going. The rule is absolute: the target language only. No translation, no first language, meaning conveyed by pictures, objects, gesture and context. Grammar is meant to be absorbed inductively — you see enough examples that the pattern emerges without anyone stating it.
Rosetta Stone is the pure app expression of this: you match photographs to sentences, and the app famously refuses to tell you what anything means. Advocates argue this forces you to build direct associations between concepts and target-language words, rather than routing everything through a mental translation step you'll later have to unlearn.
The honest assessment is that it does build those direct associations, and that it is slow and sometimes maddening for abstract vocabulary. It is easy to show 'apple' with a photograph. It is very hard to show 'nevertheless', 'owe', or 'lease' — and the direct method's answer is essentially that you'll get there eventually with enough context. Learners who like puzzles thrive on this; learners in a hurry mostly don't.
4. Comprehensible input: understand things slightly beyond you
The most influential idea in modern language teaching, and the most argued about. Stephen Krashen's claim, from the early 1980s, is that we acquire language in essentially one way — by understanding messages pitched a little above our current level, what he called i+1. Conscious study, in this view, produces something different and lesser: knowledge you can monitor your speech with, but not the fluent system underneath. The practical prescription is enormous amounts of interesting, understandable listening and reading, with a relaxed attitude and no forced speaking early on.
Lingopie is this school in app form: you watch real Spanish or Korean television with dual subtitles, tapping any word you don't know. LingQ and Beelinguapp work the same seam with text. When we tested Lingopie the mechanic was exactly what the theory prescribes — comprehension supported just enough to keep the input understandable, with the content itself doing the motivating.
Almost everyone in the field now agrees that abundant comprehensible input is necessary. Whether it is sufficient — Krashen's strong claim — is much less accepted, and the standard objection is that input alone tends to produce learners who understand well and speak inaccurately, because nothing ever forces them to produce and be corrected. Treat this school as the engine of your listening and vocabulary, not as the whole diet.
5. Communicative and task-based teaching: use it to do something
The dominant approach in classrooms since the 1980s. Its founding insight was that knowing a language means being able to do things with it — order the meal, apologise convincingly, argue about football — and that this ability doesn't fall out of grammar knowledge automatically. So the classroom should be organised around real communication, with fluency valued over perfect accuracy and errors treated as a normal part of learning rather than a failure of discipline. Task-based teaching is its most structured strand: complete a genuine task, and let the language you need emerge from it.
In apps this shows up as scenario dialogues and, increasingly, AI conversation partners. Babbel's dialogue trainer drops you into a workplace conversation; Busuu builds its lessons around functional goals; Speak's whole product is talking to a machine that talks back. italki isn't an app in this sense at all — it's a marketplace for exactly this school, delivered by a human being.
This is the tradition with the strongest claim on what most people actually want, which is to speak. Its limitation in app form is that a conversation with a scripted or synthetic partner is a rehearsal, not a performance: the machine is endlessly patient, never confused, and never asks the thing you didn't prepare for. It is excellent practice and it flatters you slightly.
6. Peer and tutor correction: someone who knows more fixes your output
This one comes from a different intellectual family — Vygotsky's sociocultural psychology, and the interactionist tradition in second-language research. The claim is that learning happens in the gap between what you can do alone and what you can do with help, and that being pushed to produce language and then having it corrected is where real progress is made. Producing forces you to notice what you don't know, in a way that comprehension never does.
The purest app version is Busuu's community: our test account declared it was learning English and spoke Spanish, and within a minute a real learner had sent an exercise for us to correct, while our own writing went to native speakers for the same treatment. italki delivers the same mechanism professionally, with a tutor who notices your errors and works on them.
The evidence for corrective feedback is good, especially when it's specific and immediate. The catch is entirely logistical: this school requires other people, which means scheduling, money, or the social nerve to post something imperfect where strangers will see it. It is the least convenient school and, for speaking and writing accuracy, probably the most valuable.
7. Retrieval practice and spaced repetition: the memory school
This isn't really a language-teaching philosophy at all — it's cognitive psychology that language learners adopted, and it comes with the best experimental evidence in this entire article. Two findings drive it. The spacing effect: material reviewed at widening intervals is retained far better than the same material studied in one block, a result that goes back to Hermann Ebbinghaus in 1885 and has survived every attempt to knock it down. And the testing effect: the act of retrieving something from memory strengthens it more than re-reading it does.
Anki is the school in its pure form. Every card you grade Again, Hard, Good or Easy is you telling an algorithm when to show it next, and our testing found intervals stretching from under a minute to five days on a single first session. Memrise applies the same scheduling to a friendlier course, and Lingopie quietly does it too — the words you tap while watching become a spaced deck.
What it does superbly is make vocabulary stay. What it cannot do is teach you what to do with the words, and this is where enthusiasts go wrong: a learner with four thousand perfectly retained flashcards and no conversational practice is a common and slightly tragic figure. Retrieval practice is the memory layer beneath a method, not a method by itself.
A note on gamification: it is not a teaching method
Streaks, points, leagues, gems and treasure chests belong to none of these schools. They are a motivation layer wrapped around whichever pedagogy the app already had — usually grammar–translation — and they are worth naming separately because they are easy to mistake for the product.
That distinction has a practical use when you're choosing. The reward machinery tells you how hard the app will work to bring you back tomorrow; the school underneath tells you what you'll actually be doing when you arrive. An app can be superb at the first and mediocre at the second, and the app-store rating will not separate them for you.
How to identify your app's school in about thirty seconds
Open it and do one exercise, then ask: was I told a rule, or left to infer it? Was my own language on screen at any point? Did I have to produce something, or only recognise the right option? Was there a human being anywhere in the process? Did anything get scheduled for later review?
Rule stated plus your own language present means grammar–translation. No translation anywhere and meaning carried by images means the direct method. Repeat-after-me with a fixed correct answer means audiolingual drilling. Long stretches of listening or reading you actually enjoy means comprehensible input. A task to accomplish with a partner, human or synthetic, means communicative teaching. Someone correcting what you produced means the sociocultural school. A card coming back tomorrow because you found it hard means retrieval practice.
What the evidence supports, and the practical upshot
The research picture, compressed honestly: nobody seriously disputes that large amounts of comprehensible input are necessary. The evidence for retrieval practice and spacing is the strongest and least controversial in the field. Corrective feedback on what you produce is well supported. Explicit grammar instruction helps adults more than the input purists concede. And pure drilling, while a poor theory of language, remains genuinely useful for pronunciation and set phrases.
What follows from that is unfashionable but simple: the schools are not rivals to choose between, they are components to stack. Each one is strong exactly where another is weak. Input builds comprehension and vocabulary but tolerates sloppy production. Correction fixes production but requires other people. Retrieval makes things stick but doesn't teach use. Communication rehearses use but can drift without accuracy checks. Grammar explains the machinery but doesn't build fluency.
In practice the stack most learners converge on, whatever apps they use, looks like this: something with lots of input for volume, something with retrieval for retention, and something with a human for correction — plus, if you're an adult who likes knowing why, an explanation of the rules to hang it all on. The app names change every few years. The seven schools don't.
Where these ideas come from
Surveys of the traditions themselves: Richards, J. C., & Rodgers, T. S. (2014). Approaches and Methods in Language Teaching, 3rd ed. Cambridge University Press — the standard reference work for this whole taxonomy. Lightbown, P. M., & Spada, N. How Languages are Learned. Oxford University Press — the most readable overview of what the research actually shows.
On the individual schools: Chomsky, N. (1959). Review of B. F. Skinner's Verbal Behavior. Language, 35(1), 26–58 — the critique that ended audiolingualism's claim to be a theory of language. Krashen, S. D. (1982). Principles and Practice in Second Language Acquisition, and Krashen, S. D., & Terrell, T. D. (1983). The Natural Approach — the input hypothesis in the authors' own words. Hymes, D. (1972). On Communicative Competence — the paper that gave communicative teaching its founding idea. Ellis, R. (2003). Task-based Language Learning and Teaching. Oxford University Press. Vygotsky, L. S. (1978). Mind in Society. Harvard University Press — the zone of proximal development. Long, M. H. (1996). The Role of the Linguistic Environment in Second Language Acquisition — the interaction hypothesis. Swain, M. (1985). Communicative Competence: Some Roles of Comprehensible Input and Comprehensible Output in its Development — the case for pushed output.
On the evidence claims made above: Norris, J. M., & Ortega, L. (2000). Effectiveness of L2 Instruction: A Research Synthesis and Quantitative Meta-analysis. Language Learning, 50(3), 417–528 — synthesised 49 studies and found explicit instruction produced larger and more durable gains than implicit instruction. Li, S. (2010). The Effectiveness of Corrective Feedback in SLA: A Meta-analysis. Language Learning, 60(2), 309–365 — 33 studies, a medium overall effect (d = 0.64) that held up over time.
On memory: Ebbinghaus, H. (1885). Über das Gedächtnis, translated as Memory: A Contribution to Experimental Psychology — the original forgetting curve. Cepeda, N. J., Pashler, H., Vul, E., Wixted, J. T., & Rohrer, D. (2006). Distributed Practice in Verbal Recall Tasks: A Review and Quantitative Synthesis. Psychological Bulletin, 132(3), 354–380 — the spacing effect, synthesised. Roediger, H. L., & Karpicke, J. D. (2006). Test-Enhanced Learning: Taking Memory Tests Improves Long-Term Retention. Psychological Science, 17(3), 249–255 — the testing effect. Woźniak, P. A. (1990). Optimization of Learning — the SM-2 algorithm that every modern flashcard scheduler, Anki included, descends from.
Everything said here about what the apps do on screen comes from our own hands-on testing on a real Android phone, documented in the individual reviews.
Frequently asked questions
What teaching method does Duolingo use?
Primarily grammar–translation — you convert sentences between your language and the target one, with rules made explicit — plus elements of drilling and spaced review. Its streaks, gems and leagues are a motivation layer, not a teaching method.
Which language-teaching method works best?
No single one is complete. The evidence is strongest for spaced retrieval practice (for retention), abundant comprehensible input (for comprehension and vocabulary), and corrective feedback on what you produce (for accuracy). They are complements, not competitors — the schools are weak exactly where the others are strong.
Why does Rosetta Stone never translate anything?
Because it belongs to the direct method tradition, which holds that routing a new language through your first one builds a translation habit you later have to unlearn. Meaning is conveyed by pictures and context instead. It builds direct associations well, and it struggles with abstract vocabulary.
Is comprehensible input enough on its own?
Almost everyone agrees it is necessary; far fewer accept Stephen Krashen's stronger claim that it is sufficient. Input-only learners often understand well and speak inaccurately, because nothing pushes them to produce language and have it corrected.
Are flashcard apps like Anki a teaching method?
Not by themselves. Spaced retrieval is a memory technique with excellent experimental support, but it teaches you to recall items, not to use them. It works best as the memory layer underneath a method that supplies context and practice.