Basics

The exact formula behind the ranking — what your languages, your goal and your proficiency level each do to an app's score

A transparent walkthrough of the BestLanguageApp scoring model: the hard gates that remove apps outright, the weighted blend that ranks the rest, and the arithmetic that decides which app lands at #1.

Published 2026-08-05 · Updated 2026-08-05

How it works

From your selections to a ranked list
YOUR INPUTHARD GATESSOFT WEIGHTSOUTPUTremovescoreLanguagesknown + learningGoaltravel, work, exam…Proficiencybeginner → fluentLanguage taught?else excludedLive humanpractice?if requestedUsable at yourlevel?score > 3Editorialbase0–100Language fit0–10Goal fitmean 0–10Level fit0–10Ranked listscore 0–100, sorted

Selections first remove apps that cannot do the job, then reweight the survivors.

Two stages: gates, then weights

Every ranking on this site is produced the same way, from the same numbers, in two stages. First a set of hard gates removes apps that cannot do what you asked for at all. Then the apps that survive are scored by a weighted blend and sorted.

The distinction matters. A gate is binary and unforgiving: an app that does not teach Icelandic cannot be the best Icelandic app, no matter how good it is at everything else, so it is removed and shown separately under 'excluded' with the reason. A weight is a matter of degree: an app that is merely good at your goal is not disqualified, it just scores lower than one that is excellent at it.

Nothing in the model is personalised beyond the selections you make on the page — there is no history, no profile and no paid placement. Two people making identical selections see an identical list.

Stage one: the hard gates

Three conditions can remove an app before scoring begins.

The first is the language you are learning. Each app carries an explicit list of languages it does not teach. If your target language is on that list, the app is excluded. Note that this is separate from teaching quality: a language that is taught badly stays in the ranking with a low language-fit score, while a language that is not taught at all is a gate.

The second applies only if you ask for practice with real people. Apps are flagged for whether they offer live human practice — tutors, live classes or community correction — and if you request it, apps without it are removed rather than merely penalised, because no weighting can conjure a human on the other end.

The third is your proficiency level. Each app has a fit score from 0 to 10 for beginner, intermediate and fluent learners. If the app scores 3 or less at the level you selected, it is excluded as unsuitable rather than shown near the bottom. This is the gate that most changes the list: several apps that dominate the beginner ranking disappear entirely for fluent learners, and vice versa.

Stage two: the scoring formula

Surviving apps are scored on a 0–100 scale by blending four components: our editorial base score, how well the app teaches your specific target language, how well it serves the goal you selected, and how well it suits your proficiency level.

The base score is the app's overall quality independent of anything you selected — teaching design, content depth, pricing transparency and how it held up in hands-on testing. It is already on a 0–100 scale. The other three components are stored as 0–10 fit scores and multiplied by ten to bring them onto the same scale before blending.

The result is a weighted mean, not a sum. That choice is deliberate and has a consequence worth understanding: selecting more goals does not inflate the scores of generalist apps. It sharpens the ranking instead, because an app now has to be good at every goal you picked rather than merely good at one of them.

S_a = the app's final score, rounded to a whole number. B_a = editorial base (0–100). L_a(ℓ) = language fit (0–10). G_a = goal fit (0–10). P_a(p) = proficiency fit (0–10). δ_g and δ_p are 1 when you have selected a goal or a level and 0 otherwise.
The four weights. The denominator re-normalises whenever a dimension is unselected, so scores stay on the same 0–100 scale however much of the page you fill in.

What your languages do

The language you already speak and the language you are learning enter the model in different ways.

The language you are learning does the heavy lifting. It drives the exclusion gate, and it sets the language-fit term, which carries a weight of 0.24 — as much as your goal. Language fit is stored per app as a default with per-language overrides, so an app with a strong Spanish course and a thin Korean one is scored differently depending on which you picked, rather than being averaged into a single misleading number.

The language you already speak matters mainly for interface and explanation quality: some courses are built from scratch for a specific language pair, while others translate one course into many source languages. Where that difference is material to teaching quality it is already reflected in the per-language override for that app.

Language fit is sparse by design: only languages where an app is notably better or worse than its own average carry an override.

What your goal does

Every app is scored 0–10 against each of the five learner goals on the page: travel, conversation, exam preparation, work and family. Selecting one or more goals switches on the goal term in the formula, at a weight of 0.24.

The goal score is the arithmetic mean over the goals you selected, not the sum. Selecting travel alone asks 'which app is best for travel'. Selecting travel and exam together asks 'which app is best at both' — and an app that is superb at one and hopeless at the other lands in the middle, which is the personalized answer.

Goals never remove an app. An app that is poorly suited to your goal still appears, lower down, tagged with a short reason such as 'Not built for your goals'.

C is the set of goals you selected. When C is empty the term drops out of the formula entirely and its weight is redistributed across the remaining dimensions.

What your proficiency level does

Proficiency is the only dimension that acts as both a gate and a weight, and it is the one that reshapes the list most.

As a gate, a fit score of 3 or below at your selected level removes the app. This reflects a real property of the market: a gamified beginner course is not a weak choice for a C1 learner, it is the wrong tool, and burying it at position eleven would imply it was merely a bit worse rather than unsuitable.

As a weight, the surviving apps' level fit contributes at 0.18 — the smallest of the four weights, because the gate has already done the coarse filtering and the remaining differences are matters of degree. An app scoring 9 for your level earns the 'Built for your level' tag; one scoring 5 or below is flagged as aimed at a different level.

Levels map roughly onto the CEFR scale: beginner is A0–A2, intermediate B1–B2, and fluent C1–C2. If you are unsure which to pick, choose the lower one — the gate is stricter at the top of the scale, so an over-optimistic selection removes more useful apps than a conservative one does.

A worked example

Suppose an app has an editorial base of 88, a language fit of 9 for the language you selected, a goal fit averaging 7.5 across the two goals you picked, and a level fit of 8 for intermediate learners. It teaches your language, so it passes the first gate, and 8 is comfortably above 3, so it passes the third.

All four dimensions are active, so the denominator is the full 1.00 and the numerator is 0.34 × 88 + 0.24 × 90 + 0.24 × 75 + 0.18 × 80, giving 83.9 — displayed as 84.

Now change one thing: clear your goal selection. The goal term disappears, the denominator falls to 0.76, and the same app scores 85. This is why scores shift slightly as you fill in the page. They are not being adjusted after the fact; a different question is being asked, and the mean is taken over a different set of dimensions.

Ties are broken by the editorial base score, so when two apps blend to the same number the one we rated higher overall on its own merits is listed first.

All four dimensions active.
The same app with no goal selected: the weight is re-normalised rather than treated as a zero.

What the model deliberately does not do

It does not use download counts, star ratings or revenue as inputs. Those measure marketing reach, not teaching quality, and they are the reason most 'best app' lists read like a popularity chart.

It does not accept payment for position. There is no field in the data model that a company could buy.

It does not hide the losers. Excluded apps are listed with the gate that removed them, so you can see whether the exclusion was about your language, your level, or your request for live practice — and change your selection if you disagree with the call.

And it does not pretend to more precision than it has. The underlying 0–10 fit scores are editorial judgements from hands-on testing, not measurements. The formula makes those judgements consistent and auditable; it does not make them objective. Two whole points of difference in a final score is meaningful. One is noise.

Frequently asked questions

Why does an app's score change when I select more options?

Because the score is a weighted mean, and selecting a dimension adds it to both the numerator and the denominator. An unselected dimension is not scored as zero — it is left out and its weight redistributed, so the scale stays 0–100 throughout.

Why is an app missing from the list entirely?

It was removed by one of the three hard gates: it does not teach your target language, it has no live human practice when you asked for that, or it scores 3 or below at your proficiency level. Excluded apps are shown separately with the specific reason.

Does selecting more goals help generalist apps?

No. The goal term is the mean across the goals you select, not the sum, so an app has to be good at all of them. Selecting several goals typically narrows the field rather than flattening it.

Why does proficiency remove apps instead of just lowering their score?

Because level mismatch is a category error rather than a quality difference. A course designed for absolute beginners is not a slightly worse choice for an advanced learner; it is the wrong product, and ranking it eleventh would understate that.

Can a company pay to rank higher?

No. Position is a pure function of the editorial base score, the per-language fit scores, the goal fit scores and the level fit scores, all set by us. There is no commercial input to the formula.

Apps with this feature

← All guides