How Langly Rates Material Difficulty: The Unknown-Word Encounter Model
A technical overview of Langly's material difficulty rating: simulating a learner's unknown-word encounter rate, then adjusting videos for how the language is actually delivered — measured speaking speed and pauses.
Yuya Uzu
Every video on Langly carries one of five difficulty levels — Super beginner, Beginner, Low intermediate, High intermediate, and Advanced (roughly A1 through C1+ in CEFR terms). This document describes how those levels are assigned. The rating is derived entirely from statistics over the material itself — there is no manual tagging and no LLM judgment involved — and it was designed around one goal: the level should match how difficult the material actually feels to a learner, not how it scores on a readability formula.
The key ideas
- We rate difficulty by unknown-word encounters, not coverage. A hard word costs once; repetitions are free. Material that keeps reusing its vocabulary rates easier — the way it actually feels.
- The level is the learner who understands ~95% of the material — comfortable to follow, with something left to learn.
- Word difficulty comes from conversational usage, not written text. Frequency ranks are counted over Langly's own library of spoken content, because news-corpus rankings bury everyday spoken words.
- Delivery counts, not just vocabulary. Measured speaking speed adjusts the level: native-speed speech rates harder, learner-slow delivery rates easier.
- Everything is calibrated against felt difficulty — thresholds are set by actually watching materials, not by formula aesthetics.
Why coverage percentages weren't enough
The conventional way to rate content difficulty is vocabulary coverage: find the vocabulary size needed to know, say, 90% of the material's words, and map that size to a level. We built this first, and it has one blind spot: it can't take repetition into account — and repetition is a big part of how difficult a material actually feels. A learner podcast that keeps reusing the same handful of topic words scores badly on coverage, but a real learner absorbs those words within the first minute. A word you've already met three times in this video is not the same obstacle as a brand-new one.
The encounter-rate simulation
The core of the rating is a simple simulation. Picture a learner who knows the most frequent words of the language, consuming the opening of the material. Every word outside their vocabulary costs 1 — but only the first time. Once they've met it, meeting it again is free. So a material that keeps reusing its hard words stays cheap, while a constant stream of new unknowns keeps costing — which is exactly how it feels.
We assume a learner at each level already knows a certain base vocabulary: the 500 most frequent words for a super beginner, 1,000 for a beginner, and so on up to 12,000 for an advanced learner. Then we run the simulation for each of them and ask: for whom does about 95% of this material read as known? That's the sweet spot — comfortable enough to follow, with enough new words to learn from — and that learner's level becomes the material's level:
| Base vocabulary (words already known) | Learner level |
|---|---|
| 500 | Super beginner |
| 1,000 | Beginner |
| 2,000 / 3,000 | Low intermediate |
| 5,000 / 8,000 | High intermediate |
| 12,000 | Advanced |
What counts as a word
The frequency ranks originally came from wordfreq, a general-purpose frequency dataset — but general corpora are dominated by written web and news text, and that register skew misranks exactly the words that matter here: everyday spoken vocabulary gets buried as "rare" because newspapers don't use it. So we switched to Langly's own corpus — word frequencies counted over the materials in our library, which is spoken, conversational content: the same register the learner is consuming. Words are counted by lemma, so conjugated forms count under their dictionary form and an inflection never masquerades as a rare word.
The delivery adjustment
Vocabulary statistics see what a material says, but not how it's delivered — and delivery is a real part of felt difficulty. A native-speed debate built from easy words is not a beginner experience, and a slow, deliberately paced vlog can be far friendlier than its vocabulary suggests. So for videos, a second signal adjusts the vocabulary verdict.
We measure speaking speed from the subtitle timings — both how fast the speaker talks in their fluent stretches and the average pace across the whole video — along with how much pause time the delivery leaves between lines. Native-speed speech nudges the rating harder, even when the words are easy; slow, deliberate delivery nudges it easier. And speech that is dramatically slower than anything native — which in practice only exists in learner-directed content — can ease the rating further still.
The speed thresholds are set per language (Mandarin syllables and English words don't tick at the same rate), each calibrated against materials whose felt difficulty we verified by watching them.
Where you see it
The resulting level appears as the badge on every material card, and the material detail page shows the receipt behind the judgment — the vocabulary size the material was judged at, the encounter rate there, the measured delivery verdict, and whether it adjusted the result — so the rating is always inspectable rather than a black box.
About the author
Related posts
Langly as a Language Reactor Alternative: An Honest Comparison (2026)
Language Reactor and Langly both turn native video into learning material, but they solve different problems. A concise, accurate feature comparison — including where Language Reactor is still the better pick.
10+ Best YouTube Channels for Learning Chinese with Comprehensible Input
A curated guide to the best YouTube channels for learning Chinese with comprehensible input — grouped by level, from absolute beginner to intermediate.