The evidence
Why frequency order gets you speaking fastest
Not all words earn their keep equally. A few hundred do an enormous amount of the work, and the rest tail off fast. That imbalance is the whole argument β so it's worth looking at the actual numbers.
What the research says
The first thousand words do most of the lifting
For Danish specifically, the best published figure comes from a 2018 study that built lemma-based lists of the 2,000 most frequent Danish words:
The first 1,000 lemmas covered 76.96% of a corpus of general written Danish β and 63.04% of written academic Danish.
Jakobsen, Coxhead & Henriksen (2018)
Roughly three words in four, from a thousand entries. And the shape of the curve matters as much as the height: the same study found the second thousand words contributed far less coverage than the first β a pattern that also holds in English and French.
That is the real case for frequency order. Not that vocabulary stops mattering, but that the first words you learn are worth several times what the later ones are, so the order you meet them in changes how quickly you can use the language at all.
Speech versus writing
Spoken language is the easier target
Speech leans on a smaller vocabulary than writing does. Working from the British National Corpus, Nation (2006) put the vocabulary needed for unassisted comprehension at 6,000β7,000 word families for spoken English, against 8,000β9,000 for written text.
For informal speech the numbers are friendlier still. A 2022 study of movies, TV programmes and soap operas β the closest published analogue to the subtitle data underneath this project β found:
| Word families known | Coverage of informal spoken English |
|---|---|
| 1,000 | 90.6β91.4% |
| 2,000 | 94.7β96.7% |
| 3,000 | 96.4β97.6% |
Ha (2022). Figures include proper nouns, marginal words, transparent compounds and acronyms as known.
That study puts 95% coverage β a common comfort threshold β at 2,000β3,000 word families, and 98% at 4,000β5,000.
The caveat
Why we don't put the English numbers on the Danish list
It would be easy, and wrong, to take "1,000 words gets you 91% of speech" and print it next to a Danish word list. Two things break in that move.
The unit is different. The English studies count word families β a headword plus its inflections and its derivations, so national, nationalisere and nationalitet would ride along with nation for free. The Danish study counts lemmas: a headword plus inflections only. A word-family list always reports higher coverage at the same size, because each entry is doing more work.
This list is lemma-based β our pipeline folds inflections into a lemma but never derivations β so the Danish lemma figure is the comparable one, and the flattering English figure is measuring something we don't have.
And the language is different. Danish compounds productively, which affects how much ground a fixed number of entries can cover. English results don't transfer over unexamined.
This field has already been burned once by an over-optimistic number. A 1956 study of Australian oral English reported that 2,000 word families gave about 99% coverage of speech. When Adolphs & Schmitt re-ran the question in 2003 against a modern spoken corpus, the real figure came in under 95%.
The open question
What nobody has published
There is no published coverage figure for spoken Danish. The Danish study measured written corpora; the spoken-language studies are English.
Because speech is consistently less lexically demanding than writing, the spoken-Danish number is very likely higher than the 76.96% measured for written Danish. We are not going to tell you what it is, because nobody knows yet.
It is a gap we're in an unusually good position to close β the largest source under this list is film and television subtitles, which is dialogue. That is a future project, and if it happens the number will appear here with its method attached.
The takeaway
So what does this mean for you
- Early words are worth more than later ones. Front-load them. That is what frequency order does.
- Comprehension improves in steps, not smoothly. The jump from nothing to the first few hundred words is unlike any later stretch of equal size.
- Coverage is not comprehension. Knowing 95% of the words still means one word in twenty is unfamiliar β several per paragraph. High coverage makes a text workable, not transparent.
References
Sources
- Jakobsen, A. S., Coxhead, A., & Henriksen, B. (2018). General and academic high frequency vocabulary in Danish. Nordand β Nordisk tidsskrift for andresprΓ₯ksforskning, 2018(1), 64β89. doi:10.18261/issn.2535-3381-2018-01-04
- Ha, H. T. (2022). Vocabulary Demands of Informal Spoken English Revisited: What Does It Take to Understand Movies, TV Programs, and Soap Operas? Frontiers in Psychology, 13:831684. doi:10.3389/fpsyg.2022.831684
- Nation, I. S. P. (2006). How Large a Vocabulary Is Needed for Reading and Listening? The Canadian Modern Language Review, 63(1), 59β82.
- Adolphs, S., & Schmitt, N. (2003). Lexical Coverage of Spoken Discourse. Applied Linguistics, 24(4), 425β438.
The first words are the ones that pay
Start at rank 1 and work down β the order the argument above is about.
Start with the first 50 words