How MostUsedWords Builds a Frequency Dictionary

How MostUsedWords Builds a Frequency Dictionary

MostUsedWords frequency dictionaries rank the 10,000 most common words in a language from large, mostly subtitle-based corpora, then split that list into four 2,500-word books. Each entry is built through a combined process: computational frequency ranking, linguistic tagging, AI-assisted drafting, and a final pass by professional native speaker editors. Our earlier, fully human-edited books were independently checked and came in under a 1.87% error rate.

The problem with most frequency lists

Word frequency lists have existed for decades. Most of them are just that: a list. A word, a rank number, maybe a translation. No pronunciation guide, no gender or conjugation pattern, no sentence showing how the word actually gets used. Founder Edmond ran into this in 2015 while learning Spanish himself. He built a better version for his own study first, then turned it into MostUsedWords once he realized every other learner hit the same wall.

Why 10,000 words, split into four books of 2,500

Word frequency in every language follows a steep curve: a small set of words carries most everyday use, and each additional word contributes less than the one before it. By our own figures, the most common 2,500 words cover roughly 95% of daily use in a language, and the top 10,000 gets to about 98% of spoken language and 97% of written language β€” the point where you can usually work out the rest from context. That's exactly why we don't sell one 10,000-word block. Splitting the list into four sequential books of 2,500 words each gives learners a fixed, disclosed unit to work through and a clear sense of progress, instead of an undifferentiated wall of vocabulary:

  • Book 1 (words 1–2,500) β€” the core vocabulary that carries most daily conversation
  • Book 2 (words 2,501–5,000) β€” everyday vocabulary beyond the basics
  • Book 3 (words 5,001–7,500) β€” broader, more specific vocabulary
  • Book 4 (words 7,501–10,000) β€” advanced and lower-frequency vocabulary

Learners can buy one book, all four, or the full bundle, and always know exactly which 2,500 words they're getting and where those words sit in the language's actual frequency ranking.

How each word is verified

Every entry goes through four distinct steps, not one:

  1. Corpus and ranking (computer science). Word frequency is computed from large corpora, mainly subtitle data, per language β€” subtitles capture spoken, everyday language rather than formal written text, which matters for a list meant to teach how a language is actually used.
  2. Linguistic tagging (linguistics). Each word gets its grammatical detail: gender, conjugation pattern, part of speech, and IPA pronunciation.
  3. Drafting (AI). Example sentences and supporting content are AI-drafted from the tagged data, not written from scratch by hand for every one of thousands of entries per language.
  4. Editing (human). Professional native speaker proofreaders and editors review and correct every entry before publication.

This is a computer science + linguistics + AI + human editing pipeline, not an AI-only or human-only one. The AI step handles the volume; the human step catches what AI gets wrong.

Track record on accuracy

Before AI was part of the process, our books were built entirely by hand. Those fully human-edited books were independently checked and came in under a 1.87% error rate β€” the bar the current hybrid process is held to.

How this compares

MostUsedWords Bare frequency list (free/generic) Anki deck (community) General app (Duolingo-style)
Ranked by real usage frequency Yes, subtitle-based corpus, top 10,000 Sometimes Varies by deck No β€” curriculum-ordered, not frequency-ordered
IPA pronunciation included Yes Rarely Rarely No (audio only, no IPA)
Grammar detail (gender, conjugation) Yes No Varies Taught separately, not per-word
Example sentence per word Yes No Varies Yes, but not frequency-ordered
Human-edited Yes, professional native speakers No No, crowd-sourced Editorial team, general content

Frequently asked questions

How many words do I need to know to understand a language?

By our own figures, the most common 2,500 words cover roughly 95% of everyday use in a language. The top 10,000 gets you to about 98% of spoken language and 97% of written language β€” the practical point where you can usually work out remaining unfamiliar words from context. That's why our books are ordered by actual frequency rank rather than by topic or difficulty level.

Is a frequency dictionary better than Anki or a general app?

They solve different problems. Anki and spaced-repetition apps are a study method β€” they don't tell you which words are worth learning first. A frequency dictionary answers that question: it's the source list, ranked by real usage, that you then study with whatever method you prefer, including Anki.

Why split the 10,000-word list into four books instead of one?

Because word frequency drops off sharply after the most common words, lumping all 10,000 into one undifferentiated list obscures how much value each block actually carries. Four fixed 2,500-word units let learners buy and track progress against a known, disclosed unit instead of an open-ended pile of vocabulary.

What corpus is the frequency ranking based on?

Mainly subtitle corpora per language, chosen because subtitles reflect spoken, everyday usage rather than formal written registers.

Who edits the content?

Every entry is drafted using a combination of computational frequency data and AI, then reviewed and corrected by professional native speaker proofreaders and editors before publication.