How many Spanish words do you know?
Five minutes and you will have a number. Einlang shows you Spanish words drawn from every level of frequency, you mark the ones you know, and the result comes back as a vocabulary size, a CEFR level, and a list of books you could read at that size. Free, no sign-up, and nothing you answer leaves your browser.
Drawn from 280 vetted Spanish words and 40 invented ones. The invented words are the part that makes the number mean something: they measure how often a half-familiar word feels known, and that rate is subtracted before you see a figure.
How the Spanish test works
- You get three or four short screens of Spanish words, drawn from every level of frequency from the commonest thousand to the 16,000th.
- Tap the ones you know - meaning you could give the sense of the word if you met it in a sentence. Leave the rest.
- Some of the words are invented. They are not there to catch you out; they are how the estimate is corrected, so mark only what you really know and the number comes back honest.
It looks like this:
About five minutes. No sign-up, no account, and your answers stay in this browser tab.
How many Spanish words is enough?
A vocabulary size only means something next to what it covers. Reading with occasional look-ups needs about 95% of the running words known; reading straight through needs about 98%. This is where each vocabulary size lands against everyday spoken Spanish and against the 45 literary works in Einlang's corpus.
| Word families | Roughly | Everyday speech | Literary prose |
|---|---|---|---|
| 1,000 | A1 | 77.2% | 69.5% |
| 2,000 | A2 | 83.5% | 75.8% |
| 3,000 | B1 | 86.7% | 79.5% |
| 5,000 | B2 | 90.1% | 83.9% |
| 8,000 | C1 | 92.8% | 87.4% |
| 16,000 | C2 | 96.3% | 91.2% |
The steepness at the top is the point: the first 1,000 families already carry 77.2% of spoken Spanish, and the next 7,000 add only 15.6 points. That is why progress feels fast for a year and slow afterwards, and why how many words you need to be fluent has no single answer.
How the estimate is calculated
The test is a sample, not an inventory. Einlang holds a frequency-ranked index of Spanish word families and cuts it into 7 bands, from the commonest thousand down to the 16,000th. You are shown a handful of words from each band; the share you know in a band is taken as the share you know of the whole band, and the bands are added up. Knowing four of five words drawn from the 1,000-2,000 band means roughly 800 of those thousand families, and so on down.
This is why the invented words are there. A yes/no test with nothing to check it against measures confidence, not vocabulary - and people are not being dishonest when they overclaim, a half-familiar word genuinely feels known. So about a third of what you are shown does not exist. Those words are built from the letter patterns of real Spanish - which letters follow which, how words in this language begin and end - and anything that comes out a real word, or within one letter of one, is discarded. The rate at which you mark them known is subtracted from every band before the estimate is formed. Mark none, and nothing changes; mark a third of them, and the estimate drops by roughly a third of a band each time.
It stops when it has learned enough. If you knew at most one of the ten words drawn from the 3,000-8,000 range, the answer is already found and the hardest screen is skipped rather than asked. The bar is deliberately that high: every band above a stop is scored as nothing known without being asked about, so a stop that fires early is far more expensive than a screen that runs long. Whatever the test saves goes on the band where your answers are actually changing, which is the only place extra precision buys anything.
Word families, not words. tanta and its inflections count once between them, the same unit the reading research uses and the same one the text level checker counts in. A figure quoted in individual forms would be two to three times larger and mean less.
How accurate is it?
Against simulated readers of known vocabulary size, the average estimate lands within about 5% of the truth - so the method is not biased high or low. A single run is much looser than that: nine runs in ten fall within about a third either side of the true figure, which is what sampling fifty-odd words out of 16,000 costs you. That is why the result is given as a range as well as a number, and why the range is worth more attention than the number.
Two things it cannot do. It cannot tell recognition from use - you may know reflejo when you read it and never reach for it when you speak, and this counts that as known, as every test of this design does. And it cannot see whether you know the right 16,000 words for what you actually read: a vocabulary built on crime novels and one built on economics can total the same and share very little.
The CEFR level is a translation, not a grade. Einlang maps sizes to levels with the same frequency cutoffs the text level checker publishes - A1 at 1,000 families through C2 at 16,000. The CEFR describes far more than vocabulary, and no exam board would grade you on a word list. Treat the level as shorthand for the size, not as a result.
Where the words come from
The same index the text level checker runs on: ranks blended from a frequency list built from film and television subtitles, which carries modern everyday vocabulary, and 45 public-domain books from Project Gutenberg, which carry the written vocabulary subtitles miss. The subtitle data is the OpenSubtitles-derived list published by hermitdave/FrequencyWords, used under CC BY-SA 4.0.
Which of those ranks become questions is decided separately, and strictly. A frequency list is full of names, brands, and untranslated song lyrics, and the checker can afford to rank them - a test cannot ask about them. So a candidate has to be far more frequent in Spanish than in English, French, Spanish or German, which removes names and loans in one move; it has to appear across at least six books in the corpus, which removes the character names that dominate a single novel; it has to still be current in the modern list, which removes nineteenth-century spelling; and in French and Spanish it must not be capitalised mid-sentence. What survives is 280 words per language, and each test draws a different sample from them.
Nothing you answer leaves your browser. Einlang publishes this site as static files - there is no server here to send anything to. The word list is downloaded to your machine and every count is done there. Nothing is stored, so there is no score to look up later; if you want to keep it, copy it.
What to do with the number
A vocabulary size is only useful when it changes what you pick up next. The book difficulty checker will tell you how many unknown words a page any measured classic would give you at your size, and the Spanish text level checker does the same for anything you paste in. For how vocabulary actually accumulates, comprehensible input and spaced repetition are the two ideas that matter most.
Other languages
Add the next thousand Spanish words
Photograph any Spanish page you want to read and Einlang turns it into vocabulary, grammar notes, and drills for that exact page - so the words you learn are the ones standing between you and the book in your hands.
Try Einlang free