Skip to main content

German CEFR text level checker

Paste any German text and Einlang estimates its CEFR level from A1 to C2, marks the words that will slow a reader down, and gives you the hard-word list to hand out. Every word is checked against 25,000 German word families. Free, no sign-up, and nothing you paste leaves your browser.

For reference, everyday spoken German sits 12.5% outside its own 2,000-family core, and that is the yardstick your passage is measured against. If the text turns out to sit above your level, Einlang turns a photo of any page into vocabulary and grammar notes for that page.

0 words
Or try an example:

How many German words does each CEFR level need?

Einlang treats each CEFR level as a vocabulary size: A1 is the first 1,000 German word families, C2 is 16,000. The two coverage columns show how far each of those vocabularies actually gets you - what share of the running words it accounts for in everyday spoken German, and in the 48 literary works in the corpus. The gap between them is why a learner who follows a film can still stall on a novel.

LevelWord familiesEveryday speechLiterary proseWords you meet around here
A11,00082.8%75.3%geschickt, geschrieben, führt, bescheid
A22,00087.5%81.2%mönch, höflich, stirn, glocke
B13,00090%84.5%verpflichtet, versprich, käfig, just
B25,00092.5%88%jüngling, verführt, portier, zwischenfall
C18,00094.6%90.5%insgeheim, rapper, getrennte, kurzsichtig
C216,00097.6%93.4%gorilla, herrschers, geschwärzt, übermannt
  • A1 - Beginner. Built almost entirely from the first thousand words of the language, in short sentences.
  • A2 - Elementary. Everyday vocabulary and past tenses, still close to the common core.
  • B1 - Intermediate. Ordinary written language - news, blogs, straightforward fiction. A dictionary helps.
  • B2 - Upper intermediate. Abstract vocabulary and longer sentences. Readable with regular look-ups.
  • C1 - Advanced. Specialist or literary vocabulary, dense sentences. Slow going without support.
  • C2 - Mastery. Rare vocabulary or long, layered sentences. Ambitious - take it slowly.

How the level is calculated

Einlang reads two things from a passage: how far down the frequency list its words sit, and how long its sentences run. Every word is reduced to its family - so parle, parlait and parlé count once between them - and looked up in an index of 25,000 families. The headline measure is the share of the passage that falls outside the 2,000 commonest families, compared against the same figure for everyday speech in that language. The verdict is the harder of the two readings, vocabulary or sentence length, because a passage written in common words can still be difficult if the sentences never end.

Why the comparison is against that language, not across languages. Everyday spoken German sits 12.5% outside its own core vocabulary. The equivalent figure differs by several points between French, Spanish, and German, entirely because of how each language inflects rather than because one is harder. Normalising against each language's own baseline is what lets one set of thresholds work for all three.

Names are skipped. A capitalised word in the middle of a sentence, with no rank in the index, is treated as a person or a place and left out of the maths - otherwise a novel full of characters reads as a novel full of vocabulary you are missing. German is the hard case here, and worth being straight about: because German capitalises every noun, there is no way to tell Hamburg from Handbuch by capitalisation alone. A German place or surname common enough to have a rank will occasionally appear in the hard-word list. It is obvious when it happens - the word is right there - but it is a real limit, and French and Spanish do not have it.

It checks that the passage is actually in German. A text in the wrong language looks identical to one full of very rare words - every word unrecognised - so without this the checker would answer C2 to an English paragraph and sound certain about it. Einlang instead measures how much of the passage is built from the 200 commonest German words, the articles, prepositions and auxiliaries that every German text uses whatever its subject. Real German of any register stays above 35%; a passage in another language does not reach 20%. When that check fails you get told which language it looks like instead of a level - and you can still force the reading if you know better.

Where the word list comes from. Ranks are blended from two corpora so neither register dominates: a frequency list built from film and television subtitles, which carries modern everyday vocabulary, and 48 public-domain books from Project Gutenberg, which carry the literary and formal written vocabulary subtitles miss. A list built only from books would report ordinateur and week-end as words nobody knows; one built only from subtitles would lose half of written German. The subtitle data is the OpenSubtitles-derived list published by hermitdave/FrequencyWords, used under CC BY-SA 4.0.

Is this an official CEFR assessment?

No, and it is worth being exact about why. This measures vocabulary load and sentence length. The CEFR describes far more than that - grammar, cohesion, register, how familiar the subject matter is - and none of it is read here. There is also no official CEFR word list published in machine-readable form for French, Spanish, and German, so the bands are frequency-rank cutoffs chosen to sit in the range the vocabulary-size research reports for each level, and they are printed above so you can argue with them.

Calibrated against textbook-register passages, the estimate lands within one level, and A1 and A2 overlap too much for any measure of this kind to separate reliably. Two further limits worth knowing: poetry and song lyrics score easier than they read, because line breaks make sentences look short; and a passage under 150 words gives noisy percentages simply because there is not much to count.

Used for what it is good at - comparing two texts, spotting which words will stop a class, checking whether a chapter is a reasonable next step - it is reliable. Used as a certificate, it is not. That distinction is the whole reason the method is on the page.

What to do with a text that is too hard

A text one level above you is the one worth reading - that is where the learning happens, as long as you have support for the gap. If the hard-word list above runs long, pre-teach the ten most frequent entries before starting rather than looking them up mid-page. Our post on comprehensible input explains why that ordering matters, and how many words you need to be fluent covers where the coverage thresholds come from. To pick a whole book rather than a passage, the book difficulty checker ranks public-domain classics the same way.

Checking a whole German book instead?

Pasting a chapter tells you about that chapter. For a whole title, Einlang has already measured 48 public-domain German classics the same way - the book difficulty checker ranks them from easiest to hardest, with the vocabulary you would need for each. The easiest of the 48 is Nathan der Weise by Gotthold Ephraim Lessing.

Other languages

Read the German that is one level too hard

Photograph any page and Einlang turns it into vocabulary, grammar explanations, and drills scoped to that exact page - so the passage above your level becomes the one you finish.

Try Einlang free