Skip to main content

French CEFR text level checker

Paste any French text and Einlang estimates its CEFR level from A1 to C2, marks the words that will slow a reader down, and gives you the hard-word list to hand out. Every word is checked against 25,000 French word families. Free, no sign-up, and nothing you paste leaves your browser.

For reference, everyday spoken French sits 12.3% outside its own 2,000-family core, and that is the yardstick your passage is measured against. If the text turns out to sit above your level, Einlang turns a photo of any page into vocabulary and grammar notes for that page.

0 words
Or try an example:

How many French words does each CEFR level need?

Einlang treats each CEFR level as a vocabulary size: A1 is the first 1,000 French word families, C2 is 16,000. The two coverage columns show how far each of those vocabularies actually gets you - what share of the running words it accounts for in everyday spoken French, and in the 76 literary works in the corpus. The gap between them is why a learner who follows a film can still stall on a novel.

LevelWord familiesEveryday speechLiterary proseWords you meet around here
A11,00082.2%74.2%exemple, époque, dents, lire
A22,00087.7%81.5%poing, ordonné, penche, embarrassant
B13,00090.5%85.7%roger, hantée, grec, solitaire
B25,00093.3%90.1%décida, fabuleux, exaltant, venez-vous
C18,00095.4%93.2%instance, partielle, prévisions, avalanche
C216,00098.3%95.9%señorita, virement, débiteur, tarr
  • A1 - Beginner. Built almost entirely from the first thousand words of the language, in short sentences.
  • A2 - Elementary. Everyday vocabulary and past tenses, still close to the common core.
  • B1 - Intermediate. Ordinary written language - news, blogs, straightforward fiction. A dictionary helps.
  • B2 - Upper intermediate. Abstract vocabulary and longer sentences. Readable with regular look-ups.
  • C1 - Advanced. Specialist or literary vocabulary, dense sentences. Slow going without support.
  • C2 - Mastery. Rare vocabulary or long, layered sentences. Ambitious - take it slowly.

How the level is calculated

Einlang reads two things from a passage: how far down the frequency list its words sit, and how long its sentences run. Every word is reduced to its family - so parle, parlait and parlé count once between them - and looked up in an index of 25,000 families. The headline measure is the share of the passage that falls outside the 2,000 commonest families, compared against the same figure for everyday speech in that language. The verdict is the harder of the two readings, vocabulary or sentence length, because a passage written in common words can still be difficult if the sentences never end.

Why the comparison is against that language, not across languages. Everyday spoken French sits 12.3% outside its own core vocabulary. The equivalent figure differs by several points between French, Spanish, and German, entirely because of how each language inflects rather than because one is harder. Normalising against each language's own baseline is what lets one set of thresholds work for all three.

Names are skipped. French capitalises nothing mid-sentence except proper nouns - not nationalities, months, or weekdays - so a capital letter in the middle of a sentence is the surest sign of a name there is. Those words are left out of the maths, and a name spotted once is skipped everywhere in the passage, including where it opens a sentence. Without that rule a novel full of characters reads as a novel full of vocabulary you are missing, and any frequency list built from subtitles is full of place names besides.

It checks that the passage is actually in French. A text in the wrong language looks identical to one full of very rare words - every word unrecognised - so without this the checker would answer C2 to an English paragraph and sound certain about it. Einlang instead measures how much of the passage is built from the 200 commonest French words, the articles, prepositions and auxiliaries that every French text uses whatever its subject. Real French of any register stays above 35%; a passage in another language does not reach 20%. When that check fails you get told which language it looks like instead of a level - and you can still force the reading if you know better.

Where the word list comes from. Ranks are blended from two corpora so neither register dominates: a frequency list built from film and television subtitles, which carries modern everyday vocabulary, and 76 public-domain books from Project Gutenberg, which carry the literary and formal written vocabulary subtitles miss. A list built only from books would report ordinateur and week-end as words nobody knows; one built only from subtitles would lose half of written French. The subtitle data is the OpenSubtitles-derived list published by hermitdave/FrequencyWords, used under CC BY-SA 4.0.

Is this an official CEFR assessment?

No, and it is worth being exact about why. This measures vocabulary load and sentence length. The CEFR describes far more than that - grammar, cohesion, register, how familiar the subject matter is - and none of it is read here. There is also no official CEFR word list published in machine-readable form for French, Spanish, and German, so the bands are frequency-rank cutoffs chosen to sit in the range the vocabulary-size research reports for each level, and they are printed above so you can argue with them.

Calibrated against textbook-register passages, the estimate lands within one level, and A1 and A2 overlap too much for any measure of this kind to separate reliably. Two further limits worth knowing: poetry and song lyrics score easier than they read, because line breaks make sentences look short; and a passage under 150 words gives noisy percentages simply because there is not much to count.

Used for what it is good at - comparing two texts, spotting which words will stop a class, checking whether a chapter is a reasonable next step - it is reliable. Used as a certificate, it is not. That distinction is the whole reason the method is on the page.

What to do with a text that is too hard

A text one level above you is the one worth reading - that is where the learning happens, as long as you have support for the gap. If the hard-word list above runs long, pre-teach the ten most frequent entries before starting rather than looking them up mid-page. Our post on comprehensible input explains why that ordering matters, and how many words you need to be fluent covers where the coverage thresholds come from. To pick a whole book rather than a passage, the book difficulty checker ranks public-domain classics the same way.

Checking a whole French book instead?

Pasting a chapter tells you about that chapter. For a whole title, Einlang has already measured 76 public-domain French classics the same way - the book difficulty checker ranks them from easiest to hardest, with the vocabulary you would need for each. The easiest of the 76 is La Dame aux camélias by Alexandre Dumas fils.

Other languages

Read the French that is one level too hard

Photograph any page and Einlang turns it into vocabulary, grammar explanations, and drills scoped to that exact page - so the passage above your level becomes the one you finish.

Try Einlang free