Method and scope
What the Text Analyser does
It estimates how demanding a reading text or listening transcript is for a learner of English, on the CEFR scale, and shows the vocabulary and sentence measures that a teacher or materials writer can act on. Every measure is shown against reference ranges: the values typical of texts at each level, taken from texts whose CEFR levels were established externally, independently of us.
The level estimate
The level comes from a statistical model trained on texts at known CEFR levels. On general-English texts it matches the level about 83% of the time for reading texts and 75% for listening transcripts, and is within one level 97–99% of the time. It is a guide for choosing and adapting texts, not a certification of a text's level.
C1 and C2 are reported together as C1+: the distinction between them is not measured reliably from a text alone. Texts easier than A2 are reported as A2 or below.
When no overall level is given
An overall level is shown only for texts like those it was validated on. You still get every measure and every reference range when it is withheld, which is when:
- The text is long. The estimate was validated on texts of typical exam length (up to about 750 words for reading and 450 for listening). Longer texts change the measures in ways the model has not been checked against. Analyse a representative section to get a level.
- The text is academic. Academic texts use far more academic vocabulary than general-English texts at the same level, and the estimate tends to overstate them. Where academic vocabulary is well above the general-English range, no level is given.
- The text is very short (under 50 words).
Long texts, section by section
For a text longer than the validated length, the analyser can split it into even sections at paragraph or turn boundaries and estimate each one. You see the level of every section and a count of them, for example "3 of 4 sections are estimated at B2". That count describes the sections. It is not an overall level: a long text asks more of a reader or listener than any single section does, and the analyser does not measure that. On long texts of known level, a section matched the level of its whole text about four times in five for reading and two times in three for listening, and was almost always within one level. All of those test texts were at B2 or above, so the evidence for long texts at lower levels is thinner.
The measures
- Word families for 95% and 98% coverage
- How many of the most frequent English word families a reader must know to understand 95% or 98% of the running words. Around 95% supports reading with help; 98% supports independent reading. Proper nouns, transparent compounds and acronyms count as known. These figures move in whole thousands and are dependable only for texts of about 300 words or more, so they are reported but not compared across levels.
- Words at B2 level or above
- The share of running words whose CEFR vocabulary level is B2, C1 or C2. Of the measures shown, this one follows a text's level most closely.
- Words outside the most frequent 2,000 families
- The share of running words beyond a basic vocabulary.
- Academic Word List
- The share of running words from Coxhead's Academic Word List, counting every member of each word family.
- Mean sentence length, Flesch Reading Ease, Flesch–Kincaid grade
- Standard readability measures.
- Vocabulary by CEFR band
- The share of words at each CEFR vocabulary level.
- Words per minute
- For listening, if you give the recording length.
Your texts
Texts are analysed in memory and are not stored, logged or used for training. For each analysis we record only the word count and the time taken. Like any website, our web server keeps a standard access log (IP address and time of each request).
Credits
- Word-family coverage: Paul Nation, BNC/COCA word family lists, Victoria University of Wellington, CC BY-SA 4.0.
- Academic Word List: Coxhead, A. (2000). A new academic word list. TESOL Quarterly, 34(2), via the Range program lists, CC BY-SA 4.0.
- CEFR vocabulary levels draw on the CEFR-Annotated WordNet (Kikuchi, Ono, Soga, Tanabe & Ozono, 2025), CC BY 4.0, which includes WordNet 3.0, © 2006 Princeton University. WordNet is a registered trademark of Princeton University; no endorsement is implied.
Text Analyser