reflectAI.systemsText Analyser

Method and scope

What the Text Analyser does

It estimates how demanding a reading text or listening transcript is for a learner of English, on the CEFR scale, and shows the vocabulary and sentence measures that a teacher or materials writer can act on. Every measure is shown against reference ranges: the values typical of texts at each level, taken from texts whose CEFR levels were established externally, independently of us.

The level estimate

The level comes from a statistical model trained on texts at known CEFR levels. On general-English texts it matches the level about 83% of the time for reading texts and 75% for listening transcripts, and is within one level 97–99% of the time. It is a guide for choosing and adapting texts, not a certification of a text's level.

C1 and C2 are reported together as C1+: the distinction between them is not measured reliably from a text alone. Texts easier than A2 are reported as A2 or below.

When no overall level is given

An overall level is shown only for texts like those it was validated on. You still get every measure and every reference range when it is withheld, which is when:

Long texts, section by section

For a text longer than the validated length, the analyser can split it into even sections at paragraph or turn boundaries and estimate each one. You see the level of every section and a count of them, for example "3 of 4 sections are estimated at B2". That count describes the sections. It is not an overall level: a long text asks more of a reader or listener than any single section does, and the analyser does not measure that. On long texts of known level, a section matched the level of its whole text about four times in five for reading and two times in three for listening, and was almost always within one level. All of those test texts were at B2 or above, so the evidence for long texts at lower levels is thinner.

The measures

Word families for 95% and 98% coverage
How many of the most frequent English word families a reader must know to understand 95% or 98% of the running words. Around 95% supports reading with help; 98% supports independent reading. Proper nouns, transparent compounds and acronyms count as known. These figures move in whole thousands and are dependable only for texts of about 300 words or more, so they are reported but not compared across levels.
Words at B2 level or above
The share of running words whose CEFR vocabulary level is B2, C1 or C2. Of the measures shown, this one follows a text's level most closely.
Words outside the most frequent 2,000 families
The share of running words beyond a basic vocabulary.
Academic Word List
The share of running words from Coxhead's Academic Word List, counting every member of each word family.
Mean sentence length, Flesch Reading Ease, Flesch–Kincaid grade
Standard readability measures.
Vocabulary by CEFR band
The share of words at each CEFR vocabulary level.
Words per minute
For listening, if you give the recording length.

Your texts

Texts are analysed in memory and are not stored, logged or used for training. For each analysis we record only the word count and the time taken. Like any website, our web server keeps a standard access log (IP address and time of each request).

Credits