Writing practice research · October 2026
B2 First essay word count and practice scores: October 2026 study
Among 134 B2 First Essay responses from 129 accounts, the median length was 187 words. The length–score rank correlation was 0.093; this association does not establish a benefit from adding words.
Published · Research publisher: Lucas Weaver
- Retained responses
- 134
- Publication
- 5 October 2026
- Evidence
- AI practice estimates
Submission window: 22 August–5 October 2026. Scores are stored AI practice estimates. This self-selected sample does not establish official results, scoring accuracy or causal learning gains.
01
What question does this study answer?
B2 First essays are written to a defined task with notes to address. A response can grow because it develops the required points, but also because it repeats an opinion or adds material outside the question. Counting words cannot distinguish those cases.
The official essay instructions specify 140–190 words. This report compares observed practice lengths with that range and examines their association with the checker’s stored overall estimates on its 0–5 scale. It does not convert those estimates into Cambridge English Scale scores or official examination results.
The sample includes only B2 First essays. Articles, reviews and letters have different purposes and reader relationships; combining them would make the length comparison less interpretable. The data identifies the chosen exam task, not independent confirmation of each writer’s CEFR proficiency.
Task and rubric reference: Cambridge English B2 First handbook for teachers.
02
How long were the evaluated responses?
The middle half of responses contained 172–223.75 words. The mean overall estimate was 3.95 and the median was 4. These are separate descriptive summaries of this cohort, rather than recommended writing targets.
60 responses (44.8%) exceeded 190 words; 74 (55.2%) were inside the 140–190 range.
The lower tail is limited by the product’s submission checks. The recent evaluated cohort contains few or no responses below the official task minimum. It cannot establish the score consequences of writing substantially under length.
03
Do longer responses receive higher practice estimates?
Across the displayed length groups, the means do not form a consistently increasing sequence.
The table shows separate groups of observed responses. It does not follow one answer as words are added or removed. Accounts can appear in more than one group, so the contributing-account column should not be summed.
| Words | Responses | Contributing accounts | Mean estimate | Median estimate |
|---|---|---|---|---|
| 140–189 | 69 | 68 | 3.88 | 4 |
| 190–229 | 34 | 33 | 4.06 | 4 |
| 230–3000 | 31 | 30 | 4 | 4 |
Groups require at least 20 responses and 10 accounts. Word bands are descriptive; some cross an official upper limit. Exact compliance counts are reported separately.
04
Does the pattern survive one response per writer?
With the first retained response per account and task, the sample contains 129 responses. Its median length is 187 words and its rank correlation is 0.065, compared with 0.093 in the full retained cohort.
Spearman’s correlation compares ranks, using average ranks for ties. A value near zero indicates little monotonic association in this sample. It does not rule out task-specific relationships, and it is not a test of whether an editing intervention works.
| Task cohort | Sampling rule | Responses | Median words | Mean estimate | Median estimate | Length–score correlation |
|---|---|---|---|---|---|---|
| Cambridge · B2 First · Essay | all responses | 134 | 187 | 3.95 | 4 | 0.093 |
| Cambridge · B2 First · Essay | first per account and task | 129 | 187 | 3.94 | 4 | 0.065 |
05
How can learners and teachers use this evidence?
For learners, a response above the official range is an invitation to inspect relevance and economy. Identify the sentence that addresses each required point, keep the reasoning that supports it, and remove material that adds no new explanation. This revision process is a practical suggestion, not a measured treatment effect.
For teachers, the share above the range can support a discussion about planning. The task asks the writer to do a specific amount of communicative work within a constraint. Our estimates do not establish an automatic penalty per extra word, and the absence of such a penalty in an AI estimate would not override official task instructions.
The smaller number of retained responses makes individual length-band summaries more sensitive to task mix. We display contributing-account counts and a first-response-per-account check so readers can judge that sensitivity alongside the main result.
Methods
How this report was prepared
The study uses the latest completed AI evaluation for each eligible submission in the 22 August–5 October 2026 window. Site, task, score and word-count checks precede normalized exact-text deduplication. Accounts contributing more than 50 distinct retained responses in the window are excluded by an exploratory audit rule.
Detailed primary cohorts require 100 responses and 30 accounts. Length and score bins require 20 responses and 10 accounts. The first-response-per-account-and-task analysis is a sensitivity check; it can contain fewer responses than the primary cohort. Account counts across tasks and bins can overlap.
The source does not identify every evaluator model and prompt version, verify examination conditions, or provide independent examiner scores. Lightly edited duplicates and external assistance may remain. Differences describe this practice sample and scoring system; they do not demonstrate official proficiency, causal improvement or universal task difficulty.
Computation and drafting were assisted by AI. This release has no independent human examiner validation, pedagogical review or journal peer review. Only aggregates are published; raw writing, prompts, feedback and learner identifiers are excluded.
Read the full dataset definition, formulas, exclusions and reporting thresholds.
Questions about the findings
Is the median word count an ideal target?
No. It describes this practice sample. Follow the task requirements and use the response’s coverage, development and clarity to judge what to include.
Does a higher mean in a longer group prove that adding words helps?
No. The groups differ in writers, prompts and practice conditions. This study does not compare matched edited versions of the same answer.
Evidence
Aggregate data, sources and citation
The downloadable files contain the report’s eligible cohort summaries, criterion profiles and unsuppressed length and score bins. CSV uses one row per measure; JSON preserves the table groupings. Neither file contains raw essays or learner identifiers.
Dataset version and integrity
Version 2026-10-05.1. SHA-256 of this report’s JSON file:
d0df28f19985a52a6fc4cef1add2c2dcfcedefe9bb54f153e1796e063bbaf26fOfficial and contextual sources
Suggested citation
Weaver, Lucas. “B2 First essay word count and practice scores: October 2026 study.” Cambridge Writing Checker, 5 October 2026. Version 2026-10-05.1. https://cambridgewritingchecker.com/research/b2-first-essay-word-count-and-practice-scores-2026
Related writing research
Cambridge Writing Checker is an independent practice service. These reports are not affiliated with or endorsed by the examination organizations.
