Writing practice research · October 2026
B2 First and C1 Advanced essay practice-score distributions
The median stored overall estimates were 4 and 4 in the two studied cohorts. These distributions describe AI-evaluated practice responses, with task cohorts and repeat-writer sensitivity reported separately.
Published · Research publisher: Lucas Weaver
- Retained responses
- 364
- Publication
- 5 October 2026
- Evidence
- AI practice estimates
Submission window: 22 August–5 October 2026. Scores are stored AI practice estimates. This self-selected sample does not establish official results, scoring accuracy or causal learning gains.
01
What population do these distributions describe?
An average hides the shape of a score distribution. A sample with many middle estimates and a smaller upper group can have the same mean as a sample concentrated near one value. This report provides score bands, medians and the middle half of the observed estimates so readers can see that shape.
All scores in this report are estimates stored by an AI writing-practice checker. They are not official examination results. No participant supplied a verified examination result for this analysis, and the corpus does not establish a pass rate or the likelihood that an individual will achieve a target score.
We keep task cohorts separate. A difference in the observed distributions may reflect who practised each task, the prompts they selected, the amount of drafting, or differences in the scoring setup. It does not rank the tasks by difficulty.
02
Where is the middle of the practice-score distribution?
The 25th and 75th percentiles bound the middle half of the observed estimates. They describe sample spread, rather than an uncertainty interval for an official examination score. Quartiles use continuous interpolation, so they can fall between values actually stored for an individual response.
The overall estimate is a separately stored output. This study does not replace it with a newly calculated average of the criterion estimates. Readers should consult the criterion study to see which dimensions sit below the overall profile.
| Task cohort | Responses | Accounts | Mean | 25th percentile | Median | 75th percentile |
|---|---|---|---|---|---|---|
| B2 First · Essay | 134 | 129 | 3.95 | 3 | 4 | 4.5 |
| C1 Advanced · Essay | 230 | 212 | 3.79 | 3 | 4 | 4 |
03
How are responses distributed across score bands?
Bands are grouped for readability and privacy. They are descriptive intervals within the native estimate scale, rather than CEFR assignments or exam pass categories. Shares use the full task cohort as the denominator.
Contributing accounts can appear in several score bands after submitting different responses. Their counts therefore overlap. A change between a person’s submissions is not analyzed here and must not be inferred from the table.
| Task cohort | Estimate band | Responses | Share of full cohort | Accounts |
|---|---|---|---|---|
| B2 First · Essay | 3 to below 4 | 33 | 24.6% | 33 |
| B2 First · Essay | 4–5 | 95 | 70.9% | 91 |
| C1 Advanced · Essay | 3 to below 4 | 70 | 30.4% | 68 |
| C1 Advanced · Essay | 4–5 | 148 | 64.3% | 137 |
Detailed bins need at least 20 responses and 10 accounts. Suppressed tails are not zero; displayed shares may total less than 100%.
04
Does the profile change with one response per account?
For B2 First Essay, the mean changes from 3.95 to 3.94 under the first-response rule; the median changes from 4 to 4.
For C1 Advanced Essay, the mean changes from 3.79 to 3.78 under the first-response rule; the median changes from 4 to 4.
The first retained response per account and task produces a contribution-balanced description. It is not a baseline test: accounts may have used other services or practised before the study window.
| Task cohort | Sampling rule | Responses | Median words | Mean estimate | Median estimate | Length–score correlation |
|---|---|---|---|---|---|---|
| Cambridge · B2 First · Essay | all responses | 134 | 187 | 3.95 | 4 | 0.093 |
| Cambridge · B2 First · Essay | first per account and task | 129 | 187 | 3.94 | 4 | 0.065 |
| Cambridge · C1 Advanced · Essay | all responses | 230 | 256 | 3.79 | 4 | 0.135 |
| Cambridge · C1 Advanced · Essay | first per account and task | 212 | 256 | 3.78 | 4 | 0.116 |
05
How should these estimates be used?
For teachers, the distribution can help choose the range of practice material to inspect next. It does not justify assigning a class to one level based on an online-checker percentile. A placement decision requires direct evidence about each learner and the relevant assessment standard.
For learners, a middle estimate describes the center of this site’s practice sample. It is not a target and is not a verdict about your ability. The most useful comparison is between your response and the task’s requirements, supported by criterion-level feedback and, when needed, a qualified teacher’s judgment.
The upper score band should not be presented as a success rate. It combines distinct stored values, includes repeat writers in the response-weighted sample and has no verified external outcome. The separate first-response summary shows how a different contribution rule changes the profile.
These results are an October publication describing a defined recent window. They are not a time-series comparison with the August benchmark. The cohorts use different date windows and cleaning decisions, so subtracting their averages would not measure improvement.
Methods
How this report was prepared
The study uses the latest completed AI evaluation for each eligible submission in the 22 August–5 October 2026 window. Site, task, score and word-count checks precede normalized exact-text deduplication. Accounts contributing more than 50 distinct retained responses in the window are excluded by an exploratory audit rule.
Detailed primary cohorts require 100 responses and 30 accounts. Length and score bins require 20 responses and 10 accounts. The first-response-per-account-and-task analysis is a sensitivity check; it can contain fewer responses than the primary cohort. Account counts across tasks and bins can overlap.
The source does not identify every evaluator model and prompt version, verify examination conditions, or provide independent examiner scores. Lightly edited duplicates and external assistance may remain. Differences describe this practice sample and scoring system; they do not demonstrate official proficiency, causal improvement or universal task difficulty.
Computation and drafting were assisted by AI. This release has no independent human examiner validation, pedagogical review or journal peer review. Only aggregates are published; raw writing, prompts, feedback and learner identifiers are excluded.
Read the full dataset definition, formulas, exclusions and reporting thresholds.
Questions about the findings
Are these official exam score distributions?
No. They are stored AI practice estimates from users who chose an online checker, not representative examination results.
Can the upper score band be reported as a pass rate?
No. There are no verified examination outcomes in this study. These estimate bands are not pass/fail categories.
Evidence
Aggregate data, sources and citation
The downloadable files contain the report’s eligible cohort summaries, criterion profiles and unsuppressed length and score bins. CSV uses one row per measure; JSON preserves the table groupings. Neither file contains raw essays or learner identifiers.
Dataset version and integrity
Version 2026-10-05.1. SHA-256 of this report’s JSON file:
38906e83fa36eee84b51552d1735c23703ca622a879e6ffd4ad1e59fc44bd9edOfficial and contextual sources
Suggested citation
Weaver, Lucas. “B2 First and C1 Advanced essay practice-score distributions.” Cambridge Writing Checker, 5 October 2026. Version 2026-10-05.1. https://cambridgewritingchecker.com/research/cambridge-writing-practice-score-distributions-october-2026
Related writing research
Cambridge Writing Checker is an independent practice service. These reports are not affiliated with or endorsed by the examination organizations.
