Writing practice research · October 2026
Cambridge essay criterion profiles: Language, ties and score gaps
Language had the lowest mean stored estimate in both studied cohorts. Shared-lowest and uniquely-lowest counts show how often that pattern applied to individual practice responses.
Published · Research publisher: Lucas Weaver
- Retained responses
- 364
- Publication
- 5 October 2026
- Evidence
- AI practice estimates
Submission window: 22 August–5 October 2026. Scores are stored AI practice estimates. This self-selected sample does not establish official results, scoring accuracy or causal learning gains.
01
What does a lowest criterion actually mean?
Cambridge Writing assessment separates Content, Communicative Achievement, Organisation and Language. The four estimates answer different questions about the same response. An overall estimate can hide a lower value on one dimension, so this study examines the full profile for B2 First and C1 Advanced essays separately.
Language covers vocabulary and grammatical resources, including control and appropriacy. It should not be reduced to a grammar-error count. This study uses stored criterion estimates; it has not counted or independently verified errors in the underlying essays.
The research question is whether a criterion repeatedly sits below the others inside this checker’s evaluations. A systematic pattern could reflect the writing, the scoring instructions, or both. Human examiner comparisons would be required before interpreting the pattern as a measure of official criterion difficulty.
Task and rubric reference: Cambridge C1 Advanced writing assessment guide.
02
How often was language control the lowest estimate?
In B2 First Essay, Language shared the lowest estimate in 124 of 134 responses (92.5%). It was the sole lowest criterion in 61 (45.5%).
In C1 Advanced Essay, Language shared the lowest estimate in 210 of 230 responses (91.3%). It was the sole lowest criterion in 92 (40.0%).
| Task cohort | Criterion | Mean | Median | Shared lowest: count / share | Unique lowest: count / share | Mean gap to highest |
|---|---|---|---|---|---|---|
| B2 First · Essay | Communicative Achievement | 3.71 | 4 | 40 / 29.9% | 0 / 0.0% | 0.22 |
| B2 First · Essay | Content | 3.78 | 4 | 35 / 26.1% | 1 / 0.7% | 0.14 |
| B2 First · Essay | Language | 3.09 | 3 | 124 / 92.5% | 61 / 45.5% | 0.84 |
| B2 First · Essay | Organisation | 3.6 | 4 | 55 / 41.0% | 2 / 1.5% | 0.33 |
| C1 Advanced · Essay | Communicative Achievement | 3.57 | 4 | 117 / 50.9% | 2 / 0.9% | 0.22 |
| C1 Advanced · Essay | Content | 3.58 | 4 | 118 / 51.3% | 2 / 0.9% | 0.2 |
| C1 Advanced · Essay | Language | 3.17 | 3 | 210 / 91.3% | 92 / 40.0% | 0.61 |
| C1 Advanced · Essay | Organisation | 3.63 | 4 | 105 / 45.7% | 2 / 0.9% | 0.16 |
Native practice-estimate scales: IELTS 0–9; Cambridge 0–5. Gap means compare each criterion with the highest criterion in the same response.
03
How much of the pattern comes from tied estimates?
All four criterion estimates were equal in 18 of 134 B2 First Essay responses (13.4%). Those responses contribute to every shared-lowest count.
All four criterion estimates were equal in 79 of 230 C1 Advanced Essay responses (34.3%). Those responses contribute to every shared-lowest count.
Shared minima overlap. The unique-lowest shares are mutually exclusive, but they do not sum to 100% because responses with tied minima do not have a unique lowest criterion. This prevents a tied profile from being reported as four separate problems.
04
What happens when repeat responses are removed?
Language remains the lowest mean criterion under the first-response rule in both cohorts. The full criterion table is provided so that the ordering can be checked directly.
A systematic ordering can also be an evaluator behavior. No blinded human-rating comparison was performed, and the historical evaluation records do not identify every model and prompt version.
| Task cohort | Criterion | Responses | Mean estimate | Shared lowest | Unique lowest |
|---|---|---|---|---|---|
| B2 First · Essay | Communicative Achievement | 129 | 3.69 | 29.5% | 0.0% |
| B2 First · Essay | Content | 129 | 3.78 | 25.6% | 0.8% |
| B2 First · Essay | Language | 129 | 3.07 | 93.0% | 45.0% |
| B2 First · Essay | Organisation | 129 | 3.57 | 41.9% | 1.6% |
| C1 Advanced · Essay | Communicative Achievement | 212 | 3.54 | 49.5% | 0.5% |
| C1 Advanced · Essay | Content | 212 | 3.56 | 50.0% | 0.9% |
| C1 Advanced · Essay | Language | 212 | 3.12 | 91.5% | 41.0% |
| C1 Advanced · Essay | Organisation | 212 | 3.59 | 44.8% | 0.9% |
05
What should change in teaching or revision?
The distinction between a shared and a unique minimum changes the practical interpretation. A criterion that frequently shares the lowest value may be part of a broad profile rather than the single obstacle holding a response back. We report both counts so readers can see that distinction instead of treating every shared minimum as a separate diagnosis.
Use a lower criterion estimate as a prompt for inspection. Review the actual feedback and the response before choosing an exercise. The tables cannot establish which grammatical structures, vocabulary choices, task omissions or organizational problems are present in an individual answer.
Means summarize the checker’s estimate pattern in this practice cohort. The native scales are ordered assessment estimates, and differences between decimal means should not be interpreted as precise amounts of language ability. We retain medians and response counts alongside the means for that reason.
The first-response-per-account check asks whether repeat contributors change the ordering. It does not remove external assistance, prompt differences or potential scorer bias. Agreement between the two summaries supports the stability of this observed pattern under one sampling change, rather than validating the scoring system.
Methods
How this report was prepared
The study uses the latest completed AI evaluation for each eligible submission in the 22 August–5 October 2026 window. Site, task, score and word-count checks precede normalized exact-text deduplication. Accounts contributing more than 50 distinct retained responses in the window are excluded by an exploratory audit rule.
Detailed primary cohorts require 100 responses and 30 accounts. Length and score bins require 20 responses and 10 accounts. The first-response-per-account-and-task analysis is a sensitivity check; it can contain fewer responses than the primary cohort. Account counts across tasks and bins can overlap.
The source does not identify every evaluator model and prompt version, verify examination conditions, or provide independent examiner scores. Lightly edited duplicates and external assistance may remain. Differences describe this practice sample and scoring system; they do not demonstrate official proficiency, causal improvement or universal task difficulty.
Computation and drafting were assisted by AI. This release has no independent human examiner validation, pedagogical review or journal peer review. Only aggregates are published; raw writing, prompts, feedback and learner identifiers are excluded.
Read the full dataset definition, formulas, exclusions and reporting thresholds.
Questions about the findings
Are shared-lowest percentages supposed to total 100%?
No. Every criterion tied at the minimum is counted. The same response can contribute to several shared-lowest counts.
Does this prove the lowest criterion is the hardest exam skill?
No. The observed ordering belongs to this checker’s practice estimates. Writing differences and scorer behavior can both produce the pattern.
Evidence
Aggregate data, sources and citation
The downloadable files contain the report’s eligible cohort summaries, criterion profiles and unsuppressed length and score bins. CSV uses one row per measure; JSON preserves the table groupings. Neither file contains raw essays or learner identifiers.
Dataset version and integrity
Version 2026-10-05.1. SHA-256 of this report’s JSON file:
38906e83fa36eee84b51552d1735c23703ca622a879e6ffd4ad1e59fc44bd9edOfficial and contextual sources
Suggested citation
Weaver, Lucas. “Cambridge essay criterion profiles: Language, ties and score gaps.” Cambridge Writing Checker, 5 October 2026. Version 2026-10-05.1. https://cambridgewritingchecker.com/research/cambridge-writing-criterion-bottlenecks-october-2026
Related writing research
Cambridge Writing Checker is an independent practice service. These reports are not affiliated with or endorsed by the examination organizations.
