IELTS Writing scoring accuracy, benchmarked

IELTS Writing Scoring Accuracy Benchmark: Five Platforms on 97 Official Responses

This is the most comprehensive public comparison of IELTS Writing scoring accuracy available online.

5 min read

The benchmark at a glance

IELTS Writing Practice is the most accurate scorer in this 97-response benchmark. It records the lowest average error (0.53 Band), keeps 74.2% of scores within 0.5 Band and 93.8% within 1 Band, and has the smallest overall bias (-0.27).

It also ranks first on both Task 1 and Task 2, and it keeps the lead on the newer Cambridge IELTS 20 and 21 responses — the subset that matters most when older material may already have appeared in model training data.

IELTSWritingChecker.org has three more exact matches (31 versus 28), but it also has more than twice as many errors above 1 Band (13 versus 6). On the measures that show how close scores stay to the official Band, IELTS Writing Practice leads.

97 officially scored published responses

Overall accuracy results

Mean absolute error, inclusive threshold shares, signed bias, and exact-match counts for all five platforms on the full 97-response set. Error metrics use two decimals in the chart; percentage shares use one decimal and counts remain exact.

Band score error

Mean absolute error

Lower is better
IELTS Writing Practice
0.53
IELTSWritingChecker.org
0.57
ChatGPT
0.88
Engnovate
0.94
Writing9
1.72

Within 0.5 Band

Higher is better
IELTS Writing Practice
74.2% · 72/97
IELTSWritingChecker.org
70.1% · 68/97
ChatGPT
46.4% · 45/97
Engnovate
39.2% · 38/97
Writing9
19.6% · 19/97

Within 1.0 Band

Higher is better
IELTS Writing Practice
93.8% · 91/97
IELTSWritingChecker.org
86.6% · 84/97
ChatGPT
76.3% · 74/97
Engnovate
73.2% · 71/97
Writing9
35.1% · 34/97

Signed bias

Closer to zero is better

Signed bias is the average signed error. Negative means the platform scored below the official Band on average.

IELTS Writing Practice
-0.27
IELTSWritingChecker.org
-0.38
ChatGPT
-0.85
Engnovate
-0.86
Writing9
-1.58

Reading the overall results

The four charts above compare each platform's scores with the published official Band. Here's how to read them.

  • MAE is the average distance between a platform score and the official Band. Lower is better.
  • Within 0.5 Band and Within 1 Band show how often a score stays no more than that far from the official Band. Higher is better.
  • Signed bias shows whether a platform usually scores higher or lower than the official Band. Closer to zero is better, and a negative value means scoring lower on average.

For example, IELTSWritingChecker.org has 31 exact matches, versus 28 for IELTS Writing Practice. Exact matches still matter, but IELTS Writing Practice has 6 errors above 1 Band while IELTSWritingChecker.org has 13, and fewer errors above 1 Band means fewer large misses in everyday use.

Mean absolute error (Band)

Task 1 versus Task 2

Mean absolute error by task type across all five platforms, with 51 Task 1 responses and 46 Task 2 responses. Error metrics use two decimals in the chart; lower values are better.

Lower is better

Task 1

n = 51
IELTS Writing Practice
0.42
IELTSWritingChecker.org
0.45
ChatGPT
0.77
Engnovate
0.75
Writing9
1.51

Task 2

n = 46
IELTS Writing Practice
0.65
IELTSWritingChecker.org
0.70
ChatGPT
0.99
Engnovate
1.14
Writing9
1.96

What the task split shows

Splitting the 97 responses by task type gives 51 Task 1 responses and 46 Task 2 responses. Mean absolute error by task:

  • Task 1: IELTS Writing Practice 0.42; IELTSWritingChecker.org 0.45; ChatGPT 0.77; Engnovate 0.75; Writing9 1.51.
  • Task 2: IELTS Writing Practice 0.65; IELTSWritingChecker.org 0.70; ChatGPT 0.99; Engnovate 1.14; Writing9 1.96.

IELTS Writing Practice records the lowest MAE on both tasks. Task 2 MAE is higher than Task 1 MAE for all five observed workflows in this sample. The benchmark does not establish why; it records the pattern without a measured cause, and we decline to speculate.

For Writing9, one preprocessing note applies. Its observed public workflow did not accept the Academic Task 1 image, and it is evaluated here as the product actually worked. When Writing9 displayed <4, the declared benchmark convention recorded it as 3.5, the upper bound of the displayed below-4 range, before errors were calculated. No result is excluded: all five platforms have complete results on both task subsets.

32 responses from Cambridge IELTS 20 and 21

Recent Cambridge stress test

Results for the 32-response Cambridge IELTS 20 and 21 subset. Error metrics use two decimals; percentage shares use one decimal and counts remain exact.

Band score error

Mean absolute error

Lower is better
IELTS Writing Practice
0.55
IELTSWritingChecker.org
0.69
ChatGPT
0.92
Engnovate
1.02
Writing9
2.05

Within 0.5 Band

Higher is better
IELTS Writing Practice
71.9% · 23/32
IELTSWritingChecker.org
62.5% · 20/32
ChatGPT
43.8% · 14/32
Engnovate
40.6% · 13/32
Writing9
6.3% · 2/32

Within 1.0 Band

Higher is better
IELTS Writing Practice
93.8% · 30/32
IELTSWritingChecker.org
81.3% · 26/32
ChatGPT
71.9% · 23/32
Engnovate
62.5% · 20/32
Writing9
15.6% · 5/32

Signed bias

Closer to zero is better

Signed bias is the average signed error. Negative means the platform scored below the official Band on average.

IELTS Writing Practice
-0.39
IELTSWritingChecker.org
-0.56
ChatGPT
-0.92
Engnovate
-0.89
Writing9
-2.02

What the recent-books test can tell us

Older IELTS responses may already have appeared in model training data, which can make results on familiar material look stronger. Cambridge IELTS 20 and 21 are the newest books in this dataset. Testing their 32 responses separately reduces that risk, so this is the subset worth paying the most attention to.

IELTS Writing Practice still leads all four measures: 0.55 MAE, 71.9% within 0.5 Band, 93.8% within 1 Band, and -0.39 signed bias. The lead therefore holds on the newer material, not only across the full dataset.

This test cannot prove that no model has seen these responses, because providers do not publish complete training data. It is a stronger check, not a guarantee.

97 responses

The 97-response dataset

Every source family in the 97-response benchmark with its count, plus Task 1/Task 2 and Academic/General Training totals. Counts are exact.

Task 1: 51 responses

Task 2: 46 responses

Academic: 54 responses

General Training: 43 responses

  1. IELTS.org Academic sample responses (2023)12
  2. IELTS.org General Training sample responses (2023)11
  3. IELTS.org computer-delivered Academic sample responses4
  4. Official IELTS Practice Materials Volume 110
  5. Official IELTS Practice Materials Volume 210
  6. Cambridge IELTS 18 Academic2
  7. Cambridge IELTS 19 Academic8
  8. Cambridge IELTS 19 General Training8
  9. Cambridge IELTS 20 Academic8
  10. Cambridge IELTS 20 General Training8
  11. Cambridge IELTS 21 Academic8
  12. Cambridge IELTS 21 General Training8

Methodology, limitations, and sources

Ground truth. Each of the 97 responses carries a published official overall Band. The benchmark uses only that published official overall Band as the reference score. No claim is made about criterion-level ground truth; the study does not assess criterion-level scores in either direction.

Measures. Plain-text formulas: signed error = platform Band minus published official Band. MAE = the mean of the absolute signed errors. Signed bias = the mean signed error. Within-0.5 and within-1.0 are inclusive threshold counts, and exact match means zero absolute error. Chart labels for error metrics use two decimals; percentage shares use one decimal and counts remain exact. Full-precision benchmark values, for example a mean absolute error of 0.5309 for IELTS Writing Practice on the full set, underlie every figure shown.

Dataset. The 97 responses comprise: IELTS.org Academic sample responses (2023): 12; IELTS.org General Training sample responses (2023): 11; IELTS.org computer-delivered Academic sample responses: 4; Official IELTS Practice Materials Volume 1: 10; Official IELTS Practice Materials Volume 2: 10; Cambridge IELTS 18 Academic: 2; Cambridge IELTS 19 Academic: 8; Cambridge IELTS 19 General Training: 8; Cambridge IELTS 20 Academic: 8; Cambridge IELTS 20 General Training: 8; Cambridge IELTS 21 Academic: 8; Cambridge IELTS 21 General Training: 8. Total: 97. Task 1: 51. Task 2: 46. Academic: 54. General Training: 43. Cambridge IELTS 18 contributes Academic responses only.

Limitations. Writing9's observed public workflow did not accept the Academic Task 1 image; when it displayed <4, the declared benchmark convention recorded it as 3.5, the upper bound of the displayed below-4 range. All five platforms have complete results on both the full n=97 set and the recent n=32 subset; no result is missing from either set. Automated scores are estimates, not official IELTS results. The raw per-response platform score table, feedback and review records, annotations, and internal model payloads are retained privately and are not published; the official source responses themselves are published materials from their publishers, not secret.

Related published comparisons. For related published comparisons, see the IELTSWritingChecker.org accuracy study and its analysis of ChatGPT for IELTS Writing.

Official IELTS resources. For official scoring guidance and writing preparation materials, see IELTS.org's resources for setting IELTS scores and IELTS.org's writing test resources.

Get an estimated Band and diagnostic feedback on your own writing

Check your own writing

An invitation to try the existing IELTS Writing Practice scoring simulator, which returns an estimated Band and diagnostic feedback. It does not provide an official IELTS result.

IELTS writing simulatorHigh-fidelity
Task 2 question loaded.
Part 2

You should spend about 40 minutes on this task. Write at least 250 words.

Words: 040:00 starts on first word