Verify quoted words, reasoning and the scope of scoring claims.
Revise and transfer
Make your own revision and practise the same skill on a new task.
Separate an official result from practice feedback
An AI practice band, a teacher's lesson feedback and an official IELTS result serve different purposes. An official result comes from the test process. Practice feedback can help you decide what to revise, but its usefulness does not establish how closely its number will match a future test. English AIdol does not publish an independently audited accuracy percentage for its estimates. Read the product limitations.
For IELTS Writing, examiners assess task fulfilment, organisation, vocabulary and grammar; Task 1 and Task 2 are assessed separately. Speaking includes pronunciation and fluency as well as language use. The official scoring explanation describes these criteria. A number produced from one paragraph or a typed transcript is not a complete assessment of both Writing tasks or of spoken performance.
A helpful comparison therefore starts with the decision you need to make. For a sentence revision, inspect whether the comment is correct and useful. For test readiness, consider a complete timed attempt and qualified feedback on uncertainties. For acceptance by an institution, use its current requirements and the official result it accepts. Do not substitute a practice dashboard for that evidence.
What would support an accuracy claim?
“Aligned with the band descriptors” describes an intended framework; it is not an accuracy measurement. Before relying on a percentage, ask what was measured, against which reference, and on whose responses. A product name or the word “calibrated” does not answer those questions. Compare claims only when their methods and populations are sufficiently alike.
Questions to ask before comparing numbers
Evidence
What to check
Task and sample
Which test version, task types, proficiency range, languages and recording conditions were included? How were responses selected?
Reference ratings
Who assessed each complete response, what qualifications and instructions did they have, and how were disagreements handled?
Evaluation separation
Were evaluation responses kept separate from development and tuning? Was overlap with publicly available samples investigated?
Model and settings
Which model/version, prompts and input settings were tested? When was the evaluation run?
Errors and uncertainty
Are exact agreement, error sizes, direction of errors and uncertainty reported, including disappointing cases?
Ownership and limits
Who funded and conducted the work, what can be reproduced, and which learners/tasks remain unevaluated?
Missing evidence should stay unknown. It is not a reason to invent a zero, assume a competitor fails, or treat two marketing percentages as directly comparable. Our open comparison method and blank worksheet help record those distinctions; they do not supply an independent product validation.
What the cited research can and cannot tell you
Koraishi’s 2024 study assessed the ChatGPT-4 version available in November 2023 on 55 publicly sourced Task 2 samples. It reported weighted kappa of 0.811 while also identifying substantial disagreement on individual responses. That coefficient is not an 81.1% accuracy rate or a guaranteed error range. The result concerns the paper's setup; it does not validate English AIdol, a different model, Speaking scores or every learner. This article is a source review, not a new replication study.
Likewise, repeated agreement is not enough to prove correctness. Two tools might share a mistaken interpretation; the same tool might repeat a mistaken answer. A changed score after rewriting does not isolate learning from task difficulty, prompt changes or the assessor. Keep the original response and the conditions alongside any number you record.
Worked exercise: check three feedback comments
The following task, paragraph and comments are original teaching material, not real user data, a recalled test question or an experiment. No band is assigned. Try deciding which comments to accept before opening the explanation.
Practice question: Should a company offer employees flexible starting times? Give a reason and explain a possible limitation.
Learner paragraph: “Flexible starting times can help employees who share a car with family members. Each employee choose a starting time within an agreed window. For example, a parent could arrive after taking a child to school, then finish later. However, the customer support desk would still need enough people at its advertised opening time.”
Comment A: Change “Each employee choose” to “Each employee chooses.”
Comment B: Replace “can help” with “will increase productivity by 40%” to make the argument stronger.
Comment C: Remove the customer-support sentence because every paragraph must support the proposal without qualification.
Check the comments and the revision
A is a justified correction: “Each employee” is singular, so the present-tense verb is “chooses.” B is unsupported: the paragraph supplies neither a productivity study nor that number, and it changes a possibility into a certainty. C misreads the task: the question specifically asks for a limitation; the support-desk sentence answers it.
One possible revision: “Flexible starting times can help employees coordinate shared transport and school journeys. Each employee chooses a start within an agreed window and finishes correspondingly later. The company would still need a rota that keeps its customer support desk staffed during advertised hours.”
The revision fixes agreement and makes the operational limit explicit. It also changes the examples and compresses the paragraph; a learner should decide whether those choices preserve their intended meaning. A polished rewrite is not automatically better evidence of the learner's own ability.
Transfer task: Write three sentences supporting remote meetings and one sentence explaining a limitation. Check every number, quotation and causal claim. If a tool suggests adding evidence, use a clearly hypothetical illustration or a source you have actually verified; do not invent a study to make the paragraph sound authoritative.
Speaking feedback needs the right input
A transcript lets you inspect wording, but it does not preserve all the information in speech. It cannot by itself show how a word sounded, where stress fell or how a hesitation affected the listener. Before accepting a pronunciation comment, check that the tool used your recording and points to an audible example. The official Speaking format and assessment explanation is the reference for the actual test.
Original practice check: record a short answer about a place you visit, then listen once before reading the transcript. Mark one passage that was hard to follow. Compare it with the tool's comment, replay the passage and record a fresh version. If the transcript says a word you did not say, flag that mismatch before accepting a vocabulary or grammar judgement based on it. A recording problem and a language problem need different responses.
Do not treat a particular accent as a defect or use a count of pauses as a complete Speaking score. Ask what the listener could understand and which concrete change improves that passage. A teacher can help investigate uncertain feedback, but the title “teacher” alone also does not guarantee IELTS-specific assessment expertise; ask about relevant experience and the scope of the service.
A practical comparison without pretending it is research
Choose one complete task and save your unchanged attempt, conditions and time used.
Give each feedback source the same material. Record the tool/version or reviewer, date, plan and any missing information.
Check a small set of comments against your actual words or audio. Label each supported, unsupported or unresolved, and explain the label.
Revise one issue yourself. Ask whether the change answers the task more clearly without changing your meaning.
Try a fresh task independently. Record whether the same issue recurs, without turning this small personal record into a population accuracy claim.
If the comments conflict, bring the exact examples to a qualified reviewer instead of averaging the two band estimates. A half-band average of incompatible judgements is not a new independent assessment. If cost is a concern, focus a review request on the unresolved issue rather than buying a service on an unsupported promise.
Official sources and the cited paper were checked on . Product information is first-party information, not independent validation. English AIdol sells practice services and has a commercial interest in this subject. This article was prepared with AI assistance; no human expert review or endorsement by an exam owner is claimed.
This correction removes unsupported product-accuracy percentages, competitor rankings and fixed score-gain promises from the earlier article. It does not establish replacement performance statistics. Report a factual error through our editorial and corrections policy.
Is English AIdol validated to match an examiner within half a band?
No independently audited accuracy percentage is published. Treat automated bands as unofficial estimates and inspect the evidence behind individual comments.
Is weighted kappa an accuracy percentage?
No. It is an agreement statistic. The cited study concerns a particular Task 2 setup and does not validate another product or Speaking assessment.
How should I respond to contradictory feedback?
Compare the original task, the actual response and each comment’s evidence. Do not assume averaging the scores resolves a disagreement. Seek qualified review for an important unresolved judgement.