Why Your IELTS Speaking Score Is a Range — and How to Train for the Strict Examiner
On this page
Your IELTS Speaking score can reasonably fall into a range rather than one fixed number because band descriptors describe a spread of performance within each band, so a response sitting near a boundary can be judged either way depending on which examiner you get and how they weigh its particular strengths and weaknesses. Training for the stricter end of that range, rather than assuming the more generous outcome, is the safer strategy because it prepares you for a standard that holds up regardless of who happens to be marking you on the day.
Key takeaways
- Band descriptors describe a spread of performance within each band, so borderline responses can reasonably be judged either way.
- Examiner training and standardisation reduce this variation but do not remove it entirely, especially near a boundary.
- Consistency of errors and coherence under unscripted follow-up questions push a borderline performance up or down more than one strong or weak moment.
- Training for the stricter, lower-bound outcome is safer than hoping for a lenient one, since it holds up regardless of who marks you.
- Mock practice that shows both a best-case and a worst-case reading of the same response makes this idea concrete instead of abstract.
Why the same performance can get different scores
IELTS Speaking is marked against public band descriptors covering four criteria, and each band’s descriptor is a description of a type of performance, not a single fixed transcript. A real response is messy — it might show Band 7 level flexibility in its grammar but Band 6 level hesitation in its delivery — and an examiner has to weigh those mixed signals into one score per criterion. Near a boundary, two trained examiners can reasonably weigh the same mixed performance slightly differently, one rounding up because the strengths stood out to them, another rounding down because the weaknesses did.
This is not the same as saying scoring is inconsistent or unreliable. Examiners go through standardised training and certification specifically to keep this variation as small as possible, and for a performance that is clearly and consistently at one band level, most examiners will agree closely. The variation this post is about is specifically the narrower band of genuinely borderline performances, where a response could honestly be described as sitting at the top of one band or the bottom of the next.
What pushes a borderline performance up or down
Consistency of errors, not isolated moments
A response with one genuinely impressive sentence surrounded by frequent basic errors tends to read differently from a response that is steady throughout, even if both contain a similar mix of strong and weak language somewhere in the transcript. Examiners are listening for a pattern across the whole response, not scoring based on the single best or single worst sentence. A recurring error that shows up in nearly every answer is a much stronger signal than an occasional slip, because it tells the examiner what your speech actually looks like on a typical sentence, not just your best one.
Coherence under follow-up questions
Part 1 and Part 3 both include questions you cannot fully prepare for in advance, and how coherently you handle an unscripted follow-up is a strong signal near a boundary. A candidate who can only maintain fluent, well-organised speech on rehearsed material but becomes noticeably less coherent the moment a question goes slightly off-script gives the examiner useful information about how stable that performance actually is. This is part of why memorised answers tend to backfire, a point covered in more detail in how to answer any Part 2 cue card — a script that cannot flex under a follow-up question exposes exactly the instability that pushes a borderline score down rather than up.
Range used naturally versus range used to impress
Near a boundary, examiners are also weighing whether more advanced vocabulary or grammar was used naturally, where it fit, or inserted in a way that reads as an attempt to sound more advanced than the rest of the response supports. A complex structure used once, awkwardly, in an otherwise simple response tends to read as reaching rather than as evidence of genuine range, which is a distinction covered in more depth in what actually separates Band 6 from Band 7.
Why training for the strict examiner is the safer strategy
You cannot know in advance how strictly your particular examiner will weigh a borderline performance, and you cannot choose who you get on test day. Because of that, treating the lenient, generous reading of your own performance as the expected outcome is a risky strategy — it means your actual result depends on a factor entirely outside your control. Treating the stricter, more critical reading as the standard you need to clear is safer, because a performance that holds up under a strict reading will also hold up under a generous one, but the reverse is not true.
This does not mean assuming the worst about your ability or becoming discouraged. It means being honest, in practice, about where a response has a recurring weakness rather than giving yourself the benefit of the doubt the way a lenient examiner might. A recurring grammar error you excuse in practice because “the examiner will understand what I meant” is exactly the kind of thing that can push a borderline response down rather than up on the actual test.
Making this concrete: best case versus worst case
The idea of a scoring range is easiest to apply in practice by treating a mock response two ways: score it once as a generous, friendly examiner would, focusing on what went well, and once as a strict examiner would, focusing specifically on recurring weaknesses and how they held up under any unscripted follow-up. The gap between those two readings is informative — a small gap suggests a stable performance at that level, while a large gap suggests a borderline one that depends heavily on how it happens to be judged. Practising against the stricter of the two readings, and working specifically to close that gap, is a more reliable way to prepare than practising against whichever reading feels more encouraging in the moment. IELTS Mentor AI scores mock responses this way directly, showing a best-case and worst-case band so you can see exactly how wide that gap is on a given response rather than estimating it yourself.
What this idea is not saying
This is not a claim about how often examiners disagree, or about any particular margin separating a lenient and a strict reading — there is no published figure for that, and treating this as a precise, measurable gap you can look up would be misleading. It is a way of thinking about a genuinely borderline performance: such a performance can honestly be described from two angles, a more forgiving one and a more critical one, and the useful question is not “which one is correct” but “does my performance hold up under the more critical angle.” Framed that way, the idea is a preparation strategy, not a statistic about examiners.
It is also worth separating this from simply feeling nervous about test day. Nerves affect delivery on the day itself, which is a different issue from whether a performance, delivered as intended, sits solidly within a band or right on its edge. Training for the strict reading is about the underlying performance, not about managing exam-day anxiety, though a performance that is solid under a strict reading tends to leave more room for error if nerves do affect delivery on the day.
A practical routine for training toward the stricter reading
Record a full response, then review it twice with a deliberate change in mindset each time. On the first pass, note everything that went well — this is the encouraging, best-case reading, and it matters because it tells you what to keep doing. On the second pass, specifically hunt for the same weakness appearing more than once, and for how the response held up the moment a question moved slightly away from anything you had prepared. That second pass is the strict reading, and it is usually where the actual, actionable feedback lives. Over several weeks, track whether the gap between your two readings narrows — that narrowing is a more meaningful sign of progress than either reading in isolation, because it means your performance is becoming less dependent on how generously it happens to be judged.
The range is information, not bad news
A wide gap between a best-case and worst-case reading of your own speaking is not a discouraging result — it is a more precise diagnosis than a single number would give you. A single band score can hide exactly how borderline a performance actually is; a range makes that visible, and a visible gap is something you can specifically close through practice, rather than something you can only hope gets judged kindly on the day.
Frequently asked questions
Why can two examiners give slightly different scores for the same performance?
Band descriptors describe a range of performance within each band, not a single fixed pattern, so a response sitting near the boundary between two bands can reasonably be judged either way depending on how an individual examiner weighs its strengths and weaknesses. Examiners are trained and standardised to reduce this, but it does not eliminate borderline judgement calls entirely.
Should you rely on getting a lenient examiner on test day?
No. You cannot know in advance how strictly your performance will be judged, so the safer strategy is preparing to a standard that holds up even under a stricter reading, rather than hoping for a generous one. A performance that only works with a lenient examiner is not a stable Band 7, it is a borderline one that got lucky.
What pushes a borderline Speaking performance up rather than down?
Consistency matters more than isolated strong moments - a response with one impressive sentence surrounded by frequent basic errors reads differently from one that is steady throughout. Coherence under follow-up questions, where you cannot rely on rehearsed language, is also a strong signal examiners weigh heavily near a boundary.
Does practising with a strict standard risk lowering your confidence unfairly?
It can feel that way at first, but the goal is not to feel worse - it is to see the specific gaps a lenient self-assessment would hide. Most learners find that once they know precisely what is holding a criterion at the lower end, the fix is more concrete and less discouraging than a vague sense of being "not quite Band 7 yet."
More from the blog
IELTS Speaking Part 3: How to Structure an Opinion
How to answer IELTS Speaking Part 3 discussion questions using a position-reason-example-concession structure, and how the section differs from Part 2.
Do AI IELTS Speaking Scorers Actually Work? What They Get Right and Wrong
Where AI IELTS Speaking scoring is genuinely reliable, where it still falls short of a human examiner, and how to use any automated scorer without being misled by it.