Published measurement evidence
A score that tells you how sure it is.
Two practice papers are never the same paper. Ours are drawn from one bank and deliberately avoid repeating what you have already seen, so one sitting can run harder than the next and report the same number for it. This page is what we do about that: we measure how hard each question actually is from real attempts, compare your paper against our own bank, and publish an error band with the score instead of a bare figure. Your score itself does not move — we tell you which way the paper leaned and let you read the number accordingly. Where the evidence is too thin, we publish the refusal rather than a number.
Where the measurement stands today
Section I
371 questions · 0 measured (0%)
Not enough of them have been answered often enough yet to publish an average difficulty — we need 30 measured questions in a section before that number means anything, and a question counts as measured at 10 attempts. Until then, a paper in this section carries no statement about how it drew, and the results screen says so. The score is the same either way.
Section III
906 questions · 0 measured (0%)
Not enough of them have been answered often enough yet to publish an average difficulty — we need 30 measured questions in a section before that number means anything, and a question counts as measured at 10 attempts. Until then, a paper in this section carries no statement about how it drew, and the results screen says so. The score is the same either way.
Difficulty equating is internal to our own question bank, and not enough of our questions have been answered often enough to equate our practice papers against each other yet. So far: Section I 0/371 (0%); Section III 0/906 (0%). Of 44 practice sections sat in the last 90 days, at most 0 (0%) carried enough measured questions to be equated. Until that changes, a section estimate carries no statement about how its paper drew, and says so. It carries the same conversion it always has — equating never moves the number either way. Measured over question statistics over all recorded attempts, recomputed nightly; a question is measured at 10 attempts.
read 25 August 2026 · sittings counted over the last 90 days · method internal-mean-equating/v2
What we measure
Two things, both from real attempts by real students on our own questions, recomputed nightly. Facility is the share of attempts that got a question right — our measure of how hard it is. We do not store one until a question has been answered 10 times, because below that it is noise wearing a decimal point. Discrimination is the correlation between getting a question right and being strong in that section otherwise (a point-biserial), and it needs 50 answer-and-ability pairs before we store one.
Only facility is used to equate a paper. Discrimination is used to find questions that are behaving badly — one that strong students miss and weak students get is usually a broken question, not a hard one — and it does not move anybody’s score. Both describe our cohort attempting our bank. They are not the real exam’s question statistics and are never presented as them.
What we equate — and what we refuse to do with it
Paper difficulty is equated internally: this paper's measured questions are compared against the mean difficulty of our own question bank for the same section, in raw-percent terms. What that comparison changes is what we say about the estimate — whether this paper ran harder or easier than our bank's average, and therefore whether to read the estimate as conservative or generous. It does not change the estimate. The number stays the one our raw-to-scale conversion gives, and the conversion itself is unchanged.
The comparison happens in raw-percent space: we take the average difficulty of the measured questions on your paper and set it against the average difficulty of our whole bank for the same section, and the gap between them is how much harder or easier your paper drew. Only the measured part of a paper counts, and the gap is scaled by how much of the paper that is — we can speak for the questions we have measured and for no others. We will not state a gap larger than 8 raw percentage points either way: past that the measurement is likelier to be broken than the paper is to be extreme.
How hard a paper drew never changes the score we report. We could apply it — the arithmetic is simple, and it would let us tell you a kinder number on a hard paper — but the evidence is not strong enough to spend that way: question difficulty here is measured on our own students attempting our own bank, and those students chose to sit that paper and chose to finish it. The correction it would produce is not small either. At the largest difference we would state, it moves a reported score by 2 to 7 scaled points across the range competitive candidates sit in, in both directions, depending only on where you sit on the curve. So we publish the direction as context instead: whether this paper ran harder or easier than our bank's average, and what that means for reading your estimate.
So what you get on your results screen is a sentence, not a second number. If your paper ran harder than our bank’s average, it says so and tells you to read your estimate as conservative; if it ran easier, it says to read it as generous; if it ran level, it says that too. There is no adjusted score anywhere in the product to compare against the one you were given, because we do not compute one.
This is internal, and the word matters: it is our questions against our own questions. It is not a comparison with the real exam, with another provider, or with any external standard.
How to read a band
The band is one standard error for a test of this length and this spread of question difficulties, converted through the same scale conversion as the score. Questions we have not measured widen it. It is an estimate of how much this instrument's own number would move on a comparable paper — it does not model the real exam's marking, scaling or cohort.
A band is how far this instrument's own number would be expected to move if you sat a comparable paper drawn from the same bank — it is not a range your real result will land in, and nothing here can tell you that. Questions we have not measured make it wider, never narrower, so a young bank publishes cautious bands rather than confident ones. Below our floors we publish no band at all and say why.
We round it outward and never publish one smaller than a single point, because we would rather overstate our own uncertainty than understate it. It also assumes questions are independent of each other, which ours are not quite: they arrive in groups sharing one stimulus, so the real error is somewhat larger than one standard error of this kind. Treat a band as a floor on the uncertainty, not a ceiling.
It sits with the score on your results, and on a trend line it does not appear at all: a band drawn over a climbing line reads as a claim about where you are heading, and it is not one.
The floors, and what happens below them
- A question is measured at 10 attempts. Below that it counts as unmeasured — which widens a band, and never narrows it.
- We will only say how a paper drew when 60% of it is measured and at least 8 of its questions are. Below that we say nothing about its difficulty rather than guess at it — and the score, which that comparison never touches, is unchanged either way.
- Our bank needs 30 measured questions in a section before it can be a baseline for anything.
- A band needs 8 questions and 25% of them measured. Below that we publish no band and say which floor we missed.
- Section II has no band here at all. An essay is marked by examiners, not measured from question statistics, so a sitting that includes one publishes no combined band on the overall — its marking evidence is on the essay page instead.
What we do not do
Nothing here is fitted to real ACER results. Per-sitting, per-question ACER ground truth is not obtainable — not by us, not by anyone outside ACER, and not later — so every comparison on this page is internal: our own questions measured against our own questions, on our own students' attempts.
Our raw-percent-to-score conversion is hand-set, and none of this moves it. Nor does any of it move the raw percentage that goes in: your score is that conversion applied to the questions you actually got right. What the difficulty measurement changes is what we say around the number, never the number.
So there are two separate things on your results screen, and it is worth keeping them apart. The conversion from raw percentage to a scaled number is a judgement we made by hand, and this work does not validate it — nor does it quietly correct it behind your back. What this work measures is narrower and real: which of our papers ran harder than which, and how much our own number would move on a comparable paper. Anyone who tells you their practice score is calibrated to the real exam is telling you something they cannot know.
Honest limits: these are practice estimates from our own bank, measured on our own students. Only ACER issues real GAMSAT results, on exam day. Reported figures are floored, windowed and dated; where a figure has not cleared its floor this page says so rather than rounding it into existence. GAMSAT® is a registered trademark of ACER, which is not affiliated with and does not endorse Aptavia.
Looking for the essay side?
Section II is marked, not measured from question statistics, so its evidence is a different instrument entirely — re-mark consistency, agreement between examiners, and how our marks track reported results.
How our essay marking is calibrated