Published calibration evidence
A mark you can test.
Students test essay markers the obvious way: submit the same essay twice and see if the number holds. We think that's exactly right — so we run that test against a frozen benchmark and publish the results, stamped with the date and engine version they were measured on. To our knowledge, no other GAMSAT marker publishes any calibration evidence at all.
These numbers describe an older version of our marking engine.
They were measured on engine je-2026-08a; we are now running je-2026-08g. The figures below are real and unedited, but until we re-run the benchmark on the current engine they are evidence about the older one. We would rather say that here than let you assume otherwise.
±0
Re-mark consistency
Repeat-marked probe essays land within 2 points of themselves (median 0) in our frozen-benchmark runs. Measured 2 August 2026, engine je-2026-08a.
6 pts
Cross-examiner agreement
Average gap between our two independent AI examiners on the same essay, across the benchmark. Every mark is this cross-check; a real disagreement calls in a third examiner. Measured 2 August 2026, engine je-2026-08a.
47/48
Benchmark accuracy
Essays landing inside their expected score range on the frozen benchmark — a consistency check against the engine's own reviewed baseline, not against official ACER scores (that comparison is the tracker beside this). Measured 2 August 2026, engine je-2026-08a.
1
Tracking against real ACER results
1 student has reported a real ACER Section II score against essays we marked. We hold any accuracy claim until 30, and publish it then either way. Students report their official S2 after each sitting and each sees privately how close their pre-sitting average came.
The method, plainly
ACER publishes exactly two assessment criteria for Section II — the quality of the thinking, and the control of language in expressing it — and no marking rubric. Our instrument is six sub-criteria built from those two published criteria, with what each level on our reported scale means fixed by written reference essays the examiners are shown — not by any real ACER result, because ACER releases none — and watched by regression evals: a frozen benchmark of 48 essays with expected score ranges, run on a drift schedule and re-run deliberately around engine changes. When we deliberately change marking behaviour, the benchmark is re-frozen in a recorded commit so score movement always comes with a recorded reason. When a change has shipped ahead of its re-run — as one has now — this page carries a notice at the top saying so, because a statistic that does not name the engine it was measured on is a claim about a configuration that may no longer exist.
Every essay, on every plan, is marked by two independent AI examiners and cross-checked; ACER itself states that Section II responses are “scored independently three times” — the same principle. When our two examiners genuinely disagree, a third examiner from a different model family is called in and each criterion is decided by the middle mark, disclosed on the grade.
The real-ACER comparison beside the benchmark works like this: when a student reports their official Section II score, we compare it to the average of the essays we had already marked for them before that sitting. We compute nothing until 10 students have reported, and we make no public accuracy claim until 30 — then we publish it whether or not it flatters us. Its limits are real and we state them rather than bury them: an official S2 comes from two essays written under exam pressure while ours come from practice written whenever the student liked, so this measures how well our marking PREDICTS a real result, not a re-mark of the same script; and reported scores are unverified.
Honest limits: these are calibrated practice estimates, not official ACER marks — only ACER issues those, on exam day. Consistency is measured on our frozen benchmark, not on your individual essay. GAMSAT® is a registered trademark of ACER, which is not affiliated with and does not endorse Aptavia.
Looking for Sections I and III?
Those are multiple-choice, so they are measured rather than marked: question difficulty from real attempts, papers equated against our own bank, and an error band published with the score.
How our practice questions are measured