parkrun's age grade tells you how you ran for your age — but makes no allowance for the course beneath your feet. Using 3.4 million Australian results, we measure the difficulty of every Australian parkrun and build a grade that lets you compare a run on a hard trail with a run on a fast flat course, fairly.
Section 1
Every Saturday morning, tens of thousands of Australians run a free, timed 5km at one of more than five hundred parkrun events, from city parklands to coastal paths to bush trails. Every finisher receives an age grade — a percentage score comparing their time to a world-class standard for their age and gender, so that runners of different ages can be compared on a common scale.
Age grading is a genuinely good idea, but parkrun is candid about one thing it leaves out. In its own words, age grading "makes no allowance for different weather conditions or the varying terrains of our courses." A 28-minute run on a flat, fast course and a 28-minute run on a hilly trail receive the same age grade, even though the second is plainly the harder effort. This paper takes that gap as its starting point and asks a simple question: how much harder is one Australian parkrun than another, and can we measure it well enough to compare performances fairly across courses?
Using 3,446,488 Australian parkrun results from the 2025 calendar year, we derive a difficulty factor for every one of 530 courses, and combine it with the existing age and gender adjustment into a single number — the Terrain-Adjusted Grade, or TAG. A runner's TAG is comparable across every Australian parkrun, whatever the terrain. The full table of course factors is published in Section 5, and a free calculator lets any parkrunner compute their own.
The approach here was first developed and validated on the United Kingdom's parkrun results. The mathematics is the same, with two differences worth stating plainly at the outset. First, the Australian factors are anchored to the fastest Australian parkrun rather than the UK reference, so they describe difficulty within Australia — comparisons between Australian courses are exact, while comparing the absolute grade level between countries should be treated as approximate. Second, the two independent methods used to derive the factors agree slightly less closely in Australia than in the UK, for reasons we explain honestly in Section 4: Australia is a far larger country with fewer, smaller, more scattered courses, which makes one of the methods work with thinner data. Neither difference changes the method; both are simply the conditions it operates under here.
Section 2
Age grading converts a finish time into a percentage by comparing it to a standard time — the best a runner of that age and gender might achieve:
The standard times come from world-record performances scaled by an age factor reflecting the expected decline with age. parkrun's tables are loosely based on those published by World Masters Athletics, in the widely used 2020 revision maintained by the statistician Alan Jones (the "ALJ 2020" tables). It is a genuine advance — without it, older runners and women would have no way to compare their efforts against younger men. But it is not perfectly fair, and it is worth being honest about where it falls short before building anything on top of it.
We first confirmed that Australia uses the same age-grading standard as the UK, so the method carries across cleanly. Working backwards from each published grade and finish time recovers the standard time parkrun assigned to each single year of age, and these land on the same table the UK uses, with the open-class anchors matching to the second (12:53 for men, 14:48 for women). The standard is shared.
So is its main weakness. When we compare the standard parkrun actually applies against the published ALJ figures, a clear pattern emerges. For men, the two track closely across the whole age range — parkrun's standard is well calibrated, never more than about twenty seconds from ALJ even into the eighties. For women, the two agree at younger ages but then diverge: from around the mid-fifties, parkrun's standard for women becomes progressively more generous than ALJ, and the gap widens steeply with age. By 65 the women's standard is roughly forty-five seconds slower than ALJ would set it; by 80 it is more than two minutes. A slower standard means an easier grade, so older women receive systematically higher age grades than the published standard implies — not because of anything they did, but because of how the table is built.
How far parkrun's standard departs from ALJ, by age
Seconds by which Australia's parkrun standard is slower (more generous) than the published ALJ standard. Men stay close throughout; the women's standard pulls steadily away from the mid-fifties on.
Australian-derived parkrun standard minus ALJ standard, ages 25–84. Positive = parkrun more generous.
This is the same pattern found in the United Kingdom analysis, reproduced here on a second continent and a different field — which strengthens the original conclusion. The divergence is a property of the grading tables themselves, not of any one country's runners. The full reverse-engineered standard is given below, with the ALJ figures alongside so the gap can be read directly. The women's column pulls steadily away down the page; the men's barely moves.
| Age | ALJ men | parkrun men | Diff | ALJ women | parkrun women | Diff |
|---|
parkrun standard reverse-engineered from 3.4M Australian results; ALJ is the published 2020 standard. Diff = parkrun − ALJ (seconds); positive means parkrun is more generous. Figures highlighted where the female gap exceeds 30 seconds.
This matters for what follows. Reverse-engineering parkrun's standard was the diagnostic that exposed the female-veteran flaw; it is not the standard we build TAG on. The Terrain-Adjusted Grade uses the published ALJ 2020 age factors for its age-and-gender correction, not the more generous figures parkrun applies. So TAG does not inherit the inflation — it quietly corrects it. An older woman's TAG is measured against the same auditable standard as everyone else's, which is precisely why a TAG can be lower than the age grade parkrun shows her: the published standard is the stricter, fairer one. TAG therefore improves on parkrun's grade in two distinct ways at once — it adopts a published, even-handed age standard in place of parkrun's over-generous one, and it adds the course correction parkrun makes no attempt at.
In short
parkrun's own grading has a real, documented flaw: its standard runs increasingly generous for older women, so two runners of equal merit do not always receive equal grades. TAG sidesteps this by using the published ALJ tables rather than parkrun's reverse-engineered ones, and then adds the thing parkrun does not adjust for at all — the difficulty of the course. The rest of this paper is about that second correction.
The Question
Age grading corrects for who you are. It says nothing about where you run. Across 530 Australian courses, that turns out to matter a great deal.
Section 3
The difference between Australian courses is large. A runner finishing the mountainous Mundy Regional course near Perth — a bush trail climbing more than two hundred metres — in a given time is working far harder than a runner posting the same time on a flat, fast coastal course like Port Sorell in Tasmania. Yet both receive the identical age grade. Across the full set of courses, the difficulty difference between the easiest and hardest amounts to roughly thirty per cent in finishing time for an equivalent effort.
This is exactly the variation age grading sets aside, and it is not a minor correction. A keen runner who happens to live near a hard course will see lower age grades for the rest of their parkrun life than an identical runner near a fast one — not because they are slower, but because of where they run. A fair comparison across courses requires knowing how much of a runner's grade is the course and how much is the runner.
The challenge is that course difficulty is not something parkrun publishes, and it cannot be read off a map. Elevation, surface, exposure, the number of turns, even the typical congestion at the start all feed into how fast a course runs. Rather than try to model these directly, we let the results speak: with millions of finishes across hundreds of courses, the data itself reveals how much harder one course is than another. The next section describes two completely different ways of extracting that signal.
Section 4
We derive each course's difficulty factor two completely independent ways. The value of using two methods built on different data and different assumptions is that where they agree, we can be confident the factor reflects real terrain rather than a quirk of one approach.
Because age grading already removes age and gender, what is left in a course's distribution of grades — comparing like with like across everyone who runs there — is the effect of the course itself. We take the profile of grades a course produces, compare it against the best any course achieves across that profile, and read the typical shortfall as the course's factor. This method uses the entire field at each venue.
The second method ignores the field and tracks individuals. When the same runner records times at two different courses, the difference between their performances is a direct measure of the difference between the courses, because their own ability is held constant. Across every runner who visited two or more courses, we gather all such comparisons and solve for the single set of factors that best fits them all at once. Each runner is their own control, which removes ability from the picture in a way the first method cannot.
Across the 530 courses the two methods correlate at r = 0.733, with a mean absolute difference of about 0.05. They agree closely on the extremes — both rank the same handful of flat coastal courses fastest and the same trail courses hardest — and they agree well in the dense, diverse middle. Where they part company is at small and remote courses, and the reason is instructive: each method has a weakness that surfaces there, and the two weaknesses pull in opposite directions.
The two methods at eight illustrative courses
Each row spans the gap between the grade-distribution factor (red) and the runner-matching factor (blue); the dark marker is the blended value used. Note the close agreement at the fast and slow extremes, and the wider gaps at the regional courses between.
Grade-distribution, runner-matching and blended factors for eight courses. Fastest course (Port Sorell) = 1.000.
The grade-distribution method tends to read a remote course as harder than it is, because its field is slow and recreational and may contain no fast runners at all — the absence of fast times looks like difficulty. The runner-matching method tends to read the same course as easier than it is, because the few runners linking it to the wider network are mostly locals who know the course and run it relatively well, which looks like speed. Both effects are real, both are larger in Australia's sparse network than in the UK, and they point opposite ways.
We checked whether a single cause could be isolated and corrected. Course descriptions confirm the genuinely hard courses are genuinely hard. Seasonal heat is real — courses in the tropical north run several per cent slower in summer — but it averages out of a full-year factor and does not drive the disagreement. Controlling for the month each run took place barely moves the matching factors at all. In the end, neither bias can be cleanly removed, and there is no third measurement to break the tie. When two independent estimates carry opposing biases of similar, uncertain size, the principled response is to average them, so each method's error partly cancels the other's.
How the factors are combined
The final factor for each course is the even average of the two methods, anchored so the fastest course equals 1.000. We considered weighting the blend toward one method and rejected it: any departure from an even split amplifies one of the two known biases without good evidence to justify it. The fifty-fifty blend is the honest choice and keeps Australia consistent with the UK work.
Section 5
The factor for each course can be read as: "at this course, runners post grades roughly (1 − factor) × 100% below what they would on the fastest course." A factor of 0.90 means grades about ten per cent below the reference. The fastest Australian parkrun anchors the scale at 1.000; every other factor is below it.
The distribution is clustered, with a long tail of hard courses. Most Australian parkruns sit between about 0.90 and 0.97 — they are fast or moderate. A smaller number of genuinely challenging trail and hill courses trail off toward 0.70 and below.
Distribution of course speed factors — 530 Australian parkruns
Number of courses in each factor band. The bulk cluster around 0.94–0.96, with a long left tail of harder courses. Fastest course = 1.000.
Blended factors across 530 courses. Median 0.945, minimum 0.694 (Mundy Regional).
The two method columns below let you see the inputs to each factor: the grade-distribution figure, the runner-matching figure, and the blended value used in TAG. Where the two methods diverge widely — almost always at smaller, more remote courses — the blended factor is a reasoned compromise rather than a precise answer, and should be read with that in mind. Search by name or sort by any column.
| Course | Grade dist. | Runner match | Factor | Speed |
|---|
530 courses. Factor of 1.000 = fastest Australian parkrun (Port Sorell). Blended evenly from the grade-distribution and runner-matching methods.
Section 6
Combining the age and gender correction with the course factor gives a single number — the Terrain-Adjusted Grade — that accounts for who you are and where you run:
Because the course factor enters the denominator, a hard course raises a runner's TAG to compensate exactly for the difficulty it imposed. Two runners of equal ability — one on a flat city course, one on a tough trail — receive the same TAG for equivalent efforts, where the plain age grade would have rewarded the first and penalised the second. That is the whole point: TAG is a measure of the runner, not of where they happened to run.
The same factors also turn TAG into a course predictor. Because the factor describes how much slower one course is than another, a time at one course implies an expected time at any other:
What TAG gives an Australian parkrunner
A single grade that is comparable across all 530 courses — so you can see your true progress as you travel between parkruns, compare a run on a hard regional trail with one on a fast city course on equal terms, and predict what you would run anywhere else. It is offered as a complement to parkrun's own age grade, not a replacement, and is not affiliated with parkrun.
Section 7
The factors for small and remote courses carry more uncertainty than those for large, diverse ones, because that is where the two methods diverge and the blend is a compromise. A meaningful minority of Australian courses see no fast runners all year, which weakens the grade-distribution signal there. Seasonal heat is a real effect at individual events — a summer run in the tropical north can be several per cent slower than the year-round factor implies — though it averages out of the factor itself, and a seasonal adjustment is a natural future addition. Some courses run multiple tight-cornered laps where runners may cover slightly less than the full distance, mildly flattering their factors. And the factors reflect the 2025 course layouts; a significant change to a course would require recalculation. Finally, because the factors are anchored within Australia, comparisons of the absolute grade level against other countries should be treated as approximate.
The full set of factors and a free calculator — compute your TAG, compare any of your times across all 530 courses, and predict your expected time anywhere — are at parkrun-calculator-anz.openair.tools. TAG also underpins the cross-distance race predictor at skamper.openair.tools, which estimates times from a single recent result across distances from parkrun to ultramarathon.
This is a working paper and we welcome feedback — from Australian parkrun regulars on whether the factors match your experience of relative course difficulty, from anyone with local knowledge of the harder regional courses, and from statisticians on the method. The same approach is being prepared next for New Zealand.
Contact
Openair Research · parkrun calculator · Skamper TAG calculator · Independent analysis. Not affiliated with or endorsed by parkrun.
1. parkrun is a series of free, weekly, timed 5km events. "parkrun" (lowercase p) is used throughout, consistent with the organisation's own style. This analysis uses publicly recorded results and is not affiliated with or endorsed by parkrun.
2. "Gender" is used throughout, matching parkrun's own categorisation: runners register as female, male, prefer not to say, or another gender identity, and results are presented accordingly. The ALJ age-grading standards underlying the age grade (and TAG) are themselves derived from male and female physiological performance and so are sex-based proxies; parkrun applies them according to a runner's self-declared gender, and age-grade data is consequently available only for the female and male categories.
3. Dataset: 3,727,763 raw Australian results for the 2025 calendar year, of which 3,446,488 remain after removing data errors (finishes faster than 13:00 or with implausible grades). Factors were derived for the 530 courses with sufficient data.
4. parkrun's published grade rests on tables loosely based on the ALJ standard; reverse-engineering them from the Australian data shows they match ALJ closely for men but run progressively generous for older women, identical to the UK. TAG instead uses the published ALJ 2020 age factors directly. Open-class 5km standards: 12:53 men, 14:48 women.
5. Course factors blend two independent methods — a grade-distribution method across each course's whole field, and a runner-matching method tracking the same runners across courses — averaged evenly and anchored so the fastest Australian course equals 1.000. The two methods correlate at r = 0.733 across 530 courses (mean absolute difference 0.05).
6. The method follows the approach developed for the United Kingdom; see the companion UK paper for the age-grading analysis in fuller detail.