A national federation looking back at its junior pipeline often poses a question that sounds elementary. Were the under-fourteens of 2018 stronger or weaker than the under-fourteens of 2024? The question matters operationally. If the recent cohorts are weaker, the federation’s investment in development needs revisiting. If they are stronger, the trajectory is favourable and the existing programme is working. If they are roughly equal, the federation’s attention can stay on individual prospects rather than cohort-level interventions.
The question sounds elementary. The data is in the federation’s database. The query is one line of SQL.
The query produces an answer. The answer is not credible.
The reason the answer is not credible — and the deeper reason most federations stop asking the question after one or two attempts — is that the underlying scoring system has changed across the years. Rule changes, age-class shifts, tournament format updates, scoring updates, junior ranking algorithm revisions, surface mix changes in regional events, eligibility changes, qualification rule changes. Each change is small in its own year. Compound across six or eight years and the data from 2018 is not the same kind of data as the data from 2024. The cohorts cannot be compared by direct comparison of their stored scores.
This article describes why the data has the problem, what credible cross-year normalisation looks like, and what most federations end up doing instead. It connects to the federation talent pipelines article covered earlier, which dealt with cross-regional normalisation in a single year. This one deals with cross-year normalisation within a single region.
What changes between years
A short, non-exhaustive list of the kinds of change that erode comparability.
Age-class shifts. A federation that runs under-fourteens, under-sixteens and under-eighteens may at some point reconfigure to under-thirteens, under-fifteens, under-seventeens. The change is administratively reasonable; the data from before and after the change is not directly comparable.
Rule changes. The introduction of no-let, the shift to seven-point tie-break at six-all, the adjustment of let calls on serve, the changes in shot-clock duration. Each rule change subtly shifts the distribution of match outcomes. Over a decade, the cumulative effect is large.
Scoring updates. Some federations modify the junior ranking algorithm itself. A points-allocation change at the junior level rebases the entire ranking and renders the historical numbers non-comparable.
Tournament format updates. A federation may shift from single-elimination to round-robin-then-elimination for under-twelves, or vice versa. The number of matches per player per event changes, the variance in match outcomes changes, the data shape changes.
Surface mix. A region that played 60% clay events in 2018 and 30% clay events in 2024 has produced different players, in part because of the surface exposure. The cohort comparison conflates surface mix with talent quality if it is not controlled for.
Eligibility and qualification rules. Changes in birth-month cutoffs, in regional qualification rules, in scholarship eligibility — all subtly alter who enters the data in each year.
Singly, each change is manageable. Cumulatively, they are the reason the cross-year comparison defaults to undefendable.
What normalisation strategies actually work
Three strategies, in increasing rigour and cost.
Strategy one: anchor cohort comparison. Take the small set of players whose careers span both eras — players who were under-twelves in 2016 and under-eighteens in 2022, for example. Use them as anchors. Compute the cumulative ranking each player achieved in their first three years of senior pipeline. The anchors are imperfect but they provide an indirect measure: if the same player’s trajectory is steeper in one cohort than another, the cohorts produced different opportunity sets. This is the cheapest method and the most operationally accessible. The output is suggestive, not definitive.
Strategy two: cross-year ranking re-construction. Reconstruct the ranking algorithm in its various historical forms and compute, for each year, the equivalent of “what would the 2018 ranking algorithm output for the 2024 cohort using their match results re-scored under 2018 rules?” The exercise is laborious. It requires the historical algorithm specifications to be documented. Most federations cannot do this because they did not retain the algorithm specifications across versions. Where it is feasible, it produces the cleanest comparison.
Strategy three: independent-signal anchoring. Use a signal that is invariant across rule changes — match length, average rally length, serve speed (where measured), break-point conversion rate — and compare those signals across cohorts. The signals are not the full picture. Match length, for example, is a function of both player quality and rule changes (let-rule abolition shortens matches). But carefully chosen signals provide a robustness check against the ranking-based comparison. The literature surveyed at the Sloan Sports Analytics Conference and in adjacent IEEE Xplore papers covers the methodology in detail.
In practice, federations that do this work use a combination of strategies one and three, with strategy two reserved for the cases where the federation has retained sufficient algorithm documentation.
Why most federations give up
Three reasons.
Operational priority. The cross-year comparison is, by its nature, retrospective. The federation that has limited analytical capacity will prioritise current-year decisions over retrospective questions. The cross-year project sits on a back-burner until a director with a specific need asks for it.
Lack of clean documentation. Strategy two requires the historical algorithm specifications. Most federations have not maintained them. The institutional memory walks out with retiring administrators. Rebuilding the documentation is itself a substantial project.
Risk of an unwanted answer. A federation that has been claiming “the pipeline is producing stronger juniors than ever before” risks a comparison that does not support the claim. The political incentives to leave the question unanswered are real. The director who pushes the comparison through is making an institutional bet.
The SportBusiness coverage of federation development programmes documents this dynamic across multiple sports. The pattern is not specific to tennis. It is structural to federation operations.
What the comparison enables
When done credibly, the cross-year comparison answers a small set of high-value questions.
Are the federation’s investment decisions producing measurable improvement in cohort quality over time, controlling for rule changes?
Where in the pipeline — under-twelves, under-fourteens, under-sixteens, under-eighteens — has the federation’s improvement (or decline) been most pronounced?
Which regions have produced increasing or decreasing relative contribution to the national pool over time, and what does that imply for regional resource allocation?
These are the questions that, when answered with defensible data, can move federation budgets and shape multi-year strategy. The work is worth doing precisely because the answers can be acted on at the highest level.
The cross-regional version of this work — comparing regions within a single year — is the work covered in the federation talent pipelines article. The two layers complement each other. A federation that does both has, for the first time, a defensible picture of its pipeline across both region and time.
The longitudinal data problem more generally
This is a specific instance of a general problem in longitudinal data analysis. Any dataset collected over many years with evolving collection methods, evolving entity definitions, and evolving rules suffers the same erosion of comparability. The methodology is borrowed from adjacent fields — educational testing, epidemiology, economic time-series — where the problem has been studied for longer.
The Tennis Industry Magazine technology coverage has not yet covered this specific federation problem in depth, but the methodology pieces from adjacent sports apply. The work is not novel mathematics; it is novel application.
This sits inside the talent identification and scouting cluster.
What this means for your federation
Three implications.
First, before commissioning a cross-year comparison, decide what decision it is meant to inform. The comparison is expensive to do credibly. A federation that does it without a decision in mind ends up with a report that satisfies no one.
Second, the work is best scoped in stages. Start with strategy one — the anchor cohort comparison — which is cheap and operationally accessible. Use the suggestive output to decide whether the more expensive strategies are warranted. Most federations stop at strategy one because the suggestive output is already actionable.
Third, the institutional barriers are larger than the technical ones. A federation that wants this work done has to commit to potentially unwelcome findings. Without that commitment, the work will produce a report that gets quietly shelved.
To scope an anchor-cohort comparison against your federation’s data, book a 60-minute call. Bring a list of the rule and algorithm changes in the period you want to compare. We will tell you which strategy fits your data and what the realistic output looks like.