A national federation, in most countries, runs talent identification through regional coordinators. Three to twelve regions, each with a head scout, each with a stable of regional coaches who feed her information about promising juniors. The system has worked, in roughly this form, since the seventies. It still produces players. It also, almost universally, produces a particular kind of failure that is hard to see from inside.
The failure is this: the federation cannot answer a basic question on demand. Who are the top twenty under-fourteens nationally on transferable signals? Not on tournament-result ranking — that question is easy and the answer is in the federation database. The deeper question, the one that informs national-squad selection and the allocation of scarce training spots, requires a comparable data layer that does not exist in most federations.
This article describes the gap, why it persists, and what the architecture of closing it actually looks like. It draws on the training-load context covered separately, and on the broader literature surveyed at the Sloan Sports Analytics Conference and SportBusiness.
What the regional reports actually contain
Pull the file for any nominated under-fourteens prospect from any regional coordinator and you will find — in some combination, varying in completeness — three to five pages of notes. Service motion observations. Forehand mechanical signature. Movement quality on the deuce side. Match-temperament anecdotes. A scoring rubric, often a five- or seven-point scale on technique, athleticism, tactics and competitiveness. Sometimes video clips. Sometimes only the rubric.
The reports are not bad. Many are excellent. They reflect the judgement of an experienced coach watching a player she has watched for two years. The information density is high.
The problem is not the regional report. It is what happens when the federation tries to compare reports from two different regions.
Why the rubrics are not comparable
Each regional coordinator uses a rubric that has evolved locally. The “movement quality 7/10” rubric in one region is calibrated against the prospects that coordinator has personally seen over a fifteen-year career. The “movement quality 7/10” rubric in a neighbouring region is calibrated against a different population of players, by a different coordinator, with a different threshold for what 7 means versus 8.
The numbers are not on the same scale. They look like they are, because they are both expressed on a 1-10 line. They are not.
A federation that aggregates these scores — averages them, ranks them, builds policy off them — is doing what statisticians call combining incommensurable measurements. The aggregate is meaningless. The federation knows it is meaningless. The federation does it anyway, because the alternative is admitting that the national ranking question cannot be answered, which is operationally unacceptable.
This is the gap. It is not a technology problem. It is a measurement-design problem that technology can help solve.
The architecture of closure: federated normalisation
The solution does not require centralising scouting at the federation level. Regional coordinators continue to do what they do, with the tools they have, in the format they prefer. The solution is a normalisation layer that sits above the regional reports and converts them into a shared schema.
The normalisation works in three steps.
Step one: rubric reconciliation. Each regional rubric is mapped, dimension by dimension, onto a shared rubric. “Movement quality” in region A is mapped onto the same axis as “court coverage” in region B and “athletic baseline” in region C. The mapping is done by a working group of coordinators, not by federation HQ alone. The output is a translation table.
Step two: anchor-player calibration. A small set of anchor players is rated by every regional coordinator independently. The anchors are usually current under-eighteens whose trajectory is already partly known. The cross-regional ratings on these anchors reveal each coordinator’s calibration. A coordinator who rates the anchors consistently a point higher than the others is calibrated to apply a one-point offset to her downstream ratings.
Step three: scoring on the normalised rubric. Each prospect’s regional report is scored in the regional rubric. The normalisation layer translates it into the federated rubric and applies the calibration offset. The output is a comparable score that aggregates across regions.
This is not novel mathematics. It is the same family of methods used in educational testing for international score comparability and in cross-broadcast judging in figure skating. The framing matters more than the technique. Federations resist the framing because it implies their existing data layer is not comparable, which is true.
Where junior tournament data fits
Tournament results are the only signal currently comparable across regions. The federation knows who won what, against whom, at what age. The problem with tournament results in isolation is that they are dominated by physical maturation. The biggest under-fourteen in a region wins more matches in February than the technically best under-fourteen because his service motion can be powered by his height. By under-eighteens, the maturation advantage has redistributed. The federation that selects on tournament results selects biased toward early developers and against late developers.
A defensible data layer combines normalised scouting scores with tournament results, weighted to control for maturation. The mathematics is described well enough in the Sloan papers on junior development and in adjacent work catalogued at Tennis Industry Magazine. The implementation is the harder part. It requires the rubric reconciliation and the anchor calibration to be done first.
Computer-vision augmentation
A subset of regions are now layering computer vision over match video from regional under-fourteens tournaments. The same measurements covered in the computer vision primer — shot direction, depth, contact-point class — applied at scale produce a quantitative signal that supplements the coordinator’s qualitative rubric. The signal is most useful for technical dimensions (consistency, contact-point efficiency) and least useful for tactical dimensions (intent, point construction), as the primer makes clear.
Where computer vision matters most: it provides a calibration check on the coordinator’s rubric. A coordinator who rates a player’s forehand mechanics 8/10 but whose CV-derived contact-point efficiency sits in the cohort’s bottom quartile is calibrated differently than the other coordinators. The CV data is not a replacement for the rubric. It is an independent check on it.
The institutional barrier
The technical work is the easy part. The institutional work — getting regional coordinators to agree on a shared rubric, to have their calibration assessed against anchor players, to have their scores adjusted by a federated translation layer — is the harder part. Most federations stall here, not in the modelling.
The framing that tends to work, from operations that have made the transition: the normalisation does not take authority from the regional coordinator. It makes her scores legible to the federation. The coordinator continues to evaluate. The federation, for the first time, can compare. The two functions are not in competition. Articulating that clearly is what makes the institutional barrier passable.
This work fits inside the broader talent identification and scouting cluster. The cross-year normalisation problem — how to compare under-fourteens cohorts in 2024 against under-fourteens cohorts in 2027 — is a related but distinct problem covered separately later in this series.
What this means for your federation
Three questions before any technology procurement.
First, are your regional rubrics documented? If they are not written down, the rubric reconciliation step cannot start. The data layer cannot be built on top of undocumented intuition.
Second, is there appetite at the head-of-development level for cross-regional calibration? Without that appetite, the anchor-player exercise will not happen and the normalisation will fail. The appetite has to come from the top.
Third, what is the actual operational use case? “Comparable national ranking” sounds compelling. The federation that builds the data layer and then has no clear decision it wants to inform with it — squad selection? scholarship allocation? regional resource distribution? — produces a dashboard nobody reads. Define the decision first.
If the answers to those three are clear, the architecture is buildable in nine to twelve months. To scope what that looks like in your specific federation, book a 60-minute call. We will work backwards from the decision you want to inform.