Why Sports Analytics Needs an Ontology
By Scott Krotee
Every sports organization is drowning in data and starving for answers. The scores are logged. The film is graded. The trackers are running. And yet the single question that matters — how good is this player, really, and what should we do about it? — still takes days to answer, if it gets answered at all. The problem is almost never that the data is missing. The problem is that nothing connects it.
There is a quiet assumption in sports technology that more data is the same as more insight. Buy another tracking system, add another scouting app, export another spreadsheet, and clarity will follow. It rarely does. What follows instead is fragmentation — dozens of tools, each with its own format, its own definitions, its own idea of what a "possession" or a "grade" or a "player" even is. The data accumulates. The understanding does not.
The data isn't the problem. The disconnect is.
Consider a single athlete. Their identity lives in one system. Their game outcomes live in another. A scout's film grades sit in a shared drive. Combine numbers are in a PDF someone emailed last spring. A third-party tracker holds their movement data behind an API. Every one of those sources is describing the same player — but none of them agree on how to say so, and none of them are wired together.
So when a decision-maker asks a simple question, an analyst has to become a translator. They pull exports, reconcile mismatched names, guess at which grade maps to which event, and stitch it together by hand in a spreadsheet that will be stale by the time it's finished. The intelligence was always in the data. It just wasn't reachable.
"Fragmented data isn't just inconvenient. It's inert. Until the pieces are connected by a shared model, most of what an organization collects will never influence a single decision."
What an ontology actually is
An ontology is not a database, a dashboard, or another data lake to dump things into. It is a shared model of reality — a precise definition of the objects that matter, the relationships between them, and the rules for how they fit together. In an operational context, an ontology answers questions like: What is a player? What is a team? What is an observation of performance, and how does one observation roll up into an evaluation, a ranking, a decision?
Once those objects and relationships are defined, every incoming source can be mapped onto them. A film grade, a tracked event, and a manual submission stop being three unrelated records and become three observations of the same thing, in the same language. The messy, real-world inputs are still messy — but the model on top of them is clean, and analytics can finally run on a clean model.
Fragmented in. Standardized here. Decisions out.
Why fragmented data stays unused
The uncomfortable truth is that most sports data is collected and then quietly abandoned. Not because it lacked value, but because the cost of making it usable was higher than any single decision could justify. When every question requires a fresh act of manual integration, people stop asking. The data becomes an archive, not an asset.
An ontology inverts that economics. The integration work happens once, at the level of the model, instead of over and over at the level of each question. After that, new data flows in and immediately joins everything already there. The marginal cost of asking a hard question drops toward zero — and that is when an organization actually starts using what it has been collecting all along.
From data to decisions
This is the shift that matters. Traditional analytics stops at description: here is a chart, here is a table, here is what happened. An ontology-backed system is built for decisions: here is the player, connected to every observation of them, ranked against every comparable, with the reasoning traceable all the way back to the source event. The model doesn't just store what happened — it makes what happened operational.
- One definition, enforced everywhere. A "player" or a "performance" means the same thing in every report, every ranking, every model.
- Every source, one language. Film, tracking, and manual grades become comparable observations instead of incompatible files.
- Traceable intelligence. Every ranking and projection can be followed back to the underlying observations that produced it.
- Compounding value. Each new data source makes every existing question easier to answer, not harder.
This is what the StatLink Ontology is built to do
StatLink's advantage isn't another tracker or another dashboard. It's an ontology that maps fragmented sports data onto one standardized model, so analytics can actually run on top of it. At its center is the Trial — a single contextualized observation with an outcome. Every tagged outcome, from any source, becomes exactly one Trial, and trials aggregate upward into scorecards, rankings, and the intelligence that RAVEN reasons over.
It's the difference between owning data and being able to use it. The whole model — the accounts, teams, scorecards, metrics, trials, and the hierarchy that turns observations into rankings — is laid out object by object, with examples, on the ontology page.
See the StatLink Ontology
Explore the model that turns fragmented sports data into one connected system — the objects, the relationships, and the Trial at the center of it all.
Explore the StatLink Ontology →More data was never the answer. A shared model of what that data means is. That's the ontology advantage — and it's the foundation everything StatLink and IMPACTCAP are built on.