A code graph is not a software ontology
AI coding agents need more than source code and graph edges. They need a system model that preserves meaning, evidence, uncertainty and the boundaries of what it knows.
Recently, my car told me that two things were due: Ölwechsel and Fahrzeugcheck. An oil change and a vehicle check. Nothing particularly complicated.
I went to the workshop’s website and booked an Ölwechsel and a Mobilitätscheck. The names seemed to match what the car was asking for closely enough.
When I picked up the car, everything I had booked was done, but the computer was still warning me that the Fahrzeugcheck was overdue. I assumed the workshop had simply forgotten to reset it. No big deal. I was going to be back soon to change the tires anyway.
So when booking the next appointment, I left a comment asking them to reset the Fahrzeugcheck they had forgotten the previous time.
They called me five minutes later.
“You didn’t do a Fahrzeugcheck.”
Now I was confused. I gave them the date and explained what had been done. They found the appointment and confirmed it.
“Yes, but that’s not a Fahrzeugcheck.”
So I asked the obvious question: what exactly should I have booked?
“Inspektion.”
Fine. I opened the website and looked at Inspektion. Its description talked about changing the oil, filters and doing several other things, including work I had just paid them to do separately.
The person from the workshop looked at the same information and eventually admitted that, yes, this was confusing. He wasn’t entirely sure what I should book either. We agreed to figure it out when I brought the car in.
This is, in a surprisingly mundane way, an ontology problem.
There was no shortage of data. The car knew what it needed. The workshop knew which services it provided. The booking system had categories and descriptions. The employee had access to my previous appointment. What was missing was a shared model of what those things meant and how they related to each other.
The car had a concept called Fahrzeugcheck. The booking system had
Mobilitätscheck and Inspektion. An Inspektion apparently satisfied the
car’s requirement for a Fahrzeugcheck, while a Mobilitätscheck did not. At
the same time, an Inspektion included or overlapped with an Ölwechsel, which
was also available as a separate service.
The vocabulary existed. The relationship that mattered did not.
An AI agent could have made exactly the same mistake I did. Fahrzeugcheck to
Mobilitätscheck is a perfectly plausible linguistic match. Giving the model a
larger context window would not necessarily have fixed it. The workshop website,
my service history and the text from the car would have provided more data, but
none of them explicitly described the relationship the decision depended on.
Software systems have this problem at a much larger scale. Instead of four confusing service names, they contain hundreds of thousands of symbols, packages, APIs, events, databases, services and dependencies spread across many repositories. Now we are asking AI agents to work inside them.
What those agents need is not simply more code or a larger graph. They need a reproducible model that says what the system’s entities and relationships mean, where each claim came from, and where the model’s ability to reason stops.
From vocabulary to taxonomy to ontology
These words are used differently across disciplines, so it is useful to say what I mean by them here.
A vocabulary gives things stable names. A taxonomy organizes them into categories and hierarchies. An ontology goes further: it describes what those things mean, how they relate and which conclusions those relationships support.
In software, the vocabulary might include repositories, modules, symbols, routes, storage systems and dependencies. The ontology says that a module declares a symbol, a symbol calls another symbol, a route is handled by a function, a function writes to storage, or a service publishes an event consumed somewhere else.
Once those relationships have stable meanings, we no longer have only a catalogue of software objects. We have a model of the software system.
Real systems, however, are not written in one language, using one framework and living in one repository. A Go function, a Java method, a C++ function and a JavaScript function have different semantics. Frameworks introduce their own concepts on top. Systems communicate through HTTP, events, RPC, databases, shared libraries and combinations of all of them.
If the taxonomy simply mirrors every language and framework, we have several language-specific models sitting next to each other. Before building useful relationships across them, we need a shared conceptual contract.
A symbol should be something that makes sense independently of whether it came
from Go, C++, Ruby or JavaScript. The same applies to modules, routes, storage and
dependencies, and to relationships such as calls, implements, declares,
reads_from, writes_to, publishes and consumes.
That generalisation has limits. Generalise too little and every ecosystem
remains an island. Generalise too much and we end up with a beautifully universal
model consisting of thing -> relates_to -> thing, which tells us almost
nothing.
Finding the middle ground is part of designing the ontology. It is not something an abstract syntax tree or a language server can decide for us.
ASTs and language servers understand islands
This problem bothered me long before coding agents existed.
I have been a JetBrains user for years. Imagine a web application where the frontend is open in WebStorm and the backend is open in GoLand. Both tools understand their respective codebases remarkably well, but they have no intelligent way to agree with each other.
From my perspective, I am working on one software system. From the tooling’s perspective, I have two projects.
The frontend makes an HTTP request. Somewhere in the backend there is a route handling that request. From there the system might call several functions, access storage and publish an event. As a human, I understand that all of this belongs to one interaction. The tools understand their individual islands, but the system exists across those islands.
This distinction matters when abstract syntax trees and the Language Server Protocol are proposed as foundations for AI coding tools. An AST gives us a structured representation of code. A language server can add definitions, references, types, implementations and other semantic information. These are extremely valuable sources of facts.
But AST plus language-server data across ten repositories does not automatically produce an ontology of the system. It can produce ten very well-understood islands. The ontology requires connecting them.
The frontend request needs to be related to the backend route. An event producer needs to be related to its consumers. A database model needs to be related to the storage it represents. A generated client needs to be connected to its API. Framework concepts need to become concepts in the system rather than merely symbols or annotations in source code.
Some of those relationships can be observed directly. Some can be derived from other relationships. Some require framework-specific knowledge. ASTs and language servers are therefore not alternatives to a software ontology. They are observers from which parts of it can be constructed.
Adding more observers does not automatically create the model.
A code graph is not automatically an ontology
When coding agents started consuming enormous amounts of source code, graphs became an obvious response. Instead of repeatedly feeding raw files to the model, extract their structure, put it into a graph and retrieve the relevant pieces.
That can save tokens. It does not necessarily create understanding.
A graph can still effectively tell the agent:
Here are some nodes and edges that look relevant. You figure out what they mean.
The agent may still have to decide whether two concepts in different repositories correspond to each other, what a relationship implies, which paths matter and which conclusions can safely be drawn from the retrieved data. We have compressed the input without necessarily reducing the uncertainty.
A graph is a representation. An ontology supplies a contract for the meaning of what is represented. The distinction matters more than the choice of graph storage, as I also found when I compared four code-graph systems that made very different choices around superficially similar nodes and edges.
Deterministic extraction does not remove this distinction. An AST parser might deterministically identify a function, and a language server might deterministically resolve a reference. Those are deterministic observations. It does not follow that every semantic conclusion built from them is certain.
Consider dead code. An analyzer might establish that it found no incoming calls to a symbol:
observed incoming calls = 0
That is not the same statement as:
this code is dead
C++ macros, conditional compilation, templates, function pointers and dynamic linking complicate that conclusion. JavaScript presents different problems through dynamic imports, runtime property access and framework conventions. The observation can be reproducible while the conclusion remains conditional on what the analyzer could see.
A deterministic graph can still be deterministically incomplete.
A useful model represents what it does not know
Consider a cross-repository fixture where a web service calls an api
service. Static analysis detects three outbound HTTP calls. Two can be composed
and matched to routes served by the API. The third builds part of its path from
a tenant identifier available only at runtime:
http.Get(apiBase + "/api/v2/" + tenant + "/orders")
The path does not exist as one static string anywhere in the source. Guessing which route it means would produce an edge that looks as authoritative as the two real ones.
The useful result is therefore not merely two edges:
| Observation | Count |
|---|---|
| Outbound calls detected | 3 |
| Calls matched to a served route | 2 |
| Calls left unresolved | 1 |
The unresolved count is part of the model. It distinguishes “we found no relationship” from “there is no relationship.”
The same distinction matters when an analyzer reports no consumers of an API. It may have analyzed every relevant repository and found none. It may have skipped two repositories written in an unsupported language. It may have encountered dynamic behaviour that cannot be resolved statically. Or part of the system may not have been analyzed at all.
If every one of these situations produces the same empty graph, absence of evidence has silently become evidence of absence. For an agent, that can be worse than explicit uncertainty because the answer looks authoritative.
This is why confidence and coverage are different and both matter. Confidence describes how strongly a particular conclusion is supported by the available evidence. Coverage describes how much of the relevant world was available to reason about. We can have high confidence that no callers were found within the analyzed code while having poor coverage of all possible callers.
Not every claim should have the same epistemic status either. A dependency cycle either exists in the measured import graph or it does not. Calling a class a “god class” depends on a threshold or a comparison with the rest of the repository. Both conclusions can be computed reproducibly, but only one is a structural fact about the measured graph. Reasonable engineers can disagree about the other.
A useful ontology should distinguish what was observed, what was derived, what was inferred and what could not be determined. That distinction can be represented with labels, evidence, derivation paths or carefully defined confidence values. The representation matters less than refusing to turn all four into equivalent facts.
Reproducible means reproducible from stated inputs
The goal is not necessarily a deterministic ontology. A more useful goal is a reproducible ontology with explicit uncertainty and explicit knowledge boundaries.
Even that requires a stronger contract than it first appears to.
“Run it again against the same software system” is underspecified. Which commit? Was the working tree dirty? Which extractor versions ran? Which configuration and ignore rules were active? Which files failed to parse? Which repositories were not present?
A reproducible graph therefore needs a receipt recording the source revision, tooling, effective configuration, exclusions, parse failures and coverage. Its identity should include the graph, tool version and configuration. If two snapshots were produced with different extractors or ignore rules, their difference describes both a software change and a measurement change, so they should not be treated as directly comparable.
Any system that claims reproducibility needs to preserve the inputs that make its results reproducible. The ontology should be able to answer not only “what do you know?” but also “what did you inspect, how did you inspect it, and what did you leave out?”
Being precise about those limits is as valuable as being precise about the facts themselves.
An ontology should fit the questions it answers
There is no single ontology suitable for every domain.
A product-catalogue ontology might need to describe that an object is a shoe, that a running shoe is a type of athletic shoe, that it has a size and material, and that differently named products refer to the same commercial item. It might prioritise classification, equivalence and interoperability between organisations.
A software ontology is built to answer different questions. What calls this function? Which route reaches it? What storage does it modify? Which service consumes the event it publishes? What might be affected if its contract changes?
Those relationships form paths, and the paths often matter more than the individual entities. Software puts unusual weight on behaviour, reachability, dependency, change impact and time: the relationship that exists at one commit may not exist at the next.
The difference in purpose should shape the implementation. A shared catalogue may benefit from standardized vocabularies and the Semantic Web stack: RDF as a graph data model, RDFS or OWL for expressing richer semantics, and SPARQL for querying RDF data.
A software-analysis system with a bounded vocabulary and known questions may need none of them. Typed records, documented relation semantics, provenance, coverage and purpose-built graph traversals may be the better contract.
Calling a model an ontology does not commit us to a serialization format, query language or reasoning engine. Those are implementation choices. The ontology is the conceptual agreement: what kinds of things exist, what its relationships mean, which conclusions they support and where its knowledge ends.
Loading source entities into a generic knowledge graph is therefore not enough. But adopting an elaborate ontology stack is not proof that the resulting model understands software either.
The ontology is not the context
A software ontology can contain hundreds of thousands of entities and millions of relationships. That is fine. The mistake would be expecting the agent to consume all of it.
If an agent is changing one function, it might need to know that the function implements an interface, is called from three places, handles two HTTP routes, writes to a particular storage system and eventually causes an event consumed by another repository. It probably does not need to know about the other 200,000 symbols.
Change the task to modifying a database schema and another part of the ontology becomes relevant. Change an API contract and another part becomes relevant again. The engineering problem is extracting the small subgraph relevant to the task while preserving the semantics, evidence and coverage needed to interpret it correctly.
The same principle applies to adding more data. Git history, documentation, tickets, runtime traces, logs and database schemas can all enrich the model. Collecting them does not automatically improve it.
My car problem was not caused by having too little information about the
workshop’s services. Another ten paragraphs describing an Inspektion would not
necessarily have told me that it was the thing my car called Fahrzeugcheck.
One correct relationship would have been more valuable:
Inspektion -> satisfies -> Fahrzeugcheck
Knowing that another 50,000 symbols exist does not necessarily help an agent understand the consequences of changing an API. Knowing that three services in two other repositories consume it probably does.
The ontology is not the context. It is what allows us to construct the right context for the question in front of us.
From linters to explainers
LLMs are probabilistic systems. That is one of their strengths: they can reason about ambiguous problems where deterministic software struggles. It does not follow that we should make deterministic problems probabilistic.
If static analysis can establish that function A calls function B, the agent does not need to infer the relationship from two source files. If framework analysis can establish that function B handles an HTTP route, the model does not need to guess from an annotation. If cross-repository analysis can connect an event producer to its consumers, the agent should not reconstruct that relationship every time it works on the producer.
Every reliable relationship already represented removes one thing the model has to rediscover correctly. When a relationship cannot be established reliably, saying so is useful information too.
This changes what static analysis can provide. We have had excellent linters, compiler warnings and IDE inspections for decades, but most traditionally operate within a bounded context: a project, repository, module, language or particular analysis model.
Imagine a local analyzer reports that a function appears unused. That may be correct according to the relationships it can see. But perhaps the function is registered through a framework convention, implements an interface used dynamically, handles a route, or backs an API ultimately called from another repository.
Conversely, a dependency that looks reasonable inside one repository might introduce a relationship between architectural domains that were previously independent. The individual analyzer has not necessarily become smarter. The world around its finding has become richer.
This is why I increasingly prefer the idea of explainers to linters for this kind of analysis.
A linter tends to make a bounded statement: this code violates this rule. An explainer can say: this is what we observed, this is how this part of the system relates to the rest, this is the conclusion the evidence supports, and this is what we could not establish.
That output is useful to both humans and agents.
Architecture becomes shared data
Ask an engineer joining a large organisation what happens when a customer presses a particular button. They might start in a frontend repository, discover an API, jump into another repository, find a service, follow an event, search for its consumers and eventually discover where the data is stored.
The architecture was there all along. It was distributed across source code, build systems, framework conventions, infrastructure, documentation and people’s heads.
Once those relationships are represented explicitly, architecture becomes data. We can query it, diff it, ask what changed between two versions, define constraints against it, identify unexpected dependencies and calculate impact before changing something.
Humans and agents can use the same underlying model. A developer can ask what depends on a service before changing it. An agent can retrieve the same neighborhood before generating code. CI can determine whether a change introduced an architectural violation. The interfaces differ; the underlying representation does not have to.
This is why software ontologies become particularly interesting in the agentic era. Machines are actively modifying systems built and maintained by humans, continuously and at machine speed. They should not have to rediscover those systems from raw text every time they perform a task.
Nor do we need to respond by constructing the largest or most formally elaborate ontology possible. We need a shared taxonomy that survives languages and frameworks. We need meaningful relationships across the boundaries where the software actually operates. We need to distinguish observations from conclusions. We need confidence where certainty is not possible, and coverage that tells us where our ability to reason stops.
The complete ontology may contain millions of relationships. The agent should see the ones that matter. The human should be able to inspect the same relationships. And when the ontology does not know something, both should know that too.
Because knowing that you do not know whether Mobilitätscheck means
Fahrzeugcheck is considerably safer than confidently booking the wrong one.
Author’s note: I work on Enola, an open-source project that is one implementation of what this essay describes. The cross-repository fixture and snapshot receipt discussed above come from it and are public and reproducible. I did not write this essay because of Enola, though. It is a high-level reflection on a problem that needs solving, whichever tools end up solving it.