Rasterflow, Earth Intelligence & inference engine now in public preview Learn More

What is an Ontology? Definition, Parts, Spatial Ontologies

An ontology is a formal model of a domain: the kinds of things it contains, their properties, and how they relate, written precisely enough for software to reason over it. A spatial ontology adds what matters about location: geometry, coordinate reference systems, and relations such as contains, touches, and located in.

Key takeaways

  • An ontology defines classes, properties, relations, and rules for a domain, so people and software use the same meanings.
  • In information science and AI, ontology means a data model. In philosophy, it means the study of what exists.
  • A taxonomy is a hierarchy. An ontology adds other relations and logic. A knowledge graph is the data organized by an ontology.
  • A spatial ontology treats geometry as a first-class concept and defines topological relations, CRSs, scale, and time.
  • AI agents need typed entities, stable IDs, and explicit spatial relations to answer questions about real places.
  • Much of spatial analysis already runs on graphs, from spatial weights to road networks. A spatial ontology names their nodes and edges.

What is an ontology?

In philosophy, ontology is the branch of metaphysics that studies what exists, and most searches for the word come from that sense. Computer science borrowed the term for a model you can write down and share. Tom Gruber's widely cited 1993 definition calls it an "explicit specification of a conceptualization" (see his page about the definition of ontology), and later work added "formal" and "shared": agreed on by a community and written in a language a machine can parse.

The practical definition in Stanford's Ontology Development 101 guide, by Natalya Noy and Deborah McGuinness, is the one most teams work from: a formal explicit description of the concepts in a domain (classes), the properties of each concept (slots), and restrictions on those properties (facets).

Why ontology is back in AI

Language models were trained on text, documents, databases, and the internet. They produce fluent answers, and they have no built-in agreement about what an "asset", a "site", or a "customer" means in a given company. Ontologies supply that agreement. Microsoft describes the ontology item in Fabric IQ as a shared, machine-understandable representation of a business, bound to enterprise data so that people, applications, and AI agents use the same vocabulary. Palantir, Salesforce, and Databricks use the word in the same sense.

The parts of an ontology

Diagram of an ontology and a knowledge graph. The ontology layer has the classes Place, Building, and Division with their properties, Bakery as a subclass of Place, the relations located_in and contains, and two rules. The knowledge graph layer has Acme Bread Company located_in the San Francisco Ferry Building, San Francisco containing the Ferry Building, and an inferred link placing Acme Bread Company in San Francisco
The ontology defines classes, relations, and rules. The knowledge graph holds the individuals. The green link is inferred from a rule.

Classes

Classes are the kinds of things in the domain: Place, Building, Address, Road segment, Administrative division. Classes form hierarchies. A Bakery is a subclass of Place, so every bakery is also a place.

Properties and relations

Datatype properties attach values, such as a building's height in meters. Object properties connect two things, such as located_in from a Place to a Building. Each relation has a domain (what it starts from) and a range (what it points to), and those constraints let software catch data that makes no sense.

Individuals

Individuals, or instances, are the actual things: the San Francisco Ferry Building, a specific bakery, a specific street address. The ontology defines the classes. The instances live in the data.

Axioms and rules

Axioms state what must be true. Two classes can be disjoint (nothing is both a road and a building). A relation can be transitive: if a bakery is inside the Ferry Building and the Ferry Building is inside San Francisco, the bakery is inside San Francisco. Reasoners use axioms to check consistency and infer facts that were never written down.

How ontologies are written

The W3C standards stack is the common format:

  • RDF stores facts as subject, predicate, object triples.
  • RDF Schema adds classes, subclasses, domains, and ranges.
  • OWL 2, a W3C Recommendation since 2012, adds richer logic such as disjointness, cardinality, and property characteristics.
  • SPARQL queries RDF data.

Many teams also express an ontology as typed tables and edge lists in a lakehouse, which is how the example later in this article is stored.

Ontology vs taxonomy vs schema vs knowledge graph

TaxonomySchemaOntologyKnowledge graph
What it holdsA hierarchy of termsTable or document structureClasses, relations, properties, rulesEntities and links (instance data)
RelationsBroader and narrower onlyForeign keysAny named relationEdges typed by the ontology
Meaning and inferenceLittleStorage constraintsLogic a reasoner can useInherits the ontology's meaning
Typical technologyCategory lists, SKOSSQL DDL, JSON SchemaRDF, OWL, typed edge tablesGraph databases, edge tables
Spatial exampleBuilding, then Commercial, then OfficeA buildings table with a geometry columnPlace located_in BuildingAcme Bread Company located_in Ferry Building

A semantic layer, a term common in business intelligence, overlaps with an ontology: both name business concepts consistently. A semantic layer usually centers on metrics and joins for reporting. An ontology covers entity types, relations, and rules more broadly.

What makes an ontology spatial?

A general ontology can say that a store is in a city. A spatial ontology can say where, with what geometry, in which coordinate system, and in what topological relation to everything around it.

Geometry as a first-class concept

OGC GeoSPARQL defines an RDF vocabulary for geospatial data and an extension to SPARQL for querying it. Its core separates a Feature (the thing, such as a building) from its Geometry (the shape), linked by properties such as geo:hasGeometry, geo:hasCentroid, and geo:hasBoundingBox. One feature can have several geometries, such as a point and a footprint. The current version, GeoSPARQL 1.1, was published on 29 January 2024.

Topological relations

Spatial ontologies define how regions relate, and those relations carry logic a reasoner can use. GeoSPARQL includes three families and states how they correspond:

  • Simple Features relations: equals, disjoint, intersects, touches, crosses, within, contains, and overlaps. These are the same predicates a spatial join uses in SQL, such as ST_Contains and ST_Intersects.
  • Egenhofer relations: based on Max Egenhofer's intersection models, which compare the interiors, boundaries, and exteriors of two geometries. The dimensionally extended version (DE-9IM) is what ST_Relate returns.
  • RCC8: the Region Connection Calculus (Randell, Cui, and Cohn, 1992) defines eight relations between regions: disconnected (DC), externally connected (EC), equal (EQ), partially overlapping (PO), tangential proper part (TPP), non-tangential proper part (NTPP), and the inverses TPPi and NTPPi. Its composition table says which relations can hold between A and C given the relations from A to B and from B to C.

GeoSPARQL's query rewrite extension goes one step further: it derives a relation between two features, such as geo:sfIntersects, from a geometry function applied to their shapes. The geometry implies the relation, so no one has to assert it.

Coordinate reference systems

A geometry has no location without a coordinate reference system. A GeoSPARQL WKT literal can start with a CRS identifier, and when it does not, GeoSPARQL 1.1 assumes WGS 84 longitude and latitude (CRS84). Mixing CRSs breaks both topology and distance.

Scale and hierarchy

A city is a point on a world map and a polygon on a regional one. Spatial ontologies handle this with multiple geometries per feature and with containment hierarchies. Overture's divisions theme runs from country down to microhood, each division inside its parent, and the GeoNames ontology models the same hierarchy with parentFeature.

Time

Buildings are demolished, roads are rerouted, and boundaries are redrawn. W3C OWL-Time provides classes for instants and intervals, so facts can carry the period when they were true.

Stable identity: Overture GERS

Relations only stay useful if every dataset refers to the same thing by the same ID. Overture's Global Entity Reference System (GERS) assigns a stable identifier to each feature in its reference map, so anyone who matches their data to a feature once can join on the identifier afterward. Wherobots shows how in How to use GERS IDs in Wherobots.

Why spatial ontologies matter for AI agents

Language models hold a great deal of geographic knowledge, and they still struggle with raw coordinates. The GeoLLM paper (ICLR 2024) found that naively querying language models using geographic coordinates alone is ineffective. An agent working on the physical world needs entities with types, stable IDs, geometry, and explicit relations it can query, which is what a spatial ontology and its knowledge graph provide. This grounding is the core of physical AI and geospatial AI.

Spatial knowledge graphs and Graph RAG

Knowledge graphs built from text capture what someone wrote down. Many of the relations that matter in the physical world are never written: which buildings sit inside a flood zone, which stores share a block, which address a delivery route reaches. A spatial knowledge graph computes those edges from geometry. Each spatial predicate in a spatial join produces a typed relation: ST_Contains gives located_in, ST_Touches gives adjacent_to, and ST_DWithin or ST_KNN gives near. Wherobots describes the approach in Spatial Graph RAG for the physical world. Graph RAG then retrieves a relevant subgraph as context for a model.

The Overture Maps Foundation calls the cost of rebuilding these links in every project the "conflation tax", and it is testing a shared answer. In Grounding AI and LLMs with Overture's cross-theme knowledge graph, Overture describes ORATOR (Overture Maps Foundation Knowledge Graph), a prototype built by Overture member Wherobots on the principle that geometry is the foreign key: when a place falls inside a building polygon, a spatial predicate turns that fact into an edge. Three design choices in the prototype are worth copying in any spatial ontology:

  • Confidence on every edge. Strict containment scores 1.0 and weaker spatial evidence scores lower. Address linking runs containment and street-name matching first at 1.0, then a proximity fallback scored 0.6 to 0.95.
  • Provenance on every edge, so a model can trace why two entities were connected.
  • Inherited relations. A place inherits the road access of the building that contains it, which cut San Francisco access edges from a possible 610,000 to about 172,000.

Overture is asking its community whether cross-theme relations belong in the core schema or a downstream product, and how confidence and provenance should be standardized.

Ontologies in Wherobots

Wherobots publishes the ORATOR prototype graph in its open data catalog as wherobots_open_data.spatial_knowledge_graph, with node and edge tables for San Francisco (nodes_sf, edges_sf) and Manhattan (nodes_manhattan, edges_manhattan). Nodes taken from Overture use GERS IDs, so they join back to Overture data in the Havasu catalog. The San Francisco tables hold 170,965 buildings, 55,134 places, 387,872 addresses, 83,765 road connectors, and 102 divisions from Overture, plus 16,151 snap points the graph adds itself: points on the road network that buildings and addresses connect to, with graph-generated IDs instead of GERS IDs. The header image above shows the Ferry Building and its neighbourhood in this graph.

These are the main relations in the San Francisco edges table, with the node types each one connects, as returned by a query run in Wherobots. Together they are the ontology's domain and range, read straight from the data:

RelationFromToEdges
has_addressBuildingAddress387,826
nearBuildingBuilding301,433
accessBuildingSnap point144,320
located_inPlaceBuilding54,675
has_addressPlaceAddress53,308
roadConnectorConnector48,800
same_asPlacePlace11,839
containsDivisionAddress, building, connector, place4,330,832

A two-hop question shows why typed relations help. This query finds the San Francisco Ferry Building and counts the places located in it, by category:

SELECT p.category, COUNT(*) AS places
FROM wherobots_open_data.spatial_knowledge_graph.nodes_sf b
JOIN wherobots_open_data.spatial_knowledge_graph.edges_sf e
  ON e.dst = b.id AND e.type = 'located_in'
JOIN wherobots_open_data.spatial_knowledge_graph.nodes_sf p
  ON p.id = e.src AND p.type = 'place'
WHERE b.type = 'building' AND b.name = 'San Francisco Ferry Building'
GROUP BY p.category
ORDER BY places DESC
LIMIT 5
categoryplaces
bakery8
farmers_market7
(none)5
coffee_shop5
food_truck4

The building has 112 places located in it in total, and those edges carry different levels of evidence. Each edge stores how it was made and how sure the match is in its properties column, for example {"method":"proximity","confidence":0.956}. For the Ferry Building, 86 places fall inside the building polygon (containment, confidence 1.0), and 26 were matched by proximity, with confidence from 0.554 to 1.0. Across San Francisco, 36,351 located_in edges come from containment and 18,324 from proximity. A query or an agent can filter on that confidence instead of treating every edge as certain.

Network diagram of the San Francisco Ferry Building in a spatial knowledge graph: 29 places in five categories linked to the building by located_in edges, same_as links between duplicate places, a neighbouring building linked by near, access edges from both buildings to one snap point on a primary road, and United States, California and San Francisco linked by contains
The Ferry Building and its neighbourhood in the San Francisco graph. Dashed located_in edges are proximity matches, labelled with their confidence.

The same pattern answers questions such as which addresses a building has, which buildings sit near it, and how it connects to the road network. With the Wherobots MCP server, an AI coding tool can find these tables in the catalog and run traversals like this one in plain language.

Query the graph with GraphFrames

Every extra hop in SQL adds another join. GraphFrames, the DataFrame-based graph library for Apache Spark, writes the hops as a path pattern instead, and it runs in the same Wherobots session as the spatial SQL that built the graph. The knowledge graph tables already use the column names GraphFrames expects, id on nodes and src and dst on edges, so they load as a graph without any renaming:

from graphframes import GraphFrame
from pyspark.sql import functions as F
from sedona.spark import SedonaContext

config = SedonaContext.builder().getOrCreate()
sedona = SedonaContext.create(config)

skg = "wherobots_open_data.spatial_knowledge_graph"
g = GraphFrame(sedona.table(f"{skg}.nodes_sf"), sedona.table(f"{skg}.edges_sf"))

# One hop: places located in the Ferry Building, by category
ferry = g.find("(p)-[e]->(b)").filter(
    "e.type = 'located_in' AND b.name = 'San Francisco Ferry Building'"
)
ferry.groupBy("p.category").count().orderBy(F.desc("count")).show(5)

# Two hops: from each place, through its building, to the road network
reach = g.find("(p)-[e1]->(b); (b)-[e2]->(s)").filter(
    "e1.type = 'located_in' AND e2.type = 'access' "
    "AND b.name = 'San Francisco Ferry Building'"
)
reach.groupBy("s.id").count().show()  # one snap point, 112 paths

The motif (p)-[e1]->(b); (b)-[e2]->(s) reads as the question itself: a place located in a building that has road access. For the Ferry Building, every one of its 112 places resolves to the same snap point, about 60 meters away on a primary road, which is ORATOR's inherited access applied in a query. GraphFrames also provides connected components, shortest paths, and PageRank over the same tables.

To build your own graph, start from Overture tables and compute edges with spatial joins: located_in from ST_Contains, near from ST_KNN or ST_DWithin, and containment from the divisions hierarchy. The wkls library gives quick access to the division hierarchy by name.

Graphs and spatial analysis

Graphs and spatial analysis grew up together. Many methods in geospatial analysis already compute over a graph of neighbors, and a spatial ontology gives that graph typed nodes and named relations.

Spatial statistics run on neighbor graphs

Tobler's first law of geography (1970) says that near things are more related than distant things. Spatial statistics put that idea to work with a spatial weights matrix, which records which observations count as neighbors and how strongly. That matrix is the adjacency matrix of a graph: rook and queen contiguity link polygons that share an edge or a corner, distance bands link everything within a threshold, and k-nearest neighbors links each point to its closest k.

Map of San Francisco census tracts drawn as a graph: each tract is a node, 641 links join tracts that share an edge and 98 dashed links join tracts that touch only at a corner, with one downtown tract and its 12 neighbours highlighted
San Francisco’s census tracts as a contiguity graph, from an ST_Intersects self-join in Wherobots: 240 tracts and 739 neighbour pairs.

Moran's I (1950) and Geary's C (1954) measure spatial autocorrelation over that graph. Luc Anselin's Local Indicators of Spatial Association (1995) compute one value per node from its neighbors, which is how hot spot and outlier maps are made, and spatial regression models put the same matrix inside a regression. PySAL's libpysal Graph class is documented as encoding spatial weights matrices. In Wherobots, ST_BinaryDistanceBandColumn and ST_WeightedDistanceBandColumn build distance band neighbor lists, and ST_GLocal computes Getis-Ord Gi and Gi* hot spot statistics over them.

Graph queries with spatial predicates

Once a graph holds locations, many questions mix a traversal with a spatial filter, such as which entities linked to a building lie within a region. Graph databases are built for traversals and spatial databases for spatial filters, and running both efficiently in one query is hard. Research by Yuhan Sun and Mo Sarwat at Arizona State University's Data Systems Lab works on that gap:

  • GeoReach (2016) answers reachability queries with a spatial range predicate, such as whether a vertex can reach any vertex inside a region, by storing compact spatial reachability information with each vertex.
  • GeoExpand (GeoInformatica, 2019) adds a spatially pruned expansion operator to the Neo4j graph database.
  • Riso-Tree (ACM Transactions on Spatial Algorithms and Systems, 2021) indexes the spatial entities in a graph database so that queries combining graph patterns and spatial predicates prune the traversal early, evaluated on Wikidata.
  • Spindra (Sun, Jia Yu, and Sarwat, ICDE 2019) packages these ideas as a geographic knowledge graph management system with a map interface.

Computing the spatial relations at scale is the other half. Jia Yu, Jinxuan Wu, and Mo Sarwat introduced GeoSpark (ACM SIGSPATIAL 2015), and Yu, Zongsi Zhang, and Sarwat described its spatial partitioning, indexing, and spatial join design in Spatial data management in Apache Spark (GeoInformatica, 2019). GeoSpark became Apache Sedona, Yu and Sarwat went on to found Wherobots, and the distributed spatial joins that produce a graph like ORATOR come from that work.

Roads are graphs

Road networks are the most studied spatial graphs. Marc Barthélemy's review Spatial networks (Physics Reports, 2011) describes them as close to planar, Porta, Crucitti, and Latora's primal approach (2006) measures centrality with intersections as nodes and street segments as edges, and Geoff Boeing's OSMnx builds these graphs from OpenStreetMap.

Map of 5-minute and 10-minute drive-time areas from the San Francisco Ferry Building computed with ST_Isochrone, extending along freeways to Treasure Island and across the Bay Bridge to Oakland, compared with a dashed straight-line circle that covers open bay
Drive-time areas from the Ferry Building, computed on the road network with ST_Isochrone in Wherobots: 28.5 km² in 5 minutes and 173 km² in 10.

Distance along the network often differs from distance on the map, and Okabe and Sugihara's Spatial Analysis Along Networks (Wiley, 2012) adapts density and clustering statistics to road distance. An isochrone is the area reachable within a travel time, and Wherobots computes it with ST_Isochrone. In the ORATOR graph, the access edges to snap points connect the place graph to the road graph.

Grids are graphs too

Discrete global grids turn the planet into a graph of cells. In H3, gridDisk returns every cell within k steps and gridDistance returns the number of steps between two cells; Wherobots exposes these as ST_H3KRing and ST_H3CellDistance. Cells also work as a shared join key: KnowWhereGraph precomputes topological relations between its features and S2 cells to integrate data.

Conflation is a graph problem

Deciding which records describe the same real-world place is conflation, and it is usually solved on a graph. Records become nodes, candidate matches scored by distance, name, and category become edges, and connected components group the records that refer to one place. He, Li, and Zhang (2024) fuse semantic and spatial signals for this matching. GraphFrames provides connected components at Spark scale, and Wherobots adds spatial versions such as ST_IntersectsCC and ST_DWithinCC. The same_as edges in the Ferry Building figure are this step: candidate links that conflation can confirm or reject.

Machine learning on spatial graphs

Graph neural networks learn from a node's neighbors, which suits spatial data. DCRNN (Li et al., 2018) and STGCN (Yu, Yin, and Zhu, 2018) forecast traffic on graphs of road sensors, and Google reported ETA prediction with graph neural networks in production in Google Maps (CIKM 2021). Knowledge graph embeddings can be location-aware too: SE-KGE (Mai et al., 2020) encodes each entity's location and extent for geographic question answering. Janowicz and colleagues call this direction spatially explicit AI.

Geographic knowledge graphs

ORATOR joins a line of geographic knowledge graphs, including YAGO2, WorldKG (built from OpenStreetMap), KnowWhereGraph, and UUKG for urban prediction tasks. ORATOR's nodes are Overture features with stable GERS IDs, so the graph joins back to the full feature tables.

What spatial context adds to an ontology

  • Edges no one wrote down. Spatial predicates compute relations from geometry at any scale, each with its method and confidence.
  • Inference. Containment chains across levels, and the RCC8 composition table limits which relations can hold between two regions linked through a third.
  • Distance and connectivity. Neighbor graphs carry spatial dependence into statistics and models, and road graphs measure distance the way people travel.
  • Identity. Conflation over a candidate graph turns duplicate records into one entity with a stable ID.

Read more from Wherobots

Query the spatial knowledge graph with a Wherobots free trial at cloud.wherobots.com.

Frequently asked questions

What is an ontology in simple terms?

An ontology is a shared, written-down model of a subject: the kinds of things it contains, the properties those things have, and the ways they relate to each other. It is written precisely enough that software can check data against it and draw conclusions from it, for example that if a restaurant is inside a building and the building is inside a city, the restaurant is inside that city too.

What is ontology in AI?

In AI, an ontology gives models and agents a fixed vocabulary for a domain: which entity types exist, which relations connect them, and which rules hold. Agents use it to interpret data consistently, to plan queries, and to ground answers in real entities. Vendors such as Microsoft describe ontologies as a machine-understandable representation of a business that people, applications, and AI agents share.

What does Palantir mean by ontology?

Palantir uses ontology for the layer in its software that connects an organization’s data to the real-world things that data describes, such as customers, assets, or orders, and the links between them. It is the information science sense of the word, applied to one company’s operational data.

What is an example of an ontology?

A small spatial example: the classes Place, Building, Address, and Road segment; the relations located_in (Place to Building), has_address (Building to Address), and access (Building to the road network); and a rule that located_in is transitive. Public examples include schema.org, GeoNames, OGC GeoSPARQL, and the Gene Ontology.

What is the difference between an ontology and a taxonomy?

A taxonomy arranges terms in a hierarchy of broader and narrower categories, such as Building, then Commercial building, then Office. An ontology includes that hierarchy and adds other relations between concepts, properties with types, and logical rules. Every taxonomy can be part of an ontology, but a taxonomy alone cannot say that a place is located in a building.

Is a knowledge graph an ontology?

Not exactly. An ontology defines the types and relations, and a knowledge graph is the data: the actual entities and the links between them, organized by an ontology. In a spatial knowledge graph, the ontology says that places can be located in buildings, and the graph records that the Acme Bread Company is located in the San Francisco Ferry Building.

How are graphs used in spatial analysis?

Many spatial methods compute over a graph of neighbors. Spatial statistics such as Moran’s I and local hot spot measures use a spatial weights matrix, which is the adjacency matrix of a neighbor graph. Road networks are graphs used for routing, isochrones, and centrality. Spatial knowledge graphs compute edges such as located_in and near from geometry, and graph neural networks learn from those neighbor relations for tasks such as traffic forecasting.

What is GeoSPARQL?

GeoSPARQL is an OGC standard for geospatial data on the Semantic Web. It defines an RDF vocabulary for features and geometries, such as geo:Feature, geo:Geometry, and geo:hasGeometry, plus spatial relation properties and functions that extend the SPARQL query language. Version 1.1 was published in January 2024.