Planetary-scale answers, unlocked.
A Hands-On Guide for Working with Large-Scale Spatial Data. Learn more.
Authors
An ontology is a formal model of a domain: the kinds of things it contains, their properties, and how they relate, written precisely enough for software to reason over it. A spatial ontology adds what matters about location: geometry, coordinate reference systems, and relations such as contains, touches, and located in.
In philosophy, ontology is the branch of metaphysics that studies what exists, and most searches for the word come from that sense. Computer science borrowed the term for a model you can write down and share. Tom Gruber's widely cited 1993 definition calls it an "explicit specification of a conceptualization" (see his page about the definition of ontology), and later work added "formal" and "shared": agreed on by a community and written in a language a machine can parse.
The practical definition in Stanford's Ontology Development 101 guide, by Natalya Noy and Deborah McGuinness, is the one most teams work from: a formal explicit description of the concepts in a domain (classes), the properties of each concept (slots), and restrictions on those properties (facets).
Language models were trained on text, documents, databases, and the internet. They produce fluent answers, and they have no built-in agreement about what an "asset", a "site", or a "customer" means in a given company. Ontologies supply that agreement. Microsoft describes the ontology item in Fabric IQ as a shared, machine-understandable representation of a business, bound to enterprise data so that people, applications, and AI agents use the same vocabulary. Palantir, Salesforce, and Databricks use the word in the same sense.
Classes are the kinds of things in the domain: Place, Building, Address, Road segment, Administrative division. Classes form hierarchies. A Bakery is a subclass of Place, so every bakery is also a place.
Datatype properties attach values, such as a building's height in meters. Object properties connect two things, such as located_in from a Place to a Building. Each relation has a domain (what it starts from) and a range (what it points to), and those constraints let software catch data that makes no sense.
Individuals, or instances, are the actual things: the San Francisco Ferry Building, a specific bakery, a specific street address. The ontology defines the classes. The instances live in the data.
Axioms state what must be true. Two classes can be disjoint (nothing is both a road and a building). A relation can be transitive: if a bakery is inside the Ferry Building and the Ferry Building is inside San Francisco, the bakery is inside San Francisco. Reasoners use axioms to check consistency and infer facts that were never written down.
The W3C standards stack is the common format:
Many teams also express an ontology as typed tables and edge lists in a lakehouse, which is how the example later in this article is stored.
A semantic layer, a term common in business intelligence, overlaps with an ontology: both name business concepts consistently. A semantic layer usually centers on metrics and joins for reporting. An ontology covers entity types, relations, and rules more broadly.
A general ontology can say that a store is in a city. A spatial ontology can say where, with what geometry, in which coordinate system, and in what topological relation to everything around it.
OGC GeoSPARQL defines an RDF vocabulary for geospatial data and an extension to SPARQL for querying it. Its core separates a Feature (the thing, such as a building) from its Geometry (the shape), linked by properties such as geo:hasGeometry, geo:hasCentroid, and geo:hasBoundingBox. One feature can have several geometries, such as a point and a footprint. The current version, GeoSPARQL 1.1, was published on 29 January 2024.
Spatial ontologies define how regions relate, and those relations carry logic a reasoner can use. GeoSPARQL includes three families and states how they correspond:
GeoSPARQL's query rewrite extension goes one step further: it derives a relation between two features, such as geo:sfIntersects, from a geometry function applied to their shapes. The geometry implies the relation, so no one has to assert it.
A geometry has no location without a coordinate reference system. A GeoSPARQL WKT literal can start with a CRS identifier, and when it does not, GeoSPARQL 1.1 assumes WGS 84 longitude and latitude (CRS84). Mixing CRSs breaks both topology and distance.
A city is a point on a world map and a polygon on a regional one. Spatial ontologies handle this with multiple geometries per feature and with containment hierarchies. Overture's divisions theme runs from country down to microhood, each division inside its parent, and the GeoNames ontology models the same hierarchy with parentFeature.
Buildings are demolished, roads are rerouted, and boundaries are redrawn. W3C OWL-Time provides classes for instants and intervals, so facts can carry the period when they were true.
Relations only stay useful if every dataset refers to the same thing by the same ID. Overture's Global Entity Reference System (GERS) assigns a stable identifier to each feature in its reference map, so anyone who matches their data to a feature once can join on the identifier afterward. Wherobots shows how in How to use GERS IDs in Wherobots.
Language models hold a great deal of geographic knowledge, and they still struggle with raw coordinates. The GeoLLM paper (ICLR 2024) found that naively querying language models using geographic coordinates alone is ineffective. An agent working on the physical world needs entities with types, stable IDs, geometry, and explicit relations it can query, which is what a spatial ontology and its knowledge graph provide. This grounding is the core of physical AI and geospatial AI.
Knowledge graphs built from text capture what someone wrote down. Many of the relations that matter in the physical world are never written: which buildings sit inside a flood zone, which stores share a block, which address a delivery route reaches. A spatial knowledge graph computes those edges from geometry. Each spatial predicate in a spatial join produces a typed relation: ST_Contains gives located_in, ST_Touches gives adjacent_to, and ST_DWithin or ST_KNN gives near. Wherobots describes the approach in Spatial Graph RAG for the physical world. Graph RAG then retrieves a relevant subgraph as context for a model.
The Overture Maps Foundation calls the cost of rebuilding these links in every project the "conflation tax", and it is testing a shared answer. In Grounding AI and LLMs with Overture's cross-theme knowledge graph, Overture describes ORATOR (Overture Maps Foundation Knowledge Graph), a prototype built by Overture member Wherobots on the principle that geometry is the foreign key: when a place falls inside a building polygon, a spatial predicate turns that fact into an edge. Three design choices in the prototype are worth copying in any spatial ontology:
Overture is asking its community whether cross-theme relations belong in the core schema or a downstream product, and how confidence and provenance should be standardized.
Wherobots publishes the ORATOR prototype graph in its open data catalog as wherobots_open_data.spatial_knowledge_graph, with node and edge tables for San Francisco (nodes_sf, edges_sf) and Manhattan (nodes_manhattan, edges_manhattan). Nodes taken from Overture use GERS IDs, so they join back to Overture data in the Havasu catalog. The San Francisco tables hold 170,965 buildings, 55,134 places, 387,872 addresses, 83,765 road connectors, and 102 divisions from Overture, plus 16,151 snap points the graph adds itself: points on the road network that buildings and addresses connect to, with graph-generated IDs instead of GERS IDs. The header image above shows the Ferry Building and its neighbourhood in this graph.
wherobots_open_data.spatial_knowledge_graph
nodes_sf
edges_sf
nodes_manhattan
edges_manhattan
These are the main relations in the San Francisco edges table, with the node types each one connects, as returned by a query run in Wherobots. Together they are the ontology's domain and range, read straight from the data:
A two-hop question shows why typed relations help. This query finds the San Francisco Ferry Building and counts the places located in it, by category:
SELECT p.category, COUNT(*) AS places FROM wherobots_open_data.spatial_knowledge_graph.nodes_sf b JOIN wherobots_open_data.spatial_knowledge_graph.edges_sf e ON e.dst = b.id AND e.type = 'located_in' JOIN wherobots_open_data.spatial_knowledge_graph.nodes_sf p ON p.id = e.src AND p.type = 'place' WHERE b.type = 'building' AND b.name = 'San Francisco Ferry Building' GROUP BY p.category ORDER BY places DESC LIMIT 5
The building has 112 places located in it in total, and those edges carry different levels of evidence. Each edge stores how it was made and how sure the match is in its properties column, for example {"method":"proximity","confidence":0.956}. For the Ferry Building, 86 places fall inside the building polygon (containment, confidence 1.0), and 26 were matched by proximity, with confidence from 0.554 to 1.0. Across San Francisco, 36,351 located_in edges come from containment and 18,324 from proximity. A query or an agent can filter on that confidence instead of treating every edge as certain.
properties
{"method":"proximity","confidence":0.956}
The same pattern answers questions such as which addresses a building has, which buildings sit near it, and how it connects to the road network. With the Wherobots MCP server, an AI coding tool can find these tables in the catalog and run traversals like this one in plain language.
Every extra hop in SQL adds another join. GraphFrames, the DataFrame-based graph library for Apache Spark, writes the hops as a path pattern instead, and it runs in the same Wherobots session as the spatial SQL that built the graph. The knowledge graph tables already use the column names GraphFrames expects, id on nodes and src and dst on edges, so they load as a graph without any renaming:
id
src
dst
from graphframes import GraphFrame from pyspark.sql import functions as F from sedona.spark import SedonaContext config = SedonaContext.builder().getOrCreate() sedona = SedonaContext.create(config) skg = "wherobots_open_data.spatial_knowledge_graph" g = GraphFrame(sedona.table(f"{skg}.nodes_sf"), sedona.table(f"{skg}.edges_sf")) # One hop: places located in the Ferry Building, by category ferry = g.find("(p)-[e]->(b)").filter( "e.type = 'located_in' AND b.name = 'San Francisco Ferry Building'" ) ferry.groupBy("p.category").count().orderBy(F.desc("count")).show(5) # Two hops: from each place, through its building, to the road network reach = g.find("(p)-[e1]->(b); (b)-[e2]->(s)").filter( "e1.type = 'located_in' AND e2.type = 'access' " "AND b.name = 'San Francisco Ferry Building'" ) reach.groupBy("s.id").count().show() # one snap point, 112 paths
The motif (p)-[e1]->(b); (b)-[e2]->(s) reads as the question itself: a place located in a building that has road access. For the Ferry Building, every one of its 112 places resolves to the same snap point, about 60 meters away on a primary road, which is ORATOR's inherited access applied in a query. GraphFrames also provides connected components, shortest paths, and PageRank over the same tables.
(p)-[e1]->(b); (b)-[e2]->(s)
To build your own graph, start from Overture tables and compute edges with spatial joins: located_in from ST_Contains, near from ST_KNN or ST_DWithin, and containment from the divisions hierarchy. The wkls library gives quick access to the division hierarchy by name.
Graphs and spatial analysis grew up together. Many methods in geospatial analysis already compute over a graph of neighbors, and a spatial ontology gives that graph typed nodes and named relations.
Tobler's first law of geography (1970) says that near things are more related than distant things. Spatial statistics put that idea to work with a spatial weights matrix, which records which observations count as neighbors and how strongly. That matrix is the adjacency matrix of a graph: rook and queen contiguity link polygons that share an edge or a corner, distance bands link everything within a threshold, and k-nearest neighbors links each point to its closest k.
Moran's I (1950) and Geary's C (1954) measure spatial autocorrelation over that graph. Luc Anselin's Local Indicators of Spatial Association (1995) compute one value per node from its neighbors, which is how hot spot and outlier maps are made, and spatial regression models put the same matrix inside a regression. PySAL's libpysal Graph class is documented as encoding spatial weights matrices. In Wherobots, ST_BinaryDistanceBandColumn and ST_WeightedDistanceBandColumn build distance band neighbor lists, and ST_GLocal computes Getis-Ord Gi and Gi* hot spot statistics over them.
Graph
ST_BinaryDistanceBandColumn
ST_WeightedDistanceBandColumn
ST_GLocal
Once a graph holds locations, many questions mix a traversal with a spatial filter, such as which entities linked to a building lie within a region. Graph databases are built for traversals and spatial databases for spatial filters, and running both efficiently in one query is hard. Research by Yuhan Sun and Mo Sarwat at Arizona State University's Data Systems Lab works on that gap:
Computing the spatial relations at scale is the other half. Jia Yu, Jinxuan Wu, and Mo Sarwat introduced GeoSpark (ACM SIGSPATIAL 2015), and Yu, Zongsi Zhang, and Sarwat described its spatial partitioning, indexing, and spatial join design in Spatial data management in Apache Spark (GeoInformatica, 2019). GeoSpark became Apache Sedona, Yu and Sarwat went on to found Wherobots, and the distributed spatial joins that produce a graph like ORATOR come from that work.
Road networks are the most studied spatial graphs. Marc Barthélemy's review Spatial networks (Physics Reports, 2011) describes them as close to planar, Porta, Crucitti, and Latora's primal approach (2006) measures centrality with intersections as nodes and street segments as edges, and Geoff Boeing's OSMnx builds these graphs from OpenStreetMap.
Distance along the network often differs from distance on the map, and Okabe and Sugihara's Spatial Analysis Along Networks (Wiley, 2012) adapts density and clustering statistics to road distance. An isochrone is the area reachable within a travel time, and Wherobots computes it with ST_Isochrone. In the ORATOR graph, the access edges to snap points connect the place graph to the road graph.
Discrete global grids turn the planet into a graph of cells. In H3, gridDisk returns every cell within k steps and gridDistance returns the number of steps between two cells; Wherobots exposes these as ST_H3KRing and ST_H3CellDistance. Cells also work as a shared join key: KnowWhereGraph precomputes topological relations between its features and S2 cells to integrate data.
Deciding which records describe the same real-world place is conflation, and it is usually solved on a graph. Records become nodes, candidate matches scored by distance, name, and category become edges, and connected components group the records that refer to one place. He, Li, and Zhang (2024) fuse semantic and spatial signals for this matching. GraphFrames provides connected components at Spark scale, and Wherobots adds spatial versions such as ST_IntersectsCC and ST_DWithinCC. The same_as edges in the Ferry Building figure are this step: candidate links that conflation can confirm or reject.
Graph neural networks learn from a node's neighbors, which suits spatial data. DCRNN (Li et al., 2018) and STGCN (Yu, Yin, and Zhu, 2018) forecast traffic on graphs of road sensors, and Google reported ETA prediction with graph neural networks in production in Google Maps (CIKM 2021). Knowledge graph embeddings can be location-aware too: SE-KGE (Mai et al., 2020) encodes each entity's location and extent for geographic question answering. Janowicz and colleagues call this direction spatially explicit AI.
ORATOR joins a line of geographic knowledge graphs, including YAGO2, WorldKG (built from OpenStreetMap), KnowWhereGraph, and UUKG for urban prediction tasks. ORATOR's nodes are Overture features with stable GERS IDs, so the graph joins back to the full feature tables.
Query the spatial knowledge graph with a Wherobots free trial at cloud.wherobots.com.
An ontology is a shared, written-down model of a subject: the kinds of things it contains, the properties those things have, and the ways they relate to each other. It is written precisely enough that software can check data against it and draw conclusions from it, for example that if a restaurant is inside a building and the building is inside a city, the restaurant is inside that city too.
In AI, an ontology gives models and agents a fixed vocabulary for a domain: which entity types exist, which relations connect them, and which rules hold. Agents use it to interpret data consistently, to plan queries, and to ground answers in real entities. Vendors such as Microsoft describe ontologies as a machine-understandable representation of a business that people, applications, and AI agents share.
Palantir uses ontology for the layer in its software that connects an organization’s data to the real-world things that data describes, such as customers, assets, or orders, and the links between them. It is the information science sense of the word, applied to one company’s operational data.
A small spatial example: the classes Place, Building, Address, and Road segment; the relations located_in (Place to Building), has_address (Building to Address), and access (Building to the road network); and a rule that located_in is transitive. Public examples include schema.org, GeoNames, OGC GeoSPARQL, and the Gene Ontology.
A taxonomy arranges terms in a hierarchy of broader and narrower categories, such as Building, then Commercial building, then Office. An ontology includes that hierarchy and adds other relations between concepts, properties with types, and logical rules. Every taxonomy can be part of an ontology, but a taxonomy alone cannot say that a place is located in a building.
Not exactly. An ontology defines the types and relations, and a knowledge graph is the data: the actual entities and the links between them, organized by an ontology. In a spatial knowledge graph, the ontology says that places can be located in buildings, and the graph records that the Acme Bread Company is located in the San Francisco Ferry Building.
Many spatial methods compute over a graph of neighbors. Spatial statistics such as Moran’s I and local hot spot measures use a spatial weights matrix, which is the adjacency matrix of a neighbor graph. Road networks are graphs used for routing, isochrones, and centrality. Spatial knowledge graphs compute edges such as located_in and near from geometry, and graph neural networks learn from those neighbor relations for tasks such as traffic forecasting.
GeoSPARQL is an OGC standard for geospatial data on the Semantic Web. It defines an RDF vocabulary for features and geometries, such as geo:Feature, geo:Geometry, and geo:hasGeometry, plus spatial relation properties and functions that extend the SPARQL query language. Version 1.1 was published in January 2024.
What is H3? Uber’s Hexagonal Spatial Index
H3 is Uber's open-source hexagonal grid for indexing locations on Earth. How the H3 index works, its 16 resolutions, H3 vs S2 and geohash, and H3 in SQL.
What is Geospatial Data? Types, Formats, Examples
Geospatial data describes things tied to a place on Earth. The two main types, vector and raster, plus formats, sources, and examples.
What is Geospatial Analysis? Methods and Examples
Geospatial analysis, or spatial analysis, examines where things are and how location shapes patterns. Methods, examples, GIS vs spatial analysis, and SQL.
share this article
Awesome that you’d like to share our articles. Where would you like to share it to: