Planetary-scale answers, unlocked.
A Hands-On Guide for Working with Large-Scale Spatial Data. Learn more.
Authors
Context engineering is the practice of choosing, formatting, and maintaining the information a large language model receives when it runs: instructions, tool definitions, retrieved data, memory, and conversation history. Anthropic describes it in Effective context engineering for AI agents as the set of strategies for curating and maintaining the optimal set of tokens during inference. Context engineering for physical AI applies the same practice to questions about real places, where the context is geometry, location, imagery, and the spatial relations between them.
A model's context is the set of tokens it reads when it generates a response. Anthropic's Applied AI team calls context engineering "the natural progression of prompt engineering": the work moves from writing one good prompt to managing everything that enters the context window over many turns of an agent loop. Its guiding principle is to find "the smallest possible set of high-signal tokens that maximize the likelihood of some desired outcome."
Andrej Karpathy backed the term on 25 June 2025, when he described context engineering as "the delicate art and science of filling the context window" with the right information for the next step. Phil Schmid's breakdown of context lists the parts that fill that window:
IBM's explainer frames context engineering as the framework that brings prompt engineering and retrieval-augmented generation together, with data from structured sources counted as context.
Context windows now hold hundreds of thousands of tokens, and filling them has a cost. Liu and colleagues showed in Lost in the Middle (TACL, 2024) that accuracy is highest when the relevant passage sits at the start or end of a long input and drops when it sits in the middle. Chroma's Context Rot report (Hong, Troynikov, and Huber, 2025) tested 18 models, held task difficulty constant, varied only input length, and found that performance grew less reliable as input grew, even on simple tasks. Anthropic calls this a finite attention budget: every extra row or paragraph has to earn its place.
Context engineering contains the other two. Prompt engineering is one input to it, and retrieval-augmented generation (RAG) is one way to fill it.
RAG comes from Lewis and colleagues' Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (NeurIPS 2020), which paired a language model with a dense vector index of Wikipedia searched by a neural retriever. Context engineering keeps retrieval and adds the decisions around it: what to keep, what to drop, and when to fetch more.
LangChain's context engineering guide groups the techniques into four strategies: write, select, compress, and isolate. Anthropic's article describes the same moves:
Tools are how an agent selects context. The Model Context Protocol, which Anthropic open-sourced on 25 November 2024, is a standard for connecting AI assistants to the systems where data lives, so any MCP client can call a data source's tools.
Physical AI covers AI that perceives, reasons about, and acts in the physical world, from robots to models that read satellite imagery. Language models were trained on text, documents, databases, and the internet, and those sources rarely state which building sits in which flood zone. The GeoLLM study (Manvi and colleagues, ICLR 2024) found that prompting with coordinates alone gives poor estimates, and that adding nearby OpenStreetMap features to the prompt improved results by 70% over baselines. Context engineering for physical AI is the work of producing that map context for each question. Five things change when the context is places, and scale, covered after the research, changes the cost of all five.
Location is the join key. A text agent matches context by keyword or meaning. A place agent matches it by position: the places inside a building, the nearest road, the county around it. Each relation is a spatial join computed with predicates such as ST_Contains, ST_Intersects, and ST_KNN.
The layers come in different shapes. One building draws on polygons (footprints, parcels, flood zones, boundaries), lines (roads), points (businesses), and rasters (elevation, imagery). Geospatial data mixes these types, and the context step reduces each one to a few facts a model can read.
Freshness. Each layer updates on its own cycle. Overture publishes a map release every month, parcel records follow county schedules, flood maps carry an effective date, and weather warnings carry an expiration time. A context packet that records each source's date lets the reader judge what is current, and the ORATOR comparison below shows how far two dates for one building can drift.
Coordinate systems. A latitude and longitude only have meaning in a stated coordinate reference system. Overture stores geometry in longitude and latitude (EPSG:4326), so a planar distance on it comes out in degrees, and a degree of longitude shrinks with latitude: about 111 km at the Equator and 88.1 km at the Ferry Building's 37.8° N, where a degree of latitude is 111.0 km. The gap between the Ferry Building and The Embarcadero, its nearest primary road, is 0.00021 degrees. Multiplying by a fixed 111.32 km per degree gives 23.4 m. Measured between the closest points found in UTM zone 10N, the gap is 20.4 m, and the WGS 84 ellipsoid distance between those same points is also 20.4 m. The fixed conversion overstates the gap by about 15%, and the error grows toward the poles. The context step should do this math in the engine, with ST_DistanceSphere, ST_DistanceSpheroid, or a projected CRS, and pass the model meters.
Stable IDs. Each fact in the context needs an ID that a reader can cite and an agent can use to fetch more. Overture's Global Entity Reference System gives each feature an ID that stays stable across releases, so a building ID in the context can be fetched again next month. Grid systems such as H3 add a shared cell ID for data with no common key.
A spatial knowledge graph stores these relations ahead of time as typed edges, and an ontology defines what each edge means. Spatial Graph RAG for the physical world works through that approach: spatial SQL builds the edges, and graph traversal retrieves a subgraph as context.
Wherobots is the AI Context Engine for the Physical World, built for context engineering for physical AI: WherobotsDB runs the spatial joins, and the Wherobots MCP server lets an AI agent find tables, write spatial SQL, and run it. The Havasu catalog holds the layers, including Overture buildings, Overture places, road segments, divisions, and the Copernicus 30 m elevation model.
This query assembles context for the San Francisco Ferry Building. It finds the footprint, counts the places inside it, measures the distance to the nearest road in meters in UTM zone 10N, lists the divisions that contain it from country down, and reads the elevation model at its centroid. Each fact comes back with the ID or source record it came from. Bounding box filters keep each table read to the blocks around the building:
WITH b AS ( SELECT id, names.primary AS name, height, geometry, ST_Centroid(geometry) AS c, ST_Transform(ST_SetSRID(geometry, 4326), 'EPSG:4326', 'EPSG:32610') AS utm, sources[0].dataset AS footprint_source, sources[0].update_time AS footprint_updated FROM wherobots_open_data.overture_maps_foundation.buildings_building WHERE bbox.xmin > -122.396 AND bbox.xmax < -122.391 AND bbox.ymin > 37.794 AND bbox.ymax < 37.797 AND names.primary = 'San Francisco Ferry Building' ), places AS ( SELECT COUNT(*) AS n FROM wherobots_open_data.overture_maps_foundation.places_place p JOIN b ON ST_Contains(b.geometry, p.geometry) WHERE p.bbox.xmin > -122.396 AND p.bbox.xmax < -122.391 AND p.bbox.ymin > 37.794 AND p.bbox.ymax < 37.797 ), road AS ( SELECT s.id AS road_id, s.names.primary AS road_name, s.class AS road_class, ST_Distance(b.utm, ST_Transform(ST_SetSRID(s.geometry, 4326), 'EPSG:4326', 'EPSG:32610')) AS road_m FROM wherobots_open_data.overture_maps_foundation.transportation_segment s CROSS JOIN b WHERE s.bbox.xmin > -122.400 AND s.bbox.xmax < -122.387 AND s.bbox.ymin > 37.792 AND s.bbox.ymax < 37.799 AND s.subtype = 'road' AND s.class IN ('motorway', 'trunk', 'primary', 'secondary', 'tertiary', 'residential') ORDER BY road_m LIMIT 1 ), divs AS ( SELECT array_sort(collect_list(struct( CASE d.subtype WHEN 'country' THEN 1 WHEN 'region' THEN 2 WHEN 'county' THEN 3 ELSE 4 END AS rank, d.names.primary AS name, d.division_id AS division_id))) AS levels FROM wherobots_open_data.overture_maps_foundation.divisions_division_area d JOIN b ON ST_Contains(d.geometry, b.c) WHERE d.bbox.xmin < -122.393 AND d.bbox.xmax > -122.394 AND d.bbox.ymin < 37.795 AND d.bbox.ymax > 37.796 AND d.class = 'land' ), dem AS ( SELECT t.name AS dem_tile, RS_Value(t.rast, b.c) AS surface_m FROM wherobots_open_data.copernicus_dem.glo_30m t JOIN b ON ST_Intersects(t.footprint, b.c) ) SELECT b.id, b.name, b.height, ROUND(ST_AreaSpheroid(b.geometry)) AS footprint_m2, b.footprint_source, b.footprint_updated, places.n AS places_inside, road.road_id, road.road_name, road.road_class, ROUND(road.road_m, 1) AS road_m, concat_ws(' > ', transform(divs.levels, x -> x.name)) AS divisions, transform(divs.levels, x -> x.division_id) AS division_ids, dem.dem_tile, ROUND(dem.surface_m, 1) AS surface_m FROM b CROSS JOIN places CROSS JOIN road CROSS JOIN divs CROSS JOIN dem
It returns one row, the context packet for the building:
The 205 places span 49 categories, led by financial services (30), law firms (24), restaurants (19), and legal services (19). The Copernicus DEM is a digital surface model, so the surface value includes the roof as well as the ground. A second query adds a hazard layer from NWS watches and warnings: 102 National Weather Service alert polygons from the San Francisco Bay Area office (MTR) cover the building's centroid, including 76 areal flood advisories (VTEC code FA.Y) issued between February 2009 and November 2025:
SELECT PHENOM, SIG, COUNT(*) AS warnings, MIN(ISSUED) AS first_issued, MAX(ISSUED) AS last_issued FROM wherobots_open_data.noaa.nws_watch_warnings WHERE WFO = 'MTR' AND ST_Intersects(geometry, ST_Point(-122.39344, 37.79553)) GROUP BY PHENOM, SIG ORDER BY warnings DESC
Within 400 m of the building's centroid sit 79 buildings, 1,469 places, and 429 road segments, 1,977 rows in all. Passing all of those rows to a model would spend its attention budget on records the question does not need. The packet keeps the facts the question needs. Each feature fact carries its Overture ID, and the footprint and elevation values carry their source, so the agent can cite them and fetch more. The 205 places are a count over the building's geometry, so the building ID is the key that fetches them. The same pattern extends to parcels and flood zones, each one more join on the same geometry.
The packet above reduced 1,977 rows to 8 facts. Two more choices set how a packet like that grows and ages: how far from the building to search, and how current each source is. The radius counts below come from Overture Maps release 2026-09-23.1, read from Overture's public bucket, with distances computed in UTM zone 10N. The graph counts come from Wherobots queries on wherobots_open_data.
Context for a building starts with what is near it, and the count of nearby features grows fast with distance from the footprint.
The growth is uneven: places jump from 335 to 554 between 150 m and 175 m. The radius is a property of the question. Road access needs a few tens of meters, foot traffic needs a few hundred, and a flood or wildfire question needs the extent of the hazard. Tobler's first law of geography (1970) holds that everything is related to everything else, and near things more than distant things, which supports starting small and widening only when the question calls for it. No general rule picks the distance, and choosing it remains an open problem.
Each fact in a packet has its own date, and derived data ages too. Overture places carry their own source records and confidence scores. The ORATOR knowledge graph that Wherobots publishes as wherobots_open_data.spatial_knowledge_graph carries an update time of 18 June 2026 on its Ferry Building node. It links 112 places to the building with located_in edges, 86 of them at confidence 1.0, the containment matches described in What is an ontology?. The September 2026 Overture release has 205 place points inside the same footprint. The counts differ in date and in method, since ORATOR also adds proximity matches. The building's GERS ID, 68b6339c-54c7-4736-a228-01345e78bcdd, is the same in both, so an agent holding the graph's ID can fetch the current record.
wherobots_open_data.spatial_knowledge_graph
Context engineering for physical AI draws on three research threads: retrieval for language models, GeoAI, and distributed spatial computing.
The retriever in RAG came from Karpukhin and colleagues' Dense Passage Retrieval (EMNLP 2020), which beat a BM25 keyword baseline by 9 to 19 points of top-20 passage accuracy on open-domain question answering. As context windows grew, Lost in the Middle and Context Rot shifted the question from how much fits in the window to where the relevant passage sits and how much of the input helps. Edge and colleagues at Microsoft Research introduced GraphRAG (2024), which builds a knowledge graph from a text corpus, summarizes its communities, and answers global questions from those summaries. The Model Context Protocol (2024) then standardized how agents call data sources, and Anthropic's 2025 article set out on-demand retrieval, compaction, notes, and sub-agents as the working methods.
Each step moved work out of the model and into the system around it: first retrieval, then selection, then the decision of what to keep across a long run.
Geography research reached the same conclusion from the other side: location carries information that text alone does not. Tobler's first law is the reason context for a place starts with what is around it. Janowicz, Gao, McKenzie, Hu, and Bhaduri argued in GeoAI: spatially explicit artificial intelligence techniques for geographic knowledge discovery and beyond (IJGIS, 2020) that models for geographic problems need location, distance, and spatial relations built in.
Geographic knowledge in language models is uneven. Yamada and colleagues' Evaluating Spatial Understanding of Large Language Models (TMLR, 2024) tested navigation over grids, rings, and trees and found that accuracy varied with the structure and with how it was described. GeoLLM, cited above, measured what nearby map features in the prompt add to a model's estimates. Mai and colleagues' vision paper On the Opportunities and Challenges of Foundation Models for GeoAI (ACM TSAS, 2024) surveys where general models fall short on geospatial tasks and calls for multimodal models that combine text, imagery, and vector data.
Two lines of work turn those findings into systems. KnowWhereGraph, described by Janowicz and colleagues in Know, Know Where, KnowWhereGraph (AI Magazine, 2022), pairs a cross-domain geographic knowledge graph with a geo-enrichment service. For a region of interest, the service returns linked data on natural hazards, soils, climate, and demographics: a context packet for a place, served as a graph. Yu and colleagues' Spatial-RAG (2025) answers geospatial questions by pairing a spatial database with a language model. A sparse spatial filter finds candidates by location, dense semantic matching ranks them by meaning, and the final answers balance both.
One address takes a handful of joins. A portfolio of a million addresses takes the same joins a million times over tables with billions of rows, so a packet that is cheap for one building is expensive for a national portfolio. The joins have to run in a distributed engine, with results cached by stable ID. The radius multiplies the cost: for the Ferry Building, widening the search from 50 m to 300 m raises the place count from 242 to 1,090, and a portfolio repeats that choice for every building.
Jia Yu, Jinxuan Wu, and Mo Sarwat introduced GeoSpark (ACM SIGSPATIAL 2015) to run spatial joins like these across a cluster, and Yu, Zongsi Zhang, and Sarwat described its spatial partitioning, indexing, and join design in Spatial data management in Apache Spark: the GeoSpark perspective and beyond (GeoInformatica, 2019). GeoSpark became Apache Sedona, and Yu and Sarwat went on to found Wherobots.
Connect an AI agent to physical world data through the Wherobots MCP server with a Wherobots free trial at cloud.wherobots.com.
Context engineering is deciding what information an AI model reads before it answers: the instructions, the tools it can call, the documents or database rows it retrieves, what it remembers, and the conversation so far. The goal is the smallest set of relevant information that leads to a correct answer, refreshed at every step of an agent’s work.
A context engineer designs the systems that feed an AI model or agent its information: retrieval pipelines, tool definitions, memory, and the rules for what to keep, summarize, or drop as a task runs. The work sits between software engineering and data engineering, and it is often part of an AI engineer or machine learning engineer role.
Context engineering is a skill within AI engineer, machine learning engineer, and applied AI roles, so pay follows those job titles, which salary surveys report by level and location. The title context engineer is new, and published salary data for it on its own is thin.
Context engineering includes prompt engineering. Writing clear instructions still matters, and context engineering adds the other inputs an agent needs across many steps: retrieval, tools, memory, and compaction. Anthropic describes context engineering as the natural progression of prompt engineering.
No standards body defines a fixed set of layers. Most breakdowns cover the same parts: instructions (the system prompt), the user’s request, short-term state and history, long-term memory, retrieved information, and available tools, often with a required output format. Phil Schmid’s list names seven, and LangChain groups the techniques into four strategies: write, select, compress, and isolate.
Retrieval-augmented generation (RAG) is one technique inside context engineering: it searches an index and adds the matching passages or rows to the prompt. Context engineering also sets which tools the model can call, what it remembers between steps, when to summarize, and when to hand work to a sub-agent.
It is context engineering for questions about real places and objects. The context comes from spatial joins across buildings, parcels, roads, administrative boundaries, hazards, and imagery, each with its own scale, update cycle, coordinate system, and identifiers. A spatial engine computes those relations and returns a compact, sourced summary of one place for the model to read.
Language models were trained on text, documents, databases, and the internet, which rarely state relations such as which building sits in which flood zone. The GeoLLM study (ICLR 2024) found that prompting with coordinates alone gives poor geographic estimates, and that adding nearby map features from OpenStreetMap to the prompt improved results by 70% over baselines.
What is an Ontology? Definition, Parts, Spatial Ontologies
An ontology is a formal model of the things in a domain and how they relate. Its parts, how it differs from a taxonomy or knowledge graph, and what makes one spatial.
What is H3? Uber’s Hexagonal Spatial Index
H3 is Uber's open-source hexagonal grid for indexing locations on Earth. How the H3 index works, its 16 resolutions, H3 vs S2, geohash and quadkeys, and H3 in SQL.
What is Geospatial Data? Types, Formats, Examples
Geospatial data describes things tied to a place on Earth. The two main types, vector and raster, plus formats, sources, and examples.
share this article
Awesome that you’d like to share our articles. Where would you like to share it to: