Rasterflow, Earth Intelligence & inference engine now in public preview Learn More

What is Context Engineering for Physical AI?

Authors

Context engineering is the practice of choosing, formatting, and maintaining the information a large language model receives when it runs: instructions, tool definitions, retrieved data, memory, and conversation history. Anthropic describes it in Effective context engineering for AI agents as the set of strategies for curating and maintaining the optimal set of tokens during inference. Context engineering for physical AI applies the same practice to questions about real places, where the context is geometry, location, imagery, and the spatial relations between them.

Key takeaways

  • Context is every token a model reads at inference time. Context engineering sets which tokens those are, step by step.
  • Prompt engineering writes the instructions. Context engineering also covers retrieval, tools, memory, and compression across an agent's whole run.
  • Model accuracy becomes less reliable as input grows, so the goal is the smallest set of high-signal tokens.
  • For physical AI, the context comes from spatial joins across buildings, places, roads, boundaries, hazards, and imagery, each with its own scale, update cycle, coordinate system, and identifiers.
  • A spatial engine computes that context in a query, and the model receives a compact, sourced summary of one place.
  • Distances belong in meters. At the Ferry Building, a fixed degree-to-meter conversion overstates the gap to the nearest road by about 15%.

What is context engineering?

A model's context is the set of tokens it reads when it generates a response. Anthropic's Applied AI team calls context engineering "the natural progression of prompt engineering": the work moves from writing one good prompt to managing everything that enters the context window over many turns of an agent loop. Its guiding principle is to find "the smallest possible set of high-signal tokens that maximize the likelihood of some desired outcome."

Andrej Karpathy backed the term on 25 June 2025, when he described context engineering as "the delicate art and science of filling the context window" with the right information for the next step. Phil Schmid's breakdown of context lists the parts that fill that window:

  • Instructions: the system prompt, rules, and examples.
  • The user prompt: the task at hand.
  • State and history: the current conversation.
  • Long-term memory: facts and preferences kept across sessions.
  • Retrieved information: documents, database rows, and API results.
  • Available tools: the functions the model can call, with their descriptions.
  • Structured output: the format the answer must follow.

IBM's explainer frames context engineering as the framework that brings prompt engineering and retrieval-augmented generation together, with data from structured sources counted as context.

Why longer context can hurt

Context windows now hold hundreds of thousands of tokens, and filling them has a cost. Liu and colleagues showed in Lost in the Middle (TACL, 2024) that accuracy is highest when the relevant passage sits at the start or end of a long input and drops when it sits in the middle. Chroma's Context Rot report (Hong, Troynikov, and Huber, 2025) tested 18 models, held task difficulty constant, varied only input length, and found that performance grew less reliable as input grew, even on simple tasks. Anthropic calls this a finite attention budget: every extra row or paragraph has to earn its place.

Context engineering vs prompt engineering vs RAG

Context engineering contains the other two. Prompt engineering is one input to it, and retrieval-augmented generation (RAG) is one way to fill it.

Prompt engineeringRAGContext engineering
What it shapesThe instructionsRetrieved passages or rowsEvery token in the window, at every step
When it happensWritten once, before the runAt query timeAt each step of an agent loop
Main techniquesRole, rules, examples, output formatEmbedding search, rerankingRetrieval, tool design, memory, compaction, sub-agents
Typical failureAmbiguous instructionsRelevant passage missedToo much, stale, or conflicting context

RAG comes from Lewis and colleagues' Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (NeurIPS 2020), which paired a language model with a dense vector index of Wikipedia searched by a neural retriever. Context engineering keeps retrieval and adds the decisions around it: what to keep, what to drop, and when to fetch more.

How context engineering works

LangChain's context engineering guide groups the techniques into four strategies: write, select, compress, and isolate. Anthropic's article describes the same moves:

  1. Retrieve on demand. The agent keeps lightweight identifiers, such as file paths, stored queries, or links, and loads the data behind them with a tool call when a step needs it. Claude Code loads its CLAUDE.md files up front and uses glob and grep to pull other files as it works.
  2. Compact. When a conversation nears the window limit, the agent summarizes it and continues from the summary.
  3. Write notes. The agent saves progress and open questions to memory outside the window and reads them back later.
  4. Isolate with sub-agents. A sub-agent runs one subtask in its own clean window and returns a condensed result.

Tools are how an agent selects context. The Model Context Protocol, which Anthropic open-sourced on 25 November 2024, is a standard for connecting AI assistants to the systems where data lives, so any MCP client can call a data source's tools.

Context engineering examples

  • Coding agents. An agent reads the project's instruction file, searches the repository with grep, opens only the matching files, and runs tests to check its change.
  • Customer support. A support agent retrieves the customer's account, the matching help articles, and the ticket history.
  • Deep research. A lead agent sends sub-agents to read sources in parallel, and each returns a short summary with citations.
  • Property questions. An agent asked about one building needs its footprint, the businesses inside it, the road that reaches it, and the boundaries and hazards around it. No document holds those facts together, so a spatial query has to compute them. The example below does this for one building.

Context engineering for physical AI

Physical AI covers AI that perceives, reasons about, and acts in the physical world, from robots to models that read satellite imagery. Language models were trained on text, documents, databases, and the internet, and those sources rarely state which building sits in which flood zone. The GeoLLM study (Manvi and colleagues, ICLR 2024) found that prompting with coordinates alone gives poor estimates, and that adding nearby OpenStreetMap features to the prompt improved results by 70% over baselines. Context engineering for physical AI is the work of producing that map context for each question. Five things change when the context is places, and scale, covered after the research, changes the cost of all five.

Location is the join key. A text agent matches context by keyword or meaning. A place agent matches it by position: the places inside a building, the nearest road, the county around it. Each relation is a spatial join computed with predicates such as ST_Contains, ST_Intersects, and ST_KNN.

The layers come in different shapes. One building draws on polygons (footprints, parcels, flood zones, boundaries), lines (roads), points (businesses), and rasters (elevation, imagery). Geospatial data mixes these types, and the context step reduces each one to a few facts a model can read.

Freshness. Each layer updates on its own cycle. Overture publishes a map release every month, parcel records follow county schedules, flood maps carry an effective date, and weather warnings carry an expiration time. A context packet that records each source's date lets the reader judge what is current, and the ORATOR comparison below shows how far two dates for one building can drift.

Coordinate systems. A latitude and longitude only have meaning in a stated coordinate reference system. Overture stores geometry in longitude and latitude (EPSG:4326), so a planar distance on it comes out in degrees, and a degree of longitude shrinks with latitude: about 111 km at the Equator and 88.1 km at the Ferry Building's 37.8° N, where a degree of latitude is 111.0 km. The gap between the Ferry Building and The Embarcadero, its nearest primary road, is 0.00021 degrees. Multiplying by a fixed 111.32 km per degree gives 23.4 m. Measured between the closest points found in UTM zone 10N, the gap is 20.4 m, and the WGS 84 ellipsoid distance between those same points is also 20.4 m. The fixed conversion overstates the gap by about 15%, and the error grows toward the poles. The context step should do this math in the engine, with ST_DistanceSphere, ST_DistanceSpheroid, or a projected CRS, and pass the model meters.

Line chart of the length of one degree of latitude and one degree of longitude on the WGS 84 ellipsoid from 0 to 80 degrees latitude. Latitude stays near 111 kilometers, and longitude falls from 111 kilometers at the Equator to 88.1 kilometers at the Ferry Building's 37.8 degrees north. A card compares the distance from the Ferry Building to The Embarcadero as 0.00021 degrees, 23.4 meters by a fixed conversion, 20.4 meters geodesic, and 20.4 meters in UTM
At 37.8° N a degree of longitude is 88.1 km and a degree of latitude is 111.0 km. The Ferry Building to The Embarcadero gap is 0.00021° in EPSG:4326 and about 20 m on the ground.

Stable IDs. Each fact in the context needs an ID that a reader can cite and an agent can use to fetch more. Overture's Global Entity Reference System gives each feature an ID that stays stable across releases, so a building ID in the context can be fetched again next month. Grid systems such as H3 add a shared cell ID for data with no common key.

A spatial knowledge graph stores these relations ahead of time as typed edges, and an ontology defines what each edge means. Spatial Graph RAG for the physical world works through that approach: spatial SQL builds the edges, and graph traversal retrieves a subgraph as context.

Context engineering in Wherobots

Wherobots is the AI Context Engine for the Physical World, built for context engineering for physical AI: WherobotsDB runs the spatial joins, and the Wherobots MCP server lets an AI agent find tables, write spatial SQL, and run it. The Havasu catalog holds the layers, including Overture buildings, Overture places, road segments, divisions, and the Copernicus 30 m elevation model.

This query assembles context for the San Francisco Ferry Building. It finds the footprint, counts the places inside it, measures the distance to the nearest road in meters in UTM zone 10N, lists the divisions that contain it from country down, and reads the elevation model at its centroid. Each fact comes back with the ID or source record it came from. Bounding box filters keep each table read to the blocks around the building:

WITH b AS (
  SELECT id, names.primary AS name, height, geometry, ST_Centroid(geometry) AS c,
         ST_Transform(ST_SetSRID(geometry, 4326), 'EPSG:4326', 'EPSG:32610') AS utm,
         sources[0].dataset AS footprint_source, sources[0].update_time AS footprint_updated
  FROM wherobots_open_data.overture_maps_foundation.buildings_building
  WHERE bbox.xmin > -122.396 AND bbox.xmax < -122.391
    AND bbox.ymin > 37.794 AND bbox.ymax < 37.797
    AND names.primary = 'San Francisco Ferry Building'
),
places AS (
  SELECT COUNT(*) AS n
  FROM wherobots_open_data.overture_maps_foundation.places_place p
  JOIN b ON ST_Contains(b.geometry, p.geometry)
  WHERE p.bbox.xmin > -122.396 AND p.bbox.xmax < -122.391
    AND p.bbox.ymin > 37.794 AND p.bbox.ymax < 37.797
),
road AS (
  SELECT s.id AS road_id, s.names.primary AS road_name, s.class AS road_class,
         ST_Distance(b.utm, ST_Transform(ST_SetSRID(s.geometry, 4326), 'EPSG:4326', 'EPSG:32610')) AS road_m
  FROM wherobots_open_data.overture_maps_foundation.transportation_segment s
  CROSS JOIN b
  WHERE s.bbox.xmin > -122.400 AND s.bbox.xmax < -122.387
    AND s.bbox.ymin > 37.792 AND s.bbox.ymax < 37.799
    AND s.subtype = 'road'
    AND s.class IN ('motorway', 'trunk', 'primary', 'secondary', 'tertiary', 'residential')
  ORDER BY road_m
  LIMIT 1
),
divs AS (
  SELECT array_sort(collect_list(struct(
           CASE d.subtype WHEN 'country' THEN 1 WHEN 'region' THEN 2
                WHEN 'county' THEN 3 ELSE 4 END AS rank,
           d.names.primary AS name, d.division_id AS division_id))) AS levels
  FROM wherobots_open_data.overture_maps_foundation.divisions_division_area d
  JOIN b ON ST_Contains(d.geometry, b.c)
  WHERE d.bbox.xmin < -122.393 AND d.bbox.xmax > -122.394
    AND d.bbox.ymin < 37.795 AND d.bbox.ymax > 37.796
    AND d.class = 'land'
),
dem AS (
  SELECT t.name AS dem_tile, RS_Value(t.rast, b.c) AS surface_m
  FROM wherobots_open_data.copernicus_dem.glo_30m t
  JOIN b ON ST_Intersects(t.footprint, b.c)
)
SELECT b.id, b.name, b.height, ROUND(ST_AreaSpheroid(b.geometry)) AS footprint_m2,
       b.footprint_source, b.footprint_updated,
       places.n AS places_inside,
       road.road_id, road.road_name, road.road_class, ROUND(road.road_m, 1) AS road_m,
       concat_ws(' > ', transform(divs.levels, x -> x.name)) AS divisions,
       transform(divs.levels, x -> x.division_id) AS division_ids,
       dem.dem_tile, ROUND(dem.surface_m, 1) AS surface_m
FROM b CROSS JOIN places CROSS JOIN road CROSS JOIN divs CROSS JOIN dem

It returns one row, the context packet for the building:

FieldValue
id68b6339c-54c7-4736-a228-01345e78bcdd
height15 m
footprint_m29,900
footprint_source, footprint_updatedOpenStreetMap, 31 August 2026
places_inside205
road_id3bbd1042-26cc-4251-b785-0f610be864d8
road_name, road_class, road_mThe Embarcadero, primary, 20.4 m
divisionsUnited States > California > San Francisco
division_idsf39eb4af-5206-481b-b19e-bd784ded3f05, 4c555166-326b-48b4-8a33-155eb7c23e8d, 713e9077-600c-4f06-84b9-6bae6e559c34
dem_tile, surface_mCopernicus_DSM_COG_10_N37_00_W123_00_DEM.tif, 10.8

The 205 places span 49 categories, led by financial services (30), law firms (24), restaurants (19), and legal services (19). The Copernicus DEM is a digital surface model, so the surface value includes the roof as well as the ground. A second query adds a hazard layer from NWS watches and warnings: 102 National Weather Service alert polygons from the San Francisco Bay Area office (MTR) cover the building's centroid, including 76 areal flood advisories (VTEC code FA.Y) issued between February 2009 and November 2025:

SELECT PHENOM, SIG, COUNT(*) AS warnings, MIN(ISSUED) AS first_issued, MAX(ISSUED) AS last_issued
FROM wherobots_open_data.noaa.nws_watch_warnings
WHERE WFO = 'MTR'
  AND ST_Intersects(geometry, ST_Point(-122.39344, 37.79553))
GROUP BY PHENOM, SIG
ORDER BY warnings DESC
Map of a 400 meter circle around the San Francisco Ferry Building with building footprints, place points, and road segments drawn faintly, the Ferry Building footprint highlighted in purple with 205 place points inside it, and the nearest primary road, The Embarcadero, highlighted in yellow. A card beside the map lists the building's GERS ID, footprint area, place count, top categories, nearest road, divisions, surface height, and footprint source
Within 400 m of the Ferry Building sit 1,977 buildings, places, and road segments. The context packet keeps 8 facts about the building. Source: Overture Maps release 2026-09-23.1 and Copernicus GLO-30 in wherobots_open_data.

Within 400 m of the building's centroid sit 79 buildings, 1,469 places, and 429 road segments, 1,977 rows in all. Passing all of those rows to a model would spend its attention budget on records the question does not need. The packet keeps the facts the question needs. Each feature fact carries its Overture ID, and the footprint and elevation values carry their source, so the agent can cite them and fetch more. The 205 places are a count over the building's geometry, so the building ID is the key that fetches them. The same pattern extends to parcels and flood zones, each one more join on the same geometry.

Radius and freshness on the Ferry Building

The packet above reduced 1,977 rows to 8 facts. Two more choices set how a packet like that grows and ages: how far from the building to search, and how current each source is. The radius counts below come from Overture Maps release 2026-09-23.1, read from Overture's public bucket, with distances computed in UTM zone 10N. The graph counts come from Wherobots queries on wherobots_open_data.

Context for a building starts with what is near it, and the count of nearby features grows fast with distance from the footprint.

Three bar charts of features within 25 to 300 meters of the San Francisco Ferry Building footprint. Places rise from 242 at 50 meters to 1,090 at 300 meters, road segments from 75 to 360, and other buildings from 4 to 68
Features within a distance of the Ferry Building footprint. From 50 m to 300 m, places grow from 242 to 1,090 and road segments from 75 to 360. Source: Overture Maps release 2026-09-23.1.
Distance from the footprintPlacesRoad segmentsOther buildings
50 m242754
100 m26713311
200 m68524728
300 m1,09036068

The growth is uneven: places jump from 335 to 554 between 150 m and 175 m. The radius is a property of the question. Road access needs a few tens of meters, foot traffic needs a few hundred, and a flood or wildfire question needs the extent of the hazard. Tobler's first law of geography (1970) holds that everything is related to everything else, and near things more than distant things, which supports starting small and widening only when the question calls for it. No general rule picks the distance, and choosing it remains an open problem.

How current each source is

Each fact in a packet has its own date, and derived data ages too. Overture places carry their own source records and confidence scores. The ORATOR knowledge graph that Wherobots publishes as wherobots_open_data.spatial_knowledge_graph carries an update time of 18 June 2026 on its Ferry Building node. It links 112 places to the building with located_in edges, 86 of them at confidence 1.0, the containment matches described in What is an ontology?. The September 2026 Overture release has 205 place points inside the same footprint. The counts differ in date and in method, since ORATOR also adds proximity matches. The building's GERS ID, 68b6339c-54c7-4736-a228-01345e78bcdd, is the same in both, so an agent holding the graph's ID can fetch the current record.

Context engineering research

Context engineering for physical AI draws on three research threads: retrieval for language models, GeoAI, and distributed spatial computing.

From retrieval to context engineering

The retriever in RAG came from Karpukhin and colleagues' Dense Passage Retrieval (EMNLP 2020), which beat a BM25 keyword baseline by 9 to 19 points of top-20 passage accuracy on open-domain question answering. As context windows grew, Lost in the Middle and Context Rot shifted the question from how much fits in the window to where the relevant passage sits and how much of the input helps. Edge and colleagues at Microsoft Research introduced GraphRAG (2024), which builds a knowledge graph from a text corpus, summarizes its communities, and answers global questions from those summaries. The Model Context Protocol (2024) then standardized how agents call data sources, and Anthropic's 2025 article set out on-demand retrieval, compaction, notes, and sub-agents as the working methods.

Each step moved work out of the model and into the system around it: first retrieval, then selection, then the decision of what to keep across a long run.

Why place questions need computed context

Geography research reached the same conclusion from the other side: location carries information that text alone does not. Tobler's first law is the reason context for a place starts with what is around it. Janowicz, Gao, McKenzie, Hu, and Bhaduri argued in GeoAI: spatially explicit artificial intelligence techniques for geographic knowledge discovery and beyond (IJGIS, 2020) that models for geographic problems need location, distance, and spatial relations built in.

Geographic knowledge in language models is uneven. Yamada and colleagues' Evaluating Spatial Understanding of Large Language Models (TMLR, 2024) tested navigation over grids, rings, and trees and found that accuracy varied with the structure and with how it was described. GeoLLM, cited above, measured what nearby map features in the prompt add to a model's estimates. Mai and colleagues' vision paper On the Opportunities and Challenges of Foundation Models for GeoAI (ACM TSAS, 2024) surveys where general models fall short on geospatial tasks and calls for multimodal models that combine text, imagery, and vector data.

Spatial retrieval and geo-enrichment

Two lines of work turn those findings into systems. KnowWhereGraph, described by Janowicz and colleagues in Know, Know Where, KnowWhereGraph (AI Magazine, 2022), pairs a cross-domain geographic knowledge graph with a geo-enrichment service. For a region of interest, the service returns linked data on natural hazards, soils, climate, and demographics: a context packet for a place, served as a graph. Yu and colleagues' Spatial-RAG (2025) answers geospatial questions by pairing a spatial database with a language model. A sparse spatial filter finds candidates by location, dense semantic matching ranks them by meaning, and the final answers balance both.

What scale changes

One address takes a handful of joins. A portfolio of a million addresses takes the same joins a million times over tables with billions of rows, so a packet that is cheap for one building is expensive for a national portfolio. The joins have to run in a distributed engine, with results cached by stable ID. The radius multiplies the cost: for the Ferry Building, widening the search from 50 m to 300 m raises the place count from 242 to 1,090, and a portfolio repeats that choice for every building.

Jia Yu, Jinxuan Wu, and Mo Sarwat introduced GeoSpark (ACM SIGSPATIAL 2015) to run spatial joins like these across a cluster, and Yu, Zongsi Zhang, and Sarwat described its spatial partitioning, indexing, and join design in Spatial data management in Apache Spark: the GeoSpark perspective and beyond (GeoInformatica, 2019). GeoSpark became Apache Sedona, and Yu and Sarwat went on to found Wherobots.

Open problems

  • Conflicting sources. Overture, parcel records, and a county's own data can disagree on a building's height or use. A packet that keeps each value with its source puts the conflict in plain view.
  • Time. Most layers hold one snapshot. Questions about change need versioned sources and a date on every fact.
  • Evaluation. Benchmarks for spatial reasoning test models on puzzles and map questions. Few test whether an agent fetched the right context for a real place and used it correctly.
  • Physical AI: AI that perceives, reasons about, and acts in the physical world
  • Geospatial AI: machine learning and AI agents applied to spatial data
  • Ontology: the classes and relations that give spatial context its meaning
  • Spatial join: the operation that computes spatial relations between datasets
  • Building footprint: the polygon at the center of most building context
  • POI data: the businesses and services inside buildings
  • Coordinate reference system: what gives a coordinate its location and units

Read more from Wherobots

Connect an AI agent to physical world data through the Wherobots MCP server with a Wherobots free trial at cloud.wherobots.com.

Frequently asked questions

What is context engineering in simple terms?

Context engineering is deciding what information an AI model reads before it answers: the instructions, the tools it can call, the documents or database rows it retrieves, what it remembers, and the conversation so far. The goal is the smallest set of relevant information that leads to a correct answer, refreshed at every step of an agent’s work.

What is a context engineer?

A context engineer designs the systems that feed an AI model or agent its information: retrieval pipelines, tool definitions, memory, and the rules for what to keep, summarize, or drop as a task runs. The work sits between software engineering and data engineering, and it is often part of an AI engineer or machine learning engineer role.

How much does a context engineer make?

Context engineering is a skill within AI engineer, machine learning engineer, and applied AI roles, so pay follows those job titles, which salary surveys report by level and location. The title context engineer is new, and published salary data for it on its own is thin.

Is context engineering replacing prompt engineering?

Context engineering includes prompt engineering. Writing clear instructions still matters, and context engineering adds the other inputs an agent needs across many steps: retrieval, tools, memory, and compaction. Anthropic describes context engineering as the natural progression of prompt engineering.

What are the five layers of context engineering?

No standards body defines a fixed set of layers. Most breakdowns cover the same parts: instructions (the system prompt), the user’s request, short-term state and history, long-term memory, retrieved information, and available tools, often with a required output format. Phil Schmid’s list names seven, and LangChain groups the techniques into four strategies: write, select, compress, and isolate.

What is the difference between context engineering and RAG?

Retrieval-augmented generation (RAG) is one technique inside context engineering: it searches an index and adds the matching passages or rows to the prompt. Context engineering also sets which tools the model can call, what it remembers between steps, when to summarize, and when to hand work to a sub-agent.

What is context engineering for physical AI?

It is context engineering for questions about real places and objects. The context comes from spatial joins across buildings, parcels, roads, administrative boundaries, hazards, and imagery, each with its own scale, update cycle, coordinate system, and identifiers. A spatial engine computes those relations and returns a compact, sourced summary of one place for the model to read.

Why do AI agents need spatial context?

Language models were trained on text, documents, databases, and the internet, which rarely state relations such as which building sits in which flood zone. The GeoLLM study (ICLR 2024) found that prompting with coordinates alone gives poor geographic estimates, and that adding nearby map features from OpenStreetMap to the prompt improved results by 70% over baselines.