Planetary-scale answers, unlocked.
A Hands-On Guide for Working with Large-Scale Spatial Data. Learn more.
Authors
Ground truth is information known to be true because it was observed or measured directly, used as the reference that a model is trained on and checked against. In machine learning it is usually a set of labels. For physical world data it is an observation with a known place and time: a rain gauge total, a surveyed field plot, a logged vehicle trajectory. A model is only as trustworthy as the ground truth it has been scored on, and for forecasts, maps, and world models of real places, scoring means matching each prediction to the observation at the same location.
Ground truth is the answer a prediction is judged against. The term comes from remote sensing, where analysts compared what a satellite or aerial image appeared to show with what a visit to the ground found. Russell Congalton's A review of assessing the accuracy of classifications of remotely sensed data (Remote Sensing of Environment, 1991) set out how to sample those reference sites and build the error matrix that compares a classified map with them, and that method is still how land cover maps report accuracy.
Three properties make a reference usable as ground truth:
In supervised machine learning, ground truth is the labeled dataset: each example paired with its correct answer. A model is fit to a training split, tuned on a validation split, and scored once on a test split it has never seen. The test score means something only when the test labels are correct.
They often are not. Curtis Northcutt, Anish Athalye, and Jonas Mueller's Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks (NeurIPS 2021) found an average of at least 3.3% label errors across the test sets of 10 widely used computer vision, language, and audio benchmarks, including at least 6% of the ImageNet validation set. A model that matches a wrong label is counted as correct, and one that gets the true answer is counted as wrong.
For imagery, labels are drawn on a map: a building outline traced over an aerial photo, a land cover class assigned to a field plot. The building footprint and land cover classification articles cover how those reference sets are built and how a model's output polygons are scored against them.
For a forecast or a world model of a real place, ground truth is an observation of the physical world: what a station, a gauge, a survey, or a sensor recorded at a known location and time. Allan Murphy and Robert Winkler's A general framework for forecast verification (Monthly Weather Review, 1987) defined verification as the study of the joint distribution of forecasts and observations, which only exists once each forecast is paired with the observation it predicted.
Two kinds of data look like ground truth and are not:
The same five steps apply to a weather forecast, a land cover map, and a driving simulator.
Wherobots is the AI Context Engine for Physical World Data. WherobotsDB runs spatial SQL across vector and raster data and reads public archives in place, the Havasu catalog holds open datasets as Apache Iceberg tables, and the Wherobots MCP server lets AI agents run the same queries. RasterFlow (Public Preview) builds mosaics from imagery, runs computer vision inference on them, and writes the results as vectorized outputs on the same platform.
This walkthrough scores a learned weather model against ground truth for one extreme event. Hurricane Helene caused catastrophic inland flooding in western North Carolina in September 2024 (NHC Tropical Cyclone Report). The world model article assembled the observed state of Buncombe County from GHCN-Daily gauges for that storm. Here the prediction is GraphCast, published by Remi Lam and colleagues as Learning skillful medium-range global weather forecasting (Science, 2023), from the run NOAA's National Weather Service made from GFS initial conditions at 00 UTC on 25 September 2024, the day the heavy rain began.
NOAA publishes each GraphCast GFS run on AWS as GRIB2 files, one per 6-hour forecast step, each about 266 MB, with an index that lists the byte range of every field. WherobotsDB reads GeoTIFF and NetCDF rasters and has no GRIB2 reader, and NOAA's machine learning weather prediction archive, which holds GraphCast runs from the same GFS initial conditions, stores each run as one 9.34 GB NetCDF file, which failed to load as a single binary value. So the 16 forecast cells around the county were taken from the 6-hourly total precipitation fields of the 25 September run, fetched by byte range (about 3.2 MB each) and decoded with ECMWF's ecCodes library:
# 0-96 h total precipitation (mm) from NOAA GraphCastGFS, run 2024-09-25 00 UTC import requests, eccodes base = ("https://noaa-nws-graphcastgfs-pds.s3.amazonaws.com/graphcastgfs.20240925/00/" "forecasts_13_levels/graphcastgfs.t00z.pgrb2.0p25.f") total = 0 for h in range(6, 97, 6): idx = requests.get(f"{base}{h:03d}.idx").text.splitlines() i = next(n for n, l in enumerate(idx) if f"APCP:surface:{h-6}-{h} hour acc" in l) start, end = int(idx[i].split(":")[1]), int(idx[i + 1].split(":")[1]) - 1 msg = requests.get(f"{base}{h:03d}", headers={"Range": f"bytes={start}-{end}"}).content g = eccodes.codes_new_from_message(msg) total = total + eccodes.codes_get_values(g).reshape(721, 1440) # 0.25° grid, 90N to 90S, 0 to 360E eccodes.codes_release(g) # Grid point (lat, lon): total[round((90 - lat) / 0.25), round((lon % 360) / 0.25)]
The 16 values go into the query as a table. Everything after that, including every number below, comes from one WherobotsDB query.
The query reads the GHCN-Daily station list and the 2024 daily file in place from NOAA's public bucket, keeps the stations inside the TIGER/Line county polygon, and joins each gauge to the forecast cell that contains it. It keeps gauges with quality-controlled rain on all eight days from 21 to 28 September, so each has a complete storm total for 25 to 28 September and a persistence forecast: the gauge's own total for the four days before.
-- Ground truth for a world model: Hurricane Helene rain at every complete Buncombe County gauge, -- 25 to 28 September 2024, scored against the GraphCast forecast from 00 UTC 25 September -- (NOAA GraphCastGFS, 0 to 96 h total precipitation on its 0.25-degree grid) and against -- persistence (the gauge's own total for the previous four days) WITH gc(lat, lon, gc_mm) AS ( -- GraphCast grid points around the county, decoded from NOAA's GraphCastGFS GRIB2 files VALUES (35.25, -83.0, 217.3), (35.25, -82.75, 212.8), (35.25, -82.5, 203.1), (35.25, -82.25, 190.2), (35.5, -83.0, 191.5), (35.5, -82.75, 192.7), (35.5, -82.5, 182.8), (35.5, -82.25, 184.9), (35.75, -83.0, 169.3), (35.75, -82.75, 163.3), (35.75, -82.5, 166.2), (35.75, -82.25, 176.3), (36.0, -83.0, 144.2), (36.0, -82.75, 153.0), (36.0, -82.5, 157.9), (36.0, -82.25, 155.5) ), cell AS ( -- Each grid point's cell: 0.25 degrees on a side, centered on the point SELECT gc_mm, ST_MakeEnvelope(lon - 0.125, lat - 0.125, lon + 0.125, lat + 0.125) AS geom FROM gc ), st AS ( SELECT substr(value, 1, 11) AS station_id, CAST(trim(substr(value, 13, 8)) AS DOUBLE) AS lat, CAST(trim(substr(value, 22, 9)) AS DOUBLE) AS lon, trim(substr(value, 42, 30)) AS name FROM text.`s3://noaa-ghcn-pds/ghcnd-stations.txt` WHERE substr(value, 39, 2) = 'NC' ), stn AS ( SELECT /*+ BROADCAST(c) */ st.* FROM st JOIN wherobots_open_data.us_census.tiger_county c ON ST_Intersects(c.geometry, ST_Point(st.lon, st.lat)) WHERE c.STATEFP = '37' AND c.NAME = 'Buncombe' ), obs AS ( -- Quality-controlled daily rain (tenths of mm) for 21 to 28 September SELECT /*+ BROADCAST(s) */ o._c0 AS station_id, to_date(o._c1, 'yyyyMMdd') AS d, CAST(o._c3 AS DOUBLE) / 10 AS mm FROM csv.`s3://noaa-ghcn-pds/csv/by_year/2024.csv` o JOIN stn s ON o._c0 = s.station_id WHERE o._c2 = 'PRCP' AND o._c5 IS NULL AND o._c1 BETWEEN '20240921' AND '20240928' ), tot AS ( SELECT station_id, SUM(IF(d >= DATE '2024-09-25', 1, 0)) AS event_days, SUM(IF(d >= DATE '2024-09-25', mm, 0)) AS obs_mm, SUM(IF(d < DATE '2024-09-25', 1, 0)) AS prior_days, SUM(IF(d < DATE '2024-09-25', mm, 0)) AS persist_mm FROM obs GROUP BY station_id ), scored AS ( -- Point-in-cell spatial join: each gauge meets the forecast cell it sits in SELECT /*+ BROADCAST(c) */ s.station_id, s.name, t.obs_mm, c.gc_mm, t.persist_mm, c.gc_mm - t.obs_mm AS gc_err, t.persist_mm - t.obs_mm AS persist_err FROM stn s JOIN tot t ON s.station_id = t.station_id JOIN cell c ON ST_Intersects(c.geom, ST_Point(s.lon, s.lat)) WHERE t.event_days = 4 AND t.prior_days = 4 ) SELECT COUNT(*) AS gauges, COUNT(DISTINCT station_id) AS distinct_gauges, ROUND(PERCENTILE(obs_mm, 0.5), 1) AS median_obs_mm, ROUND(PERCENTILE(gc_mm, 0.5), 1) AS median_graphcast_mm, ROUND(AVG(ABS(gc_err)), 1) AS graphcast_mae_mm, ROUND(AVG(gc_err), 1) AS graphcast_bias_mm, ROUND(SQRT(AVG(gc_err * gc_err)), 1) AS graphcast_rmse_mm, ROUND(AVG(ABS(persist_err)), 1) AS persistence_mae_mm, ROUND(SQRT(AVG(persist_err * persist_err)), 1) AS persistence_rmse_mm, ROUND(1 - AVG(gc_err * gc_err) / AVG(persist_err * persist_err), 2) AS mse_skill_vs_persistence, SUM(IF(gc_err < 0, 1, 0)) AS gauges_underforecast, ROUND(MIN(gc_mm / obs_mm), 2) AS min_share_forecast, ROUND(PERCENTILE(gc_mm / obs_mm, 0.5), 2) AS median_share_forecast, ROUND(MAX(gc_mm / obs_mm), 2) AS max_share_forecast, MAX_BY(name, -gc_err) AS worst_gauge, ROUND(MAX(obs_mm), 1) AS worst_obs_mm, ROUND(MAX_BY(gc_mm, -gc_err), 1) AS worst_forecast_mm, ROUND(MAX(IF(station_id = 'USW00003812', obs_mm, NULL)), 1) AS airport_obs_mm, ROUND(MAX(IF(station_id = 'USW00003812', gc_mm, NULL)), 1) AS airport_forecast_mm FROM scored
It ran in 21 seconds and returned one row:
GraphCast placed heavy rain over the county four days out, and against persistence, which forecast little rain because the days before the storm were nearly dry, it scores a skill of 0.71. Against the gauges it forecast about half of what fell, and less than fell at every one of them. The bias equals the mean absolute error, so every error has the same sign. EXPLAIN FORMATTED shows R-tree index joins (BroadcastIndexJoin) for the stations to the county and the gauges to the cells, a BroadcastHashJoin from the stations to the CSV rows, and no CartesianProduct.
EXPLAIN FORMATTED
BroadcastIndexJoin
BroadcastHashJoin
CartesianProduct
The map shows the scale mismatch. A 0.25° cell is about 28 km north to south here and carries one forecast value, while the gauges inside it recorded different totals: the cell containing Black Mountain 2.1 W forecast 184.9 mm where that gauge recorded 516.4 mm.
Scoring is the start. The error measured at known places and times goes back into the model in three ways.
Retraining and fine-tuning. Supervised models, including learned weather models such as GraphCast, are trained to reduce the gap between prediction and target. When the target is reanalysis, errors at stations can survive training, which is why the observation-based benchmarks above exist. Ground truth for extremes like Helene shows where a model needs more or better training examples.
Data assimilation. Forecast systems correct their starting state with new observations every cycle. Geir Evensen's Sequential data assimilation with a nonlinear quasi-geostrophic model using Monte Carlo methods to forecast error statistics (Journal of Geophysical Research, 1994) introduced the ensemble Kalman filter, which systems such as the HRRR model use to blend forecasts with radar and surface observations.
Training inside a world model. In model-based reinforcement learning, an agent learns a policy inside a learned world model. Danijar Hafner and colleagues' DreamerV3, in Mastering diverse control tasks through world models (Nature, 2025), trains behavior on imagined rollouts. The policy learns whatever the model predicts, including its mistakes. David Ha and Jürgen Schmidhuber's World Models (2018) describe a controller that found an adversarial policy inside an imperfect model, moving so that monsters in the simulated game never fired, an exploit absent from the real game. Ground truth measures that gap before training starts. For driving world models, the reference is real logged behavior, as in the Waymo Open Sim Agents Challenge, and for world models of real places it extends to the place itself: road geometry, terrain, and the weather recorded there on the day a scene depicts. The physical AI article covers how these models train robots and vehicles.
Ground truth as a formal practice grew from two fields. In remote sensing, Congalton's 1991 review made the error matrix and reference sampling the standard for map accuracy. In weather, Murphy and Winkler's 1987 framework treated verification as the joint distribution of forecasts and observations, and Murphy's 1988 skill score tied it to a reference forecast. Machine learning added held-out test sets and, with Northcutt and colleagues in 2021, measured how often the labels in them are wrong. Learned weather models brought the two traditions together: WeatherBench 2, by Stephan Rasp and colleagues in WeatherBench 2: A benchmark for the next generation of data-driven global weather models (Journal of Advances in Modeling Earth Systems, 2024), scores them on shared grids, and station benchmarks such as WeatherReal score them where people live.
What scale changes. One county, one storm, and one forecast run make a single query. A season of twice-daily runs scored at every gauge in a country, or every simulated scene matched to the road it depicts, multiplies both sides of the join, and the matching has to run where the observations and predictions are stored. Spatial partitioning and per-partition indexes, introduced for Apache Spark by Jia Yu, Jinxuan Wu, and Mohamed Sarwat in GeoSpark (ACM SIGSPATIAL 2015), keep each comparison to nearby candidates. The partitioning tax article covers what happens to joins like this when the data is split into tiles instead.
Score your own models against ground truth with a Wherobots free trial at cloud.wherobots.com.
Ground truth is information known to be true because it was observed or measured directly, used as the reference that a model is trained on or checked against. In machine learning it is usually a set of labels. For physical world data it is an observation with a known place and time, such as a rain gauge reading, a surveyed field plot, or a logged vehicle trajectory.
The term comes from remote sensing, where it meant information collected on the ground to check what a satellite or aerial image appeared to show. It now means any trusted reference for judging a prediction: if a model says a pixel is forest, the ground truth is what a visit to that spot found.
In machine learning, ground truth data is the set of correct answers for a task: the label for each image, the class for each pixel, or the value a model should predict. Models are trained on part of it and evaluated on a held-out part. Labels come from people or instruments and contain errors: Curtis Northcutt, Anish Athalye, and Jonas Mueller found an average of at least 3.3% label errors across the test sets of 10 widely used benchmarks in Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks (NeurIPS 2021).
A ground truth dataset is a collection of trusted reference values used to train or score a model. Examples include labeled image sets, field survey plots used to check land cover maps, and networks of weather stations such as NOAA’s Global Historical Climatology Network, which holds daily records from land stations worldwide.
A rain gauge total checks a weather forecast, a surveyed field plot checks a land cover classification, a hand-traced roof outline checks a building footprint model, an insurance claim checks a damage estimate, and a logged trajectory from a real vehicle checks a driving simulator. Each pairs a prediction with an observation of the same thing at the same place and time.
A label is one recorded answer, usually made by a person or an instrument. Ground truth is the reference those labels are meant to capture. Labels approximate ground truth: two annotators can disagree, and a label can be wrong. Treating labels as perfect ground truth hides their error inside the model’s measured error.
Define exactly what the model predicts, collect independent observations of the same quantity, match each prediction to the observation at the same place and time, and compute an error against a simple reference forecast such as persistence. For spatial data the matching step is a spatial join: a gauge point inside a forecast cell, or a building outline inside an image tile.
Reanalysis such as ERA5 is a gridded reconstruction of past weather that combines a model with historical observations, so it is a model product. Learned weather models are trained and often scored on it, and it can differ from what a station measured: Vivek Ramavajjala and Peetak Mitra found that FourCastNet outperforms IFS when tested against ERA5 but shows no benefit when tested against real-world observations, in Verification against in-situ observations for Data-Driven Weather Prediction (2023).
A world model predicts how an environment changes, and agents can be trained inside it. If its predictions drift from reality, anything trained inside it learns the drift. Ground truth at known places and times measures that drift so the model can be corrected before it is used for training or decisions.
Wherobots is the AI Context Engine for Physical World Data. WherobotsDB reads observations such as NOAA station records in place, joins them to boundaries, forecast cells, buildings, and imagery in spatial SQL, and scores predictions at every location in one query. RasterFlow (Public Preview) builds mosaics from imagery, runs computer vision inference on them, and writes the results as vectorized outputs on the same platform, where they can be checked against reference data.
What is a World Model? How AI Predicts the Physical World
A world model is an AI model that predicts how an environment changes, so a system can plan before it acts. How world models work, the main types and examples, and why a world model of real places is only as good as the observations it is checked against.
What is a Context Engine? Why Physical World Data Needs One
A context engine is the system that computes, stores, and serves the context AI models and agents receive. How it works, examples, and why physical world data needs one that computes spatial relations.
What Is the Partitioning Tax? Tiles, Seams, and Reruns
The partitioning tax is the cost of splitting spatial data into tiles when it outgrows one machine: tile loops, seam errors at tile edges, and reruns. How it shows up and how distributed spatial engines remove it.
share this article
Awesome that you’d like to share our articles. Where would you like to share it to: