Rasterflow, Earth Intelligence & inference engine now in public preview Learn More

What is Ground Truth Data? Meaning and How to Check AI

Authors

Ground truth is information known to be true because it was observed or measured directly, used as the reference that a model is trained on and checked against. In machine learning it is usually a set of labels. For physical world data it is an observation with a known place and time: a rain gauge total, a surveyed field plot, a logged vehicle trajectory. A model is only as trustworthy as the ground truth it has been scored on, and for forecasts, maps, and world models of real places, scoring means matching each prediction to the observation at the same location.

Key takeaways

  • Ground truth is the trusted reference for a prediction. A forecast, a warning, or another model's output is a prediction and never counts as ground truth.
  • In machine learning, ground truth is a labeled dataset, and labels carry their own errors.
  • For physical world data, ground truth is an observation with coordinates and a timestamp, and checking a model against it is a spatial join plus a time window.
  • In Wherobots, one query scored GraphCast's four-day rain forecast for Hurricane Helene at 29 rain gauges in Buncombe County, North Carolina, in 21 seconds: the forecast was below the observed total at all 29, with a median of 51% of what fell.
  • Errors measured against ground truth feed back into models through retraining, data assimilation, and the decision to trust a world model as a training environment.

What is ground truth?

Ground truth is the answer a prediction is judged against. The term comes from remote sensing, where analysts compared what a satellite or aerial image appeared to show with what a visit to the ground found. Russell Congalton's A review of assessing the accuracy of classifications of remotely sensed data (Remote Sensing of Environment, 1991) set out how to sample those reference sites and build the error matrix that compares a classified map with them, and that method is still how land cover maps report accuracy.

Three properties make a reference usable as ground truth:

  1. Independent. It was measured separately from the model. A forecast scored against a later run of the same forecast system measures consistency, and says nothing about the world.
  2. Located and timed. It carries the place and time it describes, so it can be paired with the prediction for that place and time.
  3. Quality controlled. Its own error is known and small compared with the model error being measured.

Ground truth in machine learning

In supervised machine learning, ground truth is the labeled dataset: each example paired with its correct answer. A model is fit to a training split, tuned on a validation split, and scored once on a test split it has never seen. The test score means something only when the test labels are correct.

They often are not. Curtis Northcutt, Anish Athalye, and Jonas Mueller's Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks (NeurIPS 2021) found an average of at least 3.3% label errors across the test sets of 10 widely used computer vision, language, and audio benchmarks, including at least 6% of the ImageNet validation set. A model that matches a wrong label is counted as correct, and one that gets the true answer is counted as wrong.

For imagery, labels are drawn on a map: a building outline traced over an aerial photo, a land cover class assigned to a field plot. The building footprint and land cover classification articles cover how those reference sets are built and how a model's output polygons are scored against them.

Ground truth for physical world data

For a forecast or a world model of a real place, ground truth is an observation of the physical world: what a station, a gauge, a survey, or a sensor recorded at a known location and time. Allan Murphy and Robert Winkler's A general framework for forecast verification (Monthly Weather Review, 1987) defined verification as the study of the joint distribution of forecasts and observations, which only exists once each forecast is paired with the observation it predicted.

Two kinds of data look like ground truth and are not:

  • Warnings and forecasts. A flood warning polygon is a prediction of where flooding will happen. The gauge reading inside it is the observation.
  • Reanalysis. ERA5, described by Hans Hersbach and colleagues in The ERA5 global reanalysis (Quarterly Journal of the Royal Meteorological Society, 2020), reconstructs past weather on a grid by combining a forecast model with historical observations. Learned weather models train on it. At a station it can differ from what was measured: Vivek Ramavajjala and Peetak Mitra's Verification against in-situ observations for Data-Driven Weather Prediction (2023) found that FourCastNet outperforms the IFS model when tested against ERA5 but shows no benefit when tested against real-world observations, and Microsoft researchers built WeatherReal (2024) on station observations because reanalysis diverges from them for near-surface temperature, wind, precipitation, and clouds.
ReferenceWhat it isUse as ground truth
Station or gauge recordA measurement at a point and timeYes, for that point; quality flags show which values passed checks
Field survey or traced labelA person's record of what is at a placeYes, with its labeling error measured
Logged sensor data and trajectoriesWhat a vehicle or robot recordedYes, for the places and conditions it covered
Reanalysis gridModel plus observations, on a gridA training target and a gridded reference, with known differences from stations
Forecast, warning, or model outputA predictionNo

How to check a model against ground truth

The same five steps apply to a weather forecast, a land cover map, and a driving simulator.

  1. Define the prediction. State the quantity, place, and time window exactly: total rain from 25 to 28 September 2024 in one 0.25° forecast cell, or the class of one pixel.
  2. Collect independent observations of the same quantity. Keep only quality-controlled values and record which places have gaps.
  3. Match prediction to observation by place and time. For spatial data this is a spatial join: a gauge point inside a forecast cell, a building outline inside an image tile. Align time windows to the observation's own day or interval.
  4. Score against a reference. Compute the error, then compare it with a simple reference forecast such as persistence (the recent past repeats) or climatology. Murphy's skill score based on mean square error (Monthly Weather Review, 1988) is one minus the ratio of the model's mean squared error to the reference's: 1 is perfect, 0 is no better than the reference.
  5. Feed the error back. Retrain, correct the next starting state, or restrict where the model is trusted.
Diagram of five steps for checking a model against ground truth. Prediction: GraphCast forecast 182.8 millimeters in the airport's cell, or a simulated rare driving scenario. Ground truth: the airport gauge recorded 355.1 millimeters from 25 to 28 September, or logged trajectories and sensor data. Match place and time: a gauge point inside a 0.25 degree cell, or a scene matched to the road it depicts. Score: mean absolute error 185.2 millimeters and skill 0.71 against persistence, or the likelihood of logged real behavior. A fifth bar, feed the error back, lists retraining, assimilation, and training inside the model, with an arrow back to the first step
Checking a model against ground truth, with the Helene numbers from this article and driving examples from the Waymo Open Sim Agents Challenge. Schematic.

Ground truth examples

  • Weather forecasts. Rain gauges and weather stations. NOAA's Global Historical Climatology Network Daily, described by Matthew Menne and colleagues in An overview of the Global Historical Climatology Network-Daily database (Journal of Atmospheric and Oceanic Technology, 2012), publishes quality-controlled daily records from land stations worldwide, with a flag on every value that failed a check.
  • Land cover maps. Field plots and photo-interpreted reference sites, scored with Congalton's error matrix.
  • Building and object detection. Traced building outlines. How well does SAM3 detect building footprints? compared roughly 312,000 detected roofs with Overture building footprints as an independent reference, and notes that Overture is not ground truth in rural areas, where it misses buildings.
  • Driving simulators. Logged trajectories from real vehicles. Nico Montali and colleagues' The Waymo Open Sim Agents Challenge (NeurIPS 2023) scores a simulator by the likelihood of real logged trajectories under the behavior its simulated agents produce.
  • Hazard and loss models. Observed hazard footprints and insurance claims, the references a catastrophe model is calibrated against.

Ground truth in Wherobots

Wherobots is the AI Context Engine for Physical World Data. WherobotsDB runs spatial SQL across vector and raster data and reads public archives in place, the Havasu catalog holds open datasets as Apache Iceberg tables, and the Wherobots MCP server lets AI agents run the same queries. RasterFlow (Public Preview) builds mosaics from imagery, runs computer vision inference on them, and writes the results as vectorized outputs on the same platform.

This walkthrough scores a learned weather model against ground truth for one extreme event. Hurricane Helene caused catastrophic inland flooding in western North Carolina in September 2024 (NHC Tropical Cyclone Report). The world model article assembled the observed state of Buncombe County from GHCN-Daily gauges for that storm. Here the prediction is GraphCast, published by Remi Lam and colleagues as Learning skillful medium-range global weather forecasting (Science, 2023), from the run NOAA's National Weather Service made from GFS initial conditions at 00 UTC on 25 September 2024, the day the heavy rain began.

Getting the forecast cells

NOAA publishes each GraphCast GFS run on AWS as GRIB2 files, one per 6-hour forecast step, each about 266 MB, with an index that lists the byte range of every field. WherobotsDB reads GeoTIFF and NetCDF rasters and has no GRIB2 reader, and NOAA's machine learning weather prediction archive, which holds GraphCast runs from the same GFS initial conditions, stores each run as one 9.34 GB NetCDF file, which failed to load as a single binary value. So the 16 forecast cells around the county were taken from the 6-hourly total precipitation fields of the 25 September run, fetched by byte range (about 3.2 MB each) and decoded with ECMWF's ecCodes library:

# 0-96 h total precipitation (mm) from NOAA GraphCastGFS, run 2024-09-25 00 UTC
import requests, eccodes
base = ("https://noaa-nws-graphcastgfs-pds.s3.amazonaws.com/graphcastgfs.20240925/00/"
        "forecasts_13_levels/graphcastgfs.t00z.pgrb2.0p25.f")
total = 0
for h in range(6, 97, 6):
    idx = requests.get(f"{base}{h:03d}.idx").text.splitlines()
    i = next(n for n, l in enumerate(idx) if f"APCP:surface:{h-6}-{h} hour acc" in l)
    start, end = int(idx[i].split(":")[1]), int(idx[i + 1].split(":")[1]) - 1
    msg = requests.get(f"{base}{h:03d}", headers={"Range": f"bytes={start}-{end}"}).content
    g = eccodes.codes_new_from_message(msg)
    total = total + eccodes.codes_get_values(g).reshape(721, 1440)  # 0.25° grid, 90N to 90S, 0 to 360E
    eccodes.codes_release(g)
# Grid point (lat, lon): total[round((90 - lat) / 0.25), round((lon % 360) / 0.25)]

The 16 values go into the query as a table. Everything after that, including every number below, comes from one WherobotsDB query.

Scoring the forecast at every gauge

The query reads the GHCN-Daily station list and the 2024 daily file in place from NOAA's public bucket, keeps the stations inside the TIGER/Line county polygon, and joins each gauge to the forecast cell that contains it. It keeps gauges with quality-controlled rain on all eight days from 21 to 28 September, so each has a complete storm total for 25 to 28 September and a persistence forecast: the gauge's own total for the four days before.

-- Ground truth for a world model: Hurricane Helene rain at every complete Buncombe County gauge,
-- 25 to 28 September 2024, scored against the GraphCast forecast from 00 UTC 25 September
-- (NOAA GraphCastGFS, 0 to 96 h total precipitation on its 0.25-degree grid) and against
-- persistence (the gauge's own total for the previous four days)
WITH gc(lat, lon, gc_mm) AS (
  -- GraphCast grid points around the county, decoded from NOAA's GraphCastGFS GRIB2 files
  VALUES
    (35.25, -83.0, 217.3), (35.25, -82.75, 212.8), (35.25, -82.5, 203.1), (35.25, -82.25, 190.2),
    (35.5, -83.0, 191.5), (35.5, -82.75, 192.7), (35.5, -82.5, 182.8), (35.5, -82.25, 184.9),
    (35.75, -83.0, 169.3), (35.75, -82.75, 163.3), (35.75, -82.5, 166.2), (35.75, -82.25, 176.3),
    (36.0, -83.0, 144.2), (36.0, -82.75, 153.0), (36.0, -82.5, 157.9), (36.0, -82.25, 155.5)
),
cell AS (
  -- Each grid point's cell: 0.25 degrees on a side, centered on the point
  SELECT gc_mm, ST_MakeEnvelope(lon - 0.125, lat - 0.125, lon + 0.125, lat + 0.125) AS geom
  FROM gc
),
st AS (
  SELECT substr(value, 1, 11) AS station_id,
         CAST(trim(substr(value, 13, 8)) AS DOUBLE) AS lat,
         CAST(trim(substr(value, 22, 9)) AS DOUBLE) AS lon,
         trim(substr(value, 42, 30)) AS name
  FROM text.`s3://noaa-ghcn-pds/ghcnd-stations.txt`
  WHERE substr(value, 39, 2) = 'NC'
),
stn AS (
  SELECT /*+ BROADCAST(c) */ st.*
  FROM st JOIN wherobots_open_data.us_census.tiger_county c
    ON ST_Intersects(c.geometry, ST_Point(st.lon, st.lat))
  WHERE c.STATEFP = '37' AND c.NAME = 'Buncombe'
),
obs AS (
  -- Quality-controlled daily rain (tenths of mm) for 21 to 28 September
  SELECT /*+ BROADCAST(s) */ o._c0 AS station_id, to_date(o._c1, 'yyyyMMdd') AS d,
         CAST(o._c3 AS DOUBLE) / 10 AS mm
  FROM csv.`s3://noaa-ghcn-pds/csv/by_year/2024.csv` o
  JOIN stn s ON o._c0 = s.station_id
  WHERE o._c2 = 'PRCP' AND o._c5 IS NULL AND o._c1 BETWEEN '20240921' AND '20240928'
),
tot AS (
  SELECT station_id,
         SUM(IF(d >= DATE '2024-09-25', 1, 0)) AS event_days,
         SUM(IF(d >= DATE '2024-09-25', mm, 0)) AS obs_mm,
         SUM(IF(d < DATE '2024-09-25', 1, 0)) AS prior_days,
         SUM(IF(d < DATE '2024-09-25', mm, 0)) AS persist_mm
  FROM obs GROUP BY station_id
),
scored AS (
  -- Point-in-cell spatial join: each gauge meets the forecast cell it sits in
  SELECT /*+ BROADCAST(c) */ s.station_id, s.name, t.obs_mm, c.gc_mm, t.persist_mm,
         c.gc_mm - t.obs_mm AS gc_err, t.persist_mm - t.obs_mm AS persist_err
  FROM stn s
  JOIN tot t ON s.station_id = t.station_id
  JOIN cell c ON ST_Intersects(c.geom, ST_Point(s.lon, s.lat))
  WHERE t.event_days = 4 AND t.prior_days = 4
)
SELECT COUNT(*) AS gauges,
       COUNT(DISTINCT station_id) AS distinct_gauges,
       ROUND(PERCENTILE(obs_mm, 0.5), 1) AS median_obs_mm,
       ROUND(PERCENTILE(gc_mm, 0.5), 1) AS median_graphcast_mm,
       ROUND(AVG(ABS(gc_err)), 1) AS graphcast_mae_mm,
       ROUND(AVG(gc_err), 1) AS graphcast_bias_mm,
       ROUND(SQRT(AVG(gc_err * gc_err)), 1) AS graphcast_rmse_mm,
       ROUND(AVG(ABS(persist_err)), 1) AS persistence_mae_mm,
       ROUND(SQRT(AVG(persist_err * persist_err)), 1) AS persistence_rmse_mm,
       ROUND(1 - AVG(gc_err * gc_err) / AVG(persist_err * persist_err), 2) AS mse_skill_vs_persistence,
       SUM(IF(gc_err < 0, 1, 0)) AS gauges_underforecast,
       ROUND(MIN(gc_mm / obs_mm), 2) AS min_share_forecast,
       ROUND(PERCENTILE(gc_mm / obs_mm, 0.5), 2) AS median_share_forecast,
       ROUND(MAX(gc_mm / obs_mm), 2) AS max_share_forecast,
       MAX_BY(name, -gc_err) AS worst_gauge,
       ROUND(MAX(obs_mm), 1) AS worst_obs_mm,
       ROUND(MAX_BY(gc_mm, -gc_err), 1) AS worst_forecast_mm,
       ROUND(MAX(IF(station_id = 'USW00003812', obs_mm, NULL)), 1) AS airport_obs_mm,
       ROUND(MAX(IF(station_id = 'USW00003812', gc_mm, NULL)), 1) AS airport_forecast_mm
FROM scored

It ran in 21 seconds and returned one row:

MeasureValue
Complete gauges scored, each matched to one cell29
Median rain observed, 25 to 28 September356.7 mm
Median GraphCast forecast at those gauges182.8 mm
GraphCast mean absolute error; bias185.2 mm; -185.2 mm
Persistence mean absolute error356.5 mm
Skill against persistence (mean squared error)0.71
Gauges where GraphCast forecast less than fell29 of 29
GraphCast share of observed rain: lowest, median, highest36%, 51%, 70%
Largest missBlack Mountain 2.1 W: 516.4 mm observed, 184.9 mm forecast
Asheville Regional Airport355.1 mm observed, 182.8 mm forecast

GraphCast placed heavy rain over the county four days out, and against persistence, which forecast little rain because the days before the storm were nearly dry, it scores a skill of 0.71. Against the gauges it forecast about half of what fell, and less than fell at every one of them. The bias equals the mean absolute error, so every error has the same sign. EXPLAIN FORMATTED shows R-tree index joins (BroadcastIndexJoin) for the stations to the county and the gauges to the cells, a BroadcastHashJoin from the stations to the CSV rows, and no CartesianProduct.

Bar chart of 29 Buncombe County rain gauges sorted by rain observed from 25 to 28 September 2024, from 516.4 millimeters at Black Mountain 2.1 W down to 238.5 millimeters. A purple mark on each bar shows the GraphCast forecast for the gauge's cell, between about 163 and 193 millimeters, below every bar. Orange dots near zero show the persistence forecast. Callouts mark Black Mountain 2.1 W, 516.4 fell and GraphCast 184.9, and Asheville Regional Airport, 355.1 fell and GraphCast 182.8. A side panel lists median observed 356.7 millimeters, median GraphCast 182.8, GraphCast mean absolute error 185.2, persistence 356.5, and skill 0.71
GraphCast’s Helene forecast scored at every complete gauge: below the observed total at all 29, median 51% of what fell. Source: NOAA GHCN-Daily and NOAA GraphCastGFS in Wherobots; one query, 21 s.

The map shows the scale mismatch. A 0.25° cell is about 28 km north to south here and carries one forecast value, while the gauges inside it recorded different totals: the cell containing Black Mountain 2.1 W forecast 184.9 mm where that gauge recorded 516.4 mm.

Map of Buncombe County, North Carolina, overlaid on a grid of GraphCast 0.25 degree cells colored by forecast rain from 163 to 193 millimeters, all in shades of blue. Dots mark 29 rain gauges colored on the same scale by observed rain, from light blue near 240 millimeters to orange and red near 500, each warmer than the cell around it. Labels mark Asheville Regional Airport, 355.1 millimeters fell with a cell forecast of 183, and Black Mountain 2.1 W, 516.4 fell with a cell forecast of 185
GraphCast cells and the gauges inside them on one color scale: every gauge recorded more than its cell’s forecast. Source: NOAA GraphCastGFS, NOAA GHCN-Daily, US Census TIGER/Line in Wherobots.

How ground truth improves models

Scoring is the start. The error measured at known places and times goes back into the model in three ways.

Retraining and fine-tuning. Supervised models, including learned weather models such as GraphCast, are trained to reduce the gap between prediction and target. When the target is reanalysis, errors at stations can survive training, which is why the observation-based benchmarks above exist. Ground truth for extremes like Helene shows where a model needs more or better training examples.

Data assimilation. Forecast systems correct their starting state with new observations every cycle. Geir Evensen's Sequential data assimilation with a nonlinear quasi-geostrophic model using Monte Carlo methods to forecast error statistics (Journal of Geophysical Research, 1994) introduced the ensemble Kalman filter, which systems such as the HRRR model use to blend forecasts with radar and surface observations.

Training inside a world model. In model-based reinforcement learning, an agent learns a policy inside a learned world model. Danijar Hafner and colleagues' DreamerV3, in Mastering diverse control tasks through world models (Nature, 2025), trains behavior on imagined rollouts. The policy learns whatever the model predicts, including its mistakes. David Ha and Jürgen Schmidhuber's World Models (2018) describe a controller that found an adversarial policy inside an imperfect model, moving so that monsters in the simulated game never fired, an exploit absent from the real game. Ground truth measures that gap before training starts. For driving world models, the reference is real logged behavior, as in the Waymo Open Sim Agents Challenge, and for world models of real places it extends to the place itself: road geometry, terrain, and the weather recorded there on the day a scene depicts. The physical AI article covers how these models train robots and vehicles.

Ground truth research and what changes at scale

Ground truth as a formal practice grew from two fields. In remote sensing, Congalton's 1991 review made the error matrix and reference sampling the standard for map accuracy. In weather, Murphy and Winkler's 1987 framework treated verification as the joint distribution of forecasts and observations, and Murphy's 1988 skill score tied it to a reference forecast. Machine learning added held-out test sets and, with Northcutt and colleagues in 2021, measured how often the labels in them are wrong. Learned weather models brought the two traditions together: WeatherBench 2, by Stephan Rasp and colleagues in WeatherBench 2: A benchmark for the next generation of data-driven global weather models (Journal of Advances in Modeling Earth Systems, 2024), scores them on shared grids, and station benchmarks such as WeatherReal score them where people live.

What scale changes. One county, one storm, and one forecast run make a single query. A season of twice-daily runs scored at every gauge in a country, or every simulated scene matched to the road it depicts, multiplies both sides of the join, and the matching has to run where the observations and predictions are stored. Spatial partitioning and per-partition indexes, introduced for Apache Spark by Jia Yu, Jinxuan Wu, and Mohamed Sarwat in GeoSpark (ACM SIGSPATIAL 2015), keep each comparison to nearby candidates. The partitioning tax article covers what happens to joins like this when the data is split into tiles instead.

Open problems

  • Points against cells. A gauge measures a point and a 0.25° cell here averages about 630 km². Part of the 185.2 mm error is the difference between a point and an area, and separating it from model error needs denser observations or downscaling.
  • Uneven coverage. Of the 139 stations ever listed in Buncombe County, 29 had complete records for these eight days, and gauges cluster near towns. Ground truth is thinnest in the mountains, where the most rain fell.
  • Time windows. Many volunteer gauges report the 24 hours ending at 7 a.m. local time. Totals over several days reduce the mismatch with a forecast window measured in UTC, and single-day scores need each gauge's observation time.
  • One event. A single storm is an example. A verdict on a model needs many events and the stochastic view of how often extremes like this occur.

Read more from Wherobots

Score your own models against ground truth with a Wherobots free trial at cloud.wherobots.com.

Frequently asked questions

What is ground truth?

Ground truth is information known to be true because it was observed or measured directly, used as the reference that a model is trained on or checked against. In machine learning it is usually a set of labels. For physical world data it is an observation with a known place and time, such as a rain gauge reading, a surveyed field plot, or a logged vehicle trajectory.

What does ground truth mean?

The term comes from remote sensing, where it meant information collected on the ground to check what a satellite or aerial image appeared to show. It now means any trusted reference for judging a prediction: if a model says a pixel is forest, the ground truth is what a visit to that spot found.

What is ground truth data in machine learning?

In machine learning, ground truth data is the set of correct answers for a task: the label for each image, the class for each pixel, or the value a model should predict. Models are trained on part of it and evaluated on a held-out part. Labels come from people or instruments and contain errors: Curtis Northcutt, Anish Athalye, and Jonas Mueller found an average of at least 3.3% label errors across the test sets of 10 widely used benchmarks in Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks (NeurIPS 2021).

What is a ground truth dataset?

A ground truth dataset is a collection of trusted reference values used to train or score a model. Examples include labeled image sets, field survey plots used to check land cover maps, and networks of weather stations such as NOAA’s Global Historical Climatology Network, which holds daily records from land stations worldwide.

What are examples of ground truth?

A rain gauge total checks a weather forecast, a surveyed field plot checks a land cover classification, a hand-traced roof outline checks a building footprint model, an insurance claim checks a damage estimate, and a logged trajectory from a real vehicle checks a driving simulator. Each pairs a prediction with an observation of the same thing at the same place and time.

What is the difference between ground truth and labels?

A label is one recorded answer, usually made by a person or an instrument. Ground truth is the reference those labels are meant to capture. Labels approximate ground truth: two annotators can disagree, and a label can be wrong. Treating labels as perfect ground truth hides their error inside the model’s measured error.

How do you check a model against ground truth?

Define exactly what the model predicts, collect independent observations of the same quantity, match each prediction to the observation at the same place and time, and compute an error against a simple reference forecast such as persistence. For spatial data the matching step is a spatial join: a gauge point inside a forecast cell, or a building outline inside an image tile.

Is reanalysis ground truth?

Reanalysis such as ERA5 is a gridded reconstruction of past weather that combines a model with historical observations, so it is a model product. Learned weather models are trained and often scored on it, and it can differ from what a station measured: Vivek Ramavajjala and Peetak Mitra found that FourCastNet outperforms IFS when tested against ERA5 but shows no benefit when tested against real-world observations, in Verification against in-situ observations for Data-Driven Weather Prediction (2023).

Why do world models need ground truth?

A world model predicts how an environment changes, and agents can be trained inside it. If its predictions drift from reality, anything trained inside it learns the drift. Ground truth at known places and times measures that drift so the model can be corrected before it is used for training or decisions.

How does Wherobots help with ground truth?

Wherobots is the AI Context Engine for Physical World Data. WherobotsDB reads observations such as NOAA station records in place, joins them to boundaries, forecast cells, buildings, and imagery in spatial SQL, and scores predictions at every location in one query. RasterFlow (Public Preview) builds mosaics from imagery, runs computer vision inference on them, and writes the results as vectorized outputs on the same platform, where they can be checked against reference data.