Rasterflow, Earth Intelligence & inference engine now in public preview Learn More

What is a World Model? How AI Predicts the Physical World

Authors

A world model is an AI model that predicts how an environment changes over time, so a system can test actions or produce a forecast before anything happens in the real world. It encodes what it observes into a state, predicts the next state from that state and an action, and chains the predictions forward. Game-playing agents, robots, driving simulators, and learned weather models all use the idea. For real places, a world model is only as good as the observations it learns from and is checked against: measurements with known coordinates and timestamps.

Key takeaways

  • A world model predicts the next state of an environment. Planning, imagination, and forecasting all come from chaining those predictions.
  • The term has three common meanings today: learned dynamics models inside agents, world foundation models that generate video and 3D scenes, and the mental model idea from cognitive science it grew out of.
  • Learned weather models such as GraphCast and GenCast are world models of the atmosphere, and numerical weather prediction is the physics-based version.
  • Every world model is scored against observations it never saw. In Wherobots, one query assembled the observed state of Buncombe County, North Carolina, during Hurricane Helene from 61 rain gauges in 23 seconds, and showed a persistence forecast at Asheville Regional Airport missing by 98.6 mm on the gauge's first day of heavy rain.

What is a world model?

A world model is a model of an environment's dynamics: given the current state, and an action when there is an agent, it predicts what comes next. The prediction can be a compressed vector, a video frame, a 3D scene, or a grid of atmospheric variables. Search results for the term show three meanings.

Machine learning world models. In reinforcement learning, a world model is the learned part of an agent that predicts the consequences of actions. David Ha and Jürgen Schmidhuber's World Models (2018) compressed game frames into a small code, learned to predict the next code, and trained a controller entirely inside that model's "dream" of the VizDoom game before transferring it back to the real game.

World foundation models and world generators. Since 2024 the term also names large models that generate interactive video or 3D environments. Google DeepMind's Genie 3 (August 2025) generates interactive environments from a text prompt at 720p and 24 frames per second, NVIDIA's Cosmos platform (2025) offers pretrained world foundation models for robots and vehicles, and World Labs introduced Atlas on 1 September 2026 as an omni world model, pretrained to operate on text, images, video, and 3D, that generates, reconstructs, and simulates 3D worlds. The physical AI article covers how these models train robots and autonomous vehicles.

The mental model origin. The idea is older than computing. Kenneth Craik argued in The Nature of Explanation (Cambridge University Press, 1943) that an organism carrying a "small-scale model" of external reality and of its own possible actions can try out alternatives and react to situations before they arise. Ha and Schmidhuber open their paper with Jay Forrester's 1971 observation that the image of the world we carry in our heads is a model built from selected concepts and the relationships between them, and Philip Johnson-Laird's Mental Models (Harvard University Press, 1983) built a theory of reasoning on the same idea.

How a world model works

Diagram of six steps in a world model. Observe takes camera frames and joint angles for an agent, and stations, satellites, radar and reanalysis for the Earth system. Encode state produces a latent vector per frame or a grid of variables on a 0.25 degree mesh. Predict next state gives the next latent after an action or the atmosphere 6 hours later. Plan or roll out tries actions in imagination or steps forward to a 10-day forecast. Act or forecast picks the best action or issues guidance for real places. A sixth bar, compare with what happened, loops back to observe
The world model loop, with agent examples from Dreamer and Earth system examples from GraphCast. Step 6 needs observations the model never saw. Schematic.
  1. Observe. The model takes in sensor data and records: camera frames and joint angles for a robot, or station, radar, and satellite measurements for the atmosphere. Each observation has a time, and for real places, a location.
  2. Encode the state. An encoder compresses observations into a state the model can update. Agents use a latent vector; weather models use a grid of variables such as temperature and wind on a global mesh. Rudolf Kalman's A New Approach to Linear Filtering and Prediction Problems (Journal of Basic Engineering, 1960) set out the state-transition view of estimating a hidden state from noisy measurements. Its descendants, such as the ensemble Kalman filter in the HRRR data assimilation system, combine model predictions with new observations every forecast cycle.
  3. Predict the next state. A transition or dynamics model predicts the state one step later, given the current state and an action. In a learned model this function comes from training data; in a physics simulator it comes from equations.
  4. Plan or roll out. Chaining predictions produces an imagined trajectory. An agent compares many candidate actions inside the model, which is cheaper and safer than trying each one; a weather model steps forward to produce a forecast days ahead.
  5. Act or forecast. The agent takes the best action, or the forecast is issued for real places.
  6. Compare with what happened. New observations arrive, and the gap between prediction and observation, measured against ground truth, is the error that retraining and assimilation work to reduce.

Richard Sutton's Dyna architecture (1991) already combined these steps: it learned a model of the effects of actions, planned with that model even when it was imperfect, and acted reactively.

Types of world models

Latent dynamics models learn a compact state from pixels and predict forward in that latent space. PlaNet, Dreamer, DreamerV3, and MuZero belong here. They train on one environment's observations, actions, and rewards.

Video and 3D world generators learn from large video collections and output frames or scenes. Genie, Cosmos, Genie 3, the Waymo World Model, and Atlas belong here. Their state is a sequence of images or a 3D scene, and control comes from actions, text, or camera paths.

Physics simulators and numerical weather prediction step equations of fluid motion and thermodynamics forward from an analysis of current observations. NOAA's HRRR model runs every hour on a 3 km grid over the contiguous United States. Peter Bauer, Alan Thorpe, and Gilbert Brunet describe the steady gains of this approach in The quiet revolution of numerical weather prediction (Nature, 2015). Digital twins of a factory, city, or planet apply simulation to one specific place.

Learned weather and Earth system models replace the hand-written equations with a neural network trained on decades of reanalysis, a gridded reconstruction of past weather. Remi Lam and colleagues' GraphCast, published as Learning skillful medium-range global weather forecasting (Science, 2023), predicts hundreds of variables 10 days ahead at 0.25° resolution in under a minute and outperformed the leading operational deterministic system on 90% of 1,380 verification targets.

TypeExamplesState it learns fromWhat it is checked against
Latent dynamics modelWorld Models, PlaNet, Dreamer, MuZeroFrames, actions, and rewards from one environmentHeld-out episodes and task returns
Video and 3D world generatorGenie 3, Cosmos, Waymo World Model, AtlasLarge video collections; camera and lidar logs for drivingHeld-out real video and sensor logs
Physics simulator and NWPHRRR, ECMWF IFS, digital twinsEquations, plus radar, satellite, aircraft, and surface observations assimilated each cycleStation, radar, and satellite observations; later analyses
Learned weather and Earth modelGraphCast, Pangu-Weather, GenCast, AuroraReanalysis grids built from decades of observationsReanalysis and observations at stations, scored against a reference forecast

Sources: the papers and releases linked in this section and in the research section below. The last column names the data each model type is evaluated on in its own publications.

World model vs LLM

A large language model predicts the next token of text and was trained on text, documents, databases, and the internet. A world model predicts the next state of an environment. The two overlap at the edges: Genie 3 takes a text prompt, and Meta's V-JEPA 2 aligns a video world model with a language model for video question answering. Yann LeCun's position paper A Path Towards Autonomous Machine Intelligence (2022) argues that machines that reason and plan need a predictive world model trained by self-supervised learning on observation, built on joint embedding predictive architectures (JEPA). A language model can describe a flood; a world model of a river basin predicts where the water goes next, and that prediction can be scored against a gauge.

World model examples

  • Minecraft. Danijar Hafner and colleagues' DreamerV3, published as Mastering diverse control tasks through world models (Nature, 2025), used one configuration across more than 150 tasks and, by the authors' account, was the first algorithm to collect diamonds in Minecraft from scratch without human data.
  • Board games and Atari. Julian Schrittwieser and colleagues' MuZero, in Mastering Atari, Go, chess and shogi by planning with a learned model (Nature, 2020), matched AlphaZero in Go, chess, and shogi without being given the rules.
  • Robot arms. Meta's V-JEPA 2 (2025) pretrained on more than 1 million hours of internet video, then post-trained an action-conditioned world model on less than 62 hours of robot video and used it to plan pick-and-place tasks on Franka arms in two labs.
  • Driving simulation. The Waymo World Model (February 2026) adapts Genie 3 to generate camera and lidar output for rare driving scenarios.
  • Global weather. GraphCast, Pangu-Weather, GenCast, and Aurora forecast the atmosphere days ahead. NOAA runs several learned models, including GraphCast, Pangu-Weather, and FourCastNet, twice a day from NOAA GFS and ECMWF IFS initial conditions and publishes the forecasts in the NOAA machine learning weather prediction archive.
  • Regional weather. The HRRR model is a physics-based world model of the US atmosphere, rerun hourly from new radar and surface observations.

To build one, collect timestamped observations, choose a state representation, train a transition model, roll it forward, and score it on held-out observations.

Why world models of real places need physical world data

A world model of a real place is trained on observed state and checked against observed state. Both come from physical world data: geometry, imagery, elevation, and weather records, each with known coordinates and timestamps. A learned weather model trains on reanalysis, and its forecast for a town is judged by the gauge in that town. Without the location of each observation, a prediction cannot be matched to what happened there.

Forecasts, warnings, and model outputs are predictions, so they never count as the observation a model is scored against. Verification compares a forecast with observations at the same place and time, and a skill score measures how much it improves on a reference forecast such as climatology or persistence. Allan Murphy formalized skill scores based on mean square error in Skill scores based on the mean square error and their relationships to the correlation coefficient (Monthly Weather Review, 1988).

That evaluation is a spatial join. Each gauge is a point, each forecast is a grid or a polygon, and each place of interest is a building, a road, or a county. Assembling observed state for one area means joining point records to boundaries, rasters to points, and warnings to locations, the same computation a context engine runs for AI agents.

World model in Wherobots

Wherobots is the AI Context Engine for Physical World Data. WherobotsDB runs spatial SQL across vector and raster data, the Havasu catalog holds open datasets as Apache Iceberg tables, and the Wherobots MCP server lets AI agents list tables and run queries. RasterFlow (Public Preview) builds mosaics from imagery, runs computer vision inference on them, and writes the results as vectorized outputs on the same platform.

This walkthrough builds the observed state a world model of Buncombe County, North Carolina, would be checked against during Hurricane Helene, which caused catastrophic inland flooding there in September 2024 (NHC Tropical Cyclone Report). The observations come from NOAA's Global Historical Climatology Network Daily (GHCN-Daily), described by Matthew Menne and colleagues in An overview of the Global Historical Climatology Network-Daily database (Journal of Atmospheric and Oceanic Technology, 2012) and published on AWS through the Registry of Open Data. WherobotsDB reads the station list and the 2024 CSV file in place, keeps the stations inside the TIGER/Line county polygon, and joins them to the NWS flash flood warnings issued during the storm. Each gauge's record holds its storm rain, its 2024 total, a persistence forecast error for daily maximum temperature, and its warning count.

-- Observed state for a world model of Buncombe County, North Carolina, in 2024:
-- every GHCN-Daily station inside the county, its quality-controlled rain and
-- maximum temperature, the error of a persistence forecast (tomorrow = today),
-- and the NWS flash flood warnings (forecasts) that covered it during Helene
WITH st AS (
  -- GHCN-Daily station list: fixed-width text with ID, latitude, longitude, elevation, state, name
  SELECT substr(value, 1, 11) AS station_id,
         CAST(trim(substr(value, 13, 8)) AS DOUBLE) AS lat,
         CAST(trim(substr(value, 22, 9)) AS DOUBLE) AS lon,
         CAST(trim(substr(value, 32, 6)) AS DOUBLE) AS elev_m,
         trim(substr(value, 42, 30)) AS name
  FROM text.`s3://noaa-ghcn-pds/ghcnd-stations.txt`
  WHERE substr(value, 39, 2) = 'NC'
),
stn AS (
  -- Stations whose coordinates fall inside the TIGER/Line Buncombe County polygon
  SELECT /*+ BROADCAST(c) */ st.*
  FROM st JOIN wherobots_open_data.us_census.tiger_county c
    ON ST_Intersects(c.geometry, ST_Point(st.lon, st.lat))
  WHERE c.STATEFP = '37' AND c.NAME = 'Buncombe'
),
obs AS (
  -- 2024 daily observations, read once: PRCP in tenths of mm, TMAX in tenths of degrees C
  -- A blank quality flag (_c5) means the value passed NOAA's quality checks
  SELECT /*+ BROADCAST(s) */ o._c0 AS station_id, to_date(o._c1, 'yyyyMMdd') AS d,
         o._c2 AS element, CAST(o._c3 AS DOUBLE) / 10 AS v
  FROM csv.`s3://noaa-ghcn-pds/csv/by_year/2024.csv` o
  JOIN stn s ON o._c0 = s.station_id
  WHERE o._c2 IN ('PRCP', 'TMAX') AND o._c5 IS NULL
),
lagged AS (
  -- Persistence forecast: the next day's value equals the previous day's observed value
  SELECT obs.*,
         LAG(v) OVER (PARTITION BY station_id, element ORDER BY d) AS prev_v,
         LAG(d) OVER (PARTITION BY station_id, element ORDER BY d) AS prev_d
  FROM obs
),
agg AS (
  SELECT station_id,
         SUM(IF(element = 'PRCP', 1, 0)) AS prcp_days,
         ROUND(SUM(IF(element = 'PRCP', v, 0)), 1) AS prcp_2024_mm,
         SUM(IF(element = 'PRCP' AND d BETWEEN DATE '2024-09-25' AND DATE '2024-09-28', 1, 0)) AS helene_days,
         ROUND(SUM(IF(element = 'PRCP' AND d BETWEEN DATE '2024-09-25' AND DATE '2024-09-28', v, 0)), 1) AS helene_mm,
         SUM(IF(element = 'TMAX' AND datediff(d, prev_d) = 1, 1, 0)) AS tmax_pairs,
         SUM(IF(element = 'TMAX' AND datediff(d, prev_d) = 1, ABS(v - prev_v), 0)) AS tmax_abs_err_sum,
         ROUND(MAX(IF(element = 'TMAX' AND datediff(d, prev_d) = 1, ABS(v - prev_v), NULL)), 1) AS persist_max_err_c
  FROM lagged
  GROUP BY station_id
),
ffw AS (
  -- NWS flash flood warnings from the Greenville-Spartanburg office issued 25 to 28 September 2024
  SELECT ISSUED, geometry
  FROM wherobots_open_data.noaa.nws_watch_warnings
  WHERE VTEC_YEAR = 2024 AND WFO = 'GSP' AND PHENOM = 'FF' AND SIG = 'W'
    AND ISSUED >= TIMESTAMP '2024-09-25 00:00:00' AND ISSUED < TIMESTAMP '2024-09-29 00:00:00'
),
ffw_s AS (
  SELECT /*+ BROADCAST(w) */ s.station_id, COUNT(*) AS ffw_helene
  FROM stn s JOIN ffw w ON ST_Intersects(w.geometry, ST_Point(s.lon, s.lat))
  GROUP BY s.station_id
),
rec AS (
  SELECT s.station_id, s.name, s.lat, s.lon, s.elev_m,
         a.prcp_days, a.prcp_2024_mm, a.helene_days, a.helene_mm,
         a.tmax_pairs, a.tmax_abs_err_sum,
         ROUND(a.tmax_abs_err_sum / NULLIF(a.tmax_pairs, 0), 2) AS persist_mae_c,
         a.persist_max_err_c,
         COALESCE(f.ffw_helene, 0) AS ffw_helene
  FROM stn s
  JOIN agg a ON s.station_id = a.station_id
  LEFT JOIN ffw_s f ON s.station_id = f.station_id
  WHERE a.prcp_days > 0
)
SELECT station_id, name, lat, lon, elev_m, prcp_days, prcp_2024_mm, helene_days, helene_mm,
       tmax_pairs, persist_mae_c, persist_max_err_c, ffw_helene,
       (SELECT COUNT(*) FROM stn) AS stations_ever_in_county,
       COUNT(*) OVER () AS stations_reporting_prcp_2024,
       SUM(IF(helene_days = 4, 1, 0)) OVER () AS complete_helene,
       SUM(IF(helene_days = 4 AND helene_mm >= 254, 1, 0)) OVER () AS complete_10in_plus,
       SUM(IF(helene_days = 4 AND ffw_helene > 0, 1, 0)) OVER () AS complete_under_ffw,
       SUM(IF(tmax_pairs > 0, 1, 0)) OVER () AS tmax_stations,
       ROUND(SUM(tmax_abs_err_sum) OVER () / SUM(tmax_pairs) OVER (), 2) AS pooled_persist_mae_c
FROM rec
ORDER BY helene_days DESC, helene_mm DESC

The query ran in 23 seconds and read the worldwide 2024 file once from the public bucket. The GHCN-Daily station list places 139 stations inside the county polygon, and 61 of them reported quality-controlled rain in 2024. About two thirds of GHCN-Daily stations measure precipitation only, according to the AWS registry entry, and here only 6 also reported daily maximum temperature. 35 gauges reported all four days from 25 to 28 September, and 33 of those recorded 254 mm (10 inches) or more. The highest four-day total was 561.9 mm at Black Mountain 5.5 SE. All 35 complete gauges sat inside at least one flash flood warning issued in those days; the warnings are forecasts, so the gauge totals are the observations a flood model would be scored on. EXPLAIN FORMATTED on a draft with the same joins showed a broadcast R-tree join (BroadcastIndexJoin) for the station, county, and warning lookups and a BroadcastHashJoin from the stations to the CSV rows, with no CartesianProduct.

The record for Asheville Regional Airport shows what one row of observed state holds:

FieldValue
GHCN-Daily IDUSW00003812
Location, elevation35.4317, -82.5378; 645.6 m
Rain, 25 to 28 September 2024355.1 mm, all four days reported
Rain, 20241,678.4 mm over 366 reported days
Persistence forecast of daily maximum temperaturemean absolute error 2.99 °C over 365 day pairs; largest error 13.8 °C
NWS flash flood warnings covering the gauge, 25 to 28 September4
Map of Buncombe County, North Carolina, with GHCN-Daily rain gauges colored by rain from 25 to 28 September 2024 on a scale from 0 to 570 millimeters. Filled dots mark 35 gauges with all four days reported, with the highest, 561.9 millimeters at Black Mountain 5.5 SE, in the east. Hollow dots mark gauges with gaps. Labels mark Asheville downtown at 365.0 millimeters and Asheville Regional Airport at 355.1 millimeters. A side panel lists 139 stations ever in the county, 61 reporting rain in 2024, 35 complete during the storm, and 33 at 254 millimeters or more
The observed state of Buncombe County during Helene: 35 complete gauges, 33 with 254 mm or more over 25 to 28 September 2024. Source: NOAA GHCN-Daily and US Census TIGER/Line in Wherobots; one query, 23 s.

Scoring a forecast against the gauges

Persistence, the forecast that tomorrow equals today, is the simplest world model and the usual reference forecast. A second query lines up each day's observation at the airport with the previous day's value, and summarizes all county gauges by day:

-- Daily observed state across Buncombe County stations, 15 September to 5 October 2024,
-- with the Asheville Regional Airport value and its persistence forecast error
WITH st AS (
  SELECT substr(value, 1, 11) AS station_id,
         CAST(trim(substr(value, 13, 8)) AS DOUBLE) AS lat,
         CAST(trim(substr(value, 22, 9)) AS DOUBLE) AS lon
  FROM text.`s3://noaa-ghcn-pds/ghcnd-stations.txt`
  WHERE substr(value, 39, 2) = 'NC'
),
stn AS (
  SELECT /*+ BROADCAST(c) */ st.station_id
  FROM st JOIN wherobots_open_data.us_census.tiger_county c
    ON ST_Intersects(c.geometry, ST_Point(st.lon, st.lat))
  WHERE c.STATEFP = '37' AND c.NAME = 'Buncombe'
),
obs AS (
  SELECT /*+ BROADCAST(s) */ o._c0 AS station_id, to_date(o._c1, 'yyyyMMdd') AS d,
         o._c2 AS element, CAST(o._c3 AS DOUBLE) / 10 AS v
  FROM csv.`s3://noaa-ghcn-pds/csv/by_year/2024.csv` o
  JOIN stn s ON o._c0 = s.station_id
  WHERE o._c2 IN ('PRCP', 'TMAX') AND o._c5 IS NULL
    AND o._c1 BETWEEN '20240914' AND '20241005'
),
daily AS (
  SELECT d, element, COUNT(*) AS stations,
         ROUND(percentile(v, 0.5), 1) AS median_v, ROUND(MAX(v), 1) AS max_v,
         MAX(IF(station_id = 'USW00003812', v, NULL)) AS airport_v
  FROM obs
  GROUP BY d, element
),
lagged AS (
  -- Persistence: each day's forecast is the previous day's observation (14 September feeds 15 September)
  SELECT d, element, stations, median_v, max_v, airport_v,
         LAG(airport_v) OVER (PARTITION BY element ORDER BY d) AS airport_persistence_forecast
  FROM daily
)
SELECT d, element, stations, median_v, max_v, airport_v, airport_persistence_forecast,
       ROUND(airport_v - airport_persistence_forecast, 1) AS airport_persistence_error
FROM lagged
WHERE d >= DATE '2024-09-15'
ORDER BY element, d

It ran in 15.8 seconds. The airport gauge recorded 103.9 mm on 25 September, 146.8 mm on 26 September, 104.4 mm on 27 September, and 0.0 mm on 28 September. Persistence forecast 5.3 mm for 25 September and missed by 98.6 mm, then forecast 104.4 mm for 28 September, when no rain fell. Across the county, the median gauge recorded 141.2 mm on 26 September, and the highest single-day reading was 276.1 mm on 27 September. For daily maximum temperature, persistence had a mean absolute error of 3.17 °C across the 6 temperature stations over 2024 (first query). A learned or physics-based model has to beat those numbers at these gauges on days like these.

Bar chart of daily rain at Asheville Regional Airport from 15 September to 5 October 2024, with faint bars for the highest gauge in Buncombe County and an orange line for the persistence forecast. The 25 to 28 September window is shaded and labeled 355.1 millimeters observed. Callouts mark the persistence forecast missing by 98.6 millimeters on 25 September, forecasting 5.3 when 103.9 fell, and being off by 104.4 millimeters on 28 September, forecasting 104.4 when 0.0 fell
A persistence forecast scored against the Asheville Regional Airport gauge during Helene: errors of 98.6 mm on 25 September and 104.4 mm on 28 September. Source: NOAA GHCN-Daily in Wherobots; 15.8 s.

What a learned forecast adds to the data

The forecasts to compare against are large. A third query lists the NOAA archive's files for the run initialized at 00 UTC on 26 September 2024, between the first and second days of heavy rain at the airport:

-- How big one learned weather forecast is: the 00 UTC run of 26 September 2024
-- from the four machine learning models in NOAA's MLWP archive (file listing only),
-- plus the Buncombe County outline used for the station map
SELECT regexp_extract(path, '/([A-Z]{4}_v[0-9]{3}_GFS)/', 1) AS model_run,
       ROUND(length / 1e9, 2) AS gb, CAST(NULL AS STRING) AS county_wkt
FROM binaryFile.`s3://noaa-oar-mlwp-data/*_GFS/2024/0926/*_2024092600_f000_f240_06.nc`
UNION ALL
SELECT 'Buncombe County outline', NULL, ST_AsText(ST_SimplifyPreserveTopology(geometry, 0.002))
FROM wherobots_open_data.us_census.tiger_county
WHERE STATEFP = '37' AND NAME = 'Buncombe'

In 7.6 seconds it returned three 10-day global forecasts at 6-hour steps (the archive held no Aurora file for that run): 9.34 GB from GraphCast (GRAP_v100_GFS), 7.47 GB from FourCastNet v2 (FOUR_v200_GFS), and 7.31 GB from Pangu-Weather (PANG_v100_GFS), one NetCDF file per run. Scoring one of them at the 35 gauges means reading the forecast cells over one county from a global file and joining them to the gauge points. The HRRR article shows the same pattern for physics-based forecasts: fetch the fields once, convert them to a raster format WherobotsDB reads, and join with raster and vector data in one query.

World model research and what changes at scale

Timeline with four lanes from 1940 to 2026. Mental models and control: Craik 1943, Kalman 1960, Forrester 1971, Johnson-Laird 1983. Agents and robots: Sutton's Dyna 1991, Ha and Schmidhuber's World Models 2018, PlaNet 2019, Dreamer 2020, MuZero 2020, LeCun's JEPA path 2022, DreamerV3 and V-JEPA 2 in 2025. Video and 3D world generators: Genie 2024, Cosmos and Genie 3 in 2025, Waymo World Model and World Labs Atlas in 2026. Earth system models: Bauer's numerical weather prediction review 2015, Earth digital twin 2021, Pangu-Weather and GraphCast 2023, WeatherBench 2 2024, GenCast and Aurora 2025
Four lines of work that use the term world model, from Craik (1943) to Atlas (2026). Dates follow the versions cited in this article.

From mental models to learned dynamics

Craik's 1943 small-scale model and Kalman's 1960 filter gave the two halves of the idea: predict with an internal model, and correct it with measurements. After Dyna (1991) and Ha and Schmidhuber (2018), Danijar Hafner and colleagues learned dynamics directly from pixels and planned in latent space with PlaNet, in Learning Latent Dynamics for Planning from Pixels (ICML 2019), and trained behavior by imagined rollouts with Dreamer, in Dream to Control: Learning Behaviors by Latent Imagination (ICLR 2020). MuZero (2020) showed planning with a learned model at superhuman level, and DreamerV3 (2025) generalized one configuration across domains. LeCun's 2022 position paper and the JEPA models that followed, such as V-JEPA 2, predict in a learned representation space instead of pixel space.

World foundation models

Jake Bruce and colleagues' Genie: Generative Interactive Environments (ICML 2024) trained an 11 billion parameter model on 30,000 hours of internet gameplay video without action labels, and learned a latent action space that lets a user control generated worlds frame by frame. NVIDIA's Cosmos World Foundation Model Platform for Physical AI (2025) packaged video tokenizers and pretrained world models for robots and vehicles, Genie 3 raised interactive generation to 720p at 24 frames per second, and Atlas produces minute-long 1440p video with camera control and reconstructs real-world scenes from one to dozens of input images.

Learned Earth system models

In weather, learned world models now compete with physics-based ones. Kaifeng Bi and colleagues' Pangu-Weather, in Accurate medium-range global weather forecasting with 3D neural networks (Nature, 2023), trained on 39 years of global data. GraphCast followed in Science the same year. Ilan Price and colleagues' GenCast, in Probabilistic weather forecasting with machine learning (Nature, 2025), generates ensembles of 15-day forecasts and outperformed the leading operational ensemble on 97.2% of 1,320 targets. Cristian Bodnar and colleagues' Aurora, in A foundation model for the Earth system (Nature, 2025), trained on more than one million hours of geophysical data and was fine-tuned for air quality, ocean waves, and tropical cyclone tracks. Stephan Rasp and colleagues built WeatherBench 2 (Journal of Advances in Modeling Earth Systems, 2024) as a shared benchmark for evaluating these models.

What scale changes

One county needed one year of worldwide observations read in place and three forecast files of 7.31 to 9.34 GB each, for a single run. A season of twice-daily runs, scored at every gauge in a state, multiplies both sides, and the join between forecast cells and observation points has to run where the files sit. Spatial partitioning and per-partition indexes on a cluster, the approach Jia Yu, Jinxuan Wu, and Mohamed Sarwat introduced with GeoSpark (ACM SIGSPATIAL 2015), keep that to one query per question. Imagery adds a third side: AI satellite imagery analysis turns pixels into the buildings, flood extents, and land cover a world model of a place needs as state, and geospatial AI covers the models that do it.

Open problems

  • Extremes. At the airport, 355.1 mm of the year's 1,678.4 mm fell in the four Helene days. Rare days like these are where a model trained mostly on ordinary days is least tested.
  • Uneven observations. Of 139 stations ever listed in Buncombe County, 35 reported all four storm days, and gauges cluster near towns.
  • Physical consistency. Generated scenes can look right and still break physics, and benchmarks for physical plausibility are still forming.
  • Matching predictions to places. A 0.25° forecast cell covers many gauges, buildings, and elevations. Scoring at a point needs the digital elevation model and the geometry of the place, joined in one coordinate reference system.

Read more from Wherobots

Assemble observed state for your own places with a Wherobots free trial at cloud.wherobots.com.

Frequently asked questions

What is a world model?

A world model is an AI model that predicts how an environment will change, given its current state and, for an agent, the action it takes. It encodes observations into a state, predicts the next state, and chains those predictions so a system can test actions or produce a forecast before anything happens. David Ha and Jürgen Schmidhuber popularized the term in machine learning with World Models (2018).

What is the difference between an LLM and a world model?

A large language model (LLM) predicts the next token of text and was trained on text, documents, databases, and the internet. A world model predicts the next state of an environment, such as the next video frame, robot position, or atmospheric field, often conditioned on an action. Some systems combine them: vision-language models feed world models, and world models generate scenes from text prompts. Yann LeCun’s A Path Towards Autonomous Machine Intelligence (2022) argues that machines need predictive world models to reason and plan.

How do I build a world model?

Collect observations of the environment with timestamps, plus the actions taken if there is an agent. Choose a state representation (a latent vector, a grid, or a scene). Train a transition model that predicts the next state from the current one. Roll it forward to plan or forecast. Then score its predictions against observations it never saw during training, and compare them with a simple reference such as persistence or climatology. For real places, the observations need coordinates and times so predictions can be matched to the right location.

Does Waymo use world models?

Yes. In February 2026 Waymo introduced the Waymo World Model, a generative world model for driving simulation built on Google DeepMind’s Genie 3 and adapted to produce camera and lidar outputs. Engineers can control its scenes with language prompts, driving inputs, and scene layouts to simulate rare events.

What is a world foundation model?

A world foundation model is a world model pretrained on large video collections and then fine-tuned for a robot, vehicle, or scene. NVIDIA’s Cosmos World Foundation Model Platform for Physical AI (2025) describes a video curation pipeline, video tokenizers, and pretrained models released with open weights. See What is physical AI? for how these models fit into robot and vehicle training.

What are examples of world models?

Dreamer and DreamerV3 learn a latent model of a game or control task and train behavior inside it. MuZero plans with a learned model in Atari, Go, chess, and shogi. Genie and Genie 3 generate interactive environments, NVIDIA Cosmos and the Waymo World Model generate driving and robot scenes, and World Labs’ Atlas generates and reconstructs 3D worlds. In weather, GraphCast, Pangu-Weather, GenCast, and Aurora are learned models that step the atmosphere forward in time.

Is a digital twin a world model?

A digital twin is a model of one specific place or asset, kept in step with sensor data from it. When the twin predicts how that place will change, it serves as a world model for it. Peter Bauer, Bjorn Stevens, and Wilco Hazeleger describe a digital twin of Earth (Nature Climate Change, 2021) that combines simulation and observations to predict environmental change.

How is a world model evaluated?

By comparing its predictions with observations it did not train on. Agent world models are scored on held-out episodes and task returns. Weather models are scored against observations and analyses with metrics such as mean absolute error and skill scores, which measure improvement over a reference forecast such as persistence or climatology. A forecast or a warning is a prediction, so it never counts as a verifying observation. What is ground truth data? scores GraphCast against Hurricane Helene rain gauges step by step.

What is a mental model in psychology?

A mental model is an internal representation a person uses to anticipate events. Kenneth Craik proposed in The Nature of Explanation (1943) that the mind carries a small-scale model of external reality and of its own possible actions, and Philip Johnson-Laird developed the idea in Mental Models (1983). Machine learning world models borrow the same idea: predict the outcome inside a model before acting.

How does Wherobots relate to world models?

Wherobots is the AI Context Engine for Physical World Data. A world model of real places needs observed state with coordinates and timestamps, and WherobotsDB joins those observations, such as weather station records, boundaries, buildings, and elevation, in spatial SQL. RasterFlow (Public Preview) builds mosaics from imagery, runs computer vision inference on them, and writes the results as vectorized outputs on the same platform, and the Wherobots MCP server lets AI agents query the results.