Rasterflow, Earth Intelligence & inference engine now in public preview Learn More

What is Earth Observation Data? Types, Sources

What is Earth observation data?

Earth observation data is the imagery and measurements of the Earth’s surface, oceans, and atmosphere collected by satellites, aircraft, and drones. It includes optical and radar imagery, elevation, land cover, sea surface temperature, and atmospheric gases, delivered as georeferenced rasters with a timestamp. Open programs such as Copernicus and Landsat supply most of it free.

What Earth observation data is

Earth observation data is the output of remote sensing: any measurement of the planet made from above, processed into a georeferenced product with a known date.

The defining property is that it is systematic. A satellite images the same ground on a fixed orbit and schedule, so Earth observation data is a time series over every place, whether or not anyone asked for it.

Four families cover most of it. Optical imagery records reflected sunlight in visible and infrared bands, and radar imagery records microwave backscatter through cloud.

Thermal imagery records emitted heat. Atmospheric instruments record gases, aerosols, and clouds.

Derived products sit on top. Elevation models, land cover maps, vegetation indices, sea surface temperature, and soil moisture are all Earth observation data one processing step removed from the sensor.

Producers split into public programs, led by the European Union’s Copernicus and the United States’ Landsat and NASA missions, and commercial operators such as Planet, Maxar, ICEYE, and Airbus.

What Earth observation data does

Earth observation data replaces sampling with coverage. It measures every field, forest, coastline, and city on a schedule, and the archive shows how each one has changed.

Governments report crop area, forest loss, and urban expansion from it under climate and trade agreements.

Insurers quantify exposure before an event and damage after it, across a whole portfolio at once.

Energy companies site wind and solar plants on irradiance, terrain, and land cover, then monitor construction and vegetation encroachment.

Supply chain teams watch ports, mines, and factories for activity and disruption without a person on site.

Agencies detect wildfires within hours of ignition, map floods within a day, and track methane leaks from orbit.

Each use shares one requirement: joining the observation to the assets, boundaries, or policies it concerns, which is where most of the engineering effort goes.

How Earth observation data works

Earth observation data arrives as scenes, each covering a fixed footprint on one date, and is organized by processing level.

LevelWhat it containsReady for
0Raw sensor countsReprocessing only
1Georeferenced radiance or amplitudeVisualization
2Surface reflectance or calibrated backscatter, with quality masksAnalysis
3Composites and mosaics over timeTime series, mapping
4Model output such as biomass or soil moistureDecision products

Analysis Ready Data (ARD) is the industry term for Level 2 products on a consistent grid with masks applied, so a user can compute on them without preprocessing.

Formats are raster. Cloud Optimized GeoTIFF holds most imagery, NetCDF and Zarr hold multidimensional atmospheric and ocean data, and point clouds hold lidar.

Discovery runs through STAC, the SpatioTemporal Asset Catalog, a JSON standard that indexes scenes by footprint, date, and asset URL. Most open archives on AWS, Microsoft Planetary Computer, and the Copernicus Data Space expose a STAC API.

Volume shapes everything. Copernicus alone distributes tens of terabytes a day, so Earth observation data lives in cloud object storage and is processed where it sits.

The output of an Earth observation workflow is a table: a value per asset, per field, or per boundary, per date.

Earth observation data in Wherobots

WherobotsDB reads Earth observation data where it is stored and joins it to vector data in the same SQL query, which is the step most workflows spend their effort on.

The STAC reader queries any STAC API by area and date and returns one row per scene, with each band asset already registered as an out-db raster. Pixels are read only when a function touches them.

RS_ functions do the raster work. RS_MapAlgebra computes indices, RS_Clip cuts to a boundary, RS_ZonalStats summarizes per polygon, and RS_Value samples at a point, each in parallel across every scene.

Havasu, the Apache Iceberg table format extended with native raster and geometry types, stores the resulting time series with ACID guarantees and spatial filter pushdown.

RasterFlow supplies the imagery preparation and model layer. It builds cloud-filtered Sentinel-2 and NAIP mosaics for an area of interest and runs built-in or custom PyTorch models over them.

The Wherobots Data Hub offers ready tables, including ESA WorldCover land cover and Overture Maps buildings, so a join has a partner on day one.

Nothing moves. Data stays in the customer’s S3 bucket or Iceberg catalog, and compute comes to it.

The Wherobots MCP Server exposes the same catalog and functions to Claude Code and other agents. WherobotsDB is built by the original creators of Apache Sedona, 100% code compatible across all spatial functions.

-- Dominant land cover class and pixel count per administrative area (ESA WorldCover)
SELECT a.name,
       RS_ZonalStats(lc.rast, a.geometry, 1, 'mode')  AS dominant_class,
       RS_ZonalStats(lc.rast, a.geometry, 1, 'count') AS pixels
FROM wherobots.esa.esa_world_cover lc
JOIN admin_areas a ON RS_Intersects(lc.rast, a.geometry);

What is Earth observation data?

Earth observation data is the imagery and measurements of the Earth’s surface, oceans, and atmosphere collected by satellites, aircraft, and drones. It includes optical and radar imagery, elevation, land cover, sea surface temperature, and atmospheric gases, delivered as georeferenced rasters with a timestamp. Open programs such as Copernicus and Landsat supply most of it free.

What Earth observation data is

Earth observation data is the output of remote sensing: any measurement of the planet made from above, processed into a georeferenced product with a known date.

The defining property is that it is systematic. A satellite images the same ground on a fixed orbit and schedule, so Earth observation data is a time series over every place, whether or not anyone asked for it.

Four families cover most of it. Optical imagery records reflected sunlight in visible and infrared bands, and radar imagery records microwave backscatter through cloud.

Thermal imagery records emitted heat. Atmospheric instruments record gases, aerosols, and clouds.

Derived products sit on top. Elevation models, land cover maps, vegetation indices, sea surface temperature, and soil moisture are all Earth observation data one processing step removed from the sensor.

Producers split into public programs, led by the European Union’s Copernicus and the United States’ Landsat and NASA missions, and commercial operators such as Planet, Maxar, ICEYE, and Airbus.

What Earth observation data does

Earth observation data replaces sampling with coverage. It measures every field, forest, coastline, and city on a schedule, and the archive shows how each one has changed.

Governments report crop area, forest loss, and urban expansion from it under climate and trade agreements.

Insurers quantify exposure before an event and damage after it, across a whole portfolio at once.

Energy companies site wind and solar plants on irradiance, terrain, and land cover, then monitor construction and vegetation encroachment.

Supply chain teams watch ports, mines, and factories for activity and disruption without a person on site.

Agencies detect wildfires within hours of ignition, map floods within a day, and track methane leaks from orbit.

Each use shares one requirement: joining the observation to the assets, boundaries, or policies it concerns, which is where most of the engineering effort goes.

How Earth observation data works

Earth observation data arrives as scenes, each covering a fixed footprint on one date, and is organized by processing level.

LevelWhat it containsReady for
0Raw sensor countsReprocessing only
1Georeferenced radiance or amplitudeVisualization
2Surface reflectance or calibrated backscatter, with quality masksAnalysis
3Composites and mosaics over timeTime series, mapping
4Model output such as biomass or soil moistureDecision products

Analysis Ready Data (ARD) is the industry term for Level 2 products on a consistent grid with masks applied, so a user can compute on them without preprocessing.

Formats are raster. Cloud Optimized GeoTIFF holds most imagery, NetCDF and Zarr hold multidimensional atmospheric and ocean data, and point clouds hold lidar.

Discovery runs through STAC, the SpatioTemporal Asset Catalog, a JSON standard that indexes scenes by footprint, date, and asset URL. Most open archives on AWS, Microsoft Planetary Computer, and the Copernicus Data Space expose a STAC API.

Volume shapes everything. Copernicus alone distributes tens of terabytes a day, so Earth observation data lives in cloud object storage and is processed where it sits.

The output of an Earth observation workflow is a table: a value per asset, per field, or per boundary, per date.

Earth observation data in Wherobots

WherobotsDB reads Earth observation data where it is stored and joins it to vector data in the same SQL query, which is the step most workflows spend their effort on.

The STAC reader queries any STAC API by area and date and returns one row per scene, with each band asset already registered as an out-db raster. Pixels are read only when a function touches them.

RS_ functions do the raster work. RS_MapAlgebra computes indices, RS_Clip cuts to a boundary, RS_ZonalStats summarizes per polygon, and RS_Value samples at a point, each in parallel across every scene.

Havasu, the Apache Iceberg table format extended with native raster and geometry types, stores the resulting time series with ACID guarantees and spatial filter pushdown.

RasterFlow supplies the imagery preparation and model layer. It builds cloud-filtered Sentinel-2 and NAIP mosaics for an area of interest and runs built-in or custom PyTorch models over them.

The Wherobots Data Hub offers ready tables, including ESA WorldCover land cover and Overture Maps buildings, so a join has a partner on day one.

Nothing moves. Data stays in the customer’s S3 bucket or Iceberg catalog, and compute comes to it.

The Wherobots MCP Server exposes the same catalog and functions to Claude Code and other agents. WherobotsDB is built by the original creators of Apache Sedona, 100% code compatible across all spatial functions.

-- Dominant land cover class and pixel count per administrative area (ESA WorldCover)
SELECT a.name,
       RS_ZonalStats(lc.rast, a.geometry, 1, 'mode')  AS dominant_class,
       RS_ZonalStats(lc.rast, a.geometry, 1, 'count') AS pixels
FROM wherobots.esa.esa_world_cover lc
JOIN admin_areas a ON RS_Intersects(lc.rast, a.geometry);

Related terms

  • Remote sensing: the technique that produces Earth observation data
  • Sentinel-2: the most downloaded open Earth observation mission
  • Landsat: the longest Earth observation archive
  • SAR data: the all-weather branch of Earth observation

Read more from Wherobots

Query Earth observation data in place with WherobotsDB at cloud.wherobots.com.

Frequently asked questions

What is the difference between Earth observation data and remote sensing?

Remote sensing is the technique of measuring the surface from a distance with sensors on satellites or aircraft. Earth observation data is the product that technique delivers: the georeferenced imagery and measurements. Remote sensing is the verb; Earth observation data is the noun, and the two are often used interchangeably.

Is Earth observation data free?

Most Earth observation data is free. The European Union’s Copernicus program (Sentinel-1, -2, -3, -5P) and the United States’ Landsat and NASA missions publish open data with no license fee. Commercial Earth observation data from Planet, Maxar, ICEYE, and Airbus is licensed and offers finer resolution or daily revisits.

What are the main types of Earth observation data?

The main types of Earth observation data are optical imagery (visible and infrared reflectance), radar imagery (SAR backscatter), thermal imagery (emitted heat), atmospheric measurements (gases and aerosols), and elevation from lidar and radar altimetry. Derived products such as land cover, vegetation indices, and sea surface temperature are built from these.

What format is Earth observation data in?

Earth observation data is mostly raster. Imagery ships as GeoTIFF, increasingly Cloud Optimized GeoTIFF, with one file per band, while multidimensional atmospheric and ocean data uses NetCDF or Zarr. Discovery uses STAC, a JSON catalog standard, and derived vector products such as detected buildings ship as GeoParquet.