Planetary-scale answers, unlocked.
A Hands-On Guide for Working with Large-Scale Spatial Data. Learn more.
Authors
Satellite imagery analysis is the process of extracting information from images of the Earth's surface captured by satellites and aircraft. It turns pixels into answers: which land is forest, where every building and field sits, how tall the trees are, and what changed since last year. Most AI in satellite imagery today is computer vision: models that classify, detect, segment, and measure what the pixels show.
Key takeaways
Methods range from an analyst tracing a shoreline by hand to a neural network labeling every pixel on a continent, and they fall into four generations:
The research section below traces who introduced each generation and when. The input is always a raster: a grid of pixels with one or more spectral bands, stored as a GeoTIFF, a Cloud Optimized GeoTIFF (an OGC standard since 2023, which lets a reader fetch only the tiles it needs over HTTP), or a cloud-native array format such as Zarr. The output is often vector data, polygons and points that join to parcels, roads, and other tables, which is why raster and vector data meet in this field. Remote sensing satellite imagery analysis covers optical, radar, and thermal sensors, and the same methods apply to aerial photographs.
The model is one step in a longer pipeline, and the output is usable only when the surrounding steps work.
Steps 2, 3, 5, and 6 are where projects stall. Mosaicking a county is manageable; mosaicking a continent means managing tens of thousands of scenes, projections, and seams.
Validation borrows its measures from computer vision. Intersection over union (IoU) divides the area two shapes share by the area they cover together, so 1.0 is a perfect match, and the PASCAL VOC challenge (Everingham et al., International Journal of Computer Vision, 2010) counted a detection as correct when its IoU with the reference exceeded 0.5, now the usual cutoff. Precision is the share of detections that are real, and recall is the share of real objects found. Classified maps use an error matrix, covered under measuring accuracy below.
Most AI satellite imagery analysis falls into five computer vision tasks, each with its own question and output.
Instance segmentation sits between detection and segmentation: it returns a separate mask for every object, so two adjacent roofs stay two features. Change detection compares two or more dates of the same place, with each date run through the same classification or segmentation step. Foundation models now cover several of these tasks at once: one promptable model can detect roofs, roads, and solar panels from text prompts with no task-specific training, as the Wherobots example below shows.
Ground sample distance, the ground width of one pixel, sets the smallest object a model can find. A rule of thumb is several pixels across the smallest object of interest.
Landsat 1 launched on July 23, 1972, which makes Landsat the longest continuous record. The Sentinel-2 mission paper (Drusch et al., 2012) describes the 13-band design behind most free analysis today. Spectral depth matters as much as pixel size: near-infrared separates vegetation from everything else, which is why NAIP's fourth band and Sentinel-2's red-edge bands carry information the eye cannot see. Hyperspectral imagery extends this to hundreds of narrow bands for mineral and crop stress mapping.
Each of these well-known projects pairs one task with one sensor:
Each use joins model output to other data. A roof polygon matters once it is linked to a parcel, a policy, or a flood zone, which is where satellite imagery GIS analysis begins.
Wherobots RasterFlow runs the full pipeline as a managed service: it builds mosaics from satellite and aerial imagery, runs computer vision models on them, and vectorizes the results through a Python API. Mosaics and predictions are written as Zarr, and vector results as GeoParquet. Built-in models cover common tasks:
Custom PyTorch models run through the bring-your-own-model path. Sentinel-2 mosaics skip scenes with 75% or more cloud cover, mask cloud, shadow, and cirrus pixels with the Scene Classification Layer, and take a per-pixel median of what remains.
Pricing is calculated before a task runs, from area, resolution, bands, and time periods. Finding every solar panel array across a 500 km² county with SAM3 on 30 cm NAIP costs $35.00, mosaic and inference included.
# Detect solar panels across a county from a text prompt with SAM3 on 30 cm NAIP from datetime import datetime from rasterflow_remote import RasterflowClient from rasterflow_remote.data_models import GeometryModelRecipes rf = RasterflowClient() detections = rf.predict_mosaic_geometries_recipe( aoi="s3://your-bucket/county.parquet", start=datetime(2022, 1, 1), end=datetime(2023, 1, 1), model_recipe=GeometryModelRecipes.SAM3_TEXT_GEOMETRY, text_prompt="solar panel", confidence_threshold=0.5, ) print(detections.uri) # GeoParquet containing detected solar-panel polygons
The GeoParquet output loads into WherobotsDB, built by the original creators of Apache Sedona, where a spatial join links each polygon to parcels, buildings, or hazard zones in SQL.
In Marion County, Oregon, one RasterFlow pipeline built a 133 GB mosaic from 2022 NAIP imagery and ran SAM3 with eight text prompts in a single pass, keeping detections with a confidence of 0.5 or more. Wherobots publishes the output as wherobots_open_data.rasterflow_output_samples.marion_county_sam3_vector. This query summarizes it per prompt, with areas computed in UTM zone 10N (EPSG:32610) so they come out in square meters:
wherobots_open_data.rasterflow_output_samples.marion_county_sam3_vector
SELECT label, COUNT(*) AS n, ROUND(AVG(bbox_score), 3) AS mean_score, ROUND(SUM(ST_Area(ST_Transform(geometry, 'EPSG:4326', 'EPSG:32610'))) / 1e6, 3) AS area_km2 FROM wherobots_open_data.rasterflow_output_samples.marion_county_sam3_vector GROUP BY label ORDER BY n DESC
The table holds 1,006,232 detections, and every row carries the same time stamp, 1 January 2022, the start of the 2022 NAIP window used to build the mosaic. The 312,245 roofs cover 41.7 km².
Three patterns stand out in the per-prompt summary.
A detection is only as useful as its agreement with something independent. The county-wide evaluation of SAM3 roofs against Overture building footprints clipped both datasets to the official Marion County boundary, spatially joined SAM3 roofs to Overture Maps buildings on intersection in EPSG:32610, and computed IoU for every intersecting pair. Inside the boundary it counted 261,414 SAM3 roofs and 146,642 Overture buildings, and the join returned 207,109 candidate pairs. Total roof area agreed within 7%, and SAM3 matched the shape of three-quarters of known buildings. Many unmatched detections were real rural houses and outbuildings missing from Overture, so the authors judged real precision likely higher than 75%.
Satellite imagery analysis rests on fifty years of research, from texture statistics computed on the first Landsat scenes to promptable segmentation models. Four threads shape it: how pixels became objects, how accuracy is measured, what benchmarks and foundation models changed, and which problems remain open.
The first computer analyses of satellite images classified one pixel at a time from its spectral values. Haralick, Shanmugam, and Dinstein added spatial context in Textural features for image classification (IEEE Transactions on Systems, Man, and Cybernetics, 1973): statistics of gray-tone co-occurrence in a neighborhood, tested on aerial photographs and multispectral imagery from the Earth Resources Technology Satellite, later renamed Landsat 1. Compton Tucker's study of red and photographic infrared combinations (Remote Sensing of Environment, 1979) compared linear combinations of red and near-infrared bands for monitoring vegetation, and the normalized difference it favored became NDVI.
Comparing dates came next. Ashbindu Singh's review of digital change detection techniques (International Journal of Remote Sensing, 1989) evaluated the procedures for comparing multitemporal images that had been developed by then, and it remains the starting point for the field. The Hansen forest map above ran that idea globally on Landsat time series.
Sub-meter imagery changed the unit of analysis. Once a house spans hundreds of pixels, a pixel classifier assigns roof texture, shadow, and lawn to separate classes. Thomas Blaschke's Object based image analysis for remote sensing (ISPRS Journal of Photogrammetry and Remote Sensing, 2010) describes the response: segment the image into objects first, then classify each object by its spectra, shape, and relations to its neighbors. Instance segmentation models such as SAM3 complete that shift, returning one polygon per object straight from the network.
Russell Congalton's review of assessing the accuracy of classifications of remotely sensed data (Remote Sensing of Environment, 1991) set the error matrix as the core tool and named the choices that decide whether an accuracy number means anything: the classification scheme, the sampling design, the sample size, and spatial autocorrelation between samples. Nearby pixels are similar, so samples drawn close together are not independent, which is the first law of geography applied to validation. Object detection added the IoU threshold from PASCAL VOC, the measure behind the Marion County comparison.
The open problem is the reference itself. Overture buildings are a curated, multi-source dataset, and they still likely miss buildings in rural areas, which is why the Marion County evaluation treated many SAM3 "false positives" as real buildings. Roofs also differ from footprints: overhangs and occlusion mean a roof outline and a footprint never match exactly. Any accuracy number on overhead imagery is a statement about two datasets, and the dates, definitions, and gaps of the reference belong beside it.
Zhu, Tuia, Mou, and colleagues reviewed the arrival of neural networks in Deep learning in remote sensing: a comprehensive review and list of resources (IEEE Geoscience and Remote Sensing Magazine, 2017). Two ingredients made the shift: architectures from other fields and labeled benchmarks for overhead imagery.
U-Net (Ronneberger, Fischer, and Brox, MICCAI 2015) was designed for biomedical images and became the default segmentation network for satellite imagery, because its expanding path recovers the precise localization that building and road edges need. The benchmarks followed, alongside Fields of the World and Tile2Net from the examples above:
Every benchmark above needs labels for its task. Foundation models move most of the learning to unlabeled imagery or to large general datasets, and the research follows three lines.
Promptable segmentation. Kirillov and colleagues' Segment Anything (ICCV 2023) released over 1 billion masks on 11 million images and segmented an object from a point or a box. Meta's SAM 3 (Carion et al., 2025) replaced the click with a concept: a noun phrase such as "roofs" or an image exemplar, returning every matching instance, trained on a dataset with 4 million unique concept labels. That change is what lets one RasterFlow run detect roofs, roads, and solar panels across a county with no labels.
Pretrained Earth observation encoders. NASA and IBM's Prithvi (Jakubik et al., 2023) was pretrained on more than 1 TB of Harmonized Landsat Sentinel-2 imagery for fine-tuning on downstream tasks with small labeled datasets. Meta's canopy height model in the examples above used a self-supervised DINOv2 encoder.
Embeddings as a product. Rolf and colleagues' MOSAIKS (Nature Communications, 2021) computes image encodings once and shares them across tasks, so each new task needs only a linear regression on the user's own labels. Google DeepMind's AlphaEarth Foundations (Brown et al., 2025) takes the same approach with a learned model, compressing each location's imagery into annual global embedding layers from 2017 through 2024, so similarity and change become distance calculations.
At county scale the model is the hard part. At continental scale the data movement is. Gorelick and colleagues described one answer in Google Earth Engine: planetary-scale geospatial analysis for everyone (Remote Sensing of Environment, 2017): a catalog of analysis-ready imagery next to parallel compute. The Cloud Optimized GeoTIFF layout defined above serves the same goal on object storage, where each reader fetches only the tiles it needs.
The vector side scales too. A billion model polygons need a spatial engine to join them to parcels, buildings, and boundaries. Jia Yu, Jinxuan Wu, and Mohamed Sarwat introduced GeoSpark (ACM SIGSPATIAL 2015) for distributed spatial queries on Apache Spark, and Yu, Zongsi Zhang, and Sarwat detailed its spatial partitioning, indexing, and join design in Spatial data management in Apache Spark: the GeoSpark perspective and beyond (GeoInformatica, 2019). GeoSpark became Apache Sedona. Kanchan Chowdhury and Sarwat's GeoTorch (ACM SIGSPATIAL 2022) brought a spatiotemporal deep learning framework into the same research line.
The Fields of the World global release shows the numbers at the far end. Taylor Geospatial ran the PRUE model on RasterFlow, generating 348.7 TB across 540,794 objects to produce 8.2 billion field boundaries. At that size, the Marion County query becomes one of thousands, and the spatial engine that runs them sets how long the answer takes.
Run a built-in RasterFlow model from the Model Hub at cloud.wherobots.com.
AI analyzes satellite images with computer vision models that classify each pixel, draw boxes or outlines around objects, estimate values such as tree height, and compare dates to find change. A pipeline builds a cloud-free mosaic, splits it into patches, runs the model on each patch, stitches the predictions back together, and converts them to polygons for analysis.
Start with what the sensor records: its resolution, bands, and capture date. True color shows the scene as the eye sees it, and false color combinations such as near-infrared, red, and green make vegetation stand out. Analysts read tone, texture, shape, size, shadow, and context, and AI models are trained on the same cues from labeled examples.
Sentinel-2 is the most used free source for analysis, with 10 m bands and a revisit of about five days. Landsat adds a record back to 1972 at 30 m. In the United States, USDA NAIP aerial imagery reaches 30 cm to 1 m, fine enough for roofs and sidewalks. Sentinel-1 radar images through clouds.
Object detection finds individual things, such as cars or roofs, and returns a box or outline for each one with a confidence score. Semantic segmentation labels every pixel with a class, such as road or sidewalk, without separating one object from the next. Instance segmentation combines both and returns a separate mask for each object.
Yes, with promptable models. Meta’s SAM 3 takes a short noun phrase such as “roofs” or “solar panels” and returns a mask for every matching object, with no task-specific training. Results on overhead imagery vary by object, so teams check a sample against reference data before using the output.
The object sets the resolution. A rule of thumb is several pixels across the smallest object. Fields and forests work at 10 m, buildings need about 1 m or finer, and roofs, solar panels, sidewalks, and vehicles need 30 to 60 cm. Finer imagery costs more to store and process per square kilometer.
A well-known example is Hansen and colleagues’ 2013 global forest map, which classified Landsat images at 30 m to find 2.3 million km² of forest lost from 2000 to 2012. Other examples include grading building damage after a hurricane, mapping every farm field boundary from Sentinel-2, and detecting roofs and solar panels in aerial imagery. In Marion County, Oregon, SAM3 on Wherobots RasterFlow returned 1,006,232 detections from eight text prompts, including 312,245 roofs.
Most Maxar imagery is commercial and licensed for a fee. Its Open Data Program, now run under the Vantor name, releases before and after imagery of major sudden-onset disasters for free under a Creative Commons Attribution-NonCommercial 4.0 license, so it can be used with attribution for non-commercial purposes such as humanitarian response and research.
No public service shows a live satellite view of a house. Google Earth and Google Maps show stored imagery from satellites and aircraft, and its capture date varies by location. Free satellites such as Sentinel-2 revisit about every five days at 10 m, too coarse to show a single house in detail.
For classified maps, compare a sample of pixels against reference data in an error matrix, as Congalton’s 1991 review set out. For detected objects, compute intersection over union (IoU) between each detection and a reference shape: a detection usually counts as correct above 0.5. Precision is the share of detections that are real and recall is the share of real objects found.
What is Context Engineering for Physical AI?
Context engineering is choosing what an AI model receives at each step. How it works, how it compares to RAG, and what changes when the context is real places.
What is an Ontology? Definition, Parts, Spatial Ontologies
An ontology is a formal model of the things in a domain and how they relate. Its parts, how it differs from a taxonomy or knowledge graph, and what makes one spatial.
What is H3? Uber’s Hexagonal Spatial Index
H3 is Uber's open-source hexagonal grid for indexing locations on Earth. How the H3 index works, its 16 resolutions, H3 vs S2, geohash and quadkeys, and H3 in SQL.
share this article
Awesome that you’d like to share our articles. Where would you like to share it to: