Rasterflow, Earth Intelligence & inference engine now in public preview Learn More

What is Satellite Imagery Analysis? AI, Tasks, Uses

Authors

Satellite imagery analysis is the process of extracting information from images of the Earth's surface captured by satellites and aircraft. It turns pixels into answers: which land is forest, where every building and field sits, how tall the trees are, and what changed since last year. Most AI in satellite imagery today is computer vision: models that classify, detect, segment, and measure what the pixels show.

Key takeaways

  • Satellite imagery analysis turns pixels into maps, object counts, measurements, and change alerts.
  • Five computer vision tasks cover most of the work: classification, object detection, segmentation, regression, and change detection.
  • Resolution sets what is visible. Sentinel-2 at 10 m maps fields, and 30 cm aerial imagery maps roofs and sidewalks.
  • At scale, the hard part is the pipeline around the model: mosaics, cloud masking, patching, stitching, and vectorization.
  • Wherobots RasterFlow runs that pipeline as a managed service with built-in models such as SAM3, Tile2Net, and Fields of the World.

What satellite imagery analysis is

Methods range from an analyst tracing a shoreline by hand to a neural network labeling every pixel on a continent, and they fall into four generations:

  • Visual interpretation. A trained analyst reads tone, texture, shape, size, shadow, and context.
  • Pixel and spectral analysis. Band math and classifiers work on the pixel values of Landsat and Sentinel-2, and band ratios such as NDVI measure vegetation.
  • Object-based image analysis. Pixels are grouped into segments first, then the segments are classified by shape and context.
  • Deep learning. Neural networks are trained on labeled examples of roofs, sidewalks, or field boundaries and then label every new image.

The research section below traces who introduced each generation and when. The input is always a raster: a grid of pixels with one or more spectral bands, stored as a GeoTIFF, a Cloud Optimized GeoTIFF (an OGC standard since 2023, which lets a reader fetch only the tiles it needs over HTTP), or a cloud-native array format such as Zarr. The output is often vector data, polygons and points that join to parcels, roads, and other tables, which is why raster and vector data meet in this field. Remote sensing satellite imagery analysis covers optical, radar, and thermal sensors, and the same methods apply to aerial photographs.

How satellite imagery analysis works

The model is one step in a longer pipeline, and the output is usable only when the surrounding steps work.

  1. Select imagery. Pick the sensor, date range, and area of interest, then filter scenes by cloud cover.
  2. Build a mosaic. Combine many scenes into one seamless, georeferenced raster on a common grid.
  3. Patch. Split the mosaic into fixed-size patches that fit in GPU memory, with overlap so objects on a patch edge are seen whole.
  4. Run inference. Apply the model to every patch, often thousands to millions of them in parallel.
  5. Merge. Stitch predictions back into one raster, blending the overlaps to remove seams.
  6. Vectorize. Threshold the predictions and convert pixels to polygons, or keep the model's polygons directly.
  7. Validate. Compare a sample against reference data.
Diagram of the seven steps of a satellite imagery analysis pipeline: select imagery, build a mosaic, patch, run inference, merge, vectorize, and join and validate. Mosaic, patch, merge, and vectorize are highlighted as the steps where large projects stall. Notes under each step give the Marion County, Oregon run: a 133 GB NAIP mosaic, 1,008 pixel patches, SAM3 with eight prompts, and about one million polygons
The pipeline around the model, with the Marion County, Oregon SAM3 run on USDA NAIP imagery as the example: a 133 GB mosaic in, about one million polygons out.

Steps 2, 3, 5, and 6 are where projects stall. Mosaicking a county is manageable; mosaicking a continent means managing tens of thousands of scenes, projections, and seams.

Validation borrows its measures from computer vision. Intersection over union (IoU) divides the area two shapes share by the area they cover together, so 1.0 is a perfect match, and the PASCAL VOC challenge (Everingham et al., International Journal of Computer Vision, 2010) counted a detection as correct when its IoU with the reference exceeded 0.5, now the usual cutoff. Precision is the share of detections that are real, and recall is the share of real objects found. Classified maps use an error matrix, covered under measuring accuracy below.

The tasks: what AI does with satellite imagery

Most AI satellite imagery analysis falls into five computer vision tasks, each with its own question and output.

TaskQuestion it answersOutputExample
ClassificationWhat is this pixel or tile?A class per pixel or per tileLand cover classification
Object detectionWhere is each object?Boxes or outlines with confidence scoresRoofs, solar panels, ships, aircraft
Semantic segmentationWhich pixels belong to each class?A class mask per pixelRoads, sidewalks, water
RegressionHow much, how tall, how dense?A continuous value per pixelCanopy height in meters
Change detectionWhat changed between dates?A change mask or alertClearing, construction, damage

Instance segmentation sits between detection and segmentation: it returns a separate mask for every object, so two adjacent roofs stay two features. Change detection compares two or more dates of the same place, with each date run through the same classification or segmentation step. Foundation models now cover several of these tasks at once: one promptable model can detect roofs, roads, and solar panels from text prompts with no task-specific training, as the Wherobots example below shows.

Resolution: what each sensor records

Ground sample distance, the ground width of one pixel, sets the smallest object a model can find. A rule of thumb is several pixels across the smallest object of interest.

SourceResolutionRevisitGood for
Landsat30 m8 days with two satellitesLong-term change since 1972, land cover
Sentinel-210 to 60 mAbout 5 daysFields, forests, water, seasonal change
Sentinel-1 SARAbout 10 m6 days with two satellitesFloods, ships, any weather
USDA NAIP30 cm to 1 mAbout every 2 years per stateRoofs, pools, sidewalks, trees
Commercial opticalUnder 1 mVaries by providerVehicles, construction, defense

Landsat 1 launched on July 23, 1972, which makes Landsat the longest continuous record. The Sentinel-2 mission paper (Drusch et al., 2012) describes the 13-band design behind most free analysis today. Spectral depth matters as much as pixel size: near-infrared separates vegetation from everything else, which is why NAIP's fourth band and Sentinel-2's red-edge bands carry information the eye cannot see. Hyperspectral imagery extends this to hundreds of narrow bands for mineral and crop stress mapping.

Satellite imagery analysis examples

Each of these well-known projects pairs one task with one sensor:

  • Global forest loss. Hansen and colleagues mapped forest loss and gain for every 30 m Landsat pixel (Science, 2013) by classifying Landsat time series, and found 2.3 million km² of forest lost from 2000 to 2012. Global Forest Watch now combines optical and radar systems into integrated deforestation alerts.
  • Disaster damage. The Open Data Program started by Maxar, now run under the Vantor name, releases before and after imagery of major disasters, and models trained on the xBD benchmark score the damage to each building.
  • Agriculture. Fields of the World (Kerner et al., AAAI 2025) is a field boundary benchmark of 70,462 Sentinel-2 samples from 24 countries on four continents. Its models predict both the field interior and its edge, so neighboring fields separate cleanly.
  • Forestry and carbon. Meta's canopy height model (Tolan et al., Remote Sensing of Environment, 2024) estimates tree height from sub-meter RGB imagery, with a decoder trained against aerial lidar.
  • Urban planning. Tile2Net (Hosseini et al., Computers, Environment and Urban Systems, 2023) segments sidewalks, crosswalks, and footpaths from aerial tiles and assembles them into a pedestrian network.
  • Poverty mapping. Jean and colleagues combined satellite imagery and machine learning to predict poverty (Science, 2016): a network trained on nighttime lights produced daytime image features that explained up to 75% of the variation in local economic outcomes.
  • Insurance. Roof condition, solar panels, pools, and tree overhang for every property in a portfolio, plus damage maps after hurricanes and wildfires.
  • Mapping and defense. Building footprints and roads for open basemaps, and monitoring of sites for new structures and vehicles.

Each use joins model output to other data. A roof polygon matters once it is linked to a parcel, a policy, or a flood zone, which is where satellite imagery GIS analysis begins.

Satellite imagery analysis in Wherobots

Wherobots RasterFlow runs the full pipeline as a managed service: it builds mosaics from satellite and aerial imagery, runs computer vision models on them, and vectorizes the results through a Python API. Mosaics and predictions are written as Zarr, and vector results as GeoParquet. Built-in models cover common tasks:

  • SAM3 for text-prompted object detection on 30 cm NAIP.
  • Tile2Net for sidewalks, crosswalks, and roads on 30 cm NAIP.
  • Meta CHM v1 for canopy height on 60 cm NAIP.
  • ChesapeakeRSC for rural roads on 1 m NAIP.
  • Fields of the World for field boundaries on Sentinel-2 planting and harvest composites.

Custom PyTorch models run through the bring-your-own-model path. Sentinel-2 mosaics skip scenes with 75% or more cloud cover, mask cloud, shadow, and cirrus pixels with the Scene Classification Layer, and take a per-pixel median of what remains.

Pricing is calculated before a task runs, from area, resolution, bands, and time periods. Finding every solar panel array across a 500 km² county with SAM3 on 30 cm NAIP costs $35.00, mosaic and inference included.

# Detect solar panels across a county from a text prompt with SAM3 on 30 cm NAIP
from datetime import datetime
from rasterflow_remote import RasterflowClient
from rasterflow_remote.data_models import GeometryModelRecipes

rf = RasterflowClient()
detections = rf.predict_mosaic_geometries_recipe(
    aoi="s3://your-bucket/county.parquet",
    start=datetime(2022, 1, 1),
    end=datetime(2023, 1, 1),
    model_recipe=GeometryModelRecipes.SAM3_TEXT_GEOMETRY,
    text_prompt="solar panel",
    confidence_threshold=0.5,
)
print(detections.uri)  # GeoParquet containing detected solar-panel polygons

The GeoParquet output loads into WherobotsDB, built by the original creators of Apache Sedona, where a spatial join links each polygon to parcels, buildings, or hazard zones in SQL.

A Marion County walkthrough

In Marion County, Oregon, one RasterFlow pipeline built a 133 GB mosaic from 2022 NAIP imagery and ran SAM3 with eight text prompts in a single pass, keeping detections with a confidence of 0.5 or more. Wherobots publishes the output as wherobots_open_data.rasterflow_output_samples.marion_county_sam3_vector. This query summarizes it per prompt, with areas computed in UTM zone 10N (EPSG:32610) so they come out in square meters:

SELECT label,
       COUNT(*) AS n,
       ROUND(AVG(bbox_score), 3) AS mean_score,
       ROUND(SUM(ST_Area(ST_Transform(geometry, 'EPSG:4326', 'EPSG:32610'))) / 1e6, 3) AS area_km2
FROM wherobots_open_data.rasterflow_output_samples.marion_county_sam3_vector
GROUP BY label
ORDER BY n DESC
labelnmean_scorearea_km2mean area per detection
roads670,6250.57254.78982 m²
roofs312,2450.68241.710134 m²
solar panels18,5430.6142.101113 m²
shipping containers2,5540.5850.18271 m²
swimming pools1,4340.7250.06747 m²
airports6690.5710.254380 m²
airplanes1310.7560.00538 m²
tractors310.5500.00132 m²

The table holds 1,006,232 detections, and every row carries the same time stamp, 1 January 2022, the start of the 2022 NAIP window used to build the mosaic. The 312,245 roofs cover 41.7 km².

Bar chart on a log scale of SAM3 detections in Marion County, Oregon by text prompt: 670,625 roads, 312,245 roofs, 18,543 solar panels, 2,554 shipping containers, 1,434 swimming pools, 669 airports, 131 airplanes, and 31 tractors, with mean confidence and total area for each prompt
SAM3 detections per text prompt across Marion County, Oregon, queried in Wherobots: 1,006,232 polygons, including 312,245 roofs covering 41.7 km².

What the confidence scores show

Three patterns stand out in the per-prompt summary.

Scatter plot of the eight SAM3 text prompts in Marion County, Oregon, with mean area per detection on a log scale against mean confidence. Airplanes score 0.756 at 38 square meters, swimming pools 0.725 at 47, roofs 0.682 at 134, solar panels 0.614 at 113, shipping containers 0.585 at 71, roads 0.572 at 82, airports 0.571 at 380, and tractors 0.550 at 32
Mean confidence against mean object size for each SAM3 prompt in Marion County, from a GROUP BY query in Wherobots. Airplanes score highest (0.756) and tractors lowest (0.550).
  • Confidence varies by concept. Airplanes (0.756) and swimming pools (0.725) have distinct shapes and colors and score highest. Tractors (0.550) and airports (0.571) score lowest. Roofs sit between them at 0.682.
  • Rare prompts return few objects. Roads and roofs make up 97.7% of the rows, while the county returned 131 airplanes and 31 tractors. Any precision estimate for a rare class rests on a small sample.
  • The cutoff shapes every statistic. Detections under 0.5 were never stored, so the table cannot show the low-confidence half of the precision and recall curve.

Scoring roofs against Overture buildings

A detection is only as useful as its agreement with something independent. The county-wide evaluation of SAM3 roofs against Overture building footprints clipped both datasets to the official Marion County boundary, spatially joined SAM3 roofs to Overture Maps buildings on intersection in EPSG:32610, and computed IoU for every intersecting pair. Inside the boundary it counted 261,414 SAM3 roofs and 146,642 Overture buildings, and the join returned 207,109 candidate pairs. Total roof area agreed within 7%, and SAM3 matched the shape of three-quarters of known buildings. Many unmatched detections were real rural houses and outbuildings missing from Overture, so the authors judged real precision likely higher than 75%.

Research on satellite imagery analysis

Satellite imagery analysis rests on fifty years of research, from texture statistics computed on the first Landsat scenes to promptable segmentation models. Four threads shape it: how pixels became objects, how accuracy is measured, what benchmarks and foundation models changed, and which problems remain open.

From pixels to objects

The first computer analyses of satellite images classified one pixel at a time from its spectral values. Haralick, Shanmugam, and Dinstein added spatial context in Textural features for image classification (IEEE Transactions on Systems, Man, and Cybernetics, 1973): statistics of gray-tone co-occurrence in a neighborhood, tested on aerial photographs and multispectral imagery from the Earth Resources Technology Satellite, later renamed Landsat 1. Compton Tucker's study of red and photographic infrared combinations (Remote Sensing of Environment, 1979) compared linear combinations of red and near-infrared bands for monitoring vegetation, and the normalized difference it favored became NDVI.

Comparing dates came next. Ashbindu Singh's review of digital change detection techniques (International Journal of Remote Sensing, 1989) evaluated the procedures for comparing multitemporal images that had been developed by then, and it remains the starting point for the field. The Hansen forest map above ran that idea globally on Landsat time series.

Sub-meter imagery changed the unit of analysis. Once a house spans hundreds of pixels, a pixel classifier assigns roof texture, shadow, and lawn to separate classes. Thomas Blaschke's Object based image analysis for remote sensing (ISPRS Journal of Photogrammetry and Remote Sensing, 2010) describes the response: segment the image into objects first, then classify each object by its spectra, shape, and relations to its neighbors. Instance segmentation models such as SAM3 complete that shift, returning one polygon per object straight from the network.

Measuring accuracy

Russell Congalton's review of assessing the accuracy of classifications of remotely sensed data (Remote Sensing of Environment, 1991) set the error matrix as the core tool and named the choices that decide whether an accuracy number means anything: the classification scheme, the sampling design, the sample size, and spatial autocorrelation between samples. Nearby pixels are similar, so samples drawn close together are not independent, which is the first law of geography applied to validation. Object detection added the IoU threshold from PASCAL VOC, the measure behind the Marion County comparison.

The open problem is the reference itself. Overture buildings are a curated, multi-source dataset, and they still likely miss buildings in rural areas, which is why the Marion County evaluation treated many SAM3 "false positives" as real buildings. Roofs also differ from footprints: overhangs and occlusion mean a roof outline and a footprint never match exactly. Any accuracy number on overhead imagery is a statement about two datasets, and the dates, definitions, and gaps of the reference belong beside it.

Benchmarks and deep learning

Zhu, Tuia, Mou, and colleagues reviewed the arrival of neural networks in Deep learning in remote sensing: a comprehensive review and list of resources (IEEE Geoscience and Remote Sensing Magazine, 2017). Two ingredients made the shift: architectures from other fields and labeled benchmarks for overhead imagery.

U-Net (Ronneberger, Fischer, and Brox, MICCAI 2015) was designed for biomedical images and became the default segmentation network for satellite imagery, because its expanding path recovers the precise localization that building and road edges need. The benchmarks followed, alongside Fields of the World and Tile2Net from the examples above:

  • EuroSAT (Helber et al., 2019): 27,000 labeled Sentinel-2 patches in ten land use classes, across all 13 bands.
  • xView (Lam et al., 2018): over 1 million objects in 60 classes in 0.3 m WorldView-3 imagery.
  • SpaceNet (Van Etten, Lindenbaum, and Bacastow, 2018): building footprint and road network challenges.
  • xBD (Gupta et al., 2019): 850,736 building polygons across 19 disasters, labeled before and after.

Foundation models and embeddings

Every benchmark above needs labels for its task. Foundation models move most of the learning to unlabeled imagery or to large general datasets, and the research follows three lines.

Promptable segmentation. Kirillov and colleagues' Segment Anything (ICCV 2023) released over 1 billion masks on 11 million images and segmented an object from a point or a box. Meta's SAM 3 (Carion et al., 2025) replaced the click with a concept: a noun phrase such as "roofs" or an image exemplar, returning every matching instance, trained on a dataset with 4 million unique concept labels. That change is what lets one RasterFlow run detect roofs, roads, and solar panels across a county with no labels.

Pretrained Earth observation encoders. NASA and IBM's Prithvi (Jakubik et al., 2023) was pretrained on more than 1 TB of Harmonized Landsat Sentinel-2 imagery for fine-tuning on downstream tasks with small labeled datasets. Meta's canopy height model in the examples above used a self-supervised DINOv2 encoder.

Embeddings as a product. Rolf and colleagues' MOSAIKS (Nature Communications, 2021) computes image encodings once and shares them across tasks, so each new task needs only a linear regression on the user's own labels. Google DeepMind's AlphaEarth Foundations (Brown et al., 2025) takes the same approach with a learned model, compressing each location's imagery into annual global embedding layers from 2017 through 2024, so similarity and change become distance calculations.

Open problems

  • Domain shift. A model trained on one region, season, or sensor degrades on another. Tuia, Persello, and Bruzzone's overview of domain adaptation for remote sensing classification (IEEE Geoscience and Remote Sensing Magazine, 2016) frames the problem, and Fields of the World spans 24 countries to capture that diversity.
  • Reference data. Validation sets are often other models' outputs or volunteer maps with their own gaps, as the Marion County comparison shows.
  • Time. Imagery, labels, and reference vectors rarely share a date. A 2022 NAIP roof scored against a later building dataset mixes model error with real change.
  • Spatial dependence. Random train and test splits of neighboring patches leak information, the same autocorrelation Congalton warned about in 1991.

What scale changes

At county scale the model is the hard part. At continental scale the data movement is. Gorelick and colleagues described one answer in Google Earth Engine: planetary-scale geospatial analysis for everyone (Remote Sensing of Environment, 2017): a catalog of analysis-ready imagery next to parallel compute. The Cloud Optimized GeoTIFF layout defined above serves the same goal on object storage, where each reader fetches only the tiles it needs.

The vector side scales too. A billion model polygons need a spatial engine to join them to parcels, buildings, and boundaries. Jia Yu, Jinxuan Wu, and Mohamed Sarwat introduced GeoSpark (ACM SIGSPATIAL 2015) for distributed spatial queries on Apache Spark, and Yu, Zongsi Zhang, and Sarwat detailed its spatial partitioning, indexing, and join design in Spatial data management in Apache Spark: the GeoSpark perspective and beyond (GeoInformatica, 2019). GeoSpark became Apache Sedona. Kanchan Chowdhury and Sarwat's GeoTorch (ACM SIGSPATIAL 2022) brought a spatiotemporal deep learning framework into the same research line.

The Fields of the World global release shows the numbers at the far end. Taylor Geospatial ran the PRUE model on RasterFlow, generating 348.7 TB across 540,794 objects to produce 8.2 billion field boundaries. At that size, the Marion County query becomes one of thousands, and the spatial engine that runs them sets how long the answer takes.

Read more from Wherobots

Run a built-in RasterFlow model from the Model Hub at cloud.wherobots.com.

Frequently asked questions

How is AI used to analyze satellite images?

AI analyzes satellite images with computer vision models that classify each pixel, draw boxes or outlines around objects, estimate values such as tree height, and compare dates to find change. A pipeline builds a cloud-free mosaic, splits it into patches, runs the model on each patch, stitches the predictions back together, and converts them to polygons for analysis.

How do you interpret satellite imagery?

Start with what the sensor records: its resolution, bands, and capture date. True color shows the scene as the eye sees it, and false color combinations such as near-infrared, red, and green make vegetation stand out. Analysts read tone, texture, shape, size, shadow, and context, and AI models are trained on the same cues from labeled examples.

What is the best free satellite imagery for analysis?

Sentinel-2 is the most used free source for analysis, with 10 m bands and a revisit of about five days. Landsat adds a record back to 1972 at 30 m. In the United States, USDA NAIP aerial imagery reaches 30 cm to 1 m, fine enough for roofs and sidewalks. Sentinel-1 radar images through clouds.

What is the difference between object detection and segmentation in satellite imagery?

Object detection finds individual things, such as cars or roofs, and returns a box or outline for each one with a confidence score. Semantic segmentation labels every pixel with a class, such as road or sidewalk, without separating one object from the next. Instance segmentation combines both and returns a separate mask for each object.

Can AI detect objects in satellite images without training a model?

Yes, with promptable models. Meta’s SAM 3 takes a short noun phrase such as “roofs” or “solar panels” and returns a mask for every matching object, with no task-specific training. Results on overhead imagery vary by object, so teams check a sample against reference data before using the output.

What resolution do you need for satellite imagery analysis?

The object sets the resolution. A rule of thumb is several pixels across the smallest object. Fields and forests work at 10 m, buildings need about 1 m or finer, and roofs, solar panels, sidewalks, and vehicles need 30 to 60 cm. Finer imagery costs more to store and process per square kilometer.

What is an example of satellite imagery analysis?

A well-known example is Hansen and colleagues’ 2013 global forest map, which classified Landsat images at 30 m to find 2.3 million km² of forest lost from 2000 to 2012. Other examples include grading building damage after a hurricane, mapping every farm field boundary from Sentinel-2, and detecting roofs and solar panels in aerial imagery. In Marion County, Oregon, SAM3 on Wherobots RasterFlow returned 1,006,232 detections from eight text prompts, including 312,245 roofs.

Is Maxar free to use?

Most Maxar imagery is commercial and licensed for a fee. Its Open Data Program, now run under the Vantor name, releases before and after imagery of major sudden-onset disasters for free under a Creative Commons Attribution-NonCommercial 4.0 license, so it can be used with attribution for non-commercial purposes such as humanitarian response and research.

Can I see a satellite view of my house in real time?

No public service shows a live satellite view of a house. Google Earth and Google Maps show stored imagery from satellites and aircraft, and its capture date varies by location. Free satellites such as Sentinel-2 revisit about every five days at 10 m, too coarse to show a single house in detail.

How do you measure the accuracy of satellite imagery analysis?

For classified maps, compare a sample of pixels against reference data in an error matrix, as Congalton’s 1991 review set out. For detected objects, compute intersection over union (IoU) between each detection and a reference shape: a detection usually counts as correct above 0.5. Precision is the share of detections that are real and recall is the share of real objects found.