Rasterflow, Earth Intelligence & inference engine now in public preview Learn More

How We Delivered “Fields of The World” with RasterFlow: A Planetary-Scale GeoAI Pipeline

Author: Len Strnad

Fields of the World (FTW) is a Taylor Geospatial effort to produce globally consistent agricultural field-boundary data for land-use monitoring, food-system analysis, and model development. For the 2024-2025 global release, Taylor Geospatial partnered with Wherobots to run their latest FTW model, PRUE, on Wherobots RasterFlow; see the PRUE paper.

In this notebook, we walk through the production pipeline behind that release. The PRUE model turns seasonal Sentinel-2 composites into global field and field_boundaries predictions, and RasterFlow handles the larger systems problem: running data access, compositing, inference, and vectorization as one coordinated workflow and publishing the results as analysis-ready raster and vector products.

The focus here is practical architecture rather than model internals. We cover how the pipeline is structured, what it produced at global scale, and the implementation choices that made the run reproducible.

What we cover:

  1. RasterFlow and why it is designed for reproducible, scalable global runs.
  2. A Scale Profile with artifact sizes/object counts for this dataset.
  3. The RasterFlow workflow stages and resulting outputs:
  4. Implementation notes in Appendix — Technical Notes.

RasterFlow Execution Model for a Global-Scale GeoAI Pipeline

This pipeline runs on RasterFlow, Wherobots’ planetary-scale inference engine for Earth Intelligence. RasterFlow provides a reproducible way to run a global-scale GeoAI pipeline.

It is a geospatial workflow system backed by Flyte V2 and Ray. We chose Flyte because long-running global jobs need strong workflow semantics: versioned executions, retries, lineage, and reproducible reruns. We chose Ray Datasets because raster inference and vectorization are naturally blockwise workloads, and Ray gives us elastic parallelism and task streaming support to support the many small task problem.

The rasterflow_remote client exposes common GeoAI workflow steps while running in each organization’s isolated namespace. As a result, resource controls, failure containment, and reruns stay scoped to that organization.

We construct the client and load the model config for the PRUE model:

from rasterflow_remote.model_registry import get_model_registry_config, ModelRegistryEnum
from rasterflow_remote import RasterflowClient

client = RasterflowClient()

# see https://huggingface.co/wherobots/prue-pt2 for model details
cfg = get_model_registry_config(ModelRegistryEnum.HUGGINGFACE, "wherobots/prue-pt2")

Build Mosaic

build_mosaic converts raw satellite scenes into analysis-ready feature mosaics on a shared grid.

For the FTW production run, this stage computes planting and harvest composites, writes feature COGs, and then materializes a global feature mosaic that can be stacked with model outputs. We persist that global feature product in Zarr because chunked, cloud-native storage lets us read small spatial windows directly from object storage during validation, inference, and downstream analytics.

from rasterflow_remote import DatasetEnum
from datetime import datetime

mosaic = client.build_mosaic(
    datasets=[DatasetEnum.S2_MED_PLANTING, DatasetEnum.S2_MED_HARVEST],
    aoi="https://github.com/fieldsoftheworld/ftw-inference-app/blob/main/src/data/s2-grid.json",
    start=datetime(2023, 1, 1),
    end=datetime(2025, 1, 1),
    crs_epsg=4326,
)

For implementation details, see Appendix: Feature COG and Appendix: GDAL notes.

Inspect the published feature mosaic

The produced feature mosaic is published as a Zarr store on Source Cooperative. Therefore, we can inspect structure, bands, and a spatial subset directly.

import xarray as xr
import rasterix

mosaic_ds = xr.open_zarr(
    "https://data.source.coop/ftw/global-data/features/zarr/alpha/global.zarr"
).pipe(rasterix.assign_index)
mosaic_ds
<xarray.Dataset> Size: 502TB
Dimensions:      (time: 2, band: 10, y: 1566049, x: 4007517)
Coordinates:
  * time         (time) datetime64[ns] 16B 2024-01-01 2025-01-01
  * band         (band) object 80B 's2med_harvest:B02' ... 's2med_planting:N_...
  * y            (y) float64 13MB 83.75 83.75 83.75 ... -56.93 -56.93 -56.93
  * x            (x) float64 32MB -180.0 -180.0 -180.0 ... 180.0 180.0 180.0
    spatial_ref  int64 8B ...
Data variables:
    variables    (time, band, y, x) float32 502TB dask.array<chunksize=(1, 1, 4096, 4096), meta=np.ndarray>
Indexes:
  ┌ x        RasterIndex (crs=None)
  └ y

Persisting an intermediate global Zarr mosaic lets us query any region without reprojecting or resampling.

Training data for the PRUE model is also aligned to EPSG:4326, so feature and prediction stores remain grid-compatible for side-by-side analysis. For data layout details, see Appendix: Zarr internals.

For the examples below, we use the same Japan AOI shown later in the PMTiles preview so the raster and vector artifacts are easy to compare. For more information on xarray, see Appendix: Xarray notes.

import matplotlib.pyplot as plt
import numpy as np

bbox = {
    "x": slice(140.24, 140.34),
    "y": slice(37.36, 37.26),
}
subset = mosaic_ds.sel(x=bbox["x"], y=bbox["y"])
geo_aspect = 1 / np.cos(np.deg2rad(float(subset.y.mean())))

fig, axes = plt.subplots(1, 2, figsize=(9, 9), constrained_layout=True)
for ax, band_prefix, title in [
    (axes[0], "s2med_planting", "Planting RGB (PMTiles AOI, 2025)"),
    (axes[1], "s2med_harvest", "Harvest RGB (PMTiles AOI, 2025)"),
]:
    subset["variables"].sel(
        time="2025-01-01",
        band=[f"{band_prefix}:B04", f"{band_prefix}:B03", f"{band_prefix}:B02"],
    ).plot.imshow(ax=ax, robust=True)
    ax.set_aspect(geo_aspect)
    ax.set_title(title)

plt.show()

See additional information in Appendix: Zarr internals and Appendix: Xarray notes.

Predict Mosaic

predict_mosaic runs global model inference over the feature mosaic and writes one prediction Zarr store.

You provide the model config, and the workflow executes inference across the global grid. Keeping the result in Zarr preserves the same chunked access pattern as the feature store, so validation and downstream vectorization can stay blockwise and cloud-native instead of stitching regional outputs back together later.

predictions = client.predict_mosaic(
    store="https://data.source.coop/ftw/global-data/features/zarr/alpha/global.zarr",
    model_path=cfg.model_path,
    patch_size=cfg.patch_size,
    clip_size=cfg.clip_size,
    device=cfg.device,
    features=cfg.features,
    labels=cfg.labels,
    actor=cfg.actor,
    max_batch_size=cfg.max_batch_size,
    merge_mode=cfg.merge_mode,
)

For internals and scheduling details, see Appendix: config, Appendix: Ray usage, and Appendix: distributed inference.

For a similar inference step using text-prompted detection instead of segmentation models, see how Segment Anything 3 runs on aerial imagery with RasterFlow.

Inspect the published prediction mosaic

The prediction Zarr is also published on Source Cooperative. Let’s verify shape, bands, and value ranges.

The shape, bounds, and time remain aligned with the feature mosaic; however, the bands now represent predicted classes.

pred_ds = xr.open_zarr(
    "https://data.source.coop/ftw/global-data/predictions/zarr/alpha/global.zarr"
).pipe(rasterix.assign_index)
pred_ds
<xarray.Dataset> Size: 151TB
Dimensions:      (time: 2, band: 3, y: 1566049, x: 4007517)
Coordinates:
  * time         (time) datetime64[ns] 16B 2024-01-01 2025-01-01
  * band         (band) <U20 240B 'non_field_background' ... 'field_boundaries'
  * y            (y) float64 13MB 83.75 83.75 83.75 ... -56.93 -56.93 -56.93
  * x            (x) float64 32MB -180.0 -180.0 -180.0 ... 180.0 180.0 180.0
    spatial_ref  int64 8B ...
Data variables:
    variables    (time, band, y, x) float32 151TB dask.array<chunksize=(1, 1, 2048, 2048), meta=np.ndarray>
Indexes:
  ┌ x        RasterIndex (crs=None)
  └ y

Let’s plot the three labels for the same PMTiles-area AOI so the raster predictions line up with the vector preview later in the post:

import numpy as np

bbox = {
    "x": slice(140.24, 140.34),
    "y": slice(37.36, 37.26),
}
pred_subset = pred_ds.sel(x=bbox["x"], y=bbox["y"])
geo_aspect = 1 / np.cos(np.deg2rad(float(pred_subset.y.mean())))

grid = pred_subset["variables"].plot.imshow(col="band", row="time", figsize=(12, 8))
for ax in grid.axes.flat:
    ax.set_aspect(geo_aspect)

plt.show()

Zarr metadata for those curious is available in Appendix: Zarr internals.

Vectorize Mosaic

vectorize_mosaic converts prediction rasters into field-boundary GeoParquet outputs.

You specify feature bands, threshold, and method. Then the workflow writes distributed vector outputs for analytics and downstream publishing. We write GeoParquet because it keeps the result columnar, splittable, and easy to consume from batch analytics systems.

Conceptually, this output is a distributed geospatial table rather than an array: each row is a polygonized field candidate, the geometry column holds the boundary, and the remaining columns carry the scalar attributes you want to inspect in GeoPandas or query later in SedonaDB and WherobotsDB.

For distributed implementation details, see Appendix: vectorization.

from rasterflow_remote.data_models import VectorizeMethodEnum, SemSegRasterioConfig

vectors = client.vectorize_mosaic(
    store="https://data.source.coop/ftw/global-data/predictions/zarr/alpha/global.zarr",
    features=["field"],
    threshold=0.5,
    vectorize_method=VectorizeMethodEnum.SEMANTIC_SEGMENTATION_RASTERIO,
    vectorize_config=SemSegRasterioConfig(stats=False, medial_skeletonize=False),
)

A lightweight way to inspect that schema locally is:

import geopandas as gpd

vector_preview = gpd.read_parquet(vectors.uri)
vector_preview.head(3)

That gives you a normal GeoDataFrame view of the vector output: one geometry column plus the per-row attributes emitted by the vectorization stage.

Implementation details for distributed/blockwise vectorization are in Appendix: vectorization.

PMTiles

PMTiles packages vector outputs for fast, interactive map delivery. We use it here because a single archive is easy to publish, cache, and stream into browser maps without standing up a tile service.

We also support generating PMTiles from large GeoParquet outputs directly in WherobotsDB using our scalable PMTiles generator workflow, as described here: How to Generate PMTiles for Overture Maps with WherobotsDB VTiles.

Scale Profile

This section summarizes the scale of each pipeline step and the size of the generated outputs. Understanding the scale of this production run will help when we discuss design choices later in this blog post.

ArtifactPipeline StepS3 LocationSizeObject CountAdditional Scale Signal
Feature COGsbuild_mosaics3://us-west-2.opendata.source.coop/tge-labs/ftw-global-data/features/cogs/153 TB90,918Each COG is a median composite from ~5-10 Sentinel-2 scenes per pixel
Feature Zarr mosaicbuild_mosaics3://us-west-2.opendata.source.coop/tge-labs/ftw-global-data/features/zarr/alpha/global.zarr150 TB363,9997,499,140 logical chunks (sharded to keep object count manageable)
Prediction Zarr mosaicpredict_mosaics3://us-west-2.opendata.source.coop/tge-labs/ftw-global-data/predictions/zarr/alpha/global.zarr45 TB84,8778,982,630 logical chunks (also sharded)
Vector outputs (GeoParquet)vectorize_mosaics3://us-west-2.opendata.source.coop/tge-labs/ftw-global-data/predictions/vectors/alpha/results/675 GB1,0008,217,195,679 rows

Aggregate data output (above artifacts): ~348.7 TB across 540,794 objects.

RasterFlow summary

You can reproduce this end-to-end global-scale GeoAI pipeline with three API calls:

from datetime import datetime

from rasterflow_remote import RasterflowClient, DatasetEnum
from rasterflow_remote.model_registry import get_model_registry_config, ModelRegistryEnum
from rasterflow_remote.data_models import VectorizeMethodEnum, SemSegRasterioConfig

client = RasterflowClient()
cfg = get_model_registry_config(ModelRegistryEnum.HUGGINGFACE, "wherobots/prue-pt2")

mosaic = client.build_mosaic(
    datasets=[DatasetEnum.S2_MED_PLANTING, DatasetEnum.S2_MED_HARVEST],
    aoi="https://github.com/fieldsoftheworld/ftw-inference-app/blob/main/src/data/s2-grid.json",
    start=datetime(2023, 1, 1),
    end=datetime(2025, 1, 1),
    crs_epsg=4326,
)

predictions = client.predict_mosaic(
    store=mosaic.uri,
    model_path=cfg.model_path,
    patch_size=cfg.patch_size,
    clip_size=cfg.clip_size,
    device=cfg.device,
    features=cfg.features,
    labels=cfg.labels,
    actor=cfg.actor,
    max_batch_size=cfg.max_batch_size,
    merge_mode=cfg.merge_mode,
)

vectors = client.vectorize_mosaic(
    store=predictions.uri,
    features=["field"],
    threshold=0.5,
    vectorize_method=VectorizeMethodEnum.SEMANTIC_SEGMENTATION_RASTERIO,
    vectorize_config=SemSegRasterioConfig(stats=False, medial_skeletonize=False),
)

In short, this global-scale GeoAI pipeline with RasterFlow starts from global feature mosaics, runs one global inference workflow, and publishes both analytics-ready and map-ready outputs.

If you want a hands-on starting point, sign up for the RasterFlow Private Preview today and try the FTW notebook in the Solution Gallery.

If your team is evaluating global-scale raster + vector production workflows, this is exactly what RasterFlow private preview is for.

For API details and additional examples, see the Rasterflow documentation.

Appendix — Technical Notes

Building Feature COGs

Internally, build_mosaic orchestrates four key steps:

  1. Seasonal DOY filtering — candidate scenes are filtered by day-of-year windows derived from latitude-aware seasonal heuristics.
  2. Quality masking — cloud/quality masks suppress invalid observations.
  3. Greedy scene selection — we iteratively pick scenes that maximize uncovered valid area so we reach broad coverage with a bounded scene count.
  4. Temporal compositing — selected scenes are collapsed into median composites for downstream model features.

The current DOY heuristic and greedy objective are intentionally conservative and production-stable, but there is room for improvement (for example, region-specific phenology priors, dynamic cloud climatology, and cost-aware objective tuning).

We also produce a COG index, which is used with GDAL’s GTI driver to build the underlying global mosaic.

import geopandas as gpd
import fsspec

with fsspec.open("https://data.source.coop/ftw/global-data/features/cogs/alpha/index.parquet") as f:
    index = gpd.read_parquet(f)
index.head(3)
tile_id geometry time dataset url
0 01FBE POLYGON ((180 -50.59941, 178.76332 -50.5626, 1… 2024-01-01 00:00:00+00:00 s2med_planting s3://us-west-2.opendata.source.coop/tge-labs/f…
1 01FBE POLYGON ((180 -50.59941, 178.76332 -50.5626, 1… 2025-01-01 00:00:00+00:00 s2med_planting s3://us-west-2.opendata.source.coop/tge-labs/f…
2 01FBF POLYGON ((180 -49.70001, 178.84184 -49.66599, … 2024-01-01 00:00:00+00:00 s2med_planting s3://us-west-2.opendata.source.coop/tge-labs/f…

Mosaicking notes (GTI + Ray)

After feature COGs are written, we build a global virtual mosaic index and then materialize aligned arrays on the target grid. This relies on GDAL’s GTI driver as the index layer and mosaicking.

Grid configuration for this run:

  • CRS: EPSG:4326
  • Resolution: 8.983119e-5 degrees (~10 m at equator)
  • Resampling: cubic

The GTI-backed approach keeps COG indexing separate from downstream chunk/block execution and maps cleanly to distributed scheduling.

Zarr internals

Zarr is the storage layer that makes the global feature and prediction mosaics usable as products rather than just intermediate files.

Two layout choices matter here:

  • Logical chunks are the units we want readers and compute stages to address.
  • Shards are the larger storage objects we actually write to object storage.

We split them on purpose. Small logical chunks keep regional reads and retry boundaries precise, while larger shards prevent the store from exploding into billions of tiny objects. Larger shards also lower overall S3 PUT and LIST costs and reduce request-rate pressure.

For the feature store, the logical chunk is (1, 1, 4096, 4096) and shards are written at (1, 1, 12288, 12288), so each object packs a 3 x 3 tile of logical chunks. For the prediction store, the logical chunk is (1, 1, 2048, 2048) and shards are written at (1, 3, 8192, 8192), so each object packs all three class bands across a 4 x 4 spatial tile of logical chunks.

The tradeoff is deliberate: larger shards improve object-store efficiency and reduce listing overhead, but smaller logical chunks still give us better partial reads, safer retries, and more control over Ray execution geometry.

The feature Zarr mosaic metadata is as follows:

import zarr

group = zarr.open("https://data.source.coop/ftw/global-data/features/zarr/alpha/global.zarr")
group["variables"].metadata.__dict__
{'shape': (2, 10, 1566049, 4007517),
 'data_type': Float32(endianness='little'),
 'chunk_grid': RegularChunkGrid(chunk_shape=(1, 1, 12288, 12288)),
 'chunk_key_encoding': DefaultChunkKeyEncoding(separator='/'),
 'codecs': (ShardingCodec(chunk_shape=(1, 1, 4096, 4096), codecs=(BytesCodec(endian=<Endian.little: 'little'>), ZstdCodec(level=0, checksum=False)), index_codecs=(BytesCodec(endian=<Endian.little: 'little'>), Crc32cCodec()), index_location=<ShardingCodecIndexLocation.end: 'end'>),),
 'dimension_names': ('time', 'band', 'y', 'x'),
 'fill_value': np.float32(nan),
 'attributes': {'coordinates': 'spatial_ref', '_FillValue': 'AAAAAAAA+H8='},
 'storage_transformers': (),
 'extra_fields': {}}

The prediction zarr mosaic metadata is as follows:

group = zarr.open("https://data.source.coop/ftw/global-data/predictions/zarr/alpha/global.zarr")
group["variables"].metadata.__dict__
{'shape': (2, 3, 1566049, 4007517),
 'data_type': Float32(endianness='little'),
 'chunk_grid': RegularChunkGrid(chunk_shape=(1, 3, 8192, 8192)),
 'chunk_key_encoding': DefaultChunkKeyEncoding(separator='/'),
 'codecs': (ShardingCodec(chunk_shape=(1, 1, 2048, 2048), codecs=(BytesCodec(endian=<Endian.little: 'little'>), ZstdCodec(level=0, checksum=False)), index_codecs=(BytesCodec(endian=<Endian.little: 'little'>), Crc32cCodec()), index_location=<ShardingCodecIndexLocation.end: 'end'>),),
 'dimension_names': ('time', 'band', 'y', 'x'),
 'fill_value': np.float32(nan),
 'attributes': {'coordinates': 'spatial_ref', '_FillValue': 'AAAAAAAA+H8='},
 'storage_transformers': (),
 'extra_fields': {}}

Xarray notes

We use rasterix to attach analytical coordinates lazily so opening the global feature store does not require eagerly reading large coordinate vectors every time.

Why this matters:

  • x length: 1,566,049
  • y length: 4,007,517
  • At float64 precision, coordinates alone are roughly (1,566,049 + 4,007,517) * 8 ~= 45 MB before reading any model features.

Operationally, our xarray pattern is:

  1. Open lazily from object storage.
  2. Slice first (.sel) to bound I/O.
  3. Materialize only when plotting/validating.

This keeps exploratory analysis responsive while preserving direct comparability between feature and prediction stores on the same grid.

The Inference Config

The InferenceConfig is the contract between model packaging and distributed execution. It tells RasterFlow how to read features, how large each inference window should be, what runtime to launch, and how overlapping predictions should be merged back into the output mosaic.

In practice, the fields fall into a few groups:

  • model_path and actor select the model artifact and the inference runtime implementation.
  • patch_size and clip_size define the read window and how much border area is discarded to suppress edge artifacts.
  • features and labels define the exact tensor interface between the Zarr store and the model.
  • device and max_batch_size control throughput versus memory pressure on the target worker.
  • merge_mode defines how overlapping patch predictions are blended into the final raster.

from rasterflow_remote.data_models import MergeModeEnum, InferenceActorEnum

{
    # path to the pt2 archive on hugging face
    'model_path': 'https://huggingface.co/wherobots/prue-pt2/resolve/main/prue-efnetb7.pt2',
    # an enum used to specify how to handle the model for inference
    'actor': InferenceActorEnum.SEMANTIC_SEGMENTATION_PYTORCH,
    # patch size determines the size of the inputs provided to the model
    'patch_size': 256,
    # clip size determines the size we clip off the borders to reduce edge effects during inference
    'clip_size': 32,
    # device to run inference on, e.g. 'cuda' or 'cpu'
    'device': 'cuda',
    # list of features in order as required by the model. These names map to the input Zarr band labels.
    'features': ['s2med_harvest:B04',
    's2med_harvest:B03',
    's2med_harvest:B02',
    's2med_harvest:B08',
    's2med_planting:B04',
    's2med_planting:B03',
    's2med_planting:B02',
    's2med_planting:B08'],
    # the named labels for the predicted classes. These will be the output band labels in the output Zarr.
    'labels': ['non_field_background', 'field', 'field_boundaries'],
    # the maximum batch size assuming our default A10 GPU with 24GB of memory.
    'max_batch_size': 128,
    # the method to use when merging overlapping predictions to reduce edge effects at the patch level
    'merge_mode': MergeModeEnum.WEIGHTED_AVERAGE
}

How Ray is used for inference

We chose Ray as the data-plane engine because the workload is embarrassingly parallel at block level, but still needs data-aware scheduling and actor pools for GPU inference.

Ray is used as a bounded block-processing engine over sharded Zarr storage.

Key mapping model:

  • Logical chunks define storage math and indexability.
  • Shards reduce object count by packing multiple chunks per object.
  • Ray blocks define execution granularity and may span multiple chunks.

This decoupling lets us tune compute batch shape without rewriting storage layout. In the data pipeline, Ray Data uses Arrow-backed transport so tabular metadata/state can move across stages with minimal copying.

Practical takeaway: throughput scales with actor parallelism while per-task memory remains tied to block geometry, not global mosaic size.

Distributed inference (wherobots-rasterflow)

The InferenceConfig above feeds directly into block planning, actor setup, and overlap-aware merging. Internally, predict_mosaic orchestrates three phases:

  1. Block planning — derive execution blocks from chunk metadata and resource limits.
  2. Read/Infer/Write pipeline — execute overlap-aware reads, actor inference, and deterministic writes.
  3. Retry-safe commit — persist block outputs with stable target regions.
# conceptual : chunk/shard metadata -> Ray blocks

chunk_meta = read_zarr_chunk_metadata(store)
blocks = plan_ray_blocks(chunk_meta, target_rows_per_block)

ray_ds = ray.data.from_items(blocks)
result = (
    ray_ds
    .map(read_zarr_block_with_overlap)
    .map(run_inference_actor_pool)
    .map(clip_and_write_block)
)

Our distributed approach to vectorizing

vectorize_mosaic operates at block level so geometry generation scales linearly with partition count rather than requiring one global polygonization pass.

High-level flow:

  1. Read prediction blocks from Zarr.
  2. Apply thresholding per block.
  3. Polygonize per block.
  4. Emit GeoParquet partitions and merge metadata.

This keeps memory bounded and allows retries/reprocessing for individual partitions.

Sign up for RasterFlow Private Preview

Key takeaways

  • Taylor Geospatial partnered with Wherobots to run the PRUE model on RasterFlow for the 2024–2025 Fields of the World global release. PRUE turns seasonal Sentinel-2 composites into field and field_boundaries predictions; RasterFlow runs data access, compositing, inference, and vectorization as one workflow.
  • The pipeline is three API calls: build_mosaic (planting and harvest Sentinel-2 median composites, 2023-01-01 to 2025-01-01, EPSG:4326), predict_mosaic, and vectorize_mosaic (field band, threshold 0.5). RasterFlow is backed by Flyte V2 and Ray, running in each organization's isolated namespace.
  • Scale profile: feature COGs 153 TB (90,918 objects); feature Zarr mosaic 150 TB (363,999 objects, 7,499,140 logical chunks); prediction Zarr 45 TB (84,877 objects, 8,982,630 logical chunks); vector GeoParquet 675 GB (1,000 objects, 8,217,195,679 rows). Aggregate output ~348.7 TB across 540,794 objects.
  • Zarr layout splits logical chunks from shards on purpose: feature logical chunks (1, 1, 4096, 4096) packed into (1, 1, 12288, 12288) shards; prediction logical chunks (1, 1, 2048, 2048) packed into (1, 3, 8192, 8192) shards covering all three class bands. That keeps regional reads precise without exploding object count.

Wherobots 2024 accomplishments, and what’s on-deck in 2025

Editor’s note: The Wherobots Spatial Data Catalog is now the Havasu Catalog.

Introduction

2024 was a transformative year for Wherobots. Our mission to revolutionize how geospatial data is used took significant strides forward, positively impacting our customers and industry. Over the past year, we more than tripled the size of our team and successfully closed a $21.5M Series A funding round. We expanded accessibility to Wherobots’ industry-leading geospatial query performance, integrated Wherobots into the native AWS buying experience, and unveiled groundbreaking features like Raster Inference, Map Matching, and GeoStats—empowering users to create scalable geospatial solutions like never before.

Our Mission

Before founding Wherobots, co-founders Mo and Jia identified critical challenges limiting the potential of geospatial data. These stemmed from how geospatial data was traditionally stored, formatted, and processed, and made this data incredibly painful to utilize, particularly at scale.

Over the recent decades, data and analytics investment was mostly directed towards solutions for internet data. However compared to internet data, geospatial data is a lot more complex, which makes it harder to query. It’s polygons representing land and buildings, GPS trajectories, satellite and drone imagery, weather data, and more—all tied to Earth’s imperfect spherical surface. And querying this data generally means you need to filter and join it with other datasets (geo or non-geo). Due to this complexity, existing cloud analytics engines built for structured internet data struggle to efficiently run spatial queries at scale. They also miss features necessary to prepare this data, they lack features that make solution development productive, and simply cannot compute spatial results with high precision. As a result, solutions based on geospatial data are expensive, or otherwise shelved.

We are addressing these challenges. By reducing the cost and effort to build with geospatial data, Wherobots will enable a new wave of innovation for the physical world. This will drive breakthroughs in products, business operations, science, government, and make a positive impact on our climate.

Our mission is simple yet ambitious: make geospatial data easy to use.

Here’s what some of our customers have to say about how we’re helping them achieve their missions.

Customer Highlights

AddressCloud

Enabling insurers to calculate geographic risk with precision

“Wherobots runs our compute operations that used to take hours or days to complete, in minutes. As we provide perils information (flood, fire, etc) to insurers at the property level, we particularly appreciate the ability to be able to run combined vector/raster analysis, without having to previously transform the raster data into vector format or some other format.”
– John Powell, Senior Geospatial Data Engineer at AddressCloud

Overture Maps Foundation

Creating next-generation map products with scalable, open map data

“Overture produces a building dataset covering all buildings in the world, with 2.3B geometries and growing, that’s updated frequently. There’s a lot of data and compute that goes into producing and keeping it up to date,” said Jennings Anderson, Geoscientist at Overture and Data Engineer at Meta. “We accelerated the pipelines that produce the buildings dataset by up to 20x after we moved them to Wherobots, which required a simple redirection of our code. We retained compatibility with Apache Sedona, and the move put us into a development experience that’s made us more productive.”
– Jennings Anderson, Geoscientist at Overture and Data Engineer at Meta

Why Wherobots Stands Out

Several recurring themes highlight why customers choose Wherobots:

  • Unmatched performance and cost efficiency: Wherobots delivers up to 20x better spatial join performance compared to modern cloud data engines, at a fraction of the cost.

  • Ease of innovation: Wherobots makes it easy to build solutions with raster (e.g., satellite imagery), vector (e.g., geometry, geography) data, and your first party data regardless of scale.

  • Modern cloud architecture: Wherobots is fully compatible with Apache Sedona, and runs seamlessly on data lakes with support for Apache Iceberg and Apache Parquet.

2024 Milestones

Funding & Market Validation

In 2024, we raised $21.5M in a Series A round led by Felicis, with support from Wing Venture Capital, Clear Ventures, JetBlue Ventures, and P7 Ventures. This funding reflects confidence in our mission and the massive market opportunity for geospatial solutions in the cloud.

Team Growth

The Wherobots team—the “Botsters”—tripled in size this year. While engineering saw the most growth, we also built out go-to-market, marketing, and product teams and are actively scaling our sales team. As we head into 2025, we’re actively hiring for roles across the company to support our expanding vision.

Product Innovations

We launched several key features in 2024 that expanded the boundaries of geospatial data solutions. *The features noted with an are only available in the professional or enterprise edition of Wherobots.****

Cloud Native

  • A pay-as-you-go offering on the AWS Marketplace makes it easy to subscribe and pay on-demand using AWS Marketplace billing.
  • A storage integration for Amazon S3, to quickly and securely integrate with first or third party data.

Security and Access

  • SAML Single Sign-On, makes logging in simple, secure, and seamless for users in companies with centralized login management systems.
  • The Spatial SQL API, Typescript and Python SDKs, and a JDBC driver make it possible to query WherobotsDB using popular or custom query interfaces.
  • Continuous improvement of internal security and service availability.

Open Data Architecture

Accelerating Geospatial Solution Development

  • Raster Inference, to easily extract insights from satellite imagery at scale using SQL. (We’re hosting an upcoming panel discussion with an incredible lineup of speakers to discuss the MLM STAC Extension and Raster Inference, with a focus on optimizing Earth observation models for production. Learn more and save the date here.)
  • Distributed Map Matching is a purpose built algorithm for snapping GPS trajectories to known segments like roads, with high performance at-scale.
  • GeoStats: a geostatistics suite designed for scale, performance, and streamlining solution development.
  • Support for K-nearest neighbor joins (exact and approximate) to efficiently query for geospatial neighbors at-scale.
  • Vtiles, a vector tile solution purpose built for creating vector tiles at scale with high performance.
  • Many new vector (ST) and raster (RS) functions to accelerate the developer productivity.
  • Continuous improvement of spatial and non-spatial query performance to reduce cost and make workloads more compute efficient (reducing climate impact).

Automation

  • Job Runs integrated with Apache Airflow, to make it easy and familiar to automate new and existing processing workflows.
  • Service Principals, enable authentication and automated usage of Wherobots, decoupled from the tenure or privileges of human users.

Looking Ahead

In 2025, we plan to bring Wherobots Cloud to the EU market with support for the AWS Europe (Ireland) region, and achieve the SOC 2 Type 2 certification (currently in progress). We’ll continue to focus on:

  • Making Earth observation data easier to utilize.
  • Enhancing developer productivity and experiences.
  • Improving query engine performance and data compatibility.
  • Strengthening service availability and support for customers.
  • Delivering new administrative controls and observability.

Ready to Build?

We are currently offering a 30-day free trial covering up to $400 in usage via the AWS Marketplace. Getting started is easy. There are many example notebooks for various geospatial use cases that you can explore and run without any coding experience required. Not only do the notebooks help you get started, but we also see most of our customers use these notebooks as references for the solutions they end up building.

Join the Mission

Motivated by our mission? Join our growing team—visit our careers page for open roles. You can also share feedback at feedback@wherobots.com or contact me directly at damian@wherobots.com.

Try Wherobots Pro

Key takeaways

  • This January 2025 year-in-review covers 2024 milestones rather than a single product launch: the team more than tripled, and Wherobots closed a $21.5M Series A led by Felicis, with Wing, Clear Ventures, JetBlue Ventures, and P7 Ventures.
  • AddressCloud cut property-level flood/fire peril jobs from hours or days to minutes, including combined vector/raster analysis without converting rasters first. Overture global buildings pipeline (2.3 billion geometries in this post) accelerated up to 20x after a code redirect onto Wherobots.
  • The company cites up to 20x better spatial-join performance than modern cloud data engines, full raster+vector support, Apache Sedona compatibility, and Iceberg/Parquet lakehouse architecture.
  • 2024 product launches included AWS Marketplace pay-as-you-go, S3 storage integration, SAML SSO, Spatial SQL API plus TypeScript/Python SDKs and JDBC, Spatial Catalog/Havasu, Raster Inference, distributed map matching, GeoStats, KNN joins, VTiles, Airflow job runs, and service principals. Starred items are Professional/Enterprise only.
  • 2025 plans stated here: AWS Europe (Ireland), SOC 2 Type 2 (then in progress), easier Earth observation, developer experience, query performance, availability, and admin/observability. A 30-day AWS Marketplace trial covered up to $400 in usage.

WherobotsAI Raster Inference is GA with Support for Bring Your Own Model

Editor’s note: Raster inference is now part of RasterFlow, the Wherobots engine for running computer vision models on satellite and aerial imagery.

Introduction

UPDATE: Raster inference is included in Wherobots RasterFlow. See Wherobots Get Started with RasterFlow – Wherobots for the most up-to-date workflows.

We are excited to announce that WherobotsAI Raster Inference is now generally available! Raster Inference is a serverless, planetary-scale computer vision solution that enables data teams to extract meaningful insights from aerial imagery (raster data) sources, such as satellites or drones, and puts these insights at the fingertips of data scientists and developers.

During the preview period, customers analyzed raster data by running inference with a limited set of Wherobots-hosted open-source computer vision models. Now, you can bring your own model into Raster Inference, offload inference pipeline management effort, and apply this capability to a much broader set of use cases. We have also made significant enhancements to our compute service to accelerate inference performance.

Typical Use Cases

Data teams use WherobotsAI Raster Inference to identify information in complex, large-scale overhead imagery data. Some common use cases include:

  • Agriculture: Satellite and drone imagery are critical for monitoring land use, predicting crop yields, and improving sustainable farming practices.
  • Environment and Conservation: Raster imagery is essential for monitoring ecosystem and biodiversity changes, such as tracking glacier melting, sea level rise, temperature fluctuations, and environmental degradation (e.g., oil spills, deforestation).
  • Energy: Renewable energy developers analyze satellite data to assess land suitability, solar radiation, and wind patterns, to determine optimal locations for renewable energy projects.
  • Insurance: Insurers use raster data to calculate risk assessments based on environmental and local factors to conduct and improve damage assessments at scale.
  • Map Creation & Maintinence: Global digital, high fidelity map producers and maintainers extract features such as buildings, road networks, landcover, shipping lanes, etc., and their respective changes in order to maintain truthful connection between global map data products and the real world.

Traditional challenges with computer vision pipelines

We’ve met with many businesses struggling to get critical insights from raster data. Some are manually sifting through imagery. This method doesn’t scale, it’s expensive, error prone, and time intensive. Others are utilizing complex computer vision solutions that are not designed for overhead raster imagery. These computer visions solutions:

  • Take time to build, and require significant management to efficiently load, store, and process large raster datasets.
  • Are difficult to scale to accommodate increasing workload sizes.
  • Are fragile with multiple components involved and challenges maintaining compatibility with an evolving modeling stack.
  • Require effort to experiment, test, and integrate new models and inference runs.

Benefits of WherobotsAI Raster Inference

With WherobotsAI Raster Inference, you can:

  • Use data pipelines that are ready for small to planetary-scale raster data.
  • Deploy an on-demand solution in seconds that scales to meet workload needs without having to manage infrastructure.
  • Easily import your model or utilize any model hosted by Wherobots.
  • Experiment and generate critical insights faster.

In the following sections, we’ll provide a brief overview of Wherobots-hosted models, how to bring your own model, recent performance improvements, and how to use this feature effectively.

Choosing a Model for Raster Inference

Wherobots-Hosted Models

Wherobots-hosted models are precompiled and optimized for raster inference, enabling them to scale effortlessly and execute inference on large datasets. Below is a brief overview of the initial set of Wherobots-hosted models. We plan to host additional models based on customer feedback. For more information on these models and their performance metrics, see our docs page here.

Wherobots Hosted Models

Below is a code snippet on how to use landcover-eurosat-sentinel2 in WherobotsAI Raster Inference:

# set the model name

model_id = "landcover-eurosat-sentinel2"

# call the model in raster inference
df_predictions = df_raster_input.withColumn("preds", rs_classify(model_id, "outdb_raster"))

The STAC Machine Learning Model Extension Specification

The model import function of Raster Inference is built on the STAC Machine Learning Model (MLM) Extension Specification, an open community standard for model sharing we co-developed with CRIM and other collaborators. The MLM specification enables model portability, making it easier to use models across teams and compute platforms.

Before the MLM specification, data scientists and modelers sharing geospatial computer vision models often had to adapt existing standards, such as HuggingFace model cards, to store relevant information. This involved repurposing the cards to specify details like required raster bands, necessary data preprocessing, and post-processing functions. Without a community standard for organizing this information, these model cards were often inconsistent and documented in varying formats, making model sharing and reproducibility across different compute platforms cumbersome.

The STAC MLM extension introduces a community standard designed to simplify storing and sharing geospatial computer vision models. It achieves this by providing a comprehensive schema to:

  • Describe critical geospatial model attributes, such as geolocation and temporal range.
  • Include key model inference reproducibility details, such as required bands, model artifact locations, and pre- and post- processing steps.
  • Enable model collections to be searched alongside associated spatiotemporal datasets.

The MLM specification has already been adopted in key modeling efforts at Terradue and is proudly supported by Radiant Earth. We’ll share more in a series of blog posts and a panel discussion with our collaborators on the MLM in late January — register here on the interest form to receive an invite when the date is finalized.

Bring Your own Model

The MLM specification enables users to quickly and easily use many geospatial, deep learning-based computer vision models with Raster Inference. We leverage a model’s MLM specification to specify and integrate the required data preprocessing, model, and post processing into the larger raster inference pipeline.

To bring your own model to Raster Inference:

  1. Fill out our MLM form for your model. This will create a MLM formatted JSON file (MLM JSON) describing your model. Download the file.
  2. Upload the MLM JSON file from Step 1 to your AWS S3 bucket.
  3. Copy and save the AWS S3 URI to your MLM JSON.
  4. During runtime, use your model for inference by calling the raster inference function with your MLM’s S3 URI link.

You can find a full walkthrough in our documentation on using the MLM form to bring your own model.

Performance Improvements

In addition to enabling bring your own model, we’ve accelerated asynchronous data loading in the inference engine to boost performance. Below, you can see how performance has evolved with experiments conducted using WherobotsAI Raster Inference and a Tiny GPU Runtime.

Example: Raster Inference with Bring Your own Model

For a full tutorial example on how to bring your own model to segment solar farms in Sentinel-2 imagery, see our documentation here.

Example: Run Raster Inference with a Wherobots-hosted model

We’ll walk through an example on how to identify solar infrastructure in raster imagery using a Wherobots-hosted model. In our example, we will be using the model:

solar-satlas-sentinel2

The full Python notebook file can be found on our GitHub.

  1. Set up the WherobotsDB context
import warnings
warnings.filterwarnings('ignore')

from wherobots.inference.data.io import read_raster_table
from sedona.spark import SedonaContext
from pyspark.sql.functions import expr

from wherobots.inference.engine.register import create_semantic_segmentation_udfs
from pyspark.sql.functions import col

config = SedonaContext.builder().appName('segmentation-batch-inference')
    .getOrCreate()

sedona = SedonaContext.create(config)
  1. Load Satellite Imagery
tif_folder_path = "s3a://wherobots-benchmark-prod/data/ml/satlas"
files_df = read_raster_table(tif_folder_path, sedona, limit=400)
df_raster_input = files_df.withColumn(
        "outdb_raster", expr("RS_FromPath(path)")
    )

df_raster_input.cache().count()
df_raster_input.show(truncate=False)
df_raster_input.createOrReplaceTempView("df_raster_input")
  1. Run WherobotsAI Raster Inference

Specify a Wherobots-hosted model to run inference

model_id = "solar-satlas-sentinel2"

You can run WherobotsAI Raster Inference using either the Wherobot’s SQL API or Python API.

Using the SQL API

predictions_df = sedona.sql("""
SELECT
  outdb_raster,
  segment_result.*
FROM (
  SELECT
    outdb_raster,
    RS_SEGMENT('{model_id}', outdb_raster) AS segment_result
  FROM
    df_raster_input
) AS segment_fields
""")

predictions_df.cache().count()
predictions_df.show()
predictions_df.createOrReplaceTempView("predictions")

Using the Python API

rs_segment =  create_semantic_segmentation_udfs(batch_size = 10, sedona=sedona)
df = df_raster_input.withColumn("segment_result", rs_segment(model_id, col("outdb_raster"))).select(
                               "outdb_raster",
                               col("segment_result.confidence_array").alias("confidence_array"),
                               col("segment_result.class_map").alias("class_map")
                           )
df.show(3)
  1. Extract predicted geometries (continued from step 4 using the SQL API)
df_multipolys = sedona.sql("""
    WITH t AS (
        SELECT RS_SEGMENT_TO_GEOMS(outdb_raster, confidence_array, array(1), class_map, 0.65) result
        FROM predictions
    )
    SELECT result.* FROM t
""")

df_multipolys.cache().count()
df_multipolys.show()
df_multipolys.createOrReplaceTempView("multipolygon_predictions")

df_merged_predictions = sedona.sql("""
    SELECT
        element_at(class_name, 1) AS class_name,
        cast(element_at(average_pixel_confidence_score, 1) AS double) AS average_pixel_confidence_score,
        ST_Collect(geometry) AS merged_geom
    FROM
        multipolygon_predictions
""")
df_filtered_predictions = df_merged_predictions.filter("ST_IsEmpty(merged_geom) = False")
df_filtered_predictions.cache().count()
df_filtered_predictions.show()
  1. Visualize results
from sedona.maps.SedonaKepler import SedonaKepler
config = {
    'version': 'v1',
    'config': {
        'mapStyle': {
            'styleType': 'dark',
            'topLayerGroups': {},
            'visibleLayerGroups': {},
            'mapStyles': {}
        },
    }
}
map = SedonaKepler.create_map(config=config)

SedonaKepler.add_df(map, df=df_filtered_predictions, name="Solar Farm Detections")
map

Get started with WherobotsAI Raster Inference

Professional Edition Users

If you’re a Wherobots Professional Edition or Enterprise user, you have access to all capabilities of WherobotsAI Raster Inference! If you have access to GPU runtimes, sign in to your account now to launch a Wherobots Notebook and explore the feature. If you don’t yet have access, request it today and start using Raster Inference as soon as tomorrow. Explore how easy it is to bring your own model using our guided example on GitHub or within a Wherobots notebook instance. Start integrating WherobotsAI Raster Inference into your workflow today!

Community Edition Users

Although Raster Inference is not available in Wherobots Community Edition, we are currently offering a free trial for the Wherobots Professional Edition. You can either sign up through AWS Marketplace or upgrade your account to get started for free and integrate Raster Inference into your workflow today.

What’s next

We’re eager to hear about the models you’d like us to support and any features you’d like to see added. For product feedback, feel free to email us at feedback@wherobots.com (no request is too small). To see how others are using Raster Inference, ask questions, and share your own experiences, join the Wherobots community. We look forward to your ideas and creations!

Missed our panel discussion with our collaborators CRIM, Terradue and Radiant Earth on the MLM STAC extension? Watch the recording below.

Get Started with Wherobots

Key takeaways

  • WherobotsAI Raster Inference is generally available as a serverless, planetary-scale computer vision service that extracts features from satellite and drone raster imagery for data scientists and developers.
  • Bring-your-own-model is now supported through the STAC Machine Learning Model (MLM) Extension Specification, co-developed with CRIM and other collaborators, adopted at Terradue, and supported by Radiant Earth.
  • Wherobots-hosted models are precompiled for scale; the post walks through landcover-eurosat-sentinel2 classification and solar-satlas-sentinel2 semantic segmentation, including converting predictions to geometries at a 0.65 confidence threshold.
  • The GA release also accelerated asynchronous data loading in the inference engine, with experiments shown for a Tiny GPU Runtime (chart on the page; no numeric speedups are stated in the text).
  • Raster Inference is included for Professional and Enterprise users with GPU runtimes, not Community Edition. Community users can start a Professional trial via AWS Marketplace or an account upgrade.

Introducing GeoStats for WherobotsAI and Apache Sedona

Introducing GeoStats for WherobotsAI and Apache Sedona

We are excited to introduce GeoStats, a machine learning (ML) and statistical toolbox for WherobotsAI and Apache Sedona users. With GeoStats, you can easily identify critical patterns in geospatial datasets such as hotspots and anomalies, and quickly get critical insights from large scale data. While these algorithms are supported in other packages, we’ve optimized each algorithm to be highly performant for small to planetary scale geospatial workloads. That means, you can get results from these algorithms significantly faster, at a lower cost, and do it all more productively, through a unified development experience purpose-built for geospatial data science and ETL.

The Wherobots toolbox supports DBSCAN, Local Outlier Factor (LOF), and Getis-Ord (Gi*) algorithms. Apache Sedona users can utilize DBSCAN starting with Apache Sedona version 1.7.0, and like all other features of Apache Sedona, its fully compatible with Wherobots.

Use Cases for GeoStats

DBSCAN

DBSCAN is the most popular algorithm we see in geospatial use cases. It identifies clusters, areas of your data that are closely packed together, and outliers, areas of your data that are set apart.

Typical use cases for DBSCAN are found in:

  • Retail: Decision makers use DBSCAN with location data to understand areas of high and low pedestrian activity to decide where to setup retailing establishments.
  • City Planning: City planners use DBSCAN with GPS data to optimize transit support by identifying high usage routes, areas in need of additional transit options, or areas that have too much support.
  • Air Traffic Control: Traffic controllers use DBSCAN to identify areas with increasing weather activity to optimize flight routing.
  • Risk computation: Insurers and others can use DBSCAN to make policy decisions and calculate risk where risk is correlated to the proximity of two or more features of interest.

Local Outlier Factor (LOF)

LOF is an anomaly detection algorithm that identifies outliers present in a dataset.

Typical use cases for LOF include:

  • Data analysis and cleansing: Data teams can use LOF to identify and remove anomalies within a dataset, like removing erroneous GPS data points from a trace dataset

Getis-Ord (Gi*)

Getis-Ord is also a popular algorithm for identifying local hot and cold spots.

Typical use cases for Gi* include:

  • Public health: Officials can use disease data with Gi* to identify areas of abnormal disease outbreak
  • Telecommunications: Network administrators can use Gi* to identify areas of high demand and optimize network deployment
  • Insurance: Insurers can identify areas prone to specific claims to better manage risk
Click here to launch this interactive notebook

Traditional challenges with using these algorithms on geospatial data

Before GeoStats, teams leveraging any of the algorithms in the toolbox in a data analysis or ML pipeline would:

  1. Struggle to get performance or scale from the underlying solutions that also don’t perform well when joining geospatial data.
  2. Determine how to host and scale open source versions of popular ML and statistical algorithms, like PostGIS or scikit-learn DBSCAN, PySal Gi*, or scikit-learn LOF, to work for geospatial data types and geospatial data formats.
  3. Replicate this overhead each time they want to deploy a new algorithm for geospatial data.

Benefits of WherobotsAI GeoStats

With GeoStats in WherobotsAI, you can now:

  1. Easily run native algorithms on a cloud-based engine, optimized for producing spatial data products and insights at scale.
  2. Use these algorithms without the operational overhead associated with setup and maintenance.
  3. Leverage optimized, hosted algorithms within a single platform to easily experiment and get critical insights faster.

We’ll walk through a brief overview of each algorithm, how to use them, and show how they perform at various scales.

Diving Deeper into the GeoStats Toolbox

DBSCAN Overview

DBSCAN is a density-based clustering algorithm. Given a set of points in some space, it groups points with many nearby neighbors and marks as outlier points that lie alone in low-density regions.

How to use DBSCAN in Wherobots

The following examples assume you have already setup an organization and have an active runtime and notebook, with a dataframe of interest to run the algorithms on.

WherobotsAI GeoStats DBSCAN Python API Overview
For a full walk through see the Python API reference: dbscan(...).

  • Supported Geometries: points, linestrings, polygons
  • Hyperparameters: max distance to neighbors (epsilon), min neighbors (min_points)
  • Output: dataframe with cluster id

DBSCAN Walk Through

  1. Choose your dataset and create a Sedona DataFrame.
dataset=sedona.createDataFrame(X).select(ST_MakePoint("_1", "_2").alias("geometry"))
  1. Choose values for your hyperparameters, max distance to neighbors (epsilon) and minimum neighbors (min_points). These values will determine how DBSCAN identifies clusters.
epsilon=0.3
min_points=10
  1. Run DBSCAN on your DataFrame with your chosen hyperparameter values.
clusters_df = dbscan(df, epsilon=0.3, min_points=10, include_outliers=True)
  1. Analyze the results. For each datapoint, DBSCAN returns the cluster it’s associated with or if it’s an outlier.
+--------------------+------+-------+
|            geometry|isCore|cluster|
+--------------------+------+-------+
|POINT (1.22185277...| false|      1|
|POINT (0.77885034...| false|      1|
|POINT (-2.2744742...| false|      2|
+--------------------+------+-------+

only showing top 3 rows

There’s a complete example of how to use DBSCAN in the Wherobots user documentation.

DBSCAN Performance Overview

To show DBSCAN performance in Wherobots, we created a European sample of the Overture buildings dataset, and ran DBSCAN to identify clusters of buildings near each other, starting from the geographic center of Europe and worked outwards. For each subsampled dataset, we run DBSCAN with an epsilon of 0.005 degrees (i.e. ~30 feet) and min_points value of 4 on a Large runtime in Wherobots Cloud. As seen below, DBSCAN effectively processes an increasing number of records, with 100M records taking 1.6 hrs to process.

Local Outlier Factor (LOF)

LOF is an anomaly detection algorithms that identifies outliers present in a dataset. It does this by measuring how close a given data point is to a set of k-nearest neighbors (with k being a user chosen hyperparameter) in comparison to how close its nearest neighbors are to their nearest neighbors. LOF provides a score that represents the degree to which a record is an inlier or outlier.

How to use LOF in Wherobots

For the full example, please see this docs page.

WherobotsAI GeoStats LOF Python API Overview
For a full walk through see the Python API reference: local_outlier_factor(...).

  • Supported Geometries: points, linestrings, polygons
  • Hyperparameters: number of nearest neighbors to use
  • Output: score representing degree of inlier or outlier

LOF Walk Through

  1. Choose your dataset and create a Sedona DataFrame.
df = sedona.createDataFrame(X).select(ST_MakePoint(f.col("_1"), f.col("_2")).alias("geometry"))
  1. Choose your k value for how many nearest neighbors you want to use to measure density near a given datapoint.
k=20
  1. Run LOF on your DataFrame with your chosen k value.
outliers_df = local_outlier_factor(df, k=20)
  1. Analyze your results. LOF returns a score for each datapoint representing the degree of inlier or outlier.
+--------------------+------------------+
|            geometry|               lof|
+--------------------+------------------+
|POINT (-1.9169927...|0.9991534865548664|
|POINT (-1.7562422...|1.1370318880088373|
|POINT (-2.0107478...|1.1533763384772193|
+--------------------+------------------+
only showing top 3 rows

There’s a complete example of how to use LOF in the Wherobots user documentation.

LOF Performance Overview

We followed the same procedure with DBSCAN but ran LOF to identify clusters of buildings near each other. With each set of buildings we ran LOF with a k=20 on a large Wherobots Cloud runtime. As seen below, GeoStats LOF scales effectively with growing data size with 100M records taking 10 mins to process.

Getis-Ord (Gi*) Overview

Getis-Ord is an algorithm for identifying statistically significant local hot and cold spots.

How to use GeoStats Gi*

WherobotsAI GeoStats Gi* Python API Overview
For the full example, please see this docs g_local(...).

  • Supported Geometries: points, linestrings, polygons
  • Hyperparameters: star, neighbor weighting
  • Output: Set of statistics that indicate the degree of local hot or cold spot for a given record

Gi* Walk Through

  1. Choose your dataset and create a Sedona Dataframe.
places_df = (
    sedona.table("wherobots_open_data.overture_2024_07_22.places_place")
        .select(f.col("geometry"), f.col("categories"))
        .withColumn("h3Cell", ST_H3CellIDs(f.col("geometry"), h3_zoom_level, False)[0])
)
  1. Choose how you’d like to weight datapoints (ex: datapoints in a specific geographic area need to be weighted higher or any datapoint close to a given datapoint need to be weighted higher) and star (boolean to indicate if a record is a neighbor of itself).
star = True
neighbor_search_radius_degrees = 1.0
variable_column = "myNumericColumnName"

weighted_dataframe = add_binary_distance_band_column(
        df,
        neighbor_search_radius_degrees,
        include_self=star
)
  1. Run Gi* on your DataFrame with your chosen hyperparameters.
gi_df = g_local(
        weighted_dataframe,
    variable_column,
    star=star
)
  1. Analyze your results. For each datapoint, Gi* returns a set of statistics that indicate the degree of local hot or cold spot.
+----------+-------------------+--------------------+--------------------+------------------+--------------------+
|num_places|                  G|                  EG|                  VG|                 Z|                   P|
+----------+-------------------+--------------------+--------------------+------------------+--------------------+
|       871| 0.1397485091609774|0.013219284603421462|5.542296862370928E-5|16.995969941572465|                 0.0|
|       908|0.16097739240211956|0.013219284603421462|5.542296862370928E-5|19.847528249317246|                 0.0|
|       218|0.11812096144582315|0.013219284603421462|5.542296862370928E-5|14.090861243071908|                 0.0|
+----------+-------------------+--------------------+--------------------+------------------+--------------------+
only showing top 3 rows

There’s a complete example of how to use Gi* in the Wherobots user documentation.

Getis-Ord Performance Overview

To showcase how Gi performs in Wherobots, again we used the same example as DBSCAN, but ran Gi on the area of the buildings. With each set of buildings we ran Gi* with a binary neighbor weight and a neighborhood radius of .007 degrees (~0.5 miles) on a Large runtime in Wherobots Cloud. As seen below, the algorithm scales mostly linearly with the number of records, with 100M records taking 1.6 hours to process.

Get started with WherobotsAI GeoStats

The way we implemented these algorithms for large scale geospatial workloads, will help you make sense of your geospatial data faster. You can get started for free today.

  • If you haven’t already, create a free Wherobots Organization subscribed to the Community Edition of Wherobots.
  • Start a Wherobots Notebook
  • In the Notebook environment, explore the notebook_example/python/wherobots-ai/ folder for examples that you can use to get started.
  • Need additional help? Check out our user documentation, and send us a note if needed at support@wherobots.com.

Apache Sedona Users

Apache Sedona users will have access to GeoStats DBSCAN with the Apache Sedona 1.7.0 release. Subscribe to the Spatial Intelligence Newsletter and join the Sedona community to get notified of the release and get started!

What’s next

We’re excited to hear what ML and statistical algorithms you’d like us to support. We can’t wait for your feedback and to see what you’ll create!

Start building with Wherobots

Key takeaways

  • GeoStats is a machine-learning and statistics toolbox for WherobotsAI and Apache Sedona, with DBSCAN clustering, Local Outlier Factor (LOF) anomaly detection, and Getis-Ord Gi* hot/cold spots.
  • Apache Sedona users get DBSCAN starting in Sedona 1.7.0; the Wherobots toolbox includes all three algorithms and stays Sedona-compatible.
  • On a Large Wherobots runtime, DBSCAN on a growing European Overture buildings sample (epsilon 0.005 degrees, about 30 feet, min_points 4) processed 100 million records in 1.6 hours.
  • LOF on the same buildings progression with k=20 processed 100 million records in 10 minutes on a Large runtime.
  • Gi* with binary neighbor weights and a 0.007-degree (about 0.5 mile) radius processed 100 million records in 1.6 hours on a Large runtime, scaling mostly linearly.

WherobotsAI Map Matching: Turn GPS Tracks into Road-Aligned Trip Insights at Scale

What is Map Matching?

GPS data is inherently noisy and often lacks precision, which can make it challenging to extract accurate insights. This imprecision means that the GPS points logged may not accurately represent the actual locations where a device was. For example, GPS data from a drive around a lake may incorrectly include points that are over the water!

Click here to launch this interactive notebook

To address these inaccuracies, teams commonly use two approaches:

  1. Identifying and Dropping Erroneous Points: This method involves manually or algorithmically filtering out GPS points that are clearly incorrect. However, this approach can reduce analytical accuracy, be costly, and is time-intensive.
  2. Map Matching Techniques: A smarter and more effective approach involves using map matching techniques. These techniques take the noisy GPS data points and compute the most likely path taken based on known transportation segments such as roadways or trails.

WherobotsAI Map Matching offers an advanced solution for this problem that’s significantly more performant and 99% lower cost than Mapbox, Google Maps, or HERE maps alternatives. It performs map matching with high scale on millions or even billions of trips with ease and performance, ensuring that the GPS data aligns accurately with the actual paths most likely taken.

map matching telematics

An illustration of map matching. Blue dots: GPS samples, Green line: matched trajectory.

Map matching is a common solution for preparing GPS data for use in a wide range of applications including:

  • Sattelite & GPS based navigation
  • GPS tracking of freight
  • Assessing risk of driving behavior for improved insurance pricing
  • Post hoc analysis of self driving car trips for telematics teams
  • Transportation engineering and urban planning

The objective of map matching is to accurately determine which road or path in the digital map corresponds to the observed geographic coordinates, considering factors such as the accuracy of the location data, the density and layout of the road network, and the speed and direction of travel.

Existing Solutions for Map Matching

Most map matching implementations are variants of the Hidden Markov Model (HMM)-based algorithm described by Newson and Krumm in their seminal paper, “Hidden Markov Map Matching through Noise and Sparseness.” This foundational research has influenced a variety of map matching solutions available today.

However, traditional HMM-based approaches have notable downsides when working with large-scale GPS datasets:

  1. Significant Costs: Many commercially available map matching APIs charge substantial fees for large-scale usage.
  2. Performance Issues: Traditional map matching algorithms, while accurate, are often not optimized for large-scale computation. They can be prohibitively slow, especially when dealing with extensive GPS data, as the underlying computation struggles to handle the data scale efficiently.

These challenges highlight the need for more efficient and cost-effective solutions capable of handling large-scale GPS datasets without compromising on performance.

RESTful API Map Matching Options

The Mapbox Map Matching API, HERE Maps Route Matching API, and Google Roads API are powerful RESTful APIs for performing map matching. These solutions are particularly effective for small-scale applications. However, for large-scale applications, such as population-level analysis involving millions of trajectories, the costs can become prohibitively high.

For example, as of July 2024, the approximate costs for matching 1 million trips are:

  • Mapbox: $1,600
  • HERE Maps: $4,400
  • Google Maps Platform: $8,000
  • WherobotsAI Map Matching: ~$10 (see below)

These prices are based on public pricing pages and do not consider any potential volume-based discounts that may be available.

While these APIs provide robust and accurate map matching capabilities, organizations seeking to perform extensive analyses often must explore more cost-effective alternatives.

Open-Source Map Matching Solutions

Open-source software such as such as Valhalla or GraphHopper can also be used for map matching. However, these solutions are designed for use on a single-machine. If your map matching workload exceeds the capacity that machine, your workload will suffer from extended processing times. Furthermore, you will end up running out of headroom if you are vertically scaling up the ladder of VM sizes.

Meet WherobotsAI Map Matching

Map Matching is a high performance, low cost, and planetary scale map matching capability for your telematics pipelines.

WherobotsAI provides a scalable map matching feature designed for small to very large scale trajectory datasets. It works seamlessly with other Wherobots capabilities, which means you can implement data cleaning, data transformations, and map matching in one single (serverless) data processing pipeline. We’ll see how it works in the following sections.

How it works

WherobotsAI Map Matching takes a DataFrame containing trajectories and another DataFrame containing road segments, and produces a DataFrame containing map matched results. Here is a walk-through of using WherobotsAI Map Matching to match trajectories in the VED dataset to the OpenStreetMap (OSM) road network.

1. Preparing the Trajectory Data

First, we load the trajectory data. We’ll use the preprocessed VED dataset stored as GeoParquet files for demonstration.

dfPath = sedona.read.format("geoparquet").load("s3://wherobots-benchmark-prod/data/mm/ved/VED_traj/")

The trajectory dataset should contain the following attributes:

  • A unique ID for trips. In this example the ids attribute is the unique ID of each trip.
  • A geometry attribute containing LineStrings, in this case the geometry attribute is for trip data.

The rows in the trajectory DataFrame look like this:

+---+-----+----+--------------------+--------------------+
|ids|VehId|Trip|              coords|            geometry|
+---+-----+----+--------------------+--------------------+
|  0|    8| 706|[{0, 42.277558333...|LINESTRING (-83.6...|
|  1|    8| 707|[{0, 42.277681388...|LINESTRING (-83.6...|
|  2|    8| 708|[{0, 42.261997222...|LINESTRING (-83.7...|
|  3|   10|1558|[{0, 42.277065833...|LINESTRING (-83.7...|
|  4|   10|1561|[{0, 42.286599722...|LINESTRING (-83.7...|
+---+-----+----+--------------------+--------------------+
only showing top 5 rows

2. Preparing the Road Network Data

We’ll use the OpenStreetMap (OSM) data specific to the Ann Arbor, Michigan region to map match our trip data with. Wherobots provides an API for loading road network data from OSM XML files.

from wherobots import matcher
dfEdge = matcher.load_osm("s3://wherobots-examples/data/osm_AnnArbor_large.xml", "[car]")
dfEdge.show(5)

The loaded road network DataFrame looks like this:

+--------------------+----------+--------+----------+-----------+----------+-----------+
|            geometry|       src|     dst|   src_lat|    src_lon|   dst_lat|    dst_lon|
+--------------------+----------+--------+----------+-----------+----------+-----------+
|LINESTRING (-83.7...|  68133325|27254523| 42.238819|-83.7390142|42.2386159|-83.7390153|
|LINESTRING (-83.7...|9405840276|27254523|42.2386058|-83.7388915|42.2386159|-83.7390153|
|LINESTRING (-83.7...|  68133353|27254523|42.2385675|-83.7390856|42.2386159|-83.7390153|
|LINESTRING (-83.7...|2262917109|27254523|42.2384552|-83.7390313|42.2386159|-83.7390153|
|LINESTRING (-83.7...|9979197063|27489080|42.3200426|-83.7272283|42.3200887|-83.7273003|
+--------------------+----------+--------+----------+-----------+----------+-----------+
only showing top 5 rows

Users can also prepare the road network data from any data source using any data processing procedures, as long as the schema of the road network DataFrame conforms to the requirement of the Map Matching API.

3. Run Map Matching

Once the trajectories and road network data is ready, we can run matcher.match to match trajectories to the road network.

dfMmResult = matcher.match(dfEdge, dfPath, "geometry", "geometry")

The dfMmResult contains the trajectories snapped to the roads in matched_points attribute:

+---+--------------------+--------------------+--------------------+
|ids|     observed_points|      matched_points|       matched_nodes|
+---+--------------------+--------------------+--------------------+
|275|LINESTRING (-83.6...|LINESTRING (-83.6...|[62574078, 773611...|
|253|LINESTRING (-83.6...|LINESTRING (-83.6...|[5930199197, 6252...|
| 88|LINESTRING (-83.7...|LINESTRING (-83.7...|[4931645364, 6249...|
|561|LINESTRING (-83.6...|LINESTRING (-83.6...|[29314519, 773612...|
|154|LINESTRING (-83.7...|LINESTRING (-83.7...|[5284529433, 6252...|
+---+--------------------+--------------------+--------------------+
only showing top 5 rows

We can visualize the map matching result using SedonaKepler to see what the matched trajectories look like:

mapAll = SedonaKepler.create_map()
SedonaKepler.add_df(mapAll, dfEdge, name="Road Network")
SedonaKepler.add_df(mapAll, dfMmResult.selectExpr("observed_points AS geometry"), name="Observed Points")
SedonaKepler.add_df(mapAll, dfMmResult.selectExpr("matched_points AS geometry"), name="Matched Points")
mapAll

The following figure shows the map matching results. The red lines are original trajectories, and the green lines are matched trajectories. We can see that the noisy original trajectories are all snapped to the road network.

map matching results snaps gps points

Performance

We used WherobotsAI Map Matching to match 90 million trips across the entire US in just 1.5 hours on the Wherobots 2XL runtime, which equates to approximately 1 million trips per minute, at a cost of $873 (99% less expensive than Mapbox, while completing in a small fraction of the time). The average cost of matching 1 million trips is $9.70. Here’s how the cost of Map Matching 1 million trips in Wherobots compares to alternatives.

  • Mapbox: $1,600
  • HERE Maps: $4,400
  • Google Maps Platform: $8,000
  • WherobotsAI Map Matching: $9.70

The “optimization magic” behind WherobotsAI Map Matching lies in how Wherobots intelligently and automatically co-partitions trajectory and road network datasets based on the spatial proximity of their elements, ensuring a balanced distribution of work. As a result, the computational load is balanced evenly through this partitioning strategy, and makes map matching with Wherobots highly efficient, scalable, and affordable compared to alternatives.

Try It Out!

You can try out Map Matching by starting a notebook environment in Wherobots Cloud and running our example notebook within Wherobots Cloud.

notebook_example/python/wherobots-ai/mapmatching_example.ipynb

You can also check out the Map Matching tutorial and reference documentation for more information!

Try Map Matching

Key takeaways

  • Map matching reconstructs the most likely path on known road or trail segments from noisy GPS samples, instead of dropping outliers. WherobotsAI Map Matching uses a Hidden Markov Model approach in the lineage of Newson and Krumm, distributed across Wherobots rather than a single machine or per-trip REST API.
  • Public list prices cited as of July 2024 for matching 1 million trips: Mapbox $1,600, HERE Maps $4,400, Google Maps Platform $8,000. WherobotsAI Map Matching is about $10 per million trips.
  • A US-scale map matching job processed 90 million trips in 1.5 hours on the Wherobots 2XL runtime—about 1 million trips per minute—at a cost of $873 (average $9.70 per million trips), described as 99% less expensive than Mapbox.
  • The API takes a trajectories DataFrame and a road-segments DataFrame (OSM can be loaded with matcher.load_osm) and returns observed_points, matched_points, and matched_nodes. The walkthrough uses the VED dataset and Ann Arbor OSM.
  • Map matching prepares GPS data for satellite and GPS-based navigation, freight tracking, insurance risk, self-driving telematics, and transportation planning. Open-source options such as Valhalla and GraphHopper are single-machine; commercial REST APIs are costly at population scale.

Unlock Satellite Imagery Insights with WherobotsAI Raster Inference

Editor’s note: Raster inference is now part of RasterFlow, the Wherobots engine for running computer vision models on satellite and aerial imagery.

UPDATE: Raster inference is included in Wherobots RasterFlow. See Wherobots Get Started with RasterFlow – Wherobots for the most up-to-date workflows.

Recently we introduced WherobotsAI Raster Inference to unlock analytics on satellite and aerial imagery using SQL or Python. Raster Inference simplifies extracting insights from satellite and aerial imagery using SQL or Python, and is powered by open-source machine learning models. Below we’ll dig into the popular computer vision tasks that Raster Inference supports, describe how it works, and how you can use it to run batch inference to find and map electricity infrastructure.

Watch the live demo of these capabilities here.

The Power of Machine Learning with Satellite Imagery

Petabytes of satellite imagery are generated each day all over the world in a dizzying number of sensor types and image resolutions. The applications for satellite imagery and other remote sensing data sources are broad and diverse. For example, satellites with consistent, continuous orbits are ideal for monitoring forest carbon stocks to validate carbon credits or estimating agricultural yields.

However, this data has been inaccessible for most analysts and even seasoned ML practitioners because insight extraction required specialized skills. We’ve done the work to make insight extraction simple and accessible to more people. Raster Inference abstracts the complexity and scales to support planetary-scale imagery datasets, so you don’t need ML expertise to derive insights. In this blog, we explore the key features that make Raster Inference effective for land cover classification, solar farm mapping, and marine infrastructure detection. And, in the near future, you will be able to use Raster Inference with your own models!

Raster Inference supports the three most common kinds of computer vision models that are applied to imagery: classification, object detection, and semantic segmentation. Instance segmentation (combines object localization and semantic segmentation) is another common type of model which is not currently supported, but let us know if you need by contacting us and we can add it to the roadmap.

Computer Vision Detection Types
Computer Vision Detection Categories from Lin et al. Microsoft COCO: Common Objects in Context

The figure above illustrates these tasks. Image classification is when an image is assigned one or more text labels. In image (a), the scene is assigned the labels “person”, “sheep”, and “dog”. Image (b) is an example of object localization (or object detection). Object localization creates bounding boxes around objects of interest and assigns labels. In this image, five sheep are localized separately along with one human and one dog. Finally, semantic segmentation is when each pixel is given a category label, as shown in image (c). Here we can see all the pixels belonging to sheep are labeled blue, the dog is labeled red, and the person is labeled teal.

While these examples highlight detection tasks on regular imagery, these computer vision models can be applied to raster formatted imagery. Raster data formats are the most common data formats for satellite and aerial imagery. When objects of interest in raster imagery are localized, their bounding boxes can be georeferenced, which means that each pixel is localized to spatial coordinates, such as latitude and longitude. Therefore, georeferencing is object localization suited for spatial analytics.

https://wherobots.com/wp-content/uploads/2024/06/remotesensing-11-00339-g005.png?ver=1790894673

The example above shows various applications of object detection for localizing and classifying features in high resolution satellite and aerial imagery. This example comes from DOTA, a 15-class dataset of different objects in RGB and grayscale satellite imagery. Public datasets like DOTA are used to develop and benchmark machine learning models.

Not only are there many publicly available object detection models, but also there are many semantic segmentation models.

Semantic Segmentation
Sourced from “A Scale-Aware Masked Autoencoder for Multi-scale Geospatial Representation Learning”.

Not every machine learning model should be treated equally, and each will have their own tradeoffs. You can see the difference between the ground truth image (human annotated buildings representing the real world) and segmentation results across two models (Scale-MAE and Vanilla MAE). These results are derived from the same image at two different resolutions (referred to as GSD, or Ground Sampling Distance).

  • Scale-MAE is a model developed to handle detection tasks at various resolutions with different sensor inputs. It uses a similar MAE model architecture as the Vanilla MAE, but is trained specifically for detection tasks on overhead imagery that span different resolutions.
  • The Vanilla MAE is not trained to handle varying resolutions in overhead imagery. It’s performance suffers in the top row and especially the bottom row, where resolution is coarser, as seen by the mismatch between Vanilla MAE and the ground truth image where many pixels are incorrectly classified.

Satellite Analytics Before Raster Inference

Without Raster Inference, typically a team who is looking to extract insights from overhead imagery using ML would need to:

  1. Deploy a distributed runtime to scale out workloads such as data loading, preprocessing, and inference.
  2. Develop functionality to operate on raster metadata to easily filter it by location to run inference workloads on specific areas of interest.
  3. Optimize models to run performantly on GPUs, which can involve complex rewrites of the underlying model prediction logic.
  4. Create and manage data preprocessing pipelines to normalize, resize, and collate raster imagery into the correct data type and size required by the model.
  5. Develop the logic to run data loading, preprocessing, and model inference efficiently at scale.

Raster Inference and its SQL and Python APIs abstract this complexity so you and your team can easily perform inference on massive raster datasets.

Raster Inference APIs for SQL and Python

Raster Inference offers APIs in both SQL and Python to run inference tasks. These APIs are designed to be easy to use, even if you’re not a machine learning expert. RS_CLASSIFY can be used for scene classification, RS_BBOXES_DETECT for object detection, and RS_SEGMENT for semantic segmentation. Each function produces tabular results which can be georeferenced either for the scene, object, or segmentation depending on the function. The records can be joined or visualized with other data (geospatial or traditional) to curate enriched datasets and insights. Here are SQL and Python examples for RS_Segment.

RS_SEGMENT('{model_id}', outdb_raster) AS segment_result
df = df_raster_input.withColumn("segment_result", rs_segment(model_id, col("outdb_raster")))

Example: Mapping Electricity Infrastructure

Imagine you want to optimize the location of new EV charging stations, but you want to target locations based on the availability of green energy sources, such as local solar farms. You can use Raster Inference to detect and locate solar farms and cross-reference these locations with internal data or other vector geometries that captures demand for EV charging. This use case will be demonstrated in our upcoming release webinar on July 10th.

Let’s walk through how to use Raster Inference for this use case.

First, we run predictions on rasters to find solar farms. The following code block that calls RS_SEGMENT shows how easy this is.

CREATE OR REPLACE TEMP VIEW segment_fields AS (
    SELECT
        outdb_raster,
        RS_SEGMENT('{model_id}', outdb_raster) AS segment_result
    FROM
    az_high_demand_with_scene
)

The confidence_array column produced from RS_SEGMENT can be assigned the same geospatial coordinates as the raster input and converted to a vector that can be spatially joined and processed with WherobotsDB using RS_SEGMENT_TO_GEOMS. We select a confidence threshold of .65 so that we only georeference high confidence detections.

WITH t AS (
        SELECT RS_SEGMENT_TO_GEOMS(outdb_raster, confidence_array, array(1), class_map, 0.65) result
        FROM predictions_df
    )
    SELECT result.* FROM t
+----------+--------------------+--------------------+
|     class|avg_confidence_score|            geometry|
+----------+--------------------+--------------------+
|Solar Farm|  0.7205783606825462|MULTIPOLYGON (((-...|
|Solar Farm|  0.7273308333550763|MULTIPOLYGON (((-...|
|Solar Farm|  0.7301468510823231|MULTIPOLYGON (((-...|
|Solar Farm|  0.7180177244988899|MULTIPOLYGON (((-...|
|Solar Farm|   0.728077805771141|MULTIPOLYGON (((-...|
|Solar Farm|     0.7264981572898|MULTIPOLYGON (((-...|
|Solar Farm|  0.7044100126912517|MULTIPOLYGON (((-...|
|Solar Farm|  0.7137283466756343|MULTIPOLYGON (((-...|
+----------+--------------------+--------------------+

This allows us to integrate the vectorized model predictions with other spatial datasets and easily visualize the results with SedonaKepler.

https://wherobots.com/wp-content/uploads/2024/06/solar_farm_detection-1-1024x398.png?ver=1790894673

Here Raster Inference runs on a 85 GiB dataset with 2,200 raster scenes for Arizona. Using a Sedona (tiny) runtime, Raster Inference completed in 430 seconds, predicting solar farms for all low cloud cover satellite images for the state of Arizona for the month of October. If we scale up our runtime to a San Francisco (small) runtime, the inference speed nearly doubles. In general, average bytes processed per second by Wherobots increases as datasets scale in size because startup costs are amortized over time. Processing speed also increases as runtimes scale in size.

Inference time (seconds)Runtime Size
430Sedona
246San Francisco

We use predictions from the output of Raster Inference to derive insights about which zip codes have the most solar farms, as shown below. This statement joins predicted solar farms with zip codes by location, then ranks zip codes by the pre-computed solar farm area within each zip code. We skipped this step for brevity but you can see it and others in the notebook example.

az_solar_zip_codes = sedona.sql("""
SELECT solar_area, any_value(az_zta5.geometry) AS geometry, ZCTA5CE10
FROM predictions_polys JOIN az_zta5
WHERE ST_Intersects(az_zta5.geometry, predictions_polys.geometry)
GROUP BY ZCTA5CE10
ORDER BY solar_area DESC
""")
https://wherobots.com/wp-content/uploads/2024/06/final_analysis.png?ver=1790894673

These predictions are made possible by SATLAS, a family of machine learning models released with Apache 2.0 licensing from Allen AI. The solar model demonstrated above was derived from the SATLAS foundational model. This foundational model can be used as a building block to create models to address specific detection challenges like solar farm detection. Additionally, there are many other open source machine learning models available for deriving insights from satellite imagery, many of which are provided by the TorchGeo project. We are just beginning to explore what these models can achieve for planetary-scale monitoring.

If you have a specific model you would like to see made available, please contact us to let us know.

For detailed instructions on using Raster Inference, please refer to our example Jupyter notebooks in the documentation.

https://wherobots.com/wp-content/uploads/2024/06/Screenshot_2024-06-08_at_2.11.07_PM-1024x683.png?ver=1790894673

Here are some links to get you started:
https://docs.wherobots.com/latest/tutorials/wherobotsai/wherobots-inference/segmentation/

https://docs.wherobots.com/latest/api/wherobots-inference/pythondoc/inference/sql_functions

Getting Started

Getting started with WherobotsAI Raster Inference is easy. We’ve provided three models in Wherobots Cloud that can be used with our GPU optimized runtimes. Sign up for Wherobots Pro via the AWS Marketplace and get up to $400 in free credits. You can test out the Pro features, including Raster Inference.

Stay tuned for updates on improvements to Raster Inference that will make it possible to run more models, including your own custom models. We’re excited to hear what models you’d like us to support, or the integrations you need to make running your own models even easier with Raster Inference. We can’t wait for your feedback and to see what you’ll create!

Start building with Wherobots

Key takeaways

  • WherobotsAI Raster Inference (later folded into RasterFlow) runs open-source computer vision models on satellite and aerial rasters from SQL or Python, without requiring users to stand up GPU pipelines.
  • Supported tasks are classification (RS_CLASSIFY), object detection (RS_BBOXES_DETECT), and semantic segmentation (RS_SEGMENT). Instance segmentation is not supported in this post.
  • An Arizona solar-farm example processed an 85 GiB dataset with 2,200 scenes: 430 seconds on a Sedona (tiny) runtime and 246 seconds on a San Francisco (small) runtime, then vectorized detections at 0.65 confidence with RS_SEGMENT_TO_GEOMS.
  • The solar model comes from SATLAS (Allen AI, Apache 2.0). The post also points at TorchGeo models and says three models are hosted in Wherobots Cloud on GPU-optimized runtimes.
  • Professional Edition via AWS Marketplace includes up to $400 in free credits. The authors say custom/BYO models are on the roadmap after this preview.