Rasterflow, Earth Intelligence & inference engine now in public preview Learn More

Wherobots for QGIS: Cloud-Scale Spatial Data, from Wherobots Labs

Wherobots for QGIS is a new plugin that connects QGIS to Wherobots Cloud. From a panel inside QGIS, an analyst can run spatial SQL against WherobotsDB, load the results as a map layer, push a local layer up to a Wherobots Iceberg table, and pull raster data into the canvas. 

Why Wherobots for QGIS: Eliminating the partitioning tax

GIS analysts hit the same wall in city after city. A workflow that runs fine on a neighborhood runs fine on a district. At state scale, it crashes or runs for hours. At country scale, the analyst has to write a loop, tile the data by county, process each tile, handle the features that fall on the boundary, reassemble the output, and discover the seam artifacts two days later. That loop is not the analysis. It is overhead that the existing tools and databases impose.

This is the partitioning tax. It is not a performance problem. It is an architectural one. Desktop GIS and single-node databases process data on one machine. When the data outgrows that machine, the analyst absorbs the cost: split the extent, manage the tiles, merge the results, debug the edges. Every time a parameter changes, run it again. Testing at city scale and assuming it generalizes is not a workflow, it is a bet.

The same constraint blocks AI coding assistants. Tools like Claude Code and Cursor are fluent at writing spatial SQL for a single-pass query. Ask one to scale a workflow from a county to a state and it generates a partition loop with no overlap buffer and no dedupe logic. The model does not know about edge effects. The developer still has to do the work.

Wherobots for QGIS is a new plugin that removes the wall on the QGIS side of the stack. An analyst writes spatial SQL in a panel inside QGIS. WherobotsDB executes it in the cloud across the full extent, whether that extent is a city block or a continent, with no partition logic in the query and no tile management by the analyst. The result loads as a map layer.

What the Wherobots QGIS plugin does

The panel has three tabs, plus a connection screen where you enter an API key and pick a region and runtime.

SQL Query

Write spatial SQL in a free-form editor, or switch to Browse Tables and pick a table from the catalog. Set a row limit. Check one box to restrict results to the current map extent. Results land in the QGIS project as a layer, ready for styling, joins, or export.

Upload

Select a vector layer in your project, give it a destination table name, and send it to Wherobots. The plugin creates the table and inserts the features in batches while QGIS stays responsive.

Raster

Run raster SQL with RS_ functions against raster tables. Use the current map extent as the bounding box. A query that returns RS_AsGeoTiff loads as a raster layer. A query that returns RS_Values gives you pixel values as a table.

The plugin runs on QGIS 3.22 through QGIS 4.0, covering both Qt5 and Qt6 builds.

One copy of the data across QGIS, notebooks, and AI coding assistants

Wherobots has one job in this picture: hold the data and run the compute. The interface should follow the person doing the work.

Today the same WherobotsDB tables are reachable from a Python notebook, a SQL client, a Python Job, the Wherobots CLI, and from AI coding assistants like VS Code, Cursor, and Claude Code through the Wherobots MCP Server. QGIS joins that list. The tables an analyst opens on the canvas are the same tables an agent queries through MCP. One copy of the data. No exports.

This is what making spatial accessible means in practice. An analyst who knows QGIS and knows SQL can now run a spatial join across hundreds of millions of records and see the result on a map, without standing up a cluster or learning a new interface. WherobotsDB, built by the original creators of Apache Sedona, handles the distribution, indexing, and coordinate systems underneath.

How to install the Wherobots plugin in QGIS

  1. Download the plugin zip from the Releases page and install it in QGIS under Plugins → Manage and Install Plugins → Install from ZIP.
  2. Install the wherobots-python-dbapi package into QGIS’s bundled Python. The README has a two-line snippet for the QGIS Python Console that works on every platform.
  3. Open the Wherobots panel from the toolbar or Web → Wherobots, paste your API key, and connect.

You need a Wherobots Cloud account. Start a free trial if you do not have one. The wherobots_open_data catalog, including Overture Maps, is a good first query.

Introducing Wherobots Labs

Wherobots for QGIS is the first project published under Wherobots Labs.

Labs is where Wherobots ships the projects that live at the edge of the platform: connectors, plugins, adapters, example pipelines, and skills for AI coding assistants. They are built by the solutions architects and customer engineers who work with Wherobots customers every day, and by community contributors.

Every Labs repository follows the same rules. It is public on GitHub under the wherobots organization with a labs- prefix. It is Apache 2.0 licensed unless noted otherwise. It has a README with a plain statement of support expectations, a CONTRIBUTING.md, and GitHub Issues enabled. It gets a security review before first release, and security fixes continue under the standard Wherobots security policy.

What a Labs project does not carry is a production SLA. You are responsible for confirming a project fits your use case before you rely on it, and you are free to fork it. A Labs project that earns adoption can graduate to a supported product surface.

More Labs projects are on the way. Browse them at here and open an issue or a pull request on any of them.

If you use QGIS and you have a dataset that stopped fitting on your machine, install the plugin and tell us what you build.

Key takeaways

  • Wherobots for QGIS connects QGIS to Wherobots Cloud, so an analyst runs spatial SQL against WherobotsDB without leaving the map.
  • The compute runs in the cloud. QGIS stays the place to view the answer. One copy of the data, no exports.
  • Three tabs: SQL Query, Upload, and Raster. The plugin runs on QGIS 3.22 through QGIS 4.0, on both Qt5 and Qt6.
  • The same WherobotsDB tables reach Python notebooks, SQL clients, the Wherobots CLI, and AI coding tools through MCP. QGIS joins that list.
  • Available today from the QGIS Plugin Repository. It needs a Wherobots Cloud account, with a free trial. It is the first release from Wherobots Labs.
Get Started with Wherobots

How Wherobots builds with NVIDIA to let AI see the physical world

AI was blind to the physical world, but Wherobots is giving AI the ability to understand it using data, compute, and NVIDIA GPUs

The world and what happens in it is digitized by petabytes of raw and derivative spatial datasets of various data types and scales, and the potential for applying AI to it is immense. But the AI models and agents we use every day need connectivity to tools that turn this data into usable insights and relationships.

Wherobots gives AI the ability operate on and understand raw physical-world data, and NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs are core to it. Architecturally here’s how this works at a high level. Wherobots:

  1. Integrates with your open data architecture: by utilizing data lakes like Amazon S3 and Iceberg catalogs such as Databricks Unity Catalog and AWS Glue Data Catalog as the center of truth, you can integrate Wherobots into your existing data architecture.
  2. Brings the spatial capability: data teams and their agents have the capability to solve the most challenging physical world problems by directing Wherobots to extract insights from multi-dimensional, raster, vector, and tabular datasets of various formats and scales.
  3. Is plug-and-play: in minutes, you can enable AI coding tools like VS Code, Claude Code, and Cursor to direct Wherobots operations on this data and build innovations with AI.

Under the hood, we deploy RTX PRO 6000 Blackwell GPUs and libraries on tasks they are optimized for.

  • RTX PRO 6000 Blackwell GPUs run at the core of RasterFlow, which is the preprocessing and inference solution designed to make insight extraction from satellite and aerial sensor datasets easy. We also leverage NVIDIA CUDA™ streams to optimize performance. RasterFlow crunches pixels with a distributed fleet of NVIDIA G7e Amazon Web Services (AWS) instances, accelerated by RTX PRO 6000 Blackwell GPUs. These GPUs, built on the groundbreaking Blackwell architecture, feature the second-generation Transformer Engine to deliver unprecedented performance, efficiency, and scale for generative AI and accelerated computing.
  • SedonaDB, which WherobotsDB utilizes in a distributed mode, can run spatial joins directly on the ray tracing cores inside RTX PRO 6000 Blackwell GPUs to deliver the best price-performance for a class of spatial join types.

Let’s dig in.

RasterFlow on NVIDIA

RasterFlow prepares large scale imagery data for inference and runs computer vision models on a distributed inference architectureIts goal is to give customers the ability to generate insights from this data across physical areas of interest of any scale, enabling real-world applications such as:

  • Last-Mile Delivery Routing: Identifies sidewalks (using models like Tile2Net) which can help determine viable navigational routes for both manual and autonomous delivery systems.
  • Property Risk Analytics: Can aid in quantifying the number of properties facing flood risk across coastal cities by embedding FEMA flood zone designations into parcel-level data, helping insurers, reinsurers, and real estate developers evaluate policy and property portfolios.

Since its introduction in private preview, we continue to optimize the RasterFlow architecture, reaching a point where mosaicking and inference tasks cost 2–3x less than the next best option at scale. For instance, RasterFlow supports Meta’s Segment Anything Model (SAM3) for general-purpose object detection on Earth Observation imagery. With the efficiencies we’ve achieved, you can use SAM3 to detect objects from high-resolution 30cm imagery for less than $25 per 1,000 square kilometers.

Example of roof detections on 30cm imagery using Meta’s Segment Anything Model (SAM3)
Example of roof detections on 30cm imagery using Meta’s Segment Anything Model (SAM3)

RasterFlow uses CUDA™ streams to keep GPUs busy

Data movement can kill inference performance, so we use NVIDIA CUDA™ streams and RTX PRO 6000 Blackwell GPUs to address this bottleneck. Every batch of satellite imagery patches travels from host memory to the GPU before the model can run, and every batch of prediction patches travels back before RasterFlow can merge it into a seamless output mosaic. If we ran these steps serially one after another, our GPUs would idle on PCIe transfers with every forward pass.

CUDA™ streams remove that idle time. A stream is an ordered queue of GPU work, where different kinds of work in different streams can execute concurrently. . RTX PRO 6000 Blackwell GPUs include dedicated copy engines that move data across PCIe independently of the compute cores, which means a transfer in one stream proceeds while a model forward pass runs in another. RasterFlow drives this through PyTorch’s torch.cuda.Stream API.

This inference loop has separate upload and download streams feeding two reusable buffer slots. While the model runs batch N on the first stream, a second stream copies the next batch to the GPU, and a download stream drains the predictions from batch N-1 back to host memory, where the CPU merges patches in the mosaic accumulator. At steady state three batches are in flight at once: one uploading, one computing, one landing in the output. RasterFlow allocates these buffers once and reuses them for every batch. CUDA™ events sequence the handoff between streams, so data movement never starts before the forward pass that produced it finishes. This allows the GPU to stay busy from the first patch to the last.

What’s next: GPU-native decompression with NVIDIA nvCOMP™ and Zarr

RasterFlow stores every analysis-ready mosaic as a Zarr store of compressed chunks on Amazon S3, and today those chunks decode on the CPU.


nvCOMP™, NVIDIA’s library of GPU-accelerated compression codecs, can move that decode step onto the GPU itself, and we are looking forward to nvCOMP™-backed codec support landing in Zarr. The payoff of this feature is three-fold. First, compressed chunks cross PCIe instead of raw pixels, so each transfer carries fewer bytes. Second, decompression runs at GPU memory bandwidth (which is higher than CPU) and the decoded array materializes directly in GPU memory where the model needs it. Finally, CPU cores are free for the work only they can do: assembling patches and merging predictions into the mosaic.

RasterFlow use cases

RasterFlow continues to unblock customers who need to extract insights from global-scale imagery datasets. We are actively working with customers across multiple industries, and utilizing RasterFlow to deliver results like these:

  1. The Fields of the World global run produced globally consistent agricultural field-boundary data for land-use monitoring, food-system analysis, and model development. We executed this run using RasterFlow, and led it in partnership with Taylor Geospatial, Microsoft AI for Good Lab, and Nasa Harvest. With RasterFlow you can run this same model with 10m Sentinel-2 imagery, including cloud-free mosaic generation and model inference, for less than $1 per 10,000 square kilometers.
Agricultural field boundaries predicted by Taylor Geospatial’s Fields of the World model running on RasterFlow.
Agricultural field boundaries predicted by Taylor Geospatial’s Fields of the World model running on RasterFlow.
  1. We are actively supporting the US Forest Service in their research, and our joint goal is to bring their new FireCon model online in the second half of this calendar year. This initiative is key to delivering far more accurate daily wildfire containment suitability mapping system for active wildfires of the Western US, enabling better planning, management, and response to wildfires.
FireCon Model RasterFlow
FireCon model predictions for the best fire control points generated across the western United States by RasterFlow.

SedonaDB on NVIDIA

WherobotsDB runs a distributed version of SedonaDB, which customers use to efficiently process and join two dimensional raster and vector data. In SedonaDB 0.4, we shipped a GPU-accelerated spatial join as an extension that runs on RTX PRO 6000 Blackwell GPU ray tracing cores. In practice, almost all spatial analyses are powered by one or more spatial joins, which are computationally intensive and frequently a bottleneck in spatial queries. But SedonaDB with GPU acceleration provides relief, and delivers up to a 5.93x speedup and a 59.02% cost reduction on our standard spatial benchmark. The research behind it, RayBooster: A Ray Tracing Engine to Accelerate SedonaDB, was accepted to VLDB 2026 in the Industry Track, and we developed it with The Ohio State University.

Attempts to accelerate traditional analytical database workflows using GPUs have typically required utmost care and/or rethinking of a workflow to minimize data transfer between the existing workflow (on the CPU) and the accelerated implementation (on the GPU). The spatial join is an ideal candidate for this type of acceleration because it is computationally intensive, and relatively small amounts of data transfer can result in large numbers of computations. Furthermore, the required data transfer is only a single direction; geometries must be transferred to the GPU but need not be transferred back.

Another limitation to GPU acceleration of traditional database workflows has been a requirement for a potentially large amount of GPU memory to avoid completely rewriting a workflow with careful attention to the sizes of the inputs. SedonaDB’s join implementation was built to partition large joins into smaller ones that can be run with limited memory with its CPU-based join, and while this heuristic needs to be tuned differently for GPUs, we did not need to rewrite our GPU implementation to support arbitrarily large joins. This extends the acceleration we observed to a wider range of GPUs that customers may have available.

Deploying AI on physical world data with NVIDIA

High complexity, scale, and costs have limited the utilization of spatial data, resulting in a significant gap between AI and physical world data. This gap is closing fast and as it closes, humans are more capable of directing AI towards solutions that drive top, bottom, and the sustainability lines of their business.

Through our investments in Apache Sedona and Wherobots, and by building on NVIDIA technology, data teams and their agents are becoming increasingly capable of innovating and operating solutions with physical world data, under budget, with the talent and AI tools they already have, regardless of spatial data type and degree of complexity.

Key takeaways

  • Wherobots gives AI the ability to operate on and understand raw physical-world data, and NVIDIA GPUs are core to it. Wherobots deploys NVIDIA GPUs and libraries on tasks they are optimized for.
  • RasterFlow crunches pixels with a distributed fleet of NVIDIA G7e AWS instances, accelerated by RTX PRO 6000 Blackwell GPUs. CUDA streams keep three batches in flight at once, driving mosaicking and inference costs to typically 2-3x less than the next best option at scale. SAM3 on 30cm imagery runs for under $25 per 1,000 square kilometers.
  • SedonaDB runs spatial joins directly on the ray tracing cores inside NVIDIA GPUs. In SedonaDB 0.4, Wherobots shipped a GPU-accelerated spatial join that delivers up to a 5.9x speedup and a 59% cost reduction on the standard spatial benchmark. The research paper, RayBooster: A Ray Tracing Engine to Accelerate SedonaDB, was accepted to VLDB 2026 in the Industry Track.
  • RasterFlow is validated in production with Taylor Geospatial, Microsoft AI for Good Lab, and NASA Harvest (Fields of the World), and is actively supporting the US Forest Service (FireCon) to deliver daily wildfire containment suitability mapping across the Western US in H2 2026.
Get Started with Wherobots

Orchestrating Wherobots Jobs from AWS Step Functions: A Reference Architecture

When the execution engine is loosely coupled with orchestration, the pipeline is blind while waiting for one fact: did the job finish, fail, or die? This was the case with jobs running inside Wherobots Cloud and the AWS pipeline launching it… That was not cool with me, so we designed a reference implementation using AWS Step Functions, their callback pattern, and the Wherobots Python SDK.

The reference implementation is intentionally small and the accompanying (vibe coded) interactive app is designed to walk you through the implementation for learning purposes. There are four Lambda functions, two state machines, one API Gateway endpoint, and a simple Wherobots job script. Every AWS resource is created by plain CLI calls you can read, and a guided app drives the whole lifecycle: upload, deploy, run, teardown. The demo job counts Overture Maps buildings within one kilometer of downtown Seattle (geodesic buffer, EPSG:4326) and returns 1,084.

What the Reference Architecture Does

The pipeline submits a Wherobots job to our distributed engine through the wherobots-python-sdk, pauses for the callback, and resumes the moment the job reports its own result. The submit handler Lambda calls WherobotsJob(...).submit() with the task token passed as plain job arguments. The state machine then waits. There is no polling loop on this happy path. When the job finishes, the state machine wakes within seconds and the next state receives the job’s output as its input.

Authentication stays minimal on both sides. The Lambdas hold one Wherobots API key. The SDK uploads the job script to managed storage over short-lived STS credentials, so no long-lived AWS keys appear anywhere. Wherobots runs the job on a managed Apache Sedona runtime, built by the original creators of Apache Sedona and 100% code compatible across all spatial functions.

How the Callback Pattern Works

Step Functions supports a task type that pauses until something calls SendTaskSuccess or SendTaskFailure with a one-time task token. The catch: the job runs in Wherobots’ AWS account and cannot sign calls to your Step Functions API. The reference bridges that gap with a relay. The job POSTs plain JSON over HTTPS to an API Gateway endpoint, and a small Lambda behind it relays the payload into the Step Functions API.

The single-use task token is the security capability. A request without a live token can do nothing, which is what makes a public endpoint acceptable for a reference implementation. Production hardening adds an API key or IAM auth at the gateway. The reference uses API Gateway instead of a Lambda Function URL as blocking anonymous Function URL invocation by service control policy is a general best practice, and API Gateway is the sanctioned front door.

Failure Mode Drives the Design

In a decoupled orchestration and compute engine system, one major question to consider is “What happens when the job fails?”. How the job fails drives the resolution and code path, one layer cannot cover all cases.

Soft failures report themselves

The Wherobots job logic is wrapped in a typical try/except loop so the Python-level failures, from bad SQL to bad data, are caught and posted as a failure callback carrying the full traceback. The execution fails and within seconds the traceback appears in the AWS console with the trace. In live verification, a forced RuntimeError is surfaced in the execution history with the exact line that raised it.

A Heartbeat catches the hard death

A driver OOM or a cluster kill ends the Python process mid-instruction. No except or finally runs and no callback is ever sent. There are no silver bullets to report a death like this, which is why the job also runs a daemon thread that posts a heartbeat every 60 seconds. The task sets HeartbeatSeconds: 240; the window counts from task entry, and the job’s first heartbeat cannot arrive until the runtime finishes provisioning, so the window sits above typical provisioning time. On a rare slower cold start the timeout routes to the fallback poller, which finds the run healthy and still delivers COMPLETED. When the process dies, heartbeats stop with it, and the timeout fires within four minutes initiating the failure tracing process.

The fallback poller finds the truth

The heartbeat timeout routes to a recovery path. A FindRun Lambda recovers the run ID from the deterministic job name using WherobotsJob.list_runs, then a reusable poller state machine checks the run every 30 seconds until Wherobots returns a terminal status. The poller takes {run_id, poll_seconds} as input and carries no reference to the pipeline that submitted the job, so any other pipeline can reuse it as-is. In live verification, the simulated OOM (os._exit(137), no callback, no goodbye) reached a confirmed FAILED verdict in 4 minutes and 46 seconds, end to end: the heartbeat window plus one poll cycle (adjust HeartbeatSeconds and the polling interval to your desired detection latency).

The Guided Dashboard

The repository ships a local app to walk you through the architecture. Five stages build on each other in order:

  1. PREFLIGHT: reads a masked .env
  2. UPLOAD: pushes the job script to Wherobots storage
  3. BUILD streams every AWS CLI call as it creates the infrastructure
  4. RUN: kicks off the step function chain
  5. TEARDOWN: removes resources created in step 3

The test-mode dropdown makes the failure contract demonstrable through the live execution graph:

How to Run It

Minimal setup required to play with the dashboard and implementation. Currently the repo is private, but if you wan to run it contact us and we can give you access. We expect this to be live as an open repo very soon in Wherobots Labs.

git clone https://github.com/wherobots/labs-reference-implementations.git
cd labs-reference-implementations/aws-step-functions-python-sdk
python3 -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
cp .env.example .env   # set WHEROBOTS_API_KEY; blank AWS keys use your default chain
python3 dashboard/server.py

Open the dashboard, click through the five stages, and run the three test modes. The demo costs a few cents of Lambda and Step Functions time plus one tiny Wherobots runtime per run. Don’t forget to run (and confirm) the Teardown cell, it removes every AWS resource the build created, (each resource carries ManagedBy and TeardownBy tags for auditing) .

Design Decisions Worth Stealing

Four steps in this reference transfer to production systems directly.

Deterministic job names. The execution input carries the job name, so the fallback path can recover the run ID after the callback path dies. Recovery depends on nothing the failed job was supposed to send.

A reusable poller. Polling logic lives in its own state machine with a two-field input contract. The callback version of the pipeline demoted it to a fallback without changing a line of it.

Failure evidence over failure flags. The soft-failure callback carries the traceback, and the fallback poller reports the run’s terminal status from the Wherobots API. Whoever reads the failed execution reads the cause.

Symmetric teardown. Every create call in the deploy script has a delete call in the teardown script, and both are readable CLI commands. Reference infrastructure that cannot be removed cleanly teaches the wrong habit.

Where It Fits

This reference covers job runs: script-shaped and batch-shaped work submitted to Wherobots compute. Interactive SQL over the Spatial SQL API and Airflow-based orchestration with the Wherobots provider are separate surfaces with their own patterns. If your pipelines live in Step Functions today and your spatial workloads need managed jobs at scale, this architecture is the shortest path between them.

Key takeaways

  • The pipeline submits a Wherobots job through the wherobots-python-sdk, pauses for the callback, and resumes the moment the job reports its result. No polling loop on the happy path.
  • The single-use task token is the security capability. A request without a live token does nothing, which makes a public API Gateway endpoint acceptable for a reference implementation.
  • Two failure modes, two mechanisms: a try/except block posts the full traceback for soft failures; a 60-second heartbeat plus a fallback poller catch a hard death like a driver OOM or cluster kill.
  • Live verification: a forced RuntimeError surfaces the exact line that raised it within seconds; a simulated OOM (os._exit(137), no callback) reaches a confirmed FAILED verdict in 4 minutes and 46 seconds.
  • Four design decisions transfer to production directly: deterministic job names, a reusable two-field poller, failure evidence over failure flags, and symmetric teardown.
  • The demo job counts Overture Maps buildings within one kilometer of downtown Seattle (geodesic buffer, EPSG:4326) and returns 1,084.
Get Started with Wherobots

RasterFlow is now available in Public Preview

RasterFlow makes planetary-scale earth intelligence workflows easy and costs predictable.

We are excited to announce that RasterFlow is now in Public Preview, opening up the power of planetary scale Earth Intelligence to all Wherobots Professional Edition customers!

RasterFlow let’s you solve complex monitoring challenges with vision-language models or tailored models for specific use cases, without needing to manage complex raster preparation and inference infrastructure.

Teams are already running RasterFlow at planetary scale:

  • The USDA Forest Service uses RasterFlow to deliver wildfire containment predictions to its Wildfire Risk Management Division at an operational cadence.
  • Miraterra uses RasterFlow to predict agricultural field boundaries across the U.S. Midwest and Canada’s Prairie Provinces, then joins detailed microbial samples and geospatial embeddings to those boundaries.
  • Taylor Geospatial worked with Wherobots to produce 8.2 billion global field boundaries for the Fields of the World program.

Join our Public Preview virtual event, Pixels to Predictions: Planetary-Scale Earth Observation with RasterFlow, for a live walkthrough of built-in models, predictable pricing, and real customer pipelines. Register Here →

Join Public Preview Event: Pixels to Predictions with RasterFlow.

RasterFlow is a serverless image preparation and computer vision engine that makes it easy to extract insights from large scale raster datasets. It builds mosaics from multiple raster data sources, runs inference with computer vision models, and vectorizes the results, through a high-level API that simplifies the complexity of raster pipelines and distributed computing. You pay RasterFlow usage using a predictable, low cost pricing model that scales with your area of interest.

With RasterFlow’s built-in models, users can instantly launch tasks for common Earth Intelligence use cases. For instance, RasterFlow includes Meta’s SAM3 model for text-prompted object detection, Taylor Geospatial’s Fields of the World model for agricultural field boundaries, Meta’s CHM v1 for estimating tree canopy height. You can also bring your own, and we are expanding the set of open models we offer out-of-the box (let us know which open models you want offered and we can add them).

Market leading pricing that you can predict

RasterFlow’s predictable pricing allows you to estimate the cost of your tasks before they run. The price for each task is based on the data volume processed, so you can accurately estimate the cost of any task before execution. And pricing scales linearly with the size of your area of interest, so you can extrapolate the costs from smaller test runs to planetary scale.

Data volume = Area in km² × Pixels per km² × Bands × Time periods

The four inputs are:

  • Area: Size of your AOI in km²
  • Pixels per km²: Determined by input resolution; finer imagery means more pixels
  • Bands: For example, 4 for RGB + NIR
  • Time periods: For example, 3 annual observations

Let’s walk through a concrete example. If you wanted to find every solar panel array across a 500 km² county, you could use the built-in SAM3 model and a simple text prompt (“solar panel”) to generate detections. Here’s the code to launch this task across your area of interest using RasterFlow’s built-in support for 30cm imagery from USDA’s National Agriculture Imagery Program (NAIP):

Pythonfrom rasterflow_remote import RasterflowClient
from rasterflow_remote.data_models import GeometryModelRecipes

rf = RasterflowClient()

detections = rf.predict_mosaic_geometries_recipe(
    aoi="s3://your-bucket/county.parquet",
    start=datetime(2022, 1, 1),
    end=datetime(2023, 1, 1),
    model_recipe=GeometryModelRecipes.SAM3_TEXT_GEOMETRY,
    text_prompt="solar panel",
    confidence_threshold=0.5,
)

print(detections.uri)  # GeoParquet containing detected solar-panel polygons

To calculate the price for this task, we first determine the number of input pixel values based on the size of the AOI (500 km²), the resolution of the dataset (30cm NAIP), the number of bands (4), and the number of time periods (1).

Price = Data volume × Task-specific rate

Then, for mosaic generation and inferencing tasks, we apply a “complexity factor” to account for differences in task processing. For instance, building a mosaic from NAIP imagery is simpler than creating a cloud-free composite from Sentinel-2 imagery, so we apply a 0.1× multiplier.

Similarly, different models have different complexity factors, so we apply the relevant inference complexity factor (in this case, 1.0× for SAM3).

Finally, each RasterFlow task has a specific price that may vary by compute region. Combining these factors, we arrive at our total costs for these tasks:

TaskComplexity factorRasterFlow Spatial Units (SU)Price per SUCost
Mosaic generation (NAIP)0.1×2.22$0.75$1.67
Inference (SAM3)1.0×22.22$1.50$33.33
Total$35.00

For more details on RasterFlow pricing, see our pricing page and documentation.

To see some example solar panel detections for Marion County, Oregon, here is an interactive visualization:

Check out our viewer to explore these results further. Or, for more details about our Text to Detections support with SAM3, see this blog post: Detecting Objects From Text Prompts with RasterFlow and SAM3.

Your mosaics, predictions, and vectors in your storage

Wherobots storage integrations connect RasterFlow directly to your own S3 buckets. Tasks can read your areas of interest and proprietary imagery from your buckets, then write mosaics, predictions, and vectorized results back to them. Wherobots securely manages the required roles and credentials, so you can focus on your workflows instead of wrangling permissions.

Pythonfrom rasterflow_remote import RasterflowClient, DatasetEnum

client = RasterflowClient()

result = client.build_mosaics(
    datasets=[DatasetEnum.S2_MED_HARVEST],
    aoi="s3://my-company-data/aois/project.parquet", # one or multiple geometries
    start=datetime(2024, 1, 1),
    end=datetime(2025, 1, 1),
    bucket="s3://my-company-data/rasterflow/results",
)

print(result.first_row_mosaic)
"s3://my-company-data/rasterflow/results/mosaics/<run-id>/mosaic_index.parquet"

Outputs remain in open, interoperable formats: Zarr for mosaics and predictions, GeoParquet for vectorized results, and Iceberg tables through managed catalogs. Your data and results remain in the storage your applications already use, without a separate migration or export workflow.

Built-in visualization to inspect your Earth observation insights

Visual inspection is essential for validating inference results at scale, but large raster outputs are difficult to explore in their raw form. Wherobots lets users instantly layer mosaics, model predictions, and vectorized results on an interactive map, making it easy to assess quality, tune thresholds, and spot misaligned or spurious detections. Every completed RasterFlow task includes a one-click link to view its results in the Workload History.

RasterFlow also includes tasks that optimize existing Zarr stores for interactive viewing by adding image pyramids, downsampled overviews, and histogram statistics. The built-in map client then streams coarse tiles when zoomed out and full-resolution pixels when zoomed in, delivering responsive exploration at any scale.

For example, see this interactive visualization of model outputs from SAM3:

How RasterFlow compares to Google Earth Engine

Google Earth Engine provides a deep planetary imagery catalog and a strong environment for exploratory analysis. But for planetary scale workflows, Google Earth Engine has some significant limitations:

Limited cost predictability. Earth Engine bills in EECU-hours, so you only learn the total cost after the job finishes. With RasterFlow, you can calculate the cost of your tasks before you run them, eliminating uncertainty and potential billing surprises.

Build your own inference pipelines. Building a planetary scale earth observation pipelines with computer vision models requires integration with Vertex AI and custom pipeline development. Workflows that integrate with imagery data sources, patch tiles, handle seams, and maintains georeferences are costly to develop and maintain. RasterFlow packages mosaicking, inference, and vectorization as built-in tasks with a simple API.

Results are siloed in Earth Engine. Analysis results are stored in Earth Engine, which is separate from Google Cloud Platform or Google Cloud Storage. Using them elsewhere requires queuing export jobs, so a team building on AWS pays both export and egress costs.

Mosaicking cost. Because of these export costs, mosaicking costs on Earth Engine can be significantly higher for developers in AWS. For instance, generating and exporting a Sentinel-2 mosaic with all 12 bands for 150,000 km² costs about $50 in Earth Engine vs. $13.50 for RasterFlow.

Customer impact at planetary scale

Fields of the World: a global field-boundary layer

Fields of the World is a Taylor Geospatial effort to produce globally consistent agricultural field-boundary data for land-use monitoring, food-system analysis, and model development. Taylor Geospatial partnered with Wherobots to run their PRUE model on RasterFlow for the 2024 to 2025 global release. Read about the Fields Of The World (FTW) Project.

We achieved this with three tasks that you can run today:

  1. Building mosaics to create the seasonal Sentinel-2 composites
  2. Model inference to run the PRUE model globally
  3. Vectorization to convert the per-pixel predictions into field boundaries in GeoParquet.
StageArtifactSize
Feature COGs90,918 objects153 TB
Feature Zarr mosaic363,999 objects, 7,499,140 logical chunks150 TB
Prediction Zarr mosaic84,877 objects, 8,982,630 logical chunks45 TB
Vector output, GeoParquet1,000 objects, 8,217,195,679 rows675 GB

In total, we generated 348.7 TB across 540,794 objects, to produce 8.2 billion field boundaries. The full pipeline write-up is in Fields of the World: a GeoAI pipeline on RasterFlow.

Daily wildfire containment mapping with the USDA Forest Service

USDA firecon
FireCon model predictions for the best fire control points, generated across the western United States by RasterFlow.

The USDA Forest Service FireCon model produces containment-suitability maps for active wildfires across the Western United States. Because fuel, terrain, weather, and fire conditions change continuously, yesterday’s map may not reflect the conditions crews face today. Previously, the cost and complexity of running the full pipeline limited how frequently the team could update these maps.

RasterFlow changes that by ingesting potential control location data, terrain characteristics, weather forecasts, and daily soil moisture readings, normalizing the data and running the full inference pipeline end-to-end in a matter of hours rather than a full day. That dramatic reduction in both runtime and cost means the team can afford to run the pipeline multiple times per day, giving frontline response teams a current view of the fire landscape. The results are also delivered as multi-resolution raster layers, so they render smoothly whether crews are looking at a broad regional view or zooming in on a specific fire line. In practice, this translates directly into better-informed containment and resource decisions on the ground, especially during fast-changing fire conditions where yesterday’s map simply isn’t good enough.

Get started today

RasterFlow is available now to Wherobots Cloud Professional Edition users:

  1. Sign in to Wherobots Cloud.
  2. Build a mosaic or run a built-in model by starting with one of these notebooks:
  3. Register for the Public Preview webinar.

Key takeaways

  • RasterFlow is now in Public Preview, opening up the power of planetary-scale Earth observation to all Wherobots customers, available today to Wherobots Cloud Professional Edition users.
  • Predictable, pre-execution pricing. Pricing is based on data volume processed (Area in km² × Pixels per km² × Bands × Time periods × task-specific rate), so you can accurately estimate the cost of any task before execution and extrapolate from smaller test runs to planetary scale.
  • An alternative to Google Earth Engine. Unlike Earth Engine’s EECU-hour billing that reveals cost only after a job finishes, RasterFlow lets you calculate costs up front and keeps results in your own storage. A 150,000 km² Sentinel-2 mosaic with all 12 bands costs about $50 in Earth Engine vs. $13.50 for RasterFlow.
  • Proven at planetary scale with built-in models. RasterFlow ships with Meta’s SAM3, Taylor Geospatial’s Fields of the World, and Meta’s CHM v1. RasterFlow powers wildfire containment predictions for the USDA Forest Service, field-boundary work for Miraterra Soil, and the 8.2 billion global field boundaries Taylor Geospatial produced for Fields of the World.
See a live walkthrough of Rasterflow.

Learning Wherobots by Building a National AI Data Center Suitability Report

This guest post is from a Wherobots user George Chandeep Corea, which covers an exploration of how he used Wherobots MCP with his preferred AI coding tools to build the interactive AI data center site suitability analysis tool below. The following quote is from his own biography.

Learning by Doing

Below is a working national suitability report for AI data centers. Move the weight sliders and the ranking reorders. Toggle the tailings-dam scenario and the buildable area changes. Everything in it was computed on Wherobots over 15.91 million authoritative geometries drawn from 16 national and state portals, then published as a zero-dependency interactive document that opens in any browser with no compute environment behind it.

Have a look first. The rest of this post explains how it was built and how I shaped it, as well as where I plan to go next.

Open the full report in a new tab →

I learn by doing. I built this on a private basis to contribute to the community discourse. For me it is all about the data; any opinions, perceived or intentional, are my own and have nothing to do with my official job roles. All data is from public sources.

It covers 17 candidate industrial sites across all 8 Australian states and territories, spanning the National Electricity Market and the SWIS grid.

Why I built this AI data center suitability model

For me, this is a community and conservation-minded project: not about saying no to development, but about helping identify where development can create the most value with the least harm. The point is to make siting decisions more transparent, more constructive, and easier to discuss across stakeholders.

AI data centers are resource-intensive. They depend on power, water, land, and supporting infrastructure, and they can also create real pressures on ecosystems and communities. I wanted to build something that makes those tradeoffs visible instead of hiding them behind a single score or reports that sit on shelves and don’t let users see it through there lenses.

Why I used Wherobots for spatial suitability at scale

I wanted a platform that could handle real spatial scale without forcing me into a slow, fragile workflow. I had been hearing about it on LinkedIn and Youtube by following Matt Forrest and wanted to learn by doing.

This project queried 15.91 million authoritative geometries across 16 national and state/territory spatial datasets. At the regional modeling scale, the underlying NSW stack ingested 1,751,315 geometries across multiple cloud spatial tables, evaluating 4.92 million spatial join combinations. Using cloud-native GeoParquet inside Wherobots Havasu (Spatially Aware Apache Iceberg) tables, the storage footprint compressed from ~2.9 GB raw equivalent to ~430.7 MB — an 85.2% reduction.

That scale matters because suitability analysis only becomes useful when it can be rerun quickly as assumptions change.

Verified runtime breakdown across 15.9M geometrics

Four different operations, four different numbers. They are not alternatives to each other: 2.4 s for the spatial SQL join execution (distributed joins + net developable area overlays across 1.75M+ geometries); 18.4 s → 3.2 s for the national scan across 15.91M geometries using Hilbert space-filling curve partitioning; 200.6 s for the cold end-to-end batch ETL (uncached ingest, GDA2020 reprojection, ST_MakeValid topology repair, Iceberg writes); and < 1 ms for the client-side What-If recalculation in the browser.

What I was trying to learn

I wanted to understand how far I could push a cloud spatial workflow while still producing something that a non-technical audience could review. The answer was to combine Wherobots for the heavy computation with a document-style HTML report for presentation.

The report is the public-facing layer of the analysis. It includes an interactive map, a ranked candidate leaderboard, benchmarking tables, provenance, methodology, and a what-if sandbox for changing ranking weights.

My recent work covers GDA2020 cadastral modernization at an enterprise level, environmental monitoring at a state level, and disaster modeling and high level coordination with FEMA. As the creator of *AuraSiting Crafter* (currently hunter_spatial crafter), I built an open-source multi-criteria engine benchmarking clean energy, water security, acoustic buffers, and developable land across 15.91 million Australian geometries using Wherobots MCP as the critical foundation.
George Chandeep Corea

Builder, Learner, Public Sector

Why the interactive suitability report matters

The HTML report is where the analysis becomes useful for public review. It is not just a visualization — it is a way to interrogate the assumptions behind the ranking.

The map is interactive, searchable, and configurable. Readers can change basemaps, add WMS/WFS services, and bring in external layers to compare candidate sites against their own data or public reference layers (some functionality coming soon!). That matters because spatial decisions are rarely made from one dataset alone. Adding data will make the report more useful for collaboration, because different stakeholders can overlay the information that matters to them: local context, community knowledge, infrastructure plans, environmental constraints, or jurisdictional boundaries.

The sandbox (already available) is especially valuable because readers can adjust the ranking weights and immediately see how the candidate ordering changes. That makes the tradeoffs tangible instead of theoretical.

Suitability report overview

The report overview combines summary KPI cards, an interactive What-If simulation sandbox, an interactive siting map with layer controls, and a ranked candidate leaderboard into a single reviewable spatial document.

How the suitability workflow works

The pipeline is straightforward:

  1. Ingest authoritative layers — harvest from live government WFS/REST portals and national registers into cloud object storage
  2. Apply exclusions and setbacks — deduct 30 m riparian corridors, 20 m high-pressure gas/water pipeline easements, high-value biodiversity zones, slope >5%, and tailings dam hazards for example.
  3. Calculate network distances — topological winding distances (1.32× factor) to ≥132kV substations and wastewater outfalls for example
  4. Evaluate sensitive receptors — continuous sigmoidal buffer penalty around schools, hospitals, and residential meshblocks for example
  5. Benchmark against regional baselines — compare candidates against simulated baselines across interstate transition hubs (Latrobe Valley, Collie, Gladstone) for example. Only the hunter has detailed data. Sites outside the Development Application (DA) in LMCC has data simulated (not mocked or made up though) through analysing relevant real spatial data.
  6. Publish a zero-dependency report — compile layers, audits, and metadata into a standalone interactive HTML document

The scoring model uses many weights:

  • Power grid proximity: 50%
  • Recycled water proximity: 30%
  • Parcel size: 20%
  • Others added/will be added as the project develops and

Those weights reflect the practical requirements of AI infrastructure while keeping environmental and resource considerations visible.

How the suitability score is calculated

Suitability = 0.40·S_power + 0.25·S_sensitive + 0.20·S_water + 0.15·S_size

Power grid proximity (40%). Proximity to ≥132kV transmission substations. Full points between 100 m and 500 m; linear decay to 0.0 at 5 km; a 0.7 penalty under 100 m for EMF safety and acoustic isolation.

Sensitive receptor buffer (25%). A continuous sigmoidal decay around a 500 m compliance threshold: S_sensitive(d) = 1 / (1 + e^(-0.01·(d - 500))). Hard exclusion under 300 m; acoustic-barrier buffer penalty 300–500 m (0.10–0.50); compliant 500 m–1.5 km (0.50–1.0); optimal workforce balance 1.5–5 km (1.0); linear commute decay beyond 5 km.

Recycled water proximity (20%). Distance to wastewater treatment plants for sustainable evaporative cooling, decaying linearly from 1.0 at ≤1 km to 0.0 at ≥10 km.

Developable parcel size (15%). Contiguous flat buildable land. Parcels ≥15 ha score 1.0; parcels under 3 ha get a 0.1 baseline, with linear interpolation between.

Slope calculations exclude land above 5% grade to avoid excessive earthworks.

What the analysis is based on

National input layers

The national model integrates 16 authoritative datasets across all 8 Australian jurisdictions: Geoscape National Cadastre & G-NAF (15,420,800), ABS 2021 Meshblocks & UCL (368,290), National Sensitive Receptors from ACARA / NHSD / OSM (47,510), Geoscience Australia electricity grid (4,820), BoM & GA surface water and outfalls (42,100), and state planning/cadastral portals — NSW SEED, QLD QSpatial, Vicmap, Landgate (27,725). Total: 15,911,245 geometries.

Those totals are the available universe published across the portals — 47,510 POIs and 15.4M cadastral parcels are what ACARA, NHSD, OpenStreetMap and Geoscape publish, not what this pipeline processed. The ingested and audited cohort is 1.75M regional features in the Hunter deep-dive, 368,290 ABS meshblocks partitioned nationally, 17 indexed candidate sites, and a ground-truth QA sample of 33 sensitive receptors (19 schools via ACARA, 14 hospitals via NHSD) sitting inside the candidate industrial zones across all 8 states, used to calibrate the sigmoidal buffer decay curve.

Of the 17 candidates, 4 are micro-sited in detail in the NSW Hunter; the remaining 13 are simulated baselines across interstate transition hubs (Latrobe Valley VIC, Collie WA, Gladstone QLD).

The earlier NSW regional stack: still the measured layer

DatasetFeature countNotes
NSW Transport Network (Rail)275,421Transport context
NSW Biodiversity Constraint Zones262,258Environmental exclusion
NSW Energy Grid Infrastructure241,573Power proximity
ABS Census Meshblocks223,238Demographic context
NSW Pipeline Corridors197,247Setback constraints
TfNSW Active Transport Pathways188,576Access and corridor context
NSW Hydrography & Waterways181,501Water and riparian context
ABS Regional Demographics1,160Regional benchmarking

Storage and compression on Havasu Iceberg

All spatial tables are cataloged under org_catalog.fgsdb.* on Wherobots Cloud and persisted in cloud object storage, in GDA2020 / MGA Zone 56 (EPSG:7856).

TableGeometriesUncompressed rawGeoParquet footprintSavings
macquarie_biodiversity_constraints262,258~580.0 MB84.2 MB85.5%
macquarie_energy_infrastructure241,573~420.0 MB62.5 MB85.1%
macquarie_transport_rail275,421~390.0 MB58.1 MB85.1%
macquarie_pipeline_corridors197,247~280.0 MB41.8 MB85.1%
macquarie_abs_meshblocks223,238~650.0 MB98.4 MB84.9%
macquarie_water_hydrography181,501~310.0 MB44.6 MB85.6%
macquarie_active_transport188,576~260.0 MB39.2 MB84.9%
abs_demographics1,160~12.0 MB1.8 MB85.0%
Total (8 tables)1,751,315~2.9 GB~430.7 MB85.2%

To avoid real-time WFS/FeatureServer REST API timeouts during Spark runs, every dataset is ingested and hosted as an optimized cloud spatial Iceberg table rather than fetched live.

The spatial SQL underneath

Building the net developable area mask — unioning riparian, pipeline, and rail buffers and subtracting them from the sub-precinct boundaries:

SELECT p.precinct_key,
       ST_Difference(p.geom, ST_Union_Aggr(c.geom)) AS net_developable_geom
FROM precinct_transform p
LEFT JOIN constraints c ON ST_Intersects(p.geom, c.geom)
GROUP BY p.precinct_key, p.geom

Computing nearest distances to transmission substations and wastewater treatment outfalls:

SELECT mb.mb_code21,
       MIN(ST_Distance(mb.mb_geom, ST_Transform(p.geometry, 'EPSG:4326', 'EPSG:7856'))) / 1000.0 AS dist_to_substation_km,
       MIN(ST_Distance(mb.mb_geom, ST_Transform(w.geometry, 'EPSG:4326', 'EPSG:7856'))) / 1000.0 AS dist_to_wwtw_km
FROM industrial_meshblocks mb
CROSS JOIN org_catalog.fgsdb.macquarie_energy_infrastructure p
CROSS JOIN org_catalog.fgsdb.macquarie_water_hydrography w
GROUP BY mb.mb_code21

How I think about the planning question

My background in conservation shaped this project. I’m not trying to use technology to say “no” to development. I’m trying to use it to ask better questions about where development belongs, where it creates value, and how to reduce unnecessary conflict between infrastructure, ecosystems, and communities.

That is why the report includes explicit constraint layers, benchmarking, and scenario controls. The aim is to support better judgment, not replace it.

Benchmarking and query speed mechanics

Benchmarking and speed mechanics breakdown detailing how spatial SQL queries execute across 1.75M+ geometries via metadata envelope pruning, Hilbert curve clustering, vectorized memory execution, and parallel distributed spatial joins.

Four things account for the query speed:

Metadata envelope pruning. Havasu Iceberg stores 2D bounding box envelopes directly inside Iceberg AVRO manifest files. Queries with spatial predicates such as ST_Intersects prune the large majority of irrelevant Parquet files at the metadata layer, before any raw disk bytes are scanned.

Hilbert curve spatial clustering. Geometries are sorted with 2D Hilbert space-filling curves during ingestion, so geographically adjacent features land in the same Parquet row groups and storage partitions. That removes random disk I/O seek overhead.

Vectorized memory execution. Apache Sedona operates directly on columnar GeoParquet WKB geometry buffers, avoiding serialization costs between Python, the Spark JVM, and native spatial drivers.

Parallel distributed spatial joins. Quad-tree and R-tree spatial indexes partition the query space across worker nodes, turning expensive O(N × M) cross-joins into O(N log M) parallel bucket joins.

Candidates are benchmarked against simulated local and regional baselines — Latrobe Valley in VIC, Collie in WA, Gladstone in QLD — to position NSW development opportunities within the wider national energy market transition.

How the what-if sandbox runs with no server

The simulation sandbox runs entirely in the browser: no server calls, no network latency, no cloud API charges during a slider session.

The heavy spatial work — topological winding distances, buffer overlaps, elevation head drops, thermodynamic decay rates — is computed at build time on Wherobots and embedded in the report’s JSON payload. Moving a weight slider then normalizes the raw values so the weights sum to 1.0 and recalculates suitability across every candidate record in JavaScript, in under a millisecond. Toggling the tailings dam safety switch swaps pad areas between the declared and de-declared cases (+15.2 ha unlocked) and updates map polygons and audit cards without refetching any GeoJSON. Candidates re-sort by active score, updating the leaderboard, marker radii, and popup badges.

The payoff is threefold: no per-query cloud compute cost during interactive sessions, full interactivity when the HTML report is emailed or opened offline, and slider manipulation that stays smooth without waiting on network round-trips.

What’s next for the suitability model

The following describes the author’s planned work, not shipped functionality.

Sensitive receptor scoring. The next version will score candidates against sensitive community receptors — schools and early childhood centers, hospitals and aged care, residential meshblocks, and workforce commute bands — to prevent acoustic, thermal, electromagnetic, and visual conflicts. The model uses a continuous sigmoidal penalty around a 500m critical setback threshold, with a workforce accessibility modifier that’s neutral in the 1.5km–5.0km commute band. In practice: under 300m is a critical exclusion, 300–500m carries a buffer penalty requiring an acoustic barrier, 500m–1.5km is compliant, 1.5–5.0km is the optimal community-and-workforce balance, and beyond 5km the score falls off for commute burden.

Grid and water policy analysis. Following the National Cabinet debate on AI data center power demand and regional grid security, the framework will evaluate proximity to 132kV, 330kV, and 500kV bulk transmission, flag candidates adjacent to retiring coal-fired stations that can reuse existing heavy transmission without expensive grid upgrades, map candidates against Renewable Energy Zones and firming assets for 24/7 clean energy matching, and restrict cooling supply strictly to recycled and wastewater sources — scoring zero for any site dependent on potable drinking water reserves or vulnerable aquifers. Candidates then sort into three tiers: grid-ready and fast-track eligible, conditional pending firming storage, or constrained by congestion and potable water reliance.

Open platform integration. To let the public explore these models interactively, hunter_spatial_crafter will integrate with opengeos/GeoLibre. The design point worth noting is zero duplication: Wherobots writes suitability layers to a central cloud bucket as GeoParquet and PMTiles, and GeoLibre’s in-browser DuckDB-WASM engine reads those exact same files over HTTP range requests — fetching only the row groups it needs, with no dataset conversion, no server-side copy, and no large client download.

Engineering efficiencies: reducing compute spend

Total cloud batch compute spend across dozens of iterative development, benchmark, and QA runs was ~$36 AUD (US$24.13) — roughly $1.03 per run across ~35 automated batch runs. Analyzing the execution profile shows how a production pipeline could cut that further:

Decouple heavy geometry joins from lightweight scoring. The pipeline splits into a compute-intensive geometric tier (ingest, GDA2020 reprojection, ST_MakeValid repair, 30 m/20 m buffers, ST_Difference overlays across millions of polygons) and a compute-light scoring tier (decay curves and weighted composites over precomputed distance attributes). Structured as a DAG with intermediate materialised GeoParquet stages, tuning a weight or the sigmoidal threshold never re-runs the heavy joins — only the downstream matrix recalculates.

Fingerprint sources and memoize snapshots. Baseline layers change infrequently. Content hashing (ETags, GeoParquet file hashes, Iceberg snapshot manifest IDs) lets untouched tables be skipped, reading straight from cached Havasu Iceberg partitions.

Process delta partitions. When state portals publish quarterly cadastral updates, Iceberg’s ACID snapshot metadata lets WherobotsDB (optimized and managed Apache Sedona) isolate only modified parcel geometries instead of full continental scans.

Offload interactive compute to the client. Compiling precomputed distance topologies into the standalone report puts millions of interactive public scenario evaluations at $0.00 cloud compute cost.

Applied together, these reduce continuous CI/CD pipeline cost from ~$36 AUD to under $5 AUD.

Why I think this is useful

This project taught me that Wherobots is more than a faster way to run spatial SQL. It is a way to make large-scale spatial analysis practical for real decision support. It gave me a path from raw geodata to a polished, reviewable spatial document — one that can be opened by a stakeholder, examined by a technical reviewer, and discussed in public.

For a personal project, that was the goal: learn the platform, test it at real scale, create something useful, and contribute to the broader discussion about responsible AI infrastructure.

If you’re interested in spatial analysis, cloud geospatial workflows, or responsible AI infrastructure, feel free to connect with me on LinkedIn.

Key takeaways

  • Wherobots computed a national suitability model over 15.91 million authoritative geometries drawn from 16 national and state portals, published as a zero-dependency interactive document that opens in any browser with no compute environment behind it.
  • Cloud-native GeoParquet inside Wherobots Havasu (Spatially Aware Apache Iceberg) tables compressed the storage footprint from ~2.9 GB raw equivalent to ~430.7 MB, an 85.2% reduction.
  • Runtime at national scale: 2.4 s for the spatial SQL join execution, 18.4 s down to 3.2 s for the national scan across 15.91M geometries using Hilbert space-filling curve partitioning, 200.6 s for the cold end-to-end batch ETL, and under 1 ms for the client-side What-If recalculation in the browser.
  • Total cloud batch compute spend across dozens of iterative development, benchmark, and QA runs was ~$36 AUD (US$24.13), roughly $1.03 per run across ~35 automated batch runs.
  • The model covers 17 candidate industrial sites across all 8 Australian states and territories, scored on the formula Suitability = 0.40·S_power + 0.25·S_sensitive + 0.20·S_water + 0.15·S_size.
Get Started with Wherobots

How Bad Telemetry Data Sabotages Modern Fleets

By the Teams at Action Engine & Wherobots

Fleet monitoring is undergoing a generational shift. Fleet monitoring, the systems that ingest and analyze vehicle telemetry to track fleet health, performance, and safety, has become the foundation of how operators run vehicles, not just track them. Modern vehicles generate orders of magnitude more telemetry than even five years ago – GPS, fuel and battery signals, sensor and event streams, on-board diagnostics; that data is no longer just a record of where the fleet has been. It’s the data layer operators use to make fleets more cost-effective, more sustainable, and increasingly more autonomous.

This trajectory points in one direction: vehicles are now navigating by data. Autonomous and semi-autonomous fleets sit at the intersection of GIS, computer vision, and vision-language models – systems where the quality of the underlying data directly determines whether the vehicle stops at the right time, takes the right route, or correctly perceives the world around it.

Generic fleet management platforms were built for an earlier era. They tell you where your fleet is. They struggle with the harder question of where your fleet is going wrong. Fleet management platforms apply global thresholds, treat each ping in isolation, and miss anomalies hidden in the geographic and temporal context around a record.

Equipment doesn’t fail without warning. Vehicles don’t break down without prior warning signs. Fleets aren’t underutilized by accident. But the anomalies that precede these outcomes are almost always invisible inside generic fleet dashboards.

They hide inside telemetry that looks fine in a table – GPS, fuel consumption, temperatures, pressures, sensor readings – until they aggregate into an incident.

To close this gap, Action Engine built Aspen Fleet: an anomaly detection system engineered specifically for the future of fleet operations. Powered under the hood by Wherobots, the industry standard for distributed spatial compute, Aspen Fleet catches the patterns that generic platforms miss. By running against both historical and near-real-time telemetry, it uses spatial context as a first-class signal to surface anomalies earlier than previously possible.

Aspen finds fuel-consumption outliers across thousands of transit vehicles by checking each unit against its own baseline, and maps the deviations onto the specific Portland segments where they occurred.

What Generic Fleet Platforms Miss

Traditional fleet management platforms work well for basic visibility – locations, statuses, and last-known positions. However, they struggle as true anomaly detectors for three specific reasons.

They apply global thresholds

A fuel consumption value that’s anomalous on a flat highway is completely normal on a mountain pass. An idle event in a depot is operational; the same event at an intersection in a residential block is suspicious. When a platform alerts on a single threshold for the entire fleet, it either generates false positives that operators learn to ignore or it misses real anomalies entirely.

They treat each telemetry ping in isolation

A single GPS coordinate that places a delivery truck in the middle of a lake looks like one slightly weird record. A single fuel reading that drops three percent looks like noise. But each of these is part of a sequence – what happened before, what’s happening around it geographically, and what the same vehicle was doing on the same route last week. Without that spatial and temporal context, the signal that matters may never surface.

They don’t see the new data layer at all

Connected and autonomous fleets generate streams that legacy platforms were never designed to validate – computer vision detection logs, model outputs, and perception confidence scores. These streams need their own quality logic: how many objects of a given class the computer vision (CV) stack detected on a route segment today versus yesterday, whether two cameras are producing duplicate detections of the same physical object, or whether model v2.1 has regressed against model v2.0 on a specific stretch of road. Generic platforms don’t ask these questions because they aren’t built to.

This is a gap that Action Engine and Wherobots came together to close.

Segment-level CV regression caught: model v2.1 reports 39 objects against a baseline of 25, driven by duplicate utility-pole detections.

What Aspen Fleet Detects

Aspen Fleet ships with a library of over 100 pre-built detection rules tuned specifically for fleet telemetry. Some catch errors in the data itself, while others catch operational anomalies that traditional platforms miss because they lack spatial reasoning. The rules group into a handful of core categories:

  • Location validity: GPS teleportation (a vehicle moving 500 km in two minutes), impossible coordinates (a car “on water” or in a region it was never dispatched to), distance-to-road violations, and GPS drift (a vehicle 80 meters off any drivable surface).
  • Signal integrity: Coordinate freezes, signal loss in tunnels and dead zones, out-of-order timestamps after backfilled resumptions, and hung trackers reporting identical points sequentially.
  • Operational behavior: Idle outliers in unexpected geographies, dwell-time anomalies on familiar routes, and route deviations that only register when historical patterns are known.
  • Sensor envelopes: Fuel consumption, coolant and oil temperature, tire pressure, EV battery state of charge, and signal strength – each validated against expected envelopes per vehicle, per route segment, and per device class.
  • Computer vision and perception outputs: Detection counts per route segment, duplicate detections across cameras, and model performance comparisons between deployed software versions on the same physical road segment.

The unifying capability is that none of these checks rely on uniform thresholds. They reason about the geography, the route segment, the device history, and the temporal pattern simultaneously. Uniform thresholds are where generic platforms produce false positives and noisy dashboards, while Aspen prioritizes accuracy. Furthermore, when a fleet has its own operational logic that the pre-built library doesn’t cover, customers can easily author custom checks using the same engine.

Two Modes: Historical and Near-Real-Time Fleet Observability

aspex fleet x wherobots spatial anomaly detection

Aspen Fleet runs in two complementary modes

Historical Anomaly Detection

Historical anomaly detection can process years of stored fleet telemetry at scale to find systematic anomalies, retroactively label bad records, and produce a clean baseline that analytics and downstream models can stand on. This is where Aspen is unusually strong. Fleet operators sit on years of telemetry that nobody fully trusts the time or compute to validate it at scale was previoulsy unreachable. Analytics built on that data inherit its noise: KPIs drift, utilization metrics misrepresent reality, maintenance forecasts are anchored on contaminated baselines, and any machine learning model trained on the data inherits whatever errors were hidden within it.

Aspen processes historical datasets at scale with Wherobots distributed spaital compute to find systematic anomalies, retroactively label bad records, and produce a clean baseline that analytics and downstream models can actually stand on. The output isn’t just a list of errors – it’s a measurably more accurate version of the fleet’s own history.

Near-Real-Time Anomaly Detection

Near real-time anomaly detection runs continuously on streaming telemetry, surfacing anomalies within minutes so fleet operations can act before slow-developing patterns become incidents. This mode runs continuously as telemetry streams in. Aspen sees ping sequences in near-real time, applies the same spatial-context-aware checks to streaming data, and surfaces anomalies within minutes of the event that caused them. Fleet operations teams can act on these alerts before a slow-developing pattern – a fuel system slowly degrading, a sensor drifting out of calibration, a route consistently underperforming, or a CV model regressing on a specific stretch of road – turns into a vehicle off the road or a bad decision in production.

The combination matters. Historical analysis tells the team what they’ve been missing, while real-time analysis makes sure they stop missing it going forward.

Surfacing issues earlier, helping teams understand what actually requires action

The practical consequence of spatial-context-aware detection is that Aspen catches anomalies that classical fleet tools surface days or weeks later – if at all. For example:

  • Fuel System Degradation: A fuel system degrading over weeks shows up as a small, slowly widening gap between expected and observed consumption on specific route segments. A global threshold misses it. Aspen sees the drift early because its baseline is segment-specific.
  • Brake-System Issues: This shows up as a subtle change in deceleration patterns at specific intersections the vehicle drives repeatedly. A generic “harsh braking” alert miscalibrated for the fleet either fires constantly or never fires at all. Aspen flags the change because it knows what normal looks like at this specific intersection for this specific vehicle.
  • Perception Stack Decay: A computer vision model degrading shows up as a gradual drop in detection counts or confidence on familiar route segments, often caused by sensor obstruction, a degrading camera, or environmental shifts. Generic fleet platforms don’t see this at all because they don’t process perception outputs. Aspen flags it by comparing detections segment-by-segment and day-over-day.
  • Underutilization: An underutilized vehicle shows up as a pattern of idle outliers in non-operational geofences combined with reduced route variation. Most fleet platforms surface neither signal cleanly. Aspen combines them into a single readable indicator.

This is what “surfacing issues earlier” actually means in practice: anomalies that mature into equipment failure, vehicle breakdown, inefficient utilization, or a degraded perception stack get caught while they’re still “drifts” – not after they’ve become costly incidents.

How Aspen Fleet Runs at Scale

The Infrastructure Behind the Insight

To analyze complex spatial and temporal data at scale, you need a new kind of architecture. Aspen Fleet is cloud-native, utilizing the high-performance distributed spatial compute platform provided by Wherobots.

Wherobots provides the underlying engine that makes spatial data a first-class citizen. Because of this powerful foundation, historical scans across billions of telemetry records are complete in minutes, while streaming checks seamlessly keep pace with the highest-volume fleets in production today.

Data integration is built to be frictionless. Customers can connect their data through whatever channel best suits their architecture – whether that means leveraging existing Geotab and telematics platforms, custom REST endpoints, batch dumps to cloud object storage, or live gRPC feeds directly from on-board vehicle systems. Aspen Fleet and Wherobots handle the ingestion, normalization, and processing automatically.

The Takeaway: Fleet Observability Needs Spatial Context

The fleets that will win the next decade – the ones that are measurably more cost-effective, more sustainable, and increasingly autonomous – are the ones that can read and act on their own data accurately. Most fleet management platforms answer the basic question: “Where is the fleet?” Aspen Fleet answers a much harder, more valuable question: “What is going wrong, and how early can we catch it?”

The combination of deep spatial context, dual historical and real-time processing modes, a validation library tuned for modern perception outputs, and an engine built for massive scale changes the paradigm of fleet operations. Don’t run a fleet you only hope is healthy, run one you can prove is.

Live demo with the teams behind Aspen Fleet and Wherobots

Join us on Tuesday, August 18 at 9AM PT / 12PM ET.

See Fleet Observability in Action

Key takeaways

  • Generic fleet platforms apply global thresholds, treat each ping in isolation, and miss anomalies hidden in geographic and temporal context. Action Engine built Aspen Fleet, an anomaly detection system powered by Wherobots distributed spatial compute, to catch those patterns against historical and near-real-time telemetry.
  • Aspen ships a library of over 100 pre-built detection rules covering location validity (GPS teleportation such as 500 km in two minutes, coordinates on water, 80-meter GPS drift off a drivable surface), signal integrity, operational behavior, sensor envelopes, and computer vision / perception outputs.
  • None of the checks rely on uniform thresholds. They reason about geography, route segment, device history, and temporal pattern at once. Customers can also author custom checks on the same engine when the pre-built library does not cover their operational logic.
  • Aspen runs in two modes: historical scans that process years of stored telemetry at scale to label bad records and produce a clean baseline, and near-real-time checks that surface anomalies within minutes of the event. Integration can use Geotab and other telematics platforms, custom REST endpoints, batch dumps to object storage, or live gRPC feeds from on-board systems.

Introducing the Wherobots Innovation Edition, designed to accelerate your physical world objectives 

The Wherobots Innovation Edition helps you deliver outcomes on top of spatial data that propel your organization forward, and make this data AI-ready.

Today we are announcing the Wherobots Innovation Edition. This is an annual partnership that pairs the full Wherobots Cloud platform with our forward deployed spatial engineering expertise, developed over years of delivering solutions, supporting production workloads, and leading the Apache Sedona project.

We are announcing the Innovation Edition out of ongoing demand. The combination of our expertise in building geospatial data pipelines for production environments, and the Wherobots product has repeatedly delivered success for companies in private arrangements. It gives you a direct way to access our team of distinguished spatial data experts, on demand, to accelerate the realization of your goals.

The Innovation Edition is built for one job, which is helping you ship outcomes while making spatial data ready for AI. As a result: 

  • Risk models are fresh, have more coverage, are more precise, and drive more of the right actions
  • Services, supply chains, and logistical operations can improve continuously
  • Threats to physical ecosystems and supply chains can be mitigated before they cause harm
  • World models improve on the back of higher quality, more complete context
  • CapEx heavy investments can produce higher returns
  • Agriculture and forest management practices are more efficient

We have proven to customers that AI is already capable of co-piloting these and other innovative, physical-world solutions when it has access to the right context. We have successfully delivered on multiple mission-critical engagements where customers relied on our team to remove roadblocks and deliver outcomes on schedule, and relied on our product to deliver outcomes with data. We are not opinionated on what cloud offering you’re using. Need solutions to thrive in AWS, GCP, Azure, Databricks, and Snowflake? No problem. Wherobots is designed to bring its capabilities into these environments. 

Here is what one of our customers, Sergey Sukov at Action Engine has said about working with Wherobots and leveraging the innovation tier. We made this offering public due to customers and partners like Action Engine.

Wherobots has been more than infrastructure for us. Their team helped us design a system that treats spatial context as a first-class signal, and their engine runs it reliably in production against the highest-volume fleets running with Aspen Fleet. That combination of product and expertise is why our fleet data is ready for the AI stack our customers depend on.
Sergey Sukov

CEO, Action Engine

Is the Innovation Edition right for you?

The Innovation Edition fits your team if you:

  • Have widespread interest and investment in the physical world, and need data driven approaches to understand risk and opportunities
  • Want to unleash the potential of AI on spatial data, using the existing data estate you have hosted on Databricks, AWS, Google, Azure, or Snowflake 
  • Want to build competency and ownership of your spatial data workflows rather than outsource them
  • Need solutions fast, regardless of data scale, complexity, and data type

How engagement works

Qualification for the Innovation Edition is a lightweight, no-risk process. It is designed to quickly assess whether we are the right partner for your objectives. 

1. Reach out. We will schedule a 30-minute call to learn about your objectives and technical requirements.

2. Align on the path forward. We can guide your team, ship insights into your environment, or both. We will respond with the path forward, including timelines, milestones, a proof of value, and architecture. A follow-up call or POC can follow, as required.

3. Subscribe. Once there is mutual alignment, you subscribe to the Innovation Edition and the engagement begins. We will co-build the engagement plan with you and how it will grow over time. 

4. Innovate Together. We help you define the objectives and key results that we are achieving together, the technical path to bring results into production, and co-build with you. We will review progress with your team regularly on a weekly or bi-weekly basis, and conduct quarterly business reviews to ensure you are successful.

If you want to get started with an initial discovery call, reach out to us to discuss next steps and building together as a team. 

Key takeaways

  • The Innovation Edition is an annual partnership that pairs the full Wherobots Cloud platform with forward-deployed spatial engineering expertise, developed over years of production workloads and leading the Apache Sedona project.
  • The offering exists because private arrangements combining that expertise with the product have repeatedly delivered for customers. It gives a direct, on-demand way to access distinguished spatial data experts to accelerate outcomes and make spatial data AI-ready.
  • Wherobots is not opinionated about which cloud you run: the post states solutions can thrive in AWS, GCP, Azure, Databricks, and Snowflake. Action Engine CEO Sergey Sukov is quoted on using the innovation tier to treat spatial context as a first-class signal in production against high-volume Aspen Fleet workloads.
  • Engagement is a lightweight qualification: a 30-minute call, alignment on path (guide, ship insights, or both) with timelines, milestones, proof of value, and architecture, then subscribe. Progress is reviewed weekly or bi-weekly, with quarterly business reviews.

Building the Wherobots Mobility Solution Accelerator: A Technical Deep Dive

From Raw GPS Pings to Spatial Intelligence

In Part 1, we explored why mobility data breaks traditional spatial systems and what a modern processing architecture should look like. Now we get into the implementation.

This solution accelerator is a three-notebook pipeline built on Wherobots that processes the Microsoft Research GeoLife GPS Trajectories dataset, 182 users, 17,621 trajectories, and millions of GPS points collected in Beijing between 2007 and 2012. We chose this dataset specifically because it includes altitude data, which lets us demonstrate full XYZM (4D) geometry processing, something most spatial tutorials skip entirely because their data (or their platform) does not support it.

A Python preparation script (prepare_geolife.py) converts thousands of raw .plt files into a single CSV for cloud ingestion. From there, the three notebooks handle everything: ingestion, profiling, cleaning, enrichment, trajectory construction, map matching, spatial indexing, clustering, anomaly detection, and the creation of GeoParquet-backed analytical views that flow directly into Felt for interactive, collaborative visualization.

Notebook 1: Bronze Layer for Ingestion and Spatial Profiling

The Bronze layer loads the combined CSV from S3 into WherobotsDB (optimized Apache Sedona) and establishes the spatial foundation for everything downstream. You can access the notebook here.

Geometry Construction and Parsing

The first operation converts raw latitude and longitude columns into proper 2D point geometries. But even before that, we hit our first real-world challenge. Spark’s inferSchema option inferred the time column as a TimestampType, silently prepending today’s date to raw time values. The dates looked plausible in isolation but were completely wrong.

The fix uses DATE_FORMAT() to safely extract date and time parts regardless of the inferred type, while simultaneously constructing point geometries:

bronze_df = sedona.sql("""

    SELECT

        user_id,

        CAST(latitude AS DOUBLE) AS latitude,

        CAST(longitude AS DOUBLE) AS longitude,

        CAST(altitude_ft AS DOUBLE) AS altitude_ft,

        DATE_FORMAT(date_str, 'yyyy-MM-dd') AS date_str,

        DATE_FORMAT(time_str, 'HH:mm:ss') AS time_str,

        ST_MakePoint(

            CAST(longitude AS DOUBLE),

            CAST(latitude AS DOUBLE)

        ) AS geometry

    FROM raw_gps

""")

This is not just a formatting step, calling ST_MakePoint registers the data as a spatial type, enabling WherobotsDB’s spatial indexing and query optimization for all subsequent operations.

Lesson: Never trust Spark’s schema inference for temporal columns in mobility data. Explicitly cast and parse time values to avoid silent data corruption.

Data Quality Profiling

Before any transformation, we profile the raw data comprehensively. The dataset uses -777 as a sentinel value for missing altitude, so we quantify what percentage of records are affected, check coordinate bounds for points outside the expected Beijing-area extent, and analyze per-user distribution to understand data balance:

sedona.sql("""

    SELECT

        COUNT(*) AS total,

        SUM(CASE WHEN altitude_ft = -777 THEN 1 ELSE 0 END) AS invalid_altitude,

        ROUND(100.0 * SUM(CASE WHEN altitude_ft = -777 THEN 1 ELSE 0 END)

              / COUNT(*), 2) AS pct_invalid,

        ROUND(MIN(CASE WHEN altitude_ft != -777 THEN altitude_ft END), 1)

            AS min_valid_alt_ft,

        ROUND(MAX(CASE WHEN altitude_ft != -777 THEN altitude_ft END), 1)

            AS max_valid_alt_ft,

        ROUND(AVG(CASE WHEN altitude_ft != -777 THEN altitude_ft END), 1)

            AS avg_valid_alt_ft

    FROM bronze

""")

We also use SedonaKepler to visualize a 1% sample of points, confirming that the spatial extent covers the expected Beijing area. This visual validation catches problems that statistical profiling misses, sparse coverage zones, spatial outliers, and artifacts that only become apparent on a map.

The Bronze layer outputs raw GeoParquet to S3, giving us a spatially-typed, columnar, compressed foundation for all downstream processing.


Notebook 2: Silver Layer for The Transformation Engine

The Silver layer is where the pipeline earns its keep. This notebook handles cleaning, 4D geometry construction, trip segmentation, trajectory building, movement metric derivation, spatial indexing, and map matching. Each step has meaningful technical nuance worth examining.

4D XYZM Geometry Construction

After filtering invalid records and converting altitude from feet to meters, we construct 4D XYZM point geometries that encode position, elevation, and time into a single geometric object:

points_4d_df = sedona.sql("""

    SELECT

        user_id,

        latitude,

        longitude,

        altitude_m,

        epoch_seconds,

        date_str,

        time_str,

        ST_MakePoint(

            longitude,

            latitude,

            altitude_m,

            CAST(epoch_seconds AS DOUBLE)

        ) AS geometry_4d

    FROM cleaned

    WHERE epoch_seconds IS NOT NULL

""")

In this encoding, X = longitude, Y = latitude, Z = elevation in meters, M = Unix epoch timestamp as a measure value. We verify the construction with Sedona's dimension inspection functions:

sedona.sql("""

    SELECT

        ST_HasZ(geometry_4d) AS has_z,

        ST_HasM(geometry_4d) AS has_m,

        ST_CoordDim(geometry_4d) AS coord_dim,

        ST_Is3D(geometry_4d) AS is_3d

    FROM points_4d

    LIMIT 1

""")

These checks are not optional—they confirm that downstream operations like ST_ZMin(), ST_ZMax(), and ST_3DDistance() will have the dimensional data they need.

Trip Segmentation with PySpark Window Functions

Continuous GPS streams must be split into discrete trips. We use PySpark window functions to compute the time difference between consecutive GPS points for each user, flag gaps exceeding 20 minutes as trip boundaries, and assign composite trip IDs:

window_user = Window.partitionBy("user_id").orderBy("epoch_seconds")

segmented_df = (

    points_4d_df

    .withColumn("prev_epoch", F.lag("epoch_seconds").over(window_user))

    .withColumn("time_delta_s", F.col("epoch_seconds") - F.col("prev_epoch"))

    .withColumn(

        "is_new_trip",

        F.when(

            F.col("time_delta_s").isNull()

            | (F.col("time_delta_s") > TIME_GAP_THRESHOLD_SECONDS),

            F.lit(1)

        ).otherwise(F.lit(0))

    )

    .withColumn(

        "trip_segment",

        F.sum("is_new_trip").over(window_user)

    )

    # Create composite trip ID: user_id + segment number

    .withColumn(

        "trip_id",

        F.concat(F.col("user_id"), F.lit("_"), F.col("trip_segment"))

    )

)

This approach is clean, declarative, and executes efficiently in Spark’s distributed computation model. The TIME_GAP_THRESHOLD_SECONDS parameter (set to 1,200 seconds / 20 minutes) is easily adjustable for different use cases, delivery fleets might use 5 minutes, while long-haul trucking might use 60.

Trajectory Construction and the Ordering Problem

Once trips are defined, we build XYZM LineString trajectories per trip. This is where we hit one of the most important challenges in distributed trajectory processing: COLLECT_LIST in Spark does not guarantee order. The spatial geometry, the actual XY path, may be correct, but M values (timestamps) can be scrambled.

The solution requires a CTE-based approach that also handles the single-point edge case (ST_MakeLine requires at least 2 points, and Spark evaluates SELECT before HAVING):

trajectories_df = sedona.sql("""

    WITH trip_counts AS (

        SELECT trip_id

        FROM segmented

        GROUP BY trip_id

        HAVING COUNT(*) >= 2

    ),

    valid_points AS (

        SELECT s.*

        FROM segmented s

        INNER JOIN trip_counts tc ON s.trip_id = tc.trip_id

        ORDER BY s.user_id, s.trip_id, s.epoch_seconds

    )

    SELECT

        user_id,

        trip_id,

        ST_MakeLine(COLLECT_LIST(geometry_4d)) AS trajectory,

        COUNT(*) AS point_count,

        MIN(epoch_seconds) AS start_time,

        MAX(epoch_seconds) AS end_time,

        MIN(date_str) AS start_date,

        MAX(date_str) AS end_date

    FROM valid_points

    GROUP BY user_id, trip_id

""")

Sedona provides ST_IsValidTrajectory() to validate M-value ordering. When we applied it, every single trajectory was rejected due to the distributed collect ordering issue. Our pragmatic solution: use point_count filters instead of ST_IsValidTrajectory() as the quality gate, acknowledging that the spatial geometry is correct regardless of M ordering.

Lesson: In any distributed system, aggregation functions that collect values into arrays or lists may not preserve insertion order. Design your pipeline to tolerate this, or implement explicit sorting within the aggregation.

Movement Metrics

With trajectories constructed, we derive per-trip metrics using Sedona’s elevation and measurement functions:

trip_metrics_df = sedona.sql("""

    SELECT

        user_id,

        trip_id,

        trajectory,

        point_count,

        start_time,

        end_time,

        (end_time - start_time) AS duration_s,

        ROUND(ST_Length(trajectory), 6) AS distance_deg,

        ROUND(ST_ZMin(trajectory), 2) AS min_elevation_m,

        ROUND(ST_ZMax(trajectory), 2) AS max_elevation_m,

        ROUND(ST_ZMax(trajectory) - ST_ZMin(trajectory), 2) AS elevation_range_m

    FROM trajectories

    WHERE point_count >= 2

""")

Per-point metrics like speed, elevation delta, distance between consecutive points are computed separately using window functions over the segmented DataFrame with ST_Distance().

Spatial Indexing: H3 and GeoHash

We assign both H3 hexagon cell IDs and GeoHash values in a single pass, giving downstream Gold-layer analytics flexibility in how they aggregate spatial data:

indexed_points = sedona.sql("""

    SELECT

        *,

        EXPLODE(ST_H3CellIDs(geometry_4d, 9, false)) AS h3_cell_id,

        ST_GeoHash(geometry_4d, 7) AS geohash

    FROM enriched

""")

H3 at resolution 9 provides hexagonal cells roughly 105 meters across, ideal for urban mobility density analysis. GeoHash at precision 7 gives approximately 150-meter cells useful for range queries and segment identification.

Map Matching with Wherobots

The Silver layer’s final major operation is map matching, snapping noisy GPS traces to the actual road network. The matcher expects a simple DataFrame with IDs and geometry:

from wherobots import matcher

# Load Beijing OSM road network

roads_df = matcher.load_osm(OSM_DATA_PATH, "[car]")

# Prepare trajectories for matching

paths_df = trip_metrics_df.select(

    col("trip_id").alias("ids"),

    col("trajectory").alias("geometry")

)

# Run map matching

matched_df = matcher.match(

    roads_df,       # Road network edges

    paths_df,       # GPS trajectory LineStrings

    "geometry",     # Road geometry column name

    "geometry"      # Path geometry column name

)

The matcher produces three outputs per trajectory: observed_points (the raw GPS trace), matched_points (the road-snapped route), and matched_nodes (OSM node IDs along the matched path). Having this as a native operation within Wherobots, rather than calling an external API with rate limits and per-request pricing, eliminates an entire category of infrastructure complexity.

Notebook 3: Gold Layer for Analytics, Exploration, and Deep Dives

The Gold layer transforms Silver-layer trajectories into purpose-built analytical views. We structured these into three tiers: analytical, exploratory, and deep dive.

Analytical Views

H3 Hexbin Activity Density Heatmap. We aggregate GPS points by H3 cell and convert cell IDs back to hexagon polygons, enriching each cell with multi-dimensional metrics:

h3_density = sedona.sql("""

    SELECT

        h3_cell_id,

        ST_H3ToGeom(ARRAY(h3_cell_id))[0] AS geometry,

        COUNT(*) AS point_count,

        COUNT(DISTINCT user_id) AS unique_users,

        COUNT(DISTINCT trip_id) AS unique_trips,

        ROUND(AVG(speed_mps), 2) AS avg_speed_mps,

        ROUND(AVG(altitude_m), 2) AS avg_elevation_m,

        ROUND(MIN(altitude_m), 2) AS min_elevation_m,

        ROUND(MAX(altitude_m), 2) AS max_elevation_m

    FROM silver_points

    WHERE speed_mps IS NOT NULL

    GROUP BY h3_cell_id

    ORDER BY point_count DESC

""")

This view is the foundation for understanding spatial activity patterns, where movement concentrates, where it is sparse, and how intensity varies across the study area.

Temporal Patterns. Hourly and day-of-week aggregations reveal when mobility peaks and troughs occur:

temporal_patterns = sedona.sql("""

    SELECT

        HOUR(FROM_UNIXTIME(epoch_seconds)) AS hour_of_day,

        DAYOFWEEK(FROM_UNIXTIME(epoch_seconds)) AS day_of_week,

        COUNT(*) AS point_count,

        COUNT(DISTINCT user_id) AS active_users,

        COUNT(DISTINCT trip_id) AS active_trips,

        ROUND(AVG(speed_mps), 2) AS avg_speed_mps

    FROM silver_points

    WHERE speed_mps IS NOT NULL

    GROUP BY

        HOUR(FROM_UNIXTIME(epoch_seconds)),

        DAYOFWEEK(FROM_UNIXTIME(epoch_seconds))

    ORDER BY day_of_week, hour_of_day

""")

Trip Statistics and Map Matching Coverage. Distribution summaries characterize mobility behavior, while the ratio of successfully matched trajectories provides a quality metric for both the GPS data and the road network.

Exploratory Views

Elevation Profiles. We sample points at 5% intervals along the longest trajectories using ST_LineInterpolatePoint, then extract elevation and timestamps. This required splitting operations into separate CTEs because Spark does not allow EXPLODE() nested inside other expressions:

elevation_profiles = sedona.sql("""

    WITH top_trips AS (

        SELECT trip_id, trajectory, point_count

        FROM silver_trajectories

        WHERE point_count >= 10

        ORDER BY point_count DESC

        LIMIT 10

    ),

    raw_steps AS (

        SELECT EXPLODE(SEQUENCE(0, 100, 5)) AS step

    ),

    fractions AS (

        SELECT step / 100.0 AS fraction FROM raw_steps

    )

    SELECT

        t.trip_id,

        f.fraction,

        ST_LineInterpolatePoint(t.trajectory, f.fraction) AS point_along_route,

        ROUND(ST_Z(ST_LineInterpolatePoint(t.trajectory, f.fraction)), 2)

            AS elevation_m,

        ROUND(ST_M(ST_LineInterpolatePoint(t.trajectory, f.fraction)), 0)

            AS epoch_seconds

    FROM top_trips t

    CROSS JOIN fractions f

    ORDER BY t.trip_id, f.fraction

""")

2D vs 3D Distance Comparison. By comparing ST_Length() with ST_3DDistance(), we quantify the impact of terrain on distance calculations. In hilly areas, the difference can be significant enough to affect route planning, fuel modeling, and ETAs:

distance_comparison = sedona.sql("""

    SELECT

        trip_id,

        point_count,

        elevation_range_m,

        ROUND(ST_Length(trajectory) * 111320, 0) AS distance_2d_m,

        ROUND(

            ST_3DDistance(

                ST_StartPoint(trajectory),

                ST_EndPoint(trajectory)

            ) * 111320, 0

        ) AS straight_line_3d_m,

        ROUND(

            ST_Distance(

                ST_StartPoint(trajectory),

                ST_EndPoint(trajectory)

            ) * 111320, 0

        ) AS straight_line_2d_m

    FROM silver_trajectories

    WHERE elevation_range_m > 50

    ORDER BY elevation_range_m DESC

    LIMIT 20

""")

Map Matched vs Raw Comparison. Side-by-side SedonaKepler layers showing the original noisy GPS trace against the road-snapped route make the value of map matching immediately visible to any stakeholder.

Deep Dive Views

DBSCAN Stop Point Clustering. We filter points with speed below 0.5 m/s, then cluster using ST_DBSCAN with geodesic distance. An important implementation detail: ST_DBSCAN requires a physical column reference, not a computed expression:

# Filter to stationary points

stop_points = points_df.filter("speed_mps IS NOT NULL AND speed_mps < 0.5")

# ST_DBSCAN requires a named reference to a physical column.

# useSpheroid=true so epsilon is in meters (geodesic distance).

clusters_raw = sedona.sql("""

    SELECT

        *,

        ST_DBSCAN(geometry_4d, 100, 5, true) AS cluster_result

    FROM stop_points

""")

The 100-meter epsilon is appropriate for identifying distinct stop locations, buildings, intersections, transit stops, in an urban environment. Clusters with 10+ stop points are then aggregated into hotspots:

hotspots = sedona.sql("""

    SELECT

        cluster AS hotspot_id,

        COUNT(*) AS total_stops,

        COUNT(DISTINCT user_id) AS unique_visitors,

        COUNT(DISTINCT trip_id) AS unique_trips,

        ST_MakePoint(AVG(longitude), AVG(latitude)) AS geometry,

        ROUND(AVG(altitude_m), 1) AS avg_elevation_m

    FROM clusters

    WHERE cluster != -1 AND isCore = true

    GROUP BY cluster

    HAVING COUNT(*) >= 10

    ORDER BY total_stops DESC

""")

Trajectory Anomaly Detection. We flag trips with extreme values using a CTE pattern that avoids Spark’s “HAVING without GROUP BY” error:

anomalies = sedona.sql("""

    WITH trip_stats AS (

        SELECT

            AVG(duration_s) AS mean_duration,

            STDDEV(duration_s) AS std_duration,

            AVG(distance_deg * 111320) AS mean_distance,

            STDDEV(distance_deg * 111320) AS std_distance,

            AVG(elevation_range_m) AS mean_elev_range,

            STDDEV(elevation_range_m) AS std_elev_range

        FROM silver_trajectories

        WHERE duration_s > 0

    ),

    classified AS (

        SELECT

            t.trip_id,

            t.user_id,

            t.trajectory,

            t.duration_s,

            ROUND(t.distance_deg * 111320, 0) AS distance_m,

            t.elevation_range_m,

            CASE

                WHEN t.elevation_range_m

                    > (s.mean_elev_range + 3 * s.std_elev_range)

                    THEN 'extreme_elevation'

                WHEN t.duration_s

                    > (s.mean_duration + 3 * s.std_duration)

                    THEN 'extreme_duration'

                WHEN t.distance_deg * 111320

                    > (s.mean_distance + 3 * s.std_distance)

                    THEN 'extreme_distance'

                ELSE NULL

            END AS anomaly_type

        FROM silver_trajectories t

        CROSS JOIN trip_stats s

    )

    SELECT * FROM classified

    WHERE anomaly_type IS NOT NULL

""")

Road Segment Speed Analysis. By joining matched routes with trajectory metrics and decomposing routes into individual segments using GeoHash pairs, we calculate average speed per road segment and classify by congestion:

matched_with_metrics = sedona.sql("""

    SELECT

        m.ids AS trip_id,

        t.user_id,

        m.matched_points AS geometry,

        t.duration_s,

        ROUND(ST_Length(m.matched_points) * 111320, 0) AS matched_distance_m,

        CASE WHEN t.duration_s > 0

            THEN ROUND(

                ST_Length(m.matched_points) * 111320 / t.duration_s * 3.6, 1)

            ELSE 0

        END AS avg_speed_kmh,

        CASE WHEN t.duration_s > 0 THEN

            CASE

                WHEN (ST_Length(m.matched_points) * 111320

                      / t.duration_s * 3.6) < 15 THEN 'congested'

                WHEN (ST_Length(m.matched_points) * 111320

                      / t.duration_s * 3.6) < 40 THEN 'urban'

                WHEN (ST_Length(m.matched_points) * 111320

                      / t.duration_s * 3.6) < 80 THEN 'arterial'

                ELSE 'highway'

            END

            ELSE 'unknown'

        END AS road_class

    FROM silver_matched m

    INNER JOIN silver_trajectories t ON m.ids = t.trip_id

""")

This visualization uses the actual road LineString geometries from map matching, not H3 hexagons, providing road-level granularity that is directly actionable for traffic engineering and route optimization.

From Notebooks to Dashboards: The Wherobots + Felt Connection

Processing mobility data at scale is only half the equation. The other half is getting the results into the hands of decision-makers (operations teams, urban planners, fleet managers, logistics analysts) who need interactive, shareable maps, not Jupyter notebooks.

This is where the Wherobots and Felt integration completes the picture. Wherobots and Felt recently announced a strategic partnership that connects Wherobots’ spatial intelligence lakehouse directly with Felt’s collaborative, browser-based mapping platform. The integration lets organizations go from processing petabyte-scale geospatial datasets in the cloud to exploring insights in interactive maps, without moving large datasets between systems.

For this mobility accelerator, the workflow is straightforward. Every Gold-layer view we produce, H3 density heatmaps, hotspot clusters, road segment speed classifications, trajectory anomaly maps, is written as GeoParquet to S3. Through the native Wherobots-Felt integration, these datasets are directly accessible in Felt, where they become live, interactive, collaborative map layers. There is no export step, no format conversion, and no data movement friction.

What this means in practice:

Shareable analysis, not static screenshots. Instead of exporting a Kepler.gl map as an image or HTML file, the H3 activity density view becomes a live Felt map that operations teams can explore, filter, annotate, and share via a link, viewable from any device, no GIS software required.

Collaborative investigation. When the anomaly detection pipeline flags suspicious trajectories, an analyst does not need to walk a fleet manager through a notebook. They share a Felt map where the flagged routes are overlaid on the road speed classification layer, and the fleet manager can pan, zoom, and query the data themselves.

Operational dashboards from analytical views. The road segment speed analysis, with its congested/urban/arterial/highway classification, becomes a traffic conditions dashboard. The hotspot identification layer becomes a POI and dwell-time analysis tool. These are not one-off visualizations—they are reusable, updatable artifacts that stay connected to the processed data.

AI-assisted map creation. Felt’s AI-driven interface lets users interact with spatial data using natural language prompts, lowering the barrier for non-GIS teams to extract insights. Combined with Wherobots’ processing power, this creates what both companies describe as the “SQL-to-map” workflow, from Spatial SQL query to interactive, shareable map in seconds.

This combination is already in production. Leaf Agriculture uses Wherobots and Felt together to process millions of acres of tractor telemetry and imagery data, turning their agricultural data lake into interactive maps and dashboards distributed via links instead of in-person screen-sharings. The mobility use case follows the same pattern: process at scale with Wherobots, visualize and collaborate in Felt.

Challenges and Solutions: A Practitioner’s Reference

Every mobility data pipeline encounters edge cases. Here are the ones we solved in this accelerator, documented so you can avoid them:

1. Spark inferSchema corrupting time values. Spark silently prepended today’s date to inferred TimestampType columns. Fix: Use DATE_FORMAT() to extract date and time parts explicitly.

2. ST_MakeLine failing on single-point groups. A LineString requires at least two points, but Spark evaluates SELECT before HAVING. Fix: Use a CTE to pre-filter trips with >= 2 points before the aggregation query.

3. COLLECT_LIST not preserving order. Distributed aggregation does not guarantee array ordering. Fix: Accept unordered M values in trajectories and use point_count filters instead of ST_IsValidTrajectory() as a quality gate.

4. EXPLODE nested in expressions. Spark does not allow EXPLODE() inside other SQL expressions. Fix: Split the operation into separate CTEs.

5. ST_DBSCAN requiring physical column references. The function does not accept computed expressions as geometry input. Fix: Use the persisted GeoParquet column.

6. HAVING without GROUP BY. Applying aggregate thresholds without a GROUP BY clause. Fix: Use a CTE to compute thresholds, then apply as WHERE conditions.

Apache Sedona Spatial SQL Functions Used

This accelerator demonstrates a broad cross-section of Sedona’s spatial SQL capabilities:

CategoryFunctions
4D Point ConstructionST_MakePoint(x, y, z, m)
Dimension InspectionST_Z(), ST_M(), ST_HasZ(), ST_HasM(), ST_CoordDim(), ST_Is3D()
Elevation AnalysisST_ZMin(), ST_ZMax(), ST_3DDistance()
Trajectory BuildingST_MakeLine(), ST_IsValidTrajectory()
Trajectory SamplingST_LineInterpolatePoint()
Spatial IndexingST_H3CellIDs(), ST_H3ToGeom(), ST_GeoHash()
Spatial MeasurementST_Length(), ST_Distance(), ST_StartPoint(), ST_EndPoint()
ClusteringST_DBSCAN()
Map Matchingmatcher.load_osm(), matcher.match()
Geometry ConstructionST_MakePoint(), ST_Buffer()
VisualizationSedonaKepler.create_map(), SedonaKepler.add_df()

GeoParquet Output and Tool Interoperability

All Gold-layer views are persisted as GeoParquet files. GeoParquet has emerged as the standard columnar format for geospatial data in the cloud-native ecosystem, offering efficient compression, predicate pushdown for spatial filters, and broad tool compatibility. The analytical views produced by this accelerator are immediately consumable in Felt, Kepler.gl, QGIS, Foursquare Studio, DuckDB Spatial, and any other tool that reads GeoParquet, no export step, no format conversion, no data loss.

With the Wherobots-Felt integration, GeoParquet outputs in S3 become live data sources for interactive maps. This closes the loop from raw GPS pings to collaborative dashboards in a single, end-to-end spatial data stack.

Get Started with GPS Trajectory Processing on Wherobots

The Wherobots Mobility Solution Accelerator is designed to be a starting point, not a black box. The three notebooks are fully documented, the Spatial SQL is readable and modifiable, and every intermediate result is inspectable as a Spark DataFrame or visualizable in SedonaKepler, and from there, publishable as an interactive Felt map.

Whether you are processing fleet telematics, rideshare trajectories, maritime AIS data, or drone flight logs, the patterns demonstrated here—medallion architecture, 4D geometry processing, distributed trip segmentation, integrated map matching, GeoParquet output, and collaborative visualization through Felt—translate directly to your use case.

To explore the accelerator, visit Wherobots and get started with a notebook environment and connect your favorite IDE to the Wherobots MCP Server and Spatial AI Coding Tools. To see how the Wherobots + Felt stack works together, check out the integration documentation. If you have questions about applying these patterns to your mobility data, reach out to the Wherobots team or join the Apache Sedona community on Discord.

Get Started with the Spatial AI Coding Tools

Key takeaways

  • The accelerator is a three-notebook medallion pipeline on Wherobots that processes Microsoft Research GeoLife: 182 users, 17,621 trajectories, and millions of GPS points collected in Beijing between 2007 and 2012. GeoLife includes altitude, which lets the pipeline demonstrate full XYZM (4D) geometry processing.
  • Never trust Spark schema inference for temporal columns in mobility data: inferSchema treated time as TimestampType and silently prepended today's date. The fix is DATE_FORMAT() with explicit casts, plus ST_MakePoint so the data is a spatial type from Bronze onward. The dataset uses -777 as a sentinel for missing altitude.
  • Trip segmentation uses a 20-minute (1,200-second) gap threshold, adjustable for other domains. COLLECT_LIST does not preserve order in Spark, so ST_IsValidTrajectory() rejected every trajectory; the pipeline uses a point_count >= 2 quality gate instead. Map matching is native in Wherobots via matcher.load_osm and matcher.match, not an external rate-limited API.
  • Gold views include H3 density at resolution 9 (~105 m cells), GeoHash precision 7 (~150 m), ST_DBSCAN stop clustering (100 m epsilon, min 5 points, geodesic), and anomaly flags at mean + 3 standard deviations. Outputs are GeoParquet on S3, consumable in Felt, Kepler.gl, QGIS, DuckDB Spatial, and Foursquare Studio.

Introducing developer tools that let AI build with physical world data

Your AI can now understand and query spatial data using the Wherobots MCP server, VS Code extension, and CLI.

The physical world is a new frontier for AI, but modern AI-driven tools are limited by what they can do with this type of data. LLMs don’t understand how to use physical world data for analytical purposes. For example, when we asked ChatGPT to compute the flood risk from sea level rise for all homes along the California coastline, it came back claiming that “no one can give you a single exact number”, plus some unverifiable numbers:

wherobots-mcp-server example

That may have been true, until now!

In the past, it would take highly skilled developers who are familiar with spatial data and its query patterns, weeks to prototype this type of analysis. Now, developers and analysts, irrespective of their geospatial skills, can build it on-demand using natural language. Using your AI agent connected to Wherobots, they can build working solutions with small to very large scale geospatial data in minutes.

Wherobots now offers your AI the capability to understand spatial data to generate quality code. The pairing includes direct access to WherobotsDB, a cloud-based, secure, and distributed execution environment purpose-built to generate results efficiently at scale. Soon, the same AI will be capable of driving RasterFlow, the planetary-scale inference engine for Earth Intelligence to extract machine-generated insights from satellite, drone, and sensor datasets.

If you’re working in the energy space or you’re interested to see an agent take plain-language question and orchestrating pipelines to end results, check out the upcoming live session.

mcp rasterflow energy webinar
Spatial AI for Energy

Enabling AI to Work with Physical World Data

We are launching three new tools that let AI understand and drive the analysis of spatial data; an MCP server, a CLI, and a VS Code extension. These tools are designed to bring geospatial development into IDEs like VS Code, Claude Code, OpenCode, Cursor, Windsurf, Kiro and others. And soon the Wherobots connector will make it easy to use the same analytical capability in Claude and ChatGPT from respective marketplaces.

With these tools, LLMs and agents can discover and understand spatial data, generate and debug code, and build solutions considerably faster using natural language as an interface.

As a result, developers and analysts immediately become more productive with geospatial data, enabling them to solve more problems, and shorten the development cycle from weeks down to potentially minutes. Wherobots integrates with data lakes and lakehouses including AWS S3, Databricks Unity Catalog, and AWS Glue, allowing teams to realize massive gains using their existing data.

Getting Started

  1. Sign up for a free trial of Wherobots Pro. Use Wherobots for free for 30 days and up to $95, with an additional $250 credit if you activate your spatial AI coding tools.
  2. Create a Wherobots API key.
  3. Install the Wherobots extension for VS Code.

Optionally, establish secure integrations with datasets in Amazon S3, Databricks Unity Catalog, or AWS Glue. Wherobots will utilize these integrated datasets, along with hosted datasets within its catalog.

You can also install Wherobots on other IDEs, work with the MCP server, utilize the CLI and agent skills.

Start using the extension by typing a prompt into the chat window:

  • “Using NDVI datasets and field boundaries for California, what kind of field-level crop health insights can I compute?”
  • “Write a notebook for me to ingest the latest American Community Survey data from US Census.”
  • “I need to know how many properties in California coastal cities are facing flooding risk. I want you to come up with an analysis framework for this, help me identify the right datasets to help me solve this problem, generate some initial insights from those for me to validate, and finally (once I approve) write a notebook for me to run this analysis independently.”

Using VS Code and Claude Opus 4.6 harnessed to Wherobots, we were able to complete the California flood risk analysis that ChatGPT couldn’t in under 30 minutes and for less than $5 of Wherobots usage. The California Coastal Flood Risk notebook we built is here.

Here is a quick video interacting with a generated notebook via VSCode.

Enabling Agents to Understand the Physical World

What we announced today enables people to use AI to build solutions with physical world data, at a fraction of the time and cost. Agents are now capable of developing production-grade geospatial data applications, autonomously or semi-autonomously with a human in the loop with Wherobots acting as the natural language interface and the context engine between the two.

What’s Possible Now

Here are a few example prompts that can drive prototypes and working solutions. To build prototypes or solutions, you will need to ensure the right data is integrated with Wherobots such that your AI can use Wherobots to understand it, and execute effectively based on your directions. For many organizations, this data is already available either in their private data lake (Amazon S3, Databricks Unity Catalog, AWS Glue Data Catalog) or in the public domain (STAC, public datasets, purchased datasets).

Mobility, fleet management, and logistics: “Use Wherobots to transform the raw GPS data for the month of March [located in Databricks Unity Catalog Table X] into trips, and match it to the Overture transportation network using Wherobots map matching. Tell me which segments of road in the state of California were the most constraining for my trips. Define constrained as the speed traveled was less than 50% of the advertised speed limit, rank these segments by trips taken, and also eliminate road segments that were within 1/4 mile of an intersection.”

Marketing and advertising: “Using the fields of the world dataset generated by Wherobots RasterFlow, join fields to Regrid parcels. Also join this result with my customer database located in [S3 location]. I want to identify unique farm owners who are not customers and sell to them.”

Insurance: “Compute the flood risk score using the [flood plane raster] for all properties under general home insurance located in my [S3 bucket path]. Identify which properties have flood insurance and are at the most risk, based on [criteria X]. Separately tell me which properties are at risk, but not insured so I can target them for insurance offerings.”

Agriculture: “Use the fields of the world dataset as a filter. Use RasterFlow and the latest Sentinel 2 data in the AWS data exchange to compute NDVI for the state of California. Join these results with the fields to produce NDVI statistics in the month of July for all fields in California.” (This example will be AI-driven soon, but is feasible today with RasterFlow in private preview)

What’s Coming Next

In the coming months you can expect:

  • Wherobots Connector on the Claude and OpenAI Marketplaces
  • OAuth support for the MCP Server
  • Additional hosted datasets in the Wherobots Hub
  • New integrations with additional data sources and catalogs

Here’s a sneak peek of the experience we are planning to enable via Claude, including notebook generation and insights provided directly inside Claude Chat:

Please reach out to us at product@wherobots.com or support@wherobots.com if you have feedback, requests, or questions. Start your free trial below and start using the spatial AI coding tools.

Start your free trial.

Key takeaways

  • Wherobots launched three tools that let AI understand and drive spatial analysis: an MCP server, a CLI, and a VS Code extension, designed for IDEs including VS Code, Claude Code, OpenCode, Cursor, Windsurf, and Kiro. A Wherobots connector for Claude and ChatGPT marketplaces is described as coming soon.
  • Using VS Code and Claude Opus 4.6 harnessed to Wherobots, the California coastal flood-risk analysis that ChatGPT declined to quantify as a single number was completed in under 30 minutes for less than $5 of Wherobots usage. The notebook is linked from the post.
  • Free trial stated here: Wherobots Pro for 30 days and up to $95, with an additional $250 credit if you activate spatial AI coding tools. Create an API key, install the VS Code extension, and optionally integrate S3, Databricks Unity Catalog, or AWS Glue.
  • Wherobots integrates with AWS S3, Databricks Unity Catalog, and AWS Glue so agents work against existing lakes. RasterFlow-driven agent workflows are described as coming soon. Roadmap items: marketplace connectors, OAuth for MCP, more Hub datasets, and additional catalogs.

It takes 15 minutes for the Caltrain to get from Sunnyvale to SAP Center

That’s how long it took our MCP server to go from “how many bus stops are in Maryland” to an answer

I’ve been doing a lot of reading lately on how AI is going to transform spatial workloads and that curiosity led me to this post on geoMusings. Here, Bill is demonstrating how Claude Code and agent skills capabilities can be used to wire up a chat-to-query-results interface in a few hours. He showcased the new skill by getting the agent to query his local Postgres instance for the number of Metro bus stops in Maryland, which returned a precise 4,563.

I need to count the number of records in the metro_bus_stops table that are inside Maryland.The database is at localhost:5432, database name is “dev”,user “postgres” with password “postgres”.
Points table: public.metro_bus_stops (geometry column: geom, id column: id)Polygons table: public.maryland_boundary (geometry column: geom, name column: name)

As a dabbler of AI agents and a minor contributor to Wherobots’ very own MCP server, I immediately wondered how our MCP server would do against such a challenge. So I fired up my VS Code and just straight up asked:

“How many bus stops are in Maryland?”

Bear in mind, at the time I did not know if we have any data with bus stops in it in Wherobots’ data catalogs, I did not know what shape that data was in, I did not know if the MCP server could come up with a reasonable administrative boundary for Maryland, etc. And I fired off this query just as my CalTrain was departing Sunnyvale station.

In about 5 minutes, the MCP server already identified two tables with bus stop information called places_place under the Overture Maps Foundation database in Wherobots Open Catalog. It achieved that by exploring our catalog and running sample queries against those tables to find the right data; all with zero human intervention. We are right about Lawrence Station at the point.

In the next 5 minutes, the MCP server ran a series of queries against that table, self-identified errors (i.e., got 0 results and understood it was not expected), adjusted the query, switched tables, changed approaches until it was able to produce actual results. Our MCP server believes there are 19,740 bus stops in Maryland which is ~5 times as many as Bill’s post suggests. We just got to Santa Clara station, by the way, for those of you who are still following.

So being a good aspiring data engineer, I challenged the MCP server:

Why does this blog think there are only 4563 then?
https://blog.geomusings.com/2026/01/14/spatial-analysis-with-claude-code/ 

The MCP server went back to work and gave me the diagnosis; Bill’s query is focused on Metro bus stops and my original question did not specify that:

So in the last 5 minutes of this journey, I asked it to focus on Washington Metropolitan Area Transit Authority (WMATA) bus stops only and see what it comes up with! And just as we were about to pull into San Jose Diridon Station, the MCP server told me that there are 6,224 Metro bus stops in Maryland. 

Now, whether there are 4,563 Metro bus stops in Maryland or 6,224 ones, is a matter that shall be validated with people far more knowledgeable than myself on buses and their stops. The main point is that AI is making it possible for non-experts like myself to go from a question (expressed in natural language) to real insights in minutes (well a 15-minute train ride to be precise). Wherobots MCP is giving the AI the ability to answer questions about the real-world. 

In the real world, I would have asked the MCP server to generate a Notebook for me to reproduce this output and plot it on a map. I would then share that with my colleague to help me validate, correct and optimize my findings. What would have taken days to weeks (to go from theory to some early explorations to a shareable PoC and, finally, to production-quality code) can now be achieved in a matter of hours. 

The Caltrain experiment was just one question. In our recent office hours, we walked through the MCP server end to end, showing how it explores catalogs, generates spatial queries, debugs errors, and produces reproducible outputs. See the full workflow in action.

Want to get started with our MCP server? Check out our getting started guide. It takes less than 5 minutes to configure the server and start chatting with the physical world! 

Create your account to get started

Key takeaways

  • On a Caltrain ride from Sunnyvale to San Jose Diridon (~15 minutes), the author asked the Wherobots MCP server in VS Code 'How many bus stops are in Maryland?' with no prior knowledge of whether the catalog had bus-stop data or a Maryland boundary.
  • In about five minutes (around Lawrence Station) the MCP server identified places_place under Overture Maps in the Wherobots Open Catalog by exploring the catalog and running sample queries, with zero human intervention.
  • In the next five minutes it self-identified errors (including 0-result queries), adjusted SQL, switched tables, and produced 19,740 bus stops in Maryland — about 5× the 4,563 Metro bus stops in Bill Dollins' geoMusings Claude Code post. When challenged with that post, it diagnosed that Bill's query was Metro-only.
  • In the last five minutes, constrained to WMATA, it returned 6,224 Metro bus stops in Maryland. The author states those counts still need validation by people who know buses; the point is natural-language to insight in minutes. A notebook for validation is described as what they would do in a real workflow.