The Wherobots Spatial AI Assistant is now in the Anthropic Connectors Directory Posted on August 20, 2026October 3, 2026 by Tiffany Huynh You can now ask Claude questions about the physical world and get answers grounded in real spatial data. Which facilities sit in the path of today’s storms? Which county west of the Mississippi has the fewest pharmacies per person? How much of Austin’s rooftop area sits under canopy? The Wherobots Spatial AI Assistant, now available in the Anthropic Connectors Directory, answers these questions in plain language and returns results, maps, and reports directly in your Claude conversation. Spatial AI Assistant by Wherobots on Anthropic Marketplace Wherobots is the AI context engine for the physical world. It gives AI the spatial intelligence it needs to reason about where things are, how they relate, and what has happened on the ground. Here is an example of an insurance risk analysis of all of Colorado that we initially built with Claude Opus 5 and the Spatial AI Assistant, then expanded on with the VSCode Spatial AI Coding Assistant to be a complete hosted application. The challenge with traditional data systems for building an application like this is both the skills and the data processing cost. We were able to easily join Regrid parcels with Overtures buildings, build hexes that show various layers of risk derived from open raster data, along with the underlying geometry scored on each parcel in a matter of minutes with a Wherobots large runtime. Having the infrastructure and processing capability at you and your AI’s “fingertips” to build with makes moving from idea to production application far easier and faster than without. Screening indicators from modelled and tract-resolution inputs — not underwriting. Counts are floors. Imagery is locational context only. Open full screen ↗ Getting started with the Spatial AI Assistant takes two steps: install the connector, and sign in to your Wherobots organization, and then begin to prompt Claude. With zero additional setup, Claude can begin to utilize datasets including: Overture: Buildings, Transportation, Places Foursquare: Open Places NWS Watch/Warnings, updated hourly GLO-30/90 DEM Elevation Rasters Regrid Parcels (sample) See it in action here: Interacting with Claude Chat to create an analysis about coastal flood risk in California. Analyzing open datasetsThe Wherobots Global Hub continues to expand the open datasets it offers out of the box. In the near future we plan to add wildfire risk models to the hub in partnership with the USDA Forest Service as well as the national hydrography dataset (NHD). You can also connect to STAC collections for additional raster datasets. If there is open or proprietary data that you are interested in having access to fully supported and maintained by Wherobots while iterating with Claude’s models, reach out to us about it. Analyzing proprietary datasetsIn a few steps, customers can easily attach Wherobots to their S3 storage buckets, catalog systems like Databricks Unity Catalog and AWS Glue Data Catalog, and authoritative 3rd party datasets like Regrid’s Parcels. These integrations enable Wherobots to query proprietary mobility, customer, and asset data with Claude.Once you’ve connected Wherobots every Claude conversation can query the data Wherobots has access to, so you can test your ideas and drive progress with physical world data. How the Spatial AI Assistant Works The Spatial AI Assistant plugs into the Wherobots Spatial SQL API and uses skills for accessing the Wherobots documentation, as well as the datasets registered in the Wherobots Global Hub. Claude uses Wherobots skills packaged with the connector to write strong, valid spatial SQL with (generally) correct functions, syntax, and joins. Results come back into your Claude conversation, ready to work with, and you don’t need to interact with a roadmap or a team of spatial experts to get it. Try It Today There’s a 14 day/$95 free trial available to all first time users to Wherobots that you can utilize to test your ideas. After the trial expires, customers pay a low on-demand rate for Wherobots spatial units consumed through the Spatial SQL API. When you start your trial let us know you have connected the Spatial AI Assistant and you can qualify for an additional $250 in Wherobots credits. Getting started is easy. Install Wherobots Spatial AI Assistant for Claude, connect your Wherobots organization, and ask your first question like: How much of Austin’s rooftop area sits under canopy? Want to see what this looks like in practice before you install? Watch the MCP demos in the videos below. The workflows shown there, from natural language questions to working spatial SQL to results on a map, are the same workflows the Spatial AI Assistant puts in front of your entire organization. Key takeawaysThe Wherobots Spatial AI Assistant is now in the Anthropic Connectors Directory, so you can ask Claude questions about the physical world and get results, maps, and reports in the conversation.Getting started is two steps: install the connector and sign in to your Wherobots organization. Out of the box Claude can use Overture buildings, transportation, and places; Foursquare Open Places; NWS watches and warnings (hourly); GLO-30/90 DEM; and a Regrid parcels sample.The assistant plugs into the Wherobots Spatial SQL API and uses packaged skills so Claude writes valid spatial SQL. You don’t need a GIS team in the loop to get an answer.First-time users get a 14 day / $95 free trial. Tell Wherobots you connected the Spatial AI Assistant and you can qualify for an additional $250 in credits.
Introducing the Wherobots Python SDK Posted on May 6, 2026October 3, 2026 by Daniel Smith What is the Wherobots Python SDK? The Wherobots Python SDK is a typed Python client for submitting, monitoring, and managing Wherobots job runs. It ships on PyPI as wherobots-python-sdk. One install, one API key, and you’re running spatial jobs from any Python environment: CI/CD pipelines, notebooks, a local shell. The SDK is built for three workflows. Engineers wiring spatial jobs into production pipelines. Data scientists iterating on a Wherobots script and streaming logs back to the terminal. Ops leads watching what’s running across the organization. Wherobots Python SDK Use Cases Wherobots customers run spatial workloads on a cadence: mapping platformsrefreshes OSM-derived road networks every morning before downstream pipelines fire. Ag-tech teams pull Sentinel-2, compute NDVI, and join it to millions of crop boundaries every five days. Each of these is a Wherobots job that requires uploading scripts to shared storage, wiring up an Airflow DAG, and managing the operator that called the Runs REST API. The Wherobots Python SDK is the shorter path. One install, one API key, and three lines of Python submit a job from any environment that runs Python: a CI/CD pipeline, AWS Step Functions, a notebook, a local shell. The jobs module of the SDK assembles three workflows customers were already trying to assemble by hand. Scheduled spatial data products Mapping and data-provider customers refresh Overture, OSM, or Sentinel-derived datasets on a fixed cadence. The output is a versioned Iceberg or GeoParquet table that downstream teams query directly. With the SDK, the entire refresh is a Python function that runs on cron, GitHub Actions, or AWS EventBridge. No Airflow cluster to host. No operator package to install. Orchestrator-invoked jobs from Step Functions, Prefect, or Dagster Several Wherobots customers run their broader ETL in AWS Step Functions or Prefect and treat Wherobots as one task in a longer chain. The SDK gives those orchestrators a clean Python interface to submit a Wherobots job, wait for completion, and pass the output URI to the next step. A single call replaces the boilerplate of presigned uploads, polling, and log retrieval. CI/CD-driven analytics for production spatial pipelines Engineers maintaining property climate risk pipelines, mobility joins, or agricultural monitoring jobs want their spatial code to ship through the same review and deploy path as the rest of the codebase. The SDK fits a standard pattern: commit a script to GitHub, run tests in CI, deploy by submitting a job from the runner with WherobotsJob.submit(). Logs stream back to the build output. Failed jobs fail the build. The rest of this post walks through install, the WherobotsJob API, dependency management, and the security model. Install the Wherobots Python SDK pip install wherobots-python-sdk export WHEROBOTS_API_KEY="your-api-key" The only runtime dependency is requests. No AWS credentials, no bucket configuration. The WherobotsJob class The SDK exposes a single class today: WherobotsJob. Point it at a script, give the job a name, and call .submit(). Submit a job and stream logs from wherobots import WherobotsJob job = WherobotsJob( script="etl_pipeline.py", name="nightly-etl", runtime="large", ) job.submit() status = job.wait_for_completion(stream_logs=True) print(f"Finished with status: {status.value}") That’s the full lifecycle. The SDK uploads local scripts to Wherobots-managed storage via presigned URLs, polls for completion, and streams logs back to your terminal. When wait_for_completion returns, you get a JobStatus enum: COMPLETED, FAILED, or CANCELLED. Only an API key. No AWS credentials, no bucket setup. Pass arguments and configuration Real jobs need arguments, Spark configuration, and dependencies. Pass them through the constructor: job = WherobotsJob( script="spatial_join.py", name="q4-spatial-join", runtime="x-large-himem", timeout_seconds=7200, args=["--input", "s3://bucket/parcels/", "--output", "s3://bucket/results/"], spark_configs={ "spark.sql.shuffle.partitions": "200", "spark.executor.memory": "8g", }, dependencies=[ WherobotsJob.add_pypi_dependency("geopandas", "0.14.0"), WherobotsJob.add_file_dependency("s3://bucket/libs/custom_udfs.whl"), ], ) The SDK validates inputs at construction time. Bad runtime names, missing JAR main classes, negative disk sizes, and empty scripts all raise WherobotsValidationError before a single network call leaves the client. List and filter runs A WherobotsJob instance isn’t required to query your organization’s job runs: from wherobots import WherobotsJob, JobStatus page = WherobotsJob.list_runs( status=[JobStatus.FAILED], name_pattern="etl-*", size=10, ) for run in page.items: print(f"{run.id} {run.name} {run.status}") Cancel jobs and handle errors from wherobots import WherobotsJob, WherobotsTimeoutError job = WherobotsJob(script="long_running.py", name="cancellable-job") job.submit() try: status = job.wait_for_completion(max_wait_seconds=600) except WherobotsTimeoutError: job.cancel() print("Job cancelled after timeout") The exception hierarchy is flat. WherobotsAPIError carries the HTTP status code and request ID for debugging. WherobotsValidationError catches bad inputs at construction time. WherobotsTimeoutError fires when max_wait_seconds is exceeded. All three inherit from WherobotsJobError, so a single except block catches everything when you need it to. Run scripts from your S3 storage integrations If your script already lives in a Wherobots S3 Storage Integration, reference it directly by S3 URI and skip the upload step: job = WherobotsJob( script="s3://my-integration-bucket/scripts/pipeline.py", name="pipeline-job-001", runtime="small", auto_upload=False, ) Discover your integration paths programmatically: from wherobots.api.files import FilesAPI from wherobots.config import WherobotsConfig config = WherobotsConfig.from_env() with FilesAPI.from_config(config) as files_api: for si in files_api.list_integrations(): print(f"{si.name}: {si.path} ({si.region})") Design Principles The SDK is opinionated in four ways: requests is the only runtime dependency. No Pydantic, no boto3, no heavy frameworks. Install footprint stays small. Presigned uploads only. Local scripts go up through Wherobots API presigned URLs. No AWS credentials in user code, ever. Type-safe throughout. Typed dataclass models for every API response, JobStatus and Runtime enums, and a PEP 561 py.typed marker for downstream mypy users. Security-first defaults. HTTPS only (HTTP rejected at init). No redirect following, which prevents auth header leaks. API keys masked in repr(). Path traversal defense on uploads. POST requests are never retried. Getting Started pip install wherobots-python-sdk export WHEROBOTS_API_KEY="your-api-key" # Submit a job and watch it run python -c " from wherobots import WherobotsJob job = WherobotsJob(script='my_script.py', name='first-job', runtime='tiny') job.submit() job.wait_for_completion(stream_logs=True) " Source: github.com/wherobots/wherobots-python-sdk. Available now on PyPI: pip install wherobots-python-sdk. Start Building with Wherobots Get Started Key takeawaysThe Wherobots Python SDK is a typed Python client for submitting, monitoring, and managing Wherobots job runs. It ships on PyPI as wherobots-python-sdk. One install, one API key, and jobs run from CI/CD, notebooks, Step Functions, or a local shell.The only runtime dependency is requests. No AWS credentials and no bucket configuration are required in user code. Local scripts upload through Wherobots API presigned URLs. Scripts that already live in an S3 Storage Integration can be referenced by S3 URI with auto_upload=False.WherobotsJob is the single class today: point it at a script, give it a name and runtime, call submit(), then wait_for_completion(stream_logs=True). Status is COMPLETED, FAILED, or CANCELLED. Constructor validation raises WherobotsValidationError before any network call.Design constraints: HTTPS only (HTTP rejected at init), no redirect following, API keys masked in repr(), path-traversal defense on uploads, and POST requests are never retried. Typed dataclass models, JobStatus and Runtime enums, and a PEP 561 py.typed marker are included.
Spatial Data Processing Platforms: A Comparison of Enterprise and Cloud-Native Options Posted on May 1, 2026October 3, 2026 by Matt Forrest For Data Engineers and Architects Evaluating Spatial Workloads on Snowflake, Databricks, and PostGIS Six platforms dominate spatial data processing today: PostGIS for transactional workloads under 100GB, Snowflake and BigQuery GIS for light spatial enrichment inside a broader analytics platform, Databricks for vector spatial joins on the Lakehouse, Apache Sedona for self-managed open-source distributed spatial compute, and Wherobots for production raster and vector workflows at scale. The right choice depends on whether spatial is a side feature of your data work or the foundation of your product. If you’re processing billions of spatial records, running expensive zonal statistics, or optimizing complex spatial joins in Snowflake and watching bills climb while performance stalls, this comparison is for you. Spatial data processing has fundamentally changed. Traditional approaches (PostGIS, desktop GIS) do not scale. Modern data warehouses (Snowflake, BigQuery) added spatial as a feature, not a foundation. The cost-performance gap is wider than most teams realize Three Architectures for Spatial Data Processing Three architectural approaches dominate: Traditional Spatial Databases (PostGIS, SQL Server Spatial) General-Purpose Cloud Data Platforms (Snowflake, Databricks, BigQuery) Purpose-Built Spatial Compute (Wherobots, Apache Sedona) Each has a place. The right system depends on your workload, data volume, and whether spatial analysis is central to your business or a side feature. How to Choose a Spatial Data Processing Platform Use this table as a starting point. The detailed comparison for each platform follows below. PlatformBest ForAvoid ifPostGISDatasets under 100GBYou need distributed scaleSnowflakeSpatial as a minor enrichment step (<5% of workload)Spatial is core to your productDatabricksVector spatial joins on the LakehouseYou need production raster processingBigQuery GISSimple spatial queries on GCPYou need precision coordinate systemsApache SedonaFull control, on-premises deploymentYou lack dedicated Spark engineersWherobotsRaster and vector at scale, serverlessSpatial is incidental to your workload PostGIS (Traditional Spatial Database): Best for datasets under 100GB What it is: PostgreSQL extension for spatial operations. The industry standard for two decades. Best for: Teams already on PostgreSQL Transactional spatial applications (routing, geocoding services) Datasets under 100GB Organizations with strong Postgres DBA expertise Tradeoffs: Vertical scaling only. Distributing PostGIS across a cluster requires third-party tools (Citus, etc.). No native cloud-native format support. Reading GeoParquet or Cloud Optimized GeoTIFFs requires extensions or ETL. Raster processing is limited. PostGIS Raster exists but isn’t designed for large-scale raster analytics. Cost grows linearly. More performance requires bigger instances. A utility company processing 500GB of infrastructure data with frequent spatial joins will hit memory limits. You’ll end up partitioning manually, managing indexes carefully, and eventually looking for distributed alternatives. That is the ceiling. Most teams reach this limit sooner than they planned for. PostGIS is rock-solid for traditional GIS applications but wasn’t architected for cloud-scale distributed spatial analytics. PostGIS is not going anywhere, and it should not. Treating PostGIS as a scalable cloud analytics layer is the most common mistake teams make with it. Snowflake with Geospatial Support: Good for Casual Spatial Queries What it is: Cloud data warehouse with geometry/geography types and ~80 spatial functions. Best for: Organizations already standardized on Snowflake Spatial operations as 5-10% of total workload Point-in-polygon lookups, geocoding enrichment Teams prioritizing unified data platform over spatial performance Tradeoffs: Cost grows quickly on heavy spatial workloads. Spatial joins in Snowflake are expensive. A medium warehouse running complex polygon overlays or zonal statistics consumes thousands of credits. Limited spatial optimization. Snowflake’s architecture wasn’t designed for spatial partitioning or distributed spatial indexing. Weak raster support. No native raster data types or analysis functions. You’re on your own for satellite imagery or elevation analysis. Performance vs. specialized engines. Snowflake’s columnar architecture is not optimized for spatial predicates, so heavy spatial SQL workloads run materially slower than on purpose-built engines. A logistics company enriching 10M shipment records with census tract data (point-in-polygon join) will see reasonable performance. Daily overlay analysis on 100M parcels against zoning boundaries spirals in cost fast. The point-in-polygon job works. The moment you scale it, the bill shows up. Snowflake spatial works for casual spatial queries in a broader analytics platform. If spatial is core to your workload, you’re paying general-purpose pricing for a workload that benefits from purpose-built optimization. Databricks with Native Spatial SQL: Strong for Vector, Not Yet for Raster What it is: Native spatial support built into Databricks Runtime and SQL Serverless with GEOMETRY and GEOGRAPHY data types and 80+ spatial functions. Databricks Mosaic, the earlier open-source library, is deprecated. Native Spatial SQL is the current recommended approach. Best for: Teams already on Databricks Lakehouse Vector spatial processing at scale with serverless execution Organizations wanting distributed spatial joins without managing Spark clusters Data science workflows requiring spatial features alongside ML pipelines Key capabilities: Native GEOMETRY/GEOGRAPHY types with automatic bounding box statistics 80+ spatial SQL functions for constructing, transforming, measuring, and analyzing geometries Serverless execution available in Databricks SQL (no cluster management) Tradeoffs: Raster support is limited. Native Spatial SQL excels at vector processing, but production-grade raster analysis (satellite imagery, zonal statistics on elevation data, raster algebra) is not available in the native product today. Databricks is still gathering requirements for raster capabilities. Complexity for non-Databricks shops. If you’re not already invested in the Databricks ecosystem, onboarding requires understanding Delta Lake, Unity Catalog, and Lakehouse architecture. Cost for light spatial users. Like Snowflake, if spatial represents <5% of your workload, Databricks’ full platform might be overkill. A real estate analytics team processing 200M parcel boundaries against flood zones, zoning maps, and census tracts runs distributed spatial joins in Databricks SQL Serverless with automatic optimization. Performance is strong for vector-only workflows. Teams that need to overlay satellite imagery for vegetation analysis or run zonal statistics on elevation rasters wait for native raster support or build workarounds. Databricks native Spatial SQL is a major step forward for vector spatial processing on the Lakehouse. It is serverless and eliminates the operational complexity of managing Spark clusters. For workflows where raster analysis is central, it is not production-ready today. BigQuery GIS: A Fit for Simple Spatial Queries in GCP What it is: Google’s spatial extension for BigQuery with geography types and ~50 spatial functions. Best for: Organizations on Google Cloud Simple spatial queries at scale (geocoding, distance calculations) Integration with Earth Engine or Google Maps Platform Tradeoffs: Geography-only model (spherical geometry). No projected coordinate systems. All calculations happen on WGS84 sphere, problematic for local coordinate systems or precision work. Limited spatial indexing. BigQuery doesn’t support R-trees or spatial indexes like PostGIS. Performance depends on BigQuery’s columnar architecture, which isn’t optimized for spatial predicates. Weak raster support. Like Snowflake, no native raster data types. You can use Google EarthEngine connected to BigQuery for Zonal Stats on Raster data, but beyond this you need to customize heavily with bespoke services. A GCP-native team geocoding 50 million addresses or running distance calculations between delivery points will get the job done. The moment you need precision coordinate systems, complex polygon overlays at volume, or anything raster-related, you will start looking for alternatives. GCP teams running heavy raster or precision-coordinate work often pair BigQuery with a specialized engine for those workloads. BigQuery GIS handles simple spatial enrichment well within a GCP data warehouse. Heavy or precision spatial workloads usually move to a specialized engine. Apache Sedona Self-Managed: The Deepest Open-Source Spatial Engine What it is: Open-source distributed spatial engine built on Apache Spark. The foundation that powers Wherobots and formerly powered Databricks Mosaic. Best for: Teams with Spark expertise who want full control Organizations requiring on-premises deployment Custom spatial algorithm development Tradeoffs: You manage everything. Cluster provisioning, version upgrades, performance tuning, dependency management. No serverless execution. Cold starts, idle cluster costs, and manual scaling. Operational burden. Unless you have dedicated Spark platform engineers, operational overhead is real. Performance gap vs. managed offerings. Self-managed Sedona lacks the query optimizations (Photon-style vectorization, R-tree indexing tuned for spatial predicates) that managed platforms ship by default. Community support only: there is a great Apache Sedona community that we support and encourage grwoth, but if you want Enterprise support it’s best to come to Wherobots directly. Apache Sedona is the deepest open-source spatial engine available. Self-managing it makes sense only when you have the team and infrastructure already in place, and even then, managed alternatives deliver real performance advantages. Note that if you are a heavy Spark team but want full Enterprise support with advanced capabilities beyond Apache Sedona, but 100% API compatibility, we offer WherobotsDB Bring Your Own Spark (BYOS) as an option within our Enterprise tier at Wherobots. Note: Wherobots offers a drop in replacement for Apache Sedona that can run in any spark infrastructure, WherobotsDB “Bring Your Own Spark”. This option sits between the Managed Compute version of Wherobots Cloud (below) and Apache Sedona. For customers who want to bring their Apache Sedona pipelines to Enterprise Support, this option is a great way to get started with us. You can learn more about it here and explore all that it has to offer, and contact us if you would like access to it. Wherobots: Purpose-Built Spatial Compute for Raster and Vector at Scale What it is: The AI Context Engine for the Physical World. A fully managed, serverless spatial compute platform built by the original creators of Apache Sedona, the most widely deployed distributed spatial engine in the world. Why it exists: General-purpose platforms treat spatial as a feature, not a foundation. Wherobots architects every layer, from storage to compute to query optimization, specifically for spatial workflows across both vector and raster data. Best for: Heavy spatial ETL and analytics workloads combining vector, raster and tabular data Organizations processing hundreds of millions or billions of spatial records Teams running complex spatial joins, overlay analysis, and zonal statistics regularly Companies hitting cost or performance walls in Snowflake or need raster capabilities missing in Databricks Production-grade raster and vector processing at scale Key capabilities: Full vector and raster support. WherobotsDB and RasterFlow process satellite imagery, run zonal statistics, and perform raster algebra in the same query engine. 300+ spatial functions covering vector and raster data, with native Spark SQL for tabular operations, compared to roughly one raster function in BigQuery GIS. Production-ready today for both vector and raster workflows. Benchmark-validated performance. SpatialBench SF1000 results: 3x faster spatial query performance and 46% lower cost than general-purpose cloud data warehouses. Purpose-built spatial optimization. Distributed spatial indexing, spatial partitioning, and query optimization designed specifically for geometry and raster operations. Built by the Apache Sedona creators who understand spatial at the engine level. True serverless spatial compute. No cluster management, instant cold starts, scale-to-zero billing. You write Spatial SQL, Wherobots handles execution. Cloud-native format optimized. Reads GeoParquet, Cloud Optimized GeoTIFFs, Zarr, PMTiles natively. No ETL overhead. Apache Sedona compatibility. Wherobots is Apache Sedona, but managed, serverless, and optimized with performance enhancements. Use familiar APIs without operational overhead. Wherobots vs. Apache Sedona: What is the Difference? Apache Sedona is the open-source engine. Wherobots is Sedona managed, serverless, and performance-optimized. The core spatial APIs are identical. Wherobots adds the infrastructure layer: no cluster provisioning, instant scaling, and query optimizations the team built specifically for production spatial workflows. Teams already running Sedona lift and shift their workloads to Wherobots without rewriting code. A climate analytics company processing daily satellite imagery updates (50GB+ of COGs) with vector boundaries for wildfire risk zones. The team needs both raster analysis (vegetation indices, temperature anomalies) and vector operations (overlay analysis, zonal statistics). Wherobots handles both in a single serverless workflow. Most general-purpose platforms require custom raster solutions or have no native raster support. Snowflake has no raster capabilities. That is not a workaround. That is the workflow running as designed. Wherobots is the only serverless, purpose-built spatial compute platform with production-grade raster and vector support. If your workflows combine raster and vector and you are running them at volume, few platforms today handle both vector and raster in a single query engine without integration work. Decision Framework: Choosing Your Spatial Stack Use PostGIS if: You’re building a transactional spatial application Datasets are <100GB You need fine-grained control over indexes and query plans Use Snowflake/BigQuery if: Spatial is a minor enrichment step (<5% of workload) You’re already standardized on these platforms Queries are simple (geocoding, point-in-polygon lookups) Use Databricks if: You’re already on Databricks Lakehouse Vector spatial joins are your primary workload You want serverless spatial SQL with automatic optimization Raster processing is not required for your use cases Use Wherobots if: Raster and vector processing are both essential Spatial processing is core to your business (>20% of workload) You’re running complex spatial joins, overlay analysis, or zonal statistics You need production-grade raster analysis (satellite, climate, elevation) Cost or performance in current platforms is a pain point You want serverless spatial SQL optimized at every architectural layer How to Benchmark Spatial Performance Yourself Vendor performance claims are easy to make and hard to verify. SpatialBench is the open standard for benchmarking spatial query engines. It runs reproducible workloads (point-in-polygon, range queries, spatial joins, K-nearest neighbors) at multiple scale factors so you can compare engines on your own infrastructure. If you’re evaluating two or more platforms in this list, run SpatialBench on each before committing. Tell us what you’re building. We’ll show you what spatial processing at scale actually looks like for your workload. Talk to our team. Start Building with Wherobots Get Started Key takeawaysSix platforms dominate the comparison: PostGIS for transactional workloads under 100 GB; Snowflake and BigQuery GIS for light spatial enrichment inside a broader analytics platform; Databricks for vector spatial joins on the Lakehouse; self-managed Apache Sedona for full control; and Wherobots for production raster and vector workflows at scale.PostGIS is the two-decade industry standard but vertically scales only. A utility processing 500 GB of infrastructure data with frequent spatial joins will hit memory limits. Treating PostGIS as a scalable cloud analytics layer is described as the most common mistake teams make with it.Snowflake has ~80 spatial functions and weak raster support; Databricks Native Spatial SQL has 80+ functions and is serverless for vector, but production-grade raster is not available in the native product today (Mosaic is deprecated). BigQuery GIS is geography-only on a WGS84 sphere (~50 functions) with no R-tree indexes and weak raster.WherobotsDB plus RasterFlow process satellite imagery, zonal statistics, and raster algebra in the same query engine, with 300+ spatial functions. SpatialBench SF1000 results cited: 3x faster spatial query performance and 46% lower cost than general-purpose cloud data warehouses. Wherobots is Apache Sedona, managed, serverless, and optimized; Sedona APIs lift and shift without rewriting code. BYOS is offered in Enterprise.
Spatial Data Pipeline Architecture: PostGIS and Wherobots Together Posted on April 21, 2026October 3, 2026 by Matt Forrest In the world of data architecture, there is a dangerous myth that you have to choose “one tool to rule them all.” We often see organizations paralyzed by the debate: “Should we use a Database or a Data Lake?” A spatial data pipeline architecture built for both large-scale analytics and operational queries is one of the harder infrastructure decisions a geospatial team makes. Most organizations try to solve it with a single tool. That is where the problems start. The teams that get this right do not choose between a data lake and a spatial database. They use both, in a defined sequence known as the Geospatial Medallion Architecture. The medallion architecture, a data pipeline pattern established in the data lakehouse ecosystem, organizes data into three progressive quality layers: Bronze, Silver, and Gold. Wherobots and PostGIS each own a distinct role in that pipeline. Wherobots is a cloud-native spatial analytics platform, built by the original creators of Apache Sedona, that processes large-scale geospatial datasets using distributed compute. PostGIS is an open-source spatial extension for PostgreSQL that adds support for geographic objects and enables location-based queries in SQL. The spatial medallion architecture organizes geospatial data into three layers: Bronze: Raw ingestion. All incoming data (GPS feeds, satellite imagery, shapefiles) lands in cloud object storage (S3, Azure Blob, GCS) without transformation. The goal is preservation and speed of capture. Silver: Refinement. A distributed compute platform like Wherobots validates geometries, enriches records with spatial attributes, and converts raw formats to optimized columnar formats like GeoParquet. This layer is what data scientists use to train spatial AI models. Gold: Delivery. Aggregated, query-ready data served by PostGIS to BI tools and applications. Volume is low because the data is pre-aggregated. Response times are sub-second. Think of your data not as static files, but as raw material (like crude oil or iron ore) that must be refined through a series of stages before it is valuable to your business. Why Spatial Data Teams End Up with Swamps or Silos Before we look at the solution, let’s look at the problem. The “Data Swamp”: You dump all your raw files (CSV, Shapefiles, GeoJSON) into cloud storage. It’s cheap, but nobody can find anything. It’s a mess. The “Database Silo”: You try to force everything into your operational database (PostGIS). The data is clean, but the database becomes slow, and expensive. Your analysts crash the system with heavy queries, and your app users complain about speed. The Spatial Data Pipeline Medallion Architecture The Medallion Architecture solves this by organizing your data into three distinct layers of quality: Bronze, Silver, and Gold. This approach allows you to use the right tool for the right job: Wherobots for heavy industrial refining, and PostGIS for precision delivery. Wherobots is a cloud-native spatial analytics platform, built by the original creators of Apache Sedona, that processes large-scale geospatial datasets using distributed compute. PostGIS is an open-source spatial extension for PostgreSQL that adds support for geographic objects and enables location-based queries in SQL. 1. Bronze Layer: Raw Ingestion into Cloud Object Storage The “Landing Zone” In the spatial medallion architecture, the Bronze layer is the raw ingestion zone where all incoming data lands without transformation. This is the entry point for all your data. Whether its real-time telemetry from 10,000 delivery trucks, daily dumps of satellite imagery, or messy spreadsheets from a partner, it all lands here first. The Goal: Speed and Preservation. We don’t change the data here; we just capture it so we never lose the original source. The Technology: Cloud Object Storage (S3, Azure Blob, GCS). Why: It is incredibly cheap and infinite. You can dump petabytes here without breaking the bank. 2. The Silver Layer: The Refinery Clean, Standardize, & Enrich This is where the magic happens and where the heavy lifting is required. Raw data is rarely ready for business. It has duplicates, missing fields, or invalid geometries (like a building polygon that twists into itself). The Task: Validation: Checking if GPS points are actually on land. Enrichment: Taking a raw coordinate and adding “City,” “State,” and “Flood Risk Score” columns. Optimization: Converting raw formats like CSV and Shapefile into columnar formats like GeoParquet or Apache Iceberg tables, which Wherobots stores natively through its Havasu spatial lakehouse catalog. The Technology: Wherobots. Why: This “refining” process requires massive computing power. You might be processing billions of records. Wherobots can spin up 1,000 nodes to crunch this data in minutes, clean it, and write it back to the lake as a trusted, high-quality “Silver” dataset. This is the layer your Data Scientists love because it’s clean enough to train AI models but detailed enough to find deep patterns. AddressCloud runs property-level perils models for insurers, processing flood, fire, and climate risk data across millions of addresses. John Powell, Senior Geospatial Data Engineer at AddressCloud, describes what changed when they moved that workload into the Silver layer: “From a developer perspective, having data, algorithms and compute (and to be presented with a Spark/Sedona context in a Jupiter notebook on startup) combined in one platform is extremely powerful, comparable in many respects to Google Earth Engine, but with much greater guarantees of, and control over, job completion.” The result: operations that previously took hours or days now complete in minutes, with no preprocessing step required to combine raster and vector data. 3. Gold Layer: Optimized Data for BI Tools and Applications Aggregated & Ready for Business This is the “Showroom” layer. This data is highly polished, aggregated, and formatted for specific business questions. For example: “Total Sales by Zip Code” or “Active Drivers by City.” The Task: Serving answers to users instantly. When a CEO opens a dashboard or a customer opens an app, they can’t wait 10 seconds for a query to run. They need sub-second speed. The Technology: PostGIS. Why: By the time data reaches the “Gold” layer, the volume is much smaller because we have aggregated it. PostGIS is the perfect tool here. It is optimized for high-speed retrieval. It serves this “Gold” data to your BI tools (Tableau, Looker) and your web applications instantly. How the Three Layers Work Together: The “Better Together” Workflow By adopting this strategy, you create a data supply chain that maximizes the strengths of every tool: Lower Costs: You stop using your expensive database to store terabytes of raw, messy junk. That stays in the cheap Bronze layer. Higher Stability: Your heavy analytical jobs run in Wherobots (Silver Layer), completely separate from your operational database. Your analysts can crunch numbers all day without ever slowing down the app for your customers. Faster Innovation: Your data scientists don’t have to beg for database access. They can work directly with the Silver data in Wherobots to build advanced AI models, while the business teams continue to use the Gold data in PostGIS. Summary: A Blueprint for Success The spatial medallion architecture is not a tool choice. It is a pipeline pattern that assigns the right tool to the right job. If you are a leader looking to modernize your geospatial stack, don’t look for a “PostGIS replacement.” Look for a partner. Use Wherobots to act as your heavy-lifting factory: ingesting, cleaning, and crunching massive scale data. Use PostGIS to act as your high-speed storefront: delivering those insights to users the moment they need them. This hybrid approach, the Spatial Medallion Architecture, is how modern organizations turn location data into competitive advantage.This is part three of a series. The prior posts cover PostGIS vs Wherobots for spatial data lakehouses and spatial database cost comparisons. Start Building with Wherobots Get Started Key takeawaysThe post argues you do not choose one tool to rule them all. A spatial pipeline built for both large-scale analytics and operational queries uses both Wherobots and PostGIS in a defined sequence: the Geospatial Medallion Architecture (Bronze, Silver, Gold).Bronze is raw ingestion into cheap object storage (S3, Azure Blob, GCS) with no transformation. Silver is the refinery: Wherobots validates geometries, enriches records, and writes GeoParquet or Iceberg via Havasu. Gold is aggregated, query-ready data served by PostGIS to BI tools and apps with sub-second response times.AddressCloud's John Powell is quoted on moving property-level perils models (flood, fire, climate risk across millions of addresses) into the Silver layer: data, algorithms, and compute in one Spark/Sedona notebook, comparable to Google Earth Engine but with greater control over job completion. Operations that previously took hours or days now complete in minutes, with no preprocessing step to combine raster and vector.The payoff of the split: raw junk stays in cheap Bronze instead of an expensive database; heavy jobs run in Wherobots so they cannot slow customer-facing PostGIS; data scientists work on Silver without begging for database access. This is part three of a series after the lakehouse guide and cost comparison.
Iceberg v3 Gets Native Geo Types. It’s More Than a Format Upgrade Posted on April 15, 2026October 3, 2026 by Pranav Toggi Introduction Geospatial data touches nearly every industry, and until recently, the open lakehouse had no native way to handle it. Snowflake recently announced Iceberg v3 support with native geometry and geography types. It’s the first major engine to ship the geospatial extensions to the Iceberg spec. These types are now part of the open standard, available to every engine in the ecosystem. In their Iceberg v3 post, the Snowflake team called out where the geo work came from: “A special mention to the entire Wherobots team, which implemented geospatial support on its own fork of Iceberg before offering its expertise to the Iceberg community, providing leadership and implementing the feature for the Iceberg project.” The work started in 2022, the year Wherobots was founded. This post is the story behind it: how production experience became an open standard, and what it means for the teams and tools building on spatial data. How Geospatial Data Worked in Iceberg Before v3 Until Iceberg v3, geospatial columns did not exist as a concept in the table format. Engineers stored geometry as opaque binary blobs: Well-Known Binary (WKB) bytes in a binary column. The Iceberg catalog had no way to know the column contained spatial data. The practical consequences: No CRS metadata. The coordinate reference system did not live in the column. Engineers tracked it in documentation or external conventions. No bounding-box statistics. A query engine had no way to skip files based on spatial predicates without reading the geometry itself. No cross-engine portability. Every team needing spatial analytics on a lakehouse built custom encoding pipelines, engine-specific readers, and schema conventions living outside the format. It worked, but it was fragile. And it did not travel across engines. Havasu: Production System First Havasu was our answer to that problem: a spatial lakehouse extension built on an Iceberg fork. Not a proof of concept. A production system, running real customer workloads since 2022, where we could pressure-test the design decisions that would eventually become a standard: Spatial indexing integrated into table metadata Bounding-box statistics at the manifest level, enabling file-group pruning on spatial predicates CRS propagation through the schema, so coordinate reference systems traveled with the data Geometry column encoding with explicit format annotations (WKB, EWKB, WKT, GeoJSON) That production experience shaped our GeoLake research, which formalized the core design principles: unambiguous CRS representation, efficient geometry encoding in columnar storage, and bounding-box statistics that let spatial predicates be evaluated before any geometry data is read. Then we contributed the design, the implementation experience, and the lessons learned upstream to Apache Parquet and Apache Iceberg, so the broader ecosystem could build on it. Many of the design decisions validated in Havasu are now part of the Iceberg v3 spec. Bounding-box statistics, CRS propagation, and geometry encoding all made the transition from production system to open standard. Before and After: Geometry in the Iceberg Schema To understand why this matters, consider the difference concretely. Before (Iceberg v1/v2): A geometry column in the Iceberg schema looks like this: { "id": 5, "name": "geom", "type": "binary" } The catalog sees raw bytes. No spatial semantics. An engine reading this table has no way to know this column contains geometries, what CRS they’re in, or how to prune files spatially, unless it relies on out-of-band conventions like GeoParquet file-level metadata. After (Iceberg v3): The same column becomes: { "id": 5, "name": "geom", "type": "geometry(srid:4326)" } Now the CRS is a property of the type. The Parquet files carry GEOMETRY logical type annotations that any conforming engine can recognize. Bounding-box statistics in the manifest use a compact coordinate encoding, so engines can skip entire file groups that fall outside a query’s spatial bounds, before reading a single geometry value. Spatial pruning moves to the format level. This is the architecture we designed and validated in Havasu, now available to any engine that reads Iceberg v3. Making It Native: Apache Parquet and Iceberg Making geospatial a first-class type in the open lakehouse required coordinated changes to two foundational Apache projects. On the Parquet side, the PR to add GEOMETRY and GEOGRAPHY logical types to the format spec drew over 400 comments across months of design review: encoding formats, CRS semantics, edge-interpolation behavior, edge-case handling. Jia Yu (Wherobots Co-Founder and Apache Sedona PMC Chair) and Kristin Cowalcijk (Apache Sedona PMC member) drove core design decisions from the start. On the Iceberg side, the geo type spec defined how Iceberg catalogs spatial columns, stores bounding-box metadata, and handles spatial partitioning, with another 240+ comments of cross-community design work. The core API and implementation, authored by Kristin, followed with bounding-box types, geospatial predicates, and Parquet geo read/write. Over a year of coordinated work across both Apache communities. The result: geometry and geography as native primitive types, from the storage format through the table format – with the same level of spec support that timestamps, decimals, and every other type have always had. The First Wave of Adoption Snowflake’s announcement is a milestone, not the finish line. It’s the first major engine to ship v3 geo types. It won’t be the last. Iceberg v3 adoption is accelerating broadly. AWS Glue shipped v3 support at re:Invent 2025. Dremio followed with GA support in their cloud platform. As more engines adopt v3, spatial columns will interoperate the same way every other column type already does. The infrastructure barrier that kept geospatial data siloed from mainstream analytics is coming down. The end-to-end Parquet geo read/write path for Iceberg is under active review in the Apache Iceberg project. It connects v3 schema types to properly annotated Parquet files with GEOMETRY logical types. Once merged, any Iceberg-compatible engine will produce and consume fully v3-native geo tables. What This Means for Spatial Data Interoperability Apache Iceberg is increasingly how organizations query data across engines, and native geo types mean spatial data now travels the same way every other column type does. A geo-typed table written from one engine is readable by another with correct CRS, bounding-box stats, and spatial pruning intact. No proprietary connectors. No out-of-band conventions. Cross-engine portability also makes spatial data accessible to AI. Tools and agents working with physical-world data need consistent access across engines and catalogs. When the type contract lives in the format, access does not require per-system special-casing. Ahead of the Standard Wherobots Cloud is where teams run spatial ETL, large-scale analytics, and geospatial data engineering on Iceberg today, fully managed. As the upstream v3 pipeline completes, Wherobots Cloud will be the first to produce fully v3-native geo tables. The natural result of being the team that designed the spec and has the longest production track record on it. Apache Sedona reads and writes Iceberg tables with geospatial columns today, using the Havasu encoding that preceded the v3 spec. This means Sedona users have had production-grade spatial lakehouse capabilities since before the standard was finalized, and the migration path to v3-native types is straightforward as upstream support matures. WherobotsDB and Sedona also go beyond what the spec requires: distributed CRS-aware computation, automatic transformation across datasets in different projections, and support for CRS formats beyond SRID integers (WKT, PROJ strings, grid-based datum shifts). Most engines that adopt v3 will read the CRS tag. WherobotsDB and Sedona compute with it at scale. The end-to-end Parquet geo read/write path for Iceberg is under active review in the Apache Iceberg project. It’s the final piece connecting v3 schema types to properly annotated Parquet files with GEOMETRY logical types. Once merged, any Iceberg-compatible engine will be able to produce and consume fully v3-native geo tables. Where Wherobots Goes Beyond the v3 Spec The v3 geo types are the foundation. The next layer is making the full pipeline seamless: spatial partitioning strategies, advanced spatial indexing beyond bounding boxes, and tighter integration between spatial predicates and query planning across engines. These are problems we’ve been working on in Havasu for years. Now that the standard exists, that work can happen in the open, and the entire ecosystem benefits. The Spatial AI Coding Assistant, including the Wherobots MCP Server, VS Code Extension, and CLI, lets developers and AI agents work with spatial data through natural language. A developer describes an analysis problem, and the tools find relevant datasets, generate spatial queries, and execute them on Wherobots Cloud. Native geo types in Iceberg make this work cleanly: when the type, CRS, and spatial statistics live in the format, an agent queries physical-world data the same way it queries any other column, across any engine reading Iceberg. If you’re building your spatial data architecture on Iceberg, we’d like to talk. Key takeawaysUntil Iceberg v3, geospatial columns did not exist in the table format. Engineers stored geometry as WKB in a binary column, with no CRS in the schema, no bounding-box statistics, and no cross-engine portability. v3 adds native geometry and geography types; Snowflake was the first major engine to ship them.Snowflake's Iceberg v3 post called out the Wherobots team for implementing geospatial support on its own Iceberg fork before contributing the feature to the Iceberg project. That work started in 2022, the year Wherobots was founded, in Havasu: a production spatial lakehouse extension, not a proof of concept.The Parquet PR for GEOMETRY and GEOGRAPHY logical types drew over 400 comments. The Iceberg geo type spec had another 240+ comments. Jia Yu and Kristin Cowalcijk drove core design. After v3, a column is geometry(srid:4326) rather than binary, and manifest bbox stats can skip file groups before any geometry is read.AWS Glue shipped Iceberg v3 support at re:Invent 2025; Dremio followed with GA in their cloud platform. Apache Sedona already reads and writes Iceberg geospatial columns using the Havasu encoding that preceded the spec. WherobotsDB and Sedona also compute with CRS at scale (WKT, PROJ strings, grid-based datum shifts), beyond reading an SRID tag.
Mobility Data Processing at Scale: Why Traditional Spatial Systems Break Down Posted on March 26, 2026October 3, 2026 by Matt Forrest Mobility data is the continuous stream of GPS location records capturing how people, vehicles, and assets move through the world. Processing it at scale is fundamentally different from standard spatial data work because it carries temporal dependencies: the order and timing of observations define movement, not just position. Most organizations have discovered, often painfully, that the tools they already use were not designed with that in mind. In Part 2, we walk through the technical implementation: a three-notebook medallion architecture built on Wherobots and Apache Sedona that takes raw GPS pings and transforms them into analysis-ready, GeoParquet-backed analytical views. Why Mobility Data Is Harder to Process Than It Looks Every second, billions of GPS-equipped devices generate spatiotemporal data capturing how people, goods, and vehicles move through the physical world. The market reflects how seriously organizations are treating this: the global fleet management market is valued at approximately $27 billion in 2025 and projected to exceed $122 billion by 2035. Mobility data analytics platforms are on a similar trajectory, from $2.5 billion to over $11 billion by 2034. But collecting this data and actually extracting reliable intelligence from it are two very different things. Most organizations have discovered, often painfully, that the tools they already use were not designed with GPS mobility data in mind. The result is brittle pipelines, inconsistent methodologies, and analyses that quietly produce misleading conclusions. Researchers at ACM have characterized the field as requiring its own dedicated science because general-purpose data science pipelines consistently produce suboptimal results when applied to movement data. Understanding why requires looking at the specific properties that make mobility data uniquely difficult to process correctly. What Makes Mobility Data Different From Standard Spatial Data Mobility data is not just spatial data with timestamps attached. It is a distinct category of data that violates assumptions built into most data processing systems. Researchers at ACM have characterized the field as requiring its own dedicated science—Mobility Data Science—because general-purpose data science pipelines consistently produce suboptimal results when applied to movement data. Understanding why requires examining the specific properties that make trajectory data uniquely challenging. Why GPS Data Volume Overwhelms Traditional Databases The volume and velocity problem is the recognition that GPS data generation at fleet scale is not a batch analytics challenge, it is a continuous, high-throughput data engineering problem that demands distributed processing from the start. A single connected vehicle generating GPS pings at one-second intervals produces roughly 86,400 records per day. A fleet of 10,000 vehicles generates over 860 million data points daily. Multiply this across the millions of connected vehicles, delivery drones, rideshare fleets, and maritime vessels operating globally, and the scale becomes staggering. Traditional spatial databases like PostGIS, which excel at transactional workloads and moderate-scale analytics, were not designed for this volume. Loading hundreds of millions of GPS points into PostgreSQL, constructing geometries, and running spatial joins or trajectory reconstruction queries can take hours or days on a single node. Adding more hardware does not solve the fundamental problem: PostGIS was not built for distributed, parallel spatial computation. How GPS Signal Noise Corrupts Downstream Analysis GPS noise is error in raw location readings caused by signal reflection, satellite loss in tunnels and urban canyons, and atmospheric interference and it cascades through every downstream analysis built on top of it. Signals bounce off buildings, lose satellites in tunnels and urban canyons, and suffer from atmospheric interference. Studies have documented median GPS errors of 7 meters with standard deviations exceeding 23 meters in urban environments. Points can appear on the wrong side of a street, inside buildings, or kilometers from the actual position when signal quality degrades. Speed calculations between consecutive points can spike to physically impossible values. Distance measurements accumulate systematic errors. Clustering algorithms identify phantom stop locations. Without rigorous cleaning and validation at the earliest stages of the pipeline, every subsequent insight is built on a compromised foundation. Researchers at the University of Pennsylvania’s Computational Social Science Lab studied exactly this problem in the context of COVID-19 epidemic modeling. Using the same GPS mobility dataset, they found that different but individually reasonable preprocessing choices led to substantially different conclusions. a methodological “garden of forking paths” where reproducibility became nearly impossible. The root causes: data sparsity, sampling bias, and inconsistent algorithmic choices at the preprocessing stage. Why Temporal Ordering Is Critical in Mobility Data Processing Trip segmentation is the process of splitting a continuous GPS stream into discrete trips by detecting temporal gaps, periods where no data was recorded or the device was stationary. Mobility data is not simply geospatial—it is spatiotemporal. Every GPS point has a position and a timestamp, and the relationship between consecutive observations is what defines movement. This trip segmentation step alone introduces significant methodological complexity, because the threshold you choose (5 minutes? 20 minutes? 60 minutes?) fundamentally changes the structure of your resulting trajectories and all metrics derived from them. Beyond segmentation, ordering matters. Trajectories are sequences, not sets. Every analytical operation—speed calculation, direction changes, stop detection, map matching—depends on correct chronological ordering within each trip. Systems that do not preserve or guarantee ordering (a common challenge in distributed frameworks) can produce geometries that appear valid but contain scrambled temporal information. Why 2D Spatial Systems Fail for Mobility Analytics The dimensionality problem is the gap between how most spatial systems model location, latitude and longitude only and what mobility analysis increasingly requires: 3D and 4D geometry that encodes elevation and time directly into the geometry itself. Most spatial systems treat location as a 2D construct: latitude and longitude. But mobility analysis increasingly demands 3D and 4D processing. Elevation matters for fuel consumption modeling, route optimization in mountainous terrain, aviation and drone trajectories, and any analysis where the difference between 2D and 3D distance is materially significant. Adding a temporal measure dimension (the “M” in XYZM geometries) enables encoding timestamps directly into the geometry itself, supporting trajectory validation and interpolation operations that are impossible with 2D points. Yet 4D geometry support—constructing XYZM points, building trajectories from them, and performing analytical operations that respect all four dimensions—is rare. Most spatial SQL implementations either lack these functions entirely or implement them inconsistently. This forces practitioners to maintain separate columns for elevation and time, losing the computational advantages of integrated 4D geometry processing. Where Traditional Spatial Systems Fail for Mobility Data SystemPrimary Limitation for Mobility DataDesktop GIS (QGIS, ArcGIS Pro)Single-machine ceiling, no distributed processingPostGISSingle-node, not built for hundreds of millions of GPS pointsCloud data warehouses (Snowflake, BigQuery)Shallow spatial support, cannot handle XYZM or map matchingVanilla Apache SparkNo native spatial types, no spatial indexingExternal map matching APIsRate limits and per-request pricing make batch processing prohibitive Desktop GIS: The Single-Machine Ceiling Tools like QGIS and ArcGIS Pro are extraordinarily capable for visualization, manual analysis, and working with datasets that fit in memory. But they hit a hard wall with mobility data at scale. Loading millions of GPS trajectories, performing trip segmentation with window functions, running DBSCAN clustering on stop points, and executing map matching against a road network are not operations that desktop GIS was designed to handle. Analysts working with fleet-scale data routinely encounter out-of-memory errors, multi-hour processing times, and the inability to iterate quickly on analytical parameters. Spatial Databases: Scale Without Spatial Intelligence PostGIS remains the gold standard for spatial SQL and is an excellent choice for many use cases. However, it is fundamentally a single-node system. Scaling PostGIS to handle hundreds of millions of GPS points requires expensive vertical scaling, and even then, operations like trajectory construction across thousands of users, spatial indexing with H3 or GeoHash, and DBSCAN clustering at urban scale can exhaust available resources. Cloud data warehouses like Snowflake, BigQuery, and Redshift have added spatial capabilities, but these tend to be shallow implementations optimized for simple point-in-polygon or distance queries. Constructing XYZM trajectories from ordered GPS points, performing spatial clustering, computing 3D distances, or running map matching against a road network are either unsupported or require convoluted workarounds that sacrifice performance and maintainability. General-Purpose Distributed Frameworks: Power Without Spatial Awareness Apache Spark provides the distributed computing muscle needed for mobility-scale data, but vanilla Spark has no concept of spatial data types, spatial indexing, or geometric operations. Running a spatial join in pure Spark requires broadcasting datasets or implementing custom partitioning strategies—both of which are error-prone and perform poorly at scale compared to purpose-built spatial engines. This is precisely the gap that Apache Sedona and Wherobots were designed to fill. Sedona extends Spark (and other distributed frameworks) with native spatial data types, over 290 spatial SQL functions, spatial indexing, and optimized query planning that understands geometric predicates. Wherobots builds on Sedona to provide a fully managed, cloud-native spatial intelligence platform where teams can process GPS-scale data without managing infrastructure, configuring clusters, or bolting together fragmented toolchains. Map Matching: The Unsolved Infrastructure Problem Map matching—the process of snapping noisy GPS traces to the actual road network—is one of the most computationally demanding and methodologically complex steps in any mobility pipeline. It requires loading a complete road network graph, computing probabilistic alignments between GPS observations and candidate road segments, and resolving ambiguities at intersections, parallel roads, and complex interchanges. Most map matching solutions are either commercial APIs with strict rate limits and per-request pricing that make batch processing of historical data prohibitively expensive, or open-source tools that require significant infrastructure setup and do not scale to city-wide or fleet-wide datasets. Researchers have consistently identified scalability as the primary bottleneck: algorithms that produce accurate matches on small datasets fail to perform when confronted with millions of trajectories. Having map matching available as an integrated capability within the same distributed environment where you are already processing and analyzing your trajectory data—rather than as an external API call or a separate system—eliminates an entire category of infrastructure complexity and data movement overhead. The Real Cost of a Fragmented Mobility Data Pipeline In practice, most organizations processing mobility data have assembled a patchwork of tools: Python scripts for data cleaning, PostGIS for spatial operations, custom code for trip segmentation, an external API for map matching, a separate clustering library, and a visualization tool at the end. Each transition between tools introduces data serialization overhead, potential for schema drift, and opportunities for subtle bugs. This fragmentation carries real costs beyond engineering time. When a researcher needs to change the trip segmentation threshold from 20 minutes to 30 minutes, the entire pipeline must be re-executed across multiple systems. When a new data source arrives with a slightly different schema, adapters must be updated at each integration point. When results need to be reproduced for regulatory or academic review, reconstructing the exact sequence of operations across disparate tools is often impractical. The ideal mobility data pipeline processes GPS pings through ingestion, cleaning, enrichment, trajectory construction, map matching, spatial indexing, clustering, and analytical aggregation—all within a single, distributed, spatially-aware environment where every step is expressed in SQL or Python, every intermediate result is inspectable, and the entire pipeline can be reproduced with a single execution. What a Modern Mobility Data Architecture Looks Like The medallion architecture—Bronze, Silver, Gold—has become the standard pattern for progressive data refinement in the data lakehouse world. But applying it to mobility data requires rethinking what each layer does, because spatial data introduces transformations and enrichment steps that have no analog in conventional data engineering. Bronze is not just ingestion—it is spatial profiling. You are not only loading CSV or Parquet files; you are constructing point geometries, validating coordinate bounds, assessing data quality metrics like altitude validity and temporal coverage, and establishing the spatial extent of your dataset. Silver is where the heavy lifting happens. This is trip segmentation, 4D geometry construction, trajectory building, movement metric derivation, spatial indexing, and map matching. Each of these operations is computationally intensive, order-dependent, and requires spatial functions that most data platforms simply do not have. Gold produces the analytical views that power downstream consumption: H3 hexbin density heatmaps, temporal activity patterns, stop detection via spatial clustering, trajectory anomaly flagging, and road segment speed analysis. These views are written as GeoParquet files—compatible with Kepler.gl, QGIS, Felt, Foursquare Studio, and DuckDB Spatial—ensuring that the output of the pipeline is immediately consumable by any modern geospatial visualization or analytics tool. How Wherobots Handles Mobility Data Processing: What We Built To demonstrate this architecture in practice, we built a three-notebook Wherobots Mobility Solution Accelerator using the Microsoft Research GeoLife GPS Trajectories dataset—one of the few open mobility datasets that includes elevation data, enabling full 4D geometry processing. The dataset contains 17,621 trajectories from 182 users in Beijing, with latitude, longitude, altitude, and timestamps spanning 2007–2012. In Part 2, we walk through every notebook in detail: the Bronze layer’s ingestion and profiling pipeline, the Silver layer’s 4D trajectory construction and map matching workflow, and the Gold layer’s analytical and exploratory views. We cover the specific Apache Sedona spatial SQL functions used at each step, the PySpark window function patterns for trip segmentation and movement metrics, and the real-world challenges we encountered and solved—from Spark’s schema inference corrupting timestamp values, to COLLECT_LIST not preserving order in trajectory construction, to DBSCAN requiring physical column references. If you are building mobility analytics pipelines and hitting the limitations of your current toolchain, Part 2 will give you a concrete, reproducible blueprint for how to do it on Wherobots. Get Started with Wherobots Access Now Key takeawaysMobility data is not spatial data with timestamps attached. Order and timing define movement. ACM researchers characterize the field as requiring its own science because general-purpose pipelines consistently produce suboptimal results on movement data.Volume: a vehicle pinging at one-second intervals produces roughly 86,400 records per day; a 10,000-vehicle fleet generates over 860 million points daily. Median GPS error of 7 meters with standard deviations exceeding 23 meters has been documented in urban environments. Different but reasonable preprocessing choices on the same GPS dataset led to substantially different COVID-19 modeling conclusions (UPenn CSSLab).Most spatial systems are 2D. Elevation and XYZM (time as M) matter for fuel modeling, mountainous routing, aviation/drones, and trajectory validation, but 4D support is rare. Trip-gap thresholds (5 vs 20 vs 60 minutes) change trajectory structure and every derived metric. Distributed collect operations can scramble M-ordering.Desktop GIS hits a single-machine ceiling; PostGIS is single-node; Snowflake/BigQuery spatial support is described as shallow for XYZM and map matching; vanilla Spark has no native spatial types. Map matching as an external API hits rate limits and per-request pricing. Part 2 implements a three-notebook medallion pipeline on Wherobots/Sedona using GeoLife (17,621 trajectories, 182 users, Beijing 2007–2012).
PostGIS vs Wherobots: What It Actually Costs You to Choose Wrong Posted on March 19, 2026October 3, 2026 by Matt Forrest When building a geospatial platform, technical decisions are never just technical, they are financial. Choosing the wrong architecture for your spatial data doesn’t just frustrate your data team; it directly impacts your bottom line through large cloud infrastructure bills and, perhaps more dangerously, delayed business insights. For decision-makers, the choice between a traditional spatial database (like PostGIS, an open-source extension of PostgreSQL that adds support for storing and querying location data) and a cloud-native geospatial analytics platform (like Wherobots, built with distributed computing including Apache Spark to process massive spatial datasets in parallel across compute clusters) comes down to two fundamental metrics: Time to Insight and Total Cost of Ownership (TCO). To understand where to invest your budget, we need to look beyond the software labels and understand the economics of how these systems handle data. If you are new to the PostGIS vs Wherobots discussion, start with [Part 1: PostGIS, Wherobots, and the Spatial Data Lakehouse: A Strategic Guide for Leaders]. This post assumes you understand the architectural difference and focuses on what it actually costs you to choose wrong. Key Takeaways: PostGIS is optimized for low-latency lookups. Wherobots is optimized for high-throughput analytics and data processing. Using the wrong one for the wrong job costs you both time and money. A PostGIS server must be provisioned for peak load, so you pay for maximum capacity even when usage is low. Wherobots is designed to charge primarily for active compute time, so you are not paying for idle capacity. For industries like insurance, logistics, and urban planning, the right architecture choice can dramatically reduce query time for large-scale spatial analysis — in some cases from hours or days down to minutes. PostGIS and Wherobots are not mutually exclusive. Many enterprises use Wherobots to process data at scale, then serve results through PostGIS for live application access. Use the checklist in this post as a fast diagnostic: if you are waiting hours for spatial queries or your cloud bill is outpacing your revenue growth, you have a strong case for cloud-native spatial compute. Why Slow Spatial Queries Cost More Than You Think In the modern enterprise, the value of data decays over time. An answer delivered in 5 seconds is actionable; an answer delivered in 5 days is a post-mortem. The architecture you choose dictates how fast you can answer complex questions. Low Latency vs High Throughput: What Speed Actually Means for Each Tool It is crucial to understand the difference between “speed” for an app and “speed” for analytics. Low Latency (PostGIS): This is the speed of retrieval. When a customer opens your delivery app and asks, “Where is my driver?”, they need an answer in milliseconds. PostGIS is optimized for this. It uses heavy indexing to find a single “needle in a haystack” instantly. High Throughput (Wherobots): This is the speed of processing. When your risk analyst asks, “Which of our 50,000 retail locations are at risk of flooding based on the new 100-year climate models?”, they are not looking for a needle; they are looking for patterns across the whole haystack. The Bottleneck: If you try to run that massive climate model analysis in PostGIS, the database has to check every single location against every single flood zone, even with optimizations for search like spatial indexing. Because PostGIS typically runs on a single server, running that analysis forces your app and your analytics to compete for the same resources. The query might take 24 hours, and your customer-facing app slows down the whole time. The Solution: Wherobots breaks that same job into parallel tasks across a cluster of worker nodes, each handling a partitioned geographic slice. Because the work happens simultaneously, the job finishes in minutes instead of hours. Business Impact: Your analysts get answers before lunch, not next week. Your operational apps remain fast for customers because the heavy lifting happened elsewhere. PostGIS vs Wherobots: Why the Pricing Models Are Fundamentally Different The second major factor is how you pay for these capabilities. Query speed is only half the cost story. The other half is how each system charges you. The pricing models for databases and cloud-native engines are fundamentally different. Why PostGIS Forces You to Pay for Peak Capacity Around the Clock A high-performance database like PostGIS requires expensive hardware specifically, high-speed RAM and fast CPUs, to keep your data accessible. The Lease Model: PostGIS requires provisioning a server for peak load, which means you pay for maximum capacity even when usage is low. Think of it like leasing a Ferrari just to drive to the grocery store on Sundays. You have to provision this server for your peak usage. If you need to run a heavy report once a week that requires 64 cores of CPU, you must pay for a 64-core server 24 hours a day, 7 days a week. The Waste: For the other 6 days and 23 hours, that expensive server sits idle, burning budget. It is analogous to leasing a Ferrari just to drive to the grocery store on Sundays. How Wherobots Charges Only for Active Compute Time Wherobots uses an elastic, on-demand pricing model: you consume compute while an operation is actively running, then billing stops when it finishes. Storage and compute are decoupled, meaning your data sits in low-cost object storage and you rent processing power only when you need it. Storage is Cheap: You keep your massive datasets in low-cost Object Storage (like Amazon S3), which costs pennies per gigabyte. Compute is On-Demand: When you need to run that heavy climate model, you rent the 1,000 computers for exactly 15 minutes. The moment the job is done, the machines turn off, and the billing stops. The Savings: You convert a massive fixed Capital Expenditure (CapEx) into a lean, manageable Operational Expenditure (OpEx). Create your Wherobots account Get Started PostGIS vs Wherobots: Which One Is Right for Your Use Case? To make this concrete, let’s look at three distinct industry examples and which tool provides the best ROI for each. 1. Logistics and Delivery: Why Real-Time Tracking Needs PostGIS Scenario: You need to route drivers in real-time and show customers where their package is. The Choice: PostGIS. Why: You need transactional guarantees. If a driver marks a package as “Delivered,” that data must be instantly saved and visible. You are doing millions of tiny, fast lookups. 2. Insurance and Real Estate: Why Portfolio-Scale Risk Analysis Needs Wherobots Scenario: You need to calculate risk premiums for 10 million homes based on historical wildfire data, distance to fire stations, and vegetation density indices. The Choice: Wherobots. Why: This is a “Global Join.” You are comparing massive datasets against each other. PostGIS would take weeks to process this at a national scale. Wherobots can recalculate the entire portfolio every night, allowing you to adjust pricing dynamically. 3. Urban Planning: Why Time-Series Sensor Data Overwhelms a Standard Database Scenario: You are ingesting telemetry data from 50,000 connected traffic lights and sensors to analyze traffic congestion trends over the last 5 years. The Choice: Wherobots. Why: The volume of data (Time-Series) is too large for a standard database. A database would bloat, slow down, and become expensive to back up. Wherobots can read this data directly from cheap storage, aggregate it into trends, and output the results. PostGIS vs Wherobots: Decision Checklist PostGIS is a good choice when your primary use case is powering a user-facing application that needs fast, transactional lookups on a relatively stable dataset. Wherobots is the better choice when you are running analytical queries across complex datasets, processing historical data at scale, or need compute costs that flex with actual usage rather than peak capacity. If you are currently evaluating your data stack, use this simple checklist to guide your architecture decision. Stick with PostGIS if: [ ] Your primary goal is powering a user-facing application. [ ] You need to edit data manually (e.g., fixing property boundaries). [ ] Your dataset is relatively stable and fits comfortably on one large server. [ ] You require strict “ACID” transactions (meaning every write is confirmed and visible before the next read — no partial updates, no stale reads) Move to Wherobots if: [ ] You are waiting hours or days for analytical queries to finish. [ ] Your cloud database bill is growing faster than your revenue. [ ] You need to join two massive datasets (e.g., “All Buildings” + “All Parcels”). [ ] You are building AI/Machine Learning models that need to “learn” from all your historical data. The most competitive organizations today realize they don’t have to choose just one. They use Wherobots to crunch the data cheaply and efficiently, and then move the polished results into PostGIS for instant access—a strategy we will cover in our next post on the “Medallion Architecture”, a data design pattern where raw, refined, and production-ready data are stored in separate layers, each optimized for different workloads Ready to see what this looks like for your workload? Contact us and get a cost comparison built around your actual data volume.
WherobotsDB is 3x faster with up to 45% better price performance Posted on March 11, 2026October 3, 2026 by Damian Today we are announcing the next generation of WherobotsDB, the Apache Sedona and Spark 4 compatible engine, is now generally available. Compared to the the previous generation of WherobotsDB, this next gen (now the current) architecture and version accelerates queries by up to 3x, with up to 45% better price-performance. Previously in preview, it is now offered through the latest version of WherobotsDB. How Customers Use WherobotsDB Our customers are using WherobotsDB to create insights from spatial data at scale that result in improved products, services, and decision making in the physical world. They are realizing breakthroughs in fleet operations, improving their risk projections, increasing the accuracy of vegetative forecasts, analyzing change, and overall are more capable of innovating against physical world interests. With Wherobots on AWS, not only can we easily scale to millions of acres and continuous tractor telemetry normalization within LeafLake, but also we can rest assured that our costs won’t spiral out of control. — G. Bailey Stockdale CEO, Leaf Agriculture The workloads customers run or want to run, are becoming more ambitious too. Customers and AI alike demand better/faster/cheaper solutions for working a wide variety of raster and vector spatial datasets, and they need to fuse this data with valuable business context as well. And of course, they want to do it without scale and function limitations. WherobotsDB was always designed from the ground up to meet these needs. But solutions are now even easier to realize, because the latest version of WherobotsDB allows you to do more, faster, at a lower cost. WherobotsDB Benchmark Results: TPC-H and SpatialBench Performance The next generation of WherobotsDB delivers a substantial performance increase for both spatial and non-spatial queries. Compared to the previous engine, our benchmarking runs for of TPC-H and SpatialBench both show significant performance improvements across scale factors of 100 and 1000. Up to a 3x peak acceleration for analytical queries based on TPC-H at a scale factor of 1000. Workloads can see a nearly threefold increase in throughput. The mean observed acceleration was 1.7x. Up to a 2.5x acceleration for spatial queries and joins based on SpatialBench at a scale factor of 1000. Spatial operations, such as intersects and joins, complete significantly faster on small to massive datasets. The mean observed acceleration was 1.9x. Up to a 45% improvement in price performance. Customers can achieve more with less cost, and the added performance makes it possible for interactive workloads to downscale into smaller runtimes. Compared to the previous version of WherobotsDB, the average improvement in price performance was 25%. Price Performance vs Next Best Engine This shows the total cost of all SpatialBench queries at SF 1000 that the next best engine could finish under a timeout of 10 hours, which limited the comparison to Q1-Q5 and Q7. The remainder of the queries (Q6, Q8-Q12) could not be completed by that engine and were excluded from this analysis. The current generation of WherobotsDB is more capable and 46% lower cost than the next best engine, which is a popular Spark based serverless engine with Spatial SQL support. How the current version of WherobotsDB compares on cost to the previous, as well as the next best engine. WherobotsDB Capabilities: What No Other Engine Offers WherobotsDB is the only engine capable of meeting the following spatial data requirements that customers have. These capabilities power lakehouse architectures like the one described in The Medallion Architecture for Geospatial Data. You get: ✅ high performance, cost efficient, and scalable vector, raster, tabular data operations in a unified query environment ✅ compatibility with Spark 4 and Sedona ✅ interoperability with zero-copy on lakehouses and data lakes to keep data in your control ✅ unification with RasterFlow, to easily orchestrate planetary scale inference and analytics workflows starting with raw imagery datasets The following matrix isolates the vector data processing capabilities of the next best alternatives to WherobotsDB using SpatialBench runs at a scale factor of 1000. Raster data capabilities were not compared, because WherobotsDB was the only engine in the set that supports raster data capabilities. Apache Sedona SpatialBench query capability matrix running on WherobotsDB in comparison to multiple other spark based engines. Contact us and we can share additional details, or rerun these benchmarks on an engine of your choice. WherobotsDB Architecture: Rust-Native, Arrow-Columnar Execution The original architecture for WherobotsDB was built on a JVM-based execution model. It is an extraordinary platform for distributed computing, but its row-oriented execution model and JVM memory management introduce overhead that compounds at scale, especially for spatial-heavy workloads where every row carries complex spatial objects that must be serialized, deserialized, and processed one at a time. The new WherobotsDB version is built on a new architecture that addresses these bottlenecks by replacing the JVM-based execution layer with a Rust-native, Arrow-columnar engine, optimized for spatial data executions. This architecture builds on native geospatial support in Apache Iceberg and Parquet. Native Spatial Execution with SedonaDB The latest version of WherobotsDB takes advantage of SedonaDB, an open-source, blazing-fast analytical database engine where geospatial data is the first-class citizen. The use of Rust and SedonaDB allowed us to move spatial logic out of the JVM and directly onto the native execution layer. It provides a unified execution model that supports everything from scalar and window functions to complex spatial joins, aggregations, and geometry operations. Apache DataFusion Integration Apache DataFusion is a Rust-native query engine built from the ground up on the Apache Arrow in-memory columnar format. By integrating DataFusion’s native execution with WherobotsDB’s distributed engine, you get the best of both worlds: WherobotsDB’ battle-tested scheduling and DataFusion’s high-performance native processing. This means your existing WherobotsDB-based workflows, SQL queries, and Python notebooks continue to work exactly as before, but the actual computation accelerates from optimized native code rather than in the JVM. Zero-Copy Efficiency To remove the significant overhead of data conversion, we implemented a high-performance native geometry type based on the GeoArrow specification. This allows for “Zero-Copy” data handling, utilizing Arrow’s nested memory layout to represent geometries and geographies without the costly serialization and deserialization steps typically found in spatial databases. Spark 4 Compatability WherobotsDB is now compatible with the latest features of Spark 4 to provide a modern, robust environment that enforces ANSI SQL by default. This upgrade integrates the Wherobots engine with the newest advancements in distributed computing, including improved query planning and execution protocols. Try WherobotsDB Free: 30-Day Trial Get started now using a 30 day, $95 free trial available for the Professional Edition of Wherobots. The latest version of WherobotsDB is generally available today and is the default version for all new runtimes on Wherobots Cloud. Due to the enforcement of ANSI SQL, if you’re an existing customer we recommend testing your workloads on the latest version of WherobotsDB in a notebook, SQL session, or job run. For most workloads, the upgrade should be seamless. If you have questions about the upgrade, experience unexpected behavior, want a custom benchmark, or would like to discuss how Wherobots can benefit your business, reach out to the Wherobots team at support@wherobots.com, sales@wherobots.com, or by filling out our contact us form. Key takeawaysThe next generation of WherobotsDB (Apache Sedona and Spark 4 compatible) is generally available. Versus the previous generation: up to 3x peak acceleration on TPC-H at SF1000 (mean 1.7x); up to 2.5x on SpatialBench spatial queries and joins at SF1000 (mean 1.9x); up to 45% better price-performance (average 25%).On SpatialBench SF1000 cost, the current engine is 46% lower cost than the next-best engine — described as a popular Spark-based serverless engine with Spatial SQL support. That comparison includes only queries the other engine finished under a 10-hour timeout (Q1–Q5 and Q7); Q6 and Q8–Q12 were excluded because the other engine could not complete them.Architecture change: the JVM-based, row-oriented execution layer is replaced with a Rust-native, Arrow-columnar engine. Spatial logic runs through SedonaDB on the native layer; Apache DataFusion provides native processing on Arrow; a GeoArrow-based native geometry type enables zero-copy handling. Existing SQL and Python notebooks continue to work.WherobotsDB is described as the only engine in the compared set with raster capabilities, plus vector and tabular ops in one query environment, Spark 4 and Sedona compatibility, zero-copy lakehouse interoperability, and unification with RasterFlow. The latest version is the default for all new Wherobots Cloud runtimes. Body trial terms: 30-day, $95 Professional trial. ANSI SQL is enforced, so existing customers should test workloads before upgrading.
It takes 15 minutes for the Caltrain to get from Sunnyvale to SAP Center Posted on February 19, 2026October 3, 2026 by Pouyan Aminian That’s how long it took our MCP server to go from “how many bus stops are in Maryland” to an answer I’ve been doing a lot of reading lately on how AI is going to transform spatial workloads and that curiosity led me to this post on geoMusings. Here, Bill is demonstrating how Claude Code and agent skills capabilities can be used to wire up a chat-to-query-results interface in a few hours. He showcased the new skill by getting the agent to query his local Postgres instance for the number of Metro bus stops in Maryland, which returned a precise 4,563. I need to count the number of records in the metro_bus_stops table that are inside Maryland.The database is at localhost:5432, database name is “dev”,user “postgres” with password “postgres”.Points table: public.metro_bus_stops (geometry column: geom, id column: id)Polygons table: public.maryland_boundary (geometry column: geom, name column: name) As a dabbler of AI agents and a minor contributor to Wherobots’ very own MCP server, I immediately wondered how our MCP server would do against such a challenge. So I fired up my VS Code and just straight up asked: “How many bus stops are in Maryland?” Bear in mind, at the time I did not know if we have any data with bus stops in it in Wherobots’ data catalogs, I did not know what shape that data was in, I did not know if the MCP server could come up with a reasonable administrative boundary for Maryland, etc. And I fired off this query just as my CalTrain was departing Sunnyvale station. In about 5 minutes, the MCP server already identified two tables with bus stop information called places_place under the Overture Maps Foundation database in Wherobots Open Catalog. It achieved that by exploring our catalog and running sample queries against those tables to find the right data; all with zero human intervention. We are right about Lawrence Station at the point. In the next 5 minutes, the MCP server ran a series of queries against that table, self-identified errors (i.e., got 0 results and understood it was not expected), adjusted the query, switched tables, changed approaches until it was able to produce actual results. Our MCP server believes there are 19,740 bus stops in Maryland which is ~5 times as many as Bill’s post suggests. We just got to Santa Clara station, by the way, for those of you who are still following. So being a good aspiring data engineer, I challenged the MCP server: Why does this blog think there are only 4563 then?https://blog.geomusings.com/2026/01/14/spatial-analysis-with-claude-code/ The MCP server went back to work and gave me the diagnosis; Bill’s query is focused on Metro bus stops and my original question did not specify that: So in the last 5 minutes of this journey, I asked it to focus on Washington Metropolitan Area Transit Authority (WMATA) bus stops only and see what it comes up with! And just as we were about to pull into San Jose Diridon Station, the MCP server told me that there are 6,224 Metro bus stops in Maryland. Now, whether there are 4,563 Metro bus stops in Maryland or 6,224 ones, is a matter that shall be validated with people far more knowledgeable than myself on buses and their stops. The main point is that AI is making it possible for non-experts like myself to go from a question (expressed in natural language) to real insights in minutes (well a 15-minute train ride to be precise). Wherobots MCP is giving the AI the ability to answer questions about the real-world. In the real world, I would have asked the MCP server to generate a Notebook for me to reproduce this output and plot it on a map. I would then share that with my colleague to help me validate, correct and optimize my findings. What would have taken days to weeks (to go from theory to some early explorations to a shareable PoC and, finally, to production-quality code) can now be achieved in a matter of hours. The Caltrain experiment was just one question. In our recent office hours, we walked through the MCP server end to end, showing how it explores catalogs, generates spatial queries, debugs errors, and produces reproducible outputs. See the full workflow in action. Want to get started with our MCP server? Check out our getting started guide. It takes less than 5 minutes to configure the server and start chatting with the physical world! Create your account to get started START BUILDING Key takeawaysOn a Caltrain ride from Sunnyvale to San Jose Diridon (~15 minutes), the author asked the Wherobots MCP server in VS Code 'How many bus stops are in Maryland?' with no prior knowledge of whether the catalog had bus-stop data or a Maryland boundary.In about five minutes (around Lawrence Station) the MCP server identified places_place under Overture Maps in the Wherobots Open Catalog by exploring the catalog and running sample queries, with zero human intervention.In the next five minutes it self-identified errors (including 0-result queries), adjusted SQL, switched tables, and produced 19,740 bus stops in Maryland — about 5× the 4,563 Metro bus stops in Bill Dollins' geoMusings Claude Code post. When challenged with that post, it diagnosed that Bill's query was Metro-only.In the last five minutes, constrained to WMATA, it returned 6,224 Metro bus stops in Maryland. The author states those counts still need validation by people who know buses; the point is natural-language to insight in minutes. A notebook for validation is described as what they would do in a real workflow.
Wherobots and Felt Partner to Modernize Spatial Intelligence Posted on February 10, 2026September 1, 2026 by Ben Pruden We’re excited to announce Wherobots and Felt are partnering to enable data teams to innovate with physical world data and move beyond legacy GIS, using the modern spatial intelligence stack. The stack with Wherobots and Felt provides a cloud-native, spatial processing and collaborative mapping solution that accelerates innovation and time-to-insight across an organization. Wherobots delivers the most capable spatial query and inference engine for creating insights from physical world data (raster, vector, structured) of any scale. Felt delivers an intuitive, browser-based experience that enables business teams to explore this data, ask questions, and share insights collaboratively. The combination provides a new path to move forward for teams that are innovation-limited by their desktop-bound GIS tools, restrictive licensing arrangements, and unscalable workflows. What is Felt? Felt is the new standard for collaborative mapping, and their product is often described as “the Google Docs of GIS.” It is a cloud-native platform designed to turn complex geospatial data into actionable insights through its unique collaborative map development capabilities. Unlike traditional Geographic Information Systems (GIS) that are often desktop-bound, Felt lives entirely in the browser, allowing teams to create, analyze, and share interactive maps with the speed and ease of a modern productivity tool. Why Legacy GIS is Limiting Spatial Data Innovation Spatial data comes in various formats and sizes, and it needs to be processed and combined with other datasets for people, systems, and AI to innovate with it. Because few systems were developed to handle this wide spectrum of data complexity, format, and scale in a graceful way, users have been forced to create cumbersome, time consuming workarounds. In turn, this has led to a high degree of specialization, and a special category of GIS tools that simply can’t keep up with demand. Buyers, faced with few options, have been limited to old-guard licensing arrangements for lagging technology solutions. Such engagements plant a thorn in the side of many organizations who end up constrained in their ability to innovate unless they hire more specialized staff or contractors to develop workarounds that copy data to and from bespoke GIS tools or databases. What buyers want is to deploy modern tooling and AI that “just absorbs the complexity” such that their data practitioners can just work with this data and achieve their goals. They also want more flexibility to take the best tools over time and demand data sovereignty. If you’re working on a small yet nimble team of data scientists and engineers solving world-level problems, you want the capability to easily crunch terabytes of vector, raster, and structured data using familiar tools. Data interactivity is key because the easiest way to understand spatial data quality at scale is to inspect relationships visually with performance, and you can iterate faster towards the shared objective when in-team collaboration is seamless. Historically teams relied on piecemeal operations that extended the innovation cycles because of data and work siloes, or limited tooling support. Similarly, analysts, planners, and stakeholders want simple, AI enabled, visual-first tools to understand and ask their questions about spatial data so they can make decisions from it. Their ability to tailor analytics from a visual-first tool, was limited by the capabilities of the underlying query engine and the data available to it. Moving from Legacy GIS to a Modern Spatial Intelligence Stack Our partnership with Felt directly addresses this friction with a seamless integration between platforms to deliver the spatial intelligence stack. Using Felt’s AI-assisted and collaborative map development experience, customers can easily query, interrogate, and build insightful visualizations from multiple sources of spatial and non spatial data in their private data lakes or open repositories. This is done using Wherobots as the query engine that absorbs the heavy lifting associated with multi-modal spatial data processing and cataloging. Developers can also build directly against this stack using Wherobots’ Spatial AI Coding Tools. Key Benefits of the Wherobots and Felt Integration Easy for everyone: No user is left behind. Wherobots supports data engineers, scientists, and developers who need speed and performance, while Felt enables front end developers, GIS users, and business teams to get insight, instantly from the analysis. AI first: from inference, MCP tools, to collaborative maps driven by AI and natural language prompts, users have the modern AI tools to generate insights quickly at scale. Built for the lakehouse: The lakehouse architecture maximizes data sovereignty and agility by giving you the freedom to choose the best tools for the job over time. Spatially optimized: World class spatial capability and performance is ready out of the box. Wherobots provides the scale, performance, and capability to crunch and perform inference on multi-modal spatial data than any other cloud platform. A fully managed, friendly, open standards based deployment: on-demand deployment and pay as you go pricing lets you reach your goals through open, well understood standards and architecture. Queries against Felt-visualized data run on WherobotsDB, with 3x faster performance than previous-generation engines. Customer Example: How Leaf Agriculture Uses Wherobots and Felt The partnership started with a joint customer request from Leaf Agriculture, who has been using Wherobots and Felt for productionizing their LeafLake offering over this past year, and to harmonize large scale tractor data for their Leaf Unified API. Leaf serves customers and partners like Syngenta, Bayer and Farmers Edge, and many others in the agricultural economy. The delivery of LeafLake, supported by Wherobots and Felt, creates a high-velocity data fabric that turns fragmented agronomic data into planetary-scale intelligence. Wherobots serves as the query engine, running distributed spatial SQL operations to create insight-ready datasets from millions of acres of machine data and imagery at 5–20x the speed and at a fraction of the cost of traditional solutions. By running directly on Leaf’s unified data lake, Wherobots transforms raw telemetry into structured, “AI-ready” insights in seconds. Felt acts as the collaborative “window” into this data, providing a high-performance mapping interface that lives entirely in the browser. Through a native integration, data processed in Wherobots is made directly to Felt. This “SQL-to-map” workflow allows agronomists and decision-makers to interact with LeafLake data in real-time in Felt, enabling a “Google Docs” style of collaboration over complex agricultural insights. This map showcases farm data from Leaf Agriculture’s LeafLake platform through Felt’s interface for the purpose of precision farming based on tractor telemetry and soil data. Get Started Today This integration marks a major step forward making the creation of spatial intelligence accessible to entire organizations. We can’t wait to see what you build with the combined power of Wherobots and Felt. To learn more: Check out the documentation to get started in minutes. Join us in this upcoming session to see how Wherobots and Felt work together to build the foundation for spatial AI. Talk to our team to explore options. Key takeawaysWherobots and Felt are partnering on a cloud-native stack: Wherobots as the spatial query and inference engine for raster, vector, and structured data of any scale; Felt as the browser-based collaborative mapping layer (often described as 'the Google Docs of GIS').The pitch versus legacy GIS: desktop-bound tools, restrictive licensing, and unscalable workarounds that copy data in and out of bespoke GIS. The combined stack is lakehouse-based, AI-first (inference, MCP, natural-language maps), fully managed, pay-as-you-go, and open-standards based.Queries against Felt-visualized data run on WherobotsDB, with 3x faster performance than previous-generation engines (as stated in this post).Joint customer Leaf Agriculture uses the stack for LeafLake and the Leaf Unified API, serving customers including Syngenta, Bayer, and Farmers Edge. Wherobots runs distributed spatial SQL on millions of acres of machine data and imagery at 5–20x the speed and a fraction of the cost of traditional solutions, turning telemetry into structured insights in seconds. Felt is the collaborative window; the workflow is described as SQL-to-map.