Wherobots brought modern infrastructure to spatial data in 2025 Posted on January 26, 2026October 3, 2026 by Damian In 2026 we’re bridging the gap between AI and data from the physical world. Entering 2025, we knew we needed to prove Wherobots is fundamentally the best place to create and run spatial data workloads at scale. Last year we directed the vast majority of our energy at strengthening the core fundamentals– ease of use, cost, performance, and reliability, knowing this focus would resonate with customers and value would be amplified through what we build later on. We knew we needed to bring spatial data into the modern data architecture, which from our vantage point, is the data lakehouse. If we did nothing, much of this data would otherwise remain siloed, “special”, and out of reach of modern analytics engines that could put this data to work. This is why we led contributions of GEO type support to Iceberg and Parquet. Off-platform through the open source Apache Sedona project, we saw an opportunity to develop a lightweight query engine that would appeal to developers because it would provide the support they need out of the box and accelerate iterations with spatial data. And so SedonaDB was born. This post is a high level summary of these and other accomplishments from our team in 2025. Now, we are actively building on top of this improved foundation, to enable AI and data practitioners across industries and use cases to operate with a heightened understanding of the physical world, for any area of interest. Customer success with scalable spatial data platforms This is what matters most. Everything else written here is just supporting evidence that shows how we made our customers more successful with spatial data this year. And what better proof than their own words? “Working with Wherobots let us focus on what matters – helping our clients make better land decisions. Their platform helps us scale efficiently while keeping our attention on real-world outcomes across energy, conservation, and development” Danan Margason Founder & CEO at Aarden.ai “The fact that Wherobots can mosaic imagery over millions of square kilometers, run AI models over those mosaics, and then organize the outputs in minutes, for 10 to hundreds of dollars is nothing short of incredible, making global scale analysis a routine pipeline instead of an enormous and extremely expensive endeavor, has the potential to change monitoring tasks far beyond field boundary analysis.” Caleb Robinson Principal Research Scientist, Microsoft AI for Good “With Wherobots, we were able to merge 15+ complex vector datasets in minutes and run high-resolution ML inference at a fraction of the cost of our legacy stack. The combination of speed, scalability, and ease of integration has boosted our engineering productivity and will accelerate how quickly we can deliver new geospatial data products to market.” Rashmit Singh CTO, SatSure “We’re helping to democratize large-scale spatial analytics for everyone. With Wherobots, we can move faster, scale bigger, and help more organizations make smarter decisions about the physical world.” Eric Pollard Founder and CEO, ParGo Elevating spatial data and AI capabilities for AWS customers It’s no surprise that many of these customers are AWS customers. In late 2024, we launched Wherobots Cloud as a product AWS customers could subscribe to directly through the AWS Marketplace. This activated a key value distribution channel between Wherobots and AWS customers. It also started our partnership in earnest with AWS. We continue to work closely with the AWS team to bring world-class spatial capabilities into the hands of their customers so they can better realize their objectives with physical world data. Leadership in the open data and lakehouse ecosystem We were the team that led the introduction of GEO type support to Iceberg and Parquet, which led to the incorporation of GEO types in the Databricks Delta Lake project. Because these projects form foundational components of the modern data architecture, with GEO type support, a significant portion of spatial data could now be interpreted and safely interoperated on by common compute engines. That also meant it no longer had to live in siloed architectures. It could thrive in the common data estate – the data lake, processed using engines like Spark, Snowflake, BigQuery, or Wherobots. If you squint, there is now a clear path for making spatial data look just like “data” in the eyes of developers and AI systems, particularly when capable engines like Wherobots can crunch it without a problem. Cloud and lakehouse integrations for spatial workflows Wherobots Cloud evolved into a full-fledged spatial intelligence platform designed not just for querying geospatial data very efficiently at scale, but for building production-grade workflows that integrate high value derivatives of physical world data into customers’ existing data architectures. Our native integration with Amazon S3 makes it possible for customers to run Wherobots on spatial data in their storage. In 2025 we announced our integration with Unity Catalog, enabling Databricks customers to activate the value of Wherobots on spatial data in their Databricks lakehouse. We will continue to add integrations such that our customers can just add Wherobots’ magic to the data infrastructure they already have. Apache Sedona appeals to a wider audience The Apache Sedona community developed and launched SedonaDB, the first open-source, single-node analytical database engine that treats spatial data as a first-class citizen. They also made it significantly easier to compare query performance across engines using SpatialBench, the first benchmarking framework for spatial queries. SedonaDB makes spatial data significantly more useful and analytically accessible for a wider range of use cases and personas. SpatialBench is there to streamline the decision making process for users looking to choose an engine based on spatial query price-performance and capability. Here’s a link to the announcement for both. WherobotsDB keeps getting better We’re continuously improving the WherobotsDB engine, raising the bar we’re self-setting for spatial query price-performance and capability. In 2025 we announced multiple new functions, tools, and compatibility with GeoPandas. We also announced a preview of a new runtime version, 2.x, which contains the latest optimizations for spatial range queries, spatial filtering, and spatial joins. Rust is at the core, and it leverages vectorized execution. Compared to the first major version, 2.x further accelerates spatial queries of up to 3.3x. While Wherobots is generally known for its spatial data capability, 2.x is significantly more performant for general purpose query operations, to a degree in which it’s also TPC-H competitive to alternative managed Spark engines in the market.We will be sharing benchmarking results when we announce general availability for version 2 soon. Making remote sensing and Earth sbservation data AI-ready We launched RasterFlow in private preview, the first serverless workflow purpose built to prepare and perform inference on large scale Earth observation (EO) datasets. RasterFlow is an Earth Intelligence solution, addressing the infrastructure challenges and high costs that prevented companies from utilizing raw EO data in the first place. We’ve packaged years of GeoAI expertise into a serverless, easy to use product. RasterFlow creates inference-ready mosaics after digesting large, unprepared imagery datasets, and runs model inference on these mosaics with custom or open PyTorch models to perform tasks such as change detection, classification, and segmentation. Results are delivered as geometries in Iceberg tables in a customer’s S3 bucket or as tables in Databricks Unity Catalog, to be processed by WherobotsDB or alternative engines. Enabling AI on physical world data Physical world data is noisy, it’s large, and it’s generally semi-structured or unstructured. This data also needs more context in order for it to be useful, which requires spatial joins to other datasets. It can be publicly available, or reside as private assets within an organization’s data estate. Teams and AI systems alike need this data to be processed and contextualized with other data in order for it to be useful. At scale Wherobots provides arguably the best tools for spatial processing and contextualization at the most fundamental level, for the modern data architecture. Now Wherobots is ready to be wired up to AI systems. Late in 2025 we announced the availability of the Wherobots MCP server to give your LLM access to Wherobots’ tools. Now, LLMs can use the MCP server to efficiently design queries by understanding the spatial and non spatial data in your data estate (via the S3 and Unity Catalog integration), and run those queries on a high performance engine to answer questions about the physical world. Soon we’ll integrate the MCP server with RasterFlow. That way, an AI agent can design and trigger a workflow in Wherobots that starts with fresh EO data, prepares it for inference, generates predictions using a collection of PyTorch machine learning models, and perform additional enrichment or transforms if needed to produce the result. Join the upcoming office hour on MCP server to learn more. What’s next for Wherobots in 2026 You can reach out to the product team at product@wherobots.com, or me directly at damian@wherobots.com to share the challenges you’re facing and see how we can solve them now or with capabilities we’ll add. Here are a few discrete roadmap items we are working on, and categories of investment planned in 2026. Make RasterFlow generally available Enabling AI on physical world data Bringing Wherobots closer to developers (VS Code, Kiro, etc) Providing support for Wherobots in the compute environment you need it to run in Making WherobotsDB even faster and cost effective Broaden the capabilities of SedonaDB Key takeaways2025 focus was core fundamentals — ease of use, cost, performance, reliability — and bringing spatial data into the lakehouse. Wherobots led GEO type support in Iceberg and Parquet, which then flowed into Databricks Delta Lake, so spatial columns can live in the common data estate instead of siloed GIS stores.Wherobots Cloud grew into a spatial intelligence platform with native S3 integration and a Unity Catalog integration so Databricks customers can run Wherobots on lakehouse spatial data. Wherobots Cloud had launched on AWS Marketplace in late 2024.The Apache Sedona community launched SedonaDB (single-node analytical database with spatial as a first-class citizen) and SpatialBench. WherobotsDB 2.x preview was announced with Rust and vectorized execution, accelerating spatial queries by up to 3.3x versus the first major version, and described as TPC-H competitive with alternative managed Spark engines.RasterFlow launched in private preview as a serverless workflow to mosaic unprepared EO imagery and run custom or open PyTorch models, writing geometries to Iceberg in S3 or Unity Catalog. Late 2025: Wherobots MCP server so LLMs can design and run queries against S3 and Unity Catalog data. 2026 plans listed: RasterFlow GA, more AI enablement, VS Code/Kiro, more compute environments, faster WherobotsDB, broader SedonaDB.
The Medallion Architecture for Geospatial Data: Why Spatial Intelligence Demands a Different Approach Posted on January 9, 2026October 4, 2026 by Matt Forrest When most data engineers hear “medallion architecture,” they think of the traditional multi-hop layering pattern that powers countless analytics pipelines. The concept is sound: progressively refine raw data into analytical data and products. But geospatial data breaks conventional data engineering in ways that demand we rethink the entire pipeline. This isn’t about just storing location data in your existing medallion setup. It’s about recognizing that spatial data introduces complexity in data structures, compute patterns, and efficiency requirements that traditional architectures simply cannot handle without significant compromise. The medallion (or multi-hop) architecture, when properly adapted for geospatial workloads, becomes something more powerful: a systematic approach to managing spatial intelligence at scale. In this post, I will outline how we have taught this in the Wherobots Geospatial Data Engineering Associate course to leverage this base design and group different spatial tasks, query patterns, and outputs into each step, leveraging both raster and vector datasets. What Is Medallion Architecture? Medallion architecture is a data design pattern that organizes a lakehouse into three layers of increasing quality: Bronze, Silver, and Gold. Bronze stores raw data as it arrives. Silver stores cleaned, validated, and conformed records that downstream jobs can trust. Gold stores aggregated, business-ready tables for analytics, dashboards, and machine learning. Databricks popularized the pattern, which is also called a multi-hop architecture because data moves through the layers in sequence. LayerWhat it storesGeospatial exampleBronzeRaw source data, unchangedShapefiles, GeoTIFFs, and GPS feeds, registered in place on cloud storageSilverValidated, standardized dataRepaired geometries, one coordinate reference system, GeoParquet and cloud-optimized GeoTIFF filesGoldAggregated, analysis-ready dataParcel risk scores, H3 aggregates, and PMTiles for web maps Each layer needs changes for geospatial data. Spatial sources mix vector and raster formats, span many coordinate reference systems, and require spatial joins that cost more than tabular joins. The sections below adapt Bronze, Silver, and Gold for those workloads. Why Geospatial Data Demands Architectural Rethinking Before diving into the medallion architecture itself, let’s acknowledge what makes geospatial data different. The Fragmentation Problem In most organizations, geospatial data doesn’t arrive nicely packaged. You might pull property boundaries from a municipal data portal in shapefiles, elevation data from AWS open data as cloud-optimized GeoTIFFs, satellite imagery from NASA, and local POI data from your internal systems. Each source has different formats, coordinate systems, refresh rates, and quality levels. Traditional data lakes aren’t equipped to handle this heterogeneity efficiently. Moving terabytes of raw geospatial data through your pipeline adds cost and latency. And then there’s the format problem: legacy formats like shapefiles and JPEG2000 can’t be partially queried. If you want Seattle property data from global shapefiles, the traditional approach is to download the entire file, unzip it, filter locally, then transfer what you need. The Scale-Complexity Tradeoff Geospatial scale isn’t just about row count. It’s about the combination of volume, global distribution, mixed data types (vectors, rasters, array-based formats, point clouds), conversion / interpolation between those data types, and the inherent cost of spatial operations. A simple point-in-polygon operation across a million properties and thousands of geographical features involves computation that traditional systems struggle with. This is why you see organizations either underutilizing their spatial data or building custom solutions that work for one specific problem but don’t generalize. The efficiency gap is real, and it’s expensive. The Medallion Architecture: Bronze, Silver, Gold The medallion architecture solves these challenges by enforcing structure while maintaining flexibility. Here’s how it works for geospatial data: Bronze Layer: Ingestion and Preservation The bronze layer is deceptively simple: get raw geospatial data into your system without transformation. For Data Engineers: This layer is your staging area. You’re ingesting data from disparate sources: APIs, files, databases, cloud storage without enforcing business logic. The goal is to preserve raw data fidelity while standardizing the container. For geospatial data specifically, this means several important decisions: Leave data at source when possible. If you have a global elevation dataset sitting on AWS S3, don’t download it into your local systems. Create a remote reference to it. This is one of the biggest efficiency gains in modern geospatial data pipelines. Convert to cloud-native formats. As data arrives in your bronze layer, immediately convert legacy formats (shapefiles, JPEG2000, uncompressed GeoTIFFs) into cloud-native equivalents (GeoParquet, cloud-optimized GeoTIFFs). This isn’t optimization but a prerequisite for efficient querying downstream. Preserve lineage and versioning. Geospatial data often changes such as when property boundaries get redrawn, satellite imagery updates, elevation models improve. You need to track when data arrived, from where, and what version you’re working with. Raster tiling and storage optimization. For raster data (imagery, DEMs), bronze handles tiling and creates remote references. Instead of loading massive rasters into memory, you tile them and access only what you need. For GIS Professionals: Think of bronze as your data repository. In traditional GIS workflows, you manage multiple datasets, each in its own format, living in different folders or databases. Bronze centralizes this but respects the raw nature of the data. You’re not making decisions about coordinate systems, geometry validation, or spatial relationships yet. This is also where automation becomes critical. Instead of manually downloading datasets monthly, bronze ingestion jobs run on schedules, automatically pulling new data and versioning it. This means your analyses always reflect current reality. Data table formats such as Apache Iceberg and Delta Lake provide several advantages here since you can easily append data to the table for records that have changed. They also use underlying cloud-native formats such as GeoParquet with the new V3 Iceberg specification and within Wherobots you can read out-of-database rasters (or a reference to a specific part of an image) without moving the source raster file saving massive amounts of I/O operations. How Wherobots Handles the Bronze Layer Wherobots Cloud simplifies bronze layer operations through WherobotsDB, its cloud-native spatial processing engine. When ingesting raw geospatial data, Wherobots immediately enables critical capabilities for you: Format conversion at ingestion: Wherobots automatically converts legacy spatial formats (shapefiles, JPEG2000, standard GeoTIFFs) into cloud-native formats during ingestion. This happens transparently—you connect to data in whatever format it arrives, and it’s can be optimized for cloud storage and querying. Remote reference handling: For massive open datasets, Wherobots manages remote references natively. Global elevation datasets sitting on AWS S3 don’t get duplicated into your systems. Instead, Wherobots creates efficient references to the original data, allowing you to query it as if it were local while paying only for what you access. Apache Iceberg storage with spatial extensions: Data in bronze is stored using Havasu, Wherobots’ Apache Iceberg-based spatial table format – which is the foundation of the Iceberg v3 spec. This provides built-in versioning, time travel, and ACID transactions from day one. You can rewind to yesterday’s data for audits or reprocessing without complex manual versioning schemes. Automatic metadata tracking: Wherobots tracks lineage, data arrival times, and schema information automatically, eliminating the need for manual data cataloging in bronze. The key to Wherobots’ approach to bronze is that it enforces good practices without requiring you to build them yourself. You’re getting schema evolution, ACID compliance, and spatial optimization automatically. Silver Layer: Enrichment and Standardization Silver is where the intelligence begins to emerge. This is where you clean, enrich, and structure data into a consistent, analysis-ready state. It’s also often where Data Science teams may want to work before “publishing” analysis ready data to end users (in gold tables). For Data Engineers: Silver involves: Coordinate system harmonization. Raw geospatial data arrives in dozens of different projections. Silver transforms everything into standard coordinate systems (typically WGS 84 for global work, or local UTM zones for regional analysis). Geometry validation and repair. Geospatial data is messy. Overlapping polygons, self-intersecting lines, invalid geometries from legacy systems all appears in silver transformation. You validate geometries and fix what’s repairable. Spatial enrichment operations. This is where you add geospatial relationships. Perform spatial joins to associate properties with neighborhoods. Buffer roads to identify nearby areas. Create 3D geometries by joining elevation and additional data attributes to 2D features. Calculate spatial statistics like area, perimeter, or distance to nearest feature. Advanced spatial operations. Often times we think of spatial relationships just as a spatial join. But performing analysis such as a K-Nearest Neighbor join, zonal statistics (or a raster to vector join), spatial aggregation, distance within (i.e. features within N distance), area weighted interpolation (such as weighted population statistics), shortest path analysis, line of sight analysis, space and time proximity analysis (i.e. near miss) and more. These require not only optimized spatial data but the right functions to make them work. The key architectural principle in silver is that you’re not making final analytical decisions you’re preparing the data so downstream consumers can make those decisions efficiently. For GIS Professionals: Silver is where your spatial data processing happens. This is the layer where you answer structural questions: What coordinate system makes sense for this analysis region? Are there geometry issues I need to resolve? What spatial relationships exist between datasets that I’ll need downstream? Should I enrich my vector data with elevation? Satellite indices? Population density? In traditional GIS, this work happens ad-hoc before each analysis. In the medallion architecture, you do it once, well, and capture the work in reusable transformations. This is efficiency multiplied across teams and projects. Another key note is that you will likely want different scales of compute to handle these processes. Easier process like spatial joins on vector basic geometries may only require a smaller compute instance whereas a large scale NDVI (or other vegetation indexes) would require a larger compute instance. The ability to mix and match your compute scales is a critical advantage in a spatial platform not only for speed but cost optimization. How Wherobots Handles the Silver Layer The silver layer is where WherobotsDB spatial optimization truly shines. Wherobots provides native spatial operations that make complex transformations not just possible but efficient at scale: Coordinate system transformations at speed: Wherobots includes optimized functions for reprojecting geometries across coordinate systems. What might take hours in desktop GIS or loose scripts in Python happens in seconds on Wherobots’ distributed architecture. This is critical because spatial efficiency compounds—faster transformations mean you can afford to enrich data more completely. Geometry validation and repair: Wherobots includes native geometry validation and repair functions. You can identify invalid geometries, split self-intersecting polygons, and fix common data quality issues in SQL. This work is parallelized across the cluster, not constrained by a single machine’s memory. Spatial enrichment operations: Silver transformations in Wherobots use Spatial SQL with 300+ functions for vector and raster operations. Spatial joins that would take hours in traditional systems execute in minutes. Buffer operations, intersection calculations, proximity analysis—all happen efficiently on distributed data. Wherobots handles both vector and raster data natively in the same query, meaning you can enrich vector property data with raster elevation or satellite imagery in a single operation. Raster tiling and remote storage: For raster data, Wherobots can tile and optimize imagery while keeping it remote on S3. You’re not loading massive datasets into memory. Instead, Wherobots’ raster functions work against remote tiles, accessing only what’s needed for each query. Reusable SQL transformations: Silver transformations are written in Spatial SQL and stored as views or jobs. These are version-controlled, reproducible, and can be re-run as new data arrives in bronze. Unlike one-off Python scripts, silver logic in Wherobots is production-grade from day one. The combination of Spatial SQL and WherobotsDB’s distributed architecture means silver layer work that once required deep geospatial expertise and careful optimization is now accessible to data engineers who know SQL. Wherobots handles the spatial complexity. Gold Layer: Analytical Readiness and Delivery Gold is your analytics-ready product layer. Data here is fully prepared, enriched, and available for immediate use in downstream applications and BI or GIS systems. For Data Engineers: Gold contains: Advanced aggregates and spatial statistics. Point-in-polygon counts (how many properties in each neighborhood), zonal statistics (average elevation by region), spatial clustering (identify hotspots) updated on a regular basis, ready to join by identifier to non spatial data. Optimized indices and pre-computed relationships. Store commonly-needed spatial joins and queries pre-computed, reducing downstream query time. Tiled and multi-format outputs. Generate PMTiles for web visualization, GeoParquet for analytics, and optionally push to PostGIS or SedonaDB/DuckDB for application serving. Quality-controlled data. Remove erroneous geometries that made it through silver, apply final business logic, and ensure consistency. AI ready data. Create language based data that is ready for an LLM or agentic applications to consume removing heavy geometries which can confuse an LLM. Gold is where you can afford to be opinionated because the work upstream ensures those opinions are well-informed. For GIS Professionals: Gold is your analysis starting point. Instead of wrestling with raw data, you access gold tables that are: Geometrically valid and properly projected Enriched with relevant spatial context (elevation, proximity to features, administrative boundaries) Pre-aggregated for common questions Formatted for direct use in your tools and custom applications (QGIS, Felt, BI dashboards, machine learning pipelines) You go from “I need to combine three datasets and fix projection issues” to “Here’s the enriched regional dataset, ready for analysis.” How Wherobots Handles the Gold Layer Gold layer queries run on WherobotsDB, which delivers 3x faster performance and 45% better price-performance on spatial workloads. Gold layer work in Wherobots focuses on making spatial data immediately useful across your entire organization: Pre-computed aggregates at scale: Wherobots can compute spatial aggregates—point-in-polygon counts, zonal statistics, spatial clustering—and store them efficiently. These pre-computed metrics are small enough to serve directly to applications but powerful enough to support complex analyses. A query that aggregates a billion point observations into regional summaries completes in seconds. Multiple output formats: Gold data can be materialized in multiple forms depending on consumer needs. Export as GeoParquet for analytics, create PMTiles for web visualization, push to PostGIS for application serving, or generate standard Parquet for BI tools. Wherobots’ native format support means you write once and consume many ways. Web tile generation at scale: Wherobots includes a scalable vector tile (VTiles) generator that’s optimized for producing map tiles from massive datasets. Instead of spending weeks generating tiles for a global dataset, Wherobots produces them in hours. These tiles feed directly into web applications, Felt, QGIS, or any tile-consuming tool. RasterFlow integration: Gold layer work in Wherobots increasingly includes AI-powered enrichment. RasterFlow raster inference allows you to extract insights from satellite imagery at scale—identifying buildings, roads, vegetation patterns—and embed those insights directly into your gold tables. Machine learning models that would require custom integration work in other systems are built into Wherobots’ gold workflow. Spatial SQL API for consumption: Instead of exporting and distributing files, gold data is served through Wherobots’ Spatial SQL API. Applications, dashboards, and downstream systems query gold data directly through Python SDK, Java JDBC driver, or standard SQL endpoints. This means gold data stays fresh and you don’t distribute copies. Gold in Wherobots is less about storing static snapshots and more about providing a live, queryable product layer that adapts to consumer needs. Why This Architecture Solves Geospatial Problems Efficiency Through Format Optimization Traditional systems treat geospatial data like regular tabular data. This is inefficient. A shapefile is a collection of files in a zip that requires downloading the entire dataset to access a subset. GeoParquet, by contrast, stores data in columnar format with spatial indexing built-in. You can query it over the internet by bounding box or geometry filter and only transfer what you need. The medallion architecture enforces this progression from legacy to cloud-native formats, reducing storage costs and query latency dramatically. A terabyte global elevation dataset stays on AWS S3 as a reference. You don’t duplicate it. Separation of Concerns Bronze, silver, and gold are separate databases with separate lineage. This means: Raw data safety. Your source data is never modified. If a transformation downstream goes wrong, you can reprocess from bronze without losing the original. Independent evolution. Teams can improve silver transformations without affecting gold. Applications consuming gold don’t care about silver changes as long as contracts remain stable. Governance simplicity. Access controls are straightforward. Grant different teams access to different layers based on their role. Automation and Scalability By encoding transformations into standardized layers, you enable automation. Data ingestion becomes scheduled jobs. Updates propagate automatically from bronze through silver to gold. You stop manually managing data and start managing logic. This scales to global datasets because you’re not moving data—you’re moving queries. Spatial predicates push down to where the data lives, reducing network transfer and computation. Real-World Application: Housing Analytics at Scale Consider a concrete example: analyzing housing market patterns across Seattle using a blend of property records, elevation, transportation networks, satellite imagery, and census data. Bronze Layer: Ingest property sales records from the county (updated monthly) Pull property boundaries from municipal GIS (shapefile, converted to GeoParquet) Reference global DEM (stays on AWS S3) Ingest road network (OpenStreetMap) Store satellite imagery references (Copernicus, AWS Open Data) Silver Layer: Transform property boundaries to WGS 84 and validate geometries Perform spatial join: properties to neighborhoods Enrich with elevation by querying DEM at property centroids Calculate proximity metrics: distance to nearest transit, nearest park Tile satellite imagery for efficient access Gold Layer: Aggregate: median price by neighborhood, price trends by elevation Pre-compute spatial clusters: identify hot markets Generate PMTiles for web visualization Export clean GeoParquet for ML pipelines Push refined data to PostGIS for application serving This progression from raw, fragmented sources to a unified, enriched analytical product is exactly what the medallion architecture enables. And the efficiency gains compound as more analyses build on the same gold layer. Comparison to Other Approaches Why Not Keep Everything in PostGIS? PostGIS is powerful for transactional spatial queries and operations, but it’s a database optimized for consistency and ACID transactions, not analytics at scale. As data volume grows, PostGIS becomes expensive to operate and scale. You’re paying for transactional guarantees you don’t need for analytics. The medallion approach uses cloud storage (S3) as the primary store, which scales cheaply, and can use PostGIS only for the final gold layer serving to applications that need it. This is typically more cost-effective and performant. Why Not Use a Specialized GIS Data Warehouse? Geospatial-specific data warehouses exist but are typically proprietary, expensive, and inflexible. They don’t integrate easily with your existing data infrastructure or machine learning pipelines. The medallion architecture is platform-agnostic—it works with Spark, Flink, DuckDB, or any distributed compute framework that understands spatial operations. Why Not Just Use Desktop GIS with Cloud Storage? Desktop GIS tools (QGIS, ArcGIS) can read cloud storage but aren’t designed for production pipelines. They require manual steps, don’t automate updates, and don’t scale to the volume and frequency of modern geospatial data. The medallion architecture automates what desktop GIS does manually. Implementation Considerations Technology Stack The medallion architecture for geospatial data typically uses: Storage: Cloud object storage (AWS S3, Google Cloud Storage, Azure Blob) Table Format: Apache Iceberg with spatial extensions (Havasu), which provides schema evolution, ACID transactions, time travel, and spatial indexing Compute: Apache Sedona or Wherobots (distributed geospatial framework) or SedonaDB (single-node spatial SQL) Orchestration: Workflow tools (Airflow) to manage bronze→silver→gold pipelines Visualization: Web tiles (PMTiles), GIS tools (QGIS, Felt), or BI platforms Iceberg: The Foundation Apache Iceberg is the table format that makes this all work efficiently. It’s a metadata layer over Parquet files that provides: Schema evolution: Add columns without breaking downstream queries ACID transactions: Reliable concurrent updates Time travel: Query historical snapshots (rewind to yesterday’s data) Partitioning: Automatic data organization for efficient queries Spatial indexing: Efficient spatial predicates and pushdown optimization With Iceberg’s spatial extensions, you get native geometry storage and spatial optimizations, making it the ideal foundation for medallion pipelines. Plus with many upstream data warehouses now accepting Iceberg V3 you have a zero ETL process to all of these systems from your gold tables. Data Lineage and Governance The medallion architecture inherently supports data governance. Each layer has clear ownership and lineage. Track which source datasets feed which analyses. Implement role-based access controls at layer boundaries. Maintain data catalogs describing bronze sources, silver transformations, and gold products. Putting this into practice with Wherobots From Theory to Practice The medallion architecture isn’t a theoretical concept—it’s proven across thousands of data engineering organizations. What makes it powerful for geospatial is that it acknowledges geospatial’s unique challenges: Format diversity: Enforced conversion to cloud-native formats Computation intensity: Spatial predicates pushed down to efficient compute Scale complexity: Remote references and tiling instead of data movement Governance needs: Clear layer separation for access control and lineage Integration requirements: Gold layer can feed into GIS tools, ML pipelines, or applications For teams managing geospatial data, whether you’re data engineers building analytics platforms or GIS professionals upgrading legacy workflows, the medallion architecture provides a systematic, scalable path forward. The alternative (managing spatial data ad-hoc) becomes increasingly untenable as volume, velocity, and complexity grow. The medallion architecture makes it possible to move at scale without moving data. Key Takeaways Bronze is raw: Get all your geospatial sources into cloud storage without transformation, leaving data at source where possible Silver is structure: Standardize projections, validate geometries, enrich with spatial relationships, and convert to cloud-native formats Gold is analysis-ready: Aggregate, optimize, and prepare data for consumption across tools and applications Iceberg is the glue: Use Apache Iceberg’s spatial extensions as your table format to manage schema evolution, lineage, and efficient spatial operations Efficiency is existential: Geospatial data demands careful format choices and architectural decisions—the medallion approach systematizes those choices The geospatial industry is shifting from moving data to moving queries. The medallion architecture is how you make that shift sustainable. Medallion Architecture FAQ What are the Bronze, Silver, and Gold layers? Bronze holds raw data as it arrives from source systems. Silver holds data that has been cleaned, validated, deduplicated, and conformed to shared schemas. Gold holds aggregated tables built for a specific use: reporting, dashboards, or model features. In a geospatial pipeline, Silver is where geometries are repaired and coordinate reference systems are standardized. Is medallion architecture ETL or ELT? Medallion architecture usually follows ELT. Raw data lands in Bronze first, and transformations run inside the lakehouse as data moves to Silver and Gold. For geospatial data, the heaviest transformations happen in Silver: geometry validation, reprojection, and conversion to GeoParquet or cloud-optimized GeoTIFF. Who came up with medallion architecture? Databricks introduced the term as part of its lakehouse guidance. The pattern builds on older staged data warehouse designs that separate raw, cleaned, and presentation layers. Does Snowflake use medallion architecture? Yes. Medallion architecture is a design pattern, so data teams apply it on Snowflake, BigQuery, Databricks, and open lakehouses built on Apache Iceberg. Each layer maps to a schema, a database, or a set of tables. What are the downsides of medallion architecture? Three layers mean more copies of data, more pipelines to maintain, and added latency between ingestion and analysis. Geospatial teams cut the copy cost by registering large rasters in place in Bronze and by storing Silver and Gold as Iceberg tables backed by GeoParquet files. How does medallion architecture work for geospatial data? Bronze registers raw vector and raster sources, often without moving them. Silver validates geometries, standardizes coordinate reference systems, and adds spatial relationships with spatial joins. Gold produces aggregated tables, map tiles, and model features. Apache Iceberg tables with native geometry types hold each layer, and WherobotsDB runs the spatial transformations between them. The Geospatial Data Engineering Associate course covers each layer in depth. Try it on your own data START BUILDING
Introducing SedonaDB and SpatialBench for Apache Sedona Posted on September 24, 2025October 3, 2026 by Damian Our role at Wherobots and as leaders in the Apache Sedona community is to help more developers, organizations, and AI systems positively transform the physical world using spatial data. In order to make the scale of transformation we envision possible, we’ve had to address significant bottlenecks in how data is stored and queried. We’re excited to celebrate the availability of SedonaDB and SpatialBench for Apache Sedona. Together they represent the next phase in our plan to accelerate innovation with spatial data and bridge the intelligence gap between AI and the physical world. Intro to SedonaDB: A modern query engine that gets spatial right SedonaDB is the first open-source, single-node analytical database engine that treats spatial data as a first-class citizen. Most analytical query engines already support general-purpose operations: filtering, joins, aggregations, and APIs for SQL or Python. But when it comes to operating on spatial data those same engines fall short: support for geometry and geography types, coordinate reference systems (CRS), spatial joins, and raster or vector operations is missing. The workaround is to bolt on an extension like PostGIS (PostgreSQL), DuckDB Spatial (DuckDB), or SedonaSpark (Spark). While powerful, extensions inherit the limits, costs, and complexities of their host systems, require extra setup and tuning, and can force builders to develop around performance and usability gaps instead of developing their ideas. SedonaDB is different. It’s for builders solving problems with physical world data. Written in Rust, it’s lightweight, blazing fast, and spatial-native. Out of the box, it provides: Full support for spatial types, joins, CRS, and functions on top of industry standard query operations. Query optimizations, indexing, and data pruning features under the hood that make spatial operations just work with high performance. Pythonic and SQL interfaces familiar to developers, plus APIs for R and Rust. Flexibility to run in single-machine environments on local files or data lakes. SedonaDB uses Apache Arrow and Apache DataFusion, and provides everything you need from a modern vectorized query engine. But it delivers the unique ability to also run high performance spatial workloads easily, without requiring extensions. Read the announcement on the Apache Sedona blog to dive in and roll up your sleeves. What led to SedonaDB? In 2020, Apache Sedona was incubated to address a significant support gap in distributed geospatial data processing. Since then, Sedona has enabled companies like Uber, Amazon Last Mile Delivery, JB Hunt, and thousands of others with geographically distributed operations or interests to build and run more efficient and effective physical operations at scale. It is widely used today to bring geospatial processing support to Apache Spark, Apache Flink, and also Snowflake. But distributed systems aren’t for everyone or the right fit for every use case, and we could do more to drive innovation in lower-scale scenarios. Accelerating innovation Many ideas are bootstrapped in no-to-low cost environments where iteration cycles are fast and low risk. There’s a lot that a developer can do today using a laptop or a single virtual machine, with modern software and LLMs—without adding a dependency that adds unwanted cost and complexity to the innovation cycle. Once their ideas are viable, they may not even require a distributed compute environment in production like Spark, or one that is “fully managed” by a vendor. So the next step was pretty clear. We had to make it easier for builders to use spatial data in no-to-low cost environments so they can iterate and positively transform the physical world, faster. We also decided to address these challenges through open-source software to maximize accessibility. Making development easier If you look around the ecosystem, you’ll notice a pattern: to get the analytical support you need for geospatial data, you deploy an analytics engine without the spatial analytics support you want, and then you bolt on what you need via an extension. Extensions are great and they serve a purpose very well. After all, SedonaSpark is an extension! But that doesn’t mean the combination of engine + extension is ideal. It requires additional setup and management, can require tuning to achieve a reasonable performance, and the underlying engine may end up becoming a bottleneck. Additionally, the development experience around the engine may be overly complex or lack support for the language you prefer, and the engine itself might introduce compute, cost, and other overhead. Working from the root causes of these challenges, along with the desire to drive more innovation, our next step became obvious. We needed to create a query engine that aids spatial data solutions development out of the box with popular pythonic and SQL interfaces and is optimized for single-machine environments. But was there enough value created by a spatial-first query engine compared to general purpose query engines with spatial extensions? Optimizing for spatial data = a better future Spatial data is no longer a minor class of data. It’s everywhere, the rate at which it’s being generated is growing every day, and its use cases span numerous industries. It streams from devices, vehicles, satellites, and drones, and derivatives from this data inform automation and decision-making across business, government, and research. The solutions being developed with it are transforming how organizations operate in the physical world. Innovation is happening today with this data despite the friction above, but the pace of this innovation can be accelerated by a query engine with internals intentionally designed to help developers realize the full potential of this data. This engine is SedonaDB, and it’s backed by an open-source community (Apache Sedona) that is committed to solving physical-world challenges through data and technology. Intro to SpatialBench: The first standard for spatial query performance “Without standards, there can be no improvement” – Taiichi Ohno. This statement from the founder of the Toyota Production System is an analogy for why we built SpatialBench. There was no standard way of measuring spatial query performance, so progress couldn’t be easily quantified or query engines objectively compared on this dimension. We built SpatialBench to establish first standards. The initial release supports 12 representative queries, ranging from simple to complex workloads, and includes a data generator for scale factors 1, 10, 100, and 1000. We hope this framework and its future versions will guide innovation that leads to a greater understanding of the physical world. We also used SpatialBench to benchmark SedonaDB, DuckDB (with its spatial extension), and GeoPandas at scale factors 1 and 10. Those results are published here. Next Steps Get Started with SedonaDB: Try it out and contribute to the roadmap. Use SpatialBench: Measure spatial query performance using a consistent standard. Watch the webinar (hosted by CNG): We’ve walked through SedonaDB and SpatialBench, and introduced Wherobots’ Startup Accelerator Program. Key takeawaysSedonaDB is the first open-source, single-node analytical database engine that treats spatial data as a first-class citizen rather than a bolt-on extension like PostGIS, DuckDB Spatial, or SedonaSpark.Written in Rust on Apache Arrow and Apache DataFusion, it ships spatial types, joins, CRS, functions, indexing, and data pruning with Python and SQL interfaces plus APIs for R and Rust, and it runs on a laptop or a single VM against local files or a data lake.Apache Sedona was incubated in 2020 for distributed geospatial processing and is used by companies such as Uber, Amazon Last Mile Delivery, and JB Hunt on Spark, Flink, and Snowflake. SedonaDB covers the no-to-low-cost, single-machine side of that ecosystem.SpatialBench is introduced as the first standard for spatial query performance: 12 representative queries from simple to complex, plus a data generator at scale factors 1, 10, 100, and 1000.Wherobots used SpatialBench to compare SedonaDB, DuckDB with its spatial extension, and GeoPandas at scale factors 1 and 10; those results are published on the linked Apache Sedona announcement, not in this post.
Advancing the Integration of Map Data via Overture’s Global Entity Reference System and Wherobots Posted on June 25, 2025October 3, 2026 by Ben Pruden Editor’s note: The Wherobots Spatial Data Catalog is now the Havasu Catalog. The general availability of the Overture Maps Foundation’s Global Entity Reference System (GERS) makes it a lot easier to build intelligence about features of our physical world. What is GERS? You can think of a GERS as a system for applying a unique key to physical features in the world. Overture assigns GERS IDs to millions of features in their data products, such as office buildings, highways, countries, rivers, schools, and more. Using GERS IDs, you can more easily join datasets to build a more complete view of physical-world features in space and measure relationships over time. Key benefits of GERS IDs include: Persistent identification: The same physical location maintains the same GERS ID over time Cross-dataset compatibility: Enable joining and enrichment across multiple datasets or providers Standardized reference: Provide a common language for location data across the geospatial ecosystem You can read more about GERS IDs in Overture’s documentation. Getting started with Overture in Wherobots At Wherobots, we’re proud to be a member of the Overture Maps Foundation, supporting the project since its formation, and as an official member since 2024. We also host and manage all of Overture’s recent datasets in the Wherobots Spatial Catalog, and they are offered at no additional cost to all customers. These datasets are production ready and at your fingertips. Select * from wherobots_open_data.overture.buildings_building New Wherobots customers can get started with Overture’s datasets in the free-to-use Community Edition, and graduate to the Professional edition when they want to join Overture data with their data in cloud storage. It’s Easy to use GERS in Wherobots A key design goal of GERS is to simplify how organizations can join and enrich their own datasets using canonical Overture datasets. Wherobots makes this vision real by providing built-in support for the GERS schema, enabling users to easily query, filter, and join to valuable geospatial datasets using GERS IDs, and through the use of spatial join predicates available in Apache Sedona and Wherobots. Whether data teams are working with parcel boundaries, retail site locations, road networks, or foot traffic telemetry, Wherobots makes it easy to enrich data using GERS with SQL or Python. Overture GERS Schema Extension Paths (not exhaustive) Stay tuned for a new tutorial from Wherobots that shows you how to use GERS IDs with Overture Places. Overtures datasets are produced using Wherobots Wherobots was founded by the original creators of Apache Sedona, the open-source engine for distributed geospatial processing. Overture now runs many of their Apache Spark and Sedona based data pipelines on Wherobots because they run up to 20x faster, at a fraction of the cost, and the Overture team benefits from the spatial expertise we offer them as a customer. The Overture team uses Wherobots’ Apache Airflow support to trigger job runs that power production of their planetary scale datasets. This feature and Wherobots compatibility with Spark and Sedona, also made it very easy for Overture to redirect where their Airflow-orchestrated jobs ran. “Overture produces a building dataset covering all buildings in the world, with 2.6B geometries and growing, that’s updated frequently. There’s a lot of data, and compute that goes into producing it and keeping it up to date. We accelerated the pipelines that produce the buildings dataset by up to 20x after we moved them to Wherobots, which required a simple redirection of our code. We retained compatibility with Apache Sedona, and the move put us into a development experience that’s made us more productive.” – Jennings Anderson, Geoscientist at Overture and Data Engineer at Meta. Redefining Standards in Open Source We’re committed to improving open standards in the geospatial data ecosystem. Wherobots is a leading contributor of spatial type support in Apache Iceberg and Parquet, the most popular open table and file formats for the cloud data lakehouse, to make it easier for companies to utilize geospatial data.We’ve partnered with Overture to improve the foundation for a scalable, versioned, and queryable world of features backed by GERS. By modernizing how geospatial data is accessed in the cloud via spatial data type support in Iceberg and Parquet, and improving accessibility and utility of open map data via GERS, we believe new use cases for spatial data will emerge to improve business operations, research, and our way of life. Join us in Building the Spatial Data Stack of the Future The launch of GERS is a big step for advancing spatial intelligence, and a leap toward a more open, interoperable data ecosystem. Wherobots is proud to be part of this journey, and we’re excited to continue supporting the Overture mission through operational pipelines, open standards, and accessible tools.If you’re building with geospatial data, we invite you to explore how Wherobots can help you take full advantage of GERS by easily joining your first party data with the GERS ID system. Next steps If you’d like a one-on-one demo from our team, you can request it here. Learn about our data mirror with Overture in their documentation. We will be publishing an example notebook that shows you how to use GERS IDs with Overture Places soon. Start Building with Wherobots Access Now Key takeawaysOverture Maps Foundation Global Entity Reference System (GERS) is generally available: stable unique keys on millions of physical features (buildings, highways, countries, rivers, schools) so datasets can be joined and tracked over time.Wherobots has been an Overture member since 2024 and hosts recent Overture datasets in the Spatial Catalog at no additional cost. Community Edition can query them; Professional Edition joins them to data in your cloud storage.Wherobots supports the GERS schema so you can query, filter, and spatially join first-party parcels, retail sites, roads, or telemetry to Overture with SQL or Python.Overture produces a global buildings dataset of 2.6 billion geometries (and growing). After moving Spark/Sedona pipelines to Wherobots—with a simple code redirect and Airflow job-run support—those pipelines accelerated by up to 20x while staying Apache Sedona compatible.Wherobots also led GEO type support in Apache Iceberg and Parquet so GERS-backed features can live in an open, versioned lakehouse rather than a proprietary silo.
The Spatial Intelligence Newsletter: GeoPandas, Isochrones, Iceberg GEO, SAM2 and More Posted on May 30, 2025October 3, 2026 by Tiffany Huynh ✨ Welcome to this month’s newsletter! Whether you’re comparing performance, exploring isochrones for spatial analysis, or implementing ML models for feature detection in satellite imagery, there’s something here for everyone. Read on for expert insights, product updates, and upcoming events you won’t want to miss! 🚨Last Call: Don’t miss your chance to try Wherobots Pro—free on AWS until May 31st! If you’ve been meaning to explore Wherobots, now’s the perfect time. Get full access to the Pro tier, including unlimited performance and scalability for production workloads. Take advantage of powerful features like raster inference, map matching, and travel isochrones—only available in the Pro tier. Latest Content 💡GeoPandas vs. Wherobots: Which is better for spatial analysis? We ran a spatial join between the Overture Buildings dataset and a postal codes dataset. Wherobots consistently outperforms GeoPandas in terms of speed—especially at scale. In fact, GeoPandas was unable to handle the largest dataset due to memory limitations. That said, GeoPandas is great for smaller datasets. Its interoperability with lots of libraries makes it a great tool for lightweight spatial workflows. Wherobots & Sedona are built for scale. It supports both Python and SQL APIs. And it’s ideal for large datasets and SQL-based workflows. Combine these tools for the best of both worlds. Flexibility for small and large datasets with access to a broader ecosystem of libraries. Find out how Sedona and GeoPandas are better together. How Wherobots Generated Drive Time Isochrones for Every POI in the US 📍Imagine if your geospatial models could reason in terms of real travel time rather than just distance. ⭐ Key highlights include: What makes the Isochrones Dataset unique and how it was built How to integrate isochrones with demographic, land use, or POI datasets A live demo of geospatial weighted regression to model service coverage, accessibility, or economic potential Why time-based modeling outperforms distance-based approaches in urban analytics Learn how Wherobots generated drive-time isochrones for every POI in the U.S., and how this data can transform planning in urban analytics, retail, transportation, and more. Geospatial Tables in the Open Lakehouse: A New Era for Iceberg & Parquet 🧊With geospatial data now supported natively in Apache Iceberg and Parquet, spatial analytics becomes as scalable and performant as any other big data workload. Join leaders from Planet, Databricks, Foursquare, and Wherobots in this insightful discussion as they explore the impact of native GEO support in Iceberg, including its core functionality, performance benefits, industry adoption, strategies for avoiding vendor lock-in, and what’s ahead for the future of geospatial data. 🔮 Product Updates What if you could detect features in satellite imagery using natural language– without relying on GPU-intensive infrastructure? With Raster Inference supporting Meta’s Segment Anything 2 (SAM2) model, you can now perform object detection and feature segmentation at scale on terabytes of satellite imagery—just by describing what you’re looking for in plain text. Instead of downloading models or manually wrangling raster data, just run everything on Wherobots’ distributed cloud compute. Query massive satellite datasets using natural language: “Find container ships” 🚢 or “Segment pickleball courts” 🎾—no GPUs, no model downloads, just cloud-native inference. Because these results are stored as Iceberg tables in your S3 bucket, you can easily join with other datasets using WherobotsDB and continue your analysis inline with 300+ spatial features and functions. Currently it supports OWLv2 SAM2, and you can go from prompt to polygons in a single SQL query. 🚀 Wherobots is now available in the AWS Europe (Ireland) region! The EU leads in innovation across climate solutions, mobility, automotive design, precision agriculture, and urban development—industries that all rely heavily on geospatial data. Yet modern cloud tools haven’t kept pace, making geospatial development costly and complex, often requiring specialized teams and infrastructure. For teams working in the EU, this launch offers lower latency support for European data residency requirements and simplified scaling—with no infrastructure headaches. Apache Sedona Community Cloud Native Geospatial Analytics with Apache Sedona (published by O’Reilly) This latest chapter dives deeper intro the integration between Apache Sedona and the Python data science stack—empowering spatial workflows that scale using tools you already know and love. If you’re following the series or want to catch up: 🔹 Chapter 1: Introduction to Apache Sedona 🔹 Chapter 2: Getting Started with Apache Sedona 🔹 Chapter 3: Working With Geospatial Data at Scale (Data Loading) 🔹 Chapter 4: Vector Data Analysis with Spatial SQL (Points, Lines, and Polygons) 💡 Latest: Sedona + PyData Ecosystem 📥 Get the hands-on guide for working with large-scale spatial data. Apache Sedona Community Office Hour Missed the most recent office hour? Watch the recording here. We covered: Broadcast join support for distributed KNN Joins The new STAC reader & OpenStreetMap (OSM) PBF reader A live demo using Sedona with Overture Maps data. 📅 Mark your calendar for the next session, where we’ll cover upcoming Sedona 1.8.0 release: GeoPandas API on Sedona Sedona on PyFlink Spark 4.0 support Sedona vectorized Python UDF GeoParquet, and Parquet Geo types Upcoming Events Connect with us at the Data+AI Summit Our team is heading to the Data+AI Summit in San Francisco, and we’d love to connect in person! 🎤 Don’t miss our session: Iceberg Geo Type: Transforming Geospatial Data Management at Scale, presented by Jia Yu (Wherobots) and Szehon Ho (Databricks) 📍 Find us at booth #E406 for free SWAG and to learn more about geospatial ETL. 🍻 Join us Wednesday evening at the Apache Iceberg Meetup, featuring Iceberg GEO and building mobility knowledge graphs with PuppyGraph. 📧 Book a meeting with us here to find a time to connect. Join us at the EO Summit If you’re heading to the EO Summit in New York, we’d love to learn about how you’re working with earth observation data. Share your use cases and let’s explore how Wherobots can help. 📧 Reach out at info@wherobots.com to schedule a time to meet! Key takeawaysThis May 2025 Spatial Intelligence Newsletter is a roundup of existing posts, product notes, and events—not a standalone technical deep dive.A limited Pro trial on AWS ran through May 31st, with Pro-only features called out: raster inference, map matching, and travel isochrones.Raster Inference added Meta SAM 2 (with OWLv2) so you can prompt object detection and segmentation on terabytes of satellite imagery in SQL; results land as Iceberg tables in your S3 bucket for use with 300+ spatial functions. No local GPUs or model downloads.Wherobots became available in AWS Europe (Ireland), aimed at EU data-residency and lower-latency workloads in climate, mobility, automotive, agriculture, and urban development.Community notes include O'Reilly Cloud Native Geospatial Analytics with Apache Sedona (through the PyData chapter), the next office hour previewing Sedona 1.8.0 (GeoPandas API, PyFlink, Spark 4.0), Data+AI Summit booth #E406, and an Iceberg Geo Type session with Jia Yu and Szehon Ho.
Wherobots, the Spatial Intelligence Cloud, is Now Available in AWS Europe Posted on April 14, 2025October 3, 2026 by Tiffany Huynh The EU is widely recognized as a world leader for climate solutions, automotive design and manufacturing, mobility systems and analysis, environmental monitoring, agriculture and precision farming, and urban development. Geospatial data is foundational to the success of these innovations. However most of the new technology is developed using tools and cloud services not optimized for geospatial development. Compared to internet data, support for geospatial data in the modern cloud environments has lagged. This technology gap has made working with geospatial data expensive, and required staffing data teams with unique expertise. Introducing Wherobots for the AWS Europe (Ireland) Region We are excited to announce that Wherobots is ready for EU native workloads. This expansion makes it possible for many EU companies to approve the use of Wherobots and adhere to their data residency requirements by processing and storing data in-region. Wherobots is the Spatial Intelligence Cloud Wherobots’ mission is to make it easy for our customers to utilize geospatial data. We are delivering on it via a cloud optimized for developing and running solutions about the physical world, at any scale. Our purpose-built, cloud native approach is enabling teams at AddressCloud and Overture Maps Foundation to accelerate their pace of innovation with geospatial data. Their workloads run up to 20x faster after migrating from popular cloud-based engines, developer productivity is boosted with the most feature complete development experience for SQL and Python, and costs are reduced, putting new solutions in reach. A Cloud Native Lakehouse Architecture The architecture of Wherobots is cloud native, and is deeply rooted in open source. Apache Sedona, the open source geospatial engine for Apache Spark, Apache Flink, and Snowflake, is 100% compatible with Wherobots. Users can easily lift and shift their Apache Sedona based applications into Wherobots with zero code changes. Wherobots is also one of the leading companies bringing GEO support into popular open file and open table formats like Parquet and Iceberg, and uses these formats by default. That way, you can deploy various engines on your data, benefit from the advantages of a Lakehouse engine such as ACID transactions and table versioning, without locking your data into proprietary vendor siloes, or inelastic solutions that couple storage with compute. Getting Started Getting started is easy. To use Wherobots within the AWS Europe (Ireland) region, get srated with the Professional Edition on the AWS Marketplace. Create a notebook and explore one of many examples designed to help you realize what you can create using SQL and Python. There’s no infrastructure to manage. Teams just use and pay for Wherobots usage on-demand via Wherobots Spatial Units, which reflect the amount of serverless computation consumed. If there are other clouds or regions that you’re interested in using beyond the ones we currently support, please reach out to us at product@wherobots.com or fill out this form here. You can read more about Wherobots on the website or by exploring our product documentation. Create a Pro tier account on AWS Get Started Key takeawaysWherobots is available for EU-native workloads in the AWS Europe (Ireland) region, so teams can process and store data in-region to meet data-residency requirements.The launch targets EU strengths in climate, automotive, mobility, environmental monitoring, precision agriculture, and urban development—areas where cloud tools have lagged for geospatial compared with internet data.Customers such as AddressCloud and Overture Maps Foundation are cited as running workloads up to 20x faster after migrating from popular cloud engines, with a feature-complete SQL and Python experience and on-demand Spatial Units billing.Architecture is a cloud-native lakehouse: 100% Apache Sedona compatible (lift-and-shift with zero code changes), defaulting to open Iceberg and Parquet/GEO formats so storage stays decoupled from compute.Start on Professional Edition via AWS Marketplace, open a notebook, and run examples. There is no infrastructure to manage. Requests for additional clouds or regions go to product@wherobots.com.
Spatial Intelligence Newsletter: Location Intelligence w/ Isochrones, Overture Places, Cloud-Native Geospatial, Iceberg and More Posted on April 10, 2025October 3, 2026 by Tiffany Huynh Welcome to the April edition of the Spatial Intelligence Newsletter! This month, we’re covering the benefits of using Apache Iceberg, spatial joins, cloud-native geospatial, and new product updates like isochrones to help you make better location-based decisions. What does it all mean, and how can it help you increase data productivity? Check it out here! 👇 ⏰💰Hurry, time is running out! We’re currently offering a FREE $400 credit when you subscribe to the Professional edition of Wherobots, which includes exclusive features like GeoAI with WherobotsAI Raster Inference, map matching for cleaning messy GPS data, new travel isochrones for better location-based decisions, and the ability to bring your own cloud storage, just to name a few. In addition, with the release of our drive-time isochrones, we’re now offering free access to Overture Places data—enriched with drive-time isochrones across every location in the U.S.—through our Pro tier data catalog. And we will be maintaining that dataset with every release in the future, so you’ll be able to use it going forward for location intelligence. There’s no obligation to get started, so be sure to take advantage of this (it’s like free money). Offer ends on May 31st, so don’t wait! Latest Content Benefits of Apache Iceberg for geospatial data analysis 🧊 Apache Iceberg support for GEO data brings a significant modernization for geospatial data and solutions. This support makes it easier for you to bring geospatial data into an open data architecture that decouples compute and storage and lower your costs. By adopting Iceberg in a data lake, you’re enabling your team to leverage the right tool for the job without needing to worry about locking your data into a vendor or a database solution that doesn’t scale. Additionally, traditional file formats and row-oriented databases struggle when scaling beyond a million features, often performing poorly or only accommodating data that fits comfortably in memory. 😩 Iceberg, built on Parquet, solves this with lightning fast reads, scalability for larger-than-memory datasets, and developer friendly features like DML operations. Plus with added capabilities like versioning and time travel, users can query both current and historical data seamlessly. 🔍 Follow along this post to learn how to use Apache Iceberg with Sedona and find out how these features benefit spatial computations. Cloud-Native Geospatial: More Than Just Big Data e💡We had a very insightful discussion with Amy Rose (CTO) from Overture Maps and Eshwaran Venkat (CTO & Co-Founder) from Dotlas on cloud-native geospatial technology. Here are some highlights: Cloud-native geospatial is not just for big data; it’s more accessible than you might think. You should be able to work with spatial data the way you work with any other data type. Increasing deliverability and breaking down data silos: Non-spatial communities can now work with spatial data. Compute systems that make the process more scalable, accessible, elastic, and cost-efficient. How Dotlas and Overture Maps are optimizing their data pipelines, achieving performance gains, and improving cost efficiency. Spatial Joins at Scale: Unlocking Advanced Geospatial Analytics with Wherobots 🌎🤝 Spatial joins are essential for geospatial data analysis, but it can be slow or computationally expensive when working with large-scale datasets. Follow along in this tutorial as we walk through how easy and cost-effective it is to: Join datasets using spatial predicates like ST_Intersects to combine facilities with administrative boundaries and efficiently find the k-nearest neighbor with the ST_AKNN function. Apply spatial filters and improve performance through strategies like partitioning by geohash Take your geospatial data analytics to the next level and ensure spatial joins aren’t a bottleneck in solving your business challenges. Apache Sedona Sedona Success Story: Optimizing ETL pipelines at scale with Comcast 📊 Some of the challenges that Comcast was trying to overcome was data volume and repeatability. That’s why David Buchanan, GIS Architect, turned to Apache Sedona, which allowed him to reduce processing times from 5 hours to 30 minutes compared to GeoPandas. Watch the recording to learn more. Apache Sedona Office Hours 😎 We just released Sedona 1.7.1, with some new features : SQL interface for GeoStats (ST_DBSCAN, ST_GLocal, ST_LocalOutlierFactor) Broadcast join support for distributed KNN Join STAC catalog & OpenStreetMap (OSM) PBF reader New ST functions like ST_RemoveRepeatedPoints If you missed the office hour, check out the recording to learn more about the latest release. And don’t forget to mark your calendar for the next office hour! 🗓️ Product Updates Overture Places with Isochrones Dataset: Accelerate accessibility analysis with a ready-to-use dataset containing pre-calculated 5, 10, 15, and 20-minute driving isochrones for millions of US Overture Places (pro+). New ST Isochrones functions: Make data-driven decisions on logistics, site selection, and market reach using Wherobots’ travel isochrone functions in SQL or Python (pro+). Audit Logs: Admins gain enhanced security and accountability insights by using Wherobots’ detailed, exportable audit logs to track key Organization actions and system events (pro+). STAC Reader: Simplify workflows and accelerate queries by loading STAC geospatial datasets directly into Sedona DataFrames in Wherobots (OSS & community+). Job Run Monitoring: Visually track job execution, analyze resource usage, and manage runs directly within Wherobots for enhanced control and optimization (pro+). Idle Timeout for Notebooks: Gain control over notebook runtime costs and resource usage with customizable idle timeouts that automatically terminate inactive notebooks (community+). 🆓 Both the Overture Places with isochrones dataset and isochrone functions, as well as the audit logs and job run monitoring, are available exclusively in the Pro tier. Take advantage of the free trial (ending soon!) to try these features and see how they can help solve some of the bottlenecks you might be facing when working with spatial data. Upcoming Events Geospatial Tables in the Open Lakehouse: A New Era for Iceberg and Parquet Wednesday, May 5 at 9AM PT | Virtual It’s easier than ever to work with geospatial data, with Iceberg and Parquet now offering powerful solutions for both geospatial experts and non-spatial professionals. Join this livestream with leaders from Foursquare, Databricks, Planet, and Wherobots as they discuss the historical challenges of handling spatial data, bridging the gap, and future adoption of these advancements. Apache Sedona + Iceberg GEO Meetup Monday, May 12 at 5:00PM PT | San Francisco, California Join us for a fun and informative evening as we explore Apache Iceberg’s new native geospatial support, designed to solve major challenges in managing geospatial data at scale. This will be a great opportunity to connect with professionals in the field to learn about the latest developments in spatial data, as well as exciting projects people are working on.🌟 Featured speakers: Jia Yu, Co-Founder and Chief Architect, Wherobots Matt Forrest, Director of Customer Engineering and PLG, Wherobots Yingjun Wu, Founder and CEO, RisingWave Labs CNG Conference April 30 – May 2 | Snowbird, Utah We’re excited to attend the upcoming CNG Conference! Be sure to check out these sessions: Day 1 1:15pm-2:45pm | Workshop: Interfacing with Cloud-Native Overture Data and the GERS Ecosystem – Sean Knight Day 2 9:45am-11:15am | Track 2: Introducing geospatial support in Apache Iceberg – Matthew Powers 11:45am-1:15pm | Extract insights from satellite imagery at scale with WherobotsAI – Damian Wylie 4:30pm-5:00pm | Plenary Panel: Builders Panel – Mo Sarwat 👥 If you’ll be at the conference, we’d love to meet with and chat about how you’re working with geospatial data. Feel free to reach out if you’d like to schedule a time to connect! Key takeawaysThis April 2025 newsletter roundup points to Iceberg-for-geospatial, cloud-native geospatial with Overture and Dotlas, spatial-join tutorials, and product launches—not original benchmark figures.A Professional Edition offer included a free $400 credit through May 31st, covering GeoAI Raster Inference, map matching, travel isochrones, and bring-your-own cloud storage.Pro catalog added Overture Places enriched with 5-, 10-, 15-, and 20-minute U.S. driving isochrones, plus new ST_Isochrone functions in SQL and Python. Other Pro items: audit logs and job-run monitoring. STAC reader and notebook idle timeout landed for Community+.Sedona 1.7.1 shipped SQL GeoStats (ST_DBSCAN, ST_GLocal, ST_LocalOutlierFactor), broadcast KNN joins, STAC and OSM PBF readers, and ST_RemoveRepeatedPoints. A Comcast success story cut ETL from 5 hours to 30 minutes versus GeoPandas.Upcoming at the time: an Iceberg/Parquet livestream on May 5, a San Francisco Iceberg GEO meetup on May 12, and CNG Conference sessions in Snowbird April 30-May 2.
The Spatial Intelligence Newsletter: Map Matching, Spatial Joins, ML for EO, Cloud-Native Geospatial and More Posted on March 14, 2025October 3, 2026 by Tiffany Huynh 👋 Welcome back to the latest edition of the Spatial Intelligence Newsletter! We’ve been busy brewing up some exciting things here at Wherobots, so we have plenty of new updates and content to share! Latest Content Don’t Let Messy GPS Slow You Down. The Fastest Way to Clean Up Messy GPS Data – And Save Money Raw GPS data is messy. 😵💫 Noisy signals, lost connections, and inaccuracies make it hard to extract valuable insights. Imagine using your GPS to get to your location, only to find it telling you to drive over water instead of the road (personally, I’ve even had the map tell me to walk on water 🌊🚶🏻♀️). Wherobots’ map matching corrects trajectories by aligning them with real-world road networks (❌no more walking on water! ), all while delivering unmatched accuracy and performance (and saving money!). Apache Iceberg and Parquet now support GEO– A Huge Step Forward for Cloud Native Geo Geospatial data has always been thought of as a second class citizen because of what modernized the data ecosystem of today, leaving geospatial data mostly behind. But that’s no longer the case. Thanks to the efforts of the Apache Iceberg and Parquet communities, both Iceberg and Parquet now support geometry and geography (collectively the GEO) data types! 🎉 What does this mean? With native geospatial data type support in Apache Iceberg and Parquet, you can seamlessly run query and processing engines like Wherobots, DuckDB, Apache Sedona, Apache Spark, Databricks, Snowflake, and BigQuery on your data. All the while benefitting from faster queries and lower storage costs from Parquet formatted data. 💨 Exploring design and key features to enhance spatial data workloads with Iceberg GEO With Apache Icerberg and Parquet now supporting GEO types, this helps improve the economics of utilizing geospatial data in end solutions.This advancement allows organizations to create higher-value, lower-cost products and achieve faster results over time. Let’s take a closer look at these GEO data types in Iceberg, exploring their design, key features, and implementation considerations. Learn how leveraging these features with Apache Sedona and Wherobots can enhance cost performance and data governance, ensuring the best possible experience for spatial data workloads. 📈 Optimizing Earth Observation Models for Production with ML Model Extension What are the challenges of applying AI to geospatial problems? 🤖Join panel speakers from Wherobots, Radiant Earth, CRIM and Terradue as they discuss how this challenge led to the development of an open, portable solution for describing computer vision models trained on overhead imagery. Learn about the MLM STAC Extension, its use cases, and why model developers should adopt it, along with Raster Inference– a serverless computer vision solution that extracts valuable insights from aerial imagery. 🌎 Getting Started With Wherobots Interested in getting started with Wherobots, but unsure of where to begin? Here are some helpful resources. 👇 Wherobots 101: Mastering Scalable Geospatial Data Processing Want to take your geospatial analytics to the next level? Whether you’re just starting out or already working with spatial data, learn how to leverage valuable tools and workflows in Wherobots Cloud to analyze, visualize and interpret geospatial datasets. From setting up your account to mastering advanced analytics, this session is a helpful guide to set you up for success! Wherobots 102: Reading and Processing Cloud Native Geospatial Data Learn how to efficiently load, manage and analyze raster and vector data in Wherobots’ hosted environment. Whether you’re working with massive geospatial datasets or looking for optimized workflows to write and query GeoParquet and Cloud-Optimized GeoTIFFs (COGs), this video will equip you with the tools and techniques to scale your geospatial analysis. Working with Foursquare Places Data Which neighborhood in San Francisco has the most coffee shops? Dive into the Foursquare Open Places dataset, a free and open dataset providing 100M+ global places of interest, with our latest tutorial. ☕ You’ll be able to query using Spatial SQL, subset the data for a specific region, search for specific businesses or places, and aggregate locations by geography. By the end of this tutorial, you’ll have a choropleth map showing the number of coffee shops, sorted by neighborhood. Apache Sedona Community Sedona Success Story: Optimizing ETL pipelines at scale with Comcast 🚀 Is scaling your ETL pipeline a priority? Discover how Comcast successfully achieved this by using Apache Sedona, all while boosting productivity and improving the quality of their network operations. 🌐 Learn how Apache Sedona reduces vendor lock-in. Understand why it outperforms tools like GeoPandas and PostGIS. See how it improves the ability of the Xfinity network team to optimize their network operations through a global view of performance quality and degradation. Find out how it integrates seamlessly with Apache Spark and other distributed engines. O’Reilly: Cloud Native Geospatial Analytics with Apache Sedona – Navigating Large-Scale Spatial Data We know that handling large-scale spatial data can be daunting, which is why we’ve designed this guide to simplify geospatial data. This will help boost your spatial analytics expertise and transform the way you work with geospatial data! 💪 Our newest chapter, focusing on vector data analysis using spatial SQL, is now available. If you’ve already accessed the previous chapters, be sure to check your inbox (on a separate email) for the latest one! 📧 Engage with the Community Through Sedona Office Hours We host monthly office hours to bring you the latest news and updates to Apache Sedona. Mark your calendars for the next one. Even if you can’t make it, we’ll send you the recording and slides to make sure you don’t miss anything that might be helpful to you. 🤝 Upcoming Events Spatial Joins at Scale: Unlocking Advanced Geospatial Analytics If you’ve ever struggled with Spatial Joins (you know who you are), then this is the one to join (pun intended, courtesy of Matt Forrest 😎)! Learn how to seamlessly integrate Python and Wherobots to perform advanced spatial joins and analyses on geospatial data. Gain practical skills and best practices for processing and visualizing spatial data at scale. Don’t miss this opportunity to boost your spatial analytics expertise and transform how you work with geospatial data. Fireside Chat with Overture Maps and Dotlas on Cloud-Native Geospatial: More Than Just Big Data How is cloud-native geospatial reshaping the way organizations interact with spatial data? ☁️🌎 It prioritizes flexibility, changes how data consumers connect, removes friction, and unlocks new possibilities. Join us, alongside Amy Rose from the Overture Maps Foundation and Eshwaran Venka from Dotlas, as we explore how modern approaches enable scalability across various compute infrastructures, eliminate the need to move massive datasets, and allow users to work with data wherever they are—whether locally or in the cloud. Hear about where geospatial technology is headed. This is a conversation you definitely don’t want to miss! Getting Started 🆓 Getting started with Wherobots is easy. If you haven’t already, create a free account and dive in. If you’re looking to take your geospatial analytics to the next level—whether it’s full access to open datasets, map matching, or raster inference—try the Pro tier for free. Get started with Wherobots Try Now Key takeawaysThis March 2025 newsletter is a content and community roundup covering map matching, Iceberg/Parquet GEO, the MLM STAC Extension, getting-started videos, and upcoming events.It highlights Wherobots map matching for snapping noisy GPS to real road networks, and native geometry/geography types in Apache Iceberg and Parquet so engines such as Wherobots, DuckDB, Sedona, Spark, Databricks, Snowflake, and BigQuery can share one copy of the data.A companion technical post on Iceberg GEO design is featured, along with a panel on the MLM STAC Extension and Raster Inference for describing and running computer-vision models on overhead imagery.Getting-started pointers include Wherobots 101 and 102 sessions and a Foursquare Open Places tutorial (100M+ global POIs) that maps San Francisco coffee shops by neighborhood.Community: Comcast Sedona ETL story, O'Reilly Sedona vector-SQL chapter, monthly office hours, a spatial-joins webinar, and a fireside chat with Overture Maps and Dotlas on cloud-native geospatial.
Apache Iceberg and Parquet now support GEO Posted on February 11, 2025October 3, 2026 by Ben Pruden Geospatial data isn’t special anymore, and that’s a good thing. Geospatial solutions were thought of as “special”, because what modernized the data ecosystem of today, left geospatial data mostly behind. This changes today. Thanks to the efforts of the Apache Iceberg and Parquet communities, we are excited to share that both Iceberg and Parquet now support geometry and geography (collectively the GEO) data types. Geospatial challenges Geospatial data has been disconnected from the broader data ecosystem that modernized from open file formats like Apache Parquet, and open table formats like Apache Iceberg, Delta Lake, and Apache Hudi.The benefits of these cloud-native open file and table formats fueled widespread adoption of data lake and lakehouse architectures. Organizations moved away from the use of expensive proprietary systems, away from data siloes that coupled compute with storage and didn’t scale, and away from formats that locked them in and stifled innovation. Relative to legacy options, these cloud-native formats fundamentally change how data is stored, managed, and accessed. This in turn lowers costs, increases agency, and unlocks innovation over time. But because geospatial data was different, which led to a number of technical challenges, it wasn’t supported by these formats from the start. As a result developers building solutions with geospatial data struggled with fragmented formats, proprietary file types, and data siloes – making solutions harder and costlier to build. The silos will break down With native geospatial data type support in Apache Iceberg and Parquet, you can seamlessly run query and processing engines like Wherobots, DuckDB, Apache Sedona, Apache Spark, Databricks, Snowflake, and BigQuery on your data. All the while benefitting from faster queries and lower storage costs from Parquet formatted data.These changes improve short and long term economics for geospatial solutions. Organizations will have a new freedom to innovate with a lower cost, highly interoperable architecture. They get to choose the best tool for the job over time without having to shuttle data between systems. Their costs reduce, productivity improves, innovation accelerates, and the playing field is leveled with respect to who can provide the best solution for their data. The legacy siloes will break down, just like they’ve done for non-geospatial data. And most importantly, these changes will lead to new innovation about our physical world. Benefits of Iceberg and Parquet These changes make geospatial solutions based on a data lake a lot more attractive. Here are a few benefits. Iceberg and Parquet alone don’t separate compute from storage, but together they make it possible to utilize low cost data lake storage, along with multiple independent high performance computing solutions for different use cases ACID transactions and data versioning enable the use of multiple compute engines without conflicts Time travel allows tracking of data changes over time Query performance is higher from features like column pruning, row-group filtering, and fast file access Open data formats minimize vendor lock-in Geospatial data will be supported across a broader ecosystem of tools and services And many more… See how these formats fit into a full Bronze, Silver, Gold pipeline in The Medallion Architecture for Geospatial Data Grassroots efforts made this happen These changes were the result of grassroots initiatives, investment, and influence from community members at Planet, CARTO, Wherobots, and many others across the Cloud Native Geospatial community. This includes GeoParquet, which was a grassroots project and an extension of Parquet that proved its worth through use and popularity, countless meetups, and discussions. And we also want to give credit to the Iceberg community for working with members of the Wherobots team, to bring a solution forward while also influencing the Parquet community to make a GEO native data type.While Iceberg and Parquet communities led with support for GEO data types, we welcome compatibility and support for GEO data types in all cloud-native formats, including Apache Hudi and Delta Lake. Thoughts from Szehon Ho, Apache Iceberg PMC Member“The long-awaited incorporation of geospatial data types in the Iceberg V3 spec extends a core theme of Iceberg as a project to provide a universal ‘shared warehouse storage’ across many engines and users, and will now allow this huge, growing ecosystem to work on the same geospatial data as well, unlocking many exciting use cases. It is also a demonstration of Iceberg community’s willingness to take the time and ‘do hard things’, engaging in months of very active discussions across companies and OSS communities, finally reaching consensus on a spec that supports the largest variety of use cases in the fast-evolving geospatial data domain.” Thoughts from Chris Holmes, co-creator of GeoParquet“The community developed and rallied behind GeoParquet to make geospatial data in Parquet fully interoperable and to let the geospatial world tap into all the advantages the big data world has been getting from Parquet. I’m very excited to see Parquet and Iceberg formally support geospatial types, and look forward to the acceleration in geospatial innovation that these changes will activate across industries and for our planet.” Looking ahead Committers are already working to bring support for these changes into Apache Sedona, and will notify the community as they are introduced. At Wherobots, we’ve supported these GEO data types in Havasu (our Iceberg fork) which we built to enable geospatial lakehouse architectures with Wherobots, along with GeoParquet. We’ve begun developing native support for Iceberg and Parquet into how Wherobots operates on customer data, and will put our full support behind these native formats moving forward. WherobotsDB now runs natively on these formats, delivering 3x faster query performance.To learn more about the reasoning behind the Iceberg GEO types design, the trade-offs we navigated, and what it all means for implementers, please read our follow-up blog: Iceberg GEO: Technical Insights and Implementation Strategies. If you need support throughout your journey adopting and utilizing these cloud-native formats for geospatial use, reach out to Apache Iceberg on Slack or Apache Sedona on Discord. Watch this livestream with leaders from Foursquare, Databricks, Planet, and Wherobots as they discuss the historical challenges of handling spatial data, bridging the gap, and future adoption of these advancements. Sign up for our newsletter to stay up to date with everything we are doing to enable the spatial community to embrace the modern geospatial lake-house. Key takeawaysApache Iceberg and Apache Parquet now natively support geometry and geography (GEO) data types, so geospatial data can live in the same open lakehouse formats as the rest of the modern data stack.Engines including Wherobots, DuckDB, Apache Sedona, Apache Spark, Databricks, Snowflake, and BigQuery can run on one copy of the data, with Parquet faster queries and lower storage cost and Iceberg ACID transactions, versioning, time travel, and column/row-group pruning.The change is framed as breaking geospatial silos: no more proprietary file types and compute-coupled storage as the default path for spatial solutions.It was a grassroots effort across Planet, CARTO, Wherobots, GeoParquet, and the Iceberg community (quotes from Iceberg PMC member Szehon Ho and GeoParquet co-creator Chris Holmes). Hudi and Delta Lake are invited to add GEO as well.Wherobots already supported these types in Havasu (its Iceberg fork) and GeoParquet; WherobotsDB now runs natively on Iceberg and Parquet GEO with 3x faster query performance, and Sedona committers were adding support at the time of writing.
Wherobots 2024 accomplishments, and what’s on-deck in 2025 Posted on January 23, 2025October 3, 2026 by Damian Editor’s note: The Wherobots Spatial Data Catalog is now the Havasu Catalog. Introduction 2024 was a transformative year for Wherobots. Our mission to revolutionize how geospatial data is used took significant strides forward, positively impacting our customers and industry. Over the past year, we more than tripled the size of our team and successfully closed a $21.5M Series A funding round. We expanded accessibility to Wherobots’ industry-leading geospatial query performance, integrated Wherobots into the native AWS buying experience, and unveiled groundbreaking features like Raster Inference, Map Matching, and GeoStats—empowering users to create scalable geospatial solutions like never before. Our Mission Before founding Wherobots, co-founders Mo and Jia identified critical challenges limiting the potential of geospatial data. These stemmed from how geospatial data was traditionally stored, formatted, and processed, and made this data incredibly painful to utilize, particularly at scale. Over the recent decades, data and analytics investment was mostly directed towards solutions for internet data. However compared to internet data, geospatial data is a lot more complex, which makes it harder to query. It’s polygons representing land and buildings, GPS trajectories, satellite and drone imagery, weather data, and more—all tied to Earth’s imperfect spherical surface. And querying this data generally means you need to filter and join it with other datasets (geo or non-geo). Due to this complexity, existing cloud analytics engines built for structured internet data struggle to efficiently run spatial queries at scale. They also miss features necessary to prepare this data, they lack features that make solution development productive, and simply cannot compute spatial results with high precision. As a result, solutions based on geospatial data are expensive, or otherwise shelved. We are addressing these challenges. By reducing the cost and effort to build with geospatial data, Wherobots will enable a new wave of innovation for the physical world. This will drive breakthroughs in products, business operations, science, government, and make a positive impact on our climate. Our mission is simple yet ambitious: make geospatial data easy to use. Here’s what some of our customers have to say about how we’re helping them achieve their missions. Customer Highlights AddressCloud Enabling insurers to calculate geographic risk with precision “Wherobots runs our compute operations that used to take hours or days to complete, in minutes. As we provide perils information (flood, fire, etc) to insurers at the property level, we particularly appreciate the ability to be able to run combined vector/raster analysis, without having to previously transform the raster data into vector format or some other format.” – John Powell, Senior Geospatial Data Engineer at AddressCloud Overture Maps Foundation Creating next-generation map products with scalable, open map data “Overture produces a building dataset covering all buildings in the world, with 2.3B geometries and growing, that’s updated frequently. There’s a lot of data and compute that goes into producing and keeping it up to date,” said Jennings Anderson, Geoscientist at Overture and Data Engineer at Meta. “We accelerated the pipelines that produce the buildings dataset by up to 20x after we moved them to Wherobots, which required a simple redirection of our code. We retained compatibility with Apache Sedona, and the move put us into a development experience that’s made us more productive.” – Jennings Anderson, Geoscientist at Overture and Data Engineer at Meta Why Wherobots Stands Out Several recurring themes highlight why customers choose Wherobots: Unmatched performance and cost efficiency: Wherobots delivers up to 20x better spatial join performance compared to modern cloud data engines, at a fraction of the cost. Ease of innovation: Wherobots makes it easy to build solutions with raster (e.g., satellite imagery), vector (e.g., geometry, geography) data, and your first party data regardless of scale. Modern cloud architecture: Wherobots is fully compatible with Apache Sedona, and runs seamlessly on data lakes with support for Apache Iceberg and Apache Parquet. 2024 Milestones Funding & Market Validation In 2024, we raised $21.5M in a Series A round led by Felicis, with support from Wing Venture Capital, Clear Ventures, JetBlue Ventures, and P7 Ventures. This funding reflects confidence in our mission and the massive market opportunity for geospatial solutions in the cloud. Team Growth The Wherobots team—the “Botsters”—tripled in size this year. While engineering saw the most growth, we also built out go-to-market, marketing, and product teams and are actively scaling our sales team. As we head into 2025, we’re actively hiring for roles across the company to support our expanding vision. Product Innovations We launched several key features in 2024 that expanded the boundaries of geospatial data solutions. *The features noted with an are only available in the professional or enterprise edition of Wherobots.**** Cloud Native A pay-as-you-go offering on the AWS Marketplace makes it easy to subscribe and pay on-demand using AWS Marketplace billing. A storage integration for Amazon S3, to quickly and securely integrate with first or third party data. Security and Access SAML Single Sign-On, makes logging in simple, secure, and seamless for users in companies with centralized login management systems. The Spatial SQL API, Typescript and Python SDKs, and a JDBC driver make it possible to query WherobotsDB using popular or custom query interfaces. Continuous improvement of internal security and service availability. Open Data Architecture The first version of the Spatial Catalog (known as Havasu, with core functionality soon to be merged into Apache Iceberg). Accelerating Geospatial Solution Development Raster Inference, to easily extract insights from satellite imagery at scale using SQL. (We’re hosting an upcoming panel discussion with an incredible lineup of speakers to discuss the MLM STAC Extension and Raster Inference, with a focus on optimizing Earth observation models for production. Learn more and save the date here.) Distributed Map Matching is a purpose built algorithm for snapping GPS trajectories to known segments like roads, with high performance at-scale. GeoStats: a geostatistics suite designed for scale, performance, and streamlining solution development. Support for K-nearest neighbor joins (exact and approximate) to efficiently query for geospatial neighbors at-scale. Vtiles, a vector tile solution purpose built for creating vector tiles at scale with high performance. Many new vector (ST) and raster (RS) functions to accelerate the developer productivity. Continuous improvement of spatial and non-spatial query performance to reduce cost and make workloads more compute efficient (reducing climate impact). Automation Job Runs integrated with Apache Airflow, to make it easy and familiar to automate new and existing processing workflows. Service Principals, enable authentication and automated usage of Wherobots, decoupled from the tenure or privileges of human users. Looking Ahead In 2025, we plan to bring Wherobots Cloud to the EU market with support for the AWS Europe (Ireland) region, and achieve the SOC 2 Type 2 certification (currently in progress). We’ll continue to focus on: Making Earth observation data easier to utilize. Enhancing developer productivity and experiences. Improving query engine performance and data compatibility. Strengthening service availability and support for customers. Delivering new administrative controls and observability. Ready to Build? We are currently offering a 30-day free trial covering up to $400 in usage via the AWS Marketplace. Getting started is easy. There are many example notebooks for various geospatial use cases that you can explore and run without any coding experience required. Not only do the notebooks help you get started, but we also see most of our customers use these notebooks as references for the solutions they end up building. Join the Mission Motivated by our mission? Join our growing team—visit our careers page for open roles. You can also share feedback at feedback@wherobots.com or contact me directly at damian@wherobots.com. Try Wherobots Pro Get Started Key takeawaysThis January 2025 year-in-review covers 2024 milestones rather than a single product launch: the team more than tripled, and Wherobots closed a $21.5M Series A led by Felicis, with Wing, Clear Ventures, JetBlue Ventures, and P7 Ventures.AddressCloud cut property-level flood/fire peril jobs from hours or days to minutes, including combined vector/raster analysis without converting rasters first. Overture global buildings pipeline (2.3 billion geometries in this post) accelerated up to 20x after a code redirect onto Wherobots.The company cites up to 20x better spatial-join performance than modern cloud data engines, full raster+vector support, Apache Sedona compatibility, and Iceberg/Parquet lakehouse architecture.2024 product launches included AWS Marketplace pay-as-you-go, S3 storage integration, SAML SSO, Spatial SQL API plus TypeScript/Python SDKs and JDBC, Spatial Catalog/Havasu, Raster Inference, distributed map matching, GeoStats, KNN joins, VTiles, Airflow job runs, and service principals. Starred items are Professional/Enterprise only.2025 plans stated here: AWS Europe (Ireland), SOC 2 Type 2 (then in progress), easier Earth observation, developer experience, query performance, availability, and admin/observability. A 30-day AWS Marketplace trial covered up to $400 in usage.