Wherobots for QGIS: Cloud-Scale Spatial Data, from Wherobots Labs Posted on October 2, 2026October 3, 2026 by Matt Forrest Wherobots for QGIS is a new plugin that connects QGIS to Wherobots Cloud. From a panel inside QGIS, an analyst can run spatial SQL against WherobotsDB, load the results as a map layer, push a local layer up to a Wherobots Iceberg table, and pull raster data into the canvas. Why Wherobots for QGIS: Eliminating the partitioning tax GIS analysts hit the same wall in city after city. A workflow that runs fine on a neighborhood runs fine on a district. At state scale, it crashes or runs for hours. At country scale, the analyst has to write a loop, tile the data by county, process each tile, handle the features that fall on the boundary, reassemble the output, and discover the seam artifacts two days later. That loop is not the analysis. It is overhead that the existing tools and databases impose. This is the partitioning tax. It is not a performance problem. It is an architectural one. Desktop GIS and single-node databases process data on one machine. When the data outgrows that machine, the analyst absorbs the cost: split the extent, manage the tiles, merge the results, debug the edges. Every time a parameter changes, run it again. Testing at city scale and assuming it generalizes is not a workflow, it is a bet. The same constraint blocks AI coding assistants. Tools like Claude Code and Cursor are fluent at writing spatial SQL for a single-pass query. Ask one to scale a workflow from a county to a state and it generates a partition loop with no overlap buffer and no dedupe logic. The model does not know about edge effects. The developer still has to do the work. Wherobots for QGIS is a new plugin that removes the wall on the QGIS side of the stack. An analyst writes spatial SQL in a panel inside QGIS. WherobotsDB executes it in the cloud across the full extent, whether that extent is a city block or a continent, with no partition logic in the query and no tile management by the analyst. The result loads as a map layer. What the Wherobots QGIS plugin does The panel has three tabs, plus a connection screen where you enter an API key and pick a region and runtime. SQL Query Write spatial SQL in a free-form editor, or switch to Browse Tables and pick a table from the catalog. Set a row limit. Check one box to restrict results to the current map extent. Results land in the QGIS project as a layer, ready for styling, joins, or export. Upload Select a vector layer in your project, give it a destination table name, and send it to Wherobots. The plugin creates the table and inserts the features in batches while QGIS stays responsive. Raster Run raster SQL with RS_ functions against raster tables. Use the current map extent as the bounding box. A query that returns RS_AsGeoTiff loads as a raster layer. A query that returns RS_Values gives you pixel values as a table. The plugin runs on QGIS 3.22 through QGIS 4.0, covering both Qt5 and Qt6 builds. One copy of the data across QGIS, notebooks, and AI coding assistants Wherobots has one job in this picture: hold the data and run the compute. The interface should follow the person doing the work. Today the same WherobotsDB tables are reachable from a Python notebook, a SQL client, a Python Job, the Wherobots CLI, and from AI coding assistants like VS Code, Cursor, and Claude Code through the Wherobots MCP Server. QGIS joins that list. The tables an analyst opens on the canvas are the same tables an agent queries through MCP. One copy of the data. No exports. This is what making spatial accessible means in practice. An analyst who knows QGIS and knows SQL can now run a spatial join across hundreds of millions of records and see the result on a map, without standing up a cluster or learning a new interface. WherobotsDB, built by the original creators of Apache Sedona, handles the distribution, indexing, and coordinate systems underneath. How to install the Wherobots plugin in QGIS Download the plugin zip from the Releases page and install it in QGIS under Plugins → Manage and Install Plugins → Install from ZIP. Install the wherobots-python-dbapi package into QGIS’s bundled Python. The README has a two-line snippet for the QGIS Python Console that works on every platform. Open the Wherobots panel from the toolbar or Web → Wherobots, paste your API key, and connect. You need a Wherobots Cloud account. Start a free trial if you do not have one. The wherobots_open_data catalog, including Overture Maps, is a good first query. Introducing Wherobots Labs Wherobots for QGIS is the first project published under Wherobots Labs. Labs is where Wherobots ships the projects that live at the edge of the platform: connectors, plugins, adapters, example pipelines, and skills for AI coding assistants. They are built by the solutions architects and customer engineers who work with Wherobots customers every day, and by community contributors. Every Labs repository follows the same rules. It is public on GitHub under the wherobots organization with a labs- prefix. It is Apache 2.0 licensed unless noted otherwise. It has a README with a plain statement of support expectations, a CONTRIBUTING.md, and GitHub Issues enabled. It gets a security review before first release, and security fixes continue under the standard Wherobots security policy. What a Labs project does not carry is a production SLA. You are responsible for confirming a project fits your use case before you rely on it, and you are free to fork it. A Labs project that earns adoption can graduate to a supported product surface. More Labs projects are on the way. Browse them at here and open an issue or a pull request on any of them. If you use QGIS and you have a dataset that stopped fitting on your machine, install the plugin and tell us what you build. Key takeawaysWherobots for QGIS connects QGIS to Wherobots Cloud, so an analyst runs spatial SQL against WherobotsDB without leaving the map.The compute runs in the cloud. QGIS stays the place to view the answer. One copy of the data, no exports.Three tabs: SQL Query, Upload, and Raster. The plugin runs on QGIS 3.22 through QGIS 4.0, on both Qt5 and Qt6.The same WherobotsDB tables reach Python notebooks, SQL clients, the Wherobots CLI, and AI coding tools through MCP. QGIS joins that list.Available today from the QGIS Plugin Repository. It needs a Wherobots Cloud account, with a free trial. It is the first release from Wherobots Labs. Get Started with Wherobots Try Now
Wherobots Now Federates with the AWS Glue Data Catalog Posted on August 5, 2026October 4, 2026 by Pouyan Aminian AWS customers can now bring Wherobots’ best-in-class spatial intelligence capabilities to any dataset managed by AWS Glue Data Catalog. AWS Glue Data Catalog (GDC) customers use it to manage structured data, like user activity and revenue reports, and complex geospatial data, like parcels, buildings, and mobility datasets. AWS services like EMR, Redshift, and Athena excel at processing the former, while Wherobots is purpose-built for the latter. AWS customers have standardized tooling, governance, and processes around the GDC, which they do not want to replicate on another catalog. Having to manage two catalogs introduces friction and slows down innovation. How Wherobots Connects to the AWS Glue Data Catalog Today, Wherobots addresses this friction by making it easier for customers to use the best spatial data tool for the job while maintaining their governance layer in the AWS GDC. Using IAM policies, customers can now federate access between the AWS Glue Data Catalog and Wherobots. Wherobots will handle reading geospatial data from that catalog, processing it and writing it back to it, while Glue continues to act as the central catalog. For instance, consider a national e-commerce retailer coordinating just-in-time restocking across regional distribution centers. The workflow: Wherobots can ingest Sentinel-1 SAR imagery, which sees through cloud cover And ingest NOAA’s High-Resolution Rapid Refresh (HRRR) precipitation forecast raster Join both to the Overture Roads network connecting each center to its delivery routes Buffering the affected road segments and running zonal statistics against the flood and precipitation layers, the retail team can use Wherobots to calculate a route disruption score for each corridor and writes it to a federated Glue table. Supply chain teams at the retailer can then pull that score into their EMR restocking pipelines to trigger early shipments or reroute orders before a corridor goes down. Risk scores along road segments in Europe. Spatial Compute Built for Physical-World Data Trino, Postgres, and Spark are battle-tested engines, but they were not designed for what spatial data demands, such as coordinate reference systems (CRS), mixed raster and vector formats, and joins across billions of geometries. Built by the original creators of Apache Sedona and 100% code compatible across all spatial functions, WherobotsDB handles the spatial workloads that stall on general-purpose engines. With RasterFlow, this power extends to Earth Observation, enabling planetary-scale mosaicking and AI inference pipelines alongside high-performance vector analytics. Customers can also take advantage of Wherobots’ Spatial AI tools to get insights from spatial data in minutes using natural language, with little to no prior geospatial expertise. Open table formats and catalog services like AWS Glue Data Catalog remove the tradeoff: your governance, tooling, and data stay in Glue, while the engine built for spatial data does the spatial work. Regrid, uses Wherobots to provide comprehensive parcel data with boundaries, buildings, addresses, and geographic enrichments for all your location decisions. "We made a deliberate decision to standardize on AWS. The expertise and experience Wherobots bring to helping us navigate and architect the transition of 4,000+ geospatial pipelines to a highly scalable and efficient cloud native processing workflow has been invaluable to the success of our project. Wherobots is creating the future of cloud geospatial processing." Blake Girardot Senior Data Engineer at Regrid Get started with Wherobots and the AWS Glue Data Catalog Integrate Wherobots with your AWS Glue Data Catalog securely using AWS IAM in a few minutes, and run your first workloads on the Professional free trial: 14 days or $95 in credits, whichever comes first. We also offer hands-on spatial expertise to help you move your ideas and initiatives into production faster. We deliver this through the Wherobots Innovation Edition, which you can subscribe to by reaching out to us. Key takeawaysAWS Glue Data Catalog customers can now federate access to Wherobots using IAM policies. Wherobots reads geospatial data from the catalog, processes it, and writes it back, while Glue remains the central catalog. Customers keep existing GDC governance, tooling, and processes instead of running a second catalog.WherobotsDB is built by the original creators of Apache Sedona and is 100% code compatible across spatial functions. It handles CRS, mixed raster and vector formats, and joins across billions of geometries that stall on general-purpose engines such as Trino, Postgres, and Spark.A national e-commerce restocking example joins Sentinel-1 SAR imagery, NOAA HRRR precipitation forecast rasters, and the Overture Roads network, then buffers affected segments, runs zonal statistics, and writes a route disruption score to a federated Glue table that EMR restocking pipelines can pull.Regrid standardized on AWS and used Wherobots to help architect the transition of 4,000+ geospatial pipelines to cloud-native processing. Integration uses AWS IAM and can be stood up in a few minutes on the Professional free trial: 14 days or $95 in credits, whichever comes first.
Introducing the Wherobots Innovation Edition, designed to accelerate your physical world objectives Posted on July 8, 2026October 3, 2026 by Damian The Wherobots Innovation Edition helps you deliver outcomes on top of spatial data that propel your organization forward, and make this data AI-ready.Today we are announcing the Wherobots Innovation Edition. This is an annual partnership that pairs the full Wherobots Cloud platform with our forward deployed spatial engineering expertise, developed over years of delivering solutions, supporting production workloads, and leading the Apache Sedona project. We are announcing the Innovation Edition out of ongoing demand. The combination of our expertise in building geospatial data pipelines for production environments, and the Wherobots product has repeatedly delivered success for companies in private arrangements. It gives you a direct way to access our team of distinguished spatial data experts, on demand, to accelerate the realization of your goals. The Innovation Edition is built for one job, which is helping you ship outcomes while making spatial data ready for AI. As a result: Risk models are fresh, have more coverage, are more precise, and drive more of the right actions Services, supply chains, and logistical operations can improve continuously Threats to physical ecosystems and supply chains can be mitigated before they cause harm World models improve on the back of higher quality, more complete context CapEx heavy investments can produce higher returns Agriculture and forest management practices are more efficient We have proven to customers that AI is already capable of co-piloting these and other innovative, physical-world solutions when it has access to the right context. We have successfully delivered on multiple mission-critical engagements where customers relied on our team to remove roadblocks and deliver outcomes on schedule, and relied on our product to deliver outcomes with data. We are not opinionated on what cloud offering you’re using. Need solutions to thrive in AWS, GCP, Azure, Databricks, and Snowflake? No problem. Wherobots is designed to bring its capabilities into these environments. Here is what one of our customers, Sergey Sukov at Action Engine has said about working with Wherobots and leveraging the innovation tier. We made this offering public due to customers and partners like Action Engine. Wherobots has been more than infrastructure for us. Their team helped us design a system that treats spatial context as a first-class signal, and their engine runs it reliably in production against the highest-volume fleets running with Aspen Fleet. That combination of product and expertise is why our fleet data is ready for the AI stack our customers depend on. Sergey Sukov CEO, Action Engine Is the Innovation Edition right for you? The Innovation Edition fits your team if you: Have widespread interest and investment in the physical world, and need data driven approaches to understand risk and opportunities Want to unleash the potential of AI on spatial data, using the existing data estate you have hosted on Databricks, AWS, Google, Azure, or Snowflake Want to build competency and ownership of your spatial data workflows rather than outsource them Need solutions fast, regardless of data scale, complexity, and data type How engagement works Qualification for the Innovation Edition is a lightweight, no-risk process. It is designed to quickly assess whether we are the right partner for your objectives. 1. Reach out. We will schedule a 30-minute call to learn about your objectives and technical requirements. 2. Align on the path forward. We can guide your team, ship insights into your environment, or both. We will respond with the path forward, including timelines, milestones, a proof of value, and architecture. A follow-up call or POC can follow, as required. 3. Subscribe. Once there is mutual alignment, you subscribe to the Innovation Edition and the engagement begins. We will co-build the engagement plan with you and how it will grow over time. 4. Innovate Together. We help you define the objectives and key results that we are achieving together, the technical path to bring results into production, and co-build with you. We will review progress with your team regularly on a weekly or bi-weekly basis, and conduct quarterly business reviews to ensure you are successful. If you want to get started with an initial discovery call, reach out to us to discuss next steps and building together as a team. Key takeawaysThe Innovation Edition is an annual partnership that pairs the full Wherobots Cloud platform with forward-deployed spatial engineering expertise, developed over years of production workloads and leading the Apache Sedona project.The offering exists because private arrangements combining that expertise with the product have repeatedly delivered for customers. It gives a direct, on-demand way to access distinguished spatial data experts to accelerate outcomes and make spatial data AI-ready.Wherobots is not opinionated about which cloud you run: the post states solutions can thrive in AWS, GCP, Azure, Databricks, and Snowflake. Action Engine CEO Sergey Sukov is quoted on using the innovation tier to treat spatial context as a first-class signal in production against high-volume Aspen Fleet workloads.Engagement is a lightweight qualification: a 30-minute call, alignment on path (guide, ship insights, or both) with timelines, milestones, proof of value, and architecture, then subscribe. Progress is reviewed weekly or bi-weekly, with quarterly business reviews.
Wherobots brought modern infrastructure to spatial data in 2025 Posted on January 26, 2026October 3, 2026 by Damian In 2026 we’re bridging the gap between AI and data from the physical world. Entering 2025, we knew we needed to prove Wherobots is fundamentally the best place to create and run spatial data workloads at scale. Last year we directed the vast majority of our energy at strengthening the core fundamentals– ease of use, cost, performance, and reliability, knowing this focus would resonate with customers and value would be amplified through what we build later on. We knew we needed to bring spatial data into the modern data architecture, which from our vantage point, is the data lakehouse. If we did nothing, much of this data would otherwise remain siloed, “special”, and out of reach of modern analytics engines that could put this data to work. This is why we led contributions of GEO type support to Iceberg and Parquet. Off-platform through the open source Apache Sedona project, we saw an opportunity to develop a lightweight query engine that would appeal to developers because it would provide the support they need out of the box and accelerate iterations with spatial data. And so SedonaDB was born. This post is a high level summary of these and other accomplishments from our team in 2025. Now, we are actively building on top of this improved foundation, to enable AI and data practitioners across industries and use cases to operate with a heightened understanding of the physical world, for any area of interest. Customer success with scalable spatial data platforms This is what matters most. Everything else written here is just supporting evidence that shows how we made our customers more successful with spatial data this year. And what better proof than their own words? “Working with Wherobots let us focus on what matters – helping our clients make better land decisions. Their platform helps us scale efficiently while keeping our attention on real-world outcomes across energy, conservation, and development” Danan Margason Founder & CEO at Aarden.ai “The fact that Wherobots can mosaic imagery over millions of square kilometers, run AI models over those mosaics, and then organize the outputs in minutes, for 10 to hundreds of dollars is nothing short of incredible, making global scale analysis a routine pipeline instead of an enormous and extremely expensive endeavor, has the potential to change monitoring tasks far beyond field boundary analysis.” Caleb Robinson Principal Research Scientist, Microsoft AI for Good “With Wherobots, we were able to merge 15+ complex vector datasets in minutes and run high-resolution ML inference at a fraction of the cost of our legacy stack. The combination of speed, scalability, and ease of integration has boosted our engineering productivity and will accelerate how quickly we can deliver new geospatial data products to market.” Rashmit Singh CTO, SatSure “We’re helping to democratize large-scale spatial analytics for everyone. With Wherobots, we can move faster, scale bigger, and help more organizations make smarter decisions about the physical world.” Eric Pollard Founder and CEO, ParGo Elevating spatial data and AI capabilities for AWS customers It’s no surprise that many of these customers are AWS customers. In late 2024, we launched Wherobots Cloud as a product AWS customers could subscribe to directly through the AWS Marketplace. This activated a key value distribution channel between Wherobots and AWS customers. It also started our partnership in earnest with AWS. We continue to work closely with the AWS team to bring world-class spatial capabilities into the hands of their customers so they can better realize their objectives with physical world data. Leadership in the open data and lakehouse ecosystem We were the team that led the introduction of GEO type support to Iceberg and Parquet, which led to the incorporation of GEO types in the Databricks Delta Lake project. Because these projects form foundational components of the modern data architecture, with GEO type support, a significant portion of spatial data could now be interpreted and safely interoperated on by common compute engines. That also meant it no longer had to live in siloed architectures. It could thrive in the common data estate – the data lake, processed using engines like Spark, Snowflake, BigQuery, or Wherobots. If you squint, there is now a clear path for making spatial data look just like “data” in the eyes of developers and AI systems, particularly when capable engines like Wherobots can crunch it without a problem. Cloud and lakehouse integrations for spatial workflows Wherobots Cloud evolved into a full-fledged spatial intelligence platform designed not just for querying geospatial data very efficiently at scale, but for building production-grade workflows that integrate high value derivatives of physical world data into customers’ existing data architectures. Our native integration with Amazon S3 makes it possible for customers to run Wherobots on spatial data in their storage. In 2025 we announced our integration with Unity Catalog, enabling Databricks customers to activate the value of Wherobots on spatial data in their Databricks lakehouse. We will continue to add integrations such that our customers can just add Wherobots’ magic to the data infrastructure they already have. Apache Sedona appeals to a wider audience The Apache Sedona community developed and launched SedonaDB, the first open-source, single-node analytical database engine that treats spatial data as a first-class citizen. They also made it significantly easier to compare query performance across engines using SpatialBench, the first benchmarking framework for spatial queries. SedonaDB makes spatial data significantly more useful and analytically accessible for a wider range of use cases and personas. SpatialBench is there to streamline the decision making process for users looking to choose an engine based on spatial query price-performance and capability. Here’s a link to the announcement for both. WherobotsDB keeps getting better We’re continuously improving the WherobotsDB engine, raising the bar we’re self-setting for spatial query price-performance and capability. In 2025 we announced multiple new functions, tools, and compatibility with GeoPandas. We also announced a preview of a new runtime version, 2.x, which contains the latest optimizations for spatial range queries, spatial filtering, and spatial joins. Rust is at the core, and it leverages vectorized execution. Compared to the first major version, 2.x further accelerates spatial queries of up to 3.3x. While Wherobots is generally known for its spatial data capability, 2.x is significantly more performant for general purpose query operations, to a degree in which it’s also TPC-H competitive to alternative managed Spark engines in the market.We will be sharing benchmarking results when we announce general availability for version 2 soon. Making remote sensing and Earth sbservation data AI-ready We launched RasterFlow in private preview, the first serverless workflow purpose built to prepare and perform inference on large scale Earth observation (EO) datasets. RasterFlow is an Earth Intelligence solution, addressing the infrastructure challenges and high costs that prevented companies from utilizing raw EO data in the first place. We’ve packaged years of GeoAI expertise into a serverless, easy to use product. RasterFlow creates inference-ready mosaics after digesting large, unprepared imagery datasets, and runs model inference on these mosaics with custom or open PyTorch models to perform tasks such as change detection, classification, and segmentation. Results are delivered as geometries in Iceberg tables in a customer’s S3 bucket or as tables in Databricks Unity Catalog, to be processed by WherobotsDB or alternative engines. Enabling AI on physical world data Physical world data is noisy, it’s large, and it’s generally semi-structured or unstructured. This data also needs more context in order for it to be useful, which requires spatial joins to other datasets. It can be publicly available, or reside as private assets within an organization’s data estate. Teams and AI systems alike need this data to be processed and contextualized with other data in order for it to be useful. At scale Wherobots provides arguably the best tools for spatial processing and contextualization at the most fundamental level, for the modern data architecture. Now Wherobots is ready to be wired up to AI systems. Late in 2025 we announced the availability of the Wherobots MCP server to give your LLM access to Wherobots’ tools. Now, LLMs can use the MCP server to efficiently design queries by understanding the spatial and non spatial data in your data estate (via the S3 and Unity Catalog integration), and run those queries on a high performance engine to answer questions about the physical world. Soon we’ll integrate the MCP server with RasterFlow. That way, an AI agent can design and trigger a workflow in Wherobots that starts with fresh EO data, prepares it for inference, generates predictions using a collection of PyTorch machine learning models, and perform additional enrichment or transforms if needed to produce the result. Join the upcoming office hour on MCP server to learn more. What’s next for Wherobots in 2026 You can reach out to the product team at product@wherobots.com, or me directly at damian@wherobots.com to share the challenges you’re facing and see how we can solve them now or with capabilities we’ll add. Here are a few discrete roadmap items we are working on, and categories of investment planned in 2026. Make RasterFlow generally available Enabling AI on physical world data Bringing Wherobots closer to developers (VS Code, Kiro, etc) Providing support for Wherobots in the compute environment you need it to run in Making WherobotsDB even faster and cost effective Broaden the capabilities of SedonaDB Key takeaways2025 focus was core fundamentals — ease of use, cost, performance, reliability — and bringing spatial data into the lakehouse. Wherobots led GEO type support in Iceberg and Parquet, which then flowed into Databricks Delta Lake, so spatial columns can live in the common data estate instead of siloed GIS stores.Wherobots Cloud grew into a spatial intelligence platform with native S3 integration and a Unity Catalog integration so Databricks customers can run Wherobots on lakehouse spatial data. Wherobots Cloud had launched on AWS Marketplace in late 2024.The Apache Sedona community launched SedonaDB (single-node analytical database with spatial as a first-class citizen) and SpatialBench. WherobotsDB 2.x preview was announced with Rust and vectorized execution, accelerating spatial queries by up to 3.3x versus the first major version, and described as TPC-H competitive with alternative managed Spark engines.RasterFlow launched in private preview as a serverless workflow to mosaic unprepared EO imagery and run custom or open PyTorch models, writing geometries to Iceberg in S3 or Unity Catalog. Late 2025: Wherobots MCP server so LLMs can design and run queries against S3 and Unity Catalog data. 2026 plans listed: RasterFlow GA, more AI enablement, VS Code/Kiro, more compute environments, faster WherobotsDB, broader SedonaDB.
Introducing RasterFlow: a planetary scale inference engine for Earth Intelligence Posted on December 10, 2025October 4, 2026 by Philip Darringer We’re very excited to announce RasterFlow is now available to select customers in a private preview. If you are interested in learning more or would like to request access to the preview, contact us here! RasterFlow is a serverless image preparation and inference engine that makes it significantly easier to generate Earth Intelligence from planetary scale Earth Observation (EO) datasets. With it, customers and their AI agents will be significantly more capable of innovating with EO data and integrating earth insights into their data infrastructure. Upcoming Session: See an AI agent take plain-language question and orchestrating pipelines to end results using RasterFlow and the Wherobots MCP server. See it in action Join us live How RasterFlow Powers Earth Intelligence at Scale A few weeks ago, we announced our collaboration with the Taylor Geospatial Engine to help them evaluate their Fields of the World (FTW) machine learning model that segments agricultural field boundaries. Using an early release of RasterFlow, we were able to quickly and cost-effectively run this model at scale. Here’s a breakdown of how this works in practice. RasterFlow ingests and assembles the source imagery – in this case Sentinel-2 – into an inference-ready mosaic, generating representative features using the FTW model for planting and harvest seasons, and removing cloud cover as needed (1, 2). The FTW model is run against this mosaic using RasterFlow’s distributed inference engine to predict fields and field boundaries (3). RasterFlow predictions are then vectorized into geometries and made available as an Iceberg table (4) that can be used in WherobotsDB or other downstream applications and data systems for field-level crop insights. RasterFlow’s applicability is much wider than Sentinel 2 and FTW. It supports Zarr and COG imagery datasets and PyTorch computer vision models for inference. RasterFlow at Scale The images above represent sample outputs for a small area in Kansas, but RasterFlow can be very attractive for larger scale runs. In our collaboration with the Taylor Geospatial Engine, we executed larger scale runs including the Continental United States (CONUS), Japan, Mexico, South Africa, Switzerland and Rwanda. RasterFlow’s efficient parallel processing enabled each of these large scale workflows to complete in minutes to a few hours. RasterFlow autoscales compute resources based on expected compute and inference load, which is a function of area and time range, dataset density, and model complexity. Challenges using EO Data Most data teams do not have the expertise or the budget to build and operate the unique infrastructure and software stack required to extract insights from EO datasets using computer vision models. These barriers have prevented innovative ideas from getting off the ground. According to Gartner, only 1% of AI models today leverage physical world data, vs a projected 80% by 2029. Similarly, AI agents are projected to generate 10 times more data from physical environments than from all digital AI applications combined.1 However, AI agents can’t economically make sense of this raw data because it has to be prepared by the same costly, complex, and unique infrastructure the data teams need, but neither have access to. Here’s an example that underscores these challenges: if you or an AI agent are trying to analyze wildfire state and predicted spread to measure risk to infrastructure, developers typically need to build dedicated pipelines that: Ingest and prepare imagery for inference, minimizing noise such as cloud cover and edge effects Deploy a machine learning model on prepared imagery, trained to segment and classify fires Tune model inference for scale and efficiency, while minimizing edge and tiling effects from individual tasks Measure change over time using models that take into account wind direction, speed, vegetation, buildings and other infrastructure in the probable path of the fire Join model predictions with other important context including building footprints, land parcels and infrastructure such as powerlines and pipelines to calculate overall risk Forecast the spread of the fire In total, these steps require significant investments in both infrastructure development, operations, and talent that most businesses are unable to justify, much even accomplish. On-Demand Imagery Preparation and Inference for Earth Observation Workflows The inspiration for RasterFlow was to make it easy for any company to use large scale sensor datasets and computer vision models to unblock innovation and AI applications for the physical world. RasterFlow does this by combining decades of expertise with a fully managed, inference and mosaicking workflow and API designed for Earth Intelligence at any scale. Here are a few key capabilities: On-demand serverless operations for imagery ingestion, preparation (also known as mosaicking), and inference. Built-in support for popular open datasets and open models so you can get started quickly. Inference results that can be converted to vector geometries and integrated into a lakehouse architecture; in a customer’s cloud storage bucket as Parquet files in Apache Iceberg tables. Ability to easily postprocess these results with WherobotsDB or other lakehouse engines with support for spatial operations, such as Databricks, Snowflake, or Google BigQuery. Simple enough for any engineer, scientist, or analyst to use: just pick a model, an area of interest to deploy that model, and a time range. Advanced users can take advantage of lower-level APIs to customize their planetary-scale inference runs. RasterFlow Operators: Core Functions for Preparing Imagery, Model Inference, and Vectorization RasterFlow provides fully managed operations required for processing Earth Observation datasets, including: Imagery ingestion and preparation to remove cloud cover, edge effects, and build a high quality inference-ready mosaic Distributed inference for large scale computer vision, geospatial foundational and other PyTorch model runs Vectorization of model outputs into geometries or as analytics ready rasters For object detection workloads, RasterFlow pairs with models like Segment Anything 3, see how SAM 3 performs on aerial and satellite imagery for a full walkthrough. Ingesting and Preparing Satellite Imagery for Model Inference Satellites and drones capture imagery on a particular flight path. And it may take multiple drone flights, or days, weeks, or even months for the flight paths of a satellite constellation to capture clean imagery for a particular area of interest. Clouds and weather events may still block what you may be interested in. In these circumstances it’s important to understand the rate of coverage and define your time horizon accordingly, to build a mosaic. A mosaic is a composite image that is the result of composing high-quality pixels (e.g., cloud free) over a time range, and stitching them together for a particular area. Base satellite layers in your favorite map applications (Google Maps, Mapbox) are cloud-free mosaics composed from images over a wide time range. Many computer vision models are trained to find relatively durable things on Earth, like buildings, roads, and land cover. But when clouds, coverage, imagery edge effects, or other types of “noise” exist in the input imagery, the quality of inference suffers. The purpose of the mosaic is to correct for this noise and make imagery, inference-ready, so model inference produces the results you want. RasterFlow takes care of this heavy lifting for you, creating an inference-ready mosaic that maximizes the usefulness of today’s Earth Observation models. Distributed Geospatial Inference at Planetary Scale We’ve moved past the use of eyes to analyze imagery, and are now capable of letting machines do this work for us. With RasterFlow, today’s machine learning models can perform tasks such as object detection, segmentation, and classification, on a very large area of interest, with orders of magnitude more efficiency and scale than an analyst’s eyes can offer. The RasterFlow inference engine is designed for small to very large scale runs. It efficiently parallelizes across the input mosaic across a distributed and serverless inference architecture while minimizing tiling effects typically produced when inference pipelines operate on individual tiles. Running Hosted or Custom Geospatial AI Models with RasterFlow For convenience, RasterFlow currently hosts popular open source PyTorch geospatial computer vision models that are ready to use. These models currently include: Fields of the World (FTW) Field Boundary Delineation Meta and World Resource Institute Tree Canopy Height Prediction ChesapeakeRSC Road Segmentation Tile2Net Pathway Segmentation You can also import your own custom PyTorch model to your Wherobots Organization for private deployment. RasterFlow + TorchGeo: Simplifying PyTorch-Based Geospatial AI Wherobots actively supports the TorchGeo project which helps machine learning experts to more easily work with geospatial data within the PyTorch ecosystem. We will continue to build out RasterFlow integrations with TorchGeo, including onboarding additional TorchGeo models and further simplifying the model lifecycle for PyTorch models. While we are starting with support for PyTorch focusing on TorchGeo models, we are open to adding support for other model frameworks. Calling Geospatial Model Developers: Contribute to RasterFlow We are continually adding new, open source geospatial computer vision models to the Wherobots Model Hub. And if you’re a model developer, we’re interested in speaking with you to onboard your model and distribute the value of your work to a wider audience using Wherobots RasterFlow. Vectorizing Model Outputs: From Raster Predictions to Geospatial Geometries Many computer vision models output rasters, where each pixel in the raster represents a predicted real-world value such as height of the tree canopy, or the confidence that the pixel represents a certain feature such as an agricultural field boundary or a sidewalk. RasterFlow provides built-in support for raster vectorization, turning pixel values into rich, concise geometries. These geometries represent features of interest that can be post-processed, conflated, and integrated into your workflows because they are yours, stored in open source file (Parquet) and table (Iceberg) formats in your S3 bucket. Using RasterFlow with Geospatial Foundation Models and Embeddings Recent developments in Geospatial Foundation Models have generated tremendous interest in the research community, potentially accelerating Earth Observation applications the same way that Large Language Models (LLMs) and embeddings have transformed AI’s ability to generate language. RasterFlow can generate embeddings from the latest open Geospatial Foundation Models, including OlmoEarth from the Allen Institute for AI (Ai2) and Clay. With RasterFlow’s ability to cost-effectively generate embeddings at scale, researchers and practitioners can easily generate embeddings for their area of interest and evaluate their suitability and power. Customers and Partners Using RasterFlow for Scalable Earth Intelligence One highlight while developing RasterFlow has been our collaboration with customers and partners like SatSure, Taylor Geospatial Engine, and Spyrosoft. We’ve used feedback from these teams to ensure we are solving for customer needs. Before the Thanksgiving holiday we shared our recent learnings from working together with Taylor Geospatial Engine, who have been incredibly helpful in providing input on the types of ways their ecosystem of developers and ML engineers would want to interact with RasterFlow. SatSure is an existing Wherobots customer and an early adopter of RasterFlow, and we are excited to see what they build next with it. "RasterFlow meaningfully accelerates the work SatSure and Wherobots already do together. By automating mosaicking, preprocessing, and distributed inference into a single, on-demand workflow, it removes much of the engineering overhead required to operationalize our models at national and multi-season scale. This helps us move new geospatial AI models into production faster, iterate more quickly with customers, and deliver fresher, high-resolution insights across agriculture, banking and financial services, and infrastructure use cases." Rashmit Singh CTO and co-founder, SatSure One of the largest deployments to date processed 348 TB of satellite imagery for a global release, see how RasterFlow delivered Fields of the World at planetary scale RasterFlow Availability and Multi-Cloud Architecture Wherobots infrastructure runs natively on AWS and customers pay for use through the AWS marketplace. RasterFlow and WherobotsDB support hybrid architectures, where data is read from, and results are written to other environments such as GCP, Azure, Oracle, or on-premises. This is particularly useful when processing open datasets or using open models and the environment in which data is processed may not be a concern. On-demand pricing for RasterFlow will be announced at a later date, but can be discussed with customers participating in the private preview. Next Steps: Try RasterFlow and Explore the Wherobots Spatial Data Platform We invite anyone who wants to test out RasterFlow to request to join the private preview here. Get started building with the most capable and efficient spatial data platform using the Wherobots Professional Edition. Sign up for the newsletter to keep pace with what’s happening at Wherobots. Source – 27 August 2025, Gartner Innovation Insight: World Models Are Set to Empower AI Agents With Imagination ↩︎ Join the Private Preview Sign Up Key takeawaysRasterFlow is a serverless imagery-preparation and inference engine, now in private preview, that turns planetary-scale Earth Observation datasets into Earth Intelligence without customers having to build their own mosaicking and model-serving stack.Working with the Taylor Geospatial Engine, RasterFlow ran Fields of the World inference across CONUS, Japan, Mexico, South Africa, Switzerland, and Rwanda. Those large-scale workflows completed in minutes to a few hours as compute autoscaled with area, time range, dataset density, and model complexity.It ingests Zarr and Cloud-Optimized GeoTIFF imagery, hosts PyTorch models including FTW field boundaries, Meta/WRI tree canopy height, ChesapeakeRSC road segmentation, and Tile2Net pathway segmentation, and lets organizations import their own private PyTorch models.Predictions are vectorized into geometries and written as Parquet files in Apache Iceberg tables in the customer S3 bucket, then post-processed in WherobotsDB or other lakehouse engines such as Databricks, Snowflake, or BigQuery.Gartner projects that only 1% of AI models leverage physical-world data today versus 80% by 2029, while AI agents are projected to generate 10 times more data from physical environments than from all digital AI applications combined. One of the largest RasterFlow deployments processed 348 TB of satellite imagery for a global release.
Your Perfect Week for Geospatial at AWS re:Invent 2025 Posted on November 6, 2025October 4, 2026 by Tiffany Huynh December means it’s time for AWS re:Invent in Las Vegas! Thanks to our friends at Felt, we’ve mapped out what your perfect week for geospatial could look like — so you don’t miss a thing! 📍Find Wherobots at booth #1239 in the Expo Hall Start your conference by visiting us in the Expo Hall — we’ll be there all week! Come talk geospatial with us. Ask us about Wherobots or Apache Sedona, spatial data processing, GeoAI, satellite imagery analysis, raster and vector joins, or any exciting spatial projects you’re working on. Have a challenge? We’d love to hear about it! 🎉 Kick off your week with the Geo Party We’re back with the Geo Party, hosted by Wherobots and Felt. This is a must-attend event for anyone in geospatial at re:Invent. You won’t want to miss it! Be sure to RSVP here as space is limited! ⭐ Top Geospatial Picks for re:Invent Here’s a curated list of sessions to help you make the most of the conference week. Check out the interactive map for the full agenda of recommendations, organized by location, and bookmark it for easy access! Monday, December 1 TimeSessionTitleSpeaker(s)11AM – WynnAgentic AI for geospatial modeling Agentic AI’s new generation of Industry and Line of Business solutions (PEX202)Michael Schmidt, Ryan Thomas5PM – Caesars ForumConstruction intelligence in action Redefining Operations: Caterpillar’s Geospatial Intelligence Solution (IND322)Steve Blackwell, Alan Doty, Ash Shah, Sriram Somasuntharam Tuesday, December 2 TimeSessionTitleSpeaker(s)5PM – Mandalay BayGlobal connectivity from spaceArchitecting resilient global networks with Project Kuiper (ARC320)Nick Matthews Wednesday, December 3 TimeSessionTitleSpeakers10AM – Mandalay BayData modeling deep diveAdvanced data modeling for Amazon ElastiCache (DAT438)Kevin McGehee, Yaron Sananes5PM – MGM GrandAI-powered emergency response Weather emergency advisor with Amazon Bedrock and Location Service (WPS311)Charlotte Fondren, Bradley Wyman Thursday, December 4 TimeSessionTitleSpeaker11AM – MGM GrandHands-on geofencing that actually worksGeo-fencing & real-time geospatial alerts with ElastiCache Valkey (DAT408)Gururaj S Bayari, Damon LaCaille, Kevin McGehee, Veerendra Nayak, Allen Samuels3PM – Mandalay BaySustainable computing for climate workSustainable computing for climate solutions (AIM417)Nitin Pathak, Guyu Ye If you’d like to find a time to connect at re:Invent, fill out the form below and we’ll be in touch to set something up! Key takeawaysThis is an event roundup for AWS re:Invent 2025 in Las Vegas, mapped with Felt, not a product announcement.Find Wherobots at Expo Hall booth #1239 all week to talk Apache Sedona, spatial processing, GeoAI, satellite imagery, and raster/vector joins.Wherobots and Felt are hosting the Geo Party; RSVP is required because space is limited.Curated geospatial session picks run Monday, December 1 through Thursday, December 4 across the Wynn, Caesars Forum, Mandalay Bay, and MGM Grand, covering agentic AI, Caterpillar geospatial operations, Project Kuiper, ElastiCache data modeling, Bedrock emergency response, Valkey geofencing, and sustainable computing for climate.Use the post interactive map for the full location-organized agenda, or fill out the form in the post to book time with the Wherobots team on site.
Introducing wkls: A Python Library for Instantly Accessing Global Administrative Boundaries Posted on November 5, 2025October 3, 2026 by Pouyan Aminian Why Defining Administrative Boundaries in Spatial Data is Hard If you’ve ever worked with spatial data, you probably needed to define a geographic boundary within which to conduct your analysis. Most of the time, these are administrative boundaries such as cities, states, provinces, countries etc. For instance, if you want to scope your analysis to New York City, you’d need to look for the admin boundary online, find an authoritative source, figure out how to either download and hardcode the data into your code or build a pipeline that reads directly from their APIs (if they support it). In case of New York’s CSV file, doing this will give you a 14.5K characters long text mostly consisting of unreadable coordinates, a common challenge for developers before tools like the wkls Python library. Before wkls – 150 Lines of Unreadable Coordinates nyc = 'MULTIPOLYGON (((-74.046135 40.691125, -74.046176 40.691092,\ -74.047041 40.691041, -74.047149 40.690985, -74.047207 40.690893, \ -74.047196 40.690794, -74.047146 40.690714, -74.047026 40.690589, \ -74.047041 40.690482, -74.047183 40.690411, -74.046248 40.689319, \ [... 149 lines later!] -74.0400963 40.6989342, -74.0401502 40.6989014)))' After wkls – 1 Line Referencing the Hierarchical Admin Boundary nyc = 'MULTIPOLYGON (((-74.046135 40.691125, -74.046176 40.691092,\ -74.047041 40.691041, -74.047149 40.690985, -74.047207 40.690893, \ -74.047196 40.690794, -74.047146 40.690714, -74.047026 40.690589, \ -74.047041 40.690482, -74.047183 40.690411, -74.046248 40.689319, \ [... 149 lines later!] -74.0400963 40.6989342, -74.0401502 40.6989014)))' These boundaries are well-known and well-defined, but most geospatial tools do not include them natively. This is because getting geopolitically precise administrative boundaries is challenging and often results in very large datasets (e.g., 10K-1M points per boundary). As stated above, oftentimes data practitioners are forced to find Shapefiles for these boundaries on the internet, write code to download them from the source and include them in their projects. Alternatively, developers sometimes hardcode these strings into their projects or use inaccurate bounding boxes instead of the actual administrative boundary. This is, at best, boilerplate code that needs to be written and maintained over and over again and a possible source of inconsistencies between projects. We heard this feedback from our customers repeatedly and it lined up perfectly with our mission to make geospatial easy to work with. That is why we are very excited to introduce the Well Known Locations (wkls) library. The wkls library (pronounced “Whickles”) includes ~625K global administrative boundaries — from countries to cities — which can be referred to by name using clean, chainable Python syntax. The library reads directly from Overture Maps Foundation GeoParquet data hosted on the AWS Open Data Registry. The supported formats are WKT, WKB, HexWKB, GeoJSON, and SVG. The library is included in Wherobots core libraries and there is zero installation or configuration required to take advantage of it. How to Use the wkls Python Library Start by importing the library into your code and reference the cities using Python’s objection notation: import wkls wkls.us.wkt() # country: United States wkls.us.ny.wkt() # state: New York wkls.us.nyc.cityofnewyork.wkt() # city: New York wkls["us"]["ny"]["cityofnewyork"].wkt() # dictionary-style access wkls supports up to 3 chained attributes: Country (required) – must be a 2-letter ISO 3166-1 alpha-2 code (e.g. us, de, fr) Region (optional) – must be a valid region ISO code suffix (e.g. ca for US-CA, ny for US-NY) Place (optional) – a name match against subtypes: county, locality, or neighborhood For instance, the chained expression wkls.us.ca.sanfrancisco returns a data frame object containing all the matches to the administrative boundary for San Francisco. In most cases, the call resolves to a single admin boundary object (i.e., row). If there are name collisions (e.g., two representations of city of San Francisco, one with just the land border and the other including shorelines as well), multiple rows may be returned. Once you have the administrative boundary object, it can be used like any other geometry within Wherobots. For instance, you can calculate intersections, reference the boundary in any raster function, etc. For more information please read the wkls documentation. Want to contribute to the library? You can open issues, submit pull requests, improve documentation and more by following the instructions on this open source repository. Try Wherobots: Faster, Smarter Geospatial Queries Making administrative boundaries more accessible is not the only way we are making geospatial developers’ lives easier. Our platform runs spatial queries 5-20X faster and up to 60% more cost efficient to use compared to other industry leading solutions. We also have rich functionality that allows you to run vector and raster functions in the same query. Our Spatial AI capabilities are also industry leading. Finally, our Community tier is free to try! Get Started Now Key takeawayswkls (pronounced Whickles) is a Python library that lets you reference about 625K global administrative boundaries—from countries to cities—by name with chainable syntax instead of pasting giant WKT strings.A New York City boundary sourced as CSV is about 14.5K characters of coordinates; typical admin polygons run 10K-1M points, which is why teams hardcode WKT, hit public APIs, or fall back to bounding boxes.The library reads Overture Maps Foundation GeoParquet from the AWS Open Data Registry and returns WKT, WKB, HexWKB, GeoJSON, or SVG. It ships in Wherobots core libraries with zero extra install or configuration.Access is up to three chained attributes: country (required ISO 3166-1 alpha-2, e.g. us), optional region ISO suffix (e.g. ny), and optional place matched against county, locality, or neighborhood subtypes—for example wkls.us.nyc.cityofnewyork.wkt().Name collisions can return multiple rows (for example two San Francisco representations, land-only vs. including shorelines). The project is open source and accepts issues and pull requests.
Introducing SedonaDB and SpatialBench for Apache Sedona Posted on September 24, 2025October 3, 2026 by Damian Our role at Wherobots and as leaders in the Apache Sedona community is to help more developers, organizations, and AI systems positively transform the physical world using spatial data. In order to make the scale of transformation we envision possible, we’ve had to address significant bottlenecks in how data is stored and queried. We’re excited to celebrate the availability of SedonaDB and SpatialBench for Apache Sedona. Together they represent the next phase in our plan to accelerate innovation with spatial data and bridge the intelligence gap between AI and the physical world. Intro to SedonaDB: A modern query engine that gets spatial right SedonaDB is the first open-source, single-node analytical database engine that treats spatial data as a first-class citizen. Most analytical query engines already support general-purpose operations: filtering, joins, aggregations, and APIs for SQL or Python. But when it comes to operating on spatial data those same engines fall short: support for geometry and geography types, coordinate reference systems (CRS), spatial joins, and raster or vector operations is missing. The workaround is to bolt on an extension like PostGIS (PostgreSQL), DuckDB Spatial (DuckDB), or SedonaSpark (Spark). While powerful, extensions inherit the limits, costs, and complexities of their host systems, require extra setup and tuning, and can force builders to develop around performance and usability gaps instead of developing their ideas. SedonaDB is different. It’s for builders solving problems with physical world data. Written in Rust, it’s lightweight, blazing fast, and spatial-native. Out of the box, it provides: Full support for spatial types, joins, CRS, and functions on top of industry standard query operations. Query optimizations, indexing, and data pruning features under the hood that make spatial operations just work with high performance. Pythonic and SQL interfaces familiar to developers, plus APIs for R and Rust. Flexibility to run in single-machine environments on local files or data lakes. SedonaDB uses Apache Arrow and Apache DataFusion, and provides everything you need from a modern vectorized query engine. But it delivers the unique ability to also run high performance spatial workloads easily, without requiring extensions. Read the announcement on the Apache Sedona blog to dive in and roll up your sleeves. What led to SedonaDB? In 2020, Apache Sedona was incubated to address a significant support gap in distributed geospatial data processing. Since then, Sedona has enabled companies like Uber, Amazon Last Mile Delivery, JB Hunt, and thousands of others with geographically distributed operations or interests to build and run more efficient and effective physical operations at scale. It is widely used today to bring geospatial processing support to Apache Spark, Apache Flink, and also Snowflake. But distributed systems aren’t for everyone or the right fit for every use case, and we could do more to drive innovation in lower-scale scenarios. Accelerating innovation Many ideas are bootstrapped in no-to-low cost environments where iteration cycles are fast and low risk. There’s a lot that a developer can do today using a laptop or a single virtual machine, with modern software and LLMs—without adding a dependency that adds unwanted cost and complexity to the innovation cycle. Once their ideas are viable, they may not even require a distributed compute environment in production like Spark, or one that is “fully managed” by a vendor. So the next step was pretty clear. We had to make it easier for builders to use spatial data in no-to-low cost environments so they can iterate and positively transform the physical world, faster. We also decided to address these challenges through open-source software to maximize accessibility. Making development easier If you look around the ecosystem, you’ll notice a pattern: to get the analytical support you need for geospatial data, you deploy an analytics engine without the spatial analytics support you want, and then you bolt on what you need via an extension. Extensions are great and they serve a purpose very well. After all, SedonaSpark is an extension! But that doesn’t mean the combination of engine + extension is ideal. It requires additional setup and management, can require tuning to achieve a reasonable performance, and the underlying engine may end up becoming a bottleneck. Additionally, the development experience around the engine may be overly complex or lack support for the language you prefer, and the engine itself might introduce compute, cost, and other overhead. Working from the root causes of these challenges, along with the desire to drive more innovation, our next step became obvious. We needed to create a query engine that aids spatial data solutions development out of the box with popular pythonic and SQL interfaces and is optimized for single-machine environments. But was there enough value created by a spatial-first query engine compared to general purpose query engines with spatial extensions? Optimizing for spatial data = a better future Spatial data is no longer a minor class of data. It’s everywhere, the rate at which it’s being generated is growing every day, and its use cases span numerous industries. It streams from devices, vehicles, satellites, and drones, and derivatives from this data inform automation and decision-making across business, government, and research. The solutions being developed with it are transforming how organizations operate in the physical world. Innovation is happening today with this data despite the friction above, but the pace of this innovation can be accelerated by a query engine with internals intentionally designed to help developers realize the full potential of this data. This engine is SedonaDB, and it’s backed by an open-source community (Apache Sedona) that is committed to solving physical-world challenges through data and technology. Intro to SpatialBench: The first standard for spatial query performance “Without standards, there can be no improvement” – Taiichi Ohno. This statement from the founder of the Toyota Production System is an analogy for why we built SpatialBench. There was no standard way of measuring spatial query performance, so progress couldn’t be easily quantified or query engines objectively compared on this dimension. We built SpatialBench to establish first standards. The initial release supports 12 representative queries, ranging from simple to complex workloads, and includes a data generator for scale factors 1, 10, 100, and 1000. We hope this framework and its future versions will guide innovation that leads to a greater understanding of the physical world. We also used SpatialBench to benchmark SedonaDB, DuckDB (with its spatial extension), and GeoPandas at scale factors 1 and 10. Those results are published here. Next Steps Get Started with SedonaDB: Try it out and contribute to the roadmap. Use SpatialBench: Measure spatial query performance using a consistent standard. Watch the webinar (hosted by CNG): We’ve walked through SedonaDB and SpatialBench, and introduced Wherobots’ Startup Accelerator Program. Key takeawaysSedonaDB is the first open-source, single-node analytical database engine that treats spatial data as a first-class citizen rather than a bolt-on extension like PostGIS, DuckDB Spatial, or SedonaSpark.Written in Rust on Apache Arrow and Apache DataFusion, it ships spatial types, joins, CRS, functions, indexing, and data pruning with Python and SQL interfaces plus APIs for R and Rust, and it runs on a laptop or a single VM against local files or a data lake.Apache Sedona was incubated in 2020 for distributed geospatial processing and is used by companies such as Uber, Amazon Last Mile Delivery, and JB Hunt on Spark, Flink, and Snowflake. SedonaDB covers the no-to-low-cost, single-machine side of that ecosystem.SpatialBench is introduced as the first standard for spatial query performance: 12 representative queries from simple to complex, plus a data generator at scale factors 1, 10, 100, and 1000.Wherobots used SpatialBench to compare SedonaDB, DuckDB with its spatial extension, and GeoPandas at scale factors 1 and 10; those results are published on the linked Apache Sedona announcement, not in this post.
Wherobots Spatial Intelligence Engine Integrates with Databricks Unity Catalog for Spatial Data Posted on September 11, 2025October 4, 2026 by Damian TL;DR: Wherobots now integrates with Databricks Unity Catalog, enabling users to process spatial data up to 20x faster with 60% cost savings. This integration supports raster/vector data, 300+ spatial functions, and enterprise security—all while maintaining full compatibility with Apache Sedona and Spark. Databricks Geospatial Performance with Wherobots 5-20x faster spatial query performance 60% cost reduction on spatial workloads 300+ spatial SQL/Python/Scala functions 100% Apache Sedona compatibility (zero code changes) Databricks users can now enhance their geospatial analytics capabilities with Wherobots, a spatial intelligence engine purpose-built for processing data from the physical world. Wherobots brings advanced raster processing, computer vision ML inference, and industry-leading performance to your existing Databricks environment. With Wherobots Data Federation for Unity Catalog, you can: Expand spatial coverage to grow revenues, improve margins, and make better decisions with complete spatial intelligence Create innovative spatial data products that leverage aerial imagery, IoT sensors, and mobility data at planetary scale Build with raster, vector, and tabular data using familiar SQL, Python, or Scala interfaces Run computer vision ML models on geo-imagery and sensor datasets from local to continental scale Migrate existing workflows with zero code changes—WherobotsDB is fully compatible with Apache Sedona and Spark Customers like Dotlas, Leaf Agriculture, and Overture are achieving step-function improvements in performance, cost efficiency, and innovation by integrating Wherobots with their Databricks platforms. Wherobots Data Federation connects directly to Unity Catalog, allowing you to read from and write to Iceberg or Delta tables with Databricks service principal and OAuth or PAT token authentication—no data migration required. Get started | Schedule a demo What is Spatial Intelligence? Spatial intelligence is the understanding of features of interest, and their relationships across space and time in multi-dimensional environments. With it, bridges are formed between digital and physical worlds. Decision making can improve, and you can create better products or services with higher returns. Simple spatial intelligence: Customer visits grouped for a location or area. But simplified formats like aggregations inherently lack precision, and you need to make tradeoffs between cost and resolution, all of which limit their usefulness. These aggregations are typically grouped by cells in a grid (like H3) and most commonly used to create visualizations. While interesting to look at, visualizations can be used to support a decision, intuition, or analysis, but they are generally less actionable because precision or other context is missing. Complete spatial intelligence: Forecast, or a composite risk, opportunity, or value score associated with potentially millions of specific assets or features across a continent derived from any valuable combination of IoT, location, building, weather, road network, terrain, crops, parcel, aerial imagery, BI, or mobility datasets. By processing perspectives about features of interest from multiple valued aspects, the complete picture forms, which becomes highly actionable, and extremely useful intelligence. But even the most popular data platforms still don’t make it easy to create. Wherobots enables complete spatial intelligence at scale with Databricks Unity Catalog integration. Why Do Traditional Data Platforms Struggle with Geospatial Data? There are many cloud data engines and warehouses that support the simple form described above, including BigQuery, Snowflake, and now Databricks with its Spatial SQL support in preview. However they still lack feature and data type support, reasonable query price-performance at scale, and the solution expertise you may need to create a complete form of spatial intelligence. Here’s why. Shaped by demand, most data platforms were first designed to handle the structured data exhaust from the web and devices connected to it, not data collected from or about the physical world. Physical world data is inherently complex, unstructured, and doesn’t fit neatly into key-based joins. It takes a specialized compute engine to make it easy to create solutions from spatial data. Easy means it’s capable of fusing and transforming various spatial and non spatial data types with high accuracy, scale, performance, and low cost, while ensuring development is productive with the teams you have. The spatial extensions and APIs for today’s big data engines and warehouses provide limited support for simple workloads. But because of design bottlenecks, missing features, and limited technical support, spatial solutions on these platforms can be expensive and difficult to build, while ideas remain far-fetched. What Makes a Modern Spatial Intelligence Solution Effective? Ideally the solution for creating complete spatial intelligence just fits into your existing software development workflows, already supports your future needs and the data you want to utilize, is accessible to the teams you have, and just performs at the right scale — on-demand, at a cost that encourages innovation. It’s lakehouse ready, so you don’t need to move your data or utilize proprietary formats or data types to use it. You also have dedicated expertise in reach to unblock innovation. With this capability at your fingertips, ideas can flow and innovation takes place. Your business can reach higher levels of efficiency, reducing costs, carbon footprint, and risk. You can speed up deliveries or pickups, increase the effectiveness of CAPEX, improve consistency, grow revenue, and build in ways that were thought to be impossible. This capability is Wherobots, and it’s directly available to Databricks users via data federation with Databricks Unity Catalog. How Wherobots Solves Geospatial Data Challenges Using Wherobots you can easily build a complete picture of what’s happened, over space and time, and integrate this intelligence into your Databricks data platform to drive growth – faster and at a lower cost than ever. Our mission is to make spatial data easy to utilize, and it’s all we are focused on. The results of our focus speak for themselves. Databricks Unity Catalog Geospatial Integration: Key Wherobots Features Wherobots makes it easy and economical to produce local to planetary scale data solutions that rely on any combination of aerial and overhead imagery, IoT and mobility data, ground truth datasets, and your own business context. And using Wherobots Data Federation with Unity Catalog, you can easily integrate the data products you build with Wherobots, into your Databricks data platform while retaining custody and governance of data. What Types of Spatial Data Does Wherobots Support? (Raster & Vector) There are two main classes of spatial data supported by Wherobots: raster and vector data. You also get the support and scale you’d expect for tabular data operations from Wherobots’ Spark compatible engine. Raster data is typically a collection of sensor or imagery data, where each pixel in the image represents information about what is being captured, like temperature, elevation, infrared spectrum, etc. File formats include GeoTIFFs, Zarr, and NetCDF and more. Raster datasets are commonly GBs to TBs in scale. Vector data is a collection of multi-dimensional geometries or geographies that represent the trajectory, shape, elevation, and location of things. They can be trips, points, and outlines of features like buildings or parcel and crop boundaries. File formats include GeoParquet (soon to be Parquet), Shapefiles, and GeoJSON. Performance Benchmarks: Up to 20x faster queries, 60% cost savings Customers like Leaf Agriculture, Dotlas, Overture and others have compared the price-performance of using Wherobots for their spatial data workloads vs other managed Spark or other leading data platforms. Subscribed to the Professional Edition, they are self-reporting up to 20x better performance (5x-20x is typical) with on-demand savings reaching as high as 60%, with even higher savings from the Enterprise Edition. Data teams are equally less limited by scaling bottlenecks. This becomes apparent after workloads finish faster on smaller WherobotsDB runtimes, and after customers realize they have significant headroom to scale well past their existing needs. “Previously, our data volumes and processing requirements were increasing faster than we could keep up with, burdening our team with costly rebuilds. Now with Wherobots, not only can we easily scale to millions of acres, we also can rest assured that our costs won’t spiral out of control.” – G. Bailey Stockdale, CEO Leaf Agriculture These results are a function of specialization and a company-wide focus; WherobotsDB was built first for processing spatial data. This intentional design obviates the typical performance bottlenecks and complexities now alive in leading data platforms and warehouses, which were first designed for purposes unrelated to processing spatial data. While the quotes from our customers matter the most, we also know how important open performance benchmarks are. But currently there are no spatial query benchmarks (or at least reputable ones), which makes query performance hard to compare across platforms without trials. It’s also hard to claim progress was made on performance when standards have not been established. We’re working on this too, and soon we will release a new open source spatial query benchmarking framework for Apache Sedona, and we will release spatial query performance results for query engines, data warehouses, and data platforms. We already have preliminary results that compare WherobotsDB to Apache Sedona on various managed Spark engines along with query performance from engines with Spatial SQL APIs. Feel free to reach out and we can share these results when you contact us. Security & Compliance: Enterprise Grade Data Protection Wherobots is serverless and built for data security first. There’s no infrastructure to manage, although customers can also choose to run Wherobots in their AWS VPC for maximum control. Apache Sedona Expertise: Built by the Original Creators Wherobots was founded by the original creators of Apache Sedona, and Sedona is the most widely used geospatial extension for Apache Spark and in Databricks. With decades of research and experience with spatial data, open source, and cloud-scale systems, our product and team are ready to support Databricks customers’ solutions on the lakehouse. We’re also a team leading geospatial modernization efforts in open source. Wherobots has supported GEO types for years with our Havasu table format. But rather than keeping this support in-house, we decided these types would better serve the physical world in the open so we proactively drove support for them in Apache Iceberg and Parquet. When to Use Databricks Native Spatial SQL vs. Wherobots Choose Databricks Native Spatial SQL when you need: Exploratory geospatial analysis Basic spatial joins (ST_Intersects, ST_Contains, ST_Distance) Standard point-in-polygon queries Simple location-based aggregations No raster or satellite imagery processing Choose Wherobots for Databricks when you need: Production spatial intelligence workloads at scale Advanced raster and vector data processing Computer vision ML inference on satellite/aerial imagery Performance optimization (5-20x faster queries) Cost reduction (up to 60% savings on spatial compute) Planetary-scale datasets (TB to PB range) Complex spatial ETL pipelines Apache Sedona compatibility for existing workflows Get Started with Wherobots x Databricks Geospatial Analytics The lakehouse gives you the ability to choose the product best suited for the job. Don’t settle for the simple form of spatial intelligence or what the default provider offers, when complete is in reach with better economics, scale, capability, performance, and support. By integrating Wherobots into your Databricks workflows, organizations can reduce costs, improve operations, and realize new innovations powered by data from the physical world. Ready to enhance your Databricks geospatial capabilities? Get Started Here – Schedule a meeting with experts Read Documentation – Integration guide and API reference Processing spatial data in Databricks? Check out this spatial query benchmark on Databricks with SpatialBench. Key takeawaysWherobots now federates with Databricks Unity Catalog so you can read and write Iceberg or Delta tables in place—no data migration—using a Databricks service principal with OAuth or a PAT token.Customers such as Dotlas, Leaf Agriculture, and Overture report typical spatial-query speedups of 5-20x and on-demand cost savings as high as 60% on Professional Edition (higher on Enterprise), with 300+ spatial SQL/Python/Scala functions and 100% Apache Sedona compatibility (zero code changes).Wherobots handles raster (GeoTIFF, Zarr, NetCDF; commonly GBs to TBs) and vector (GeoParquet, Shapefile, GeoJSON) plus tabular Spark-compatible workloads, including computer-vision inference on geo-imagery from local to continental scale.Use Databricks native Spatial SQL for exploratory joins, point-in-polygon, and simple aggregations with no raster needs. Use Wherobots for production-scale spatial intelligence, raster+vector pipelines, GeoAI, TB-PB datasets, and existing Sedona workloads.Wherobots was founded by the original creators of Apache Sedona, the most widely used geospatial extension for Spark and Databricks, and is serverless—or can run its compute plane in your AWS VPC.
Advancing the Integration of Map Data via Overture’s Global Entity Reference System and Wherobots Posted on June 25, 2025October 3, 2026 by Ben Pruden Editor’s note: The Wherobots Spatial Data Catalog is now the Havasu Catalog. The general availability of the Overture Maps Foundation’s Global Entity Reference System (GERS) makes it a lot easier to build intelligence about features of our physical world. What is GERS? You can think of a GERS as a system for applying a unique key to physical features in the world. Overture assigns GERS IDs to millions of features in their data products, such as office buildings, highways, countries, rivers, schools, and more. Using GERS IDs, you can more easily join datasets to build a more complete view of physical-world features in space and measure relationships over time. Key benefits of GERS IDs include: Persistent identification: The same physical location maintains the same GERS ID over time Cross-dataset compatibility: Enable joining and enrichment across multiple datasets or providers Standardized reference: Provide a common language for location data across the geospatial ecosystem You can read more about GERS IDs in Overture’s documentation. Getting started with Overture in Wherobots At Wherobots, we’re proud to be a member of the Overture Maps Foundation, supporting the project since its formation, and as an official member since 2024. We also host and manage all of Overture’s recent datasets in the Wherobots Spatial Catalog, and they are offered at no additional cost to all customers. These datasets are production ready and at your fingertips. Select * from wherobots_open_data.overture.buildings_building New Wherobots customers can get started with Overture’s datasets in the free-to-use Community Edition, and graduate to the Professional edition when they want to join Overture data with their data in cloud storage. It’s Easy to use GERS in Wherobots A key design goal of GERS is to simplify how organizations can join and enrich their own datasets using canonical Overture datasets. Wherobots makes this vision real by providing built-in support for the GERS schema, enabling users to easily query, filter, and join to valuable geospatial datasets using GERS IDs, and through the use of spatial join predicates available in Apache Sedona and Wherobots. Whether data teams are working with parcel boundaries, retail site locations, road networks, or foot traffic telemetry, Wherobots makes it easy to enrich data using GERS with SQL or Python. Overture GERS Schema Extension Paths (not exhaustive) Stay tuned for a new tutorial from Wherobots that shows you how to use GERS IDs with Overture Places. Overtures datasets are produced using Wherobots Wherobots was founded by the original creators of Apache Sedona, the open-source engine for distributed geospatial processing. Overture now runs many of their Apache Spark and Sedona based data pipelines on Wherobots because they run up to 20x faster, at a fraction of the cost, and the Overture team benefits from the spatial expertise we offer them as a customer. The Overture team uses Wherobots’ Apache Airflow support to trigger job runs that power production of their planetary scale datasets. This feature and Wherobots compatibility with Spark and Sedona, also made it very easy for Overture to redirect where their Airflow-orchestrated jobs ran. “Overture produces a building dataset covering all buildings in the world, with 2.6B geometries and growing, that’s updated frequently. There’s a lot of data, and compute that goes into producing it and keeping it up to date. We accelerated the pipelines that produce the buildings dataset by up to 20x after we moved them to Wherobots, which required a simple redirection of our code. We retained compatibility with Apache Sedona, and the move put us into a development experience that’s made us more productive.” – Jennings Anderson, Geoscientist at Overture and Data Engineer at Meta. Redefining Standards in Open Source We’re committed to improving open standards in the geospatial data ecosystem. Wherobots is a leading contributor of spatial type support in Apache Iceberg and Parquet, the most popular open table and file formats for the cloud data lakehouse, to make it easier for companies to utilize geospatial data.We’ve partnered with Overture to improve the foundation for a scalable, versioned, and queryable world of features backed by GERS. By modernizing how geospatial data is accessed in the cloud via spatial data type support in Iceberg and Parquet, and improving accessibility and utility of open map data via GERS, we believe new use cases for spatial data will emerge to improve business operations, research, and our way of life. Join us in Building the Spatial Data Stack of the Future The launch of GERS is a big step for advancing spatial intelligence, and a leap toward a more open, interoperable data ecosystem. Wherobots is proud to be part of this journey, and we’re excited to continue supporting the Overture mission through operational pipelines, open standards, and accessible tools.If you’re building with geospatial data, we invite you to explore how Wherobots can help you take full advantage of GERS by easily joining your first party data with the GERS ID system. Next steps If you’d like a one-on-one demo from our team, you can request it here. Learn about our data mirror with Overture in their documentation. We will be publishing an example notebook that shows you how to use GERS IDs with Overture Places soon. Start Building with Wherobots Access Now Key takeawaysOverture Maps Foundation Global Entity Reference System (GERS) is generally available: stable unique keys on millions of physical features (buildings, highways, countries, rivers, schools) so datasets can be joined and tracked over time.Wherobots has been an Overture member since 2024 and hosts recent Overture datasets in the Spatial Catalog at no additional cost. Community Edition can query them; Professional Edition joins them to data in your cloud storage.Wherobots supports the GERS schema so you can query, filter, and spatially join first-party parcels, retail sites, roads, or telemetry to Overture with SQL or Python.Overture produces a global buildings dataset of 2.6 billion geometries (and growing). After moving Spark/Sedona pipelines to Wherobots—with a simple code redirect and Airflow job-run support—those pipelines accelerated by up to 20x while staying Apache Sedona compatible.Wherobots also led GEO type support in Apache Iceberg and Parquet so GERS-backed features can live in an open, versioned lakehouse rather than a proprietary silo.