Scaling Spatial Analysis: How KNN Solves the Spatial Density Problem for Large-Scale Proximity Analysis Posted on February 5, 2026October 3, 2026 by Pranav Toggi How we processed 44 million geometries across 5 US states by solving the spatial density problem that breaks traditional spatial proximity analysis When scaling spatial proximity analysis from city to state to national level, the hidden challenge isn’t computational power—it’s spatial density. The techniques that work perfectly for urban neighborhoods fail dramatically when applied across heterogeneous landscapes. The standard professional approach—using ST_DWithin with a fixed search radius—breaks down when spatial density varies. A 500-meter radius might capture 20 candidate features in Manhattan but zero in rural Wyoming. No single distance works for both. This article demonstrates how k-nearest neighbors (ST_KNN) solves this problem. Unlike fixed-radius predicates that yield density-dependent result sets, KNN applies a top-k constraint—guaranteeing bounded cardinality regardless of local feature distribution. No distance threshold tuning required. To validate this approach, we ran a buildings-to-roads proximity analysis across five US states on Wherobots Cloud: 44.4 million buildings against 535,000 road segments in 2.3 hours for $157—less than half a cent per geometry. The technique applies equally to any spatial proximity problem: customers to stores, facilities to services, properties to amenities. The Professional’s Dilemma: Static vs Adaptive Search The ST_DWithin Approach For any GIS professional, the standard approach to spatial proximity analysis is `ST_DWithin` with a fixed radius: sedona.sql(''' SELECT a.*, b.*, ST_Distance(a.geometry, b.geometry) as distance FROM query_geometries a JOIN target_geometries b ON ST_DWithin(a.geometry, b.geometry, 500, true) -- Fixed 500m radius ORDER BY distance ''') This works beautifully for spatially homogeneous regions. But scale it to a state or nation, and you hit the spatial density problem—where non-uniform feature distribution causes query behavior to become unpredictable. The Spatial Density Problem Consider the same ST_DWithin(geometry, 500m) query across different spatial contexts: Figure 1: The same ST_DWithin query produces many candidates in dense urban areas but 0 candidates in sparse rural areas The same query produces wildly different result set cardinalities based on local feature density. The Radius Paradox Attempting to solve this with radius adjustment creates new problems: Radius Urban Result Rural Result Problem 500m 20 candidates 0 candidates Rural queries fail 2km 200 candidates 3 candidates Urban over-processing 10km 2000+ candidates 10 candidates Urban becomes intractable No single radius value works for heterogeneous spatial data: increase it to capture sparse regions, and dense regions become intractable with combinatorial explosion in candidate counts. This isn’t a theoretical problem—it’s the practical limitation that prevents reliable large-scale spatial analysis using traditional methods. KNN Spatial Analysis: Solving the Density Problem at Scale K-nearest neighbors elegantly solves the spatial density problem by enforcing a cardinality constraint rather than a distance constraint—always returning exactly k candidates regardless of local feature density. A useful mental model: ST_DWithin applies a distance predicate (fixed radius, variable cardinality), while ST_KNN applies a rank-based predicate (fixed cardinality, variable distance). This produces the effect of a density-adaptive search area—though KNN doesn’t compute any radius internally. The Effect in Practice sedona.sql(''' SELECT a.*, b.*, ST_Distance(a.geometry, b.geometry) as distance FROM query_geometries a JOIN target_geometries b ON ST_KNN(a.geometry, b.geometry, 10) -- Always 10 candidates ORDER BY distance ''') The same query now produces consistent results across all spatial densities: Figure 2: KNN always returns exactly k candidates—nearby ones in dense areas, more distant ones in sparse areas. KNN doesn’t compute or adjust a radius. It simply finds the k nearest neighbors, wherever they are. In dense areas, those neighbors happen to be close; in sparse areas, they’re farther away. The “adaptive” behavior emerges naturally from asking “what’s nearest?” rather than “what’s within X meters?” The Adaptive Advantage Metric ST_DWithin ST_KNN Result cardinality Variable (0 to 1000+) Bounded (k) Time complexity Data-dependent Predictable O(n × k) Result quality Density-dependent Distribution-invariant Parameter tuning Manual, error-prone Not required The Key Insight: ST_DWithin asks “What’s within X meters?” (fixed radius, variable candidates). ST_KNN asks “What are the K nearest?” (fixed candidates, variable distance). This fundamental difference is why KNN handles heterogeneous density gracefully—no radius tuning required. Note: This “adaptive” behavior is conceptual—a way to understand why KNN outperforms fixed-radius approaches for heterogeneous data. It’s distinct from the optional distance bound parameter (see Practical Guidance), which imposes an actual maximum distance limit on candidates. The Two-Stage Pattern for Accurate Spatial Proximity KNN computes distances using geometry centroids (or bounding box representatives) for computational efficiency. For complex geometries like linestrings or large polygons, the centroid-to-centroid distance may diverge significantly from the true minimum Hausdorff distance. We address this with a two-stage refinement approach: Stage 1 (Candidate Generation): KNN selects k candidates using approximate centroid-based distance—O(nlogk)O(n\log k) Stage 2 (Exact Refinement): Precise geometric calculation (ST_ClosestPoint, ST_DistanceSpheroid) on k candidates only—O(k)O(k) per query point This decomposition gives us both fast candidate pruning and geometrically precise results. KNN Implementation Pattern Here’s the complete two-stage pattern using Wherobots SQL: Stage 1 – Candidate Generation: knn_df = sedona.sql(''' SELECT query.id AS query_id, query.geometry AS query_geometry, target.id AS target_id, target.geometry AS target_geometry, -- Exact closest point calculation ST_ClosestPoint(target.geometry, query.geometry) AS closest_point, -- Precise spheroidal distance ST_DistanceSpheroid( ST_ClosestPoint(target.geometry, query.geometry), query.geometry ) AS distance_meters FROM query_geometries AS query JOIN target_geometries AS target ON ST_AKNN(query.geometry, target.geometry, 10, false) ''') knn_df.writeTo('wherobots.pranav.knn_candidates') Stage 2 – Exact refinement: sedona.sql(''' ranked AS ( SELECT *, ROW_NUMBER() OVER ( PARTITION BY query_id ORDER BY distance_meters ASC ) AS rank FROM wherobots.pranav.knn_candidates ) SELECT * FROM ranked WHERE rank = 1; ''') Key functions: ST_AKNN(..., 10, false): Approximate KNN with k=10 candidates, Euclidean distance Why ST_AKNN over ST_KNN?: When using KNN for candidate generation (followed by exact refinement via ST_ClosestPoint + ST_DistanceSpheroid), approximate KNN is preferred. The relaxed precision bounds of ST_AKNN enable faster index traversal, and any approximation error is eliminated in the refinement stage. ST_ClosestPoint: Precise closest point on target geometry ST_DistanceSpheroid: Accurate geodetic distance in meters ROW_NUMBER(): Rank candidates and select the true closest The pattern works for all large non-point geometries. Results: Consistent Performance Across Spatial Densities We validated this approach using a buildings-to-roads proximity analysis across five US states. Each state contains a mix of dense urban, suburban, and sparse rural areas—exactly the heterogeneous density that breaks traditional methods. State Buildings (Query Geometries) Roads (Target Features) Time (sec) Cost Throughput New York 6,447,782 75,329 1,374 $25.31 4,693/sec Texas 13,289,136 166,943 2,140 $41.60 6,208/sec Colorado 2,764,970 34,432 1,074 $21.35 2,574/sec California 13,648,296 135,247 2,086 $36.95 6,541/sec Florida 8,201,965 122,048 1,490 $30.35 5,504/sec Total 44,425,199 535,538 535,538 $157.08 5,301/sec Key Observations Consistent Throughput: Average of 5,300 geometries per second across all states, despite each containing vastly different urban/rural mixtures. This consistency is the adaptive search radius in action. Cost Predictability: $0.0035 per geometry regardless of local spatial density. Budget with confidence. Linear Scaling: Processing time scales linearly with geometry count—no density-dependent surprises. Practical Guidance ST_KNN vs ST_AKNN for Proximity Analysis For the two-stage pattern described in this article, use ST_AKNN (approximate) rather than ST_KNN (exact): Function Speed Precision Best For ST_KNN Fast Exact When KNN result is the final answer ST_AKNN Faster Approximate Search space reduction before exact calculations Since we’re computing exact distances with ST_DistanceSpheroid on the candidates anyway, the approximate nature of ST_AKNN has no impact on final accuracy—only on speed. Using Distance Bounds with KNN If you know the maximum acceptable distance for your use case, add a distance bound parameter to further optimize performance: -- Only consider candidates within 5000 meters ST_AKNN(query.geometry, target.geometry, 10, true, 5000) Unlike the conceptual “adaptive search area” discussed earlier, this is an actual distance predicate pushed down to the spatial partitioning stage—not applied as post-hoc filtering. Candidates beyond this threshold are pruned during index traversal, reducing I/O and computation.This is useful when: Business logic requires a limit: e.g., “nearest hospital within 10km” Performance optimization: You know neighbors beyond X meters aren’t relevant Emergency response: Facilities must be within a critical response distance KNN Performance Optimization Tips Materialize intermediate results: Either cache or write to disk as Iceberg table, the KNN output before applying filters Process categories separately: Run KNN for each target type (e.g., motorways, then trunk roads) independently Use appropriate runtime: Medium runtime handled 13M geometries in ~35 minutes Applications of KNN-Based Spatial Proximity Analysis The adaptive search radius pattern applies to any spatial proximity problem with heterogeneous density: Infrastructure & Planning Properties → nearest utilities, transit, services Customers → nearest stores, facilities, competitors Environmental Analysis Development sites → nearest protected areas, water bodies Facilities → nearest residential areas, schools Business Intelligence Locations → nearest amenities, employment centers Assets → nearest maintenance facilities, resources General Pattern: For each query geometry A, find target geometry B that minimizes some expensive function f(A, B). Use KNN to reduce candidates, then apply exact calculation to the reduced set. Conclusion Large-scale spatial analysis requires solving the spatial density problem that causes traditional distance predicates to fail. KNN provides the solution: a cardinality-bounded query operator that delivers consistent result sets regardless of local feature distribution, with predictable computational complexity. Key Takeaways Density-Invariant Selection: KNN’s bounded cardinality constraint naturally accommodates varying feature density—no manual threshold tuning required. Predictable Performance: Consistent throughput and cost across heterogeneous spatial data, enabling reliable budgeting and planning. Two-Stage Pattern: Combine KNN’s fast candidate selection with precise geometric calculations for both speed and accuracy. At $0.0035 per geometry, spatial proximity analysis at any scale becomes economically trivial. The technique that processed 44 million buildings in 2.3 hours works equally well for customer analytics, infrastructure planning, or environmental assessment. KNN’s density-invariant behavior isn’t just a performance optimization—it’s what makes heterogeneous spatial analysis reliable. Resources Wherobots Documentation: ST_AKNN, ST_KNN, spatial function reference Wherobots Cloud: Distributed spatial analytics platform Overture Maps Foundation: Open map data used in our validation Create your Wherobots account START BUILDING
How Aarden.ai Scaled Spatial Intelligence 300× Faster for Land Investments with Wherobots Posted on November 17, 2025October 4, 2026 by Ben Pruden When Aarden.ai emerged from stealth recently with $4M in funding to “empower landowners in data center and renewable energy deals,” the company joined a new wave of data and AI startups reimagining how physical-world data drives modern business. Their mission: help institutional land investors rapidly evaluate the value and potential uses of land across the country. To do that, Aarden needed to process vast geospatial datasets such as parcels, forests, soil, endangered species, energy infrastructure, and turn them into actionable business intelligence. The problem? Their early Python stack couldn’t keep up. Aarden.ai needed faster, scalable geospatial processing and Wherobots made it possible. Results First: From Seven Days to Thirty Minutes Before adopting Wherobots, a single statewide geospatial computation — like calculating the distance from every parcel in a state to the nearest water source — took seven days to complete. After moving their geospatial data pipelines to Wherobots, that same job ran in just 30 minutes on a medium compute instance. Scripts that once took a week to run could now be developed in a day and executed in under an hour. That acceleration unlocked a new rhythm for Aarden’s engineering team: iterate daily, explore new models, and scale from a single state to a national view of land opportunity. As founding staff engineer Steven Yee put it, “We went from babysitting compute jobs for a week to getting results before lunch.” Why Startups Like Aarden Choose Wherobots Why did Aarden.ai choose Wherobots? For scalability, ease of use, and integrated raster-vector support. For startups building data applications rooted in the physical world — agriculture, climate, energy, mobility, land, or infrastructure — data scale and spatial complexity are unavoidable. The Aarden team, led by geospatial scientist Ben Hudson, knew this well. Hudson’s background processing satellite imagery for Greenland’s ice sheet and building Zillow’s Zestimate engine gave him firsthand experience in the pain of scaling geospatial pipelines. When Aarden began, they tried the standard open-source stack: GeoPandas, RasterIO, GDAL. It worked for prototypes, but not for production. The choice came down to two questions: How do we scale spatial computation without building a Spark team? How do we avoid spending months managing infrastructure instead of shipping data products? The answer was Wherobots. “Spark is incredibly powerful — but it’s also a huge learning curve, Wherobots shortened the painful part of Spark and gave us production-grade scalability without having to babysit clusters.” Ben Hudson Co-Founder and Head of Applied Science, aarden.ai Wherobots’ full support for both rasters and vectors meant Aarden could seamlessly combine terrain, vegetation, and parcel data into unified models. They could also store and query data using Iceberg tables, eliminating the need for maintaining large Postgres clusters. The alternative platforms — like Google Earth Engine or Microsoft’s Planetary Computer — weren’t built for their hybrid vector-raster workflows or the flexibility needed to prototype and deploy quickly. “Wherobots is the most proven way to do it,” Hudson said. “It’s the reliable, full-featured, tried-and-true option.” Making Land Data Useful for Decision-Makers Aarden’s customers are institutional land investors evaluating large portfolios of property — for carbon capture, solar, timber, or data center opportunities. Wherobots powers the data engine behind Aarden’s platform, turning sprawling public datasets into clear, numeric insights. End users don’t see geotiffs or coordinate grids. They see a simple interface: Which land deals have the highest alpha potential? Behind that simplicity is Wherobots’ compute layer — transforming complex geospatial and environmental data into machine-learning-ready features and business metrics. “Our customers don’t need to be geospatial experts,” said Hudson. “They just need to make smart business decisions. Wherobots helps us turn a mountain of geospatial data into simple, singular, useful numbers and cash flow analyses.” Looking Ahead As Aarden scales, their focus is on robustness, repeatability, and rapid iteration. With Wherobots, they can run production-grade geospatial pipelines without worrying about cluster management or data scaling. And as they expand nationally, their confidence is simple: “It just works.” Aarden’s story reflects a broader trend. The next generation of data startups — those whose insights are grounded in the physical world — are choosing Wherobots to get from prototype to production faster. Because when your data is as big as the planet, you need compute that scales with it. Looking to get started for your spatial data pipelines and intelligence application? Get started today in community (free) or try out pro for your team. Or reach out to sales for a demo. Key takeawaysAarden.ai uses Wherobots to help institutional land investors evaluate parcels for data centers, renewables, carbon capture, solar, and timber after raising $4M coming out of stealth.A statewide job that calculated distance from every parcel to the nearest water source dropped from seven days on Aarden early Python stack to 30 minutes on a Wherobots medium compute instance—about 300x faster.Scripts that once took a week to run can now be developed in a day and executed in under an hour, so the team iterates daily and scales from a single state to a national view of land opportunity.GeoPandas, RasterIO, and GDAL worked for prototypes but not production; Wherobots gave them Spark-scale raster and vector processing without standing up a Spark team, plus Iceberg tables instead of large Postgres clusters.Founding engineer Steven Yee: we went from babysitting compute jobs for a week to getting results before lunch.
Raster Spatial Joins at Scale: Google Earth Engine and BigQuery vs Apache Sedona and Wherobots Posted on July 31, 2025October 4, 2026 by Matt Forrest If you are deciding between Google Earth Engine + BigQuery and Apache Sedona + Wherobots for large-scale geospatial analysis, here is the short answer: Wherobots completed a statewide zonal statistics analysis across every building in Texas in 3 minutes 28 seconds. BigQuery timed out on the same query. For a Dallas metro area test using the same datasets, Apache Sedona finished in 37 seconds vs BigQuery’s 8 minutes 24 seconds, making it 13x faster. And before BigQuery could even run the spatial join, filtering the global dataset down to a single bounding box took 58 minutes and 50 seconds. This post breaks down why that gap exists, what it means architecturally, and when each platform actually makes sense for your workload. Key Takeaways Here is what we found: Apache Sedona via Wherobots is 13x faster than BigQuery for large-scale zonal statistics on a metro-area dataset BigQuery with Google Earth Engine could not complete a state-scale spatial join without timing out; Wherobots finished the same job in under 4 minutes Apache Sedona reads directly from cloud storage (S3, STAC) with no data migration required; BigQuery requires importing data into Google Cloud first Apache Sedona supports 300+ spatial functions; BigQuery offers 71, with ST_REGIONSTATS as the only raster-vector function For production pipelines at scale, Wherobots wins on speed, flexibility, and cost of data movement. Google Earth Engine remains useful for exploratory, one-off analysis within its own dataset catalog What is Zonal Statistics Analysis and Why Does Scale Matter? A zonal statistics analysis is one kind of result of a spatial join between a raster dataset and a vector dataset. This analysis can return things like: sum: Sum of the pixel values in a zone mean: Arithmetic mean of those values min or max: Limit values of the pixels in the zones For example, we could use vector boundaries of postal codes to find the average elevation from a raster: Imagine that you want to take a study area, maybe a specific bounding box or metro area, and understand the dominant land classification in that area. In this case, each grid square in the raster might have a value like ”water” or “trees” and you want to find the classification with the maximum count in each of your vector areas. At a small scale this kind of analysis is relatively easy for one-off work. There are a few engines that allow you to do this. Shapely/GeoPandas and Rasterio are three commonly used Python libraries used together to do this. PostGIS, which includes spatial SQL functions to work with raster and vector data, extending PostgreSQL with spatial functionality. Google Earth Engine, which includes the ability to perform zonal statistics and now add support to perform calculations with vector data inside BigQuery. Apache Sedona, which provides a functional layer to perform distributed spatial queries using Apache Spark and other engines. However, there are plenty of use cases where you do want to do this at large scale: Understand the land classification across thousands, hundreds of thousands, or even millions of individual buildings or properties. Analyze climate or weather data in small or large scale areas over time. Get up to date fire risk using global or country specific data. Calculate and update flood risk across thousands of properties by analyzing the height above nearest drainage points. These data sets have the potential to add massive value to analytical workflows for many industries and sectors. Unfortunately, the Python and PostGIS approaches are difficult to scale effectively. While they work up to a point, they run on a single compute instance and therefore will only scale as far as the computing power and memory of that machine. So how can you do this at scale and without disrupting your existing architecture? Let’s look at the two cloud options, Apache Sedona and Google BigQuery. How Google Earth Engine + BigQuery Handles Large-Scale Spatial Joins (And Where It Struggles) Apache Sedona and Google Earth Engine both provide solutions when you need to process larger amounts of data. Google Earth Engine, originally created in 2009, generally stores its raster data in multi-resolution pyramids in data that is not directly exposed to the user via a proprietary library. Google Earth Engine does store some vector data and it’s distributed in a similar format via the libraries themselves. In Google Earth Engine you can access planetary scale data and analyze it for small and large scale problems without the need to download the data on your computer. In recent years, Google has started to expand the capabilities for Earth Engine to integrate with the Google Cloud Suite of products, most notably BigQuery. BigQuery is an online analytical processing (or OLAP) data warehouse that allows you to store massive amounts of long-form data and analyze it efficiently using distributed computing. It uses a proprietary storage format that allows it to quickly distribute workloads across data that’s stored in the Google Cloud architecture. The downside of this proprietary format is that you must import your data into BigQuery before leveraging the distributed computing resources that make it powerful. BigQuery makes it possible to calculate zonal statistics using Earth Engine data, but unfortunately that geospatial data does not reside in the same proprietary, fast format. (this differs from the approach with Wherobots and Apache Sedona outlined later in the post). Recently, BigQuery added the function ST_REGIONSTATS that allows you to directly integrate BigQuery SQL with datasets from Earth Engine. All in all, this workflow is great for traditional tabular analytics if you already have or are planning to migrate data into BigQuery. However, if you want to leverage open formats, Apache Sedona and Wherobots allow you do do that at scale with open core tooling. How Apache Sedona + Wherobots Handles the Same Problem Apache Sedona is a distributed computing framework that brings spatial functionality into the Apache Spark ecosystem. It allows you to work with large-scale datasets, both vector and raster, and combine them into new datasets using zonal statistical functions. It uses an underlying distributed computing framework that allows you to distribute work in parallel, allowing you to scale up to planetary scale computational workloads. A key to Sedona’s scalability is that it allows you to leave the data where it sits. Many exabytes of raster data are already in an accessible storage format on the web and Sedona lets you leave that data there without spending the time or money needed to copy it yourself. This means that there are no proprietary formats, no data ingestion, and no extra steps to actually start working with this data. Platform Comparison: Apache Sedona + Wherobots vs Google Earth Engine + BigQuery Here’s how the two approaches compare across key technical and operational dimensions: Apache Sedona & WherobotsGoogle Earth Engine & Big QueryOpen formats✅ Works with standard formats like Cloud-Optimized GeoTIFF; no vendor lock-in⚠️ Data needs to be ingested into BigQuery, internal proprietary storage for vector and raster dataNo ETL needed✅ Reads directly from remote raster sources like S3 or STAC endpoints without importing❌ Only works with supported Google Earth Engine datasets and data stored in Google Cloud StorageScalable✅ Distributed compute + open-source stack enables scaling from local to cloud✅ Good for one-off, small to mid-scale analysesSpatial function support✅ Over 300 raster and vector functions supported in Apache Sedona⚠️ ST_REGIONSTATS is the only raster/vector function; 71 total spatial functions in BigQueryCloud-native✅ Fully supported cloud environment in AWS (GCP coming soon)✅ Integrated with Google Cloud; Easy to use if your stack is already in GCP To scale zonal statistics queries even more, Wherobots only materializes the values or pixels of raster data it needs. Not only does this save time on the computational workload, but it also eliminates a significant amount of moving data back and forth, which adds time and cost in other systems. Benchmark Results: Wherobots vs Google Earth Engine + BigQuery Using ST_RegionStats in BigQuery via Google Earth Engine To compare the two systems I used a common open dataset to analyze land classification near buildings in Texas. This analysis can help determine if a building is in a built-up urban environment or if it is an area with different natural resources. I used the WorldCover dataset from the European Space Agency, which contains land cover data across the entire globe. The data is in AWS S3 storage as an open dataset here, as well as Google Earth Engine. For the buildings, I used the Overture Maps dataset, which contains building footprints for the entire globe. Overture data is in both the Wherobots Data Hub and the BigQuery Open Data project. Initially, I tried to compare the query speed for every building in Texas. Wherobots finished in about 3½ minutes, but BigQuery continued to time out after several attempts. To scale back my test, I focused on buildings in a small area around Dallas, Texas. Each building was buffered by 30 meters to ensure that we are looking at the area around the buildings, not just the footprint. To make this query work, I first needed to enable Earth Engine in BigQuery in the Google Cloud console, including: Activating the Earth Engine API Adding in required permissions for my project Locating and subscribing to the required Raster dataset via the Analytics Hub Below is the code that I used to run this for the Overture Maps data in BigQuery. WITH texas AS ( SELECT st_buffer(geometry, 30) as geometry, id FROM <code>PROJECT.DATASET.overture_texas</code> where st_intersects(geometry, st_geogfromtext('POLYGON((-97.000482 33.023937, -96.463632 33.023937, -96.463632 32.613222, -97.000482 32.613222, -97.000482 33.023937))')) ) SELECT t.geometry, t.id, ST_REGIONSTATS( t.geometry, (SELECT assets.image.href FROM <code>PROJECT.DATASET.landcover</code> ), 'Map' ).mean as mean FROM texas AS t ORDER BY mean DESC; You can use the existing Overture Maps data in BigQuery, however I recommend making a new table with just the data you need for the analysis. Using a WHERE statement to filter the global Overture Maps table to the small bounding box took 58 minutes and 50 seconds. You can view the complete guide for doing this in GCP here. Using Apache Sedona and Wherobots Using Apache Sedona and Wherobots we can efficiently load in both the ESA WorldCover data from AWS S3 open data using built-in functions to read a Spatio-Temporal Asset Catalog (or STAC) endpoint that has an API to read large collections of raster data. Click here to launch this interactive notebook Launch Notebook stac_df = sedona.read.format("stac").load( "<https://services.terrascope.be/stac/collections/urn:eop:VITO:ESA_WorldCover_10m_2021_AWS_V2>" ) stac_df.printSchema() Alternatively, we can also load this directly from the S3 data source, as you can see here which is automatically retiled which sets up nicely to create our tiled out-of-database raster dataframe, which we can then turn into an Iceberg table to use in the future: esa = sedona.read.format("raster")\\ .option("retile", "true")\\ .load("s3://esa-worldcover/v200/2021/map/*.tif*") esa.createOrReplaceTempView("esa_outdb_rasters") YOUR_CATALOG_NAME = 'my_catalog' # Persist as a raster table in Iceberg sedona.sql(f""" CREATE OR REPLACE TABLE wherobots.{YOUR_CATALOG_NAME}.esa_world_cover AS SELECT * FROM esa_outdb_rasters """) # ✅ Out‑DB raster table created! From here I can set up my zonal stats query using the function RS_ZonalStatsAll which will actually return a STRUCT of all the stats instead of just the selected one as BigQuery does. The values included are: count: Count of the pixels. sum: Sum of the pixel values. mean: Arithmetic mean. median: Median. mode: Mode. stddev: Standard deviation. variance: Variance. min: Minimum value of the zone. max: Maximum value of the zone. zonal_stats_df = sedona.sql(f""" SELECT p.id, RS_ZonalStatsAll(r.rast, p.buffer, 1) AS stats FROM wherobots.{YOUR_CATALOG_NAME}.texas_buildings p JOIN wherobots.{YOUR_CATALOG_NAME}.esa_world_cover r ON RS_Intersects(r.rast, p.buffer) """) # ✅ Zonal stats computed! With this query, you could choose to write the results to an Iceberg table, or you could choose to persist them to files. In this case, I chose to persist these to files for our test. zonal_stats_df.write \\ .format("geoparquet") \\ .mode("overwrite") \\ .save(user_uri + "/results")") And that’s it. You don’t need to move any data, download anything, or spin up any additional services apart from the notebook to run this process. Results The performance gap between the two platforms was significant at both scales tested. Test 1: All of Texas The first test that I ran tried to intersect all of the Overture Maps buildings in the state of Texas against this global raster dataset. With Wherobots, I ran the iteration seven times, writing the complete dataset to disk each time. The results for the iterations were: 3 minutes 28 seconds (average for each iteration) ± 25.9 s per iteration (mean ± std. dev. of 7 runs, 1 loop each) BigQuery was unable to finish the test without timing out. Test 2: One city For this test, I limited the query to a bounding box in the area of Dallas shown below. POLYGON((-97.000482 33.023937, -96.463632 33.023937, -96.463632 32.613222, -97.000482 32.613222, -97.000482 33.023937)) For this smaller area, the BigQuery computation ran in 8 minutes and 24 seconds. Below are the results in Wherobots using the smaller area: 37.4 seconds (average over 7 iterations) ± 1.78 s per iteration (mean ± std. dev. of 7 runs, 1 loop each) Performance Summary AreaBigQueryWherobotsState of TexasN/A (timed out)3m 28sDallas, Texas metro area8m 24s37s When to Use Apache Sedona + Wherobots vs Google Earth Engine The benchmark results make the performance case clearly. But there are architectural reasons to consider Sedona and Wherobots beyond raw speed. Performance and scale Sedona scales horizontally using distributed compute across Apache Spark, with strategic partitioning and adjustable runtime size to handle metro, national, or global workloads. Apache Sedona supports 300+ raster and vector functions, compared to BigQuery’s single raster-vector function, ST_REGIONSTATS. That function coverage matters when building full analytical pipelines, not just one-off queries. Wherobots includes job management for scheduled, repeatable runs, which matters when your raster datasets change over time, such as fire risk, flood risk, or land cover updates. Architecture and flexibility Sedona reads directly from remote cloud storage like AWS S3 or STAC endpoints without copying data into a proprietary system. BigQuery requires data to live inside Google Cloud before its distributed compute kicks in. You are not limited to Earth Engine’s dataset catalog. Any open dataset accessible via cloud storage works with Sedona without ingestion or reformatting. The stack is open core. You can run Apache Sedona locally to prototype, then move to Wherobots for production without rewriting your analysis. For teams evaluating Google Earth Engine + BigQuery against Apache Sedona + Wherobots for production geospatial workloads, the benchmark is clear. At metro scale, Sedona is 13x faster. At state scale, BigQuery could not finish the job. If your data lives outside Google Cloud, or you need more than one raster-vector function, or you need repeatable scheduled pipelines, Wherobots is the stronger architectural choice. Google Earth Engine remains a capable tool for exploratory analysis within its own catalog. Try the interactive notebook Get Started
Dekart Supports Wherobots as a Spatial SQL Engine Posted on July 16, 2025October 3, 2026 by Ben Pruden We’re excited to announce that Dekart now supports WherobotsDB as a spatial SQL engine, enabling you to execute high performance spatial queries while rapidly visualizing query results on a map all within the Dekart platform. If you’d like to skip the blog post and jump right in, you can follow our Getting Started Guide here. What is Dekart? Dekart is a geospatial analytics application, available as a managed platform or self-hosted, that turns SQL queries into shareable, interactive maps. It glues a slim Golang backend onto the Kepler.gl library, letting analysts connect directly to their cloud data warehouse and visualize millions of rows without exporting files or writing code. Dekart gives data teams a fast path from SELECT… to a production-ready map, without the overhead of proprietary GIS stacks like ESRI or heavyweight desktop software. If your spatial analysis already lives in SQL, Dekart is a practical, lightweight way to visualize and share it, all from your browser. How the Wherobots-Dekart integration shines This integration allows customers to use Wherobots as the query engine to create visualizations in Dekart. Here’s why the combination matters. Price-performance: By using WherobotsDB as a spatial SQL engine to query data in S3, you can join, transform, and filter planetary-scale spatial datasets up to 20x faster and more efficiently than other lakehouse engines in the cloud. This efficiency boost improves performance and reduces compute costs. Easy visualization: Working from Dekart’s SQL editor, you can rapidly execute Spatial SQL queries against WherobotsDB, and visually interact with the results in Dekart’s interactive map experience. Modern lakehouse architecture: With the combination of Dekart and Wherobots, you can create powerful map-based visualizations from spatial data that sits securely in your S3 bucket. The Wherobots modern lakehouse architecture eliminates the need for ETL to a large data warehouse just to process or visualize location data. This architecture is especially powerful for teams who work with large, dynamic spatial datasets such as: Mobility traces from GPS or cell towers Field boundaries from drone or satellite imagery Delivery routes and transportation networks Infrastructure assets, outage zones, and signal maps These datasets can often change daily or hourly. With the integration between Wherobots and Dekart, you don’t need to rebuild maps from scratch. Just re-run your Spatial SQL query and get an updated visualization from the versioned data in your cloud lakehouse. This is what turns Kepler.gl from a static exploration tool into a repeatable, enterprise-grade intelligence system for the physical world. And because both Wherobots and Dekart are built on open standards and open source foundations (Sedona, Iceberg, Kepler), you’re not locked into a proprietary GIS ecosystem. Real-World Use Case: Telco Network Optimization Spatial data is the underpinning of the telecommunications industry. From cell tower coverage to network congestion to user demand mapping, telcos generate a continuous stream of location-rich telemetry. This data is used for maintaining quality of service, rolling out infrastructure, and responding to outages. But processing this much data at scale and making sense of it often presents a challenge. Many telcos already use Apache Sedona in data lakehouses like Databricks to process call records, analyze tower placement, and perform network optimization. These workloads can involve billions of records and highly complex spatial joins. Look no further than how Comcast is using Apache Sedona here. With Wherobots and Dekart, telco data teams can now take this one step further: Run complex spatial queries directly on cloud-native Iceberg tables, powered by Wherobots as your spatial SQL engine Visualize cell coverage, signal drop zones, or congestion patterns instantly in Dekart Layer in cluster analysis, such as Getis-Ord Gi* (G-star) statistics, to detect statistically significant hot spots and cold spots Iterate quickly, comparing performance before and after tower optimizations Share insights easily with planners, engineers, and execs—no GIS software required Here’s an example:A network operations team wants to analyze dropped call patterns across the San Francisco Bay Area. They write a Spatial SQL query in Wherobots that joins call records to a hex grid, then apply Getis-Ord Gi* to identify statistically significant clusters of poor performance. Within seconds, Dekart renders the output as an interactive map—highlighting not just where drop-offs occur, but where they represent systemic problems. No exporting data. No manual mapping. Just fast, repeatable insights from data about the physical world. Process data with Wherobots SQL Spatial Engine, and customize how its visualized Historically, you would use the same platform to both process and visualize spatial data. This often resulted in mediocre performance on both of those tasks. At Wherobots, we believe that the optimal workflow is having a dedicated spatial data processing platform (like Wherobots) that you can pair with the visualization tool that works best for your use case. That’s why we are particularly excited by this integration with Dekart. It gives our users yet another in our growing list of visualization tools they can layer on top of Wherobots to support their day-to-day spatial workflows. We believe that the modern spatial stack should use the best-in-class tool for each step in a spatial workflow. In this case, Wherobots for spatial data processing and querying, Dekart and other solutions for visualizing on a live map. That’s why we support: Kepler.gl and Deck.gl maps and visualizations embedded within our notebooks Apache Superset for building dashboards from spatial data QGIS for traditional desktop GIS users who want the processing power of Wherobots (coming soon, currently in Alpha) We also allow users to generate VTiles (a type of map tile from vector data) from any spatial dataset using SQL. Those tiles can be exported as PMTiles–a performant, portable, cloud-native tile format for serving map tiles. These PMTiles can be used with MapLibre, Leaflet, or any WebGL tile renderer. Ready to Level Up Your Spatial Intelligence? Start Today. Whether you’re optimizing communication networks, mapping field boundaries, analyzing human movement, or publishing public map layers, Wherobots and Dekart provide a fast path for converting spatial data into insights. Create your Wherobots account here. Getting Started To get started with Wherobots and Dekart, you can follow our Getting Started Guide here. Or for a quick view of what’s involved, take a look at this video where Dekart founder Volodymyr Bilonenko walks through connecting Dekart to Wherobots and running a query. Let’s build the next generation of intelligence for the physical world. Start Building with Wherobots Get Started Key takeawaysDekart, a Kepler.gl-based geospatial analytics app (managed or self-hosted), now runs Spatial SQL against WherobotsDB and renders results as shareable interactive maps in the browser.Using Wherobots as the engine, teams can join, transform, and filter planetary-scale spatial data in S3 up to 20x faster and more efficiently than other cloud lakehouse engines, then visualize without exporting files or standing up a GIS stack.The lakehouse pattern keeps data in your S3/Iceberg tables. Re-run the SQL when mobility traces, field boundaries, routes, or outage layers change daily or hourly and the map updates—no ETL into a warehouse just to draw it.A telco example joins call records to a hex grid, applies Getis-Ord Gi* hotspot analysis, and maps statistically significant dropped-call clusters in the San Francisco Bay Area from the Dekart SQL editor.Wherobots also embeds Kepler.gl/Deck.gl in notebooks, supports Apache Superset dashboards, has QGIS integration in alpha, and can emit VTiles/PMTiles for MapLibre, Leaflet, or other WebGL renderers.
Exploring the Foursquare Open Places Data with Wherobots Posted on July 9, 2025October 4, 2026 by Ben Pruden Introduction At Wherobots, we’re excited to offer access to the Foursquare Open Places dataset through the Wherobots Global Hub. We maintain a pipeline to keep this dataset updated along with the Foursquare releases. In this tutorial, we’ll show you how to work with Foursquare’s Open Places dataset using Wherobots. We’ll demonstrate how to access, query, and visualize Points of Interest (POI) data, culminating in aggregating and visualizing the data as a choropleth map. The overall process we will show is applicable to many different use cases; for this tutorial we will use it to analyze coffee shops across San Francisco and see how they are distributed by neighborhood. If you would like to follow along, you can spin up a free instance here and access this notebook for yourself in the Wherobots Jupyter Notebook environment under the examples folder. Let’s dive in. What is the Foursquare Places Dataset? Foursquare’s Places dataset is a comprehensive global collection of points of interest (POIs) containing over 100 million locations worldwide. What makes this dataset particularly valuable is that it’s continuously updated and verified through Foursquare’s Placemaker Tools, which enable community contributions to maintain accuracy. The dataset includes details such as: Business names and locations Categories and classifications Address information Social media links Opening dates and closure information Tutorial To follow along with this tutorial, you’ll need a Wherobots account. You can sign up for free at Wherobots.com. Setting Up the Environment First, let’s set up our Sedona context in Wherobots. This will start up our spark-based compute environment and give us access to the geospatial functions we need to work with the Foursquare dataset. from sedona.spark import * config = SedonaContext.builder().getOrCreate() sedona = SedonaContext.create(config) Exploring the Foursquare Data in Wherobots The Foursquare data is already accessible in the Wherobots Open Data Catalog, so there’s no need to download or import it separately. Let’s check what tables are available: sedona.sql("SHOW tables IN wherobots_open_data.foursquare").show(truncate=False) +----------+----------+-----------+ |namespace |tableName |isTemporary| +----------+----------+-----------+ |foursquare|categories|false | |foursquare|places |false | +----------+----------+-----------+ There are two tables: places which contain the actual locations, and categories which provide the classification system. Let’s look at the schema of the places table: sedona.table("wherobots_open_data.foursquare.places").printSchema() You’ll get a result that looks something like this: root |-- fsq_place_id: string (nullable = true) |-- name: string (nullable = true) |-- latitude: double (nullable = true) |-- longitude: double (nullable = true) |-- address: string (nullable = true) |-- locality: string (nullable = true) |-- region: string (nullable = true) |-- postcode: string (nullable = true) |-- admin_region: string (nullable = true) |-- post_town: string (nullable = true) |-- po_box: string (nullable = true) |-- country: string (nullable = true) |-- date_created: string (nullable = true) |-- date_refreshed: string (nullable = true) |-- date_closed: string (nullable = true) |-- tel: string (nullable = true) |-- website: string (nullable = true) |-- email: string (nullable = true) |-- facebook_id: long (nullable = true) |-- instagram: string (nullable = true) |-- twitter: string (nullable = true) |-- fsq_category_ids: array (nullable = true) | |-- element: string (containsNull = true) |-- fsq_category_labels: array (nullable = true) | |-- element: string (containsNull = true) |-- placemaker_url: string (nullable = true) |-- geometry: geometry (nullable = true) |-- bbox: struct (nullable = true) | |-- xmin: double (nullable = true) | |-- ymin: double (nullable = true) | |-- xmax: double (nullable = true) | |-- ymax: double (nullable = true) The schema includes fields like fsq_place_id, name, latitude, longitude, address, country, fsq_category_labels, and more. Importantly, it also includes a geometry field that we can use for spatial operations. Working with Versions Wherobots provides access to multiple versions of the Foursquare data. You can select a specific version using the VERSION AS OF clause: sedona.sql("SELECT * FROM wherobots_open_data.foursquare.places VERSION AS OF 'dt=2025-02-06'").show(5, truncate=True) If you don’t specify a version, you’ll get the latest one by default. Basic Analysis of the Dataset Let’s start by selecting the columns we’re interested in: places_df = sedona.sql(""" SELECT fsq_place_id, name, geometry, fsq_category_labels, country, date_refreshed, date_created FROM wherobots_open_data.foursquare.places WHERE date_closed IS NULL AND name IS NOT NULL AND geometry IS NOT NULL AND country IS NOT NULL """) places_df.createOrReplaceTempView("places") Now, let’s run a quick query to see what we’re working with and get some basic stats: sedona.sql(""" SELECT COUNT(*) as total_places, COUNT(CASE WHEN name IS NOT NULL THEN 1 END) as has_name, COUNT(CASE WHEN address IS NOT NULL THEN 1 END) as has_address, COUNT(CASE WHEN fsq_category_labels IS NOT NULL THEN 1 END) as has_categories, COUNT(CASE WHEN date_closed IS NULL THEN 1 END) as still_open FROM places """).show() When you run the above query, you’ll get a result like this: +------------+---------+-----------+--------------+----------+ |total_places| has_name|has_address|has_categories|still_open| +------------+---------+-----------+--------------+----------+ | 104635092|104635092| 67380672| 92959858| 98436217| +------------+---------+-----------+--------------+----------+ That’s over 104 million places in the dataset. About 67 million have addresses, and nearly 93 million have category labels. That’s an impressive dataset. Filtering by Region Let’s narrow our focus to a specific region. We can first use standard SQL to filter to the United States: us_df = sedona.sql(""" SELECT * FROM places WHERE country = 'US' """) This approach using standard SQL works well for cases where the region you want to filter for is already defined in the dataset. For example, the dataset has a country column, so we can filter by country. But what if we wanted to filter by a geographic region that isn’t already defined in the dataset? For example, maybe we want to filter by census block groups, cities, or internal company polygons that we have defined for trade areas. That’s where the power of spatial joins in Wherobots comes in. We can define any arbitrary polygon and use it to filter the data. Note that the actual polygon is a very long WKT string, so we truncate it here in the example code. But the full WKT polygon is defined in the notebook itself for you to access. As an alternative you could use other open boundary datasets, such as within the Overture Divisions dataset and join the Foursquare places against those divisions. # Define a polygon for San Francisco SF = "MULTIPOLYGON (((-122.4773830027518 37.8110279985324, ...))" # Filter to places within San Francisco sf_df = sedona.sql(f""" SELECT * FROM places WHERE ST_Contains( ST_GeomFromWKT('{SF}'), geometry) """) Now we have just the places located within San Francisco. I know that filtering the data for a single polygon may not be a huge task by itself, but where you really start to see the power of Wherobots is when this scales up. We could just as easily spatially join all 104 million places in the dataset into a million different polygons (buildings, cities, counties, senate districts, zip codes, etc…) based on which polygon a point falls within–and we could use that exact same ST_Contains function to do so. Visualizing the Data Wherobots includes built-in visualization capabilities through SedonaKepler. Let’s create a simple map of our San Francisco places: sf_map = SedonaKepler.create_map(sf_df, "Places") We can see from the image that there are POIs clear across San Francisco, with a dense concentration of places in the Financial District downtown. Searching for Specific Places Now say we want to search for specific places such as a brand or retail chain. For example, perhaps we want to find all Starbucks locations in San Francisco. To do so we can use the following query: sbux_df = sedona.sql(""" SELECT * FROM sf_places WHERE LOWER(name) = 'starbucks' """) sbux_map = SedonaKepler.create_map(sbux_df, "Starbucks in San Francisco") Exploring Categories The Foursquare data uses a hierarchical category system. Let’s see what categories are most common in our San Francisco dataset: sedona.sql(""" SELECT fsq_category_labels[0] as primary_category, COUNT(*) as count FROM sf_places GROUP BY fsq_category_labels[0] ORDER BY count DESC LIMIT 20 """).show(truncate=False) Here are the top categories in San Francisco by count: Primary CategoryCountNULL8597Business and Professional Services > Office6069Business and Professional Services > Office > Tech Startup3615Community and Government > Residential Building > Apartment or Condo2505Health and Medicine > Physician > Doctor’s Office2439Landmarks and Outdoors > Structure2048Business and Professional Services > Health and Beauty Service > Hair Salon1492Arts and Entertainment > Art Gallery1474Travel and Transportation > Transportation Service > Public Transportation > Bus Line1238Dining and Drinking > Restaurant1238 We can also filter to places that fall into a specific category. For example, instead of searching for Starbucks locations by name, let’s select all Points of Interest that are in the Coffee Shop category: category_places = sedona.sql(""" SELECT * FROM sf_places WHERE ARRAY_CONTAINS(fsq_category_labels, 'Dining and Drinking > Cafe, Coffee, and Tea House > Coffee Shop') """) Here’s a quick look at the table that we get after running that query. +--------------------+--------------------+--------------------+--------------------+-------+--------------+------------+ | fsq_place_id| name| geometry| fsq_category_labels|country|date_refreshed|date_created| +--------------------+--------------------+--------------------+--------------------+-------+--------------+------------+ |5622e819498ecbeed...|Four Barrel at TI...|POINT (-122.37331...|[Dining and Drink...| US| 2024-10-26| 2015-10-18| |49d6619ff964a520b...| Cafe deStijl|POINT (-122.40036...|[Dining and Drink...| US| 2022-07-29| 2009-04-03| |49dbc22af964a520f...|Battery Street Co...|POINT (-122.40133...|[Dining and Drink...| US| 2024-05-06| 2009-04-07| |5b1599d604d1ae002...| Good Mojo|POINT (-122.40049...|[Dining and Drink...| US| 2024-10-26| 2018-06-04| |4b02f2f9f964a5205...| Starbucks|POINT (-122.40102...|[Dining and Drink...| US| 2024-11-05| 2009-11-17| |49fa3e4ff964a520d...| Jackson Place Cafe|POINT (-122.40135...|[Dining and Drink...| US| 2024-07-17| 2009-05-01| |4d065736a26854819...|Réveille Coffee C...|POINT (-122.40035...|[Dining and Drink...| US| 2024-10-20| 2010-12-13| |4bafe772f964a5205...| om bucks|POINT (-122.40090...|[Dining and Drink...| US| 2023-04-04| 2010-03-28| |4ab317e7f964a5207...| Peet's Coffee|POINT (-122.40115...|[Dining and Drink...| US| 2023-09-20| 2009-09-18| |2af323a0d2d542feb...| Aroma Espresso Bar|POINT (-122.40102...|[Dining and Drink...| US| 2012-08-27| 2012-08-27| +--------------------+--------------------+--------------------+--------------------+-------+--------------+------------+ Creating a Choropleth Map: Coffee Shops by Neighborhood For a more advanced visualization, let’s create a choropleth map showing the number of coffee shops in each San Francisco neighborhood. First we will need a dataset of polygons for the SF neighborhoods. Luckily, we have one from the SF open data portal. We’ll load these from a CSV file: SF_NEIGHBORHOODS_URL = "s3://wherobots-examples/data/sf_neighborhoods.csv" neighborhoods = (sedona.read.format('csv') .option('header', 'true') .option('delimiter', ',') .option('inferSchema', 'true') .load(SF_NEIGHBORHOODS_URL) ) For context, here’s the table that results after we load those neighborhood polygon boundaries. +--------------------+--------------------+ | the_geom| neighborho| +--------------------+--------------------+ |MULTIPOLYGON (((-...| Seacliff| |MULTIPOLYGON (((-...| Haight Ashbury| |MULTIPOLYGON (((-...| Outer Mission| |MULTIPOLYGON (((-...| Inner Sunset| |MULTIPOLYGON (((-...|Downtown/Civic Ce...| |MULTIPOLYGON (((-...| Diamond Heights| |MULTIPOLYGON (((-...| Lakeshore| |MULTIPOLYGON (((-...| Russian Hill| |MULTIPOLYGON (((-...| Noe Valley| |MULTIPOLYGON (((-...| Treasure Island/YBI| +--------------------+--------------------+ As you can see, there are two columns, the geometry column that contains the multi-polygon that defines the boundaries of the neighborhood and the neighborhood column that contains the name. If you’ve spent any time in SF you’ll certainly recognize some of these names. Now that we have the boundaries we’ll use a spatial join to count coffee shops in each neighborhood: neighborhood_agg = sedona.sql(""" SELECT ST_GeomFromWKT(n.the_geom) AS geometry, n.neighborho AS neighborhood, COUNT(*) as location_count, collect_list(p.name) AS coffee_shops FROM category_places p JOIN neighborhoods n ON ST_CONTAINS(ST_GeomFromWKT(n.the_geom), p.geometry) GROUP BY n.the_geom, n.neighborho """) This gives us the following table where we can see the neighborhood name, its polygon geometry, the count of coffee shops contained in that neighborhood, and an array of the names of the coffee shops. +--------------------+--------------------+--------------+--------------------+ | geometry| neighborhood|location_count| coffee_shops| +--------------------+--------------------+--------------+--------------------+ |MULTIPOLYGON (((-...| Treasure Island/YBI| 1|[Four Barrel at T...| |MULTIPOLYGON (((-...| Potrero Hill| 25|[Starbucks, WFM C...| |MULTIPOLYGON (((-...| South of Market| 119|[W6 Coffee Bar, S...| |MULTIPOLYGON (((-...| Bayview| 20|[The Happy Vegan,...| |MULTIPOLYGON (((-...| Financial District| 186|[Starbucks, Jacks...| |MULTIPOLYGON (((-...| Visitacion Valley| 4|[Joe Leland, Miss...| |MULTIPOLYGON (((-...| Chinatown| 11|[Cafe Vivo, Latte...| |MULTIPOLYGON (((-...|Downtown/Civic Ce...| 98|[Chai Bar, Dignit...| |MULTIPOLYGON (((-...| North Beach| 30|[Cafe deStijl, Ba...| |MULTIPOLYGON (((-...| Nob Hill| 22|[Cafe Mozart, Gal...| +--------------------+--------------------+--------------+--------------------+ Finally we have what we need to create our visualization. Time to create our choropleth map. Note here that I am using a custom map config to set the color scale, style, and other properties. map = SedonaKepler.create_map(df=neighborhood_agg, name="Coffee Shop Count", config=map_config) The resulting visualization reveals interesting patterns. Downtown San Francisco (the Financial District) has the highest concentration with 186 coffee shops, while areas like the Presidio are relatively sparse with only around 10 options. The Inner Richmond neighborhood shows about 37 coffee shops. With these live Sedona Kepler maps, you can hover over the neighborhood you’re interested in to see the full list of coffee shops in that area. That wraps up our tutorial on working with Foursquare Places data in Wherobots. I hope you enjoyed learning about these powerful geospatial capabilities. Before we wrap up, let’s take a quick look at what makes this dataset from Foursquare unique. What Makes Foursquare’s Places Data Unique? Unlike some other POI datasets, Foursquare’s Places data is continually updated through their Placemaker Tools. These web-based tools allow community members to: Add new venues Update existing information Validate place data This community-driven approach helps ensure the data remains current and accurately reflects real-world changes. Conclusion The Foursquare Places dataset combined with Wherobots’ powerful geospatial capabilities offers an excellent foundation for location-based analytics. Whether you’re studying urban patterns, analyzing business distributions, or planning a coffee crawl through San Francisco, these tools make it easy to extract meaningful insights from spatial data. Ready to explore the Foursquare Places data for yourself? Sign up for a free Wherobots account and try out this notebook! Bonus: Video Walk Through By the way, a couple months ago I released this video where I walk through this exact use case in one of our Wherobots demo notebooks. Feel free to check it out. Get started for free Try Now Key takeawaysFoursquare Open Places is in the Wherobots Open Data Catalog as two tables—places and categories—kept current with Foursquare releases. You can pin a snapshot with VERSION AS OF (example: dt=2025-02-06) or read the latest by default.A catalog query in the tutorial counted 104,635,092 places; 67,380,672 had addresses, 92,959,858 had category labels, and 98,436,217 were still open (date_closed IS NULL).Filter with ordinary SQL (country = US) or with a spatial predicate such as ST_Contains against any polygon—census blocks, trade areas, or a San Francisco WKT. The same ST_Contains pattern scales to joining all 104 million places into many polygons.The worked example maps San Francisco coffee shops by neighborhood: Financial District 186, South of Market 119, Inner Richmond about 37, Presidio around 10. Top SF primary categories include Office (6,069) and Tech Startup (3,615).Places are community-updated through Foursquare Placemaker Tools. Follow along in the example notebook under the Wherobots Jupyter examples folder on a free instance.
Wherobots received its SOC 2 Type 2 attestation Posted on June 18, 2025September 1, 2026 by Jia Yu We are thrilled to announce Wherobots has received its SOC 2 Type 2 attestation report, reinforcing our commitment to data security and data privacy for our customers. SOC 2 (Systems and Organization Controls) is a standard trusted by industry leaders and a requirement for many enterprises engaging with software providers. It evaluates an organization’s information security practices, ensuring that controls are prompt and effective, and that data is kept secure and confidential. While a Type 1 attestation assesses policies and procedures at a single point in time, the Type 2 attestation we have received demands a rigorous, in-depth evaluation and audit of the effectiveness of the implemented controls over time. To view our SOC 2 Type 2 report, or more information on our security policies, please view our Trust Center or contact us at security@wherobots.com. Why this matters There are billions of devices roaming the world, logging petabytes of data from trips, activities, and events. Satellites and drones are scanning the world, capturing what’s happening on Earth, how its changing, and how humans terraforming it. This data can be highly sensitive and highly valued. With Wherobots, businesses can utilize this data from the physical world at a distinguished scale, price-performance, and ease, while keeping data secure and in their control. With Wherobots’ cloud-native Lakehouse architecture, you can bring Wherobots’ capabilities for spatial and non-spatial data analytics and AI at planetary scale – right to your data, wherever it lives. Security first Security is not just a feature, it’s part of our engineering culture and infused into how we design and build our software, our internal systems, and our production environments. Wherobots Cloud was developed from the ground up with best practices and secure-by-design principles that work backwards from data security first. In its architecture, the Wherobots Cloud control plane is isolated from its compute plane, and each workload is isolated from the cloud hypervisor up to create a trusted environment for your data. This architecture is serverless by default, but can also run its compute plane in your own cloud VPC (BYOC). Wherobots is the distinguished and obvious solution for spatial computation and AI on the lakehouse. Our SOC 2 Type 2 report, available on the Wherobots Trust Center, is now available to expedite procurement, vendor, and security reviews. Want to learn about more Wherobots security capabilities? Visit our docs on Getting Started with Wherobots to review service principles, audit logs, SAML SSO, and more. Get Started with Wherobots START BUILDING Key takeawaysWherobots has received a SOC 2 Type 2 attestation report, the industry-trusted audit of security and confidentiality controls over a period of time—not a Type 1 point-in-time snapshot of policies.The report is available from the Wherobots Trust Center or by emailing security@wherobots.com, and is intended to speed procurement, vendor, and security reviews.Wherobots Cloud isolates the control plane from the compute plane, and isolates each workload from the hypervisor up. The service is serverless by default and can also run its compute plane in the customer AWS VPC (BYOC).The lakehouse architecture brings spatial analytics and AI to data where it already lives, rather than requiring customers to surrender custody of sensitive trip, activity, satellite, or drone data.Docs cover service principles, audit logs, SAML SSO, and related security capabilities for teams going through enterprise review.
Wherobots, the Spatial Intelligence Cloud, is Now Available in AWS Europe Posted on April 14, 2025October 3, 2026 by Tiffany Huynh The EU is widely recognized as a world leader for climate solutions, automotive design and manufacturing, mobility systems and analysis, environmental monitoring, agriculture and precision farming, and urban development. Geospatial data is foundational to the success of these innovations. However most of the new technology is developed using tools and cloud services not optimized for geospatial development. Compared to internet data, support for geospatial data in the modern cloud environments has lagged. This technology gap has made working with geospatial data expensive, and required staffing data teams with unique expertise. Introducing Wherobots for the AWS Europe (Ireland) Region We are excited to announce that Wherobots is ready for EU native workloads. This expansion makes it possible for many EU companies to approve the use of Wherobots and adhere to their data residency requirements by processing and storing data in-region. Wherobots is the Spatial Intelligence Cloud Wherobots’ mission is to make it easy for our customers to utilize geospatial data. We are delivering on it via a cloud optimized for developing and running solutions about the physical world, at any scale. Our purpose-built, cloud native approach is enabling teams at AddressCloud and Overture Maps Foundation to accelerate their pace of innovation with geospatial data. Their workloads run up to 20x faster after migrating from popular cloud-based engines, developer productivity is boosted with the most feature complete development experience for SQL and Python, and costs are reduced, putting new solutions in reach. A Cloud Native Lakehouse Architecture The architecture of Wherobots is cloud native, and is deeply rooted in open source. Apache Sedona, the open source geospatial engine for Apache Spark, Apache Flink, and Snowflake, is 100% compatible with Wherobots. Users can easily lift and shift their Apache Sedona based applications into Wherobots with zero code changes. Wherobots is also one of the leading companies bringing GEO support into popular open file and open table formats like Parquet and Iceberg, and uses these formats by default. That way, you can deploy various engines on your data, benefit from the advantages of a Lakehouse engine such as ACID transactions and table versioning, without locking your data into proprietary vendor siloes, or inelastic solutions that couple storage with compute. Getting Started Getting started is easy. To use Wherobots within the AWS Europe (Ireland) region, get srated with the Professional Edition on the AWS Marketplace. Create a notebook and explore one of many examples designed to help you realize what you can create using SQL and Python. There’s no infrastructure to manage. Teams just use and pay for Wherobots usage on-demand via Wherobots Spatial Units, which reflect the amount of serverless computation consumed. If there are other clouds or regions that you’re interested in using beyond the ones we currently support, please reach out to us at product@wherobots.com or fill out this form here. You can read more about Wherobots on the website or by exploring our product documentation. Create a Pro tier account on AWS Get Started Key takeawaysWherobots is available for EU-native workloads in the AWS Europe (Ireland) region, so teams can process and store data in-region to meet data-residency requirements.The launch targets EU strengths in climate, automotive, mobility, environmental monitoring, precision agriculture, and urban development—areas where cloud tools have lagged for geospatial compared with internet data.Customers such as AddressCloud and Overture Maps Foundation are cited as running workloads up to 20x faster after migrating from popular cloud engines, with a feature-complete SQL and Python experience and on-demand Spatial Units billing.Architecture is a cloud-native lakehouse: 100% Apache Sedona compatible (lift-and-shift with zero code changes), defaulting to open Iceberg and Parquet/GEO formats so storage stays decoupled from compute.Start on Professional Edition via AWS Marketplace, open a notebook, and run examples. There is no infrastructure to manage. Requests for additional clouds or regions go to product@wherobots.com.
Spatial Intelligence Newsletter: Location Intelligence w/ Isochrones, Overture Places, Cloud-Native Geospatial, Iceberg and More Posted on April 10, 2025October 3, 2026 by Tiffany Huynh Welcome to the April edition of the Spatial Intelligence Newsletter! This month, we’re covering the benefits of using Apache Iceberg, spatial joins, cloud-native geospatial, and new product updates like isochrones to help you make better location-based decisions. What does it all mean, and how can it help you increase data productivity? Check it out here! 👇 ⏰💰Hurry, time is running out! We’re currently offering a FREE $400 credit when you subscribe to the Professional edition of Wherobots, which includes exclusive features like GeoAI with WherobotsAI Raster Inference, map matching for cleaning messy GPS data, new travel isochrones for better location-based decisions, and the ability to bring your own cloud storage, just to name a few. In addition, with the release of our drive-time isochrones, we’re now offering free access to Overture Places data—enriched with drive-time isochrones across every location in the U.S.—through our Pro tier data catalog. And we will be maintaining that dataset with every release in the future, so you’ll be able to use it going forward for location intelligence. There’s no obligation to get started, so be sure to take advantage of this (it’s like free money). Offer ends on May 31st, so don’t wait! Latest Content Benefits of Apache Iceberg for geospatial data analysis 🧊 Apache Iceberg support for GEO data brings a significant modernization for geospatial data and solutions. This support makes it easier for you to bring geospatial data into an open data architecture that decouples compute and storage and lower your costs. By adopting Iceberg in a data lake, you’re enabling your team to leverage the right tool for the job without needing to worry about locking your data into a vendor or a database solution that doesn’t scale. Additionally, traditional file formats and row-oriented databases struggle when scaling beyond a million features, often performing poorly or only accommodating data that fits comfortably in memory. 😩 Iceberg, built on Parquet, solves this with lightning fast reads, scalability for larger-than-memory datasets, and developer friendly features like DML operations. Plus with added capabilities like versioning and time travel, users can query both current and historical data seamlessly. 🔍 Follow along this post to learn how to use Apache Iceberg with Sedona and find out how these features benefit spatial computations. Cloud-Native Geospatial: More Than Just Big Data e💡We had a very insightful discussion with Amy Rose (CTO) from Overture Maps and Eshwaran Venkat (CTO & Co-Founder) from Dotlas on cloud-native geospatial technology. Here are some highlights: Cloud-native geospatial is not just for big data; it’s more accessible than you might think. You should be able to work with spatial data the way you work with any other data type. Increasing deliverability and breaking down data silos: Non-spatial communities can now work with spatial data. Compute systems that make the process more scalable, accessible, elastic, and cost-efficient. How Dotlas and Overture Maps are optimizing their data pipelines, achieving performance gains, and improving cost efficiency. Spatial Joins at Scale: Unlocking Advanced Geospatial Analytics with Wherobots 🌎🤝 Spatial joins are essential for geospatial data analysis, but it can be slow or computationally expensive when working with large-scale datasets. Follow along in this tutorial as we walk through how easy and cost-effective it is to: Join datasets using spatial predicates like ST_Intersects to combine facilities with administrative boundaries and efficiently find the k-nearest neighbor with the ST_AKNN function. Apply spatial filters and improve performance through strategies like partitioning by geohash Take your geospatial data analytics to the next level and ensure spatial joins aren’t a bottleneck in solving your business challenges. Apache Sedona Sedona Success Story: Optimizing ETL pipelines at scale with Comcast 📊 Some of the challenges that Comcast was trying to overcome was data volume and repeatability. That’s why David Buchanan, GIS Architect, turned to Apache Sedona, which allowed him to reduce processing times from 5 hours to 30 minutes compared to GeoPandas. Watch the recording to learn more. Apache Sedona Office Hours 😎 We just released Sedona 1.7.1, with some new features : SQL interface for GeoStats (ST_DBSCAN, ST_GLocal, ST_LocalOutlierFactor) Broadcast join support for distributed KNN Join STAC catalog & OpenStreetMap (OSM) PBF reader New ST functions like ST_RemoveRepeatedPoints If you missed the office hour, check out the recording to learn more about the latest release. And don’t forget to mark your calendar for the next office hour! 🗓️ Product Updates Overture Places with Isochrones Dataset: Accelerate accessibility analysis with a ready-to-use dataset containing pre-calculated 5, 10, 15, and 20-minute driving isochrones for millions of US Overture Places (pro+). New ST Isochrones functions: Make data-driven decisions on logistics, site selection, and market reach using Wherobots’ travel isochrone functions in SQL or Python (pro+). Audit Logs: Admins gain enhanced security and accountability insights by using Wherobots’ detailed, exportable audit logs to track key Organization actions and system events (pro+). STAC Reader: Simplify workflows and accelerate queries by loading STAC geospatial datasets directly into Sedona DataFrames in Wherobots (OSS & community+). Job Run Monitoring: Visually track job execution, analyze resource usage, and manage runs directly within Wherobots for enhanced control and optimization (pro+). Idle Timeout for Notebooks: Gain control over notebook runtime costs and resource usage with customizable idle timeouts that automatically terminate inactive notebooks (community+). 🆓 Both the Overture Places with isochrones dataset and isochrone functions, as well as the audit logs and job run monitoring, are available exclusively in the Pro tier. Take advantage of the free trial (ending soon!) to try these features and see how they can help solve some of the bottlenecks you might be facing when working with spatial data. Upcoming Events Geospatial Tables in the Open Lakehouse: A New Era for Iceberg and Parquet Wednesday, May 5 at 9AM PT | Virtual It’s easier than ever to work with geospatial data, with Iceberg and Parquet now offering powerful solutions for both geospatial experts and non-spatial professionals. Join this livestream with leaders from Foursquare, Databricks, Planet, and Wherobots as they discuss the historical challenges of handling spatial data, bridging the gap, and future adoption of these advancements. Apache Sedona + Iceberg GEO Meetup Monday, May 12 at 5:00PM PT | San Francisco, California Join us for a fun and informative evening as we explore Apache Iceberg’s new native geospatial support, designed to solve major challenges in managing geospatial data at scale. This will be a great opportunity to connect with professionals in the field to learn about the latest developments in spatial data, as well as exciting projects people are working on.🌟 Featured speakers: Jia Yu, Co-Founder and Chief Architect, Wherobots Matt Forrest, Director of Customer Engineering and PLG, Wherobots Yingjun Wu, Founder and CEO, RisingWave Labs CNG Conference April 30 – May 2 | Snowbird, Utah We’re excited to attend the upcoming CNG Conference! Be sure to check out these sessions: Day 1 1:15pm-2:45pm | Workshop: Interfacing with Cloud-Native Overture Data and the GERS Ecosystem – Sean Knight Day 2 9:45am-11:15am | Track 2: Introducing geospatial support in Apache Iceberg – Matthew Powers 11:45am-1:15pm | Extract insights from satellite imagery at scale with WherobotsAI – Damian Wylie 4:30pm-5:00pm | Plenary Panel: Builders Panel – Mo Sarwat 👥 If you’ll be at the conference, we’d love to meet with and chat about how you’re working with geospatial data. Feel free to reach out if you’d like to schedule a time to connect! Key takeawaysThis April 2025 newsletter roundup points to Iceberg-for-geospatial, cloud-native geospatial with Overture and Dotlas, spatial-join tutorials, and product launches—not original benchmark figures.A Professional Edition offer included a free $400 credit through May 31st, covering GeoAI Raster Inference, map matching, travel isochrones, and bring-your-own cloud storage.Pro catalog added Overture Places enriched with 5-, 10-, 15-, and 20-minute U.S. driving isochrones, plus new ST_Isochrone functions in SQL and Python. Other Pro items: audit logs and job-run monitoring. STAC reader and notebook idle timeout landed for Community+.Sedona 1.7.1 shipped SQL GeoStats (ST_DBSCAN, ST_GLocal, ST_LocalOutlierFactor), broadcast KNN joins, STAC and OSM PBF readers, and ST_RemoveRepeatedPoints. A Comcast success story cut ETL from 5 hours to 30 minutes versus GeoPandas.Upcoming at the time: an Iceberg/Parquet livestream on May 5, a San Francisco Iceberg GEO meetup on May 12, and CNG Conference sessions in Snowbird April 30-May 2.
Wherobots is ready for AWS workloads Posted on November 26, 2024October 3, 2026 by Ben Pruden Planetary-scale geospatial solutions are now accessible via Wherobots on the AWS Marketplace We’re thrilled to announce that Wherobots is generally available for AWS customers with pay-as-you-go pricing via the AWS Marketplace. AWS customers can subscribe to a 30 day, free trial of the Wherobots Professional Edition for up to $400 in usage, and discover how easy it is to create spatial solutions that propel their business forward. The integration with the AWS Marketplace simplifies the Wherbots buying and usage experience, particularly those with AWS commitments or discounts that apply to AWS Marketplace spend. Coupled with a secure integration to run Wherobots on private or public S3 buckets, Wherobots is where the next generation of geospatial solutions are developed on AWS. The potential of spatial data is very high Many companies have significant investments in assets, products, or services that are influenced by our dynamic world. To be competitive, adaptive, and profitable, companies need to accelerate the velocity of decisions about these investments. Insurance companies like State Farm need solutions for calculating asset risk in the face of a rapidly changing climate. Telecommunications providers like Comcast optimize their network operations to account for bandwidth constraints from physical barriers like buildings, tunnels, and weather. Retailers like Starbucks need to identify the next retail location to launch or sunset based on mobility data, supply and demand, and demographic trends. Logistics providers like Amazon Last Mile Delivery need to autonomously adjust distribution and delivery plans based on roadway conditions, road updates, traffic conditions, and delivery payloads. Farming and forestry operations are looking for ways to optimize yields in a way that’s sustainable and profitable. Solar and wind farm operators are identifying the next best locations to develop sustainable power sources, and also connect to expanding charging station networks across the globe. These industries are modernizing with the cloud, but services in the cloud haven’t made it easy to create these types of solutions, until now. Spatial ideas are everywhere, but solutions are sparse If these types of solutions resonate with your business, I’d wager that if you asked your teams to produce ideas that rely on geospatial data, they’d have a lot to share. But the reality is most teams are not enabled to unlock these ideas. They will say it’s really hard, if not infeasible to turn these ideas into solutions that propel your business forward. Which also means the ideas are put on the back burner or considered far-fetched. Wherobots makes spatial solutions accessible Built on and fully compatible with Apache Sedona, Wherobots is the Spatial Intelligence Cloud. Wherobots offers geospatial ETL, analytics, and AI solutions that make it easy for data scientists and engineers to create spatial data products, and the intelligence that drives their business forward. Wherobots delivers industry leading scalability and spatial computing performance on your data lake. It’s up to 20x more performant than Apache Sedona and Apache Spark using a serverless analytics and inference engine optimized for spatial operations. Development is unified across what are otherwise siloed data types – raster (satellite and drone imagery) and vector (mobility data, polygons, trips, roads) data. And solutions can be built with SQL, Python, or Scala in a notebook, putting solutions in-reach to common developers. There’s 300+ built in functions and higher level features, like Map Matching, Geostats, and Raster Inference to accelerate development. With pay-as-you-go pricing on the AWS Marketplace, Wherobots is where the next generation of spatial data solutions are built. What are customers saying? Wherobots customers like AddressCloud and Overture realized typical performance gains of 5-20x, lower costs, and objectively higher developer productivity after migrating their Apache Sedona workloads into Wherobots. AddressCloud helps insurers calculate geographic risk “Wherobots has given us the potential to run jobs that used to take hours or days to minutes and removed the need to think about provisioning compute. As we provide perils information (flood, fire, etc) to insurers at the property level, we particularly appreciate the ability to be able to run combined vector/raster analysis, without having to previously transform the raster data into vector format or some other format,” said John Powell, Senior Geospatial Data Engineer at Addresscloud.” Overture provides current and next-generation map products by creating reliable, easy-to-use, and interoperable open map data “Overture produces a building dataset covering all buildings in the world, with 2.3B geometries and growing, that’s updated frequently. There’s a lot of data and compute that goes into producing and keeping it up to date,” said Jennings Anderson, Geoscientist at Overture and Data Engineer at Meta. “We accelerated the pipelines that produce the buildings dataset by up to 20x after we moved them to Wherobots, which required a simple redirection of our code. We retained compatibility with Apache Sedona, and the move put us into a development experience that’s made us more productive.” Getting started is easy. From your AWS account, subscribe to the Professional Edition of Wherobots on the AWS marketplace. You can get started risk-free in the professional edition with features required by production workloads. We’ve built tens of example notebooks to help you go from zero to iterating with spatial data in minutes. Feedback? Our mission is to make it easy for our customers to utilize geospatial data. We cannot complete our mission without your input, and are working with a variety of customers to shape what we do next. If you are invested in the problems we are solving, and have an idea to improve our product, please contact us at feedback@wherobots.com, or contact me directly at damian@wherobots.com. I’m eager to hear from you, and we are there to help you innovate with Wherobots on AWS. Try Wherobots on AWS Marketplace Get Started Key takeawaysWherobots is generally available to AWS customers as pay-as-you-go on AWS Marketplace, with a 30-day Professional Edition trial covering up to $400 in usage.The Spatial Intelligence Cloud is built on Apache Sedona and is described as up to 20x more performant than Apache Sedona and Apache Spark, with a serverless engine for spatial analytics and inference on the customer data lake.Development covers raster (satellite and drone imagery) and vector (mobility, polygons, trips, roads) in one place, using SQL, Python, or Scala in notebooks, with 300+ built-in functions plus Map Matching, GeoStats, and Raster Inference.AddressCloud reports jobs that used to take hours or days now run in minutes, including combined vector/raster peril analysis without converting rasters first.Overture’s global buildings dataset—2.3 billion geometries and growing—accelerated by up to 20x after moving pipelines to Wherobots with a code redirect, while staying Apache Sedona compatible.
Announcing Our 21.5M Series A :: Unlocking Answers to Planetary-scale Questions. Posted on November 26, 2024October 3, 2026 by Ben Pruden Unlocking answers to planetary-scale questions. By Wherobots co-founders Mo Sarwat and Jia Yu Each day, satellites, drones, applications, and GPS devices generate petabytes of spatial data that can be used to solve real-world problems. But the majority of this data is stuck in siloed legacy systems or sits idle and disjointed. We see the potential this data can have for business, the planet, government, and societies. And we’re on a mission to help companies fully utilize it so they can tackle issues like how to manage their fleets of vessels and vehicles, where and how to build infrastructure, and determine the best methods to assess and mitigate risk of catastrophic natural disasters. To achieve our mission, we’re partnering with leading investors in the technology space and have raised $21.5M in Series A funding—led by Felicis, with continued support from Wing Venture Capital and Clear Ventures and participation from JetBlue Ventures and P7 Ventures. Aydin Senkut, Founder and Managing Partner at Felicis will also be joining Wherobots’ Board of Directors together with Peter Wagner from Wing Venture Capital. We are committed to constantly improving our technology to process and analyze geospatial data faster and more efficiently and this funding will accelerate our product development and go-to-market operations. From Research to Market Growing up in Egypt, Mo saw the impact of climate change first-hand. Rising temperatures and pollution are threatening the air quality and water supply of millions of Egyptians. These challenges, among others, are not isolated—they reflect global issues as our world changes faster than ever. Motivated by these realities, we set out to harness data that captures what’s happening in the physical world to drive innovation and empower people to tackle both large-scale and localized problems. We met when Mo was a professor and Jia was finishing his PhD at ASU. We bonded over our shared passion for geospatial data and its untapped use cases. We realized the popular data warehousing and analytics solutions available were built from the ground up to process internet data, not geospatial data. When geospatial data is forced into these systems, they underperform, lack essential features for intuitive geospatial analysis, and are often either closed-source or reliant on outdated architectures. These limitations make geospatial solutions inaccessible for most organizations. Recognizing this gap, we set out to create a solution tailored to the unique challenges of geospatial data, unlocking its power for organizations of all sizes. Apache Sedona — an open-source geospatial compute framework — was our first response to this issue. Today Apache Sedona has over 40M downloads and is now used to run planetary-scale workloads by companies like Amazon.com for last mile delivery and Land O’ Lakes for precision agriculture. After years of growing Apache Sedona, we saw a tremendous appetite for a more in-depth enterprise solution. Enter Wherobots, a fully managed, scalable cloud platform that is purpose-built to make geospatial solutions easy to create while maintaining compatibility with Apache Sedona. Wherobots also integrates well with the modern data and AI ecosystem, making it a plug-n-play option for Fortune 1000 companies to derive value from the geospatial data they collect. Putting Data to Work Wherobots’ Spatial Intelligence Cloud empowers businesses to unlock planetary-scale solutions and put their spatial data to work. Data teams are able to solve problems faster and more efficiently on a compute engine that’s optimized for spatial analytics, a broad set of functions in SQL, Python, and Java, with a variety of native geospatially specific functions, as well as the ability for customers to bring in their own AI and ML models to drive insight from the physical world. This makes Wherobots a far more productive system for data teams to get their work done without switching contexts. Using Wherobots, industries across financial services & insurance, transportation, logistics & supply chain, energy, agriculture, and social services can analyze real-world issues up to 20x faster at a planetary-scale. This means more informed, faster decision making around areas like last mile delivery, infrastructure, mobility, and agriculture. Our Community and Partners Industries need to stay ahead of an evolving planet as the climate changes, natural disasters become more prevalent, the rate in which people migrate increases, geopolitical issues become more common, and interconnected systems continue to evolve. These shifts can raise both macro and micro level challenges around everything from where to focus a businesses’ operations and infrastructure to where consumer demand is moving. Wherobots activates the data businesses already have available by making their geospatial context more complete and precise, allowing them to scale and plot an intelligent and adaptable course forward. We’re bringing this to life working with organizations like The Overture Maps Foundation, a coalition of industry leaders including Meta, Microsoft, Amazon, and TomTom, to support its global mapping initiatives, Addresscloud, to help insurers understand geographic risk, and GeoPostcodes, to support analytics for its global postal and population database. Here’s what they have to say: “Overture produces a building dataset covering all buildings in the world, with 2.3B geometries and growing, that’s updated frequently. There’s a lot of data, and compute that goes into producing it and keeping it up to date,” said Jennings Anderson, Geoscientist at Overture and Data Engineer at Meta. “We accelerated the pipelines that produce the buildings dataset by up to 20x after we moved them to Wherobots, which required a simple redirection of our code. We retained compatibility with Apache Sedona, and the move put us into a development experience that’s made us more productive.” “Our high quality data results from aggregating reliable raster population data with our curated boundaries vector database,” said Jerome Urbain, Head of Products at GeoPostcodes. “With Wherobots Spatial SQL, we’re able to analyze population data more efficiently and more accurately, reducing processing time from 39 days to less than one day, and deliver it to our customers across the globe in a far more timely manner. Not only was this a massive speed increase, the overall impact to our data team is they are able to work far more efficiently and productively, answering questions for our customers faster and helping to grow our business.” “Wherobots has given us the potential to run jobs that used to take hours or days to minutes and removed the need to think about provisioning compute. As we provide perils information (flood, fire, etc) to insurers at the property level, we particularly appreciate the ability to be able to run combined vector/raster analysis, without having to previously transform the raster data into vector format or some other format,” said John Powell, Senior Geospatial Data Engineer at Addresscloud. “From a developer perspective, having data, algorithms and compute (and to be presented with a Spark/Sedona context in a Jupyter notebook on startup) combined in one platform is extremely powerful, comparable in many respects to Google Earth Engine, but with much greater guarantees of, and control over, job completion.” AWS Marketplace Integration We’re researchers at heart and we understand that there are so many undiscovered use cases for geospatial data—our customers are the ones helping expand and retool the industry. We’re excited to reach even more teams through our availability on the AWS marketplace, allowing customers to leverage their AWS committed spend and benefit from integrated billing. We’ll be at AWS re:Invent in December (next week!) to learn what else geospatial data can take on. If you are coming to re:Invent, sign up for our GeoParty on the 4th of December, or check out our lightning talk at 12:30 on Thursday. Additionally, the Amazon Last Mile team will be showcasing how they utilize Apache Sedona at the Open Source Developer Theater. For a full overview of everything we have going on at re:Invent, checkout our overview page. What’s Next When we first started the groundwork for Wherobots, we were shocked at the disconnect between the vast amount of planetary data available and cloud data infrastructure support. Through this new round of funding, we hope to provide companies with the tools to really see the world we live in, adapt to new challenges, create intelligence, and potentially save lives. We hope you follow along for our next chapter and if this sounds like something you want to be a part of—we’re always looking for great talent. Want to keep up with the latest developer news from the Wherobots and Apache Sedona community? Sign up for the The Spatial Intelligence Newsletter: Key takeawaysThis is a Series A funding announcement rather than a product deep dive: Wherobots raised $21.5M led by Felicis, with Wing Venture Capital, Clear Ventures, JetBlue Ventures, and P7 Ventures. Aydin Senkut (Felicis) joins the board with Peter Wagner (Wing).Co-founders Mo Sarwat and Jia Yu created Apache Sedona (from GeoSpark at Arizona State University). The post says Sedona has over 40 million downloads and is used for planetary-scale workloads by companies such as Amazon.com (last-mile delivery) and Land O’Lakes (precision agriculture).Wherobots is positioned as a fully managed Spatial Intelligence Cloud compatible with Apache Sedona. Industries cited can analyze real-world issues up to 20x faster at planetary scale.Customer quotes recap existing results: Overture’s 2.3 billion building geometries accelerated up to 20x; GeoPostcodes cut population-boundary processing from 39 days to less than one day; AddressCloud moved property-level flood/fire jobs from hours or days to minutes.The round also funds go-to-market, including AWS Marketplace availability and AWS re:Invent (GeoParty on December 4 and a Thursday 12:30 lightning talk).