Learning Wherobots by Building a National AI Data Center Suitability Report Posted on August 27, 2026October 3, 2026 by George Chandeep Corea This guest post is from a Wherobots user George Chandeep Corea, which covers an exploration of how he used Wherobots MCP with his preferred AI coding tools to build the interactive AI data center site suitability analysis tool below. The following quote is from his own biography. Learning by Doing Below is a working national suitability report for AI data centers. Move the weight sliders and the ranking reorders. Toggle the tailings-dam scenario and the buildable area changes. Everything in it was computed on Wherobots over 15.91 million authoritative geometries drawn from 16 national and state portals, then published as a zero-dependency interactive document that opens in any browser with no compute environment behind it. Have a look first. The rest of this post explains how it was built and how I shaped it, as well as where I plan to go next. Open the full report in a new tab → I learn by doing. I built this on a private basis to contribute to the community discourse. For me it is all about the data; any opinions, perceived or intentional, are my own and have nothing to do with my official job roles. All data is from public sources. It covers 17 candidate industrial sites across all 8 Australian states and territories, spanning the National Electricity Market and the SWIS grid. Why I built this AI data center suitability model For me, this is a community and conservation-minded project: not about saying no to development, but about helping identify where development can create the most value with the least harm. The point is to make siting decisions more transparent, more constructive, and easier to discuss across stakeholders. AI data centers are resource-intensive. They depend on power, water, land, and supporting infrastructure, and they can also create real pressures on ecosystems and communities. I wanted to build something that makes those tradeoffs visible instead of hiding them behind a single score or reports that sit on shelves and don’t let users see it through there lenses. Why I used Wherobots for spatial suitability at scale I wanted a platform that could handle real spatial scale without forcing me into a slow, fragile workflow. I had been hearing about it on LinkedIn and Youtube by following Matt Forrest and wanted to learn by doing. This project queried 15.91 million authoritative geometries across 16 national and state/territory spatial datasets. At the regional modeling scale, the underlying NSW stack ingested 1,751,315 geometries across multiple cloud spatial tables, evaluating 4.92 million spatial join combinations. Using cloud-native GeoParquet inside Wherobots Havasu (Spatially Aware Apache Iceberg) tables, the storage footprint compressed from ~2.9 GB raw equivalent to ~430.7 MB — an 85.2% reduction. That scale matters because suitability analysis only becomes useful when it can be rerun quickly as assumptions change. Verified runtime breakdown across 15.9M geometrics Four different operations, four different numbers. They are not alternatives to each other: 2.4 s for the spatial SQL join execution (distributed joins + net developable area overlays across 1.75M+ geometries); 18.4 s → 3.2 s for the national scan across 15.91M geometries using Hilbert space-filling curve partitioning; 200.6 s for the cold end-to-end batch ETL (uncached ingest, GDA2020 reprojection, ST_MakeValid topology repair, Iceberg writes); and < 1 ms for the client-side What-If recalculation in the browser. What I was trying to learn I wanted to understand how far I could push a cloud spatial workflow while still producing something that a non-technical audience could review. The answer was to combine Wherobots for the heavy computation with a document-style HTML report for presentation. The report is the public-facing layer of the analysis. It includes an interactive map, a ranked candidate leaderboard, benchmarking tables, provenance, methodology, and a what-if sandbox for changing ranking weights. My recent work covers GDA2020 cadastral modernization at an enterprise level, environmental monitoring at a state level, and disaster modeling and high level coordination with FEMA. As the creator of *AuraSiting Crafter* (currently hunter_spatial crafter), I built an open-source multi-criteria engine benchmarking clean energy, water security, acoustic buffers, and developable land across 15.91 million Australian geometries using Wherobots MCP as the critical foundation. George Chandeep Corea Builder, Learner, Public Sector Why the interactive suitability report matters The HTML report is where the analysis becomes useful for public review. It is not just a visualization — it is a way to interrogate the assumptions behind the ranking. The map is interactive, searchable, and configurable. Readers can change basemaps, add WMS/WFS services, and bring in external layers to compare candidate sites against their own data or public reference layers (some functionality coming soon!). That matters because spatial decisions are rarely made from one dataset alone. Adding data will make the report more useful for collaboration, because different stakeholders can overlay the information that matters to them: local context, community knowledge, infrastructure plans, environmental constraints, or jurisdictional boundaries. The sandbox (already available) is especially valuable because readers can adjust the ranking weights and immediately see how the candidate ordering changes. That makes the tradeoffs tangible instead of theoretical. Suitability report overview The report overview combines summary KPI cards, an interactive What-If simulation sandbox, an interactive siting map with layer controls, and a ranked candidate leaderboard into a single reviewable spatial document. How the suitability workflow works The pipeline is straightforward: Ingest authoritative layers — harvest from live government WFS/REST portals and national registers into cloud object storage Apply exclusions and setbacks — deduct 30 m riparian corridors, 20 m high-pressure gas/water pipeline easements, high-value biodiversity zones, slope >5%, and tailings dam hazards for example. Calculate network distances — topological winding distances (1.32× factor) to ≥132kV substations and wastewater outfalls for example Evaluate sensitive receptors — continuous sigmoidal buffer penalty around schools, hospitals, and residential meshblocks for example Benchmark against regional baselines — compare candidates against simulated baselines across interstate transition hubs (Latrobe Valley, Collie, Gladstone) for example. Only the hunter has detailed data. Sites outside the Development Application (DA) in LMCC has data simulated (not mocked or made up though) through analysing relevant real spatial data. Publish a zero-dependency report — compile layers, audits, and metadata into a standalone interactive HTML document The scoring model uses many weights: Power grid proximity: 50% Recycled water proximity: 30% Parcel size: 20% Others added/will be added as the project develops and Those weights reflect the practical requirements of AI infrastructure while keeping environmental and resource considerations visible. How the suitability score is calculated Suitability = 0.40·S_power + 0.25·S_sensitive + 0.20·S_water + 0.15·S_size Power grid proximity (40%). Proximity to ≥132kV transmission substations. Full points between 100 m and 500 m; linear decay to 0.0 at 5 km; a 0.7 penalty under 100 m for EMF safety and acoustic isolation. Sensitive receptor buffer (25%). A continuous sigmoidal decay around a 500 m compliance threshold: S_sensitive(d) = 1 / (1 + e^(-0.01·(d - 500))). Hard exclusion under 300 m; acoustic-barrier buffer penalty 300–500 m (0.10–0.50); compliant 500 m–1.5 km (0.50–1.0); optimal workforce balance 1.5–5 km (1.0); linear commute decay beyond 5 km. Recycled water proximity (20%). Distance to wastewater treatment plants for sustainable evaporative cooling, decaying linearly from 1.0 at ≤1 km to 0.0 at ≥10 km. Developable parcel size (15%). Contiguous flat buildable land. Parcels ≥15 ha score 1.0; parcels under 3 ha get a 0.1 baseline, with linear interpolation between. Slope calculations exclude land above 5% grade to avoid excessive earthworks. What the analysis is based on National input layers The national model integrates 16 authoritative datasets across all 8 Australian jurisdictions: Geoscape National Cadastre & G-NAF (15,420,800), ABS 2021 Meshblocks & UCL (368,290), National Sensitive Receptors from ACARA / NHSD / OSM (47,510), Geoscience Australia electricity grid (4,820), BoM & GA surface water and outfalls (42,100), and state planning/cadastral portals — NSW SEED, QLD QSpatial, Vicmap, Landgate (27,725). Total: 15,911,245 geometries. Those totals are the available universe published across the portals — 47,510 POIs and 15.4M cadastral parcels are what ACARA, NHSD, OpenStreetMap and Geoscape publish, not what this pipeline processed. The ingested and audited cohort is 1.75M regional features in the Hunter deep-dive, 368,290 ABS meshblocks partitioned nationally, 17 indexed candidate sites, and a ground-truth QA sample of 33 sensitive receptors (19 schools via ACARA, 14 hospitals via NHSD) sitting inside the candidate industrial zones across all 8 states, used to calibrate the sigmoidal buffer decay curve. Of the 17 candidates, 4 are micro-sited in detail in the NSW Hunter; the remaining 13 are simulated baselines across interstate transition hubs (Latrobe Valley VIC, Collie WA, Gladstone QLD). The earlier NSW regional stack: still the measured layer DatasetFeature countNotesNSW Transport Network (Rail)275,421Transport contextNSW Biodiversity Constraint Zones262,258Environmental exclusionNSW Energy Grid Infrastructure241,573Power proximityABS Census Meshblocks223,238Demographic contextNSW Pipeline Corridors197,247Setback constraintsTfNSW Active Transport Pathways188,576Access and corridor contextNSW Hydrography & Waterways181,501Water and riparian contextABS Regional Demographics1,160Regional benchmarking Storage and compression on Havasu Iceberg All spatial tables are cataloged under org_catalog.fgsdb.* on Wherobots Cloud and persisted in cloud object storage, in GDA2020 / MGA Zone 56 (EPSG:7856). TableGeometriesUncompressed rawGeoParquet footprintSavingsmacquarie_biodiversity_constraints262,258~580.0 MB84.2 MB85.5%macquarie_energy_infrastructure241,573~420.0 MB62.5 MB85.1%macquarie_transport_rail275,421~390.0 MB58.1 MB85.1%macquarie_pipeline_corridors197,247~280.0 MB41.8 MB85.1%macquarie_abs_meshblocks223,238~650.0 MB98.4 MB84.9%macquarie_water_hydrography181,501~310.0 MB44.6 MB85.6%macquarie_active_transport188,576~260.0 MB39.2 MB84.9%abs_demographics1,160~12.0 MB1.8 MB85.0%Total (8 tables)1,751,315~2.9 GB~430.7 MB85.2% To avoid real-time WFS/FeatureServer REST API timeouts during Spark runs, every dataset is ingested and hosted as an optimized cloud spatial Iceberg table rather than fetched live. The spatial SQL underneath Building the net developable area mask — unioning riparian, pipeline, and rail buffers and subtracting them from the sub-precinct boundaries: SELECT p.precinct_key, ST_Difference(p.geom, ST_Union_Aggr(c.geom)) AS net_developable_geom FROM precinct_transform p LEFT JOIN constraints c ON ST_Intersects(p.geom, c.geom) GROUP BY p.precinct_key, p.geom Computing nearest distances to transmission substations and wastewater treatment outfalls: SELECT mb.mb_code21, MIN(ST_Distance(mb.mb_geom, ST_Transform(p.geometry, 'EPSG:4326', 'EPSG:7856'))) / 1000.0 AS dist_to_substation_km, MIN(ST_Distance(mb.mb_geom, ST_Transform(w.geometry, 'EPSG:4326', 'EPSG:7856'))) / 1000.0 AS dist_to_wwtw_km FROM industrial_meshblocks mb CROSS JOIN org_catalog.fgsdb.macquarie_energy_infrastructure p CROSS JOIN org_catalog.fgsdb.macquarie_water_hydrography w GROUP BY mb.mb_code21 How I think about the planning question My background in conservation shaped this project. I’m not trying to use technology to say “no” to development. I’m trying to use it to ask better questions about where development belongs, where it creates value, and how to reduce unnecessary conflict between infrastructure, ecosystems, and communities. That is why the report includes explicit constraint layers, benchmarking, and scenario controls. The aim is to support better judgment, not replace it. Benchmarking and query speed mechanics Benchmarking and speed mechanics breakdown detailing how spatial SQL queries execute across 1.75M+ geometries via metadata envelope pruning, Hilbert curve clustering, vectorized memory execution, and parallel distributed spatial joins. Four things account for the query speed: Metadata envelope pruning. Havasu Iceberg stores 2D bounding box envelopes directly inside Iceberg AVRO manifest files. Queries with spatial predicates such as ST_Intersects prune the large majority of irrelevant Parquet files at the metadata layer, before any raw disk bytes are scanned. Hilbert curve spatial clustering. Geometries are sorted with 2D Hilbert space-filling curves during ingestion, so geographically adjacent features land in the same Parquet row groups and storage partitions. That removes random disk I/O seek overhead. Vectorized memory execution. Apache Sedona operates directly on columnar GeoParquet WKB geometry buffers, avoiding serialization costs between Python, the Spark JVM, and native spatial drivers. Parallel distributed spatial joins. Quad-tree and R-tree spatial indexes partition the query space across worker nodes, turning expensive O(N × M) cross-joins into O(N log M) parallel bucket joins. Candidates are benchmarked against simulated local and regional baselines — Latrobe Valley in VIC, Collie in WA, Gladstone in QLD — to position NSW development opportunities within the wider national energy market transition. How the what-if sandbox runs with no server The simulation sandbox runs entirely in the browser: no server calls, no network latency, no cloud API charges during a slider session. The heavy spatial work — topological winding distances, buffer overlaps, elevation head drops, thermodynamic decay rates — is computed at build time on Wherobots and embedded in the report’s JSON payload. Moving a weight slider then normalizes the raw values so the weights sum to 1.0 and recalculates suitability across every candidate record in JavaScript, in under a millisecond. Toggling the tailings dam safety switch swaps pad areas between the declared and de-declared cases (+15.2 ha unlocked) and updates map polygons and audit cards without refetching any GeoJSON. Candidates re-sort by active score, updating the leaderboard, marker radii, and popup badges. The payoff is threefold: no per-query cloud compute cost during interactive sessions, full interactivity when the HTML report is emailed or opened offline, and slider manipulation that stays smooth without waiting on network round-trips. What’s next for the suitability model The following describes the author’s planned work, not shipped functionality. Sensitive receptor scoring. The next version will score candidates against sensitive community receptors — schools and early childhood centers, hospitals and aged care, residential meshblocks, and workforce commute bands — to prevent acoustic, thermal, electromagnetic, and visual conflicts. The model uses a continuous sigmoidal penalty around a 500m critical setback threshold, with a workforce accessibility modifier that’s neutral in the 1.5km–5.0km commute band. In practice: under 300m is a critical exclusion, 300–500m carries a buffer penalty requiring an acoustic barrier, 500m–1.5km is compliant, 1.5–5.0km is the optimal community-and-workforce balance, and beyond 5km the score falls off for commute burden. Grid and water policy analysis. Following the National Cabinet debate on AI data center power demand and regional grid security, the framework will evaluate proximity to 132kV, 330kV, and 500kV bulk transmission, flag candidates adjacent to retiring coal-fired stations that can reuse existing heavy transmission without expensive grid upgrades, map candidates against Renewable Energy Zones and firming assets for 24/7 clean energy matching, and restrict cooling supply strictly to recycled and wastewater sources — scoring zero for any site dependent on potable drinking water reserves or vulnerable aquifers. Candidates then sort into three tiers: grid-ready and fast-track eligible, conditional pending firming storage, or constrained by congestion and potable water reliance. Open platform integration. To let the public explore these models interactively, hunter_spatial_crafter will integrate with opengeos/GeoLibre. The design point worth noting is zero duplication: Wherobots writes suitability layers to a central cloud bucket as GeoParquet and PMTiles, and GeoLibre’s in-browser DuckDB-WASM engine reads those exact same files over HTTP range requests — fetching only the row groups it needs, with no dataset conversion, no server-side copy, and no large client download. Engineering efficiencies: reducing compute spend Total cloud batch compute spend across dozens of iterative development, benchmark, and QA runs was ~$36 AUD (US$24.13) — roughly $1.03 per run across ~35 automated batch runs. Analyzing the execution profile shows how a production pipeline could cut that further: Decouple heavy geometry joins from lightweight scoring. The pipeline splits into a compute-intensive geometric tier (ingest, GDA2020 reprojection, ST_MakeValid repair, 30 m/20 m buffers, ST_Difference overlays across millions of polygons) and a compute-light scoring tier (decay curves and weighted composites over precomputed distance attributes). Structured as a DAG with intermediate materialised GeoParquet stages, tuning a weight or the sigmoidal threshold never re-runs the heavy joins — only the downstream matrix recalculates. Fingerprint sources and memoize snapshots. Baseline layers change infrequently. Content hashing (ETags, GeoParquet file hashes, Iceberg snapshot manifest IDs) lets untouched tables be skipped, reading straight from cached Havasu Iceberg partitions. Process delta partitions. When state portals publish quarterly cadastral updates, Iceberg’s ACID snapshot metadata lets WherobotsDB (optimized and managed Apache Sedona) isolate only modified parcel geometries instead of full continental scans. Offload interactive compute to the client. Compiling precomputed distance topologies into the standalone report puts millions of interactive public scenario evaluations at $0.00 cloud compute cost. Applied together, these reduce continuous CI/CD pipeline cost from ~$36 AUD to under $5 AUD. Why I think this is useful This project taught me that Wherobots is more than a faster way to run spatial SQL. It is a way to make large-scale spatial analysis practical for real decision support. It gave me a path from raw geodata to a polished, reviewable spatial document — one that can be opened by a stakeholder, examined by a technical reviewer, and discussed in public. For a personal project, that was the goal: learn the platform, test it at real scale, create something useful, and contribute to the broader discussion about responsible AI infrastructure. If you’re interested in spatial analysis, cloud geospatial workflows, or responsible AI infrastructure, feel free to connect with me on LinkedIn. Key takeawaysWherobots computed a national suitability model over 15.91 million authoritative geometries drawn from 16 national and state portals, published as a zero-dependency interactive document that opens in any browser with no compute environment behind it.Cloud-native GeoParquet inside Wherobots Havasu (Spatially Aware Apache Iceberg) tables compressed the storage footprint from ~2.9 GB raw equivalent to ~430.7 MB, an 85.2% reduction.Runtime at national scale: 2.4 s for the spatial SQL join execution, 18.4 s down to 3.2 s for the national scan across 15.91M geometries using Hilbert space-filling curve partitioning, 200.6 s for the cold end-to-end batch ETL, and under 1 ms for the client-side What-If recalculation in the browser.Total cloud batch compute spend across dozens of iterative development, benchmark, and QA runs was ~$36 AUD (US$24.13), roughly $1.03 per run across ~35 automated batch runs.The model covers 17 candidate industrial sites across all 8 Australian states and territories, scored on the formula Suitability = 0.40·S_power + 0.25·S_sensitive + 0.20·S_water + 0.15·S_size. Get Started with Wherobots Try Now
How Bad Telemetry Data Sabotages Modern Fleets Posted on July 8, 2026August 31, 2026 by Ben Pruden By the Teams at Action Engine & Wherobots Fleet monitoring is undergoing a generational shift. Fleet monitoring, the systems that ingest and analyze vehicle telemetry to track fleet health, performance, and safety, has become the foundation of how operators run vehicles, not just track them. Modern vehicles generate orders of magnitude more telemetry than even five years ago – GPS, fuel and battery signals, sensor and event streams, on-board diagnostics; that data is no longer just a record of where the fleet has been. It’s the data layer operators use to make fleets more cost-effective, more sustainable, and increasingly more autonomous. This trajectory points in one direction: vehicles are now navigating by data. Autonomous and semi-autonomous fleets sit at the intersection of GIS, computer vision, and vision-language models – systems where the quality of the underlying data directly determines whether the vehicle stops at the right time, takes the right route, or correctly perceives the world around it. Generic fleet management platforms were built for an earlier era. They tell you where your fleet is. They struggle with the harder question of where your fleet is going wrong. Fleet management platforms apply global thresholds, treat each ping in isolation, and miss anomalies hidden in the geographic and temporal context around a record.Equipment doesn’t fail without warning. Vehicles don’t break down without prior warning signs. Fleets aren’t underutilized by accident. But the anomalies that precede these outcomes are almost always invisible inside generic fleet dashboards.They hide inside telemetry that looks fine in a table – GPS, fuel consumption, temperatures, pressures, sensor readings – until they aggregate into an incident. To close this gap, Action Engine built Aspen Fleet: an anomaly detection system engineered specifically for the future of fleet operations. Powered under the hood by Wherobots, the industry standard for distributed spatial compute, Aspen Fleet catches the patterns that generic platforms miss. By running against both historical and near-real-time telemetry, it uses spatial context as a first-class signal to surface anomalies earlier than previously possible. Aspen finds fuel-consumption outliers across thousands of transit vehicles by checking each unit against its own baseline, and maps the deviations onto the specific Portland segments where they occurred. What Generic Fleet Platforms Miss Traditional fleet management platforms work well for basic visibility – locations, statuses, and last-known positions. However, they struggle as true anomaly detectors for three specific reasons. They apply global thresholds A fuel consumption value that’s anomalous on a flat highway is completely normal on a mountain pass. An idle event in a depot is operational; the same event at an intersection in a residential block is suspicious. When a platform alerts on a single threshold for the entire fleet, it either generates false positives that operators learn to ignore or it misses real anomalies entirely. They treat each telemetry ping in isolation A single GPS coordinate that places a delivery truck in the middle of a lake looks like one slightly weird record. A single fuel reading that drops three percent looks like noise. But each of these is part of a sequence – what happened before, what’s happening around it geographically, and what the same vehicle was doing on the same route last week. Without that spatial and temporal context, the signal that matters may never surface. They don’t see the new data layer at all Connected and autonomous fleets generate streams that legacy platforms were never designed to validate – computer vision detection logs, model outputs, and perception confidence scores. These streams need their own quality logic: how many objects of a given class the computer vision (CV) stack detected on a route segment today versus yesterday, whether two cameras are producing duplicate detections of the same physical object, or whether model v2.1 has regressed against model v2.0 on a specific stretch of road. Generic platforms don’t ask these questions because they aren’t built to. This is a gap that Action Engine and Wherobots came together to close. Segment-level CV regression caught: model v2.1 reports 39 objects against a baseline of 25, driven by duplicate utility-pole detections. What Aspen Fleet Detects Aspen Fleet ships with a library of over 100 pre-built detection rules tuned specifically for fleet telemetry. Some catch errors in the data itself, while others catch operational anomalies that traditional platforms miss because they lack spatial reasoning. The rules group into a handful of core categories: Location validity: GPS teleportation (a vehicle moving 500 km in two minutes), impossible coordinates (a car “on water” or in a region it was never dispatched to), distance-to-road violations, and GPS drift (a vehicle 80 meters off any drivable surface). Signal integrity: Coordinate freezes, signal loss in tunnels and dead zones, out-of-order timestamps after backfilled resumptions, and hung trackers reporting identical points sequentially. Operational behavior: Idle outliers in unexpected geographies, dwell-time anomalies on familiar routes, and route deviations that only register when historical patterns are known. Sensor envelopes: Fuel consumption, coolant and oil temperature, tire pressure, EV battery state of charge, and signal strength – each validated against expected envelopes per vehicle, per route segment, and per device class. Computer vision and perception outputs: Detection counts per route segment, duplicate detections across cameras, and model performance comparisons between deployed software versions on the same physical road segment. The unifying capability is that none of these checks rely on uniform thresholds. They reason about the geography, the route segment, the device history, and the temporal pattern simultaneously. Uniform thresholds are where generic platforms produce false positives and noisy dashboards, while Aspen prioritizes accuracy. Furthermore, when a fleet has its own operational logic that the pre-built library doesn’t cover, customers can easily author custom checks using the same engine. Two Modes: Historical and Near-Real-Time Fleet Observability Aspen Fleet runs in two complementary modes Historical Anomaly Detection Historical anomaly detection can process years of stored fleet telemetry at scale to find systematic anomalies, retroactively label bad records, and produce a clean baseline that analytics and downstream models can stand on. This is where Aspen is unusually strong. Fleet operators sit on years of telemetry that nobody fully trusts the time or compute to validate it at scale was previoulsy unreachable. Analytics built on that data inherit its noise: KPIs drift, utilization metrics misrepresent reality, maintenance forecasts are anchored on contaminated baselines, and any machine learning model trained on the data inherits whatever errors were hidden within it. Aspen processes historical datasets at scale with Wherobots distributed spaital compute to find systematic anomalies, retroactively label bad records, and produce a clean baseline that analytics and downstream models can actually stand on. The output isn’t just a list of errors – it’s a measurably more accurate version of the fleet’s own history. Near-Real-Time Anomaly Detection Near real-time anomaly detection runs continuously on streaming telemetry, surfacing anomalies within minutes so fleet operations can act before slow-developing patterns become incidents. This mode runs continuously as telemetry streams in. Aspen sees ping sequences in near-real time, applies the same spatial-context-aware checks to streaming data, and surfaces anomalies within minutes of the event that caused them. Fleet operations teams can act on these alerts before a slow-developing pattern – a fuel system slowly degrading, a sensor drifting out of calibration, a route consistently underperforming, or a CV model regressing on a specific stretch of road – turns into a vehicle off the road or a bad decision in production. The combination matters. Historical analysis tells the team what they’ve been missing, while real-time analysis makes sure they stop missing it going forward. Surfacing issues earlier, helping teams understand what actually requires action The practical consequence of spatial-context-aware detection is that Aspen catches anomalies that classical fleet tools surface days or weeks later – if at all. For example: Fuel System Degradation: A fuel system degrading over weeks shows up as a small, slowly widening gap between expected and observed consumption on specific route segments. A global threshold misses it. Aspen sees the drift early because its baseline is segment-specific. Brake-System Issues: This shows up as a subtle change in deceleration patterns at specific intersections the vehicle drives repeatedly. A generic “harsh braking” alert miscalibrated for the fleet either fires constantly or never fires at all. Aspen flags the change because it knows what normal looks like at this specific intersection for this specific vehicle. Perception Stack Decay: A computer vision model degrading shows up as a gradual drop in detection counts or confidence on familiar route segments, often caused by sensor obstruction, a degrading camera, or environmental shifts. Generic fleet platforms don’t see this at all because they don’t process perception outputs. Aspen flags it by comparing detections segment-by-segment and day-over-day. Underutilization: An underutilized vehicle shows up as a pattern of idle outliers in non-operational geofences combined with reduced route variation. Most fleet platforms surface neither signal cleanly. Aspen combines them into a single readable indicator. This is what “surfacing issues earlier” actually means in practice: anomalies that mature into equipment failure, vehicle breakdown, inefficient utilization, or a degraded perception stack get caught while they’re still “drifts” – not after they’ve become costly incidents. How Aspen Fleet Runs at Scale The Infrastructure Behind the Insight To analyze complex spatial and temporal data at scale, you need a new kind of architecture. Aspen Fleet is cloud-native, utilizing the high-performance distributed spatial compute platform provided by Wherobots. Wherobots provides the underlying engine that makes spatial data a first-class citizen. Because of this powerful foundation, historical scans across billions of telemetry records are complete in minutes, while streaming checks seamlessly keep pace with the highest-volume fleets in production today. Data integration is built to be frictionless. Customers can connect their data through whatever channel best suits their architecture – whether that means leveraging existing Geotab and telematics platforms, custom REST endpoints, batch dumps to cloud object storage, or live gRPC feeds directly from on-board vehicle systems. Aspen Fleet and Wherobots handle the ingestion, normalization, and processing automatically. The Takeaway: Fleet Observability Needs Spatial Context The fleets that will win the next decade – the ones that are measurably more cost-effective, more sustainable, and increasingly autonomous – are the ones that can read and act on their own data accurately. Most fleet management platforms answer the basic question: “Where is the fleet?” Aspen Fleet answers a much harder, more valuable question: “What is going wrong, and how early can we catch it?” The combination of deep spatial context, dual historical and real-time processing modes, a validation library tuned for modern perception outputs, and an engine built for massive scale changes the paradigm of fleet operations. Don’t run a fleet you only hope is healthy, run one you can prove is. Live demo with the teams behind Aspen Fleet and Wherobots Join us on Tuesday, August 18 at 9AM PT / 12PM ET. See Fleet Observability in Action Save Your Seat Key takeawaysGeneric fleet platforms apply global thresholds, treat each ping in isolation, and miss anomalies hidden in geographic and temporal context. Action Engine built Aspen Fleet, an anomaly detection system powered by Wherobots distributed spatial compute, to catch those patterns against historical and near-real-time telemetry.Aspen ships a library of over 100 pre-built detection rules covering location validity (GPS teleportation such as 500 km in two minutes, coordinates on water, 80-meter GPS drift off a drivable surface), signal integrity, operational behavior, sensor envelopes, and computer vision / perception outputs.None of the checks rely on uniform thresholds. They reason about geography, route segment, device history, and temporal pattern at once. Customers can also author custom checks on the same engine when the pre-built library does not cover their operational logic.Aspen runs in two modes: historical scans that process years of stored telemetry at scale to label bad records and produce a clean baseline, and near-real-time checks that surface anomalies within minutes of the event. Integration can use Geotab and other telematics platforms, custom REST endpoints, batch dumps to object storage, or live gRPC feeds from on-board systems.
Introducing RasterFlow: a planetary scale inference engine for Earth Intelligence Posted on December 10, 2025September 1, 2026 by Philip Darringer We’re very excited to announce RasterFlow is now available to select customers in a private preview. If you are interested in learning more or would like to request access to the preview, contact us here! RasterFlow is a serverless image preparation and inference engine that makes it significantly easier to generate Earth Intelligence from planetary scale Earth Observation (EO) datasets. With it, customers and their AI agents will be significantly more capable of innovating with EO data and integrating earth insights into their data infrastructure. Upcoming Session: See an AI agent take plain-language question and orchestrating pipelines to end results using RasterFlow and the Wherobots MCP server. See it in action Join us live How RasterFlow Powers Earth Intelligence at Scale A few weeks ago, we announced our collaboration with the Taylor Geospatial Engine to help them evaluate their Fields of the World (FTW) machine learning model that segments agricultural field boundaries. Using an early release of RasterFlow, we were able to quickly and cost-effectively run this model at scale. Here’s a breakdown of how this works in practice. RasterFlow ingests and assembles the source imagery – in this case Sentinel-2 – into an inference-ready mosaic, generating representative features using the FTW model for planting and harvest seasons, and removing cloud cover as needed (1, 2). The FTW model is run against this mosaic using RasterFlow’s distributed inference engine to predict fields and field boundaries (3). RasterFlow predictions are then vectorized into geometries and made available as an Iceberg table (4) that can be used in WherobotsDB or other downstream applications and data systems for field-level crop insights. RasterFlow’s applicability is much wider than Sentinel 2 and FTW. It supports Zarr and COG imagery datasets and PyTorch computer vision models for inference. RasterFlow at Scale The images above represent sample outputs for a small area in Kansas, but RasterFlow can be very attractive for larger scale runs. In our collaboration with the Taylor Geospatial Engine, we executed larger scale runs including the Continental United States (CONUS), Japan, Mexico, South Africa, Switzerland and Rwanda. RasterFlow’s efficient parallel processing enabled each of these large scale workflows to complete in minutes to a few hours. RasterFlow autoscales compute resources based on expected compute and inference load, which is a function of area and time range, dataset density, and model complexity. Challenges using EO Data Most data teams do not have the expertise or the budget to build and operate the unique infrastructure and software stack required to extract insights from EO datasets using computer vision models. These barriers have prevented innovative ideas from getting off the ground. According to Gartner, only 1% of AI models today leverage physical world data, vs a projected 80% by 2029. Similarly, AI agents are projected to generate 10 times more data from physical environments than from all digital AI applications combined.1 However, AI agents can’t economically make sense of this raw data because it has to be prepared by the same costly, complex, and unique infrastructure the data teams need, but neither have access to. Here’s an example that underscores these challenges: if you or an AI agent are trying to analyze wildfire state and predicted spread to measure risk to infrastructure, developers typically need to build dedicated pipelines that: Ingest and prepare imagery for inference, minimizing noise such as cloud cover and edge effects Deploy a machine learning model on prepared imagery, trained to segment and classify fires Tune model inference for scale and efficiency, while minimizing edge and tiling effects from individual tasks Measure change over time using models that take into account wind direction, speed, vegetation, buildings and other infrastructure in the probable path of the fire Join model predictions with other important context including building footprints, land parcels and infrastructure such as powerlines and pipelines to calculate overall risk Forecast the spread of the fire In total, these steps require significant investments in both infrastructure development, operations, and talent that most businesses are unable to justify, much even accomplish. On-Demand Imagery Preparation and Inference for Earth Observation Workflows The inspiration for RasterFlow was to make it easy for any company to use large scale sensor datasets and computer vision models to unblock innovation and AI applications for the physical world. RasterFlow does this by combining decades of expertise with a fully managed, inference and mosaicking workflow and API designed for Earth Intelligence at any scale. Here are a few key capabilities: On-demand serverless operations for imagery ingestion, preparation (also known as mosaicking), and inference. Built-in support for popular open datasets and open models so you can get started quickly. Inference results that can be converted to vector geometries and integrated into a lakehouse architecture; in a customer’s cloud storage bucket as Parquet files in Apache Iceberg tables. Ability to easily postprocess these results with WherobotsDB or other lakehouse engines with support for spatial operations, such as Databricks, Snowflake, or Google BigQuery. Simple enough for any engineer, scientist, or analyst to use: just pick a model, an area of interest to deploy that model, and a time range. Advanced users can take advantage of lower-level APIs to customize their planetary-scale inference runs. RasterFlow Operators: Core Functions for Preparing Imagery, Model Inference, and Vectorization RasterFlow provides fully managed operations required for processing Earth Observation datasets, including: Imagery ingestion and preparation to remove cloud cover, edge effects, and build a high quality inference-ready mosaic Distributed inference for large scale computer vision, geospatial foundational and other PyTorch model runs Vectorization of model outputs into geometries or as analytics ready rasters For object detection workloads, RasterFlow pairs with models like Segment Anything 3, see how SAM 3 performs on aerial and satellite imagery for a full walkthrough. Ingesting and Preparing Satellite Imagery for Model Inference Satellites and drones capture imagery on a particular flight path. And it may take multiple drone flights, or days, weeks, or even months for the flight paths of a satellite constellation to capture clean imagery for a particular area of interest. Clouds and weather events may still block what you may be interested in. In these circumstances it’s important to understand the rate of coverage and define your time horizon accordingly, to build a mosaic. A mosaic is a composite image that is the result of composing high-quality pixels (e.g., cloud free) over a time range, and stitching them together for a particular area. Base satellite layers in your favorite map applications (Google Maps, Mapbox) are cloud-free mosaics composed from images over a wide time range. Many computer vision models are trained to find relatively durable things on Earth, like buildings, roads, and land cover. But when clouds, coverage, imagery edge effects, or other types of “noise” exist in the input imagery, the quality of inference suffers. The purpose of the mosaic is to correct for this noise and make imagery, inference-ready, so model inference produces the results you want. RasterFlow takes care of this heavy lifting for you, creating an inference-ready mosaic that maximizes the usefulness of today’s Earth Observation models. Distributed Geospatial Inference at Planetary Scale We’ve moved past the use of eyes to analyze imagery, and are now capable of letting machines do this work for us. With RasterFlow, today’s machine learning models can perform tasks such as object detection, segmentation, and classification, on a very large area of interest, with orders of magnitude more efficiency and scale than an analyst’s eyes can offer. The RasterFlow inference engine is designed for small to very large scale runs. It efficiently parallelizes across the input mosaic across a distributed and serverless inference architecture while minimizing tiling effects typically produced when inference pipelines operate on individual tiles. Running Hosted or Custom Geospatial AI Models with RasterFlow For convenience, RasterFlow currently hosts popular open source PyTorch geospatial computer vision models that are ready to use. These models currently include: Fields of the World (FTW) Field Boundary Delineation Meta and World Resource Institute Tree Canopy Height Prediction ChesapeakeRSC Road Segmentation Tile2Net Pathway Segmentation You can also import your own custom PyTorch model to your Wherobots Organization for private deployment. RasterFlow + TorchGeo: Simplifying PyTorch-Based Geospatial AI Wherobots actively supports the TorchGeo project which helps machine learning experts to more easily work with geospatial data within the PyTorch ecosystem. We will continue to build out RasterFlow integrations with TorchGeo, including onboarding additional TorchGeo models and further simplifying the model lifecycle for PyTorch models. While we are starting with support for PyTorch focusing on TorchGeo models, we are open to adding support for other model frameworks. Calling Geospatial Model Developers: Contribute to RasterFlow We are continually adding new, open source geospatial computer vision models to the Wherobots Model Hub. And if you’re a model developer, we’re interested in speaking with you to onboard your model and distribute the value of your work to a wider audience using Wherobots RasterFlow. Vectorizing Model Outputs: From Raster Predictions to Geospatial Geometries Many computer vision models output rasters, where each pixel in the raster represents a predicted real-world value such as height of the tree canopy, or the confidence that the pixel represents a certain feature such as an agricultural field boundary or a sidewalk. RasterFlow provides built-in support for raster vectorization, turning pixel values into rich, concise geometries. These geometries represent features of interest that can be post-processed, conflated, and integrated into your workflows because they are yours, stored in open source file (Parquet) and table (Iceberg) formats in your S3 bucket. Using RasterFlow with Geospatial Foundation Models and Embeddings Recent developments in Geospatial Foundation Models have generated tremendous interest in the research community, potentially accelerating Earth Observation applications the same way that Large Language Models (LLMs) and embeddings have transformed AI’s ability to generate language. RasterFlow can generate embeddings from the latest open Geospatial Foundation Models, including OlmoEarth from the Allen Institute for AI (Ai2) and Clay. With RasterFlow’s ability to cost-effectively generate embeddings at scale, researchers and practitioners can easily generate embeddings for their area of interest and evaluate their suitability and power. Customers and Partners Using RasterFlow for Scalable Earth Intelligence One highlight while developing RasterFlow has been our collaboration with customers and partners like SatSure, Taylor Geospatial Engine, and Spyrosoft. We’ve used feedback from these teams to ensure we are solving for customer needs. Before the Thanksgiving holiday we shared our recent learnings from working together with Taylor Geospatial Engine, who have been incredibly helpful in providing input on the types of ways their ecosystem of developers and ML engineers would want to interact with RasterFlow. SatSure is an existing Wherobots customer and an early adopter of RasterFlow, and we are excited to see what they build next with it. "RasterFlow meaningfully accelerates the work SatSure and Wherobots already do together. By automating mosaicking, preprocessing, and distributed inference into a single, on-demand workflow, it removes much of the engineering overhead required to operationalize our models at national and multi-season scale. This helps us move new geospatial AI models into production faster, iterate more quickly with customers, and deliver fresher, high-resolution insights across agriculture, banking and financial services, and infrastructure use cases." Rashmit Singh CTO and co-founder, SatSure One of the largest deployments to date processed 348 TB of satellite imagery for a global release, see how RasterFlow delivered Fields of the World at planetary scale RasterFlow Availability and Multi-Cloud Architecture Wherobots infrastructure runs natively on AWS and customers pay for use through the AWS marketplace. RasterFlow and WherobotsDB support hybrid architectures, where data is read from, and results are written to other environments such as GCP, Azure, Oracle, or on-premises. This is particularly useful when processing open datasets or using open models and the environment in which data is processed may not be a concern. On-demand pricing for RasterFlow will be announced at a later date, but can be discussed with customers participating in the private preview. Next Steps: Try RasterFlow and Explore the Wherobots Spatial Data Platform We invite anyone who wants to test out RasterFlow to request to join the private preview here. Get started building with the most capable and efficient spatial data platform using the Wherobots Professional Edition. Sign up for the newsletter to keep pace with what’s happening at Wherobots. Source – 27 August 2025, Gartner Innovation Insight: World Models Are Set to Empower AI Agents With Imagination ↩︎ Join the Private Preview Sign Up Key takeawaysRasterFlow is a serverless imagery-preparation and inference engine, now in private preview, that turns planetary-scale Earth Observation datasets into Earth Intelligence without customers having to build their own mosaicking and model-serving stack.Working with the Taylor Geospatial Engine, RasterFlow ran Fields of the World inference across CONUS, Japan, Mexico, South Africa, Switzerland, and Rwanda. Those large-scale workflows completed in minutes to a few hours as compute autoscaled with area, time range, dataset density, and model complexity.It ingests Zarr and Cloud-Optimized GeoTIFF imagery, hosts PyTorch models including FTW field boundaries, Meta/WRI tree canopy height, ChesapeakeRSC road segmentation, and Tile2Net pathway segmentation, and lets organizations import their own private PyTorch models.Predictions are vectorized into geometries and written as Parquet files in Apache Iceberg tables in the customer S3 bucket, then post-processed in WherobotsDB or other lakehouse engines such as Databricks, Snowflake, or BigQuery.Gartner projects that only 1% of AI models leverage physical-world data today versus 80% by 2029, while AI agents are projected to generate 10 times more data from physical environments than from all digital AI applications combined. One of the largest RasterFlow deployments processed 348 TB of satellite imagery for a global release.
How Aarden.ai Scaled Spatial Intelligence 300× Faster for Land Investments with Wherobots Posted on November 17, 2025October 3, 2026 by Ben Pruden When Aarden.ai emerged from stealth recently with $4M in funding to “empower landowners in data center and renewable energy deals,” the company joined a new wave of data and AI startups reimagining how physical-world data drives modern business. Their mission: help institutional land investors rapidly evaluate the value and potential uses of land across the country. To do that, Aarden needed to process vast geospatial datasets such as parcels, forests, soil, endangered species, energy infrastructure, and turn them into actionable business intelligence. The problem? Their early Python stack couldn’t keep up. Aarden.ai needed faster, scalable geospatial processing and Wherobots made it possible. Results First: From Seven Days to Thirty Minutes Before adopting Wherobots, a single statewide geospatial computation — like calculating the distance from every parcel in a state to the nearest water source — took seven days to complete. After moving their geospatial data pipelines to Wherobots, that same job ran in just 30 minutes on a medium compute instance. Scripts that once took a week to run could now be developed in a day and executed in under an hour. That acceleration unlocked a new rhythm for Aarden’s engineering team: iterate daily, explore new models, and scale from a single state to a national view of land opportunity. As founding staff engineer Steven Yee put it, “We went from babysitting compute jobs for a week to getting results before lunch.” Why Startups Like Aarden Choose Wherobots Why did Aarden.ai choose Wherobots? For scalability, ease of use, and integrated raster-vector support. For startups building data applications rooted in the physical world — agriculture, climate, energy, mobility, land, or infrastructure — data scale and spatial complexity are unavoidable. The Aarden team, led by geospatial scientist Ben Hudson, knew this well. Hudson’s background processing satellite imagery for Greenland’s ice sheet and building Zillow’s Zestimate engine gave him firsthand experience in the pain of scaling geospatial pipelines. When Aarden began, they tried the standard open-source stack: GeoPandas, RasterIO, GDAL. It worked for prototypes, but not for production. The choice came down to two questions: How do we scale spatial computation without building a Spark team? How do we avoid spending months managing infrastructure instead of shipping data products? The answer was Wherobots. “Spark is incredibly powerful — but it’s also a huge learning curve, Wherobots shortened the painful part of Spark and gave us production-grade scalability without having to babysit clusters.” Ben Hudson Co-Founder and Head of Applied Science, aarden.ai Wherobots’ full support for both rasters and vectors meant Aarden could seamlessly combine terrain, vegetation, and parcel data into unified models. They could also store and query data using Iceberg tables, eliminating the need for maintaining large Postgres clusters. The alternative platforms — like Google Earth Engine or Microsoft’s Planetary Computer — weren’t built for their hybrid vector-raster workflows or the flexibility needed to prototype and deploy quickly. “Wherobots is the most proven way to do it,” Hudson said. “It’s the reliable, full-featured, tried-and-true option.” Making Land Data Useful for Decision-Makers Aarden’s customers are institutional land investors evaluating large portfolios of property — for carbon capture, solar, timber, or data center opportunities. Wherobots powers the data engine behind Aarden’s platform, turning sprawling public datasets into clear, numeric insights. End users don’t see geotiffs or coordinate grids. They see a simple interface: Which land deals have the highest alpha potential? Behind that simplicity is Wherobots’ compute layer — transforming complex geospatial and environmental data into machine-learning-ready features and business metrics. “Our customers don’t need to be geospatial experts,” said Hudson. “They just need to make smart business decisions. Wherobots helps us turn a mountain of geospatial data into simple, singular, useful numbers and cash flow analyses.” Looking Ahead As Aarden scales, their focus is on robustness, repeatability, and rapid iteration. With Wherobots, they can run production-grade geospatial pipelines without worrying about cluster management or data scaling. And as they expand nationally, their confidence is simple: “It just works.” Aarden’s story reflects a broader trend. The next generation of data startups — those whose insights are grounded in the physical world — are choosing Wherobots to get from prototype to production faster. Because when your data is as big as the planet, you need compute that scales with it. Looking to get started for your spatial data pipelines and intelligence application? Get started today in community (free) or try out pro for your team. Or reach out to sales for a demo. Key takeawaysAarden.ai uses Wherobots to help institutional land investors evaluate parcels for data centers, renewables, carbon capture, solar, and timber after raising $4M coming out of stealth.A statewide job that calculated distance from every parcel to the nearest water source dropped from seven days on Aarden early Python stack to 30 minutes on a Wherobots medium compute instance—about 300x faster.Scripts that once took a week to run can now be developed in a day and executed in under an hour, so the team iterates daily and scales from a single state to a national view of land opportunity.GeoPandas, RasterIO, and GDAL worked for prototypes but not production; Wherobots gave them Spark-scale raster and vector processing without standing up a Spark team, plus Iceberg tables instead of large Postgres clusters.Founding engineer Steven Yee: we went from babysitting compute jobs for a week to getting results before lunch.
Advancing the Integration of Map Data via Overture’s Global Entity Reference System and Wherobots Posted on June 25, 2025October 3, 2026 by Ben Pruden Editor’s note: The Wherobots Spatial Data Catalog is now the Havasu Catalog. The general availability of the Overture Maps Foundation’s Global Entity Reference System (GERS) makes it a lot easier to build intelligence about features of our physical world. What is GERS? You can think of a GERS as a system for applying a unique key to physical features in the world. Overture assigns GERS IDs to millions of features in their data products, such as office buildings, highways, countries, rivers, schools, and more. Using GERS IDs, you can more easily join datasets to build a more complete view of physical-world features in space and measure relationships over time. Key benefits of GERS IDs include: Persistent identification: The same physical location maintains the same GERS ID over time Cross-dataset compatibility: Enable joining and enrichment across multiple datasets or providers Standardized reference: Provide a common language for location data across the geospatial ecosystem You can read more about GERS IDs in Overture’s documentation. Getting started with Overture in Wherobots At Wherobots, we’re proud to be a member of the Overture Maps Foundation, supporting the project since its formation, and as an official member since 2024. We also host and manage all of Overture’s recent datasets in the Wherobots Spatial Catalog, and they are offered at no additional cost to all customers. These datasets are production ready and at your fingertips. Select * from wherobots_open_data.overture.buildings_building New Wherobots customers can get started with Overture’s datasets in the free-to-use Community Edition, and graduate to the Professional edition when they want to join Overture data with their data in cloud storage. It’s Easy to use GERS in Wherobots A key design goal of GERS is to simplify how organizations can join and enrich their own datasets using canonical Overture datasets. Wherobots makes this vision real by providing built-in support for the GERS schema, enabling users to easily query, filter, and join to valuable geospatial datasets using GERS IDs, and through the use of spatial join predicates available in Apache Sedona and Wherobots. Whether data teams are working with parcel boundaries, retail site locations, road networks, or foot traffic telemetry, Wherobots makes it easy to enrich data using GERS with SQL or Python. Overture GERS Schema Extension Paths (not exhaustive) Stay tuned for a new tutorial from Wherobots that shows you how to use GERS IDs with Overture Places. Overtures datasets are produced using Wherobots Wherobots was founded by the original creators of Apache Sedona, the open-source engine for distributed geospatial processing. Overture now runs many of their Apache Spark and Sedona based data pipelines on Wherobots because they run up to 20x faster, at a fraction of the cost, and the Overture team benefits from the spatial expertise we offer them as a customer. The Overture team uses Wherobots’ Apache Airflow support to trigger job runs that power production of their planetary scale datasets. This feature and Wherobots compatibility with Spark and Sedona, also made it very easy for Overture to redirect where their Airflow-orchestrated jobs ran. “Overture produces a building dataset covering all buildings in the world, with 2.6B geometries and growing, that’s updated frequently. There’s a lot of data, and compute that goes into producing it and keeping it up to date. We accelerated the pipelines that produce the buildings dataset by up to 20x after we moved them to Wherobots, which required a simple redirection of our code. We retained compatibility with Apache Sedona, and the move put us into a development experience that’s made us more productive.” – Jennings Anderson, Geoscientist at Overture and Data Engineer at Meta. Redefining Standards in Open Source We’re committed to improving open standards in the geospatial data ecosystem. Wherobots is a leading contributor of spatial type support in Apache Iceberg and Parquet, the most popular open table and file formats for the cloud data lakehouse, to make it easier for companies to utilize geospatial data.We’ve partnered with Overture to improve the foundation for a scalable, versioned, and queryable world of features backed by GERS. By modernizing how geospatial data is accessed in the cloud via spatial data type support in Iceberg and Parquet, and improving accessibility and utility of open map data via GERS, we believe new use cases for spatial data will emerge to improve business operations, research, and our way of life. Join us in Building the Spatial Data Stack of the Future The launch of GERS is a big step for advancing spatial intelligence, and a leap toward a more open, interoperable data ecosystem. Wherobots is proud to be part of this journey, and we’re excited to continue supporting the Overture mission through operational pipelines, open standards, and accessible tools.If you’re building with geospatial data, we invite you to explore how Wherobots can help you take full advantage of GERS by easily joining your first party data with the GERS ID system. Next steps If you’d like a one-on-one demo from our team, you can request it here. Learn about our data mirror with Overture in their documentation. We will be publishing an example notebook that shows you how to use GERS IDs with Overture Places soon. Start Building with Wherobots Access Now Key takeawaysOverture Maps Foundation Global Entity Reference System (GERS) is generally available: stable unique keys on millions of physical features (buildings, highways, countries, rivers, schools) so datasets can be joined and tracked over time.Wherobots has been an Overture member since 2024 and hosts recent Overture datasets in the Spatial Catalog at no additional cost. Community Edition can query them; Professional Edition joins them to data in your cloud storage.Wherobots supports the GERS schema so you can query, filter, and spatially join first-party parcels, retail sites, roads, or telemetry to Overture with SQL or Python.Overture produces a global buildings dataset of 2.6 billion geometries (and growing). After moving Spark/Sedona pipelines to Wherobots—with a simple code redirect and Airflow job-run support—those pipelines accelerated by up to 20x while staying Apache Sedona compatible.Wherobots also led GEO type support in Apache Iceberg and Parquet so GERS-backed features can live in an open, versioned lakehouse rather than a proprietary silo.
Wherobots 2024 accomplishments, and what’s on-deck in 2025 Posted on January 23, 2025October 3, 2026 by Damian Editor’s note: The Wherobots Spatial Data Catalog is now the Havasu Catalog. Introduction 2024 was a transformative year for Wherobots. Our mission to revolutionize how geospatial data is used took significant strides forward, positively impacting our customers and industry. Over the past year, we more than tripled the size of our team and successfully closed a $21.5M Series A funding round. We expanded accessibility to Wherobots’ industry-leading geospatial query performance, integrated Wherobots into the native AWS buying experience, and unveiled groundbreaking features like Raster Inference, Map Matching, and GeoStats—empowering users to create scalable geospatial solutions like never before. Our Mission Before founding Wherobots, co-founders Mo and Jia identified critical challenges limiting the potential of geospatial data. These stemmed from how geospatial data was traditionally stored, formatted, and processed, and made this data incredibly painful to utilize, particularly at scale. Over the recent decades, data and analytics investment was mostly directed towards solutions for internet data. However compared to internet data, geospatial data is a lot more complex, which makes it harder to query. It’s polygons representing land and buildings, GPS trajectories, satellite and drone imagery, weather data, and more—all tied to Earth’s imperfect spherical surface. And querying this data generally means you need to filter and join it with other datasets (geo or non-geo). Due to this complexity, existing cloud analytics engines built for structured internet data struggle to efficiently run spatial queries at scale. They also miss features necessary to prepare this data, they lack features that make solution development productive, and simply cannot compute spatial results with high precision. As a result, solutions based on geospatial data are expensive, or otherwise shelved. We are addressing these challenges. By reducing the cost and effort to build with geospatial data, Wherobots will enable a new wave of innovation for the physical world. This will drive breakthroughs in products, business operations, science, government, and make a positive impact on our climate. Our mission is simple yet ambitious: make geospatial data easy to use. Here’s what some of our customers have to say about how we’re helping them achieve their missions. Customer Highlights AddressCloud Enabling insurers to calculate geographic risk with precision “Wherobots runs our compute operations that used to take hours or days to complete, in minutes. As we provide perils information (flood, fire, etc) to insurers at the property level, we particularly appreciate the ability to be able to run combined vector/raster analysis, without having to previously transform the raster data into vector format or some other format.” – John Powell, Senior Geospatial Data Engineer at AddressCloud Overture Maps Foundation Creating next-generation map products with scalable, open map data “Overture produces a building dataset covering all buildings in the world, with 2.3B geometries and growing, that’s updated frequently. There’s a lot of data and compute that goes into producing and keeping it up to date,” said Jennings Anderson, Geoscientist at Overture and Data Engineer at Meta. “We accelerated the pipelines that produce the buildings dataset by up to 20x after we moved them to Wherobots, which required a simple redirection of our code. We retained compatibility with Apache Sedona, and the move put us into a development experience that’s made us more productive.” – Jennings Anderson, Geoscientist at Overture and Data Engineer at Meta Why Wherobots Stands Out Several recurring themes highlight why customers choose Wherobots: Unmatched performance and cost efficiency: Wherobots delivers up to 20x better spatial join performance compared to modern cloud data engines, at a fraction of the cost. Ease of innovation: Wherobots makes it easy to build solutions with raster (e.g., satellite imagery), vector (e.g., geometry, geography) data, and your first party data regardless of scale. Modern cloud architecture: Wherobots is fully compatible with Apache Sedona, and runs seamlessly on data lakes with support for Apache Iceberg and Apache Parquet. 2024 Milestones Funding & Market Validation In 2024, we raised $21.5M in a Series A round led by Felicis, with support from Wing Venture Capital, Clear Ventures, JetBlue Ventures, and P7 Ventures. This funding reflects confidence in our mission and the massive market opportunity for geospatial solutions in the cloud. Team Growth The Wherobots team—the “Botsters”—tripled in size this year. While engineering saw the most growth, we also built out go-to-market, marketing, and product teams and are actively scaling our sales team. As we head into 2025, we’re actively hiring for roles across the company to support our expanding vision. Product Innovations We launched several key features in 2024 that expanded the boundaries of geospatial data solutions. *The features noted with an are only available in the professional or enterprise edition of Wherobots.**** Cloud Native A pay-as-you-go offering on the AWS Marketplace makes it easy to subscribe and pay on-demand using AWS Marketplace billing. A storage integration for Amazon S3, to quickly and securely integrate with first or third party data. Security and Access SAML Single Sign-On, makes logging in simple, secure, and seamless for users in companies with centralized login management systems. The Spatial SQL API, Typescript and Python SDKs, and a JDBC driver make it possible to query WherobotsDB using popular or custom query interfaces. Continuous improvement of internal security and service availability. Open Data Architecture The first version of the Spatial Catalog (known as Havasu, with core functionality soon to be merged into Apache Iceberg). Accelerating Geospatial Solution Development Raster Inference, to easily extract insights from satellite imagery at scale using SQL. (We’re hosting an upcoming panel discussion with an incredible lineup of speakers to discuss the MLM STAC Extension and Raster Inference, with a focus on optimizing Earth observation models for production. Learn more and save the date here.) Distributed Map Matching is a purpose built algorithm for snapping GPS trajectories to known segments like roads, with high performance at-scale. GeoStats: a geostatistics suite designed for scale, performance, and streamlining solution development. Support for K-nearest neighbor joins (exact and approximate) to efficiently query for geospatial neighbors at-scale. Vtiles, a vector tile solution purpose built for creating vector tiles at scale with high performance. Many new vector (ST) and raster (RS) functions to accelerate the developer productivity. Continuous improvement of spatial and non-spatial query performance to reduce cost and make workloads more compute efficient (reducing climate impact). Automation Job Runs integrated with Apache Airflow, to make it easy and familiar to automate new and existing processing workflows. Service Principals, enable authentication and automated usage of Wherobots, decoupled from the tenure or privileges of human users. Looking Ahead In 2025, we plan to bring Wherobots Cloud to the EU market with support for the AWS Europe (Ireland) region, and achieve the SOC 2 Type 2 certification (currently in progress). We’ll continue to focus on: Making Earth observation data easier to utilize. Enhancing developer productivity and experiences. Improving query engine performance and data compatibility. Strengthening service availability and support for customers. Delivering new administrative controls and observability. Ready to Build? We are currently offering a 30-day free trial covering up to $400 in usage via the AWS Marketplace. Getting started is easy. There are many example notebooks for various geospatial use cases that you can explore and run without any coding experience required. Not only do the notebooks help you get started, but we also see most of our customers use these notebooks as references for the solutions they end up building. Join the Mission Motivated by our mission? Join our growing team—visit our careers page for open roles. You can also share feedback at feedback@wherobots.com or contact me directly at damian@wherobots.com. Try Wherobots Pro Get Started Key takeawaysThis January 2025 year-in-review covers 2024 milestones rather than a single product launch: the team more than tripled, and Wherobots closed a $21.5M Series A led by Felicis, with Wing, Clear Ventures, JetBlue Ventures, and P7 Ventures.AddressCloud cut property-level flood/fire peril jobs from hours or days to minutes, including combined vector/raster analysis without converting rasters first. Overture global buildings pipeline (2.3 billion geometries in this post) accelerated up to 20x after a code redirect onto Wherobots.The company cites up to 20x better spatial-join performance than modern cloud data engines, full raster+vector support, Apache Sedona compatibility, and Iceberg/Parquet lakehouse architecture.2024 product launches included AWS Marketplace pay-as-you-go, S3 storage integration, SAML SSO, Spatial SQL API plus TypeScript/Python SDKs and JDBC, Spatial Catalog/Havasu, Raster Inference, distributed map matching, GeoStats, KNN joins, VTiles, Airflow job runs, and service principals. Starred items are Professional/Enterprise only.2025 plans stated here: AWS Europe (Ireland), SOC 2 Type 2 (then in progress), easier Earth observation, developer experience, query performance, availability, and admin/observability. A 30-day AWS Marketplace trial covered up to $400 in usage.
Accelerating Vector and Raster Analysis at GeoPostcodes Posted on November 25, 2024September 1, 2026 by Ben Pruden Accelerating Vector and Raster Analysis at GeoPostcodes How long is too long to wait for a data set to process? In the fast-paced world of data as a service, efficiency isn’t just a nice-to-have; it’s essential. But for GeoPostcodes, before Wherobots, previous updates to the global population movement datasets took about 39 days. This was running on a headless QGIS instance in combination with PostGIS, combining population rasters with postal code boundaries. For GeoPostcodes, before Wherobots, previous updates to the global population movement datasets took about 39 days. At Wherobots, we’ve made it our mission to push the boundaries of what’s possible in distributed spatial computing to not just process data faster, but to increase productivity. When GeoPostcodes, a leader in geospatial data products, sought to enhance their administrative boundaries enrichment processes with population estimates derived from the Global Human Settlement (GHS) population grid, they knew they needed a new solution. Together, we’ve managed to transform a process that once took nearly a month and a half into a task completed in hours. See how this works in an interactive notebook Launch Notebook Let’s take a closer look at how this collaboration reshaped GeoPostcodes capabilities to deliver fresh data in orders of magnitudes of less time. TLDR: The collaboration between Wherobots and GeoPostcodes yielded a significant improvement in processing time for GeoPostcodes new data products. This progress has powerful implications for development and applications in fields that rely on timely and accurate data processing. Read ahead for the details or check out the live stream recording. Why This Matters So, why is this performance gain so important? The ability to process complex geospatial data in a fraction of the time has profound implications for a wide range of industries. The benefits are not isolated to reduced analysis time and system complexity. Running analysis quickly allows data engineers and analysts to iterate faster, test out new product / data ideas, and decrease time to market. For urban planners, home builders, and retailers, having access to up-to-date population estimates can inform better decisions about infrastructure development, operations, and services. For logistics companies, understanding population distributions can enhance route planning and resource allocation. And for market analysts, demographic insights are invaluable for crafting targeted strategies and predicting consumer behavior. By reducing processing times from weeks to hours, Wherobots and GeoPostcodes are empowering these industries to make faster, more informed decisions. The days of waiting weeks for analysis results are over—now, critical data can be at your fingertips in just a matter of hours. About Wherobots and GeoPostcodes Wherobots is at the forefront of distributed spatial computing technology with hosted and highly optimized Apache Sedona (our founders are the original creators of the project). We specialize in providing a platform that empowers data teams to build and deliver data products efficiently and at scale. Through a cloud native, highly optimized service, built for developers and data engineers, Wherobots enables data teams to iterate and produce faster by running complex spatial computations at unprecedented scale and speed. Wherobots technology scales to meet any data workload, offering the flexibility needed to handle even the most demanding spatial data processing tasks efficiently. Whether it’s raster or vector analysis, or even machine learning tasks, it delivers the computational power required to get the job done faster and more effectively, while maintaining 100% code compatibility with open source Apache Sedona. GeoPostcodes has established itself as a trusted provider of high-quality geospatial data. Their datasets are vital to industries such as Logistics, Insurance, e-commerce, and market analysis, providing zip codes, cities, and administrative areas for 247 countries, available as 15.9M points (geocoding) and 880k boundaries (polygons). With their newest feature—population estimates derived from the GHS population grid—GeoPostcodes adds a valuable layer of demographic information to their datasets, giving users deeper insights and more powerful tools for decision-making. A bit about the data Mapping Units: GeoPostcodes provides high resolution administrative units at various scales (e.g. county, state, national) as well as zip code boundaries. The smallest administrative units and the zip code boundaries served as the mapping units. We enriched the administrative units with population estimates from the GHS data. Using the smallest administrative unit allowed us to roll the population estimates into the higher level administrative units reducing the need for additional processing. Global Human Settlement data (GHS) The Global Human Settlement layer, developed by the European Commission’s Joint Research Centre, provides global population estimates at a high spatial resolution. It is based on a combination of satellite imagery, census data, and other sources. The estimates are available in 5 year increments starting in 1975 and run through 2030. They are provided in a gridded format with a resolution of at various resolutions, for this analysis we chose the 100m resolution data. The Old Way: Over a month of processing in PostgreSQL Before the collaboration with Wherobots, GeoPostcodes had developed a method for deriving population estimates for their administrative boundaries dataset. This process used their PosGIS data warehouse and a dockerized headless instance of QGIS. In this setup, the GHS population grid data is copied to the docker container running QGis. This dataset, rich with global population distribution information, was then subjected to a series of complex spatial operations. These operations included spatial joins and aggregations, which were necessary to map population estimates to the corresponding administrative boundaries. Despite the robustness of this method, the entire process took thirty nine (39) days to complete. The sheer size of the datasets, coupled with the complexity of the spatial operations, meant that the database was constantly working at full capacity. This prolonged processing time posed a significant bottleneck, delaying critical processes for GeoPostcodes’ data operations teams. The Wherobots Solution: Hours Instead of Weeks This is where Wherobots came in. Leveraging the Wherobots distributed computing platform, we were able to reduce the processing time from weeks to just hours—a drastic improvement that has far-reaching implications for geospatial data analysis. Here’s how we did it. The process began with ingesting the administrative units and the GHS population grid into WherobotsDB. We leverage Wherobots Out-DB Rasters , RS_TileExplode() function, and repartitioning which allows more performant reading and workload distribution as only the required pixels for analysis are read and joined to the mapping units. From there, we partitioned the administrative units for even distribution of the workload. This partitioning allowed us to perform spatial joins and aggregations in parallel, significantly speeding up the process. While a single machine working through these operations sequentially would take days, our distributed approach meant that each task was handled simultaneously across multiple nodes, drastically reducing the overall processing time. Results Looking at a subset of the results and comparing them to the runtime of the legacy system we can clearly see that leveraging Apache Sedona on Wherobots provides increased performance and dramatically reduced run time. CountryWherobotsRun Time (minutes)Legacy SystemRun Time (minutes)Zip Code CountSpeed ratioJapan362113,29120xRussia66117042,98418xSaudi Arabia312537,28244xUnited States678030,144135xGreat Britain22510,05910xFrance3236,0359xCosta Rica1144709x Looking Ahead: The Future of Geospatial Data This collaboration between Wherobots and GeoPostcodes is just the beginning. As we continue to push the boundaries of what’s possible with distributed computing and geospatial analysis, we’re excited about the future possibilities. On the horizon are even more advanced features and capabilities. Real-time data processing and machine learning integration are among the innovations we’re exploring, promising to make geospatial data analysis not just faster, but also smarter and more adaptive to changing conditions. Together, Wherobots and GeoPostcodes are setting a new standard for efficiency and precision in the world of geospatial data. Whether you’re in urban planning, Insurance, logistics, or any other industry that relies on this data, the future’s looking brighter—and faster—than ever before. Start building with Wherobots Get Started Key takeawaysGeoPostcodes used to refresh global population movement datasets in about 39 days on a headless QGIS instance plus PostGIS, combining population rasters with postal-code boundaries.On Wherobots the same enrichment finished in hours, using out-DB rasters, RS_TileExplode(), and repartitioning so only required pixels are read and joined to mapping units.GeoPostcodes covers zip codes, cities, and administrative areas for 247 countries as 15.9 million geocoding points and 880,000 boundary polygons, now enriched with Global Human Settlement (GHS) population estimates.The analysis used GHS 100-meter resolution grids in 5-year increments from 1975 through 2030, rolling estimates from the smallest administrative units up to higher levels.Country runtimes versus the legacy system include United States 6 vs 780 minutes (135x, 30,144 zip codes), Saudi Arabia 3 vs 125 minutes (44x), Japan 3 vs 62 minutes (20x), and Russia 66 vs 1,170 minutes (18x).
🌶 Comparing taco chains :: a consumer retail cannibalization study with isochrones Posted on June 11, 2024September 1, 2026 by Ben Pruden Authors Ilya Marchenko Using Wherobots for a Retail Cannibalization Study Comparing Two Leading Taco Chains In this post, we explore how to implement a workflow from the commercial real estate (CRE) space using Wherobots Cloud. This workflow is commonly known as a cannibalization study, and we will be using WherobotsDB, POI data from OvertureMaps, the open source Valhalla API, and visualization capabilities offered by SedonaKepler. What is a retail cannibalization study? In CRE (consumer real estate), stakeholders are often interested in questions like “If we build a new fast food restaurant here, how will its performance be affected by other similar fast food locations that already exist nearby?”. The idea of the new fast food restaurant “eating into” the sales of other fast food restaurants that already exist nearby is what is known as ‘cannibalization’. The main objective of studying this phenomenon is to determine the extent to which a new store might divert sales from existing stores owned by the same company or brand and evaluate the overall impact on the company’s market share and profitability in the area. Cannibalization Study in Wherobots For this case study, we will look at two taco chains which are located primarily in Texas: Torchy’s Tacos and Velvet Taco. In general, information about the performance of individual locations and customer demographics are often proprietary information. We can, however, still learn a great deal about the potential for cannibalization both between these two chains as competitors, and between individual locations of each chain. We also know, based on our own experience, these chains compete with each other. Which taco shop to go to when we are visiting Texas is always a spicy debate. We begin by importing modules that will be useful to us as we go on. import geopandas as gpd import pandas as pd import requests from sedona.spark import * from pyspark.sql.functions import explode, array from pyspark.sql import functions as F Next, we can initiate a Sedona context. config = SedonaContext.builder().getOrCreate() sedona = SedonaContext.create(config) Identifying Points of Interest Now, we need to retrieve the locations of Torchy’s Tacos and Velvet Taco locations. In general, one can do this via a variety of both free and paid means. We will look at a simple, free approach that is made possible by the integration of Overture Maps data into the Wherobots environment: sedona.table("wherobots_open_data.overture.places_place"). \\ createOrReplaceTempView("places") We create a view of the Overture Maps places database, which contains information on points of interest (POI’s) worldwide. Now, we can select the POI’s which are relevant to this exercise: stores = sedona.sql(""" SELECT id, names.common[0].value as name, ST_X(geometry) as long, ST_Y(geometry) as lat, geometry, CASE WHEN names.common[0].value LIKE "%Torchy's Tacos%" THEN "Torchy's Tacos" ELSE 'Velvet Taco' END AS chain FROM places WHERE addresses[0].region = 'TX' AND (names.common[0].value LIKE "%Torchy's Tacos%" OR names.common[0].value LIKE '%Velvet Taco%') """) Calling stores.show() gives us a look at the spark DataFrame we created: +--------------------+--------------+-----------+----------+ | id| name| long| lat| +--------------------+--------------+-----------+----------+ |tmp_8104A79216254...|Torchy's Tacos| -98.59689| 29.60891| |tmp_D17CA8BD72325...|Torchy's Tacos| -97.74175| 30.29368| |tmp_F497329382C10...| Velvet Taco| -95.48866| 30.18314| |tmp_9B40A1BF3237E...|Torchy's Tacos| -96.805853| 32.909982| |tmp_38210E5EC047B...|Torchy's Tacos| -96.68755| 33.10118| |tmp_DF0C5DF6CA549...|Torchy's Tacos| -97.75159| 30.24542| |tmp_BE38CAC8D46CF...|Torchy's Tacos| -97.80877| 30.52676| |tmp_44390C4117BEA...|Torchy's Tacos| -97.82594| 30.4547| |tmp_8032605AA5BDC...| Velvet Taco| -96.469695| 32.898634| |tmp_0A2AA67757F42...|Torchy's Tacos| -96.44858| 32.90856| |tmp_643821EB9C104...|Torchy's Tacos| -97.11933| 32.94021| |tmp_0042962D27E06...| Velvet Taco|-95.3905374|29.7444214| |tmp_8D0E2246C3F36...|Torchy's Tacos| -97.15952| 33.22987| |tmp_CB939610BC175...|Torchy's Tacos| -95.62067| 29.60098| |tmp_54C9A79320840...|Torchy's Tacos| -97.75604| 30.37091| |tmp_96D7B4FBCB327...|Torchy's Tacos| -98.49816| 29.60937| |tmp_1BB732F35314D...| Velvet Taco| -95.41044| 29.804| |tmp_55787B14975DD...| Velvet Taco|-96.7173913|32.9758554| |tmp_7DC02C9CC1FAA...|Torchy's Tacos| -95.29544| 32.30361| |tmp_1987B31B9E24D...| Velvet Taco| -95.41006| 29.770256| +--------------------+--------------+-----------+----------+ only showing top 20 rows We’ve retrieved the latitude and longitude of our locations, as well as the name of the chain each location belongs to. We used the CASE WHEN statement in our query in order to simplify the location names. This way, we can easily select all the stores from the Torchy’s Tacos chain, for example, and not have to worry about individual locations being called things like “Torchy’s Tacos – Rice Village” or “Velvet Taco Midtown”, etc. We can also visualize these locations using SedonaKepler. First, we can create the map using the following snippet: location_map = SedonaKepler.create_map(stores, "Locations", config = location_map_cfg) Then, we can display the results by simply calling location_map in the notebook. For convenience, we included the location_map_cfg Python dict in our notebook, which stores the settings necessary for the map to be created with the locations color-coded by chain. If we wish to make modifications to the map and save the new configuration for later use, we can do so by calling location_map.config and saving the result either as a cell in our notebook or in a separate location_map_cfg.py file. Generating Isochrones Now, for each of these locations, we can generate a polygon known as an isochrone or drivetime. These polygons will represent the areas that are within a certain time’s drive from the given location. We will generate these drivetimes using the Valhalla isochrone api: def get_isochrone(lat, lng, costing, time_steps, name, location_id): url = "<https://valhalla1.openstreetmap.de/isochrone>" params = { "locations": [{"lon": lng, "lat": lat}], "contours": [{"time": i} for i in time_steps], "costing": costing, "polygons": 1, } response = requests.post(url, json=params) if response: result = response.json() if 'error_code' not in result.keys(): df = gpd.GeoDataFrame.from_features(result) df['name'] = name df['id'] = location_id return df[['name','id','geometry']] The function takes as its input a latitude and longitude value, a costing paratemeter, a location name, and a location id. The output is a dataframe which contains a Shapely polygon representing the isochrone, along with the a name and id of the location the isochrone corresponds to. We have separate columns for a location id and a location name so that we can use the id column to examine isochrones for individual restaurants and we can use the name column to look at isochrones for each of the chains. The costing parameter can take on several different values (see the API reference here), and it can be used to create “drivetimes” assuming the user is either walking, driving, or taking public transport. We create a geoDataFrame of all of the 5-minute drivetimes for our taco restaurant locations drivetimes_5_min = pd.concat([get_isochrone(row.lat, row.long, 'auto', [5], row.chain, row.id) for row in stores.select('id','chain','lat','long').collect()]) and then save it to our S3 storage for later use: drivetimes_5_min.to_csv('s3://path/drivetimes_5_min_torchys_velvet.csv', index = False) Because we are using a free API and we have to create quite a few of these isochrones, we highly recommend saving the file for later analysis. For the purposes of this blog, we have provided a ready-made isochrone file here, which we can load into Wherobots with the following snippet: sedona.read.option('header','true').format('csv') .\\ load('s3://path/drivetimes_5_min_torchys_velvet.csv') .\\ createOrReplaceTempView('drivetimes_5_min') We can now visualize our drivetime polygons in SedonaKepler. As before, we first create the map with the snippet below. map_isochrones = sedona.read.option('header','true').format('csv'). \\ load('s3://path/drivetimes_5_min_torchys_velvet.csv') isochrone_map = SedonaKepler.create_map(map_isochrones, "Isochrones", config = isochrone_map_cfg) Now, we can display the result by calling isochrone_map . The Analysis At this point, we have a collection of the Torchy’s and Velvet Taco locations in Texas, and we know the areas which are within a 5-minute drive of each location. What we want to do now is to estimate the number of potential customers that live near each of these locations, and the extent to which these populations overlap. A First Look Before we look at how these two chains might compete with each other, let’s also take a look at the extent to which restaurants within each chain might be cannibalizing each others’ sales. A quick way to do this is by using the filtering feature in Kepler to look at isochrones for a single chain: We see that locations for each chain are fairly spread out and (at least at the 5-minute drivetime level), there is not a high degree of cannibalization within each chain. Looking at the isochrones for both chains, however, we notice that Velvet Taco locations often tend to be near Torchy’s Tacos locations (or vice-versa). At this point, all we have are qualitative statements based on these maps. Next, we will show how to use H3 and existing open-source datasets to make these statements more quantitative. Estimating Cannibalization Potential As we can see by looking at the map of isochrones above, they are highly irregular polygons which have a considerable amount of overlap. In general, these polygons are not described in a ‘nice’ way by any administrative boundaries such as census block groups, census tracts, etc. Therefore, we will have to be a little creative in order to estimate the population inside them. One way of doing this using the tools provided by Apache Sedona and Wherobots is to convert these polygons to H3 hexes. We can do this with the following snippet: sedona.sql(""" SELECT ST_H3CellIds(ST_GeomFromWKT(geometry), 8, false) AS h3, name, id FROM drivetimes_5_min """).select(explode('h3'), 'name','id').withColumnRenamed('col','h3') .\\ createOrReplaceTempView('h3_isochrones') This turns our table of drivetime polygons into a table where each row represents a hexagon with sides roughly 400m long, which is a part of a drivetime polygon. We also record the chain that these hexagons are associated to (the chain that the polygon they came from belongs to). We store each hexagon in its own row because this will simplify the process of estimating population later on. Although the question of estimating population inside individual H3 hexes is also a difficult one (we will release a notebook on this soon), open-source datasets with this information are available online, and we will use one such dataset, provided by Kontur: kontur = sedona.read.option('header','true') .\\ load('s3://path/us_h3_8_pop.geojson', format="json") .\\ drop('_corrupt_record').dropna() .\\ selectExpr('CAST(CONV(properties.h3, 16, 10) AS BIGINT) AS h3', 'properties.population as population') kontur.createOrReplaceTempView('kontur') We can now enhance our h3_isochrones table with population counts for each H3 hex: sedona.sql(""" SELECT ST_H3CellIds(ST_GeomFromWKT(geometry), 8, false) AS h3, name, id FROM drivetimes_5_min """).select(explode('h3'), 'name','id').withColumnRenamed('col','h3') .\\ join(kontur, 'h3', 'left').distinct().createOrReplaceTempView('h3_isochrones') At this stage, we can also quickly compute the cannibalization potential within each chain. Using the following code, for example, we can estimate the number of people who live within a 5 minute drive of more than one Torcy’s Tacos: sedona.sql(""" SELECT ST_H3CellIds(ST_GeomFromWKT(geometry), 8, false) AS h3, name, id FROM drivetimes_5_min """).select(explode('h3'), 'name','id').withColumnRenamed('col','h3') .\\ join(kontur, 'h3', 'left').filter('name LIKE "%Torchy%"').select('h3','population') .\\ groupBy('h3').count().filter('count >= 2').join(kontur, 'h3', 'left').distinct() .\\ agg(F.sum('population')).collect()[0][0] 97903.0 We can easily change this code to compute the same information for Velvet Taco by changing filter('name LIKE "%Torchy%"') in line 4 of the above snippet to filter('name LIKE "%Velvet%"') . If we do this, we will see that 100298 people live within a 5 minute drive of more than one Velvet Taco. Thus, we see that the Torchy’s Tacos brand appears to be slightly better at avoiding canibalization among its own locations (especially given that Torchy’s Tacos has more locations than Velvet Taco). Now, we can run the following query to show the number of people in Texas who live within a 5 minutes drive of a Torchy’s Tacos: sedona.sql(""" WITH distinct_h3 (h3, population) AS ( SELECT DISTINCT h3, ANY_VALUE(population) FROM h3_isochrones WHERE name LIKE "%Torchy's%" GROUP BY h3 ) SELECT SUM(population) FROM distinct_h3 """).show() The reason we select distinct H3 hexes here is because a single hex can belong to more than one isochrone (as evidenced by the SedonaKepler visualizations above). We get the following output: +---------------+ |sum(population)| +---------------+ | 1546765.0| +---------------+ So roughly 1.5 million people in Texas live within a 5-minute drive of a Torchy’s Tacos location. Looking at our previous calculations for how many people live near more than one restaurant of the same chain, we can see that Torchy’s Tacos locations near each other cannibalize about 6.3% of the potential customers who live within 5 minutes of a Torchy’s location. Running a similar query for Velvet Taco tells us that roughly half as many people live within a 5-minute drive of a Velvet Taco: sedona.sql(""" WITH distinct_h3 (h3, population) AS ( SELECT DISTINCT h3, ANY_VALUE(population) FROM h3_isochrones WHERE name LIKE '%Velvet Taco%' GROUP BY h3 ) SELECT SUM(population) FROM distinct_h3 """).show() +---------------+ |sum(population)| +---------------+ | 750360.0| +---------------+ As before, we can also see that Velvet Taco locations near each other cannibalize about 13.4% of the potential customers who live within 5 minutes of a Velvet Taco location. Now, we can estimate the potential for cannibalization between these two chains: sedona.sql(""" WITH overlap_h3 (h3, population) AS ( SELECT DISTINCT a.h3, ANY_VALUE(a.population) FROM h3_isochrones a LEFT JOIN h3_isochrones b ON a.h3 = b.h3 WHERE a.name != b.name GROUP BY a.h3 ) SELECT sum(population) FROM overlap_h3 """).show() which gives: +---------------+ |sum(population)| +---------------+ | 415033.0| +---------------+ We can see that more than half of the people who live near a Velvet Taco location also live near a Torchy’s Tacos location and we can visualize this population overlap: isochrones_h3_map_data = sedona.sql(""" SELECT ST_H3CellIds(ST_GeomFromWKT(geometry), 8, false) AS h3, name, id FROM drivetimes_5_min """).select(explode('h3'), 'name','id').withColumnRenamed('col','h3') .\ join(kontur, 'h3', 'left').select('name','population',array('h3')).withColumnRenamed('array(h3)','h3').selectExpr('name','population','ST_H3ToGeom(h3)[0] AS geometry') isochrones_h3_map = SedonaKepler.create_map(isochrones_h3_map_data, 'Isochrones in H3', config = isochrones_h3_map_cfg) Create a Free Account Get Started Key takeawaysThe notebook studies retail cannibalization between Torchy’s Tacos and Velvet Taco in Texas using Overture Places, 5-minute drive isochrones from the Valhalla API, H3, and Kontur population.Drive-time polygons are converted to H3 cells at resolution 8 (hex sides roughly 400 meters) and joined to Kontur US H3-8 population.About 1,546,765 people live within a 5-minute drive of a Torchy’s; about 750,360 live within 5 minutes of a Velvet Taco.97,903 people live within 5 minutes of more than one Torchy’s (about 6.3% of Torchy’s 5-minute population). 100,298 people are within 5 minutes of more than one Velvet Taco (about 13.4%).415,033 people live in the overlap of both chains’ 5-minute isochrones—more than half of Velvet Taco’s 5-minute population also lives near a Torchy’s.
Raster Data Analysis, Processing Petabytes of Agronomic Data, Overview of Sedona 1.5, and Unlocking The Spatial Frontier – This Month In Wherobots Posted on March 13, 2024October 3, 2026 by Ben Pruden Welcome to This Month In Wherobots, the monthly newsletter for data practitioners in the Apache Sedona and Wherobots community. This month we’re exploring raster data analysis with Spatial SQL, processing petabytes of agronomic data with Apache Sedona, a deep dive on new features added in the 1.5 release series, and an overview of working with files in Wherobots Cloud. Raster Data Analysis With Spatial SQL & Apache Sedona One of the strengths of Apache Sedona and Wherobots Cloud is the ability to work with large scale vector and raster geospatial data together using Spatial SQL. This post (and video) takes a look at how to get started working with raster data in Sedona using Spatial SQL and some of the use cases for raster data analysis including vector / raster join operations, zonal statistics, and using raster map algebra. Read The Article: Raster Data Analysis With Spatial SQL & Apache Sedona Featured Community Member: Luiz Santana This month’s Wherobots & Apache Sedona featured community member is Luiz Santana. Luiz is Co-Founder and CTO of Leaf Agriculture. He has extensive experience as a former data architect and developer. Luiz has a PhD in Computer Science from Universidade Federal de Santa Catarina and spent time researching data processing and integration in highly scalable environments. Leaf Agriculture is building the unified API for food and agriculture by leveraging large-scale agronomic data. Luiz has given several presentations at conferences such as Apache Sedona: How To Process Petabytes of agrnomic data with Spark, and Perspectives on the use of data in Agriculture, which covers how Leaf uses Apache Sedona to analyze large-scale agricultural data and how Sedona fits into their stack alongside other technologies. Thank you Luiz for your work with the Apache Sedona and Wherobots community and sharing your knowledge and experience! Apache Sedona: How To Process Petabytes of Agronomic Data With Spark In this presentation from The Developer’s Conference Luiz Santana shares the experience of using Apache Sedona at Leaf Agriculture to process petabytes of agronomic data from satellites, agricultural machines, drones and other sensors. He discusses how Leaf uses Sedona for tasks such as geometry intersections, geographic searches, and polygon transformations with high performance and speed. Luiz also presented Perspectives On The Use of Data in Agriculture which covers some of the data challenges that Leaf handles and an overview of the technologies used to address these challenges, including Apache Sedona. See The Slides From The Presentation Working With Files – Getting Started With Wherobots Cloud This post takes a look at loading and working with our own data in Wherobots Cloud as well as creating and saving data as the result of our analysis, such as the end result of a data pipeline. It covers importing files in various formats including CSV, GeoJSON, Shapefile, and GeoTIFF in Wherobots Cloud, working with AWS S3 cloud object storage, and creating GeoParquet files using Apache Sedona. Read The Post: Working With Files – Getting Started With Wherobots Cloud Introducing Sedona 1.5: Making Sedona the most comprehensive & scalable spatial data processing and ETL engine for both raster and vector data The 1.5 series of Apache Sedona represents a leap forward in geospatial processing that adds essential features and enhancements to make Sedona a comprehensive, all-in-one cluster computing engine for geospatial vector and raster data analysis. This post covers XYZM coordinates and SRID, vector and raster joins, raster data manipulation, visualization with SedonaKepler and SedonaPyDeck, GeoParquet reading and writing, H3 hexagons, and new cluster compute engine support. Read The Post: Introducing Sedona 1.5 – Making Sedona The Most Comprehensive & Scalable Spatial Data Processing and ETL Engine For Both Raster and Vector Data Unlocking the Spatial Frontier: The Evolution and Potential of spatial technology in Apple Vision Pro and Augmented Reality Apps Apple adopted the term “spatial computing” when announcing the Apple Vision Pro to describe its new augmented reality platform. This post from Wherobots CEO Mo Sarwat examines spatial computing in the context of augmented reality experiences to explore spatial object localization and presentation and the role of spatial query processing and spatial data analytics in Apple Vision Pro. Read The Post: Unlocking The Spatial Frontier Upcoming Events Apache Sedona Community Office Hour – (Online Zoom Call – April 9, 2024) – Join the Apache Sedona community on the second Tuesday of each month for updates on the state of Apache Sedona, presentation and demo of recent features, and provide your input into the roadmap, future plans, and contribution opportunities. Training Series – Large Scale Geospatial Analytics With Graphs And The PyData Ecosystem – (Online – March 26, 2024) – For our monthly livestream this month we’re partnering with our friends at Neo4j for a live virtual training series focused on geospatial analytics, the PyData ecosystem, and how graphs and Neo4j can be used alongside Apache Sedona and Wherobots Cloud to make sense of geospatial data. Register for free here. Subsurface Conference – Havasu: A Table Format For Spatial Attributes In A Data Lake Architecture (Online – May 2, 2024) – This talk at the Subsurface Live conference will introduce the Havasu open table format that extends Apache Iceberg to support spatial data. SW2 Conference – Cloud Native Geospatial Analytics With Apache Sedona, GeoParquet and Apache Iceberg (Denver – May 15, 2024) – This talk will cover recent developments in the cloud-native data ecosystem that address analyzing geospatial data at scale including Apache Sedona for executing large scale spatial queries, GeoParquet an optimized storage format for efficiently managing spatial data, and the Havasu extension to the Apache Iceberg open table format for managing tables in a spatial data lakehouse. Subscribe to the Spatial Intelligence Newsletter to get the latest updates on Wherobots, Apache Sedona and geospatial data: Key takeawaysThis March 2024 “This Month In Wherobots” newsletter points at other posts rather than introducing new product numbers.Featured member Luiz Santana (Leaf Agriculture) describes using Apache Sedona with Spark to process petabytes of agronomic data from satellites, machines, drones, and sensors.Linked tutorials cover raster analysis with Spatial SQL (vector/raster joins, zonal stats, map algebra), working with files in Wherobots Cloud, Sedona 1.5, and Mo Sarwat’s Apple Vision Pro spatial-computing essay.Upcoming events listed: Sedona office hour April 9, a Neo4j + Sedona training on March 26, Havasu at Subsurface Live on May 2, and a cloud-native geospatial talk at SW2 in Denver on May 15.Leaf’s TDC talk is titled Apache Sedona: How To Process Petabytes of agronomic data with Spark, alongside a second talk on data in agriculture and how Sedona sits in Leaf’s stack.
Introducing The Havasu Spatial Table Format, Wherobots 2023 In Review, Analyzing Overture Maps and Real Estate Data: This Month In Wherobots Posted on January 3, 2024September 1, 2026 by Ben Pruden In the latest edition of This Month In Wherobots the latest highlights from the Wherobots & Apache Sedona community include an overview of the Havasu spatial table format, a look back at Wherobots’ 2023 in review, analyzing Overture Maps and real estate data, finding the perfect Christmas tree, plus big data analytics for sustainable smart cities. Havasu: A Table Format for Spatial Attributes In A Data Lake Architecture The Havasu spatial table format is an extension of Apache Iceberg that brings native spatial support to data lakes. This post introduces Havasu and some of its key features and technical insights such as support for primitive spatial data types and storage, spatial statistics, and spatial filter push down and indexing. Examples of creating and querying Havasu tables as well as how Havasu working with GeoParquet is also covered. If you’ve used the Wherobots open data catalog you may have already leveraged Havasu! Read More About The Havasu Spatial Table Format Featured Community Member: Muhammed O?uzhan METE Our featured community member this month is Dr. Muhammed O?uzhan METE. Muhammed is an Assistant Professor at Istanbul Technical University in the Geomatics Engineering Department. His research areas of focus are land management, real estate management, GIS, machine learning, and big data analytics. He is also an AWS Community Builder. Muhammed presented “Geospatial Big Data Analytics For Sustainable Smart Cities” at the FOSS4G 2023 conference where he shared how Apache Sedona can be used for large scale spatial analytics. Connect with Muhammed On LinkedIn Geospatial Big Data Analytics For Sustainable Smart Cities In this presentation from FOSS4G 2023 Dr. Muhammed Oguzhan Mete covers the need for large-scale geospatial data analysis for sustainable smart city project development. He covers some of the use cases for large scale data analysis relevant for smart cities such as how to make smart infrastructure decisions to meet goals for sustainability. He discusses how to work with geospatial big data in cloud computing environments for the purposes of analyzing energy performance of buildings at scale and how spatial joins and spatial clustering algorithms can be implemented using Apache Sedona and Dask GeoPandas to identify clusters of lower energy efficiency scores to inform policy decision making. Finally, he shows a benchmark comparing the performance of Apache Sedona and Dask GeoPandas for spatial joins and spatial clustering. You’ll have to watch the video to see the final results! Watch the recording of “Geospatial Big Data Analytics For Sustainable Smart Cities” Wherobots: 2023 Year In Review 2023 has been an exciting year for Wherobots and this post points out a handful of key moments for Wherobots and the Apache Sedona community. Highlights of the year included 130% growth of Apache Sedona usage, key hires for growing the Wherobots team, launching Wherobots Cloud and SedonaDB, and raising a $5.5M seed funding round. Read More About Wherobots: 2023 Year In Review Analyzing The Overture Maps Places Dataset In this article from Pranav Toggi we learn how to query and analyze the Overture Maps Places dataset using SedonaDB and Wherobots Cloud. After a review of the data schema we see how to filter for points of interest within New York City and explore businesses within walking distance of stadiums and event venues. Along the way we learn several ways to visualize the data and results of our analysis. Read “Analyzing The Overture Maps Places Dataset Using SedonaDB, Wherobots Cloud, & GeoParquet Finding The Perfect Christmas Tree With USFS Map Data, QGIS, & SedonaDB This tutorial walks through how to use US Forest Service road data and aerial imagery to find the perfect Christmas tree. It covers loading and querying USFS road data, annotating aerial imagery in QGIS, then combining them and querying using SQL to find routes in National Forests to the perfect tree stands. Read “Finding The Perfect Christmas Tree With USFS Map Data, QGIS, & SedonaDB” Analyzing Real Estate Data With SedonaDB In a livestream tutorial on the Wherobots YouTube channel we worked through how to use data from Zillow to analyze real estate values in the US. We calculated the change in real estate values at the county level over the last 5 years and created choropleth maps to help visualize the results. We also wrote up a written tutorial so you can follow along in Wherobots Cloud. Watch Analyzing Real Estate Data With SedonaDB Upcoming Events GeoBuiz Summit (Monteray California – January 9-11, 2024) – Wherobots CEO Mo Sarwat will be presenting at the GeoBuiz Summit in a plenary panel addressing the “Geospatial Value Chain for Retail and Commerce”. If you’re at the conference stop by and say hi! GeoParquet Community Day (San Francisco – January 30th, 2024) – Join us for GeoParquet Community Day to highlight the usage of spatial data in Parquet, open table formats, and cloud-native spatial analytics. Data Day Texas (Austin, TX – January 27, 2024) – Will from Wherobots will be presenting a talk on geospatial analytics with Apache Sedona and Spatial SQL Hands On With Havasu and GeoParquet (Online Livestream – January 25th, 2024) – In this livestream we’ll take a look at the Havasu spatial table format and how Havasu can be used with GeoParquet. Be sure to subscribe to the Wherobots YouTube channel to keep up to date with more Wherobots livestreams and videos! Want to receive this monthly update in your inbox? Sign up for the Spatial Intelligence Newsletter: Key takeawaysThis January 2024 “This Month In Wherobots” newsletter recaps Havasu, a 2023 year-in-review, Overture Places, a Christmas-tree tutorial, and Zillow real-estate analysis.Havasu is introduced as an Apache Iceberg extension with native spatial types, spatial statistics, filter push-down, and indexing, including GeoParquet-backed tables already used by the open data catalog.The 2023 review highlights 130% growth in Apache Sedona usage, key hires, the launch of Wherobots Cloud and SedonaDB, and a $5.5 million seed round.Featured community member Dr. Muhammed Oğuzhan Mete (Istanbul Technical University) presented Geospatial Big Data Analytics For Sustainable Smart Cities at FOSS4G 2023 using Sedona and Dask GeoPandas.Upcoming: GeoBuiz Summit January 9–11, Data Day Texas January 27, GeoParquet Community Day in San Francisco January 30, and a Havasu/GeoParquet livestream January 25.