Rasterflow, Earth Intelligence & inference engine now in public preview Learn More

What is the Havasu Catalog? A Spatial Data Catalog

Authors

The Havasu Catalog is the Wherobots spatial data catalog: one index of open geospatial datasets, RasterFlow computer vision models, solution notebooks, and demo apps, with every dataset stored as an Apache Iceberg table that WherobotsDB queries in place with spatial SQL. The name comes from Havasu, the spatial table format Wherobots built to give Iceberg tables geometry and raster columns, spatial statistics, and spatial filter pushdown before Iceberg supported geospatial types natively.

Key takeaways

  • A spatial data catalog maps table names to data files and adds geometry and raster types, a CRS per column, and per-file spatial extents.
  • Havasu extends the Apache Iceberg table spec with geometry and raster types and bounding boxes in the manifest files. Iceberg format version 3 now defines native geometry and geography types, and Wherobots upgrades its open data tables to them automatically.
  • The Havasu Catalog lists 63 entries: 42 datasets, 6 RasterFlow models, 9 solution notebooks, and 6 demo apps.
  • Every table carries Iceberg snapshots, so a query can read an earlier Overture Maps release with VERSION AS OF, and metadata tables expose snapshots, tags, and data files to SQL.
  • In the Overture buildings table, 3 of 5,141 data files have bounding boxes that overlap a Lake Havasu City query window, so a spatial filter can skip the other files before reading any data.

What is a spatial data catalog?

A data catalog is the layer that maps a table name such as wherobots_open_data.overture_maps_foundation.buildings_building to the files that hold it, the schema of those files, and the history of changes. A spatial data catalog adds three things that location data needs:

  • Spatial types. Geometry columns for points, lines, and polygons, and raster columns for imagery and elevation, stored as typed values.
  • A coordinate reference system per column. The coordinate reference system says how the coordinates map to places on Earth, so two tables in different systems are never compared by accident.
  • Spatial extents. The bounding box of every data file, recorded in metadata, so a query for one city opens only the files whose boxes touch that city.

Asset catalogs such as STAC describe files, mostly satellite scenes, with JSON records of footprint, time, and band links. Table catalogs such as Iceberg describe rows in analytic tables. The Havasu Catalog is a table catalog that can hold asset catalogs as tables: the Sentinel-2 L2A scene index is a table of STAC items, one row per scene.

How the Havasu Catalog works

Each name in the Havasu Catalog points to an Iceberg table whose metadata and Apache Parquet files sit in cloud object storage.

Catalogs, databases, and tables

Each catalog holds databases, each database holds tables, and SQL names a table as <catalog>.<database>.<table> (querying datasets in the Wherobots docs). Wherobots Cloud provides wherobots_open_data to every organization edition, wherobots_pro_data to paid editions, and an organization catalog for a team's own tables. A team can also attach tables through Unity Catalog, AWS Glue, or any Iceberg catalog, so one query can join private and open data.

Havasu: Iceberg with geometry and raster columns

Apache Iceberg is an open table format for large analytic tables. Its table spec sets the goals the Havasu Catalog inherits: reads see one committed snapshot of a table, writes add and remove files in one atomic operation, and schema evolution supports adding, dropping, reordering, and renaming columns.

The Havasu table spec (version 0.1.0) extends Iceberg format versions 1 and 2 in four ways and leaves the rest of Iceberg unchanged:

  • Primitive spatial types. geometry follows the OGC Simple Features standard and can be encoded as WKB, EWKB, WKT, or GeoJSON in Parquet. raster stores a grid with its affine geo-referencing and CRS.
  • Raster storage modes. Band values can sit in the table (in-db) or in external files such as Cloud Optimized GeoTIFFs (out-db), with only the URI and geo-referencing in the row.
  • Spatial statistics. Each data file's entry in the manifest records the minimum bounding rectangle of its geometry or raster values. Raster bounds are always stored in WGS84.
  • Spatial transforms. Functions that partition or cluster rows by location.

Havasu supports ACID transactions, schema evolution, and time travel on geometries and rasters, and every Iceberg feature except merge-on-read tables.

In WherobotsDB, a column declared as geometry(4326) is an Iceberg v3 native column with one enforced CRS, readable as geometry by other Iceberg engines with geospatial support, and geography(4326) interprets coordinates on a sphere. Older columns without a declared CRS work the same in WherobotsDB, and other engines read them as binary. Wherobots upgrades the read-only public tables it manages, such as those in wherobots_open_data, automatically.

Snapshots, time travel, and metadata tables

Every commit to an Iceberg table writes a new snapshot: the complete list of data files in the table at that moment. Older snapshots stay readable until they expire, and named references (tags and branches) point at particular snapshots.

The Havasu Catalog uses tags for Overture Maps. Wherobots keeps every stable Overture release as a tag, so VERSION AS OF '2025-05-21.0' reads the places table as Overture published it on 21 May 2025, and a query with no version reads the latest stable release. Iceberg exposes this history to SQL through metadata tables appended to the table name: .snapshots lists commits with a summary of row and file counts, .refs lists tags and branches, and .files lists every data file with its row count and column statistics.

Spatial filter pushdown

With spatial filter pushdown, before any data file opens, WherobotsDB compares each file's bounding box from the manifest with the query window and drops every file whose box does not intersect it. The Havasu spec calls the test an inclusive projection: ST_Intersects(geom, Q) becomes ST_Intersects(MBR[geom], Q) at the file level, a test that can return false positives and never false negatives. The same pushdown works for rasters with RS_Intersects, RS_Contains, and RS_Within.

Pruning only helps when nearby rows share files. CREATE SPATIAL INDEX FOR <table> USING hilbert(geom, <precision>) rewrites a table sorted by the Hilbert curve index of each geometry, so each file covers a compact area and its bounding box stays small.

Diagram of the five layers behind one Havasu Catalog table: the catalog name wherobots_open_data.overture_maps_foundation.buildings_building, table metadata with schema ID 7, 60 snapshots from 22 April 2025 to 26 September 2026, manifests with per-file bounding boxes, and 5,141 Parquet data files holding 2,533,842,612 rows and 253 GB
The layers behind the Overture buildings table in the Havasu Catalog, from its snapshot and file metadata: 60 snapshots and 5,141 Parquet files holding 2,533,842,612 rows.

Havasu Catalog examples

The catalog groups four kinds of entries: datasets of vector and raster data, RasterFlow computer vision models with sample output tables, solution notebooks, and demo apps.

Entry typeCountExamples
Datasets42Overture buildings, Overture places, USDA NAIP aerial imagery, Sentinel-2 seasonal mosaics, Copernicus DEM GLO-30
RasterFlow models6Fields of the World, Tile2Net, Meta CHM v1, ChesapeakeRSC, SAM3, bring your own model
Solution notebooks9Detecting field boundaries with RasterFlow, Reading STAC data, Spatial Joins
Demo apps6Colorado property risk explorer, California grid wildfire exposure

Dataset sources include the Overture Maps Foundation, the US Census Bureau, OpenStreetMap, ESA Copernicus Sentinel-2, USDA NAIP, and Google and Microsoft building footprints. Four entries show the range:

  • Overture buildings and places. Vector tables with a geometry column and a bbox struct per row, tagged by Overture release. Places carry names, categories, and confidence scores.
  • Copernicus DEM GLO-30. A raster table of elevation tiles, each row a tile with its footprint. GLO-30 is a digital surface model, so its heights include roofs and trees.
  • Sentinel-2 L2A scene index. STAC items as rows: footprint, acquisition time, cloud cover, and links to Cloud Optimized GeoTIFF bands that WherobotsDB reads as out-db rasters.
  • RasterFlow outputs. Fields of the World boundaries, Tile2Net sidewalks, and ChesapeakeRSC roads are outputs of the RasterFlow models, stored as vector tables. RasterFlow runs them over NAIP, Sentinel-2, or a team's own imagery.

Havasu vs GeoParquet, Iceberg v3, and STAC

Four specifications cover spatial data in cloud storage, and each answers a different question. GeoParquet defines a file. Iceberg v3 and Havasu define a table made of many files. STAC defines a catalog of assets.

GeoParquet 1.1Iceberg v3 geo typesHavasu 0.1.0STAC 1.1
Unit describedOne Parquet fileA table of many filesA table of many filesA catalog of assets (scenes, files)
Geometry encodingWKB, or native GeoArrow encodingsWKB, geometry(C) and geography(C, A)WKB, EWKB, WKT, or GeoJSONGeoJSON footprints in JSON records
RasterNoNoYes, in-db and out-dbLinks to raster assets
CRSPROJJSON in file metadata, default OGC:CRS84Column parameter, default OGC:CRS84SRID per valueProjection extension
Spatial statisticsFile bbox and optional per-row bbox covering columnPer-file bounding box in manifestsPer-file bounding box in manifestsItem and collection extents
Transactions and snapshotsNoYesYesNo
Governed byOGC Standards Working GroupApache Software FoundationWherobots (open spec)OGC Community Standard

GeoParquet 1.1, released in June 2024, added a bounding box covering column and GeoArrow-based encodings that let readers use Parquet row-group statistics. The Parquet geospatial definitions added GEOMETRY and GEOGRAPHY logical types in Parquet format 2.11 in March 2025, which Iceberg v3 files use. STAC became an OGC Community Standard in 2025, with the core specification at version 1.1.0. The formats stack: an Iceberg v3 table is a set of Parquet files with native geospatial types.

The Havasu Catalog in Wherobots

This walkthrough runs read-only SQL on wherobots_open_data and follows one place: Lake Havasu City, Arizona, inside the box from 114.38° W to 114.28° W and 34.43° N to 34.55° N.

1. List the databases

SHOW SCHEMAS IN wherobots_open_data

The open data catalog holds 11 databases: aster_gdem, copernicus_dem, foursquare, meta_canopy_height, noaa, overture_maps_foundation, partner_samples, rasterflow_output_samples, sentinel2, spatial_knowledge_graph, and us_census.

2. Count rows and files from snapshot metadata

The .snapshots metadata table returns each commit with a summary map. The latest summary gives a table's row count, file count, and size without reading a single data file.

-- Latest snapshot summary of four open data tables, read from Iceberg metadata only
WITH s AS (
  SELECT 'overture buildings_building' AS tbl, committed_at, summary
  FROM wherobots_open_data.overture_maps_foundation.buildings_building.snapshots
  UNION ALL
  SELECT 'overture places_place', committed_at, summary
  FROM wherobots_open_data.overture_maps_foundation.places_place.snapshots
  UNION ALL
  SELECT 'overture transportation_segment', committed_at, summary
  FROM wherobots_open_data.overture_maps_foundation.transportation_segment.snapshots
  UNION ALL
  SELECT 'us_census tiger_county', committed_at, summary
  FROM wherobots_open_data.us_census.tiger_county.snapshots
)
SELECT tbl,
       COUNT(*)                                            AS snapshots,
       MIN(committed_at)                                   AS first_commit,
       MAX(committed_at)                                   AS last_commit,
       MAX_BY(summary['total-records'], committed_at)      AS total_records,
       MAX_BY(summary['total-data-files'], committed_at)   AS data_files,
       MAX_BY(summary['total-files-size'], committed_at)   AS bytes
FROM s
GROUP BY tbl
ORDER BY tbl
TableSnapshotsRowsData filesSize
Overture buildings_building602,533,842,6125,141253.2 GB
Overture places_place6081,455,4231899.5 GB
Overture transportation_segment60352,054,7101,39368.5 GB
Census tiger_county23,235266.7 MB

The query ran in 4.8 seconds. The buildings table's 60 snapshots run from 22 April 2025 to 26 September 2026 (UTC).

3. Read earlier Overture releases with time travel

SELECT name, type, snapshot_id FROM wherobots_open_data.overture_maps_foundation.places_place.refs ORDER BY type, name returns 31 references: the main branch and 30 Overture release tags, from 2024-07-22.0 to 2026-09-23.1. main points at the same snapshot as 2026-09-23.1, so a query without a version reads the newest release. This query counts the places inside the Lake Havasu City box in ten of those releases:

-- Places inside the Lake Havasu City box in ten Overture releases, read with time travel
SELECT '2024-07-22.0' AS release, COUNT(*) AS places
FROM wherobots_open_data.overture_maps_foundation.places_place VERSION AS OF '2024-07-22.0'
WHERE bbox.xmin <= -114.28 AND bbox.xmax >= -114.38 AND bbox.ymin <= 34.55 AND bbox.ymax >= 34.43
  AND ST_Intersects(geometry, ST_PolygonFromEnvelope(-114.38, 34.43, -114.28, 34.55))
UNION ALL
SELECT '2024-10-23.0' AS release, COUNT(*) AS places
FROM wherobots_open_data.overture_maps_foundation.places_place VERSION AS OF '2024-10-23.0'
WHERE bbox.xmin <= -114.28 AND bbox.xmax >= -114.38 AND bbox.ymin <= 34.55 AND bbox.ymax >= 34.43
  AND ST_Intersects(geometry, ST_PolygonFromEnvelope(-114.38, 34.43, -114.28, 34.55))
UNION ALL
SELECT '2025-01-22.0' AS release, COUNT(*) AS places
FROM wherobots_open_data.overture_maps_foundation.places_place VERSION AS OF '2025-01-22.0'
WHERE bbox.xmin <= -114.28 AND bbox.xmax >= -114.38 AND bbox.ymin <= 34.55 AND bbox.ymax >= 34.43
  AND ST_Intersects(geometry, ST_PolygonFromEnvelope(-114.38, 34.43, -114.28, 34.55))
UNION ALL
SELECT '2025-04-23.0' AS release, COUNT(*) AS places
FROM wherobots_open_data.overture_maps_foundation.places_place VERSION AS OF '2025-04-23.0'
WHERE bbox.xmin <= -114.28 AND bbox.xmax >= -114.38 AND bbox.ymin <= 34.55 AND bbox.ymax >= 34.43
  AND ST_Intersects(geometry, ST_PolygonFromEnvelope(-114.38, 34.43, -114.28, 34.55))
UNION ALL
SELECT '2025-07-23.0' AS release, COUNT(*) AS places
FROM wherobots_open_data.overture_maps_foundation.places_place VERSION AS OF '2025-07-23.0'
WHERE bbox.xmin <= -114.28 AND bbox.xmax >= -114.38 AND bbox.ymin <= 34.55 AND bbox.ymax >= 34.43
  AND ST_Intersects(geometry, ST_PolygonFromEnvelope(-114.38, 34.43, -114.28, 34.55))
UNION ALL
SELECT '2025-10-22.0' AS release, COUNT(*) AS places
FROM wherobots_open_data.overture_maps_foundation.places_place VERSION AS OF '2025-10-22.0'
WHERE bbox.xmin <= -114.28 AND bbox.xmax >= -114.38 AND bbox.ymin <= 34.55 AND bbox.ymax >= 34.43
  AND ST_Intersects(geometry, ST_PolygonFromEnvelope(-114.38, 34.43, -114.28, 34.55))
UNION ALL
SELECT '2026-01-21.0' AS release, COUNT(*) AS places
FROM wherobots_open_data.overture_maps_foundation.places_place VERSION AS OF '2026-01-21.0'
WHERE bbox.xmin <= -114.28 AND bbox.xmax >= -114.38 AND bbox.ymin <= 34.55 AND bbox.ymax >= 34.43
  AND ST_Intersects(geometry, ST_PolygonFromEnvelope(-114.38, 34.43, -114.28, 34.55))
UNION ALL
SELECT '2026-04-15.0' AS release, COUNT(*) AS places
FROM wherobots_open_data.overture_maps_foundation.places_place VERSION AS OF '2026-04-15.0'
WHERE bbox.xmin <= -114.28 AND bbox.xmax >= -114.38 AND bbox.ymin <= 34.55 AND bbox.ymax >= 34.43
  AND ST_Intersects(geometry, ST_PolygonFromEnvelope(-114.38, 34.43, -114.28, 34.55))
UNION ALL
SELECT '2026-07-22.0' AS release, COUNT(*) AS places
FROM wherobots_open_data.overture_maps_foundation.places_place VERSION AS OF '2026-07-22.0'
WHERE bbox.xmin <= -114.28 AND bbox.xmax >= -114.38 AND bbox.ymin <= 34.55 AND bbox.ymax >= 34.43
  AND ST_Intersects(geometry, ST_PolygonFromEnvelope(-114.38, 34.43, -114.28, 34.55))
UNION ALL
SELECT '2026-09-23.1' AS release, COUNT(*) AS places
FROM wherobots_open_data.overture_maps_foundation.places_place VERSION AS OF '2026-09-23.1'
WHERE bbox.xmin <= -114.28 AND bbox.xmax >= -114.38 AND bbox.ymin <= 34.55 AND bbox.ymax >= 34.43
  AND ST_Intersects(geometry, ST_PolygonFromEnvelope(-114.38, 34.43, -114.28, 34.55))
ORDER BY release
Bar chart of Overture places inside the Lake Havasu City box in ten releases read with VERSION AS OF: 2,341 in July 2024, 2,552 in October 2024 and January 2025, 2,664 in April 2025, 2,600 in July 2025, 2,966 in October 2025, 3,015 in January 2026, 3,059 in April 2026, 2,963 in July 2026, and 3,606 in September 2026
One window, ten Overture releases, one query: places in the Lake Havasu City box grew from 2,341 in July 2024 to 3,606 in September 2026.

All ten versions came back in 8.9 seconds. The count rose from 2,341 in the July 2024 release to 3,606 in September 2026, with dips in July 2025 (2,600) and July 2026 (2,963). The same SQL reproduces any release in the list without a copy of the data.

4. Join two catalog tables

This join counts the places inside a building footprint in Lake Havasu City, by category, on the current release.

WITH area AS (SELECT ST_PolygonFromEnvelope(-114.38, 34.43, -114.28, 34.55) AS g)
SELECT p.basic_category, COUNT(DISTINCT p.id) AS places_in_buildings
FROM wherobots_open_data.overture_maps_foundation.places_place p
JOIN wherobots_open_data.overture_maps_foundation.buildings_building b
  ON ST_Contains(b.geometry, p.geometry)
CROSS JOIN area
WHERE ST_Intersects(p.geometry, area.g) AND ST_Intersects(b.geometry, area.g)
  AND p.basic_category IS NOT NULL
GROUP BY p.basic_category
ORDER BY places_in_buildings DESC
LIMIT 5
basic_categoryplaces_in_buildings
home_service181
real_estate_service165
personal_or_beauty_service116
restaurant92
financial_service92

It returned these counts in 10.1 seconds. COUNT(DISTINCT p.id) counts a place once where two footprints overlap.

5. See spatial pruning in the plan and the file list

EXPLAIN FORMATTED shows where the spatial filter goes:

EXPLAIN FORMATTED
SELECT COUNT(*) AS buildings
FROM wherobots_open_data.overture_maps_foundation.buildings_building
WHERE ST_Intersects(geometry, ST_PolygonFromEnvelope(-114.38, 34.43, -114.28, 34.55))

The scan node of the plan reads:

IcebergScan(table=wherobots_open_data.overture_maps_foundation.buildings_building, schemaId=7,
  snapshotId=461598865520318774, branch=null,
  filters=st_intersects(geometry, POLYGON ((-114.38 34.43, -114.38 34.55, -114.28 34.55, -114.28 34.43, -114.38 34.43))), ...)

The ST_Intersects predicate sits inside the Iceberg scan as a pushed filter, so it is applied during file planning, and an exact ST_PreparedIntersects filter runs on the rows that remain. schemaId=7 shows the table schema has changed several times since the table was created, and every snapshot still reads.

The .files metadata table shows how many files a window like this one leaves. Overture tables carry a bbox struct column, and Iceberg keeps its minimum and maximum per file:

-- Data files of the Overture buildings table whose bbox column statistics overlap the Lake Havasu City box
SELECT COUNT(*)          AS data_files,
       SUM(record_count) AS rows_in_table,
       SUM(CASE WHEN readable_metrics.`bbox.xmin`.lower_bound <= -114.28
                 AND readable_metrics.`bbox.xmax`.upper_bound >= -114.38
                 AND readable_metrics.`bbox.ymin`.lower_bound <= 34.55
                 AND readable_metrics.`bbox.ymax`.upper_bound >= 34.43 THEN 1 ELSE 0 END)            AS files_overlapping_box,
       SUM(CASE WHEN readable_metrics.`bbox.xmin`.lower_bound <= -114.28
                 AND readable_metrics.`bbox.xmax`.upper_bound >= -114.38
                 AND readable_metrics.`bbox.ymin`.lower_bound <= 34.55
                 AND readable_metrics.`bbox.ymax`.upper_bound >= 34.43 THEN record_count ELSE 0 END) AS rows_in_overlapping_files
FROM wherobots_open_data.overture_maps_foundation.buildings_building.files

Of 5,141 files holding 2,533,842,612 buildings, 3 files overlap the box, and they hold 1,576,641 rows. The query read only manifests and finished in 4.6 seconds. A second query lists the per-file boxes for the US Southwest, 125° W to 102° W and 30° N to 42° N, and returns 84 files:

-- Per-file bounding boxes of Overture buildings files that overlap the US Southwest (125 W to 102 W, 30 N to 42 N)
SELECT readable_metrics.`bbox.xmin`.lower_bound AS file_xmin,
       readable_metrics.`bbox.xmax`.upper_bound AS file_xmax,
       readable_metrics.`bbox.ymin`.lower_bound AS file_ymin,
       readable_metrics.`bbox.ymax`.upper_bound AS file_ymax,
       record_count,
       (readable_metrics.`bbox.xmin`.lower_bound <= -114.28 AND readable_metrics.`bbox.xmax`.upper_bound >= -114.38
        AND readable_metrics.`bbox.ymin`.lower_bound <= 34.55 AND readable_metrics.`bbox.ymax`.upper_bound >= 34.43) AS overlaps_lake_havasu_box
FROM wherobots_open_data.overture_maps_foundation.buildings_building.files
WHERE readable_metrics.`bbox.xmin`.lower_bound <= -102 AND readable_metrics.`bbox.xmax`.upper_bound >= -125
  AND readable_metrics.`bbox.ymin`.lower_bound <= 42 AND readable_metrics.`bbox.ymax`.upper_bound >= 30
Map of the US Southwest from 125 to 102 degrees west and 30 to 42 degrees north with the bounding boxes of 84 Overture buildings data files drawn as rectangles. Small boxes cluster along the California coast and larger boxes cover the desert. A small red query window at Lake Havasu City overlaps three file boxes, outlined in amber: one covering southern Nevada and western Arizona, and two that extend beyond the map
Per-file bounding boxes of the Overture buildings table over the US Southwest. 84 files overlap the map, and 3 of the table’s 5,141 files overlap the Lake Havasu City window.

File boxes are small where buildings are dense, along the California coast, and large over the desert. One of the three overlapping files spans 117.36° W to 112.43° W, around the lower Colorado River. The other two have continental boxes, one from 179.80° W to 90.10° W and 23.87° N to 79.42° N, so a query almost anywhere in North America opens them.

Research and what changes at scale

Two research lines meet in the Havasu Catalog: table formats that give object storage the guarantees of a database, and spatial data skipping.

From data lakes to lakehouses

Iceberg began at Netflix as a replacement for Hive tables, whose directory listings were slow and inconsistent on Amazon S3 and gave no atomic commits. Ryan Blue presented the design in Introducing Iceberg: Tables designed for object stores at Strata Data New York in September 2018. Iceberg entered the Apache Incubator in November 2018 and became a top-level Apache project in May 2020. Its central idea is to track every data file in metadata, which makes snapshots, atomic commits, and file-level statistics possible.

Michael Armbrust and colleagues at Databricks described the same move for Delta Lake in Delta Lake: High-Performance ACID Table Storage over Cloud Object Stores (PVLDB, 2020): a transaction log over Parquet files provides ACID properties, time travel, and fast metadata operations. Armbrust, Ali Ghodsi, Reynold Xin, and Matei Zaharia then named the pattern in Lakehouse: A New Generation of Open Platforms that Unify Data Warehousing and Advanced Analytics (CIDR 2021): open file formats with a transactional metadata layer, serving SQL analytics and machine learning from one copy of the data. Paras Jain, Peter Kraft, Conor Power, Tathagata Das, Ion Stoica, and Zaharia compared Delta Lake, Hudi, and Iceberg designs and released the LHBench benchmark in Analyzing and Comparing Lakehouse Storage Systems (CIDR 2023).

Spatial data on the lakehouse

None of those formats had spatial types. Jia Yu, Jinxuan Wu, and Mohamed Sarwat introduced distributed spatial processing on Apache Spark with GeoSpark (ACM SIGSPATIAL 2015) at Arizona State University, and Yu, Zongsi Zhang, and Sarwat described its spatial partitioning, indexing, and join design in Spatial data management in Apache Spark: the GeoSpark perspective and beyond (GeoInformatica, 2019). GeoSpark became Apache Sedona, which the Apache Software Foundation announced as a top-level project in February 2023. The creators of Sedona founded Wherobots, and WherobotsDB is fully code compatible with Apache Sedona across all spatial functions. Read the Apache Sedona story.

Storage was the next gap. Wherobots began work on Havasu in 2022 and announced the format in December 2023. In parallel, the GeoParquet project, started in 2022, standardized geometry in single Parquet files and reached 1.0.0 in 2023. In early 2024, the spatial and Iceberg communities began adding geospatial types to Iceberg with the Havasu design as a reference, with contributors from Wherobots, CARTO, Planet, Apple, Databricks, Snowflake, and others. The Parquet pull request for GEOMETRY and GEOGRAPHY drew more than 400 comments and the Iceberg type spec more than 240. In February 2025, Wherobots announced that Iceberg and Parquet support geo types. The Iceberg community voted to adopt the v3 spec in May 2025, and the spec now lists versions 1, 2, and 3 as complete and adopted. Read how Iceberg v3 geospatial types work.

With spatial types native to Iceberg, the Havasu name moved up a layer. The Wherobots Spatial Data Catalog, which held datasets such as Overture Maps, became the Havasu Catalog, and its scope widened to RasterFlow models, notebooks, and apps.

Data skipping and spatial layout

File-level bounding boxes descend from two older ideas. Guido Moerkotte's Small Materialized Aggregates (VLDB 1998) stored the minimum and maximum of each column per block of rows so a scan could skip blocks that cannot match, the idea behind zone maps and Parquet statistics. Antonin Guttman's R-trees: a dynamic index structure for spatial searching (SIGMOD 1984) organized geometries by nested minimum bounding rectangles, the same rectangles a Havasu manifest stores per file.

Layout decides how much skipping works. Ibrahim Kamel and Christos Faloutsos sorted rectangles by the Hilbert value of their centers to build better packed R-trees in Hilbert R-tree: An Improved R-tree using Fractals (VLDB 1994), the ordering behind CREATE SPATIAL INDEX. Ahmed Eldawy, Louai Alarabi, and Mohamed Mokbel compared grid, quadtree, STR, k-d tree, Z-curve, and Hilbert partitioning for distributed range queries and joins in Spatial partitioning techniques in SpatialHadoop (PVLDB, 2015). Liwen Sun, Michael Franklin, Sanjay Krishnan, and Reynold Xin built blocks from the query workload in Fine-grained partitioning for aggressive data skipping (SIGMOD 2014), and Zongheng Yang and colleagues learned layouts with reinforcement learning in Qd-tree: Learning Data Layouts for Big Data Analytics (SIGMOD 2020).

Learned methods now reach spatial data. Varun Pandey and colleagues tested learned indexes for spatial range queries in The Case for Learned Spatial Indexes (AIDB 2020), and Keizo Hori and colleagues used deep reinforcement learning to choose spatial partitions in Learned spatial data partitioning (aiDM 2023), testing on Apache Sedona with run times up to 59.4% lower for distance joins.

What scale changes

At catalog scale the metadata becomes the index. The Overture buildings table holds 2,533,842,612 rows in 5,141 files, and a city query is fast because 3 file boxes overlap it, a result decided in the manifests before any Parquet file opens. The same numbers show the limits. File boxes are only as tight as the layout, and the two continental boxes in the walkthrough get opened by queries across most of North America. History multiplies metadata: the places table's 60 snapshots and 30 release tags keep 30 full versions readable, and every engine that reads the table has to agree on the same bounds, CRS, and geometry encoding.

Open problems

  • One layout, many queries. A Hilbert sort serves range queries well. Workloads that mix spatial windows with time, category, or attribute filters need layouts that serve several predicates.
  • Geography bounds. Iceberg v3 lets a geography bounding box cross the antimeridian, with xmin greater than xmax. Every engine has to read those boxes the same way, or pruning drops valid rows.
  • Metadata growth. Every snapshot and tag keeps its files alive. A catalog that keeps every Overture release trades storage for reproducibility, and snapshot expiry has to respect tags.
  • Rasters in table formats. Iceberg v3 defines vector types only. Havasu's raster type and out-db bands have no upstream equivalent yet, so raster tables stay Havasu-specific.

What is a knowledge lake?

A knowledge lake is a catalog layer over a data lake that indexes data together with the context needed to use it: schemas, sample queries, models, and applications. Wherobots calls the Havasu Catalog the knowledge lake for the physical world. Language models were trained on text, documents, databases, and the internet, and an AI agent asked which buildings sit in a flood zone also needs spatial tables and an index of where they live. The Wherobots MCP server and AI coding tools connect Claude Code, Cursor, and VS Code to the catalog for physical AI and geospatial AI work.

What's next for the Havasu Catalog

  • Native Iceberg v3 tables. Open data tables move to v3 geometry columns automatically, and the docs give a copy-and-switch path for a team's own tables.
  • More engines on the same tables. Snowflake shipped Iceberg v3 geometry and geography types, AWS Glue added v3 support at re:Invent 2025, and Dremio followed with general availability.

Browse the Havasu Catalog, or open the Wherobots developer docs. Continued access after the free trial requires the Professional tier or above.

Read more from Wherobots

Start a free trial and query the Havasu Catalog at login.cloud.wherobots.com.

Frequently asked questions

What is the Havasu Catalog?

The Havasu Catalog is the Wherobots catalog of physical world data, models, and apps. It lists 42 open geospatial datasets, 6 RasterFlow computer vision models, 9 solution notebooks, and 6 demo apps. Each dataset is an Apache Iceberg table with its schema and sample SQL, so a team can query it in Wherobots without a download. Browse the Havasu Catalog.

Is the Havasu Catalog free to use?

The catalog is free to browse at wherobots.com/havasu. Querying its tables in Wherobots is free during the Wherobots free trial, and continued access after the trial requires the Professional tier or above. RasterFlow, the model inference engine behind the catalog’s model entries, is in Public Preview for Professional, Innovation, and Enterprise organizations.

What is the difference between the Havasu table format and Apache Iceberg?

Havasu is the Wherobots spatial table format built on Apache Iceberg. Its specification extends Iceberg format versions 1 and 2 with geometry and raster types, per-file bounding boxes in the manifests, and spatial filter pushdown, and leaves the rest of Iceberg unchanged. Iceberg format version 3 now defines native geometry and geography types, designed with Havasu as a reference. In WherobotsDB, a column declared as geometry(4326) is an Iceberg v3 native column, and Wherobots upgrades its open data tables automatically.

What is a knowledge lake?

A knowledge lake is a catalog layer over a data lake that indexes data together with the context needed to use it: schemas, sample queries, models, and applications. Wherobots calls the Havasu Catalog the knowledge lake for the physical world because it covers data, models, jobs, and apps about real places.

Can I add my own data to the Havasu Catalog?

Yes. A team’s own spatial tables sit in an organization catalog next to the open datasets, or connect through Unity Catalog, AWS Glue, or any Apache Iceberg catalog, so one query can join private and open data.

Which datasets are in the Havasu Catalog?

The catalog holds 42 datasets from sources including the Overture Maps Foundation, the US Census Bureau, OpenStreetMap, ESA Copernicus Sentinel-2, USDA NAIP aerial imagery, Copernicus and ASTER elevation models, and Google and Microsoft building footprints. In SQL, SHOW SCHEMAS IN wherobots_open_data lists the 11 databases of the open data catalog.

Can AI agents use the Havasu Catalog?

Yes. The Wherobots MCP server connects AI coding tools such as Claude Code, Cursor, and VS Code to the catalog, so an agent can find a dataset, read its schema, and write and run the spatial SQL from a plain-language request. Set up the AI coding tools.

What is a spatial data catalog?

A spatial data catalog maps table names to the files that hold them, like any data catalog, and adds what location data needs: geometry and raster column types, a coordinate reference system per column, and the bounding box of each data file, so a query for one area can skip files outside it.

Does the Havasu Catalog support time travel?

Yes. Every table is an Apache Iceberg table with snapshots, and Wherobots keeps each stable Overture Maps release as a tag. VERSION AS OF '2025-05-21.0' reads the Overture places table as that release published it, and the .refs metadata table lists the available tags.

How does spatial filter pushdown work?

When a table is written, the manifest records the bounding box of the geometries in each data file. A query with a spatial filter such as ST_Intersects compares the query window with those boxes first and skips every file whose box does not intersect it. In the Overture buildings table, 3 of 5,141 data files overlap a Lake Havasu City window.

What is the difference between GeoParquet and Iceberg geometry types?

GeoParquet defines how geometry is stored in a single Parquet file, with WKB or GeoArrow encodings and CRS metadata. Iceberg v3 geometry and geography types define spatial columns for a whole table of many Parquet files, with per-file bounding boxes in the manifests, snapshots, and transactions. Iceberg v3 data files use the Parquet GEOMETRY and GEOGRAPHY logical types.