Planetary-scale answers, unlocked.
A Hands-On Guide for Working with Large-Scale Spatial Data. Learn more.
In the world of geospatial data, entity matching and data integration are common challenges. In this blog post, we’ll explore how to use the Overture Maps Foundation GERS IDs within Wherobots to link information about the same physical location across different datasets.
GERS IDs (Global Entity Reference System identifiers) are persistent, unique identifiers for physical places and entities in the physical world. Created by the Overture Maps Foundation, these IDs serve as a universal reference system that allows different datasets to refer to the same physical location reliably.
Attributes of GERS IDs include:
While Overture Maps data comes with GERS IDs built in, many other datasets–open source, commercial data products, and of course internal company datasets–don’t include these identifiers. This presents a challenge: how do you match your existing location data to GERS IDs to enable integration with the broader ecosystem?
At the recent Cloud Native Geospatial Summit, I co-presented with the Overture Maps Foundation team in a workshop session on GERS. My presentation focused on how to take a non-GERSified Point of Interest (POI) dataset and join it to the Overture Places dataset to assign GERS IDs to those POIs that have one. This simple process also allows users to identify those POIs that do not appear in the Overture Places dataset and thus do not have GERS IDs associated with them.
This blog post is a blogified version of that workshop, presented as a tutorial.
If you’d like to follow along with this tutorial, you’ll need a Wherobots Cloud account. You can sign up for a free community edition here, or sign up for a paid plan through the AWS marketplace here.
First, let’s initialize our Wherobots environment with the necessary libraries:
from sedona.spark import * from pyspark.sql import functions as f import os config = (SedonaContext.builder().getOrCreate()) sedona = SedonaContext.create(config)
The heart of our solution is a pair of utility functions that perform GERS ID matching.
The first of these, gersify(), is for finding a matching point and getting the GERS ID for a single point geometry stored as a WKT or well-known-text format. It takes as input a WKT point geometry and a search string that relates to the name of the POI.
gersify(),
This is for if you need to process a one-off. However, if you have a dataframe with many points (say, 16,000), putting this function into a loop would be incredibly slow. That’s why we have the next function.
The second function, gersify_dataframe(), is for doing this same GERS matching operation on a larger scale. It takes as input a dataframe of points and a search string that relates to those points (i.e. “park” or “stadium”).
gersify_dataframe(),
Let’s examine the code for these:
def gersify(point_wkt, search_param): """ Uses a point and a search string to find the closest matching GERS ID. Returns a DataFrame with the GERS ID and other attributes from Overture Maps. """ # Create a point from WKT query = f""" WITH point AS ( SELECT ST_GeomFromWKT('{point_wkt}') as point ) SELECT p.id as gers_id, p.geometry as OMF_geom, p.names.primary, p.categories.main as category, p.websites.primary as website, p.phones.primary as phone, p.geometry FROM point, places p WHERE ST_DWithin(p.geometry, point, 500, true) ORDER BY ST_Distance(p.geometry, point) ASC """ return_df = sedona.sql(query).withColumn("distance_from_point", f.expr("ST_DistanceSpheroid(geometry, point)")).cache().where(f"names.primary like '%{search_param}%'") return return_df def gersify_dataframe(df, search_param): """ Matches and add GERS ID to any dataset. Returns all rows that have a GERS ID. """ # Register the input dataframe as a temporary view df.createOrReplaceTempView("_temp_df") # For each row in the dataframe, find the closest matching GERS ID inter_query = """ WITH points AS ( SELECT id, geometry FROM _temp_df ) SELECT p.id as gers_id, p.geometry as OMF_geom, df.*, ST_DistanceSpheroid(p.geometry, df.geometry) as distance_from_point FROM points df, places p WHERE ST_DWithin(p.geometry, df.geometry, 500, true) AND p.names.primary LIKE '%Starbucks%' QUALIFY ROW_NUMBER() OVER (PARTITION BY df.id ORDER BY ST_DistanceSpheroid(p.geometry, df.geometry)) = 1 """ # Execute the query to find matches inter_df = sedona.sql(inter_query) inter_df.createOrReplaceTempView("inter_df") # Find rows that didn't match and include them in the result remaining_rows = """ SELECT NULL as gers_id, NULL AS OMF_geom, df.*, NULL as distance_from_point FROM _temp_df df LEFT ANTI JOIN inter_df i_df ON df.id = i_df.id """ return_df = sedona.sql(remaining_rows) # Combine matched and unmatched rows final_df = inter_df.union(return_df) return final_df
For this example, we’ll work with a dataset of Starbucks locations across the United States:
df.count() # 16820 locations df.show()
Let’s take a look at our dataset. Here are the first few rows of the result:
Before enrichment, let’s visualize our base dataset to get an idea of what we are working with.
map = SedonaKepler.create_map(df, "Original Location Dataset") map
Now for the exciting part, let’s enrich our dataset with GERS IDs:
# NOTE: the string search is case sensitive df_gersified = gersify_dataframe(df, "Starbucks")
Let’s examine the results:
df_gersified.show(20, False)
Let’s visualize both datasets together to see what points matched and what points did not match to anything.
SedonaKepler.add_df(map, df_gersified.drop("geometry"), "OMF enrichment locations") map
By enriching our dataset with GERS IDs, we’ve unlocked several powerful capabilities:
GERS IDs represent a powerful tool for geospatial data integration, and Wherobots makes it easy to incorporate them into your workflows. With the functions demonstrated in this blog post, you can enrich any location dataset with GERS IDs, enabling integration with the broader geospatial data ecosystem.
Want to try it yourself? Sign up for a Wherobots account and explore the full capabilities of the platform for your spatial data processing needs.
Key takeaways
Global Entity Reference System identifiers from the Overture Maps Foundation. They are persistent, unique IDs for physical places, designed so different datasets can refer to the same real-world entity.
The tutorial gersify_dataframe() function spatially joins each input point to Overture Places within 500 meters, filters on a case-sensitive name search (e.g. Starbucks), and keeps the nearest match. Unmatched rows are returned with a null gers_id.
The sample dataset has 16,820 U.S. Starbucks locations. Enriching more than 16,000 locations finished in under a minute on a Medium Wherobots cluster.
A Wherobots Cloud account (free Community edition or a paid plan via AWS Marketplace) and the notebook/SQL in the post. Overture Places is available in the Wherobots catalog for the join.
RasterFlow is now available in Public Preview
RasterFlow makes planetary-scale earth intelligence workflows easy and costs predictable. We are excited to announce that RasterFlow is now in Public Preview, opening up the power of planetary scale Earth Intelligence to all Wherobots Professional Edition customers! RasterFlow let’s you solve complex monitoring challenges with vision-language models or tailored models for specific use cases, without […]
Wherobots for QGIS: Cloud-Scale Spatial Data, from Wherobots Labs
Wherobots for QGIS is a new plugin that connects QGIS to Wherobots Cloud. From a panel inside QGIS, an analyst can run spatial SQL against WherobotsDB, load the results as a map layer, push a local layer up to a Wherobots Iceberg table, and pull raster data into the canvas. Why Wherobots for QGIS: Eliminating […]
How Wherobots builds with NVIDIA to let AI see the physical world
The world and what happens in it is digitized by petabytes of raw and derivative spatial datasets of various data types and scales, and the potential for applying AI to it is immense. But the AI models and agents we use every day need connectivity to tools that turn this data into usable insights and relationships. Wherobots gives AI the ability operate on and understand raw physical-world data, and NVIDIA GPUs are core to it. Architecturally here’s how this works at a high level.
El Niño 2026, atmospheric rivers, and California’s burn scars: one SQL engine, two data models
7,713 mapped building footprints sit inside or within 500 m (1,640 ft) of the 42 Eaton basins USGS rates high hazard for debris flows, and 388 km (241 mi) of mapped road and path cross the low ground below the scar. WherobotsDB finds both in SQL: a join to the USGS basins for the buildings, and a raster vector join over elevation for the roads.
share this article
Awesome that you’d like to share our articles. Where would you like to share it to: