WherobotsAI Raster Inference is GA with Support for Bring Your Own Model Posted on December 17, 2024October 3, 2026 by Ben Pruden Editor’s note: Raster inference is now part of RasterFlow, the Wherobots engine for running computer vision models on satellite and aerial imagery. Introduction UPDATE: Raster inference is included in Wherobots RasterFlow. See Wherobots Get Started with RasterFlow – Wherobots for the most up-to-date workflows. We are excited to announce that WherobotsAI Raster Inference is now generally available! Raster Inference is a serverless, planetary-scale computer vision solution that enables data teams to extract meaningful insights from aerial imagery (raster data) sources, such as satellites or drones, and puts these insights at the fingertips of data scientists and developers. During the preview period, customers analyzed raster data by running inference with a limited set of Wherobots-hosted open-source computer vision models. Now, you can bring your own model into Raster Inference, offload inference pipeline management effort, and apply this capability to a much broader set of use cases. We have also made significant enhancements to our compute service to accelerate inference performance. Typical Use Cases Data teams use WherobotsAI Raster Inference to identify information in complex, large-scale overhead imagery data. Some common use cases include: Agriculture: Satellite and drone imagery are critical for monitoring land use, predicting crop yields, and improving sustainable farming practices. Environment and Conservation: Raster imagery is essential for monitoring ecosystem and biodiversity changes, such as tracking glacier melting, sea level rise, temperature fluctuations, and environmental degradation (e.g., oil spills, deforestation). Energy: Renewable energy developers analyze satellite data to assess land suitability, solar radiation, and wind patterns, to determine optimal locations for renewable energy projects. Insurance: Insurers use raster data to calculate risk assessments based on environmental and local factors to conduct and improve damage assessments at scale. Map Creation & Maintinence: Global digital, high fidelity map producers and maintainers extract features such as buildings, road networks, landcover, shipping lanes, etc., and their respective changes in order to maintain truthful connection between global map data products and the real world. Traditional challenges with computer vision pipelines We’ve met with many businesses struggling to get critical insights from raster data. Some are manually sifting through imagery. This method doesn’t scale, it’s expensive, error prone, and time intensive. Others are utilizing complex computer vision solutions that are not designed for overhead raster imagery. These computer visions solutions: Take time to build, and require significant management to efficiently load, store, and process large raster datasets. Are difficult to scale to accommodate increasing workload sizes. Are fragile with multiple components involved and challenges maintaining compatibility with an evolving modeling stack. Require effort to experiment, test, and integrate new models and inference runs. Benefits of WherobotsAI Raster Inference With WherobotsAI Raster Inference, you can: Use data pipelines that are ready for small to planetary-scale raster data. Deploy an on-demand solution in seconds that scales to meet workload needs without having to manage infrastructure. Easily import your model or utilize any model hosted by Wherobots. Experiment and generate critical insights faster. In the following sections, we’ll provide a brief overview of Wherobots-hosted models, how to bring your own model, recent performance improvements, and how to use this feature effectively. Choosing a Model for Raster Inference Wherobots-Hosted Models Wherobots-hosted models are precompiled and optimized for raster inference, enabling them to scale effortlessly and execute inference on large datasets. Below is a brief overview of the initial set of Wherobots-hosted models. We plan to host additional models based on customer feedback. For more information on these models and their performance metrics, see our docs page here. Below is a code snippet on how to use landcover-eurosat-sentinel2 in WherobotsAI Raster Inference: # set the model name model_id = "landcover-eurosat-sentinel2" # call the model in raster inference df_predictions = df_raster_input.withColumn("preds", rs_classify(model_id, "outdb_raster")) The STAC Machine Learning Model Extension Specification The model import function of Raster Inference is built on the STAC Machine Learning Model (MLM) Extension Specification, an open community standard for model sharing we co-developed with CRIM and other collaborators. The MLM specification enables model portability, making it easier to use models across teams and compute platforms. Before the MLM specification, data scientists and modelers sharing geospatial computer vision models often had to adapt existing standards, such as HuggingFace model cards, to store relevant information. This involved repurposing the cards to specify details like required raster bands, necessary data preprocessing, and post-processing functions. Without a community standard for organizing this information, these model cards were often inconsistent and documented in varying formats, making model sharing and reproducibility across different compute platforms cumbersome. The STAC MLM extension introduces a community standard designed to simplify storing and sharing geospatial computer vision models. It achieves this by providing a comprehensive schema to: Describe critical geospatial model attributes, such as geolocation and temporal range. Include key model inference reproducibility details, such as required bands, model artifact locations, and pre- and post- processing steps. Enable model collections to be searched alongside associated spatiotemporal datasets. The MLM specification has already been adopted in key modeling efforts at Terradue and is proudly supported by Radiant Earth. We’ll share more in a series of blog posts and a panel discussion with our collaborators on the MLM in late January — register here on the interest form to receive an invite when the date is finalized. Bring Your own Model The MLM specification enables users to quickly and easily use many geospatial, deep learning-based computer vision models with Raster Inference. We leverage a model’s MLM specification to specify and integrate the required data preprocessing, model, and post processing into the larger raster inference pipeline. To bring your own model to Raster Inference: Fill out our MLM form for your model. This will create a MLM formatted JSON file (MLM JSON) describing your model. Download the file. Upload the MLM JSON file from Step 1 to your AWS S3 bucket. Copy and save the AWS S3 URI to your MLM JSON. During runtime, use your model for inference by calling the raster inference function with your MLM’s S3 URI link. You can find a full walkthrough in our documentation on using the MLM form to bring your own model. Performance Improvements In addition to enabling bring your own model, we’ve accelerated asynchronous data loading in the inference engine to boost performance. Below, you can see how performance has evolved with experiments conducted using WherobotsAI Raster Inference and a Tiny GPU Runtime. Example: Raster Inference with Bring Your own Model For a full tutorial example on how to bring your own model to segment solar farms in Sentinel-2 imagery, see our documentation here. Example: Run Raster Inference with a Wherobots-hosted model We’ll walk through an example on how to identify solar infrastructure in raster imagery using a Wherobots-hosted model. In our example, we will be using the model: solar-satlas-sentinel2 The full Python notebook file can be found on our GitHub. Set up the WherobotsDB context import warnings warnings.filterwarnings('ignore') from wherobots.inference.data.io import read_raster_table from sedona.spark import SedonaContext from pyspark.sql.functions import expr from wherobots.inference.engine.register import create_semantic_segmentation_udfs from pyspark.sql.functions import col config = SedonaContext.builder().appName('segmentation-batch-inference') .getOrCreate() sedona = SedonaContext.create(config) Load Satellite Imagery tif_folder_path = "s3a://wherobots-benchmark-prod/data/ml/satlas" files_df = read_raster_table(tif_folder_path, sedona, limit=400) df_raster_input = files_df.withColumn( "outdb_raster", expr("RS_FromPath(path)") ) df_raster_input.cache().count() df_raster_input.show(truncate=False) df_raster_input.createOrReplaceTempView("df_raster_input") Run WherobotsAI Raster Inference Specify a Wherobots-hosted model to run inference model_id = "solar-satlas-sentinel2" You can run WherobotsAI Raster Inference using either the Wherobot’s SQL API or Python API. Using the SQL API predictions_df = sedona.sql(""" SELECT outdb_raster, segment_result.* FROM ( SELECT outdb_raster, RS_SEGMENT('{model_id}', outdb_raster) AS segment_result FROM df_raster_input ) AS segment_fields """) predictions_df.cache().count() predictions_df.show() predictions_df.createOrReplaceTempView("predictions") Using the Python API rs_segment = create_semantic_segmentation_udfs(batch_size = 10, sedona=sedona) df = df_raster_input.withColumn("segment_result", rs_segment(model_id, col("outdb_raster"))).select( "outdb_raster", col("segment_result.confidence_array").alias("confidence_array"), col("segment_result.class_map").alias("class_map") ) df.show(3) Extract predicted geometries (continued from step 4 using the SQL API) df_multipolys = sedona.sql(""" WITH t AS ( SELECT RS_SEGMENT_TO_GEOMS(outdb_raster, confidence_array, array(1), class_map, 0.65) result FROM predictions ) SELECT result.* FROM t """) df_multipolys.cache().count() df_multipolys.show() df_multipolys.createOrReplaceTempView("multipolygon_predictions") df_merged_predictions = sedona.sql(""" SELECT element_at(class_name, 1) AS class_name, cast(element_at(average_pixel_confidence_score, 1) AS double) AS average_pixel_confidence_score, ST_Collect(geometry) AS merged_geom FROM multipolygon_predictions """) df_filtered_predictions = df_merged_predictions.filter("ST_IsEmpty(merged_geom) = False") df_filtered_predictions.cache().count() df_filtered_predictions.show() Visualize results from sedona.maps.SedonaKepler import SedonaKepler config = { 'version': 'v1', 'config': { 'mapStyle': { 'styleType': 'dark', 'topLayerGroups': {}, 'visibleLayerGroups': {}, 'mapStyles': {} }, } } map = SedonaKepler.create_map(config=config) SedonaKepler.add_df(map, df=df_filtered_predictions, name="Solar Farm Detections") map Get started with WherobotsAI Raster Inference Professional Edition Users If you’re a Wherobots Professional Edition or Enterprise user, you have access to all capabilities of WherobotsAI Raster Inference! If you have access to GPU runtimes, sign in to your account now to launch a Wherobots Notebook and explore the feature. If you don’t yet have access, request it today and start using Raster Inference as soon as tomorrow. Explore how easy it is to bring your own model using our guided example on GitHub or within a Wherobots notebook instance. Start integrating WherobotsAI Raster Inference into your workflow today! Community Edition Users Although Raster Inference is not available in Wherobots Community Edition, we are currently offering a free trial for the Wherobots Professional Edition. You can either sign up through AWS Marketplace or upgrade your account to get started for free and integrate Raster Inference into your workflow today. What’s next We’re eager to hear about the models you’d like us to support and any features you’d like to see added. For product feedback, feel free to email us at feedback@wherobots.com (no request is too small). To see how others are using Raster Inference, ask questions, and share your own experiences, join the Wherobots community. We look forward to your ideas and creations! Missed our panel discussion with our collaborators CRIM, Terradue and Radiant Earth on the MLM STAC extension? Watch the recording below. Get Started with Wherobots Try Now Key takeawaysWherobotsAI Raster Inference is generally available as a serverless, planetary-scale computer vision service that extracts features from satellite and drone raster imagery for data scientists and developers.Bring-your-own-model is now supported through the STAC Machine Learning Model (MLM) Extension Specification, co-developed with CRIM and other collaborators, adopted at Terradue, and supported by Radiant Earth.Wherobots-hosted models are precompiled for scale; the post walks through landcover-eurosat-sentinel2 classification and solar-satlas-sentinel2 semantic segmentation, including converting predictions to geometries at a 0.65 confidence threshold.The GA release also accelerated asynchronous data loading in the inference engine, with experiments shown for a Tiny GPU Runtime (chart on the page; no numeric speedups are stated in the text).Raster Inference is included for Professional and Enterprise users with GPU runtimes, not Community Edition. Community users can start a Professional trial via AWS Marketplace or an account upgrade.
Unlock Satellite Imagery Insights with WherobotsAI Raster Inference Posted on June 24, 2024October 3, 2026 by Ben Pruden Editor’s note: Raster inference is now part of RasterFlow, the Wherobots engine for running computer vision models on satellite and aerial imagery. UPDATE: Raster inference is included in Wherobots RasterFlow. See Wherobots Get Started with RasterFlow – Wherobots for the most up-to-date workflows. Recently we introduced WherobotsAI Raster Inference to unlock analytics on satellite and aerial imagery using SQL or Python. Raster Inference simplifies extracting insights from satellite and aerial imagery using SQL or Python, and is powered by open-source machine learning models. Below we’ll dig into the popular computer vision tasks that Raster Inference supports, describe how it works, and how you can use it to run batch inference to find and map electricity infrastructure. Watch the live demo of these capabilities here. The Power of Machine Learning with Satellite Imagery Petabytes of satellite imagery are generated each day all over the world in a dizzying number of sensor types and image resolutions. The applications for satellite imagery and other remote sensing data sources are broad and diverse. For example, satellites with consistent, continuous orbits are ideal for monitoring forest carbon stocks to validate carbon credits or estimating agricultural yields. However, this data has been inaccessible for most analysts and even seasoned ML practitioners because insight extraction required specialized skills. We’ve done the work to make insight extraction simple and accessible to more people. Raster Inference abstracts the complexity and scales to support planetary-scale imagery datasets, so you don’t need ML expertise to derive insights. In this blog, we explore the key features that make Raster Inference effective for land cover classification, solar farm mapping, and marine infrastructure detection. And, in the near future, you will be able to use Raster Inference with your own models! Introduction to Popular and Supported Machine Learning Tasks Raster Inference supports the three most common kinds of computer vision models that are applied to imagery: classification, object detection, and semantic segmentation. Instance segmentation (combines object localization and semantic segmentation) is another common type of model which is not currently supported, but let us know if you need by contacting us and we can add it to the roadmap. Computer Vision Detection Categories from Lin et al. Microsoft COCO: Common Objects in Context The figure above illustrates these tasks. Image classification is when an image is assigned one or more text labels. In image (a), the scene is assigned the labels “person”, “sheep”, and “dog”. Image (b) is an example of object localization (or object detection). Object localization creates bounding boxes around objects of interest and assigns labels. In this image, five sheep are localized separately along with one human and one dog. Finally, semantic segmentation is when each pixel is given a category label, as shown in image (c). Here we can see all the pixels belonging to sheep are labeled blue, the dog is labeled red, and the person is labeled teal. While these examples highlight detection tasks on regular imagery, these computer vision models can be applied to raster formatted imagery. Raster data formats are the most common data formats for satellite and aerial imagery. When objects of interest in raster imagery are localized, their bounding boxes can be georeferenced, which means that each pixel is localized to spatial coordinates, such as latitude and longitude. Therefore, georeferencing is object localization suited for spatial analytics. The example above shows various applications of object detection for localizing and classifying features in high resolution satellite and aerial imagery. This example comes from DOTA, a 15-class dataset of different objects in RGB and grayscale satellite imagery. Public datasets like DOTA are used to develop and benchmark machine learning models. Not only are there many publicly available object detection models, but also there are many semantic segmentation models. Sourced from “A Scale-Aware Masked Autoencoder for Multi-scale Geospatial Representation Learning”. Not every machine learning model should be treated equally, and each will have their own tradeoffs. You can see the difference between the ground truth image (human annotated buildings representing the real world) and segmentation results across two models (Scale-MAE and Vanilla MAE). These results are derived from the same image at two different resolutions (referred to as GSD, or Ground Sampling Distance). Scale-MAE is a model developed to handle detection tasks at various resolutions with different sensor inputs. It uses a similar MAE model architecture as the Vanilla MAE, but is trained specifically for detection tasks on overhead imagery that span different resolutions. The Vanilla MAE is not trained to handle varying resolutions in overhead imagery. It’s performance suffers in the top row and especially the bottom row, where resolution is coarser, as seen by the mismatch between Vanilla MAE and the ground truth image where many pixels are incorrectly classified. Satellite Analytics Before Raster Inference Without Raster Inference, typically a team who is looking to extract insights from overhead imagery using ML would need to: Deploy a distributed runtime to scale out workloads such as data loading, preprocessing, and inference. Develop functionality to operate on raster metadata to easily filter it by location to run inference workloads on specific areas of interest. Optimize models to run performantly on GPUs, which can involve complex rewrites of the underlying model prediction logic. Create and manage data preprocessing pipelines to normalize, resize, and collate raster imagery into the correct data type and size required by the model. Develop the logic to run data loading, preprocessing, and model inference efficiently at scale. Raster Inference and its SQL and Python APIs abstract this complexity so you and your team can easily perform inference on massive raster datasets. Raster Inference APIs for SQL and Python Raster Inference offers APIs in both SQL and Python to run inference tasks. These APIs are designed to be easy to use, even if you’re not a machine learning expert. RS_CLASSIFY can be used for scene classification, RS_BBOXES_DETECT for object detection, and RS_SEGMENT for semantic segmentation. Each function produces tabular results which can be georeferenced either for the scene, object, or segmentation depending on the function. The records can be joined or visualized with other data (geospatial or traditional) to curate enriched datasets and insights. Here are SQL and Python examples for RS_Segment. RS_SEGMENT('{model_id}', outdb_raster) AS segment_result df = df_raster_input.withColumn("segment_result", rs_segment(model_id, col("outdb_raster"))) Example: Mapping Electricity Infrastructure Imagine you want to optimize the location of new EV charging stations, but you want to target locations based on the availability of green energy sources, such as local solar farms. You can use Raster Inference to detect and locate solar farms and cross-reference these locations with internal data or other vector geometries that captures demand for EV charging. This use case will be demonstrated in our upcoming release webinar on July 10th. Let’s walk through how to use Raster Inference for this use case. First, we run predictions on rasters to find solar farms. The following code block that calls RS_SEGMENT shows how easy this is. CREATE OR REPLACE TEMP VIEW segment_fields AS ( SELECT outdb_raster, RS_SEGMENT('{model_id}', outdb_raster) AS segment_result FROM az_high_demand_with_scene ) The confidence_array column produced from RS_SEGMENT can be assigned the same geospatial coordinates as the raster input and converted to a vector that can be spatially joined and processed with WherobotsDB using RS_SEGMENT_TO_GEOMS. We select a confidence threshold of .65 so that we only georeference high confidence detections. WITH t AS ( SELECT RS_SEGMENT_TO_GEOMS(outdb_raster, confidence_array, array(1), class_map, 0.65) result FROM predictions_df ) SELECT result.* FROM t +----------+--------------------+--------------------+ | class|avg_confidence_score| geometry| +----------+--------------------+--------------------+ |Solar Farm| 0.7205783606825462|MULTIPOLYGON (((-...| |Solar Farm| 0.7273308333550763|MULTIPOLYGON (((-...| |Solar Farm| 0.7301468510823231|MULTIPOLYGON (((-...| |Solar Farm| 0.7180177244988899|MULTIPOLYGON (((-...| |Solar Farm| 0.728077805771141|MULTIPOLYGON (((-...| |Solar Farm| 0.7264981572898|MULTIPOLYGON (((-...| |Solar Farm| 0.7044100126912517|MULTIPOLYGON (((-...| |Solar Farm| 0.7137283466756343|MULTIPOLYGON (((-...| +----------+--------------------+--------------------+ This allows us to integrate the vectorized model predictions with other spatial datasets and easily visualize the results with SedonaKepler. Here Raster Inference runs on a 85 GiB dataset with 2,200 raster scenes for Arizona. Using a Sedona (tiny) runtime, Raster Inference completed in 430 seconds, predicting solar farms for all low cloud cover satellite images for the state of Arizona for the month of October. If we scale up our runtime to a San Francisco (small) runtime, the inference speed nearly doubles. In general, average bytes processed per second by Wherobots increases as datasets scale in size because startup costs are amortized over time. Processing speed also increases as runtimes scale in size. Inference time (seconds)Runtime Size430Sedona246San Francisco We use predictions from the output of Raster Inference to derive insights about which zip codes have the most solar farms, as shown below. This statement joins predicted solar farms with zip codes by location, then ranks zip codes by the pre-computed solar farm area within each zip code. We skipped this step for brevity but you can see it and others in the notebook example. az_solar_zip_codes = sedona.sql(""" SELECT solar_area, any_value(az_zta5.geometry) AS geometry, ZCTA5CE10 FROM predictions_polys JOIN az_zta5 WHERE ST_Intersects(az_zta5.geometry, predictions_polys.geometry) GROUP BY ZCTA5CE10 ORDER BY solar_area DESC """) These predictions are made possible by SATLAS, a family of machine learning models released with Apache 2.0 licensing from Allen AI. The solar model demonstrated above was derived from the SATLAS foundational model. This foundational model can be used as a building block to create models to address specific detection challenges like solar farm detection. Additionally, there are many other open source machine learning models available for deriving insights from satellite imagery, many of which are provided by the TorchGeo project. We are just beginning to explore what these models can achieve for planetary-scale monitoring. If you have a specific model you would like to see made available, please contact us to let us know. For detailed instructions on using Raster Inference, please refer to our example Jupyter notebooks in the documentation. Here are some links to get you started:https://docs.wherobots.com/latest/tutorials/wherobotsai/wherobots-inference/segmentation/ https://docs.wherobots.com/latest/api/wherobots-inference/pythondoc/inference/sql_functions Getting Started Getting started with WherobotsAI Raster Inference is easy. We’ve provided three models in Wherobots Cloud that can be used with our GPU optimized runtimes. Sign up for Wherobots Pro via the AWS Marketplace and get up to $400 in free credits. You can test out the Pro features, including Raster Inference. Stay tuned for updates on improvements to Raster Inference that will make it possible to run more models, including your own custom models. We’re excited to hear what models you’d like us to support, or the integrations you need to make running your own models even easier with Raster Inference. We can’t wait for your feedback and to see what you’ll create! Start building with Wherobots Get Started Key takeawaysWherobotsAI Raster Inference (later folded into RasterFlow) runs open-source computer vision models on satellite and aerial rasters from SQL or Python, without requiring users to stand up GPU pipelines.Supported tasks are classification (RS_CLASSIFY), object detection (RS_BBOXES_DETECT), and semantic segmentation (RS_SEGMENT). Instance segmentation is not supported in this post.An Arizona solar-farm example processed an 85 GiB dataset with 2,200 scenes: 430 seconds on a Sedona (tiny) runtime and 246 seconds on a San Francisco (small) runtime, then vectorized detections at 0.65 confidence with RS_SEGMENT_TO_GEOMS.The solar model comes from SATLAS (Allen AI, Apache 2.0). The post also points at TorchGeo models and says three models are hosted in Wherobots Cloud on GPU-optimized runtimes.Professional Edition via AWS Marketplace includes up to $400 in free credits. The authors say custom/BYO models are on the roadmap after this preview.
Introducing WherobotsAI for planetary inference, and capabilities that modernize spatial intelligence at scale Posted on June 5, 2024October 3, 2026 by Ben Pruden Editor’s note: Raster inference is now part of RasterFlow, the Wherobots engine for running computer vision models on satellite and aerial imagery. UPDATE: Raster inference is included in Wherobots RasterFlow. See Wherobots Get Started with RasterFlow – Wherobots for the most up-to-date workflows. We are excited to announce a preview of WherobotsAI, our new suite of AI and ML powered capabilities that unlock spatial intelligence in satellite imagery and GPS location data. Additionally, we are bringing the high-performance of WherobotsDB to your favorite data applications with a Spatial SQL API that integrates WherobotsDB with more interfaces including Apache Airflow for Spatial ETL. Finally, we’re introducing the most scalable vector tile generator on earth to make it easier for teams to produce engaging and interactive map applications. All of these new features are capable of operating on planetary-scale data. Watch the walkthrough of this release here. Wherobots Mission and Vision Before we dive into this release, we think it’s important to understand how these capabilities fit into our mission, our product principles, and vision for the Spatial Intelligence Cloud so you can see where we are headed. Our MissionThese new capabilities are core to Wherobots’ mission, which is to unlock spatial intelligence of earth, society, and business, at a planetary scale. We will do this by making it extremely easy to utilize data and AI technology purpose-built for creating spatial intelligence that’s cloud-native and compatible with modern open data architectures. Our Product Principles We’re building the spatial intelligence platform for modern organizations. Every organization with a mission directly linked to the performance of tangible assets, goods and services, or data products about what’s happening in the physical world, will need a spatial intelligence platform to be competitive, sustainable, and climate adaptive. It delivers intelligence for the greater good. Teams and their organizations want to analyze their worlds to create a net positive impact for business, society, and the earth. It’s purpose-built yet simple. Spatial intelligence won’t scale through in-house ‘spatial experts’, or through general purpose architectures that are not optimized for spatial workloads or development experiences. It’s efficient at any scale. Maximal performance, scale, and cost efficiency can only be achieved through a cloud-native, serverless solution. It creates intelligence with AI. Every organization will need AI alongside modern analytics to create spatial intelligence. It’s open by default. Pace of innovation depends on choice. Organizations that adopt cloud-native, open source compatible, and modern open data architectures will innovate faster because they have more choices in the solutions they can use. Our VisionWe exist because creating spatial intelligence at-scale is hard. Our contributions to Apache Sedona, leadership in the open geospatial domain, and investments in Wherobots Cloud have, and will make it easier. Users of Apache Sedona, Wherobots customers, and ultimately any AI application will be enabled to support better decisions about our physical and virtual worlds. They will be able to create solutions to improve these worlds that were otherwise infeasible or too costly to build. And the solutions developed will have a positive impact on society, business, and earth — at a planetary scale. Introducing WherobotsAI There are petabytes of satellite or aerial imagery produced every day. Yet for most analysts, scientists, and developers, these datasets are analytically inaccessible outside of the naked eye. As a result most organizations still rely on humans and their eyes, to analyze satellite or other forms of aerial imagery. Wherobots can already perform analytics of overhead imagery (also known as raster data) and geospatial objects (known as vector data) simultaneously at scale. But organizations also want to use modern AI and ML technologies to streamline and scale otherwise visual, single threaded tasks like object detection, classification, and segmentation from overhead imagery. Like satellite imagery that is generally hard to analyze, businesses also find it hard to analyze GPS data in their applications because it’s too noisy; points don’t always correspond to the actual path taken. Teams need an easy solution for snapping noisy GPS data to road or other segment types, at any scale. Today we are announcing WherobotsAI which offers fully managed AI and machine learning capabilities that accelerate the development of spatial insights, for anyone familiar with SQL or Python. WherobotsAI capabilities include: [new] Raster Inference (preview): A first of its kind, Raster Inference unlocks the analytical potential of satellite or aerial imagery at a planetary scale, by integrating AI models with WherobotsDB to make it extremely easy to detect, classify, and segment features of interest in satellite and aerial images. You can see how easy it is to detect and georeference solar farms here, with just a few lines of SQL: SELECT outdb_raster, RS_SEGMENT(‘solar-satlas-sentinel2’, outdb_raster) AS solar_farm_result FROM df_raster_input These georeferenced predictions can be queried with WherobotsDB and can be interactively explored in a Wherobots notebook. Below is an example of detection of solar panels in SedonaKepler. The models and AI infrastructure powering Raster Inference are fully managed, which means there’s nothing to set up or configure. Today, you can use Raster Inference to detect, segment, and classify solar farms, land cover, and marine infrastructure from terabyte-scale Sentinel-2 true color and multispectral imagery datasets in under half an hour, on our GPU runtimes available in the Wherobots Professional Edition. Soon we will be making the inference metadata for the models public, so if your own models meet this standard, they are supported by Raster Inference. These models and datasets are just the starting point for WherobotsAI. We are looking forward to hearing from you to help us define the roadmap for what we should build support for next. Map Matching: If you need to analyze trips at scale, but struggle to wrangle noisy GPS data, Map Matching is capable of turning billions of noisy GPS pings into signal, by snapping shared points to road or other vector segments. Teams are using Map Matching to process hundreds of millions of vehicle trips per hour. This speed surpasses any current commercial solutions, all for a cost of just a few hundred dollars. Here’s an example of what WherobotsAI Map Matching does to improve the quality of your trip segments. Red and yellow line segments were created from raw, noisy GPS data. Green represents Map Matched segments. Visit the user documentation to learn more and get started with WherobotsAI. A Spatial SQL API for WherobotsDB WherobotsDB, our serverless, highly efficient compute engine compatible with Apache Sedona is up to 60x more performant for spatial joins than popular general purpose big data engines and warehouses, and up to 20x faster than Apache Sedona on its own. It will remain the most performant, earth-friendly solution for your spatial workloads at any scale. Until today, teams had two options for harnessing WherobotsDB: they could write and run queries in Wherobots managed notebooks, or run spatial ETL pipelines using the Wherobots jobs interface. Today, we’re enabling you to bring the utility of WherobotsDB to more interfaces with the new Spatial SQL API. Using this API, teams can remotely execute Spatial SQL queries using a remote SQL editor, build first-party applications using our client SDKs in Python (WherobotsDB API driver) and Java (Wherobots JDBC driver), or orchestrate spatial ETL pipelines using a Wherobots Apache Airflow provider. Run spatial queries with popular SQL IDEs The following is an example of how to integrate Harlequin, a popular SQL IDE with WherobotsDB. You’ll need a Wherobots API key to get started with Harlequin (or any remote client). API keys allow you to authenticate with Wherobots Cloud for programmatic access to Wherobots APIs and services. API keys can be created following a few steps in our user documentation. We will query WherobotsDB using Harlequin in the Airflow example later in this blog. $ pip install harlequin-wherobots $ harlequin -a wherobots --api-key $(< api.key) You can find more information on how to use Harlequin in its documentation, and on the WherobotsDB adapter on its GitHub repository. The Wherobots Python driver enables integration with many other tools as well. Here’s an example of using the Wherobots Python driver in the QGIS Python console to fetch points of interest from the Overture Maps dataset using Spatial SQL API. from wherobots.db import connect from wherobots.db.region import Region from wherobots.db.runtime import Runtime import geopandas from shapely import wkt with connect( token=os.environ.get("WBC_TOKEN"), runtime=Runtime.SEDONA, region=Region.AWS_US_WEST_2, host="api.cloud.wherobots.com" ) as conn: curr = conn.cursor() curr.execute(""" SELECT names.common[0].value AS name, categories.main AS category, geometry FROM wherobots_open_data.overture.places_place WHERE ST_DistanceSphere(ST_GeomFromWKT("POINT (-122.46552 37.77196)"), geometry) < 10000 AND categories.main = "hiking_trail" """) results = curr.fetchall() print(results) results["geometry"] = results.geometry.apply(wkt.loads) gdf = geopandas.GeoDataFrame(results, crs="EPSG:4326",geometry="geometry") def add_geodataframe_to_layer(geodataframe, layer_name): # Create a new memory layer layer = QgsVectorLayer(geodataframe.to_json(), layer_name, "ogr") # Add the layer to the QGIS project QgsProject.instance().addMapLayer(layer) add_geodataframe_to_layer(gdf, "POI Layer") Visit the Wherobots user documentation to get started with the Spatial SQL API, or see our latest blog post that goes deeper into how to use our database drivers with the Spatial SQL API. Automating Spatial ETL workflows with the Apache Airflow provider for Wherobots ETL (extract, transform, load) workflows are oftentimes required to prepare spatial data for interactive analytics, or to refresh datasets automatically as new data arrives. Apache Airflow is a powerful and popular open source orchestrator of data workflows. With the Wherobots Apache Airflow provider, you can now use Apache Airflow to convert your spatial SQL queries into automated workflows running on Wherobots Cloud. Here’s an example of the Wherobots Airflow provider in use. In this example we identify the top 100 buildings in the state of New York with the most places (facilities, services, business, etc.) registered within them using the Overture Maps dataset, and we’ll eventually auto-refresh the result daily. The initial view can be generated with the following SQL query: CREATE TABLE wherobots.test_db.top_100_hot_buildings_daily AS SELECT buildings.id AS building, first(buildings.names), count(places.geometry) AS places_count, '2023-07-24' AS ts FROM wherobots_open_data.overture.places_place places JOIN wherobots_open_data.overture.buildings_building buildings ON ST_CONTAINS(buildings.geometry, places.geometry) WHERE places.updatetime >= '2023-07-24' AND places.updatetime < '2023-07-25' AND ST_CONTAINS(ST_PolygonFromEnvelope(-79.762152, 40.496103, -71.856214, 45.01585), places.geometry) AND ST_CONTAINS(ST_PolygonFromEnvelope(-79.762152, 40.496103, -71.856214, 45.01585), buildings.geometry) GROUP BY building ORDER BY places_count DESC LIMIT 100 A place in Overture is defined as real-world facilities, services, businesses or amenities. We used an arbitrary date of 2023-07-24. New York is defined by a simple bounding box polygon (79.762152, 40.496103, -71.856214, 45.01585) (we could alternatively join with its appropriate administrative boundary polygon) We use two WHERE clauses on places.updatetime to filter one day’s worth of data. The query creates a new table wherobots.test_db.top_100_hot_buildings_daily to store the query result. Note that it will not directly return any records because we are loading directly into a table. Now, lets use Harlequin as described earlier to inspect the outcome of creating this table with the above query: SELECT * FROM wherobots.test_db.top_100_hot_buildings_daily Apache Airflow and the Airflow Provider for Wherobots allow you to schedule and execute this query each day, injecting the appropriate date filters into your templatized query. In your Apache Airflow instance, install the airflow-providers-wherobots library. You can either execute pip install airflow-providers-wherobots, or add the library to the dependency list of your Apache Airflow runtime. Create a new “generic” connection for Wherobots called wherobots_default, using api.cloud.wherobots.com as the “Host” and your Wherobots API key as the “Password”. The next step is to create an Airflow DAG. The Wherobots Provider exposes the WherobotsSqlOperator for executing SQL queries. Update the hardcoded “2023-07-24” in your query into the Airflow template macros {ds} and {next_ds}, which will be rendered as the DAG schedule date on the fly: import datetime from airflow import DAG from airflow_providers_wherobots.operators.sql import WherobotsSqlOperator with DAG( dag_id="example_wherobots_sql_dag", start_date=datetime.datetime.strptime("2023-07-24", "%Y-%m-%d"), schedule="@daily", catchup=True, max_active_runs=1, ): operator = WherobotsSqlOperator( task_id="execute_query", wait_for_downstream=True, sql=""" INSERT INTO wherobots.test_db.top_100_hot_buildings_daily SELECT buildings.id AS building, first(buildings.names), count(places.geometry) AS places_count, '{{ ds }}' AS ts FROM wherobots_open_data.overture.places_place places JOIN wherobots_open_data.overture.buildings_building buildings ON ST_CONTAINS(buildings.geometry, places.geometry) WHERE places.updatetime >= '{{ ds }}' AND places.updatetime < '{{ next_ds }}' AND ST_CONTAINS(ST_PolygonFromEnvelope(-79.762152, 40.496103, -71.856214, 45.01585), places.geometry) AND ST_CONTAINS(ST_PolygonFromEnvelope(-79.762152, 40.496103, -71.856214, 45.01585), buildings.geometry) GROUP BY building ORDER BY places_count DESC LIMIT 100 """, return_last=False, ) You can visualize the status of the and log of the DAG’s execution in the Apache Airflow UI. As shown below, the operator prints out the exact query rendered and executed when you run your DAG. Please visit the Wherobots user documentation for more details on how to set up your Apache Airflow instance with the Wherobots Provider. Generate Vector Tiles — formatted as PMTiles — at Global Scale Vector tiles are high resolution representations of features optimized for visualization, computed offline and displayed in map applications. This decouples dataset preparation from client side rendering driven by zooming and panning. By decoupling dataset preparation from the interactive experience, map developers use vector tiles to significantly improve the utility, clarity, and responsiveness of feature rich interactive map applications. Traditional vector tiles generators like Tippecanoe are limited to the processing capability of a single VM and require the use of limited formats. These solutions are great for small-scale tile generation workloads when data is already in the right file format. But if you’re like the teams we’ve worked with, you may start small and need to scale past the limits of a single VM, or have a variety of file formats. You just want to generate vector tiles with the data you have, at any scale without having to worry about format conversion steps, configuring infrastructure, partitioning your workload around the capability of a VM, or waiting for workloads to complete. Vector Tile Generation, or VTiles for WherobotsDB generates vector tiles in PMTiles format across common data lake formats, incredibly quickly and at a planetary scale, so you can start small and know you have the capability to scale without having to look for another solution. VTiles is incredibly fast because serverless computation is parallelized, and the WherobotsDB engine is optimized for vector tile generation. This means your development teams can spend less time building map applications that matter to your customers. Using a Tokyo runtime, we generated vector tiles with VTiles for all buildings in the Overture dataset, from zoom levels 4-15 across the entire planet, in 23 minutes. That’s fast and efficient for a planetary scale operation. You can run the tile-generation-example notebook in the Wherobots Pro tier to experience the speed and simplicity of Vtiles yourself. Here’s what this looks like: [videopress 3DYMNV9r] Visit our user documentation to start generating vector tiles at-scale. Try Wherobots Pro Get Started Key takeawaysThis release previews WherobotsAI: Raster Inference for detect/classify/segment on satellite imagery via SQL/Python, plus Map Matching that snaps noisy GPS to road segments at trip scale.Raster Inference is fully managed on GPU runtimes in Professional Edition. The post says you can detect, segment, and classify solar farms, land cover, and marine infrastructure from terabyte-scale Sentinel-2 imagery in under half an hour.WherobotsDB is described as up to 60x more performant for spatial joins than popular general-purpose big-data engines and warehouses, and up to 20x faster than Apache Sedona alone.The new Spatial SQL API exposes WherobotsDB to remote SQL IDEs (Harlequin), Python and Java client SDKs, QGIS, and an Apache Airflow provider (WherobotsSqlOperator) for scheduled spatial ETL.VTiles generated PMTiles for all Overture buildings, zooms 4–15, planet-wide, in 23 minutes on a Tokyo runtime.