Resources

Introducing Strabo: A Geosimilarity Engine for Earth Data

Jeff

3 min. read

One of LGND’s missions is to build the best platform for the creation, storage, indexing, and serving of massive earth embedding datasets. The pursuit of this mission has led us to build Strabo, the world’s first geosimilarity engine.

Why is this engine needed?

Today’s vector indexes are fantastic for semantic similarity but do not take into account spatial or temporal structure. A solar farm in Japan is semantically similar to a solar farm in the US but we know these are not the same. They exist under disparate contexts, are subject to distinct regulations, and are built from different materials. A vector index views these as the same because the qualifying difference is in their location.

For teams doing global-scale earth observation, that blind spot means expensive workarounds: manual filtering, regional re-indexing, or bolting geospatial logic onto infrastructure that wasn't built for it.

Strabo is built from the ground up to treat space and time as first-class dimensions of search, not metadata applied after the fact. To show what that unlocks, we're releasing a demo that runs lightweight landcover classifiers across a Strabo index of ~9.2 billion Sentinel-2 embeddings.

The name is a nod to Strabo of Amasia, the ancient Greek geographer whose Geographica was among the first attempts to systematically unify knowledge of the world's places, peoples, and regions into a single coherent structure. That's the same instinct behind Strabo: rather than treating location as scattered metadata bolted onto a search index, it organizes the world's earth observation data around space and time.

What is a Geosimilarity Engine?

A geosimilarity engine extends traditional vector search with a geospatial tiling system to allow indexing, search, and visualization of earth embeddings across various spatial and temporal scales. Where a general-purpose vector database treats location as a filter applied to the search, a geosimilarity engine carries space and time in the physical layout of the index and in the scoring function. This allows the engine to support a variety of downstream use cases; spatio-temporal-semantic search, change detection, and agentic reasoning to name a few.

Strabo is a modified IVF index backed by object storage that has the following features:

  • Spatial and temporal overviews (coarse, searchable summaries).
  • Multiple embedding channels in one index.
  • Automated compaction and index maintenance.
  • Space and time as a rank, not just a filter.
  • Row-level access policies (logical tenancy).
  • Query planner, telemetry, metrics.
  • Distributed queries for large index scans.
  • Read-through NVMe cache.
  • GPU accelerated distributed builds for large indexes.
  • Append support, as imagery is collected.
  • Raster provenance.
  • 50-120x cheaper than existing commercial solutions.

It draws inspiration from Turbopuffer, Lance, Ray, BigQuery, Lucene, RocksDB, Iceberg, COG, and VirtualiZarr.

What’s Next?

The demo above is a preview of what's possible with Strabo, but it's just the beginning. Next week, we'll share how you can start using Strabo directly within the LGND platform. 

We'll also be releasing several more blog posts in the coming weeks discussing Strabo's capabilities and the applications we're building on top of it.