GeoParquet inventories
Save STAC metadata for your study area, search it locally, and update it when needed.
If you often search the same region, save its STAC metadata to GeoParquet and reuse it across queries. You can change dates and filters within the saved coverage without fetching the metadata again. Imagery and previews still come from the provider.
For a runnable Python example, open the notebook.
The CLI examples below use a catalog named earth-search in superstac.yml and
its configured collection name, sentinel-2-l2a.
Python wheels include GeoParquet. To build the CLI with it:
cargo build -p superstac-cli --features geoparquetSave your first inventory
Start with a small area and date range. Downloading an entire catalog can take a long time and use substantial storage.
superstac ingest --catalog earth-search --name madrid-feb \
--collection sentinel-2-l2a --datetime 2025-02-01/2025-03-01 \
--bbox=-4.5,39.5,-3,41 --output ./data --max-dataset-mib 1024This saves metadata under ./data and shows progress as pages arrive. A percentage
and ETA appear when the provider supplies a total. Use --no-progress to hide
progress; JSON output stays on stdout and progress goes to stderr.
madrid-feb names this saved area and date range, called a scope. Use another
name to add a different area or month to the same dataset. Repeat --catalog to
include more catalogs. Ingestion only runs when you request it.
If the download stops, repeat the command with --resume. Keep the name, filters,
and page size unchanged. If a provider’s continuation token has expired, restart
without --resume. Resuming an already complete scope does not refresh it.
Search saved metadata
superstac --mode snapshot --dataset ./data --json search \
--collection sentinel-2-l2a --datetime 2025-02-10/2025-02-15 \
--bbox=-4,40,-3.5,40.5 --limit 20Your search can differ from the ingestion query, as long as the saved collections, area, and dates cover it. Several scopes can cover a query together, but gaps count as missing coverage. Snapshot mode never contacts providers; check per-catalog response metadata for coverage failures.
Add --sortby=-datetime for newest first. Sorting applies within each catalog,
not across the combined response. Collection discovery describes the saved data,
rather than the provider’s entire catalog.
Fall back to live catalogs
Use auto when you want local searches where possible and live searches otherwise:
superstac --mode auto --dataset ./data --max-snapshot-age-seconds 86400 \
search --collection sentinel-2-l2a --datetime 2025-02-10/2025-02-15 \
--bbox=-4,40,-3.5,40.5 --limit 20| Mode | Where the results come from |
|---|---|
snapshot | Saved data only, with no age limit. |
auto | Complete, fresh local coverage; otherwise the catalog’s API. |
live | The catalog’s API. |
Auto mode defaults to a maximum age of one day, measured from the start of ingestion. It also falls back if the dataset is missing, unreadable, or belongs to a different source URL. A valid local search with no matches stays local.
Live results are not saved automatically. Run ingest to add coverage you want
to reuse. Auto mode reads the latest manifest on each search; reopen snapshot
clients after updating their dataset.
Without --mode, supplying a dataset or file selects snapshot mode; otherwise
SuperSTAC uses live mode. To search an existing standalone file:
superstac --geoparquet earth-search=./items.parquet search --limit 20Repeat --geoparquet CATALOG=PATH for more catalogs. Standalone files do not carry
the managed dataset’s coverage and freshness guarantees.
Use Python
from superstac import Client
client = Client.from_yaml("superstac.yml")
client.ingest(
"earth-search", "./data", name="madrid-feb",
collections=["sentinel-2-l2a"], datetime="2025-02-01/2025-03-01",
bbox=[-4.5, 39.5, -3, 41],
progress=lambda event: print(event["phase"], event["items_saved"]),
)
local = Client.from_yaml("superstac.yml", mode="snapshot", dataset="./data")
results = local.search(
collections=["sentinel-2-l2a"], datetime="2025-02-10/2025-02-15",
bbox=[-4, 40, -3.5, 40.5], sortby=["-datetime"], limit=20,
)Use mode="auto" and max_snapshot_age_seconds=86400 for live fallback.
AsyncClient supports the same options, with awaitable ingest, search, and
maintenance methods. Constructors and from_yaml() remain synchronous.
Progress callbacks run on a worker thread and should return promptly. Callback exceptions are reported without stopping ingestion. See the Python API for method signatures.
Update an inventory
Rerun ingestion with the same name to replace that scope. Other scopes remain unchanged. Changing its filters replaces its previous coverage too.
To fetch only a recent acquisition window, use --incremental-since:
superstac ingest --catalog earth-search --name madrid-feb \
--collection sentinel-2-l2a --datetime 2025-02-01/2025-03-05 \
--bbox=-4.5,39.5,-3,41 --output ./data \
--incremental-since 2025-02-25T00:00:00ZKeep the existing scope’s source, collections, area, and start date. You may extend the end date; the update window must overlap the previous coverage without a gap. New records replace saved copies with the same collection and item ID.
This checks acquisition dates, not the provider’s change history. It cannot detect
deletions or changes outside the window. Run a full refresh without
--incremental-since when you need to reconcile those changes. Incremental updates
keep the original freshness timestamp because older records haven’t been rechecked.
Manage disk space
Ingestion has a default 1,024 MiB storage limit for Parquet files, including old
files and unfinished downloads. Set --max-dataset-mib to raise it, or 0 to remove
the limit. If ingestion hits the limit, raise it and resume. Metadata and filesystem
overhead are outside the limit.
A dataset can contain many Parquet files; you search the dataset directory as one
inventory. Use --items-per-file to adjust the approximate rows per file.
--page-size controls each provider request, not the total number of items saved.
Old files are kept until you explicitly clean them up. Allow room for both old and new files during a refresh or compaction.
superstac compact --output ./data --items-per-file 10000 --max-dataset-mib 2048
superstac cleanup --output ./data # preview what can be removed
superstac cleanup --output ./data --apply # remove unreferenced Parquet filesCompaction combines updates and removes duplicate records while preserving coverage
and freshness. If it fails, the previous inventory remains readable. Cleanup
protects files used by the current manifest and saved checkpoints. Close snapshot
clients before applying cleanup; active readers or writers prevent it from running.
shutdown() stops monitoring but does not release a client’s snapshot lock.
Rust
Enable the engine’s geoparquet feature. Use ingest_catalog with IngestOptions
to save data, SuperSTACEngine::from_dataset for local search, or
SuperSTACEngine::automatic(storage, path, max_age) for live fallback.
The engine also exports compact_dataset and cleanup_dataset.
IngestOptions::progress accepts an Arc<dyn Fn(&IngestProgress) + Send + Sync>;
callbacks should return promptly and not panic. Datasets currently use local
storage, not remote object stores.
Performance and version information
Selective queries can skip irrelevant Parquet row groups. Broad queries and sorting read more data; searches spanning multiple scopes or uncompacted updates also do more work. Compaction can help after incremental updates. Deduplication uses memory proportional to the unique items scanned. See the benchmark for measurements and how to reproduce them.
Newly written files record the SuperSTAC version automatically, including after compaction. Existing files keep their metadata until rewritten. To inspect a file:
import json
import pyarrow.parquet as pq
footer = pq.read_metadata("path/to/part.parquet")
print(footer.created_by)
print(json.loads(footer.metadata[b"superstac"]))For bug reports, include superstac --version or Python’s superstac.__version__.
Search diagnostics also include metadata.superstac_version.
Version-1 inventories remain readable. Their next successful ingestion writes a version-2 manifest, which older SuperSTAC builds cannot read.