{
 "cells": [
  {
   "cell_type": "markdown",
   "id": "fc3d9e9d",
   "metadata": {},
   "source": [
    "# Build your satellite-data inventory. Search it again and again.\n",
    "\n",
    "Save Sentinel-2 metadata for a small study area around Madrid, then explore different dates and scenes without querying the provider every time. You will also try live fallback, refresh the inventory, and inspect its writer version. **Imagery is not downloaded.** Previews and live catalog queries require internet.\n",
    "\n",
    "Run cells from top to bottom in Jupyter or a standard Google Colab Python runtime. The example makes a few small public catalog requests. Results can change or be empty; failures are displayed explicitly. Colab files disappear when its runtime ends\u2014download the archive at the end.\n",
    "\n",
    "This notebook requires SuperSTAC **0.3.x**. Before the release is published, install your locally built v0.3 wheel and skip the installation cell. [Open this notebook in Google Colab](https://colab.research.google.com/drive/1JIx8j0Vi7jypKe5Q6klgCzYQMFHMaKGc?usp=sharing)."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "c5f44bfe",
   "metadata": {
    "tags": [
     "installation"
    ]
   },
   "outputs": [],
   "source": [
    "%pip install -q \"superstac>=0.3,<0.4\" pyarrow"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "445bc7dc",
   "metadata": {
    "tags": []
   },
   "outputs": [],
   "source": [
    "import gc\n",
    "import json\n",
    "import os\n",
    "from pathlib import Path\n",
    "import shutil\n",
    "import tempfile\n",
    "import time\n",
    "\n",
    "import pyarrow.parquet as pq\n",
    "import superstac\n",
    "from superstac import Client, AsyncClient\n",
    "from IPython.display import display, FileLink, Image\n",
    "\n",
    "print(\"SuperSTAC:\", superstac.__version__)\n",
    "print(\"GeoParquet available:\", superstac.geoparquet_available)\n",
    "assert superstac.geoparquet_available, \"Install a GeoParquet-enabled v0.3 build.\"\n",
    "assert hasattr(Client, \"ingest\"), \"Restart the kernel after installing v0.3.\"\n"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "d571088f",
   "metadata": {},
   "source": [
    "## 1. Choose a small area and date window\n",
    "\n",
    "A unique temporary directory prevents this demo from overwriting an existing inventory. Keep the default bounds for a quick run. Raising the search limit does not affect ingestion: ingestion follows all pages within its scope."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "d075d504",
   "metadata": {
    "tags": []
   },
   "outputs": [],
   "source": [
    "CATALOG_ID = \"earth-search\"\n",
    "CATALOG_URL = os.environ.get(\"SUPERSTAC_DEMO_CATALOG_URL\", \"https://earth-search.aws.element84.com/v1\")\n",
    "COLLECTION = \"sentinel-2-l2a\"\n",
    "BBOX = [-3.8, 40.3, -3.6, 40.5]\n",
    "WEEK_ONE = \"2025-02-01T00:00:00Z/2025-02-08T00:00:00Z\"\n",
    "WEEK_TWO = \"2025-02-08T00:00:00Z/2025-02-15T00:00:00Z\"\n",
    "DATASET = Path(tempfile.mkdtemp(prefix=\"superstac-geoparquet-demo-\", dir=Path.cwd()))\n",
    "CATALOGS = [{\"id\": CATALOG_ID, \"url\": CATALOG_URL}]\n",
    "SETTINGS = {\"logging_enabled\": False, \"enable_background_health_monitor\": False,\n",
    "            \"search_healthy_catalogs_only\": False, \"max_retry_attempts\": 1,\n",
    "            \"per_catalog_timeout_seconds\": 30}\n",
    "client = Client(catalogs=CATALOGS, settings=SETTINGS)\n",
    "print(\"Dataset:\", DATASET)\n"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "cbbc293d",
   "metadata": {},
   "source": [
    "## 2. Ingest with progress\n",
    "\n",
    "The callback reports received and durably saved metadata records. A provider may omit a total count, so a percentage is not always available. If interrupted, repeat the call with `resume=True` and the same scope, name, and page size."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "f9901b98",
   "metadata": {
    "tags": []
   },
   "outputs": [],
   "source": [
    "last_progress = 0.0\n",
    "\n",
    "def progress(event):\n",
    "    global last_progress\n",
    "    now = time.monotonic()\n",
    "    if now - last_progress >= 1 or event[\"phase\"] in (\"completed\", \"failed\"):\n",
    "        print(f\"{event['phase']}: {event['items_received']} received, \"\n",
    "              f\"{event['items_saved']} saved, {event['files']} files\")\n",
    "        last_progress = now\n",
    "\n",
    "snapshot = client.ingest(\n",
    "    CATALOG_ID, str(DATASET), name=\"madrid-week-one\",\n",
    "    collections=[COLLECTION], bbox=BBOX, datetime=WEEK_ONE,\n",
    "    page_size=100, items_per_file=1000, max_dataset_mib=128,\n",
    "    progress=progress,\n",
    ")\n",
    "assert snapshot[\"complete\"]\n",
    "print(\"Saved metadata rows:\", snapshot[\"items\"])\n"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "5b408de1",
   "metadata": {},
   "source": [
    "## 3. Search and discover locally\n",
    "\n",
    "This search stays within saved coverage and requires no catalog API request. `sortby` applies before the per-catalog limit. Collection descriptions summarize the saved inventory, not the provider's full collection."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "88de8101",
   "metadata": {
    "tags": []
   },
   "outputs": [],
   "source": [
    "local = Client(catalogs=CATALOGS, settings=SETTINGS,\n",
    "               mode=\"snapshot\", dataset=str(DATASET))\n",
    "query = dict(collections=[COLLECTION], bbox=BBOX, datetime=WEEK_ONE,\n",
    "             sortby=[\"-datetime\"], limit=5)\n",
    "started = time.perf_counter()\n",
    "result = local.search(**query)\n",
    "print(f\"Local search: {(time.perf_counter() - started) * 1000:.1f} ms\")\n",
    "print(json.dumps(result.metadata, indent=2))\n",
    "assert result.metadata[\"catalogs_failed\"] == 0, result.metadata[\"failures\"]\n",
    "print(\"Item IDs:\", [item[\"id\"] for item in result.items()])\n",
    "print(\"Collections:\", local.list_collections())\n"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "6ef8f0a0",
   "metadata": {},
   "source": [
    "### Optional online preview\n",
    "\n",
    "Saved metadata contains asset URLs, not image pixels. Some providers require signing or authentication. If no thumbnail is advertised, this cell simply reports that."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "4655ad0c",
   "metadata": {
    "tags": []
   },
   "outputs": [],
   "source": [
    "thumbnail = next((asset[\"href\"] for item in result.items()\n",
    "                  for key, asset in item.get(\"assets\", {}).items()\n",
    "                  if key == \"thumbnail\" or \"thumbnail\" in asset.get(\"roles\", [])), None)\n",
    "if thumbnail:\n",
    "    display(Image(url=thumbnail, width=500))\n",
    "else:\n",
    "    print(\"No thumbnail URL in these results; inspect item assets for available previews.\")\n"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "1fb0741e",
   "metadata": {},
   "source": [
    "## 4. Inspect the file's writer version\n",
    "\n",
    "The Parquet footer includes both standard GeoParquet metadata and SuperSTAC provenance. The file's writer version can differ from the version currently installed. Empty scopes have no Parquet files."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "8b9b4317",
   "metadata": {
    "tags": []
   },
   "outputs": [],
   "source": [
    "manifest = json.loads((DATASET / \"manifest.json\").read_text())\n",
    "files = manifest[\"catalogs\"][CATALOG_ID][\"madrid-week-one\"][\"files\"]\n",
    "if files:\n",
    "    footer = pq.read_metadata(DATASET / files[0][\"path\"])\n",
    "    print(\"created_by:\", footer.created_by)\n",
    "    producer = json.loads(footer.metadata[b\"superstac\"])\n",
    "    print(\"SuperSTAC provenance:\", producer)\n",
    "    assert producer[\"version\"] == superstac.__version__\n",
    "    assert b\"geo\" in footer.metadata and b\"stac-geoparquet\" in footer.metadata\n",
    "else:\n",
    "    print(\"Complete empty inventory: no Parquet files to inspect.\")\n"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "6163aa25",
   "metadata": {},
   "source": [
    "## 5. Retain a second scope and search their union\n",
    "\n",
    "Adjacent scopes can jointly cover a larger query; gaps cannot. Reopen snapshot clients to see a newly published manifest. Duplicate collection/item IDs are returned once."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "6e0e3dae",
   "metadata": {
    "tags": []
   },
   "outputs": [],
   "source": [
    "client.ingest(CATALOG_ID, str(DATASET), name=\"madrid-week-two\",\n",
    "              collections=[COLLECTION], bbox=BBOX, datetime=WEEK_TWO,\n",
    "              max_dataset_mib=128, progress=progress)\n",
    "local.shutdown()\n",
    "del local\n",
    "gc.collect()\n",
    "local = Client(catalogs=CATALOGS, settings=SETTINGS,\n",
    "               mode=\"snapshot\", dataset=str(DATASET))\n",
    "union = local.search(collections=[COLLECTION], bbox=BBOX,\n",
    "                     datetime=\"2025-02-01T00:00:00Z/2025-02-15T00:00:00Z\",\n",
    "                     sortby=[\"-datetime\"], limit=10)\n",
    "assert union.metadata[\"catalogs_failed\"] == 0, union.metadata[\"failures\"]\n",
    "print(\"Named scopes:\", list(json.loads((DATASET / \"manifest.json\").read_text())[\"catalogs\"][CATALOG_ID]))\n",
    "print(\"Union returned:\", len(union))\n"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "5dffad0a",
   "metadata": {},
   "source": [
    "## 6. Automatic API fallback\n",
    "\n",
    "Auto uses fresh, complete local coverage where available. March is outside these saved scopes, so this query goes to the API. A zero freshness limit also forces live requests. Auto does not save its API results as complete inventories."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "c98cdbbe",
   "metadata": {
    "tags": []
   },
   "outputs": [],
   "source": [
    "auto = Client(catalogs=CATALOGS, settings=SETTINGS, mode=\"auto\",\n",
    "              dataset=str(DATASET), max_snapshot_age_seconds=86400)\n",
    "manifest_before = (DATASET / \"manifest.json\").read_bytes()\n",
    "live_result = auto.search(collections=[COLLECTION], bbox=BBOX,\n",
    "                          datetime=\"2025-03-01T00:00:00Z/2025-03-08T00:00:00Z\", limit=3)\n",
    "print(json.dumps(live_result.metadata, indent=2))\n",
    "assert live_result.metadata[\"catalogs_failed\"] == 0, live_result.metadata[\"failures\"]\n",
    "assert manifest_before == (DATASET / \"manifest.json\").read_bytes()\n",
    "auto.shutdown()\n"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "1e693a07",
   "metadata": {},
   "source": [
    "## 7. Incremental acquisition-window refresh\n",
    "\n",
    "Extend week two by a day while re-fetching an overlapping window. New records override older records with the same collection/item ID. This does **not** detect deletions or late changes outside that window. Omit `incremental_since` for a full refresh. Untouched history retains its original freshness timestamp."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "32d7c36e",
   "metadata": {
    "tags": []
   },
   "outputs": [],
   "source": [
    "updated = client.ingest(\n",
    "    CATALOG_ID, str(DATASET), name=\"madrid-week-two\",\n",
    "    collections=[COLLECTION], bbox=BBOX,\n",
    "    datetime=\"2025-02-08T00:00:00Z/2025-02-16T00:00:00Z\",\n",
    "    incremental_since=\"2025-02-14T00:00:00Z\",\n",
    "    max_dataset_mib=128, progress=progress,\n",
    ")\n",
    "print(\"Overlay present:\", updated[\"overlays\"])\n",
    "assert updated[\"overlays\"]\n"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "bf5a91d4",
   "metadata": {},
   "source": [
    "## 8. Compact and reclaim retired files\n",
    "\n",
    "Compaction resolves duplicate versions and rewrites bounded files. Old readers remain usable. Applied cleanup refuses active snapshot readers, so release them first; `shutdown()` alone does not release their inventory handle. The calls below only affect the temporary demo dataset."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "c048fa46",
   "metadata": {
    "tags": []
   },
   "outputs": [],
   "source": [
    "client.compact_dataset(str(DATASET), items_per_file=1000, max_dataset_mib=128)\n",
    "local.shutdown()\n",
    "del local\n",
    "gc.collect()\n",
    "preview = client.cleanup_dataset(str(DATASET))\n",
    "print(\"Reclaimable bytes:\", preview[\"bytes\"])\n",
    "print(\"Retired files:\", preview[\"files\"])\n",
    "cleaned = client.cleanup_dataset(str(DATASET), apply=True)\n",
    "print(\"Cleanup applied:\", cleaned[\"applied\"])\n"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "a7a30163",
   "metadata": {},
   "source": [
    "## 9. Async clients use the same backend\n",
    "\n",
    "Jupyter and Colab support top-level `await`; there is no need for `asyncio.run()` here. Async ingestion and maintenance use the same keyword options as their synchronous equivalents."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "b5d31b72",
   "metadata": {
    "tags": []
   },
   "outputs": [],
   "source": [
    "async_client = AsyncClient(catalogs=CATALOGS, settings=SETTINGS,\n",
    "                           mode=\"snapshot\", dataset=str(DATASET))\n",
    "async_result = await async_client.search(**query)\n",
    "assert async_result.metadata[\"catalogs_failed\"] == 0\n",
    "print(\"Async local results:\", len(async_result))\n",
    "await async_client.shutdown()\n",
    "del async_client\n",
    "gc.collect()\n"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "0385f7cd",
   "metadata": {},
   "source": [
    "## 10. Download the inventory and support information\n",
    "\n",
    "The archive includes the manifest, metadata files, and a small environment report. It contains no imagery. Download it before ending a Colab runtime. When filing a bug, include the package version, a minimal query, and failure metadata; remove credentials or sensitive data from any attachment."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "bca187c8",
   "metadata": {
    "tags": []
   },
   "outputs": [],
   "source": [
    "client.shutdown()\n",
    "support = {\"superstac_version\": superstac.__version__,\n",
    "           \"query\": query, \"metadata\": result.metadata}\n",
    "(DATASET / \"demo-report.json\").write_text(json.dumps(support, indent=2))\n",
    "archive = shutil.make_archive(str(DATASET), \"zip\", root_dir=DATASET)\n",
    "print(\"Saved:\", archive)\n",
    "try:\n",
    "    from google.colab import files as colab_files\n",
    "except ImportError:\n",
    "    display(FileLink(Path(archive).name))\n",
    "else:\n",
    "    colab_files.download(archive)\n"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "b5b81161",
   "metadata": {},
   "source": [
    "## Next steps\n",
    "\n",
    "- [GeoParquet guide](https://spatialnode.com/superstac/docs/guides/geoparquet)\n",
    "- [Measured local-scan benchmark](https://spatialnode.com/superstac/docs/guides/benchmarks)\n",
    "- [What\u2019s new](https://spatialnode.com/superstac/docs/releases)\n",
    "\n",
    "This notebook's single timing is a demo, not a benchmark. The documented benchmark compares the same scanner with and without pruning, asserts equal results, and reports repeated measurements."
   ]
  }
 ],
 "metadata": {
  "colab": {
   "provenance": []
  },
  "kernelspec": {
   "display_name": "Python 3",
   "language": "python",
   "name": "python3"
  },
  "language_info": {
   "name": "python",
   "version": "3.12"
  }
 },
 "nbformat": 4,
 "nbformat_minor": 5
}
