{
  "cells": [
    {
      "cell_type": "markdown",
      "id": "introduction",
      "metadata": {},
      "source": [
        "# SuperSTAC — Many catalogs. One search.\n",
        "\n",
        "Search Sentinel-2 item metadata across **Earth Search** and **Microsoft Planetary Computer**, inspect the combined response, map scene footprints, and export GeoJSON. No API key or GPU is needed for these public metadata searches.\n",
        "\n",
        "**In Colab:** choose a standard Python CPU runtime, then run cells from top to bottom. You can upload this `.ipynb` using **File → Upload notebook**. In Jupyter, use a Python 3.9+ environment compatible with the current SuperSTAC wheel.\n",
        "\n",
        "This notebook describes SuperSTAC's current alpha API. Public catalogs and released package versions can change, so results and counts are not fixed. Internet access is required for package installation, catalog requests, and map tiles.\n",
        "\n",
        "[Python guide](https://spatialnode.com/superstac/docs/python/overview) · [Source repository](https://github.com/Spatialnode/superstac)"
      ]
    },
    {
      "cell_type": "markdown",
      "id": "installation-notes",
      "metadata": {},
      "source": [
        "## 1. Install and import\n",
        "\n",
        "SuperSTAC provides the search engine. Pandas displays tables; Folium displays the map. The install cell targets the active notebook kernel.\n",
        "\n",
        "If a compatible wheel is unavailable, follow the repository's [source installation guide](https://github.com/Spatialnode/superstac/blob/main/docs/content/docs/start/installation.md). If you replace an already-imported version, restart the runtime before continuing. The version printed below helps when reporting an issue."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "id": "install",
      "metadata": {},
      "outputs": [],
      "source": [
        "%pip install --quiet superstac pandas folium"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "id": "imports",
      "metadata": {},
      "outputs": [],
      "source": [
        "import json\n",
        "import platform\n",
        "from importlib.metadata import version\n",
        "from pathlib import Path\n",
        "from zipfile import ZipFile\n",
        "\n",
        "import folium\n",
        "import pandas as pd\n",
        "from IPython.display import FileLink, display\n",
        "from superstac import Client\n",
        "\n",
        "print(f\"Python {platform.python_version()}\")\n",
        "print(f\"SuperSTAC {version('superstac')}\")"
      ]
    },
    {
      "cell_type": "markdown",
      "id": "query-notes",
      "metadata": {},
      "source": [
        "## 2. Choose the catalogs and query\n",
        "\n",
        "The bounding box covers Luxembourg and its surroundings, in **[west, south, east, north] longitude/latitude degrees (EPSG:4326)**. Change it and the date interval to explore another area.\n",
        "\n",
        "`limit=5` is a **per-catalog cap**, so two catalogs can contribute up to ten items before deduplication. It is not a global limit, a count of every matching scene, or a guarantee of the newest or least cloudy scenes. Deduplication compares item IDs; the same observation with different IDs can remain in the output.\n",
        "\n",
        "For a short notebook session, each catalog disables its background health monitor. Startup health checks and collection discovery still run."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "id": "configuration",
      "metadata": {},
      "outputs": [],
      "source": [
        "CATALOGS = [\n",
        "    {\n",
        "        \"id\": \"earth-search\",\n",
        "        \"url\": \"https://earth-search.aws.element84.com/v1\",\n",
        "        \"settings\": {\n",
        "            \"health_check_strategy\": \"hourly\",\n",
        "            \"healthy_status_code_range\": [200, 299],\n",
        "            \"enable_background_health_monitor\": False,\n",
        "        },\n",
        "    },\n",
        "    {\n",
        "        \"id\": \"microsoft\",\n",
        "        \"url\": \"https://planetarycomputer.microsoft.com/api/stac/v1\",\n",
        "        \"settings\": {\n",
        "            \"health_check_strategy\": \"hourly\",\n",
        "            \"healthy_status_code_range\": [200, 299],\n",
        "            \"enable_background_health_monitor\": False,\n",
        "        },\n",
        "    },\n",
        "]\n",
        "\n",
        "QUERY = {\n",
        "    \"collections\": [\"sentinel-2-l2a\"],\n",
        "    \"bbox\": [6.0, 49.0, 7.0, 50.0],\n",
        "    \"datetime\": \"2024-01-01T00:00:00Z/2024-01-31T23:59:59Z\",\n",
        "    \"limit\": 5,\n",
        "}\n",
        "\n",
        "SETTINGS = {\n",
        "    \"max_concurrent_catalogs\": 2,\n",
        "    \"per_catalog_timeout_seconds\": 30,\n",
        "    \"max_retry_attempts\": 2,\n",
        "    \"max_items_per_catalog\": 5,\n",
        "    \"deduplicate_items\": True,\n",
        "    \"unify_response\": True,\n",
        "}\n",
        "\n",
        "print(json.dumps(QUERY, indent=2))"
      ]
    },
    {
      "cell_type": "markdown",
      "id": "search-notes",
      "metadata": {},
      "source": [
        "## 3. Discover collections and run one federated search\n",
        "\n",
        "SuperSTAC starts the engine, discovers which catalogs advertise the collection, and searches the selected sources concurrently. Startup can take longer than the search itself; the per-catalog search timeout does not bound the entire startup process.\n",
        "\n",
        "The `finally` block closes the client even if discovery or search raises an exception. Results are materialized before shutdown and remain available to the later cells. To try a different query, edit the configuration above and rerun this cell and the result cells below."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "id": "federated-search",
      "metadata": {},
      "outputs": [],
      "source": [
        "# Clear previous results so a failed rerun cannot accidentally reuse them.\n",
        "search = None\n",
        "items = []\n",
        "metadata = {}\n",
        "geojson = {\"type\": \"FeatureCollection\", \"features\": []}\n",
        "\n",
        "client = Client(catalogs=CATALOGS, settings=SETTINGS)\n",
        "try:\n",
        "    available = client.list_collections()\n",
        "    selected = [entry for entry in available if entry[\"id\"] in QUERY[\"collections\"]]\n",
        "    print(\"Discovery: catalogs advertising the requested collection\")\n",
        "    display(pd.DataFrame(selected, columns=[\"id\", \"catalogs\"]))\n",
        "\n",
        "    search = client.search(**QUERY)\n",
        "    items = search.items()\n",
        "    metadata = search.metadata\n",
        "    geojson = search.to_geojson()\n",
        "finally:\n",
        "    client.shutdown()\n",
        "\n",
        "print(f\"Returned {search.matched()} items after aggregation.\")"
      ]
    },
    {
      "cell_type": "markdown",
      "id": "metadata-notes",
      "metadata": {},
      "source": [
        "## 4. Check completeness before using the data\n",
        "\n",
        "A response can contain useful items and catalog failures at the same time. A healthy/capability filter can also exclude a catalog before execution without recording a search failure. Compare the configured catalog count, discovery table, and metadata rather than assuming every source was searched.\n",
        "\n",
        "`matched()` and `total_items` count returned items after aggregation, not all scenes available upstream. Python exposes ordinary item dictionaries and run-level metadata; it does not expose the Rust per-item `catalog_id` / `seen_in` provenance wrapper."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "id": "inspect-metadata",
      "metadata": {},
      "outputs": [],
      "source": [
        "if search is None:\n",
        "    raise RuntimeError(\"Run the search cell successfully before inspecting results.\")\n",
        "\n",
        "print(json.dumps(metadata, indent=2))\n",
        "\n",
        "if metadata.get(\"catalogs_queried\", 0) < len(CATALOGS):\n",
        "    print(\"Some configured catalogs were not selected. Check discovery and catalog health.\")\n",
        "for failure in metadata.get(\"failures\", []):\n",
        "    print(f\"Catalog failure: {failure['catalog_id']}: {failure['reason']}\")\n",
        "if not items:\n",
        "    print(\"No items returned. Check failures and unsupported collections before widening the date range or area.\")"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "id": "item-table",
      "metadata": {},
      "outputs": [],
      "source": [
        "rows = []\n",
        "for item in items:\n",
        "    properties = item.get(\"properties\", {})\n",
        "    rows.append({\n",
        "        \"id\": item[\"id\"],\n",
        "        \"collection\": item.get(\"collection\"),\n",
        "        \"datetime\": properties.get(\"datetime\") or properties.get(\"start_datetime\"),\n",
        "        \"cloud_cover_percent\": properties.get(\"eo:cloud_cover\"),\n",
        "        \"asset_count\": len(item.get(\"assets\", {})),\n",
        "    })\n",
        "\n",
        "item_table = pd.DataFrame(rows, columns=[\n",
        "    \"id\", \"collection\", \"datetime\", \"cloud_cover_percent\", \"asset_count\",\n",
        "])\n",
        "display(item_table)"
      ]
    },
    {
      "cell_type": "markdown",
      "id": "map-notes",
      "metadata": {},
      "source": [
        "## 5. Map the scene footprints\n",
        "\n",
        "The green rectangle is the search area; the blue polygons are returned scene footprints. A footprint can extend beyond the query area. Hover over a scene to inspect its ID. This displays metadata geometry over OpenStreetMap tiles, not satellite image pixels."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "id": "footprint-map",
      "metadata": {},
      "outputs": [],
      "source": [
        "west, south, east, north = QUERY[\"bbox\"]\n",
        "scene_map = folium.Map(\n",
        "    location=[(south + north) / 2, (west + east) / 2],\n",
        "    tiles=\"OpenStreetMap\",\n",
        "    zoom_start=7,\n",
        ")\n",
        "folium.Rectangle(\n",
        "    bounds=[[south, west], [north, east]],\n",
        "    color=\"#285d46\",\n",
        "    weight=3,\n",
        "    fill=False,\n",
        "    tooltip=\"Search area\",\n",
        ").add_to(scene_map)\n",
        "\n",
        "# Keep the display layer small; the exported GeoJSON retains the full STAC items.\n",
        "footprints = {\n",
        "    \"type\": \"FeatureCollection\",\n",
        "    \"features\": [\n",
        "        {\n",
        "            \"type\": \"Feature\",\n",
        "            \"geometry\": item[\"geometry\"],\n",
        "            \"properties\": {\"id\": item[\"id\"], \"collection\": item.get(\"collection\", \"\")},\n",
        "        }\n",
        "        for item in items if item.get(\"geometry\")\n",
        "    ],\n",
        "}\n",
        "if footprints[\"features\"]:\n",
        "    folium.GeoJson(\n",
        "        footprints,\n",
        "        name=\"Scene footprints\",\n",
        "        style_function=lambda feature: {\n",
        "            \"color\": \"#2563eb\", \"weight\": 2, \"fillOpacity\": 0.08,\n",
        "        },\n",
        "        tooltip=folium.GeoJsonTooltip(fields=[\"id\", \"collection\"]),\n",
        "    ).add_to(scene_map)\n",
        "else:\n",
        "    print(\"No scene geometry to draw; showing the search area only.\")\n",
        "\n",
        "scene_map.fit_bounds([[south, west], [north, east]])\n",
        "display(scene_map)"
      ]
    },
    {
      "cell_type": "markdown",
      "id": "asset-notes",
      "metadata": {},
      "source": [
        "## 6. Inspect available assets\n",
        "\n",
        "A STAC item links to data assets. Listing them does not download imagery. Asset keys differ by provider; some URLs require provider-specific authentication or signing. This example does not sign Planetary Computer URLs or fetch image pixels."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "id": "inspect-assets",
      "metadata": {},
      "outputs": [],
      "source": [
        "if items:\n",
        "    first_item = items[0]\n",
        "    print(f\"Assets for {first_item['id']}\")\n",
        "    display(pd.DataFrame([\n",
        "        {\"key\": key, \"type\": asset.get(\"type\"), \"roles\": asset.get(\"roles\", []), \"href\": asset.get(\"href\")}\n",
        "        for key, asset in first_item.get(\"assets\", {}).items()\n",
        "    ], columns=[\"key\", \"type\", \"roles\", \"href\"]))\n",
        "else:\n",
        "    print(\"No returned item to inspect.\")"
      ]
    },
    {
      "cell_type": "markdown",
      "id": "export-notes",
      "metadata": {},
      "source": [
        "## 7. Export and download\n",
        "\n",
        "Save full STAC items as a GeoJSON FeatureCollection, plus the search metadata and query separately. GeoJSON alone does not include the run-level failure counts. The map HTML uses online tiles and browser libraries when opened.\n",
        "\n",
        "The output directory and ZIP below are overwritten when this cell is rerun. In Colab, runtime files are temporary: download the ZIP to keep them."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "id": "export-results",
      "metadata": {},
      "outputs": [],
      "source": [
        "if search is None:\n",
        "    raise RuntimeError(\"Run the search cell successfully before exporting.\")\n",
        "\n",
        "output_dir = Path(\"superstac-output\")\n",
        "output_dir.mkdir(exist_ok=True)\n",
        "exports = {\n",
        "    \"items.geojson\": geojson,\n",
        "    \"search-metadata.json\": metadata,\n",
        "    \"query.json\": {\"catalogs\": CATALOGS, \"settings\": SETTINGS, \"query\": QUERY},\n",
        "}\n",
        "for filename, value in exports.items():\n",
        "    (output_dir / filename).write_text(json.dumps(value, indent=2), encoding=\"utf-8\")\n",
        "scene_map.save(str(output_dir / \"footprints.html\"))\n",
        "\n",
        "archive = Path(\"superstac-results.zip\")\n",
        "with ZipFile(archive, \"w\") as bundle:\n",
        "    for filename in [*exports, \"footprints.html\"]:\n",
        "        bundle.write(output_dir / filename, arcname=filename)\n",
        "print(f\"Saved {len(items)} items and search diagnostics to {archive.resolve()}\")"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "id": "download-results",
      "metadata": {},
      "outputs": [],
      "source": [
        "# Download in Colab, or display a file link in a local Jupyter notebook.\n",
        "try:\n",
        "    from google.colab import files\n",
        "except ImportError:\n",
        "    display(FileLink(str(archive)))\n",
        "else:\n",
        "    files.download(str(archive))"
      ]
    },
    {
      "cell_type": "markdown",
      "id": "next-steps",
      "metadata": {},
      "source": [
        "## Where to go next\n",
        "\n",
        "- Change the bounding box or month, then rerun from the configuration cell onward.\n",
        "- Add another compatible STAC API to `CATALOGS` and inspect discovery before searching.\n",
        "- Increase both `QUERY[\"limit\"]` and `SETTINGS[\"max_items_per_catalog\"]` to collect more items per catalog.\n",
        "- Learn about [collection and asset aliases](https://spatialnode.com/superstac/docs/guides/aliases), the [async client](https://spatialnode.com/superstac/docs/python/async), and [current limitations](https://spatialnode.com/superstac/docs/reference/status).\n",
        "\n",
        "If a source is unavailable, inspect `metadata[\"failures\"]`, then rerun the search cell to create a fresh client and repeat startup checks. A successful search with zero results is different from a failed search. No specific scene IDs, source ordering, or duplicate counts are guaranteed."
      ]
    }
  ],
  "metadata": {
    "colab": {
      "name": "superstac-python-quickstart.ipynb",
      "provenance": []
    },
    "kernelspec": {
      "display_name": "Python 3",
      "language": "python",
      "name": "python3"
    },
    "language_info": {
      "file_extension": ".py",
      "mimetype": "text/x-python",
      "name": "python",
      "pygments_lexer": "ipython3"
    }
  },
  "nbformat": 4,
  "nbformat_minor": 5
}
