Skip to content

Add PRR collection generation; point OSC item assets to PRR (v0.2.0) - #23

Open
TejasMorbagal wants to merge 59 commits into
mainfrom
TejasMorbagal/generate-prr-collection
Open

TejasMorbagal wants to merge 59 commits into
mainfrom
TejasMorbagal/generate-prr-collection

Conversation

@TejasMorbagal

@TejasMorbagal TejasMorbagal commented Jul 8, 2026 •

Copy link
Copy Markdown
Member

Summary

Adds a generate-prr-collection command that builds an ESA EarthCODE Project Results Repository (PRR) STAC collection from the existing dataset config. It also changes OSC publishing to link each dataset straight to its PRR collection, the same way other PRR-hosted OSC products are linked. The S3 STAC catalog becomes opt-in. The version goes to 0.2.0 because this includes breaking config and workflow changes (see below).

Changes

New: PRR collection generation

  • New CLI command deep-code generate-prr-collection dataset.yaml [-o DIR] and Python helper deep_code.tools.prr.generate_prr_collection. It writes a self-contained Collection → Item → Assets tree to local files and needs no GitHub or S3 credentials.
  • The Item carries the datacube extension (cube:dimensions / cube:variables read from the Zarr) and the file extension (file:size for the store and .zmetadata). Its asset hrefs are relative (./{dataset_id}), because PRR ingests the Zarr next to the item. The Collection declares the OSC, Scientific, Processing, Themes and CF extensions plus the PRR-mandatory fields.
  • New PRR config fields: osc_initiative, osc_missions, osc_contract_number, osc_project_website, osc_project_description, thumbnail, thumbnail_media_type, sci_doi, sci_citation, prr_output_dir. If PRR-required fields are missing, the command logs a warning and still runs.
  • New coord_position option (center | left | right) so the bbox and cube extents cover full grid cells.
  • generate-config templates document the PRR fields.

Changed: OSC collection links to PRR by default

  • stac_catalog_s3_root is now optional. Without it, the OSC collection links to the dataset's PRR collection:

    • child → https://eoresults.esa.int/stac/collections/{collection_id}
    • via (title "Access") → https://eoresults.esa.int/browser/#/external/eoresults.esa.int/stac/collections/{collection_id}

    Nothing is written to S3 and no S3 credentials are needed. Publishing checks that the PRR collection exists and fails with a clear error if it doesn't.

  • For datasets not in PRR, setting stac_catalog_s3_root keeps the S3 catalog + item behaviour. In that case:

    • The S3 item's asset hrefs are now absolute: the new optional access_link, or else the zarr-data asset looked up from the PRR item. Before, they were ./{dataset_id}, which resolved to a path on S3 where no Zarr exists.
    • The S3 item's self link now matches where the file is written (…/{collection_id}/items/{item_id}.json, not …/{collection_id}/item.json).

Config handling

  • dataset_status renamed to osc_status (default completed). The old key still works but logs a deprecation warning.
  • An incomplete items_config entry now gives a clear error (dataset_id is required in items_config entry 0.) instead of a KeyError. publish and generate-prr-collection share one parser (build_items_config).

Housekeeping

  • Type hints updated to built-in generics / X | None. Fixed a mutable default argument in ogc_api_record.py.
  • CI pins ruff==0.9.10 for now.
  • Docs updated: configuration.md has a new "Data links" section, and cli.md / python-api.md describe the PRR command, the PRR default, osc_status and the S3 alternative.

⚠️ Breaking changes / migration

  • Ingest into PRR before publishing to OSC. Without stac_catalog_s3_root, publish fails if the PRR collection doesn't exist.
  • dataset_status → osc_status. It still works for now with a warning; please rename it in existing configs.
  • dataset_id must be given inside items_config. A top-level dataset_id is ignored.
  • Existing configs that set stac_catalog_s3_root keep publishing to S3. Remove it to link PRR instead.

Testing

  • 165 unit tests pass locally. New tests cover:
    • the PRR tool (output tree, relative asset hrefs, required-field validation)
    • the PRR collection links, including a missing PRR collection and an S3 root skipping PRR
    • publish not writing to S3 when no root is set
    • the S3-item Zarr lookup and self link
    • the osc_status fallback and items_config validation
  • Checked against the live PRR collection pikart-atmospheric-river-catalog-v1-1: both the child and via links return 200.

Notes for reviewers

  • The OSC collection now depends on the PRR collection URL layout. For the S3 option, the item also depends on the PRR item's assets.zarr-data.href.
  • ruff format --check flags a few lines in dataset_stac_generator.py and test_custom_xrlint_rules.py that were already unformatted before this PR.

@codecov

codecov Bot commented Aug 5, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 97.11359% with 31 lines in your changes missing coverage. Please review.
✅ Project coverage is 91.69%. Comparing base (e0e26db) to head (8a272cc).

Files with missing lines Patch % Lines
deep_code/utils/dataset_stac_generator.py 95.57% 15 Missing ⚠️
deep_code/cli/prr.py 0.00% 8 Missing ⚠️
deep_code/tools/publish.py 94.11% 3 Missing ⚠️
deep_code/cli/generate_config.py 0.00% 2 Missing ⚠️
deep_code/cli/main.py 0.00% 2 Missing ⚠️
deep_code/tests/tools/test_prr.py 99.09% 1 Missing ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##             main      #23      +/-   ##
==========================================
+ Coverage   88.02%   91.69%   +3.66%     
==========================================
  Files          25       29       +4     
  Lines        2339     3238     +899     
==========================================
+ Hits         2059     2969     +910     
+ Misses        280      269      -11     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@TejasMorbagal TejasMorbagal changed the title Tejas morbagal/generate prr collection Add PRR collection generation; point OSC item assets to PRR (v0.2.0) Sep 25, 2026
Fix AttributeError on new variables; extract variable metadata once so GCMD URLs are prompted only once, Link workflow/experiment records to osc_project instead of always DeepESDL

experiment_links = []
for link in links:
if (

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

but it adds any other link if not in _EXPERIMENT_JSON_ONLY_RELS. Is this okay?

Comment on lines +250 to 263
@staticmethod
def _get_temporal_extent(dataset: xr.Dataset) -> TemporalExtent:
"""Extract temporal extent from the dataset."""
if "time" in self.dataset.coords:
dataset = dataset
if "time" in dataset.coords:
try:
# Convert the time bounds to datetime objects
time_min = pd.to_datetime(
self.dataset.time.min().values
).to_pydatetime()
time_max = pd.to_datetime(
self.dataset.time.max().values
).to_pydatetime()
time_min = pd.to_datetime(dataset.time.min().values).to_pydatetime()
time_max = pd.to_datetime(dataset.time.max().values).to_pydatetime()
return TemporalExtent([[time_min, time_max]])
except Exception as e:
raise ValueError(f"Failed to parse temporal extent: {e}")
else:
raise ValueError("Dataset does not have a 'time' coordinate.")

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
@staticmethod
def _get_temporal_extent(dataset: xr.Dataset) -> TemporalExtent:
"""Extract temporal extent from the dataset."""
time_coords = [name for name in dataset.coords if name.startswith("time")]
if time_coords:
try:
time_min = min(
pd.to_datetime(dataset[name].min().values)
for name in time_coords
).to_pydatetime()
time_max = max(
pd.to_datetime(dataset[name].max().values)
for name in time_coords
).to_pydatetime()
return TemporalExtent([[time_min, time_max]])
except Exception as e:
raise ValueError(f"Failed to parse temporal extent: {e}")
else:
raise ValueError("Dataset does not have a 'time*' coordinate.")

extracting (and prompting for) the metadata again.
"""
if variables_metadata is None:
variables_metadata = self.get_variables_metadata(dataset)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think we should just return the dataset variable names as is for better trackability & findability. From the OSC STAC extension this is valid. Something like variable_ids = list(dataset.data_vars). Not sure what to do with variables_metadata.

Comment thread deep_code/utils/dataset_stac_generator.py Outdated
osc_extension.osc_region = self.osc_region
osc_extension.osc_variables = variables
osc_extension.osc_missions = self.osc_missions
osc_extension.cf_parameter = self.cf_params or [{"name": self.collection_id}]

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The default is still very weird to me. Also Uni Leipzig DeepFeatures are not happy with this. Shall we maybe check for standard_name in the attrs for each variable. IF non is given, take "unknown". Then with the list in osc:varaibles one would at least have a mapping via the list index.

Also just a list of variable names is not correct I think. Here is an example: https://github.com/stac-extensions/cf/blob/v0.2.0/examples/collection.json

And there is already a newer version: https://github.com/stac-extensions/cf/blob/main/examples/collection.json

Maybe it would be a good idea, to make it optional.

Comment thread CHANGES.md
- `osc_project` is now omitted from `OscDatasetStacGenerator` when not provided, preserving the callee's default instead of passing `None`.

## Changes in 0.1.10 (in Development)
## Changes in 0.2.0 (in Development)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Log change is quiet massive. Maybe summarize a bit.

Co-authored-by: Konstantin Ntokas <38956538+konstntokas@users.noreply.github.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants