Add PRR collection generation; point OSC item assets to PRR (v0.2.0) - #23
TejasMorbagal wants to merge 59 commits into
Conversation
Codecov Report❌ Patch coverage is Additional details and impacted files@@ Coverage Diff @@
## main #23 +/- ##
==========================================
+ Coverage 88.02% 91.69% +3.66%
==========================================
Files 25 29 +4
Lines 2339 3238 +899
==========================================
+ Hits 2059 2969 +910
+ Misses 280 269 -11 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
Konstntokas/update prr
Co-authored-by: Tejas Morbagal Harish <tejas.morbagalharish@brockmann-consult.de>
Co-authored-by: Tejas Morbagal Harish <tejas.morbagalharish@brockmann-consult.de>
Co-authored-by: Tejas Morbagal Harish <tejas.morbagalharish@brockmann-consult.de>
Correct extent in bbox, polygon and cube:dimensions
Add missing parts to be conformant with PRR specs
…eld to osc_status
Fix AttributeError on new variables; extract variable metadata once so GCMD URLs are prompted only once, Link workflow/experiment records to osc_project instead of always DeepESDL
|
|
||
| experiment_links = [] | ||
| for link in links: | ||
| if ( |
There was a problem hiding this comment.
but it adds any other link if not in _EXPERIMENT_JSON_ONLY_RELS. Is this okay?
| @staticmethod | ||
| def _get_temporal_extent(dataset: xr.Dataset) -> TemporalExtent: | ||
| """Extract temporal extent from the dataset.""" | ||
| if "time" in self.dataset.coords: | ||
| dataset = dataset | ||
| if "time" in dataset.coords: | ||
| try: | ||
| # Convert the time bounds to datetime objects | ||
| time_min = pd.to_datetime( | ||
| self.dataset.time.min().values | ||
| ).to_pydatetime() | ||
| time_max = pd.to_datetime( | ||
| self.dataset.time.max().values | ||
| ).to_pydatetime() | ||
| time_min = pd.to_datetime(dataset.time.min().values).to_pydatetime() | ||
| time_max = pd.to_datetime(dataset.time.max().values).to_pydatetime() | ||
| return TemporalExtent([[time_min, time_max]]) | ||
| except Exception as e: | ||
| raise ValueError(f"Failed to parse temporal extent: {e}") | ||
| else: | ||
| raise ValueError("Dataset does not have a 'time' coordinate.") |
There was a problem hiding this comment.
| @staticmethod | |
| def _get_temporal_extent(dataset: xr.Dataset) -> TemporalExtent: | |
| """Extract temporal extent from the dataset.""" | |
| time_coords = [name for name in dataset.coords if name.startswith("time")] | |
| if time_coords: | |
| try: | |
| time_min = min( | |
| pd.to_datetime(dataset[name].min().values) | |
| for name in time_coords | |
| ).to_pydatetime() | |
| time_max = max( | |
| pd.to_datetime(dataset[name].max().values) | |
| for name in time_coords | |
| ).to_pydatetime() | |
| return TemporalExtent([[time_min, time_max]]) | |
| except Exception as e: | |
| raise ValueError(f"Failed to parse temporal extent: {e}") | |
| else: | |
| raise ValueError("Dataset does not have a 'time*' coordinate.") |
| extracting (and prompting for) the metadata again. | ||
| """ | ||
| if variables_metadata is None: | ||
| variables_metadata = self.get_variables_metadata(dataset) |
There was a problem hiding this comment.
I think we should just return the dataset variable names as is for better trackability & findability. From the OSC STAC extension this is valid. Something like variable_ids = list(dataset.data_vars). Not sure what to do with variables_metadata.
| osc_extension.osc_region = self.osc_region | ||
| osc_extension.osc_variables = variables | ||
| osc_extension.osc_missions = self.osc_missions | ||
| osc_extension.cf_parameter = self.cf_params or [{"name": self.collection_id}] |
There was a problem hiding this comment.
The default is still very weird to me. Also Uni Leipzig DeepFeatures are not happy with this. Shall we maybe check for standard_name in the attrs for each variable. IF non is given, take "unknown". Then with the list in osc:varaibles one would at least have a mapping via the list index.
Also just a list of variable names is not correct I think. Here is an example: https://github.com/stac-extensions/cf/blob/v0.2.0/examples/collection.json
And there is already a newer version: https://github.com/stac-extensions/cf/blob/main/examples/collection.json
Maybe it would be a good idea, to make it optional.
| - `osc_project` is now omitted from `OscDatasetStacGenerator` when not provided, preserving the callee's default instead of passing `None`. | ||
|
|
||
| ## Changes in 0.1.10 (in Development) | ||
| ## Changes in 0.2.0 (in Development) |
There was a problem hiding this comment.
Log change is quiet massive. Maybe summarize a bit.
Co-authored-by: Konstantin Ntokas <38956538+konstntokas@users.noreply.github.com>
Summary
Adds a
generate-prr-collectioncommand that builds an ESA EarthCODE Project Results Repository (PRR) STAC collection from the existing dataset config. It also changes OSC publishing to link each dataset straight to its PRR collection, the same way other PRR-hosted OSC products are linked. The S3 STAC catalog becomes opt-in. The version goes to 0.2.0 because this includes breaking config and workflow changes (see below).Changes
New: PRR collection generation
deep-code generate-prr-collection dataset.yaml [-o DIR]and Python helperdeep_code.tools.prr.generate_prr_collection. It writes a self-containedCollection → Item → Assetstree to local files and needs no GitHub or S3 credentials.datacubeextension (cube:dimensions/cube:variablesread from the Zarr) and thefileextension (file:sizefor the store and.zmetadata). Its asset hrefs are relative (./{dataset_id}), because PRR ingests the Zarr next to the item. The Collection declares the OSC, Scientific, Processing, Themes and CF extensions plus the PRR-mandatory fields.osc_initiative,osc_missions,osc_contract_number,osc_project_website,osc_project_description,thumbnail,thumbnail_media_type,sci_doi,sci_citation,prr_output_dir. If PRR-required fields are missing, the command logs a warning and still runs.coord_positionoption (center|left|right) so the bbox and cube extents cover full grid cells.generate-configtemplates document the PRR fields.Changed: OSC collection links to PRR by default
stac_catalog_s3_rootis now optional. Without it, the OSC collection links to the dataset's PRR collection:child→https://eoresults.esa.int/stac/collections/{collection_id}via(title "Access") →https://eoresults.esa.int/browser/#/external/eoresults.esa.int/stac/collections/{collection_id}Nothing is written to S3 and no S3 credentials are needed. Publishing checks that the PRR collection exists and fails with a clear error if it doesn't.
For datasets not in PRR, setting
stac_catalog_s3_rootkeeps the S3 catalog + item behaviour. In that case:access_link, or else thezarr-dataasset looked up from the PRR item. Before, they were./{dataset_id}, which resolved to a path on S3 where no Zarr exists.selflink now matches where the file is written (…/{collection_id}/items/{item_id}.json, not…/{collection_id}/item.json).Config handling
dataset_statusrenamed toosc_status(defaultcompleted). The old key still works but logs a deprecation warning.items_configentry now gives a clear error (dataset_id is required in items_config entry 0.) instead of aKeyError.publishandgenerate-prr-collectionshare one parser (build_items_config).Housekeeping
X | None. Fixed a mutable default argument inogc_api_record.py.ruff==0.9.10for now.configuration.mdhas a new "Data links" section, andcli.md/python-api.mddescribe the PRR command, the PRR default,osc_statusand the S3 alternative.stac_catalog_s3_root,publishfails if the PRR collection doesn't exist.dataset_status→osc_status. It still works for now with a warning; please rename it in existing configs.dataset_idmust be given insideitems_config. A top-leveldataset_idis ignored.stac_catalog_s3_rootkeep publishing to S3. Remove it to link PRR instead.Testing
publishnot writing to S3 when no root is setselflinkosc_statusfallback anditems_configvalidationpikart-atmospheric-river-catalog-v1-1: both thechildandvialinks return 200.Notes for reviewers
assets.zarr-data.href.ruff format --checkflags a few lines indataset_stac_generator.pyandtest_custom_xrlint_rules.pythat were already unformatted before this PR.