Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions docs/Explanations/manifests.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,6 +25,7 @@ Validation at load time:
- `extract`, when present, must be `"tar"` or `"zip"`.
- A dataset table must not contain sub-tables, and arrays of tables are rejected; ambiguous structures fail loudly instead of being silently dropped.
- A dataset table declares only the fields above, and a grouping level declares none: anything else raises `ManifestSchemaError`, naming the field and the manifest schema this fwl-io implements. A model ships its manifest with its own code, so a manifest can be newer than the installed fwl-io; ignoring an unknown field silently would leave the manifest asking for something it never gets, and a `required_by` written one level above its dataset would leave the dataset claiming no model needs it. The message names the action that fits the case: move a dataset field that sits too high, delete a field this fwl-io no longer takes, and for a name it does not know at all, check the spelling or upgrade. Two things are outside the check. A scalar at the manifest root is reserved for a manifest's own settings and is ignored, unless it names a dataset field or `subdir`. And a table is recognised as a dataset by its `zenodo` key, so a misspelt `zenodo` is reported as a table with no pin rather than as an unknown field.
- A manifest may declare the schema it was written against with a root `manifest_schema = <n>`. It is optional, and a manifest that declares one is held to it: only the schema the installed fwl-io implements is accepted. A higher number means the reader is too old, so the error says to upgrade. A lower one means the manifest was written for a schema that stopped loading when the number rose, so the error names both numbers and points here. Accepting only the implemented number is what sharpens the rest: an unknown field in a manifest that declares its schema is reported as a misspelling alone, with no second reading to weigh. The value must be a whole number of at least 1; `true` is rejected rather than read as 1. A table *named* `manifest_schema` is an ordinary directory level, as with `subdir`, and the key written inside a table is reported as misplaced rather than misspelt.
- A manifest that fails to load takes its whole provider with it: `discover_manifests` skips that package and logs a warning, so its other datasets disappear from the result too. Use `fwl-io list` to see the error.
- A `subdir` field is rejected on any table, dataset or grouping level, and at the manifest root: the location comes from the key. A table *named* `subdir` is an ordinary directory level.
- Within one manifest, two keys that differ only in case are rejected: they would share one directory and one registry file on a case-insensitive filesystem. Two installed packages declaring keys that collide is a separate check, tracked in [#18](https://github.com/FormingWorlds/fwl-io/issues/18).
Expand All @@ -37,6 +38,8 @@ Validation at load time:

An error raised while reading a manifest names the schema the running code implements, so a mismatch between a manifest and an installed fwl-io can be placed against this table.

A manifest that declares `manifest_schema` is checked against it directly, and only the implemented number is accepted. Incrementing the schema is therefore a breaking change for any manifest that declares the old one, which is the intent: the number rises precisely when manifests written for the previous one stop loading, so they should fail at the increment with a message naming both numbers rather than part way through a load with a message about some individual field. A manifest that declares nothing is read on a best-effort basis, as before.

## Archive datasets

A deposit packaged as a single archive sets `extract = "tar"` or `"zip"`. Its registry lists the one archive file and its checksum; the fetcher downloads and verifies the archive, then extracts the members into the dataset directory and discards the archive, so consumers see the extracted tree rather than a tarball. Extraction is staged and the tree is moved into place atomically, so an interrupted fetch never leaves a half-populated dataset, and any member that escapes the directory (an absolute path or a `..` component) or is not a plain file or directory (a symlink, hardlink, or device node) is rejected before anything is written.
Expand Down
4 changes: 4 additions & 0 deletions docs/How-to/add_dataset.md
Original file line number Diff line number Diff line change
Expand Up @@ -20,12 +20,16 @@ Choose the manifest:
Add a table for the dataset:

```toml
manifest_schema = 1

[interior.eos.wolf_bower_2018]
name = "Wolf & Bower (2018) MgSiO3 equation of state"
zenodo = "10.5281/zenodo.1234567"
required_by = ["aragog", "zalmoxis", "spider"]
```

The root `manifest_schema` names the schema the file is written against. It is optional and worth declaring: it lets fwl-io tell a manifest written for a different schema from a misspelt field, so a load failure names the one that applies instead of offering both. Declaring it also means the manifest has to be updated when the schema number rises, which is the point, since that is when manifests written for the previous number stop loading. The current schema is in the [schema versions](../Explanations/manifests.md#schema-versions) table.

The dotted key is the location below `FWL_DATA`, so this dataset lands in `interior/eos/wolf_bower_2018/r<record-id>`, the version directory named for its Zenodo record. Choose the key to follow the [target layout](../Explanations/manifests.md#the-fwl_data-layout), using only letters, digits, `_` and `-` per segment, each starting with a letter, digit or `_`. `required_by` lists the models whose `fwl-io fetch <model>` should include this dataset.

If the deposit is a single archive that consumers expect unpacked, add `extract = "tar"` or `extract = "zip"`; the archive is downloaded, checksum-verified, and unpacked into the dataset directory. See [Archive datasets](../Explanations/manifests.md#archive-datasets).
Expand Down
6 changes: 5 additions & 1 deletion src/fwl_io/data/shared_manifest.toml
Original file line number Diff line number Diff line change
@@ -1,6 +1,8 @@
# Datasets shared by several models of the PROTEUS ecosystem.
#
# One TOML table per dataset:
# One TOML table per dataset, below an optional schema declaration:
#
# manifest_schema = 1
#
# [group.dataset_key]
# name = "Human-readable dataset name"
Expand All @@ -23,3 +25,5 @@
# No dataset is shared across several models yet, so this manifest declares
# none. The Baraffe stellar tracks ship with the MORS package, which owns their
# manifest and registry.

manifest_schema = 1
116 changes: 107 additions & 9 deletions src/fwl_io/manifest.py
Original file line number Diff line number Diff line change
Expand Up @@ -8,13 +8,23 @@

Manifest schema, one table per dataset, identified by its ``zenodo`` key::

manifest_schema = 1 # optional, see below

[interior.eos.wolf_bower_2018]
name = "Wolf & Bower (2018) MgSiO3 equation of state"
zenodo = "10.5281/zenodo.1234567" # version DOI, never a concept DOI
dataverse = "10.34894/ABCDEF" # optional download mirror
required_by = ["aragog", "zalmoxis", "spider"]
extract = "tar" # optional: unpack a single-archive deposit

The optional root ``manifest_schema`` names the schema the file was written
against. A manifest that declares one is held to it: only the schema the
installed fwl-io implements is accepted, a higher number meaning the reader is
too old and a lower one meaning the manifest was written for a schema that
stopped loading when the number rose. That is what sharpens the diagnosis
elsewhere, since an unknown field in such a manifest can only be a
misspelling.

The dotted table key is the dataset location below the data root: the table
above resolves into ``interior/eos/wolf_bower_2018``. Key segments are
restricted to letters, digits, ``_`` and ``-``, each starting with a letter,
Expand Down Expand Up @@ -79,6 +89,12 @@
# checkout's recorded version does not.
_MANIFEST_SCHEMA = 1

# The root key by which a manifest states the schema it was written against.
# Declaring it is optional and turns an ambiguous diagnosis into a definite
# one: a manifest that names a schema this code does not implement is from the
# future and says so, rather than being reported as a possible typo.
_SCHEMA_KEY = 'manifest_schema'


class ManifestSchemaError(ValueError):
"""A manifest and the installed fwl-io disagree about the manifest schema.
Expand All @@ -101,18 +117,86 @@ def _reading_version() -> str:
return f'manifest schema {_MANIFEST_SCHEMA} (distribution {distribution})'


def _unknown_field_error(what: str) -> ManifestSchemaError:
def _unknown_field_error(what: str, declared_schema: int | None = None) -> ManifestSchemaError:
"""Build the error for a field this fwl-io does not know.

Both readings are offered because both are common: a misspelt field, and a
manifest written against a schema newer than the fwl-io reading it.
Without a declared schema both readings are offered, because both are
common: a misspelt field, and a manifest written against a schema newer
than the fwl-io reading it. A manifest that declares one has ruled the
second reading out by the time this is reached, since the reader admits
only the schema this code implements, so the message names the typo alone.
The equality is restated here rather than assumed, so that relaxing the
reader later cannot silently sharpen the message for a schema it should
not apply to.
"""
if declared_schema == _MANIFEST_SCHEMA:
return ManifestSchemaError(
f'{what}. The manifest declares {_SCHEMA_KEY} {declared_schema}, the schema '
f'this fwl-io implements, so the field is misspelt rather than newer than '
f'this code.'
)
return ManifestSchemaError(
f'{what}. This fwl-io reads {_reading_version()}: check the spelling, or '
f'upgrade fwl-io if the manifest was written against a newer schema.'
)


def _read_declared_schema(tree: dict) -> int | None:
"""Return the schema a manifest declares at its root, if it declares one.

Only the schema this code implements is admitted. A higher number means the
reader is too old. A lower one means the manifest is written for a schema
that, by the rule the number follows, stopped loading when it was
incremented; refusing it fails at the increment rather than part way
through a load. Either way the message names both numbers. Raises too when
the value is not a schema number at all.
"""
if _SCHEMA_KEY not in tree:
return None
declared = tree[_SCHEMA_KEY]
# A table of this name is a directory level like any other, the same rule
# `subdir` follows: the reserved name applies to the scalar, not to a
# dataset an author happens to have called this. Leave it to the walk.
if isinstance(declared, dict) or (
isinstance(declared, list) and any(isinstance(item, dict) for item in declared)
):
return None
# bool is an int subclass, and `manifest_schema = true` is a mistake worth
# naming rather than reading as schema 1.
if isinstance(declared, bool) or not isinstance(declared, int) or declared < 1:
raise ManifestSchemaError(
f'the manifest root declares {_SCHEMA_KEY} {declared!r}; it must be a whole '
f'number of at least 1, naming the manifest schema the file was written '
f'against.'
)
if declared > _MANIFEST_SCHEMA:
raise ManifestSchemaError(
f'the manifest declares {_SCHEMA_KEY} {declared}, but this fwl-io reads '
f'{_reading_version()}: upgrade fwl-io to read this manifest.'
)
if declared < _MANIFEST_SCHEMA:
raise ManifestSchemaError(
f'the manifest declares {_SCHEMA_KEY} {declared}, but this fwl-io reads '
f'manifest schema {_MANIFEST_SCHEMA}: the schema number rises when a '
f'manifest written for the previous one stops loading, so update the '
f'manifest against the schema versions table in the manifests '
f'documentation.'
)
return declared


def _misplaced_schema_error(what: str) -> ManifestSchemaError:
"""Build the error for the reserved schema key written below the root.

The name is spelt correctly and sits at the wrong level, so reporting it as
an unknown field would give the one reading that is certainly wrong.
"""
return ManifestSchemaError(
f'{what}. {_SCHEMA_KEY!r} is a manifest-root key naming the schema the file was '
f'written against: move the line above the first table.'
)


def _misplaced_field_error(what: str) -> ManifestSchemaError:
"""Build the error for a known field declared outside a dataset table."""
return ManifestSchemaError(
Expand Down Expand Up @@ -195,7 +279,9 @@ def _reject_declared_subdir(where: str, table: dict, derived: str | None = None)
)


def _walk_tables(tree: dict, prefix: str = '') -> list[tuple[str, dict]]:
def _walk_tables(
tree: dict, prefix: str = '', declared_schema: int | None = None
) -> list[tuple[str, dict]]:
"""Return (dotted-key, table) pairs for the dataset tables of a manifest.

A table is a dataset when it carries the ``zenodo`` key. Grouping tables
Expand All @@ -217,12 +303,18 @@ def _walk_tables(tree: dict, prefix: str = '') -> list[tuple[str, dict]]:
)
raise _misplaced_field_error(where)
if prefix:
if name == _SCHEMA_KEY:
raise _misplaced_schema_error(
f'grouping table {prefix.rstrip(".")!r} declares {name!r}'
)
raise _unknown_field_error(
f'grouping table {prefix.rstrip(".")!r} declares {name!r}, and a '
f'grouping level takes no fields'
f'grouping level takes no fields',
declared_schema,
)
# A root scalar that names no dataset field is a manifest's own
# setting.
# setting. The schema declaration is one of those, already read
# and validated before the walk began.
continue
dotted = f'{prefix}{name}'
_validate_key_segment(name, prefix.rstrip('.'))
Expand All @@ -237,15 +329,18 @@ def _walk_tables(tree: dict, prefix: str = '') -> list[tuple[str, dict]]:
if has_subtables:
raise ValueError(f'dataset {dotted!r}: dataset tables must not contain sub-tables')
unknown = sorted(set(value) - _DATASET_FIELDS)
if _SCHEMA_KEY in unknown:
raise _misplaced_schema_error(f'dataset {dotted!r} declares {_SCHEMA_KEY!r}')
if unknown:
raise _unknown_field_error(
f'dataset {dotted!r} declares {", ".join(repr(f) for f in unknown)}, '
f'which is not a dataset field '
f'(known fields: {", ".join(sorted(_DATASET_FIELDS))})'
f'(known fields: {", ".join(sorted(_DATASET_FIELDS))})',
declared_schema,
)
leaves.append((dotted, value))
elif has_subtables:
leaves.extend(_walk_tables(value, prefix=f'{dotted}.'))
leaves.extend(_walk_tables(value, prefix=f'{dotted}.', declared_schema=declared_schema))
else:
raise ValueError(
f'table {dotted!r} has no "zenodo" key, so it is neither a dataset nor a '
Expand All @@ -261,10 +356,13 @@ def load_manifest(path: str | Path) -> list[Dataset]:
with path.open('rb') as fh:
tree = tomllib.load(fh)

# Read the declaration first: a manifest above this code's schema cannot be
# judged by this code's rules, so it must be refused before they are applied.
declared_schema = _read_declared_schema(tree)
_reject_declared_subdir('the manifest root', tree)
datasets: list[Dataset] = []
folded: dict[str, str] = {}
for key, table in _walk_tables(tree):
for key, table in _walk_tables(tree, declared_schema=declared_schema):
clash = folded.setdefault(key.lower(), key)
if clash != key:
raise ValueError(
Expand Down
Loading
Loading