Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
39 changes: 39 additions & 0 deletions .github/workflows/mirror-publish.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,39 @@
name: Publish an existing Dataverse draft

on:
workflow_dispatch:
inputs:
persistent_id:
description: "Persistent id (DOI) of the existing draft to publish, e.g. doi:10.34894/EXAMPLE"
required: true
version_type:
description: "Dataverse publish version bump"
required: true
default: "major"
type: choice
options:
- major
- minor

jobs:
mirror-publish:
runs-on: ubuntu-latest
environment: dataverse
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: "3.12"
- name: Install fwl-io
run: pip install -e .
- name: Publish the draft
env:
DATAVERSE_TOKEN: ${{ secrets.DATAVERSE_TOKEN }}
# Bind the user inputs to environment variables rather than
# interpolating them into the shell script, so a crafted persistent id
# cannot inject commands into a job that holds the token.
PERSISTENT_ID: ${{ inputs.persistent_id }}
VERSION_TYPE: ${{ inputs.version_type }}
run: |
fwl-io mirror-publish "${PERSISTENT_ID}" \
--version-type "${VERSION_TYPE}"
4 changes: 4 additions & 0 deletions docs/How-to/mirror_dataset.md
Original file line number Diff line number Diff line change
Expand Up @@ -22,6 +22,10 @@ Mirroring runs from the **Mirror a Zenodo deposit to Dataverse** GitHub Actions

Commit that change in a pull request, like any other data change.

## Publishing a reviewed draft

A draft created with **publish** unchecked stays private until it is published. Run the **Publish an existing Dataverse draft** GitHub Actions workflow, supplying the draft's persistent id (the DOI printed by the mirror run, with a `doi:` prefix, for example `doi:10.34894/XXXXXX`). It only publishes; it never creates a dataset, so it cannot mint a duplicate one. Add the DOI to the manifest as in step 3 above once it is published.

## What the mirror does

For the given Zenodo version DOI, the mirror downloads and checksum-verifies every file, creates a Dataverse dataset whose title, authors, and description come from the Zenodo record (with a note recording the source DOI), uploads the files byte-identically with tabular ingest disabled, and publishes the dataset unless asked not to. A concept DOI is rejected, so the mirror always tracks a specific pinned deposit.
Expand Down
11 changes: 10 additions & 1 deletion docs/Reference/cli.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
# CLI reference

The `fwl-io` command has six subcommands. Failures are reported as concise messages on stderr (never a traceback) and exit with status 1; success exits 0. `sync` and `fetch` aggregate per-dataset failures into a multi-line report, and a download failure lists every mirror attempt.
The `fwl-io` command has seven subcommands. Failures are reported as concise messages on stderr (never a traceback) and exit with status 1; success exits 0. `sync` and `fetch` aggregate per-dataset failures into a multi-line report, and a download failure lists every mirror attempt.

## fwl-io sync

Expand Down Expand Up @@ -68,6 +68,15 @@ DATAVERSE_TOKEN=... fwl-io mirror <zenodo-doi> --collection <alias> \

Mirrors a pinned Zenodo deposit to a Dataverse collection: it downloads and checksum-verifies the deposit's files, creates a matching Dataverse dataset with citation metadata taken from the Zenodo record, uploads the files byte-identically (tabular ingest disabled), and by default publishes the dataset, then prints the Dataverse DOI to add to the consuming manifest. The API token is read from the `DATAVERSE_TOKEN` environment variable, never a command-line argument. A contact email (`--contact-email`) is required to create a dataset; only `--dry-run`, which makes no Dataverse writes, is exempt. `--subject` is validated by the server when the dataset is created, so a value outside the target installation's citation vocabulary is rejected then. `--dry-run` performs the download and metadata mapping only, making no Dataverse changes; `--no-publish` leaves the created dataset as a private draft. See [Mirror a deposit to Dataverse](../How-to/mirror_dataset.md).

## fwl-io mirror-publish

```bash
DATAVERSE_TOKEN=... fwl-io mirror-publish <persistent-id> \
[--dataverse-url URL] [--version-type VERSION_TYPE]
```

Publishes an existing Dataverse draft by its persistent id: it never creates a dataset, so it is the second step of a create-draft-then-publish workflow, run once a draft created by `fwl-io mirror --no-publish` has been reviewed. `<persistent-id>` must be of the form `doi:<prefix>/<suffix>`, for example `doi:10.34894/EXAMPLE`. The API token is read from the `DATAVERSE_TOKEN` environment variable, never a command-line argument. `--version-type` is `major` by default and accepts only `major` or `minor`. Fails clearly if the dataset is already published or the persistent id does not resolve to a draft. See [Mirror a deposit to Dataverse](../How-to/mirror_dataset.md).

## fwl-io --version

Prints the installed version.
38 changes: 36 additions & 2 deletions src/fwl_io/cli.py
Original file line number Diff line number Diff line change
@@ -1,4 +1,5 @@
"""Command-line interface: ``fwl-io sync | list | fetch | check | relocate | mirror``.
"""Command-line interface for ``fwl-io``: sync, list, fetch, check, relocate, mirror,
mirror-publish.

Failures from the package's own error types exit with status 1 and a
one-line message on stderr instead of a traceback.
Expand All @@ -12,6 +13,8 @@

from fwl_io import __version__

DEFAULT_DATAVERSE_URL = 'https://dataverse.nl'


def _cmd_sync(args: argparse.Namespace) -> int:
from fwl_io.sync import ZENODO_API, sync_manifest
Expand Down Expand Up @@ -99,6 +102,23 @@ def _cmd_mirror(args: argparse.Namespace) -> int:
return 0


def _cmd_mirror_publish(args: argparse.Namespace) -> int:
from fwl_io.mirror import publish_existing_dataverse_draft

token = os.environ.get('DATAVERSE_TOKEN', '')
if not token:
print('fwl-io: set DATAVERSE_TOKEN to publish', file=sys.stderr)
return 1
publish_existing_dataverse_draft(
args.persistent_id,
dataverse_url=args.dataverse_url,
token=token,
version_type=args.version_type,
)
print(f'published {args.persistent_id}')
return 0


def main(argv: list[str] | None = None) -> int:
parser = argparse.ArgumentParser(
prog='fwl-io',
Expand Down Expand Up @@ -141,7 +161,7 @@ def main(argv: list[str] | None = None) -> int:
p_mirror.add_argument('zenodo_doi', help='Zenodo version DOI to mirror')
p_mirror.add_argument('--collection', required=True, help='target Dataverse collection alias')
p_mirror.add_argument(
'--dataverse-url', default='https://dataverse.nl', help='Dataverse base URL'
'--dataverse-url', default=DEFAULT_DATAVERSE_URL, help='Dataverse base URL'
)
p_mirror.add_argument('--contact-name', default='PROTEUS Framework', help='dataset contact')
p_mirror.add_argument(
Expand All @@ -160,6 +180,20 @@ def main(argv: list[str] | None = None) -> int:
)
p_mirror.set_defaults(func=_cmd_mirror)

p_mirror_publish = sub.add_parser(
'mirror-publish', help='publish an existing Dataverse draft (never creates a dataset)'
)
p_mirror_publish.add_argument('persistent_id', help='persistent id (DOI) of the draft')
p_mirror_publish.add_argument(
'--dataverse-url', default=DEFAULT_DATAVERSE_URL, help='Dataverse base URL'
)
p_mirror_publish.add_argument(
'--version-type',
default='major',
help="Dataverse publish version bump: 'major' or 'minor'",
)
p_mirror_publish.set_defaults(func=_cmd_mirror_publish)

args = parser.parse_args(argv)
try:
return args.func(args)
Expand Down
64 changes: 62 additions & 2 deletions src/fwl_io/mirror.py
Original file line number Diff line number Diff line change
@@ -1,11 +1,14 @@
"""Mirror a pinned Zenodo deposit to a Dataverse.nl collection.

Zenodo is the primary source of every dataset; Dataverse is a download
mirror used as the second link in the fetch fallback chain. This module
mirror used as the second link in the fetch fallback chain. :func:`mirror_to_dataverse`
takes a Zenodo version DOI, downloads and checksum-verifies its files, then
creates a matching Dataverse dataset, uploads the files byte-identically,
and (optionally) publishes it, printing the Dataverse DOI to add to the
consuming manifest.
consuming manifest. Called with ``publish=False``, it leaves the created
dataset as a private draft instead; :func:`publish_existing_dataverse_draft`
is the second step of that workflow, publishing an existing draft by its
persistent id without ever creating a dataset.

The Dataverse writes go through the native API
(https://guides.dataverse.org/en/latest/api/native-api.html):
Expand Down Expand Up @@ -410,3 +413,60 @@ def mirror_to_dataverse(
)
raise
return persistent_id


def publish_existing_dataverse_draft(
persistent_id: str,
*,
dataverse_url: str,
token: str,
version_type: str = 'major',
) -> None:
"""Publish an existing Dataverse draft dataset by its persistent id.

This never calls :meth:`DataverseClient.create_dataset`, so it cannot
mint a duplicate dataset: it is the second step of a create-draft ->
review -> publish workflow, run once the draft created by
:func:`mirror_to_dataverse` (with ``publish=False``) has been reviewed.

Parameters
----------
persistent_id : str
Persistent id (DOI) of the existing draft, for example
``'doi:10.34894/EXAMPLE'``.
dataverse_url : str
Base URL of the Dataverse installation (for example
``https://dataverse.nl``).
token : str
Dataverse API token.
version_type : str
Dataverse publish version bump: ``'major'`` or ``'minor'``.

Raises
------
ValueError
If ``persistent_id`` is not of the form ``'doi:<prefix>/<suffix>'``,
or ``version_type`` is not ``'major'`` or ``'minor'``.
DataverseError
If the publish request fails: for example the dataset is already
published, does not exist, or the server returns a non-2xx status.
"""
if version_type not in ('major', 'minor'):
raise ValueError(
f"{version_type!r} is not a valid Dataverse version type: use 'major' or 'minor'"
)
prefix, sep, suffix = persistent_id.removeprefix('doi:').partition('/')
if (
not persistent_id.startswith('doi:')
or not sep
or not prefix
or not suffix
or any(ch.isspace() for ch in persistent_id)
):
raise ValueError(
f'{persistent_id!r} is not a Dataverse persistent id of the form '
"'doi:<prefix>/<suffix>'"
)
client = DataverseClient(dataverse_url, token)
client.publish(persistent_id, version_type=version_type)
log.info('published %s', persistent_id)
Loading
Loading