Piddiplatsch

Persistent identifiers for ESGF-NG publication records

Carsten Ehbrecht, Katharina Berger (DKRZ)

What is Piddi?

Piddiplatsch maintains persistent identifiers (PIDs) for ESGF publication records.

  • Named after Pittiplatsch, the curious TV puppet known as Pitti.
  • PID gives “Piddiplatsch” its wordplay.
  • piddi is the command-line tool, echoing “Pitti”.

Curious by nature. Persistent by design.

Pittiplatsch on Wikipedia

Pittiplatsch statue sitting on a railing in Erfurt

Pittiplatsch statue in Erfurt. Photo: Freddy123Hamster, CC BY-SA 4.0.

Piddi and ESGF-NG

Project plugins map STAC publication records to validated Handle records.

  • Connects the ESGF-NG publication stream (with Kafka) to persistent identifier services.
  • Prepares dataset and file records, including their relationships.
  • Built-in plugins: CMIP6, CMIP6Plus, CMIP7, and CORDEX-CMIP6.
  • Used in production at DKRZ; open source under Apache-2.0.

CLI: harvest → map + validation → publish, linked by JSONL.

ESGF/piddiplatsch

Piddi in ESGF-NG

Piddi maintains PIDs alongside catalog indexing; data files stay on data nodes.

Dataset and file identifiers

Dataset Handle

  • Identifies a published dataset version.
  • Records metadata and access information.
  • Links to its file Handles with HAS_PARTS.

File Handle

  • Uses the file’s tracking identifier.
  • Records file metadata and checksums.
  • Links to the dataset with IS_PART_OF.

Project adapters resolve source PIDs; shared mapping builds the Handle records.

Project plugins: shared workflow, project rules

  • Plugins own project identity, PID fields, and project-specific behavior.
  • Core shares routing, mapping machinery, validation, and JSONL persistence.
  • Select one, several, or all built-in plugins; each event reaches at most one.

Three commands, connected by JSONL

Command Input Work and output
harvest Kafka events Save raw messages as JSONL
map Raw JSONL Route to plugins, map + validate → Handle JSONL
publish Handle JSONL Write to Handle Service → receipt JSONL
piddi harvest --limit 100
piddi map --date last --project cmip6
piddi publish --project cmip6 --date last --limit 100

JSONL = one JSON record per line: inspect, retain, and replay each stage.

Using Piddiplatsch

Configure Kafka and Handle credentials in custom.toml, then check the setup:

piddi config validate
piddi config explain --project cmip6

Prepare a small batch and inspect its Handle records:

piddi harvest --limit 100
piddi map --date last --project cmip6

Publish up to 100 prepared Handles when ready:

piddi publish --project cmip6 --date last --limit 100

Outlook: a STAC lookup DB for Rook

Proposed extension: reuse Piddi’s plugin architecture beyond the PID service.

  • Origin: Piddi was built to maintain persistent identifiers via Handle services.
  • Extension: a Rook plugin could maintain dataset-to-site and local-asset lookups.
  • Benefit: the broker could choose a processing site using a local database.

Summary

Shared core, project plugins, replayable JSONL.

CMIP6 · CMIP6Plus · CMIP7 · CORDEX-CMIP6

github.com/ESGF/piddiplatsch

Appendix: configuration

Packaged defaults → site configuration → local override

src/piddiplatsch/config/default.toml
/etc/piddi/piddi.toml
./custom.toml
piddi config validate
piddi config explain --project cmip6
  • --config PATH replaces the final local override.
  • Named Handle profiles hold service settings and credentials.
  • Keep real credentials in ignored local configuration.

Appendix: consumer groups

Kafka assigns partitions before Piddi filters projects.

Process arrangement Consumer group choice
One process, several projects One group
Separate processes for different projects Distinct groups

Sharing a group between separate CMIP6 and CMIP7 consumers can send an event to the process that filters it out.

Changing project selection requires an explicit offset or replay decision.

Appendix: output and recovery

outputs/
  dump/dump_messages_<date>.jsonl
  <project>/
    handles/handles_<date>.jsonl
    failures/r<N>/failed_items_<date>.jsonl
    skipped/skipped_items_<date>.jsonl
  published/<run-scoped receipt>.jsonl
  • map reprocesses raw dumps with the chosen project selection.
  • retry writes retry_handles_<timestamp>.jsonl in the project’s handles folder.
  • publish preserves its inputs and reports each attempted Handle’s outcome.

Appendix: recovery with retry

Remap a saved failure file after fixing the underlying problem:

piddi retry outputs/cmip6/failures/r0/failed_items_2026-09-22.jsonl

Inspect the recovered Handle JSONL, then publish the path printed by retry:

# Example output path; use the timestamp from your retry run.
piddi publish --project cmip6 \
  outputs/cmip6/handles/retry_handles_2026-09-22_14-30-00-123456.jsonl

retry prepares a separate batch; publication is explicit. Logs and JSONL receipts record publication outcomes.

Appendix: service mode and Ansible

Continuous harvest + map: the systemd service prepares Handle JSONL.

# Command run by the service (publication remains a separate step).
piddi --config /etc/piddi/piddi.toml --silent consume

On a prepared Linux host, configure custom.toml, then deploy:

cd /opt/piddiplatsch
make deploy  # create/update .conda, install Piddi, run Ansible

After a manual trial, set piddi_enable_service: true in deploy/ansible/custom.yml, then enable and inspect the service:

make play  # Ansible-only configuration and service startup
systemctl status piddi
piddi top