Operations

Runtime output

Piddiplatsch writes daily JSONL files beneath consumer.output_dir:

Directory Contents
dump/ Original Kafka messages written by every harvest or consume run.
<plugin>/handles/ Handle records produced by that plugin. Direct REST/pyhandle publication writes the JSONL record before contacting the server; each line also includes project.
published/ One run-scoped JSONL receipt per publish run, containing every successful or failed Handle outcome.
<plugin>/skipped/ Records deferred after transient external failures.
<plugin>/failures/r<N>/ Records that failed processing, grouped by retry count.
skipped/, failures/r<N>/ Legacy or unresolved-project recovery records.

The raw dump is intentionally global and always written before routing. It preserves the consumed Kafka order and remains suitable for replay or investigation; creating filtered per-plugin dumps would lose that simple ordering guarantee.

JSONL Handle output is always enabled, including direct publication. If the audit record cannot be appended, that Handle is not sent to the server. A server-side failure leaves the JSONL record available for inspection or later publication.

Deferred publication writes each outcome immediately to a readable, unique published/published_<project>_handles_YYYY-MM-DD_HH-MM-SS.jsonl file using UTC and prints its path in the final CLI summary. If no project is known when the receipt opens, the filename is published_handles_...jsonl. If another run starts during the same second, _2, _3, and so on are appended. Parallel completion order may differ from input order; position, batch_index, source_file, and source_line provide stable ordering and provenance.

Staged commands

# Kafka -> raw JSONL only
piddi harvest --limit 100

# raw JSONL -> project-scoped Handle JSONL only
piddi map --project cmip6 --date 2026-08-27

# Handle JSONL -> REST Handle Service
piddi publish --project cmip6 --date 2026-08-27

# Kafka -> raw JSONL -> Handle JSONL (default production ingestion)
piddi consume

# All three stages in one process
piddi consume --publish

harvest --limit N stops normally after dumping N messages, which is useful for bounded tests against a live topic. map accepts files or directories plus --project, --all-projects, --limit, --offset, and --force. It never contacts Kafka or a Handle Service and does not modify its input dumps. It opens each source once and maps one record at a time in source order, without loading the selected files into memory. Offsets and limits apply across files and count nonblank input lines; source keys use physical line numbers. A malformed selected JSON record stops the run with its source location; output from earlier records remains written. The existing processing-error and transient-failure stop policies still apply.

The --date convenience accepts YYYY-MM-DD, today, yesterday, today-N, or last. map uses dump/dump_messages_<date>.jsonl; publish requires a project and uses <project>/handles/handles_<date>.jsonl. When neither an explicit path nor --date is supplied, both commands default to last. This selects the greatest valid date found in the relevant filenames, regardless of file modification time. Explicit paths remain available for both commands, and a path and --date are mutually exclusive.

Deferred publication

Publish completed files after mapping, using the project's configured Handle profile. Profiles contain the REST server, prefix, and credentials; see Handle configuration. Keep credentials in the ignored site configuration file.

# Publish yesterday's prepared CMIP6 Handles
piddi publish --project cmip6 --date yesterday

# Try the first 1,000 records of a specific completed file
piddi publish --project cmip6 --date 2026-08-25 --limit 1000

# Continue with the next 1,000 records, allowing transient retries
piddi publish --project cmip6 --date 2026-08-25 \
  --offset 1000 --limit 1000 --retries 3

Explicit inputs can be one file, several files, or a directory. Offsets count nonblank input lines across the selected files; limits cap attempted records, including failures. Publication never changes or deletes its inputs. Handle writes use overwrite semantics, so a completed, immutable input can be replayed after interruption. The command exits non-zero if any record fails.

--retries covers transient connection errors, timeouts, rate limiting, and server errors. The delay starts at one second and doubles up to 60 seconds; change its initial value with --retry-delay. Permanent client errors such as invalid credentials are not retried. --workers N enables concurrent requests for different Handles, while updates to the same Handle retain source order within the run.

publish --project NAME validates each Handle immediately before sending it. Missing or different projects are recorded as failures, while valid records continue. Without --project, the first parsed record selects the service configuration; subsequent records must match its project. Each JSONL file is opened once and processed in batches of at most 256 records. Writes for the same Handle stay in input order, including across batches. Malformed lines receive failure receipts; earlier writes are not rolled back. No staging database or preliminary scan is used, so progress has no known total and receipt batch_total is null. Parent references are sent as provided; receipt metadata comes only from the current record. The final summary provides per-project counts and up to 100 error messages; receipts retain all failures.

Each receipt line records the outcome, action, PID, URL, project, dataset, asset, source location, position, retry count, and error. Missing dataset or asset metadata remains null. batch_index and position refer to the entire selected input, independently of internal batch boundaries. Receipts are written regardless of log verbosity; INFO logging additionally records each publication outcome in the configured log file.

Real Handle service contract test

The opt-in live tests create uniquely named Handles and do not delete them, so use a disposable test prefix. Configure the service without storing credentials in the repository:

export PIDDI_LIVE_HANDLE_SERVER_URL=https://handle-test.example/api-root
export PIDDI_LIVE_HANDLE_PREFIX=21.TEST
export PIDDI_LIVE_HANDLE_USERNAME='300:21.TEST/testuser'
export PIDDI_LIVE_HANDLE_PASSWORD='...'
pytest -m live tests/live/test_real_handle_service.py

Set PIDDI_LIVE_HANDLE_VERIFY_HTTPS=false only for a trusted test service with a self-signed certificate. The live suite checks create, overwrite/update, read-back, bounded parallel publication, and same-PID update ordering.

The Docker mock adds a 50 ms delay to each valid PUT, approximating a serial rate of 20 Handle registrations per second. Override it when starting Docker if you need a different latency:

PIDDI_MOCK_HANDLE_PUT_DELAY_SECONDS=0.1 docker compose up -d

The delay occurs outside the mock store lock, so concurrent publisher workers can overlap requests as they would with a real service connection/database pool. GET requests and requests rejected before storage are not delayed.

These files, pid.log, and piddi.db are ignored by Git and preserved by the project's cleanup targets. There is no automatic retention policy. Dump and Handle JSONL files can grow quickly, so monitor disk usage and archive or remove old files according to the site's operational policy.

Safe inspection and retry

See Recovery & retry in the advanced documentation for recovery file locations and the retry workflow.

Retry remaps failed events without contacting the Handle Service. Each command creates and prints a distinct project-scoped retry_handles_<timestamp>.jsonl, keeping late recovery work separate from normal daily mapping output:

piddi retry outputs/cmip6/failures/r0
piddi publish --project cmip6 \
  outputs/cmip6/handles/retry_handles_<timestamp>.jsonl

Use retry --publish only for intentional immediate publication. Recovery records store the canonical project in __infos__.project; retry uses it instead of the current configured project selection. Older records without this metadata still use the configured selection. Retry opens each source once and processes batches of at most 256 records in source order, including when projects are interleaved. Retry counts are incremented once per record. The reader stops at the file size captured when it opens, so records appended during recovery are left for a later run.

Malformed JSON and invalid retry metadata are logged with their source location; valid records continue to process. The summary retains at most 100 input error messages, while all input errors are logged. A missing input is a reported failure rather than an empty successful run.

--delete-after removes an input file only when all records succeed. Malformed JSONL and skipped records are failures for this decision, so the source remains available for inspection. Inputs changed during processing are also retained, so newly appended recovery records are not deleted. Earlier successful output remains written if a later record fails.

Terminal progress

Progress is displayed by default for harvest, map, consume, and publish. Use the global --silent (or -s) or --no-progress option to hide it:

piddi --no-progress map --project cmip6 --date yesterday
piddi -v publish --project cmip6 --date yesterday

The mapping and consumption display reports message and Handle counts/rates, errors, filtered records, warnings, retractions, replicas, skips, patches, time since the last error, and elapsed time. Publication shows completed record counts and the absolute input position; its total is unknown until the stream finishes. Final command summaries are still printed with progress off.

Logging and statistics

The CLI writes WARNING and above to pid.log by default. Use -v for INFO, -vv or its memorable --debug alias for DEBUG, and --log PATH to override the configured file. --silent remains an alias for hiding progress; --progress/--no-progress provides the explicit form. At INFO level Piddiplatsch records the selected plugins, the first occurrence of every filtered project identity, publication outcomes, and periodic aggregate counts. Per-message filter decisions are available at DEBUG level.

The [logging] configuration provides the service defaults:

[logging]
level = "WARNING"
file = "/var/log/piddi/piddi.log"

File logging uses a watched handler: after logrotate renames the active file and creates a replacement, Piddiplatsch switches to the new file on its next log write. Full recovery and skipped details remain available in JSONL independently of the selected logging level. The optional SQLite reporter is controlled by [stats] and is opened by piddi consume and piddi map. harvest and the read-only top command never update the database.

The database contains a versioned current-run status model. A heartbeat is updated independently of Kafka traffic, and processing outcomes are split by canonical project. Inspect it without contacting Kafka or the Handle service:

# live view (Ctrl-C exits)
piddi top

# one project, one terminal snapshot
piddi top --project cmip7 --once

# machine-readable status
piddi top --json

# use a six-hour history window
piddi top --history 360

top marks unfinished runs stale when their heartbeat exceeds stats.stale_after_seconds. An idle topic is healthy while the process keeps heartbeating. The heartbeat interval is configured with stats.heartbeat_interval_seconds. Both default to 5 and 15 seconds, respectively. top is read-only and reports a clear error for a missing or invalid monitoring database.

The database appends cumulative per-project samples every stats.sample_interval_seconds (15 seconds by default), including samples at run startup and shutdown. The history table shows counter changes, average message throughput, and a compact throughput trend for the last stats.history_minutes (60 minutes by default). --history MINUTES overrides that window. Samples older than stats.sample_retention_days (30 days by default) are removed; set it to 0 to retain them indefinitely. Samples contain counters only; raw messages and log lines are never stored.

Shutdown behavior

SIGINT and keyboard interruption close the Kafka consumer, progress display, and statistics reporters. Processing stops with a non-zero status after the configured error limit or a fail-fast transient external failure.

Local cleanup

make clean removes build, bytecode, and test artifacts only. make clean-dist also removes other ignored development artifacts, but explicitly preserves runtime output, logs, databases, local configuration, virtual environments, and editor settings.

Deployment

After Piddi is checked out and installed into its Conda environment manually, the Ansible playbook configures one service per VM. It renders the site configuration from the same custom.toml used by manual runs, adds defaults for omitted production paths, and uses /etc/piddi/piddi.toml, /var/lib/piddi, and /var/log/piddi/piddi.log and configures systemd plus hourly log rotation. For the short production and Vagrant procedures, see deploy/README.md.

Mount an external data disk at /var/lib/piddi; the systemd service waits for a configured mount. If several isolated Piddi installations share one VM, prefer one Podman or Docker container per workflow.