Operations
Runtime output
Piddiplatsch writes daily JSONL files beneath consumer.output_dir:
| Directory | Contents |
|---|---|
dump/ |
Original Kafka messages written by every harvest or consume run. |
<plugin>/handles/ |
Handle records produced by that plugin. Direct REST/pyhandle publication writes the JSONL record before contacting the server; each line also includes project. |
published/ |
One run-scoped JSONL receipt per publish run, containing every successful or failed Handle outcome. |
<plugin>/skipped/ |
Records deferred after transient external failures. |
<plugin>/failures/r<N>/ |
Records that failed processing, grouped by retry count. |
skipped/, failures/r<N>/ |
Legacy or unresolved-project recovery records. |
The raw dump is intentionally global and always written before routing. It preserves the consumed Kafka order and remains suitable for replay or investigation; creating filtered per-plugin dumps would lose that simple ordering guarantee.
JSONL Handle output is always enabled, including direct publication. If the audit record cannot be appended, that Handle is not sent to the server. A server-side failure leaves the JSONL record available for inspection or later publication.
Deferred publication writes each outcome immediately to a readable, unique
published/published_<project>_handles_YYYY-MM-DD_HH-MM-SS.jsonl file using UTC
and prints its path in the final CLI summary. If no project is known when the
receipt opens, the filename is published_handles_...jsonl. If another run starts during the same second,
_2, _3, and so on are appended. Parallel completion order may differ from
input order; position, batch_index, source_file, and source_line provide
stable ordering and provenance.
Staged commands
# Kafka -> raw JSONL only
piddi harvest --limit 100
# raw JSONL -> project-scoped Handle JSONL only
piddi map --project cmip6 --date 2026-08-27
# Handle JSONL -> REST Handle Service
piddi publish --project cmip6 --date 2026-08-27
# Kafka -> raw JSONL -> Handle JSONL (default production ingestion)
piddi consume
# All three stages in one process
piddi consume --publish
harvest --limit N stops normally after dumping N messages, which is useful
for bounded tests against a live topic. map accepts files or directories plus
--project, --all-projects,
--limit, --offset, and --force. It never contacts Kafka or a Handle
Service and does not modify its input dumps. It opens each source once and maps
one record at a time in source order, without loading the selected files into
memory. Offsets and limits apply across files and count nonblank input lines;
source keys use physical line numbers. A malformed selected JSON record stops
the run with its source location; output from earlier records remains written.
The existing processing-error and transient-failure stop policies still apply.
The --date convenience accepts YYYY-MM-DD, today, yesterday, today-N,
or last. map uses dump/dump_messages_<date>.jsonl; publish requires a
project and uses <project>/handles/handles_<date>.jsonl. When neither an
explicit path nor --date is supplied, both commands default to last.
This selects the greatest valid date found in the relevant filenames,
regardless of file modification time. Explicit paths remain available for both
commands, and a path and --date are mutually exclusive.
Deferred publication
Publish completed files after mapping, using the project's configured Handle profile. Profiles contain the REST server, prefix, and credentials; see Handle configuration. Keep credentials in the ignored site configuration file.
# Publish yesterday's prepared CMIP6 Handles
piddi publish --project cmip6 --date yesterday
# Try the first 1,000 records of a specific completed file
piddi publish --project cmip6 --date 2026-08-25 --limit 1000
# Continue with the next 1,000 records, allowing transient retries
piddi publish --project cmip6 --date 2026-08-25 \
--offset 1000 --limit 1000 --retries 3
Explicit inputs can be one file, several files, or a directory. Offsets count nonblank input lines across the selected files; limits cap attempted records, including failures. Publication never changes or deletes its inputs. Handle writes use overwrite semantics, so a completed, immutable input can be replayed after interruption. The command exits non-zero if any record fails.
--retries covers transient connection errors, timeouts, rate limiting, and
server errors. The delay starts at one second and doubles up to 60 seconds;
change its initial value with --retry-delay. Permanent client errors such as
invalid credentials are not retried. --workers N enables concurrent requests
for different Handles, while updates to the same Handle retain source order
within the run.
publish --project NAME validates each Handle immediately before sending it.
Missing or different projects are recorded as failures, while valid records
continue. Without --project, the first parsed record selects the service
configuration; subsequent records must match its project. Each JSONL file is
opened once and processed in batches of at most 256 records. Writes for the same
Handle stay in input order, including across batches. Malformed lines receive
failure receipts; earlier writes are not rolled back. No staging database or
preliminary scan is used, so progress has no known total and receipt
batch_total is null. Parent references are sent as provided; receipt metadata
comes only from the current record. The final summary provides per-project
counts and up to 100 error messages; receipts retain all failures.
Each receipt line records the outcome, action, PID, URL, project, dataset, asset,
source location, position, retry count, and error. Missing dataset or asset
metadata remains null. batch_index and position refer to the entire
selected input, independently of internal batch boundaries. Receipts are
written regardless of log verbosity; INFO logging additionally records each
publication outcome in the configured log file.
Real Handle service contract test
The opt-in live tests create uniquely named Handles and do not delete them, so use a disposable test prefix. Configure the service without storing credentials in the repository:
export PIDDI_LIVE_HANDLE_SERVER_URL=https://handle-test.example/api-root
export PIDDI_LIVE_HANDLE_PREFIX=21.TEST
export PIDDI_LIVE_HANDLE_USERNAME='300:21.TEST/testuser'
export PIDDI_LIVE_HANDLE_PASSWORD='...'
pytest -m live tests/live/test_real_handle_service.py
Set PIDDI_LIVE_HANDLE_VERIFY_HTTPS=false only for a trusted test service with
a self-signed certificate. The live suite checks create, overwrite/update,
read-back, bounded parallel publication, and same-PID update ordering.
The Docker mock adds a 50 ms delay to each valid PUT, approximating a serial rate of 20 Handle registrations per second. Override it when starting Docker if you need a different latency:
PIDDI_MOCK_HANDLE_PUT_DELAY_SECONDS=0.1 docker compose up -d
The delay occurs outside the mock store lock, so concurrent publisher workers can overlap requests as they would with a real service connection/database pool. GET requests and requests rejected before storage are not delayed.
These files, pid.log, and piddi.db are ignored by Git and preserved by the
project's cleanup targets. There is no automatic retention policy. Dump and
Handle JSONL files can grow quickly, so monitor disk usage and archive or remove
old files according to the site's operational policy.
Safe inspection and retry
See Recovery & retry in the advanced documentation for recovery file locations and the retry workflow.
Retry remaps failed events without contacting the Handle Service. Each command
creates and prints a distinct project-scoped
retry_handles_<timestamp>.jsonl, keeping late recovery work separate from
normal daily mapping output:
piddi retry outputs/cmip6/failures/r0
piddi publish --project cmip6 \
outputs/cmip6/handles/retry_handles_<timestamp>.jsonl
Use retry --publish only for intentional immediate publication.
Recovery records store the canonical project in __infos__.project; retry
uses it instead of the current configured project selection. Older records
without this metadata still use the configured selection. Retry opens each
source once and processes batches of at most 256 records in source order,
including when projects are interleaved. Retry counts are incremented once per
record. The reader stops at the file size captured when it opens, so records
appended during recovery are left for a later run.
Malformed JSON and invalid retry metadata are logged with their source location; valid records continue to process. The summary retains at most 100 input error messages, while all input errors are logged. A missing input is a reported failure rather than an empty successful run.
--delete-after removes an input file only when all records succeed. Malformed
JSONL and skipped records are failures for this decision, so the source remains
available for inspection. Inputs changed during processing are also retained,
so newly appended recovery records are not deleted. Earlier successful output
remains written if a later record fails.
Terminal progress
Progress is displayed by default for harvest, map, consume, and publish.
Use the global --silent (or -s) or --no-progress option to hide it:
piddi --no-progress map --project cmip6 --date yesterday
piddi -v publish --project cmip6 --date yesterday
The mapping and consumption display reports message and Handle counts/rates, errors, filtered records, warnings, retractions, replicas, skips, patches, time since the last error, and elapsed time. Publication shows completed record counts and the absolute input position; its total is unknown until the stream finishes. Final command summaries are still printed with progress off.
Logging and statistics
The CLI writes WARNING and above to pid.log by default. Use -v for INFO,
-vv or its memorable --debug alias for DEBUG, and --log PATH to override
the configured file. --silent remains an alias for hiding progress;
--progress/--no-progress provides the explicit form. At INFO level Piddiplatsch
records the selected plugins, the first occurrence of every filtered project
identity, publication outcomes, and periodic aggregate counts. Per-message
filter decisions are available at DEBUG level.
The [logging] configuration provides the service defaults:
[logging]
level = "WARNING"
file = "/var/log/piddi/piddi.log"
File logging uses a watched handler: after logrotate renames the active file and
creates a replacement, Piddiplatsch switches to the new file on its next log
write. Full recovery and skipped details remain available in JSONL independently
of the selected logging level.
The optional SQLite reporter is controlled by [stats] and is opened by
piddi consume and piddi map. harvest and the read-only top command never
update the database.
The database contains a versioned current-run status model. A heartbeat is updated independently of Kafka traffic, and processing outcomes are split by canonical project. Inspect it without contacting Kafka or the Handle service:
# live view (Ctrl-C exits)
piddi top
# one project, one terminal snapshot
piddi top --project cmip7 --once
# machine-readable status
piddi top --json
# use a six-hour history window
piddi top --history 360
top marks unfinished runs stale when their heartbeat exceeds
stats.stale_after_seconds. An idle topic is healthy while the process keeps
heartbeating. The heartbeat interval is configured with
stats.heartbeat_interval_seconds. Both default to 5 and 15 seconds,
respectively. top is read-only and reports a clear error for a missing or
invalid monitoring database.
The database appends cumulative per-project samples every
stats.sample_interval_seconds (15 seconds by default), including samples at
run startup and shutdown. The history table shows counter changes, average
message throughput, and a compact throughput trend for the last
stats.history_minutes (60 minutes by default). --history MINUTES overrides
that window. Samples older than stats.sample_retention_days (30 days by
default) are removed; set it to 0 to retain them indefinitely. Samples contain
counters only; raw messages and log lines are never stored.
Shutdown behavior
SIGINT and keyboard interruption close the Kafka consumer, progress display, and statistics reporters. Processing stops with a non-zero status after the configured error limit or a fail-fast transient external failure.
Local cleanup
make clean removes build, bytecode, and test artifacts only. make clean-dist
also removes other ignored development artifacts, but explicitly preserves
runtime output, logs, databases, local configuration, virtual environments, and
editor settings.
Deployment
After Piddi is checked out and installed into its Conda environment manually,
the Ansible playbook configures one service per VM. It renders the site
configuration from the same custom.toml used by manual runs, adds defaults
for omitted production paths, and uses
/etc/piddi/piddi.toml, /var/lib/piddi, and /var/log/piddi/piddi.log and
configures systemd plus hourly log rotation. For the short production and
Vagrant procedures, see deploy/README.md.
Mount an external data disk at /var/lib/piddi; the systemd service waits for a
configured mount. If several isolated Piddi installations share one VM, prefer
one Podman or Docker container per workflow.