Backup & restore¶
Almost all of an Etminan deployment is re-installable from the package and re-derivable by re-enrolling. Exactly one thing is not: the verifier trust store. This chapter is the disaster-recovery plan for that host.
Back up by blast radius, not by host count¶
| Data | Where | If lost | Priority |
|---|---|---|---|
| Verifier trust store | verifier host — baseline.db (SQLite/WAL, which now also holds the UID→identity registry), its .audit-head anchor, state.json, tls/, /etc/etminan-verifier |
Catastrophic — baselines, audit history, the identity registry, and every pinned host fingerprint are gone; re-enroll every host and rebuild every baseline. | Must back up |
| Daemon signing key | verifier host — /var/lib/etminan-verifier/daemon-signing.key (0600). The one key etminan-verifierd signs operator actions with. |
Past daemon signatures cannot be reproduced. It lives on the verifier by design, so back it up with the trust store, not separately. | Must back up (with the DB) |
###DROP###ty is UID-mapped in baseline.db. |
Optional | ||
| Agent state | each monitored host — tls/, assigned_profile, /etc/etminan-agent |
Re-derivable: re-enroll or rotate-tls from the verifier. |
Optional / nice |
| TPM AK/EK key material | inside each host's TPM | Not backup-able by design. A TPM clear forces re-enrollment of that host regardless. | Cannot back up |
baseline.db is SQLite in WAL mode — never copy the live file
A naive cp of a live baseline.db yields a stale or corrupt restore, and
its external audit-chain anchor
can end up ahead of the copied DB — which the verifier reports as
tampering. Always back up a consistent snapshot whose anchor is derived
from the same generation.
The three responsibilities¶
Whatever backup tool you use — Bacula, restic, Borg, Veeam, ZFS snapshots — a correct Etminan DR backup must do all three of these. Port the intent, not any particular syntax:
- Snapshot the WAL database before reading it. Run a pre-backup hook that
produces a consistent copy of
baseline.dband a matching audit anchor, then back up the snapshot. Exclude the livebaseline.db,-wal,-shm, and live.audit-headfrom the backup set. - Capture the whole verifier trust-store set together. The snapshot + its
.audit-head,state.json,tls/, thedaemon-signing.key, and/etc/etminan-verifier(plaintext integration secrets and the plugin allowlist pins — the conf pins and the pinned binaries must restore together) form one consistent set. The identity registry needs no separate job: it lives insidebaseline.dband is already captured by the DB snapshot. - Include the daemon signing key with the trust store.
/var/lib/etminan-verifier/daemon-signing.keyis the one keyetminan-verifierdsigns operator actions with; losing it means past daemon signatures can no longer be reproduced. It lives on the verifier by design (the daemon, not any operator, holds it), so it belongs in the trust-store backup — there is no off-host operator-key job to keep separate any more, because operators hold no keys at all on this line.
The WAL-safe snapshot pre-job¶
The load-bearing piece is a pre-backup hook that produces a point-in-time copy and derives the anchor from that copy, so the DB and its anchor are the same generation by construction — zero race window:
#!/usr/bin/env bash
set -euo pipefail
DB="${ETMINAN_BASELINE_DB:-/var/lib/etminan-verifier/baseline.db}"
SNAP_DIR="${ETMINAN_BACKUP_DIR:-/var/backups/etminan}"
install -d -m 0700 "$SNAP_DIR"
[ -f "$DB" ] || { echo "no baseline.db at $DB — nothing to snapshot" >&2; exit 0; }
SNAP="$SNAP_DIR/baseline.db.snapshot"
# 1) Consistent point-in-time copy of the whole DB (commits the WAL internally).
sqlite3 "$DB" ".backup '$SNAP'"
# 2) Derive the audit anchor FROM the snapshot, so DB and anchor match exactly.
head_row="$(sqlite3 "$SNAP" \
"SELECT seq || ' ' || entry_hash FROM audit_log ORDER BY seq DESC LIMIT 1;")"
if [ -n "$head_row" ]; then
printf '%s\n' "$head_row" > "$SNAP.audit-head"
else
rm -f "$SNAP.audit-head" # empty audit_log (fresh verifier): no anchor yet
fi
chmod 0600 "$SNAP" "$SNAP.audit-head" 2>/dev/null || true
Why derive the anchor from the snapshot
verify_chain fails — and reports it as tampering — if the DB is behind
its anchor. If you copied the live DB and the live anchor at two different
instants, the anchor could be ahead of the DB snapshot, so the restore would
look tampered. Deriving the anchor from the snapshot itself makes them the
same generation.
state.json, the tls/ directory, daemon-signing.key, and
/etc/etminan-verifier are written atomically (temp + fsync + rename) or are
static key/config material, so a live copy of those is always consistent — no
pre-job needed for them.
What to capture, concretely¶
Back up:
/var/backups/etminan/baseline.db.snapshotand.snapshot.audit-head(from the pre-job — the snapshot also carries the UID→identity registry)/var/lib/etminan-verifier/daemon-signing.key(the daemon's one signing key,0600— custody-critical)/var/lib/etminan-verifier/state.json/var/lib/etminan-verifier/tls//etc/etminan-verifier/(config,verifier.env, plugin allowlist pins and the pinned plugin binaries)- optionally
/var/lib/etminan-verifier/identity.keyif you want to preserve the integrity-summary signing identity
Exclude the live baseline.db, baseline.db-wal, baseline.db-shm, and live
baseline.db.audit-head — the snapshot is the consistent copy.
There is no separate off-host operator-key job: operators hold no key files,
and the identity registry lives inside baseline.db.
The Bacula example set¶
Etminan provides three heavily-commented Bacula example files you copy onto your own backup infrastructure and adapt. They are examples — not installed by the package and not part of the deliverable:
etminan-verifier-snapshot.sh— the WAL-safe snapshot pre-job (shown in full above); wire it into the verifier Job asClient Run Before Job.etminan-filesets.conf— the FileSets:Etminan-Verifier-FileSet(the trust store — the only catastrophic-loss host, and now the home of the daemondaemon-signing.keytoo),Etminan-Agent-FileSet(one per monitored host, mostly re-derivable). The oldEtminan-Operator-Keys-FileSetis retired: operators hold no keys and the identity registry rides along insidebaseline.db.etminan-jobs.conf— Schedules, separate pools per blast radius, JobDefs (the snapshot wired in asClient Run Before Job), Backup/Restore Jobs, and the Console-ACL roles —standard-operator(agent pool only) anddr-operator(verifier + agent — the trust store, which includes the daemon signing key). The separatekey-custodianrole/pool is retired: with the signing key living on the verifier by design, there is no off-host operator-key pool to fence off.
Copy both .conf files onto your Bacula Director (FileSets first), adapt the
<-- site resource names and paths to your environment, and place the snapshot
script on the verifier host.
etminan-filesets.conf — FileSets (what to back up)
# Bacula FileSets for Etminan disaster recovery.
#
# EXAMPLE / REFERENCE ONLY — this is not installed by the etminan package and is
# not part of the deliverable. Copy it onto your Bacula Director and adapt the
# paths and the "<-- site resource" names to your own deployment.
#
# A generic backup config for an Etminan deployment — Etminan has exactly two
# installable parts, so there are two installation FileSets:
# * Etminan-Verifier-FileSet — the verifier install (the trust store; the only
# catastrophic-loss host).
# * Etminan-Agent-FileSet — an agent install, per monitored host; mostly
# re-derivable. Cloned per host in the Jobs file.
#
# Plus one OPTIONAL, off-host FileSet that is not an installation but is a
# critical operator secret:
# * Etminan-Operator-Keys-FileSet — the operator's baseline-approval signing
# key(s), held on an operator workstation.
# Security-sensitive: read the warning above
# that FileSet before enabling it.
#
# NOTE: nothing here is created by any install script — an Etminan deployment is
# installed from the package; the operator hand-creates the real config, TLS
# identity, and (on the verifier) the DB/state. That operator-created state is
# exactly what these FileSets capture, since the package ships only *.example
# templates and empty directories.
#
# IMPORTANT — pair the verifier FileSet with the snapshot pre-job. baseline.db
# is SQLite/WAL; this FileSet backs up the *snapshot* produced by
# etminan-verifier-snapshot.sh, never the live DB. The Jobs, Schedules, Pools,
# and the ClientRunBeforeJob wiring live in the companion etminan-jobs.conf.
# Include both from your Director config, FileSets first:
#
# @/etc/bacula/etminan/etminan-filesets.conf # example path — put them
# @/etc/bacula/etminan/etminan-jobs.conf # wherever your Director reads
FileSet {
Name = "Etminan-Verifier-FileSet"
Include {
Options {
signature = SHA256
compression = GZIP9
# Never back up the live SQLite DB or its sidecars — the snapshot below
# is the consistent copy. Excluding these avoids a corrupt/stale restore.
WildFile = "/var/lib/etminan-verifier/baseline.db"
WildFile = "/var/lib/etminan-verifier/baseline.db-wal"
WildFile = "/var/lib/etminan-verifier/baseline.db-shm"
WildFile = "/var/lib/etminan-verifier/baseline.db.audit-head"
Exclude = yes
}
Options {
signature = SHA256
compression = GZIP9
}
# --- The trust store (CATASTROPHIC if lost) ---------------------------
# WAL-safe SQLite snapshot + its consistent audit-chain anchor, produced by
# ClientRunBeforeJob = etminan-verifier-snapshot.sh. Restore: copy
# baseline.db.snapshot -> baseline.db and baseline.db.snapshot.audit-head ->
# baseline.db.audit-head (see the exclude block above — never the live DB).
File = /var/backups/etminan/baseline.db.snapshot
File = /var/backups/etminan/baseline.db.snapshot.audit-head
# Per-host enrollment + replay state: pinned AK/EK/TLS fingerprints, PCR10
# cursor, signed enrollment/rotation accountability. Written atomically
# (temp+fsync+rename) so a live copy is always consistent — safe as-is.
# Losing it forces re-enrollment of every host from scratch.
File = /var/lib/etminan-verifier/state.json
# This verifier's mTLS identity. Its fingerprint is pinned into every agent
# (ETMINAN_VERIFIER_CERT_FINGERPRINT). Losing it forces keygen-tls + re-pin
# on every agent by hand.
File = /var/lib/etminan-verifier/tls
# --- Config + plugin allowlists (recoverable, but back up) ------------
# verifier.env holds plaintext integration secrets (RT/FreeITSM/PagerDuty/
# Slack/webhook). The two *-plugins.conf files carry the SHA256 pins that
# authorize plugin execution; the *-plugins/ dirs hold the pinned binaries.
# The conf pins and the binaries must restore together.
File = /etc/etminan-verifier
# --- Convenience: fast-restore binary (re-installable from package) ---
WildFile = "/usr/bin/etminan-verifier"
File = /usr/bin
# --- NOT captured here (off-host / out of band): --------------------------
# * Operator baseline-approval signing keys — no default path; created by
# `keygen --out <path>`, held per-operator OFF the verifier host. Their
# private halves sign approve/reject/enroll; a lost key cannot re-sign
# existing baseline_batches. Back up per operator (their responsibility).
# * Plugin-catalog signing key (~/.etminan-catalog-signing.key) — offline,
# never on a verifier host (HANDBOOK §5).
# * The host TPM is the irreplaceable factor for AGENTS: a TPM clear/reset
# invalidates the AK/EK pinned in state.json and forces re-enroll. TPM
# key material cannot be backed up.
}
}
FileSet {
Name = "Etminan-Agent-FileSet"
Include {
Options {
signature = SHA256
compression = GZIP9
}
# Agent mTLS identity — fingerprint is pinned in the verifier's state.json.
# Loss requires re-enroll or rotate-tls (verifier-side action), not fatal.
File = /var/lib/etminan-agent/tls
# Last profile pushed by assign-profile; re-derivable (authoritative record
# is in the verifier's state.json + audit_log), included for a clean restore.
File = /var/lib/etminan-agent/assigned_profile
# agent.env (holds ETMINAN_VERIFIER_CERT_FINGERPRINT, the pin) + profiles.conf
File = /etc/etminan-agent
WildFile = "/usr/bin/etminan-agent"
File = /usr/bin
# NOTE: TPM AK/EK are in-TPM only — nothing on disk to back up. A TPM clear
# forces re-enrollment regardless of this backup.
}
}
# ── Etminan-Operator-Keys-FileSet — OFF-HOST, read this before enabling ───────
# The crown-jewel private keys that are deliberately NOT stored on any verifier
# or agent (that separation is the whole point — a compromised verifier host must
# not also yield the keys that authorize its baselines):
#
# * Operator baseline-approval signing key(s) — sign enroll/rotate/approve/
# reject/assign-profile. No default path; created per-operator with
# `etminan-verifier keygen --out <path>` (writes <path> + <path>.pub). A
# lost private half cannot re-sign existing baseline_batches. This is the
# key a customer deployment actually has.
# * (Only if you run your OWN private plugin catalog) a catalog signing key
# from `etminan-verifier plugins keygen-catalog`. Most deployments don't —
# they consume the vendor catalog, whose public half is compiled in and
# whose private half never exists on a customer host at all. Skip unless
# you self-sign a catalog.
#
# SECURITY — do NOT undo the separation this design relies on:
# * Back these up ONLY from the operator workstation that legitimately holds
# them, into their OWN pool (EtminanOperatorKeysPool in etminan-jobs.conf) —
# NEVER the verifier pool. A single pool restore must never hand someone both
# the baselines AND the keys that authorize them.
# * Prefer an encrypted, offline medium. If you do use Bacula, enable client-
# side Data Encryption (PKI Encryption/PKI Keypair on the operator
# workstation's FileDaemon) so the keys are encrypted at rest in the volume.
# * This FileSet is the Bacula option for shops that already back up operator
# workstations — not a recommendation to put signing keys on the same tape
# as everything else.
FileSet {
Name = "Etminan-Operator-Keys-FileSet"
Include {
Options {
signature = SHA256
compression = GZIP9
# See the SECURITY note above — configure PKI Encryption on this client's
# FileDaemon so these private keys are encrypted at rest in the volume.
}
# Fill in the ACTUAL paths on this operator's workstation. Operator signing
# keys have no default location (keygen --out <path>) — these are examples:
File = /home/<operator>/.etminan/operator-signing.key
File = /home/<operator>/.etminan/operator-signing.key.pub
# Only if you self-sign your own plugin catalog (uncommon — see above):
# File = /home/<operator>/.etminan-catalog-signing.key
}
}
etminan-jobs.conf — Schedules, Pools, Jobs, Console ACLs (when/how)
# Bacula Jobs for Etminan disaster recovery.
#
# EXAMPLE / REFERENCE ONLY — not installed by the etminan package, not part of
# the deliverable. Copy onto your Bacula Director and adapt to your deployment.
#
# This is the companion to etminan-filesets.conf — that file defines *what* to
# back up (Etminan-Verifier-FileSet, Etminan-Agent-FileSet, and the optional
# Etminan-Operator-Keys-FileSet); this file defines *when* and *how* (Schedules,
# Pools, JobDefs, and the Backup/Restore Jobs that reference those FileSets).
# Without this file the FileSets are never exercised.
#
# Include both from your Director config, FileSets first:
#
# @/etc/bacula/etminan/etminan-filesets.conf # example path — put them
# @/etc/bacula/etminan/etminan-jobs.conf # wherever your Director reads
#
# The Director, Catalog, Messages, Client (FileDaemon), and Storage resources
# are site-specific and NOT defined here — supply them from your environment.
# The names referenced below and marked "<-- site resource" must match yours:
#
# Client = etminan-verifier-fd , etminan-agent-fd , etminan-operator-fd
# Storage = File
# Messages = Standard
# Catalog = MyCatalog
#
# etminan-operator-fd is the OPERATOR WORKSTATION that holds the signing keys —
# deliberately a different client from any verifier (see the security note on
# Etminan-Operator-Keys-FileSet in etminan-filesets.conf).
#
# See Bacula-Forgejo/docs/bacula-director.conf.example for a complete Director/
# Catalog/Storage/Messages scaffold to adapt.
# ── Schedules ─────────────────────────────────────────────────────────────────
# Weekly Full + daily Incremental. The verifier snapshot pre-job regenerates
# baseline.db.snapshot on every run, so Incrementals reliably capture DB changes.
Schedule {
Name = "EtminanVerifierSchedule"
Run = Full 1st sun at 01:00
Run = Incremental mon-sat at 01:00
}
Schedule {
Name = "EtminanAgentSchedule"
# Agent state is small and mostly re-derivable — a nightly Full is cheap and
# keeps restore dead-simple (no chain to replay).
Run = Full daily at 02:00
}
Schedule {
Name = "EtminanOperatorKeysSchedule"
# Signing keys change rarely (only on rotation). A nightly Full is trivially
# cheap and means a fresh key is captured within a day of being generated.
Run = Full daily at 04:00
}
# ── Pools ─────────────────────────────────────────────────────────────────────
# The verifier trust store is the only catastrophic-loss data AND contains
# plaintext secrets (verifier.env) plus mTLS/signing material. It gets its own
# restricted pool so restore access can be locked down via Console ACLs below.
Pool {
Name = EtminanVerifierPool
Pool Type = Backup
Recycle = yes
AutoPrune = yes
Volume Retention = 365 days
Maximum Volume Bytes = 20G
Label Format = "EtminanVerifierVol-"
}
Pool {
Name = EtminanAgentPool
Pool Type = Backup
Recycle = yes
AutoPrune = yes
Volume Retention = 180 days
Maximum Volume Bytes = 20G
Label Format = "EtminanAgentVol-"
}
# The operator-keys pool holds the private signing keys and is kept in its OWN
# pool so restore can be locked to a key custodian (see Console ACLs below) and
# NEVER mixed with the verifier trust store. Long retention: these keys are
# long-lived, so keep history for the working life of the key, not just weeks.
Pool {
Name = EtminanOperatorKeysPool
Pool Type = Backup
Recycle = yes
AutoPrune = yes
Volume Retention = 3650 days
Maximum Volume Bytes = 2G
Label Format = "EtminanOperatorKeysVol-"
}
# ── JobDefs ───────────────────────────────────────────────────────────────────
JobDefs {
Name = "EtminanVerifierDefs"
Type = Backup
Level = Incremental
Client = etminan-verifier-fd # <-- site resource
FileSet = "Etminan-Verifier-FileSet" # from etminan-filesets.conf
Storage = File # <-- site resource
Messages = Standard # <-- site resource
Pool = EtminanVerifierPool
Schedule = "EtminanVerifierSchedule"
Priority = 10
# CRITICAL: take the WAL-safe snapshot BEFORE the FileSet is read, so the
# backup captures baseline.db.snapshot (+ .audit-head) instead of the live
# SQLite DB. The FileSet excludes the live DB; this produces the copy it wants.
# Place etminan-verifier-snapshot.sh anywhere on the verifier host and point
# this at it (example path shown — it runs on the client/FileDaemon).
Client Run Before Job = "/usr/local/lib/etminan/etminan-verifier-snapshot.sh"
}
JobDefs {
Name = "EtminanAgentDefs"
Type = Backup
Level = Full
Client = etminan-agent-fd # <-- site resource
FileSet = "Etminan-Agent-FileSet" # from etminan-filesets.conf
Storage = File # <-- site resource
Messages = Standard # <-- site resource
Pool = EtminanAgentPool
Schedule = "EtminanAgentSchedule"
Priority = 15
# No pre-job: agent state (tls/, assigned_profile, /etc/etminan-agent) is
# written atomically (temp+fsync+rename), so a live copy is always consistent.
}
JobDefs {
Name = "EtminanOperatorKeysDefs"
Type = Backup
Level = Full
Client = etminan-operator-fd # <-- OPERATOR WORKSTATION, never a verifier
FileSet = "Etminan-Operator-Keys-FileSet" # from etminan-filesets.conf
Storage = File # <-- site resource
Messages = Standard # <-- site resource
Pool = EtminanOperatorKeysPool
Schedule = "EtminanOperatorKeysSchedule"
Priority = 25
# No pre-job: signing keys are static files written 0600 O_EXCL at keygen.
}
# ── Backup Jobs ───────────────────────────────────────────────────────────────
Job {
Use = "EtminanVerifierDefs"
Name = "Etminan-Verifier"
}
# Define one agent Job per monitored host (each with its own Client). This is
# the template; clone it per host, overriding Name + Client.
Job {
Use = "EtminanAgentDefs"
Name = "Etminan-Agent"
Client = etminan-agent-fd # <-- site resource (one per host)
}
# Optional — enable only if you back up the operator workstation that holds the
# signing keys. Read the security note on Etminan-Operator-Keys-FileSet first.
Job {
Use = "EtminanOperatorKeysDefs"
Name = "Etminan-Operator-Keys"
}
# ── Restore Jobs ──────────────────────────────────────────────────────────────
# NOTE on verifier restore: Bacula writes baseline.db.snapshot and
# baseline.db.snapshot.audit-head back to /var/backups/etminan. To bring the
# verifier live, copy them into place (see the Verifier FileSet's restore note):
# cp /var/backups/etminan/baseline.db.snapshot \
# /var/lib/etminan-verifier/baseline.db
# cp /var/backups/etminan/baseline.db.snapshot.audit-head \
# /var/lib/etminan-verifier/baseline.db.audit-head
# Never restore the live DB directly — the FileSet never captured it.
Job {
Name = "Etminan-Verifier-Restore"
Type = Restore
Client = etminan-verifier-fd # <-- site resource
FileSet = "Etminan-Verifier-FileSet"
Storage = File # <-- site resource
Messages = Standard # <-- site resource
Pool = EtminanVerifierPool
Where = / # files carry absolute paths; restore in place
}
Job {
Name = "Etminan-Agent-Restore"
Type = Restore
Client = etminan-agent-fd # <-- site resource
FileSet = "Etminan-Agent-FileSet"
Storage = File # <-- site resource
Messages = Standard # <-- site resource
Pool = EtminanAgentPool
Where = /
}
# Restores the signing keys to a STAGING dir, not in place — so a restore can't
# silently clobber a live key. The custodian inspects them, then moves each into
# its real location by hand (keys are 0600; preserve that on the move).
Job {
Name = "Etminan-Operator-Keys-Restore"
Type = Restore
Client = etminan-operator-fd # <-- OPERATOR WORKSTATION
FileSet = "Etminan-Operator-Keys-FileSet"
Storage = File # <-- site resource
Messages = Standard # <-- site resource
Pool = EtminanOperatorKeysPool
Where = /var/tmp/etminan-key-restore # staging; move into place by hand
}
# ── Console ACLs (restrict restores by blast radius) ──────────────────────────
# The verifier pool holds secrets + trust-store material; the operator-keys pool
# holds the signing keys. These are split across three roles so no single console
# can restore both the baselines and the keys that authorize them:
# * standard-operator — agent pool only.
# * dr-operator — verifier + agent pools (the trust store), but NOT the
# operator-keys pool.
# * key-custodian — operator-keys pool ONLY. Deliberately cannot touch the
# verifier pool, preserving the separation of duties.
Console {
Name = etminan-standard-operator
Password = "standard-console-password" # <-- CHANGE
CommandACL = status, messages, show, run
JobACL = Etminan-Agent, Etminan-Agent-Restore
PoolACL = EtminanAgentPool
}
Console {
Name = etminan-dr-operator
Password = "dr-console-password" # <-- CHANGE
CommandACL = restore, run, status, messages, show, cancel
JobACL = Etminan-Verifier, Etminan-Verifier-Restore, Etminan-Agent, Etminan-Agent-Restore
PoolACL = EtminanVerifierPool, EtminanAgentPool
}
Console {
Name = etminan-key-custodian
Password = "key-custodian-console-password" # <-- CHANGE
CommandACL = restore, run, status, messages, show, cancel
JobACL = Etminan-Operator-Keys, Etminan-Operator-Keys-Restore
PoolACL = EtminanOperatorKeysPool
}
Restore¶
The whole restore is a copy of the trust store into place, then a health check — there is no re-import step.
- Reinstall the package on the replacement verifier host (or reuse the existing one).
- Restore the trust-store set. Your backup tool writes the snapshot files
back (e.g. under
/var/backups/etminan) alongsidestate.json,tls/,daemon-signing.key, and/etc/etminan-verifier. The identity registry comes back withbaseline.db— it was inside the snapshot. -
Bring the database live from the snapshot — never the live DB, which was never captured:
-
Fix ownership/permissions so the verifier's unprivileged service user owns
/var/lib/etminan-verifier, and keepbaseline.dbanddaemon-signing.keynon-world-readable (daemon-signing.keystays0600). -
Verify before trusting the box:
op pingconfirms the daemon comes back up against the restored state.baseline verify-signaturesre-checks every signature and the full audit-log chain, including its anchor — if the restore lost or reordered rows, this is what says so. -
Restart the timer/service and confirm the next
runcycle succeeds and reaches every enrolled host.
Because state.json restored the pinned AK/EK/TLS fingerprints and PCR-10
cursors, the verifier resumes attesting the existing fleet with no re-enrollment.
Rehearsing the restore¶
A backup nobody has restored is a hypothesis. Rehearse it on a schedule, and write down the date and the result — a customer's security questionnaire asks how often, not whether, and the honest answer comes from a record rather than from a recollection.
Quarterly is the interval this handbook recommends, and after any change to the backup job, the storage target or the verifier version. That is a recommendation, not a setting: nothing in the product enforces it, and nothing in the product knows whether you did it.
The rehearsal is the restore above, performed onto a spare host or a throwaway VM — never onto the live verifier — and it is finished when these two answer:
That command is the pass/fail. It re-checks every signature, the full audit-log
hash chain and its anchor — which is what proves the restored database and
its .audit-head are the same generation, the single most likely thing to be
wrong about a restore and the one a "the files copied fine" check does not
catch.
What to record, per rehearsal: the date, the backup generation used, which checks passed, and how long the whole thing took. The last one is the number you will be asked for when somebody wants a recovery-time estimate.
A rehearsal on a spare host does not disturb the fleet: the restored verifier is not the one the agents are pinned to, so it will report enrolled hosts it cannot reach. That is expected — the chain check is the pass/fail, not reachability from a lab machine.
What a restore cannot recover¶
- The daemon signing key, if you did not capture it with the trust store — a
lost
daemon-signing.keymeans past daemon signatures cannot be reproduced (existing audited actions remain valid and verifiable; the daemon simply signs new actions with a fresh key once re-seeded). This is why it belongs in the trust-store backup, not a separate off-host job. - A host whose TPM was cleared. The AK/EK pinned in
state.jsonno longer reside in that TPM, so it must be re-enrolled regardless of any backup — TPM key material cannot be backed up.
See also¶
- The audit log — why the anchor must match the DB generation.
- Hardening the deployment — disk encryption and access control for the backed-up secrets.
- Verifier configuration — the daemon signing key and the UID→identity registry this backup must preserve.