Skip to content

Troubleshooting

Common failure modes and how to diagnose them.

Start here

Both binaries fail loudly rather than degrading quietly, so the first move is always to read what they actually said.

# On a monitored host — the agent refuses to start on a bad config and says why:
systemctl status etminan-agent
journalctl -u etminan-agent -n 50

# On the verifier — is the daemon up, and who does it think you are?
etminan-verifier op ping
etminan-verifier op whoami

# Attest one host now, printing the full verdict:
etminan-verifier check web-01 --addr 10.0.0.5:7620

# Re-check every signature and the whole audit-log hash chain:
etminan-verifier baseline verify-signatures

op ping failing means the daemon is not running or its socket is not reachable; op whoami failing means it is running but your UID holds no registered identity (RBAC).

Enrollment: AK fingerprint mismatch

Symptom: enroll refuses, or a later check reports the host's AK fingerprint no longer matches the enrollment record.

  • The AK template must never change. The AK is created via TPM2_CreatePrimary from a fixed template; a different template derives a different key and breaks the fingerprint match. This is by design — it is what lets enrollment pin a stable identity with no persistent key storage.
  • A genuinely changed AK (TPM owner-clear or firmware change re-derived it) is not an error to force past — use rotate-ak, which re-binds the AK within the same TPM using the EK as the continuity anchor. A different EK means a different TPM and is refused as a re-enroll, not a rotation.
  • Re-enrolling an already-enrolled host is refused unless you pass --force, so a fat-fingered re-run cannot silently re-pin a possibly-compromised host via trust-on-first-use. Only --force after re-confirming the new AK fingerprint out of band.

PcrDigestMismatch on check / run

Symptom: check or run reports a finding: the replayed PCR 10 does not equal the value the TPM signed.

This is the verifier's core check — replay the IMA log, recompute PCR 10, and confirm it equals the signed quote. A mismatch is a real finding by default, but a few benign or environmental causes are worth ruling out:

  • The host rebooted. A reboot clears PCR 10 and starts a fresh IMA log, so the cumulative value the verifier carries belongs to a log that no longer exists and no replay can reach the signed digest again. Since 0.11 the verifier establishes this instead of guessing: on a mismatch it asks the host once more for its log from the first entry, and if replaying all of it reproduces the signed digest, the host rebooted. The cycle then carries on by itself and the reboot is reported as a host-rebooted warning, not an alarm. You do not have to do anything, and there is nothing to re-run.

  • Log growth between quote and read. The IMA log keeps growing from background system activity in the small window between the quote being taken and the log being read. Etminan handles this with prefix-matching — a valid quote's digest must match some prefix of what was received; anything after belongs to the next cycle. If you see transient mismatches that clear on the next cycle, this is the mechanism working.

  • Wrong IMA policy / PCR bank. The SHA-256 PCR bank must be read from ascii_runtime_measurements_sha256, not the generic ascii_runtime_measurements (which uses legacy SHA1-sized template hashes). A persistent mismatch across cycles on one host points at a genuine drift or a tampered log — treat it as a finding and review it.

Since 0.11 the verifier tells the agent where the delta must start, so it also knows when the answer did not cover what it asked for. A cursor problem is reported as log-cursor-gap, and a reboot clears itself and is reported as host-rebooted. So a pcr-mismatch today is what is left after the benign explanations above have been ruled out — and the finding text names each one it ruled out. Treat it as a finding: keep the host's IMA log and the quote, and do not re-enroll, which resets the AK binding and destroys the evidence.

log-cursor-gap on check / run

Symptom: a finding says the agent's log cursor is some number of bytes ahead of the verifier's.

This is not evidence of tampering, and the wording says so. It means measurements were consumed that this verifier never received, so no replay can reach the signed digest again. It comes from a host running an older agent, or one whose log was truncated or rotated underneath it.

Recovery: none needed on a 0.11 verifier — it names the starting offset in every request, so the two sides cannot drift apart, and a host whose log was truncated or rotated underneath it is re-read from the first entry on the next cycle. Re-enrolment is not needed and would reset the AK binding for a bookkeeping problem.

Then look at what failed before it — the gap is a symptom, and the cycle that caused it is the thing worth reading.

Agent and verifier upgrade together

The wire protocol carries a version and both sides must match; a mismatch is reported as a clear error rather than being papered over. 0.11 speaks v3. Upgrading only one side gives you a version error on every check — which is the intended outcome, because the alternative is an agent that quietly ignores the offset the verifier asked for.

Host stuck "catching up"

Symptom: a host reports CatchingUp cycle after cycle and never returns Genuine.

A host with a genuine backlog (a large log delta) clears in a finite number of cycles. state.json tracks consecutive_catching_up and resets it to 0 on every Genuine cycle. A host still catching up past ETMINAN_MAX_CATCHING_UP_CYCLES is a wedged agent — or a compromised host feeding a fabricated log_truncated delta to suppress the mismatch alarm forever — and Etminan alarms rather than staying silent. Investigate the agent process and the IMA log on that host.

TPM device: /dev/tpmrm0

Symptom: the agent logs that it cannot open the TPM.

  • The agent runs as an unprivileged tss-group user and must use /dev/tpmrm0 (the resource-managed device), never /dev/tpm0 (the raw device, group-owned by root). Never rely on the library's DeviceConfig::default(), which points at /dev/tpm0. Running the packaged agent as its real systemd user surfaces this; running manual tests as root masks it.
  • /dev/tpmrm0 does not exist → the TPM is absent or the kernel resource manager is not enabled.
  • exists but cannot be opened → check the device is group tss, mode 0660, and the agent user is in the tss group.
  • The device path is not configurable on this line — the agent always opens /dev/tpmrm0. A host with no TPM cannot attest.

IMA log permission denied

Symptom: the agent logs a permission error reading the IMA measurement log.

The IMA measurement log is root:root 0440 by kernel design — no group membership grants an unprivileged user read access. The fix is the systemd unit granting AmbientCapabilities=CAP_DAC_READ_SEARCH (the narrowest capability that bypasses the read-permission check), not running the agent as root. Check the unit actually carries that grant. Run by hand as root the log reads fine because root holds every capability — which is exactly why the packaged, unprivileged run is the honest test.

IMA policy did not load / measurement count is 0

Symptom: /sys/kernel/security/ima/runtime_measurements_count reads 0, and quotes carry an empty log.

  • IMA accepts exactly one successful policy write per boot (failed writes do not consume the slot). Write the complete policy in one shot, early at boot, via etminan-agent-ima-policy.service — not incrementally.
  • Once loaded, the policy write node can disappear from securityfs entirely (ENOENT), not just start rejecting writes with EACCES. write-ima-policy treats either failure mode as "already loaded" when runtime_measurements_count is already nonzero, so a second run is a no-op, not an error.

Agent won't start: verifier pinning

Symptom: the agent refuses to start, logging a verifier-pinning / TLS-mode error.

  • Set ETMINAN_VERIFIER_CERT_FINGERPRINT to the verifier's fingerprint (exactly 32 bytes / 64 hex chars). A wrong length or non-hex value is a FAIL.
  • The only alternative is ETMINAN_ALLOW_PERMISSIVE_TLS=1, which accepts a client certificate from any caller and is intended only for the one-time enrollment bootstrap. Pin the fingerprint and drop the flag once enrolled.

Certificate expiry

Symptom: connections fail and the log names an expired certificate.

Rotate: on the agent run etminan-agent keygen-tls, then re-pin on the verifier with op rotate-tls. For the verifier's own certificate, run etminan-verifier op keygen-tls, update ETMINAN_VERIFIER_CERT_FINGERPRINT in every agent's agent.env, then restart each agent — in that order (restarting an agent before its env is updated locks it out until fixed).

Audit-log chain broken

Symptom: run or baseline verify-signatures reports the audit-log chain is broken.

The error text names the exact failure — a seq gap (a row deleted or reordered), a prev_hash mismatch (chain broken), a stored-hash mismatch (a row edited after being written), a high-water above the surviving head (tail truncation), or a missing/altered externally-anchored head row (rollback / DB restore). See The audit log. On a verifier you administer, the most common innocent cause is an incorrect backup/restore — see Backup & restore for restoring a snapshot whose anchor matches the DB generation. A break you cannot explain from your own operations is a finding: treat the host as potentially compromised.

A finding won't leave the box

Symptom: you expect alerts but none arrive.

The alarm path is the only place a finding leaves the verifier, so test it directly rather than waiting for a real finding:

etminan-verifier op notify-preview            # render templates against samples (sends nothing)
etminan-verifier op notify-test --channel <name>   # send one synthetic finding through the REAL plugin

plugins verify confirms a channel is allowlisted, root-owned, and hash-matched but never executes it; notify-test closes that gap by actually attempting delivery and reporting the plugin's exit status. run also re-verifies every plugin each cycle, so a broken channel alarms on its own rather than being discovered only when needed.

See also