Troubleshooting¶
Common failure modes and how to diagnose them.
Start here¶
Both binaries fail loudly rather than degrading quietly, so the first move is always to read what they actually said.
# On a monitored host — the agent refuses to start on a bad config and says why:
systemctl status etminan-agent
journalctl -u etminan-agent -n 50
# On the verifier — is the daemon up, and who does it think you are?
etminan-verifier op ping
etminan-verifier op whoami
# Attest one host now, printing the full verdict:
etminan-verifier check web-01 --addr 10.0.0.5:7620
# Re-check every signature and the whole audit-log hash chain:
etminan-verifier baseline verify-signatures
op ping failing means the daemon is not running or its socket is not
reachable; op whoami failing means it is running but your UID holds no
registered identity (RBAC).
Enrollment: AK fingerprint mismatch¶
Symptom: enroll refuses, or a later check reports the host's AK fingerprint
no longer matches the enrollment record.
- The AK template must never change. The AK is created via
TPM2_CreatePrimaryfrom a fixed template; a different template derives a different key and breaks the fingerprint match. This is by design — it is what lets enrollment pin a stable identity with no persistent key storage. - A genuinely changed AK (TPM owner-clear or firmware change re-derived it) is not
an error to force past — use
rotate-ak, which re-binds the AK within the same TPM using the EK as the continuity anchor. A different EK means a different TPM and is refused as a re-enroll, not a rotation. - Re-enrolling an already-enrolled host is refused unless you pass
--force, so a fat-fingered re-run cannot silently re-pin a possibly-compromised host via trust-on-first-use. Only--forceafter re-confirming the new AK fingerprint out of band.
PcrDigestMismatch on check / run¶
Symptom: check or run reports a finding: the replayed PCR 10 does not equal
the value the TPM signed.
This is the verifier's core check — replay the IMA log, recompute PCR 10, and confirm it equals the signed quote. A mismatch is a real finding by default, but a few benign or environmental causes are worth ruling out:
-
The host rebooted. A reboot clears PCR 10 and starts a fresh IMA log, so the cumulative value the verifier carries belongs to a log that no longer exists and no replay can reach the signed digest again. Since 0.11 the verifier establishes this instead of guessing: on a mismatch it asks the host once more for its log from the first entry, and if replaying all of it reproduces the signed digest, the host rebooted. The cycle then carries on by itself and the reboot is reported as a
host-rebootedwarning, not an alarm. You do not have to do anything, and there is nothing to re-run. -
Log growth between quote and read. The IMA log keeps growing from background system activity in the small window between the quote being taken and the log being read. Etminan handles this with prefix-matching — a valid quote's digest must match some prefix of what was received; anything after belongs to the next cycle. If you see transient mismatches that clear on the next cycle, this is the mechanism working.
- Wrong IMA policy / PCR bank. The SHA-256 PCR bank must be read from
ascii_runtime_measurements_sha256, not the genericascii_runtime_measurements(which uses legacy SHA1-sized template hashes). A persistent mismatch across cycles on one host points at a genuine drift or a tampered log — treat it as a finding and review it.
Since 0.11 the verifier tells the agent where the delta must start, so it also
knows when the answer did not cover what it asked for. A cursor problem is
reported as log-cursor-gap, and a reboot clears itself and is reported as
host-rebooted. So a pcr-mismatch today is what is left after the benign
explanations above have been ruled out — and the finding text names each one it
ruled out. Treat it as a finding: keep the host's IMA log and the quote, and do
not re-enroll, which resets the AK binding and destroys the evidence.
log-cursor-gap on check / run¶
Symptom: a finding says the agent's log cursor is some number of bytes ahead of the verifier's.
This is not evidence of tampering, and the wording says so. It means measurements were consumed that this verifier never received, so no replay can reach the signed digest again. It comes from a host running an older agent, or one whose log was truncated or rotated underneath it.
Recovery: none needed on a 0.11 verifier — it names the starting offset in every request, so the two sides cannot drift apart, and a host whose log was truncated or rotated underneath it is re-read from the first entry on the next cycle. Re-enrolment is not needed and would reset the AK binding for a bookkeeping problem.
Then look at what failed before it — the gap is a symptom, and the cycle that caused it is the thing worth reading.
Agent and verifier upgrade together
The wire protocol carries a version and both sides must match; a mismatch is reported as a clear error rather than being papered over. 0.11 speaks v3. Upgrading only one side gives you a version error on every check — which is the intended outcome, because the alternative is an agent that quietly ignores the offset the verifier asked for.
Host stuck "catching up"¶
Symptom: a host reports CatchingUp cycle after cycle and never returns
Genuine.
A host with a genuine backlog (a large log delta) clears in a finite number of
cycles. state.json tracks consecutive_catching_up and resets it to 0 on every
Genuine cycle. A host still catching up past
ETMINAN_MAX_CATCHING_UP_CYCLES is a wedged agent — or a compromised host
feeding a fabricated log_truncated delta to suppress the mismatch alarm forever
— and Etminan alarms rather than staying silent. Investigate the agent process and
the IMA log on that host.
TPM device: /dev/tpmrm0¶
Symptom: the agent logs that it cannot open the TPM.
- The agent runs as an unprivileged
tss-group user and must use/dev/tpmrm0(the resource-managed device), never/dev/tpm0(the raw device, group-owned by root). Never rely on the library'sDeviceConfig::default(), which points at/dev/tpm0. Running the packaged agent as its real systemd user surfaces this; running manual tests as root masks it. /dev/tpmrm0 does not exist→ the TPM is absent or the kernel resource manager is not enabled.exists but cannot be opened→ check the device is grouptss, mode0660, and the agent user is in thetssgroup.- The device path is not configurable on this line — the agent always opens
/dev/tpmrm0. A host with no TPM cannot attest.
IMA log permission denied¶
Symptom: the agent logs a permission error reading the IMA measurement log.
The IMA measurement log is root:root 0440 by kernel design — no group membership
grants an unprivileged user read access. The fix is the systemd unit granting
AmbientCapabilities=CAP_DAC_READ_SEARCH (the narrowest capability that bypasses
the read-permission check), not running the agent as root. Check the unit
actually carries that grant. Run by hand as root the log reads fine because root
holds every capability — which is exactly why the packaged, unprivileged run is
the honest test.
IMA policy did not load / measurement count is 0¶
Symptom: /sys/kernel/security/ima/runtime_measurements_count reads 0, and
quotes carry an empty log.
- IMA accepts exactly one successful policy write per boot (failed writes do
not consume the slot). Write the complete policy in one shot, early at boot, via
etminan-agent-ima-policy.service— not incrementally. - Once loaded, the policy write node can disappear from securityfs entirely
(
ENOENT), not just start rejecting writes withEACCES.write-ima-policytreats either failure mode as "already loaded" whenruntime_measurements_countis already nonzero, so a second run is a no-op, not an error.
Agent won't start: verifier pinning¶
Symptom: the agent refuses to start, logging a verifier-pinning / TLS-mode error.
- Set
ETMINAN_VERIFIER_CERT_FINGERPRINTto the verifier's fingerprint (exactly 32 bytes / 64 hex chars). A wrong length or non-hex value is a FAIL. - The only alternative is
ETMINAN_ALLOW_PERMISSIVE_TLS=1, which accepts a client certificate from any caller and is intended only for the one-time enrollment bootstrap. Pin the fingerprint and drop the flag once enrolled.
Certificate expiry¶
Symptom: connections fail and the log names an expired certificate.
Rotate: on the agent run etminan-agent keygen-tls, then re-pin on the
verifier with op rotate-tls. For the verifier's own certificate, run
etminan-verifier op keygen-tls, update ETMINAN_VERIFIER_CERT_FINGERPRINT in every agent's
agent.env, then restart each agent — in that order (restarting an agent
before its env is updated locks it out until fixed).
Audit-log chain broken¶
Symptom: run or baseline verify-signatures reports the audit-log chain is
broken.
The error text names the exact failure — a seq gap (a row deleted or reordered),
a prev_hash mismatch (chain broken), a stored-hash mismatch (a row edited after
being written), a high-water above the surviving head (tail truncation), or a
missing/altered externally-anchored head row (rollback / DB restore). See
The audit log. On a verifier you administer, the most common
innocent cause is an incorrect backup/restore — see
Backup & restore for restoring a snapshot whose anchor matches
the DB generation. A break you cannot explain from your own operations is a
finding: treat the host as potentially compromised.
A finding won't leave the box¶
Symptom: you expect alerts but none arrive.
The alarm path is the only place a finding leaves the verifier, so test it directly rather than waiting for a real finding:
etminan-verifier op notify-preview # render templates against samples (sends nothing)
etminan-verifier op notify-test --channel <name> # send one synthetic finding through the REAL plugin
plugins verify confirms a channel is allowlisted, root-owned, and hash-matched
but never executes it; notify-test closes that gap by actually attempting
delivery and reporting the plugin's exit status. run also re-verifies every
plugin each cycle, so a broken channel alarms on its own rather than being
discovered only when needed.