Skip to content

Knowing the verifier is alive

Every host Etminan watches is watched by the verifier. This page is about the one machine that watches nothing else.

The verifier reports on your fleet. When it stops, it stops reporting that too — and a monitoring system that has gone quiet looks exactly like a fleet with nothing wrong. This is the one failure the product cannot tell you about itself, so it has to be arranged from outside.

The old advice, and why it was not enough

Until 0.11 the hardening guide suggested a dead man's switch:

etminan-verifier run && curl -fsS https://<your-endpoint>/<check-id>

That is still worth having, and it is still true that if it stops firing, you should look. But consider what it proves: a process exited 0. It does not say which verifier reported, how far it had got, or that the report came from the verifier at all — anything that can reach your monitoring endpoint can send that ping.

What the verifier emits now

Set one variable:

# /etc/etminan-verifier/verifier.env
ETMINAN_SELF_ATTEST_TO=/var/log/etminan/integrity-summary.jsonl

and every run appends a single line:

{"counter":418,"emitted_at":"2026-08-30T11:00:04Z","head_seq":9127,
 "head_hash":"6a1f…","identity_pub":"c4b2…","signature":"9d0e…"}

Four facts and a signature over all of them:

identity_pub which verifier this is
counter how many cycles it has completed — it only ever rises
head_seq / head_hash how far its audit log had got
signature that it really is that verifier saying so

With no ETMINAN_SELF_ATTEST_TO there is no file, no write, and no change in behaviour of any kind.

A file, not a mail

The audit anchor is mailed once a day and is evidence somebody keeps. This arrives every sweep and is operational data. Hourly mail is noise, noise gets filtered, and a filtered control is switched off while still appearing to be on. Ship the file with whatever you already use to ship files — a log shipper, rsync, a nightly scp.

The alarm is a different matter, and it does go by mail. See below.

Checking it, somewhere else

The check belongs on the receiving side, because a verifier that vouches for itself on its own hardware has said nothing.

Install etminan-witness there — not the verifier:

apt-get install etminan-witness      # or: dnf install etminan-witness

That package is a single binary and its manual page. No daemon, no timer, no service user, no key, no configuration file. It coexists with anything already on the machine and needs neither verifier package; one build serves both editions, because checking a signature is not an Enterprise feature.

Why a separate package at all

Both commands also exist, unchanged, inside etminan-verifier — nothing was moved. What did not exist was a way to install them anywhere else, and that made the claim they rest on untestable in practice: getting them onto a mail host meant installing a second verifier, with its daemon, its three systemd units, its service user and its mTLS setup step. Nobody does that to read a line of JSON, so the check that matters most was the one nobody ran.

tail -1 integrity-summary.jsonl | \
  etminan-witness verify-summary \
    --expect-identity <the daemon's public key hex> \
    --max-age-seconds 5400

--state defaults to /var/lib/etminan-witness/last-seen.json, a directory the package creates, so it can usually be left out. Keep that file, and keep it here rather than anywhere the verifier can reach: it is the memory that makes a rollback visible at all.

verify-summary opens no database, holds no key, contacts no daemon. It asks three things:

  1. Is this signed by the verifier I pinned? --expect-identity is required and is not read out of the summary — one that verified against its own embedded key would verify perfectly after somebody re-signed it with theirs. The daemon prints its public half at every start.
  2. Did the counter go up? So last month's "all fine" cannot be sent again.
  3. Did the audit log stay where it was or move forward? A head that goes backwards is a rollback; the same entry number returning with a different hash is a rewrite in place.

--max-age-seconds adds a fourth: is this recent enough. Set it a little above your sweep interval — an hourly sweep with a 90-minute window tolerates one missed run and alarms on two.

--state is the memory that makes checks 2 and 3 possible. Keep it on the receiving side, and keep it: it is not advanced when a summary is refused, because accepting a rolled-back head as the new normal would launder exactly the rollback that was caught.

Exit status

0 verified
3 integrity alarm — tampering, a rollback, or a summary too old
1 the input could not be evaluated (unreadable file, malformed JSON)

3 and 1 are separate on purpose. "The history was rewritten" and "you pointed me at a truncated file" should not wake the same person the same way.

Where the alarm goes

On a fresh receiving host: to standard error, and to the exit code. That is not a limitation to work around, it is the integration — wire whatever runs etminan-witness to treat exit 3 as a page:

# /etc/systemd/system/etminan-witness.service
[Service]
Type=oneshot
ExecStart=/bin/sh -c 'tail -1 /var/log/etminan/integrity-summary.jsonl | \
    /usr/bin/etminan-witness verify-summary --expect-identity <hex> --max-age-seconds 5400'
# exit 3 fails the unit; OnFailure= is then your alerting path
OnFailure=your-alerting@%n.service

The witness package installs no mail configuration, no notify channels and no /etc/ directory at all, and it deliberately does not: a package that arrived on a mail host and started assuming things about how that host alerts would be the same overreach that made installing the verifier there unacceptable.

A refusal is a full CRITICAL finding, kind integrity-summary-rejected. On a freshly installed receiving host it goes to standard output and the exit code and nowhere else — there is no mail on that machine until you put it there. Where the host already has the machinery it takes the ordinary route: mail, every notify channel configured there — PagerDuty, Slack, a webhook of your own — and the SIEM. Configure those on the machine that runs etminan-witness, never on the verifier: the verifier is the thing being doubted, and an alarm that it has to deliver about itself is not an alarm.

That the finding reuses the ordinary route rather than a private one is deliberate. A separate, second alerting path for the one alarm that says the trust root itself is wrong would be the wrong way round — the alarm least likely to be wired up and most important to hear.

It is the only finding whose host_id is verifier rather than a monitored host, because it is the only one about the machine doing the monitoring.

How silence becomes an alarm

Two mechanisms, and the first needs no configuration at all.

The counter catches a stopped verifier by itself. When the verifier stops emitting, the next scheduled check on the receiving side reads the same last line it read before, and refuses it: the counter did not advance. Exit 3. Nothing had to be tuned to the emitting cadence, so nothing can be tuned wrong. A verifier that has been switched off, wedged, or cut off from whatever ships its file all produce this within one receiver interval.

--max-age-seconds catches the other case: a summary that is new but late. The counter went up, so the first check passes it; the clock says it should have arrived an hour ago. Set the window a little above the emitting cadence — an hourly sweep with a 90-minute window tolerates one missed run and alarms on two.

etminan-witness verify-summary … --max-age-seconds 5400
INTEGRITY ALARM - this summary is 5480s old and the window is 5400s
                  — it is late, or the verifier stopped emitting

Both were checked on the packaged build, 2026-08-30.

The gap this does not close

Nothing on either machine reports that the receiver itself stopped running. The checks above are what the receiving side says when it runs. If the thing that runs it dies — the timer disabled, the host down, the log shipper broken so the file never arrives — then nothing is refused, because nothing is checked.

Alert on the schedule of the receiving check:

if no verify-summary run has succeeded in 2 hours → page

This is the same discipline the dead man's switch asked for, and it is still required — one layer further out than before. What changed is everything inside it: a stopped verifier is now caught by the check rather than by the absence of one, and what arrives when things are working is not a ping anything could have sent but a signed statement of who is alive and how far they have got.

And what a live compromise can still do

The signing key lives on the verifier. A verifier that is compromised while it runs — not tampered with and left, but owned and still executing — can keep emitting summaries with a rising counter, a fresh timestamp and a valid signature, over an audit head that has stopped moving because it has stopped looking at anything. To a receiver, that is indistinguishable from a quiet week.

This mechanism catches tampering, rollback and silence. It does not catch that. Closing it needs a second Etminan instance attesting the verifier host itself, so the statement is signed by a TPM rather than by a key the attacker holds. That is not built.

See also