Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
243 changes: 184 additions & 59 deletions .github/workflows/diagnostics.yml

Large diffs are not rendered by default.

61 changes: 57 additions & 4 deletions docs/CI-CD.md
Original file line number Diff line number Diff line change
Expand Up @@ -1736,10 +1736,10 @@ scratch twice, reaching the same blocker both times.

**No workflow edit is needed.** Every `secrets.VPS_*` / `secrets.DOMAIN`
reference already sits inside a job that declares
`environment: production-vps` — `deploy.yml`'s `vps` job (`:207`, environment
at `:210`), `diagnostics.yml`'s `vps` job (`:252`/`:255`), and
`environment: production-vps` — `deploy.yml`'s `vps` job (`:233`, environment
at `:236`), `diagnostics.yml`'s `vps` job (`:439`/`:442`), and
`vps-start-blackhole.yml`'s `start-blackhole-profile` job (`:22`/`:24`). The
`home` jobs (`deploy.yml:21`, `diagnostics.yml:76`) read none of the five.
`home` jobs (`deploy.yml:21`, `diagnostics.yml:112`) read none of the five.
Environment secrets also shadow repository secrets of the same name, so
*writing* the environment copies is non-breaking on its own.

Expand Down Expand Up @@ -1812,12 +1812,65 @@ Two checks depend on host provisioning rather than on the workflow (#3312):
`LIBVIRT_DEFAULT_URI=qemu:///system`, because a non-root `virsh` otherwise
talks to the empty per-user session and reports every network missing.
Group membership only takes effect across a runner restart, so unlike the
helper grant it does need a full installer run.
helper grant it does need a full installer run. The audit's own host-side
expectation — that the sandbox stack is *up* — is suspended by a dated
declaration while it is deliberately down; see
[honeypot-network-isolation.md](honeypot-network-isolation.md#5-declared-stand-down).

The OIDC discovery probe runs **from the VPS** over the job's SSH key.
Cloudflare answers 403 to GitHub-hosted runner address ranges, so the runner's
own result is printed for information only and never fails the job.

### Every finding is categorised (#3312)

This workflow failed 160 consecutive scheduled runs without anybody triaging
it, which means it carried no signal at all. The reason was not that the
findings were wrong — several of them were correct — but that a real host
fault, a lane that had never been able to see anything, a check asking from a
vantage point the endpoint answers differently to, and a deliberate stand-down
all produced the same thing: a red X, an `::error::` line with nothing on its
subject, and a body nobody had time to read. A run that lists five unrelated
things under one heading gets triaged by ignoring it.

So every finding now carries one of four categories. The category is the
annotation's own title, and each job ends with a ledger — one table, one row
per finding, plus the counts — so triage is reading six lines rather than
reconstructing a run from its log:

| category | Meaning | Fatal on a scheduled run |
|---|---|---|
| `fault` | A real regression: the pipeline or the host is broken | yes |
| `runner-config` | The check could not run because of how the runner or this repository's environment is configured — a helper not installed, a grant not applied, a secret unset | yes |
| `unmeasured` | The check did not run, and that is not a pass | yes |
| `expected` | A deliberate, declared absence | no |

`runner-config` being fatal is deliberate, and it is the same position
`scripts/verify-deploy.sh` already takes with its exit 2: folding "could not
tell" into a pass is the outcome that makes a check worse than not having it.
#3283 is what that costs when it is wrong — Elasticsearch at 1000/1000 shards
with every sensor's events dead-lettered for six days, in a lane whose only
question is whether the pipeline is flowing. What changes is that it is its own
category with its own title, so "this lane has never been able to measure
anything" is tellable apart from "the pipeline is broken" without opening the
log, and the fix is the operator command the finding names rather than the
symptom.

The vocabulary lives in `scripts/diagnostics-lib.sh`, which every step sources;
`alert` takes a category and refuses a non-fatal one, `note` records a finding
that must not redden the run, and `diag_ledger_report` prints the table. The
isolation audit has its own finer-grained labels and its own footer, and the
step carries that line into the ledger verbatim rather than re-deriving it — one
source of truth for the counts.

Both jobs now check the repository out (#2908). Every other step still reads
the deployed stack under `/opt/stacks/apiary`; the checkout is for the scripts,
so a check asks its question with the code that was just fixed. Running
`isolation-audit.sh` from a copy refreshed only by a `workflow_dispatch`-only
deploy is why the red X kept naming things that had already been fixed. The
deployed copy is still diffed against `origin/main` and reported when it drifts,
because drift is a real finding for everything else on the host that runs a
deployed script.

### Diagnostics vs. mutating deploy

```mermaid
Expand Down
78 changes: 76 additions & 2 deletions docs/honeypot-network-isolation.md
Original file line number Diff line number Diff line change
Expand Up @@ -109,7 +109,10 @@ here to implement and no issue to open. Verified by
`.github/workflows/diagnostics.yml` `home` job.

- **sshd** bound to the management address only, never the honeypot-facing one.
`ss -tlnp | grep :22` is the check.
The check asks `ss -tlnp` for a port-22 listener, and it distinguishes the two
answers that used to be the same one: "nothing is listening on :22" (`OK`) and
"the listen address could not be read at all" (`UNMEAS`, fatal). `ss … |
grep :22` exits 1 for both, which is why the audit does not use it.
- **libvirt socket** — `/var/run/libvirt/libvirt-sock` as `srwxrwx---
root:libvirt`, and `listen_tcp = 0`. The TCP socket is off by default; the
failure mode is someone enabling it while debugging remotely and leaving it.
Expand All @@ -119,7 +122,78 @@ here to implement and no issue to open. Verified by
- **Docker socket** never mounted into a honeypot-facing container. It is root
equivalent — a container that has it has the host.

## 5. Pitfalls
## 5. Declared stand-down

Freeing host RAM and CPU for a heavy leg is standing practice here: #3135
paused honeypot containers for a training run, and the sandbox VMs are a much
bigger claim on the same memory. The problem is not the stand-down. The problem
is that a stand-down and a broken host are *absence*, so the audit reported
both identically — `FAIL`, exit 1, no difference a reader could act on. That is
the "cries wolf on a healthy host" case, and it is the same defect as running
no check at all, because a run that is always red is a run nobody reads.

So absence gets a name, and the name is a dated declaration on the host:

```bash
sudo scripts/sandbox-standdown.sh declare \
--issue '#NNNN' --until YYYY-MM-DD --reason 'why, in one line'
sudo scripts/sandbox-standdown.sh show # ACTIVE or INACTIVE, with the reason
sudo scripts/sandbox-standdown.sh clear # when the stack comes back
```

The file is `/etc/apiary/sandbox-standdown` (root-owned, `0644`, because the
audit runs unprivileged as `github-deploy-runner`; overridable with
`APIARY_STANDDOWN_FILE` for testing). While a live declaration is in force the
audit reports the sandbox objects it covers as `EXPECT` and exits 0, and it says
so out loud — owner issue, expiry, and the reason — in the body, in its footer,
and in the diagnostics ledger.

The declaration is deliberately hard to leave lying around, and every one of
these failures is itself a `FAIL` rather than a shrug:

| Declaration | Why it does not count |
|---|---|
| no `issue:`, `until:` or `reason:` | Nothing to hold the exception to a window or an owner |
| `issue: 1` rather than `issue: #1` | Not a reference. An exception nobody can be held to is the thing that rots into a permanent excuse |
| `until:` in the past | A window that has run out excuses nothing |
| `until:` unparseable | The window is unknown, and an unknown window is not a long one |
| `until:` more than `APIARY_STANDDOWN_MAX_DAYS` (14) out | That is a permanent posture change, not a stand-down, and it wants an issue and a decision rather than a file |

What a declaration does **not** cover, deliberately: anything that is present
and wrong. A libvirt network that exists and *forwards*, or an `ACCEPT` rule in
the `FORWARD` chain for `virbr-sandbox`, is a measured fault with a live
declaration in force — the exception is for absence, not for amnesty.

## 6. What the audit's labels mean

Every line the audit prints carries one, so a reader never has to count columns
or infer intent from a wording change:

| Label | Meaning | Exit contribution |
|---|---|---|
| `OK` | Measured, and it agrees with the invariant | none |
| `FAIL` | A measured violation, an exception that no longer applies, or a barrier that could not be read | **exit 1** |
| `EXPECT` | Absent on purpose, behind a live dated declaration | none |
| `UNMEAS` | The check did not run. Counted and named, never folded into a pass | **exit 1** for the isolation barriers |
| `WARN` | A triaged gap with a named owner issue, tracked and visible | none |
| `--` | Not applicable, or a note | none |

The footer prints the counts and one verdict line, so the answer to "is this
host safe" is the last line of the output rather than a reading of the whole
body:

```
isolation-audit: categories -- 0 unmeasured, 0 expected-by-declaration, 0 triaged gap(s), 0 fault(s)
isolation-audit: VERDICT PASS -- every check that could be measured agrees with the invariants (0 object(s) excused by a live declaration)
```

`UNMEAS` being fatal for the barriers is the load-bearing part. "The `FORWARD`
chain could not be read" and "the `FORWARD` chain is not DROP" are different
findings, and only the second one says the host is unprotected. Reporting
"could not tell" as agreement is the outcome that makes a check worse than not
having it.

## 7. Pitfalls

| Symptom | Cause |
|---|---|
Expand Down
Loading
Loading