Skip to content

feat(honeypot-http): bounded binary evidence with capture completeness (#3212) - #3408

Merged
Xore merged 3 commits into
mainfrom
oc/3212-binary-evidence
Sep 27, 2026
Merged

Xore merged 3 commits into
mainfrom
oc/3212-binary-evidence

Conversation

@Xore

@Xore Xore commented Sep 27, 2026 •

Copy link
Copy Markdown
Owner

Closes #3212.

The HTTP event now describes what it captured rather than implying it. The 64 KiB capture cap is unchanged; what is new is which bound was hit, and a hash that says exactly which bytes it covered.

Contract

New fields on the http-honeypot event, all under honeypot.* after filebeat's honeypot ndjson target:

field type meaning
body_capture_state string complete / truncated / unknown — set on every event, so an aggregate has no missing-value bucket
body_captured_bytes int raw captured length (pre-redaction)
body_read_error string the previously-discarded io.ReadAll error, when the read failed
body_declared_bytes int the Content-Length the client claimed — a claim, never a fact
body_encoding string base64
body_b64 string bounded (4 KiB) byte-safe head of the redacted captured body
body_sha256 string hash of the redacted captured prefix
body_sha256_scope string captured-prefix-redacted — names both bounds
java_marker string stream-magic / stream-magic-base64 — a marker observed, nothing more

classifyPayload is untouched: the ordered serialized-object class still covers both Java and PHP. java_marker is a separate field precisely so an earlier-priority match cannot erase it.

Layer trace

sensor (Go) -> filebeat ndjson target honeypot -> enrich_line normalization -> flattened mapping with ignore_above: 32000 -> backend record. Tests pin the two hops that were previously unpinned: the normalization pass retains the fields, and the record carries them to the browser.

Tests

Offline and inert throughout. No object graph, class name or executable content in any fixture; no deserializer is imported (pinned by an AST test); the sole decode is base64 over eight characters compared against four literals. No network, no live stack, no production query.

  • byte-for-byte survival of a Java stream header plus invalid-UTF-8 filler
  • marker negatives: ordinary text, unrelated binary, a PHP fixture, 4/5/6-char coincidental base64 prefixes
  • exactly-at-cap / over-cap / read-error -> complete / truncated / unknown
  • an sqli + Java body keeps payload_class: sqli and java_marker: stream-magic
  • every new field present under honeypot.* and inside the 32000-char flattened ceiling
  • credential canaries: the decoded evidence, not the encoded text, must not carry the secret

@github-actions

Copy link
Copy Markdown

Dependency Review

✅ No vulnerabilities or license issues or OpenSSF Scorecard issues found.

Scanned Files

None

@Xore
Xore enabled auto-merge (squash) September 27, 2026 13:58
@Xore
Xore force-pushed the oc/3212-binary-evidence branch 2 times, most recently from 6d46061 to 243f75a Compare September 27, 2026 17:00
The http-honeypot event made exactly one statement about a request body:
`body`, a Go string of whatever one io.ReadAll behind a 64 KiB LimitReader
returned, with its error discarded. Two defects follow from that, and no
consumer can repair either one.

A JSON string does not preserve arbitrary bytes. json.Marshal replaces
every byte that is not valid UTF-8 with U+FFFD, so a Java serialized
object, a PNG, or a lone 0x80 was stored as a different byte string from
the one received, with nothing in the document saying so. And a 40-byte
body and the first 64 KiB of a 4 MB one serialized to the same shape, so
"the attacker sent 40 bytes" and "we kept 64 KiB of something much
larger" were the same record.

The capture is now described rather than implied.

- body_capture_state is the authoritative answer to "is this the whole
  body": complete, truncated (the cap was hit and more existed), or
  unknown (the read failed). captureBody reads one byte past the cap, so
  "there was more" is knowable. The state is set on every request event
  including zero-byte ones, so an aggregate over it has no missing-value
  bucket and a genuinely short body answers "complete".
- body_read_error is the io error the old `body, _ :=` threw away, because
  "truncated" and "we never found out" are different failures.
- body_b64 plus body_encoding is a bounded 4 KiB head of the captured
  prefix, base64, so the received bytes survive exactly. A consumer
  decodes it rather than guessing. The head stays well under the 32000
  ignore_above on the flattened honeypot field so it is still indexed.
- body_sha256 with body_sha256_scope "captured-prefix" reuses the field
  name galah's body_sha256 already established. It is deliberately not
  the complete body's hash: that would mean draining whatever the client
  claims to send, which is unbounded work for a fact nobody reads.
- body_declared_bytes records the Content-Length the client claimed. It
  is recorded as a claim and never trusted as a fact.

`body` itself is unchanged -- the classifier reads it and analysts grep
it -- but is now documented as the lossy half, with body_b64 as the
byte-exact one.

Java evidence is a SEPARATE field, java_marker, alongside the existing
ordered payload_class and not replacing it. classifyPayload's
"serialized-object" is one ordered class shared with PHP; splitting Java
out of it would repurpose a value existing queries already mean and
would lose the PHP case whenever a Java-looking body arrived first. Kept
separate, a body that trips "sqli" three cases earlier and also starts
with the Java header keeps both observations, and neither can erase the
other. The marker is byte patterns only -- raw magic at offset 0, or
eight base64 characters that decode to those four bytes. No
ObjectInputStream, no class loading, nothing received over the socket is
ever instantiated; the only decode is base64's, on eight characters, to
compare against four literals. The values are "stream-magic" and
"stream-magic-base64" and mean the header was seen -- never that a gadget
resolved or anything executed.

The 64 KiB cap is unchanged. A bounded capture is the point; what is new
is the record of which bound was hit.

Reconciled with #3213, which invalidated one of this branch's premises.

This branch was written against a `body` that held the raw capture, and the
test for the new byte-safe field argued that body_b64 "adds no new class of
disclosed material" on exactly that basis: the same bytes `body` already
published, so nothing new was disclosed. #3213 (merged via #3412) redacted
`body`, and with it that argument. Had the evidence kept being built from the
raw capture, body_b64 would have become the only place a submitted credential
survived the event -- while every no-leak test still passed, because base64
does not preserve substrings and so the encoded blob cannot contain a raw
secret as a literal.

So the security property wins and the evidence follows the redaction. Both
`body` and `body_b64` are built from the same string, creds.redactedBody,
which is what makes them mutually consistent by construction rather than by
convention. `bodyCapture.apply` now takes that string as an argument instead
of reading its own copy of the capture, which is also what stops this from
becoming a second redaction path with its own rules: the scrubber stays the
one in credentials.go, called once.

The capture and the published bytes are now different things, and the fields
are split accordingly rather than blurred together. BodyCaptureState,
BodyCapturedBytes and BodyReadError describe the CAPTURE and are unaffected
by redaction, because redaction is not a fact about the wire -- so
truncation stays distinguishable from a short body, and a 40-byte POST still
answers "complete" while a clipped 4 MB upload answers "truncated".
BodyB64 and BodySHA256 describe the PUBLISHED bytes.

That distinction costs the hash label its old value, so the value changed.
`body_sha256_scope` was "captured-prefix", which says which part of the body
is covered and says nothing about the fact that the bytes covered are the
scrubbed ones. A consumer comparing the hash against a body held elsewhere
would have got an unexplained mismatch on every request that carried a
credential -- the "hash whose scope is ambiguous is worse than no hash"
defect this field exists to prevent, reintroduced one axis over. It is now
"captured-prefix-redacted". The label vocabulary is new in this branch and
nothing outside it consumes the value, so this is a rename within the change
rather than a contract break; galah's existing body_sha256 field name is
untouched, and the scope label's own wording is asserted against a literal
rather than against the constant, so reverting the value fails the suite.

What did not change: the 64 KiB cap, the one-byte-past-the-cap read that
makes truncation knowable, the 4 KiB bounded evidence head, the
ignore_above margin, the separate java_marker alongside the ordered
payload_class, and detection as byte patterns with no deserialization. The
classifier and the marker still read the raw bytes, as #3213 intended --
redaction must not cost the fleet a payload signature, and what leaves the
process there is a fixed marker name rather than a byte of the payload.

Tests: 84 in http-honeypot, plus cisco-asa's 100 and 586 in tests/docs, all
green; gofmt and go vet clean. #3212's no-credential test asserted the
opposite of what is now true -- that decode(body_b64) equals the raw form --
so it is rewritten to assert the redacted behaviour, and the replacement is
strictly stronger: the old version only checked that the secret was absent
from the base64 TEXT, which is nearly vacuous, while the new one decodes the
field and checks the bytes, and compares against inspectCredentials itself
rather than a second hand-written definition of "redacted".

Writing that second test found a real gap in the guard that was being relied
on. credentials_test.go's TestPasswordNeverReachesTheEvent searches for four
encodings of the secret, one of them the base64 of the secret, which catches
a body_b64 leak only when the bytes before the secret in the body are a
multiple of three -- base64 of a whole body contains the base64 of an
interior substring verbatim only when that substring starts on a 3-byte
boundary. Measured with the evidence deliberately built from the raw
capture, it caught 2 of its own 14 channels (24 and 6 bytes of prefix) and
passed the other 12 with the credential sitting in the event, JSON and
multipart among them. TestBinaryEvidenceNeverCarriesCredentialMaterial checks
the decoded bytes instead, which does not care where the secret falls, and
catches 7 of the 8 shapes where the secret is actually within the captured
region; the eighth sits past the 64 KiB cap and was never read at all.
Both mutations -- raw evidence, reverted scope label -- were applied and
confirmed to fail the suite.

OpenAPI: untouched. The event struct is a Go sensor type serialized to
Elasticsearch through Filebeat, not a utoipa contract type, and
`cargo run --bin openapi | diff` shows no drift. No workflow file was
touched, so zizmor is unchanged and nothing new is allowlisted.

Refs #3212, #3213
#3212)

The HTTP event now describes what it captured instead of implying it:
body_capture_state (complete/truncated/unknown), body_captured_bytes,
body_read_error, body_declared_bytes, and a byte-safe head (body_b64 +
body_encoding) whose sha256 carries an explicit scope label naming both
bounds -- captured prefix, post-redaction. The 64 KiB cap is unchanged;
what is new is which bound was hit and a hash that says which bytes it
covered. java_marker records the Java stream-magic observation as its
own field, separate from the ordered serialized-object class it shares
with PHP, so an earlier-priority classifier match cannot erase it.

Contract traced end to end: the Go sensor emits the fields, filebeat's
ndjson target nests them under honeypot.*, the generic enrich_line
normalization pass is proven to retain them, the flattened honeypot
mapping with ignore_above 32000 keeps them indexed, and the backend's
record carries them to the presentation layer, where the secrets
boundary leaves the evidence intact because the sensor redacted before
building it.

Offline tests only: byte-for-byte survival of a Java header plus inert
text with invalid UTF-8, marker negatives (ordinary text, unrelated
binary, a PHP fixture, coincidental base64 prefixes), cap boundary and
read-error cases, non-erasure by an earlier-priority class, and the
credential canaries across every new field.
@Xore
Xore force-pushed the oc/3212-binary-evidence branch from 243f75a to 62f33c8 Compare September 27, 2026 23:02
@Xore Xore changed the title feat(honeypot-http): say what was captured, and keep the bytes (#3212) feat(honeypot-http): bounded binary evidence with capture completeness (#3212) Sep 27, 2026
@Xore
Xore disabled auto-merge September 27, 2026 23:22
@Xore
Xore enabled auto-merge (squash) September 27, 2026 23:26
@Xore
Xore merged commit 5725ca0 into main Sep 27, 2026
120 checks passed
@Xore
Xore deleted the oc/3212-binary-evidence branch September 27, 2026 23:37
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Preserve bounded binary HTTP evidence and capture completeness alongside existing serialization classification

1 participant