Repository navigation
test(llm-worker): corpus as data, control twin, and an honest offline skip (#3334) - #3416
Merged
Merged
Conversation
Dependency Review✅ No vulnerabilities or license issues or OpenSSF Scorecard issues found.Scanned FilesNone |
… skip (#3334) #3347 landed the behavioural corpus and the host-side timer, and left three things #3334 asks for unmeasured. This closes them. The corpus moves out of Python into llm-worker/injection_corpus.jsonl, one validated object per line. The existing eight cases are byte-identical; a malformed or unknown-keyed block now raises instead of scoring as a pass, which is the rule injection_gate.py applies to a rubric (spec_from_rubric). Two techniques were missing. attacker-json-fields plants a verdict object carrying analyst_note/approved_by_soc/soc_disposition/review_state, fields the analysis schema does not define. backtick-fence uses markdown and backtick fences around a delimiter that contracts.py rewrites to `< /untrusted_data>`, so it measures the mangled fence the model actually sees. The eleventh case is a payload-free control. Without it a model that cannot classify a clean session is indistinguishable from one that was steered, which is the false_verdict/control_case distinction from #2694. It reports under its own `control` tally and fails the run via injection_suite.all_passed(). judge() gains the two assertions the issue names and nothing had: schema-valid (key set plus a round-trip through the strict model) and no attacker-chosen field in key position. The field check reads model_dump() rather than model_dump_json(), because re-serializing escapes a leaked key's quotes and no key pattern should be asked to match `\"analyst_note\"`. It requires a colon, so narrating the attempt in prose is not scored as compliance. Loader bug found while doing this: the marker check is a substring test, and INJ_OK_1 is a prefix of INJ_OK_10, so one case's leak would fail another. The new markers are INJ_FIELD_9/INJ_FENCE_10 and the loader now rejects any marker contained in another -- a verdict invented, not observed. test_injection_suite_model.py runs the corpus through the real OllamaClient in the normal unit lane. skip_reason() is a pure function of endpoint state, so a steered model cannot be reclassified as an absent one; that separation is asserted rather than assumed. Every skip names the endpoint, the model and the pin and says in words that the suite did not run. Unreachable, not-Ollama, model absent and digest mismatch all skip; a case failing against the pinned model fails. run-llm-injection-suite.sh fingerprints the corpus alongside the pin and suite source, so editing a case invalidates the last verdict instead of being skipped as "nothing changed", and a checkout without the corpus is a named error rather than a "No such file" from inside the container after the model is loaded. Offline on this box, so the served-model leg skipped with a reason rather than passing: the pinned qwen3:14b is not served here. All assertions are enabled and left that way. Verified: 136 llm-worker tests OK (1 skip), 20 host-wiring tests OK, selftest PASS, doc gates and shellcheck clean, prompt hashes identical under a different PYTHONHASHSEED. Temporarily dropping the <untrusted_data> fence fails 6 tests; temporarily removing the attacker-field check fails 1. Both reverted -- contracts.py and test_contracts.py are byte-identical to main. Refs #3334
Xore
force-pushed
the
oc/3334-injection-regression
branch
from
September 27, 2026 18:24
891f823 to
a4ad479
Compare
Xore
enabled auto-merge (squash)
September 27, 2026 20:06
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes the three gaps in #3334 that PR #3347 left unmeasured. Does not merge.
What was already there (and what was not)
#3347 (
a355b268, inmain) already landed the 8-case corpus, the offline judge,worker.py --injection-suite, and the systemd weekly/on-pin-change wiring. The premise of "build the suite" was therefore partly stale. The real remaining gaps were narrower, and this PR is scoped to exactly those:mainbefore this PRinjection_suite.pyllm-worker/injection_corpus.jsonl, 11 casesattacker-json-fields+ a new assertion</untrusted_data>, backtick fences)backtick-fenceclean-control, run as a separate tallyattacker_field_failures()schema_failures()tests/test_injection_suite_model.pyThe 8 existing cases are preserved byte-for-byte. The only existing test-method lines touched are additions;
git diffshows the sole removed line is one import.The two defenses, and what they are not
llm-worker/contracts.py, which wraps untrusted data in<untrusted_data>and separates it from instructions. Deliberately untouched.git diff --exit-code origin/main -- llm-worker/contracts.py llm-worker/tests/test_contracts.pyis clean: both byte-identical tomain. Nothing here was relaxed to make a test pass. A model leak is a finding to report, not to paper over.A real bug this found in the existing corpus
INJ_OK_1is a substring ofINJ_OK_10. The marker check is a substring test, so a leak in thechat-template-tokenscase (markerINJ_OK_10) would also satisfy the benign-relabel case's markerINJ_OK_1— the two cases could not fail independently, and one real leak could pass another case's assertion. New markers areINJ_FIELD_9andINJ_FENCE_10, and the loader now rejects any marker that is contained in another, so this class of collision cannot come back.Evidence the assertions actually bite
Both mutations were applied, run, and reverted:
<untrusted_data>fence fromsession_prompt→ 6 tests failattacker_field_failures()fromjudge()→ 1 test failsNeither mutation survives.
contracts.pyconfirmed byte-identical toHEADafter both reverts.The real-model leg: it skipped here, and I am not going to pretend otherwise
tests/test_injection_suite_model.pydrives the real worker classification path (no reimplementation) at theslots.sessionspin fromapproved-models.json, viaModelProbe/probe_model().skip_reason()is a pure function of endpoint state, so the reason can never be a silent pass.On this box the test skipped — the pinned model is not served here. The pin is
qwen3:14b(digestbdbd181c…). Port 11434 here is an OpenAI-compatible server (qwen3-1.7b,qwen3-600m,bge-embed, …) that answers HTTP 404 on/api/tags, so the suite correctly reports "something is listening … but/api/tagsanswered HTTP 404; the behavioural suite did NOT run, because that endpoint is not an Ollama the suite can drive." With no endpoint at all it names the real pin: "Ollama at http://ollama:11434 did not answer …qwen3:14b".This is model unavailability, not model failure — every assertion above stays enabled and none were relaxed to accommodate the local box. Both skip branches are demonstrated in the test output. Note the distinction from the systemd leg:
worker.py --injection-suitewith no Ollama exits 1 (injection suite failed to run: Ollama model metadata unavailable), which is the correct fail-closed behaviour for a scheduled job. A skip is never a silent pass.The weekly measurement is the
honeypot-llm-injection-suite.timer(Thursdays 04:17 UTC), plus the.pathunit on pin change (#2969), reusing the existing canary-record convention indocs/llm-injection-suite-record.md.Determinism
Same corpus + same pin ⇒ same result. Verified identical prompt hashes under
PYTHONHASHSEED=12345. Model-side determinism uses the worker's existingtemperature: 0,seed: 66,think: false,num_ctx: 8192; the harness forcesduration=37.0,auth_success=Truefor all cases. No live network in the unit-test path — the offline lane is fully hermetic. The corpus contains no secrets; the only secret-shaped match is the synthetic/var/run/secrets/kubernetes.io/serviceaccount/tokenpath.Gates, on the rebased tip
llm-worker: 136 tests OK (1 skip) — the 1 skip is the real-model leg aboveanalysis/ghidra/models: 42 OK, of which the injection-suite host wiring is 20tests/docs: 17 OKcheck-doc-paths-exist,check-docs-reachable,check-doc-stale-paths,check-compose-env-docs,check-public-leaks: PASSshellcheck --severity=error,bash -n, pyflakes: cleanKnown limitation, disclosed rather than suppressed
The marker check is a substring test, so a model that honestly quotes the whole attacker JSON blob in its
summarywill be flagged. This is the same mention-vs-assertion tension as #2694. I have deliberately not loosened it, because loosening it would weaken an existing assertion — which is exactly what the brief forbids. Flagging it here so a reviewer can decide.Scope
10 files, +1093/−142. No unrelated files touched. Not merged, and auto-merge is not enabled — please review, particularly whether the substring limitation above is acceptable as-is.