Skip to content

test(ghidra): pin the Tier B loader to real Ghidra output, not a mock shaped like it - #3445

Merged
Xore merged 2 commits into
mainfrom
oc/1805-ghidra-fidelity
Sep 28, 2026
Merged

Xore merged 2 commits into
mainfrom
oc/1805-ghidra-fidelity

Conversation

@Xore

@Xore Xore commented Sep 28, 2026

Copy link
Copy Markdown
Owner

Pins the corpus scorer's Tier B path to real recorded Ghidra output instead of a mock shaped like it, and measures the standing gap between the qualification gate's hand-written fixtures and production's renderer.

Refs #1805. Does not close it -- the gate-vs-production convergence is #2641's job and needs the GPU; see the last section.

The gap

Every Tier B test in the tree synthesises its own cache entry, so they agree with whatever load_tier_b_evidence() happens to expect:

  • test_record_baseline.py builds its decompiled map inline
  • test_ghidra_cache.py hand-writes its entries

That means a drift in the real service's response shape -- the pseudocode key, the s key on a string, the decompiled map's value type -- passes CI against a mock that was updated along with the code. load_tier_b_evidence() had never been driven against a real Ghidra response in any test.

The issue body also claims the gate feeds hand-written prose; that part is already covered. ghidra_cache.py and the Tier B path exist, and worker/tests/test_ghidra_worker.py::test_approved_contract pins the shared prompt contract. The uncovered piece was this.

The fixture is real

Recorded Tier B transcript for process_and_injection from run 20260825T190415Z-39f27b7b (docs/benchmarks/runs/2026-08-25-20260825T190415Z-39f27b7b/transcripts.jsonl) -- the injection case, so the issue's highest-risk item.

It keeps the artefacts real Ghidra produces on a relocatable .o and a hand-written fixture would not, which is what makes it a fidelity check rather than a second mock:

halt_baddata()          /* WARNING: Unknown calling convention */    __pid_t _Var1

cache_key / ghidra_version / post_scripts_sha256 are null because the recording predates keyed Tier B caching -- the transcript's own reproducibility.ghidra_cache_key is null too. Full provenance in the file, including the corpus manifest sha256.

Round trip, byte-exact: fixture written back as a cache entry -> load_tier_b_evidence() -> build_prompt(tier="B") reproduces the recorded user_prompt exactly, sha256 98a9fb915dd6d6a19fa17eb048625eb58a8a1757a2c044a990d83e9f434fa830.

The file is inert test data. Nothing deserialises it into an object or executes it -- it is json.loads'd and formatted into a prompt string, and a test asserts the top-level key set is exactly the expected data keys.

The divergence, measured rather than asserted away

The gate's TRIAGE_CASES are hand-written prose and are not in the format _evidence() emits. Three concrete differences:

gate production _evidence()
header IMPORTS (6/6): IMPORTS (6 shown of 6):
imports human's order sorted
function line int main(void) main @ 0x401000 (420 bytes)

Closing this is not a formatting fix -- it re-scores the approved cohort, which is #2641's job. So these tests pin the distance in both directions and fail if it widens. test_the_two_scorers_do_not_agree_today fails loudly if the gap ever closes, since then the file owes an update.

This recording predates #2643, so it has no STRINGS block and the injection payload genuinely never reached the model. assert_injection_present() correctly reports the case as uncovered rather than passing, and that is asserted -- plus the positive direction, so the gate cannot be blind while still looking honest.

Verification

  • test_ghidra_fidelity.py -- 18 passed
  • 12/12 mutations caught across load_tier_b_evidence, build_prompt, assert_injection_present, _evidence and the gate
  • benchmark suite: 16 files OK, 0 failed

Three mutations survived a first version of this file, and closing them is the part worth reviewing:

  • dropping the numeric sort key (sorted(decompiled, key=lambda a: int(a, 16))) -- every address in the fixture is 6 hex digits, so a plain sorted() produces byte-identical output. No real Ghidra output in the tree has a mixed-width decompiled map, so this needed its own assertion with a documented short address rather than pretending the shuffle test covered it.
  • silently .get()-ing pseudocode -- would shrink the exam instead of failing.
  • emptying assert_injection_present's haystack -- the negative-only test was satisfied by an always-False implementation, which turns every injection case into "uncovered" rather than "failed".

No production code changed. No model, no Ghidra, no service, no network, no GPU.

Not done here

Gate-vs-production convergence, and the cohort re-score that would validate it -- #2641. Also unchanged: the recorded evidence's missing STRINGS block is a property of that recording, not a defect asserted away.

@github-actions

Copy link
Copy Markdown

Dependency Review

✅ No vulnerabilities or license issues or OpenSSF Scorecard issues found.

Scanned Files

None

… shaped like it

Every Tier B test in the tree synthesises its own cache entry --
test_record_baseline.py builds its decompiled map inline, and
test_ghidra_cache.py hand-writes its entries -- so they agree with
whatever load_tier_b_evidence() happens to expect. A drift in the real
service's response shape (the `pseudocode` key, the `s` key on a
string, the decompiled map's value type) would pass CI against a mock
that had been updated along with the code.

The fixture is real recorded Ghidra output: the Tier B transcript for
process_and_injection from run 20260825T190415Z-39f27b7b, the
highest-risk case in the issue because it carries the injection
payload. It keeps the artefacts real Ghidra produces on a relocatable
.o and a hand-written fixture would not -- halt_baddata(), "Unknown
calling convention", __pid_t _Var1 -- which is what makes it a fidelity
check rather than a second mock. cache_key/ghidra_version/
post_scripts_sha256 are null because the recording predates keyed Tier
B caching; the transcript's own reproducibility.ghidra_cache_key is
null too.

Round trip, byte-exact: the fixture written back as a cache entry ->
load_tier_b_evidence() -> build_prompt(tier="B") reproduces the
recorded user_prompt exactly, sha256
98a9fb915dd6d6a19fa17eb048625eb58a8a1757a2c044a990d83e9f434fa830.

TriageSlotDivergenceTest pins the still-open gap instead of asserting it
away. The gate's TRIAGE_CASES are hand-written prose and are not in the
format _evidence() emits: no truncation disclosure, no ordering claim,
imports in a human's order. Closing that is not a formatting fix -- it
re-scores the approved cohort, which is #2641's job and needs the GPU.
So the tests name the distance in both directions and fail if it
widens; test_the_two_scorers_do_not_agree_today fails loudly if the gap
ever closes and this file owes an update.

No production code changed. No model, no Ghidra, no service, no network.

Verified: test_ghidra_fidelity.py 18 passed. 12/12 mutations caught
across load_tier_b_evidence, build_prompt, assert_injection_present,
_evidence and the gate -- including three that survived an earlier
version of this file and whose absence is the point: dropping the
numeric sort key, silently .get()-ing `pseudocode`, and emptying
assert_injection_present's haystack. Those three now have their own
assertions, and the address-order docstring no longer claims the shuffle
test covers the numeric sort when it does not. benchmark suite: 16
files OK, 0 failed.

Refs #1805
@Xore
Xore force-pushed the oc/1805-ghidra-fidelity branch from e6ea720 to 4622502 Compare September 28, 2026 07:43
@Xore
Xore merged commit b5e2db8 into main Sep 28, 2026
118 checks passed
@Xore
Xore deleted the oc/1805-ghidra-fidelity branch September 28, 2026 08:15
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant