Skip to content

fix(ghidra): raise pinned benchmark output_tokens 512 -> 4096 (#1804 follow-up) - #3308

Merged
Xore merged 1 commit into
mainfrom
fix/benchmark-output-tokens
Sep 25, 2026
Merged

Xore merged 1 commit into
mainfrom
fix/benchmark-output-tokens

Conversation

@Xore

@Xore Xore commented Sep 25, 2026 •

Copy link
Copy Markdown
Owner

The 512-token cap was truncating Ghidra triage and benchmark answers, not limiting them.

Evidence

Measured on the Tier A corpus, 2026-09-25, with #159's own harness:

cap base cut off fine-tune cut off
512 14 of 14 9 of 14
2048 1 of 17 0 of 17
4096 0 of 17 0 of 17

The case that survived 2048 was file_write_persist, done_reason: length at 5127 chars (~1900 tokens), so 2048 is not enough either. 4096 is validated by measurement, not chosen for tidiness.

This is not new. #2694 already recorded the same cap truncating injection conclusions in 23 of 30 Tier B answers. Any baseline recorded at 512 is suspect in the same direction — and the base model is penalised hardest, because it writes longer answers. That is what made the #1804 candidate-2 A/B look like a 6-point win for the fine-tune when it is 2.

The change, and why it is three files

ghidra moves to 4096 in approved-models.json (qualification_request and runtime_request), and ghidra-worker.py's TRIAGE_OUTPUT_TOKENS default moves with it. Those are independent defaults, and CI already pins them together — the Ghidra output cap matches the approved runtime request check failed on the manifest-only change, which is the invariant working. Had the manifest been changed alone, the live worker would have kept truncating while the manifest claimed otherwise.

revdeck.qualification_request also moves to 4096: it is the benchmark's own measurement slot, and the largest measured prompt is 2008 tokens, so 2008 + 4096 = 6104 < its 8192 context.

sessions deliberately stays at 512. That slot is llm-worker session classification, not reverse engineering, and llm-worker clamps LLM_OUTPUT_TOKENS to max=2048 under an 8192 context. Its 512 was never a truncation defect, and raising it would have made the manifest value unrepresentable in the consumer that serves it.

Tests: the one hardcoded 512 expectation in test_ghidra_worker.py is updated. Nothing else needed loosening — the governance and baseline tests carry their own local request dicts, and 76 pass.

Re-measured at 4096

tierA_approved_qwen3_14b_cap4096_run1.json is committed: qwen3:14b scores 76/83 (91.6%) across 17 cases, 0 truncated, 0 done_reason: length. That is the model approved-models.json names for all three slots, so it is the baseline most likely to be quoted anywhere, and it has now been measured clean.

Known and not fixed here

Refs #1804, #1947, #2694

@github-actions

Copy link
Copy Markdown

Dependency Review

✅ No vulnerabilities or license issues or OpenSSF Scorecard issues found.

Scanned Files

None

The 512 cap truncated answers rather than limiting them. Measured
2026-09-25 on the Tier A corpus: qwen2.5-coder:7b-instruct-q4_K_M was
cut off on 14/14 cases at 512, still 1/17 at 2048 (done_reason: length
on file_write_persist, 5127 chars), and 0/17 at 4096. #2694 had already
recorded the same cap truncating injection conclusions in 23/30 Tier B
answers, so any baseline recorded at 512 is suspect in the same
direction -- and the base model is penalised hardest because it writes
longer answers.

Two sites, because they are independent defaults:

- approved-models.json: ghidra qualification_request and runtime_request.
  CI's 'Ghidra output cap matches the approved runtime request' pins the
  worker constant to this value, so leaving it would have kept the live
  worker truncating while the manifest claimed otherwise.
- ghidra-worker.py: TRIAGE_OUTPUT_TOKENS default. This is the production
  triage path, not only the benchmark.

revdeck.qualification_request also moves to 4096; it is the benchmark's
own measurement slot, and largest measured prompt is 2008 tokens so
2008 + 4096 = 6104 < its 8192 context.

sessions stays at 512. That slot is llm-worker session classification, not
reverse engineering, and llm-worker clamps LLM_OUTPUT_TOKENS to
max=2048 with an 8192 context -- the manifest value must stay
representable there. Its 512 was not a truncation defect.

Refs #1804, #1947, #2694
@Xore
Xore force-pushed the fix/benchmark-output-tokens branch from 21b07e4 to 17ae9f9 Compare September 25, 2026 19:58
@Xore
Xore merged commit 3dca445 into main Sep 25, 2026
110 checks passed
@Xore
Xore deleted the fix/benchmark-output-tokens branch September 25, 2026 20:14
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant