fix(ghidra): raise pinned benchmark output_tokens 512 -> 4096 (#1804 follow-up) - #3308
Merged
Merged
Conversation
Dependency Review✅ No vulnerabilities or license issues or OpenSSF Scorecard issues found.Scanned FilesNone |
The 512 cap truncated answers rather than limiting them. Measured 2026-09-25 on the Tier A corpus: qwen2.5-coder:7b-instruct-q4_K_M was cut off on 14/14 cases at 512, still 1/17 at 2048 (done_reason: length on file_write_persist, 5127 chars), and 0/17 at 4096. #2694 had already recorded the same cap truncating injection conclusions in 23/30 Tier B answers, so any baseline recorded at 512 is suspect in the same direction -- and the base model is penalised hardest because it writes longer answers. Two sites, because they are independent defaults: - approved-models.json: ghidra qualification_request and runtime_request. CI's 'Ghidra output cap matches the approved runtime request' pins the worker constant to this value, so leaving it would have kept the live worker truncating while the manifest claimed otherwise. - ghidra-worker.py: TRIAGE_OUTPUT_TOKENS default. This is the production triage path, not only the benchmark. revdeck.qualification_request also moves to 4096; it is the benchmark's own measurement slot, and largest measured prompt is 2008 tokens so 2008 + 4096 = 6104 < its 8192 context. sessions stays at 512. That slot is llm-worker session classification, not reverse engineering, and llm-worker clamps LLM_OUTPUT_TOKENS to max=2048 with an 8192 context -- the manifest value must stay representable there. Its 512 was not a truncation defect. Refs #1804, #1947, #2694
Xore
force-pushed
the
fix/benchmark-output-tokens
branch
from
September 25, 2026 19:58
21b07e4 to
17ae9f9
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The 512-token cap was truncating Ghidra triage and benchmark answers, not limiting them.
Evidence
Measured on the Tier A corpus, 2026-09-25, with #159's own harness:
The case that survived 2048 was
file_write_persist,done_reason: lengthat 5127 chars (~1900 tokens), so 2048 is not enough either. 4096 is validated by measurement, not chosen for tidiness.This is not new. #2694 already recorded the same cap truncating injection conclusions in 23 of 30 Tier B answers. Any baseline recorded at 512 is suspect in the same direction — and the base model is penalised hardest, because it writes longer answers. That is what made the #1804 candidate-2 A/B look like a 6-point win for the fine-tune when it is 2.
The change, and why it is three files
ghidramoves to 4096 inapproved-models.json(qualification_requestandruntime_request), andghidra-worker.py'sTRIAGE_OUTPUT_TOKENSdefault moves with it. Those are independent defaults, and CI already pins them together — theGhidra output cap matches the approved runtime requestcheck failed on the manifest-only change, which is the invariant working. Had the manifest been changed alone, the live worker would have kept truncating while the manifest claimed otherwise.revdeck.qualification_requestalso moves to 4096: it is the benchmark's own measurement slot, and the largest measured prompt is 2008 tokens, so 2008 + 4096 = 6104 < its 8192 context.sessionsdeliberately stays at 512. That slot is llm-worker session classification, not reverse engineering, andllm-workerclampsLLM_OUTPUT_TOKENStomax=2048under an 8192 context. Its 512 was never a truncation defect, and raising it would have made the manifest value unrepresentable in the consumer that serves it.Tests: the one hardcoded
512expectation intest_ghidra_worker.pyis updated. Nothing else needed loosening — the governance and baseline tests carry their own local request dicts, and 76 pass.Re-measured at 4096
tierA_approved_qwen3_14b_cap4096_run1.jsonis committed:qwen3:14bscores 76/83 (91.6%) across 17 cases, 0 truncated, 0done_reason: length. That is the modelapproved-models.jsonnames for all three slots, so it is the baseline most likely to be quoted anywhere, and it has now been measured clean.Known and not fixed here
hf.co/mradermacher/GPT-OSS-Cybersecurity-20B-Merged-heretic-GGUF:Q4_K_Mreturns empty content on every request, at everyreasoning_effortandfrequency_penalty, including a bare request with no options.gpt-oss:20banswers the same prompt normally, so this is the checkpoint, not the runtime or the cap. It scores 0/83 and that number is meaningless. This is a model problem for One benchmark, every model, every aspect: rebuild the model matrix on current hardware with error bars #1947, not a cap problem, and is not addressed by this PR.Refs #1804, #1947, #2694