fix(evaluation): prevent synthetic diagnostics from authorizing policy promotion - #1119
seonghobae wants to merge 6 commits into
Conversation
Child of #1000 preserving its complete delta. Reproduced 11 failing cases where mutable labels, invalid RMSE or synthetic reports authorized defaults. The diagnostic compatibility API is now non-authorizing. Twelve focused regressions pass; the existing self-declared-approval assertion is repaired. Research conformance records the real Fugu-Ultra five-step training scope. Full repository gates and empirical learned-policy delivery remain open.
|
Important Draft PR not reviewedDraft PRs are not automatically reviewed by default.
To automatically review draft PRs, update your CodeRabbit configuration: reviews:
auto_review:
drafts: trueThanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
@coderabbitai review Please independently review the exact child delta at |
🧩 Analysis chain🏁 Script executed: #!/bin/bash
set -euo pipefail
child='bfd39a6b8cef2e70604dce82b2b9f26449aac99a'
parent='37cf3f6bb0810ff3bcf0132ac08f093523809667'
printf '%s\n' '== Commit relationship =='
git merge-base --is-ancestor "$parent" "$child" && echo 'parent is an ancestor of child'
git show -s --format='child: %H%nparents: %P%nsubject: %s' "$child"
git show -s --format='parent: %H%nsubject: %s' "$parent"
printf '%s\n' '== Exact child changed files and statistics =='
git diff --stat "$parent" "$child"
git diff --name-status "$parent" "$child"
printf '%s\n' '== Exact child patch =='
git diff --find-renames --find-copies --no-ext-diff --unified=80 "$parent" "$child"
printf '%s\n' '== Changed Python module and test outlines =='
git diff --name-only "$parent" "$child" | while IFS= read -r path; do
case "$path" in
*.py)
echo "-- $path"
ast-grep outline "$path" --items all --type function,class --view expanded || true
;;
esac
done
printf '%s\n' '== Fugu-related text in exact child delta =='
git diff --find-renames --find-copies --no-ext-diff --unified=12 "$parent" "$child" \
| rg -n -i -C 8 'fugu|ultra|workflow|five|5 step|training|memory|isolation' || trueLength of output: 65509 🏁 Script executed: #!/bin/bash
set -euo pipefail
printf '%s\n' '== Function and threshold usages =='
rg -n -C 5 --glob '*.py' \
'\bproduction_default_change_allowed\s*\(|\bPRODUCTION_RMSE_IMPROVEMENT_THRESHOLD\b' .
printf '%s\n' '== Relevant existing test ranges =='
sed -n '220,365p' tests/test_reasoning_effort_profile.py
sed -n '430,490p' tests/test_reasoning_effort_profile.py
printf '%s\n' '== Changed lines with current line numbers =='
git diff --unified=0 37cf3f6bb0810ff3bcf0132ac08f093523809667 bfd39a6b8cef2e70604dce82b2b9f26449aac99a \
| awk '
/^\+\+\+ b\// { file=substr($0,7); next }
/^@@/ { print file " " $0; next }
/^\+/ && !/^\+\+\+/ { print file ": " $0 }
'
printf '%s\n' '== Repository working tree state =='
git status --shortLength of output: 46109 🤖 web_search executed:
Length of output: 20026
I found no blocking issue in the child delta.
I did not treat pre-existing parent-stack conditions as child-delta findings. I did not run repository tests in this review environment. The stated 12 focused regression cases remain local evidence, not independently reproduced evidence.
|
|
Bounded continuation on exact Complete source reconstructed to blob Fresh candidate verification: 14 new cases plus the unchanged 12 promotion-authority cases = 26 passed with warnings treated as errors. The changed function has 18/18 executable lines and 10/10 branch arcs covered; whole-module coverage is 75%, not 100%. This is a public-helper contract repair, not proof of the production provider-failure cause or of a learned Fugu implementation. Publishing as a normal descendant of this existing child PR; no force push or model/routing/permission change. |
The exported helper must remove existing reasoning_effort when support is false and the validated fallback is omit. Keep abstain/error fail-closed, supported native controls and the diagnostic non-authorization boundary. RED: 5 failed / 9 passed. GREEN: 14 new + 12 unchanged authority tests pass. Changed function covered 18/18 lines, 10/10 branch arcs; whole module 75%. No live provider, full repository, learned-policy or deployment claim.
|
@coderabbitai review Please review current child head Fresh combined local tests: 26 pass with warnings as errors; the changed function covers 18/18 executable lines and 10/10 branch arcs. Whole-module coverage is 75%, not full coverage. Source/test blobs match Bounded effort-adapter source writer RELEASED. Full parent integration, empirical learned-policy validation and actual deployed request behavior remain unverified. Do not claim that this public-helper repair alone implements Fugu or resolves large-PR review saturation. |
🧩 Analysis chain🏁 Script executed: #!/bin/bash
set -euo pipefail
parent='37cf3f6bb0810ff3bcf0132ac08f093523809667'
child='ea0166818ea2ad7c00874b284e613199c1722ad7'
echo '== Commit relationship =='
git merge-base --is-ancestor "$parent" "$child" && echo 'parent is an ancestor of child'
git show -s --format='child: %H%nparents: %P%nsubject: %s' "$child"
git show -s --format='parent: %H%nsubject: %s' "$parent"
echo '== Exact changed files =='
git diff --name-status "$parent" "$child"
git diff --stat "$parent" "$child"
echo '== Exact patch =='
git diff --find-renames --find-copies --no-ext-diff --unified=50 "$parent" "$child"
echo '== Python symbol outlines for changed files =='
git diff --name-only "$parent" "$child" | while IFS= read -r path; do
case "$path" in
*.py)
echo "-- $path"
ast-grep outline "$path" --items all --type function,class --view expanded || true
;;
esac
done
echo '== Definitions and call sites =='
rg -n -C 8 --glob '*.py' \
'\bapply_request_profile\s*\(|\bproduction_default_change_allowed\s*\(|\bPRODUCTION_RMSE_IMPROVEMENT_THRESHOLD\b' .
echo '== Changed lines with current locations =='
git diff --unified=0 "$parent" "$child" | awk '
/^\+\+\+ b\// { file=substr($0,7); next }
/^@@/ { print file " " $0; next }
/^\+/ && !/^\+\+\+/ { print file ": " $0 }
'Length of output: 130031 🏁 Script executed: #!/bin/bash
set -euo pipefail
echo '== Request adapter integration =='
ast-grep outline contextual_orchestrator/orchestrator.py --match 'request|profile|effort' --items all --type function,method --view expanded || true
sed -n '1915,1975p' contextual_orchestrator/orchestrator.py
echo '== Existing integration tests around profile application =='
sed -n '355,470p' tests/test_reasoning_effort_profile.py
echo '== Documentation validation test =='
sed -n '440,485p' tests/test_reasoning_effort_profile.py
echo '== Child-only changes in integration surfaces =='
git diff --unified=8 37cf3f6bb0810ff3bcf0132ac08f093523809667 ea0166818ea2ad7c00874b284e613199c1722ad7 -- \
contextual_orchestrator/orchestrator.py \
tests/test_no_heuristic_reasoning_effort_contract.py \
scripts/ci/repair_pr1000_reasoning_effort.py
echo '== Exact child patch for omission function only =='
git diff --unified=20 37cf3f6bb0810ff3bcf0132ac08f093523809667 ea0166818ea2ad7c00874b284e613199c1722ad7 -- \
contextual_orchestrator/reasoning_effort_profile.pyLength of output: 28603 🤖 web_search executed:
Length of output: 19253
I found no blocking issue in the reviewed child delta.
I did not execute repository tests because this review environment prohibits execution of repository code. The stated 26 focused passing cases remain local evidence. Hosted CI, full integration coverage, and empirical policy validation remain unverified.
|
Locally reproduced against unchanged source blob f876a17: 48 failed / 23 passed across new capability cases and unchanged omission suite. Covers strings, numbers, containers and hostile truth/equality/render hooks. No provider request or credentials used.
…ontrol Only literal True establishes reasoning-effort support. False/None retain the explicit abstain/error/omit contract. Malformed evidence is rejected without mutation or truth/equality/render hooks. Focused new plus unchanged omission/promotion tests: 83 passed; changed function 19/19 statements and 12/12 branch arcs. Full repository/hosted validation remains required.
Align the issue-568 guidance with the existing non-authorizing promotion API and typed-capability repair. Preserve all unrelated governance text. This does not claim that parent numerical migration or an empirical policy release is complete.
Successor preparation for #1119 without inheriting or retiring #1000. Retain the three existing regression blobs and the two-line legacy-test correction; add protected-main nullable-output preservation cases. Complete main source 7f8fedd yields 67 failed / 22 passed on the four standalone modules. This is local scoped RED, not hosted or whole-repository evidence.
…output omission Carry every unique #1119 source/test/doctoring/AGENTS delta onto protected main 012beaa without #1000's broad unresolved tree. Preserve main's int-or-None ceiling and newer AGENTS guidance. Scoped complete-leaf RED: 67 failed/22 passed; GREEN: 89 passed, changed helper 20/20 statements and 14/14 branches; whole module 75%. Full hosted integration remains required. Neither predecessor is closed.
Superseded by protected mainClosing without restack/force-push: a rebase of this child's six unique commits onto current Evidence
No admin-merge. No gate weakening. This predecessor Draft is closed as superseded-by-main. |
|
Superseded by merged #1136 on protected main; restack would yield empty unique tree (see supersession comment). No force-push of head==base. |
Current lineage and successor — 2026-09-12
This original child remains open / Draft at
4caac373ae861f5892acb2075cd97f38b097573f, onfix/effort-promotion-authority-20260910, still stacked on #1000 (fix/no-heuristic-batch-routingat37cf3f6bb0810ff3bcf0132ac08f093523809667).Main-based integration is now #1136, head
b4343b2241ca25cb6b39dc68f98065f3e4103395from protectedmain@012beaacd0631f8cd3391c77744eeb626269b5de. It carries every unique delta in this PR's nine changed paths: seven test/doctoring blobs identical, source and AGENTS combined with current-main-only changes.docs/doctoring/effort_main_integration_20260912.mdcontains the complete path/hash ledger. No parent source, obsolete workflow or dependency delta was imported.This is an integration repair for the old stack's absent CI, not a concurrent replacement of #1000's numerical/routing work. Keep this predecessor and #1000 open until the successor's carryover, checks/review and protected delivery are verified. Neither was closed, force-pushed or rebased.
Retained public-helper contracts
production_default_change_allowedremains non-authorizing and does not execute supplied mapping methods. It does not replace the former arbitrary threshold with another threshold.omitfallback removes a previously populatedreasoning_effortfield. Supported effort and independently valid sampling/output controls remain intact; abstain/error refuse before mutation."false", integers, containers and custom truth hooks previously crossed this boundary. Only literal booleanTrueis positive support;False/Noneretain explicit abstain/error/omit behavior. Other types fail before payload mutation without truth/equality/render hooks.The synthetic theta/shrinkage/token arithmetic and hand-authored profile defaults remain unresolved owner migration work, not empirical estimates. No improvement in Fugu/TRINITY/Conductor performance is claimed by these validation tests. Actual latent-quality estimation belongs to released Rust/fast-mlsirm contracts; CI/release approval is not psychometric evidence.
Child RED → repair commits and source identities
9967a0ce1df905c203b0592bb0a39f85c47efd8c: executable capability-evidence regressions.062173a8e3fa5f375c141e0c4efafcd7f559d5dd: strict capability type/identity checks and existing fallback preservation.98b7393bd13e83771a4b7eda3cc5d484fe572882: RCA, standards, alternatives, scenarios and verification limits.4caac373ae861f5892acb2075cd97f38b097573f: correct the issue-568 AGENTS guidance; preserve unrelated text.Original child source blob:
f876a176492d3ab51ce860bb5060a242c2f14613.Repaired child source blob:
8a98fd7a8be9e61595645c46a10500559deb3ac9.Capability tests:
ca40abf2f13747e5cdd543a480f6de6ef5c9424b.Omission tests:
f3f99a5940764eac6d64df27cfc91e86701de6fe.Promotion tests:
c815feae94153e9b32637893a999b3df2b6a821d.Local verification scopes — not hosted GREEN
The complete source was reconstructed from connector reads and matched to its Git blob before testing. No provider call, credential or paid service was used.
Child source:
Main-based successor #1136 has a separate fresh run, not transferred evidence. It retains main's
default_max_output_tokens: int | Noneand no-profile omission when the ceiling is unknown. Added six-case preservation suite: original main 67 failed / 22 passed across all four standalone modules; integrated source 89 passed. Changed helper 20/20 statements and 14/14 branch arcs, module 75%. No full package-import or whole-repository suite was run in the local partial checkout.CI repair and remaining delivery gates
Exact original child
4caac373...returned zero check runs and zero workflow runs, not queued/passing evidence. Its inheritedci.ymladmits only PRs targetingmain. Current main has replaced separate CI/Fuzz workflows withsecurity.ymland new required job identities; its jobs exclude Draft PRs. #1136 uses that unmodified current-main workflow and Ready solely for current-head review/CI admission. It does not revive obsolete CI or fabricate passing statuses.Full current-head repository tests/package quality/fuzz/security, organization-required checks, valid findings, qualifying independent review, ordinary protected merge, immutable release and consumer adoption remain open gates. Parent #1000 migration and current full-tree baseline reconciliation are not claimed complete.
Research and reconstruction records
docs/doctoring/learned_policy_authority_20260910.md: Fugu v2/TRINITY v3/Conductor v5/RLM v3 distinctions; learned selection and native effort are separate. Fugu-Ultra's five-step training setting is not a universal request/partition limit; preserve within-workflow history isolation and cross-workflow memory.docs/doctoring/effort_omission_contract_20260910.md: retained omission fault and predecessor verification.docs/doctoring/effort_capability_evidence_20260912.md: typed-evidence fault, RFC 8259 / JSON Schema 2020-12 basis, failure scenes and exact blobs.docs/doctoring/effort_main_integration_20260912.md, complete carryover and main-specific output behavior.CO#1117 owns whole-request partitioning and central .github#2068 reviewer-local state. This child does not claim their delivery, complete numerical Rust migration, a new empirical evaluation pipeline or a consumer release.