Skip to content

fix(evaluation): prevent synthetic diagnostics from authorizing policy promotion - #1119

Closed
seonghobae wants to merge 6 commits into
fix/no-heuristic-batch-routingfrom
fix/effort-promotion-authority-20260910
Closed

seonghobae wants to merge 6 commits into
fix/no-heuristic-batch-routingfrom
fix/effort-promotion-authority-20260910

Conversation

@seonghobae

@seonghobae seonghobae commented Sep 10, 2026 •

Copy link
Copy Markdown
Contributor

Current lineage and successor — 2026-09-12

This original child remains open / Draft at 4caac373ae861f5892acb2075cd97f38b097573f, on fix/effort-promotion-authority-20260910, still stacked on #1000 (fix/no-heuristic-batch-routing at 37cf3f6bb0810ff3bcf0132ac08f093523809667).

Main-based integration is now #1136, head b4343b2241ca25cb6b39dc68f98065f3e4103395 from protected main@012beaacd0631f8cd3391c77744eeb626269b5de. It carries every unique delta in this PR's nine changed paths: seven test/doctoring blobs identical, source and AGENTS combined with current-main-only changes. docs/doctoring/effort_main_integration_20260912.md contains the complete path/hash ledger. No parent source, obsolete workflow or dependency delta was imported.

This is an integration repair for the old stack's absent CI, not a concurrent replacement of #1000's numerical/routing work. Keep this predecessor and #1000 open until the successor's carryover, checks/review and protected delivery are verified. Neither was closed, force-pushed or rebased.

Retained public-helper contracts

  1. Caller-editable diagnostic RMSE/labels cannot authorize learned-policy promotion. production_default_change_allowed remains non-authorizing and does not execute supplied mapping methods. It does not replace the former arbitrary threshold with another threshold.
  2. An explicit unsupported-provider omit fallback removes a previously populated reasoning_effort field. Supported effort and independently valid sampling/output controls remain intact; abstain/error refuse before mutation.
  3. 2026-09-12 repair: Python truthiness cannot establish capability support. Strings such as "false", integers, containers and custom truth hooks previously crossed this boundary. Only literal boolean True is positive support; False/None retain explicit abstain/error/omit behavior. Other types fail before payload mutation without truth/equality/render hooks.

The synthetic theta/shrinkage/token arithmetic and hand-authored profile defaults remain unresolved owner migration work, not empirical estimates. No improvement in Fugu/TRINITY/Conductor performance is claimed by these validation tests. Actual latent-quality estimation belongs to released Rust/fast-mlsirm contracts; CI/release approval is not psychometric evidence.

Child RED → repair commits and source identities

  • 9967a0ce1df905c203b0592bb0a39f85c47efd8c: executable capability-evidence regressions.
  • 062173a8e3fa5f375c141e0c4efafcd7f559d5dd: strict capability type/identity checks and existing fallback preservation.
  • 98b7393bd13e83771a4b7eda3cc5d484fe572882: RCA, standards, alternatives, scenarios and verification limits.
  • 4caac373ae861f5892acb2075cd97f38b097573f: correct the issue-568 AGENTS guidance; preserve unrelated text.

Original child source blob: f876a176492d3ab51ce860bb5060a242c2f14613.
Repaired child source blob: 8a98fd7a8be9e61595645c46a10500559deb3ac9.
Capability tests: ca40abf2f13747e5cdd543a480f6de6ef5c9424b.
Omission tests: f3f99a5940764eac6d64df27cfc91e86701de6fe.
Promotion tests: c815feae94153e9b32637893a999b3df2b6a821d.

Local verification scopes — not hosted GREEN

The complete source was reconstructed from connector reads and matched to its Git blob before testing. No provider call, credential or paid service was used.

Child source:

  • New capability cases plus unchanged omission suite against the unmodified child: 48 failed / 23 passed, intended evidence/coercion/mutation defects.
  • Same cases after repair: 71 passed.
  • New plus unchanged omission and promotion suites: 83 passed, warnings treated as errors.
  • Changed helper: 19/19 executable statements and 12/12 branch arcs; whole module 75%, not 100%.

Main-based successor #1136 has a separate fresh run, not transferred evidence. It retains main's default_max_output_tokens: int | None and no-profile omission when the ceiling is unknown. Added six-case preservation suite: original main 67 failed / 22 passed across all four standalone modules; integrated source 89 passed. Changed helper 20/20 statements and 14/14 branch arcs, module 75%. No full package-import or whole-repository suite was run in the local partial checkout.

CI repair and remaining delivery gates

Exact original child 4caac373... returned zero check runs and zero workflow runs, not queued/passing evidence. Its inherited ci.yml admits only PRs targeting main. Current main has replaced separate CI/Fuzz workflows with security.yml and new required job identities; its jobs exclude Draft PRs. #1136 uses that unmodified current-main workflow and Ready solely for current-head review/CI admission. It does not revive obsolete CI or fabricate passing statuses.

Full current-head repository tests/package quality/fuzz/security, organization-required checks, valid findings, qualifying independent review, ordinary protected merge, immutable release and consumer adoption remain open gates. Parent #1000 migration and current full-tree baseline reconciliation are not claimed complete.

Research and reconstruction records

  • docs/doctoring/learned_policy_authority_20260910.md: Fugu v2/TRINITY v3/Conductor v5/RLM v3 distinctions; learned selection and native effort are separate. Fugu-Ultra's five-step training setting is not a universal request/partition limit; preserve within-workflow history isolation and cross-workflow memory.
  • docs/doctoring/effort_omission_contract_20260910.md: retained omission fault and predecessor verification.
  • docs/doctoring/effort_capability_evidence_20260912.md: typed-evidence fault, RFC 8259 / JSON Schema 2020-12 basis, failure scenes and exact blobs.
  • Successor only: docs/doctoring/effort_main_integration_20260912.md, complete carryover and main-specific output behavior.

CO#1117 owns whole-request partitioning and central .github#2068 reviewer-local state. This child does not claim their delivery, complete numerical Rust migration, a new empirical evaluation pipeline or a consumer release.

Child of #1000 preserving its complete delta. Reproduced 11 failing cases
where mutable labels, invalid RMSE or synthetic reports authorized defaults.
The diagnostic compatibility API is now non-authorizing. Twelve focused
regressions pass; the existing self-declared-approval assertion is repaired.
Research conformance records the real Fugu-Ultra five-step training scope.
Full repository gates and empirical learned-policy delivery remain open.
@coderabbitai

coderabbitai Bot commented Sep 10, 2026 •

Copy link
Copy Markdown

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Copy link
Copy Markdown
Contributor Author

@coderabbitai review

Please independently review the exact child delta at bfd39a6b8cef2e70604dce82b2b9f26449aac99a against parent #1000 37cf3f6bb0810ff3bcf0132ac08f093523809667, not the cumulative historical parent diff. Focus on the diagnostic-versus-authorization boundary, retained compatibility, caller-controlled mapping side effects and the corrected interpretation of Fugu-Ultra's training setup. Local evidence is 12 focused regression cases only; full hosted gates and empirical policy validity are not claimed. This request is not a self-approval or a production promotion.

@coderabbitai

coderabbitai Bot commented Sep 10, 2026 •

Copy link
Copy Markdown
🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
set -euo pipefail

child='bfd39a6b8cef2e70604dce82b2b9f26449aac99a'
parent='37cf3f6bb0810ff3bcf0132ac08f093523809667'

printf '%s\n' '== Commit relationship =='
git merge-base --is-ancestor "$parent" "$child" && echo 'parent is an ancestor of child'
git show -s --format='child: %H%nparents: %P%nsubject: %s' "$child"
git show -s --format='parent: %H%nsubject: %s' "$parent"

printf '%s\n' '== Exact child changed files and statistics =='
git diff --stat "$parent" "$child"
git diff --name-status "$parent" "$child"

printf '%s\n' '== Exact child patch =='
git diff --find-renames --find-copies --no-ext-diff --unified=80 "$parent" "$child"

printf '%s\n' '== Changed Python module and test outlines =='
git diff --name-only "$parent" "$child" | while IFS= read -r path; do
  case "$path" in
    *.py)
      echo "-- $path"
      ast-grep outline "$path" --items all --type function,class --view expanded || true
      ;;
  esac
done

printf '%s\n' '== Fugu-related text in exact child delta =='
git diff --find-renames --find-copies --no-ext-diff --unified=12 "$parent" "$child" \
  | rg -n -i -C 8 'fugu|ultra|workflow|five|5 step|training|memory|isolation' || true

Length of output: 65509


🏁 Script executed:

#!/bin/bash
set -euo pipefail

printf '%s\n' '== Function and threshold usages =='
rg -n -C 5 --glob '*.py' \
  '\bproduction_default_change_allowed\s*\(|\bPRODUCTION_RMSE_IMPROVEMENT_THRESHOLD\b' .

printf '%s\n' '== Relevant existing test ranges =='
sed -n '220,365p' tests/test_reasoning_effort_profile.py
sed -n '430,490p' tests/test_reasoning_effort_profile.py

printf '%s\n' '== Changed lines with current line numbers =='
git diff --unified=0 37cf3f6bb0810ff3bcf0132ac08f093523809667 bfd39a6b8cef2e70604dce82b2b9f26449aac99a \
  | awk '
      /^\+\+\+ b\// { file=substr($0,7); next }
      /^@@/ { print file " " $0; next }
      /^\+/ && !/^\+\+\+/ { print file ": " $0 }
    '

printf '%s\n' '== Repository working tree state =='
git status --short

Length of output: 46109


🤖 web_search executed:

Open the primary source https://arxiv.org/html/2606.21228v2 and verify the statements in Fugu-Ultra Sections 3.2.1 through 3.2.3 about (1) learned workflows and access lists, (2) agent-specific tool histories inside a workflow and memory across workflows, and (3) workflows of up to five steps and unrestricted environment interactions for multi-turn agents. Return concise quotations or section-grounded summaries with URLs.

Length of output: 20026


@seonghobae Reviewed the exact child delta bfd39a6b8cef2e70604dce82b2b9f26449aac99a against direct parent 37cf3f6bb0810ff3bcf0132ac08f093523809667.

I found no blocking issue in the child delta.

  • production_default_change_allowed retains its boolean API and always returns False.
  • The gate does not read report. Caller-controlled mapping methods cannot run at the authorization boundary.
  • PRODUCTION_RMSE_IMPROVEMENT_THRESHOLD remains importable. It has no authorization role.
  • The added tests cover mutable labels, numeric coercion, modified synthetic reports, and hostile mapping access.
  • The Fugu-Ultra documentation is consistent with the cited primary source. Up-to-five-step workflows are a training setting. They are not a universal root-call, retry, or context-partition limit. The document also preserves the distinction between intra-workflow isolation and cross-workflow shared memory.

I did not treat pre-existing parent-stack conditions as child-delta findings. I did not run repository tests in this review environment. The stated 12 focused regression cases remain local evidence, not independently reproduced evidence.

⚠️ Action not completed

Review rate limited.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

Copy link
Copy Markdown
Contributor Author

Bounded continuation on exact bfd39a6b8cef2e70604dce82b2b9f26449aac99a: repair the public effort adapter's existing unsupported+omit contract without changing the retired promotion gate or parent #1000.

Complete source reconstructed to blob 02e25e87e4ceb5e00584ac3445c59f5e602b3fa9. apply_request_profile leaves an already-present reasoning_effort field in the payload when capability support becomes false and the explicit fallback is omit. Five new cases fail on the published code; nine preservation cases pass. Minimal candidate removes only that stale field on the already-authorized omission branch. Supported native effort, configured sampling/output controls, tools/messages/stream fields and abstain/error fail-closed behavior stay intact.

Fresh candidate verification: 14 new cases plus the unchanged 12 promotion-authority cases = 26 passed with warnings treated as errors. The changed function has 18/18 executable lines and 10/10 branch arcs covered; whole-module coverage is 75%, not 100%. This is a public-helper contract repair, not proof of the production provider-failure cause or of a learned Fugu implementation. Publishing as a normal descendant of this existing child PR; no force push or model/routing/permission change.

The exported helper must remove existing reasoning_effort when support is
false and the validated fallback is omit. Keep abstain/error fail-closed,
supported native controls and the diagnostic non-authorization boundary.

RED: 5 failed / 9 passed. GREEN: 14 new + 12 unchanged authority tests pass.
Changed function covered 18/18 lines, 10/10 branch arcs; whole module 75%.
No live provider, full repository, learned-policy or deployment claim.

Copy link
Copy Markdown
Contributor Author

@coderabbitai review

Please review current child head ea0166818ea2ad7c00874b284e613199c1722ad7 against its #1000 parent, including the new omission repair. The first commit's diagnostic non-authorization boundary remains unchanged; this follow-up removes a stale native effort field only when an explicitly configured omit fallback permits it.

Fresh combined local tests: 26 pass with warnings as errors; the changed function covers 18/18 executable lines and 10/10 branch arcs. Whole-module coverage is 75%, not full coverage. Source/test blobs match f876a176492d3ab51ce860bb5060a242c2f14613 and f3f99a5940764eac6d64df27cfc91e86701de6fe. No PR-triggered workflow runs were returned for the current head when inspected; that is missing hosted evidence, not success.

Bounded effort-adapter source writer RELEASED. Full parent integration, empirical learned-policy validation and actual deployed request behavior remain unverified. Do not claim that this public-helper repair alone implements Fugu or resolves large-PR review saturation.

@coderabbitai

coderabbitai Bot commented Sep 10, 2026 •

Copy link
Copy Markdown
🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
set -euo pipefail

parent='37cf3f6bb0810ff3bcf0132ac08f093523809667'
child='ea0166818ea2ad7c00874b284e613199c1722ad7'

echo '== Commit relationship =='
git merge-base --is-ancestor "$parent" "$child" && echo 'parent is an ancestor of child'
git show -s --format='child: %H%nparents: %P%nsubject: %s' "$child"
git show -s --format='parent: %H%nsubject: %s' "$parent"

echo '== Exact changed files =='
git diff --name-status "$parent" "$child"
git diff --stat "$parent" "$child"

echo '== Exact patch =='
git diff --find-renames --find-copies --no-ext-diff --unified=50 "$parent" "$child"

echo '== Python symbol outlines for changed files =='
git diff --name-only "$parent" "$child" | while IFS= read -r path; do
  case "$path" in
    *.py)
      echo "-- $path"
      ast-grep outline "$path" --items all --type function,class --view expanded || true
      ;;
  esac
done

echo '== Definitions and call sites =='
rg -n -C 8 --glob '*.py' \
  '\bapply_request_profile\s*\(|\bproduction_default_change_allowed\s*\(|\bPRODUCTION_RMSE_IMPROVEMENT_THRESHOLD\b' .

echo '== Changed lines with current locations =='
git diff --unified=0 "$parent" "$child" | awk '
  /^\+\+\+ b\// { file=substr($0,7); next }
  /^@@/ { print file " " $0; next }
  /^\+/ && !/^\+\+\+/ { print file ": " $0 }
'

Length of output: 130031


🏁 Script executed:

#!/bin/bash
set -euo pipefail

echo '== Request adapter integration =='
ast-grep outline contextual_orchestrator/orchestrator.py --match 'request|profile|effort' --items all --type function,method --view expanded || true
sed -n '1915,1975p' contextual_orchestrator/orchestrator.py

echo '== Existing integration tests around profile application =='
sed -n '355,470p' tests/test_reasoning_effort_profile.py

echo '== Documentation validation test =='
sed -n '440,485p' tests/test_reasoning_effort_profile.py

echo '== Child-only changes in integration surfaces =='
git diff --unified=8 37cf3f6bb0810ff3bcf0132ac08f093523809667 ea0166818ea2ad7c00874b284e613199c1722ad7 -- \
  contextual_orchestrator/orchestrator.py \
  tests/test_no_heuristic_reasoning_effort_contract.py \
  scripts/ci/repair_pr1000_reasoning_effort.py

echo '== Exact child patch for omission function only =='
git diff --unified=20 37cf3f6bb0810ff3bcf0132ac08f093523809667 ea0166818ea2ad7c00874b284e613199c1722ad7 -- \
  contextual_orchestrator/reasoning_effort_profile.py

Length of output: 28603


🤖 web_search executed:

Open the primary source at https://arxiv.org/html/2606.21228v2. Verify these claims from Fugu-Ultra Sections 3.2.1–3.2.3: (1) it learns workflows and access lists; (2) agents have separate tool histories within a workflow and memory is shared across workflows; and (3) training workflows have up to five steps while multi-turn agents can interact with the environment without a fixed interaction limit. Return concise section-grounded evidence and the source URL.

Length of output: 19253


@seonghobae Reviewed child ea0166818ea2ad7c00874b284e613199c1722ad7 against parent #1000 at 37cf3f6bb0810ff3bcf0132ac08f093523809667.

I found no blocking issue in the reviewed child delta.

  • apply_request_profile removes stale reasoning_effort only after it validates the profile and only for unsupported capability with unsupported_provider_fallback == "omit".
  • Unsupported abstain and error profiles still raise before the function mutates the payload.
  • Supported profiles still set native reasoning_effort.
  • The helper retains the existing behavior when profile is None.
  • TaskOrchestrator.apply_effort_profile converts reasoning_effort to the Responses API reasoning field only when the helper retained the field. The omission branch therefore cannot forward stale effort through that adapter.
  • production_default_change_allowed preserves its boolean interface, does not read caller-controlled report data, and always returns False.
  • PRODUCTION_RMSE_IMPROVEMENT_THRESHOLD remains import-compatible and has no authorization effect.
  • The new tests cover pre-existing effort values, capability downgrade on a reused payload, supported forwarding, explicit omission idempotence, and non-mutation on unsupported non-omit fallbacks.
  • The Fugu-Ultra statements in docs/doctoring/learned_policy_authority_20260910.md match the cited primary source. The paper describes learned workflows with worker assignments and access lists, intra-workflow agent isolation for function-call trajectories, shared memory across workflows, and a training instruction of up to five workflow steps. It does not establish a universal root-request or context-partition cap. Source: https://arxiv.org/html/2606.21228v2

I did not execute repository tests because this review environment prohibits execution of repository code. The stated 26 focused passing cases remain local evidence. Hosted CI, full integration coverage, and empirical policy validation remain unverified.

⚠️ Action not completed

Review rate limited.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

Locally reproduced against unchanged source blob f876a17: 48 failed / 23 passed across new capability cases and unchanged omission suite. Covers strings, numbers, containers and hostile truth/equality/render hooks. No provider request or credentials used.
…ontrol

Only literal True establishes reasoning-effort support. False/None retain the explicit abstain/error/omit contract. Malformed evidence is rejected without mutation or truth/equality/render hooks. Focused new plus unchanged omission/promotion tests: 83 passed; changed function 19/19 statements and 12/12 branch arcs. Full repository/hosted validation remains required.
Align the issue-568 guidance with the existing non-authorizing promotion API and typed-capability repair. Preserve all unrelated governance text. This does not claim that parent numerical migration or an empirical policy release is complete.
seonghobae added a commit that referenced this pull request Sep 12, 2026
Successor preparation for #1119 without inheriting or retiring #1000. Retain the three existing regression blobs and the two-line legacy-test correction; add protected-main nullable-output preservation cases. Complete main source 7f8fedd yields 67 failed / 22 passed on the four standalone modules. This is local scoped RED, not hosted or whole-repository evidence.
seonghobae added a commit that referenced this pull request Sep 12, 2026
…output omission

Carry every unique #1119 source/test/doctoring/AGENTS delta onto protected main 012beaa without #1000's broad unresolved tree. Preserve main's int-or-None ceiling and newer AGENTS guidance. Scoped complete-leaf RED: 67 failed/22 passed; GREEN: 89 passed, changed helper 20/20 statements and 14/14 branches; whole module 75%. Full hosted integration remains required. Neither predecessor is closed.
@seonghobae seonghobae added bug Something isn't working priority: high labels Sep 12, 2026 — with ChatGPT Codex Connector
@seonghobae

Copy link
Copy Markdown
Contributor Author

Superseded by protected main

Closing without restack/force-push: a rebase of this child's six unique commits onto current origin/main (21c0b32e) immediately drops bfd39a6b as already upstream, then conflicts in reasoning_effort_profile.py against the merged main-based successor #1136 (merged 2026-09-17; merge commit is an ancestor of main). Faithful conflict resolution would retain main's stricter integration (default_max_output_tokens: int | None, nested reasoning.effort omission, ModelAgent identity checks) and yield an empty unique tree — so force-pushing head==base is refused per undraft policy.

Evidence

No admin-merge. No gate weakening. This predecessor Draft is closed as superseded-by-main.

@seonghobae

Copy link
Copy Markdown
Contributor Author

Superseded by merged #1136 on protected main; restack would yield empty unique tree (see supersession comment). No force-push of head==base.

@seonghobae seonghobae closed this Sep 17, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working priority: high

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant