Skip to content

fix(nim): retain failed tasks in paired quality and latency evidence - #1074

Draft
seonghobae wants to merge 50 commits into
codex/psychometric-kpi-successorfrom
codex/paired-policy-outcomes-20260905
Draft

fix(nim): retain failed tasks in paired quality and latency evidence#1074
seonghobae wants to merge 50 commits into
codex/psychometric-kpi-successorfrom
codex/paired-policy-outcomes-20260905

Conversation

@seonghobae

@seonghobae seonghobae commented Sep 5, 2026

Copy link
Copy Markdown
Contributor

Current exact-head checkpoint — 9f5d0ce (2026-09-07)

  • Exact head: 9f5d0ce2f7e95eb3bfcaf5cbd4e2a165dab7c86c; base: ec1c4e66512615ea1f00fa044f3bfba787aab567; lifecycle: Draft / Proposed.
  • Hosted quality job 101628312201 failed after a green full suite (3618 passed, 2 skipped) because nim_benchmark.py coverage was 99% (2791, 2798, 2809, 2814, 2819, 2827 in validate_report_schema).
  • Source 9f5d0ce2 adds test_report_schema_rejects_invalid_evaluation_contract. Local three-file coverage command: 169 passed, 1271/480 statements/branches, 0 missed, interrogate 100%.
  • Fresh source, fuzz, and supply-chain Checks must run on this SHA. No predecessor coverage result is transferred. No production policy or release is authorized.

Current exact-head checkpoint — 4351af9 (2026-09-07)

  • Exact head: 4351af9f1b6dca2f97526a46ac6dca28048ff623; base: ec1c4e66512615ea1f00fa044f3bfba787aab567; lifecycle: Draft / Proposed.
  • This ordinary non-force integration preserves aa4ff156f3001927: tied observed worker maxima yield no hindsight winner while measurement-only artifacts remain available. Concurrent commit e94dd035 adds prospective observation-coverage validation; the merge retains both deltas and their regression fixtures.
  • Fresh source, fuzz, and supply-chain Checks are running on this exact SHA. No predecessor result is transferred.
  • Fixed 2,000-resample bootstrap, fixed 95% interval, hard-coded policy comparison subset, and repository-authored inference/token/workflow budgets remain substantive no-heuristics blockers. No production policy or release is authorized.

Current exact-head checkpoint — f300192 (2026-09-07)

  • Exact head: f3001927ad28852c5cca874eb3bf61920107e42f; base: ec1c4e66512615ea1f00fa044f3bfba787aab567; lifecycle: Draft / Proposed.
  • Regression-first commits f1662dcdd217c053, followed by integration repair aa4ff156f3001927, remove the lexicographic model-id tie-break from best_single_worker_hindsight. A unique observed maximum is required; equal maxima now produce no hindsight winner while the measurement-only report remains available.
  • The existing fixed 2,000-resample percentile interval, hard-coded policy comparison subset, and repository-authored inference/token/workflow budgets remain unresolved. This checkpoint does not transfer predecessor GREEN or claim statistical release readiness.
  • The predecessor exact-head test run failed because tied dry-run fixtures aborted the whole report; f3001927 repairs that propagation. Fresh exact-head hosted Checks have not materialized yet. Do not merge until source, fuzz, security, and independent model-review gates are terminal on this SHA.

Current verified checkpoint — 621cf4d (2026-09-07)

Current pushed head: 621cf4d9cd47d166bd568cfb07a25692c47df210, base ec1c4e66512615ea1f00fa044f3bfba787aab567, Draft. This checkpoint supersedes old current-head wording below without discarding historical evidence.

  • Count/completion-only promotion is removed, and the separate reasoning-effort 55% declaration-only authority is now removed too. Its retained Boolean function returns false for every report; the removed threshold export and formerly true outcomes are explicitly documented breaking changes. Profiles, provider propagation and route/conduct defaults are unchanged.
  • Exact-head local full regression: 3612 passed, 2 skipped / 1110.73s, exit 0, matching clean start/end HEAD and JUnit3614/0failures/0errors/2skips. Focused96 passed; profile module168statements/62branches and public docstrings100%. These are software checks, not buyer KPI evidence.
  • Wheel built; separate-cwd Python-I archive import and45Python-file source-byte parity verified. This is not an independent environment installation or release.
  • New hosted run34077430336 is pending. Skipped bot reviews are not independent approval. No deployment is established.
  • Real held-out task/policy evidence, calibrated scoring, failure denominators, dependence-aware uncertainty, a released Rust statistical-owner contract and independent operational acceptance remain incomplete. No synthetic fixture or arbitrary replacement cutoff substitutes for these.

See the exact verification receipt and prospective design. Actual Edge visual inspection confirms the rendered receipt, Draft, pending checks and no-deployment status; this is not product-admin/mobile E2E.

Historical endpoint collection repair — 93bd66d

Historical head: 93bd66d494f8f44a15b9989fe21b075412103da4, historical base a4f693c41f642f960d59bb5a9849eadedd412bd5. All results in this section retain that revision scope and do not describe the current head.

Whole-model-ID URL encoding returned HTTP 404 for a public model while the documented author/slug path returned 200. Source 98cdc3ec validates exactly two nonempty non-dot segments and encodes each separately using the existing stdlib. Malformed IDs cause no request. Fixed origin, no-update behavior, authentication, routing weights and statistical-owner boundaries remain unchanged.

  • Committed RED aee1e497: 16 failed / 61 passed / exit 1 / 1.28s.
  • Source 98cdc3ec: 147 related tests passed in 4.80s, exit 0. Separate 77-test coverage run: changed fetch method 18/18 statements and 6/6 branches; default Ruff passes. Not exhaustive input or whole-repository coverage.
  • Current child focused: 270 passed in 63.75s, exit 0. The initial parent command named a nonexistent test file and ran no tests (exit 4); it is excluded from passing evidence and was corrected.
  • One actual isolated current-source collector poll at 2026-09-06 12:05:50 UTC successfully added transport window mass for openai/gpt-4o, with observed attempts still zero. No model inference, background service, or supplied provider credential. This does not establish fleet access, calibrated delivery probability, deployed behavior or buyer accuracy/latency.
  • Exact-head full suite completed: 3597 passed / 2 skipped in 1762.13s, exit 0. Matching clean start/end 93bd66d494f8f44a15b9989fe21b075412103da4, parsed JUnit 3599 cases, zero failures/errors. Session 28306 is terminal, not running. Artifacts: /tmp/co-uptime-path.T7v9Rj/child-full-*. The verification JSON's seconds field measures the outer command (1782.32s); the pytest duration above is from the terminal test log. Neither duration is a buyer latency KPI.
  • Current local Trivy filesystem scans both exited 0 with zero HIGH/CRITICAL vulnerability/secret findings in four lock targets under ignore-unfixed and default dev/test exclusions. These are not hosted Security acceptance.
  • Prior input-admission fulls are terminal: parent db4d21da 3566 passed/two skipped/739.03s and child 5eebac47 3581 passed/two skipped/754.61s, both exit 0, matching clean heads and parsed JUnit. Those and prior Trivy results predate this path repair.
  • Visual Inspection: the exact parent a4f693c4 rendered doctoring section was inspected in a real Edge desktop screenshot; the complete new section, counts and limitations are legible without observed clipping or overlap. Current parent and child PR headers/checkpoint bodies were also inspected in real Edge screenshots: Draft status, main/parent target, matching head/base, focused counts and running-full wording are readable without observed overlap. This is not product-admin/mobile E2E.

The normal child merge retains parent source/tests/documentation and the child's NIM comparison, observed audit and licensed paper delta. Current collection evidence records alternatives, live scope, prior fulls and remaining gates. RED/GREEN logs, coverage and live collector output are in /tmp/co-uptime-path.T7v9Rj; original public HTTP baseline is retained in /tmp/co-uptime-input.rjxaB9.

Rolling-window dependence, endpoint-route mix, benchmark-prior calibration, judge validation and released-owner/buyer evidence remain open. ADR0034 stays Proposed; ADR0021 acceptance and production defaults are unchanged. Both PRs stay Draft. Independent exact-head review, required hosted checks, parent-first protected merge and immutable release remain gates.

Problem and result

The policy summary counted failures in its denominator, but paired confidence intervals excluded any task where either policy failed. A policy delivering one of two answers could therefore appear tied with a policy delivering both.

The comparison now retains every shared locked task, treats unsuccessful delivery as zero reward, and reports mean delivered-score and terminal-outcome-time differences with paired bootstrap intervals. Raw failed answer scores remain null. Success counts and unmatched task counts expose the comparison denominator; duplicates and invalid measurements fail closed.

In the hand-checked unit fixture, the old apparent tie becomes a delivered-score difference of -0.5. Time differences of -50 and 1950 ms produce mean 950 ms with interval [-50, 1950]. This proves a calculation repair, not improved model accuracy or live latency.

Contract and scope

  • Report schema 2.0.0 identifies the new estimand. Version 1 comparisons must be regenerated from their original observations rather than pooled with version 2.
  • The existing paired mean-bootstrap implementation is reused. No new statistical dependency or provider call is added.
  • The former count/completion promotion thresholds have been removed. Current output is measurement-only with null recommendation/thresholds; retained paired statistics do not authorize deployment.
  • Mean intervals are conditional on the shared task set and selected policies. They do not establish p95 improvement, dependent-task validity, or uncertainty from hindsight model selection.
  • Product and technical requirements, UML sequence, research reference, migration, and remaining owner prerequisites are recorded in the NIM guide, doctoring record, research inventory, and gap baseline.

Historical identity-batch integration at fa8eef9

Historical head: fa8eef97cd7368e8985a367dc5a7a8e0147fe4d5.
Parent #1067: d740602fdd8c0e4f7d55e4d3ad37b9f560c09e01.

This normal two-parent merge retains previous child 9207412e and the new parent identity-batch change. The baseline conflict preserves both complete additions and records the new lineage. Explicit diffs show child NIM source/tests, comparison guide, observed-data audit, and papers are unchanged from 9207412e; the parent's orchestration source, psychometric tests, doctoring, and complete profile artifact are byte-identical to d740602f.

  • 243 focused tests passed in 35.24 seconds, terminal exit 0 on clean fa8eef97: NIM comparisons, release/workflow contracts, profiles, paper/planning checks, and psychometric routing. Evidence: /tmp/co-1074-catalog-batch.MIcuwO/focused-pytest.log and focused-junit.xml.
  • Exact current fa8eef9 full suite completed: 3,481 passed, two skipped, terminal exit 0 in 664.63 seconds. Start/end SHA both equal fa8eef97cd7368e8985a367dc5a7a8e0147fe4d5, and the tracked tree stayed clean. JUnit contains 3,483 cases, zero failures, zero errors, and two skips. Evidence: /tmp/co-1074-catalog-batch.MIcuwO/full-pytest.log and full-junit.xml. Session 79929 is terminal; do not restart it. Parent d740602f completed 3,466 passed, two skipped, exit 0 in 733.63 seconds on its separate clean tree. Hosted acceptance remains separate.
  • Exact fa8eef97 Trivy 0.74.0 filesystem scan passed (exit 0), default vulnerability/secret scanners, HIGH/CRITICAL and --ignore-unfixed: zero findings in four lock targets. Evidence: /tmp/co-1074-catalog-batch.MIcuwO/trivy.json and trivy.log. DB UpdatedAt 2026-09-06 07:00:11 UTC, DownloadedAt 07:08:45 UTC, NextUpdate 2026-09-07 07:00:11 UTC; this run reused that fresh DB. Default dev/test exclusions apply, and complete hosted Security acceptance remains separate.
  • The parent eliminates repeated catalog validation per batch, preserves real repeated attempts, and observes later configuration changes without a persistent cache. Its controlled profile is internal unit-fixture computation only, not a buyer latency/accuracy result. Both favorable and earlier regressing samples remain in the artifact.

Draft status, parent-first protected delivery, and independent approval remain unchanged.

Previous integration at 9207412

Historical head: 9207412ea8ac529d7d2622ab298989ef7899befb.
Historical parent #1067: bfeb73a6c58add7a23df052110593cdf43c0b0db.

This normal two-parent merge preserves prior child 402fd63a and the updated parent. The one baseline-document conflict retains both complete additions and adds the current lineage note. There were no source conflicts. Explicit diffs show the child's NIM source changes only the inherited reviewed-cost dates; its test fixture gains the parent's wall-time isolation. The observed-data audit, paper inventory, and separately licensed LLMRouter paper are byte-identical to the preceding child. Parent transport, psychometric source, both harnesses, transport tests, and psychometric doctoring are byte-identical to bfeb73a6.

  • 211 focused tests passed in 2.33 seconds, terminal exit 0, on clean 9207412e: NIM comparisons, release acceptance, workflow contracts, reasoning profiles, paper inventory, and planning contracts. Log/JUnit: /tmp/co-1074-current-integration.PBuxSQ.
  • Exact historical 9207412 full suite completed: 3,476 passed, two skipped, terminal exit 0 in 790.27 seconds. Start/end heads match and the tracked tree stayed clean. JUnit has 3,478 cases, zero failures, zero errors, and two skips. Evidence: /tmp/co-1074-current-integration.PBuxSQ. Session 90398 is terminal and must not be restarted.
  • Same-head coverage run: 211 tests passed in 7.83 seconds, with NIM 1,223 statements / 446 branches and reasoning profiles 185 statements / 68 branches, all 100%. Coverage JSON/log are in the same evidence directory.
  • Same-head Trivy 0.74.0 default vulnerability/secret filesystem scan: exit 0, zero HIGH/CRITICAL findings under --ignore-unfixed in four reported lock targets. A fresh DB download completed at 2026-09-06 07:08:45 UTC (UpdatedAt 07:00:11 UTC, NextUpdate 2026-09-07 07:00:11 UTC). JSON, scan log, and download log are in the same evidence directory. Default development/test dependency exclusions apply; this does not replace the complete hosted Security workflow.
  • Parent bfeb73a6 completed 3,461 passed, two skipped, terminal exit 0 in 675.10 seconds, with matching start/end head and clean tracked tree. Its JUnit has 3,463 cases and no failures/errors. Session 72900 is terminal. This parent result is not a full-suite result for the current child.
  • Both PRs remain Draft. Parent-first protected delivery, exact-head hosted acceptance, and independent approval remain required. No paper removed by the rights audit was restored, and no production policy was promoted.

Historical verification at 402fd63

Historical head: 402fd63a4f6a16558e8cb4e4cd8a3378da3827f6.
Historical parent #1067: ae704491fc24cd4618c709b036cc9353ea9923e5.

Ordinary merge 4fc184a020f1d63816f9e847d98a306ff27797f8 preserves both preceding child e2f3e2af8e6b723c6870fb6c4dcee2b4f7bc31af and parent histories. Both changelog and baseline additions were retained. Follow-up 402fd63a corrects only a document-location word. The child's NIM source, tests, observed-data audit, and licensed LLMRouter paper retain their preceding bytes; parent profile guards, tests, benchmark harnesses, and IRT interpretation document retain their parent bytes.

  • Exact historical 402fd63 focused verification: 210 tests passed in 2.66 seconds, covering NIM comparisons, release acceptance, workflow contracts, reasoning profiles, paper inventory, and product planning.
  • Exact historical 402fd63 full suite: 3,475 passed, two skipped, terminal exit 0 in 677.72 seconds. Start/end SHA both equal 402fd63a4f6a16558e8cb4e4cd8a3378da3827f6, and the tracked tree stayed clean. JUnit has 3,477 cases, zero failures, zero errors, and two skips. Evidence: /tmp/co-1074-measurement-integration.pCjFWW. Session 50435 completed; do not treat it as an active run.
  • Command: .venv/bin/python -m pytest -q --junitxml=<evidence-directory>/junit.xml.
  • Exact-head coverage recheck: 210 passed in 2.93 seconds, with NIM 1,223 statements / 446 branches and reasoning profiles 185 statements / 68 branches, all 100%. Configured public docstring coverage is 100%; that scope excludes private/nested functions and tests and is not the wider CodeRabbit changed-function scope.
  • git diff --check passes. Explicit source comparisons verify preservation of the separate parent and child changes.
  • Exact-head Trivy 0.74.0 filesystem scan: exit 0, no HIGH/CRITICAL findings under --ignore-unfixed, with default vulnerability and secret scanners. All four reported lock targets have zero findings. trivy fs --download-db-only checked the existing DB, updated 2026-09-05 07:05:47 UTC (cached download 10:07:59 UTC; next update 2026-09-06). This does not claim a new DB download. Default development/test dependency exclusions apply. JSON: /tmp/co-1074-measurement-integration.pCjFWW/trivy.json.
  • Hosted snapshot still has only CodeRabbit/Devin status contexts; no independent approval or complete repository Security/Quality acceptance was established. Draft status and parent-first protected delivery remain unchanged.
  • Parent executable/document head 47ae9d65 completed 3,460 passed / 2 skipped, exit 0, in 675.71 seconds on an unchanged clean tree; later parent documentation-only ae704491 passed 61 checks in 0.62 seconds. Neither result is a current-child full-suite result.

Historical child e2f3e2af completed 3,447 passed / 2 skipped, exit 0, in 677.14 seconds, with matching start/end head and clean tree; JUnit contains 3,449 cases, zero failures, zero errors, and two skips. That head also passed 149 NIM checks in 3.55 seconds with 1,223 statements / 446 branches at 100% and configured public docstrings at 100%. Its refreshed Trivy 0.74.0 scan exited 0 under HIGH/CRITICAL + ignore-unfixed with vulnerability and secret scanners; default development/test dependency exclusions apply. Those source/security results remain historical, not evidence of a scan or full run on the new integration.

Historical evidence: /tmp/co-1074-integration.naZc1Z. Earlier 1600b1d5 had 3,432 passed / 2 skipped in 664.58 seconds. None substitutes for exact-head hosted dependency review, pip-audit, CodeQL, the complete Security workflow, independent review, or protected delivery.

Inherited paper distribution correction

The parent audited its six bundled papers. FrugalGPT, RouteLLM, Hybrid LLM, and the fuzzing survey are removed from this current tree too; citations, source links, summaries, and historical hashes remain. HELM and IRT-Router retain explicit CC BY 4.0 sources and attribution. The independently licensed child LLMRouter PDF remains with its own record. No deleted copy was restored by merging, no source-code license was repurposed as paper permission, and Git history or prior release artifacts were not purged.

Public response-time evidence audit

  • Distinguish Feng et al. LLMRouter / xRouteBench from Li et al. LLMRouterBench; the latter estimates latency from tokens and serving statistics, not observed request-level p95.
  • Record immutable dataset and collector revisions. Three isolated source-function checks preserve seven-second successful and returned-error records, while an escaped exception records zero. This is a controlled source-contract finding, not an audited error count in the published dataset.
  • Dataset lineage, outcome completeness, sampling validity, and dataset redistribution rights remain unverified. No new estimator, dependency, provider call, production default, or performance claim is introduced.
  • Attach the unmodified CC BY 4.0 LLMRouter v1 PDF with SHA-256 and APA 7 attribution. A repeat download is byte-identical; Poppler's original dictionary warnings and the HTML reading alternative are documented. No raw dataset rows are redistributed.

Observed generic-matrix follow-up

Read and hash-verified all 147,924 rows from the immutable generic train/test files, using the existing PyArrow 25.0.1 reader with network denied during decoding. The committed aggregate JSON was checked again against every row. No dependency was installed and no raw prompt or response is redistributed.

  • Zero, negative, and non-finite duration counts are 0 in both inspected files; the source-code exception path does not prove zero-time contamination in this dataset.
  • 2,068 empty response strings have zero output tokens and scores, positive input/total tokens, and positive recorded time. Their terminal cause is unknown.
  • Train contains 18 repeated scored-task/model rows across one ARC-Challenge task and all 18 models. They are not exact duplicate records: 6 groups differ in output and 17 in elapsed time. No automatic deletion or independent-item interpretation is authorized.
  • No exact stimulus fingerprint overlaps train/test, and the observed task-by-model rectangles are complete. Neither result proves semantic independence or the full attempted-cell denominator.
  • Explicit terminal state, attempt identity, timing provenance, and dataset redistribution rights remain unresolved. This is an observation-contract audit, not a p95, ability, or accuracy result.

Dependency and review boundary

Draft, stacked on #1067. The parent remains open. Local Security and Quality now runs on this stacked PR, including current run 34077430336. This does not establish complete central required-workflow admission, independent review, or protected release; those owner boundaries remain separately tracked. Local validation is not protected-main or hosted acceptance. Require parent integration, complete validation on the then-current commit, and independent review before protected merge.

The coverage audit also repaired two pricing tests that previously accepted an unrelated hosted-access-expiry error. That independent test repair is already carried by the parent; this PR retains its history without duplicating its effective delta.

Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
@coderabbitai

coderabbitai Bot commented Sep 5, 2026

Copy link
Copy Markdown

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Record pinned source and dataset provenance, distinguish returned API errors from zero-duration escaped exceptions, and reject estimated latency as observed p95 proof. Attach the CC BY 4.0 paper without redistributing dataset rows.

Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
Read every row in the pinned generic train and test files using the existing Parquet reader. Record verified aggregate counts and hashes without redistributing prompts or responses. Distinguish blank outputs and repeated attempts from inferred provider failures.

Signed-off-by: Seongho Bae <me@seonghobae.me>
…ixes

Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>

# Conflicts:
#	docs/product-technical-gap-baseline.md
Signed-off-by: Seongho Bae <me@seonghobae.me>

# Conflicts:
#	docs/product-technical-gap-baseline.md
Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>

# Conflicts:
#	CHANGELOG.md
#	docs/product-technical-gap-baseline.md
Signed-off-by: Seongho Bae <me@seonghobae.me>

# Conflicts:
#	CHANGELOG.md
#	docs/product-technical-gap-baseline.md
Signed-off-by: Seongho Bae <me@seonghobae.me>

# Conflicts:
#	CHANGELOG.md
#	docs/product-technical-gap-baseline.md
Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
Keep role profiles and propagation unchanged. Require a future validated decision contract before any report can authorize defaults.

BREAKING CHANGE: remove PRODUCTION_RMSE_IMPROVEMENT_THRESHOLD; production_default_change_allowed now always returns false.

Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
@seonghobae

Copy link
Copy Markdown
Contributor Author

621cf4d9cd47d166bd568cfb07a25692c47df210을 정상 push했습니다. 55% 선언값만으로 운영 기본값 변경을 허용하던 경로를 제거했습니다. 기존 Boolean 함수는 유지하지만 검증된 의사결정 계약이 구현되기 전까지 모든 보고서에 false를 반환합니다. 제거한 상수 export와 이전 true 결과의 변경은 의도된 호환성 변경으로 커밋과 연구 문서에 명시했습니다. 역할 설정·provider 전달·실제 route/conduct 기본값은 바꾸지 않았습니다.

  • 테스트 선행 31ecc698: 출처 누락/synthetic/live 선언 × 후보 RMSE 0/0.1의 6건 모두 실패, 0.68초, 종료 코드 1.
  • 수정 후 관련 프로필·CLI·혼합 모델 선택·퍼징 96개 통과, 26.79초.
  • 해당 모듈 168문장·62분기 누락 0, coverage 100%; 공개 docstring 100%. 관련 58개 검사는 1.92초에 종료 코드 0.
  • 현재 커밋 전체 회귀 3,612 passed, 2 skipped / 1,110.73초, 종료 코드 0. JUnit도 3,614개 중 실패·오류 0, 제외 2이며 검사 후 HEAD와 tracked clean 상태를 확인했습니다.
  • 기존 backend로 wheel 빌드 후 별도 cwd/Python -I에서 wheel 내부 모듈 import와 45개 Python 파일의 원본 바이트 일치를 확인했습니다. 독립 환경 pip 설치나 배포 검증은 아닙니다.

이 수정은 근거 없는 승인을 닫는 선행 조치입니다. 실제 작업-정책 행렬, 채점자 보정, 실패 분모, 불확실성, Rust 통계 owner의 released contract와 독립 운영 수용은 여전히 필요합니다. 합성 추정식의 오차 감소를 실측 정확도 향상으로 주장하지 않습니다. 최신 HEAD의 hosted 검사·독립 리뷰·부모 우선 보호 병합이 필요하며 Draft를 유지합니다.

Signed-off-by: Seongho Bae <me@seonghobae.me>

Copy link
Copy Markdown
Contributor Author

Published documentation-only successor 77545c3a10041d65a7ce12d3d77268a5b7895f8d by normal push after predecessor run 34077430336 reached terminal SUCCESS. Two documents connect CO #86 to existing RankWeave #41/#45, distinguish released v0.18.0 from proposed native v0.19.0, and require immutable release/artifact conformance before consumption. No statistical kernel, dependency, runtime, provider request, or production-default change.

Local successor verification: paper/docstring contract tests 6 passed in 1.98 seconds; diff check and tracked working tree clean. This is not a new-head full-suite claim.

Preserved predecessor evidence: exact 621cf4d9cd47d166bd568cfb07a25692c47df210, run https://github.com/ContextualWisdomLab/contextual-orchestrator/actions/runs/34077430336 — tests/package 101606188600, fuzz 101606188752, security 101606188784 all completed successfully. Raw full suite: 3612 passed, 2 skipped, 796.56 seconds. Benchmark-only verification: 157 passed; 1225 statements and 448 branches at 100%, public docstrings 100%; wheel build/install/import step passed. This is neither whole-repository 100% coverage nor evidence of real buyer accuracy/latency improvement. The run was not cancelled or restarted.

New-head hosted checks, independent review, protected parent integration, immutable release and prospective measurement remain required. Draft status is unchanged.

Copy link
Copy Markdown
Contributor Author

Exact-head hosted verification completed for 77545c3a10041d65a7ce12d3d77268a5b7895f8d.

Run https://github.com/ContextualWisdomLab/contextual-orchestrator/actions/runs/34078472161 is terminal SUCCESS: tests/package 101609086026, security 101609085887, fuzz 101609086062 all succeeded. Raw tests/package log: 3612 passed, 2 skipped in 763.56 seconds; benchmark verification 157 passed in 11.75 seconds, 1225 statements/448 branches at 100% and public docstrings 100%; wheel build/install/import step passed. Coverage scope is the CI-selected benchmark module, not the entire repository.

Re-read PR state: head remains 77545c3, base ec1c4e6, Draft, no submitted reviews. This is current-head software verification, not independent approval, protected parent integration, release, real buyer accuracy/latency improvement, or psychometric validity. No source change, rerun, cancellation, or provider request was made.

BREAKING CHANGE: benchmark reports require schema 3 with a prospective evaluation identity plan.

Signed-off-by: Seongho Bae <me@seonghobae.me>
…60905' into codex/paired-policy-outcomes-20260905

Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
…60905' into codex/paired-policy-outcomes-20260905

Signed-off-by: Seongho Bae <me@seonghobae.me>

# Conflicts:
#	tests/test_nim_benchmark.py
@seonghobae

Copy link
Copy Markdown
Contributor Author

독립 Visual Inspection: 원격 head 4351af9f1b6dca2f97526a46ac6dca28048ff623을 API로 확인한 뒤, 동일 SHA의 docs/nim_benchmark.md#comparison-report-version-3를 실제 Edge Preview 화면과 스크린샷으로 검사했습니다. 버전 3의 계획/관측 정합성, 내부 검증과 독립 사전등록의 구분, 동점일 때 관측·다른 정책 비교를 보존하는 설명, 구버전 보고서 처리 문단이 데스크톱 화면에서 잘림·겹침 없이 표시됐습니다. 이번 검증은 문서 렌더링에 한정하며 모바일·실제 공급자 실행·hosted CI·배포를 입증하지 않습니다.

Hosted quality coverage failed at 99% because validate_report_schema's
empty-identity, invalid-identity, count, worker-mismatch, and unknown
skip-reason branches were untested. Keep those incomplete plans from
publishing as complete paired evidence.
@seonghobae

Copy link
Copy Markdown
Contributor Author

Hosted Tests and package quality failed after a green full suite because validate_report_schema coverage was 99% (lines 2791, 2798, 2809, 2814, 2819, 2827).

Exact head is now 9f5d0ce2. The added contract tests reject empty evaluation identities, invalid identity fields, non-positive counts, catalog worker mismatches, and unknown cheapest-worker skip reasons. Local three-file coverage: 169 passed, 100% statements/branches, interrogate 100%. Fresh hosted Checks on this SHA are the remaining quality evidence; this does not authorize a production default change.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working priority: medium Normal-priority or P2 work status: draft type: bug Defect or incorrect behavior

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant