Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
69 commits
Select commit Hold shift + click to select a range
fe17df9
test(noema): add observed defect false-negative corpus
seonghobae Sep 1, 2026
95b66f9
ci(temp): apply and verify PR1641 review-corpus repair
seonghobae Sep 1, 2026
e27d5c8
fix(ci): repair PR1641 temporary writer workflow
seonghobae Sep 1, 2026
eae62d1
ci(temp): trigger repaired PR1641 writer
seonghobae Sep 1, 2026
f702789
ci(temp): gate PR1641 writer while fixing review findings
seonghobae Sep 1, 2026
fe9b6da
test(noema): add generic class-evidence relabel regression
seonghobae Sep 1, 2026
468c450
fix(noema): stage concrete class-observation repair
seonghobae Sep 1, 2026
6e766c0
ci(temp): execute repaired PR1641 writer
seonghobae Sep 1, 2026
0dcfa8f
test(noema): reject vacuous class evidence
seonghobae Sep 1, 2026
d8eb254
test(noema): preserve workflow-trigger review regression
seonghobae Sep 1, 2026
10f6d11
fix(noema): bind class evidence to exact source
seonghobae Sep 1, 2026
d0c8b1f
fix(noema): harden nonvacuous review evidence
seonghobae Sep 1, 2026
8c1c82a
ci(temp): execute repaired PR1641 writer
seonghobae Sep 1, 2026
e1a4406
test(noema): avoid accidental source-token match
seonghobae Sep 1, 2026
b501d26
test(noema): keep duplicate observation regression precise
seonghobae Sep 1, 2026
b046568
ci(temp): execute repaired PR1641 writer
seonghobae Sep 1, 2026
579d5a3
fix(noema): add exact-source follow-up repair
seonghobae Sep 1, 2026
397d9d2
ci(temp): execute repaired PR1641 writer
seonghobae Sep 1, 2026
cc3c980
fix(noema): replace lexical causal admission with structural roles
seonghobae Sep 1, 2026
33b6966
ci(temp): execute repaired PR1641 writer
seonghobae Sep 1, 2026
f3ae6d2
ci(temp): execute repaired PR1641 writer
seonghobae Sep 2, 2026
8da5c3f
fix(noema): repair exact-source evidence fixtures
seonghobae Sep 2, 2026
1c3346e
ci(temp): execute repaired PR1641 writer
seonghobae Sep 2, 2026
33ec0a7
test(noema): isolate repair deadline from external DNS
seonghobae Sep 2, 2026
f33a22e
ci(temp): stage repaired PR1641 retrigger
seonghobae Sep 2, 2026
39ae210
ci(temp): execute repaired PR1641 writer
seonghobae Sep 2, 2026
e8e08cc
test(noema): cover observed-evidence validator branches
seonghobae Sep 2, 2026
02724d3
ci(temp): execute repaired PR1641 writer
seonghobae Sep 2, 2026
6a4ff1d
test(noema): cover final class-evidence failure branches
seonghobae Sep 2, 2026
f5cf383
ci(temp): execute repaired PR1641 writer
seonghobae Sep 2, 2026
0918cb3
ci(temp): execute repaired PR1641 writer
seonghobae Sep 2, 2026
249f28c
ci(temp): execute repaired PR1641 writer
seonghobae Sep 2, 2026
7c7e2fc
fix(noema): enforce structural observed-defect evidence
Sep 2, 2026
b196d3f
docs(noema): record successor-check proof
seonghobae Sep 2, 2026
7b145bd
ci(temp): repair Noema source-evidence edge cases
seonghobae Sep 2, 2026
0443565
ci: move PR 1641 edge repair off saturated runner pool
seonghobae Sep 2, 2026
812853e
ci: add robust PR 1641 edge repair
seonghobae Sep 2, 2026
f445324
ci: add fail-closed PR 1641 exact-source edge repair
seonghobae Sep 2, 2026
168f053
ci: add executable PR 1641 exact-source edge repair
seonghobae Sep 2, 2026
cdcd220
ci: make PR 1641 edge repair executable
seonghobae Sep 2, 2026
361d9fb
ci(noema): execute and self-clean exact-source edge repair
seonghobae Sep 2, 2026
2136826
chore(noema): remove superseded PR1641 repair workflow
seonghobae Sep 2, 2026
a78131a
chore(noema): remove superseded PR1641 edge workflow
seonghobae Sep 2, 2026
ec53e9c
chore(noema): remove superseded PR1641 repair helper
seonghobae Sep 2, 2026
70bc213
ci(noema): retrigger exact-source edge repair
seonghobae Sep 2, 2026
468d0ed
ci(noema): supersede flawed PR1641 source writer
seonghobae Sep 2, 2026
e4eeaa5
ci(noema): replace indentation-fragile PR1641 writer
seonghobae Sep 2, 2026
2224c45
ci(noema): run robust PR1641 exact-source GREEN repair
seonghobae Sep 2, 2026
d73afbd
ci(noema): keep broad source-evidence regression suite causal
seonghobae Sep 2, 2026
ba2ee0a
fix(ci): make PR1641 GREEN writer executable
seonghobae Sep 2, 2026
df0f735
fix(noema): expose structured failure kind
seonghobae Sep 5, 2026
2a412b8
test(noema): verify failure-kind log boundaries
seonghobae Sep 5, 2026
37f090f
Merge remote-tracking branch 'origin/main' into codex/pr1898-evidence
seonghobae Sep 5, 2026
49ae54e
fix(noema): preserve canonical gateway error code
seonghobae Sep 5, 2026
719c91b
Merge remote-tracking branch 'origin/main' into codex/pr1898-evidence
seonghobae Sep 5, 2026
0db01c2
test(noema): cover sparse gateway failure receipts
seonghobae Sep 5, 2026
c759f23
fix(noema): bind claimed verification to trusted receipts
seonghobae Sep 5, 2026
fbe2602
fix(noema): unify exact diff evidence parsing
seonghobae Sep 5, 2026
2946017
fix(noema): close whitespace evidence provenance gaps
seonghobae Sep 5, 2026
a7afb39
Merge Noema evidence provenance root into failure telemetry
seonghobae Sep 5, 2026
43a1fdc
fix(noema): align structured probe contracts
seonghobae Sep 5, 2026
9df1ea4
test(noema): cover malformed provenance containers
seonghobae Sep 5, 2026
46a0178
Merge Noema strict-schema root into failure telemetry
seonghobae Sep 5, 2026
409638a
Merge protected main into Noema evidence root
seonghobae Sep 5, 2026
f1cd875
Merge protected main through Noema evidence root
seonghobae Sep 5, 2026
80fc255
fix(noema): align decision status schema and close HTTP errors
seonghobae Sep 6, 2026
c41cc45
Merge Noema schema and HTTP response repair into telemetry stack
seonghobae Sep 6, 2026
ad48dd6
fix(noema): 응답 정리 오류와 원래 통신 실패를 분리
seonghobae Sep 6, 2026
4cf6feb
Merge Noema cleanup repair into telemetry stack
seonghobae Sep 6, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
13 changes: 13 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -52,6 +52,14 @@
- Raised `hourly-review-repair.yml`'s discovery ceiling from 50 to 200 while rotating deterministic 50-PR deep-inspection windows by hourly run number. The scheduler hydrates only the selected window and stops immediately after its single dispatch, preserving access to newer PRs without quadrupling expensive review/check/comment work. See `docs/doctoring/hourly-review-repair-single-file-consolidation.md`'s 2026-09-03 follow-up.

## [Unreleased]
- Preserve the gateway's bounded error classification in failed Noema review
diagnostics, helping maintainers select the relevant investigation without
exposing free-form response bodies. Missing classifications remain unknown;
this does not resolve the historical gateway failure. See PR #1898 and
`docs/doctoring/noema-repair-attempt-telemetry.md`.
- Failed Noema reviews retain the original network failure even when response
cleanup also fails, while process cancellation still stops the review. Direct
redirect-rejection tests now release their responses explicitly.
- Include merge-scheduler entrypoint, core, and regression-test changes in
the existing runtime-quality workflow's trigger and suite selector. Scheduler
workflow edits retain queue checks and also select the full review-repair
Expand Down Expand Up @@ -151,6 +159,9 @@ this file. The format follows Keep a Changelog, and versioned releases follow
Semantic Versioning where the repository publishes a release.

## [Unreleased]
- **Keep Noema's strict output schema and deterministic probe validator identical (#1641).** Each structured probe now declares its closed `probe_kind` together with the exact required `class_evidence` witness roles and source receipt fields. Nested `anyOf` variants preserve strict OpenAI-compatible required/additional-property semantics, so a realistic verdict cannot be rejected merely because the outbound schema and local admission contract disagree. The single-request invalid-location regression now reaches and asserts the intended changed-side rejection instead of passing on an earlier status mismatch.
- **Require source-bound observed defect classes in Noema formal reviews (#1641).** Canonical changed-line coordinates now reject JSON booleans, material reviews must cover distinct classes from the executable external-finding corpus, and class witnesses bind to exact changed-side source text (including lexical-shape-independent blank/non-ASCII lines) with non-vacuous causal observations. A single parser now owns both source text and coordinates; bounded truncation drops the incomplete line instead of synthesizing a changed-line marker, so genuine source equal to the old marker remains reviewable. The prompt explicitly attacks workflow-event authority plus mutable-alias, TOCTOU, identity, oracle, contract, authority, dependency-context, coercion, and state-machine failure shapes without fabricating benchmark claims.
- **Fail closed on fabricated Noema execution and external-source provenance (#1641).** Model-authored claims that runtime behavior, command output, toolchain help, or authoritative external documentation confirmed a conclusion now require an out-of-band typed receipt and an exact receipt citation. The isolated reviewer may still reason from changed source and recommend toolchain-specific verification; it cannot present that recommendation as executed evidence. This regression is grounded in `ConceptWeave#35@a31ae0c2`, where review `5120903874` claimed Cargo runtime/documentation confirmation although required Noema run `33938445009` executed no Cargo or documentation lookup step.
- **Pin `opencode-review-dispatch.yml` off the starved floating `ubuntu-latest` image.**
The 2026-09-01 floating-image fix (see that entry below) pinned `strix.yml`,
`opencode-review.yml`, and `noema-review.yml` -- the three required-check
Expand Down Expand Up @@ -1458,3 +1469,5 @@ Semantic Versioning where the repository publishes a release.
- Added an organization-owned reusable exact-artifact SBOM attestation boundary that validates inert six-file wheel/sdist evidence, binds CycloneDX 1.7 predicates to exact SHA-256 subjects, signs through least-privilege GitHub artifact attestations, and exports online and offline verification bundles.
- Hardened exact-artifact SBOM verification with strict finite RFC 8259 JSON, integer CycloneDX document versions, deterministic UUIDv5 subject identities, exact filename properties and single SHA-256 root bindings, environment-only shell input transfer, pinned Ubuntu 24.04 quality runners, and checksum-sealed beginner-readable offline evidence. The decision record now cites Bray (2017) so NaN and Infinity cannot be treated as sealed SBOM numbers.
- Recorded the org control-plane architecture, including exact-artifact SBOM attestation, so agents reconstruct the signing trust boundary from the repo instead of private memory.

- Noema review evidence now uses exact class-and-field claim roles and source excerpts instead of a fixed English causal-word heuristic, preserving non-ASCII and symbol-only review evidence without treating keywords as proof.
27 changes: 27 additions & 0 deletions docs/doctoring/noema-observed-defect-corpus-current-main.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,27 @@
# Noema observed-defect review corpus

The trusted Noema review gate treats externally demonstrated review misses as executable regression evidence, not as benchmark claims. Material source/test reviews must exercise at least two distinct observed defect classes and every admitted class witness remains bound to an exact changed-side source coordinate.

The current closed taxonomy is: `mutable_alias`, `time_of_check_time_of_use`, `execution_identity`, `coercion_boundary`, `test_oracle`, `cross_contract`, `authority_boundary`, `dependency_context`, and `state_machine_race`. Each class has class-specific witness keys. Witness values are `{path,line,side,source_excerpt,claim_role,observation}` records bound to the probe location. `source_excerpt` must equal the exact changed-side line, and `observation` must quote the exact source line (or `<blank>`) plus a causal/behavioral relation beyond taxonomy labels; ASCII token shape is not admission authority; repeated or differently worded generic labels do not satisfy the deterministic validator.

The outbound strict structured-output schema and the local validator share that same closed contract. Every probe is one nested `anyOf` variant that correlates a single `probe_kind` with exactly its required `class_evidence` keys; every witness field is required and unknown fields are rejected. Only the containing `adversarial_validation` value is nullable for a non-formal comment. This follows the strict structured-output rule that object properties are required (nullable when truly optional) and prevents the gateway from accepting a probe shape that deterministic admission must reject.

The model is explicitly asked to attack mutable/immutability escapes, changing getters/TOCTOU, request or tenant identity confusion, weak/vacuous oracles, cross-contract contradictions, authority overreach, missing causal dependency context, and reliability/security state-machine races. A falsified hypothesis is valid evidence and must not be promoted into a finding merely to satisfy taxonomy diversity. For CI/automation changes, the review prompt also requires checking whether the mutation credential can create the downstream events/checks the state machine depends on.

JSON booleans are rejected as line coordinates even though Python considers `True == 1`: changed-line evidence requires `type(line) is int` and a positive value. Production review calls always provide the complete changed-path manifest, which activates the observed taxonomy; direct validator unit tests may omit that manifest to exercise lower-level generic schema boundaries independently.

This repair is a narrow current-main successor to the heavily diverged PR #1589 evidence lineage. It does not copy CodeRabbitAI or Devin wording and makes no superiority claim.

Exact-head follow-up removes synthetic bounded-diff omission lines from the diff grammar entirely: truncation drops the incomplete final line and carries the separate `truncated` control flag. A genuine source line equal to the historical marker remains admissible, as do short identifiers, symbol-only lines, blank changed lines, and non-ASCII source through exact string equality rather than lexical guessing. Coordinates and source text now come from one parser so future diff fixes cannot desynchronize their trust boundaries.

The exact-head structural follow-up removes the fixed English relation-word list. Formal evidence now carries a schema-derived `claim_role` for each defect-class witness, while the deterministic gate verifies exact source identity, canonical coordinates, role identity, and distinct observations. Semantic causal adequacy remains a reviewer/evaluation responsibility; the validator does not pretend English keyword presence proves causality.

Workflow-local bootstrap or generated commits are not accepted as final review/check proof merely because their source transaction verified locally. The merge candidate must be a workflow-starting successor writer head produced through ordinary owner-side mutation, with the required review and quality checks observed on that exact unchanged head before merge.

Failed HTTP responses belong to the requesting transport. After bounded telemetry
extraction, close the response there; a secondary cleanup exception must not replace
the original typed transport failure. Process cancellation still propagates. Tests
that invoke a redirect handler directly own the resulting HTTPError and must close
it themselves rather than relying on garbage collection. Run the Noema regression
tests with `-W error`; tests in other HTTP consumers do not become passing evidence
merely because this transport was repaired.
46 changes: 45 additions & 1 deletion docs/doctoring/noema-repair-attempt-telemetry.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,7 +14,7 @@ The later review established a second ownership error: `contextual-orchestrator`

Noema now sends exactly one structured-output request to the configured gateway. GitHub Actions fixes the model alias to `orchestrator/free`; the caller declares no provider, paid fallback, sampling temperature, or fixed inference timeout. `contextual-orchestrator` owns provider discovery, capability routing, structured-output repair, failover, and upstream completion. The repository remains responsible for deterministic local validation and exact-head publication.

Every gateway call emits exactly one passive Actions annotation. Success and failure annotations include caller attempt count, elapsed duration, active phase (`connecting`, `reading`, `decoding`, or `validating`), and a best-effort serving-model identifier. Serving-model text is secret-scrubbed, control-character-normalized, UTF-8 printable, and bounded before it can reach an annotation. Raw model output is never logged.
Every gateway call emits exactly one passive Actions annotation. Success and failure annotations include caller attempt count, elapsed duration, active phase (`connecting`, `reading`, `decoding`, `validating`, or `response_error`), and a best-effort serving-model identifier. The identifier validator trims exterior whitespace, then accepts only 1–200 ASCII characters in its restricted identifier alphabet. Invalid text is omitted, not repaired. This is format and length validation, not arbitrary secret detection; the gateway must supply non-sensitive identifiers. Raw model output is never logged.

The local trailing-comma parser remains a deterministic syntax transform only. It may remove a genuine trailing comma after a complete JSON value, but missing-value forms such as `[,]`, `{,}`, `[1,,]`, and `{"a":,}` remain invalid. The transform emits no second attempt-level annotation and never bypasses semantic verdict validation.

Expand All @@ -32,3 +32,47 @@ If the gateway cannot produce a valid structured verdict, Noema fails closed aft
## Verification

The permanent contract test forbids `NOEMA_REPAIR_DEADLINE_SECONDS`, `_repair_wall_clock_deadline`, `NoemaRepairDeadlineExceeded`, `signal.setitimer`, retry-only parameters/recursion, and caller-specified `temperature`. Focused regressions prove one request on success and failure, one annotation per attempt, safe serving-model telemetry, strict missing-value rejection, accepted genuine trailing commas, and preserved exact changed-line diagnostics.

## 2026-09-05 follow-up: retain the gateway's error classification

Status: proposed consumer repair in [central PR #1898](https://github.com/ContextualWisdomLab/.github/pull/1898), not protected delivery or a resolved provider incident. This extends the evidence/control-plane requirement and G-02/G-03; ADR-0003's single-request ownership is unchanged.

### Observed failure and source trace

[Naruon #1244's failed job](https://github.com/ContextualWisdomLab/naruon/actions/runs/33933793278/job/101247827882) at head `50351e8cacc65b4124ba2145e00d41aeceef0775` reported HTTP 502, one caller attempt, `duration=1469.1s`, `phase=response_error`, and `served_model=deepseek-ai/deepseek-v4-flash-0731`. It did not preserve an error code or failure kind. The exception label `Noema gateway transport failed` therefore does not establish a network failure, nor does it establish structured-output exhaustion. That historical cause remains unknown.

Protected contextual-orchestrator source at `a080297d2546bb61e89520d637cabc202db331ec` already maps `ProviderResponseError` to HTTP 502 and the literal `invalid_structured_output` in [`server.py:7978`](https://github.com/ContextualWisdomLab/contextual-orchestrator/blob/a080297d2546bb61e89520d637cabc202db331ec/contextual_orchestrator/server.py#L7978). [`_send_error`](https://github.com/ContextualWisdomLab/contextual-orchestrator/blob/a080297d2546bb61e89520d637cabc202db331ec/contextual_orchestrator/server.py#L8189) adds a request identifier; [`_error_payload`](https://github.com/ContextualWisdomLab/contextual-orchestrator/blob/a080297d2546bb61e89520d637cabc202db331ec/contextual_orchestrator/server.py#L686) places the classification in canonical `error.code`. This path has no `failure_kind`, model, or attempt list. Those source facts identify an observable envelope, not the cause of the earlier Naruon run or an immutable release.

The original #1898 delta at `df0f735f42adbb44d45f3c3a4e503e400b47ed79` retains optional `error.detail.failure_kind`, but still discards `error.code`. Waiting only for contextual-orchestrator [#1004](https://github.com/ContextualWisdomLab/contextual-orchestrator/pull/1004), whose proposed `36133c8ab85d44fc4be2356edbdd56d9fc09f0d8` adds `structured_output_exhausted`, would leave the existing protected-source envelope unclassified. Noema must not infer that kind from a status code, copy owner repair logic, or consume the proposed branch as a released dependency.

### Chosen repair and security boundary

The existing bounded reader and formatter now preserve both independent fields: `error.code` as `error_code`, and `error.detail.failure_kind` as `failure_kind`, when present and valid. The consumer repair is [commit `49ae54e789c0a6951b3212e2182e4c64d0348a81`](https://github.com/ContextualWisdomLab/.github/commit/49ae54e789c0a6951b3212e2182e4c64d0348a81). Both the single failure annotation and the raised diagnostic use the same extracted receipt. Missing fields remain absent; the failure still fails. No request, retry, provider selection, credential, timeout, or verdict-approval rule changes.

The reader requires canonical mapping envelopes including `error.detail`, reads at most 16 KiB plus one oversize-detection byte, and rejects malformed or oversized bodies. Each new field reuses the existing identifier validator; neither free-form messages, request identifiers, arbitrary detail, nor flattened compatibility aliases are logged. Embedded CR/LF, terminal escape sequences, surrogates, delimiter injection, non-string values, and overlength identifiers cannot enter the new fields. A syntactically valid secret placed in an allowlisted field would not be detected by this validator: non-sensitive canonical classifications remain a producer obligation.

This follows OWASP's advice to define log field types and lengths, validate data crossing trust zones, prevent log injection, and exclude credentials and sensitive payloads. It does not claim that a regular expression supplies complete redaction (OWASP Foundation, n.d.).

### Reproduction, integration, and remaining delivery gates

The existing failed-call regression covers 17 values for each independent field, plus those 17 values with the sparse current gateway envelope: 51 cases. It verifies annotation/exception output, absent sibling fields, one request/annotation, and exclusion of unrelated payload text. The sparse cases have a request identifier and compatibility aliases but no model, attempts, terminal reason, or failure kind. They address an independent static review's missing-fixture finding; they must preserve the code, report an unknown model, and exclude request-ID/message text. Unit-only HTTP doubles replace the external gateway; parsing, formatting, and `call_llm` execute normally.

Removing the original four `failure_kind` lines produced 3 failures; restoring them produced 138 focused passes. Adding canonical-code assertions to the old implementation reproduced 4 failures and 30 passes, including the actual protected-source code. Before the sparse-fixture extension, the consumer fix plus focused Noema edge coverage and the environment regression produced **159 passed, zero failures/skips**. With the sparse fixture added, removing just the four new code-extraction/formatting lines reproduced **8 failed / 43 passed**; restoring the implementation produced **51 passed**. These are chronological receipts, not totals for the final candidate.

The first broader run had **2897 passed, one skipped, 21 subtests passed, and one failure**: the task-local uv environment lacked `pip`, needed by `test_materialized_bounded_include_is_resolvable_by_pip`. The project already declares `pip==26.2.1`; installing that exact declared tool fixed the test without changing source, skipping it, or modifying a shared environment. The environment combines hash-locked review requirements and that separately declared pip pin; it is not claimed as an entirely hash-locked clean install.

The original PR delta and validation commit were preserved by ordinary merges. Protected main `f250638827f8252b0d9e5cb2601f4d333f96162f` (merged prerequisite #1922) is integrated at `719c91b1f678de6da3029b8f5920d6a245520e2e`. A preliminary normal run returned **2923 passed, one skipped, 21 subtests passed** before the sparse-fixture follow-up. The existing LLVM 19 admission test was skipped because its reviewed tools are absent on this macOS host; that path remains unverified, not passed. The separate maintainer exception reported for #1922 is not authorization to bypass #1898's gates. Full normal and `GITHUB_ACTIONS=true` verification must finish on the final integrated candidate, followed by fresh current-head hosted checks and qualifying independent review. Capture exact head/base and final command results in #1898; do not transfer old-head passes.

Run from the isolated repository root:

```sh
.venv/bin/python -m pytest -q -W error tests/test_noema_review_gate.py tests/test_noema_model_output_edge_coverage.py
PATH="$PWD/.venv/bin:$PATH" .venv/bin/python -m pytest tests -q -W error -rs
GITHUB_ACTIONS=true PATH="$PWD/.venv/bin:$PATH" .venv/bin/python -m pytest tests -q -W error -rs
```

After protected delivery, confirm a real consumer run uses the exact released central revision and gateway contract. If it fails, retain the returned classification and investigate that owner path. Do not reroute to a paid model, repeat an active model request, weaken semantic validation, or declare the historical 502 repaired from unit evidence. Product runtime, real PostgreSQL, browser, release, and deployed-gateway verification are not covered by this consumer diagnostic test.

### Reference

OWASP Foundation. (n.d.). *Logging cheat sheet*. OWASP Cheat Sheet Series. Retrieved September 5, 2026, from https://cheatsheetseries.owasp.org/cheatsheets/Logging_Cheat_Sheet.html
Loading
Loading