Skip to content

Make central PR validation recoverable with exact-head evidence #2477

Description

@seonghobae

Goal

Make central PR validation recoverable and auditable for fast-mlsirm and other consumers. Each required Strix, Noema, OpenCode, and CodeQL verdict must be tied to the current PR head and real producer evidence. Provider and runner-capacity failures must remain non-passing, with bounded same-head continuation where the existing contract permits it.

Current evidence

Acceptance

  1. Diagnose each canary check from its exact-head run and artifact, separating code defects, provider outages, and scheduler delay.
  2. Repair confirmed central workflow defects with focused tests. Keep security gates fail-closed, preserve findings, and use bounded same-head continuation only for explicit transport failures.
  3. Verify the five checks on a live current-head canary or record a typed terminal infrastructure result. Do not describe a queued, skipped, synthetic, or bypassed check as green.
  4. Keep the operator receipt current as runs advance; close only after a genuine hosted validation cycle proves the contract.

Active owner

Codex, 2026-09-28T03:49:36Z. First increment: restore bounded continuation for Strix runtime transport failures.

Activity

  1. seonghobae commented on Sep 28, 2026

    @seonghobae
    ContributorAuthor

    2026-09-28 UTC receipt for the active goal:

    • First repair: fix(ci): continue Strix runtime transport failures on same PR head #2479 at cd04e6dacd97e735c8a44fd6bef3d24ab8aa4269 connects Strix runtime connection/rate-limit failures to the existing bounded same-head continuation. The scan still fails closed. A review-identified mixed sandbox/connection log is excluded. Focused tests: 83 passed under GITHUB_ACTIONS=true; the new module has 100% focused statement/branch coverage. Hosted checks are queued, so this is not hosted acceptance.
    • 🛡️ Sentinel: [HIGH] CSP 'unsafe-inline' 제거 및 동적 해시 적용 fast-mlsirm#2238 remains at 8053586b33145d9f4a4231f8193e89a4e73d1068. Its earlier direct CodeQL compatibility jobs (108701353093/108701353056) deliberately failed after dispatch while waiting for authenticated central verdicts; its direct OpenCode job 108701609296 failed while waiting for an exact-head reviewer verdict. Those logs do not identify a source defect.
    • Central Noema run 36360063791 continuation job 108765382600 was still queued with runner_id 0. At the runner snapshot, six self-hosted runners were online and busy; the central repository had 619 queued runs. This supports a scheduling-capacity diagnosis but does not establish that a runner label is misconfigured.
    • The local full Strix fixture script was stopped after more than eight minutes to avoid duplicate resource use while another worktree ran the same broad script. Its terminal result is unverified. The focused tests above are the completed local evidence.

    Next: follow exact-head hosted checks and reviews on #2479, then verify deployed continuation on a genuine current-head canary. Keep failed, queued, and skipped results distinct.

  2. seonghobae commented on Sep 28, 2026

    @seonghobae
    ContributorAuthor

    2026-09-28 follow-up: The #2479 superseded Strix and runtime-quality runs for heads d73e0ee2a and e770858cd reached completed/cancelled; only current head cd04e6dac remains queued. The CodeRabbit mixed sandbox/connection thread is resolved after commit cd04e6dac.

    For ContextualWisdomLab/fast-mlsirm#2238, central CodeQL run 36355503206 completed validate-dispatch successfully at 04:04:40 UTC, then queued real CodeQL dispatch scan (python) job 108782399240 and (actions) job 108782399244. This is admission progress, not a terminal CodeQL verdict. Central Noema continuation job 108765382600 and OpenCode coverage-evidence job 108764444766 remain queued. All six org self-hosted runners were online and busy at the latest snapshot; 573 central runs were queued. No runner-label defect is established by that evidence.

  3. seonghobae commented on Sep 28, 2026

    @seonghobae
    ContributorAuthor

    2026-09-28 control-runner RCA snapshot: continue-noema-transport job 108765382600 has been queued since 02:35:39 UTC with labels [self-hosted,linux,x64,cwlab-control] and runner_id 0. Live org runner group 6 (CWL central control) is visibility=all, allows_public_repositories=true, restricted_to_workflows=true, and includes the central main noema-review.yml path. Its one registered worker is online and busy. The workflow and runner labels match. This evidence supports a group-capacity wait; it does not support widening the security boundary or changing labels.

    The apparently long-running central jobs were admitted only recently: Noema #2464 model job started 03:46, OpenCode appguardrail coverage started 04:04, and naruon CodeQL's three analysis jobs started 04:13. Their earlier run timestamps mostly reflect queue age. All three CodeQL workers were active at the snapshot. Recheck job completions before changing allocation.

  4. seonghobae commented on Sep 28, 2026

    @seonghobae
    ContributorAuthor

    Correction to my preceding control-runner RCA: at one live read immediately after clearfolio job 108750194665 completed successfully at 04:22:15 UTC, cwlab-s2-01 was online idle while #2238 Noema continuation job 108765382600 remained queued. A later read found the runner busy again and the Noema job still queued. The idle observation was transient, so it does not prove persistent runner eligibility failure; it does disprove treating a busy snapshot alone as a complete explanation. The job has no pending deployment approval, group 6 permits the central main Noema workflow, and its labels match. Workflow-level concurrency is keyed by target repo and PR with cancel-in-progress. Current cause is admission/scheduling unresolved. I will not change group access, labels, or cancel another PR's job without stronger evidence.

  5. seonghobae commented on Sep 28, 2026

    @seonghobae
    ContributorAuthor

    2026-09-28 local verification update for #2479 head cd04e6dacd97e735c8a44fd6bef3d24ab8aa4269: the exact central Strix quality script used by Agent Review Runtime Quality CI completed with exit code 0 and test_strix_quick_gate: PASS under STRIX_TEST_PROCESS_TIMEOUT_SECONDS=3 STRIX_TEST_FAKE_SLEEP_SECONDS=5. This extends the earlier 83 affected pytest passes and the new module's 100% focused statement/branch coverage. It is local contract evidence only; hosted current-head checks and a complete vulnerability report are still required.

  6. seonghobae commented on Sep 28, 2026

    @seonghobae
    ContributorAuthor

    2026-09-28 04:49–04:51 UTC admission follow-up for fast-mlsirm #2238 (8053586b33145d9f4a4231f8193e89a4e73d1068): central Noema continuation job 108765382600 remained queued, runner_id=0, since 02:35:39 UTC. The CWL central control group includes online cwlab-s2-01; it was idle in both samples (~45 s apart). The job requests self-hosted,linux,x64,cwlab-control; the runner has those labels; the group is visible to all public repositories and allowlists central noema-review.yml@refs/heads/main. Thus continuous runner busyness is not a sufficient RCA. This is an observed admission/scheduling gap under backlog; GitHub's assignment cause is still unproven. No broad rerouting or duplicate dispatch has been made.

    Central CodeQL dispatch jobs python 108782399240 and actions 108782399244 were also queued. CWL central CodeQL briefly had an idle online runner at 04:49 UTC but both group members were busy at 04:51 UTC. Their workflow and runner labels match the group allowlist; this sample alone does not prove an eligibility defect. All five canary gates remain without terminal exact-head evidence.

  7. seonghobae commented on Sep 28, 2026

    @seonghobae
    ContributorAuthor

    2026-09-28 04:56–04:57 UTC milestone: central Strix runtime-continuation repair PR #2479 was admin-squash-merged at 1d908b76a984518393dc46d7e0181b19b97c50bc while hosted required checks were still queued. This is an authorized queue bypass after the local 83-test contract suite and test_strix_quick_gate.sh passed; it is not a clean hosted-validation pass.

    I dispatched one new default-branch Strix canary for open, ready fast-mlsirm #2238, exact head 8053586b33145d9f4a4231f8193e89a4e73d1068, base 8c1361a794176f040e7b999ca8ffc1fb4c0b5607, retry attempt 0. Central run 36379830577 was created from the repaired central main commit 1d908b76a984518393dc46d7e0181b19b97c50bc and is queued. No scan result, report, or continuation verdict exists yet. Do not count the pre-fix run 36361585404 as validation of the repair; its scan failed on a provider connection timeout and the continuation skipped.

  8. seonghobae commented on Sep 28, 2026

    @seonghobae
    ContributorAuthor

    2026-09-28 05:00 UTC Strix evidence audit for fast-mlsirm #2238, exact head 8053586b33145d9f4a4231f8193e89a4e73d1068: native run 36343272866 ended success and uploaded strix-reports artifact 10941945761. The scan metadata says completed/success and SARIF has zero results, but the generated report asserts a generic “Python OpenSSH client RCE vulnerability” with no changed path or concrete finding. The scanner log confirms zero vulnerability reports; this is contradictory narrative, not proof of a real OpenSSH defect or a clean, scoped review.

    RCA: that native run checked out trusted central source f5864568aee369f7567bd7b8b0ac249f9ec4c513, before the current strix_report_scope.py guard was added. Replaying its downloaded run.json and report against the current guard with #2238's changed paths (python/fast_mlsirm/report.py, python/fast_mlsirm/scoring/essay/report_html.py) fails as intended: scan report does not identify a changed source file. Therefore the old green check is historical execution evidence but does not establish a scoped clean verdict under today's contract. The repaired-main canary run 36379830577 is still queued; await its own artifact and terminal result before counting Strix.

  9. seonghobae commented on Sep 28, 2026

    @seonghobae
    ContributorAuthor

    2026-09-28 05:07 UTC consumer-failure log audit for open fast-mlsirm #2238 at 8053586b33145d9f4a4231f8193e89a4e73d1068:

    • OpenCode consumer job 108701609296 failed because no authenticated APPROVED or CHANGES_REQUESTED review from opencode-agent existed on that exact head. This is intentional fail-closed behavior while the central coverage-evidence job 108764444766 is still queued, not evidence of a fast-mlsirm source defect.
    • CodeQL Python job 108701353056 and Actions job 108701353093 both logged DISPATCH_OUTCOME=success, VERDICT_STATE=pending; their failed compatibility results explicitly await the authenticated central scan and subsequent rerun. Central scan jobs 108782399240/108782399244 remain queued. Neither consumer failure proves a CodeQL finding.
    • Native Noema job 108700905202 crashed in the offline loopback sidecar body-limit self-test: CPython 3.12.14 reported Fatal Python error: Segmentation fault, with a request thread in contextual_orchestrator.server._send_error / logging, main thread waiting in http.client.getresponse, and _cffi_backend loaded. This is a local sidecar process crash before a model review, not a fast-mlsirm source verdict or proof of provider 504. A separate later central Noema run 36360063791 reached the model and failed on HTTP 504 capacity; its bounded continuation job 108765382600 remains queued. The native crash is one observed occurrence; a deterministic code or runner cause is not yet established.
  10. seonghobae commented on Sep 28, 2026

    @seonghobae
    ContributorAuthor

    2026-09-28 05:11 UTC central .github queued-run census: one paginated read returned 589 queued runs (six pages). Major workflow IDs/counts: Code Quality dynamic 109; central CodeQL dispatch 72; CodeQL PR 62; OpenCode dispatch 58; agent-review-runtime-quality 52; Security Scan 37; Python Security 33; Semgrep 30; merge scheduler 25; required Noema 25; required OpenCode 16; Strix 16.

    Of 183 queue groups whose run names explicitly bind workflow + target repository + PR number + target head, zero had duplicate runs for the same workflow/repository/PR. That does not prove all queue entries are deduplicated: the GitHub-managed Code Quality: PR #... dynamic workflow had 109 queued runs, with 37 runs spread over 15 PRs that had multiple queued heads (22 runs beyond one per PR). This is a bounded queue inefficiency, but even removing every surplus dynamic run would reduce the total by only ~3.7%, and those jobs have no verified impact on the current #2238 gate admission. No run was canceled or workflow routing changed from this census. The direct #2238 Strix, Noema, OpenCode, and CodeQL handles remain live and queued; the dominant bottleneck is still unproven at the GitHub assignment/organization-capacity boundary.

  11. seonghobae commented on Sep 28, 2026

    @seonghobae
    ContributorAuthor

    2026-09-28 05:17 UTC correction/follow-up to the native Noema crash classification: #2238's failed native job used older central control source f5864568aee369f7567bd7b8b0ac249f9ec4c513. Current central main already contains two relevant changes absent from that source: 0bea6aeee binds the sidecar process to the selected CPython interpreter's matching shared library, and PR #2473, merged as d1785f7c553d877f1a88c6dbdbfab9be8988c5e9, replaced process-global stderr redirection in the threaded HTTP 413 self-test with a temporary handler on the server logger. #2473's own verification did not prove the native segfault's precise cause or deterministic resolution.

    The later central Noema job 108748091705 executed on d1785f7c5: Provision contextual-orchestrator review sidecar succeeded (02:11:38–02:18:05 UTC), then Prepare Noema model verdict failed at 02:35:27 UTC on HTTP 504 provider_capacity_unavailable. This proves the updated native Linux path got past the previous sidecar-startup failure in that run. The model verdict is still missing; bounded same-head continuation job 108765382600 remains queued. No new code patch is justified by the historical native crash alone.

  12. seonghobae commented on Sep 28, 2026

    @seonghobae
    ContributorAuthor

    2026-09-28 follow-up: fast-mlsirm #2238 was closed by its author at 06:25 UTC as a CSP regression: the policy removes inline-style permission while report bar widths still use style attributes. This canary must not be used as acceptance evidence or merged.

    Its central Noema run 36360063791 is now cancelled. Before cancellation, continuation job 108765382600 succeeded at 05:53 UTC in scheduling bounded same-head retry after HTTP 504; that is not a model verdict. OpenCode coverage job 108764444766 failed at 05:38 UTC: the log reports it could not materialize base Rust dependencies because trusted cargo vendor raised FileNotFoundError for Cargo.lock, crates/fast-mlsirm-py/Cargo.lock, and fuzz/Cargo.lock. Reproduce this coverage-preparation failure on the next valid PR.

    To limit obsolete queue load, cancellation was requested for exact #2238 central runs Strix 36379830577, OpenCode 36351939775, and CodeQL 36355503206. Requests are not terminal conclusions. The durable goal remains exact-head terminal Strix, Noema, OpenCode, and CodeQL Python/Actions evidence on a valid future PR, with provider, runner, and source failures classified from logs. #2114's earlier admin merge is not a clean-green CI result.

  13. seonghobae commented on Sep 28, 2026

    @seonghobae
    ContributorAuthor

    2026-09-28 23:45 UTC: queue evidence for contextual-orchestrator (CO) repo-local security.yml PR jobs, gathered while pushing CO#1227. There is one reversible org change, which is now rolled back, and its result was null.

    Symptom. Repo-local security.yml jobs for pull_request events on hosted ubuntu-24.04 drain about 22h after creation. The queue is slow; it is not deadlocked.

    • Run 36361499677 was created at 00:14:56Z and its four jobs started at 22:42:53Z.
    • Run 36365088454 was created at 01:12Z and its jobs started at 23:08–23:14Z.
    • 34 security.yml runs are queued. The oldest was created at 03:24Z, so the 24h auto-cancel will start hitting them soon.

    What it is not.

    • It is not hosted capacity overall. At 14:43Z, central Semgrep and Strix for CO#1344 started on hosted ubuntu-24.04 within seconds. In the same minute, CO security.yml for that same PR was queued and still is.
    • It is not the runs-on expression. The scheduled security.yml run 36406314413 got a hosted runner 90s after creation.
    • It does not affect every repo. disksage repo-local hosted jobs start within seconds.
    • It is not the ubuntu-24.04/ubuntu-latest aliases on the self-hosted runners cwlab-s1-02/03/04/06, which sit in groups 3/4/5 that admit CO only for security.yml@refs/heads/main. I tested this directly. I removed both aliases from all four runners at 23:18:11Z. A queued job stayed unassigned for 11 min. I then cancelled and reran CO run 36409607933 at 23:30Z, and its fresh jobs stayed unassigned for 8 min. I restored the aliases at 23:39:38Z. The label lists before and after are identical: self-hosted,Linux,X64,cwlab,ubuntu-24.04,ubuntu-latest,cwlab-s1-0N. Please don't repeat this experiment. The older runbook note that hosted ubuntu-latest was starved may also describe this same repo-scoped slowness.

    Still unexplained. Something specific to CO repo-local pull_request jobs is delaying runner assignment, and the API doesn't show what. The next step is a GitHub Support ticket with the run IDs above.

    CO#1227 status at head 0b4d2503. Security and Quality run 36409607933 is back at the end of the queue because of my rerun; its 24h window ends around 2026-09-29T23:30Z. Required strix failed only on the report-scope basename match: 0 vulnerabilities, and the report named orchestrator.py. Reproduction evidence is on #2500, which fixes it. After #2500 merges, strix needs a rerun.

  14. 7 remaining items

  15. seonghobae commented on Oct 1, 2026

    @seonghobae
    ContributorAuthor

    Additional Orgmetra Ready materialization canaries — 2026-10-01

    Two further source-complete protected-base PRs made one ordinary Draft → Ready transition after live head/review/thread revalidation:

    • ContextualWisdomLab/Orgmetra#259@f1f152b0838e11cba1cf583706eb0983d56af373: Foundation 34057542130 and SAST 34057542122 succeeded; zero reviews and zero threads. The exact-head workflow inventory still contains only historical CodeQL failure 34057542155; no fresh Ready-triggered CodeQL identity materialized.
    • ContextualWisdomLab/Orgmetra#448@6a06062e9d4f6e07bf4780a60384f264f7983353: Foundation 36796513822 and SAST 36796513856 succeeded; both COMMENTED review findings are repaired and zero threads remain. The exact-head inventory still contains only Draft-skipped CodeQL 36796513815; no fresh Ready-triggered CodeQL identity materialized.
    • Canonical Gap writer ContextualWisdomLab/Orgmetra#100@301f843c01f301088234c27f758c2a074cfa9aed independently shows the same result after one Ready transition: 91 COMMENTED reviews, 65 resolved threads, Foundation 36796830293 and SAST 36796830323 success, but only Draft-skipped CodeQL 36796830370 and no fresh CodeQL run.

    These unchanged-head observations extend the earlier #235 specimen across three independent Orgmetra root lanes. They do not transfer predecessor verdicts or authorize merge. All remain merge HOLD for failed Dependency Review, absent fresh CodeQL evidence, and zero qualifying approvals. No lifecycle cycling, empty/no-op commit, manual rerun, synthetic status, or bypass was used.

  16. seonghobae commented on Oct 1, 2026

    @seonghobae
    ContributorAuthor

    Canonical CodeQL owner repair opened as Draft .github#2548, exact head 8f6a870bca516d34a737bc8bd4abf5d3573573ae, tree 5cc31dec201f7b5298e15e787ee998e8f2804b54. It removes only CodeQL's Draft predicate because ruleset consumers never receive unchanged-head ready_for_review; review workflows remain Ready-gated. The branch ordinary-merges prerequisite security owner #2531 and is stacked on that branch. RED reproduced the false assumption; integrated warnings-fatal focused evidence is 78 passed. Orgmetra #235/#259/#448/#100 remain unchanged-head canaries; post-merge proof still requires a new Draft consumer head reaching terminal authenticated CodeQL evidence.

  17. seonghobae commented on Oct 1, 2026

    @seonghobae
    ContributorAuthor

    Hosted exact-head proof for Draft .github#2548 at 8f6a870bca516d34a737bc8bd4abf5d3573573ae: CodeQL PR run 36830753593 materialized Detect CodeQL languages successfully, created both CodeQL compatibility analysis (python|actions) shards, and completed Dispatch current-head CodeQL scan successfully. Both shards then failed closed at the expected pending-verdict boundary. This is direct evidence that the owner repair removes the permanent Draft skip without manufacturing terminal success. Terminal authenticated dispatch verdict/rerun and a consumer ruleset canary are still required before closure.

  18. seonghobae commented on Oct 1, 2026

    @seonghobae
    ContributorAuthor

    2026-10-01 Orgmetra #160 unchanged-head Ready canary

    • Consumer PR: ContextualWisdomLab/Orgmetra#160
    • Exact head: 34ef29f76e46289515afeabad0bdec0357f61485
    • Lifecycle: Draft → Ready once, after exact-head Foundation 36847994756, Security 36847994744, and SAST 36847994839 were terminal-success and every review thread was resolved.
    • Pre-transition CodeQL 36847994749 is terminal skipped; it is not GREEN evidence.
    • Immediate post-transition re-read returned no new CodeQL run for the unchanged head. No manual rerun, no-op commit, Draft/Ready toggle loop, or synthetic status was used.
    • Merge remains HOLD: exact-head CodeQL acceptance is absent and qualifying APPROVED reviews remain 0.

    This is another consumer canary for the central event-materialization contract; repair remains owned by the active .github writer stack rather than Orgmetra.

  19. seonghobae commented on Oct 1, 2026

    @seonghobae
    ContributorAuthor

    semantic-data-portal #82/#107: independently reproduced central CodeQL historical-configuration mismatch (2026-10-01 UTC / 2026-10-02 KST).

    Current central main 37b1024 retains the same codeql_ghas_configuration_identity.py and dispatch workflow as failed producer source 3295c25.
    Consumer #82 head 53805c8df26e84520fcb214aa418caa6c6b3ca88, #107 head 6ace6aff7998f1ed15ca1507bfb25139358e1fea; protected base e48aa13c4af7a4875d4b53e6a60b50405c265a2f. Failed central producers: 36498724870 / 36499911224; bound required runs: 36443229889 / 36444353130.

    Offline replay of authenticated analysis snapshots with the current-main pairing_ready function reproduces all four language/PR failures: the sole absent identity is dynamic/github-code-scanning/codeql:upload /language:python or /language:actions. Actual Default Setup :analyze identities exist on exact base and both exact heads. Historical base uploads 1592375545/1592375546 are from August 9; their union with active configurations causes permanent failure after 30 attempts. Analysis-read credential works; this is not a 403 or missing-language diagnosis. Analysis/SARIF gates succeeded but do not substitute for required terminal proof.

    Parent independently exercised 8 offline regression assertions: four exact-snapshot reproductions, two missing-active-Default-Setup fail-closed controls, two identical-identity controls; 8 passed. No historical evidence deleted, upload fabricated, status synthesized, protection changed, dispatch/rerun performed, or approval claimed.

    Canonical repair needs a tested trusted active-configuration lifecycle policy that preserves real base/head pairing and historical evidence. Existing Draft-materialization and wake/permission repairs do not repair this failure. After normal protected central merge, recovery needs a fresh authorized bot producer consuming new main (not an old-source rerun). I continue the independent consumer dependency-gate repair locally; this receipt identifies the central repair seam and should be reconciled with the active central writer before edits overlap.

  20. seonghobae commented on Oct 1, 2026

    @seonghobae
    ContributorAuthor

    semantic-data-portal consumer repair update (not merge acceptance): #107 now has substantive dependency repair head 147d00d1135ffdc4a098fc126c064f321ecb1858 and made one ordinary Ready transition after local exact-content independent review. Its synchronize event was emitted while Draft: CodeQL PR 36909953510 has skipped language/shard/dispatch jobs; OpenCode and Strix are also Draft-skipped. No fresh required CodeQL identity was observed on the current head during one scoped read after Ready. Native API/Analyze and Security Scan results do not replace these central gates; independent APPROVED remains absent. #82 policy-guidance head69a73bc6f012a9b7c230fcaba4f5dbbaad380646 remains Draft until its latest PyJWT repair is independently reviewed.

    The existing historical-upload GHAS pairing RCA remains separate from Draft materialization. A newly identified Python3.10 dev-lock regression is being repaired with a real follow-up commit in both consumer lanes, not an empty wake commit; #107's follow-up will synchronize an already-Ready PR. After central identity-policy repair is normally merged, require the authenticated fresh-source producer with exact new heads/base/jobs. No state cycling, synthetic status, admin merge, self approval or force push. This receipt updates the current consumer handles for the canonical owner; it does not transfer old-head results.

  21. seonghobae commented on Oct 1, 2026

    @seonghobae
    ContributorAuthor

    Final consumer dependency repairs pushed normally after separate-baseline independent local review and actual native Python3.10 hash-required installs:

    • semantic-data-portal#82 head93321a9833a35d7611c775e7724b6fc3b6ed64da, now Ready once. Hosted API36914699600 and trivy-fs job110545769414 succeeded; initial Draft synchronize CodeQL36914699585 remains skipped, not acceptance.
    • semantic-data-portal#107 headb802cfbaa1c2fdb505d8647d706fa1e2e8c3050e, already Ready at its genuine Python3.10 corrective synchronize. Hosted API36914707751 and trivy-fs job110545796968 succeeded. Required CodeQL detect was admitted (not Draft-skipped) and Strix scan started in the fresh snapshot.
      Both independent GitHub approvals still absent. Security findings were repaired, not ignored; local native suites294/291,8skips and24cross-target resolver checks passed. Original dirty evidence preserved.
      Canonical historical-upload identity RCA remains; I am exercising a scratch-only lifecycle candidate with negative controls, without editing central owner checkouts or altering gates. Recovery/merge requires normal protected canonical repair plus authenticated fresh producer and exact-head independent approval. No bypass, synthetic result, empty wake commit or Draft/Ready cycling.
  22. seonghobae commented on Oct 1, 2026

    @seonghobae
    ContributorAuthor

    Scratch-only historical CodeQL lifecycle candidate independently replayed by parent: 76 tests pass (47 new lifecycle cases, 21 unchanged central identity tests, 8 original RCA controls). Existing source was RED on four captured historical PR/language pairs. No production predicate, CLI or workflow has changed.

    Safe candidate only models centrally reviewed retirement of exact analysis IDs1592375545/1592375546 bound to exact repository/baseSHA/ref/setup, retaining original history, unknown/duplicate uploads and all active identities. It is UNAPPROVED and not a blanket :upload filter. Wrong head/ref, missing/failed active configuration, malformed or changed setup and unproved pagination remain fail-closed. Timestamp is consistency evidence, not retirement authority.

    Production gaps remain explicit: authoritative retirement approval and native GHAS convergence; complete pagination (captured base is exactly100rows); fresh trusted setup/base/head evidence. Parent confirmed unchanged configured actions/python extended setup by authenticated GET; full analyses retrieval hit rate limit and remains unverified. No new credentials or permission widening.

    Fresh current Noema jobs82/107 (110545764138/110545816498) now fail HTTP400 after working sidecar preflight, served_model meta/llama-3.2-90b-vision-instruct. Earlier provider429 probes are not the final response. Strix107 job110545792629 also failed; direct log has no causal error text. I am performing scoped artifact/source RCA rather than rerunning unchanged jobs or manufacturing success. Consumer heads remain93321a9/b802cfb, hosted source tests/security/fuzz pass but no independent APPROVED. No merge claimed.

    Please reconcile lifecycle-policy ownership with the canonical validation writer before promoting this scratch seam. Local candidate intentionally does not satisfy the required check and cannot be used to bypass it.

  23. seonghobae commented on Oct 1, 2026

    @seonghobae
    ContributorAuthor

    Current-head GHAS evidence gap closed after rate-limit reset, using the existing authenticated identity and bounded gh api --paginate --slurp GETs (no uploads/deletes/config changes):

    • protected mainref yielded50 CodeQLanalysis rows,18 exactbasee48aa13 rows;
    • pull82headref yielded10 rows,2 exact93321a9 rows;
    • pull107headref yielded6 rows,2 exactb802cfb rows.
      All terminal page responses collected; analysisIDs deduplicated. Both new exact heads have successful native :analyze actions/python. The protected exact base still contains distinct historical :upload actions/python identities. Executing the unchanged identity predicate against these fresh complete captures reproduces four current-head missing-upload failures; five parent assertions passed including IDuniqueness. This is no longer only an older-head replay, but it remains a failure proof, NOT a passing lifecycle verdict.

    DefaultSetup remains configured actions/python extended with unchanged captured setup. Retirement of old uploads still has no canonical protected policy approval or native convergence proof, so no lifecycle filtering was enabled. The scratch candidate does not weaken the required gate. Current analyses are complete for the enumerated refs at this observation; future live settlement must re-read/bind freshness. No token rotation, extra permissions, synthetic status, rerun or dispatch performed.

    Fresh Noema/Strix artifact RCA continues independently; consumer source/security corrections and original dirty files remain preserved. Canonical owner can use this complete current-head specimen to decide and test lifecycle policy without transferring old-head results.

  24. seonghobae commented on Oct 1, 2026

    @seonghobae
    ContributorAuthor

    Actual artifact RCA now independently replayed on current semantic-data-portal heads:
    Strix107 artifact11189203580/run36914702651 reports scanner rc0/completed but an unrelated generic report with no changed source path; zeroSARIF results does not establish source review. Replaying exact artifact with trusted strix_report_scope.py fails exit1 at missing changed-source scope; six taxonomy/artifact controls pass. Keep this gate intact. Repair actual source-grounded inspection/report generation or governed regeneration; do not insert paths to simulate review.
    Noema82/107 artifacts11190315068/11189803045 show successful plain-chat admission then7primary11B+2secondary11B+1primary90B attempts in one request. Terminal90B HTTP400 is nonretryable invalid_request_error and returns within~1.1/.6seconds, not a multi-minute90B timeout. Aggregate362.7/227seconds includes preceding work. Main request has strictjson_schema absent fromplain preflight. Exact rejected field is not preserved, so schema dialect/message ordering/context limit remain hypotheses, not confirmed cause. Preserve privacy redaction, strictverdict/exacthead checks and gatewayone-request policy.
    Parent is implementing a separate, narrow complete-analyses pagination repair in isolated central worktree to remove the proven first-page completeness flaw. This changes retrieval safety only; no historical retirement, no identity filtering or lifecycle-policy adoption. All open overlapping repair scope must be checked before promotion and ordinary independent review/protectedmerge required. No credentials, deployments, gateweakening or old-head result transfer.

  25. seonghobae commented on Oct 1, 2026

    @seonghobae
    ContributorAuthor

    Published the bounded GHAS complete-history retrieval repair as #2555, exact head 331512f44460e98852c4028980af0f73d8c046b0, against protected main base 37b10243cec3d160ecc9c1be75c71428b160a703.

    The implementation follows every validated Link page, rejects incomplete terminal-last evidence, combines repeated physical Link fields before validation, and binds bearer requests to canonical GitHub authority and the original repository/ref/tool/query. Existing configuration identity selection, pairing, retirement policy and redirect prohibition are unchanged.

    Evidence: 234 affected tests passed; helper 219 statements/82 branches and docstrings all 100%. The committed head also passed all 234 affected tests with GITHUB_ACTIONS=true. macOS workflow fixtures use a scratch-only Bash 5 substitution; these are not hosted gate results. Independent local full-diff review passed at diff SHA-256 cf5a1b2c643b0a30b4d5476473f6170c5d4305912b6de9c4d33db10aae14f7fc; it is NOT GitHub formal approval.

    Fresh #2555 readback confirms the published exact head and active required review/security workflows. They are pending, not passed. CodeRabbit's rate-limited status and Devin's expired-trial status do not constitute actual reviews.

    Consumer heads remain unchanged and unmerged: semantic-data-portal#82 93321a9833a35d7611c775e7724b6fc3b6ed64da and #107 b802cfbaa1c2fdb505d8647d706fa1e2e8c3050e, both REVIEW_REQUIRED without exact-head approval. This retrieval prerequisite does not by itself resolve historical upload/analyze lifecycle convergence, Noema HTTP 400 or the other hosted review failures. No retirement policy, credential changes, deployments or protection bypass were applied.

  26. seonghobae commented on Oct 2, 2026

    @seonghobae
    ContributorAuthor

    New source adoption and actual hosted acceptance delta for the existing canonical validation owner; no duplicate producer request.

    #2555 is now Ready at exact 1cda48bf9eb3aaf2eff1387c0cbceb03bd73e45a, protected base 37b10243cec3d160ecc9c1be75c71428b160a703. The existing sole writer normally integrated the existing #2530 prerequisites, #2558 pypdf security correction and retained independently reviewed HTTPError real-close test fixture correction. Complete96file substantive independent source verdict PASS binds tree 44d8fb1f8786cb31b75bb5326503fbc89b6e47ee and full-index patch SHA256 c18009e7d150a26aa0400a2e858284e4f43d51be7d62e67d789c5edd0bd66c5a; old negative fixture verdict preserved. Normal remote ref and fresh PR head were both verified; no peer branch, protected main or original semantic-data-portal dirty6 was edited.

    Fresh hosted exact-head results: Python3.14 full quality completed with5427passed,5existing skips,40subtests;18830statements/7684branches100%coverage and100%docstrings. Python3.10compatibility/runtimequality/quality/validate passed. All11 actual pip-audit input results report no known vulnerabilities; Trivy's logged expected/actualhead both match1cda48bf and actualCRITICAL/HIGH/MEDIUM findings0. osv-scan/dependency-review/Bandit/Semgrep/gitleaks are terminalSUCCESS. These are executed results, not scanner exemptions or local verdict substitution.

    The newly recovered remaining execution boundary is runner admission before producer work, not a demonstrated source/SARIF finding:

    • CodeQL producer run37037785105 has exact run-name target2555/head1cda48bf/base37b10243/requiredrun37037648338/merge-source77e2d4f8171636cab28c880b1695d4130e0c9b30. validate-dispatch job110940160534 is queued, runner_id0, steps[], labels[self-hosted,linux,x64]. Required actions/python jobs110939787274/110939787772 correctly failclosed withDISPATCH_OUTCOME=success/VERDICT_STATE=pending.
    • OpenCode producer run37037756733 exact target2555/head1cda48bf: validate-pr-metadata job110940076787 queued, runner_id0, steps[], same labels. Required opencode-review job110939958247 has no authenticated exact-head APPROVED/CHANGES_REQUESTED yet.

    Please use the existing execution/capacity owner to advance these already-enqueued producer jobs and their normal terminal publication/required-job settlement; do not launch another scanner/reviewer, duplicate dispatch, change credentials or weaken admission/gates. An API started_at on these queued rows does not establish an executed step. The existing run/required-job binding above is the bounded actionable target.

    Noema/Strix remain in progress. Fresh Devin current-head review is COMMENTED/no issues, not countedAPPROVED; no exact-head qualifying approval or protectedmerge is claimed. Consumer semantic-data-portal#82/#107 and their historical upload/analyze lifecycle convergence remain unfinished, and#108 remains after both normal merges only. The producer boundary is recorded once with this new material delta; notification is not owner acceptance or recovery proof.

  27. seonghobae commented on Oct 3, 2026

    @seonghobae
    ContributorAuthor

    Fresh early-stage consumer evidence — ELUNVERA PR #1 exact head 41219e0cf112688852875376b00b8fa3917d5b55 (2026-10-03)

    Acceptance remains a recoverable runner-start path that materializes jobs/steps/logs and produces exact-head terminal evidence.

  28. seonghobae commented on Oct 5, 2026

    @seonghobae
    ContributorAuthor

    Exact-head execution delta — 2026-10-05, observation through 09:54 KST

    PR #2577 is published at c616da7c51113405c508f8957010ee03888e3d3b, base 37b10243cec3d160ecc9c1be75c71428b160a703. Publication is not CI acceptance, qualifying approval or protected integration. The previously completed source/quality results are retained rather than rerun.

    Existing OpenCode producer failed before model review

    • Required consumer run 37246380517, job 111565110213: request step SUCCESS, immediate exact-head verdict check FAILURE. The inspected protected-base/head receipt paths are byte-identical. This sequence is compatible with the intentional asynchronous handoff: a genuine downstream formal receipt precedes the existing required-job wake. It is not, by itself, a new source defect or justification for another dispatch.
    • Existing repository-dispatch run 37246461390 is bound by its immutable title to ContextualWisdomLab/.github#2577@c616da7c51113405c508f8957010ee03888e3d3b; its top-level control revision is not target-head evidence. Job 111565255052 passed live metadata binding and merge-tree materialization, then failed Upload materialized pull request merge tree with Artifact storage quota. Coverage 111565329348 and model-review 111565329380 were SKIPPED with zero steps. No formal review/wake executed.
    • Earlier same-target run 37246452437 was CANCELLED during materialization. Cancellation actor/trigger remains unknown. The current observation searched the most recent 100 control-head suites out of 3,499, not an exhaustive organization inventory; both selected runs' job/step/annotation connections were complete.

    Noema has two distinct failures

    • Existing run 37246380405, job 111565062211, binds to the same target head. Sidecar provisioning succeeded. Prepare Noema model verdict failed with an executed HTTPError 400 in phase=response_error; Upload contextual-orchestrator sidecar evidence separately failed with Artifact storage quota/Failed to CreateArtifact. Verdict publication was skipped.
    • The six actual annotation digests were reconciled across independent projections. Provider-attempt count, upstream status and terminal-reason fields are absent from the warning/error; deeper request/provider cause remains unknown. HTTP400 is outside the existing capacity continuation set {429,500,502,503,504}; no retry eligibility is invented.
    • A five-case offline exact-selected-function discriminator admitted Noema's unchanged JSON-schema envelopes (probe floors one/two) through the gateway default source pin's format validator and rejected three malformed controls. This does not attest the historical serving revision, full routing/provider behavior or a repair of the hosted400. No schema/probe/free/ZDR gate was weakened.

    Existing owner boundaries and next transition

    The original Coordinator received the new target-bound artifact-transport failure and separate Noema400 evidence; registration is not management approval or execution. The existing source owner was asked once for an already-retained safe exact-run capture locator, not a new model run. The existing artifact-capacity/retention custodian and existing scheduler/dispatch owner remain responsible for their respective operational transitions. No artifact deletion, retention/payment/credential/runner-ACL/service change, manual dispatch/rerun, selfapproval or protected-gate bypass was performed.

    Related to #2560 and #2565. Their original single writers and unfinished acceptance remain intact. Keep this issue OPEN until genuine current-head producer execution, substantive review/security verdicts and required-consumer settlement satisfy its original acceptance. A skipped, cancelled, quota-rejected or synthetic result is not a green review.

  29. seonghobae commented on Oct 9, 2026

    @seonghobae
    ContributorAuthor

    Fresh exact-head central-validation specimen from #2283 at e7eacd952847fc149851a40ef50f61d90cb4e12c (tree b3bf209eaaeebe3714c89d1426f65073292d3195): required bootstrap, exact-head admission, cancellation, coverage source-tree, and coverage evidence are terminal success. Noema job 113733734203 passed credential/live-head/draft checks, correctly skipped model provisioning because the PR is Draft, then failed at actions/upload-artifact@043fb46… with Failed to CreateArtifact: Artifact storage quota has been hit; quota usage is recalculated every 6–12 hours. This is an organization storage-capacity blocker, not a model/source verdict. The workflow’s existing contract intentionally uploads sanitized evidence after any outcome, so the failure was not masked or blindly rerun. Dynamic GitHub Code Quality job 113732888604 failed before steps materialized (steps=[]) and its log endpoint returns no artifact, so it remains a typed pre-run/platform failure pending causal evidence. Local exact-tree verification remains GREEN: 204 focused contracts in both normal and GITHUB_ACTIONS=true environments; full warnings-as-errors suite 5186 passed, 11 skipped, 40 subtests passed. #2283 remains Draft with 0 approvals and 0 unresolved threads.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area: ci-cdCI, GitHub Actions, checks, release, or supply chainarea: securitySecurity boundary, hardening, or vulnerability preventionbugSomething isn't workingmaintenancepriority: highHigh-priority or P1 workstatus: blockedBlocked by conflict, dependency, or required prerequisite

    Type

    No type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions