Skip to content

fix(strix): record cached sandbox lifecycle diagnostics - #2573

Open
seonghobae wants to merge 55 commits into
fix/2261-inherited-quality-gatesfrom
fix/strix-sandbox-state-diagnostics-aa11d478
Open

seonghobae wants to merge 55 commits into
fix/2261-inherited-quality-gatesfrom
fix/strix-sandbox-state-diagnostics-aa11d478

Conversation

@seonghobae

Copy link
Copy Markdown
Contributor

RCA boundary

Current-head runs #2261 and #2550 passed gateway health, provider-route, and chat preflight, then failed guest authentication against in-container Caido at 127.0.0.1:48080. The evidence identifies a sandbox-bootstrap failure, not provider 429 or an LLM failure. Historical logs do not contain an OCI image digest or lifecycle state, so they cannot distinguish slow startup from a terminal process failure.

Change

The trusted Strix launcher now records only already-cached, SDK-bound Docker metadata around guest authentication: phase, freshness=cached, image configuration ID, lifecycle status, exit code, and OOM flag. It does not reload Docker, query the daemon, read logs, look up images, change authentication, skip terminal containers, alter retries/deadlines/backend selection, or call settings loading. Unsupported SDK/ABI bindings emit a fixed unavailable marker and leave the scan behavior unchanged.

Verification

  • 43 focused tests passed twice, including original token parsing, exact exception/cancellation identity, non-Docker non-interference, malformed metadata redaction, and no reload/remove behavior.
  • scripts/ci/strix_timeout_compat.py: statement·branch coverage 100%; docstrings 100%; diff check passed.
  • The fixture preserves the hashed pinned Strix 1.5.3 guest-authentication function and its Apache-2.0 attribution.
  • Current OCI tag metadata is documented as current-tag evidence only, not historical execution identity or root-cause proof.

This diagnostic does not claim Caido recovery, a green Strix scan, hosted acceptance, or merge authorization.

Developer experience: future failed scans emit bounded lifecycle context instead of forcing operators to infer sandbox state from a connection refusal.
User experience: incomplete sandbox scans remain failed evidence; no zero-finding result is promoted to a security pass.

🤖 Generated with Claude Code

seonghobae and others added 30 commits September 18, 2026 08:56
fast-mlsirm 0.11.3's published PyPI description carried the internal commercial
boundary: an enterprise sales gate, a KRW 2,000,000,000 product gate, buyer and
procurement evidence links, and 16 repo-relative links that are 404s on the
registry page because only LICENSE and the package sources ship. Checking the
live sdist finds 68 such findings; threadweave 0.1.0 and rankweave 0.1.0 carry
3 each, one of them a source module path.

The gate reads the description a registry renders - PKG-INFO from an sdist,
METADATA from a wheel - rather than the README on disk, because a release is
built from a commit and packaging config decides what is included. A repository
with no release yet can point it at the README instead.

A repository README may link internal design records; that is public
development history. The same text on a package page is different, because only
the distribution's files exist there and the reader is installing rather than
developing. So docs/adr links stay legitimate and only have to be absolute,
while docs/superpowers, docs/product, docs/commercial, docs/planning and
docs/doctoring are findings. Hard-coded deal values follow the rule the
organization already set in fast-mlsirm's acquisition_readiness_gate doctoring
record: product quality evidence must not depend on a monetary target.

Product vocabulary that merely looks commercial is deliberately not a finding.
appguardrail ships a real buyer-diligence subcommand and scopeweave really does
check procurement packages; the rule targets internal framing, not a domain.

--allow lets a repository adopt the gate before its README is fully converted
instead of landing a red check it cannot fix in one PR.

No repository calls this yet, so this PR cannot break one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YHVBDaZS5NZT9aQcbRg9Av
The first run against the three repositories that had just been corrected
produced 16, 1 and 2 findings, and nearly all of them were wrong.
contextual-orchestrator genuinely ships /api/v1/commercial_readiness/latest,
/api/v1/saleability_decisions/latest and /api/v1/commercial_due_diligence_rooms/latest,
with tests named after them; wardnet's crate genuinely computes commercial
readiness snapshots; semantic-data-portal's only finding was the sentence
explaining that PRD/TRD records are excluded. The exception for product domain
vocabulary was stated in the prose and absent from the regex.

So go-to-market-vocabulary and requirement-map now report without failing, and
--strict promotes them for a repository that wants them enforced. The
mechanical rules keep blocking because they need no judgement: a relative link
is dead on the registry page, a quoted module path is plumbing, and a
hard-coded deal value is never a product feature.

The same pass showed ADRs filed under docs/planning/adrs/ being caught by the
docs/planning/ prefix. An ADR is a public design record wherever a repository
files it, so the path check exempts it and a test pins both directions.

Verified after the change: the live fast_mlsirm-0.11.3.tar.gz still reports 22
blocking findings and 63 under --strict, the corrected fast-mlsirm README is
clean, and the three corrected READMEs pass with advisory notes only.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YHVBDaZS5NZT9aQcbRg9Av
pg-llm-batch's README links docs/doctoring/bootstrap-dsn-precedence.md,
cli-secret-input.md and postgres-logical-restore.md, and those are operator
documentation a package user needs. That repository files operational guidance
under the same directory name this one uses for incident records. A worker
removed the links to satisfy the gate and restored them in the next commit
because they were legitimate - the gate causing damage rather than preventing
it.

So internal-working-record advises. Nothing is lost: a link into such a
directory that is also repo-relative is still blocked by relative-link, which
is the mechanical defect, since the page cannot resolve it. An absolute URL to
the same file resolves, and whether that audience wants it is judgement.

That is now the line across both corrections in this branch: block only what is
broken regardless of context, advise on anything that needs to know what the
product is.

Re-verified: the live fast_mlsirm-0.11.3.tar.gz still reports 22 blocking
findings, OriginWeave's seven repo-relative links still block, pg-llm-batch's
corrected README passes with 18 advisory notes, and the corrected fast-mlsirm
README is clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YHVBDaZS5NZT9aQcbRg9Av
The first revision of this reusable workflow took a free-form `build-command`
string input and interpolated it straight into a `run:` block. A caller could
have passed arbitrary shell into a workflow running in its own repository
context, and it is what ADR 0023 forbids outright: reusable inputs are data and
capability flags, not shell source.

The contract test for the fast-mlsirm reusable workflow in this same branch
series asserts that no input reaches a `run:` body. I did not apply that rule
here and did not notice until the org's own Semgrep gate failed the PR with
yaml.github-actions.security.run-shell-injection. The gate was right.

The input is now `build: sdist | wheel | none`, mapped to a fixed `python -m
build` invocation and validated in-shell so an unexpected value exits 2 rather
than silently building nothing. `dist-path` is passed through env to that step
instead of being read from a scope where it was undefined. A contract test pins
the absence of `build-command` and the absence of any `${{ inputs.* }}` inside
any `run:` body.

Local Semgrep on this file after the change: 0 findings. Contract tests: 19
passed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YHVBDaZS5NZT9aQcbRg9Av
Use the called job's workflow repository and exact workflow SHA for the central checkout, remove the unhashed runtime build install in favor of a pinned uv action and version, validate every sdist/wheel description in an upload set, and reject release-contract links to moving GitHub branches.

Regression tests cover workflow identity, immutable build tooling, mutable blob/tree links, and matching/divergent artifact metadata. The operator note now matches the implemented blocking/advisory policy.
Preserve the exact #2261 package-description gate delta while integrating protected main after #2279. The combined tree passes 117 focused owner contracts and the complete warnings-as-errors repository suite.
…e the consumer root

Green step for a8d6261. The 24 specialized cases in
test_strix_quick_gate.sh installed the trusted gate/model/binder into
$repo_root_dir/scripts/ci and ran ./scripts/ci/strix_quick_gate.sh, so a
consumer-root binder lookup could never fail there and masked the #2292
defect. Each case now materializes into
$tmp_dir/trusted-source/scripts/ci and runs the gate from that directory
with STRIX_REPO_ROOT=$repo_root_dir, which keeps the old repo-root
semantics (the gate defaults REPO_ROOT to SCRIPT_DIR/../..).

Evidence:
- tests/test_strix_trusted_fixture_boundary.py: fails on a8d6261 (CI
  job 106083294309), passes here.
- bash scripts/ci/test_strix_quick_gate.sh on Linux, umask 022:
  a8d6261 PASS (rc=0, 727s) and this commit PASS (rc=0, 726s).
- strix-related pytest (8 files): 242 passed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01P5o6j4zfxGPdRaH4Lug8UY
seonghobae and others added 23 commits September 27, 2026 02:03
…tack

Preserve the package-description boundary delta while adopting #2291,
including the canonical AnyIO, CodeQL, and Strix owner repairs.
Keep the registry-description gate, called-workflow checkout identity, and
the Job Analysis trusted-base context. The Strix harness keeps a single
trusted fixture helper that copies the report-scope module, and it runs
the gate with STRIX_REPO_ROOT pointed at the consumer workspace.
Co-Authored-By: Claude Code <noreply@anthropic.com>
Preserve every valid package-description delta from #2261 while ordinary-merging the current canonical security-owner repair. The reusable workflow now separates the untrusted caller build from the trusted inspector, transfers only a fixed nonhidden artifact input, requires successful producer completion, and runs the central gate under isolated Python. The inherited Maturin downloader no longer uses URL suppressions.

Exact local tree evidence: 34 focused package/readiness contracts passed; full predecessor integration suite passed 5,210 tests, 8 skipped, and 40 subtests; git diff --check passed. Independent re-review reported zero Critical, Important, or Minor findings. Fresh hosted exact-head evidence and qualifying approval remain mandatory.
Advance the package-description stack through an ordinary two-parent merge with the current #2531 owner. The package boundary retains its isolated untrusted producer/trusted inspector design while adopting the restored Gap baseline, Semgrep-clean one-hop downloader, and explicit ambient proxy/auth exclusion.

Exact-tree focused owner/package/readiness verification: 50 passed; git diff --check passed. Fresh hosted exact-head evidence and qualifying approval remain mandatory.
Co-Authored-By: Claude Code <noreply@anthropic.com>
Co-Authored-By: Claude Code <noreply@anthropic.com>
Co-Authored-By: Claude Code <noreply@anthropic.com>
Co-Authored-By: Claude Code <noreply@anthropic.com>
Co-Authored-By: Claude Code <noreply@anthropic.com>
Co-Authored-By: Claude Code <noreply@anthropic.com>
Co-Authored-By: Claude Code <noreply@anthropic.com>
Co-Authored-By: Claude Code <noreply@anthropic.com>
Co-Authored-By: Claude Code <noreply@anthropic.com>
Co-Authored-By: Claude Code <noreply@anthropic.com>
Co-Authored-By: Claude Code <noreply@anthropic.com>
Co-Authored-By: Claude Code <noreply@anthropic.com>
Co-Authored-By: Claude Code <noreply@anthropic.com>
Co-Authored-By: Claude Code <noreply@anthropic.com>
Co-Authored-By: Claude Code <noreply@anthropic.com>
Co-Authored-By: Claude Code <noreply@anthropic.com>
Co-Authored-By: Claude Code <noreply@anthropic.com>
Co-Authored-By: Claude Code <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration
  • Configuration used: Organization UI
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: 10673ceb-5fe2-4c48-9e34-921413d8f001

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

seonghobae and others added 2 commits October 4, 2026 13:19
Co-Authored-By: Claude Code <noreply@anthropic.com>
Co-Authored-By: Claude Code <noreply@anthropic.com>
@seonghobae

Copy link
Copy Markdown
Contributor Author

Current-head hosted admission RCA

Exact #2573 head 41e9c45407e3f72d138da2ad24cdf934aa42fbe1 has common initial failures across CodeQL, Semgrep, and Security detection jobs. These jobs started and ended in about four seconds with no steps or retained run logs. The retained GitHub Actions check-run annotations provide the cause:

The job was not started because your account is locked due to a billing issue.

Observed check-run ids: 111362130332 (CodeQL detect), 111362130039 (Semgrep), 111362130130 (Security detect). This is account-level Actions admission, not a source finding in #2573, #2572, #2571, or #2531. No source change, retry, status synthesis, gate relaxation, or state toggle can make an unadmitted job valid.

The local exact-tree verification remains recorded separately; it does not replace hosted acceptance. The actionable recovery is to resolve the GitHub account billing lock, then trigger normal current-head admission on the active successor stack. After real Actions jobs start, current-head security results and non-author formal approval remain required. This comment is evidence only, not approval or merge authorization.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant