Skip to content

Harden safety gates: end-to-end attack tests, eval in CI, coverage floors - #18

Merged
Eilen6316 merged 2 commits into
masterfrom
feature/harden-safety-gates
Sep 19, 2026
Merged

Eilen6316 merged 2 commits into
masterfrom
feature/harden-safety-gates

Conversation

@Eilen6316

Copy link
Copy Markdown
Owner

摘要

关闭深度审查中发现的几个安全缺口:安全保证此前要么被测浅了一层,要么根本没进 CI。

第一优先 · 端到端攻击拦截 harness 测试

新增 5 个场景(tests/harness/scenarios/attack_*.yaml)驱动完整 LangGraph 流程,断言混淆/包装型破坏命令(bash -c、sudo、env)、敏感读取、以及一个字符混淆的 LOLBin 被端到端拦截或拒绝,并带审计断言。此前 red-team 只孤立地断言了策略分类。

第二优先 a · eval 回归门禁

CI 新增独立 job 跑 make eval;把"无录制时静默 skip"改为硬失败,使 prompt 行为回归真正被捕获而非空过。

第二优先 b · per-module 覆盖率下限

scripts/check_coverage_floors.py 对安全关键模块强制下限(核心逻辑 70% / 平台运行时 48%),接入 make test 与 CI,防止单个模块躲在全局 80% 平均值背后。

第二优先 c · SSRF 加固

补充 DNS rebinding TOCTOU(连接已验证 IP、单次解析)、编码 IP 字面量形式、0.0.0.0、IPv4-mapped 元数据、userinfo URL 的红队测试。

第三优先 · 配置清理

删除未启用的 ruff PLR2004 忽略项与冗余的 coverage omit=[]。

附加 · 修复存量 integration 失败

test_agent_runtime_stops_repairing_failed_command_plan_at_limit 写于停滞检测特性之前;它对每次修复都提议相同的失败命令,触发停滞检测在首次修复后就停转。本 PR 对该用例关闭停滞检测,让 max_repair_attempts 上限成为终止条件(停滞检测本身有专门的单元测试覆盖)。

…oors

Closes gaps found in a deep review where safety guarantees were tested one
layer too shallow or not gated in CI at all.

Tier 1 - end-to-end attack-blocked harness tests
  (tests/harness/scenarios/attack_*.yaml): 5 scenarios drive the full
  LangGraph flow and assert obfuscated/wrapped destructive commands
  (bash -c, sudo, env), sensitive reads, and an obfuscated LOLBin are blocked
  or refused end-to-end, with audit assertions. Previously red-team only
  asserted policy classification in isolation.

Tier 2a - eval regression gate (.github/workflows/ci.yml, tests/eval):
  add a dedicated CI job running `make eval`, and convert the silent
  skip-on-missing-recordings into a hard failure so prompt-behavior
  regressions are actually caught instead of passing vacuously.

Tier 2b - per-module coverage floors (scripts/check_coverage_floors.py,
  Makefile, ci.yml): enforce per-module floors on security-critical modules
  (70% core logic / 48% platform runtime), wired into `make test` and CI, so
  a single module cannot hide behind the global 80% average.

Tier 2c - SSRF hardening (tests/red_team/test_network_fetch_ssrf.py):
  add DNS-rebinding TOCTOU (connect to the validated IP, single resolution),
  encoded IP literal forms, 0.0.0.0, IPv4-mapped metadata, and userinfo URLs.

Tier 3 - config cleanup (pyproject.toml): drop dead ruff PLR2004 ignore (PL
  ruleset not enabled) and redundant coverage omit=[].
test_agent_runtime_stops_repairing_failed_command_plan_at_limit predates the
repair stall-detection feature (added in the failure-signature commits). It
proposes the identical failing command on every repair, so stall detection now
stops the loop after the first repair and analyze consumes a plan response
instead of the analysis, leaving the raw plan JSON as the final message.

Disable stall detection for this case so the max_repair_attempts limit is what
terminates the loop, which is what this test is named for. Stall detection has
its own dedicated unit coverage.
@Eilen6316
Eilen6316 merged commit f66c04d into master Sep 19, 2026
24 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant