Skip to content

feat: cw 1.0.1 stall ladder, lean loop and session notifier - #4

Open
victorkzam wants to merge 9 commits into
mainfrom
feat/stall-handling-1-0-1
Open

victorkzam wants to merge 9 commits into
mainfrom
feat/stall-handling-1-0-1

Conversation

@victorkzam

@victorkzam victorkzam commented Sep 22, 2026 •

Copy link
Copy Markdown
Owner

Summary

  • Stall prevention in the agents. reviewer, design-reviewer and direction-reviewer drop memory: project (fresh context is their point; memory writes cost turns). Reviewer turn caps are sized to measured need (40 for deep review, 15 for direction), both reviewer schemas gain an unchecked list so a partial report is structured rather than a stall, and every agent's Output section is a bounded contract: full report to the file the prompt names, final message of at most about 1,500 tokens.
  • Stall ladder with one home. rules/orchestration.md replaces the "retry once fresh" bullet with the ladder: ground truth first (git log and the verification command for an implementer, the report file for a reviewer), two SendMessage nudges to the same agent, one briefed respawn, then escalate. The build and design skills point to it in one line each.
  • Lean orchestration loop (plan C). /cw:design gains the lite path, a two-round review cap, delegated synthesis research and bounded briefs. /cw:build gains bounded briefs, commit by pathspec, and the stranded-implementer gate.
  • Session-labelled desktop notifier. hooks/notify.sh on Notification, Stop and StopFailure: one ping titled <folder> · <session> only when a human is needed; the built-in idle_prompt is suppressed while background agents still run; presence policy ships behind a config flag, off by default; opt-in via a config file, silent under the SDK entrypoint and CW_NOTIFY=0.
  • Docs and release. README, WORKFLOW and the design draft describe 1.0.1; plugin version bumped to 1.0.1 with the marketplace description.
  • Test fix (added after the first CI run). The four branch-guard speed cases in tests/run.sh judged the hook with bash's integer SECONDS counter against a budget of 1 tick; the heaviest payloads take about 0.65-0.75 s locally, and the macOS runner read 2 ticks on case 85 with the correct exit code, failing the first pull_request run while the push run on the same commit passed. The budget is now 2 ticks (under 3 s of wall clock) with one retry on a timing-only miss; a wrong exit code still fails at once.

Why: three measured problems since 1.0.0, recorded in the plan's Context section. Deep-review agents ended without a verdict in 4 of 4 and then 2 of 4 runs, while a single resumed-turn nudge recovered every one. One four-phase session cost about 1.6M sub-agent tokens plus a full orchestrator window. The notification hook produced 110 banners in a day, most of them idle_prompt while agents were still running, with no session label.

Test Plan

  • bash tests/run.sh all on e4b797c: hooks 154 passed, 0 failed; size 64824 bytes (87.9% of baseline); dedupe, settings-keys, trailers and scan all PASS. shellcheck -x hooks/*.sh tests/run.sh clean. /bin/bash tests/run.sh hooks under macOS system bash 3.2: 154 passed.
  • Speed-case helper exercised with fake hooks under bash 3.2: fast hook passes first try; wrong exit code fails on attempt 1; a 3.2 s hook is invoked twice and reported as a timing failure with attempt=2.
  • Hands-on check: run a deep cw:reviewer review under the new turn cap and confirm a verdict lands, or that nudge 1 of the ladder recovers it in one turn.
  • Hands-on check: with the notifier enabled, end a turn while a background agent is running and confirm no idle_prompt banner until the agent finishes; confirm the banner title reads <folder> · <session>.
  • Hands-on check: /cw:design lite <description> runs a single review round and skips the direction pass.

Design Reference

docs/plans/stall-handling-1-0-1.md

🤖 Generated with Claude Code

https://claude.ai/code/session_01STufi6776QVKvuDcJv4RPM

victorkzam and others added 9 commits September 22, 2026 21:52
Co-Authored-By: Claude <noreply@anthropic.com>
… bounded reports

Co-Authored-By: Claude <noreply@anthropic.com>
… design path

Co-Authored-By: Claude <noreply@anthropic.com>
…d briefs

Co-Authored-By: Claude <noreply@anthropic.com>
… gate and the ladder pointer

Co-Authored-By: Claude <noreply@anthropic.com>
…Failure with harness cases

Co-Authored-By: Claude <noreply@anthropic.com>
…, WORKFLOW and the draft

Co-Authored-By: Claude <noreply@anthropic.com>
Co-Authored-By: Claude <noreply@anthropic.com>
…ming miss once

The four speed cases judged the branch guard with bash's integer SECONDS
counter against a budget of 1, so a healthy run of about 0.65-0.75 s (measured
locally under /bin/bash 3.2 on the 5000-segment and 12000-token payloads) had
no headroom on a slower runner: the macOS job read 2 ticks with the right exit
code on case 85 and failed the pull_request run on 2026-09-22, while the push
run on the same commit passed. The cases exist to catch pathological
backtracking, not to hold a latency target, so the budget is now 2 ticks
(under 3 s of wall clock) and a timing-only miss with the expected exit code
gets one retry. A wrong exit code still fails on the first attempt.

Co-Authored-By: Claude <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant