Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions .claude-plugin/marketplace.json
Original file line number Diff line number Diff line change
Expand Up @@ -6,13 +6,13 @@
},
"metadata": {
"description": "The operations layer for agentic engineering: portable skills and evidence contracts connecting intent, coding agents, software factories, and independent judgment.",
"version": "3.6.0"
"version": "4.0.0"
},
"plugins": [
{
"name": "agentops",
"description": "The operations layer for agentic engineering: portable skills and evidence contracts connecting intent, coding agents, software factories, and independent judgment.",
"version": "3.6.0",
"version": "4.0.0",
"source": "./",
"author": {
"name": "Boden Fuller",
Expand Down
2 changes: 1 addition & 1 deletion .claude-plugin/plugin.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "agentops",
"version": "3.6.0",
"version": "4.0.0",
"description": "The operations layer for agentic engineering: portable skills and evidence contracts connecting intent, coding agents, software factories, and independent judgment.",
"author": {
"name": "Boden Fuller",
Expand Down
2 changes: 1 addition & 1 deletion .codex-plugin/plugin.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "agentops",
"version": "3.6.0",
"version": "4.0.0",
"description": "The operations layer for agentic engineering: portable skills and evidence contracts connecting intent, coding agents, software factories, and independent judgment.",
"skills": "./skills-codex",
"interface": {
Expand Down
8 changes: 6 additions & 2 deletions .github/workflows/nightly.yml
Original file line number Diff line number Diff line change
Expand Up @@ -84,8 +84,12 @@ jobs:

- name: Install scanner tools
run: |
sudo apt-get update -qq
sudo apt-get install -y shellcheck
python -m pip install --upgrade pip
python -m pip install semgrep==1.169.0 ruff==0.15.21 radon==6.0.1 pytest==9.1.1
python -m pip install semgrep==1.169.0 ruff==0.15.21 radon==6.0.1 pytest==9.1.1 \
-r evals/skills-rpi/requirements.txt \
-r evals/skills-rpi/requirements-readout.txt

GOBIN=/usr/local/bin go install github.com/securego/gosec/v2/cmd/gosec@v2.27.1
GOBIN=/usr/local/bin go install github.com/zricethezav/gitleaks/v8@v8.30.1
Expand Down Expand Up @@ -113,7 +117,7 @@ jobs:
- name: Run security gate (full)
run: |
chmod +x scripts/security-gate.sh
./scripts/security-gate.sh --mode full
./scripts/security-gate.sh --mode full --require-tools
env:
SECURITY_GATE_OUTPUT_DIR: ${{ runner.temp }}/agentops-security
TOOLCHAIN_OUTPUT_DIR: ${{ runner.temp }}/agentops-tooling
Expand Down
27 changes: 25 additions & 2 deletions .github/workflows/release.yml
Original file line number Diff line number Diff line change
Expand Up @@ -54,11 +54,34 @@ jobs:

- name: Install scanner tools
run: |
sudo apt-get update -qq
sudo apt-get install -y shellcheck
python -m pip install --upgrade pip
python -m pip install semgrep==1.169.0
python -m pip install semgrep==1.169.0 ruff==0.15.21 radon==6.0.1 pytest==9.1.1 \
-r evals/skills-rpi/requirements.txt \
-r evals/skills-rpi/requirements-readout.txt

GOBIN=/usr/local/bin go install github.com/securego/gosec/v2/cmd/gosec@v2.27.1
GOBIN=/usr/local/bin go install github.com/zricethezav/gitleaks/v8@v8.30.1

# golangci-lint
GOBIN=/usr/local/bin go install github.com/golangci/golangci-lint/v2/cmd/golangci-lint@v2.13.1

# Use the same scanner pins as the nightly full-security lane.
# govulncheck covers known-CVE reachability in dependencies and stdlib.
GOBIN=/usr/local/bin go install golang.org/x/vuln/cmd/govulncheck@v1.6.0

# trivy — pinned tag, download-then-execute. Piping a mutable-branch
# install script into sh was the semgrep gha-curl-pipe-shell CRITICAL
# that blocked the v3.2.0 publisher (supply-chain: a hijacked script
# would run unread).
curl -sSfL -o /tmp/trivy-install.sh https://raw.githubusercontent.com/aquasecurity/trivy/v0.72.0/contrib/install.sh
sh /tmp/trivy-install.sh -b /usr/local/bin v0.72.0

# hadolint
curl -sSfL -o /usr/local/bin/hadolint "https://github.com/hadolint/hadolint/releases/download/v2.14.0/hadolint-Linux-x86_64"
chmod +x /usr/local/bin/hadolint

- name: Generate pre-publish release evidence
run: |
mkdir -p release-artifacts
Expand All @@ -71,7 +94,7 @@ jobs:
# but fail before GoReleaser when the gate returns findings.
chmod +x scripts/security-gate.sh
set +e
./scripts/security-gate.sh --mode full --json > release-artifacts/security-gate-summary.json
./scripts/security-gate.sh --mode full --json --require-tools > release-artifacts/security-gate-summary.json
SECURITY_RC=$?
set -e
jq empty release-artifacts/security-gate-summary.json
Expand Down
101 changes: 87 additions & 14 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,24 +7,97 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0

## [Unreleased]

### Fixed
## [4.0.0] - 2026-09-13

AgentOps 4.0 makes native coding the default: accepted intent, implementation
and checks, fresh independent judgment, then finish. No AgentOps skill,
bootstrap, hook or orchestration service is required. The optional skill menu
is consolidated from the 52 roots shipped in 3.6.0 to 34, and the CLI gains
recovery and exact-content evidence helpers. This is a major release because
published command, skill and scripted workflow entry points were removed.

- Code writers measure physical target lines with a metadata-only counter after writing and checking instead of inferring counts from rendered tool output.
- Code-writer child schemas preserve the caller's exact key and target path so successful writes do not become unknown receipts after path normalization. Direct writers return one JSON receipt with boolean check status and no copied test output.
- Claude context-budget guidance uses the plugin-qualified Agent and Workflow names; readers treat slice limits as per-call budgets and continue through EOF despite answer caps, early matches or truncated output.
- Read-budget shell parsing now preserves literal quoted paths, rejects uncertain shell constructs without false attribution, honors end-of-options, and counts negative head limits correctly without integer overflow.
- Opt-in read-budget installation preserves unique settings backups and checks the matcher and handler type before declaring the guard installed.
- Context-budget workflows validate bounded worker returns, remove raw check output, report unknown state after worker failures, and preflight target identities before serial batch writes.
See the [curated release notes](https://github.com/boshu2/agentops/blob/main/docs/releases/2026-09-13-v4.0.0-notes.md)
for upgrade instructions, the complete product-area summary and known limits.

### Added

- Codex-native `bulk-reader` and `code-writer` role templates pinned to `gpt-5.6-luna`, generated with the existing skill bundle, explicit project registrations, and an opt-in personal/project installer that preserves existing configuration.
- An opt-in Codex `PreToolUse` Bash adapter and installer reuse the read-budget guard's refusal, waivers and hashed telemetry; hook trust remains with the runtime. Native role instructions and runtime limitations are documented with live reader, writer and refusal evidence in `docs/design/codex-context-budget.md`.
- The opt-in read-budget guard, `skills/cc-hooks/hooks/read-budget-guard.sh` (policy `core.context:unbounded-read`): a PreToolUse `Read|Bash` hook that blocks an unbounded `Read`, `cat`, `head` or `tail` of a file over the line budget (`AOP_READ_BUDGET_LINES`, default 350), names the two correct moves (a bounded slice or `bulk-reader` delegation), honors `AOP_WAIVE`, the waiver file and `AGENTOPS_HOOKS_DISABLED`, and appends hashed telemetry. It ships inert; `scripts/install-read-budget-guard.sh` is the opt-in installer (user, `--project` or `SETTINGS` scope).
- The `bulk-read` workflow (`workflows/bulk-read.js`): one cheap reader agent per file, in parallel, reading in guard-compatible slices and returning line-referenced bullets with truthful `lines_covered` / `complete`; the file bytes never enter the caller's context.
- The `code-write` workflow (`workflows/code-write.js`): one cheap writer agent per item from a spec plus a required reference file, matching the reference's patterns, writing only its distinct target and returning a receipt (path, line count, check result) the caller never reads back.
- The `bulk-reader` and `code-writer` plugin subagents (`agents/bulk-reader.md`, `agents/code-writer.md`): the same reader and writer modes as Agent-tool `subagent_type` targets, `haiku` by default; `bulk-reader` is read-only.
- Docs for the context-budget pattern: the `cc-hooks` `READ-BUDGET-GUARD.md` recipe, its `GUARDRAIL-VALUE-PROOF.md` entry and skill-spec reference, the `agent-native` `context-budget-delegation.md` reference plus a Reader / Writer note in its Roles, the `workflows/README.md` shapes and context-budget paragraph, and one pointer in `docs/agent-workflow-reference.md`.
- Native Go evidence helpers under `ao provenance`: immutable intent snapshots,
subject manifests and digests, verdict storage, exact-subject verification,
required native judgment receipts, and orphaned-evidence inspection. Storage
uses an explicitly selected protected non-Git destination; mechanical
verification does not issue a semantic verdict.
- `ao config context` resolves caller-owned external context routes, including
recovery from a selected native BD maintenance anchor. `ao session read-source`
returns explicitly bounded source spans and integrity facts for synthetic or
already-cleared sources; restricted-source access remains unsupported.
- Bounded `ao provenance mine-session --view excerpts` extracts cited transcript
evidence without writing a checkpoint. Optional Memory guidance covers recall,
mining and reviewed topic curation; Skill Eval measures a named skill decision.
- Claude `bulk-reader` and `code-writer` subagents, pinned to Haiku, and parallel
`bulk-read` / serial `code-write` workflows return compact findings or receipts
while keeping source and generated code out of the parent context.
- Codex-native `bulk-reader` and `code-writer` roles pinned to `gpt-5.6-luna`,
generated with the portable bundle, plus explicit project registrations and
an opt-in personal/project installer that preserves existing configuration.
- Separate opt-in Claude and Codex read-budget guards refuse oversized unbounded
file reads, suggest bounded slices or reader delegation, honor scoped waivers,
and record hashed telemetry. The default line budget is 350; installing skills
does not activate these hooks.

### Changed

- `ao quick-start`, `ao demo`, the README and runtime onboarding describe native
execution with zero mandatory skills. `ao demo --rpi` retains the explicitly
selected workflow; `ao init` remains optional. `ao skills link --skill NAME`
supports a selected subset while no-selector linking still installs the menu.
- The 34-skill menu groups intent, implementation and judgment; engineering
specialists; Memory; deliberate review strategies; and optional tool/runtime
adapters. Consolidated names and replacements are documented in `docs/MIGRATION.md`.
- Plan uses existing conversation or tracker acceptance, Implement repairs known
defects directly, and Validate judges exact content from a fresh author-distinct
context. RPI is optional; final review is assigned once with every required leg
preserved. Requested retrospectives follow the known outcome and judgment.
- Memory and instruction-improvement guidance uses authorized source episodes,
bounded CASS/MS retrieval, protected drafts and independent support/disclosure
review. Later work must demonstrate benefit; a generated lesson is not proof.
- The CLI and CI build with the Go 1.27.1 toolchain; CI Python moves to 3.14.
Dependency and pinned GitHub Actions updates are included across the full
3.6.0-to-4.0.0 interval.

### Fixed

- Full release security checks scan the whole repository and treat Python collection or runtime errors as blocking failures. Nightly checks install the declared evaluator dependencies; test collection supports duplicate source/projection module names.
- Claude writers capture a supplied check's status in its original invocation, avoiding the observed status-confirmation rerun. Native success, failing-check and direct-writer trials each ran the check once; the direct child returned plain JSON.
- Plugin conformance verifies exact skill membership and link destinations. Clean Claude and Codex plugin installs and upgrades are exercised separately from source linking; Codex metadata checks use the current Skills-only manifest contract.
- Doctor keeps unique per-action backups and preflights undo integrity before
restoration. Session mining protects checkpoint paths and concurrent writers;
handoffs preserve work/session associations and use correct newest-first ordering.
- Evidence status, skill search, changed-file gate routing, constraint checks,
provenance graph validation and scenario-result publication handle previously
missed corruption, path, matching and recovery cases.
- The read-budget parser preserves quoted paths, handles end-of-options and
negative `head` limits, and avoids attributing uncertain shell constructs to
the wrong file. Repeat refusals stay short; installers preserve unique backups.
- Context-budget workers count physical lines, preserve the caller's target
identity, remove raw check output from returns and report unknown state after
worker failure. Batch writers preflight distinct target identities.
- Codex role registration points to real source-owned TOML files instead of
symlink paths rejected by the installed runtime.
- Generated skill projections, executable entry points, documentation checks,
seeded-defect probes and contamination detection received conformance repairs.

### Removed

- The published `ao eval` command family and `ao redact`. Use a repository-selected
evaluator and owner-authorized disclosure review; generic provenance can retain
the resulting evidence.
- `workflows/rpi.js` and its scripted retry machinery. Invoke the optional RPI
skill when selected, or execute natively. AgentOps does not own an aggregate
retry controller, queue, work ownership, Git or delivery transition.
- Twenty skill roots published in 3.6.0, including `learn`, `swarm`,
`codebase-recon`, `bootstrap`, `handoff`, `standards` and `workflow-builder`,
after useful behavior moved to surviving owners or native work. `memory` and
`skill-eval` are the two new roots relative to that tag.

## [3.6.0] - 2026-08-17

Expand Down
25 changes: 22 additions & 3 deletions agents/code-writer.md
Original file line number Diff line number Diff line change
Expand Up @@ -18,9 +18,23 @@ only your receipt, and independent validation happens elsewhere. When invoked:
3. Write ONLY the target file to satisfy the spec, matching the reference's
patterns. Code only: no markdown fences, no prose outside normal code
comments
4. If the caller gives a check command, run it ONCE with Bash after writing and
record whether it passed. Keep all output in your context: diagnostics can
echo source code, so never return the raw output or a tail
4. If the caller gives a check command, run it ONCE with Bash after writing.
Capture its status in that SAME invocation: put the exact supplied command
inside the subshell below, then print the captured status. The subshell keeps
a check's `exit` or shell options from skipping status capture:

set +e
(
SUPPLIED_CHECK_COMMAND
)
agentops_check_status=$?
printf '\nAGENTOPS_CHECK_STATUS=%s\n' "$agentops_check_status"

A zero status means `check_ok: true`; any other status means false. Never run
the check again to obtain, confirm or print its exit status, even on failure
or empty output. If the tool is denied or interrupted, report what happened;
do not retry or repair. Keep all output in your context: diagnostics can echo
source code, so never return the raw output or a tail
5. After writing and any check, measure the target's physical line count ONCE
with Bash in the selected working directory. Run the metadata-only counter
`awk 'END { print NR }'` with stdin redirected from the safely shell-quoted
Expand All @@ -40,6 +54,11 @@ Return exactly one JSON object with these fields and no others:
- `summary`: one-line string of at most 300 characters saying what was written,
with no code or copied command output

Your final response is the JSON text itself, starting with `{` and ending with
`}`. Do not wrap it in a Markdown code block, even a block labelled `json`.
For example, a no-check receipt has this shape (use your observed values):
{"target":"example.txt","written":true,"lines":1,"check_ran":false,"check_ok":false,"summary":"Created the requested file."}

No markdown fences, preamble, trailing prose or extra fields. Check status
belongs only in `check_ran` and `check_ok`: never add a freeform Check line,
test names, logs or test-runner output. Even a short success line is command
Expand Down
2 changes: 1 addition & 1 deletion cli/cmd/ao/main.go
Original file line number Diff line number Diff line change
Expand Up @@ -5,7 +5,7 @@ package main
// version is set at build time via ldflags (goreleaser: -X main.version={{ .Version }}).
// The fallback identifies untagged source builds for the next release;
// published binaries override it from the release tag via GoReleaser.
var version = "3.6.0"
var version = "4.0.0"

func main() {
Execute()
Expand Down
Loading
Loading