Status: initial version (2026-09-12), mapped against the OWASP Agentic/LLM Top 10 checklist (issue #72). Every item resolves to control (where it lives) / gap (tracked) / N/A (with rationale).
| Asset | Where | Notes |
|---|---|---|
| Provider API key | TOLE_/OPENAI_* env |
Never persisted; scrubbed from wire errors/messages |
| Session logs (JSONL) | sessions dir | Full transcript incl. tool results — local only |
| Job logs | <workspace>/tole-jobs/ |
Child stdout/stderr, attacker-influenceable content |
| Workspace files | --workspace dir (default cwd) |
Readable AND writable by the model via tools |
| Host env of child processes | inherited by spawned commands | See ENV-1 |
| Subprocess execution authority | run_command / job_start / git / gh | The model chooses argv within argv-split + jail rules |
- Provider responses → context. The model's output becomes intent entries; tool args are argv-validated but otherwise untrusted.
- Tool results → context. File contents, job logs, command stdout flow back into the transcript and are sent to the provider. A malicious workspace file or a compromised job subprocess can inject instructions here (OWASP: prompt injection).
- Workspace → filesystem. File tools jail to
--workspace(component-validated walk usingsymlink_metadata, final component opened withO_NOFOLLOWon unix, #68; a parent-directory swap between walk and open remains a residual race — see Deliberate limitations);run_command/job_startrun with cwd jailed but FULL process authority — the jail bounds cwd, not capability. - Child env. Spawned commands inherit the host environment by default.
| OWASP item | Status | Where / rationale |
|---|---|---|
| Prompt injection (direct/indirect) | control + documented residual | Tool results on the wire are wrapped in TOOL_RESULT_BEGIN/END fences with inner fence markers neutralized (this fix); secrets scrubbed (E11). Residual: an injected instruction inside data can still steer the model — no framework fully solves this; delimiters + user-visible tool audit lines (what: …) limit blast radius. |
| Insecure output handling | control | Same fence treatment; result sizes capped (read_file 1 MiB file cap and 20,000-char head cap per call, run_command 4000/2000 chars, job_poll 8 KiB window). |
| Excessive agency / tool misuse | control (defense-in-depth) + documented residual | Primary control: the per-call approval gate (every non-ReadOnly spawn needs consent). Secondary: check_destructive_argv blocklist refuses headline catastrophic patterns (teardown tools, recursive rm escaping the workspace, wrapper-nested variants) in BOTH run_command and job_start. The blocklist is best-effort by nature — wrapper/interpreter bypasses (sudo apt ..., timeout 5 dd ..., find / -delete, python -c rmtree) are expected and accepted residual risk; a real argv sandbox is a different product and explicitly out of scope. |
| Sensitive data disclosure (env) | gap → fixed here (ENV-1) | Children now inherit a scrubbed env: any var whose name contains API_KEY, _SECRET, _TOKEN, or PASSWORD is removed (GITHUB_TOKEN kept for gh/git auth — documented trade-off). |
| Supply chain (deps) | control | Cargo Audit + Trivy FS Scan required by rulesets on every PR and push. |
| Resource exhaustion | control + fixed here (JOB-2) | Provider 120 s timeout, subprocess 30 s ceiling, step budget 32, loop guard, read caps; job logs are read via bounded windows on poll (never a whole-log load). Since the 2026-09-18 scan triage, a ReadOnly poll no longer rewrites the log file — runaway log disk growth is bounded by the job's lifetime and clearing a runaway log is an operator action. |
| Session/message tampering | control | Append-only JSONL, CAS state transitions, seq monotonicity, torn-line recovery with truncation (#68); compaction never drops entries. |
| Injection via config | control (MCP shipped) | .cora.yaml/session headers are owner-controlled. MCP servers (#74) add an untrusted surface: server-supplied tool metadata is never trusted for risk tier (everything is Write → approval gate), tool results flow through the same fence/scrub path as native results, and the connection env is scrubbed. Residual: a malicious MCP server controls its own tool descriptions (model-visible) — treat server config as operator trust. |
| Identity & authz of sub-agents | N/A | tole v0 is single-agent, single-tenant. MCP servers are external tools, not sub-agents. |
| Human oversight | control | Risk-tiered approval gate (every non-ReadOnly call; Destructive never auto-allowed), #68 closed the replay-without-consent hole. |
GITHUB_TOKENstays in child env — gh/git push auth needs it; the token can therefore be read by any command the model runs. Mitigating context: the model already holds provider credentials for its own API calls by necessity; both are operator-scoped secrets on a single-operator host.- Workspace width is the operator's choice —
--workspace /is technically possible and a terrible idea; the flag validates existence, not sanity. Documented in CONTRIBUTING. - Symlink swap inside
tole-jobs/— requires a local attacker with write access to the operator's workspace, at which point host compromise is already achieved. Jobs dirs are created 0700 to reduce exposure; no further control planned. write_fileparent-component swap (residual TOCTOU) — the walk checks each component withsymlink_metadata, then the finalopen()usesO_NOFOLLOW, which protects only the FINAL component. A leaf symlink swap is refused; replacing an already-verified parent directory with a symlink between the walk and theopen()is not caught and can escape the jail. Exploiting it requires a concurrent local process with write access to the workspace, and the model already has full process authority throughrun_command/job_start(the jail bounds cwd, not capability), so severity is low. A per-componentopenatdescent (O_DIRECTORY | O_NOFOLLOW, e.g. viacap-std) could close it as a possible future hardening; none is planned or promised. Non-unix hosts have noO_NOFOLLOWequivalent wired (best-effortsymlink_metadatacheck only).git status/diffare repo-wide by design — they take no pathspec and run with cwd = the tool's workdir, so they show the whole repository containing it; in a monorepo subdirectory that includes modified files of sibling directories. This is not a jail escape: the jail (#247) applies toaddpathspecs (no absolute paths,.., or pathspec magic), and the whole git tool isRisk::Write, so every call — reads included — goes through the approval gate.- Injection residual (above) — fences are a mitigation, not immunity; the model must still treat fenced content as data.
- Per-call Destructive classification heuristics (from RC-1) — needs a design discussion before code.
- MCP (#74, SHIPPED): stdio client live; trust-model extensions applied
(metadata never trusted for risk, scrubbed env, fenced results). The
former "remaining surface" items shipped: HTTP transport + server auth
are
tole serve --transport mcp(#96/#137, hardened in #136) — see the Server surfaces section below. Resource subscriptions remain unshipped; that would re-open this document.
All three faces share one rule set: non-interactive hosts never gain
interactive powers. Concretely: Destructive registration is refused
behind a non-interactive approver (structurally absent from tole mcp
and serve; ACP is the exception by design — its approvals are genuine
per-call human decisions routed to the editor, so delete_file may
register there). The daemon adds (both transports): mandatory bearer token (refuses to
start without one), per-IP auth-failure rate limiting, a connection
cap (32), IO timeouts (30s: header, auth-header, and request-body
phases on the MCP transport — SSE response streaming is never cut),
and the jail-of-jails — a client-supplied session
cwd must canonicalize inside the server workspace root (default the
server cwd, --workspace override), or a remote client could jail a
session to /. Multi-session MCP routing strips the session_id key
before the tool sees its arguments, and a no-id registry call with 2+
open sessions is refused (the #138-documented ambiguity refusal,
implemented 2026-10-05) rather than silently executing against the
server-level registry.
SKILL.md files are operator-supplied prompt content, loaded into the
system prompt (--skill) or served on demand via the ReadOnly
load_skill tool (discovery index only — name + description — until
loaded). The surface is the same as --system: whoever controls the
workspace skills/ dir or ~/.codecora/tole/skills/ controls prompt
content; discovery from a compromised repo is prompt injection by
another name. Frontmatter is validated loudly (name/dir mismatch is a
hard error) and bodies are capped (16 KiB) — but content itself is
trusted by definition, same tier as the system prompt.
The ReadOnly registry check before any install answers the
slopsquatting class: hallucinated names surface NOT FOUND with registry
candidates, one-character candidates carry a typo-squat warning, and
registry 429/5xx report rate_limited honestly (never "not found").
It is advisory to the model — installs still go through the normal
approval gate (run_command Write tier).
--trust / TOLE_TRUST presets are pure sugar over --allow globs —
they widen nothing structurally: Destructive is never auto-allowed,
write-capable native tools keep prompting under internal, and unknown
preset names (flag or env) fail loudly. A typo'd env value breaks every
invocation with a clear error — deliberately, per the
typo-silently-narrowing-trust rule.
The session summary stored by the memory loop keeps its content contract
(first prompt + final answer only — never raw tool output), so the
prompt-injection-into-memory surface is unchanged. What #143 adds is a
TYPE: sessions that executed a Write/Destructive tool are stored with
--type decision and a wrote tag. The wrote bit is derived from the
harness's own tool-risk accounting (durable, session-scoped
fact.wrote_this_turn register — set on fresh execution AND
crash-replay, never reset by turn machinery), not from model-asserted
content — the model cannot claim or suppress the decision typing.
--on-turnend gates run owner-controlled scripts at the final-message
boundary. Two surfaces are deliberate and bounded: (1) the deny reason is
owner-authored input (the gate script is trusted the same way
--on-pretool scripts are — argv-executed, env-scrubbed, timed out), and
it enters the transcript as a user-role entry the model will read; (2) the
final-text preview handed to the gate is capped at 8 KiB and never leaves
the host process. Denials are capped at 3 per turn, so a permanently
failing gate cannot livelock the loop — the turn settles durably as
StopGateBlocked, visible to replay.
All three server faces expose cooperative turn cancellation: ACP
clients send session/cancel, REST callers POST /sessions/{id}/cancel,
MCP clients call tole_session_cancel. Every path sets the same
per-session token; a blocking turn observes it at checkpoints (between
provider calls and tool executions) and settles
stopReason: "cancelled" as a normal, durable turn end — the session
resumes as usual. Per the ACP and MCP specs the receiver MAY ignore
cancellation for work that cannot be stopped: a single in-flight tool
call (including run_command) is not interrupted; the token is
checked before the next one. Unknown session ids 404; cancelling an
idle session is an idempotent no-op. The tole_session_* glob in the
--trust internal preset covers the new MCP tool.
agent_start spawns a child tole session; agent_poll reads its
result. Both are Risk::Write (#300): a settled poll consumes the
mailbox (uteke forget + mailbox_consumed in meta.json), so the tier
matches the side effect and plan mode drops it (plan mode also drops
agent_start, so no child exists to poll). --trust internal
allowlists agent_poll by exact name to keep the poll loop unattended;
agent_start still prompts. The registry enforces a structural depth cap (no
grandchildren), results travel via ephemeral uteke mailboxes, and the
parent-only --agents-worktree flag gives each child its own git
worktree. A child's prompt is model/operator-supplied — the same trust
tier as the session system prompt. Child sessions are ordinary durable
sessions: approval gates, risk tiers, and secret redaction apply
unchanged inside them.
ReadOnly tool, active only when SYSTEMONE_API_KEY is set
(SYSTEMONE_BASE_URL selects the backend — hosted Jev default,
self-hosted compatible). The decision payload sent to the backend is
model-controlled context: treat the System One backend as an external
data flow. The tool executes no writes; results are advisory input to
the session like any other ReadOnly tool.
todo_write mutates only the session's task list, which lives in the
write-once session log itself (the tool's result entry is the record; no
side channel, no file). The threat surface equals any Write tool: prompt
injection could rewrite the plan, but the list is data, never executed —
and every revision is auditable in the replay. todo_read is ReadOnly.
tole mission chains normal durable turns autonomously. The risk frame
is unchanged and deliberate: the same approval gates, risk tiers, and
write-once audit apply per chained turn — autonomy does not widen the
boundary. Budgets bound blast radius in time, steps, and tokens
(#201); exhaustion is resumable, never a dead session. The durable cost
report keeps autonomous spend honest and comparable — an unmeasured
mission is an unaudited one. --verify gives the operator a
machine-checkable completion condition stronger than the model's own
claim. Destructive tools remain structurally un-auto-allowable in
missions.
The /approvals decision endpoint is the remote trust boundary: it is
behind the same bearer-token auth + rate limiter as every serve route,
and a decision is a one-shot for exactly one queued effect (fingerprint
of tool + canonical input) — never a blanket allow. Decisions expire to
denied so a lost connection cannot strand a mission, and every decision
writes a durable audit register naming the approval id, tool, and
verdict. The phone/CLI holder is therefore a full approver: treat the
token as approval authority and scope it accordingly.
uteke-mobile consumes the #200 queue: the phone becomes a remote
approver. The boundary is the serve token — it carries approval
authority, so device compromise equals write access to the box. Locked
down accordingly: the token is a revocable secret (rotate = serve
restart), decisions are one-shot per (session, effect) with expiry to
denied (a stolen device cannot bank future approvals), every decision
is audited on the session, and there is no push endpoint in tole — the
phone pulls, so the attack surface tole exposes is exactly the
authenticated REST face. Scoped device tokens (per-device, revocable,
read-only vs approver roles) are the recognized follow-up; until then
the deployment guidance is a dedicated OS user + minimal
--allow/--trust on the serve process.
web_fetch/web_search are ReadOnly: results enter context as model
content and are never executed — the exfiltration framing is identical
to any tool output. Fetch is text-only (no JS, no browser), size-capped
(512 KB), content-type allowlisted; search requires an explicit
TOLE_WEB_SEARCH_URL backend (probe-first — no keyless scraping). A
crafted page CAN steer a mission via prompt injection; the mitigation
is the same as every other untrusted input: tool-result fencing,
approval gates on anything that matters, and mission budgets bounding
the blast radius.
<cwd>/.tole/config.toml is untrusted input from a cloned repo: it may
carry allow, trust, hooks, skill and mcp_server. Part 2 added the
trust machinery, part 3a wired the startup gate and the low-risk keys, part 3b
applies the security-sensitive ones (trust, allow, mcp_server,
on_pretool/on_posttool/on_turnend, skill, plan_mode, no_auto_mcp,
no_skills, [mission] verify/verify_timeout):
- Content-bound, outside the repo. The approval is a snapshot of the exact
file bytes in
$CODECORA_HOME/tole/trusted-configs.json(default~/.codecora), keyed by canonical project directory. Any byte change (including whitespace or CRLF) makes it untrusted again; the repo cannot pre-trust itself. There is no home-less fallback: with noCODECORA_HOMEorHOMEtole errors instead of writing a store next to the project. The home is resolved by ONE shared function (tole_core::paths, #344) also used for user-global skills and the update-check cache: empty = unset, a relative value is refused, and skills / the update check are skipped (never read from or written to the cwd) when it is unresolvable. - Terminal-safe prompt. The path, the full content and the diff against the
previously trusted version are attacker-controlled text shown to a human at
the decision. Control characters (ESC, CR, NUL, DEL, C1), bidi
overrides/isolates, line/paragraph separators and zero-width characters are
printed as visible
\u{hex}escapes, so the prompt cannot be rewritten. - Fail closed. Without a terminal there is no question: the answer is an
error naming the file and
tole config trust. A corrupt, unreadable or unsupported-version store is a hard error and is never overwritten.tole config trustrefuses a file that does not validate. - Explicit path = intent. A file named on the command line (
--config) is user intent and does not consult the store.--no-configskips discovery, the gate and any output. - Startup gating. Every command except
config,upgradeandapprovalsruns the gate before anything else (before the update banner, session creation or provider access). Onlyrun/chat/resume/missionwith a terminal on stdin AND stderr (and not--prompt-file -) may ask;sessions,statusand the protocol facesserve/acp/mcpNEVER prompt — their stdin/stdout are protocol channels — and fail closed with the exacttole config trustinstruction while the command does not run. - Vetted bytes only. The gate reads the file once; that exact content is what gets parsed into the settings. The file is never re-read after vetting, so a swap between check and use cannot smuggle in other content.
- No config, no change. Without a config file nothing is read, printed or looked up (not even the trust store location).
- Secrets stay out. The API key and serve token are env-only; the schema
rejects secret-like keys, and
config checkprints only model/base_url values. - Sensitive keys apply only after trust.
allow,trust, the hooks,mcp_serverandskilltake effect through the SAME single startup path as every other key: nothing from the file is read into the settings before the gate passed, and--no-configdrops all of them. A configallow/trustlist feeds the very same allowlist machinery as the flags. - Destructive is never allowlistable via the config.
allow = ["*"],allow = ["delete_file"]or a trust preset cannot skip the Destructive prompt on the interactive CLI (the approver checksDestructivebefore any pattern), and on the non-interactive faces the Destructive tools remain structurally unregistered. Tests drive the real binary with such a config and assert the file survives. - Faces that refuse hooks refuse config hooks too.
serve/acp/mcprefuse--on-pretool/--on-posttool(andmcp--on-turnend),missionrefuses the hooks,--plan-mode,--memoryand the client-session flags (--skill,--no-skills,--mcp-server) on every one of these faces. The check runs on the EFFECTIVE (post-config) values, so a safety hook or deny policy arriving from the file makes the command fail loudly (with a--no-confighint) instead of being silently ignored. - Replace, never merge; booleans only up. A higher layer replaces a whole
list key, so a flag cannot be "extended" by a hostile file; a boolean key is
flag || config(a flag cannot switch a configtrueoff,--no-configcan). - Unsupported keys are loud. A build without the
mcpfeature rejectsmcp_server/no_auto_mcp, one withoutshell-toolsthe hook keys andmemory, naming the key and the feature.
Residual risks: a trusted config is trusted — a user who approves a malicious
file at the prompt (or with tole config trust --yes, or names it with
--config) grants exactly what the equivalent flags could grant (allowlisted
Write tools, hooks that run as the user, MCP servers spawned as the user,
skills injected into the prompt); it is shown in full, but a human decides; the trust store is a plain 0600 file in the
user's home, so anything running as that user can edit it; the store update is
an unlocked read-modify-write, so two concurrent tole config trust runs can
lose one of the two records (the loser is simply asked again; the file itself
is replaced atomically and never corrupted).