Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 8 additions & 2 deletions docs/ARCHITECTURE.md
Original file line number Diff line number Diff line change
Expand Up @@ -604,8 +604,14 @@ system (`agent_engine/runtime/hooks/`,
start/stop, run start/end/error, tool error, and `transform_tool_result` (lets
trusted code truncate/redact/normalize a tool result before it reaches the
model) — for auth, policy, audit, and context enrichment, distinct from
LLM-invoked tools. Per-tool `input_policy` (trusted parameter injection) is
still not implemented — see §11.
LLM-invoked tools. Tool results cross the runtime through a provider-independent
normalized value whose model text, structured data, and safe artifact metadata
remain separate and survive versioned, concurrency-safe idempotent replay. Only
the text projection enters the LangChain conversation; structured values remain
available to trusted hooks and the execution ledger (see
[`ADR 0004`](adr/0004-normalized-structured-tool-results.md)). Per-tool
`input_policy` (trusted parameter injection) is still not implemented — see
§11.

**Shared tool usage (✅ done, not in the original task list):**
`agent_engine/tool_usage/` owns execution metadata: a `ToolUsageRepository` port
Expand Down
64 changes: 64 additions & 0 deletions docs/MCP_AND_TOOLS.md
Original file line number Diff line number Diff line change
Expand Up @@ -155,6 +155,66 @@ stops requesting tools. Each call is recorded in the run's tool-usage repository
with its `provider` (`"local"` or `"mcp"`), so the origin is tracked for tracing
even though it is hidden from the model.

### Normalized tool results

A successful tool call is represented inside Extra as a provider-independent
`NormalizedToolResult`, not as only a string:

```python
NormalizedToolResult(
text="Found 2 invoices",
structured={"count": 2},
artifact={"source": "billing"},
)
```

The fields have separate responsibilities:

- `text` is the existing model-facing result. MCP text blocks are joined with
deterministic normalization rules. A structured-only result uses canonical
JSON text; a completely empty result uses a stable placeholder.
- `structured` is machine-readable output, including MCP
`structuredContent`. It must be JSON-like and is not copied into a model
message as metadata. For a structured-only result, its canonical JSON is
deliberately used as the model-facing `text` fallback. The shared JSON-safe
value policy limits nesting to 64 levels, total values to 10,000, cumulative
string data to 1,000,000 UTF-8 bytes, individual keys to 1,024 bytes,
cumulative key data to 256,000 bytes, and final canonical JSON to 1 MiB.
- `artifact` carries bounded, relevant artifact metadata. Raw in-memory binary
bodies and oversized values are omitted and replaced by type/size metadata.
The adapter limits depth to 8, each collection to 128 entries, total visited
values to 1,024, individual strings to 8,192 characters, cumulative string
content to 32,768 characters, keys to 256 characters, and final canonical
metadata to 65,536 characters.

`langchain-mcp-adapters` currently exposes MCP `structuredContent` through the
`ToolMessage.artifact.structured_content` path. Extra reads that provider
contract once at the tool adapter boundary and requires version 0.2 or newer,
where that artifact contract is available. It passes only its own normalized
result deeper into the runtime. Local dictionary/list results are normalized by
the same abstraction; plain string tools remain unchanged.

Only `text` is appended to the LangChain conversation. The complete normalized
result is available to trusted result hooks and is serialized into the
idempotency ledger as a versioned, JSON-primitive payload. A replay therefore
restores the same text, structured value, and artifact metadata without calling
the provider again. None of the non-text values are added to tool-usage records,
logs, model messages, callbacks, traces, or errors automatically.
The structured-only text fallback is the explicit exception: because that JSON
becomes model text, it is visible wherever ordinary model messages are visible.

The ledger's atomic claim has one execution owner. Concurrent duplicate callers
wait for that owner and replay its immutable terminal result; terminal rows
cannot be overwritten. Legacy string ledger values restore as text-only
results. Custom repository adapters implement the same claim/wait/complete
contract and must durably serialize the versioned primitive payload.

Structured values are preserved without application-specific schema validation,
but the runtime enforces its generic JSON-safe shape. A malformed provider
result becomes a controlled failed tool result and is never converted with
`repr()` or an arbitrary `str()` fallback. See
[`ADR 0004`](adr/0004-normalized-structured-tool-results.md).

The engine is driven as an async context manager: `build()` connects MCP servers
and discovers tools; `close()` (on context exit) releases them. `run()` does not
connect MCP servers on its own — `build()` must run first.
Expand Down Expand Up @@ -223,6 +283,10 @@ resume updates the existing record instead of adding a second one. Arguments and
results are never stored: they may carry sensitive or oversized data, and no
consumer of tool usage needs them.

This tool-usage repository is distinct from the private tool-execution
idempotency ledger. Tool usage stores no result values; the execution ledger
stores `NormalizedToolResult` so a replay can reproduce the completed call.

`ToolUsageRepository` is an abstract base class with three operations (`record`,
`list_for_run`, `list_for_conversation`). The engine ships a process-local adapter
(`InMemoryToolUsageRepository`); a distributed deployment supplies a Redis- or
Expand Down
24 changes: 17 additions & 7 deletions docs/RUNTIME_HOOKS.md
Original file line number Diff line number Diff line change
Expand Up @@ -40,7 +40,7 @@ explicit-ref mode) and whether a returned value is used.
| `on_run_error` | when a run fails | the `BaseException` | ignored (never masks the error) |
| `before_tool_call` | before every local or MCP tool call | `ToolRequestContext` | ignored (observe-only) |
| `after_tool_call` | after a local or MCP tool call **succeeds** | `ToolCallContext` (status `succeeded`) | ignored |
| `transform_tool_result` | after a tool **succeeds**, before its result is appended to the conversation | `ToolResultContext` (carries the `result`) | updated `ToolResultContext` (or `None`) |
| `transform_tool_result` | after a tool **succeeds**, before its result is appended to the conversation | `ToolResultContext` (carries text, structured result, and artifact metadata) | updated `ToolResultContext` (or `None`) |
| `on_tool_error` | when a local or MCP tool call **fails** | `ToolCallContext` (status `failed`) | ignored |
| `before_mcp_request` | before every outgoing MCP HTTP request | `McpRequestContext` | updated `McpRequestContext` (or `None`) |
| `after_mcp_response` | after every MCP HTTP response | `McpResponseContext` | ignored (observe-only) |
Expand Down Expand Up @@ -343,16 +343,26 @@ EngineContext(system_name, metadata)
RunEndContext(run_id, system_name, status, visited, used_tool_count, metadata)
ToolRequestContext(agent_id, tool_name, provider, server_id, metadata)
ToolCallContext(agent_id, tool_name, provider, server_id, status, latency_ms, error, metadata)
ToolResultContext(agent_id, tool_name, provider, result, server_id, latency_ms, metadata)
ToolResultContext(agent_id, tool_name, provider, result, server_id, latency_ms, metadata, structured_result, artifact)
McpRequestContext(server_id, url, operation, tool_name, headers, metadata)
McpResponseContext(server_id, url, status_code, operation, tool_name, latency_ms, metadata)
```

All are frozen dataclasses. `RunContext.replace(**changes)` and
`McpRequestContext.with_headers({...})` and `ToolResultContext.with_result(...)`
return updated copies; the `headers` and `metadata` dicts may also be mutated in
place. Hooks never receive raw graph
state.
All are frozen dataclasses. `RunContext.replace(**changes)`,
`McpRequestContext.with_headers({...})`, and the
`ToolResultContext.with_result(...)`, `with_structured_result(...)`, and
`with_artifact(...)` helpers return updated copies; the `headers` and `metadata`
dicts may also be mutated in place. `with_result(...)` changes only model-facing
text and preserves structured output and artifact metadata. Hooks never receive
raw graph state.

`ToolResultContext.result` remains a string for existing hooks.
`structured_result` carries provider-independent machine-readable output, and
`artifact` contains structurally bounded metadata after binary and oversized
bodies have been omitted. Structural bounds do not identify application
secrets: hook implementations may inspect, redact, or replace these trusted
runtime values but must not log their contents. See
[`ADR 0004`](adr/0004-normalized-structured-tool-results.md).

---

Expand Down
94 changes: 94 additions & 0 deletions docs/adr/0004-normalized-structured-tool-results.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,94 @@
# ADR 0004 — Preserve structured tool results separately from model text

- **Status:** Accepted
- **Date:** 2026-09-01

## Context

The tool execution path reduced every successful local or MCP result to a
string. Text extraction kept MCP content blocks readable for the model, but it
discarded MCP `structuredContent`, LangChain artifact metadata, and structured
local-tool returns. The idempotency ledger consequently replayed only text, and
`transform_tool_result` hooks could not inspect machine-readable output.

## Decision

Extra owns a provider-agnostic frozen `NormalizedToolResult` with independent
`text`, `structured`, and `artifact` fields.

- The LangChain tool adapter is the normalization boundary. It extracts text
from content blocks, reads MCP structured output from the current
`ToolMessage.artifact.structured_content` contract (and the compatible
`structuredContent` spelling), and removes LangChain-specific objects before
the result enters the runtime.
- Local JSON object/array results that LangChain serializes into
`ToolMessage.content` are parsed back into `structured` while the serialized
model-facing text remains unchanged.
- The model loop consumes only `NormalizedToolResult.text`. It creates a plain
`ToolMessage` without an artifact, so structured values do not enter model
context, LangChain callbacks, or tracing as hidden message metadata.
- A structured-only result receives deterministic canonical JSON as its model
text. A completely empty result receives a stable placeholder. Unsupported
values fail with controlled model text; provider objects are never rendered
through `repr()` or arbitrary `str()` fallback. The structured-only JSON is
ordinary model text and therefore has the same model/callback visibility as
any other tool text.
- `ToolResultContext.result` remains the text field for hook compatibility and
gains additive `structured_result` and `artifact` fields. `with_result()`
changes text only, so truncating text cannot discard structured output.
- The value object validates JSON-like values, stores canonical JSON privately,
and returns copies from its accessors. Provider-owned nested objects therefore
cannot mutate the authoritative result after normalization. Generic depth,
value-count, string/key byte, integer-size, and final encoded-size budgets
prevent structured payloads from multiplying memory use across hooks,
persistence, and replay.
- The tool-execution repository persists a versioned JSON-primitive payload.
An atomic claim identifies one execution owner; concurrent duplicates wait
for a terminal result and replay it without re-invoking the provider. Terminal
rows cannot be overwritten, and legacy string records remain readable as
text-only results.
- Artifact mappings retain structurally safe, bounded metadata under explicit
per-value, aggregate-size, collection-count, key-length, and nesting limits.
Binary and oversized values are replaced by type/size/omission metadata and
are never copied into model context. This is structural safety, not semantic
secret classification; trusted hooks remain responsible for application-level
redaction before persistence when needed.
- Structured values are not validated against application output schemas in
v1, but they must satisfy the generic JSON-safe runtime contract. Malformed
provider results become controlled failed tool results.

## Contract changes

- `ToolExecutionRepository.complete()` accepts the versioned persisted payload,
and the port adds `wait_for_completion()`. Custom repository adapters must
serialize all three fields, atomically create claims, reject terminal
overwrites, and wake duplicate waiters after every owner outcome.
- Tool execution states use `ToolExecutionStatus`, a `StrEnum` that remains
wire-compatible with the existing string values while removing duplicated
state literals from the ledger implementation.
- `ToolResultContext` adds optional `structured_result` and `artifact` fields
plus immutable update helpers. Existing text-only hooks remain source
compatible.
- `ToolInvoker.invoke()` returns `NormalizedToolResult`; this is an internal
engine contract consumed by the model loop.

There is no YAML, HTTP API, widget, or model-facing text contract change.
The runtime dependency floor for `langchain-mcp-adapters` is 0.2 because the
normalizer relies on that release's structured-content artifact contract.

## Consequences

- Text-only local tools behave exactly as before.
- MCP text, structured output, and safe artifact metadata can coexist and
survive idempotent replay.
- Hooks can inspect or deliberately replace structured values without parsing
text, while ordinary text transforms preserve them by default.
- Structured values are runtime data and are not automatically written to tool
usage, hidden model-message metadata, logs, traces, or error messages. The
documented structured-only text fallback is the deliberate model-visible
exception.
- Existing persisted string results remain readable. The shipped execution
repository is process-local, so no database migration is required; external
repository adapters must adopt the expanded port and versioned payload.
- Full multimodal artifact rendering, binary persistence, and output-schema
enforcement remain out of scope.
2 changes: 1 addition & 1 deletion pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -49,7 +49,7 @@ dependencies = [
"langchain>=0.3",
"langchain-core>=0.3",
"mcp>=1.27,<2",
"langchain-mcp-adapters>=0.1",
"langchain-mcp-adapters>=0.2",
"pyjwt>=2.9",
"python-dotenv>=1.0",
"pydantic-settings>=2.0",
Expand Down
6 changes: 6 additions & 0 deletions src/agent_engine/approvals/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -68,12 +68,15 @@
RunRecord,
RunStatus,
ToolExecutionRecord,
ToolExecutionStatus,
)
from agent_engine.approvals.sanitization import mask_arguments, mask_sensitive
from agent_engine.approvals.session_approval_repository import SessionApprovalRepository
from agent_engine.approvals.session_approval_store import SessionApprovalStore
from agent_engine.approvals.tool_execution_manager import (
ToolExecutionClaim,
ToolExecutionManager,
ToolExecutionStateError,
execution_id_for,
)
from agent_engine.approvals.tool_execution_repository import ToolExecutionRepository
Expand Down Expand Up @@ -110,9 +113,12 @@
"SessionApprovalRepository",
"SessionApprovalScope",
"SessionApprovalStore",
"ToolExecutionClaim",
"ToolExecutionManager",
"ToolExecutionRecord",
"ToolExecutionRepository",
"ToolExecutionStateError",
"ToolExecutionStatus",
"ToolInvocation",
"ToolNoLongerExists",
"UnauthorizedApprover",
Expand Down
52 changes: 43 additions & 9 deletions src/agent_engine/approvals/in_memory_tool_execution_repository.py
Original file line number Diff line number Diff line change
Expand Up @@ -3,31 +3,65 @@
from __future__ import annotations

import asyncio
import copy
import dataclasses

from agent_engine.approvals.models import ToolExecutionRecord
from agent_engine.approvals.models import ToolExecutionRecord, ToolExecutionStatus
from agent_engine.approvals.tool_execution_repository import ToolExecutionRepository
from agent_engine.runtime.tool_results import PersistedToolResult


class InMemoryToolExecutionRepository(ToolExecutionRepository):
def __init__(self) -> None:
self._records: dict[str, ToolExecutionRecord] = {}
self._completed: dict[str, asyncio.Event] = {}
self._lock = asyncio.Lock()

async def get(self, execution_id: str) -> ToolExecutionRecord | None:
async with self._lock:
return self._records.get(execution_id)
record = self._records.get(execution_id)
return _copy_record(record) if record is not None else None

async def start(self, record: ToolExecutionRecord) -> tuple[ToolExecutionRecord, bool]:
async with self._lock:
existing = self._records.get(record.execution_id)
if existing is not None:
return existing, False
self._records[record.execution_id] = record
return record, True
return _copy_record(existing), False
stored = _copy_record(record)
self._records[record.execution_id] = stored
self._completed[record.execution_id] = asyncio.Event()
return _copy_record(stored), True

async def complete(self, execution_id: str, status: str, result: str) -> None:
async def wait_for_completion(self, execution_id: str) -> ToolExecutionRecord:
async with self._lock:
record = self._records.get(execution_id)
if record is not None:
record.status = status
record.result = result
if record is None:
raise KeyError(f"tool execution not found: {execution_id}")
if record.status != ToolExecutionStatus.STARTED:
return _copy_record(record)
completed = self._completed[execution_id]
await completed.wait()
async with self._lock:
return _copy_record(self._records[execution_id])

async def complete(
self,
execution_id: str,
status: ToolExecutionStatus,
result: PersistedToolResult,
) -> None:
if status == ToolExecutionStatus.STARTED:
raise ValueError("completed tool execution must be terminal")
async with self._lock:
record = self._records.get(execution_id)
if record is None:
raise KeyError(f"tool execution not found: {execution_id}")
if record.status != ToolExecutionStatus.STARTED:
raise ValueError(f"tool execution is already terminal: {execution_id}")
record.status = status
record.result = copy.deepcopy(result)
self._completed[execution_id].set()


def _copy_record(record: ToolExecutionRecord) -> ToolExecutionRecord:
return dataclasses.replace(record, result=copy.deepcopy(record.result))
11 changes: 9 additions & 2 deletions src/agent_engine/approvals/models.py
Original file line number Diff line number Diff line change
Expand Up @@ -21,6 +21,7 @@

from agent_engine.approvals.errors import InvalidStateTransition
from agent_engine.runtime.tool_models import ToolProviderName
from agent_engine.runtime.tool_results import PersistedToolResult


class RunStatus(StrEnum):
Expand All @@ -39,6 +40,12 @@ class ApprovalStatus(StrEnum):
REJECTED = "rejected"


class ToolExecutionStatus(StrEnum):
STARTED = "started"
SUCCEEDED = "succeeded"
FAILED = "failed"


# Allowed forward transitions. Anything not listed is rejected. Terminal run
# states have no outgoing path, and approvals cannot move REJECTED -> APPROVED.
_RUN_TRANSITIONS: dict[RunStatus, frozenset[RunStatus]] = {
Expand Down Expand Up @@ -154,6 +161,6 @@ class ToolExecutionRecord:
tool_call_id: str
run_id: str
tool_name: str
status: str = "started" # started | succeeded | failed
result: str | None = None
status: ToolExecutionStatus = ToolExecutionStatus.STARTED
result: PersistedToolResult | None = None
created_at: float = field(default_factory=time.time)
Loading
Loading