diff --git a/CHANGELOG.md b/CHANGELOG.md
index 3927a44..3c18028 100644
--- a/CHANGELOG.md
+++ b/CHANGELOG.md
@@ -22,6 +22,8 @@ First public-release candidate. This version is prepared but has not been pushed
- `jev_record_execution` and append-only JSONL routing/execution receipts joined by `correlation_id`.
- Offline replay/evaluation without host execution calls.
- Opt-in supervision judgments with deterministic host policy.
+- Opt-in, report-only `jev_model_route` with correlated execution outcomes and a local evidence report.
+- Hermes plugin integration with profile-scoped OpenRouter secret resolution, replay receipts, model-route correlation, and a local-agent handoff guide.
- Opt-in deterministic context filtering (`shadow` and `conservative`).
- Explicit capability discovery for skills, MCP, CLI, DSH, tools, subagents, and models.
- Experimental, opt-in browser fast-path over host-supplied observations.
diff --git a/README.md b/README.md
index bbf59c9..d1c2814 100644
--- a/README.md
+++ b/README.md
@@ -5,13 +5,11 @@
jev-layer
- Portable System-1 decision layer for agent harnesses.
- Host-owned routing, receipts, replay, and fail-open integrations.
+ Let an agent make a bounded choice without handing it control.
+ Jev recommends. Your harness still checks permissions and executes.
-
- Hermes · OMP · Codex · generic MCP
-
+Hermes · OMP · Codex · generic MCP
@@ -22,66 +20,57 @@
[English](README.md) · [Русский](README.ru.md) · [简体中文](README.zh-CN.md)
-jev-layer routes bounded choices and records evidence; the host keeps execution, permissions, approvals, retries, recovery, and final results.
-
-> Integrating jev-layer into a harness? Start with the [Agent implementation guide](docs/AGENT-IMPLEMENTATION.md), not this README alone.
-
-## Architecture
+**Install:** `npm install --global jev-layer` · [Integrate a harness](docs/AGENT-IMPLEMENTATION.md) · [Security model](SECURITY.md)
-
-
-
+## What changes
-Jev never executes a selected capability. A provider can be deterministic `demo`, OpenRouter Decisions, or TypeSafe; provider-backed tests are not required for normal CI.
+| Without Jev | With Jev |
+| --- | --- |
+| Your harness follows its existing path to choose a capability. | The harness can ask `jev_route` to choose from a bounded set it supplies. |
+| Your harness owns permissions, approvals, and execution. | Your harness still owns permissions, approvals, and execution. |
+| Execution results stay in the host's normal workflow. | The host can attach the result to the decision with `jev_record_execution` and replay cases offline. |
-## Quick Start
+Jev never executes a selected capability. If it is disabled, unavailable, invalid, or inconclusive, control returns to the host's normal path.
-Requirements: Node.js 20 or newer. There are no mandatory runtime dependencies.
+## Quick start
-Install the published CLI:
+Requires Node.js 20 or newer. There are no mandatory runtime dependencies.
```sh
npm install --global jev-layer
-```
-
-Or use a local clone:
-
-```sh
-npm install
-npm link
jev install --project /path/to/workspace
jev add generic --project /path/to/workspace
jev doctor --project /path/to/workspace
```
-`npm link` is local only. It does not publish the package. Use `node /path/to/jev-layer/bin/jev.mjs ...` instead if a global link is not wanted. The default `demo` provider is offline and deterministic.
-
-To call the stdio MCP server directly:
+The default `demo` provider is deterministic and works offline. To run the stdio MCP server directly:
```sh
jev mcp
```
-To use a provider with credentials, keep keys outside the repository:
+## How it fits into a harness
-```sh
-export JEV_LAYER_PROVIDER=openrouter
-export OPENROUTER_API_KEY='provided-by-your-secret-store'
-jev doctor --project /path/to/workspace
-```
+
+
+
-All three modes (`demo`, `openrouter`, and direct `typesafe`), their endpoints, and configuration precedence are documented in the [provider guide](docs/PROVIDERS.md).
+1. The host sends Jev a request and the candidate capabilities it already allows.
+2. Jev returns a bounded recommendation. The host checks it against its own registry and permissions.
+3. The host decides whether to execute, then can record what happened against the original `correlation_id`.
-## Core surfaces
+A provider can be deterministic `demo`, OpenRouter Decisions, or TypeSafe. Provider-backed tests are not required for normal CI; see the [provider guide](docs/PROVIDERS.md).
-- **Routing:** `jev_route` selects one capability from the host-supplied candidate set. Selection is advisory; the host validates the id and permissions.
-- **Receipts/replay:** `jev_record_execution` joins the host result to the original `correlation_id`. JSONL cases live in `.jev/replay/cases.jsonl` and can be evaluated offline with `npm run replay:evaluate`.
-- **Supervision:** `jev_supervise` returns bounded work-state judgments; deterministic host policy maps them to `continue`, `verify`, `retry`, `finish`, or `escalate`. Jev does not perform those actions.
+## What else it can do
+
+- **Supervision:** `jev_supervise` returns bounded work-state judgments. The host decides whether to continue, verify, retry, finish, or escalate.
+- **Model routing:** `jev_model_route` recommends one host-declared model profile for a future call. It is advisory only; the host measures outcomes before changing provider or model settings. Correlated receipts can be reviewed with `npm run model-route:report -- /path/to/cases.jsonl`.
+- **Shadow compaction:** `jev_shadow_compaction` produces report-only keep/drop candidates for host-supplied context. It does not summarize, mutate, or delete context, and keeps pinned evidence on provider failure.
- **Context filtering:** optional deterministic `shadow` or `conservative` filtering reduces stale context without LLM summarization.
-- **Experimental browser fast-path:** `jev_browser_step` chooses one bounded action from a host observation. The host supplies observations, approval, native execution, and recovery. It is opt-in and does not start a browser worker.
-- **Fail-open:** disabled, unavailable, invalid, or inconclusive Jev calls return control to the host's normal path. Jev never widens permissions or guesses execution.
+- **Experimental browser fast-path:** `jev_browser_step` recommends one bounded action from a host observation. The host supplies approval, native execution, and recovery; Jev does not start a browser worker.
+- **Fail open:** optional Jev surfaces are disabled by default. Jev never widens permissions or guesses execution.
-All optional surfaces are disabled by default:
+Enable optional surfaces explicitly:
```sh
JEV_BROWSER_FAST_PATH=1 jev mcp
@@ -89,43 +78,35 @@ JEV_SUPERVISION=1 jev mcp
JEV_CONTEXT_FILTER=shadow jev cli --input examples/route-request.json
```
-## Harness adapters
-
-Current examples live under `integrations/`:
-
-- `integrations/hermes/`
-- `integrations/omp/`
-- `integrations/codex/`
-- `integrations/template/`
-
-The release baseline records OMP `18.2.6`, Hermes `0.21.3` (`b675e6de`), and Codex CLI `0.155.1` observed in the preparation environment. This is a version/contract baseline, not a claim of full provider/model coverage; see [docs/COMPATIBILITY.md](docs/COMPATIBILITY.md).
-
-Adapters are intentionally thin. They may call the CLI or stdio MCP, but the host must retain native capability lookup, permissions, approvals, execution, retries, recovery, and final output.
+For provider credentials, keep keys outside the repository:
-### Adding a new harness
+```sh
+export JEV_LAYER_PROVIDER=openrouter
+export OPENROUTER_API_KEY='provided-by-your-secret-store'
+jev doctor --project /path/to/workspace
+```
-The shortest PR path is:
+## Pick an integration
-1. copy `integrations/template/adapter.mjs`;
-2. add `integrations//` and a secret-free config/example;
-3. call `jev_route`, preserve `correlation_id`, execute only through the host registry, then call `jev_record_execution`;
-4. add an offline smoke fixture for success, fail-open, approval denial, and execution receipt;
-5. document supported versions and run CI.
+Examples and adapters live under `integrations/`:
-See [CONTRIBUTING.md](CONTRIBUTING.md) for the adapter contract and [docs/SCHEMA-VERSIONING.md](docs/SCHEMA-VERSIONING.md) for compatibility rules.
+- Hermes: `integrations/hermes/` (see the [local-agent handoff guide](docs/HERMES-LOCAL-AGENT-HANDOFF.ru.md))
+- OMP: `integrations/omp/`
+- Codex: `integrations/codex/`
+- Another harness: start from `integrations/template/`
-## Browser status
+Adapters stay thin. The host retains native capability lookup, permissions, approvals, execution, retries, recovery, and final output. For the adapter contract, see [CONTRIBUTING.md](CONTRIBUTING.md). For compatibility rules, see [docs/SCHEMA-VERSIONING.md](docs/SCHEMA-VERSIONING.md).
-Browser fast-path **reliability is validated against the current real-browser fixtures**, including action sequencing, visible-link navigation, native select execution, approval denial, and recovery. **Performance optimization remains experimental**. No browser speedup claim is made.
+The release baseline records OMP `18.2.6`, Hermes `0.21.3` (`b675e6de`), and Codex CLI `0.155.1` observed in the preparation environment. This is a version/contract baseline, not a claim of full provider/model coverage; see [docs/COMPATIBILITY.md](docs/COMPATIBILITY.md).
-## Security and compatibility
+## Boundaries
-- MIT licensed; see [LICENSE](LICENSE).
-- jev-layer is not a security boundary. Host permissions and approvals are authoritative; see [SECURITY.md](SECURITY.md).
+- jev-layer is not a security boundary. Host permissions and approvals remain authoritative; see [SECURITY.md](SECURITY.md).
- Schema, MCP tool, receipt, replay, and adapter contracts are currently version 1. Prefer additive changes; do not break v1 silently.
+- Browser fast-path reliability is validated against current real-browser fixtures. Performance optimization remains experimental; no browser speedup claim is made.
- Do not commit credentials, logs containing secrets, `.env` files, or machine-specific paths.
-## Verification
+## Verify locally
```sh
npm test
@@ -135,12 +116,8 @@ npm run clean-install-smoke
npm pack --dry-run
```
-The GitHub Actions matrix runs these checks on Node.js 20, 22, and 24. Provider-backed tests require an explicitly configured secret-managed environment and are not part of ordinary PR CI.
+GitHub Actions runs these checks on Node.js 20, 22, and 24. Provider-backed tests require a secret-managed environment and are not part of ordinary PR CI.
-## Release documents
+## Project docs
-- [CONTRIBUTING.md](CONTRIBUTING.md)
-- [SECURITY.md](SECURITY.md)
-- [RELEASE.md](RELEASE.md)
-- [CHANGELOG.md](CHANGELOG.md)
-- [Agent implementation guide](docs/AGENT-IMPLEMENTATION.md)
+[Agent implementation guide](docs/AGENT-IMPLEMENTATION.md) · [Providers](docs/PROVIDERS.md) · [Compatibility](docs/COMPATIBILITY.md) · [Contributing](CONTRIBUTING.md) · [Security](SECURITY.md) · [Release](RELEASE.md) · [Changelog](CHANGELOG.md)
diff --git a/README.ru.md b/README.ru.md
index 3abeb27..6ca87ea 100644
--- a/README.ru.md
+++ b/README.ru.md
@@ -76,7 +76,9 @@ jev doctor --project /path/to/workspace
- **Routing:** `jev_route` выбирает одну capability из набора, предоставленного host. Host повторно проверяет id и permissions.
- **Receipts/replay:** `jev_record_execution` связывает результат host с исходным `correlation_id`. JSONL-файлы находятся в `.jev/replay/cases.jsonl` и проверяются офлайн через `npm run replay:evaluate`.
-- **Supervision:** `jev_supervise` возвращает ограниченные judgments о состоянии работы; детерминированная host policy преобразует их в `continue`, `verify`, `retry`, `finish` или `escalate`. Jev эти действия не выполняет.
+- **Supervision:** `jev_supervise` возвращает ограниченные judgments о состоянии работы; детерминированная host policy преобразует их в `continue`, `verify`, `retry`, `finish` или `escalate`. Receipt хранит детерминированный `evidence_state`: `present`, `missing` или `contradictory`. Противоречивые evidence не могут привести к `finish`. Jev эти действия не выполняет.
+- **Model routing:** `jev_model_route` возвращает одну рекомендацию из model profiles, которые объявил host, для следующего model call. Это только shadow: прежде чем менять provider/model setting, host обязан измерить outcome, retry, latency и cost. Решение связывается с последующим `jev_record_execution` по `correlation_id`; host записывает `result.model_route = { actual_model_id, retry_count, outcome }`, затем `npm run model-route:report -- /path/to/cases.jsonl` строит read-only evidence report.
+- **Shadow compaction:** `jev_shadow_compaction` делает пакетный консервативный report с keep/drop-кандидатами для context, который передал host. Он никогда не меняет, не суммаризирует и не удаляет context; path, error, command и requirement pin-ятся до provider review, а provider failure оставляет все остальные items.
- **Context filtering:** опциональная детерминированная фильтрация `shadow` или `conservative` убирает устаревший context без LLM-суммаризации.
- **Experimental browser fast-path:** `jev_browser_step` выбирает одно ограниченное действие из observation host. Host предоставляет observation, approval, native execution и recovery.
- **Fail-open:** при отключённом, недоступном, ошибочном или неубедительном Jev вызове управление возвращается в обычный host path. Jev не расширяет permissions и не угадывает execution.
@@ -93,7 +95,7 @@ JEV_CONTEXT_FILTER=shadow jev cli --input examples/route-request.json
Примеры находятся в `integrations/`:
-- `integrations/hermes/`
+- `integrations/hermes/` — для active plugin topology и safe change workflow см. [handoff локальному агенту (RU)](docs/HERMES-LOCAL-AGENT-HANDOFF.ru.md).
- `integrations/omp/`
- `integrations/codex/`
- `integrations/template/`
diff --git a/README.zh-CN.md b/README.zh-CN.md
index 643d452..2924467 100644
--- a/README.zh-CN.md
+++ b/README.zh-CN.md
@@ -76,7 +76,8 @@ jev doctor --project /path/to/workspace
- **Routing:** `jev_route` 从 host 提供的候选集合中选择一个 capability。host 会再次验证 id 和权限。
- **Receipts/replay:** `jev_record_execution` 使用原始 `correlation_id` 关联 host 结果。JSONL 位于 `.jev/replay/cases.jsonl`,可用 `npm run replay:evaluate` 离线评估。
-- **Supervision:** `jev_supervise` 返回有界的工作状态判断;确定性的 host policy 将其映射为 `continue`、`verify`、`retry`、`finish` 或 `escalate`。Jev 不执行这些动作。
+- **Supervision:** `jev_supervise` 返回有界的工作状态判断;确定性的 host policy 将其映射为 `continue`、`verify`、`retry`、`finish` 或 `escalate`。receipt 记录确定性的 `evidence_state`:`present`、`missing` 或 `contradictory`;矛盾证据不能导致 `finish`。Jev 不执行这些动作。
+- **Shadow compaction:** `jev_shadow_compaction` 为 host 提供的 context 生成批量、保守、只报告的 keep/drop 候选。它不会修改、总结或删除 context;path、error、command 和 requirement 在 provider review 前被保留,provider 失败时保留所有其他 items。
- **Context filtering:** 可选的确定性 `shadow` 或 `conservative` 过滤器减少过期 context,不使用 LLM 摘要。
- **Experimental browser fast-path:** `jev_browser_step` 根据 host observation 选择一个有界浏览器动作。observation、审批、原生执行和恢复都由 host 提供。
- **Fail-open:** Jev 被禁用、不可用、出错或无法确定时,控制权返回 host 的正常路径。Jev 不扩大权限,也不猜测执行。
diff --git a/docs/HERMES-LOCAL-AGENT-HANDOFF.ru.md b/docs/HERMES-LOCAL-AGENT-HANDOFF.ru.md
new file mode 100644
index 0000000..2ed5bc5
--- /dev/null
+++ b/docs/HERMES-LOCAL-AGENT-HANDOFF.ru.md
@@ -0,0 +1,177 @@
+# Hermes integration: handoff для локального агента
+
+Этот документ описывает **текущую рабочую интеграцию** `jev-layer` в Hermes. Он нужен локальному агенту, чтобы сначала понять границы и точки входа, а не повторно «интегрировать Jev» через глобальный конфиг, произвольные shell-команды или второй executor.
+
+## Что уже установлено
+
+- Исходный репозиторий: `/home/hermes/jev-layer`.
+- Активная plugin-копия Hermes: `~/.hermes/plugins/jev-layer`.
+- Plugin регистрирует ровно шесть tools:
+ - `jev_route`;
+ - `jev_record_execution`;
+ - `jev_supervise`;
+ - `jev_browser_step`;
+ - `jev_shadow_compaction`;
+ - `jev_model_route`.
+- Replay evidence активного профиля: `~/.hermes/jev-layer/replay/cases.jsonl`.
+- Gateway загружает Python plugin при старте. После изменения `integrations/hermes/*`, `plugin.yaml` или списка tools нужен пользовательский `/restart` gateway.
+
+Не путай surfaces:
+
+```text
+/home/hermes/jev-layer source checkout, тесты и документация
+~/.hermes/plugins/jev-layer установленная plugin-копия текущего профиля
+~/.hermes/jev-layer/replay/cases.jsonl profile-scoped evidence, не source artifact
+~/.hermes/config.yaml глобальный Hermes config; не менять для этой plugin без явного задания
+```
+
+## Главный принцип
+
+```text
+Hermes registry/permissions/approval/executor/recovery
+ │ closed, bounded input
+ ▼
+ jev-layer decision
+ │ correlation_id, advice only
+ ▼
+Hermes validates → native execution → jev_record_execution → JSONL evidence
+```
+
+Jev **никогда** не получает право исполнять tool, shell, URL или browser action. Любое решение — advisory. Hermes остаётся владельцем discovery, policy, approval, native execution, retry, recovery и финального ответа.
+
+Если Jev выключен, недоступен, даёт невалидный ответ или низкую уверенность, текущий normal Hermes path продолжается. Не добавляй обходной executor и не меняй retry semantics.
+
+## Файловая карта
+
+| Зачем | Файл |
+| --- | --- |
+| Hermes tool registration + profile-scoped secret bridge | `integrations/hermes/__init__.py` |
+| JSON schemas шести tool surfaces | `integrations/hermes/schemas.py` |
+| Plugin manifest | `integrations/hermes/plugin.yaml`, `plugin.yaml` |
+| Один JSONL request/response adapter process | `src/hermes-adapter.mjs` |
+| Общая bounded routing contract | `src/route.mjs`, `src/contract.mjs` |
+| Receipt/replay JSONL format | `src/receipts.mjs` |
+| Browser action space, progress и recovery | `src/browser.mjs` |
+| Model recommendation | `src/model-routing.mjs` |
+| Model-route evidence report | `src/model-route-metrics.mjs`, `scripts/model-route-report.mjs` |
+| Context keep/drop report | `src/shadow-compaction.mjs` |
+| Work-state judgement | `src/supervision.mjs` |
+| Harness-neutral integration contract | `docs/AGENT-IMPLEMENTATION.md` |
+
+## Как работают tools
+
+### `jev_route`
+
+Hermes сначала собирает **закрытый** `capabilities[]` из своего registry. Jev выбирает один известный id. До исполнения Hermes снова проверяет: `status`, точное совпадение id, availability, permission, risk и approval. После результата используется `jev_record_execution` с тем же `correlation_id`.
+
+Нельзя передавать в candidates hidden tools, raw shell command, credential, неограниченную историю или capability, для которой у Hermes нет native executor.
+
+### `jev_record_execution`
+
+Принимает результат уже выполненного/отклонённого host action. `correlation_id` должен принадлежать decision того же gateway process. Статусы: `completed`, `failed`, `not_started`.
+
+Receipt — evidence, не источник авторизации. В `result` нельзя писать токены, пароли, номера карт и полный неочищенный лог.
+
+### `jev_model_route`
+
+Возвращает рекомендацию одного из **host-declared** `models[]`. Всегда `route_mode: "shadow"`; он не меняет Hermes provider, model или reasoning effort.
+
+Когда host закончил реальный model call, он записывает outcome через `jev_record_execution`:
+
+```json
+{
+ "correlation_id": "из jev_model_route",
+ "capability_id": "recommended model id",
+ "status": "completed",
+ "result": {
+ "model_route": {
+ "actual_model_id": "фактически использованный id",
+ "retry_count": 0,
+ "outcome": "verified"
+ }
+ }
+}
+```
+
+Отчёт только читает receipts:
+
+```bash
+npm run model-route:report -- ~/.hermes/jev-layer/replay/cases.jsonl
+```
+
+Пока не накоплены репрезентативные outcomes/retries/latency/cost, **не включать** auto-switching моделей.
+
+### `jev_browser_step`
+
+Принимает только snapshot, который сделал host: `url`, visible text, targets, tabs, scroll и progress. Возвращает один bounded action: scroll, switch tab, visible link navigation, click, native select или handoff.
+
+- `scroll` и `switch_tab` могут быть исполнены только через native Hermes browser executor.
+- `click`, `select`, `navigate` требуют host approval.
+- `submit`, login, payment, delete, секретный ввод и `TYPE_TEXT` не являются быстрым Jev execution path.
+- Progress не даёт повторять уже выполненный target и блокирует повторный scroll без измеримого состояния страницы.
+- После каждого host шага можно записать routing case и execution receipt; это позволяет replay без повторного browser execution.
+
+`browser-use/jev-ultrafast` был использован как архитектурный reference (динамическая индексированная action space), но его cloud/browser worker **не установлен и не запущен**. Не подменяй Hermes browser security model его executor'ом.
+
+### `jev_shadow_compaction`
+
+Возвращает консервативный report keep/drop для переданного context. Не меняет prompt, не удаляет сообщения и не заменяет host compaction. Pinned requirements, paths, errors и commands сохраняются без provider review; provider error удерживает всё.
+
+### `jev_supervise`
+
+Делает bounded judgement о work state. Hermes, не Jev, преобразует его в `continue`, `verify`, `retry`, `finish` или `escalate`. Contradictory evidence не может закончиться `finish`.
+
+## Secrets и providers
+
+`integrations/hermes/__init__.py` получает `OPENROUTER_API_KEY` только через Hermes profile-scoped secret store (`agent.secret_scope.get_secret`). Ключ передаётся лишь в короткоживущий локальный Node adapter; не возвращается в tool output и не логируется.
+
+Для локальных/offline tests используй `provider: "demo"`. Provider-backed проверки — отдельные, opt-in; ключи никогда не клади в repo, plugin manifest, test fixture или обычный shell history.
+
+## Безопасный change workflow
+
+1. Прочитай этот документ и `docs/AGENT-IMPLEMENTATION.md`.
+2. Проверь target surface: source checkout, installed plugin copy, active profile или global config.
+3. Измени source в `/home/hermes/jev-layer` и сначала добавь/измени offline test.
+4. Запусти узкий test, затем полный suite.
+5. Синхронизируй **только нужные** source files в `~/.hermes/plugins/jev-layer`.
+6. Запусти `hermes plugins doctor jev-layer`.
+7. Если менялась регистрация/handler/schema/manifest, попроси пользователя о `/restart`.
+8. После restart проверь живой gateway tool path и read back exact receipt/report.
+9. Не делай `git push`, npm publish, GitHub release или изменение глобального Hermes config без отдельной команды пользователя.
+
+## Проверки
+
+Из source checkout:
+
+```bash
+npm test
+npm run receipt:mcp-smoke
+npm run browser:e2e
+npm run model-route:report -- /tmp/nonexistent-cases.jsonl
+python3 -m py_compile integrations/hermes/__init__.py integrations/hermes/schemas.py
+git diff --check
+hermes plugins doctor jev-layer
+```
+
+Текущее доказанное состояние: `npm test` — 35/35; MCP receipt smoke и browser fixture E2E проходят; gateway-live model route был записан и успешно связан с execution receipt/report. Это не доказательство автоматического роутинга моделей и не утверждение о real-browser speedup.
+
+## Готовый prompt локальному агенту
+
+```text
+Работай только с jev-layer в указанном target surface. Сначала прочитай:
+- docs/HERMES-LOCAL-AGENT-HANDOFF.ru.md
+- docs/AGENT-IMPLEMENTATION.md
+
+Сохрани host ownership: Jev даёт bounded advisory decision; Hermes владеет registry,
+permissions, approval, execution, retry/recovery и финальным результатом. Не добавляй
+автоматическое переключение моделей, browser executor, глобальный config write, provider
+credential или network publication без отдельного явного требования.
+
+Перед изменением назови source files, installed plugin files и gateway impact. Используй TDD:
+сначала узкий offline test, затем implementation, потом npm test + релевантные smoke tests.
+Если touch'нуты registration/handler/schema/manifest, синхронизируй plugin copy, прогоняй
+hermes plugins doctor jev-layer и попроси /restart. После рестарта сделай live tool call,
+запиши/прочитай exact receipt и только затем заявляй, что integration активна.
+
+Не делай git push, npm publish, release или удаление replay/artifact без отдельной команды.
+```
diff --git a/docs/SCHEMA-VERSIONING.md b/docs/SCHEMA-VERSIONING.md
index 37c0cfe..31ec475 100644
--- a/docs/SCHEMA-VERSIONING.md
+++ b/docs/SCHEMA-VERSIONING.md
@@ -21,6 +21,8 @@ The MCP server currently exposes these stable tool names:
- `jev_route`
- `jev_browser_step` (experimental, opt-in)
- `jev_supervise` (experimental, opt-in)
+- `jev_model_route` (report-only; experimental, opt-in)
+- `jev_shadow_compaction` (report-only; experimental, opt-in)
- `jev_record_execution`
Their input and structured output are v1. Additive optional properties are compatible. Renaming a tool, changing a required property, changing the meaning of a status, or changing who owns execution requires a new tool/schema version and adapter migration. Keep `tools/list`, `initialize`, and stdio JSON-RPC behavior backward compatible for v1 clients.
@@ -66,7 +68,7 @@ A v1 replay file is JSONL. Supported records include:
- `record_type: "execution_receipt"`
- `record_type: "supervision_case"`
-Routing cases retain a sanitized request and a decision summary. Replay reads cases without calling host tools. New optional record fields are compatible; changing record type, correlation semantics, or the meaning of a recorded status requires a new replay schema and migration.
+Routing cases retain a sanitized request and a decision summary. `supervision_case.supervision.evidence_state` is an optional v1 object with `state` (`present`, `missing`, or `contradictory`) and bounded source labels. Replay reads cases without calling host tools. New optional record fields are compatible; changing record type, correlation semantics, or the meaning of a recorded status requires a new replay schema and migration.
## Adapter contract
diff --git a/integrations/hermes/__init__.py b/integrations/hermes/__init__.py
index db0c10c..cb20d4b 100644
--- a/integrations/hermes/__init__.py
+++ b/integrations/hermes/__init__.py
@@ -16,7 +16,7 @@
from pathlib import Path
from typing import Any
-from .schemas import BROWSER_STEP, RECORD_EXECUTION, ROUTE, SUPERVISE
+from .schemas import BROWSER_STEP, MODEL_ROUTE, RECORD_EXECUTION, ROUTE, SHADOW_COMPACTION, SUPERVISE
ROOT = Path(os.environ.get("JEV_LAYER_ROOT", Path(__file__).resolve().parents[2])).expanduser().resolve()
CORE = ROOT / "src" / "hermes-adapter.mjs"
@@ -38,6 +38,17 @@ def _fallback(reason: str) -> str:
def _runtime_env() -> dict[str, str]:
env = os.environ.copy()
+ # Resolve only through Hermes's profile-scoped store, then pass it only to
+ # this short-lived local Node adapter. It is never returned or logged.
+ try:
+ from agent.secret_scope import get_secret
+ openrouter_key = get_secret("OPENROUTER_API_KEY")
+ if openrouter_key:
+ env["OPENROUTER_API_KEY"] = openrouter_key
+ except Exception:
+ # The core remains fail-open when the optional provider credential is unavailable.
+ pass
+ # Keep replay evidence in the active Hermes profile, never in the source checkout.
env.setdefault("JEV_REPLAY_CASES", str(Path(env.get("HERMES_HOME", "~/.hermes")).expanduser() / "jev-layer" / "replay" / "cases.jsonl"))
return env
@@ -101,6 +112,16 @@ def jev_browser_step(args: dict, **kwargs) -> str:
return _route(args, operation="browser_step")
+def jev_model_route(args: dict, **kwargs) -> str:
+ """Recommend a declared model profile; this never changes Hermes's model route."""
+ return _route(args, operation="model_route")
+
+
+def jev_shadow_compaction(args: dict, **kwargs) -> str:
+ """Report Jev's conservative keep/drop candidates without mutating host context."""
+ return json.dumps(_call_core("shadow_compaction", dict(args)))
+
+
def jev_supervise(args: dict, **kwargs) -> str:
"""Return a bounded work-state judgment; Hermes maps it to its own next action."""
request = dict(args)
@@ -122,6 +143,8 @@ def jev_record_execution(args: dict, **kwargs) -> str:
def register(ctx):
ctx.register_tool(name="jev_route", toolset="jev_layer", schema=ROUTE, handler=jev_route)
+ ctx.register_tool(name="jev_model_route", toolset="jev_layer", schema=MODEL_ROUTE, handler=jev_model_route)
+ ctx.register_tool(name="jev_shadow_compaction", toolset="jev_layer", schema=SHADOW_COMPACTION, handler=jev_shadow_compaction)
ctx.register_tool(name="jev_record_execution", toolset="jev_layer", schema=RECORD_EXECUTION, handler=jev_record_execution)
ctx.register_tool(name="jev_supervise", toolset="jev_layer", schema=SUPERVISE, handler=jev_supervise)
ctx.register_tool(name="jev_browser_step", toolset="jev_layer", schema=BROWSER_STEP, handler=jev_browser_step)
diff --git a/integrations/hermes/plugin.yaml b/integrations/hermes/plugin.yaml
index babb00e..8f87fd3 100644
--- a/integrations/hermes/plugin.yaml
+++ b/integrations/hermes/plugin.yaml
@@ -1,8 +1,10 @@
name: jev-layer
version: 0.1.0
-description: Full host-owned Jev routing, receipts, supervision, and browser decisions
+description: Full host-owned Jev routing, receipts, supervision, shadow compaction, model routing, and browser decisions
provides_tools:
- jev_route
+ - jev_model_route
+ - jev_shadow_compaction
- jev_record_execution
- jev_supervise
- jev_browser_step
diff --git a/integrations/hermes/schemas.py b/integrations/hermes/schemas.py
index 851920d..ea73c10 100644
--- a/integrations/hermes/schemas.py
+++ b/integrations/hermes/schemas.py
@@ -31,6 +31,32 @@
},
}
+MODEL_ROUTE = {
+ "type": "object",
+ "required": ["intent", "models"],
+ "properties": {
+ "intent": {"type": "string"},
+ "context": {"type": "object"},
+ "models": {"type": "array", "minItems": 1, "items": {"type": "object"}},
+ "harness": {"type": "string"},
+ "policy": {"type": "object"},
+ "provider": {"type": "string", "enum": ["demo", "typesafe", "openrouter"]},
+ },
+}
+
+SHADOW_COMPACTION = {
+ "type": "object",
+ "required": ["intent", "context"],
+ "properties": {
+ "intent": {"type": "string"},
+ "context": {"type": "object"},
+ "provider": {"type": "string", "enum": ["demo", "typesafe", "openrouter"]},
+ "batch_size": {"type": "integer", "minimum": 1, "maximum": 8},
+ "keep_threshold": {"type": "number", "minimum": 0, "maximum": 1},
+ "min_confidence": {"type": "number", "minimum": 0, "maximum": 1},
+ },
+}
+
SUPERVISE = {
"type": "object",
"required": ["job", "observation"],
diff --git a/package.json b/package.json
index 1c6a72a..ffc17af 100644
--- a/package.json
+++ b/package.json
@@ -48,6 +48,7 @@
"codex:mcp-smoke": "node scripts/codex-mcp-smoke.mjs",
"replay:evaluate": "node scripts/replay-eval.mjs",
"receipt:mcp-smoke": "node scripts/mcp-receipt-smoke.mjs",
+ "model-route:report": "node scripts/model-route-report.mjs",
"browser:e2e": "node scripts/browser-e2e.mjs",
"browser:benchmark": "node scripts/browser-benchmark.mjs",
"supervision:e2e": "node scripts/supervision-e2e.mjs",
diff --git a/plugin.yaml b/plugin.yaml
index babb00e..8f87fd3 100644
--- a/plugin.yaml
+++ b/plugin.yaml
@@ -1,8 +1,10 @@
name: jev-layer
version: 0.1.0
-description: Full host-owned Jev routing, receipts, supervision, and browser decisions
+description: Full host-owned Jev routing, receipts, supervision, shadow compaction, model routing, and browser decisions
provides_tools:
- jev_route
+ - jev_model_route
+ - jev_shadow_compaction
- jev_record_execution
- jev_supervise
- jev_browser_step
diff --git a/scripts/mcp-receipt-smoke.mjs b/scripts/mcp-receipt-smoke.mjs
index 1b31e03..99b4919 100644
--- a/scripts/mcp-receipt-smoke.mjs
+++ b/scripts/mcp-receipt-smoke.mjs
@@ -30,6 +30,8 @@ try {
await notification("notifications/initialized", {});
const listed = await request(2, "tools/list", {});
assert.ok(listed.result.tools.some((tool) => tool.name === "jev_route"));
+ assert.ok(listed.result.tools.some((tool) => tool.name === "jev_model_route"));
+ assert.ok(listed.result.tools.some((tool) => tool.name === "jev_shadow_compaction"));
assert.ok(listed.result.tools.some((tool) => tool.name === "jev_record_execution"));
const routed = await request(3, "tools/call", {
@@ -70,6 +72,19 @@ try {
assert.equal(receipt.host.exit_status, 0);
assert.equal(receipt.host.duration_ms, 3.2);
+ const shadow = await request(5, "tools/call", {
+ name: "jev_shadow_compaction",
+ arguments: {
+ provider: "demo",
+ intent: "inspect context before compaction",
+ context: { messages: ["stale note", "Requirement: preserve /workspace/config.json"] },
+ },
+ });
+ const shadowReport = shadow.result.structuredContent;
+ assert.equal(shadowReport.mode, "jev_shadow");
+ assert.equal(shadowReport.changed, false);
+ assert.equal(shadowReport.protected_count, 1);
+
const records = (await readFile(casesPath, "utf8")).trim().split(/\r?\n/).map(JSON.parse);
assert.deepEqual(records.map((record) => record.record_type), ["routing_case", "execution_receipt"]);
console.log(JSON.stringify({ ok: true, cases_path: casesPath, receipt }, null, 2));
diff --git a/scripts/model-route-report.mjs b/scripts/model-route-report.mjs
new file mode 100644
index 0000000..e800846
--- /dev/null
+++ b/scripts/model-route-report.mjs
@@ -0,0 +1,15 @@
+#!/usr/bin/env node
+import { readFile } from "node:fs/promises";
+import { buildModelRouteReport } from "../src/model-route-metrics.mjs";
+
+const path = process.argv[2] ?? process.env.JEV_REPLAY_CASES ?? ".jev/replay/cases.jsonl";
+let text = "";
+try {
+ text = await readFile(path, "utf8");
+} catch (error) {
+ if (error?.code !== "ENOENT") throw error;
+}
+const records = text.split(/\r?\n/).filter(Boolean).flatMap((line) => {
+ try { return [JSON.parse(line)]; } catch { return []; }
+});
+console.log(JSON.stringify(buildModelRouteReport(records), null, 2));
diff --git a/scripts/supervision-e2e.mjs b/scripts/supervision-e2e.mjs
index 1eb0029..39fc63a 100644
--- a/scripts/supervision-e2e.mjs
+++ b/scripts/supervision-e2e.mjs
@@ -66,9 +66,38 @@ try {
assert.equal(result.metrics.jev_calls, 1);
assert.ok(result.receipt.correlation_id);
+ const contradictory = await request(5, "tools/call", {
+ name: "jev_supervise",
+ arguments: {
+ enabled: true,
+ provider: process.env.JEV_SUPERVISION_PROVIDER ?? "demo",
+ harness: "supervision-e2e",
+ job: { requirements: ["run tests"] },
+ observation: { status: "claimed-complete" },
+ evidence: {
+ tests_passed: true,
+ tests_failed: true,
+ judgments: {
+ requirements_addressed: 0.9,
+ verification_needed: 0.1,
+ meaningful_progress: 0.9,
+ worker_stuck: 0.05,
+ work_off_track: 0.05,
+ completion: 0.9,
+ },
+ },
+ },
+ });
+ const contradictoryResult = contradictory.result.structuredContent;
+ assert.equal(contradictoryResult.status, "judged");
+ assert.equal(contradictoryResult.action, "verify");
+ assert.equal(contradictoryResult.evidence_state.state, "contradictory");
+
const records = (await readFile(casesPath, "utf8")).trim().split(/\r?\n/).map(JSON.parse);
- assert.deepEqual(records.map((record) => record.record_type), ["supervision_case"]);
+ assert.deepEqual(records.map((record) => record.record_type), ["supervision_case", "supervision_case"]);
assert.equal(records[0].supervision.action, "finish");
+ assert.equal(records[1].supervision.action, "verify");
+ assert.equal(records[1].supervision.evidence_state.state, "contradictory");
const replayed = deterministicSupervisionPolicy({
assessment: records[0].supervision.assessment,
evidence: records[0].request.context.evidence,
diff --git a/src/browser.mjs b/src/browser.mjs
index 668cf36..f63773b 100644
--- a/src/browser.mjs
+++ b/src/browser.mjs
@@ -125,6 +125,7 @@ export async function runBrowserFastPath({
harness = "browser",
start_url = null,
provider = "demo",
+ config,
enabled,
maxSteps = 8,
maxSeconds = 30,
@@ -169,7 +170,7 @@ export async function runBrowserFastPath({
if (elapsed(started) > maxSeconds * 1_000) return finish("handoff", "browser_fast_path_timeout", observation);
let routed;
try {
- routed = await decideBrowserStep({ goal, observation, harness, start_url, progress }, { enabled: true, provider, policy });
+ routed = await decideBrowserStep({ goal, observation, harness, start_url, progress }, { enabled: true, provider, config, policy });
} catch (error) {
metrics.failures += 1;
return finish("handoff", "browser_decision_failed", observation);
diff --git a/src/hermes-adapter.mjs b/src/hermes-adapter.mjs
index 3050047..1616024 100644
--- a/src/hermes-adapter.mjs
+++ b/src/hermes-adapter.mjs
@@ -12,6 +12,8 @@ import { configuredProvider, configuredReplayPath, loadConfig } from "./config.m
import { appendExecutionReceipt, appendRoutingCase, buildExecutionReceipt, replayCasePath } from "./receipts.mjs";
import { routeRequest } from "./route.mjs";
import { superviseWork } from "./supervision.mjs";
+import { buildShadowCompactionReport } from "./shadow-compaction.mjs";
+import { recommendModelRoute } from "./model-routing.mjs";
const { config } = await loadConfig();
const casesPath = replayCasePath(configuredReplayPath(config));
@@ -34,6 +36,7 @@ async function handle({ operation, args = {}, decision = null } = {}) {
const { provider: _provider, engine = "native", ...request } = args;
const result = await routeRequest(request, {
provider,
+ config,
engine,
contextFilterMode: request.policy?.context_filter_mode ?? config.features?.context_filter,
});
@@ -43,8 +46,9 @@ async function handle({ operation, args = {}, decision = null } = {}) {
if (operation === "browser_step") {
const { provider: _provider, enabled, ...input } = args;
const routed = await decideBrowserStep(input, {
- enabled: enabled ?? config.features?.browser_fast_path,
- provider,
+ enabled: enabled ?? config.features?.browser_fast_path,
+ provider,
+ config,
});
routed.decision.browser_action = routed.action
? { id: routed.action.id, operation: routed.action.operation, target_id: routed.action.target_id ?? null,
@@ -57,11 +61,36 @@ async function handle({ operation, args = {}, decision = null } = {}) {
const { provider: _provider, enabled, ...input } = args;
return superviseWork({
...input,
+ config,
enabled: enabled ?? config.features?.supervision,
provider,
receiptPath: casesPath,
});
}
+ if (operation === "model_route") {
+ const { provider: _provider, ...input } = args;
+ const result = await recommendModelRoute({ ...input, provider, config });
+ await persistRoutingCase({
+ schema_version: 1,
+ harness: input.harness ?? "hermes",
+ intent: input.intent,
+ context: { ...(input.context ?? {}), model_route: { mode: "shadow" } },
+ capabilities: (input.models ?? []).map((profile) => ({
+ id: profile.id,
+ kind: "model",
+ name: profile.model ?? profile.id,
+ description: profile.description ?? "",
+ source: profile.provider ?? null,
+ risk: "low",
+ })),
+ policy: input.policy ?? {},
+ }, result);
+ return result;
+ }
+ if (operation === "shadow_compaction") {
+ const { provider: _provider, ...input } = args;
+ return buildShadowCompactionReport({ ...input, provider, config });
+ }
if (operation === "record_execution") {
if (!decision || typeof decision !== "object") throw new TypeError("decision is required for record_execution");
if (args.correlation_id !== decision.correlation_id) throw new TypeError("correlation_id does not match the routed decision");
diff --git a/src/mcp-server.mjs b/src/mcp-server.mjs
index e311528..92c1be6 100755
--- a/src/mcp-server.mjs
+++ b/src/mcp-server.mjs
@@ -4,6 +4,8 @@ import { appendExecutionReceipt, appendRoutingCase, buildExecutionReceipt, repla
import { configuredProvider, configuredReplayPath, loadConfig } from "./config.mjs";
import { decideBrowserStep } from "./browser.mjs";
import { superviseWork } from "./supervision.mjs";
+import { buildShadowCompactionReport } from "./shadow-compaction.mjs";
+import { recommendModelRoute } from "./model-routing.mjs";
import { routeRequest } from "./route.mjs";
const { config } = await loadConfig();
@@ -65,6 +67,38 @@ const tools = [
},
},
},
+ {
+ name: "jev_model_route",
+ description: "Recommend one host-declared model profile for the next call. This is shadow-only: it never changes provider, model, reasoning, or execution.",
+ inputSchema: {
+ type: "object",
+ required: ["intent", "models"],
+ properties: {
+ intent: { type: "string" },
+ context: { type: "object" },
+ models: { type: "array", minItems: 1, items: { type: "object" } },
+ harness: { type: "string" },
+ policy: { type: "object" },
+ provider: { type: "string", enum: ["demo", "typesafe", "openrouter"] },
+ },
+ },
+ },
+ {
+ name: "jev_shadow_compaction",
+ description: "Report conservative Jev keep/drop candidates for host-supplied context. This tool never changes, summarizes, or deletes context.",
+ inputSchema: {
+ type: "object",
+ required: ["intent", "context"],
+ properties: {
+ intent: { type: "string" },
+ context: { type: "object" },
+ provider: { type: "string", enum: ["demo", "typesafe", "openrouter"] },
+ batch_size: { type: "integer", minimum: 1, maximum: 8 },
+ keep_threshold: { type: "number", minimum: 0, maximum: 1 },
+ min_confidence: { type: "number", minimum: 0, maximum: 1 },
+ },
+ },
+ },
{
name: "jev_record_execution",
description: "Attach a host execution result to a Jev decision and persist the unified execution receipt.",
@@ -115,6 +149,8 @@ async function handle(message) {
if (params.name === "jev_route") return route(args, id);
if (params.name === "jev_browser_step") return browserStep(args, id);
if (params.name === "jev_supervise") return supervise(args, id);
+ if (params.name === "jev_model_route") return modelRoute(args, id);
+ if (params.name === "jev_shadow_compaction") return shadowCompaction(args, id);
if (params.name === "jev_record_execution") return recordExecution(args, id);
return jsonRpcError(id, -32602, `unknown tool: ${params.name}`);
}
@@ -162,6 +198,24 @@ async function supervise(args, id) {
return toolResult(id, result, false);
}
+async function modelRoute(args, id) {
+ const { provider: requestedProvider, ...input } = args;
+ const result = await recommendModelRoute({
+ ...input,
+ provider: configuredProvider(config, requestedProvider),
+ });
+ return toolResult(id, result, false);
+}
+
+async function shadowCompaction(args, id) {
+ const { provider: requestedProvider, ...input } = args;
+ const result = await buildShadowCompactionReport({
+ ...input,
+ provider: configuredProvider(config, requestedProvider),
+ });
+ return toolResult(id, result, false);
+}
+
async function persistDecision(request, decision, id) {
pendingDecisions.set(decision.correlation_id, { request, decision });
try {
diff --git a/src/model-route-metrics.mjs b/src/model-route-metrics.mjs
new file mode 100644
index 0000000..50bf75a
--- /dev/null
+++ b/src/model-route-metrics.mjs
@@ -0,0 +1,89 @@
+/**
+ * Read-only analysis for model-route shadow cases. Recommendations stay
+ * advisory: this module joins evidence; it never changes host model settings.
+ */
+export function buildModelRouteReport(records = []) {
+ const routes = new Map();
+ const executions = new Map();
+ for (const record of Array.isArray(records) ? records : []) {
+ if (!record || typeof record !== "object" || typeof record.correlation_id !== "string") continue;
+ if (record.record_type === "routing_case" && record.request?.context?.model_route?.mode === "shadow") routes.set(record.correlation_id, record);
+ if (record.record_type === "execution_receipt") executions.set(record.correlation_id, record);
+ }
+
+ const rows = [...routes.values()].map((route) => toRow(route, executions.get(route.correlation_id))).sort((a, b) => a.recommended_model_id.localeCompare(b.recommended_model_id));
+ return {
+ schema_version: 1,
+ mode: "shadow",
+ generated_at: new Date().toISOString(),
+ summary: summarize(rows),
+ models: groupByRecommendation(rows),
+ };
+}
+
+function toRow(route, execution) {
+ const result = execution?.host?.result?.model_route;
+ const recommended = string(route.decision?.selected, "unrecommended");
+ const actual = string(result?.actual_model_id, null);
+ const outcome = string(result?.outcome, execution ? string(execution.host?.status, "unknown") : null);
+ return {
+ correlation_id: route.correlation_id,
+ recommended_model_id: recommended,
+ actual_model_id: actual,
+ recommendation_match: actual !== null && actual === recommended,
+ outcome,
+ retry_count: boundedInteger(result?.retry_count),
+ host_duration_ms: finite(execution?.host?.duration_ms),
+ jev_latency_ms: finite(route.decision?.receipt?.latency_ms),
+ jev_cost_usd: finite(route.decision?.receipt?.cost_usd),
+ };
+}
+
+function summarize(rows) {
+ const matched = rows.filter((row) => row.recommendation_match).length;
+ const actual = rows.filter((row) => row.actual_model_id !== null).length;
+ return {
+ decisions: rows.length,
+ with_host_outcome: rows.filter((row) => row.outcome !== null).length,
+ with_actual_model: actual,
+ recommendation_matches: matched,
+ recommendation_match_rate: actual ? matched / actual : null,
+ verified_outcomes: rows.filter((row) => row.outcome === "verified" || row.outcome === "completed").length,
+ failed_outcomes: rows.filter((row) => row.outcome === "failed").length,
+ total_retry_count: rows.reduce((sum, row) => sum + (row.retry_count ?? 0), 0),
+ total_jev_cost_usd: sum(rows.map((row) => row.jev_cost_usd)),
+ average_jev_latency_ms: average(rows.map((row) => row.jev_latency_ms)),
+ average_host_duration_ms: average(rows.map((row) => row.host_duration_ms)),
+ };
+}
+
+function groupByRecommendation(rows) {
+ const grouped = new Map();
+ for (const row of rows) {
+ const group = grouped.get(row.recommended_model_id) ?? {
+ recommended_model_id: row.recommended_model_id,
+ decisions: 0,
+ with_host_outcome: 0,
+ recommendation_matches: 0,
+ verified_outcomes: 0,
+ failed_outcomes: 0,
+ total_retry_count: 0,
+ host_actual_model_ids: {},
+ };
+ group.decisions += 1;
+ if (row.outcome !== null) group.with_host_outcome += 1;
+ if (row.recommendation_match) group.recommendation_matches += 1;
+ if (row.outcome === "verified" || row.outcome === "completed") group.verified_outcomes += 1;
+ if (row.outcome === "failed") group.failed_outcomes += 1;
+ group.total_retry_count += row.retry_count ?? 0;
+ if (row.actual_model_id) group.host_actual_model_ids[row.actual_model_id] = (group.host_actual_model_ids[row.actual_model_id] ?? 0) + 1;
+ grouped.set(row.recommended_model_id, group);
+ }
+ return [...grouped.values()].sort((a, b) => a.recommended_model_id.localeCompare(b.recommended_model_id));
+}
+
+function string(value, fallback) { return typeof value === "string" && value.trim() ? value.trim() : fallback; }
+function finite(value) { return typeof value === "number" && Number.isFinite(value) && value >= 0 ? value : null; }
+function boundedInteger(value) { return Number.isInteger(value) && value >= 0 ? value : null; }
+function sum(values) { return values.reduce((total, value) => total + (value ?? 0), 0); }
+function average(values) { const present = values.filter((value) => value !== null); return present.length ? sum(present) / present.length : null; }
diff --git a/src/model-routing.mjs b/src/model-routing.mjs
new file mode 100644
index 0000000..fc63221
--- /dev/null
+++ b/src/model-routing.mjs
@@ -0,0 +1,87 @@
+import { routeRequest } from "./route.mjs";
+
+/**
+ * Recommend one host-declared model profile for the next model call.
+ * This is a shadow decision: it never mutates the host model/provider setting.
+ */
+export async function recommendModelRoute({
+ intent,
+ context = {},
+ models,
+ harness = "unknown",
+ policy = {},
+ provider = "demo",
+ config,
+} = {}) {
+ if (!Array.isArray(models) || models.length === 0) throw new TypeError("models[] is required");
+ const profiles = models.map(normalizeProfile);
+ const profileById = new Map(profiles.map((profile) => [profile.id, profile]));
+ const decision = await routeRequest({
+ harness,
+ intent,
+ context: { ...context, model_route: { mode: "shadow", profiles: profiles.map(publicProfile) } },
+ capabilities: profiles.map(toCapability),
+ policy: {
+ ...policy,
+ // A recommendation must never create an approval boundary by itself.
+ confirmation_risk_levels: [],
+ },
+ }, { provider, config });
+ return {
+ ...decision,
+ route_mode: "shadow",
+ recommended_model: decision.selected ? publicProfile(profileById.get(decision.selected)) : null,
+ };
+}
+
+function normalizeProfile(value, index) {
+ if (!value || typeof value !== "object") throw new TypeError(`models[${index}] must be an object`);
+ const id = required(value.id, `models[${index}].id`);
+ const model = required(value.model ?? value.id, `models[${index}].model`);
+ const provider = string(value.provider, "unknown");
+ const reasoning_effort = string(value.reasoning_effort, "default");
+ const description = string(value.description, `${provider}/${model}; reasoning=${reasoning_effort}`);
+ return {
+ id,
+ provider,
+ model,
+ reasoning_effort,
+ description,
+ available: value.available !== false && value.availability?.available !== false,
+ availability_reason: value.availability?.reason ?? null,
+ };
+}
+
+function toCapability(profile) {
+ return {
+ id: profile.id,
+ kind: "model",
+ name: profile.model,
+ description: profile.description,
+ risk: "low",
+ available: profile.available,
+ availability: { available: profile.available, reason: profile.availability_reason },
+ source: profile.provider,
+ metadata: { model: profile.model },
+ };
+}
+
+function publicProfile(profile) {
+ if (!profile) return null;
+ return {
+ id: profile.id,
+ provider: profile.provider,
+ model: profile.model,
+ reasoning_effort: profile.reasoning_effort,
+ available: profile.available,
+ };
+}
+
+function required(value, label) {
+ if (typeof value !== "string" || !value.trim()) throw new TypeError(`${label} is required`);
+ return value.trim();
+}
+
+function string(value, fallback) {
+ return typeof value === "string" && value.trim() ? value.trim() : fallback;
+}
diff --git a/src/receipts.mjs b/src/receipts.mjs
index 6a69d17..1693dea 100644
--- a/src/receipts.mjs
+++ b/src/receipts.mjs
@@ -101,6 +101,7 @@ export function buildSupervisionReceipt({ request, result, recordedAt = new Date
action: result.action ?? "continue",
reason: result.reason ?? null,
assessment: result.assessment ?? null,
+ evidence_state: result.evidence_state ?? request.context?.supervision?.evidence_state ?? null,
policy: result.policy ?? null,
jev: {
provider: result.receipt.provider ?? null,
diff --git a/src/shadow-compaction.mjs b/src/shadow-compaction.mjs
new file mode 100644
index 0000000..193b365
--- /dev/null
+++ b/src/shadow-compaction.mjs
@@ -0,0 +1,170 @@
+import { sanitizeContext } from "./context-filter.mjs";
+import { isPinnedEvidence } from "./relevance-filter.mjs";
+import { resolveProvider } from "./providers/index.mjs";
+
+const LIST_KEYS = new Set(["messages", "events", "logs", "tool_results", "history", "transcript"]);
+const DEFAULT_BATCH_SIZE = 8;
+const DEFAULT_KEEP_THRESHOLD = 0.5;
+const DEFAULT_MIN_CONFIDENCE = 0.7;
+const MAX_ITEMS = 64;
+const MAX_PREVIEW = 320;
+
+/**
+ * Produce advisory Jev compaction decisions without mutating context.
+ * Pinned evidence is never sent as droppable and provider failure retains all items.
+ */
+export async function buildShadowCompactionReport({
+ intent = "Assess which context records must remain available to a later agent.",
+ context = {},
+ provider = "demo",
+ config,
+ batchSize = DEFAULT_BATCH_SIZE,
+ keepThreshold = DEFAULT_KEEP_THRESHOLD,
+ minConfidence = DEFAULT_MIN_CONFIDENCE,
+} = {}) {
+ const items = collectContextItems(context);
+ const normalizedBatchSize = boundedInteger(batchSize, DEFAULT_BATCH_SIZE, 1, DEFAULT_BATCH_SIZE);
+ const normalizedThreshold = boundedNumber(keepThreshold, DEFAULT_KEEP_THRESHOLD);
+ const normalizedConfidence = boundedNumber(minConfidence, DEFAULT_MIN_CONFIDENCE);
+ const protectedItems = items.filter((item) => item.pinned);
+ const evaluable = items.filter((item) => !item.pinned);
+ const decisions = new Map(protectedItems.map((item) => [item.id, {
+ decision: "keep",
+ probability: 1,
+ confidence: 1,
+ reason: `pinned:${item.pinned_reasons.join(",")}`,
+ }]));
+ let costUsd = 0;
+ let failed = null;
+
+ try {
+ const resolved = resolveProvider(provider, { config });
+ if (!resolved || typeof resolved.evaluate !== "function") throw new Error("provider does not support batched evaluation");
+ for (const batch of batches(evaluable, normalizedBatchSize)) {
+ const raw = await resolved.evaluate({
+ state: {
+ intent: boundedText(intent),
+ compaction: {
+ mode: "shadow",
+ instruction: "For each item, decide whether dropping it would lose a requirement, exact identifier, path, command, error, user decision, execution receipt, or other evidence needed later. Be conservative: keep when uncertain.",
+ items: batch.map(publicItem),
+ },
+ },
+ questions: Object.fromEntries(batch.map((item, index) => [`item_${index}`, keepQuestion()])),
+ });
+ costUsd += finiteCost(raw?.usage?.cost) ?? 0;
+ for (const [index, item] of batch.entries()) {
+ const answer = readNoul(raw?.answers?.[`item_${index}`]);
+ if (!answer) throw new Error(`provider response has no valid item_${index} judgment`);
+ const decision = answer.confidence < normalizedConfidence || answer.probability >= normalizedThreshold ? "keep" : "drop_candidate";
+ decisions.set(item.id, {
+ decision,
+ probability: answer.probability,
+ confidence: answer.confidence,
+ reason: answer.confidence < normalizedConfidence ? "low_confidence" : decision === "keep" ? "provider_keep" : "provider_drop_candidate",
+ });
+ }
+ }
+ } catch (error) {
+ failed = error instanceof Error ? error.message : String(error);
+ for (const item of evaluable) {
+ decisions.set(item.id, { decision: "keep", probability: null, confidence: null, reason: "provider_error" });
+ }
+ }
+
+ const reportedItems = items.map((item) => ({ ...publicItem(item), ...decisions.get(item.id) }));
+ const droppableCount = reportedItems.filter((item) => item.decision === "drop_candidate").length;
+ return {
+ schema_version: 1,
+ mode: "jev_shadow",
+ status: failed ? "fallback" : "judged",
+ changed: false,
+ candidate_count: reportedItems.length,
+ evaluated_count: evaluable.length,
+ protected_count: protectedItems.length,
+ retained_count: reportedItems.length - droppableCount,
+ droppable_count: droppableCount,
+ keep_threshold: normalizedThreshold,
+ min_confidence: normalizedConfidence,
+ cost_usd: failed ? null : Number(costUsd.toFixed(8)),
+ fallback: failed ? { type: "provider_error", reason: failed } : null,
+ items: reportedItems,
+ };
+}
+
+export function collectContextItems(context = {}) {
+ const source = context && typeof context === "object" ? context : {};
+ const items = [];
+ for (const [list, value] of Object.entries(source)) {
+ if (!LIST_KEYS.has(list) || !Array.isArray(value)) continue;
+ for (const [index, original] of value.entries()) {
+ if (items.length >= MAX_ITEMS) return items;
+ const sanitized = sanitizeContext(original);
+ const evidence = isPinnedEvidence(sanitized);
+ items.push({
+ id: `${list}[${index}]`,
+ list,
+ index,
+ value: sanitized,
+ pinned: evidence.pinned,
+ pinned_reasons: evidence.reasons,
+ });
+ }
+ }
+ return items;
+}
+
+function publicItem(item) {
+ return {
+ id: item.id,
+ list: item.list,
+ index: item.index,
+ pinned: item.pinned,
+ pinned_reasons: item.pinned_reasons,
+ preview: boundedText(typeof item.value === "string" ? item.value : JSON.stringify(item.value)),
+ };
+}
+
+function keepQuestion() {
+ return {
+ type: "noul",
+ instructions: "Would dropping this context record lose information needed by a later agent? Return true to keep it and false only when it is safely stale or redundant.",
+ criteria: {
+ true: "Keep: it contains a requirement, decision, identifier, path, command, error, execution evidence, or context that may matter later; uncertainty means keep.",
+ false: "Drop candidate: it is safely stale or redundant and losing it would not affect later work.",
+ },
+ };
+}
+
+function readNoul(answer) {
+ if (!answer || typeof answer !== "object") return null;
+ const probability = ["probability", "p_true", "score", "value", "noul"].map((key) => answer[key]).find(validProbability);
+ if (!validProbability(probability)) return null;
+ const confidence = validProbability(answer.confidence) ? answer.confidence : probability;
+ return { probability, confidence };
+}
+
+function batches(items, size) {
+ return Array.from({ length: Math.ceil(items.length / size) }, (_, index) => items.slice(index * size, (index + 1) * size));
+}
+
+function boundedText(value) {
+ const text = String(value ?? "");
+ return text.length > MAX_PREVIEW ? `${text.slice(0, MAX_PREVIEW)}…` : text;
+}
+
+function boundedNumber(value, fallback) {
+ return typeof value === "number" && Number.isFinite(value) ? Math.max(0, Math.min(1, value)) : fallback;
+}
+
+function boundedInteger(value, fallback, min, max) {
+ return Number.isInteger(value) ? Math.max(min, Math.min(max, value)) : fallback;
+}
+
+function validProbability(value) {
+ return typeof value === "number" && Number.isFinite(value) && value >= 0 && value <= 1;
+}
+
+function finiteCost(value) {
+ return typeof value === "number" && Number.isFinite(value) && value >= 0 ? value : null;
+}
diff --git a/src/supervision.mjs b/src/supervision.mjs
index 561ba21..d3347c1 100644
--- a/src/supervision.mjs
+++ b/src/supervision.mjs
@@ -41,6 +41,7 @@ export function buildSupervisionRequest({
evidence: bounded(evidence),
supervision: {
judgments: bounded(evidence?.judgments ?? {}),
+ evidence_state: deterministicEvidenceState(evidence),
},
},
};
@@ -70,12 +71,42 @@ export function normalizeAssessment(raw) {
return assessment;
}
+export function deterministicEvidenceState(evidence = {}) {
+ const sources = [];
+ const positive = evidence?.verification_passed === true || evidence?.tests_passed === true;
+ const negative = evidence?.verification_failed === true || evidence?.tests_failed === true;
+ if (evidence?.verification_passed === true) sources.push("verification_passed");
+ if (evidence?.tests_passed === true) sources.push("tests_passed");
+ if (evidence?.verification_failed === true) sources.push("verification_failed");
+ if (evidence?.tests_failed === true) sources.push("tests_failed");
+ if (Array.isArray(evidence?.receipts)) {
+ for (const receipt of evidence.receipts) {
+ if (!receipt || typeof receipt !== "object") continue;
+ const id = typeof receipt.id === "string" && receipt.id ? receipt.id : null;
+ if (receipt.status === "completed") {
+ sources.push(id ? `receipt:${id}` : "receipt:completed");
+ } else if (receipt.status === "failed") {
+ sources.push(id ? `receipt_failed:${id}` : "receipt:failed");
+ }
+ }
+ }
+ const hasPositiveReceipt = sources.some((source) => source === "receipt:completed" || source.startsWith("receipt:"));
+ const hasNegativeReceipt = sources.some((source) => source === "receipt:failed" || source.startsWith("receipt_failed:"));
+ if ((positive || hasPositiveReceipt) && (negative || hasNegativeReceipt)) return { state: "contradictory", sources };
+ if (negative || hasNegativeReceipt) return { state: "contradictory", sources };
+ return { state: positive || hasPositiveReceipt ? "present" : "missing", sources };
+}
+
export function deterministicSupervisionPolicy({ assessment, evidence = {}, attempts = 0, policy = {} } = {}) {
const values = assessment && typeof assessment === "object" ? assessment : {};
const threshold = Number.isFinite(policy.threshold) ? Math.max(0, Math.min(1, policy.threshold)) : 0.7;
const maxRetries = Number.isInteger(policy.max_retries) && policy.max_retries >= 0 ? policy.max_retries : 2;
- const verificationEvidence = evidence.verification_passed === true || evidence.tests_passed === true;
+ const evidenceState = deterministicEvidenceState(evidence);
+ const verificationEvidence = evidenceState.state === "present";
+ if (evidenceState.state === "contradictory") {
+ return { action: "verify", reason: "verification evidence is contradictory", policy_source: "deterministic_host_policy", evidence_state: evidenceState };
+ }
if (value(values.worker_stuck) >= threshold) {
return attempts < maxRetries
? { action: "retry", reason: "worker appears stuck and retry budget remains", policy_source: "deterministic_host_policy" }
@@ -150,6 +181,7 @@ export async function superviseWork({
action: selectedPolicy.action,
reason: selectedPolicy.reason,
assessment,
+ evidence_state: deterministicEvidenceState(evidence),
policy: selectedPolicy,
request,
metrics: {
diff --git a/test/hermes-adapter.test.mjs b/test/hermes-adapter.test.mjs
index bfce125..9f48c9a 100644
--- a/test/hermes-adapter.test.mjs
+++ b/test/hermes-adapter.test.mjs
@@ -44,15 +44,19 @@ test("Hermes adapter routes a closed candidate set and records a correlated rece
assert.equal(records.filter((record) => record.record_type === "execution_receipt").length, 1);
});
-test("Hermes adapter exposes supervision and browser surfaces without execution", async () => {
+test("Hermes adapter exposes supervision, shadow-compaction, and browser surfaces without execution", async () => {
const dir = await mkdtemp(join(tmpdir(), "jev-hermes-adapter-"));
const replayPath = join(dir, "cases.jsonl");
- const [supervision, browser] = await run([
+ const [supervision, shadow, browser] = await run([
{ operation: "supervise", args: { provider: "demo", enabled: true, harness: "hermes-test", job: { goal: "verify" }, observation: { state: "done" }, evidence: { tests_passed: true } } },
+ { operation: "shadow_compaction", args: { provider: "demo", intent: "review context", context: { messages: ["old note", "Requirement: preserve /workspace/project/config.json"] } } },
{ operation: "browser_step", args: { provider: "demo", enabled: true, harness: "hermes-test", goal: "Open the visible docs link", observation: { url: "https://example.test", targets: [{ id: "docs", name: "Docs", href: "/docs", visible: true, clickable: true }], tabs: [], scroll: { down: false, up: false } } } },
], replayPath);
assert.equal(supervision.status, "judged");
assert.ok(["continue", "verify", "retry", "finish", "escalate"].includes(supervision.action));
+ assert.equal(shadow.mode, "jev_shadow");
+ assert.equal(shadow.changed, false);
+ assert.equal(shadow.protected_count, 1);
assert.ok(["selected", "fallback", "no_decision", "needs_confirmation"].includes(browser.status));
assert.equal(browser.execution.enabled, false);
});
diff --git a/test/model-route-metrics.test.mjs b/test/model-route-metrics.test.mjs
new file mode 100644
index 0000000..173df0e
--- /dev/null
+++ b/test/model-route-metrics.test.mjs
@@ -0,0 +1,61 @@
+import assert from "node:assert/strict";
+import { test } from "node:test";
+import { buildModelRouteReport } from "../src/model-route-metrics.mjs";
+
+const rows = [
+ {
+ record_type: "routing_case",
+ correlation_id: "one",
+ request: { context: { model_route: { mode: "shadow", profiles: [{ id: "fast" }, { id: "frontier" }] } } },
+ decision: { status: "selected", selected: "fast", receipt: { latency_ms: 4, cost_usd: 0.001 } },
+ },
+ {
+ record_type: "execution_receipt",
+ correlation_id: "one",
+ host: {
+ status: "completed",
+ duration_ms: 20,
+ result: { model_route: { actual_model_id: "fast", retry_count: 1, outcome: "verified" } },
+ },
+ },
+ {
+ record_type: "routing_case",
+ correlation_id: "two",
+ request: { context: { model_route: { mode: "shadow", profiles: [{ id: "fast" }, { id: "frontier" }] } } },
+ decision: { status: "selected", selected: "frontier", receipt: { latency_ms: 6, cost_usd: 0.002 } },
+ },
+ {
+ record_type: "execution_receipt",
+ correlation_id: "two",
+ host: {
+ status: "failed",
+ duration_ms: 30,
+ result: { model_route: { actual_model_id: "fast", retry_count: 2, outcome: "failed" } },
+ },
+ },
+];
+
+test("model route report joins shadow recommendation to host outcome without treating it as a route change", () => {
+ const report = buildModelRouteReport(rows);
+ assert.equal(report.schema_version, 1);
+ assert.equal(report.summary.decisions, 2);
+ assert.equal(report.summary.with_host_outcome, 2);
+ assert.equal(report.summary.recommendation_match_rate, 0.5);
+ assert.equal(report.summary.verified_outcomes, 1);
+ assert.equal(report.summary.failed_outcomes, 1);
+ assert.equal(report.summary.total_retry_count, 3);
+ assert.deepEqual(report.models.map((row) => row.recommended_model_id), ["fast", "frontier"]);
+ assert.equal(report.models[0].host_actual_model_ids.fast, 1);
+ assert.equal(report.models[1].host_actual_model_ids.fast, 1);
+ assert.equal(report.models[1].recommendation_matches, 0);
+});
+
+test("model route report fails open on incomplete or unrelated receipts", () => {
+ const report = buildModelRouteReport([
+ { record_type: "routing_case", correlation_id: "unrelated", request: { context: {} }, decision: { selected: "tool" } },
+ { record_type: "routing_case", correlation_id: "pending", request: { context: { model_route: { mode: "shadow" } } }, decision: { selected: null } },
+ ]);
+ assert.equal(report.summary.decisions, 1);
+ assert.equal(report.summary.with_host_outcome, 0);
+ assert.equal(report.models[0].recommended_model_id, "unrecommended");
+});
diff --git a/test/model-routing.test.mjs b/test/model-routing.test.mjs
new file mode 100644
index 0000000..eadb7a8
--- /dev/null
+++ b/test/model-routing.test.mjs
@@ -0,0 +1,59 @@
+import assert from "node:assert/strict";
+import { test } from "node:test";
+import { recommendModelRoute } from "../src/model-routing.mjs";
+
+const models = [
+ {
+ id: "fast",
+ provider: "openrouter",
+ model: "small-model",
+ reasoning_effort: "low",
+ description: "fast and cheap for simple classification or file inspection",
+ },
+ {
+ id: "frontier",
+ provider: "openai-codex",
+ model: "frontier-model",
+ reasoning_effort: "high",
+ description: "strong model for difficult implementation and ambiguous reasoning",
+ },
+];
+
+test("model routing returns an advisory profile and never changes the host route", async () => {
+ let captured;
+ const result = await recommendModelRoute({
+ intent: "classify a small batch of already structured feedback",
+ context: { task_size: "small", external_side_effects: false },
+ models,
+ provider: {
+ name: "model-route-test-provider",
+ async decide(input) {
+ captured = input;
+ return {
+ answers: { tool: { type: "choice", choice: "fast", probabilities: { fast: 0.92, frontier: 0.08 }, confidence: 0.92 } },
+ usage: { cost: 0.0001 },
+ };
+ },
+ },
+ });
+
+ assert.equal(result.route_mode, "shadow");
+ assert.equal(result.status, "selected");
+ assert.equal(result.selected, "fast");
+ assert.equal(result.recommended_model.model, "small-model");
+ assert.equal(result.execution.enabled, false);
+ assert.equal(captured.candidates.every((candidate) => candidate.kind === "model"), true);
+});
+
+test("model routing excludes unavailable profiles and fails open with no selected profile", async () => {
+ const result = await recommendModelRoute({
+ intent: "inspect a file",
+ models: [{ ...models[0], available: false }],
+ provider: "demo",
+ });
+
+ assert.equal(result.status, "no_decision");
+ assert.equal(result.selected, null);
+ assert.equal(result.recommended_model, null);
+ assert.equal(result.execution.enabled, false);
+});
diff --git a/test/shadow-compaction.test.mjs b/test/shadow-compaction.test.mjs
new file mode 100644
index 0000000..1dd17af
--- /dev/null
+++ b/test/shadow-compaction.test.mjs
@@ -0,0 +1,67 @@
+import assert from "node:assert/strict";
+import { test } from "node:test";
+import { buildShadowCompactionReport } from "../src/shadow-compaction.mjs";
+
+const context = {
+ messages: [
+ "stale casual discussion",
+ "Requirement: preserve /workspace/project/config.json exactly",
+ "recent working note",
+ ],
+ tool_results: [
+ "plain old tool output",
+ ],
+};
+
+test("shadow compaction is report-only, preserves pinned evidence, and exposes candidate decisions", async () => {
+ const before = JSON.stringify(context);
+ let received;
+ const report = await buildShadowCompactionReport({
+ intent: "prepare a safe compaction report",
+ context,
+ provider: {
+ name: "shadow-test-provider",
+ async evaluate(input) {
+ received = input;
+ return {
+ answers: {
+ item_0: { type: "noul", probability: 0.1, confidence: 0.9 },
+ item_1: { type: "noul", probability: 0.9, confidence: 0.9 },
+ item_2: { type: "noul", probability: 0.4, confidence: 0.2 },
+ },
+ usage: { cost: 0.0002 },
+ };
+ },
+ },
+ });
+
+ assert.equal(JSON.stringify(context), before);
+ assert.equal(report.mode, "jev_shadow");
+ assert.equal(report.status, "judged");
+ assert.equal(report.changed, false);
+ assert.equal(report.candidate_count, 4);
+ assert.equal(report.protected_count, 1);
+ assert.equal(report.droppable_count, 1);
+ assert.equal(report.retained_count, 3);
+ assert.ok(received.questions.item_0);
+ assert.equal(report.items.find((item) => item.id === "messages[1]").decision, "keep");
+ assert.equal(report.items.find((item) => item.id === "messages[1]").reason, "pinned:path,requirement");
+ assert.equal(report.items.find((item) => item.id === "messages[0]").decision, "drop_candidate");
+ assert.equal(report.items.find((item) => item.id === "tool_results[0]").decision, "keep");
+ assert.equal(report.items.find((item) => item.id === "tool_results[0]").reason, "low_confidence");
+});
+
+test("shadow compaction fails open and marks all unresolved items as retained", async () => {
+ const report = await buildShadowCompactionReport({
+ intent: "prepare a safe compaction report",
+ context: { messages: ["old note"] },
+ provider: { name: "offline", async evaluate() { throw new Error("provider unavailable"); } },
+ });
+
+ assert.equal(report.status, "fallback");
+ assert.equal(report.changed, false);
+ assert.equal(report.droppable_count, 0);
+ assert.equal(report.retained_count, 1);
+ assert.equal(report.items[0].decision, "keep");
+ assert.equal(report.items[0].reason, "provider_error");
+});
diff --git a/test/supervision.test.mjs b/test/supervision.test.mjs
index 330cb05..16e891d 100644
--- a/test/supervision.test.mjs
+++ b/test/supervision.test.mjs
@@ -4,6 +4,7 @@ import { tmpdir } from "node:os";
import { join } from "node:path";
import { test } from "node:test";
import {
+ deterministicEvidenceState,
deterministicSupervisionPolicy,
normalizeAssessment,
superviseWork,
@@ -52,11 +53,26 @@ test("supervision contract exposes bounded dimensions and parses Noul answers",
});
});
+test("deterministic evidence state distinguishes proof, absence, and contradiction", () => {
+ assert.deepEqual(deterministicEvidenceState({}), { state: "missing", sources: [] });
+ assert.deepEqual(deterministicEvidenceState({ tests_passed: true, receipts: [{ id: "test:1", status: "completed" }] }), {
+ state: "present",
+ sources: ["tests_passed", "receipt:test:1"],
+ });
+ assert.deepEqual(deterministicEvidenceState({ tests_passed: true, tests_failed: true }), {
+ state: "contradictory",
+ sources: ["tests_passed", "tests_failed"],
+ });
+});
+
test("deterministic host policy owns the supervision action", () => {
assert.equal(deterministicSupervisionPolicy({ assessment: assessment({ verification_needed: 0.9 }) }).action, "verify");
assert.equal(deterministicSupervisionPolicy({ assessment: assessment({ worker_stuck: 0.9 }), attempts: 0 }).action, "retry");
assert.equal(deterministicSupervisionPolicy({ assessment: assessment({ worker_stuck: 0.9 }), attempts: 2 }).action, "escalate");
assert.equal(deterministicSupervisionPolicy({ assessment: assessment(), evidence: { tests_passed: true } }).action, "finish");
+ const contradiction = deterministicSupervisionPolicy({ assessment: assessment(), evidence: { tests_passed: true, tests_failed: true } });
+ assert.equal(contradiction.action, "verify");
+ assert.match(contradiction.reason, /contradictory/);
});
test("supervision is opt-in, fail-open, and records replayable judgments", async () => {
@@ -91,9 +107,11 @@ test("supervision is opt-in, fail-open, and records replayable judgments", async
assert.equal(result.action, "finish");
assert.equal(result.metrics.jev_calls, 1);
assert.equal(result.metrics.cost_usd, 0.0003);
+ assert.deepEqual(result.evidence_state, { state: "present", sources: ["tests_passed"] });
const cases = await readSupervisionCases(path);
assert.equal(cases.length, 1);
assert.equal(cases[0].record_type, "supervision_case");
assert.equal(cases[0].supervision.action, "finish");
+ assert.deepEqual(cases[0].supervision.evidence_state, { state: "present", sources: ["tests_passed"] });
assert.equal(JSON.parse(await readFile(path, "utf8")).record_type, "supervision_case");
});