Skip to content

Repository files navigation

vanth

Event-driven background jobs for agents.

Vanth is a localhost background-job daemon with a Model Context Protocol (MCP) interface. It runs detached, non-interactive shell commands; captures their output durably; parses optional AGENT_EVENT structured events into progress bars, metric series, and checkpoints; and can wake a Codex or OpenCode session when a job needs attention. It is built for one trusted user on one machine.

  • Any command: downloads, image/audio processing, ETL, ML training — if it runs in a shell, Vanth can run it detached and track it.
  • Durable: jobs and events live in SQLite (WAL, busy-timeout) and survive daemon, MCP, and machine restarts.
  • Event-first: agents job_wait for meaningful events instead of polling logs.
  • Wake-on-attention: durable at-least-once deliveries resume a Codex thread or OpenCode session when a job needs a human or agent.
  • Terminal dashboard: the native Go monitor renders a live W&B-LEET-style dashboard of jobs, metrics, and plots.

Out of scope: TLS, multi-user tenancy/RBAC, a web UI, and distributed workers. Supported: interactive stdin (job_send), concurrent-job quotas, automatic retention, cron/interval schedules, pools/priority/pause queues, readiness triggers, kill attribution + secret masking, duration/flaky analytics, managed artifacts, and remote SSH execution (beta).

For agents: when MCP is available, start work with job_start, then job_wait for progress/checkpoint/completed events instead of polling. When MCP is unavailable, use the vanth start CLI fallback. Make jobs emit AGENT_EVENT lines (below) so progress, metrics, and checkpoints appear live in the vanth-monitor dashboard; and let long jobs resume you via wake targets instead of you checking in.

Wake me when it finishes

vanth start --wake-me -- <command>

The MCP equivalent is job_start(command="...", wake_me=True). --wake-me defaults to all terminal outcomes (completed, failed, timeout, cancelled, orphaned); use --wake-me=completed,failed,checkpoint to override the event list. Never use the relay client id opencode-<pid>-<rand> as session_id: use the ses_... destination in vanth doctor --json, or omit session_id to resolve the live relay automatically.


Quick start

Install with uv (Python 3.11+):

uv tool install vanth

This installs the vanth MCP server, vanthd daemon, vanth-monitor, and the ops CLI as standalone tools (the wheel bundles the native Go monitor, so no Go toolchain is needed). Wheels are published for Windows x86_64, Linux x86_64/arm64, and macOS x86_64/arm64.

From a source checkout (development), install the project environment with uv sync and run commands as uv run vanth .... The rest of this guide shows the installed vanth command unless a row says uv run.

git clone https://github.com/abhim-dv/vanth.git
cd vanth
uv sync

The daemon autostarts on demand for MCP calls and operational CLI commands. vanth status is an observational check: it reports DOWN if the daemon is stopped and does not start it. Use vanth doctor for health details; a normal MCP call or command such as vanth list will start the daemon on demand.

  1. Register the MCP server in opencode, Codex, and Claude-style clients:

    vanth setup

    This detects client configs and supported client commands on PATH, then configures them with prompts. When a supported client's config file is missing, setup creates it before adding Vanth. Use vanth setup --yes for scripts, or name clients such as vanth setup opencode codex.

  2. Refresh and verify the client connection. MCP clients usually load server configuration at startup. Restart or reload the client after setup, then confirm its tool picker/list includes job_start, job_wait, and job_doctor. In OpenCode, opencode mcp list checks the connection. If the server is missing, run vanth status and vanth doctor; for a source checkout, make sure the client config launches uv run vanth from the repo.

  3. Check health:

    vanth status
    vanth doctor

    status shows daemon state and jobs without starting the daemon; doctor prints the full health report and may start it on demand.

  4. Pick up updates — after upgrading, restart the daemon so it runs the new code. In-flight jobs survive (runners are detached):

    vanth restart

End-to-end: run a tracked job

Once the MCP client is connected, this is the whole loop:

Goal MCP tools
Start a job job_start
Wait for progress or completion job_wait
Inspect output or status job_status, job_run_summary, job_tail, job_events
Stop or retry work job_stop, job_rerun
Diagnose a wake job_deliveries, job_delivery_attempts, job_retry_delivery

The complete MCP reference is docs/agent-tools.md.

job_start(
  command="uv run python examples\\long_job.py",
  name="demo run",
  wake_me=True,
)
# -> job_<id>

job_wait(job_id="job_<id>", filters=["checkpoint"], timeout_seconds=120)
# -> returns the first checkpoint event + current status

job_wait(job_id="job_<id>", filters=["completed", "failed", "timeout", "cancelled", "orphaned"], timeout_seconds=300)
# -> returns the terminal event + exit code

# If it failed, inspect the summary and recent output:
job_run_summary(job_id="job_<id>")
job_tail(job_id="job_<id>", stream="stderr", max_bytes=8192)

And in a third terminal, watch it live:

vanth-monitor

Command-line entry points

Installed command Source checkout command Purpose
vanth uv run vanth MCP stdio server and human CLI
vanthd uv run vanthd Background HTTP daemon
vanth-monitor uv run vanth-monitor Live terminal dashboard (Go binary, bundled in the wheel)
vanth-codex-notify uv run vanth-codex-notify Delivery adapter: reads a wake payload on stdin and dispatches it to Codex

The MCP stdio server exits when its launching client closes stdin or dies, and reaps itself after VANTH_WATCH_IDLE seconds of idle (default 1800; 0 disables). It does not exit during a blocking tool call. vanth doctor --reap-orphans cleans up orphaned MCP servers left by older versions.

Human CLI

vanth doubles as a human-facing operations CLI. Operational commands autostart the daemon on demand; bare vanth status is read-only and reports DOWN when it is stopped. vanth status <job-id> inspects a job and may start the daemon. Every flag-based command supports --json where noted for scripts.

Command Purpose
vanth --version / vanth version Print the installed version
vanth status [<job-id>] Bare command: read-only daemon up/down, pid, schema, running jobs, deliveries. With a job id: inspect that job's status/exit/runtime/last event (may start daemon). (--json)
vanth doctor Full health report (same as job_doctor, human-readable; --json)
vanth restart Gracefully stop + start the daemon (jobs survive)
vanth setup [opencode] [codex] [claude] [desktop] [--remove] [--yes] Register/unregister the MCP server in your clients' configs
vanth start [options] [--] <command...> Start a background job without MCP. Options: --name, --cwd, --timeout, --env K=V, --wake JSON, --wake-me[=EVENTS], --interactive, --priority, --pool, --tag, --notes, --secret-env, --trigger JSON, --policy JSON; -- passes command flags verbatim
vanth sleep <seconds> Start a trivial sleep job
vanth list (ps alias) List jobs (--status, --limit, --all, --json); defaults to in-flight (launching/queued/running/…), --all shows finished jobs; running jobs show DURATION and AGE
vanth deliveries [--status S] [--job JOB_ID] [--limit N] [--json] List wake deliveries, attempts, and last errors
vanth wake <job_id> [--now] [--type T] [--events a,b] [--cwd DIR] [--config JSON] [--target JSON] Add a wake target to a job after it started (in-flight or finished). Fires on future events; --now surfaces a synthetic wake immediately. CLI counterpart of job_add_wake_target / job_wake_now
vanth api Print the loopback HTTP surface and authentication details
vanth logs <job_id> (tail alias) Show a job's output (--stream stdout|stderr|all, --offset, --max-bytes, --grep, --json)
vanth wait <job_id> Block until an event fires (the CLI job_wait): defaults to terminal outcomes (completed, failed, timeout, cancelled, orphaned); --events narrows the set. Also accepts --timeout SECONDS, --since-event-id ID. Exits 0 on an event, 3 on timeout
vanth diff <job_id> <other> Compare two jobs' run specs
vanth stop <job_id> Stop a running job (--signal, --kill-after)
vanth artifacts <job_id> List a job's artifacts (--limit, --json)
vanth prune Manual retention cleanup; dry-run by default (--older-than N, --yes)
vanth backup [--out PATH] [--include-logs] Write an archive of jobs, artifacts, and events
vanth restore <archive> --yes [--force] Restore an archive
vanth remote <action> Pair, list, inspect, remove, or retry remote execution hosts
vanth autostart enable|disable|status Daemon survives reboots (Windows Task Scheduler / macOS launchd / Linux systemd user unit)
vanth help <command> Show command-specific help; vanth --help lists commands

Run vanth --help for the command list. Command-specific help is available with vanth <command> --help for status, doctor, list, start, logs, wait, stop, sleep, deliveries, wake, artifacts, diff, api, remote, backup, restore, prune, restart, setup, autostart, and version. Job-id arguments accept an unambiguous prefix, and an unknown id suggests near matches.

For vanth start, <command...> begins at the first non-option argument. A single argument is used verbatim (so a whole quoted command string works); multiple arguments are re-quoted for the host shell, so the program and its arguments — including empty ones — survive, and shell metacharacters in an argument stay data (on Windows %, !, and " cannot be encoded safely, so pass the whole command as one quoted string for those). Repeated --env, --wake, --tag, and --secret-env flags are allowed;

If the command contains shell operators (&&, |, >, <), pass it as ONE quoted string or put the steps in a script file and start that. Reassembling operators from separate arguments is where shell quoting goes wrong (and the classic PowerShell-5.1 failure mode is that a single-quoted string with inner quotes arrives split into garbage argv). vanth start refuses this with exit 2:

vanth start "cmd /c echo done && timeout /t 30 /nobreak >nul"
vanth start -- run.cmd

The first command passes one quoted shell command; the second starts a script.

For a trivial delay, use vanth sleep <seconds> rather than a shell sleep idiom.

--wake takes the same wake-target JSON shape documented in docs/agent-tools.md. The global --json flag is recognized only before --, so a wrapped command keeps its own literal --json.

Examples:

vanth status --json
vanth list --status running --limit 20 --json
vanth logs job_abc123 --stream stderr --max-bytes 65536
vanth stop job_abc123 --signal terminate --kill-after 10
vanth artifacts job_abc123 --json
vanth prune --older-than 604800 --yes
vanth backup
vanth remote list
vanth autostart enable

vanth prune --yes performs deletion; without --yes, it only previews. vanth autostart enable configures the daemon to start at login.

vanth restart is the reliable way to pick up a code/version update: it sends the daemon a graceful shutdown over loopback, waits for the old process to fully release the home lock, then starts a fresh daemon. In-flight jobs are owned by detached runners, so they continue across the restart.

vanth autostart installs a start-at-login mechanism per platform — a Windows Task Scheduler task, a macOS launchd agent, or a Linux systemd user unit — then vanth status reports whether it is active.


How it works

MCP client / HTTP client
        |
        v
   vanthd (localhost HTTP daemon, bearer-token auth)
        |                 |                    |
        |                 |                    +---> wake adapters
        |                 |                          (local_command / codex_thread / opencode_thread / webhook)
        |                 |
        |                 +----> jobs.sqlite (durable source of truth)
        |
        +----> vanth.runner (detached worker process)
                    |
                    +----> your command (own process group)
                              |
                              +----> stdout/stderr -> logs/ + AGENT_EVENT parsing

Ownership rules:

  • the runner owns the real command, its timeout, and stream draining;
  • the daemon owns maintenance, delivery dispatch, API requests, and recovery;
  • SQLite is the source of truth across process restarts;
  • the MCP and HTTP clients never need to stay alive for jobs to continue.

A job is not considered terminal until both output streams have reached EOF and all structured events have been persisted.

Job lifecycle

A job moves through a small set of states. Terminal states are permanent.

State Meaning
running Workload launched; runner is streaming output and heartbeating
completed Command exited 0, streams drained, events persisted
failed Command exited non-zero
timeout Command exceeded timeout_seconds; runner terminated it
cancelled job_stop was issued and the process tree actually terminated
orphaned Runner died unexpectedly (crash); never silently dropped

The runner enforces timeout_seconds even across daemon restarts. On recovery, a running job whose runner is gone is marked cancelled (if a stop was requested) or orphaned (if not) — never left as a zombie running row.


Installing the MCP server

vanth is the MCP stdio server. It talks to the daemon, starting it automatically on first use if it is not already running.

One-shot setup

After installing the tool, connect it to the MCP clients on your machine in a single step:

uv tool install vanth && vanth setup

vanth setup detects existing client configs and supported client commands on PATH (OpenCode, Codex, and Claude Code), creates a missing config file for a selected client when needed, and shows what it found. It backs up each existing config before changing it (.vanth-setup-<ts>.bak), and upserts the Vanth MCP entry — leaving every other setting and comment untouched. OpenCode deep-merges config.json, opencode.json, and opencode.jsonc in that order (later files win). Setup reads JSONC read-only (comments are stripped in memory, never rewritten) and edits a .jsonc only when it has no comments to lose; it writes at most one file (never a duplicate entry) and refuses to write a lower-precedence file that a higher-precedence one would override. A registration it cannot safely make is reported as a skipped client with manual instructions rather than as success, and --remove clears every safely-editable registration. vanth status reports the effective state of the deep-merged entry: an entry with enabled: false is disabled, and a file that cannot be parsed is unreadable rather than not configured.

vanth setup
vanth setup --yes
vanth setup opencode codex
vanth setup --json
vanth setup --remove

By default setup detects clients and prompts before applying changes. --yes skips prompts, client names limit the targets, --json prints machine-readable output, and --remove unregisters Vanth.

Configs it manages:

Client File Section
opencode ~/.config/opencode/config.json, opencode.json, or opencode.jsonc mcp.vanth
Codex ~/.codex/config.toml [mcp_servers.vanth]
Claude Code / Cursor ~/.claude.json mcpServers.vanth

Manually, the same entries are:

opencode

Add to ~/.config/opencode/opencode.json:

{
  "$schema": "https://opencode.ai/config.json",
  "mcp": {
    "vanth": {
      "type": "local",
      "command": ["vanth"],
      "enabled": true,
      "timeout": 15000
    }
  }
}

From a source checkout, use uv directly instead of a bare vanth:

{
  "mcp": {
    "vanth": {
      "type": "local",
      "command": ["uv", "run", "--directory", "/path/to/vanth", "vanth"],
      "enabled": true,
      "timeout": 15000
    }
  }
}

Verify the connection and tools:

opencode mcp list

Claude-style MCP clients (mcpServers)

Published wheel:

{
  "mcpServers": {
    "vanth": { "command": "vanth", "env": { "VANTH_HOME": "C:/Users/you/.vanth" } }
  }
}

From a source checkout:

{
  "mcpServers": {
    "vanth": {
      "command": "uv",
      "args": ["--directory", "/path/to/vanth", "run", "vanth"],
      "env": { "VANTH_HOME": "C:/Users/you/.vanth" }
    }
  }
}

Configuring the daemon home

Both the MCP server and the daemon resolve the same state root from VANTH_HOME (default %USERPROFILE%\.vanth on Windows, ~/.vanth on Unix; AGENT_BG_HOME is accepted as an alias). If both are set they must resolve to the same directory.


Instrumenting jobs with agent_event

Any Python script can emit structured events to stdout (or stderr) that Vanth parses and the monitor charts. This is optional — plain scripts still run and log — but it is what turns a job into a first-class tracked object.

from vanth.agent_events import agent_event, progress

# A checkpoint: something meaningful happened.
agent_event("checkpoint", "epoch complete", epoch=10, val_loss=0.42)

# A progress update: drives the progress bar and progress.* plots.
progress(10, 100, unit="epoch", stage="train", message="10/100 epochs")

# Arbitrary scalar metrics: become their own line plots.
agent_event("metric", _step=10, loss=0.42, acc=0.88, mbps=12.4)

Notes:

  • the helper prints AGENT_EVENT {json} with flush=True (flush matters);
  • progress(current, total, unit=..., stage=...) computes percent for you;
  • metric payloads: numeric fields become series; _step (if present and numeric) is the x-axis, otherwise the event sequence number is used; keys starting with _ other than _step are ignored; booleans are not metrics; NaN/Infinity/null values are skipped and counted in the monitor's warning badge;
  • any other field (e.g. file, stage, phase) is preserved and visible in the exact event table.

Example: a tracked downloader

# downloader.py
import os
from vanth.agent_events import agent_event, progress

files = ["a.bin", "b.bin", "c.bin"]
total = sum(os.path.getsize(f) for f in files)
done = 0

for f in files:
    agent_event("checkpoint", f"starting {f}", file=f)
    # ... download f ...
    done += os.path.getsize(f)
    progress(done, total, unit="bytes", stage="download",
             message=f"{done}/{total} bytes")

Example: an image-processing batch

from vanth.agent_events import agent_event, progress

images = list(find_images("input/"))
for i, img in enumerate(images, 1):
    out = process(img)                    # resize, denoise, ...
    agent_event("metric", _step=i, sharpness=out.sharpness, size_mb=out.size_mb)
    progress(i, len(images), unit="images", stage="process", message=img.name)

Timestamped, leveled logging with loguru

Vanth ships a loguru wrapper that routes every record into a structured AGENT_EVENT log line, so logs appear as timestamped, level-aware events in the event table (with the level badge and exact timestamps) instead of bare text:

from vanth.agent_logger import logger, log_with_context

logger.info("training started", lr=8e-5, batch_size=8)     # event type "log", level info
logger.warning("low disk", free_gb=2.5)
log_with_context("error", "failed to load checkpoint", path="best.pt")

Each call emits AGENT_EVENT {"type":"log","level":"info","message":"...","data":{...}} which the daemon persists as a durable event. data carries extra context. The monitor shows these in the exact event table alongside metric/progress events.


Tool reference (MCP tools)

Tool Purpose
job_start Launch a command as a detached job
job_rerun Re-launch a job, optionally overriding command/env/cwd/timeout_seconds/name/tags/notes/interactive
job_send Feed stdin to an interactive job (interactive=True first)
job_wait Block until a matching event (or timeout) — the preferred way to await jobs (return_progress optional)
job_status One job's status, command, env, progress, last event, linkage, tags
job_status_batch Many jobs' status in one call (job_ids, limit)
job_list Recent jobs, filterable by status / thread_id / name / tags
job_view Agent-facing summaries sorted by attention priority
job_events Structured events for a job (forward via since_event_id, or latest-first via reverse)
job_tail Bounded stdout/stderr log tail with byte offsets (follow/timeout_seconds/grep optional)
job_metrics_query Read stored scalar metric series (loss, acc, progress.percent, ...)
job_metric_compare Compare one metric across jobs (latest/mean/min/max/sum/count)
job_duration_stats Per-job p50/p95 duration + queue time, success rate, flaky score, slowest-N, trend
job_run_summary One-call "did it work?" — status, runtime, progress, metrics, artifacts
job_diff Diff the run specs of two jobs (command/env/cwd/tags/wake targets)
job_artifact_add Attach an artifact (checkpoint, CSV, output) to a job
job_artifacts List artifacts attached to a job
job_dashboard Downsampled chart-data view for any renderer
job_deliveries Wake deliveries for a job, filterable by status
job_mark_delivery Manually set a delivery's status
job_retry_delivery Requeue a failed delivery for dispatch
job_delivery_attempts Attempt/lease history for one delivery
job_stop Stop a running job (terminate process tree)
job_pause / job_resume Hold / release a queued (pool or trigger) job
job_request_decision Ask a human to decide something about a job; notifies its wake targets and returns a durable decision_id
job_resolve / job_withdraw_decision Answer a pending decision with one of its options, or withdraw it
job_decisions List decisions, filterable by job_id / status
pool_configure / pool_list Per-pool max_parallel + pause state, and live queue depths
schedule_create Create a cron or interval schedule that launches a job per fire
schedule_list / schedule_update / schedule_delete Manage schedules in place
schedule_next Preview a schedule's next N fire times
job_doctor Daemon health, schema, tables, binary availability
job_cleanup Dry-run or real removal of old terminal jobs
daemon_wake Schedule a self-resume wake target — full target dict or events/type/...config shorthand (Python API)

For the full parameter contract and response shapes of every tool, see docs/agent-tools.md.

AGENT_EVENT protocol

Jobs can emit AGENT_EVENT <json> lines on stdout (or stderr) to create typed events — progress, metric, checkpoint, completed, and more — that the daemon persists durably, the dashboard charts, and wake deliveries carry to agents. See vanth/agent_events.py for the Python helpers and vanth/agent_logger.py for the loguru integration.

job_start

job_start(
  command="uv run python examples\\long_job.py",
  name="training run",
  cwd="F:\\git\\project",            # optional
  env={"CUDA_VISIBLE_DEVICES": "0"}, # optional
  timeout_seconds=3600,              # optional; None = no timeout
  interactive=True,                  # optional; open stdin for job_send
  notify_on=["progress","checkpoint","failed","completed"],
  origin_thread_id="019f...",        # the agent thread that launched it
  tags=["training","gpu"],           # optional
  secret_env=["HF_TOKEN"],           # optional; mask these env values in logs/events
  wake_targets=[...],                # optional, see below
  trigger={"job_id": "job_A", "status": "completed"}  # optional DAG: start after job_A completes
)

Returns job_id, status, worker_pid, and the log/event paths. With trigger set, the job is created queued and starts automatically when the parent job reaches that status (or is cancelled if the parent ends differently).

Readiness triggers — wait for a condition, not just a job

A trigger may also carry a readiness probe; when the DAG gate and the probe are both present they are ANDed. The job stays queued until the probe passes (or is cancelled once an optional timeout_seconds elapses):

job_start(command="migrate.sh", trigger={"probe": {"type": "port",
           "host": "127.0.0.1", "port": 5432, "timeout_seconds": 120}})

# probe types
{"probe": {"type": "port",     "host": "127.0.0.1", "port": 5432}}
{"probe": {"type": "http",     "url": "http://127.0.0.1:8080/health", "expect_status": 200}}
{"probe": {"type": "log_line", "job_id": "job_B", "pattern": "ready", "stream": "stdout"}}
{"probe": {"type": "file",     "path": "/tmp/ready"}}

This lets one job orchestrate a stack — "start the DB, wait until the port accepts, then migrate" — instead of sleep hacks. Probes run on the daemon host (direct connection; no proxy) at interval_seconds cadence (default 1s), bounded per dispatcher pass so blocked probes can't stall other work. timeout_seconds is measured from when the dependency gate is satisfied (or from queue creation with no dependency gate); a missed deadline is attributed (actor="daemon" + reason) on the cancelled event.

job_send — feed stdin to an interactive job

job_send(job_id="job_...", input="y", eof=False)

Appends input to a running job's stdin. Start the job with interactive=True first. eof=True closes the job's stdin (the child sees EOF). Non-blocking: returns immediately; the input is queued to the runner. Rejects jobs that are not interactive, not running, or unknown. job_rerun preserves the interactive flag.

job_request_decision — ask a human and wait durably

job_request_decision(job_id="job_...", prompt="Ship the release?",
                     options=["approve", "deny"], timeout_seconds=3600)
job_wait(job_id="job_...", filters=["decision_resolved"])   # await the answer
job_resolve(job_id="job_...", token="dec_...", choice="approve")

Records a durable "needs a decision" request against a non-terminal job and notifies the job's wake targets, so the owning thread learns a human is needed. The job keeps running (its status is untouched) — the answer arrives as a decision_resolved event. timeout_seconds expires the request (emitting decision_expired); job_withdraw_decision cancels it; job_decisions lists pending/answered requests.

job_status — see what a job is running

job_status(job_id="job_...")

Returns status, command, cwd, env, timeout_seconds, notes, run (author, hostname, OS, Python version, CPU/GPU, git repo/branch/commit), runtime_seconds, progress, last event, thread linkage, tags, and exit code. This is the fastest way for an agent to answer "what is this job doing?" — and mirrors the run-overview you'd see for a run in W&B.

Pass notes="..." to job_start to annotate a run ("what makes this run special?"), which is preserved on job_rerun and shown in the monitor.

job_rerun — relaunch a failed job

job_rerun(job_id="job_...")

Re-launches the job with its original command, cwd, env, timeout, name, tags, origin thread, and wake targets — a new job_id is returned. Use it to retry a failed download, flaky processing batch, or transient failure without reconstructing the request.

job_list — filter by name or tag

job_list(status=["running"], name="train", tags=["gpu"], limit=20)

Filters: status (list), thread_id, name (substring), tags (must contain all listed tags).

job_events — forward or latest-first

job_events(job_id="job_...", since_event_id="evt_...", limit=20)      # events after the cursor
job_events(job_id="job_...", reverse=true, limit=20)                   # the 20 newest events, newest first

reverse: true returns the most recent events (newest first) — ideal for "what happened recently?" — and can be combined with since_event_id to page backward.

job_wait — the heart of agent usage

job_wait(job_id="job_...", filters=["checkpoint","completed","failed","timeout","cancelled","orphaned"], timeout_seconds=3600)
  • waits for the first event matching any filter, returning it with the current status;
  • pass since_event_id to wait only for events newer than one you already saw;
  • on timeout returns result: "timeout"; on daemon shutdown returns result: "shutdown".

job_view — what to show the user

job_view(thread_id="019f...", limit=20)

Returns compact summaries sorted by attention priority: running and failed jobs first, then jobs with pending/failed deliveries, then everything else. Each entry includes status, progress, the latest event, thread linkage, tags, and delivery counts.

job_stop — stop a running job

job_stop(job_id="job_...", signal="terminate", kill_after_seconds=10, reason="superseded by run #2")

Terminates the job's process tree. A graceful signal (default terminate) is sent first; if the job has not exited within kill_after_seconds, it is killed. The job becomes cancelled only after the workload tree actually terminated; otherwise it stays running and the stop is retryable.

Kill attribution. The cancelled event carries data={"actor": ..., "reason": ...} so a stop is never an unattributable "killed". Actors are tool (an MCP call), user (vanth stop, or a human HTTP call), watchdog (recovery / heartbeat reconciliation), and timeout (the runner's timeout); job_status exposes the persisted stop_actor / stop_reason. vanth stop <id> --reason "..." sets the user reason.

secret_env — mask declared secrets in captured output

job_start(command="python train.py", env={"HF_TOKEN": "hf_..."}, secret_env=["HF_TOKEN"])

Every value named in secret_env is replaced with *** before it is written to the job's captured stdout/stderr logs or parsed into structured events (the GitHub Actions ::add-mask:: pattern), so masked output never leaks through logs, events, deliveries, or the monitor. Note this protects emitted output: as with any env value, a declared secret is still stored in the job's environment in the owner-only jobs.sqlite (the single-user state directory is protected by owner-only permissions). Masking applies to local jobs (remote jobs ignore secret_env).

job_mark_delivery / job_retry_delivery — manual delivery control

job_mark_delivery(delivery_id="del_...", status="delivered", error="optional reason")
job_retry_delivery(delivery_id="del_...")   # requeue a failed delivery

job_mark_delivery sets a delivery's status by hand (e.g. after resolving an adapter problem); job_retry_delivery requeues a failed one for the next dispatch pass. job_delivery_attempts shows the claim/lease history.

job_cleanup — remove old terminal jobs

job_cleanup(older_than_seconds=86400, dry_run=true)   # preview
job_cleanup(older_than_seconds=86400, dry_run=false)  # delete

Removes terminal jobs older than the cutoff: logs, event mirrors, specs, deliveries, attempts, wake targets, events, stdin channels, then the job row. Running jobs are never selected. Dry-run is fully read-only. Cleanup is safe to repeat.

Automatic retention: set VANTH_RETENTION_SECONDS on the daemon to purge old terminal jobs in the background (polled every VANTH_RETENTION_INTERVAL_SECONDS, default 3600; safe by default — VANTH_RETENTION_DRY_RUN=0 to actually delete).

Concurrency quota: set VANTH_MAX_RUNNING_JOBS (default 0 = unlimited) to cap how many jobs may run at once; job_start returns a clear error when the cap is reached.

job_metrics_query — read stored scalar series

job_metrics_query(job_id="job_...", metric="loss", from_ms=..., to_ms=..., limit=1000)

Returns the stored series for one job, grouped by metric name. metric filters to a single series (e.g. loss, acc, progress.percent); from_ms/to_ms filter by event timestamp (epoch milliseconds). Points are ordered by event sequence. This is the read side of the terminal monitor's data.

job_metric_compare — compare a metric across runs

job_metric_compare(job_ids=["job_a", "job_b"], metric="val_loss", aggregation="min")

Compares one metric across jobs (e.g. val_loss across seeds or configs). aggregation is latest, mean, min, max, sum, or count; the result includes the per-job value plus the first/last points. This is the W&B-style "which run won?" primitive.

job_duration_stats — did it get slower?

job_duration_stats(name="nightly backup", tags=["prod"], slowest=10)

Groups terminal runs by logical job (name, falling back to the command) and returns p50/p95 runtime and queue time, success rate, a flaky score (a failed run that has a success both before and after it — real intermittency, not a first-attempt failure), each group's slowest recent runs, and a trend flag (regressing / stable / improving). The trend compares the newer half's p50 against the older half's, so it catches "this backup crept 40min → 2h over 6 weeks". The top-level slowest list is the slowest-N runs across all groups.

job_run_summary — did it work?

job_run_summary(job_id="job_...")

One call returns status, name, runtime, exit code, latest progress, notes, per-metric overview (latest/first/min/max/count), and attached artifacts — the fastest way for an agent to report on a finished job.

job_artifact_add / job_artifacts — attach outputs

job_artifact_add(job_id="job_...", name="best.pt", uri="file:///...", kind="checkpoint",
                 size_bytes=..., sha256="...", meta={"epoch": 5})
job_artifacts(job_id="job_...")

Attach artifacts (checkpoints, CSVs, rendered outputs) to a job so they are listed in job_run_summary and retrievable later. meta is free-form JSON.

job_dashboard — chart data for any renderer

job_dashboard(job_ids=["job_..."], limit=5000)

Returns the job list plus every stored metric series, downsampled to limit points per series — the same data the Go terminal monitor charts, exposed over HTTP/MCP so any client (a future web/cloud dashboard) can render it.


Schedules and queues

Vanth has no external scheduler process: the daemon's existing 0.2s maintenance loop fires schedules and launches queued jobs, and everything lives in jobs.sqlite.

Schedules (cron or interval)

A schedule launches a fresh job per fire (optionally masked with secret_env, tagged scheduled, and linked by job_status's schedule_id):

schedule_create(name="nightly backup", command="backup.sh",
                cron="0 3 * * *", timezone_name="America/New_York",
                overlap="skip")          # cron: 5 fields or @daily/@hourly/...
schedule_create(name="poll", command="poll.sh", interval_seconds=300)
schedule_next(schedule_id="sched_...", count=5)   # preview fire times
schedule_update(schedule_id="sched_...", changes={"cron": "0 4 * * *"})
schedule_update(schedule_id="sched_...", changes={"enabled": False})  # pause
schedule_list()
schedule_delete(schedule_id="sched_...")
  • Cron is 5-field (minute hour day-of-month month day-of-week), numeric values with *, ranges (1-5), lists (1,13), and steps (*/15), plus the @hourly/@daily/@weekly/@monthly/@yearly shorthands.
  • Timezones are IANA names, matched against the local wall clock. DST is handled by construction: a nonexistent local time (spring forward) is skipped; an ambiguous one (fall back) matches once per UTC minute that maps to it. UTC needs no timezone database; named zones use the OS database on Linux/macOS and the bundled tzdata package on Windows.
  • overlap: skip (default) skips the fire (advancing to the next) while a job from the same schedule is still active; allow always launches.
  • Missed fires while the daemon was down are not backfilled — the schedule resumes at the next future match (the dead-man's-switch policy already alerts on missed runs).

Queues: pools, priority, pause

Start a job into a named pool instead of launching it immediately:

pool_configure(pool="gpu", max_parallel=1)     # 0 = unlimited
pool_list()                                    # queued/running per pool
job_start(command="train.py", pool="gpu", priority=5)
job_pause(job_id="job_...")                    # hold a queued job
job_resume(job_id="job_...")
pool_configure(pool="gpu", paused=True)        # hold the whole pool

Queued jobs (pool, trigger, or both) launch from the one dispatcher ordered by priority (higher first), oldest first, once the trigger is satisfied, the pool is not paused and is under max_parallel, and the global VANTH_MAX_RUNNING_JOBS quota allows. Pausing affects queued jobs only; running jobs are untouched.


Wake targets (wake an agent when a job needs attention)

When a job emits a matching event, the daemon creates a durable delivery and dispatches it through the adapter. Delivery is at-least-once; every payload carries a delivery_id for deduplication.

Targets are not fixed at start time. POST /jobs/{id}/wake (MCP job_add_wake_target, CLI vanth wake) registers a target on a job that is already running or finished — it fires on events after registration. POST /jobs/{id}/wake-now (MCP job_wake_now, CLI vanth wake --now) registers the target AND enqueues a synthetic "wake_now" delivery at once, so a wake reaches the session even if the triggering event already fired; it never fabricates a "completed"/"failed" event.

Thread identity. codex_cli_thread / codex_desktop targets omit the id and inherit the calling task's (CODEX_THREAD_ID / VANTH_CODEX_DESKTOP_THREAD); an explicit thread_id always wins. opencode_thread does not inherit the calling session (OpenCode never injects OPENCODE_SESSION_ID into MCP subprocesses): it resolves the session from the newest registered plugin relay for the job's cwd (see vanth doctor). Omitting session_id therefore works whenever a relay is registered for that cwd; if none is, target creation fails with an actionable error. Use vanth doctor --json for the full relay/session list. A target naming an explicit session can be resumed by a subprocess (opencode run --session); the plugin relay is what lands the prompt in the TUI you are watching.

local_command

Runs an arbitrary command, passing the delivery payload as JSON on stdin:

{
  "type": "local_command",
  "events": ["checkpoint", "failed", "completed"],
  "command": ["python", "deliver.py"]
}

Exit 0 marks the delivery delivered; any other exit marks it failed.

codex_thread / codex_cli_thread

Resumes a Codex thread through the local app-server (for an unloaded CLI task):

{
  "type": "codex_thread",
  "thread_id": "019f...",
  "events": ["checkpoint", "failed", "completed"],
  "codex_command": ["C:\\codex\\codex.exe"]
}

Protocol: initialize -> thread/resume -> turn/start.

The target thread must already have had at least one turn (a persisted "rollout"). Resuming a brand-new, zero-turn thread fails with no rollout found for thread id <id> — the intended wake target is an existing/active conversation, not a never-started one.

codex_desktop

Wakes a RUNNING Codex Desktop task through the native app-tools host pipe (codex_app/send_message_to_thread on CODEX_APP_TOOLS_PIPE_PATH). This never spawns a second app-server and never falls back to the CLI thread bridge:

{
  "type": "codex_desktop",
  "thread_id": "019f...",
  "events": ["checkpoint", "failed", "completed"]
}

Desktop wake is delivered by a client-side relay: the Vanth MCP integration registers the task id it can wake, long-polls the daemon for due deliveries, submits the follow-up into the already-running Desktop task through the pipe, and acknowledges only after admission succeeds. The private pipe stays inside the Codex MCP process and is provisioned through a supported handoff — run vanth setup desktop inside a Codex Desktop session with the app-tools capability active (it writes a per-home codex_desktop.json capability file), or launch Vanth with VANTH_CODEX_DESKTOP_PIPE / VANTH_CODEX_DESKTOP_THREAD set. Without a pipe capability the delivery fails closed with an actionable "Desktop integration unavailable" error and is never routed to the CLI.

Experimental. The private host-pipe contract is not documented by official Codex material. This integration is scoped to ONE provisioned task per Desktop lifetime: one per-home codex_desktop.json stores one pipe/thread tuple, provisioning a second Desktop task overwrites the first, and a Desktop restart invalidates the private pipe. A stale capability (older than 24h) is detected and fails closed with a diagnostic asking you to re-run vanth setup desktop inside an active Desktop session. If Desktop restarts mid-operation, the relay reloads a re-provisioned capability and retries the pending wake; otherwise the wake is released back to pending (never terminally consumed) until you re-provision. Automatic or durable multi-task Desktop wake is NOT supported yet. The task must still be running in the current Desktop host lifetime; arbitrary historical/unloaded Desktop threads are not supported. Live tests found that the private host can accept sends to some such threads while producing no usable turn, so admission alone is not a delivery guarantee.

opencode_thread

Resumes an OpenCode session:

{
  "type": "opencode_thread",
  "session_id": "ses_...",              # optional if a plugin is registered
  "events": ["checkpoint", "failed", "completed"],
  "cwd": "F:\\git\\project",
  "opencode_command": ["opencode"],     # override the binary (attach path only)
  "attach": "http://127.0.0.1:4096",    # only for an `opencode serve` instance
  "timeout_seconds": 120
}

A plain opencode TUI is woken by the plugin relay, not by attach. A TUI binds no TCP port and injects no session id, so there is no server URL to attach to and an external opencode run --session writes to a backend the visible session never sees. vanth setup therefore installs a small plugin (~/.config/opencode/plugins/vanth.ts); it registers the session it lives in with the daemon and injects a wake prompt through its own in-process client, so the prompt lands in the session you are watching. Restart opencode once after installing it.

With the plugin loaded, vanth start/job_start can name just {"type": "opencode_thread", "events": [...]}: Vanth resolves the session from the newest relay registered for the job's cwd (an explicit session_id always wins). vanth doctor summarizes registered relays and their liveness; vanth doctor --json includes the full relay/session list. If none is live, opencode_thread wakes cannot be delivered. attach remains supported for headless opencode serve deployments and takes the direct-subprocess path.

The default OpenCode turn timeout is 30 seconds; raise it for long turns.

On Windows, Vanth resolves the standard npm opencode.cmd shim to the native opencode.exe shipped in the same package. This avoids cmd.exe truncating a multiline wake prompt to its first line. Explicit/nonstandard batch shims remain supported; Vanth flattens their prompt line breaks so all wake fields arrive.

Before dispatching to a plain (non-attach) session, Vanth runs a cheap opencode session list probe to confirm the session still exists — a confirmed-missing session fails fast (dead-lettered immediately, no retry burn) with opencode session not found: <id>. The probe never blocks a valid dispatch; on any ambiguity it proceeds. Opt out per-target with "skip_probe": true or globally with VANTH_OPENCODE_SKIP_PROBE=1.

webhook

POSTs the delivery payload as JSON to any HTTP(S) endpoint — a generic channel that covers ntfy, Gotify, Telegram bots, Slack/Discord webhooks, PagerDuty Events, and more:

{
  "type": "webhook",
  "url": "https://hooks.slack.com/services/...",
  "events": ["failed", "completed"],
  "headers": { "Authorization": "Bearer <token>" },
  "timeout_seconds": 10
}

The payload is the same delivery payload every adapter receives (event, prompt, delivery_id, target), POSTed with Content-Type: application/json. 2xx responses (200/201/202/204) mark the delivery delivered; anything else is a failed delivery (retried per max_attempts/retry_delay_seconds, then dead-lettered). headers lets you add auth tokens or presets for specific services.

Shared delivery options

{
  "type": "codex_thread",
  "thread_id": "019f...",
  "events": ["checkpoint"],
  "auto_dispatch": false,      // leave the delivery pending for manual inspection
  "max_attempts": 3,           // default 1
  "retry_delay_seconds": 5,    // default 5
  "timeout_seconds": 30        // adapter timeout; also sizes the delivery lease
}

With auto_dispatch: false, deliveries stay pending until an agent either dispatches them manually or changes the target.

A target that omits events inherits the job's top-level notify_on list. notify_on is not a notification switch: it only supplies default events for an already-supplied wake_targets entry. Without wake_targets it notifies nobody — the job still starts, and the start response carries a warnings entry saying so. An explicit target events always wins:

job_start(command="...", notify_on=["checkpoint","failed"],
          wake_targets=[{"type": "local_command", "command": ["deliver.py"]}])

Concurrent adapter dispatches are capped (default 4) so a burst of events doesn't spawn unlimited adapter processes; excess deliveries stay queued and are picked up on the next dispatch pass. Set VANTH_DELIVERY_MAX_CONCURRENT to tune.

Delivery operations

job_deliveries(job_id="job_...")
job_delivery_attempts(delivery_id="del_...")
job_retry_delivery(delivery_id="del_...")     # requeue a failed OR retrying delivery
job_mark_delivery(delivery_id="del_...", status="delivered")

Attempt history records the claim token, start/end times, status, and whether the attempt was reclaimed after an expired lease. If the daemon crashes after an adapter accepts a wake but before Vanth records success, the delivery is reclaimed and retried — surfaced as a reclaimed attempt rather than claimed as exactly-once delivery.

job_retry_delivery requeues a delivery immediately for dispatch — including one that is currently retrying on backoff (resets next_attempt_at). If a delivery has exhausted max_attempts, it is dead-lettered: vanth doctor reports dead_letter_count and the most recent dead_lettered deliveries (each with delivery_id, job_id, attempts, last_error) so you can see which wakes were never delivered and why.


Running the daemon

Foreground (for development or diagnosis):

uv run vanthd

Start-at-login options:

  • Windows: the daemon is started from the user Startup folder (startup_commands.bat) alongside other startup commands; a Task Scheduler action template is also in deploy/vanthd.cmd.
  • Unix: deploy/vanthd.service is a systemd user service.

Enable only one daemon per VANTH_HOME. A second daemon for the same home exits immediately (OS-level lock). The daemon binds only to loopback (127.0.0.1 / ::1 / localhost); a non-loopback VANTH_DAEMON_HOST is rejected.

Security

  • Every data route requires Authorization: Bearer <token>; the token is generated per home and never logged. GET /health is the only unauthenticated route (a cheap liveness probe for supervisors).
  • On daemon start the state directory is re-tightened to the owner: Unix chmod 0700/0600; Windows disables ACL inheritance and grants only the owner, SYSTEM, and Administrators via icacls. This blocks other accounts (e.g. sandbox/CI users that inherit read from the user profile) from reading the token or per-job env/spec data.
  • On Windows, socket SO_REUSEADDR is disabled so a second daemon cannot become a phantom listener on the same port; a failed bind releases the home lock and exits cleanly.

The Go terminal monitor

The native Go dashboard reads the same home read-only and renders live plots, progress bars, the exact event table, and log tails:

vanth-monitor

Published platform wheels include the native Go binary, so vanth-monitor runs without a Go toolchain. In a source checkout on Windows, build it once from the repo root using cmd.exe (requires Go on PATH), then run the generated executable:

mkdir dist 2>nul & go build -o dist\vanth-monitor.exe .\cmd\vanth && dist\vanth-monitor.exe monitor

The source checkout's vanth-monitor wrapper does not build or discover that output automatically; run the executable produced by go build directly.

Keys: up/down or j/k select jobs · enter pins a job's series · e event table · l log tail · s slowest-runs table · +/- zoom a chart · [/] pan · t back to live tail · ? help · q or Ctrl+C quit.


Configuration reference

Environment variables (defaults live in src/vanth/server.py, src/vanth/daemon.py, src/vanth/migrations.py):

Variable Default Purpose
VANTH_HOME ~/.vanth State root (alias: AGENT_BG_HOME)
VANTH_DAEMON_URL http://127.0.0.1:8765 Where clients reach the daemon
VANTH_DAEMON_HOST 127.0.0.1 Bind address (loopback only)
VANTH_DAEMON_PORT 8765 Bind port
VANTH_CLIENT_TIMEOUT 30s Client socket timeout; <=0 disables (long polls use their own budget)
VANTH_REQUEST_TIMEOUT 30s Daemon blocking socket read/write timeout
VANTH_MAX_REQUEST_BYTES 1 MiB HTTP request body cap
VANTH_MAX_RESPONSE_BYTES 4 MiB HTTP response cap
VANTH_MAX_EVENT_BYTES 64 KiB Single event payload cap
VANTH_MAX_EVENT_LINE_BYTES 1 MiB AGENT_EVENT line cap
VANTH_MAX_LOG_BYTES 10 MiB Per-stream log cap (drain continues)
VANTH_MAX_EVENTS_PER_JOB 100000 Structured event cap per job
VANTH_DELIVERY_POLL_INTERVAL 0.2s Maintenance loop cadence
VANTH_DELIVERY_LEASE_MARGIN 5s Extra lease time beyond adapter timeout
VANTH_RUNNER_HEARTBEAT_INTERVAL 1s Runner liveness heartbeat
VANTH_RUNNER_HEARTBEAT_STALE_AFTER 10s Heartbeat staleness threshold
VANTH_CODEX_BIN codex / C:\codex\codex.exe Codex binary
VANTH_OPENCODE_BIN opencode (via shutil.which) OpenCode binary
VANTH_LOG_LEVEL INFO Daemon log level
VANTH_LOG_MAX_BYTES 5 MiB Rotating daemon log size
VANTH_LOG_BACKUP_COUNT 3 Daemon log rotation count
VANTH_BUSY_TIMEOUT_MS 30000 SQLite write-lock wait

Key knobs in one glance:

Variable Purpose
VANTH_HOME State root (alias: AGENT_BG_HOME)
VANTH_DAEMON_HOST / VANTH_DAEMON_PORT Where the daemon binds (loopback only; default 127.0.0.1:8765)
VANTH_MAX_RUNNING_JOBS Concurrency quota; 0 = unlimited
VANTH_RETENTION_SECONDS / VANTH_RETENTION_INTERVAL / VANTH_RETENTION_DRY_RUN Automatic background retention of old terminal jobs (dry-run by default)
VANTH_NO_SETUP_HINT Suppress the stderr "MCP server not configured" hint
VANTH_OPENCODE_SKIP_PROBE Skip the opencode session list probe before dispatch
VANTH_DELIVERY_MAX_CONCURRENT Cap on concurrent adapter dispatches (default 4)
VANTH_MAX_REQUEST_BYTES HTTP request body cap (default 1 MiB)
VANTH_PROBE_BUDGET Max readiness-probe I/O calls per dispatcher pass (default 8)
VANTH_OUTBOUND_ALLOW Strict allowlist of host/ip/cidr for webhooks + http probes (link-local/metadata always denied)
VANTH_OUTBOUND_BLOCK_PRIVATE 1 also denies loopback + private destinations
VANTH_ALERT_WEBHOOK Edge-triggered operator alert destination
VANTH_ALERT_DISK_FREE_BYTES Free-disk threshold that raises a disk_low alert
VANTH_ALERT_INTERVAL Alert evaluation cadence in seconds (default 30)

Operations

State layout

~/.vanth/
  jobs.sqlite      durable jobs (incl. env, notes, run-overview) / events / deliveries / targets / attempts / tombstones
  artifacts.sqlite managed-artifact catalog (separate DB)
  artifacts-store/ content-addressed artifact blobs (+ staging)
  remote.sqlite    remote-host pairing + transfer journals (when remote is used)
  token            bearer token (owner-only permissions)
  daemon.lock      single-daemon OS lock
  daemon.json      discovery metadata (url, pid, started_at, schema, auth, token_path) — written atomically, removed on graceful shutdown
  logs/            daemon.log + per-job runner/stdout/stderr logs
  events/          per-job JSONL event mirrors (monitor fallback source)
  specs/           per-job launch specs (removed once the runner starts)
  backups/         pre-migration snapshots AND `vanth backup` archives

vanth backup archives jobs.sqlite, artifacts.sqlite, the artifacts-store/ blobs and the events/ mirrors into one verified zip (manifest.json with a SHA-256 per file); vanth restore <archive> --yes verifies, snapshots the current state, and swaps it back (refuses while the daemon looks running, and refuses a backup from a newer schema unless --force).

Health, readiness, and diagnosis

job_doctor()

Reports the state directory, database tables, delivery counts by status, schema version, PRAGMA quick_check, stale delivery leases, free disk, token path, and whether the Codex/OpenCode binaries resolve. It never reveals the token.

The HTTP daemon also exposes:

  • GET /health — cheap, unauthenticated liveness probe for supervisors;
  • GET /ready — authenticated readiness (doctor report; 503 when not ok);
  • GET /metrics — authenticated Prometheus text exposition (jobs by status, running/queued, pools, deliveries, dead letters, stale leases, disk/db size, schema, maintenance aliveness).

Alerting

Set VANTH_ALERT_WEBHOOK to receive edge-triggered operational alerts (one POST per state change, not per tick): the dead-letter queue becoming non-empty, and free disk crossing VANTH_ALERT_DISK_FREE_BYTES. Destinations go through the same outbound policy as webhooks. The payload is {type, condition, active, severity, message, details, at}.

Upgrades and backups

Schema changes are ordered SQLite migrations. Before the first migration of an existing database, a timestamped backup is written under backups/ via SQLite's backup API (never a raw file copy while WAL is active). A future database schema is rejected without touching the files.

For a full off-host copy, vanth backup writes one verified archive of every durable store (see State layout); vanth restore <archive> --yes puts it back after snapshotting the current state. Run vanth backup while the daemon is up (SQLite online backup), but stop the daemon before vanth restore.


HTTP API (equivalent of the MCP tools)

The loopback API base URL is the url in <VANTH_HOME>/daemon.json and the token is stored at <VANTH_HOME>/token. Every route except GET /health needs Authorization: Bearer <token>. Run vanth api for the current surface and route details.

Method Path Purpose
GET /metrics Prometheus text exposition
GET /jobs List jobs (status, limit, thread_id, name, tags); includes lifecycle timestamps, exit code, and runtime
POST /jobs Start a job
POST /jobs/{id}/rerun Rerun a job with its original configuration
POST /jobs/{id}/wake Register a wake target for the job's FUTURE events, after the job started (target)
POST /jobs/{id}/wake-now Register a wake target and surface a synthetic wake immediately (target)
GET /jobs/{id}/status Job status (includes command/env/cwd)
GET /jobs/{id}/events Events (since_event_id, types, limit, reverse)
GET /jobs/{id}/metrics Metric series (metric, from_ms, to_ms, limit)
GET /jobs/{id}/summary Run summary (status, runtime, metrics, artifacts)
GET /jobs/{id}/artifacts Artifacts (limit)
POST /jobs/{id}/artifacts Add an artifact
GET /metrics/compare Compare metric across jobs (job_ids, metric, aggregation)
GET /analytics/durations Duration/flakiness analytics (name, tags, limit, since_ms, slowest)
GET /dashboard Chart data (job_ids, limit)
GET /jobs/{id}/tail Log tail (stream, max_bytes, offset)
POST /jobs/{id}/wait Wait for an event
POST /jobs/{id}/stop Stop a job
POST /jobs/{id}/pause / /resume Hold / release a queued job
GET /schedules List schedules
POST /schedules Create a schedule
POST /schedules/{id}/update / /delete Edit in place / delete
GET /schedules/{id}/next Next fire times (count)
GET /pools List pools with queue depths
POST /pools Configure a pool (pool, max_parallel, paused)
GET /view Agent view (thread_id, limit)
GET /deliveries Deliveries (job_id, status, limit)
GET /deliveries/{id}/attempts Attempt history
POST /deliveries/{id}/mark Mark a delivery
POST /deliveries/{id}/retry Retry a delivery
POST /cleanup Cleanup (older_than_seconds, dry_run)
GET /doctor Health report
GET /health Unauthenticated liveness
GET /remotes Paired remote hosts
GET /remotes/doctor SSH binaries + remote state (remote_id)
GET /remotes/{id}/jobs Remote jobs from the controller's shadow (limit)
GET /remotes/{id}/status/{job} One remote job's status
GET /remotes/{id}/jobs/{job}/tail Remote log range (stream, offset, size)
POST /artifacts/push-remote Publish an artifact to a remote (remote_id, version_id)
POST /artifacts/pull-remote Fetch a remote artifact (remote_id, version_id, dest_path)

Remote execution

Running jobs on another host is beta (POSIX targets). Pairing is interactive, so it lives in the CLI; everything after that is available from MCP, the CLI, or HTTP.

vanth remote pair user@host        # one-time; writes ~/.vanth/remote.sqlite
vanth remote list                  # -> remote ids
vanth remote doctor                # SSH binaries + per-host state

Then target that host with the remote_id:

Task MCP HTTP
See hosts + ids remote_list GET /remotes
Start a job there job_start(remote_id=...) POST /jobs with "remote_id"
List its jobs job_list(remote_id=...) GET /remotes/{id}/jobs
Status / wait / stop / rerun job_status / job_wait / job_stop / job_rerun with remote_id /remotes/{id}/status/{job} etc.
Read its logs job_tail(job_id, remote_id=...) GET /remotes/{id}/jobs/{job}/tail

Notes: remote mutations require a caller-supplied idempotency_key (8–128 chars of [A-Za-z0-9_-]) — that is what makes a lost response safe to retry — and the daemon rejects a missing one (a local start must not pass one). Read costs differ: job_list(remote_id=...) reads the controller's shadow with no SSH round trip (and accepts no filters beyond limit), while job_status(remote_id=...) makes a live status request to the host, and job_tail(..., remote_id=...) reads one byte range of the remote log over the protocol (no follow/grep). vanth remote pending / retry reconcile requests whose response was lost.

Remote wakes. job_start(remote_id=..., wake_me=True) (or wake_targets=) is supported: the wake targets are registered on the LOCAL daemon (never sent to the host). The local daemon polls the host's retained structured events, so a remote wake can match any event type (including checkpoint, progress, and metric) after the binding is registered. This is best-effort for non-terminal events, subject to the remote event cap/retention and the poll interval; terminal wakes remain durable via the change feed. The controller polls every VANTH_REMOTE_WAKE_SYNC_SECONDS (default 5; set 0 to disable). Settled controller request/journal rows from that polling are pruned after VANTH_REMOTE_REQUEST_TTL_SECONDS (default 604800 = 7 days; set 0 to disable), checked every VANTH_REMOTE_PRUNE_INTERVAL_SECONDS (default 3600).


Agent usage tips

  1. Wait, don't poll. Use job_wait(job_id, filters=[...], timeout_seconds=...) instead of looping job_status. The daemon wakes the wait immediately when a matching event is persisted.
  2. Pass since_event_id to the next job_wait after handling an event, so you never re-process an old one.
  3. Tag and thread your jobs. Set origin_thread_id (the agent thread that launched the job) and tags; use job_view(thread_id=...) to summarize.
  4. Prefer job_view over job_status when presenting a situation to a user — it is already sorted by attention priority.
  5. Make jobs self-describing. Emit AGENT_EVENT progress / checkpoint / metric lines (see above). Jobs that are silent still work, but tracked jobs are far easier to reason about.
  6. Use wake targets for long jobs. If a training run or long download needs a decision at a checkpoint, add a codex_thread or opencode_thread target with events: ["checkpoint", "failed", "completed"] so the agent is resumed instead of polling.
  7. Inspect delivery failures. job_delivery_attempts shows the lease/claim history; job_retry_delivery requeues a failed one after fixing the cause.
  8. Set a sane timeout_seconds on job_start so a hung command becomes a timeout (terminal) state instead of running forever; the runner enforces it even across daemon restarts. Pass the whole shell command as ONE quoted string: a shell operator passed as its own argument (&&, |, >nul, ...) is refused. Prefer --wake-me over hand-written wake JSON.
  9. Clean up old state with job_cleanup(older_than_seconds=..., dry_run=false) so the SQLite store and log files stay bounded.
  10. Rerun failed jobs, don't rebuild them. job_rerun(job_id=...) relaunches with the original command, env, cwd, and wake targets — ideal for retrying a transiently failed download or batch.
  11. Ask "what is this job?" with job_status. It now returns the command, cwd, env, and timeout, so you can explain a job to a user without reading logs.
  12. Filter lists by name/tag. job_list(name="train", tags=["gpu"]) narrows a growing job list without paging through everything.
  13. Use reverse=true for "what happened recently." job_events(job_id, reverse=true, limit=20) returns the newest events first, and you can page further back with since_event_id set to the oldest id you've seen.
  14. A job survives the daemon. The runner is detached; jobs continue across daemon/MCP restarts. If a runner is gone at recovery, the job is marked orphaned (never silently dropped).

Examples

uv run python examples\long_job.py    # emits progress + checkpoints

examples/long_job.py is a small reference job that uses vanth.agent_events. Start it through job_start and watch it in vanth monitor.


Troubleshooting

  • Unauthorized (401): the bearer token in ~/.vanth/token is what the daemon expects. Confirm VANTH_HOME is the same for the daemon and client.
  • Second daemon won't start: another vanthd already owns this VANTH_HOME. One daemon per home by design.
  • Daemon startup reports a bind/port error: another process is using the configured loopback port (default 8765). Stop the conflicting process or choose another port by setting VANTH_DAEMON_PORT consistently for the daemon and its clients, then retry. A different VANTH_HOME prevents state collisions between isolated instances but does not free a port shared by them; assign each concurrent home a distinct port.
  • Job stuck running then orphaned: the runner process died. Check logs/<job_id>.runner.log and the heartbeat thresholds.
  • No charts in the monitor: the job isn't emitting AGENT_EVENT metric or progress lines — add them (optional).
  • OpenCode wake never arrives (TUI session): check vanth doctor for the relay summary, or vanth doctor --json for all relay/session details. If no opencode_thread relay is live, the plugin is not loaded. Re-run vanth setup opencode and restart opencode. Without a live relay an opencode_thread target is rejected at creation rather than silently lost.
  • Wake will never fire: inspect vanth deliveries --status pending (or vanth deliveries --status failed); vanth doctor reports pending_deliveries and undeliverable_wakes.
  • OpenCode wake timing out: increase timeout_seconds on the wake target beyond the expected turn length.
  • OpenCode wake failed with Session not found: the wake target's session_id is stale or was removed. Vanth now probes the session before dispatching (opencode session list) and fails fast with opencode session not found: <id> instead of retrying a dead session. Refresh the wake target's session_id (or start a new session) and job_retry_delivery to re-dispatch. Per-target opt-out: skip_probe: true; global opt-out: VANTH_OPENCODE_SKIP_PROBE=1.
  • Codex wake failed with no rollout found for thread id: the target thread has never had a turn. Start a first turn in that thread (or target an existing, active conversation) before waking it.
  • Monitor shows nothing / empty state: confirm VANTH_HOME points at the daemon's home, and that jobs.sqlite exists there.

Development

uv run pytest -q
uv run python -m compileall -q src tests examples
uv build
go vet ./...
go test ./...

uv build produces the sdist and wheel; the wheel bundles the Go monitor.

The wheel build runs a hatchling build hook (build-hooks/bundle_monitor.py) that compiles the Go monitor for the host platform and bundles it under vanth/monitor-bin/ so vanth-monitor needs no Go toolchain at runtime. go must be on PATH when building the wheel; it is not needed to install or run it. Wheels are platform-tagged (py3-none-<platform>) because they contain the native binary.

Release-gate automation lives in scripts/:

  • scripts/chaos_matrix.py — heavy synthetic workloads and kill/restart matrix;
  • scripts/real_adapter_smoke.py — opt-in live Codex/OpenCode wake smokes (set VANTH_SMOKE_CODEX_THREAD / VANTH_SMOKE_OPENCODE_SESSION);
  • scripts/generate_go_fixture.py — regenerates the deterministic schema-v5 conformance fixture in testdata/;
  • scripts/demo_jobs.py — starts demo jobs (training run, quick task, failing task) for the monitor.

Limitations (v1)

  • Delivery is at-least-once; a crash after an adapter accepts a wake but before Vanth records success is a documented, surfaced ambiguity.
  • Remote SSH execution and managed artifacts are beta (POSIX targets; single-node controller) and are not continuously live-tested against real SSH hosts. Codex Desktop wake is experimental (see Wake targets).
  • Outbound webhook/probe destinations are governed by the policy in the configuration table; link-local/cloud-metadata addresses are always refused.
  • TLS, multi-user policy/RBAC, distributed workers, and a web UI are out of scope.

About

Event-driven background jobs for agents — a localhost job daemon with an MCP interface. Start detached commands, track them via structured AGENT_EVENT progress/metrics/checkpoints, wait on durable SQLite state, and wake Codex or OpenCode sessions when a job needs attention.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages