Build task-execution agents from Kubernetes-shaped resources, pluggable runners, and optional HTTP or gRPC boundaries.
Solti is a modular Rust SDK. It provides the resource model, routing, reconciliation, execution backends, APIs, discovery, TLS, logging, and metrics used by an agent binary.
Your binary selects the required crates or enables them through the solti umbrella crate.
Your binary still owns configuration, deployment, and the final security boundary.
Solti uses Taskvisor for supervised attempt lifecycles.
| Quick start | Architecture | Platform limits | Examples |
A task agent needs more than process spawning. It must validate desired state, select a runtime, reconcile changes, supervise attempts, expose status, and shut down cleanly.
Solti separates those responsibilities:
| Concern | SDK boundary |
|---|---|
| Resource contract | Kubernetes-shaped Task, metadata, spec, status, and watches |
| Workload selection | GVK routing plus optional runner label selectors |
| Desired state | Asynchronous latest-wins reconciliation |
| Attempt lifecycle | Taskvisor restart, timeout, cancellation, and admission |
| Execution | Subprocess, native containerd 2.x, or application-owned runner |
| Public API | HTTP/JSON or gRPC |
| Agent registration | Outbound HTTP or gRPC discovery |
| Operations | TLS, logging, Prometheus, and live output |
Each layer has a direct crate. The umbrella crate only forwards features and namespaces.
Use Solti when a binary needs versioned Task resources, pluggable workload kinds, desired-state reconciliation, or a public agent API.
If the process only needs to supervise ordinary async functions, use Taskvisor directly. It has a smaller API and no resource or network layer.
Solti is not a durable job system or a control plane. Core state and live output are process-local. A compatible control plane, persistent storage, authorization policy, and deployment topology remain separate concerns.
Add the umbrella crate with only the required capabilities:
[dependencies]
solti = { version = "0.0.4", features = ["core", "exec-subprocess"] }
tokio = { version = "1", features = ["macros", "rt", "time"] }Put this in src/main.rs, then run cargo run.
The parent registers a subprocess runner and supervises one execution of the same binary.
use std::{env, io, time::Duration};
use solti::{
core::SupervisorApi,
exec::subprocess::register_subprocess_runner,
model::{
Flag, RestartPolicy, SubprocessMode, SubprocessSpec, TaskEnv, TaskManifest, TaskSpec,
TaskWorkload,
},
runner::RunnerRouter,
};
const CHILD_MODE: &str = "--solti-quick-start-child";
#[tokio::main(flavor = "current_thread")]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
if env::args().nth(1).as_deref() == Some(CHILD_MODE) {
println!("hello from the supervised subprocess");
return Ok(());
}
let mut router = RunnerRouter::new();
register_subprocess_runner(&mut router, "default")?;
let supervisor = SupervisorApi::builder(router).start().await?;
let command = env::current_exe()?.to_string_lossy().into_owned();
let workload = TaskWorkload::Subprocess(SubprocessSpec::new(
SubprocessMode::Command {
command,
args: vec![CHILD_MODE.into()],
},
TaskEnv::new(),
None,
Flag::enabled(),
));
let spec = TaskSpec::builder("quick-start", workload, 30_000_u64)
.restart(RestartPolicy::Never)
.build()?;
let committed = supervisor
.create_task(TaskManifest::new("quick-start", spec)?)
.await?;
let name = committed.name().clone();
println!("committed {name}");
let phase = tokio::time::timeout(Duration::from_secs(35), async {
loop {
let task = supervisor
.get_task(&name)
.ok_or_else(|| io::Error::other("task disappeared"))?;
let phase = task.status().phase();
if phase.is_terminal() {
break Ok::<_, io::Error>(phase);
}
tokio::time::sleep(Duration::from_millis(25)).await;
}
})
.await??;
println!("finished {name}: {phase}");
supervisor.shutdown().await?;
Ok(())
}create_task returns after desired state is committed.
Reconciliation and execution continue asynchronously.
The complete lifecycle with live output and history is in task_subprocess.rs.
All solti features are disabled by default.
Higher-level features enable their required lower layers.
| Binary requirement | Start with |
|---|---|
| Resource types and JSON Schema | model |
| Custom runner registration | runner |
| Conditional sequential workloads | chain |
| In-process desired-state runtime | core |
| Subprocess task runtime | core, exec-subprocess |
| Native containerd task runtime | core, exec-containerd |
| HTTP Task API | api-core-adapter, api-http, exec-subprocess |
| gRPC Task API | api-core-adapter, api-grpc, exec-subprocess |
| gRPC Task API with TLS or mTLS | api-core-adapter, api-grpc-tls, exec-subprocess |
| HTTP agent with discovery | api-core-adapter, api-http, discover-http, exec-subprocess |
| Complete standard integration set | full |
exec-containerd is a native containerd 2.x adapter.
It is not a Docker, CRI, or Docker Compose integration.
The component graph is acyclic. Arrows in the diagram point toward the dependency or runtime contract being consumed.
solti-core never depends on solti-exec.
Execution backends implement the solti-runner contract and are registered by the binary.
solti-api depends on core only through core-adapter.
Its handler and transport boundaries can be used with another implementation.
solti-discover is an outbound client.
It does not run the agent API or depend on core.
solti-tls has no SDK dependencies.
solti-prometheus connects to producers only through feature-selected adapters.
| Crate | Owns |
|---|---|
solti |
Umbrella feature forwarding and canonical namespaces |
solti-model |
Resources, workloads, policies, selectors, capabilities, and tokens |
solti-runner |
Runner contract, GVK routing, selectors, and execution context |
solti-chain |
Conditional sequential composition of nested workloads |
solti-core |
Desired state, reconciliation, watches, history, and live output |
solti-exec |
Execution backends and host-process controls |
solti-api |
HTTP/JSON and gRPC Task APIs |
solti-discover |
Agent registration and heartbeat client |
solti-tls |
TLS and mTLS identities, trust roots, and rustls configuration |
solti-observe |
Structured logging and supervised timezone refresh |
solti-prometheus |
Metrics adapters, collectors, and exporter endpoint |
Depend on one component crate when one boundary is enough.
Use solti when a binary composes several components.
Solti follows Kubernetes resource conventions.
A Task has apiVersion, kind, metadata, spec, and status.
The caller owns desired fields. Core owns UID, resource version, generation, timestamps, status, and conditions.
Built-in workload kinds are Subprocess, Container, Wasm, and Embedded.
Wasm is a model contract; this repository does not provide a built-in WASM runner.
Application-owned workload kinds use ExtensionWorkload with a non-solti.io GVK and strict runner-side decoding.
Runner routing uses workload GVK and an optional Kubernetes-style label selector.
Embedded carries an in-process TaskRef supplied by the binary.
It bypasses runner routing and has no HTTP or gRPC representation.
The optional solti-chain runner represents a Chain as one ordinary Task.
Its steps are nested workloads, and exactly one step is active at a time.
Each successful or failed step may select one next step.
Chain uses a regular extension workload under Task.spec.workload.
Existing HTTP and gRPC Task operations carry it without a new resource API.
The outer Task owns timeout, restart, backoff, admission, cancellation, status, history, and output. Steps are not child Task resources and do not have independent lifecycle policies. Restarting the outer Task starts the chain again from its entry step.
Create and apply commit desired state before runtime work begins.
The resource status reports the observed result through phase, generation, attempt, and the Reconciled condition.
Reconciliation is latest-wins. An older generation is discarded before expensive runner construction when a newer generation is already committed.
There is no staged rollout or availability guarantee. In-flight side effects may finish before the next reconciliation compensates for them.
Core does not run an infinite reconciliation retry queue.
Applying an identical manifest schedules one manual retry only when Reconciled=False.
Taskvisor owns attempt restart, backoff, timeout, admission, and cancellation.
Core retains bounded TaskRun history separately from the current Task status.
solti-api exposes the same task operations over HTTP/JSON and gRPC:
| Operation | Result |
|---|---|
| Create | Commit a new named Task |
| Apply | Create or update desired state |
| Get | Read one Task |
| List | Filter and paginate a stable collection snapshot |
| Watch | Stream retained changes and then live changes |
| Run history | Read retained attempts |
| Logs | Stream live stdout and stderr |
| Delete | Stop the runtime and remove the Task and its history |
HTTP uses the fixed root /apis/solti.io/v1 in the current solti-api release.
HttpApi::build returns the router and its generated OpenAPI 3.1 document.
gRPC uses the solti.task.v1 protobuf package.
Generated server and client types are available from the crate.
Each SDK binary serves the API version compiled into its selected solti-api version.
A control plane that manages different binary generations must route each advertised version to its matching contract.
Bearer authentication is disabled until the binary calls with_auth or with_authenticator.
An optional with_authorizer policy runs before each validated Task API operation.
The HTTP and gRPC boundaries enforce a 4 MiB request or message limit.
Read the complete route, message, pagination, watch, and error contract in solti-api/CONTRACT.md.
solti-discover advertises one agent endpoint, API version, runner capabilities, identity, and uptime to a control plane.
The discovery loop is returned as an Embedded manifest and TaskRef.
The binary decides whether to submit it to core.
HTTP and gRPC transports are independent features. Retryable transport failures use the generated task's Taskvisor policy. Permanent configuration or authentication failures stop the task.
Discovery does not expose an inbound server. It does not persist registration state.
Read the versioned wire contract in solti-discover/CONTRACT.md.
solti-model::Token is a redacted bearer secret with constant-time comparison.
solti-api can verify it on every route and RPC.
solti-discover can send it with every heartbeat.
The static token is authentication only.
solti-api also exposes application hooks for bearer authentication and operation-level authorization.
The SDK does not provide users, tenants, RBAC rules, tenant filtering, policy storage, secret rotation, or secret persistence.
solti-tls separates server identity, client identity, and trust roots.
It supports TLS and mandatory client-certificate authentication.
api-grpc-tls converts the shared server configuration for tonic.
HTTP server TLS is owned by the server that hosts the axum router.
discover-tls applies custom roots or mTLS to outbound HTTPS connections.
Bearer authentication and TLS are independent.
Native containerd execution is Linux-only. macOS can build the adapter and inspect its configuration, but it cannot start a native container attempt.
| Runtime path | Current platform contract | Isolation boundary |
|---|---|---|
Embedded |
Application's supported Tokio targets | Same process; no operating-system isolation |
Subprocess |
Linux and macOS | Child process; Unix session and process group |
| Non-Unix subprocess path | Implementation exists; Windows is not currently supported | Child process only |
| Custom container engine | Defined by the application adapter | Defined by that engine |
| Native containerd 2.x | Linux host and Linux container images | OCI runtime plus containerd task lifecycle |
The subprocess path always validates configuration before runner registration. Configured controls fail closed when the current platform cannot enforce them.
| Control | Platform |
|---|---|
| Session, process group, signal reset, umask | Unix |
RLIMIT_NOFILE, RLIMIT_FSIZE, core dumps |
Unix |
| Pinned working directory | Unix |
| Explicit descriptor passlist | Linux |
| Descriptor snapshot and close-on-exec checks | Other Unix |
| cgroup v2 CPU, memory, and process limits | Linux |
| Mount, network, IPC, UTS, and cgroup namespaces | Linux |
| UID, GID, supplementary groups, capabilities | Linux |
no_new_privs and seccomp denylist |
Linux |
The default subprocess backend does not enable optional resource or security controls. It clears the inherited environment, pins the working directory on Unix, restricts descriptor inheritance, and owns child cleanup.
These controls harden a host process. They do not form a complete sandbox for untrusted code.
The built-in adapter connects to one explicit Unix socket and namespace. It does not start or discover containerd.
It requires:
- a Linux host;
- containerd major version 2;
- configured snapshotter and OCI runtime plugins;
- a cached or reachable Linux image;
- an I/O root visible at the same path to the SDK process and containerd;
protocduring builds that compilecontainerd-clientbindings.
Network mode is either none or host.
none creates an OCI network namespace without configuring interfaces.
host shares the host network namespace.
The adapter does not provide bridge networking, CNI, CRI, port publishing, volumes, or Docker Compose semantics. The final binary must add those layers if its product requires them.
Run task_containerd.rs on a prepared Linux host. On other platforms the example prints the prerequisite and exits without contacting a daemon.
Keep these boundaries explicit:
- Core stores Tasks, runs, watch history, and runtime bindings in memory.
- Process restart loses all core state.
- Live output is bounded and lossy.
- Core does not persist or replay output by itself.
- Optional core hooks can forward task, run, and output events to an application-owned store.
- A slow output subscriber receives
Laggedafter events are dropped. - Watch history is bounded by change count and serialized Task bytes.
- A watch can resume only while its resource version remains retained.
- Snapshot pagination is consistent only while its continuation remains valid.
- Reconciliation is latest-wins and has no staged availability guarantee.
- Discovery registration state is not persisted.
- Static bearer authentication alone does not provide authorization or tenant isolation.
- Host-process controls are hardening, not a complete untrusted-code sandbox.
- The native container adapter provides no CRI or CNI implementation.
- The SDK contains no durable log sink.
- The SDK contains no control-plane server.
Use the persistence hooks with an external store when tasks or logs must survive process termination. The application owns delivery retries and its restart recovery flow. The SDK does not load persisted state at startup. Install application authorization and sandbox policy at the binary or service boundary.
All umbrella features are off by default.
| Feature or family | Adds |
|---|---|
model |
solti-model with JSON Schema support |
runner |
Runner contract, model, and Taskvisor |
core |
Desired-state supervisor and Taskvisor controller |
exec |
Base solti-exec namespace |
exec-host-process |
Low-level host-process policy |
exec-subprocess |
Subprocess runner and required lower layers |
exec-container |
Engine-neutral container runner |
exec-containerd |
Native containerd 2.x adapter |
exec-seccomp |
Linux host-process seccomp renderer |
api |
API handler and model boundary |
api-http, api-grpc |
HTTP or gRPC transport |
api-core-adapter |
Adapter from the API handler to SupervisorApi |
api-grpc-tls |
Shared TLS conversion for tonic |
discover |
Base discovery contracts |
discover-http, discover-grpc |
Outbound discovery transport |
discover-tls |
Custom roots or mTLS for discovery |
observe |
Logging configuration |
observe-* |
Journald, log compatibility, or timezone refresh |
prometheus-base |
Prometheus namespace and base contracts |
prometheus |
Runner and Taskvisor-controller metrics bundle |
prometheus-* |
API, discovery, process, server, state, runner, or Taskvisor adapters |
prometheus-full |
Every Prometheus adapter |
taskvisor-* |
Forwarded Taskvisor integrations |
tls |
Shared TLS and mTLS types |
full |
Complete standard integration set |
exec-seccomp provides the filter implementation.
Combine it with exec-subprocess to apply the filter to subprocess attempts.
api-http and api-grpc do not enable core.
Add api-core-adapter when the public API delegates to SupervisorApi.
full compiles the native containerd adapter.
Container execution remains Linux-only.
From a cloned checkout, start with the direct Task lifecycle:
cargo run -p solti --example task_subprocess \
--features core,exec-subprocessNames identify the boundary:
task_*calls the in-process Task lifecycle directly;agent_*assembles an API or discovery boundary;operations_*composes metrics, logging, and maintenance.
| Example | Features | Result |
|---|---|---|
| task_chain.rs | chain,core,exec-subprocess |
Conditional steps with failure recovery |
| task_subprocess.rs | core,exec-subprocess |
Output, reconciliation, terminal status, and history |
| task_custom_workload.rs | core |
Application-owned TcpProbe GVK and runner |
| task_containerd.rs | core,exec-containerd |
Native containerd 2.x attempt; Linux runtime required |
| Example | Features | Result |
|---|---|---|
| agent_http.rs | api-core-adapter,api-http,exec-subprocess |
HTTP Task API, OpenAPI, and runnable curl calls |
| agent_grpc.rs | api-core-adapter,api-grpc,exec-subprocess |
gRPC Task API, bearer auth, and grpcurl calls |
| agent_grpc_mtls.rs | api-core-adapter,api-grpc-tls,exec-subprocess |
Anonymous rejection and authenticated mTLS client |
| agent_http_discovery.rs | api-core-adapter,api-http,discover-http,exec-subprocess |
Inbound Task API and outbound discovery heartbeat |
| Example | Features | Result |
|---|---|---|
| operations_prometheus.rs | core,exec-subprocess,prometheus,prometheus-server,prometheus-state |
Supervised /metrics with real runtime samples |
| operations_observe.rs | core,exec-subprocess,observe-timezone-sync |
Logging and supervised timezone maintenance |
agent_http and agent_grpc remain active until Ctrl-C.
They print commands that can be run from a second terminal.
operations_prometheus serves http://127.0.0.1:9090/metrics until Ctrl-C.
Set SOLTI_METRICS_ADDR to use another listen address.
task_containerd requires Linux and an accessible containerd 2.x daemon for a real attempt.
Vendor the pinned protobuf contracts after a fresh checkout:
task proto/vendorThe task fetches the revision pinned in Taskfile.yml.
Generated protobuf trees are ignored by Git and included in published transport crates.
Run the workspace checks:
task ci/fmt
task ci/clippy
task ci/test
task ci/docs
task ci/publish-dry-runThe release order is declared in .github/crates.txt.
Component crates are published before the solti umbrella crate.
Issues and pull requests are welcome. Start with the relevant crate README.
Read the architecture guides before changing the model, core, or execution boundaries. Read the Task API contract or discovery contract before changing wire behavior.
Read the contributing guide before a large change.
If Solti helps your project, a GitHub star helps other Rust developers find it.

