Skip to content
pity11Public

About

Evidence-first agent framework for authorized web and cloud-native vulnerability research

Resources

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

VulnLoom

VulnLoom is an in-development, evidence-first autonomous vulnerability research and adversarial validation platform for software and systems that organizations own or are contracted to assess. Its two primary capability lines are Source Hunt and Authorized Red Team. The target product accepts source code, URLs, domains, IP addresses, networks, or combinations of them; plans bounded research tasks; validates Candidates in isolated environments; challenges them through an independent review step; and produces auditable reports. Pre-release acceptance and production-safe scheduled testing package these capabilities into operational workflows. The current implementation includes a bounded, resumable local Source Hunt V1 and fixed native Benchmark, shared Candidate-to-report gates, a closed Hybrid R10 evidence/retest/release-gate slice, and an Authorized Red Team R11 bounded attack-chain plus B1 observation-driven bounded replanning. B2.1 supports offline-tested, operator-sealed path GET observations: redirects, headers, bodies and credentials are disabled, while ordinary state receives only response metadata and digests. B2.2 can reduce the redacted Evidence from exactly one such completed GET into bounded OpenAPI 3.x path/method discoveries; server metadata and $ref values are never followed, and discoveries carry no execution authority. B2.3 adds a reviewed promotion gate: only explicitly selected, read-only discoveries whose templates have been concretized by an operator can atomically become a new Endpoint Seed Set, still without request execution authority. General public Recon, crawler, dictionary enumeration, arbitrary target expansion, lateral movement, and persistence are not enabled.

The product supports autonomous testing only inside an explicit, approved Scope. It must not scan or exploit unauthorized public targets. Its four planned entry points are source vulnerability research, pre-release security acceptance, production-safe scheduled testing, and authorized red-team simulation. Source Hunt and Authorized Red Team are the primary capability lines; the other two entries are controlled delivery workflows. Scope, network boundaries, credential isolation, evidence requirements, and human approval for consequential effects are enforced in code.

Current implementation scope

  • White-box analysis of Python Web and API projects.
  • Controlled dynamic validation in exact local/private fixtures; broader authorized live Web target execution remains staged and disabled by default.
  • IDOR/BOLA, SSRF, path traversal, injection, insecure deserialization, authorization flaws, and sensitive data exposure.
  • Planned human-reviewable report drafts for vendors, EduSRC, CNVD/CNNVD, and similar disclosure channels.
  • No scanning of unauthorized public targets, automatic platform submission, or automatic CVE requests.

The long-term product design review draft and phased migration plan live in docs/PRODUCT-ARCHITECTURE-ROADMAP.md. Capabilities described there remain planned until their explicit acceptance stage passes.

Core principles

  1. A Candidate is not a Finding. An agent may propose a candidate, but it only becomes a finding after reproducible evidence and an independent disproof check pass deterministic gates.
  2. The Control Plane owns privileged actions. Workers do not receive platform credentials, change authorization scope, or submit reports.
  3. Every validation uses an ephemeral sandbox. Source code is mounted read-only, output is stored separately, and network access is denied by default.
  4. Tools go through a Broker. Agents receive typed, limited, and auditable capabilities instead of an unrestricted host shell.
  5. Evidence comes before narrative. Every impact claim in a report must trace back to code, an observed request and response, or a reproducible test.
  6. Human approval is mandatory where it matters. State-changing tests, external callbacks, real credentials, and report submission require explicit approval.

Documentation

Project layout

VulnLoom/
├── src/vulnloom/
│   ├── adapters/           # Model-provider configuration boundary
│   ├── agent_runtime/      # Typed offline model replay and proposal boundary
│   ├── analyzers/          # Python AST and optional Semgrep analysis
│   ├── benchmark/          # Deterministic offline metrics and regression gates
│   ├── broker/             # Typed tool mediation and HTTP policy enforcement
│   ├── critic/             # Deterministic independent counterevidence review
│   ├── domain/             # Domain objects, state machines, and protocols
│   ├── evidence/           # Redaction, hashing, and evidence storage
│   ├── hypotheses/         # Deterministic Candidate generation
│   ├── ingestion/          # Archive, Git, and OCI target ingestion
│   ├── policy/             # Scope and approval enforcement
│   ├── reporting/          # Evidence-consistent offline report drafts
│   ├── red_team/           # Authorized Red Team Flow, RoE, and Recon control plane
│   ├── runners/            # Offline and Docker sandbox runners
│   ├── source_hunt/        # Resumable white-box investigation and validation chain
│   ├── storage/            # Event and validation persistence
│   ├── validation/         # Plans, orchestration, and deterministic judging
│   ├── workflows/          # Shared visibility and execution-mode contracts
│   └── cli.py              # Current command-line entry point
├── benchmarks/             # Sealed local ground-truth fixtures and baselines
├── docs/                   # Architecture, workflow, security, and roadmap
├── schemas/                # Exported JSON Schema contracts
├── scripts/                # Schema and development utilities
└── tests/                  # Offline tests and opt-in integration probes

An HTTP API and disclosure submission adapters remain planned components. A bounded live HTTPS provider adapter exists; CUC live acceptance covers the fixed PONG and fixed JSON probes. The target design treats it as the first Provider Profile within a multi-model Provider Center, rather than as a product-wide special case. CUC/DeepSeek remains the default provider for early development and explicitly authorized live acceptance; its existing probe, code-review, and Candidate-recommendation paths stay available until the generic adapter passes differential and live compatibility acceptance. A feature-gated OpenAI-compatible Chat Completions codec now has an offline fake-transport slice with strict typed decisions, model and usage checks, rejection, timeout, and cleanup tests. The admitted CUC probe identities are frozen in a content-addressed migration baseline. This new codec is not yet the default CUC path and has not received live CUC acceptance. The first trusted Profile assembly path now resolves only an explicitly allowlisted Control Plane endpoint reference, rechecks the admitted Flow and current Provider lifecycle, prepares content-addressed transport/codec contracts without network access, and creates a model registration only after the authoritative Egress Store proves a grant is active. That assembly now drives complete local fake-process Agent turns for two distinct Provider/model configurations. Negative turns cover authentication rejection, rate limiting, timeout, malformed output, zeroed credential leases, cleared wire buffers, and absence of receipts on failure. The generic path has also passed one authorized synthetic CUC Agent-decision acceptance run with a short-lived revoked grant, strict response identity and usage validation, zero tools, and verified buffer cleanup. Existing CUC feature paths remain unchanged while default-route migration is reviewed separately. The read-only code-review and Candidate Recommendation services can now bind their task-specific schemas to that trusted Profile path. A second synthetic Provider covers routed success, identity rejection, timeout, cleanup failure, replay, secret-buffer cleanup, and Candidate immutability; the CUC configuration remains the CLI default. Provider Center's trusted local CLI/application service supports referenced configuration, lifecycle, offline probing, revision-bound catalogs, audit, and default Route switching; its thin HTTP API and Web UI remain incomplete. M7.1a-M8.12 include deterministic replay, fixed provider messages, scoped credentials, isolated pinned HTTPS transport, typed Broker handoff, and a fixed two-tool Session ledger, human-gated Validation/Critic/Finding Intakes, and Approval-gated promotion. Benchmark and analyzer imports consume only sealed, pre-obtained local data and never fetch suites, rules, databases, or images.

Historical white-box vertical slice

The first implementation vertical slice was deliberately narrow:

Approved scope
→ Import one Python Web repository
→ Produce static security signals
→ Build a cross-file SourceGraph
→ Let a human select one Candidate
→ Start the test application in Docker
→ Validate the Candidate
→ Run an independent Critic
→ Produce a Markdown report draft

The current implementation now also exposes the Source Hunt V1 application service and nested CLI: bounded multi-language indexing, observation-driven cross-file investigation, integrity-checked redacted source windows, Candidate materialization, deterministic five-stage execution planning, resumable validation, authoritative Critic/Finding promotion, and report drafting. R9 fixed-Benchmark acceptance adds a sealed native coverage/ASAN adapter, normalized persistent Crash deduplication, and independently replayed PoV qualification. A1 closes the content-addressed Project Recipe Registry through network-disabled Docker admission and an authoritative non-fixture Recipe→Candidate→Validation binding; Recipe success never substitutes for vulnerability reproduction. Blind-holdout acceptance remains later depth work. External disclosure remains a separate future stage.

Unauthorized Internet-wide asset discovery, automatic submission, and unbrokered host shell access remain out of scope. The first isolated local HTTP HEAD Recon admission is implemented; broader authorized Web reconnaissance remains staged in the product architecture roadmap.

Current implementation

Phase 0: domain and safety foundation

  • Immutable Pydantic domain models and exported JSON Schema contracts.
  • Multi-provider lifecycle, capability manifests, role routes, bounded fallback policies, and immutable Flow model snapshots; Provider endpoint and credential values remain outside these contracts.
  • Separate Candidate state machine and deterministic Candidate-to-Finding gate.
  • Scope Policy Engine and approvals bound to specific action digests.
  • Evidence redaction, content addressing, and integrity verification.
  • Checkpointed SQLite Control Plane event log with atomic tamper-evident audit records, rollback/fork detection, redacted projections, and owner-only filesystem checkpoint custody; remote signing and WORM remain deployment-stage work.
  • Typed Control Plane/Worker protocol and an explicit Worker environment allowlist.
  • engagement-create, scope-approve, and status CLI commands.
  • An OpenAI-compatible model-provider configuration boundary; no network model call is made yet.

M1: secure target ingestion

  • Streaming, size-limited quarantine for ZIP and TAR artifacts.
  • Member-by-member archive extraction without extractall().
  • Validation of normalized paths, symbolic links, special files, file counts, individual sizes, total expanded size, and compression ratio.
  • Exact commit pinning for local Git repositories by reading Git objects without checkout, hooks, or target-code execution.
  • Static classification of Kubernetes, Helm, Terraform, Dockerfile, and Compose files.
  • Registration of OCI image references by sha256 digest without pulling images or connecting to Docker.
  • Atomic, read-only Target Snapshots with file-level SHA-256 manifests and idempotent reuse.
  • Idempotent TargetIngested events; failures and timeouts do not leave partial snapshots.

M2: Python Web Source Mapper

  • Python AST indexing without importing or executing target code.
  • Route discovery for Flask, FastAPI, Starlette, and Django.
  • Structured functions, cross-file calls, authentication and authorization guards, ownership checks, dangerous sinks, and input-propagation paths.
  • Content-addressed SourceGraph and StaticSignal output. A signal is a hypothesis for validation, never a Finding.
  • File size and SHA-256 verification before analysis; tampering, path escape, resource-limit violations, and timeouts fail closed.
  • Optional Semgrep adapter restricted to pre-registered local rules, with metrics and version checks disabled and no inherited API keys.
  • Idempotent SourceGraphBuilt summary events. Full graphs are stored separately as read-only objects instead of being copied into the normal event log.
  • Scope identity, version, and validity are rechecked before every source-mapping run.
  • CI runs lint, schema-drift checks, and the full test suite on Python 3.12, 3.13, and 3.14.

M3: deterministic Candidate generation

  • Converts integrity-checked SourceGraph objects into typed, human-reviewable Candidates.
  • Merges complementary signals for the same route and sink without treating analyzer output as a Finding.
  • Maps supported sink classes to CWE, preconditions, a security invariant, and the cheapest disproof task.
  • Binds every Candidate to the exact Target version, SourceGraph digest, Scope identity, and Scope version.
  • Uses stable Candidate UUIDs and SHA-256 duplicate fingerprints; repeated generation is byte-stable.
  • Excludes parse failures, visibly guarded object lookups, and external matches that cannot be classified safely.
  • Stores each CandidateSet as an immutable content-addressed object and records only a redacted summary event.

M4.1: sandbox contracts and offline runner

  • Defines immutable Static, Validation, and Report sandbox profiles with non-root identities, fixed mount slots, network modes, and resource ceilings.
  • Rejects writable roots, Linux capabilities, host-path mounts, unregistered writable paths, and profile-purpose mismatches at schema validation time.
  • Binds every Worker task to an exact Target version, Scope identity, policy digest, and sandbox-profile digest.
  • Accepts only typed registered-tool invocations with argument arrays and logical working directories—never arbitrary shell command strings.
  • Provides a SandboxRunner adapter contract and a deterministic offline implementation for success, refusal, timeout, cancellation, resource exhaustion, checkpoint/resume, idempotency, and cleanup paths.

M4.2: Tool Broker and typed HTTP

  • Adds an immutable capability registry whose digest is bound into every queued Worker task.

  • Requires a tool to be present in the Registry, Task allowlist, and Sandbox Profile while also matching the current Scope, policy digest, and Worker role.

  • Accepts HTTP methods, normalized credential-free URLs, safe headers, opaque credential/body references, and explicit time/size/redirect budgets—never raw credentials or request bodies.

  • Derives state-changing behavior from the HTTP method in trusted code and requires exact, unexpired approvals for mutations and credential use.

  • Reauthorizes every redirect, resolves every hop, pins the selected IP, verifies the reported peer, and rejects loopback, link-local metadata, multicast, unspecified, mixed-dangerous, and configured host-gateway addresses.

  • Returns only policy records, URL digests, peer metadata, Evidence IDs, final response-body SHA-256 digests, and budget usage; raw response bodies and sensitive headers stay outside the normal result path.

  • Uses deterministic offline resolver and transport adapters. No real HTTP request or network isolation claim is introduced in M4.2.

M4.3: ephemeral Docker Runner

  • Adds a trusted Docker CLI adapter that uses argument arrays only; Workers never receive the Docker socket or the host process environment.
  • Resolves content mounts through a Control Plane-owned registry, pins exact image IDs, disables pulls, and replaces image entrypoints with registered absolute tool executables.
  • Applies and re-inspects a read-only root, non-root UID/GID, dropped capabilities, no-new-privileges, a content-bound versioned seccomp contract, no network, resource limits, read-only content, and bounded noexec,nosuid,nodev tmpfs mounts.
  • Kills timed-out Workers, removes containers and anonymous storage, and refuses to report a normal result unless absence is verified.
  • Includes opt-in real-container probes for isolation, secret non-inheritance, timeout, and cleanup.
  • Adds a live Broker-owned HTTP/HTTPS transport that connects directly to the policy-selected IP, preserves the authorized hostname for HTTP Host and TLS verification, ignores proxy environment variables, verifies the actual peer, and enforces response limits.
  • Binds the selected resolver and transport implementation digest into the Tool Registry; queued work is rejected if offline and live adapters are swapped.
  • Resolves request bodies by content digest and credentials through separate opaque providers; neither raw material is copied into Broker results or HTTP Evidence metadata.
  • Stores only redacted response transcripts in the Evidence Store. Sensitive response headers, raw URLs, credential material, binary bodies, email addresses, and JSON-shaped secrets are excluded or redacted.
  • Requires rootless mode, seccomp, cgroup v2, and enforceable memory, CPU-quota, and PID controls by default. Engines that only advertise partial isolation fail closed before container creation.
  • Discovers daemon-managed network gateways for the Broker denylist. Docker Workers reject direct target_only networking and remain network-disabled; authorized target access belongs to the trusted Broker.
  • A dedicated Ubuntu 24.04 admission workflow runs Docker Engine 29.7.2 as a delegated rootless user service and proves Worker isolation from a live sibling container and daemon gateway, host-gateway denial before Broker transport, redirect-time DNS rebinding rejection, deterministic validation, timeout handling, and cleanup.

M4.4: transactional Validation Orchestrator

  • Seals each human-selected Candidate content digest, exact Target/Scope provenance, networkless Runner request, and bounded Broker calls into a content-addressed ValidationPlan.
  • Rechecks Candidate state, current Scope validity, policy digest, Validator role, profile digest, and Candidate input binding before writing a STARTED checkpoint.
  • Runs the sandbox step before Broker calls, stops on the first non-completed result, and maps denials, missing approvals, timeouts, and failures to fail-closed domain outcomes.
  • Persists one idempotent ValidationOutcome in SQLite. Completed plans replay without re-execution; an interrupted STARTED plan requires explicit recovery and is never retried automatically.
  • Separates execution from verdict. The production default remains INCONCLUSIVE; only a trusted deterministic judge may return REPRODUCED, and it may cite only Evidence IDs collected by that run.
  • Produces an EvidenceBundle and advances Candidate only through the existing state machine. It cannot create or promote a Finding.
  • Adds validation-run-offline, which exercises the control-plane path without target execution, Broker calls, sockets, or a reproduced claim.

M4.5: deterministic HTTP assertions

  • Adds a content-addressed HttpResponseAssertion selected before execution and bound to one exact Broker call.
  • Requires both an exact HTTP status and SHA-256 of the raw final response body. Status-only checks cannot produce a reproduced verdict.
  • Keeps raw response bodies out of Broker results; only the body digest, bounded metadata, and redacted Evidence reference cross the trusted boundary.
  • Adds DeterministicHttpJudge: by default it trusts only the live pinned HTTP Registry; an exact match returns the precommitted REPRODUCED or NOT_REPRODUCED result, while an offline Registry or any mismatch remains INCONCLUSIVE.
  • Verifies every Evidence object through a no-follow, size-bounded, content-integrity read before judging or sealing an EvidenceBundle.
  • Adds an opt-in composition probe covering a real ephemeral Docker Validator, a Broker-owned pinned HTTP connection to a temporary authorized fixture, Evidence capture, exact verdict, state transition, and cleanup.
  • The composition probe also passes in the dedicated rootless Linux admission workflow. Local Docker Desktop runs retain an explicit rootful test-only exception and cannot independently qualify production.

M5.1: deterministic Critic and independent disproof review

  • Seals Candidate, reproduced ValidationRun, EvidenceBundle, Scope version, validation context, and a distinct review context into a content-addressed CriticPlan.
  • Requires separate validation and review producers and assesses security controls, reachability, environment parity, and version binding exactly once.
  • Uses a fixed reducer: confirmed counterevidence rejects; any inconclusive angle leaves the Candidate validated but unpromotable; only four evidence-backed ruled-out angles advance it to CRITIC_REVIEWED.
  • Rechecks every referenced Evidence object with no-follow, size, digest, and Target-version validation before changing state.
  • Persists STARTED/COMPLETED SQLite checkpoints, returns completed outcomes idempotently, and refuses unfinished automatic replay.
  • Performs no target execution, Broker call, network access, report submission, or Finding promotion. The final promotion gate separately rechecks current Scope, reproduced-run Evidence coverage, Critic binding, and duplicate review.

M5.2: Evidence-consistent offline report drafts

  • Seals the Finding, promoted Candidate, approved EvidenceBundle, Scope version, channel, bounded narrative, and exact section citations into a content-addressed ReportDraftPlan.
  • Requires code-location, request/response, reproduction, and impact claims to cite Evidence IDs from the Finding's bundle; every bundled Evidence object is rechecked for no-follow access, size, digest, and Target version.
  • Redacts report text before persistence, escapes active HTML and Markdown image/link syntax, and never copies Evidence bodies into the draft.
  • Renders deterministic generic, EduSRC, CNVD, vendor, and CVE-draft headings to immutable local Markdown and JSON artifacts.
  • Uses STARTED/COMPLETED SQLite checkpoints, content-addressed artifact directories, bounded writes, idempotent completed replay, fail-closed recovery, and temporary-output cleanup.
  • Produces only draft review status. It has no network adapter, platform credential, approval mutation, or submission path.

M5.3: human review, revision diff, and approved local export

  • Groups report revisions into a stable Finding/channel family and binds every version after the first to the exact preceding Report digest.
  • Produces deterministic structured diffs for consecutive redacted revisions, including text and Evidence-reference changes; unchanged, unrelated, skipped, or unredacted revisions are rejected.
  • Seals the exact Report digest, artifact digest, EvidenceBundle, Scope version, reviewer, diff, decision deadline, and approval expiry into typed review protocol objects.
  • Applies only explicit approve, request_changes, or reject commands through the Control Plane state machine. Any content or citation change invalidates the sealed request.
  • Allows local export only from human_approved, before approval expiry, with an exact ReviewRecord and artifact match. Local export writes a new immutable Markdown/JSON artifact with exported status.
  • Adds offline report-review-diff, report-review-offline, and report-export-local CLI paths. None has a network or Submission adapter.

M6.1: deterministic offline benchmark and regression gate

  • Seals local benchmark cases, ground-truth Findings, pipeline observations, policies, and baselines as typed content-addressed objects.
  • Rejects any observed Finding that did not pass reproduced Validation, accepted independent Critic review, Candidate promotion, and complete Evidence gates.
  • Computes Candidate recall, Finding precision, duplicate rate, Evidence completeness, policy violations, elapsed time, total cost, and cost per Finding with deterministic reducers.
  • Applies absolute thresholds and exact-suite baseline comparisons, emitting stable violation codes and a failing CLI exit status for CI.
  • Persists STARTED/COMPLETED SQLite checkpoints and immutable local JSON/Markdown results with bounded no-follow reads and temporary-output cleanup.
  • Includes a generated local microbenchmark and baseline in benchmarks/m6_1; ordinary CI verifies fixture drift and runs the offline regression gate.
  • Adds benchmark-evaluate-offline. It consumes only sealed local files and has no Runner, Broker, network, credential, or Submission dependency.

M6.2: external benchmark local-snapshot adapters

  • Adds versioned BountyBench and AutoPenBench adapters for pre-obtained local directory snapshots; neither adapter accepts a URL or downloads data.
  • Seals every regular file by normalized path, size, and SHA-256, rejects symlinks and special files, and enforces file-count, per-file, total-size, and deadline budgets before and after normalization.
  • Reads only bounty_metadata.json labels from BountyBench. Prompt, report, exploit, setup, patch, and verification contents are never copied into normalized suites.
  • Reads AutoPenBench data/games.json in trusted code but persists only safe identities; task text and flags are discarded. CWE labels must come from a sealed vulnloom-autopenbench-cwe.json sidecar.
  • Emits typed exclusions for unsupported or missing labels and rejects malformed JSON, duplicate keys, stale mappings, ambiguous identities, adapter drift, and snapshot mutation.
  • Stores normalized suites as immutable local objects with transactional import checkpoints and adds benchmark-snapshot-manifest-local and benchmark-import-offline.

M6.3a: unified precomputed analyzer observations

  • Normalizes local CodeQL SARIF 2.1.0, Trivy JSON, Checkov JSON, and Kubesec JSON through versioned adapters without running those tools.
  • Seals the Target/version, tool version, rules digest, result digest, optional CWE mapping, adapter digest, resource limits, deadline, and idempotency key.
  • Persists only rule/message digests, normalized CWEs, severity, and safe relative locations; it discards raw messages, secret matches, and Kubernetes object identities.
  • Keeps analyzer observations structurally separate from pipeline BenchmarkObservation: they cannot carry Candidate, Validation, Critic, or Finding state.
  • Adds analyzer-result-manifest-local and analyzer-observations-import-offline; neither command accepts a URL or executes a binary.

M6.3b: explicit cross-analyzer evaluation

  • Requires a sealed, explicit Observation-to-ground-truth alignment; matching CWE labels alone never count as a detection.
  • Revalidates case, Target version, ObservationSet digest, truth ownership, and CWE compatibility before creating a checkpoint.
  • Computes overall and per-analyzer truth recall, observation precision, duplicate rate, and exclusion rate for CodeQL, Trivy, Checkov, and Kubesec.
  • Supports required-analyzer and full case-matrix gates plus exact-suite baseline regressions; per-analyzer checks prevent aggregate results from hiding one tool's regression.
  • Includes benchmarks/m6_3, fixture-drift checks, and run_m6_3_regression_gate.py in ordinary CI.
  • Adds analyzer-evaluate-offline, which creates only local immutable JSON/Markdown results and cannot alter Candidate or Finding state.
  • Accepts directories only. ZIP/TAR acquisition and extraction remain outside this adapter; callers must use an independently hardened quarantine path before presenting a directory.

M6.4: sealed analyzer execution

  • Runs exact Checkov, Kubesec, Trivy, and CodeQL registrations through the hardened Docker Runner with inspected image IDs, --pull never, network=none, read-only inputs/root, non-root identity, bounded resources, and mandatory cleanup.
  • Captures analyzer output within a trusted byte limit and completes the outer checkpoint only after the existing M6.3a adapter creates redacted Observations.
  • Restricts Trivy to its sealed read-only DB v2 and vulnerability scanner; DB/check/Java/VEX updates, secret scanning, telemetry, and version checks are disabled.
  • Binds CodeQL 2.26.2 to a Target/version/Manifest and one sealed prebuilt DB/query snapshot. A narrow wrapper copies the DB into bounded tmpfs because CodeQL writes query results, while the original DB and query pack remain read-only and are reverified after cleanup.
  • Does not expose analyzer/package downloads, arbitrary commands, Target builds, Candidate/Finding promotion, or Submission. CodeQL database construction remains separately RUN_UNTRUSTED_BUILD Approval-gated.

M7.1a: offline typed Agent Runtime

  • Seals an exact offline replay implementation, provider/model identity, supported Worker roles, and output ceiling in a content-addressed registration.
  • Binds each run to an exact TaskEnvelope, context and decision-schema digests, step/token/wall budgets, deadlines, and an idempotency key.
  • Validates untrusted structured decisions and returns only terminal summary digests or a typed, argument-digest-only tool intent. It never executes the proposed tool.
  • Persists transactional STARTED/COMPLETED checkpoints without raw model output or raw tool arguments; interrupted calls require explicit recovery and are not replayed automatically.
  • Uses no model socket, SDK, endpoint, credential, Runner, Broker, Approval, domain transition, or Submission path.

M7.1b: credential lease and local fake provider

  • Replaces direct API-key string resolution with an initialization-time allowlisted, content-addressed reference to one explicit Control Plane environment variable.
  • Holds secret bytes only in a non-serializable lease that is zeroed on success, exception, and timeout-result paths.
  • Adds a registration-bound local fake adapter to test provider identity and credential lifecycle without a socket, URL, SDK, proxy, or inherited environment.
  • Keeps the credential value and unrelated environment entries out of Worker requests, outcomes, checkpoints, schemas, and error messages.

M7.2: sealed and bounded model context

  • Requires transient context sources to match the Task's ordered input references exactly.
  • Normalizes and redacts content in trusted code, rejects unsafe controls, and enforces fragment, total-byte, count, deadline, and wall-clock limits.
  • Stores only explicitly untrusted, redacted fragments in immutable content-addressed snapshots bound to the exact Task, Target, Scope, references, and redaction policy.
  • Revalidates no-follow, regular-file, read-only, size, schema, identity, and digest properties on every stored read; failed publication removes temporary files.
  • Binds only the context snapshot ID into Agent run plans and step requests. It performs no provider, network, tool, Approval, or domain-state action; the Runtime reloads the snapshot before creating its STARTED checkpoint.

M7.3: fixed provider-message envelopes

  • Maps every Worker role to one built-in, content-addressed system template; callers cannot supply system text or template versions.
  • Renders canonical strict JSON with typed control metadata separated from escaped, explicitly untrusted context fragments.
  • Binds the plan, Task/context/model/template/schema, Target/Scope digests, tool allowlist, budgets, step, and messages into one envelope ID sealed into the step request.
  • Rejects duplicate keys, template/system/control/trust drift, request mismatch, byte overages, and rendering timeouts before the first model call.
  • Passes messages transiently to offline adapters while checkpoints and adapter audit lists retain only digests. Prompt text remains non-authoritative; Runtime/Broker enforce permissions.

M7.4: provider transport admission protocol

  • Adds a content-addressed admission object for one exact provider hostname, TLS port, canonical path, credential reference, adapter digest, request/response limits, and timeout.
  • Fixes redirects and proxies off, DNS revalidation on, raw-response persistence off, one attempt, and network_enabled=false; schema validation rejects any relaxation.
  • Derives a digest-only transport request from the exact StepRequest and Message Envelope, while the serialized message body exists only in a zeroed transient buffer.
  • Exercises credential acquisition, bounded response capture, strict JSON parsing, identity checks, typed rejection/timeout outcomes, receipts, and cleanup through an in-memory admission fake.
  • Stores only request/response digests, counts, identities, stable status, and cleanup proof. It opens no DNS, socket, HTTP, SDK, proxy, tool, Approval, or Submission path.

M7.5: subprocess-pinned HTTPS provider transport

  • Adds a fixed subprocess_https_provider adapter with a content-bound implementation digest; no caller-supplied executable, argv, header, URL, proxy, retry, or SDK is accepted.
  • Separates production live_https admissions (exact port 443 and global-only DNS) from loopback_https_probe admissions (.test, loopback-only, and a sealed test CA).
  • Re-resolves the exact hostname for every request, rejects mixed or forbidden answers, pins one numeric IP, and verifies the connected peer, TLS 1.2+ version, SNI, and hostname.
  • Sends credential and message bytes over a fixed binary stdin frame to an isolated Python process launched with -I, / cwd, closed file descriptors, a tiny environment, no shell, bounded stdout, discarded stderr, resource limits, and forced process-group cleanup.
  • Allows one POST to one canonical path, forbids redirects and encoded responses, applies header/body and parent/child wall limits, and records only peer digests, TLS version, counters, and cleanup.
  • Enforces a sealed per-minute rate limit with no automatic retry. Live transport remains library-only and no public provider is contacted by CI; Phase 3 uses a real loopback TLS subprocess probe.

M7.6: issued provider-egress lifecycle

  • Adds content-addressed issuer policies that limit exact provider IDs, networked transport modes, and grant lifetimes; no-network fake modes cannot receive an egress grant.
  • Issues immutable grants bound to one transport Admission, credential reference, adapter, purpose, issuer, validity window, and idempotency key, then binds the exact grant ID into model registration.
  • Atomically publishes read-only grant/revocation objects and records STARTED/COMPLETED issuance and revocation checkpoints in a transactional lifecycle ledger.
  • Reopens and verifies the grant, ledger state, expiry, and exact Admission binding before every DNS lookup, rate slot, credential lease, or child process. Revoked, expired, unfinished, linked, writable, malformed, or drifted grants fail closed.
  • Keeps the authority local and library-only. It adds no remote signer, provider SDK/codec, public provider call, arbitrary URL, tool execution, Approval consumption, or Submission capability.

M7.7: sealed OpenAI Responses codec

  • Adds a content-addressed openai-responses-v1 codec bound into every subprocess HTTPS model registration and to the exact admitted /v1/responses path.
  • Emits only fixed non-streaming, non-stored requests with disabled truncation and the registered strict Agent decision JSON Schema; callers cannot add tools, metadata, sessions, or parameters.
  • Accepts only one completed assistant output_text with exact model identity and typed usage, then strictly parses it as AgentDecisionPayload. Incomplete, refusal, tool-call, duplicate-key, oversized, timed-out, and protocol-drift responses fail closed.
  • Keeps provider-native tool execution disabled. A structured tool proposal remains an inert intent that must pass the existing Runtime and Broker enforcement boundaries.
  • Uses offline golden fixtures and the existing opt-in loopback TLS process probe; CI does not call a public provider or use a real provider credential.

M7.8: typed Agent intent handoff to Tool Broker

  • Binds one authoritative completed Agent run and digest-only tool intent to an independently built, exact typed BrokerCall; no model text is translated into executable parameters.
  • Restricts handoff to Validator tasks and rechecks Task, Scope, Policy, Sandbox Profile, Tool Registry, tool budget, call commitment, deadline, and Agent checkpoint before dispatch.
  • Leaves all network, DNS pinning, credentials, side-effect, and Approval enforcement inside the existing Tool Broker. The Agent never receives a socket, Docker handle, secret, or adapter.
  • Adds transactional STARTED/COMPLETED handoff checkpoints. Only an approval-required first attempt can be retried once with a new exact Broker call and independently verified Approval.
  • Converts every completed Broker result into a digest-only AgentToolObservation containing typed counts and Evidence references, never raw Agent arguments, URLs, credentials, or response bodies.
  • Tests both offline state-machine paths and an opt-in live pinned-Broker composition against a temporary authorized service; no public target, Candidate/Finding transition, or Submission is added.

M7.9: sealed Tool Observation continuation

  • Derives one new Validator Task from an authoritative completed Agent → Broker chain while preserving the exact engagement, Target/version, Scope/version, Policy, Profile, Registry, model, and absolute deadline bindings.
  • Fixes the derived Task to an empty tool allowlist and tool_calls=0; the one-step continuation may finish as complete or blocked, while another tool proposal becomes a terminal failure.
  • Reopens each exact Evidence ref through the content-addressed Evidence Store, re-redacts it, and rebuilds the same bounded read-only context snapshot before any provider call. Observation and Evidence text remain explicitly untrusted model context.
  • Accounts for tokens, prior steps, the consumed Broker call, and remaining wall time in a typed budget ledger. Exhaustion, expiry, authority drift, missing Evidence, or incomplete cleanup fails before the continuation checkpoint.
  • Adds a unique-Observation STARTED/COMPLETED SQLite lifecycle with idempotent completed replay and fail-closed conflict/recovery. Its rows contain only typed outcomes and digests, never Evidence, provider response, URL, credential, Candidate/Finding, or Submission content.
  • Phase 3 composes the real isolated loopback TLS provider, real pinned Broker, redacted Evidence Store, and zero-tool continuation against a temporary authorized target; no public egress is used.

M7.10: sealed fixed two-tool Agent session

  • Adds a content-addressed AgentSessionPlan, cumulative token/step/tool/provider/Broker budget ledger, and SQLite lifecycle around one already completed tool round, one optional second tool round, and one mandatory zero-tool terminal continuation.
  • Derives the second Validator Task with the exact inherited Target/Scope/Policy/Profile/Registry, model, and absolute deadline, while shrinking all budgets and fixing the total session to at most three provider turns and two consumed tool calls.
  • Exposes only a trusted, content-addressed AgentAuthorizedCallSet in message control. Each option is a Control-Plane-built exact read-only BrokerCall commitment; the model cannot construct or alter URL, method, headers, body, credential, network, or authorization fields.
  • Reopens every Agent, handoff, Observation, Evidence, and context checkpoint before the next action. Unlisted or duplicate commitments, drift, exhausted budgets, missing cleanup, and a third tool proposal fail closed without another Broker call.
  • Pauses an Approval-required second handoff without polling or approving it. One explicit, Approval-bound M7.8 retry may resume the session; its extra Broker attempt is counted even though the total successful tool-call budget remains exactly two.
  • Phase 3 uses three isolated loopback provider subprocesses and two exact pinned-Broker reads of a temporary authorized target. No public target/provider, arbitrary loop, target build, domain-state transition, report export, or Submission path is introduced.

M7.11: immutable Agent session audit and deterministic projection

  • Reopens the authoritative Session, Agent run, handoff, continuation and Evidence stores before producing one content-addressed AgentSessionAuditBundle; callers cannot provide a transcript or substitute a model-generated summary.
  • Recomputes ordered round identities, exact call commitments, Approval decision digests, Target/Scope provenance, cumulative token/step/tool/provider/Broker budgets and cleanup proofs.
  • Projects only completed, blocked, failed or timed_out with a stable reason code and verified Evidence refs. The recommendation is not a Candidate/Finding/Report transition or an authorization.
  • Publishes bounded, read-only JSON/Markdown containing only digests, IDs, typed counts and statuses; it never copies Evidence bodies, URLs, credentials, provider requests/responses or tool parameters.
  • Uses a separate SQLite STARTED/COMPLETED checkpoint with idempotent completed replay and fail-closed conflict/recovery. Artifact failure cleans temporary files and refuses automatic replay.
  • Extends the loopback Phase 3 composition by creating the audit from the real M7.10 session and proving a tampered chain is rejected without adding runtime network or execution authority.

M8.1: human Validation Intake and sealed plan binding

  • Reopens the immutable M7.11 Audit artifact and CandidateSet, then binds their exact digests to a Control-Plane-built typed ValidationPlan; no Agent prose or Evidence body becomes a request.
  • Accepts only a human accept, reject, or defer command bound to the exact Audit, Candidate and Validation plan. A blocked, failed or timed-out recommendation cannot be accepted.
  • Produces only a digest-only immutable decision record. The service has no Runner or Broker and does not queue Validation, mutate Candidate state, consume Approval, build a Target or submit data.
  • Uses an independent STARTED/COMPLETED SQLite ledger; drift, expiry, duplicate consumption, conflicting decisions and unfinished recovery fail closed.

M8.2: completed Validation outcome provenance binding

  • Runs only after an explicit existing Validation entry point has completed; it cannot execute, resume, retry, queue, approve, or alter that Validation.
  • Reopens the accepted Intake, Audit artifact, CandidateSet, exact Validation checkpoint, current Scope and every referenced Evidence object before creating a binding checkpoint.
  • Recomputes Runner/Broker identities, ordered calls, forced timeout/policy results, run accounting, final Candidate state and EvidenceBundle consistency; drift fails closed.
  • Persists only IDs, digests, typed result/state and timestamps in a unique STARTED/COMPLETED ledger. Idempotent replay is read-only and conflicting Intake/plan/outcome consumption is rejected.
  • Does not call a Runner, Broker, provider, Docker, network, Approval, Target build, Critic, Finding, report export, public target, or Submission path.

M8.3: human Critic Intake and sealed plan binding

  • Reopens a reproduced M8.2 binding, Audit artifact, CandidateSet, completed Validation checkpoint and Evidence before binding one independently constructed exact CriticPlan.
  • Accepts only explicit human accept, reject or defer commands. Acceptance records selection but does not execute the Critic, mutate Candidate state, create a review/Finding, or export a report.
  • Persists only IDs, digests, typed decisions and timestamps; conflict, expiry, drift and unfinished recovery fail closed without Runner, Broker, provider, network, Approval or Submission access.

M8.4: completed Critic outcome provenance binding

  • Runs only after the accepted M8.3 exact CriticPlan has completed through the existing explicit deterministic Critic entry point; the binding service cannot invoke, retry or recover Critic.
  • Reopens the M8.2 binding, Validation checkpoint, Evidence, Critic plan/outcome and current Scope, validating review identity, independent context, ruleset, counterevidence, rationale and time.
  • Enforces the exact verdict/state mapping: accepted to CRITIC_REVIEWED, rejected to REJECTED, inconclusive to unchanged VALIDATED.
  • Persists only IDs, digests, typed verdict/state and timestamps in a unique STARTED/COMPLETED ledger. It does not mutate the original Candidate or create a Finding, report, build, Approval or Submission.

M8.5: human Finding promotion Intake

  • Requires an accepted M8.4 binding, exact critic-reviewed Candidate, reproduced Validation and Evidence provenance, accepted CriticReview, and the authoritative store's unique latest, content-addressed duplicate-clear proof.
  • Binds one trusted-control-plane-built exact FindingPromotionPlan; Agent prose, Critic rationale and Evidence content cannot construct root-cause, affected-version, impact or severity fields.
  • Records only an explicit human accept, reject or defer selection in a digest-only checkpoint. Acceptance does not call Candidate promotion, create a Finding, mutate state or draft a report.
  • Drift, expiry, duplicate results, non-accepted Critic verdicts, conflicting consumption and unfinished recovery fail closed without Runner, Broker, provider, network or Submission access.

M8.6: Approval-gated deterministic Finding promotion

  • Requires both an unexpired accepted M8.5 record and a granted MUTATE_TARGET_STATE Approval bound to the exact Intake record, PromotionPlan, Candidate, Finding ID, Scope and two declared effects.
  • Reopens the complete Validation, Evidence, Critic, duplicate-check and Intake provenance chain, then calls only the existing pure Candidate-to-Finding state transition.
  • Atomically records the promoted immutable Candidate and verified Finding in a unique STARTED/COMPLETED result ledger. Completed replay is idempotent; conflicts and unfinished recovery fail closed.
  • Agent output cannot supply Approval or promotion fields, and the service has no Runner, Broker, provider, build, public-network, report-generation or Submission capability.

M8.7: human Report Intake

  • Requires one completed, sealed M8.6 promotion outcome containing a promoted Candidate and verified Finding, plus the authoritative reproduced Validation and EvidenceBundle provenance.
  • Binds a trusted-control-plane-built exact ReportDraftPlan, including report family, channel and complete Evidence-cited sections. This first version accepts only version 1 drafts and persists only IDs, digests, typed decision and timestamps.
  • Accepts only explicit human accept, reject or defer. Acceptance does not draft a Report, read report prose into an Agent context, export an artifact, approve disclosure or create a Submission.
  • Plan drift, missing or altered promotion checkpoints, Evidence corruption, expiry, conflicts and unfinished recovery fail closed without Runner, Broker, provider, target or network calls.

M8.8: accepted Intake deterministic local draft binding

  • Requires an unexpired human-accepted M8.7 record and reopens the complete M8.6, Critic, Validation, Evidence and exact ReportDraftPlan provenance chain before drafting.
  • Seals the ordered typed Evidence catalog into a digest-only execution plan, then invokes only the existing deterministic offline report service with trusted control-plane inputs.
  • Produces one immutable local report artifact that remains DRAFT, plus a prose-free outcome binding. It cannot approve, export or submit the report, mutate Candidate/Finding state, build a target, or access Runner, Broker, provider or network capabilities.
  • Drift, expiry, non-accepted Intake, pre-existing unbound drafts, conflicts and unfinished checkpoints fail closed. Completed execution is read-only and idempotent.

M8.9: human Report review Intake

  • Reopens one completed M8.8 binding, its exact DRAFT Report and immutable artifact, plus the same ordered typed Evidence catalog used for drafting.
  • Binds a trusted-control-plane-built exact ReportReviewPlan; human input is limited to accepting, rejecting or deferring entry into the later review operation.
  • Persists only IDs, digests, typed selection and timestamps. Acceptance does not call the review service or change the Report from DRAFT, and grants no report approval or export authority.
  • Artifact/Evidence corruption, plan drift, expiry, conflicts and unfinished recovery fail closed without Runner, Broker, provider, target, network, export or Submission access.

M8.10: Approval-gated deterministic Report review

  • Requires an unexpired accepted M8.9 record, an independently issued exact human ReportReviewCommand, and a granted REVIEW_REPORT Approval bound to that command and its single expected state effect.
  • Reopens the M8.8 DRAFT, artifact, Evidence catalog and M8.9 checkpoint before invoking only the existing deterministic human-review state machine.
  • Records the resulting review in the existing authoritative store and publishes a prose-free outcome binding. The source DRAFT remains immutable; completed replay is read-only and idempotent.
  • Missing/revoked Approval, decision or provenance drift, expiry, pre-existing unbound reviews and unfinished recovery fail closed. No report export, public network or Submission capability is present.

M8.11: human local Report export Intake

  • Reopens one completed M8.10 HUMAN_APPROVED binding, its authoritative review outcome and immutable artifact, then binds them to an independently constructed exact ReportExportPlan.
  • Human input is limited to accepting, rejecting or deferring that exact local export plan. Even an accepted record does not call LocalReportExportService or change the Report to EXPORTED.
  • Persists only IDs, digests, typed selection and timestamps; report prose and review rationale stay outside the Agent Intake ledger.
  • Non-approved reports, artifact or plan drift, expiry, conflicts and unfinished recovery fail closed. Runner, Broker, provider, target, public network, arbitrary destination and Submission capabilities are absent.

M8.12: Approval-gated deterministic local Report export

  • Requires a still-valid accepted M8.11 record and an independently granted EXPORT_REPORT Approval bound to the exact export plan, human-approved Report, review, artifact, Scope and fixed local state/artifact effects.
  • Reopens every authoritative M8.10/M8.11 checkpoint before invoking only the existing LocalReportExportService; neither Agent output nor Approval prose can supply export parameters.
  • Produces an immutable local EXPORTED Report artifact and a prose-free outcome binding. The source HUMAN_APPROVED Report and artifact remain immutable, and completed replay is read-only.
  • Missing/revoked Approval, drift, expiry, pre-existing unbound exports, conflicts and unfinished recovery fail closed. No destination path/URL, public network, platform token or Submission capability exists.

M9.1: closed Agent workflow quality and safety regression

  • Evaluates one content-addressed observation over the fixed 13-stage M7.11–M8.12 chain, including six explicit human Intake decisions and three exact Approvals.
  • Requires reproduced Validation, accepted Critic, immutable source Candidate/Report states and one final local EXPORTED artifact with bound Evidence.
  • Compares effect counters captured after Validation with the final export boundary. Any later provider, Broker, Runner or target call fails the gate, as does any public-network access, target build, automatic Approval or Submission.
  • The evaluator is pure and offline; STARTED/COMPLETED checkpoints and content-addressed JSON/ Markdown results are bounded, no-follow verified and replayed read-only. The Phase 3 Admission harness constructs the observation from the real local composition rather than Agent prose.

M9.2: sealed negative-path mutation corpus

  • Seals one known-good M9.1 observation plus 23 fixed scenarios covering every current quality and safety violation branch; the corpus requires every mutation exactly once and in canonical order.
  • Each scenario has an immutable expected PASS/FAIL status and exact violation-code tuple. Missing cases or changed expectations fail schema validation instead of silently updating the baseline.
  • The offline corpus evaluator deterministically mutates only typed observations and requires all 23 actual results to match the sealed baseline. It has no Runner, Broker, provider, target or Approval dependency.
  • Fixture generation and the M9.2 gate run in every Python 3.12/3.13/3.14 CI job. Public-network, target-build, automatic-Approval and Submission regressions are explicit failing scenarios.

Local development

VulnLoom requires Python 3.12 or later.

python3 -m venv .venv
.venv/bin/pip install -e '.[dev]'
.venv/bin/pytest --cov=vulnloom --cov-report=term-missing
.venv/bin/python scripts/export_schemas.py
.venv/bin/python scripts/export_benchmark_fixtures.py
.venv/bin/python scripts/run_m6_1_regression_gate.py
.venv/bin/python scripts/export_analyzer_evaluation_fixtures.py
.venv/bin/python scripts/run_m6_3_regression_gate.py
.venv/bin/python scripts/export_agent_workflow_regression_fixtures.py
.venv/bin/python scripts/run_m9_2_agent_workflow_regression_gate.py

# Optional; requires a local Alpine 3.22 image and Docker engine.
VULNLOOM_DOCKER_INTEGRATION=1 .venv/bin/pytest tests/test_runner_docker_integration.py

# Optional; opens one temporary loopback-only HTTP server.
VULNLOOM_SOCKET_INTEGRATION=1 .venv/bin/pytest tests/test_live_http_integration.py

# Optional; combines one real Docker Worker and one temporary authorized HTTP fixture.
VULNLOOM_COMPOSITION_INTEGRATION=1 \
  .venv/bin/pytest tests/test_validation_composition_integration.py

# Explicitly opt in to the isolated Hybrid source-to-preprod acceptance.
VULNLOOM_HYBRID_E2E_INTEGRATION=1 \
  .venv/bin/pytest tests/test_hybrid_e2e_integration.py

Basic usage

Target ingestion requires an approved Scope that is still valid and explicitly includes either the artifact digest or the Git URL and commit.

vulnloom --db .vulnloom/events.db --store .vulnloom/targets \
  artifact-quarantine --engagement-id <uuid> --source source.zip

# Add the returned source_name, kind, and artifact_id to a Scope,
# then have that Scope approved by an authorized reviewer.
vulnloom --db .vulnloom/events.db --store .vulnloom/targets \
  target-ingest-archive --scope-file scope.json --source source.zip

vulnloom --db .vulnloom/events.db --store .vulnloom/targets \
  target-ingest-git --scope-file scope.json --source /local/repo \
  --repository-url https://example.invalid/project.git --commit <full-commit>

vulnloom --db .vulnloom/events.db --store .vulnloom/targets \
  target-register-image --scope-file scope.json \
  --image-ref ghcr.io/example/app --digest sha256:<digest>

# Build an offline Python Web SourceGraph from a verified filesystem snapshot.
vulnloom --db .vulnloom/events.db --store .vulnloom/targets \
  source-map --snapshot-id <manifest-sha256> --scope-file scope.json \
  --analysis-store .vulnloom/analysis

# Generate deterministic validation hypotheses. This does not execute or validate them.
vulnloom --db .vulnloom/events.db \
  candidate-generate --graph-id <graph-sha256> --scope-file scope.json \
  --analysis-store .vulnloom/analysis --candidate-store .vulnloom/candidates

# Rebuild trusted static results for one already-ingested authorized Snapshot and
# emit a local human-review summary. This does not select or validate a Candidate.
vulnloom --store .vulnloom/targets shadow-pilot-local \
  --snapshot-id <manifest-sha256> --scope-file scope.json \
  --analysis-store .vulnloom/analysis \
  --candidate-store .vulnloom/candidates \
  --readiness-db .vulnloom/pilot-readiness.db \
  --readiness-store .vulnloom/pilot-readiness

# Record one explicit human Candidate choice from a completed passing shadow pilot.
# The timestamp is part of the immutable command; no ValidationPlan is created.
vulnloom --store .vulnloom/targets pilot-candidate-select-local \
  --readiness-plan-id <readiness-plan-sha256> \
  --candidate-set-id <candidate-set-sha256> --candidate-id <candidate-uuid> \
  --scope-file scope.json --reviewer <human-identity> \
  --decided-at <ISO-8601-with-timezone>

# Bind that completed selection to pre-existing, independently constructed M8.1
# plan/command files. This records Intake only and never runs Validation.
vulnloom pilot-validation-intake-bind-local \
  --selection-readiness-plan-id <readiness-plan-sha256> \
  --scope-file scope.json --audit-artifact-file audit-artifact.json \
  --intake-plan-file intake-plan.json --intake-command-file intake-command.json \
  --validation-plan-file validation-plan.json

# Bind one accepted model Recommendation selection to a prebuilt offline
# ValidationPlan. This writes Intake only and still requires RUN_VALIDATION Approval.
vulnloom candidate-recommendation-validation-intake-local \
  --scope-file scope.json --generation-db generations.db \
  --recommendation-db recommendations.db --selection-db selections.db \
  --graph-store .vulnloom/graphs --candidate-store .vulnloom/candidates \
  --validation-intake-db validation-intakes.db \
  --selection-record-id <selection-record-sha256> \
  --validation-plan-file validation-plan.json \
  --idempotency-key recommendation-intake-001

# Execute that exact binding only after a separate human RUN_VALIDATION Approval.
# This first pilot execution path is offline and refuses every Broker call.
vulnloom pilot-validation-run-offline \
  --pilot-intake-plan-id <pilot-intake-plan-sha256> \
  --intake-plan-id <m8.1-intake-plan-sha256> \
  --scope-file scope.json --validation-plan-file validation-plan.json \
  --approval-file run-validation-approval.json

# Exercise a pre-sealed ValidationPlan through the offline control-plane path.
# The plan must not contain Broker calls.
vulnloom --db .vulnloom/events.db \
  validation-run-offline --scope-file scope.json \
  --candidate-store .vulnloom/candidates \
  --candidate-set-id <candidate-set-sha256> --candidate-id <candidate-uuid> \
  --plan-file validation-plan.json --validation-db .vulnloom/validation.db \
  --evidence-store .vulnloom/evidence

# Compare two consecutive, already-redacted Report revisions without network access.
vulnloom report-review-diff --before report-v1.json --after report-v2.json

# Apply a pre-sealed human decision to an immutable local Report artifact.
vulnloom --db .vulnloom/events.db report-review-offline \
  --scope-file scope.json --artifact-file report-artifact.json \
  --evidence-bundle-file evidence-bundle.json --evidence-catalog-file evidence.json \
  --review-plan-file review-plan.json --review-command-file review-command.json

# Mark an exactly approved Report as locally exported. This never sends it anywhere.
vulnloom --db .vulnloom/events.db report-export-local \
  --scope-file scope.json --artifact-file approved-artifact.json \
  --review-record-file review-record.json --export-plan-file export-plan.json

# Evaluate a sealed local benchmark. Exit status 2 means the regression gate failed.
vulnloom benchmark-evaluate-offline \
  --suite-file suite.json --observations-file observations.json \
  --plan-file benchmark-plan.json --benchmark-db .vulnloom/benchmarks.db \
  --result-store .vulnloom/benchmark-results

# Seal an already-present local BountyBench directory without executing its files.
vulnloom benchmark-snapshot-manifest-local \
  --source /local/bountytasks --kind bountybench \
  --upstream-revision <full-commit> --license-spdx Apache-2.0

# Normalize a sealed local external snapshot. The plan binds the manifest and adapter digest.
vulnloom benchmark-import-offline \
  --source /local/bountytasks --snapshot-file snapshot.json \
  --plan-file import-plan.json --import-db .vulnloom/benchmark-imports.db \
  --suite-store .vulnloom/benchmark-suites

# Seal a precomputed local analyzer output and optional explicit CWE map.
vulnloom analyzer-result-manifest-local \
  --output codeql.sarif --analyzer codeql --target-id <uuid> \
  --target-version <exact-version> --tool-version <version> \
  --rules-digest <sha256> --cwe-map analyzer-cwe.json

# Normalize the sealed file without running the analyzer.
vulnloom analyzer-observations-import-offline \
  --output codeql.sarif --cwe-map analyzer-cwe.json \
  --snapshot-file analyzer-snapshot.json --plan-file analyzer-plan.json

# Evaluate sealed analyzer observations against an explicit reviewed alignment.
vulnloom analyzer-evaluate-offline \
  --suite-file suite.json --alignment-file alignment.json \
  --observation-set-file codeql-observations.json \
  --observation-set-file trivy-observations.json \
  --plan-file analyzer-evaluation-plan.json

# Validate a sealed source-only analyzer execution plan without running the tool.
vulnloom --store .vulnloom/targets analyzer-execution-check-offline \
  --scope-file scope.json --snapshot-id <manifest-sha256> \
  --registration-file analyzer-registration.json \
  --plan-file analyzer-execution-plan.json

shadow-pilot-local accepts only an already-ingested local Snapshot and its exact approved Scope. Trusted code rebuilds the SourceGraph and CandidateSet, checks the repository-owned admitted M9.4 baseline, stores immutable static/readiness artifacts, and prints Candidates for human review. An empty CandidateSet is a valid result, not evidence that the target has no vulnerabilities. The command has no Candidate-selection, Validation, Runner, Broker, provider, build, network, Approval, export, or Submission operation.

pilot-candidate-select-local reopens a completed passing readiness checkpoint and its immutable artifact, then verifies the exact SourceGraph, CandidateSet, Snapshot and Scope before recording one human-selected proposed Candidate. The digest-only record is suitable for a later M8.1 binding, but does not construct a ValidationPlan or authorize or execute Validation. One readiness result may select at most one Candidate.

pilot-validation-intake-bind-local requires that completed selection before it will apply an exact accepted M8.1 human command. All Audit, Candidate and ValidationPlan files must already exist and are reverified through the existing M8.1 service; the pilot bridge only adds an immutable provenance binding. It never derives Runner/Broker arguments, runs Validation, changes the Candidate, or emits a domain event.

pilot-validation-run-offline accepts only a completed M9.8 binding, its accepted M8.1 record, the identical pre-existing ValidationPlan and a separately granted exact RUN_VALIDATION Approval. The first version refuses every Broker call and all network activity, uses the offline Runner adapter, and records a digest-only execution binding. It never constructs a ValidationPlan or Approval from Agent output and never writes a Submission.

validation-run-offline accepts an already sealed, typed ValidationPlan. It rejects plans containing Broker calls, does not execute target code or open sockets, and defaults to an INCONCLUSIVE verdict. Live Broker/Docker composition currently exists as a library and opt-in integration-test path, not as a production CLI or HTTP API.

The report review commands accept only sealed JSON contracts and content-addressed local objects. report-export-local changes the Report to exported only inside the local store; it has no destination URL, disclosure adapter, or submitted transition. Benchmark evaluation consumes precomputed typed pipeline observations. External benchmark and analyzer imports only normalize pre-obtained local data. analyzer-execution-check-offline validates the sealed M6.4a control-plane contract but deliberately produces no analyzer output. M6.4b-d add library-only, exact-image, network-disabled Checkov/Kubesec/Trivy/CodeQL execution, and every output must pass the same M6.3a import boundary. M6.5 requires a complete, content-bound execution matrix before invoking M6.3b; M6.6 proves that four-analyzer composition in rootless Admission, including missing-cell and drift refusal. No CLI path installs or runs analyzers.

Model credential boundary

ModelProviderConfig.credential_reference identifies one explicit Control Plane environment variable without containing its value. M7.1b-M7.6 resolve it into a short-lived, zeroed Control Plane lease. M7.5 passes its bytes only to a one-shot isolated HTTPS child over a bounded stdin frame; M7.6 requires an active, exact, operator-issued egress grant before the credential is read. The Worker never receives the reference or value. Keys never appear in TaskEnvelope, Worker environments, Agent requests, checkpoints, event logs, or Evidence summaries. .env.example lists variable names only, and real .env files are ignored by Git.

Safety status

VulnLoom is under active development. The current release provides the trusted domain foundation, secure local target ingestion, offline static source mapping, deterministic Candidate generation, a hardened Docker adapter, live pinned Broker transport, transactional validation orchestration, deterministic HTTP assertions, redacted Evidence storage, offline benchmark gates, precomputed multi-analyzer normalization, sealed Checkov/Kubesec/Trivy/CodeQL execution, execution-to-evaluation qualification, and a typed Agent Runtime with scoped credential leases, sealed context/messages, subprocess-pinned HTTPS transport, operator-issued egress lifecycle enforcement, exact Agent-to-Broker handoff, sealed Observation continuation, and a fixed two-tool Session ledger, plus opt-in probes for real containers, analyzers, sockets, and full validation composition.

Live Docker/Broker validation, the report workflow, and provider transport are exposed through typed library paths, not a production HTTP API or general provider CLI. The rootless Linux and OS-level egress admission gate passes; transport is qualified against a local TLS fixture and bounded CUC probes, while a general research adapter remains unavailable. External disclosure/CVE submission workflows, provider-specific general adapters, and dedicated Kubernetes, Terraform, or Helm vulnerability analyzers are not implemented yet.

Use VulnLoom only on systems, source code, and test environments for which you have explicit authorization.

The M9.10 pilot outcome gate consumes a completed M9.9 execution binding before recording M8.2 outcome provenance. Supply presealed plans created by trusted control-plane code:

vulnloom pilot-validation-outcome-bind-local \
  --plan-file pilot-outcome-plan.json \
  --outcome-plan-file agent-validation-outcome-plan.json \
  --validation-plan-file validation-plan.json \
  --audit-artifact-file audit-artifact.json \
  --scope-file scope.json

The command uses local stores (override the --*-db and --*-store flags as needed). Identical sealed plans replay read-only. Missing or unfinished execution, provenance drift, expired windows, and a pre-existing bare M8.2 checkpoint fail closed. It does not execute Validation, mutate a Candidate, grant Approval, build a Target, access the network or create a Submission.

M9.11 adds pilot-bound human Critic Intake. An operator supplies a sealed M9.11 plan, completed M9.10 provenance, and an independently prepared M8.3 human accept command and CriticPlan:

vulnloom pilot-critic-intake-bind-local \
  --plan-file pilot-critic-intake-plan.json \
  --pilot-outcome-plan-file pilot-outcome-plan.json \
  --outcome-plan-file agent-validation-outcome-plan.json \
  --intake-plan-file agent-critic-intake-plan.json \
  --intake-command-file agent-critic-intake-command.json \
  --critic-plan-file critic-plan.json \
  --validation-plan-file validation-plan.json \
  --audit-artifact-file audit-artifact.json \
  --scope-file scope.json

Use the --*-db and --*-store flags for the authoritative local stores. Only reproduced Validation with a validated result Candidate and complete Evidence is eligible; the default offline pilot's inconclusive result remains rejected. The command records Intake without running Critic or Validation, changing Candidate state, granting Approval, building a Target, accessing the network or submitting anything. Completed replay is read-only; missing/STARTED provenance and bare pre-existing M8.3 checkpoints fail closed. A subsequent pilot Critic execution stage is still required.

M9.12 connects approved pilot Critic execution and M8.4 result binding. Prepare a sealed execution plan through trusted control-plane code and obtain an independent exact human RUN_CRITIC Approval:

vulnloom pilot-critic-run-local \
  --plan-file pilot-critic-execution-plan.json \
  --approval-file critic-approval.json \
  --pilot-intake-plan-file pilot-critic-intake-plan.json \
  --pilot-outcome-plan-file pilot-outcome-plan.json \
  --outcome-plan-file agent-validation-outcome-plan.json \
  --intake-plan-file agent-critic-intake-plan.json \
  --intake-command-file agent-critic-intake-command.json \
  --critic-plan-file critic-plan.json \
  --evidence-catalog-file evidence-catalog.json \
  --validation-plan-file validation-plan.json \
  --audit-artifact-file audit-artifact.json \
  --scope-file scope.json

Override the --*-db/--*-store flags to use the same authoritative stores as previous stages. The catalog is a bounded JSON array of typed Evidence metadata in Bundle-reference order, without Evidence body text. The command performs a local deterministic review of sealed assessments and records M8.4 and pilot bindings; it does not run Validation or target code, create Approval, build, access the network or submit. Replays do not review again. Incomplete checkpoints require explicit recovery. Accepted Critic outcomes still require separate duplicate-check and Finding-promotion gates.

M9.13 adds pilot-finding-intake-bind-local for human Finding Intake provenance. Supply the same upstream files and authoritative stores as M9.12, replacing --plan-file with the sealed PilotFindingIntakePlan, and add:

--critic-execution-plan-file pilot-critic-execution-plan.json
--critic-binding-plan-file agent-critic-outcome-plan.json
--finding-intake-plan-file agent-finding-intake-plan.json
--finding-intake-command-file agent-finding-intake-command.json
--promotion-plan-file finding-promotion-plan.json
--duplicate-check-file finding-duplicate-check.json
--duplicate-check-db .vulnloom/finding-duplicate-checks.db
--finding-intake-db .vulnloom/agent-finding-intakes.db
--pilot-finding-db .vulnloom/pilot-finding-intakes.db

Prepare sealed plans through trusted control-plane services. --approval-file carries the historical M9.12 RUN_CRITIC Approval, not a promotion grant. The command requires accepted Critic provenance, a current latest CLEAR duplicate check, an exact PromotionPlan and an independent human ACCEPT command. It records M8.5 Intake and a digest-only pilot binding. It does not create a Finding, change Candidate state, run Critic/Validation/target code, approve an operation, build, access the network or submit. Completed replay is read-only; incomplete checkpoints require explicit recovery.

M9.14 adds pilot-finding-promote-local for an explicitly approved local Finding state transition. Supply the same upstream files and stores as M9.13, replace --plan-file with the sealed PilotFindingPromotionPlan, and add:

--pilot-finding-intake-plan-file pilot-finding-intake-plan.json
--promotion-execution-plan-file finding-promotion-execution-plan.json
--promotion-approval-file finding-promotion-approval.json
--promotion-db .vulnloom/finding-promotions.db
--pilot-promotion-db .vulnloom/pilot-finding-promotions.db

Trusted control-plane code prepares the M8.6 execution plan and pilot plan after completed M9.13 Intake and a separate exact human promotion Approval. The historical Critic --approval-file grants no promotion authority. Execution rechecks the current Scope, latest CLEAR duplicate proof and the entire accepted provenance chain. The result contains a verified Finding and promoted Candidate in the M8.6 outcome store plus a digest-only pilot binding. The source CandidateSet remains unchanged. Completed replay only verifies persisted results; incomplete checkpoints require explicit recovery. The command does not run Validation, Critic or target code, approve operations, derive Runner/Broker parameters, build, access the network or submit anything.

M9.15 provides a standalone fixed-message Provider probe. provider-probe-prepare seals a bounded plan without reading credentials or opening a connection; provider-probe-run requires an existing operator-issued egress grant and explicit --allow-provider-network. It uses only synthetic content and no tools or research targets. Each grant is consumed once in the authoritative probe ledger, including failed attempts; replay is read-only and STARTED requires explicit recovery.

See Provider probe setup and remaining admission requirements for the configuration types and commands. A real fixed PONG probe and a separate structured JSON probe have passed against CUC; these calls do not establish vulnerability-research quality or authorize target testing.

M9.16 adds a CUC-only Chat PONG codec for the standalone probe. It binds request model cuc/deepseek to exactly deepseek-v4-flash or deepseek-v4-flash-0731, uses direct pinned HTTPS and only the real CUC_DEEPSEEK_API_KEY reference. provider-cuc-probe-config prints bounded configuration for an already-issued operator grant; it cannot issue one. No shim, generic Chat workflow or relaxed Responses identity check is added. The fixed CUC PONG and structured probes have passed real calls; the generic research adapter remains unavailable. See CUC setup.

About

Evidence-first agent framework for authorized web and cloud-native vulnerability research

Resources

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages