VulnLoom is an in-development, evidence-first autonomous vulnerability research and adversarial validation platform for software and systems that organizations own or are contracted to assess. Its two primary capability lines are Source Hunt and Authorized Red Team. The target product accepts source code, URLs, domains, IP addresses, networks, or combinations of them; plans bounded research tasks; validates Candidates in isolated environments; challenges them through an independent review step; and produces auditable reports. Pre-release acceptance and production-safe scheduled testing package these capabilities into operational workflows. The current implementation includes a bounded, resumable local Source Hunt V1 and fixed native Benchmark, shared Candidate-to-report gates, a closed Hybrid R10 evidence/retest/release-gate slice, and an Authorized Red Team R11 bounded attack-chain plus B1 observation-driven bounded replanning. B2.1 supports offline-tested, operator-sealed path GET observations: redirects, headers, bodies and credentials are disabled, while ordinary state receives only response metadata and digests. B2.2 can reduce the redacted Evidence from exactly one such completed GET into bounded OpenAPI 3.x path/method discoveries; server metadata and $ref values are never followed, and discoveries carry no execution authority. B2.3 adds a reviewed promotion gate: only explicitly selected, read-only discoveries whose templates have been concretized by an operator can atomically become a new Endpoint Seed Set, still without request execution authority. General public Recon, crawler, dictionary enumeration, arbitrary target expansion, lateral movement, and persistence are not enabled.
The product supports autonomous testing only inside an explicit, approved Scope. It must not scan or exploit unauthorized public targets. Its four planned entry points are source vulnerability research, pre-release security acceptance, production-safe scheduled testing, and authorized red-team simulation. Source Hunt and Authorized Red Team are the primary capability lines; the other two entries are controlled delivery workflows. Scope, network boundaries, credential isolation, evidence requirements, and human approval for consequential effects are enforced in code.
- White-box analysis of Python Web and API projects.
- Controlled dynamic validation in exact local/private fixtures; broader authorized live Web target execution remains staged and disabled by default.
- IDOR/BOLA, SSRF, path traversal, injection, insecure deserialization, authorization flaws, and sensitive data exposure.
- Planned human-reviewable report drafts for vendors, EduSRC, CNVD/CNNVD, and similar disclosure channels.
- No scanning of unauthorized public targets, automatic platform submission, or automatic CVE requests.
The long-term product design review draft and phased migration plan live in docs/PRODUCT-ARCHITECTURE-ROADMAP.md. Capabilities described there remain planned until their explicit acceptance stage passes.
- A Candidate is not a Finding. An agent may propose a candidate, but it only becomes a finding after reproducible evidence and an independent disproof check pass deterministic gates.
- The Control Plane owns privileged actions. Workers do not receive platform credentials, change authorization scope, or submit reports.
- Every validation uses an ephemeral sandbox. Source code is mounted read-only, output is stored separately, and network access is denied by default.
- Tools go through a Broker. Agents receive typed, limited, and auditable capabilities instead of an unrestricted host shell.
- Evidence comes before narrative. Every impact claim in a report must trace back to code, an observed request and response, or a reproducible test.
- Human approval is mandatory where it matters. State-changing tests, external callbacks, real credentials, and report submission require explicit approval.
- CONTEXT.md: domain terminology and shared language.
- AGENTS.md: mandatory engineering constraints for coding agents.
- docs/DEVELOPMENT-PLAN.md: active development order, current baseline, two capability-line plans, shared S1 security qualification, deferred work, and cross-chat handoff checklist.
- docs/PRODUCT-ARCHITECTURE-ROADMAP.md: review draft for the four product entry points, modular engines, long-term architecture, migration, and phased acceptance plan.
- docs/PRODUCT-POSITIONING-REVIEW.md: adversarial self-review of product scope, priorities, autonomy claims, and failure conditions.
- docs/MODEL-PROVIDER-CONTROL-PLANE.md: multi-provider setup, capability probing, role routing, Flow pinning, fallback, credential isolation, and Provider Center UX.
- docs/ARCHITECTURE.md: layers, components, and deployment topology.
- docs/WORKFLOW.md: agent orchestration, state machines, and verdict rules.
- docs/SECURITY.md: sandbox, network, credential, attachment, and evidence security.
- docs/DATA-MODEL.md: core entities and domain events.
- docs/ROADMAP.md: milestones and acceptance criteria.
- docs/PHASE3-ADMISSION.md: reproducible M4.3 production-isolation admission evidence.
- docs/EXTERNAL-BENCHMARKS.md: supported upstream layouts and local-snapshot safety boundary.
- docs/REFERENCE-PROJECTS.md: reference projects, adopted ideas, and rejected assumptions.
- docs/CUC-DEEPSEEK-CONFIG.md: local, secret-free setup and connectivity checks for the CUC DeepSeek gateway.
- docs/CODE-REVIEW-ASSIST.md: independent, approved, read-only model commentary on a manually selected Python snippet.
- docs/CANDIDATE-RECOMMENDATIONS.md: approved no-tool generation and deterministic admission for advisory Candidate recommendations.
- docs/SOURCE-HUNT.md: Source Hunt V1 contracts, CLI, security boundary, and remaining R9 depth.
- docs/PROJECT-RECIPES.md: A1 trusted, versioned project build/test recipe registry and execution boundary.
- docs/HYBRID-VALIDATION.md: R10 source-to-live evidence, dual remediation retest, report, and fail-closed CI/CD release gate contracts.
- docs/AUTHORIZED-RED-TEAM.md: fourth-entry Flow/RoE contract, offline Recon control slice, and safety boundary.
VulnLoom/
├── src/vulnloom/
│ ├── adapters/ # Model-provider configuration boundary
│ ├── agent_runtime/ # Typed offline model replay and proposal boundary
│ ├── analyzers/ # Python AST and optional Semgrep analysis
│ ├── benchmark/ # Deterministic offline metrics and regression gates
│ ├── broker/ # Typed tool mediation and HTTP policy enforcement
│ ├── critic/ # Deterministic independent counterevidence review
│ ├── domain/ # Domain objects, state machines, and protocols
│ ├── evidence/ # Redaction, hashing, and evidence storage
│ ├── hypotheses/ # Deterministic Candidate generation
│ ├── ingestion/ # Archive, Git, and OCI target ingestion
│ ├── policy/ # Scope and approval enforcement
│ ├── reporting/ # Evidence-consistent offline report drafts
│ ├── red_team/ # Authorized Red Team Flow, RoE, and Recon control plane
│ ├── runners/ # Offline and Docker sandbox runners
│ ├── source_hunt/ # Resumable white-box investigation and validation chain
│ ├── storage/ # Event and validation persistence
│ ├── validation/ # Plans, orchestration, and deterministic judging
│ ├── workflows/ # Shared visibility and execution-mode contracts
│ └── cli.py # Current command-line entry point
├── benchmarks/ # Sealed local ground-truth fixtures and baselines
├── docs/ # Architecture, workflow, security, and roadmap
├── schemas/ # Exported JSON Schema contracts
├── scripts/ # Schema and development utilities
└── tests/ # Offline tests and opt-in integration probes
An HTTP API and disclosure submission adapters remain planned components. A bounded live HTTPS provider adapter exists; CUC live acceptance covers the fixed PONG and fixed JSON probes. The target design treats it as the first Provider Profile within a multi-model Provider Center, rather than as a product-wide special case. CUC/DeepSeek remains the default provider for early development and explicitly authorized live acceptance; its existing probe, code-review, and Candidate-recommendation paths stay available until the generic adapter passes differential and live compatibility acceptance. A feature-gated OpenAI-compatible Chat Completions codec now has an offline fake-transport slice with strict typed decisions, model and usage checks, rejection, timeout, and cleanup tests. The admitted CUC probe identities are frozen in a content-addressed migration baseline. This new codec is not yet the default CUC path and has not received live CUC acceptance. The first trusted Profile assembly path now resolves only an explicitly allowlisted Control Plane endpoint reference, rechecks the admitted Flow and current Provider lifecycle, prepares content-addressed transport/codec contracts without network access, and creates a model registration only after the authoritative Egress Store proves a grant is active. That assembly now drives complete local fake-process Agent turns for two distinct Provider/model configurations. Negative turns cover authentication rejection, rate limiting, timeout, malformed output, zeroed credential leases, cleared wire buffers, and absence of receipts on failure. The generic path has also passed one authorized synthetic CUC Agent-decision acceptance run with a short-lived revoked grant, strict response identity and usage validation, zero tools, and verified buffer cleanup. Existing CUC feature paths remain unchanged while default-route migration is reviewed separately. The read-only code-review and Candidate Recommendation services can now bind their task-specific schemas to that trusted Profile path. A second synthetic Provider covers routed success, identity rejection, timeout, cleanup failure, replay, secret-buffer cleanup, and Candidate immutability; the CUC configuration remains the CLI default. Provider Center's trusted local CLI/application service supports referenced configuration, lifecycle, offline probing, revision-bound catalogs, audit, and default Route switching; its thin HTTP API and Web UI remain incomplete. M7.1a-M8.12 include deterministic replay, fixed provider messages, scoped credentials, isolated pinned HTTPS transport, typed Broker handoff, and a fixed two-tool Session ledger, human-gated Validation/Critic/Finding Intakes, and Approval-gated promotion. Benchmark and analyzer imports consume only sealed, pre-obtained local data and never fetch suites, rules, databases, or images.
The first implementation vertical slice was deliberately narrow:
Approved scope
→ Import one Python Web repository
→ Produce static security signals
→ Build a cross-file SourceGraph
→ Let a human select one Candidate
→ Start the test application in Docker
→ Validate the Candidate
→ Run an independent Critic
→ Produce a Markdown report draft
The current implementation now also exposes the Source Hunt V1 application service and nested CLI: bounded multi-language indexing, observation-driven cross-file investigation, integrity-checked redacted source windows, Candidate materialization, deterministic five-stage execution planning, resumable validation, authoritative Critic/Finding promotion, and report drafting. R9 fixed-Benchmark acceptance adds a sealed native coverage/ASAN adapter, normalized persistent Crash deduplication, and independently replayed PoV qualification. A1 closes the content-addressed Project Recipe Registry through network-disabled Docker admission and an authoritative non-fixture Recipe→Candidate→Validation binding; Recipe success never substitutes for vulnerability reproduction. Blind-holdout acceptance remains later depth work. External disclosure remains a separate future stage.
Unauthorized Internet-wide asset discovery, automatic submission, and unbrokered host shell access remain out of scope. The first isolated local HTTP HEAD Recon admission is implemented; broader authorized Web reconnaissance remains staged in the product architecture roadmap.
- Immutable Pydantic domain models and exported JSON Schema contracts.
- Multi-provider lifecycle, capability manifests, role routes, bounded fallback policies, and immutable Flow model snapshots; Provider endpoint and credential values remain outside these contracts.
- Separate Candidate state machine and deterministic Candidate-to-Finding gate.
- Scope Policy Engine and approvals bound to specific action digests.
- Evidence redaction, content addressing, and integrity verification.
- Checkpointed SQLite Control Plane event log with atomic tamper-evident audit records, rollback/fork detection, redacted projections, and owner-only filesystem checkpoint custody; remote signing and WORM remain deployment-stage work.
- Typed Control Plane/Worker protocol and an explicit Worker environment allowlist.
engagement-create,scope-approve, andstatusCLI commands.- An OpenAI-compatible model-provider configuration boundary; no network model call is made yet.
- Streaming, size-limited quarantine for ZIP and TAR artifacts.
- Member-by-member archive extraction without
extractall(). - Validation of normalized paths, symbolic links, special files, file counts, individual sizes, total expanded size, and compression ratio.
- Exact commit pinning for local Git repositories by reading Git objects without checkout, hooks, or target-code execution.
- Static classification of Kubernetes, Helm, Terraform, Dockerfile, and Compose files.
- Registration of OCI image references by
sha256digest without pulling images or connecting to Docker. - Atomic, read-only Target Snapshots with file-level SHA-256 manifests and idempotent reuse.
- Idempotent
TargetIngestedevents; failures and timeouts do not leave partial snapshots.
- Python AST indexing without importing or executing target code.
- Route discovery for Flask, FastAPI, Starlette, and Django.
- Structured functions, cross-file calls, authentication and authorization guards, ownership checks, dangerous sinks, and input-propagation paths.
- Content-addressed
SourceGraphandStaticSignaloutput. A signal is a hypothesis for validation, never a Finding. - File size and SHA-256 verification before analysis; tampering, path escape, resource-limit violations, and timeouts fail closed.
- Optional Semgrep adapter restricted to pre-registered local rules, with metrics and version checks disabled and no inherited API keys.
- Idempotent
SourceGraphBuiltsummary events. Full graphs are stored separately as read-only objects instead of being copied into the normal event log. - Scope identity, version, and validity are rechecked before every source-mapping run.
- CI runs lint, schema-drift checks, and the full test suite on Python 3.12, 3.13, and 3.14.
- Converts integrity-checked
SourceGraphobjects into typed, human-reviewable Candidates. - Merges complementary signals for the same route and sink without treating analyzer output as a Finding.
- Maps supported sink classes to CWE, preconditions, a security invariant, and the cheapest disproof task.
- Binds every Candidate to the exact Target version, SourceGraph digest, Scope identity, and Scope version.
- Uses stable Candidate UUIDs and SHA-256 duplicate fingerprints; repeated generation is byte-stable.
- Excludes parse failures, visibly guarded object lookups, and external matches that cannot be classified safely.
- Stores each
CandidateSetas an immutable content-addressed object and records only a redacted summary event.
- Defines immutable Static, Validation, and Report sandbox profiles with non-root identities, fixed mount slots, network modes, and resource ceilings.
- Rejects writable roots, Linux capabilities, host-path mounts, unregistered writable paths, and profile-purpose mismatches at schema validation time.
- Binds every Worker task to an exact Target version, Scope identity, policy digest, and sandbox-profile digest.
- Accepts only typed registered-tool invocations with argument arrays and logical working directories—never arbitrary shell command strings.
- Provides a
SandboxRunneradapter contract and a deterministic offline implementation for success, refusal, timeout, cancellation, resource exhaustion, checkpoint/resume, idempotency, and cleanup paths.
-
Adds an immutable capability registry whose digest is bound into every queued Worker task.
-
Requires a tool to be present in the Registry, Task allowlist, and Sandbox Profile while also matching the current Scope, policy digest, and Worker role.
-
Accepts HTTP methods, normalized credential-free URLs, safe headers, opaque credential/body references, and explicit time/size/redirect budgets—never raw credentials or request bodies.
-
Derives state-changing behavior from the HTTP method in trusted code and requires exact, unexpired approvals for mutations and credential use.
-
Reauthorizes every redirect, resolves every hop, pins the selected IP, verifies the reported peer, and rejects loopback, link-local metadata, multicast, unspecified, mixed-dangerous, and configured host-gateway addresses.
-
Returns only policy records, URL digests, peer metadata, Evidence IDs, final response-body SHA-256 digests, and budget usage; raw response bodies and sensitive headers stay outside the normal result path.
-
Uses deterministic offline resolver and transport adapters. No real HTTP request or network isolation claim is introduced in M4.2.
- Adds a trusted Docker CLI adapter that uses argument arrays only; Workers never receive the Docker socket or the host process environment.
- Resolves content mounts through a Control Plane-owned registry, pins exact image IDs, disables pulls, and replaces image entrypoints with registered absolute tool executables.
- Applies and re-inspects a read-only root, non-root UID/GID, dropped capabilities,
no-new-privileges, a content-bound versioned seccomp contract, no network, resource limits, read-only content, and boundednoexec,nosuid,nodevtmpfs mounts. - Kills timed-out Workers, removes containers and anonymous storage, and refuses to report a normal result unless absence is verified.
- Includes opt-in real-container probes for isolation, secret non-inheritance, timeout, and cleanup.
- Adds a live Broker-owned HTTP/HTTPS transport that connects directly to the policy-selected IP, preserves the authorized hostname for HTTP Host and TLS verification, ignores proxy environment variables, verifies the actual peer, and enforces response limits.
- Binds the selected resolver and transport implementation digest into the Tool Registry; queued work is rejected if offline and live adapters are swapped.
- Resolves request bodies by content digest and credentials through separate opaque providers; neither raw material is copied into Broker results or HTTP Evidence metadata.
- Stores only redacted response transcripts in the Evidence Store. Sensitive response headers, raw URLs, credential material, binary bodies, email addresses, and JSON-shaped secrets are excluded or redacted.
- Requires rootless mode, seccomp, cgroup v2, and enforceable memory, CPU-quota, and PID controls by default. Engines that only advertise partial isolation fail closed before container creation.
- Discovers daemon-managed network gateways for the Broker denylist. Docker Workers reject direct
target_onlynetworking and remain network-disabled; authorized target access belongs to the trusted Broker. - A dedicated Ubuntu 24.04 admission workflow runs Docker Engine 29.7.2 as a delegated rootless user service and proves Worker isolation from a live sibling container and daemon gateway, host-gateway denial before Broker transport, redirect-time DNS rebinding rejection, deterministic validation, timeout handling, and cleanup.
- Seals each human-selected Candidate content digest, exact Target/Scope provenance, networkless Runner request, and bounded Broker calls into a content-addressed
ValidationPlan. - Rechecks Candidate state, current Scope validity, policy digest, Validator role, profile digest, and Candidate input binding before writing a
STARTEDcheckpoint. - Runs the sandbox step before Broker calls, stops on the first non-completed result, and maps denials, missing approvals, timeouts, and failures to fail-closed domain outcomes.
- Persists one idempotent
ValidationOutcomein SQLite. Completed plans replay without re-execution; an interruptedSTARTEDplan requires explicit recovery and is never retried automatically. - Separates execution from verdict. The production default remains
INCONCLUSIVE; only a trusted deterministic judge may returnREPRODUCED, and it may cite only Evidence IDs collected by that run. - Produces an
EvidenceBundleand advancesCandidateonly through the existing state machine. It cannot create or promote aFinding. - Adds
validation-run-offline, which exercises the control-plane path without target execution, Broker calls, sockets, or a reproduced claim.
- Adds a content-addressed
HttpResponseAssertionselected before execution and bound to one exact Broker call. - Requires both an exact HTTP status and SHA-256 of the raw final response body. Status-only checks cannot produce a reproduced verdict.
- Keeps raw response bodies out of Broker results; only the body digest, bounded metadata, and redacted Evidence reference cross the trusted boundary.
- Adds
DeterministicHttpJudge: by default it trusts only the live pinned HTTP Registry; an exact match returns the precommittedREPRODUCEDorNOT_REPRODUCEDresult, while an offline Registry or any mismatch remainsINCONCLUSIVE. - Verifies every Evidence object through a no-follow, size-bounded, content-integrity read before judging or sealing an
EvidenceBundle. - Adds an opt-in composition probe covering a real ephemeral Docker Validator, a Broker-owned pinned HTTP connection to a temporary authorized fixture, Evidence capture, exact verdict, state transition, and cleanup.
- The composition probe also passes in the dedicated rootless Linux admission workflow. Local Docker Desktop runs retain an explicit rootful test-only exception and cannot independently qualify production.
- Seals Candidate, reproduced ValidationRun, EvidenceBundle, Scope version, validation context, and a distinct review context into a content-addressed
CriticPlan. - Requires separate validation and review producers and assesses security controls, reachability, environment parity, and version binding exactly once.
- Uses a fixed reducer: confirmed counterevidence rejects; any inconclusive angle leaves the Candidate validated but unpromotable; only four evidence-backed ruled-out angles advance it to
CRITIC_REVIEWED. - Rechecks every referenced Evidence object with no-follow, size, digest, and Target-version validation before changing state.
- Persists STARTED/COMPLETED SQLite checkpoints, returns completed outcomes idempotently, and refuses unfinished automatic replay.
- Performs no target execution, Broker call, network access, report submission, or Finding promotion. The final promotion gate separately rechecks current Scope, reproduced-run Evidence coverage, Critic binding, and duplicate review.
- Seals the Finding, promoted Candidate, approved EvidenceBundle, Scope version, channel, bounded narrative, and exact section citations into a content-addressed
ReportDraftPlan. - Requires code-location, request/response, reproduction, and impact claims to cite Evidence IDs from the Finding's bundle; every bundled Evidence object is rechecked for no-follow access, size, digest, and Target version.
- Redacts report text before persistence, escapes active HTML and Markdown image/link syntax, and never copies Evidence bodies into the draft.
- Renders deterministic generic, EduSRC, CNVD, vendor, and CVE-draft headings to immutable local Markdown and JSON artifacts.
- Uses STARTED/COMPLETED SQLite checkpoints, content-addressed artifact directories, bounded writes, idempotent completed replay, fail-closed recovery, and temporary-output cleanup.
- Produces only
draftreview status. It has no network adapter, platform credential, approval mutation, or submission path.
- Groups report revisions into a stable Finding/channel family and binds every version after the first to the exact preceding Report digest.
- Produces deterministic structured diffs for consecutive redacted revisions, including text and Evidence-reference changes; unchanged, unrelated, skipped, or unredacted revisions are rejected.
- Seals the exact Report digest, artifact digest, EvidenceBundle, Scope version, reviewer, diff, decision deadline, and approval expiry into typed review protocol objects.
- Applies only explicit
approve,request_changes, orrejectcommands through the Control Plane state machine. Any content or citation change invalidates the sealed request. - Allows local export only from
human_approved, before approval expiry, with an exact ReviewRecord and artifact match. Local export writes a new immutable Markdown/JSON artifact withexportedstatus. - Adds offline
report-review-diff,report-review-offline, andreport-export-localCLI paths. None has a network or Submission adapter.
- Seals local benchmark cases, ground-truth Findings, pipeline observations, policies, and baselines as typed content-addressed objects.
- Rejects any observed Finding that did not pass reproduced Validation, accepted independent Critic review, Candidate promotion, and complete Evidence gates.
- Computes Candidate recall, Finding precision, duplicate rate, Evidence completeness, policy violations, elapsed time, total cost, and cost per Finding with deterministic reducers.
- Applies absolute thresholds and exact-suite baseline comparisons, emitting stable violation codes and a failing CLI exit status for CI.
- Persists STARTED/COMPLETED SQLite checkpoints and immutable local JSON/Markdown results with bounded no-follow reads and temporary-output cleanup.
- Includes a generated local microbenchmark and baseline in
benchmarks/m6_1; ordinary CI verifies fixture drift and runs the offline regression gate. - Adds
benchmark-evaluate-offline. It consumes only sealed local files and has no Runner, Broker, network, credential, or Submission dependency.
- Adds versioned BountyBench and AutoPenBench adapters for pre-obtained local directory snapshots; neither adapter accepts a URL or downloads data.
- Seals every regular file by normalized path, size, and SHA-256, rejects symlinks and special files, and enforces file-count, per-file, total-size, and deadline budgets before and after normalization.
- Reads only
bounty_metadata.jsonlabels from BountyBench. Prompt, report, exploit, setup, patch, and verification contents are never copied into normalized suites. - Reads AutoPenBench
data/games.jsonin trusted code but persists only safe identities; task text and flags are discarded. CWE labels must come from a sealedvulnloom-autopenbench-cwe.jsonsidecar. - Emits typed exclusions for unsupported or missing labels and rejects malformed JSON, duplicate keys, stale mappings, ambiguous identities, adapter drift, and snapshot mutation.
- Stores normalized suites as immutable local objects with transactional import checkpoints and adds
benchmark-snapshot-manifest-localandbenchmark-import-offline.
- Normalizes local CodeQL SARIF 2.1.0, Trivy JSON, Checkov JSON, and Kubesec JSON through versioned adapters without running those tools.
- Seals the Target/version, tool version, rules digest, result digest, optional CWE mapping, adapter digest, resource limits, deadline, and idempotency key.
- Persists only rule/message digests, normalized CWEs, severity, and safe relative locations; it discards raw messages, secret matches, and Kubernetes object identities.
- Keeps analyzer observations structurally separate from pipeline
BenchmarkObservation: they cannot carry Candidate, Validation, Critic, or Finding state. - Adds
analyzer-result-manifest-localandanalyzer-observations-import-offline; neither command accepts a URL or executes a binary.
- Requires a sealed, explicit Observation-to-ground-truth alignment; matching CWE labels alone never count as a detection.
- Revalidates case, Target version, ObservationSet digest, truth ownership, and CWE compatibility before creating a checkpoint.
- Computes overall and per-analyzer truth recall, observation precision, duplicate rate, and exclusion rate for CodeQL, Trivy, Checkov, and Kubesec.
- Supports required-analyzer and full case-matrix gates plus exact-suite baseline regressions; per-analyzer checks prevent aggregate results from hiding one tool's regression.
- Includes
benchmarks/m6_3, fixture-drift checks, andrun_m6_3_regression_gate.pyin ordinary CI. - Adds
analyzer-evaluate-offline, which creates only local immutable JSON/Markdown results and cannot alter Candidate or Finding state. - Accepts directories only. ZIP/TAR acquisition and extraction remain outside this adapter; callers must use an independently hardened quarantine path before presenting a directory.
- Runs exact Checkov, Kubesec, Trivy, and CodeQL registrations through the hardened Docker Runner with inspected image IDs,
--pull never,network=none, read-only inputs/root, non-root identity, bounded resources, and mandatory cleanup. - Captures analyzer output within a trusted byte limit and completes the outer checkpoint only after the existing M6.3a adapter creates redacted Observations.
- Restricts Trivy to its sealed read-only DB v2 and vulnerability scanner; DB/check/Java/VEX updates, secret scanning, telemetry, and version checks are disabled.
- Binds CodeQL 2.26.2 to a Target/version/Manifest and one sealed prebuilt DB/query snapshot. A narrow wrapper copies the DB into bounded tmpfs because CodeQL writes query results, while the original DB and query pack remain read-only and are reverified after cleanup.
- Does not expose analyzer/package downloads, arbitrary commands, Target builds, Candidate/Finding promotion, or Submission. CodeQL database construction remains separately
RUN_UNTRUSTED_BUILDApproval-gated.
- Seals an exact offline replay implementation, provider/model identity, supported Worker roles, and output ceiling in a content-addressed registration.
- Binds each run to an exact
TaskEnvelope, context and decision-schema digests, step/token/wall budgets, deadlines, and an idempotency key. - Validates untrusted structured decisions and returns only terminal summary digests or a typed, argument-digest-only tool intent. It never executes the proposed tool.
- Persists transactional STARTED/COMPLETED checkpoints without raw model output or raw tool arguments; interrupted calls require explicit recovery and are not replayed automatically.
- Uses no model socket, SDK, endpoint, credential, Runner, Broker, Approval, domain transition, or Submission path.
- Replaces direct API-key string resolution with an initialization-time allowlisted, content-addressed reference to one explicit Control Plane environment variable.
- Holds secret bytes only in a non-serializable lease that is zeroed on success, exception, and timeout-result paths.
- Adds a registration-bound local fake adapter to test provider identity and credential lifecycle without a socket, URL, SDK, proxy, or inherited environment.
- Keeps the credential value and unrelated environment entries out of Worker requests, outcomes, checkpoints, schemas, and error messages.
- Requires transient context sources to match the Task's ordered input references exactly.
- Normalizes and redacts content in trusted code, rejects unsafe controls, and enforces fragment, total-byte, count, deadline, and wall-clock limits.
- Stores only explicitly untrusted, redacted fragments in immutable content-addressed snapshots bound to the exact Task, Target, Scope, references, and redaction policy.
- Revalidates no-follow, regular-file, read-only, size, schema, identity, and digest properties on every stored read; failed publication removes temporary files.
- Binds only the context snapshot ID into Agent run plans and step requests. It performs no provider, network, tool, Approval, or domain-state action; the Runtime reloads the snapshot before creating its STARTED checkpoint.
- Maps every Worker role to one built-in, content-addressed system template; callers cannot supply system text or template versions.
- Renders canonical strict JSON with typed control metadata separated from escaped, explicitly untrusted context fragments.
- Binds the plan, Task/context/model/template/schema, Target/Scope digests, tool allowlist, budgets, step, and messages into one envelope ID sealed into the step request.
- Rejects duplicate keys, template/system/control/trust drift, request mismatch, byte overages, and rendering timeouts before the first model call.
- Passes messages transiently to offline adapters while checkpoints and adapter audit lists retain only digests. Prompt text remains non-authoritative; Runtime/Broker enforce permissions.
- Adds a content-addressed admission object for one exact provider hostname, TLS port, canonical path, credential reference, adapter digest, request/response limits, and timeout.
- Fixes redirects and proxies off, DNS revalidation on, raw-response persistence off, one attempt,
and
network_enabled=false; schema validation rejects any relaxation. - Derives a digest-only transport request from the exact StepRequest and Message Envelope, while the serialized message body exists only in a zeroed transient buffer.
- Exercises credential acquisition, bounded response capture, strict JSON parsing, identity checks, typed rejection/timeout outcomes, receipts, and cleanup through an in-memory admission fake.
- Stores only request/response digests, counts, identities, stable status, and cleanup proof. It opens no DNS, socket, HTTP, SDK, proxy, tool, Approval, or Submission path.
- Adds a fixed
subprocess_https_provideradapter with a content-bound implementation digest; no caller-supplied executable, argv, header, URL, proxy, retry, or SDK is accepted. - Separates production
live_httpsadmissions (exact port 443 and global-only DNS) fromloopback_https_probeadmissions (.test, loopback-only, and a sealed test CA). - Re-resolves the exact hostname for every request, rejects mixed or forbidden answers, pins one numeric IP, and verifies the connected peer, TLS 1.2+ version, SNI, and hostname.
- Sends credential and message bytes over a fixed binary stdin frame to an isolated Python process
launched with
-I,/cwd, closed file descriptors, a tiny environment, no shell, bounded stdout, discarded stderr, resource limits, and forced process-group cleanup. - Allows one POST to one canonical path, forbids redirects and encoded responses, applies header/body and parent/child wall limits, and records only peer digests, TLS version, counters, and cleanup.
- Enforces a sealed per-minute rate limit with no automatic retry. Live transport remains library-only and no public provider is contacted by CI; Phase 3 uses a real loopback TLS subprocess probe.
- Adds content-addressed issuer policies that limit exact provider IDs, networked transport modes, and grant lifetimes; no-network fake modes cannot receive an egress grant.
- Issues immutable grants bound to one transport Admission, credential reference, adapter, purpose, issuer, validity window, and idempotency key, then binds the exact grant ID into model registration.
- Atomically publishes read-only grant/revocation objects and records STARTED/COMPLETED issuance and revocation checkpoints in a transactional lifecycle ledger.
- Reopens and verifies the grant, ledger state, expiry, and exact Admission binding before every DNS lookup, rate slot, credential lease, or child process. Revoked, expired, unfinished, linked, writable, malformed, or drifted grants fail closed.
- Keeps the authority local and library-only. It adds no remote signer, provider SDK/codec, public provider call, arbitrary URL, tool execution, Approval consumption, or Submission capability.
- Adds a content-addressed
openai-responses-v1codec bound into every subprocess HTTPS model registration and to the exact admitted/v1/responsespath. - Emits only fixed non-streaming, non-stored requests with disabled truncation and the registered strict Agent decision JSON Schema; callers cannot add tools, metadata, sessions, or parameters.
- Accepts only one completed assistant
output_textwith exact model identity and typed usage, then strictly parses it asAgentDecisionPayload. Incomplete, refusal, tool-call, duplicate-key, oversized, timed-out, and protocol-drift responses fail closed. - Keeps provider-native tool execution disabled. A structured tool proposal remains an inert intent that must pass the existing Runtime and Broker enforcement boundaries.
- Uses offline golden fixtures and the existing opt-in loopback TLS process probe; CI does not call a public provider or use a real provider credential.
- Binds one authoritative completed Agent run and digest-only tool intent to an independently built,
exact typed
BrokerCall; no model text is translated into executable parameters. - Restricts handoff to Validator tasks and rechecks Task, Scope, Policy, Sandbox Profile, Tool Registry, tool budget, call commitment, deadline, and Agent checkpoint before dispatch.
- Leaves all network, DNS pinning, credentials, side-effect, and Approval enforcement inside the existing Tool Broker. The Agent never receives a socket, Docker handle, secret, or adapter.
- Adds transactional STARTED/COMPLETED handoff checkpoints. Only an approval-required first attempt can be retried once with a new exact Broker call and independently verified Approval.
- Converts every completed Broker result into a digest-only
AgentToolObservationcontaining typed counts and Evidence references, never raw Agent arguments, URLs, credentials, or response bodies. - Tests both offline state-machine paths and an opt-in live pinned-Broker composition against a temporary authorized service; no public target, Candidate/Finding transition, or Submission is added.
- Derives one new Validator Task from an authoritative completed Agent → Broker chain while preserving the exact engagement, Target/version, Scope/version, Policy, Profile, Registry, model, and absolute deadline bindings.
- Fixes the derived Task to an empty tool allowlist and
tool_calls=0; the one-step continuation may finish as complete or blocked, while another tool proposal becomes a terminal failure. - Reopens each exact Evidence ref through the content-addressed Evidence Store, re-redacts it, and rebuilds the same bounded read-only context snapshot before any provider call. Observation and Evidence text remain explicitly untrusted model context.
- Accounts for tokens, prior steps, the consumed Broker call, and remaining wall time in a typed budget ledger. Exhaustion, expiry, authority drift, missing Evidence, or incomplete cleanup fails before the continuation checkpoint.
- Adds a unique-Observation STARTED/COMPLETED SQLite lifecycle with idempotent completed replay and fail-closed conflict/recovery. Its rows contain only typed outcomes and digests, never Evidence, provider response, URL, credential, Candidate/Finding, or Submission content.
- Phase 3 composes the real isolated loopback TLS provider, real pinned Broker, redacted Evidence Store, and zero-tool continuation against a temporary authorized target; no public egress is used.
- Adds a content-addressed
AgentSessionPlan, cumulative token/step/tool/provider/Broker budget ledger, and SQLite lifecycle around one already completed tool round, one optional second tool round, and one mandatory zero-tool terminal continuation. - Derives the second Validator Task with the exact inherited Target/Scope/Policy/Profile/Registry, model, and absolute deadline, while shrinking all budgets and fixing the total session to at most three provider turns and two consumed tool calls.
- Exposes only a trusted, content-addressed
AgentAuthorizedCallSetin message control. Each option is a Control-Plane-built exact read-onlyBrokerCallcommitment; the model cannot construct or alter URL, method, headers, body, credential, network, or authorization fields. - Reopens every Agent, handoff, Observation, Evidence, and context checkpoint before the next action. Unlisted or duplicate commitments, drift, exhausted budgets, missing cleanup, and a third tool proposal fail closed without another Broker call.
- Pauses an Approval-required second handoff without polling or approving it. One explicit, Approval-bound M7.8 retry may resume the session; its extra Broker attempt is counted even though the total successful tool-call budget remains exactly two.
- Phase 3 uses three isolated loopback provider subprocesses and two exact pinned-Broker reads of a temporary authorized target. No public target/provider, arbitrary loop, target build, domain-state transition, report export, or Submission path is introduced.
- Reopens the authoritative Session, Agent run, handoff, continuation and Evidence stores before
producing one content-addressed
AgentSessionAuditBundle; callers cannot provide a transcript or substitute a model-generated summary. - Recomputes ordered round identities, exact call commitments, Approval decision digests, Target/Scope provenance, cumulative token/step/tool/provider/Broker budgets and cleanup proofs.
- Projects only
completed,blocked,failedortimed_outwith a stable reason code and verified Evidence refs. The recommendation is not a Candidate/Finding/Report transition or an authorization. - Publishes bounded, read-only JSON/Markdown containing only digests, IDs, typed counts and statuses; it never copies Evidence bodies, URLs, credentials, provider requests/responses or tool parameters.
- Uses a separate SQLite STARTED/COMPLETED checkpoint with idempotent completed replay and fail-closed conflict/recovery. Artifact failure cleans temporary files and refuses automatic replay.
- Extends the loopback Phase 3 composition by creating the audit from the real M7.10 session and proving a tampered chain is rejected without adding runtime network or execution authority.
- Reopens the immutable M7.11 Audit artifact and CandidateSet, then binds their exact digests to a
Control-Plane-built typed
ValidationPlan; no Agent prose or Evidence body becomes a request. - Accepts only a human
accept,reject, ordefercommand bound to the exact Audit, Candidate and Validation plan. A blocked, failed or timed-out recommendation cannot be accepted. - Produces only a digest-only immutable decision record. The service has no Runner or Broker and does not queue Validation, mutate Candidate state, consume Approval, build a Target or submit data.
- Uses an independent STARTED/COMPLETED SQLite ledger; drift, expiry, duplicate consumption, conflicting decisions and unfinished recovery fail closed.
- Runs only after an explicit existing Validation entry point has completed; it cannot execute, resume, retry, queue, approve, or alter that Validation.
- Reopens the accepted Intake, Audit artifact, CandidateSet, exact Validation checkpoint, current Scope and every referenced Evidence object before creating a binding checkpoint.
- Recomputes Runner/Broker identities, ordered calls, forced timeout/policy results, run accounting, final Candidate state and EvidenceBundle consistency; drift fails closed.
- Persists only IDs, digests, typed result/state and timestamps in a unique STARTED/COMPLETED ledger. Idempotent replay is read-only and conflicting Intake/plan/outcome consumption is rejected.
- Does not call a Runner, Broker, provider, Docker, network, Approval, Target build, Critic, Finding, report export, public target, or Submission path.
- Reopens a reproduced M8.2 binding, Audit artifact, CandidateSet, completed Validation checkpoint
and Evidence before binding one independently constructed exact
CriticPlan. - Accepts only explicit human accept, reject or defer commands. Acceptance records selection but does not execute the Critic, mutate Candidate state, create a review/Finding, or export a report.
- Persists only IDs, digests, typed decisions and timestamps; conflict, expiry, drift and unfinished recovery fail closed without Runner, Broker, provider, network, Approval or Submission access.
- Runs only after the accepted M8.3 exact
CriticPlanhas completed through the existing explicit deterministic Critic entry point; the binding service cannot invoke, retry or recover Critic. - Reopens the M8.2 binding, Validation checkpoint, Evidence, Critic plan/outcome and current Scope, validating review identity, independent context, ruleset, counterevidence, rationale and time.
- Enforces the exact verdict/state mapping: accepted to
CRITIC_REVIEWED, rejected toREJECTED, inconclusive to unchangedVALIDATED. - Persists only IDs, digests, typed verdict/state and timestamps in a unique STARTED/COMPLETED ledger. It does not mutate the original Candidate or create a Finding, report, build, Approval or Submission.
- Requires an accepted M8.4 binding, exact critic-reviewed Candidate, reproduced Validation and Evidence provenance, accepted CriticReview, and the authoritative store's unique latest, content-addressed duplicate-clear proof.
- Binds one trusted-control-plane-built exact
FindingPromotionPlan; Agent prose, Critic rationale and Evidence content cannot construct root-cause, affected-version, impact or severity fields. - Records only an explicit human accept, reject or defer selection in a digest-only checkpoint. Acceptance does not call Candidate promotion, create a Finding, mutate state or draft a report.
- Drift, expiry, duplicate results, non-accepted Critic verdicts, conflicting consumption and unfinished recovery fail closed without Runner, Broker, provider, network or Submission access.
- Requires both an unexpired accepted M8.5 record and a granted
MUTATE_TARGET_STATEApproval bound to the exact Intake record, PromotionPlan, Candidate, Finding ID, Scope and two declared effects. - Reopens the complete Validation, Evidence, Critic, duplicate-check and Intake provenance chain, then calls only the existing pure Candidate-to-Finding state transition.
- Atomically records the promoted immutable Candidate and verified Finding in a unique STARTED/COMPLETED result ledger. Completed replay is idempotent; conflicts and unfinished recovery fail closed.
- Agent output cannot supply Approval or promotion fields, and the service has no Runner, Broker, provider, build, public-network, report-generation or Submission capability.
- Requires one completed, sealed M8.6 promotion outcome containing a promoted Candidate and verified Finding, plus the authoritative reproduced Validation and EvidenceBundle provenance.
- Binds a trusted-control-plane-built exact
ReportDraftPlan, including report family, channel and complete Evidence-cited sections. This first version accepts only version 1 drafts and persists only IDs, digests, typed decision and timestamps. - Accepts only explicit human accept, reject or defer. Acceptance does not draft a Report, read report prose into an Agent context, export an artifact, approve disclosure or create a Submission.
- Plan drift, missing or altered promotion checkpoints, Evidence corruption, expiry, conflicts and unfinished recovery fail closed without Runner, Broker, provider, target or network calls.
- Requires an unexpired human-accepted M8.7 record and reopens the complete M8.6, Critic,
Validation, Evidence and exact
ReportDraftPlanprovenance chain before drafting. - Seals the ordered typed Evidence catalog into a digest-only execution plan, then invokes only the existing deterministic offline report service with trusted control-plane inputs.
- Produces one immutable local report artifact that remains
DRAFT, plus a prose-free outcome binding. It cannot approve, export or submit the report, mutate Candidate/Finding state, build a target, or access Runner, Broker, provider or network capabilities. - Drift, expiry, non-accepted Intake, pre-existing unbound drafts, conflicts and unfinished checkpoints fail closed. Completed execution is read-only and idempotent.
- Reopens one completed M8.8 binding, its exact DRAFT Report and immutable artifact, plus the same ordered typed Evidence catalog used for drafting.
- Binds a trusted-control-plane-built exact
ReportReviewPlan; human input is limited to accepting, rejecting or deferring entry into the later review operation. - Persists only IDs, digests, typed selection and timestamps. Acceptance does not call the review
service or change the Report from
DRAFT, and grants no report approval or export authority. - Artifact/Evidence corruption, plan drift, expiry, conflicts and unfinished recovery fail closed without Runner, Broker, provider, target, network, export or Submission access.
- Requires an unexpired accepted M8.9 record, an independently issued exact human
ReportReviewCommand, and a grantedREVIEW_REPORTApproval bound to that command and its single expected state effect. - Reopens the M8.8 DRAFT, artifact, Evidence catalog and M8.9 checkpoint before invoking only the existing deterministic human-review state machine.
- Records the resulting review in the existing authoritative store and publishes a prose-free outcome binding. The source DRAFT remains immutable; completed replay is read-only and idempotent.
- Missing/revoked Approval, decision or provenance drift, expiry, pre-existing unbound reviews and unfinished recovery fail closed. No report export, public network or Submission capability is present.
- Reopens one completed M8.10
HUMAN_APPROVEDbinding, its authoritative review outcome and immutable artifact, then binds them to an independently constructed exactReportExportPlan. - Human input is limited to accepting, rejecting or deferring that exact local export plan. Even an
accepted record does not call
LocalReportExportServiceor change the Report toEXPORTED. - Persists only IDs, digests, typed selection and timestamps; report prose and review rationale stay outside the Agent Intake ledger.
- Non-approved reports, artifact or plan drift, expiry, conflicts and unfinished recovery fail closed. Runner, Broker, provider, target, public network, arbitrary destination and Submission capabilities are absent.
- Requires a still-valid accepted M8.11 record and an independently granted
EXPORT_REPORTApproval bound to the exact export plan, human-approved Report, review, artifact, Scope and fixed local state/artifact effects. - Reopens every authoritative M8.10/M8.11 checkpoint before invoking only the existing
LocalReportExportService; neither Agent output nor Approval prose can supply export parameters. - Produces an immutable local
EXPORTEDReport artifact and a prose-free outcome binding. The sourceHUMAN_APPROVEDReport and artifact remain immutable, and completed replay is read-only. - Missing/revoked Approval, drift, expiry, pre-existing unbound exports, conflicts and unfinished recovery fail closed. No destination path/URL, public network, platform token or Submission capability exists.
- Evaluates one content-addressed observation over the fixed 13-stage M7.11–M8.12 chain, including six explicit human Intake decisions and three exact Approvals.
- Requires reproduced Validation, accepted Critic, immutable source Candidate/Report states and one
final local
EXPORTEDartifact with bound Evidence. - Compares effect counters captured after Validation with the final export boundary. Any later provider, Broker, Runner or target call fails the gate, as does any public-network access, target build, automatic Approval or Submission.
- The evaluator is pure and offline; STARTED/COMPLETED checkpoints and content-addressed JSON/ Markdown results are bounded, no-follow verified and replayed read-only. The Phase 3 Admission harness constructs the observation from the real local composition rather than Agent prose.
- Seals one known-good M9.1 observation plus 23 fixed scenarios covering every current quality and safety violation branch; the corpus requires every mutation exactly once and in canonical order.
- Each scenario has an immutable expected PASS/FAIL status and exact violation-code tuple. Missing cases or changed expectations fail schema validation instead of silently updating the baseline.
- The offline corpus evaluator deterministically mutates only typed observations and requires all 23 actual results to match the sealed baseline. It has no Runner, Broker, provider, target or Approval dependency.
- Fixture generation and the M9.2 gate run in every Python 3.12/3.13/3.14 CI job. Public-network, target-build, automatic-Approval and Submission regressions are explicit failing scenarios.
VulnLoom requires Python 3.12 or later.
python3 -m venv .venv
.venv/bin/pip install -e '.[dev]'
.venv/bin/pytest --cov=vulnloom --cov-report=term-missing
.venv/bin/python scripts/export_schemas.py
.venv/bin/python scripts/export_benchmark_fixtures.py
.venv/bin/python scripts/run_m6_1_regression_gate.py
.venv/bin/python scripts/export_analyzer_evaluation_fixtures.py
.venv/bin/python scripts/run_m6_3_regression_gate.py
.venv/bin/python scripts/export_agent_workflow_regression_fixtures.py
.venv/bin/python scripts/run_m9_2_agent_workflow_regression_gate.py
# Optional; requires a local Alpine 3.22 image and Docker engine.
VULNLOOM_DOCKER_INTEGRATION=1 .venv/bin/pytest tests/test_runner_docker_integration.py
# Optional; opens one temporary loopback-only HTTP server.
VULNLOOM_SOCKET_INTEGRATION=1 .venv/bin/pytest tests/test_live_http_integration.py
# Optional; combines one real Docker Worker and one temporary authorized HTTP fixture.
VULNLOOM_COMPOSITION_INTEGRATION=1 \
.venv/bin/pytest tests/test_validation_composition_integration.py
# Explicitly opt in to the isolated Hybrid source-to-preprod acceptance.
VULNLOOM_HYBRID_E2E_INTEGRATION=1 \
.venv/bin/pytest tests/test_hybrid_e2e_integration.pyTarget ingestion requires an approved Scope that is still valid and explicitly includes either the artifact digest or the Git URL and commit.
vulnloom --db .vulnloom/events.db --store .vulnloom/targets \
artifact-quarantine --engagement-id <uuid> --source source.zip
# Add the returned source_name, kind, and artifact_id to a Scope,
# then have that Scope approved by an authorized reviewer.
vulnloom --db .vulnloom/events.db --store .vulnloom/targets \
target-ingest-archive --scope-file scope.json --source source.zip
vulnloom --db .vulnloom/events.db --store .vulnloom/targets \
target-ingest-git --scope-file scope.json --source /local/repo \
--repository-url https://example.invalid/project.git --commit <full-commit>
vulnloom --db .vulnloom/events.db --store .vulnloom/targets \
target-register-image --scope-file scope.json \
--image-ref ghcr.io/example/app --digest sha256:<digest>
# Build an offline Python Web SourceGraph from a verified filesystem snapshot.
vulnloom --db .vulnloom/events.db --store .vulnloom/targets \
source-map --snapshot-id <manifest-sha256> --scope-file scope.json \
--analysis-store .vulnloom/analysis
# Generate deterministic validation hypotheses. This does not execute or validate them.
vulnloom --db .vulnloom/events.db \
candidate-generate --graph-id <graph-sha256> --scope-file scope.json \
--analysis-store .vulnloom/analysis --candidate-store .vulnloom/candidates
# Rebuild trusted static results for one already-ingested authorized Snapshot and
# emit a local human-review summary. This does not select or validate a Candidate.
vulnloom --store .vulnloom/targets shadow-pilot-local \
--snapshot-id <manifest-sha256> --scope-file scope.json \
--analysis-store .vulnloom/analysis \
--candidate-store .vulnloom/candidates \
--readiness-db .vulnloom/pilot-readiness.db \
--readiness-store .vulnloom/pilot-readiness
# Record one explicit human Candidate choice from a completed passing shadow pilot.
# The timestamp is part of the immutable command; no ValidationPlan is created.
vulnloom --store .vulnloom/targets pilot-candidate-select-local \
--readiness-plan-id <readiness-plan-sha256> \
--candidate-set-id <candidate-set-sha256> --candidate-id <candidate-uuid> \
--scope-file scope.json --reviewer <human-identity> \
--decided-at <ISO-8601-with-timezone>
# Bind that completed selection to pre-existing, independently constructed M8.1
# plan/command files. This records Intake only and never runs Validation.
vulnloom pilot-validation-intake-bind-local \
--selection-readiness-plan-id <readiness-plan-sha256> \
--scope-file scope.json --audit-artifact-file audit-artifact.json \
--intake-plan-file intake-plan.json --intake-command-file intake-command.json \
--validation-plan-file validation-plan.json
# Bind one accepted model Recommendation selection to a prebuilt offline
# ValidationPlan. This writes Intake only and still requires RUN_VALIDATION Approval.
vulnloom candidate-recommendation-validation-intake-local \
--scope-file scope.json --generation-db generations.db \
--recommendation-db recommendations.db --selection-db selections.db \
--graph-store .vulnloom/graphs --candidate-store .vulnloom/candidates \
--validation-intake-db validation-intakes.db \
--selection-record-id <selection-record-sha256> \
--validation-plan-file validation-plan.json \
--idempotency-key recommendation-intake-001
# Execute that exact binding only after a separate human RUN_VALIDATION Approval.
# This first pilot execution path is offline and refuses every Broker call.
vulnloom pilot-validation-run-offline \
--pilot-intake-plan-id <pilot-intake-plan-sha256> \
--intake-plan-id <m8.1-intake-plan-sha256> \
--scope-file scope.json --validation-plan-file validation-plan.json \
--approval-file run-validation-approval.json
# Exercise a pre-sealed ValidationPlan through the offline control-plane path.
# The plan must not contain Broker calls.
vulnloom --db .vulnloom/events.db \
validation-run-offline --scope-file scope.json \
--candidate-store .vulnloom/candidates \
--candidate-set-id <candidate-set-sha256> --candidate-id <candidate-uuid> \
--plan-file validation-plan.json --validation-db .vulnloom/validation.db \
--evidence-store .vulnloom/evidence
# Compare two consecutive, already-redacted Report revisions without network access.
vulnloom report-review-diff --before report-v1.json --after report-v2.json
# Apply a pre-sealed human decision to an immutable local Report artifact.
vulnloom --db .vulnloom/events.db report-review-offline \
--scope-file scope.json --artifact-file report-artifact.json \
--evidence-bundle-file evidence-bundle.json --evidence-catalog-file evidence.json \
--review-plan-file review-plan.json --review-command-file review-command.json
# Mark an exactly approved Report as locally exported. This never sends it anywhere.
vulnloom --db .vulnloom/events.db report-export-local \
--scope-file scope.json --artifact-file approved-artifact.json \
--review-record-file review-record.json --export-plan-file export-plan.json
# Evaluate a sealed local benchmark. Exit status 2 means the regression gate failed.
vulnloom benchmark-evaluate-offline \
--suite-file suite.json --observations-file observations.json \
--plan-file benchmark-plan.json --benchmark-db .vulnloom/benchmarks.db \
--result-store .vulnloom/benchmark-results
# Seal an already-present local BountyBench directory without executing its files.
vulnloom benchmark-snapshot-manifest-local \
--source /local/bountytasks --kind bountybench \
--upstream-revision <full-commit> --license-spdx Apache-2.0
# Normalize a sealed local external snapshot. The plan binds the manifest and adapter digest.
vulnloom benchmark-import-offline \
--source /local/bountytasks --snapshot-file snapshot.json \
--plan-file import-plan.json --import-db .vulnloom/benchmark-imports.db \
--suite-store .vulnloom/benchmark-suites
# Seal a precomputed local analyzer output and optional explicit CWE map.
vulnloom analyzer-result-manifest-local \
--output codeql.sarif --analyzer codeql --target-id <uuid> \
--target-version <exact-version> --tool-version <version> \
--rules-digest <sha256> --cwe-map analyzer-cwe.json
# Normalize the sealed file without running the analyzer.
vulnloom analyzer-observations-import-offline \
--output codeql.sarif --cwe-map analyzer-cwe.json \
--snapshot-file analyzer-snapshot.json --plan-file analyzer-plan.json
# Evaluate sealed analyzer observations against an explicit reviewed alignment.
vulnloom analyzer-evaluate-offline \
--suite-file suite.json --alignment-file alignment.json \
--observation-set-file codeql-observations.json \
--observation-set-file trivy-observations.json \
--plan-file analyzer-evaluation-plan.json
# Validate a sealed source-only analyzer execution plan without running the tool.
vulnloom --store .vulnloom/targets analyzer-execution-check-offline \
--scope-file scope.json --snapshot-id <manifest-sha256> \
--registration-file analyzer-registration.json \
--plan-file analyzer-execution-plan.jsonshadow-pilot-local accepts only an already-ingested local Snapshot and its exact approved Scope. Trusted code rebuilds the SourceGraph and CandidateSet, checks the repository-owned admitted M9.4 baseline, stores immutable static/readiness artifacts, and prints Candidates for human review. An empty CandidateSet is a valid result, not evidence that the target has no vulnerabilities. The command has no Candidate-selection, Validation, Runner, Broker, provider, build, network, Approval, export, or Submission operation.
pilot-candidate-select-local reopens a completed passing readiness checkpoint and its immutable artifact, then verifies the exact SourceGraph, CandidateSet, Snapshot and Scope before recording one human-selected proposed Candidate. The digest-only record is suitable for a later M8.1 binding, but does not construct a ValidationPlan or authorize or execute Validation. One readiness result may select at most one Candidate.
pilot-validation-intake-bind-local requires that completed selection before it will apply an exact accepted M8.1 human command. All Audit, Candidate and ValidationPlan files must already exist and are reverified through the existing M8.1 service; the pilot bridge only adds an immutable provenance binding. It never derives Runner/Broker arguments, runs Validation, changes the Candidate, or emits a domain event.
pilot-validation-run-offline accepts only a completed M9.8 binding, its accepted M8.1 record, the identical pre-existing ValidationPlan and a separately granted exact RUN_VALIDATION Approval. The first version refuses every Broker call and all network activity, uses the offline Runner adapter, and records a digest-only execution binding. It never constructs a ValidationPlan or Approval from Agent output and never writes a Submission.
validation-run-offline accepts an already sealed, typed ValidationPlan. It rejects plans containing Broker calls, does not execute target code or open sockets, and defaults to an INCONCLUSIVE verdict. Live Broker/Docker composition currently exists as a library and opt-in integration-test path, not as a production CLI or HTTP API.
The report review commands accept only sealed JSON contracts and content-addressed local objects. report-export-local changes the Report to exported only inside the local store; it has no destination URL, disclosure adapter, or submitted transition. Benchmark evaluation consumes precomputed typed pipeline observations. External benchmark and analyzer imports only normalize pre-obtained local data. analyzer-execution-check-offline validates the sealed M6.4a control-plane contract but deliberately produces no analyzer output. M6.4b-d add library-only, exact-image, network-disabled Checkov/Kubesec/Trivy/CodeQL execution, and every output must pass the same M6.3a import boundary. M6.5 requires a complete, content-bound execution matrix before invoking M6.3b; M6.6 proves that four-analyzer composition in rootless Admission, including missing-cell and drift refusal. No CLI path installs or runs analyzers.
ModelProviderConfig.credential_reference identifies one explicit Control Plane environment variable without containing its value. M7.1b-M7.6 resolve it into a short-lived, zeroed Control Plane lease. M7.5 passes its bytes only to a one-shot isolated HTTPS child over a bounded stdin frame; M7.6 requires an active, exact, operator-issued egress grant before the credential is read. The Worker never receives the reference or value. Keys never appear in TaskEnvelope, Worker environments, Agent requests, checkpoints, event logs, or Evidence summaries. .env.example lists variable names only, and real .env files are ignored by Git.
VulnLoom is under active development. The current release provides the trusted domain foundation, secure local target ingestion, offline static source mapping, deterministic Candidate generation, a hardened Docker adapter, live pinned Broker transport, transactional validation orchestration, deterministic HTTP assertions, redacted Evidence storage, offline benchmark gates, precomputed multi-analyzer normalization, sealed Checkov/Kubesec/Trivy/CodeQL execution, execution-to-evaluation qualification, and a typed Agent Runtime with scoped credential leases, sealed context/messages, subprocess-pinned HTTPS transport, operator-issued egress lifecycle enforcement, exact Agent-to-Broker handoff, sealed Observation continuation, and a fixed two-tool Session ledger, plus opt-in probes for real containers, analyzers, sockets, and full validation composition.
Live Docker/Broker validation, the report workflow, and provider transport are exposed through typed library paths, not a production HTTP API or general provider CLI. The rootless Linux and OS-level egress admission gate passes; transport is qualified against a local TLS fixture and bounded CUC probes, while a general research adapter remains unavailable. External disclosure/CVE submission workflows, provider-specific general adapters, and dedicated Kubernetes, Terraform, or Helm vulnerability analyzers are not implemented yet.
Use VulnLoom only on systems, source code, and test environments for which you have explicit authorization.
The M9.10 pilot outcome gate consumes a completed M9.9 execution binding before recording M8.2 outcome provenance. Supply presealed plans created by trusted control-plane code:
vulnloom pilot-validation-outcome-bind-local \
--plan-file pilot-outcome-plan.json \
--outcome-plan-file agent-validation-outcome-plan.json \
--validation-plan-file validation-plan.json \
--audit-artifact-file audit-artifact.json \
--scope-file scope.jsonThe command uses local stores (override the --*-db and --*-store flags as needed). Identical
sealed plans replay read-only. Missing or unfinished execution, provenance drift, expired windows,
and a pre-existing bare M8.2 checkpoint fail closed. It does not execute Validation, mutate a
Candidate, grant Approval, build a Target, access the network or create a Submission.
M9.11 adds pilot-bound human Critic Intake. An operator supplies a sealed M9.11 plan, completed M9.10 provenance, and an independently prepared M8.3 human accept command and CriticPlan:
vulnloom pilot-critic-intake-bind-local \
--plan-file pilot-critic-intake-plan.json \
--pilot-outcome-plan-file pilot-outcome-plan.json \
--outcome-plan-file agent-validation-outcome-plan.json \
--intake-plan-file agent-critic-intake-plan.json \
--intake-command-file agent-critic-intake-command.json \
--critic-plan-file critic-plan.json \
--validation-plan-file validation-plan.json \
--audit-artifact-file audit-artifact.json \
--scope-file scope.jsonUse the --*-db and --*-store flags for the authoritative local stores. Only reproduced Validation
with a validated result Candidate and complete Evidence is eligible; the default offline pilot's
inconclusive result remains rejected. The command records Intake without running Critic or Validation,
changing Candidate state, granting Approval, building a Target, accessing the network or submitting
anything. Completed replay is read-only; missing/STARTED provenance and bare pre-existing M8.3
checkpoints fail closed. A subsequent pilot Critic execution stage is still required.
M9.12 connects approved pilot Critic execution and M8.4 result binding. Prepare a sealed execution
plan through trusted control-plane code and obtain an independent exact human RUN_CRITIC Approval:
vulnloom pilot-critic-run-local \
--plan-file pilot-critic-execution-plan.json \
--approval-file critic-approval.json \
--pilot-intake-plan-file pilot-critic-intake-plan.json \
--pilot-outcome-plan-file pilot-outcome-plan.json \
--outcome-plan-file agent-validation-outcome-plan.json \
--intake-plan-file agent-critic-intake-plan.json \
--intake-command-file agent-critic-intake-command.json \
--critic-plan-file critic-plan.json \
--evidence-catalog-file evidence-catalog.json \
--validation-plan-file validation-plan.json \
--audit-artifact-file audit-artifact.json \
--scope-file scope.jsonOverride the --*-db/--*-store flags to use the same authoritative stores as previous stages. The
catalog is a bounded JSON array of typed Evidence metadata in Bundle-reference order, without Evidence
body text. The command performs a local deterministic review of sealed assessments and records M8.4
and pilot bindings; it does not run Validation or target code, create Approval, build, access the
network or submit. Replays do not review again. Incomplete checkpoints require explicit recovery.
Accepted Critic outcomes still require separate duplicate-check and Finding-promotion gates.
M9.13 adds pilot-finding-intake-bind-local for human Finding Intake provenance. Supply the same
upstream files and authoritative stores as M9.12, replacing --plan-file with the sealed
PilotFindingIntakePlan, and add:
--critic-execution-plan-file pilot-critic-execution-plan.json
--critic-binding-plan-file agent-critic-outcome-plan.json
--finding-intake-plan-file agent-finding-intake-plan.json
--finding-intake-command-file agent-finding-intake-command.json
--promotion-plan-file finding-promotion-plan.json
--duplicate-check-file finding-duplicate-check.json
--duplicate-check-db .vulnloom/finding-duplicate-checks.db
--finding-intake-db .vulnloom/agent-finding-intakes.db
--pilot-finding-db .vulnloom/pilot-finding-intakes.db
Prepare sealed plans through trusted control-plane services. --approval-file carries the historical
M9.12 RUN_CRITIC Approval, not a promotion grant. The command requires accepted Critic provenance,
a current latest CLEAR duplicate check, an exact PromotionPlan and an independent human ACCEPT
command. It records M8.5 Intake and a digest-only pilot binding. It does not create a Finding, change
Candidate state, run Critic/Validation/target code, approve an operation, build, access the network or
submit. Completed replay is read-only; incomplete checkpoints require explicit recovery.
M9.14 adds pilot-finding-promote-local for an explicitly approved local Finding state transition.
Supply the same upstream files and stores as M9.13, replace --plan-file with the sealed
PilotFindingPromotionPlan, and add:
--pilot-finding-intake-plan-file pilot-finding-intake-plan.json
--promotion-execution-plan-file finding-promotion-execution-plan.json
--promotion-approval-file finding-promotion-approval.json
--promotion-db .vulnloom/finding-promotions.db
--pilot-promotion-db .vulnloom/pilot-finding-promotions.db
Trusted control-plane code prepares the M8.6 execution plan and pilot plan after completed M9.13
Intake and a separate exact human promotion Approval. The historical Critic --approval-file grants
no promotion authority. Execution rechecks the current Scope, latest CLEAR duplicate proof and the
entire accepted provenance chain. The result contains a verified Finding and promoted Candidate in
the M8.6 outcome store plus a digest-only pilot binding. The source CandidateSet remains unchanged.
Completed replay only verifies persisted results; incomplete checkpoints require explicit recovery.
The command does not run Validation, Critic or target code, approve operations, derive Runner/Broker
parameters, build, access the network or submit anything.
M9.15 provides a standalone fixed-message Provider probe. provider-probe-prepare seals a bounded
plan without reading credentials or opening a connection; provider-probe-run requires an existing
operator-issued egress grant and explicit --allow-provider-network. It uses only synthetic content
and no tools or research targets. Each grant is consumed once in the authoritative probe ledger,
including failed attempts; replay is read-only and STARTED requires explicit recovery.
See Provider probe setup and remaining admission requirements for the configuration types and commands. A real fixed PONG probe and a separate structured JSON probe have passed against CUC; these calls do not establish vulnerability-research quality or authorize target testing.
M9.16 adds a CUC-only Chat PONG codec for the standalone probe. It binds request model cuc/deepseek
to exactly deepseek-v4-flash or deepseek-v4-flash-0731, uses direct pinned HTTPS and only the real
CUC_DEEPSEEK_API_KEY reference. provider-cuc-probe-config prints bounded configuration for an
already-issued operator grant; it cannot issue one. No shim, generic Chat workflow or relaxed
Responses identity check is added. The fixed CUC PONG and structured probes have passed real calls;
the generic research adapter remains unavailable. See CUC setup.