A runnable operating system that turns PRD and Jira-style inputs into an execution-ready roadmap, dependency graph, RAID log, executive brief, metrics, and trace evidence.
One realistic input produces six reviewable artifacts in milliseconds, offline, with no API key and no model cost. The browser demo now runs the same contract checks client-side and downloads the full artifact bundle as a ZIP:
PRD / Jira input
│
├── roadmap.md sequencing, ownership, exit criteria
├── dependencies.mmd renderable Mermaid dependency graph
├── raid-log.md risks, assumptions, issues, decisions
├── executive-brief.md decision-focused leadership summary
├── run-metrics.json latency, completeness, cost, review state
└── trace.jsonl observable run event
The default output is explicitly marked DRAFT — HUMAN REVIEW REQUIRED. The system refuses unknown dependencies, duplicate IDs, and dependency cycles before creating a plausible-looking plan.
git clone https://github.com/christiancaviedes/agentic-program-ops.git
cd agentic-program-ops
python3 -m venv .venv && source .venv/bin/activate
python -m pip install -e .
program-ops examples/launch-input.json --output build/launch-planInspect the generated artifacts:
ls build/launch-plan
cat build/launch-plan/executive-brief.mdAfter an accountable human reviews the plan, regenerate it with an approval marker:
program-ops examples/launch-input.json --output build/approved-plan --approveTry the live interactive compiler: paste or upload a PRD/Jira-style JSON input, compile the plan, inspect every artifact, record the human approval state, and download the bundle. Input never leaves the browser.
The CLI accepts structured JSON with program intent, target date, named decisions, risks, and Jira-style workstreams. Each workstream has an ID, owner, team, delivery window, status, exit criteria, and dependencies.
See examples/launch-input.json for a complete enterprise AI launch scenario.
| Control | Evidence |
|---|---|
| Input integrity | Required-field, unique-ID, reference, and cycle validation |
| Human accountability | Draft-by-default review gate; explicit --approve action |
| Regression protection | Unit tests across Python 3.10, 3.11, and 3.12 |
| Evaluation gate | Completeness, dependency recall, risk visibility, review-state, and cost checks |
| Observability | Per-run metrics plus an OpenTelemetry-compatible JSONL span |
| Cost discipline | Deterministic local core; estimated model cost is $0 |
| Failure behavior | Invalid plans exit non-zero with actionable errors |
| Supply chain | No runtime dependencies; least-privilege GitHub Actions permissions |
Run the same gates as CI:
python -m unittest discover -s tests -v
python -m evals.run_evalsThe first release uses a deterministic compiler rather than an LLM. That is deliberate: dependency logic, risk ownership, and approval state should be testable before generative enrichment is introduced.
flowchart LR
A[PRD + Jira-style JSON] --> B[Contract validation]
B --> C[Dependency graph validation]
C --> D[Artifact compiler]
D --> E[Roadmap]
D --> F[RAID log]
D --> G[Executive brief]
D --> H[Metrics + trace]
E & F & G --> I{Human review}
I -->|approved| J[Planning baseline]
I -->|changes| A
What this optimizes for: reliable structure, inspectability, fast iteration, and a clean contract for future model-backed synthesis.
What it does not claim: autonomous program management, correct business judgment, or a substitute for accountable owners.
Read the architecture, ADR, threat model, and sanitized production case study.
The read-only Jira adapter fetches a bounded issue set using JQL and maps keys, owners,
status, due dates, labels, and Blocks links into the validated contract:
export JIRA_BASE_URL="https://example.atlassian.net"
export JIRA_EMAIL="program-owner@example.com"
export JIRA_API_TOKEN="..."
program-ops-jira \
--jql 'project = AIP AND fixVersion = "Pilot" ORDER BY Rank' \
--metadata examples/jira-program-metadata.json \
--output build/jira-program.jsonThe adapter performs no writes and never places the API token in command arguments or generated output. See the Jira adapter guide.
The deterministic core remains the source of truth. A separate enrichment command can ask an OpenAI-compatible endpoint—including a local gateway—for bounded narrative, questions, and risk notes without changing owners, dates, dependencies, status, or approval:
export LLM_API_KEY="..."
program-ops-enrich examples/launch-input.json \
--base-url https://api.example.com \
--model provider-model-name \
--output build/enrichment.jsonThe provider contract is swappable, credentials remain in environment variables, output is always review-only, and provider failure never blocks deterministic compilation. See the enrichment boundary.
- Missing/invalid fields: the CLI stops before creating artifacts; correct the input and rerun.
- Unknown dependency: the CLI names the workstream and missing reference.
- Dependency cycle: the CLI rejects the plan instead of inventing an execution order.
- Unreviewed output: every artifact remains visibly marked as requiring human review.
- Interrupted run: outputs are generated from source input and can be safely regenerated in a clean directory.
- Future model-provider failure: the deterministic core remains the fallback path; model enrichment must never bypass validation or approval.
v0.1: deterministic CLI, CI, evals, trace/metrics, Pages demo, governance docsv0.2: interactive browser compiler, JSON upload, artifact preview, approval gate, ZIP exportv0.3: read-only Jira Cloud adapter, provider-neutral LLM enrichment boundary, and OpenTelemetry-compatible span recordsv0.4: Linear adapter and diff-aware plan updatesv1.0: policy-controlled multi-program workspace with approval audit trail
AI programs rarely fail because teams lack another summary. They fail because inputs are ambiguous, dependencies stay hidden, decisions have no owner, and generated output looks more certain than the evidence allows.
This project demonstrates the operating discipline I bring to AI platform and technical program leadership: make the contract explicit, instrument the workflow, put quality gates before launch, and keep human accountability visible.