Skip to content

Repository files navigation

armada

armada

Loop engineering for software development.

Turn any repository into a self-organizing AI engineering team.
One command. 8 specialists. Evidence-gated delivery.

npm version npm downloads License: MIT Node.js >= 22 CI

Website · Getting Started · User Guide · Why armada?


The problem with AI coding agents

AI coding agents are fast. They are also unsupervised, unverified, and amnesiac.

  • They guess when they should ask. "Build the login page" becomes a framework choice, an auth strategy, and a color palette — all decided without you.
  • They skip verification. The same agent that wrote the bug declares it fixed. There is no maker/checker split.
  • They forget everything. Kill the terminal, lose the context. A 5-phase feature starts from scratch.
  • They have no boundaries. A solo agent rewrites your CI config while fixing a CSS bug. Nothing stops it.
  • They clobber your working tree. Direct edits on your active branch. Parallel tasks collide. Recovery is manual.

These are not model failures. They are environment failures. The fix is not a smarter model — it is a smarter harness and a tighter loop.

The fix: loop engineering

Loop engineering replaces one-shot prompting with control loops that prompt agents for you. You define the goal. The loop handles dispatch, verification, gating, and iteration — agents are components in the loop, not autonomous actors.

armada implements loop engineering as a concrete system. You write the contract (what to build and how to know it works), and the fleet runs the loop until a reviewed Pull Request lands in your repo.

armada workflow

   contract  →  dispatch  →  build  →  test  →  review  →  PR
        ↑           |                                        |
        |      [phases with no          [defects loop back   |
        |       dependency run           to the developer]   |
        |       in parallel]                                 |
        └──────── evidence gates every transition ──────────┘

The loop has mechanical properties that make it reliable:

  • Maker/checker split. Developers write code. QA and the adversary check it. A maker never passes its own work.
  • Parallel phases. Independent phases dispatch simultaneously as background subagents with disjoint file scope. Only phases that depend on each other serialize.
  • Evidence, not reports. Every gate requires proof you can read — a passing test run, a screenshot, a file:line citation. Nothing advances on "trust me."
  • Crash-proof state. Every transition writes to disk. Kill the session, reopen, and the loop continues where it left off.

Parallel feature work

armada isolates each feature in its own Git worktree (sandbox). Multiple features run simultaneously in the same repo without colliding — each with its own contract, state, and branch. One fleet, many voyages.

armada voyage auth-system              # boots a lane for feature "auth-system"
armada voyage dashboard                # boots another lane — runs in parallel
armada fleet                           # dashboard: one row per active lane

Features in separate worktrees cannot collide. main stays pristine. Every voyage ends in a PR, never a local merge.

Why this works →

Quick start

# Install globally
npm install -g @rafamacalaba/armada

# Existing repo — detects your stack, scaffolds the team in place.
cd your-repo && armada init

# New project — questionnaire, scaffold, ready to ship.
armada new my-app && cd my-app

# Zero-install trial.
npx @rafamacalaba/armada@latest new my-app

Requires Node.js 22+ and an authenticated opencode install. Run armada doctor to confirm your environment is ready.

OpenRouter Provider Discounts

Save up to 20x on OpenRouter models by routing to discounted providers (Novita, StreamLake, Xiaomi):

# Check live OpenRouter provider prices and savings multipliers
armada models --discounts

# Init repo with a preferred discounted provider
armada init --openrouter-provider Novita

After armada init, open the repo in opencode. The fleet is loaded. Describe your goal in plain language — a wish, a ticket, a PRD — and the Commodore (orchestrator) interviews you for the missing details. Together you produce the contract (armada/REQUIREMENTS.md) with phases, dependencies, and measurable success criteria. Once you approve, the fleet runs the loop autonomously until a PR opens.

The fleet

Role Codename What it does
You Admiral Sets the mission, signs the contract, merges the PR
Orchestrator Commodore Co-writes the contract, dispatches specialists, gates evidence
Backend Galleon Server logic, APIs, databases, backend tests
Frontend Clipper UI, styling, responsive pages, client tests
QA Corvette E2E tests, screenshots, owns the defect ledger
Adversary Xebec Hostile review — hunts edge cases, vulns, UI flaws
Security Frigate Auth, permissions, data leaks, dependency audit
Docs Caravel READMEs, API docs, changelogs, user manuals
Architect Bark Code review, refactoring risk, pattern compliance (read-only)

Boundaries are enforced by SDK permissions, not prompt politeness. The Commodore cannot edit source code. Security, adversary, and architect can only write their own review artifacts. Full per-role detail in the user guide.

What you get

your-repo/
├── opencode.json
├── AGENTS.md
├── armada/
│   ├── armada.yaml               # manifest: re-runnable source of truth
│   ├── REQUIREMENTS.md           # contract: phases + success criteria
│   ├── state/                    # restart-proof loop memory
│   ├── ledgers/<feature>/        # DEFECTS.md, ADVERSARIAL_REVIEW.md, SECURITY_FINDINGS.md
│   ├── e2e/<feature>/            # per-feature E2E evidence
│   └── screenshots/<feature>/    # per-feature visual evidence
└── .opencode/
    ├── agent/                    # 8 native agents with SDK-enforced permissions
    └── commands/                 # slash commands (/voyage, /patrol, /fleet, /status)

armada init never clobbers existing opencode.json or AGENTS.md. Re-scaffold any time from the manifest:

armada init --from-armada armada/armada.yaml --restart

Built with armada

armada uses itself. The fleet builds armada's own features through the same contract/dispatch/gate loop that any user would run.

  • The fleet built armada's session-based state system in ~26 minutes, at a cost of $0.18, running fully autonomously. Blank contract to working code with passing tests.
  • It surfaced a real permission deadlock — a case where the Commodore's deny-all-edit rule conflicted with a state-write. The fleet asked the right question instead of silently failing.
  • QA caught and the loop self-corrected 3 test failures the developers introduced. The gate sent them back; they fixed them.

Every feature armada ships was built by armada. Read the full story →

Learn more

  • Website — visual walkthrough of the loop, the fleet, and what each phase produces.
  • Getting started — your first feature, end to end.
  • User guide — fleet concepts, roles, day-to-day usage.
  • Why armada? — the case for loop engineering over one-shot prompting.
  • Architecture — the full technical deep dive.
  • Operator guide — CLI reference, upgrades, rollback.

License

MIT. See LICENSE.

About

Loop engineering for software development. Turn any repository into a self-organizing AI engineering team — 8 specialists, contract-driven, evidence-gated, parallel feature voyages.

Topics

Resources

Contributing

Stars

66 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages