Built from the raw OpenAI-compatible chat-completions API (no Agent SDK, no framework) so every piece of "agent orchestration" is visible and editable. Talks to Hugging Face's Inference Providers router by default — any provider speaking the same wire format (Fireworks, a Modal-hosted vLLM server, a local Ollama/vLLM instance) works by changing one environment variable, not the code. This isn't meant to replace Claude Code — it's meant to teach you what Claude Code (and every other agent product) is actually doing underneath.
| File | Answers |
|---|---|
agent/core.py |
What is an agent loop, mechanically? |
agent/tools.py |
What is this agent allowed to do? |
agent/prompts.py |
What should it do with those tools? (your policy) |
agent/schemas.py |
What shape should its final answer take? |
agent/orchestrator.py |
How many agents run, over what, in what order? |
That split is the actual lesson. core.py doesn't know what a PR is.
tools.py doesn't know when to merge vs. fix. prompts.py doesn't know
how to call the model's API. Each piece is replaceable independently —
that's what "build an agent for a product" means in practice: you're
almost always recombining these five kinds of things, not inventing a new
kind of loop each time.
Every "agent" you'll ever build is: send the conversation + tool schemas
to the model → if it asks for a tool, run it locally and append the
result as a new message → repeat → stop when the model stops asking for
tools (or you hit a turn cap). Read Agent.run() in core.py top to
bottom — the loop itself is still about 25 lines once you separate it
from streaming and structured-output logic. Everything else in this repo
exists to configure that loop, not to extend it.
Each turn is sent with stream=True — see Agent._send_turn(). Text
prints token-by-token as delta.content arrives, and the message is
reassembled from chunks into an object with .content and .tool_calls,
so the rest of the loop doesn't need to know a turn was streamed at all.
The part worth knowing: under the OpenAI-compatible format, a tool call's
arguments string arrives fragmented across multiple chunks,
identified by an index field — not delivered whole the way a
single-vendor SDK's stream helper might hand it back. _send_turn()
accumulates fragments by index and concatenates each one's arguments
piece until the stream ends. This is the single most fragile-feeling part
of the whole file, and it's covered by a dedicated test with a
deliberately multi-chunk-split tool call — see the "Implement streaming
turns" story in ISSUES.md for why that test exists.
agent/schemas.py defines PROutcome — a Pydantic model with
pr_number, status (an enum: merged / fix_pushed_pending_ci /
needs_human_review), and a one-line summary.
Once the tool-use loop ends, Agent._extract_structured_outcome() makes
one more, separate, tool-free call with
response_format={"type": "json_schema", "json_schema": {...}}, asking
the model to restate the outcome it already reached, constrained to that
schema. orchestrator.py's summary now reads result.outcome.status
directly instead of string-matching free text.
Why a separate call instead of attaching the schema to every turn: a
turn where the model still needs to call merge_pr or read_file isn't
supposed to produce a final JSON answer yet, and forcing a structured
response format onto those turns fights the tool-use turns instead of
complementing them. One clean extraction call after the loop naturally
ends is simpler to reason about — and cheap, since the conversation it's
built from is already sitting in the prompt cache from the turns that
just ran.
Real caveat, not a hypothetical one: Hugging Face's router fans out to
many different backend providers and models, and strict JSON-schema
adherence isn't guaranteed equally across all of them the way it would be
against a single vendor's own models. _extract_structured_outcome()
wraps the whole attempt in a try/except for exactly this reason — a
provider that doesn't honor strict mode degrades to outcome=None
(with final_text still populated) rather than crashing the run. Verify
this against whatever model you actually settle on; don't assume it from
one provider's behavior generalizing to another.
Two environment variables control where requests go:
INFERENCE_PROVIDER_TOKEN— who you are to the providerINFERENCE_PROVIDER_BASE_URL— which provider, defaulting tohttps://router.huggingface.co/v1if unset
Any provider speaking the OpenAI chat-completions wire format works by
changing INFERENCE_PROVIDER_BASE_URL alone — Fireworks, a Modal-hosted
vLLM server, a local Ollama or vLLM instance. This does not extend to
Anthropic's own API — it's a genuinely different shape (different
tool-call format, different streaming, different structured-output
mechanism), not a base_url swap. Supporting both would mean a real
adapter layer — a ModelBackend interface with one implementation per
wire format — which this project deliberately doesn't build until there's
an actual need for it.
Don't confuse the two:
- The agent loop (
core.py) is what one agent does within a single task — its own tool calls, its own back-and-forth with the model. - Orchestration (
orchestrator.py) is what happens above that — how many agents you spin up, over what units of work, sequential or parallel, and what you do with each one's result.
orchestrator.py now has two strategies:
One fresh Agent per PR, run one at a time. Simplest to debug — one
transcript at a time, nothing running concurrently to reason about.
This is where the earlier "sequential because gh pr checkout mutates
the one shared directory" constraint gets lifted properly, instead of
worked around. _create_worktree() does, per PR:
git fetch origin pull/<pr_number>/head:pr-agent/pr-<pr_number>
git worktree add .pr-agent-worktrees/pr-<pr_number> pr-agent/pr-<pr_number>— one repo, several working directories, each on its own branch, sharing
the same object store so you aren't cloning the repo N times. Each
worktree gets its own Agent, and concurrent.futures.ThreadPoolExecutor
runs them at the same time. Worktrees are torn down in a finally block
so a crashed agent doesn't leave one orphaned and blocking a re-run.
The refactor this required, and why it's the actual lesson:
tools.py used to read REPO_DIR = Path.cwd() as a module-level
constant — completely fine with one agent running at a time, silently
broken the moment two agents in two threads both call checkout_pr and
both expect to be "the" working directory. build_tool_impls(repo_dir)
now builds a fresh dict of closures bound to one specific directory, so
each Agent gets tools that only ever touch its own worktree. This is
the same reason you don't keep request state in module-level globals in
a web server — concurrency turns a convenience into a bug, and the fix is
always the same: stop relying on ambient/global state, pass the context
in explicitly instead.
run_all_parallel() also defaults verbose=False. Several agents
streaming tokens into one terminal at once, even behind a shared print
lock with per-agent [PR#N] prefixes, is hard to follow — good for one
debug run, not for routine use. A per-PR log file instead of stdout is
the natural next step if you want live output and readability at scale.
pip install -r requirements.txt
# clone a scratch copy — the agent switches branches in this checkout
git clone https://github.com/alvincrespo/glypto.git /tmp/glypto-agent
cd /tmp/glypto-agent
export INFERENCE_PROVIDER_TOKEN=hf_...
export PR_AGENT_DRY_RUN=1 # first runs: does everything except push/merge
python /path/to/pr-agent/main.py # sequential, streamed
python /path/to/pr-agent/main.py --parallel # worktrees + thread pool
python /path/to/pr-agent/main.py --parallel --workers 8Sequential mode streams tokens live as the model reasons — that stream
is the agent's reasoning trace, there's no hidden state anywhere else.
Either mode ends with a summary built from each PR's structured
PROutcome, not parsed prose.
Once you trust it on a couple of low-stakes repos, drop
PR_AGENT_DRY_RUN and let it actually push fixes and merge.
Not in the system prompt — in the code:
run_diagnostic_commandchecks a command allowlist before running anything (tools.py). The model can ask forrm -rf /; the tool implementation refuses it regardless of what the prompt says.commit_and_push/merge_prare the only two tools that touch upstream state, and they're the only two gated byPR_AGENT_DRY_RUN.max_turnsincore.pybounds worst-case cost and stops runaway fix-retry loops, per PR.- Worktree teardown runs in a
finallyblock, so a crashed agent still releases its checkout instead of leaving.pr-agent-worktrees/in a state that blocks the next run. - The system prompt's "stop after two failed fix attempts" is a
policy, and policies are suggestions to the model — they reduce bad
outcomes, they don't guarantee against them. Anything you actually
can't afford to have go wrong belongs in a tool-level check like the
ones above, not in prose.
PROutcome.statusis a report of what the model believes happened, not a safety mechanism — trust the dry-run logs and git history over the summary line if they ever disagree.
- A second agent type — add a
prompts.pyvariant for "triage only, never merge" and reuse the samecore.py, most oftools.py, and a newschemas.pyoutcome type. This is the moment the split really starts paying for itself. - Per-PR log files instead of interleaved stdout, so
--parallelstays legible withverbose=True. strict: trueon the tool schemas — the same structured-outputs system that gives youPROutcomecan also guaranteemerge_pr'smethodargument is always one of the three valid strings, with grammar-constrained sampling instead of relying on the model reading theenumin the schema description.- Compare to Claude Code's Dynamic Workflows — once this feels
familiar, look at what
/workflowsgenerates for the same task. You'll recognizeagent()/pipeline()/parallel()as the same "one isolated worker per unit of work" pattern asorchestrator.py, running in a managed runtime instead of your ownThreadPoolExecutor.