Skip to content

Repository files navigation

pr-agent — a hand-built agent for merging pull requests

Built from the raw OpenAI-compatible chat-completions API (no Agent SDK, no framework) so every piece of "agent orchestration" is visible and editable. Talks to Hugging Face's Inference Providers router by default — any provider speaking the same wire format (Fireworks, a Modal-hosted vLLM server, a local Ollama/vLLM instance) works by changing one environment variable, not the code. This isn't meant to replace Claude Code — it's meant to teach you what Claude Code (and every other agent product) is actually doing underneath.

The five pieces, and why they're separate files

File Answers
agent/core.py What is an agent loop, mechanically?
agent/tools.py What is this agent allowed to do?
agent/prompts.py What should it do with those tools? (your policy)
agent/schemas.py What shape should its final answer take?
agent/orchestrator.py How many agents run, over what, in what order?

That split is the actual lesson. core.py doesn't know what a PR is. tools.py doesn't know when to merge vs. fix. prompts.py doesn't know how to call the model's API. Each piece is replaceable independently — that's what "build an agent for a product" means in practice: you're almost always recombining these five kinds of things, not inventing a new kind of loop each time.

The agent loop, in one paragraph

Every "agent" you'll ever build is: send the conversation + tool schemas to the model → if it asks for a tool, run it locally and append the result as a new message → repeat → stop when the model stops asking for tools (or you hit a turn cap). Read Agent.run() in core.py top to bottom — the loop itself is still about 25 lines once you separate it from streaming and structured-output logic. Everything else in this repo exists to configure that loop, not to extend it.

Streaming

Each turn is sent with stream=True — see Agent._send_turn(). Text prints token-by-token as delta.content arrives, and the message is reassembled from chunks into an object with .content and .tool_calls, so the rest of the loop doesn't need to know a turn was streamed at all.

The part worth knowing: under the OpenAI-compatible format, a tool call's arguments string arrives fragmented across multiple chunks, identified by an index field — not delivered whole the way a single-vendor SDK's stream helper might hand it back. _send_turn() accumulates fragments by index and concatenates each one's arguments piece until the stream ends. This is the single most fragile-feeling part of the whole file, and it's covered by a dedicated test with a deliberately multi-chunk-split tool call — see the "Implement streaming turns" story in ISSUES.md for why that test exists.

Structured final output

agent/schemas.py defines PROutcome — a Pydantic model with pr_number, status (an enum: merged / fix_pushed_pending_ci / needs_human_review), and a one-line summary.

Once the tool-use loop ends, Agent._extract_structured_outcome() makes one more, separate, tool-free call with response_format={"type": "json_schema", "json_schema": {...}}, asking the model to restate the outcome it already reached, constrained to that schema. orchestrator.py's summary now reads result.outcome.status directly instead of string-matching free text.

Why a separate call instead of attaching the schema to every turn: a turn where the model still needs to call merge_pr or read_file isn't supposed to produce a final JSON answer yet, and forcing a structured response format onto those turns fights the tool-use turns instead of complementing them. One clean extraction call after the loop naturally ends is simpler to reason about — and cheap, since the conversation it's built from is already sitting in the prompt cache from the turns that just ran.

Real caveat, not a hypothetical one: Hugging Face's router fans out to many different backend providers and models, and strict JSON-schema adherence isn't guaranteed equally across all of them the way it would be against a single vendor's own models. _extract_structured_outcome() wraps the whole attempt in a try/except for exactly this reason — a provider that doesn't honor strict mode degrades to outcome=None (with final_text still populated) rather than crashing the run. Verify this against whatever model you actually settle on; don't assume it from one provider's behavior generalizing to another.

Provider — Hugging Face by default, anything OpenAI-compatible in general

Two environment variables control where requests go:

  • INFERENCE_PROVIDER_TOKEN — who you are to the provider
  • INFERENCE_PROVIDER_BASE_URL — which provider, defaulting to https://router.huggingface.co/v1 if unset

Any provider speaking the OpenAI chat-completions wire format works by changing INFERENCE_PROVIDER_BASE_URL alone — Fireworks, a Modal-hosted vLLM server, a local Ollama or vLLM instance. This does not extend to Anthropic's own API — it's a genuinely different shape (different tool-call format, different streaming, different structured-output mechanism), not a base_url swap. Supporting both would mean a real adapter layer — a ModelBackend interface with one implementation per wire format — which this project deliberately doesn't build until there's an actual need for it.

Orchestration vs. the loop

Don't confuse the two:

  • The agent loop (core.py) is what one agent does within a single task — its own tool calls, its own back-and-forth with the model.
  • Orchestration (orchestrator.py) is what happens above that — how many agents you spin up, over what units of work, sequential or parallel, and what you do with each one's result.

orchestrator.py now has two strategies:

run_all() — sequential, one shared working directory

One fresh Agent per PR, run one at a time. Simplest to debug — one transcript at a time, nothing running concurrently to reason about.

run_all_parallel() — one git worktree per PR, run concurrently

This is where the earlier "sequential because gh pr checkout mutates the one shared directory" constraint gets lifted properly, instead of worked around. _create_worktree() does, per PR:

git fetch origin pull/<pr_number>/head:pr-agent/pr-<pr_number>
git worktree add .pr-agent-worktrees/pr-<pr_number> pr-agent/pr-<pr_number>

— one repo, several working directories, each on its own branch, sharing the same object store so you aren't cloning the repo N times. Each worktree gets its own Agent, and concurrent.futures.ThreadPoolExecutor runs them at the same time. Worktrees are torn down in a finally block so a crashed agent doesn't leave one orphaned and blocking a re-run.

The refactor this required, and why it's the actual lesson: tools.py used to read REPO_DIR = Path.cwd() as a module-level constant — completely fine with one agent running at a time, silently broken the moment two agents in two threads both call checkout_pr and both expect to be "the" working directory. build_tool_impls(repo_dir) now builds a fresh dict of closures bound to one specific directory, so each Agent gets tools that only ever touch its own worktree. This is the same reason you don't keep request state in module-level globals in a web server — concurrency turns a convenience into a bug, and the fix is always the same: stop relying on ambient/global state, pass the context in explicitly instead.

run_all_parallel() also defaults verbose=False. Several agents streaming tokens into one terminal at once, even behind a shared print lock with per-agent [PR#N] prefixes, is hard to follow — good for one debug run, not for routine use. A per-PR log file instead of stdout is the natural next step if you want live output and readability at scale.

Run it

pip install -r requirements.txt

# clone a scratch copy — the agent switches branches in this checkout
git clone https://github.com/alvincrespo/glypto.git /tmp/glypto-agent
cd /tmp/glypto-agent

export INFERENCE_PROVIDER_TOKEN=hf_...
export PR_AGENT_DRY_RUN=1   # first runs: does everything except push/merge

python /path/to/pr-agent/main.py                    # sequential, streamed
python /path/to/pr-agent/main.py --parallel          # worktrees + thread pool
python /path/to/pr-agent/main.py --parallel --workers 8

Sequential mode streams tokens live as the model reasons — that stream is the agent's reasoning trace, there's no hidden state anywhere else. Either mode ends with a summary built from each PR's structured PROutcome, not parsed prose.

Once you trust it on a couple of low-stakes repos, drop PR_AGENT_DRY_RUN and let it actually push fixes and merge.

Where the safety actually lives

Not in the system prompt — in the code:

  • run_diagnostic_command checks a command allowlist before running anything (tools.py). The model can ask for rm -rf /; the tool implementation refuses it regardless of what the prompt says.
  • commit_and_push / merge_pr are the only two tools that touch upstream state, and they're the only two gated by PR_AGENT_DRY_RUN.
  • max_turns in core.py bounds worst-case cost and stops runaway fix-retry loops, per PR.
  • Worktree teardown runs in a finally block, so a crashed agent still releases its checkout instead of leaving .pr-agent-worktrees/ in a state that blocks the next run.
  • The system prompt's "stop after two failed fix attempts" is a policy, and policies are suggestions to the model — they reduce bad outcomes, they don't guarantee against them. Anything you actually can't afford to have go wrong belongs in a tool-level check like the ones above, not in prose. PROutcome.status is a report of what the model believes happened, not a safety mechanism — trust the dry-run logs and git history over the summary line if they ever disagree.

Extending it further

  • A second agent type — add a prompts.py variant for "triage only, never merge" and reuse the same core.py, most of tools.py, and a new schemas.py outcome type. This is the moment the split really starts paying for itself.
  • Per-PR log files instead of interleaved stdout, so --parallel stays legible with verbose=True.
  • strict: true on the tool schemas — the same structured-outputs system that gives you PROutcome can also guarantee merge_pr's method argument is always one of the three valid strings, with grammar-constrained sampling instead of relying on the model reading the enum in the schema description.
  • Compare to Claude Code's Dynamic Workflows — once this feels familiar, look at what /workflows generates for the same task. You'll recognize agent()/pipeline()/parallel() as the same "one isolated worker per unit of work" pattern as orchestrator.py, running in a managed runtime instead of your own ThreadPoolExecutor.

About

A custom agent that reviews, fixes, and merges pull requests. Built to leverage Hugging Face's model inference providers for agent orchestration fundamentals.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages