Run Claude Code as an optimization loop that keeps working after you close your laptop.
You tell it what to improve and how to measure it. It then repeats one cycle, on its own, for hours or days:
change something → measure it → better? commit it. worse? undo it. → repeat
Every result is written to a file, so you can check on it whenever you like, and so the loop can pick up where it left off after a crash or a restart.
Typical things to point it at: test suite runtime, bundle size, a model's accuracy on an eval set, build times, p95 latency. Anything you can measure with a script that prints a number.
Based on pi-autoresearch, which is based on karpathy/autoresearch. This version is built for Claude Code.
Contents
- How it works
- Getting started
- Checking on a running loop
- Why not just tell Claude Code to keep going?
- Running it on a server
- Using tmux
- The
arcommand - Session files
- Writing a good benchmark
- Limits and cost
- Troubleshooting
- Tests
- Safety
- Uninstall
Two simple ideas.
1. The state lives in files, not in the chat.
Everything the loop knows sits in a folder called .auto/ in your project: what it is trying to
improve, every result so far, and notes it writes to itself. A brand new Claude Code session can
read those files and carry on from exactly where the last one stopped.
This matters because long conversations get summarised, and summaries lose detail. If the record lives in a file instead, nothing is lost.
2. A hook restarts the agent each time it finishes.
Claude Code lets you run a script whenever the agent finishes replying. That is called a Stop hook. This tool installs one. When the agent finishes a reply, the hook answers with something like:
Here is the current state: baseline 41.2, best 33.5, 12 runs so far, last three were reverted. Pick the next thing to try and measure it.
Claude Code then continues the conversation on its own. Nobody types anything. That is the whole loop — no background daemon, no separate process watching it.
When the loop should genuinely stop (target reached, budget spent), the hook stays quiet and the session ends normally.
Four steps. Claude Code does most of the work; you review it before letting it run unattended.
git clone https://github.com/rishabhpoddar/autoresearch-with-claude-code.git
cd autoresearch-with-claude-code
./install.sh /path/to/your/projectThis copies a small program into your-project/.auto/ar, adds three skills Claude Code can use,
and registers the hooks. It also tells git to ignore its own files, so nothing it creates ends up
in your commits.
Your project has to be a git repository. The loop is built on git: keeping a result makes a
commit, and rejecting one resets the files. Without a repo it could not undo anything, so
setup refuses to start rather than run in a state where rejected experiments quietly pile up. If
you are starting from a plain folder, run git init && git add -A && git commit -m initial before
step 2.
You have to name a project. There is no "install it everywhere" option, so the loop can never start in a project you did not set up.
The first time you open Claude Code in that folder it will ask whether you trust the project's hooks. Say yes.
cd /path/to/your/project
claudeThen type:
/autoresearch-create
It will ask you a few questions:
- What are you trying to improve?
- How do you measure it? (a command it can run)
- Which files may it change?
- Is anything off limits?
- What number would count as done?
Then it reads your code and writes three files for you:
| File | What it is |
|---|---|
.auto/prompt.md |
The brief. What to improve, the metric, which files it may touch, what to avoid. |
.auto/measure.sh |
The benchmark. A script that runs your workload and prints the score. |
.auto/checks.sh |
Optional. Tests that must keep passing, so a change cannot be kept if it breaks something. |
It finishes by recording the settings, running the benchmark once to get a starting score, and handing back to you.
You do not type any ar commands during any of this. Claude Code runs them — ar init to record
what is being optimized, then ar run and ar log for the first measurement. The only ar
commands you are likely to type yourself are status, dashboard, on and off.
If you have already written your goal down somewhere, skip the questions: "set up autoresearch using GOAL.md".
Spend five minutes here. measure.sh is a script that will run hundreds of times, on its own,
with your permissions. And the loop will improve whatever that script measures — so if the script
measures the wrong thing, you get a lot of very efficient progress in the wrong direction.
Run it once yourself:
bash .auto/measure.shThen ask five questions about it:
Is it measuring what you actually care about? The loop takes the number literally.
Does it fail loudly when something is broken? Say your benchmark calls a server and that server is down. If the script returns a bad score instead of an error, the loop will think your last change caused it and undo a perfectly good change. The script should exit with an error instead. This is the most common way an overnight run gets wasted.
Could the score be raised by cheating? If editing the benchmark itself, or the test data,
would improve the number, list those files under Off limits in .auto/prompt.md.
How long does one run take? Multiply that by 200. If the answer is unacceptable, make the benchmark smaller now — a sample is usually fine.
Does it print anything useful besides the score? Whatever else it prints goes back to the agent. Per-category results, error counts, or slow-step timings are what help it work out where to look next, rather than guessing.
There is more detail in Writing a good benchmark.
Open a fresh Claude Code session for the actual run:
tmux new -s ar
cd /path/to/your/project
claudeThen type:
/autoresearch-resume
It reads the brief, picks something to try, and starts. Press ctrl-b then d to leave it
running and get your terminal back.
Two reasons to start a new session rather than continue the setup one. The setup conversation is
full of questions, answers and file listings, and none of that helps the loop — you want a session
whose memory is the brief on disk. And tmux keeps the session alive when you disconnect or your
laptop sleeps; without it, closing the terminal kills the loop.
If you just want to try it out while watching, continuing in the same session is fine.
From any terminal, in the project folder:
.auto/ar statusA short summary: what it is optimizing, the starting score, the best score, how many runs, and the last several runs with the notes the agent wrote about each one.
.auto/ar dashboardThe full table — every run, whether it was kept or undone, the score, and the commit.
You do not need to attach to the session to run these. They read the files on disk.
To stop the loop:
.auto/ar off # stops after the current step; everything is preserved
.auto/ar on # start it againYou can, and it works for a while. Open Claude Code, describe the benchmark, and finish with "keep trying things until the score stops improving, and don't stop to ask me." For the first half hour there is not much difference.
Things drift after that, in five ways.
It stops. However you word the instruction, the agent eventually finishes its reply and the turn ends. You come back, type "continue", and get a few more experiments. Running overnight is not something you can fix by wording the prompt better — something outside the conversation has to start the next turn. That is what the hook does.
It forgets what it tried. Long conversations get summarised. The summary keeps the gist and drops the details: which twelve ideas were tried, which four were undone, and why. Later on it tries an old dead end again and pays the full cost. Here the record is a file, and a summary of that file is handed back on every single turn.
The score gets remembered instead of recorded. Left to itself, the agent reads the number off the screen, rounds it, and recalls it forty turns later to compare against a baseline it also half-remembers. No single slip is visible, and they only go one way. Here the score is pulled out of the output by a pattern match and written to a file.
Undo gets skipped. The problem is not usually "forgot to undo it". It is doing experiment 41 on top of experiment 40, which was never undone, and then not knowing which change produced the score. Here, keeping a change makes a commit and rejecting one resets the files. Every time, without the agent deciding.
Noise gets mistaken for progress. If your benchmark wobbles by ±0.02 between identical runs, an agent watching raw numbers will find plenty of imaginary improvements and keep them. This tool reports each improvement as a multiple of how much the benchmark normally wobbles, so "that is just noise" is visible rather than something to intuit.
During one run, a wrong setting caused 307 of 1,280 test cases to error out. The eval counted each failed case as the default answer and printed a believable score — it looked like a mediocre experiment, not a broken one. An agent keeping score in its head would write that off as "that idea didn't work" and move on, and everything after it would be measured against a bad number.
What caught it was having every case written to a file next to a recorded score: the number did not match the shape of the results. The fix became a commit, and a check was added so that kind of failure now stops the run instead of quietly scoring it.
- Short runs you are watching. Ten experiments over a coffee? Skip all of this. You will spot a bad number yourself.
- It does not make the agent cleverer. Every idea still comes from the model. If it cannot work out why your benchmark is stuck, this will help it reject forty wrong ideas very tidily. That has value, but it is not insight.
- Some of it is just discipline. "Commit wins, undo losses, keep notes" in a prompt works too — for a while. It is in code here because instructions fade as the conversation gets summarised and code does not.
- A bad benchmark gets worse, not better. Point this at a metric that can be gamed and you will get a very efficient search for ways to game it.
The short version: it does not make any single experiment better. It makes the fiftieth experiment as careful as the first, and it takes you out of the loop in between. If you plan to run ten experiments, skip it. If you plan to run a hundred overnight, it is the difference between results you can trust and a pile of edits you cannot explain.
active: true only means nothing told ar to stop. The session driving the loop can die without
ar ever hearing about it, and the dashboard then reports a healthy loop for hours.
The most common cause is a Claude Code safety cap: a Stop hook may block a turn from ending only
9 consecutive times, after which the harness force-ends the turn and the session drops to an idle
prompt. The autoresearch loop is a Stop hook that blocks every turn deliberately, so it trips that
cap. install.sh now writes
"env": { "CLAUDE_CODE_STOP_HOOK_BLOCK_CAP": "1000000" }into the project's .claude/settings.json. If you started a loop before this existed, export the
variable in the shell you launch claude from, or add it to that file by hand.
ar status prints last run: Nh ago and flags <-- STALE once nothing has been logged for two
hours while the loop claims to be active — check the session is still alive before assuming the
agent is merely thinking.
A laptop that sleeps stops the loop. For runs longer than a few hours, use a machine that stays on.
The setup below has two machines: a loop host that runs Claude Code and holds the repo, and optionally a worker it reaches over SSH (a GPU box, a build machine — whatever your benchmark needs). The loop host can be small; the work happens elsewhere. If your benchmark runs locally, ignore the worker parts.
It needs Node (for Claude Code), Python 3, git and tmux:
ssh -p <port> -i ~/.ssh/<key>.pem root@<loop-host>
npm install -g @anthropic-ai/claude-code
claude --version
tmux -V ; git --version ; python3 --versionIf the loop needs a cloud CLI, install it here too:
curl -sSL https://sdk.cloud.google.com > /tmp/gc.sh
bash /tmp/gc.sh --disable-prompts --install-dir=/root
echo 'export PATH=$PATH:/root/google-cloud-sdk/bin' >> ~/.bashrcOn a cloud VM the CLI often authenticates as the machine's own service account, so you may not need a key file at all. Check before copying one over.
If the loop host can reach your git remote, git clone is easiest. Otherwise send a tarball —
leaving out virtualenvs (they do not work on a different OS), large model files, and anything
secret:
tar -czf /tmp/transfer.tgz \
--exclude='*/venv' --exclude='*/venv-*' --exclude='__pycache__' --exclude='.DS_Store' \
--exclude='.env' --exclude='*.pem' --exclude='*service-account*.json' \
.git <folder-you-need> <other-folder-you-need>
scp -P <port> -i ~/.ssh/<key>.pem /tmp/transfer.tgz root@<loop-host>:/root/
ssh -p <port> -i ~/.ssh/<key>.pem root@<loop-host> \
'mkdir -p /root/<repo> && cd /root/<repo> && tar -xzf /root/transfer.tgz'Three things go wrong almost every time:
- File ownership. tar keeps your local user id, so git says "detected dubious ownership".
Fix it with
chown -R root:root /root/<repo>andgit config --global --add safe.directory /root/<repo>. - Mac metadata files. A tarball made on a Mac contains
._*files that git will happily commit. Delete them:find . -name '._*' -type f -exec unlink {} \; - Only part of the repo. If you sent some folders and not others, git thinks the missing ones
were deleted — and keeping a result runs
git add -A, which would commit those deletions. Tell git to only track what you sent:git sparse-checkout init --cone git sparse-checkout set <folder-you-need> <other-folder-you-need> git restore -- <any-top-level-files-you-skipped> git status --short # should show nothing unexpected
Then install and check the tree is clean:
cd /root/<repo>/autoresearch && bash install.sh /root/<repo>/<project>
cd /root/<repo>/<project> && git status --short # blank means readyMake a key on the loop host rather than copying your personal one over, so you can revoke it separately:
# on the loop host
cat ~/.ssh/id_ed25519.pub # run ssh-keygen -t ed25519 first if it does not existAdd that line to the worker's ~/.ssh/authorized_keys, then give it a short name so the agent can
just write ssh gpu:
# loop host ~/.ssh/config
Host gpu
HostName <worker-ip>
Port <port>
User root
IdentityFile /root/.ssh/id_ed25519
StrictHostKeyChecking no
ServerAliveInterval 30
Check it works: ssh gpu hostname.
If both machines are on the same cloud network, point the benchmark at the worker's internal IP. It is faster, free, and not affected by firewall rules.
Copy only what the benchmark needs, chmod 600 it, and make sure git is ignoring it. Then test it
before starting the loop. A credential that fails on experiment 30 looks exactly like a
regression, and the agent will treat it as one.
cd /root/<repo>/<project>
bash .auto/measure.sh # should print METRIC lines, or fail with a clear error
.auto/ar status # "no autoresearch session" is correct before the first run
git status --short # blankClaude Code is an interactive program, so it dies when your SSH connection drops — unless it is running inside tmux. For a long run this is not optional.
tmux new -s ar # start a session called "ar"
tmux attach -t ar # come back to it later, from any machine
tmux ls # list sessions
tmux kill-session -t ar # end itInside tmux you press ctrl-b first, then a key:
| Keys | What it does |
|---|---|
ctrl-b then d |
Detach. The loop keeps running and you get your shell back. |
ctrl-b then [ |
Scroll back through output. Press q to stop scrolling. |
ctrl-b then c |
Open a second window, handy for running ar status. |
ctrl-b then n / p |
Next / previous window. |
Normal routine:
ssh -p <port> -i ~/.ssh/<key>.pem root@<loop-host>
tmux attach -t ar || tmux new -s ar
cd /root/<repo>/<project> && IS_SANDBOX=1 claude --dangerously-skip-permissions
# start the loop, then ctrl-b d and log offTo look in on it without attaching, either run .auto/ar status in a second SSH session, or read
the screen directly:
tmux capture-pane -t ar -p | tail -30Reading from disk is safer than attaching, because it is easy to type into an attached session by accident.
.auto/ar is the small program that runs the loop. It is plain Python with no dependencies. Run
it from your project folder.
.auto/ar status # short summary + recent runs and the agent's notes
.auto/ar dashboard # full table of every run.auto/ar off # stop the loop; all results are kept
.auto/ar on # start it againYou do not type these. Claude Code runs them during /autoresearch-create and then on every
iteration of the loop. They are listed so you can read the log and understand what happened:
| Command | What it does |
|---|---|
ar init --name N --metric M [--direction lower|higher] [--target T] |
Starts a session. Records the settings and turns the loop on. Does not run anything. |
ar run |
Runs .auto/measure.sh, times it, reads the METRIC lines, then runs .auto/checks.sh if it exists. |
ar log --status keep|discard|crash|checks_failed --desc "..." [--asi k=v] |
Records the result. keep commits the change; anything else undoes it. |
ar hook-stop, ar hook-pretool |
The two hook entry points. Claude Code calls them; you never do. |
ar --version prints the version, which is worth knowing when the same file has been copied into
several projects and one of them is stale.
--metric has to match the name printed by the benchmark exactly. --direction says whether
bigger or smaller is better. --target is a score that means "done" — leave it out to run until
the budget runs out.
Two things ar log will refuse, because a result that is not true is worse than no result:
- Recording the same measurement twice. After a result is logged, its measurement is marked as
used, so logging again without running the benchmark is refused rather than quietly repeating
the previous number. Pass
--metricif you measured it some other way. - Keeping a run that crashed or failed its checks. You can record it honestly as
crashorchecks_failed; you cannot record it as a win.
Two rules the loop depends on:
ar run is the only way to measure. A benchmark run by hand produces a number that exists
nowhere, so the log, the noise estimate and the next agent never see it.
--asi notes matter most on failures. When a result is rejected the code is undone and gone.
The note is the only surviving record that the idea was tried, which is what stops the loop
rediscovering the same dead end fifty experiments later.
After three runs, ar log also reports how big the best improvement is compared with how much the
benchmark normally varies. Above 2× is probably real; below 1× is probably noise. It is advice
only — it never rejects anything by itself.
Everything lives in <project>/.auto/, which the installer tells git to ignore.
| File | What it is |
|---|---|
prompt.md |
The brief: goal, metric, which files are in scope, what to avoid, what has been tried. A new session can run the whole loop from this alone. |
measure.sh |
The benchmark. Prints METRIC name=value. |
log.jsonl |
One line per run, including the agent's own notes. |
ideas.md |
Ideas it wants to come back to. |
checks.sh |
Optional. Tests that must pass for a result to be kept. |
config.json |
Optional. Limits — see below. |
ar |
The program itself. |
runtime.json, last_run.json, results/ |
Internal state. You can ignore these. |
/autoresearch-create writes measure.sh for you, so most people read this to check its work.
Write one by hand if you prefer.
There are only three rules:
- Exit 0 if the run was valid. Any other exit code means "something broke", and the result is thrown away rather than recorded as a bad score.
- Print
METRIC <name>=<number>. The one whose name matchesar init --metricis the score. Any others are recorded too, as things to keep an eye on. - Anything else you print is context for the agent. The last 60 lines are shown to it, so print whatever helps it decide what to try next.
#!/bin/bash
set -euo pipefail
cd "$(dirname "$0")/.."
[ -f .auto/env.sh ] && . .auto/env.sh
# 1. Check the environment first. If something is down, fail here rather than
# returning a bad score - otherwise a good change gets undone for no reason.
curl -sf --max-time 10 "$SERVICE/health" >/dev/null || { echo "service down" >&2; exit 1; }
# 2. A quick sanity check, so a broken edit fails in a second instead of an hour.
python3 -c "import mymodule" >/dev/null
# 3. Run the workload. If it takes under 5 seconds, run it a few times and take the median.
./run_benchmark.sh > /tmp/out.json
# 4. Print the score, plus anything useful.
python3 -c "
import json; r = json.load(open('/tmp/out.json'))
print(f\"METRIC score={r['score']:.4f}\")
print(f\"METRIC latency_p95={r['p95']:.0f}\")
for name, v in sorted(r['per_category'].items())[:5]: print(f'worst: {name} {v:.3f}')
"Things that make a real difference:
- Speed. Whatever it takes, multiply by 200. Ninety seconds becomes five hours of waiting.
- Consistency. If the same code gives 41 then 47, the loop will "find" improvements that are not there. For quick benchmarks, run the workload several times and print the median.
- Failing loudly. Covered above, and worth repeating: it is the most common cause of a wasted overnight run.
- Being hard to cheat. Whatever raises the number will get raised. If editing the benchmark or
the test data would do it, list those files under Off limits in
prompt.md.
The loop stops by itself under any of these conditions. Change them in .auto/config.json:
{
"maxIterations": 40,
"maxConsecutiveFailures": 20,
"autoResumeTurnLimit": 500,
"maxStalls": 6,
"runTimeoutSeconds": 3600,
"checksTimeoutSeconds": 600
}| Setting | Default | What it means |
|---|---|---|
| target reached | — | The score passed --target. This is the intended ending. |
maxIterations |
0 (no limit) |
Maximum number of experiments. |
maxConsecutiveFailures |
20 |
This many rejected results in a row means it has run out of ideas. |
autoResumeTurnLimit |
500 |
A final backstop. |
maxStalls |
6 |
Turns where nothing at all changed — no result, no commit, no edited file. |
These are read fresh every time, so you can change them while the loop is running. No restart needed.
maxStalls is there to catch an agent that talks instead of working. Without it, a reply with no
work in it gets a "carry on" from the hook, which produces another reply with no work in it, and
so on. Note it counts any change as progress — editing a file or making a commit both count —
so time spent debugging does not trip it.
Cost is the real limit on a long run. Set maxIterations before you leave it alone, and set a
spending cap on your account.
It does one experiment and stops. The hook is not firing. Check that
.claude/settings.json has a Stop entry pointing at this project's .auto/ar, and that
.auto/runtime.json says "active": true. Test it directly:
echo '{"cwd":"'$PWD'"}' | python3 .auto/ar hook-stopYou should get JSON back containing additionalContext. No output means the loop is deliberately
allowing it to stop — run .auto/ar status to see why.
"Hook JSON output validation failed." For a Stop hook, Claude Code accepts only
hookEventName and additionalContext inside hookSpecificOutput. There is no decision field.
"Loop paused — no run, no commit and no file change." The agent replied several times without
doing anything. .auto/ar on starts it again. If it keeps happening, the brief in prompt.md is
probably unclear about what to do next.
The program disappears mid-run. .auto/ was committed to git, and something reset the repo to
an earlier commit. Re-run install.sh — it tells git to ignore those files — and check that
git status --short shows nothing from .auto/ or .claude/.
Every experiment fails. Something in the environment is broken rather than the code. Run
bash .auto/measure.sh yourself and read the error.
"not a git repository." The loop needs git to keep and undo changes. Run
git init && git add -A && git commit -m initial in your project, then run /autoresearch-create
again.
"REVERT FAILED." A rejected result could not be undone, so the change is still in your working tree and nothing was logged. Fix the tree by hand before running anything else — otherwise the next experiment measures both changes together and the result is attributed to the wrong one.
pkill -f "something" over SSH kills your own session. Your command contains the text you are
searching for, so it matches itself. Write pkill -f "some[t]hing" instead.
Copies drift apart. Once this is installed on more than one machine, the same file exists in
several places. Treat the machine running the loop as the real one and copy from it. If the
agent has edited its own checks.sh there, overwriting it from your laptop throws away real work.
./test.sh # everything
./test.sh test_hooks # one module127 tests covering metric parsing and the noise floor, the full init / run / log cycle
against real git repos, both hooks including every guard and every bypass the off-switch blocks,
and install.sh install/uninstall round trips. No dependencies beyond python3, bash and
git.
The loop edits files, makes commits, undoes changes, and runs .auto/measure.sh and
.auto/checks.sh unattended, with your permissions. So:
- Run it on its own branch, starting from a clean tree.
- Read
measure.shandchecks.shbefore each session. They are code that will run hundreds of times. - Keep credentials out of the files listed in
prompt.md, and list secret paths under Off limits.
One specific thing to know: rejecting a result runs git restore and git clean -fd. That
deletes untracked files that git is not already ignoring. Ignored files (caches, model weights)
are safe, and so is .auto/ — but anything else you left lying around untracked is not.
./install.sh --uninstall /path/to/your/projectRemoves the skills, the program, its hooks, and the git-ignore entries it added. Your own hooks and settings are left alone.
Your results stay: log.jsonl, prompt.md, measure.sh and everything else in .auto/. Delete
that folder yourself when you no longer want them.