Skip to content

About

Autoresearch framework to optimise anything (ML model, package size, passing tests, build time)

Resources

Stars

2 stars

Watchers

0 watching

Forks

Latest commit

 

History

13 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

autoresearch

Run Claude Code as an optimization loop that keeps working after you close your laptop.

You tell it what to improve and how to measure it. It then repeats one cycle, on its own, for hours or days:

change something  →  measure it  →  better? commit it.  worse? undo it.  →  repeat

Every result is written to a file, so you can check on it whenever you like, and so the loop can pick up where it left off after a crash or a restart.

Typical things to point it at: test suite runtime, bundle size, a model's accuracy on an eval set, build times, p95 latency. Anything you can measure with a script that prints a number.

Based on pi-autoresearch, which is based on karpathy/autoresearch. This version is built for Claude Code.

Contents


How it works

Two simple ideas.

1. The state lives in files, not in the chat.

Everything the loop knows sits in a folder called .auto/ in your project: what it is trying to improve, every result so far, and notes it writes to itself. A brand new Claude Code session can read those files and carry on from exactly where the last one stopped.

This matters because long conversations get summarised, and summaries lose detail. If the record lives in a file instead, nothing is lost.

2. A hook restarts the agent each time it finishes.

Claude Code lets you run a script whenever the agent finishes replying. That is called a Stop hook. This tool installs one. When the agent finishes a reply, the hook answers with something like:

Here is the current state: baseline 41.2, best 33.5, 12 runs so far, last three were reverted. Pick the next thing to try and measure it.

Claude Code then continues the conversation on its own. Nobody types anything. That is the whole loop — no background daemon, no separate process watching it.

When the loop should genuinely stop (target reached, budget spent), the hook stays quiet and the session ends normally.


Getting started

Four steps. Claude Code does most of the work; you review it before letting it run unattended.

1. Install it into your project

git clone https://github.com/rishabhpoddar/autoresearch-with-claude-code.git
cd autoresearch-with-claude-code
./install.sh /path/to/your/project

This copies a small program into your-project/.auto/ar, adds three skills Claude Code can use, and registers the hooks. It also tells git to ignore its own files, so nothing it creates ends up in your commits.

Your project has to be a git repository. The loop is built on git: keeping a result makes a commit, and rejecting one resets the files. Without a repo it could not undo anything, so setup refuses to start rather than run in a state where rejected experiments quietly pile up. If you are starting from a plain folder, run git init && git add -A && git commit -m initial before step 2.

You have to name a project. There is no "install it everywhere" option, so the loop can never start in a project you did not set up.

The first time you open Claude Code in that folder it will ask whether you trust the project's hooks. Say yes.

2. Ask Claude Code to set it up

cd /path/to/your/project
claude

Then type:

/autoresearch-create

It will ask you a few questions:

  • What are you trying to improve?
  • How do you measure it? (a command it can run)
  • Which files may it change?
  • Is anything off limits?
  • What number would count as done?

Then it reads your code and writes three files for you:

File What it is
.auto/prompt.md The brief. What to improve, the metric, which files it may touch, what to avoid.
.auto/measure.sh The benchmark. A script that runs your workload and prints the score.
.auto/checks.sh Optional. Tests that must keep passing, so a change cannot be kept if it breaks something.

It finishes by recording the settings, running the benchmark once to get a starting score, and handing back to you.

You do not type any ar commands during any of this. Claude Code runs them — ar init to record what is being optimized, then ar run and ar log for the first measurement. The only ar commands you are likely to type yourself are status, dashboard, on and off.

If you have already written your goal down somewhere, skip the questions: "set up autoresearch using GOAL.md".

3. Read what it wrote

Spend five minutes here. measure.sh is a script that will run hundreds of times, on its own, with your permissions. And the loop will improve whatever that script measures — so if the script measures the wrong thing, you get a lot of very efficient progress in the wrong direction.

Run it once yourself:

bash .auto/measure.sh

Then ask five questions about it:

Is it measuring what you actually care about? The loop takes the number literally.

Does it fail loudly when something is broken? Say your benchmark calls a server and that server is down. If the script returns a bad score instead of an error, the loop will think your last change caused it and undo a perfectly good change. The script should exit with an error instead. This is the most common way an overnight run gets wasted.

Could the score be raised by cheating? If editing the benchmark itself, or the test data, would improve the number, list those files under Off limits in .auto/prompt.md.

How long does one run take? Multiply that by 200. If the answer is unacceptable, make the benchmark smaller now — a sample is usually fine.

Does it print anything useful besides the score? Whatever else it prints goes back to the agent. Per-category results, error counts, or slow-step timings are what help it work out where to look next, rather than guessing.

There is more detail in Writing a good benchmark.

4. Start the loop in a new session

Open a fresh Claude Code session for the actual run:

tmux new -s ar
cd /path/to/your/project
claude

Then type:

/autoresearch-resume

It reads the brief, picks something to try, and starts. Press ctrl-b then d to leave it running and get your terminal back.

Two reasons to start a new session rather than continue the setup one. The setup conversation is full of questions, answers and file listings, and none of that helps the loop — you want a session whose memory is the brief on disk. And tmux keeps the session alive when you disconnect or your laptop sleeps; without it, closing the terminal kills the loop.

If you just want to try it out while watching, continuing in the same session is fine.


Checking on a running loop

From any terminal, in the project folder:

.auto/ar status

A short summary: what it is optimizing, the starting score, the best score, how many runs, and the last several runs with the notes the agent wrote about each one.

.auto/ar dashboard

The full table — every run, whether it was kept or undone, the score, and the commit.

You do not need to attach to the session to run these. They read the files on disk.

To stop the loop:

.auto/ar off      # stops after the current step; everything is preserved
.auto/ar on       # start it again

Why not just tell Claude Code to keep going?

You can, and it works for a while. Open Claude Code, describe the benchmark, and finish with "keep trying things until the score stops improving, and don't stop to ask me." For the first half hour there is not much difference.

Things drift after that, in five ways.

It stops. However you word the instruction, the agent eventually finishes its reply and the turn ends. You come back, type "continue", and get a few more experiments. Running overnight is not something you can fix by wording the prompt better — something outside the conversation has to start the next turn. That is what the hook does.

It forgets what it tried. Long conversations get summarised. The summary keeps the gist and drops the details: which twelve ideas were tried, which four were undone, and why. Later on it tries an old dead end again and pays the full cost. Here the record is a file, and a summary of that file is handed back on every single turn.

The score gets remembered instead of recorded. Left to itself, the agent reads the number off the screen, rounds it, and recalls it forty turns later to compare against a baseline it also half-remembers. No single slip is visible, and they only go one way. Here the score is pulled out of the output by a pattern match and written to a file.

Undo gets skipped. The problem is not usually "forgot to undo it". It is doing experiment 41 on top of experiment 40, which was never undone, and then not knowing which change produced the score. Here, keeping a change makes a commit and rejecting one resets the files. Every time, without the agent deciding.

Noise gets mistaken for progress. If your benchmark wobbles by ±0.02 between identical runs, an agent watching raw numbers will find plenty of imaginary improvements and keep them. This tool reports each improvement as a multiple of how much the benchmark normally wobbles, so "that is just noise" is visible rather than something to intuit.

A real example

During one run, a wrong setting caused 307 of 1,280 test cases to error out. The eval counted each failed case as the default answer and printed a believable score — it looked like a mediocre experiment, not a broken one. An agent keeping score in its head would write that off as "that idea didn't work" and move on, and everything after it would be measured against a bad number.

What caught it was having every case written to a file next to a recorded score: the number did not match the shape of the results. The fix became a commit, and a check was added so that kind of failure now stops the run instead of quietly scoring it.

When it is not worth it

  • Short runs you are watching. Ten experiments over a coffee? Skip all of this. You will spot a bad number yourself.
  • It does not make the agent cleverer. Every idea still comes from the model. If it cannot work out why your benchmark is stuck, this will help it reject forty wrong ideas very tidily. That has value, but it is not insight.
  • Some of it is just discipline. "Commit wins, undo losses, keep notes" in a prompt works too — for a while. It is in code here because instructions fade as the conversation gets summarised and code does not.
  • A bad benchmark gets worse, not better. Point this at a metric that can be gamed and you will get a very efficient search for ways to game it.

The short version: it does not make any single experiment better. It makes the fiftieth experiment as careful as the first, and it takes you out of the loop in between. If you plan to run ten experiments, skip it. If you plan to run a hundred overnight, it is the difference between results you can trust and a pile of edits you cannot explain.


The loop stopped and ar status still says active

active: true only means nothing told ar to stop. The session driving the loop can die without ar ever hearing about it, and the dashboard then reports a healthy loop for hours.

The most common cause is a Claude Code safety cap: a Stop hook may block a turn from ending only 9 consecutive times, after which the harness force-ends the turn and the session drops to an idle prompt. The autoresearch loop is a Stop hook that blocks every turn deliberately, so it trips that cap. install.sh now writes

"env": { "CLAUDE_CODE_STOP_HOOK_BLOCK_CAP": "1000000" }

into the project's .claude/settings.json. If you started a loop before this existed, export the variable in the shell you launch claude from, or add it to that file by hand.

ar status prints last run: Nh ago and flags <-- STALE once nothing has been logged for two hours while the loop claims to be active — check the session is still alive before assuming the agent is merely thinking.

Running it on a server

A laptop that sleeps stops the loop. For runs longer than a few hours, use a machine that stays on.

The setup below has two machines: a loop host that runs Claude Code and holds the repo, and optionally a worker it reaches over SSH (a GPU box, a build machine — whatever your benchmark needs). The loop host can be small; the work happens elsewhere. If your benchmark runs locally, ignore the worker parts.

1. Set up the loop host

It needs Node (for Claude Code), Python 3, git and tmux:

ssh -p <port> -i ~/.ssh/<key>.pem root@<loop-host>
npm install -g @anthropic-ai/claude-code
claude --version
tmux -V ; git --version ; python3 --version

If the loop needs a cloud CLI, install it here too:

curl -sSL https://sdk.cloud.google.com > /tmp/gc.sh
bash /tmp/gc.sh --disable-prompts --install-dir=/root
echo 'export PATH=$PATH:/root/google-cloud-sdk/bin' >> ~/.bashrc

On a cloud VM the CLI often authenticates as the machine's own service account, so you may not need a key file at all. Check before copying one over.

2. Get your code onto it

If the loop host can reach your git remote, git clone is easiest. Otherwise send a tarball — leaving out virtualenvs (they do not work on a different OS), large model files, and anything secret:

tar -czf /tmp/transfer.tgz \
  --exclude='*/venv' --exclude='*/venv-*' --exclude='__pycache__' --exclude='.DS_Store' \
  --exclude='.env' --exclude='*.pem' --exclude='*service-account*.json' \
  .git <folder-you-need> <other-folder-you-need>

scp -P <port> -i ~/.ssh/<key>.pem /tmp/transfer.tgz root@<loop-host>:/root/
ssh -p <port> -i ~/.ssh/<key>.pem root@<loop-host> \
  'mkdir -p /root/<repo> && cd /root/<repo> && tar -xzf /root/transfer.tgz'

Three things go wrong almost every time:

  • File ownership. tar keeps your local user id, so git says "detected dubious ownership". Fix it with chown -R root:root /root/<repo> and git config --global --add safe.directory /root/<repo>.
  • Mac metadata files. A tarball made on a Mac contains ._* files that git will happily commit. Delete them: find . -name '._*' -type f -exec unlink {} \;
  • Only part of the repo. If you sent some folders and not others, git thinks the missing ones were deleted — and keeping a result runs git add -A, which would commit those deletions. Tell git to only track what you sent:
    git sparse-checkout init --cone
    git sparse-checkout set <folder-you-need> <other-folder-you-need>
    git restore -- <any-top-level-files-you-skipped>
    git status --short          # should show nothing unexpected

Then install and check the tree is clean:

cd /root/<repo>/autoresearch && bash install.sh /root/<repo>/<project>
cd /root/<repo>/<project> && git status --short     # blank means ready

3. Let the loop host reach the worker

Make a key on the loop host rather than copying your personal one over, so you can revoke it separately:

# on the loop host
cat ~/.ssh/id_ed25519.pub        # run ssh-keygen -t ed25519 first if it does not exist

Add that line to the worker's ~/.ssh/authorized_keys, then give it a short name so the agent can just write ssh gpu:

# loop host ~/.ssh/config
Host gpu
  HostName <worker-ip>
  Port <port>
  User root
  IdentityFile /root/.ssh/id_ed25519
  StrictHostKeyChecking no
  ServerAliveInterval 30

Check it works: ssh gpu hostname.

If both machines are on the same cloud network, point the benchmark at the worker's internal IP. It is faster, free, and not affected by firewall rules.

4. Credentials

Copy only what the benchmark needs, chmod 600 it, and make sure git is ignoring it. Then test it before starting the loop. A credential that fails on experiment 30 looks exactly like a regression, and the agent will treat it as one.

5. Check before you start

cd /root/<repo>/<project>
bash .auto/measure.sh          # should print METRIC lines, or fail with a clear error
.auto/ar status                # "no autoresearch session" is correct before the first run
git status --short             # blank

Using tmux

Claude Code is an interactive program, so it dies when your SSH connection drops — unless it is running inside tmux. For a long run this is not optional.

tmux new -s ar          # start a session called "ar"
tmux attach -t ar       # come back to it later, from any machine
tmux ls                 # list sessions
tmux kill-session -t ar # end it

Inside tmux you press ctrl-b first, then a key:

Keys What it does
ctrl-b then d Detach. The loop keeps running and you get your shell back.
ctrl-b then [ Scroll back through output. Press q to stop scrolling.
ctrl-b then c Open a second window, handy for running ar status.
ctrl-b then n / p Next / previous window.

Normal routine:

ssh -p <port> -i ~/.ssh/<key>.pem root@<loop-host>
tmux attach -t ar || tmux new -s ar
cd /root/<repo>/<project> && IS_SANDBOX=1 claude --dangerously-skip-permissions
# start the loop, then ctrl-b d and log off

To look in on it without attaching, either run .auto/ar status in a second SSH session, or read the screen directly:

tmux capture-pane -t ar -p | tail -30

Reading from disk is safer than attaching, because it is easy to type into an attached session by accident.


The ar command

.auto/ar is the small program that runs the loop. It is plain Python with no dependencies. Run it from your project folder.

The two you will actually use

.auto/ar status      # short summary + recent runs and the agent's notes
.auto/ar dashboard   # full table of every run

Starting and stopping

.auto/ar off      # stop the loop; all results are kept
.auto/ar on       # start it again

What Claude Code runs for you

You do not type these. Claude Code runs them during /autoresearch-create and then on every iteration of the loop. They are listed so you can read the log and understand what happened:

Command What it does
ar init --name N --metric M [--direction lower|higher] [--target T] Starts a session. Records the settings and turns the loop on. Does not run anything.
ar run Runs .auto/measure.sh, times it, reads the METRIC lines, then runs .auto/checks.sh if it exists.
ar log --status keep|discard|crash|checks_failed --desc "..." [--asi k=v] Records the result. keep commits the change; anything else undoes it.
ar hook-stop, ar hook-pretool The two hook entry points. Claude Code calls them; you never do.

ar --version prints the version, which is worth knowing when the same file has been copied into several projects and one of them is stale.

--metric has to match the name printed by the benchmark exactly. --direction says whether bigger or smaller is better. --target is a score that means "done" — leave it out to run until the budget runs out.

Two things ar log will refuse, because a result that is not true is worse than no result:

  • Recording the same measurement twice. After a result is logged, its measurement is marked as used, so logging again without running the benchmark is refused rather than quietly repeating the previous number. Pass --metric if you measured it some other way.
  • Keeping a run that crashed or failed its checks. You can record it honestly as crash or checks_failed; you cannot record it as a win.

Two rules the loop depends on:

ar run is the only way to measure. A benchmark run by hand produces a number that exists nowhere, so the log, the noise estimate and the next agent never see it.

--asi notes matter most on failures. When a result is rejected the code is undone and gone. The note is the only surviving record that the idea was tried, which is what stops the loop rediscovering the same dead end fifty experiments later.

The confidence score

After three runs, ar log also reports how big the best improvement is compared with how much the benchmark normally varies. Above 2× is probably real; below 1× is probably noise. It is advice only — it never rejects anything by itself.


Session files

Everything lives in <project>/.auto/, which the installer tells git to ignore.

File What it is
prompt.md The brief: goal, metric, which files are in scope, what to avoid, what has been tried. A new session can run the whole loop from this alone.
measure.sh The benchmark. Prints METRIC name=value.
log.jsonl One line per run, including the agent's own notes.
ideas.md Ideas it wants to come back to.
checks.sh Optional. Tests that must pass for a result to be kept.
config.json Optional. Limits — see below.
ar The program itself.
runtime.json, last_run.json, results/ Internal state. You can ignore these.

Writing a good benchmark

/autoresearch-create writes measure.sh for you, so most people read this to check its work. Write one by hand if you prefer.

There are only three rules:

  1. Exit 0 if the run was valid. Any other exit code means "something broke", and the result is thrown away rather than recorded as a bad score.
  2. Print METRIC <name>=<number>. The one whose name matches ar init --metric is the score. Any others are recorded too, as things to keep an eye on.
  3. Anything else you print is context for the agent. The last 60 lines are shown to it, so print whatever helps it decide what to try next.
#!/bin/bash
set -euo pipefail
cd "$(dirname "$0")/.."
[ -f .auto/env.sh ] && . .auto/env.sh

# 1. Check the environment first. If something is down, fail here rather than
#    returning a bad score - otherwise a good change gets undone for no reason.
curl -sf --max-time 10 "$SERVICE/health" >/dev/null || { echo "service down" >&2; exit 1; }

# 2. A quick sanity check, so a broken edit fails in a second instead of an hour.
python3 -c "import mymodule" >/dev/null

# 3. Run the workload. If it takes under 5 seconds, run it a few times and take the median.
./run_benchmark.sh > /tmp/out.json

# 4. Print the score, plus anything useful.
python3 -c "
import json; r = json.load(open('/tmp/out.json'))
print(f\"METRIC score={r['score']:.4f}\")
print(f\"METRIC latency_p95={r['p95']:.0f}\")
for name, v in sorted(r['per_category'].items())[:5]: print(f'worst: {name} {v:.3f}')
"

Things that make a real difference:

  • Speed. Whatever it takes, multiply by 200. Ninety seconds becomes five hours of waiting.
  • Consistency. If the same code gives 41 then 47, the loop will "find" improvements that are not there. For quick benchmarks, run the workload several times and print the median.
  • Failing loudly. Covered above, and worth repeating: it is the most common cause of a wasted overnight run.
  • Being hard to cheat. Whatever raises the number will get raised. If editing the benchmark or the test data would do it, list those files under Off limits in prompt.md.

Limits and cost

The loop stops by itself under any of these conditions. Change them in .auto/config.json:

{
  "maxIterations": 40,
  "maxConsecutiveFailures": 20,
  "autoResumeTurnLimit": 500,
  "maxStalls": 6,
  "runTimeoutSeconds": 3600,
  "checksTimeoutSeconds": 600
}
Setting Default What it means
target reached — The score passed --target. This is the intended ending.
maxIterations 0 (no limit) Maximum number of experiments.
maxConsecutiveFailures 20 This many rejected results in a row means it has run out of ideas.
autoResumeTurnLimit 500 A final backstop.
maxStalls 6 Turns where nothing at all changed — no result, no commit, no edited file.

These are read fresh every time, so you can change them while the loop is running. No restart needed.

maxStalls is there to catch an agent that talks instead of working. Without it, a reply with no work in it gets a "carry on" from the hook, which produces another reply with no work in it, and so on. Note it counts any change as progress — editing a file or making a commit both count — so time spent debugging does not trip it.

Cost is the real limit on a long run. Set maxIterations before you leave it alone, and set a spending cap on your account.


Troubleshooting

It does one experiment and stops. The hook is not firing. Check that .claude/settings.json has a Stop entry pointing at this project's .auto/ar, and that .auto/runtime.json says "active": true. Test it directly:

echo '{"cwd":"'$PWD'"}' | python3 .auto/ar hook-stop

You should get JSON back containing additionalContext. No output means the loop is deliberately allowing it to stop — run .auto/ar status to see why.

"Hook JSON output validation failed." For a Stop hook, Claude Code accepts only hookEventName and additionalContext inside hookSpecificOutput. There is no decision field.

"Loop paused — no run, no commit and no file change." The agent replied several times without doing anything. .auto/ar on starts it again. If it keeps happening, the brief in prompt.md is probably unclear about what to do next.

The program disappears mid-run. .auto/ was committed to git, and something reset the repo to an earlier commit. Re-run install.sh — it tells git to ignore those files — and check that git status --short shows nothing from .auto/ or .claude/.

Every experiment fails. Something in the environment is broken rather than the code. Run bash .auto/measure.sh yourself and read the error.

"not a git repository." The loop needs git to keep and undo changes. Run git init && git add -A && git commit -m initial in your project, then run /autoresearch-create again.

"REVERT FAILED." A rejected result could not be undone, so the change is still in your working tree and nothing was logged. Fix the tree by hand before running anything else — otherwise the next experiment measures both changes together and the result is attributed to the wrong one.

pkill -f "something" over SSH kills your own session. Your command contains the text you are searching for, so it matches itself. Write pkill -f "some[t]hing" instead.

Copies drift apart. Once this is installed on more than one machine, the same file exists in several places. Treat the machine running the loop as the real one and copy from it. If the agent has edited its own checks.sh there, overwriting it from your laptop throws away real work.


Tests

./test.sh              # everything
./test.sh test_hooks   # one module

127 tests covering metric parsing and the noise floor, the full init / run / log cycle against real git repos, both hooks including every guard and every bypass the off-switch blocks, and install.sh install/uninstall round trips. No dependencies beyond python3, bash and git.


Safety

The loop edits files, makes commits, undoes changes, and runs .auto/measure.sh and .auto/checks.sh unattended, with your permissions. So:

  • Run it on its own branch, starting from a clean tree.
  • Read measure.sh and checks.sh before each session. They are code that will run hundreds of times.
  • Keep credentials out of the files listed in prompt.md, and list secret paths under Off limits.

One specific thing to know: rejecting a result runs git restore and git clean -fd. That deletes untracked files that git is not already ignoring. Ignored files (caches, model weights) are safe, and so is .auto/ — but anything else you left lying around untracked is not.


Uninstall

./install.sh --uninstall /path/to/your/project

Removes the skills, the program, its hooks, and the git-ignore entries it added. Your own hooks and settings are left alone.

Your results stay: log.jsonl, prompt.md, measure.sh and everything else in .auto/. Delete that folder yourself when you no longer want them.

About

Autoresearch framework to optimise anything (ML model, package size, passing tests, build time)

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages