An autonomous trading strategy optimizer and real-time signal generator, forked from Karpathy's autoresearch and adapted for quantitative trading.
Instead of iterating on an LLM training loop, an AI agent here iterates on a stock screener and position management strategy — modifying train.py, running a walk-forward backtest, checking if the result improved, keeping or discarding, and repeating. You come back to a log of experiments and (hopefully) a better strategy. The same strategy then powers a real-time screener that produces entry signals and manages stops on open positions.
This repo is open for forks, issues, and pull requests. Whether you want to contribute a better baseline strategy, improve the harness, add new signal types, or share results — all of it helps. The goal is a shared, community-improved strategy optimizer and co-trader that everyone can benefit from. Open an issue or PR, or fork it and take it in your own direction.
Three files do all the work:
prepare.py— one-time data download and caching for all tickers. Fetches OHLCV history via yfinance. Not modified during optimization.train.py— the single file the agent edits. Contains the full strategy: screener (screen_day), position manager (manage_position), and the walk-forward evaluation harness. This is what the agent iterates on.program.md— instructions for the optimization agent: objective, keep/discard criteria, experiment sequence, closed directions, and harness configuration. This is what you iterate on as the human.
The agent runs a walk-forward backtest over a historical window, computes fold-level metrics, decides whether to keep or revert the change, and logs the result to results.tsv. Repeat for N iterations.
The same screen_day and manage_position functions in train.py are used directly by the real-time screener to generate daily entry signals and stop updates on live positions.
Optimization harness:
- Backtest window: September 2024 → March 2026 (∼19 months)
- Universe: ∼400 tickers across sectors (large-cap US equities)
- Evaluation: 6-fold walk-forward cross-validation, 60 business days per test fold
- Primary objective:
mean_test_pnl— arithmetic mean of out-of-sample P&L across all folds (higher is better) - Floor constraint:
min_test_pnl > −$30(worst single fold must not exceed this loss) - Latest baseline: mean_test_pnl ≈ $251, min_test_pnl ≈ −$3 (commit
141aa8e)
Real-time screener:
- History window: 300 days of daily OHLCV per ticker
- Universe: Full Russell 1000 + select mega-caps and major indexes
- Output: Daily entry signals (bull continuation and recovery paths) + intraday stop levels for open positions
All of the above — universe, window, fold count, objective — are configuration values that can be updated in prepare.py, train.py, and program.md to suit different goals or time periods.
Requirements: Python 3.10+, uv. No GPU required — backtests run on CPU.
# 1. Install uv (if you don't already have it)
curl -LsSf https://astral.sh/uv/install.sh | sh
# 2. Install dependencies
uv sync
# 2b. Install the Chromium build Playwright drives (the pip package does not include it;
# needed by reports/ — Tradovate / PickMyTrade report downloads)
uv run playwright install chromium
# PickMyTrade's login has a reCAPTCHA: log in ONCE per machine with a visible browser;
# the saved profile (<global>/general/browser_profiles/pickmytrade) keeps later runs headless
uv run reports/get_pickmytrade_alerts.py --headed
# 3. Download and cache ticker data (one-time, takes a few minutes)
uv run prepare.py
# 4. Run a single backtest manually to verify setup
uv run train.pyThe session-analysis skill (through get-reports) downloads the PickMyTrade alerts CSV
with a HEADLESS browser. PickMyTrade's login page shows a reCAPTCHA, which a headless
browser cannot solve. So on every new machine, and again whenever the saved login expires,
do one headed login with the user present, before relying on session-analysis:
uv run reports/get_pickmytrade_alerts.py --headedA browser window opens. The user logs in and solves the reCAPTCHA, and the script then
exports the CSV and exits. The login is kept in the persistent profile
<global>/general/browser_profiles/pickmytrade, so later headless runs (needs_login=False)
complete without the user. If a headless run instead reports that the saved profile is not
logged in, the session has expired: repeat the headed login. Always use the project's uv /
.venv Python, since the system Python does not have playwright.
On some machines — notably Windows with Smart App Control enabled — the OS blocks the
unsigned _ssl.pyd / OpenSSL DLLs that ship with uv's managed CPython, so import ssl
fails and nothing can run (An Application Control policy has blocked this file). Fix it with:
python scripts/bootstrap_env.pyIt's a no-op on machines that aren't affected (Linux, macOS, Windows without Smart App
Control). Where the block exists, it pins uv to a PSF-signed system Python 3.10 via a
machine-level uv config and rebuilds .venv. If no signed Python 3.10 is installed it prints
the one manual step (install python.org 3.10),
then re-run. The fix is intentionally machine-scoped (not committed), since the block is a
property of the machine, not the project — run this script once per affected machine.
For isolated runs, use the prepare-optimization skill first to create a dedicated git worktree so experiments don't touch the master strategy.
Open Claude Code in this repo and prompt:
Have a look at program.md and run 30 iterations of the equity strategy optimizer.
The agent reads program.md, sets up walk-forward constants, runs the baseline, and iterates — modifying train.py above the boundary, running the backtest, keeping or reverting each change. Results logged to results.tsv.
Requires IB-Gateway running on port 4002. Download futures data first:
uv run prepare_futures.pyThen open Claude Code and prompt:
Have a look at program_smt.md and run 30 iterations of the SMT divergence strategy optimizer.
The agent reads program_smt.md, runs the baseline on MNQ1!/MES1! 1m data, and iterates — modifying train_smt.py above the boundary. Both strategies are fully isolated; optimizing one never touches the other.
The daemon starts signal_smt.py at 09:00 ET on NYSE trading days, relays its output to
sessions/YYYY-MM-DD/signals.log, restarts once on crash, terminates at 13:35 ET, and calls
the Claude API to write a post-session summary.md with metrics, narrative, and parameter
recommendations.
Prerequisites:
-
Add your Anthropic API key to
.env(copy from.env.example):ANTHROPIC_API_KEY=sk-ant-... -
Verify the setup:
# Loads key from .env then runs --check (exits 0 if everything is wired up) set -a && source .env && set +a && uv run python -m orchestrator.main --check
Run the daemon:
macOS / Linux (foreground terminal or tmux):
set -a && source .env && set +a && uv run python -m orchestrator.mainWindows — run hidden in the background (no terminal window):
$proc = Start-Process -FilePath "uv" `
-ArgumentList "run", "python", "-m", "orchestrator.main" `
-WorkingDirectory $PWD -WindowStyle Hidden -PassThru
Write-Host "Orchestrator PID: $($proc.Id)".env is loaded automatically by the orchestrator at startup — no need to source it first.
Save the PID so you can stop it later.
Stop the daemon:
macOS / Linux: Press Ctrl+C. The daemon catches the interrupt and exits cleanly.
Windows:
Stop-Process -Id <PID> -ForceSession files are written to sessions/YYYY-MM-DD/ (gitignored).
Install the python-performance-optimization skill alongside this repo. It profiles and vectorizes the Python code the agent writes, keeping hot paths fast. The optimization agent is instructed to invoke it after every train.py edit — numpy vectorization matters when backtesting hundreds of tickers across dozens of folds per iteration.
prepare.py — equity data download and caching (yfinance)
prepare_futures.py — MNQ/MES 1m futures data download (IB-Gateway)
train.py — equity strategy: screener, position manager, walk-forward harness
train_smt.py — SMT divergence strategy: signal logic + intraday backtest harness
program.md — equity optimizer instructions: objective, experiment sequence
program_smt.md — SMT optimizer instructions: kill zone, divergence tuning
screen.py — real-time screener (uses screen_day / manage_position from train.py)
data/sources.py — data source abstraction (yfinance + IB-Gateway)
pyproject.toml — dependencies
results.tsv — experiment log (untracked)
strategies/ — registry of named strategy snapshots for reuse across runs
.agents/plans/ — implementation plans
- Single file to modify. The agent only touches
train.pyabove the# DO NOT EDIT BELOW THIS LINEboundary. Diffs are small and reviewable. - Walk-forward evaluation. Each iteration runs N folds of out-of-sample backtests. The fold structure prevents in-sample overfitting and gives a realistic picture of how the strategy would have performed across different market regimes.
- Harness and screener are decoupled. The evaluation harness in
train.pyand the real-time screener inscreen.pyshare the same strategy functions, so an improvement validated by the harness automatically applies to live signalling. - Multi-strategy potential. Currently there is a single global strategy applied to all tickers. Splitting into niche strategies — each optimized for a specific sector, signal type, or market regime — would likely produce better metrics and more accurate signals. The harness structure supports this; it's an open direction.
The optimization harness is effective at tweaking and improving an existing strategy that already shows a positive edge. Given a proven baseline, it reliably finds parameter improvements, entry filters, and exit refinements that hold out-of-sample.
It is less effective at building a strategy from scratch. Without a clear starting signal and a well-specified objective, the agent tends to find local improvements that don't generalize. Attempts to optimize for total P&L from a blank slate without constraining the strategy direction have so far not converged to consistently meaningful results. The harness is best used as an optimizer, not a strategy designer.
MIT