Skip to content

Repository files navigation

trAIde

Autonomous multi-agent AI crypto trader powered by Azure OpenAI and KuCoin APIs.

Three specialized agents collaborate in a continuous loop: a Trading Agent that executes orders with full risk management, a Research Agent that scouts opportunities and market intelligence in parallel, and a Supervisor Agent you can talk to via Telegram to monitor and control the system.

Contents

Architecture

trAIde runs as a single Python process: a continuous poll loop drives a Trading Agent and a Research Agent (both backed by Azure OpenAI), a code-driven ProtectionManager, and a sanitized public dashboard — while an interactive Supervisor Agent lets you steer the system from Telegram. The diagrams below show the same system from six angles.

System components

flowchart TB
    operator(["Operator"]) <-->|Telegram| sup["Supervisor Agent<br/>(read-only + note injection)"]

    subgraph core["trAIde process"]
        direction TB
        mainloop["Main Loop<br/>(polls every POLL_INTERVAL_SEC)"]
        trade["Trading Agent<br/>(46 tools)"]
        research["Research Agent<br/>(web + news scout)"]
        prot["ProtectionManager<br/>(code-driven profit-lock)"]
        mem[("MemoryStore<br/>.agent_memory.json")]
        dash["Dashboard Publisher"]
    end

    subgraph azure["Azure"]
        aoai["Azure OpenAI<br/>(LLM inference)"]
        blob[("Blob + Table storage")]
    end

    kucoin["KuCoin<br/>Spot + Futures API"]
    sandbox["SandBox web app<br/>(read-only dashboard)"]
    spectators(["Public spectators"])

    mainloop --> trade
    mainloop --> prot
    mainloop --> dash
    trade <--> aoai
    research <--> aoai
    research -->|notes| mem
    trade <-->|orders, positions| kucoin
    prot -->|breakeven, close| kucoin
    trade <-->|context, outcomes| mem
    sup -->|inject notes| mem
    sup -. read-only .-> kucoin
    dash -->|sanitized %| blob
    blob --> sandbox
    sandbox --> spectators
Loading

Trading Agent -- Places bracketed futures limits, manages positions, sets TP/SL, runs multi-timeframe analysis, and explains code-enforced risk decisions. New spot exposure and market entries are disabled; spot tools remain for existing-position protection and closes.

Research Agent -- Runs as an explicit, flat-book handoff when research is stale or repeated no-trade runs justify a wider market scan. It owns broad web discovery and logs reusable findings; ordinary position-management runs do not repeat expensive web searches.

Supervisor Agent -- Interactive Telegram bot with read access to the entire system. Can query positions, balances, performance, logs, source code, and config. Can inject temporary (one-shot, highest priority) or permanent notes into the Trading Agent's system prompt to influence its behavior.

Runtime: the poll loop

Every POLL_INTERVAL_SEC the loop rebuilds a full account snapshot, runs profit protection, checks circuit breakers, and considers agent invocation when a trigger fires or the idle threshold is reached. The model runs in a single background worker, so slow inference never pauses polling or deterministic protection. Model calls are additionally throttled by book state (FLAT_AGENT_COOLDOWN_SEC / ACTIVE_AGENT_COOLDOWN_SEC). While flat, price-move magnitude and market breadth automatically shorten the quiet cooldown toward the active cadence.

sequenceDiagram
    autonumber
    participant L as Main Loop
    participant K as KuCoin
    participant M as MemoryStore
    participant P as ProtectionManager
    participant A as Trading + Research Agents
    participant D as Dashboard

    loop Every POLL_INTERVAL_SEC
        L->>K: build snapshot (tickers, balances, positions, stops, fills)
        L->>M: update drawdown + position extremes (peak/trough PnL)
        L->>P: run(snapshot)
        P->>K: ratchet stop to breakeven / close on give-back
        L->>M: record triggered TP/SL closes (with exit price)
        L->>L: check circuit breakers (drawdown, losses, heat)
        alt triggers fired OR idle threshold reached
            L-->>A: start one background agent run
            A->>K: place / cancel / bracket orders
            A->>M: log decisions, trades, research notes
        else otherwise
            L->>L: idle_polls++
        end
        L->>D: publish sanitized snapshot (throttled)
        L-->>L: sleep until next cycle
    end
Loading

Entry decision & risk gates

Every proposed entry runs a fixed gauntlet of code-enforced gates before any order reaches the exchange. Position management (manage / hold / protect / close) bypasses the entry gates. In a hostile (bearish / RSI-exhausted) regime the confidence bar is raised and size shrunk; size also scales down with conviction (how far confidence clears the floor). A confirmed trend-aligned short can pass the anti-FOMO gate, a verified funding_carry / macro_event can pass it too (its payoff is not the trend continuing), and a daily-aligned entry can break the daily-vs-1h deadlock when a counter-bounce is stalling.

flowchart TD
    propose(["Agent proposes an entry"]) --> cb{"Circuit breaker active?"}
    cb -->|yes| reject["Rejected<br/>(no new entry)"]
    cb -->|no| pl{"Post-loss cooldown?"}
    pl -->|yes| reject
    pl -->|no| nc{"No-chase: same direction at a<br/>worse price than a recent win?"}
    nc -->|yes| reject
    nc -->|no| ti{"Trade-interval cooldown?"}
    ti -->|yes| reject
    ti -->|no| dg{"Daily gate / anti-FOMO /<br/>1h-alignment / TF-conflict?"}
    dg -->|blocked| reject
    dg -->|ok| vol{"Volatility gate<br/>(ATR + 24h range)"}
    vol -->|hard block| reject
    vol -->|soft| scale["Scale size down<br/>(quadratic)"]
    vol -->|ok| conf{"Confidence &ge; regime-adjusted MIN_CONFIDENCE<br/>and net profit after fees?"}
    scale --> conf
    conf -->|no| reject
    conf -->|yes| convsize["Scale size by conviction +<br/>concentration cap (&le; equity %)"]
    convsize --> exec["Place order +<br/>mandatory TP/SL bracket"]
    exec --> openpos(["Position open"])
Loading

Position lifecycle & profit protection

Once open, a position is protected in code on every poll: the stop ratchets to breakeven once the run clears the noise band (+0.5R at the current stop width), and a give-back of the peak run closes it to lock gains — independent of the LLM.

stateDiagram-v2
    [*] --> Flat
    Flat --> PendingEntry: limit order placed
    PendingEntry --> Flat: expired or cancelled
    PendingEntry --> Open: filled
    Flat --> Open: market entry filled
    Open --> Open: manage bracket, scale-in if gates pass
    Open --> Protected: run cleared noise band (+0.5R), stop ratcheted to breakeven
    Protected --> Protected: give-back below threshold
    Open --> Closed: TP, SL, or agent close
    Protected --> Closed: gave back 35% of peak, market close
    Protected --> Closed: TP or breakeven stop hit
    Closed --> Cooldown: post-loss or post-win no-chase
    Cooldown --> Flat: cooldown expires
    Closed --> [*]
Loading

Playbooks, measurement & risk allocation

The bot does not run one strategy. Every entry declares which playbook (setup_family) it belongs to, and each playbook keeps its own scoreboard. Nothing in the code decides that trend-following or fading is correct — that is the market's call and it changes. The code measures, and capital follows whatever currently pays.

Playbook Thesis Needs the direction call to be right?
continuation Trade with an established trend; timeframes agree and you expect persistence Yes
fade_extreme Fade a stretched move back toward value (RSI at a 30/70 extreme) Yes
breakout Enter on a break of a range or level, expecting expansion Yes
range_edge Buy support / sell resistance inside a defined range Yes
funding_carry Take the side the funding transfer pays, on the contract's own settlement clock (1h, 4h or 8h) No — the transfer happens whichever way price moves

funding_carry is the odd one out on purpose. Every other playbook is a prediction; funding is a mechanical transfer, and it is the best-documented edge in the perpetuals literature.

How a playbook earns (or loses) its capital

A family is never vetoed for being unfashionable — it is sized by its own measured forward return against the round-trip cost it has to clear.

flowchart TD
    declare(["Model declares setup_family<br/>+ the side it is entering"]) --> probe["Probe recorded AT THE CALL<br/>(market price, model, confidence)"]
    probe --> row{"Which row judges it?<br/>that SIDE's own row if it has &ge; 20 probes,<br/>else the pooled family"}
    row --> gate{"Scored yet?<br/>(&ge; 20 non-overlapping probes)"}
    gate -->|"no — unproven"| explore["Explore size &times;0.4<br/>trade small, gather evidence"]
    gate -->|yes| tstat{"t = net of cost &divide; its own SE"}
    tstat -->|"t &lt; 1: 'no edge' (net &le; 0)<br/>or 'unproven' (0 &lt; net &lt; SE)"| aside["STAND ASIDE — zero stake<br/>refusal says which of the two"]
    tstat -->|"1 &le; t &lt; 2"| ramp["Ramp &times;0.4 &rarr; &times;1.0<br/>with the evidence"]
    tstat -->|"t &ge; 2"| full["Full measured size<br/>(never sized UP beyond budget)"]
    explore --> probe
    ramp --> probe
    aside -.->|"probe still recorded,<br/>so it can recover"| probe
    full --> probe
Loading

The dotted line is the important one. A skipped trade still records its probe, because the probe is written when the call is made rather than when an order is placed. Without that, a stood-aside family would starve of the very evidence that could reinstate it — a deadlock this bot hit twice for real. (A side still thin on its own is judged on the pooled family for the stand-aside but is capped at the explore size, whatever the pooled family has earned.)

Judged per side, and told the truth (edge.family_stake_status / family_evidence_row, Sep 25 2026). Two things were wrong with the stand-aside, and neither was the rule itself.

  • It lied about why. The release bar is a stateless t = net / SE ≥ 1 — the zero point of uncertainty-shrunk Kelly, (1 − 1/t²)⁺ — recomputed every run. It is not the "hysteresis" the docstring promised; nothing is remembered, so near the line it can flip back and forth on noise (the explore floor bounds what a flip costs: 0 → 0.4×, never 0 → 1.0×). On Sep 24 all nine continuation refusals fired with the family's verdict at edge and net +0.41% to +0.54% (t ≈ 0.7–0.9), yet every refusal said "NO EDGE … coin-flip minus fees" — and the model repeated "edge does not clear costs" in its own decline. The refusal and the STAND ASIDE log now say which of two different things fired, with the net, SE and t: no edge (net ≤ 0) keeps the old wording; a positive net inside its own noise is called exactly that, and re-opens once net clears one SE. Before the model proposes, edgeReport.signalEdge.by_family now carries t_stat, standAside, stakeReason and stake (0 when stood aside, else the worse of the measured and explore factors) from the same functions the order path calls — it used to show only the raw verdict, so the model re-proposed the benched family nine times in about three hours and learned the bench from the refusals. Stateful hysteresis was considered and rejected: it would have re-opened all nine of those refusals, a loosening that needs its own evidence.
  • One side's record sized the other. The Kelly stake is a bet on a side of a playbook, but the family was scored pooled. On Sep 24 continuation's edge was carried by its longs (n=40, +1.24% at 240m) while its shorts sat at n=9, −0.18% ± 1.05% — and three SPX continuation shorts were sized 0.52–0.79 on the longs' record. signal_edge_stats now also reports by_family_side (split after the per-symbol de-overlap, at the family's own horizon), and the limit-entry path judges a bet on its side's own row once that row has 20 probes — stand-aside, no-edge shrink and explore ramp all read it. A side still thin on its own falls back to the pooled row for the stand-aside and the shrink, but is capped at the explore floor, so an unproven side never inherits the other side's upsizing. Symmetric across regimes (in a downtrend shorts earn their own verdict), no new constant; side=None keeps the pooled behaviour for every other caller. The refusal names the row that judged it ("continuation:short n=9 is unproven, judged on pooled continuation"), and its "where to go" hint lists a family that is open on only one side as that side only.

The measurement loop

Win rate and PnL conflate three different things: whether the direction was right, whether the fill was any good, and whether the exit was managed well. signalEdge isolates the first.

flowchart LR
    dcall(["Direction call"]) --> stamp["Stamp market price<br/>+ taker-flow reading<br/>at signal time"]
    stamp --> wait["Poll loop settles the forward FUTURES mark<br/>(+ funding received) at 5m / 15m / 60m / 240m<br/>(verdict + sizing still on 60m / 240m)"]
    wait --> dedupe["Collapse OVERLAPPING probes<br/>(one per symbol per window)"]
    dedupe --> score["Mean forward return<br/>signed by traded direction"]
    score --> hurdle{"&gt; measured round-trip cost?"}
    hurdle -->|yes| edge["verdict: edge"]
    hurdle -->|no| noedge["verdict: no edge"]
    hurdle -->|"n too small"| insuf["verdict: insufficient data"]
    edge --> alloc["Risk allocation<br/>per family"]
    noedge --> alloc
    insuf --> alloc
    alloc --> dcall
Loading

Two details that are easy to get wrong and were both live bugs:

  • Measure from the market price at signal time, not the limit price. Measuring from the resting limit scores the discount as if it were prediction — on real data that showed a spurious +1.24% at a 92% hit rate where the honest figure was −0.007%.
  • Collapse overlapping probes. Thirty probes on one symbol inside four hours are ~one observation. Counting them independently once flipped the live verdict to edge at t=+3.66 when the decimated truth was t=+0.51 — below cost, i.e. nothing.
  • Settle on the market that fills, plus the funding it was paid. The probe's base is the futures mark, but until Sep 25 2026 the poll loop settled it from the spot ticker. On a coin whose perp trades away from spot that gap was scored as prediction — and it landed almost entirely on funding_carry, the playbook chosen exactly when the gap is widest (ONE-USDT: spot ~0.0050 vs perp ~0.0038, a "+25% five-minute return" on a +2.5% contract move). Stored, funding_carry read t=1.90 and was advertised as "currently paying"; settled on futures it is a coin flip on the stand-aside line. Now each probe records its priceSource and settles from the SAME market (main._settle_probes_on_futures fetches a futures mark only for the symbols with a horizon due, reusing a held position's own mark; a symbol with no mark waits and is written off as unmeasured — never filled from spot). Funding the position would have received over the window is stamped as f{h} beside the price (memory.funding_received_from_history: longs receive −rate, shorts +rate) and added to the signed return, because a carry call is a bet on the transfer as well as on price. A failed funding lookup is recorded as unknown (f{h} = None, with t{h} the time the price was read), never as zero: it is re-asked on later polls and backfilled from the exact history for the horizon's own window, and one still unknown past its window is skipped in carry scoring (counted as unknownFundingCredit) — before Sep 25 a single 503 silently scored the carry credit as zero forever, always biasing carry down. All of this settlement I/O runs on the poll-loop thread after the protection pass, so it is bounded: most-urgent rows first, a per-poll budget of a quarter of POLL_INTERVAL_SEC, a per-symbol doubling backoff capped at the row's remaining tolerance, and a 5s HTTP timeout for these reads only (main._MeasurementBudget). Exit probes settle on the same futures mark (brackets trigger on it), and one with no price by the end of its measurable life is recorded unmeasured instead of being resolved days later by whatever the ticker says then. The spot map is still what the auto-triggers read. Probes stored before the fix are re-settled once with scripts/resettle_probes_futures.py.
  • Every call names its model and how sure it said it was (edge.confidence_edge_stats, Sep 25 2026). Seven absolute confidence thresholds shape entries (the floor, the hostile-regime floor, deadlock, trend-short, reversal, relative-strength, full-conviction sizing), all implicitly calibrated on one model's scale — and the stated number drifts with the model and with the tape (median on closes: 0.82 gpt-5.4-mini, 0.74 gpt-5.4, 0.78 gpt-5.6, 0.76 gpt-6). Nothing could test whether it carries information, because probes recorded neither. They now stamp model, confidence and the effective minConfidence, and a report-only statistic ranks each model's confidence against its calls' net forward return at the family's own horizon — within each UTC day, because a pooled correlation mostly measures regime (gpt-5.6's pooled +0.20, t=2.25, came from a calm stretch of low confidence next to a rally of high confidence, not from ordering calls). It is on the dashboard and the Supervisor's get_edge_scoreboard, and in the agent's edgeReport.confidenceEdge only as the running model's own row, once it has 20 scored calls (another model's record is not this model's number). Only rows carrying the full stamp (model + confidence + minConfidence) are scored: older placed-trade rows already had model and confidence, but they are the post-gate subset (calls that also cleared the stand-aside, the RR gate and sizing), so they are excluded rather than scored as if they were every call — the 09-24 gpt-5.6-luna 'inverted' row was built entirely from them. Nothing in the trading path reads it; it is the evidence any future recalibration of those thresholds would need.

Gate scoreboard — measuring the directional gates (report-only)

The directional gates in the futures limit path (daily exhaustion "anti-FOMO", opposing daily, 1h alignment, timeframe conflict, BTC correlation, 24h move cap, symbol bench) returned before the signal probe, so a refused call left no evidence and nothing scored any gate live: Sep 22–24 had 8 hard directional refusals and 0 probes — the opposing-daily branch did not even log — against 23 of 23 probed NET RR / STAND ASIDE refusals. The gates' real footprint is the model's self-censorship (97 of 131 declines cited daily exhaustion), which no hard-refusal log can see. The offline study that said the exhaustion and 1h gates "save money" does not hold outside one week: the exhaustion gate's whole advantage came from the Sep 17–24 rally (without it, exhausted longs did better), and day-bootstrap intervals span zero for both. Honest status: value currently unmeasurable. Since Sep 25 2026 the bot records the evidence that can measure it — slowly (a ~0.17%/call gap needs months of day-level data) — and gate behaviour is unchanged.

flowchart LR
    ana["analyze_market_context<br/>(every analysis)"] --> st["gate-STATE row per side<br/>which gates would refuse a plain call<br/>(once per symbol/side per 240m)"]
    ord["place_futures_limit_order"] -->|"refusal with a scored gate code"| rf["refusal row<br/>(capped per gate)"]
    ord -->|"admitted: gatesPassed stamp"| sp["signal probe<br/>(family verdicts)"]
    st --> settle["settled on the futures mark<br/>+ funding, 5/15/60/240m"]
    rf --> settle
    settle --> fold["fully settled state rows<br/>fold into UTC-day cells"]
    fold --> board["edge.gate_scoreboard<br/>blocked vs allowed, same day"]
    settle --> board
    sp -->|"admitted baseline + hatched calls"| board
    board --> out["dashboard · Supervisor · hourly log<br/>NEVER the trading prompt"]
Loading
  • Every pre-probe refusal names its gate. Each {"rejected": true} returned before the signal probe carries a gate code — a scored gate (anti_fomo, daily_opposing, h1_align, tf_conflict, correlation, move_24h, bench, vol_limit, confidence_floor — the regime-raised floor refused 14 calls on Sep 23/24, more than every directional gate combined — plus no_chase and the two cooldowns) or a structural one (malformed bracket, stale analysis, circuit breaker, caps) that records nothing. A source test walks the AST and fails if any such return lacks a code, so a gate added later is covered by construction. The opposing-daily and 24h-move refusals now log (DAILY GATE BLOCK, 24H MOVE BLOCK).
  • Hard refusals by a scored gate are recorded by the tool wrapper (place_futures_limit_order → _place_futures_limit_order_impl) as a gate probe at the live futures mark — total try/except, WARNING on failure, and the refusal returned to the model is never altered.
  • Gate-state readings are the headline: every clean analyze_market_context records, per side, which gates a plain call would hit right now (tools.directional_gates_against, pinned by tests to the gate the order path actually returns), at the analysis's futures mark. Model-independent, so it sees what the model never proposes.
  • Admitted calls say which hatch let them through (gatesPassed on the signal probe: declared / mechanical / fade / reversal / deadlock / relative strength), so the hatches are scored too — 28 hatch admissions vs 8 refusals in the same window.
  • Kept out of every verdict. Gate probes live in their own gate_probes bucket that signal_probes() never reads, so a refused or never-proposed call can't move a family verdict or evict continuation evidence. Refusal rows are capped per gate (150, like families); state rows are folded, once every horizon is stamped, into compact per-UTC-day sums (gate_state_days, 120 days by count) — raw rows at ~6x the refusal rate would add ~1 MB a month, and a row cap would keep only the last few days, a one-regime scoreboard by construction.
  • The number is a same-day differential. edge.gate_scoreboard reports, per gate and side, blocked vs the calls the gate allowed on the same UTC day and horizon (state rows at continuation's own holding horizon; refused and hatched calls at their family's horizon, against admitted calls), net of cost, with a day-clustered SE (alts move together) and a verdict only at n ≥ 20 over ≥ 3 days, with |t| above the Student-t cut-off for that many days (df = days − 1: ~4.5 at 3 days, ~2.3 at 10) — the normal |t| ≥ 2 bar called 20% of 3-day pure-noise boards significant. An absolute "refused calls returned X" is regime-dominated — in a rally every refused long looks profitable — so it is shown but never headlined.
  • Never in the trading prompt. A line saying "gate X's refusals pay" invites relabelling a call past the declared-setup hatches. It goes to the dashboard (strategyEdge.gateScoreboard, % and counts only), the Supervisor (get_gate_scoreboard, told not to relay verdicts to the trading agent unasked) and an hourly GATE SCOREBOARD (report-only) log line. Any future change to a gate should follow a sustained differential across regimes, never one window.

From risk budget to contracts

Size is derived from risk, never guessed and then capped. Because size is stop-defined, a wider stop buys proportionally fewer contracts at the same dollar risk.

flowchart TD
    eq["Account equity"] --> frac["Coherent risk fraction<br/>min(RISK_PER_TRADE_PCT,<br/>daily drawdown &divide; tolerated losses)"]
    frac --> soft["&times; WORST of:<br/>soft quality stack<br/>vs measured family factor"]
    soft --> vol["&times; volatility scale"]
    vol --> budget["Dollar risk for this trade"]
    stop["Stop distance<br/>(floored at 2.5&times; ATR15m)"] --> notional
    budget --> notional["Notional = risk &divide; stop fraction"]
    notional --> lots["Round DOWN to contract lots"]
    lots --> caps{"Hard caps:<br/>risk budget / portfolio heat /<br/>concentration / max notional"}
    caps -->|"any binds"| shrunk["Shrink to the binding cap"]
    caps -->|clear| place["Place bracketed order"]
    shrunk --> place
Loading

The soft stack and the family factor combine by taking the worse of the two, never their product. Multiplying them once collapsed a $66 account to an $8.37 notional against a $10.24 contract minimum — 79 agent runs, zero orders placed.

Every entry records how it was sized (entryContext.sizing via tools.sizing_breakdown, Sep 25 2026): equity at placement (not at fill), the effective risk fraction, the volatility scale, every soft factor and whether SIZE_QUALITY_FLOOR lifted their minimum, both family legs (familyMeasured, familyExplore — the code keeps only their minimum) and the row that judged them, the risk/concentration scales, and the contract count after each stage (raw → lotFloor → riskCap → heatCap → concentrationCap), plus unitless riskFracTarget / riskFracActual. Before this, a sizing audit had to back equity out of today's balance and rebuild the factor stack from SIZE FACTORS log lines that rotate away in about two days — and a lot rounded down on a high-priced contract (ETH/BCH ≈ 0.53× of target) looked the same as a heat cap. It is order-path telemetry, so it is total by construction: any failure logs SIZING TELEMETRY LOST at WARNING and stores None; the order is never refused, resized or delayed. The dollar fields stay in the local, git-ignored memory file — the dashboard publisher is whitelist-based and a regression test pins that none of it is published.

The soft stack stops at the exchange's floor, it does not veto there (tools.entry_contracts_at_least_one_lot, Sep 9 2026): a size factor means bet smaller, and on a ~$70 account "smaller" routinely lands under the smallest position the venue sells. Rounding that to zero turned caution into abstention — measured live over 2026-09-07→09, five of the eight entries that reached final sizing died on it: three PUMP ($2.83 / $2.60 / $2.58 against a $4.60 lot) and two XRP ($11.26 / $11.38 against $14.54). Every one of them, taken at one lot, would have risked only 26–52% of the hard per-trade budget — they were refused as too big while being a quarter to a half of what is allowed, because the guard tested an opinion-scaled budget rather than a risk one. The soft stack now floors at one lot and hands the survival question to the guards built for it: regime.risk_capped_contracts still returns 0 when one lot genuinely exceeds the risk budget, and the heat and concentration caps still bind. Those rejections are correct and unchanged; the one they replace was not.

Regimes: why the label is not the decider

The regime classifier is information, not a gate. It has to be, because it is unreliable: on live data market_regime read trending on 68 of 69 entries, including through a two-week range. Keying playbooks off that label made three of the five families structurally unreachable — a counter-trend breakout was rejected before it could log a single probe.

flowchart LR
    subgraph label["Regime label (informational)"]
        trend["trending"]
        range["ranging"]
        squeeze["squeeze"]
    end
    subgraph reality["What actually decides"]
        decl["Model DECLARES a playbook"]
        score["Scoreboard sizes it"]
    end
    label -.->|"informs, never vetoes"| decl
    decl --> score
    score -->|"pays"| more["More capital"]
    score -->|"does not pay"| less["Less, then none"]
Loading

The declaration is the trigger; the scoreboard is the judge. This is the codebase's core rule in one line: code enforces survival, the model owns opportunity — so widening what may be proposed is safe, because what may be risked stays fully code-governed downstream.

Taker flow — measuring who is pushing price before trading it

Every playbook above reads closed candles at 15m and above. That means the bot has always been able to see what price did and never who was pushing it — the buy/sell aggressor split is not in a kline, on KuCoin or anywhere else. The public trade tape (/api/v1/trade/history) is the one endpoint that carries it: each row is a taker — someone who crossed the spread and demanded immediate execution — with a side. Resting liquidity never appears. So the buy share answers one narrow question: of the volume that insisted on trading right now, how much of it was buying?

This is being measured, not traded, and the honest prior is that it will not pay:

  • Order-flow imbalance is real and well documented (Cont, Kukanov & Stoikov 2014) — but it is measured in ten-second buckets and is largely contemporaneous: price moves because of the flow. The lagged, forecastable part is concentrated inside a minute and decays fast.
  • The retail crypto version (CVD, taker buy/sell ratio) has no rigorous public backtest behind it — no out-of-sample results, no hit rates, no ICs. It is descriptive, and it is venue-fragmented.
  • KuCoin is not where these prices are set. Its tape is one venue's slice of a market led elsewhere.
  • And the wall: a ~0.12–0.21% round trip (taker fee both sides plus measured slippage) against a few basis points of short-horizon drift. An effect can be perfectly real and still not clear that.

So instead of adding a flow gate, the repo does what it did for signal edge and exit discipline — probe → settle forward → surface the statistic → let the model decide:

  1. Sample the tape every poll for held positions first, then the rest of the watchlist ranked by the move that is about to wake the agent (TAKER_FLOW_MAX_SYMBOLS, one public REST call each). Ranking the tail alphabetically sounds neutral and is not: with a 50-coin universe and a cap of 12 it recorded AAVE…FARTCOIN for three days while every direction call went to TAO, INJ, WLD and XLM — a sample that never covers the trades cannot score them. A 100-trade window spans 2–7 minutes on a mid-cap perp against a 60s poll, so consecutive windows overlap heavily — the running level (buyShareEwma) is built from the newly arrived slice only, since smoothing the whole window would be smoothing the same trades repeatedly and claiming a confidence the tape never provided. Readings also expire (10 polls): a symbol that rotates out of the cap stops being refreshed rather than being marked dead, and a reading that outlives its shelf life is dropped everywhere it is read, because absent is honest and stale is indistinguishable from current once it has been stamped onto a call.
  2. Stamp the reading onto every direction call at the moment it is made — by settle time the tape is hours gone, so a reading not captured at the call is unrecoverable.
  3. Score it forward at 5m / 15m / 60m / 240m. The two short horizons were added because that is where flow information is documented to live; a statistic that only looked at 60m could not see the effect it is testing for.
  4. Report the spread: with-flow forward return minus against-flow forward return. Two staged verdicts, because conflating them is how a real-but-unusable effect gets traded — informative (clears its own standard error) and tradable (also clears the round trip).

Two properties keep this from quietly corrupting the existing scoreboard. Report widely, act narrowly: the new 5m/15m horizons are reported everywhere, but the live signal_edge_stats verdict and family sizing stay anchored to the horizons they were always computed on. And a forward stamp is written off as missed rather than back-filled once it is more than 20% of its horizon late — without that, adding a 5-minute horizon to a store of already-settled probes would have back-stamped every retained probe with the current price and labelled multi-day returns as five-minute returns.

Nothing in the trading path reads any of it. The model sees edgeReport.takerFlow only once the sample can answer the question — an n=5 reading in the prompt is an invitation to trade a coin flip — and even then it is a number to reason about, never a rule to obey.

Memory & dashboard data flow

The local MemoryStore is the agent's working memory (auto-pruned by RETENTION_DAYS). A whitelist sanitizer publishes a normalized, dollar-free projection to Azure, which a separate SandBox web app renders for public spectators.

The published payload includes a strategyEdge block — the honest headline for the whole system. Win rate and PnL conflate three different things (was the direction call right, was the fill any good, was the exit managed well), so they cannot say why the bot is winning or losing. strategyEdge measures the signal alone: forward return from the market price at signal time, signed by the traded direction, against the round-trip cost it must clear. It carries verdict (edge / no edge / insufficient data), byHorizon, and byFamily — the per-playbook score — plus familyRiskFactor, the multiplier each family is currently earning, so the dashboard explains why capital sits where it does rather than only reporting the result. Everything in the block is percentages, counts and verdicts: no balance, equity, position size or account identifier is involved, so it publishes in every disclosure mode including the default normalized.

"strategyEdge": {
  "verdict": "no edge", "n": 55, "costPct": 0.12, "bestHorizon": "60m",
  "byFamily": {
    "continuation":  {"n": 28, "mean_pct": -0.04, "hit_rate": 0.36, "verdict": "no edge"},
    "fade_extreme":  {"n": 21, "mean_pct": +0.19, "hit_rate": 0.57, "verdict": "insufficient data"}
  },
  "familyRiskFactor": {"continuation": 0.5, "fade_extreme": 1.0}
}

strategyEdge.gateScoreboard (Sep 25 2026) adds the gate scoreboard: per directional gate and side, what the gate blocks against what it allows on the same day — percent returns, SEs, t, counts and verdicts only.

Alongside it sits a takerFlow block — the live tape, and a running experiment on it. Until now the bot only ever saw closed candles at 15m and above: it could see what price did, never who was pushing it, because KuCoin klines carry no taker split. takerFlow.live fixes the blind spot on the panel — per symbol, the share of taker volume lifting the offer (buyShare, volume-weighted; buyTradeShare, one vote per trade — they diverge when size and count disagree, which is the whale/retail split a single CVD number hides), smoothed across polls into buyShareEwma, with an ageSec so a stalled sampler cannot pass as a calm market. takerFlow.byHorizon is the experiment: the forward return of direction calls made with the flow minus those made against it, per horizon. The spread form matters — a raw "with-flow calls made +0.1%" proves nothing if every call made +0.1%, so subtracting the against-flow group cancels the book's own directional bias. Read the verdict next to coverage (the fraction of scored calls carrying a reading): informative means the spread clears its own standard error, tradable means it also clears the round trip. Only the second would be worth acting on, and at a ~0.2% round trip against a few basis points of short-horizon drift it is expected to fail — see below for why it is being measured anyway. Shares, counts, ages and percentages only, so it publishes in every disclosure mode.

"takerFlow": {
  "enabled": true, "verdict": "no information", "n": 64, "coverage": 0.41, "costPct": 0.12,
  "byHorizon": {
    "5m":  {"with": {"n": 22, "mean_pct": +0.03}, "against": {"n": 19, "mean_pct": +0.01},
            "spread_pct": +0.02, "spread_stderr_pct": 0.05, "verdict": "no information"}
  },
  "live": {"BTC-USDT": {"buyShare": 0.61, "buyShareEwma": 0.58, "trades": 100,
                        "spanSec": 182.4, "ageSec": 47, "samples": 214}}
}
flowchart LR
    subgraph mem["MemoryStore — .agent_memory.json"]
        direction TB
        m1["trades"]
        m2["decisions<br/>entries cap 50 / outcomes cap 200"]
        m3["plans + research notes"]
        m4["sentiments / triggers / coins"]
        m5["position_extremes<br/>(peak/trough PnL)"]
        m6["supervisor notes<br/>temporary / permanent"]
        m7["limits<br/>(per-venue drawdown)"]
    end

    agents["Trading + Research agents"] -->|context in, outcomes out| mem
    mainloop["Main Loop"] -->|extremes, triggered closes| mem
    sup["Supervisor"] -->|inject notes| mem

    mem --> pub["DashboardPublisher<br/>(whitelist sanitize — never $)"]
    pub -->|normalized % only| az[("Azure Blob + Table")]
    az --> sandbox["SandBox web app<br/>(ECharts terminal UI)"]
Loading

Features

Technical Analysis

  • 12+ indicators: EMA (fast/slow), MACD (line/signal/histogram), RSI, ATR, Bollinger Bands (with BBW%), Stochastic %K/%D, VWAP, ADX, Plus/Minus DI
  • 4 timeframes: 1D (regime gate), 4H (40% weight), 1H (35%), 15m (25%) with weighted directional scoring, daily trend gate, and timeframe conflict detection
  • Market regime detection: Trending (ADX > 25), Ranging (ADX < 20), Squeeze (BBW < 2% + low ADX) -- each with confidence scores
  • Volume profile: Point of Control (POC), Value Area High/Low (VAH/VAL) for support/resistance levels
  • OI-price divergence: Classifies open interest vs price movement (strong trend, short covering, aggressive shorts, long capitulation)
  • Funding rate divergence: Detects hidden strength/weakness from funding rate misalignment

Risk Management

  • Circuit breakers: Auto-restrict to close-only mode when daily drawdown, consecutive losses, or portfolio heat exceed thresholds

  • Optional staged take-profit: Can split TP into 60%/40% tranches, but defaults off because early realization compresses the admitted reward:risk

  • Kelly criterion sizing: Quarter-Kelly position sizing from rolling trade performance (requires minimum trade history)

  • Post-loss cooldown: Blocks new entries on a symbol for a configurable period after a loss

  • Profit-lock (breakeven ratchet + give-back cap): Enforced in code every poll, independent of the LLM (src/protection.py). Once a position's favorable excursion reaches PROFIT_LOCK_BREAKEVEN_TRIGGER_R× its initial risk, the stop is ratcheted to a fee-adjusted breakeven so the trade can no longer turn into a loss. If price then gives back ≥ PROFIT_LOCK_GIVEBACK_PCT of its peak run, the position is market-closed (reduce-only) to lock the remaining gain. Stops a profitable trade from round-tripping into a loss when the agent fails to tighten protection itself. Set PROFIT_LOCK_DRY_RUN=true to log intended actions without placing orders.

  • Trailing ratchet (PROFIT_LOCK_TRAIL_ENABLED, default ON): the primary "let winners run" mechanism, which replaces the give-back market-close. Once the peak favourable run reaches PROFIT_LOCK_TRAIL_ARM_R, the stop ratchets up to lock the greater of PROFIT_LOCK_TRAIL_LOCK_FRAC × peak and peak − PROFIT_LOCK_TRAIL_DISTANCE_R × own-risk, floored at fee-breakeven and never moving backwards. Because it is a resting stop (not a market close), the trade rides shallow pullbacks to its TP or trails a runner. Self-normalizing in R units — no ATR feed, no percentage or per-symbol knob to tune. (R is anchored to the lifecycle's original risk, so it keeps working after the stop reaches breakeven.) Jul 27 2026 correction: an earlier tuning armed the trail at 0.5R and locked 50% of the peak. That was measured and it made things worse — arming at 0.5R engaged the ratchet on noise-scale excursions (the sample's median favourable excursion is 0.27R) and then booked half of it, collapsing the average win. Live over those 27 trades: 37% win rate, avg win +0.38R vs avg loss −0.60R, net −6.4R. Replaying the same entries on real 1-minute paths under the planned bracket isolates the exit rule from the agent's re-bracketing — 63% winners but avg win +0.35R against avg loss −1.05R, a 0.33 payoff ratio, net −4.5R. No win rate survives that payoff. Trail-arm sweep on the same paths — arm 0.40R −0.17R · 0.50R +0.39R · 0.75R −0.24R · 1.00R +2.72R · 1.25R +2.25R · 1.50R +1.97R. The principle the numbers point at: a trailing stop must not arm until the trade has cleared the noise band the stop was drawn around — that band was ~1.4× ATR when the sweep ran. Aug 11 2026 correction: that sweep endorsed "1R", but 1R is a distance measured in stop units, and the stop floor (STOP_ATR_FLOOR_MULT, plus the adaptive widener) has since roughly doubled — the live median stop is now 3.0× the intraday ATR vs 1.4× in July. So the same 1.0R now sits at ~3× ATR, a move the tape almost never makes: over the last 58 live closes the median favourable excursion was 0.30R and only 12% reached +1R, so the trail never armed and winners round-tripped from solid green to a full stop-out (21 trades peaked +0.95R on average and kept +0.24R; GRAM +0.73R→−0.96R, ADA +0.62R→−1.01R, XRP +0.67R→−0.93R). The parameter stayed frozen in R-units while R's meaning doubled underneath it — the same stale-constant-after-a-geometry-change failure the config warns about elsewhere. Arming at 0.5R restores the ~1.4× ATR distance the July replay actually validated (replaying the 58 closes lifts expectancy from −0.24R toward −0.13R, flat across 0.4–0.6R, so it is not a knife-edge), so it preserves that finding under the new geometry rather than overturning it. Defaults are now arm 0.5R, lock 0.33, trail 1.0R. Sep 1 2026 — the follow-up above is now implemented, and it is the trail distance that mattered most. Measuring what the book actually leaves on the table: across 81 closes the trades reached +31R of peak favourable excursion and realised −4.5R — a 35R give-back that dwarfs the entire net loss. The cause was PROFIT_LOCK_TRAIL_DISTANCE_R = 1.0: a full R behind the peak means the peak − distance branch stays negative until the peak clears 1.5R, which happened in 0 of 81 trades, so the flat lock_frac branch always won and every winner banked a fixed 33% of its run (0G-USDT peaked +0.97R and kept +0.13R). The fix anchors both the arm and the distance to the symbol's own noise: the trail now rides one noise band behind the peak, where the band is 1 / stopAtrMult in R — literally the ATR distance that trade's own entry stop was floored to, read per-position from entryContext and passed to decide_protection(noise_band_r=...). A run shorter than one band cannot arm the trail at all, since any stop inside the band is just a free exit for the chop. Replaying the real decide_protection over the 34 closes that recorded stopAtrMult: locked R improves −8.45R → −6.86R (+1.59R, +0.047R/trade) with the arm unchanged for typical trades (median stopAtrMult 2.51 ⇒ band 0.40R vs the 0.5R arm), so the change is targeted at the give-back rather than at when protection engages. Because the band is derived per trade, this cannot re-stale when the adaptive stop floor moves again — which was the whole point of the follow-up. PROFIT_LOCK_TRAIL_DISTANCE_R and PROFIT_LOCK_TRAIL_LOCK_FRAC remain the fallback for positions whose entry predates stopAtrMult being recorded. The give-back close remains available for PROFIT_LOCK_TRAIL_ENABLED=false.

  • Trend-adaptive give-back (PROFIT_LOCK_TREND_ADAPTIVE, legacy path when trailing is off): the tight give-back defaults are a mean-reversion harness — they shook the bot out of exactly the high-ATR trends it correctly identified. A trade whose peak run reaches PROFIT_LOCK_TREND_RUNNER_R× its own risk is a revealed trend winner, so the give-back cap arms later (PROFIT_LOCK_TREND_GIVEBACK_ARM_R) and tolerates a deeper pullback (PROFIT_LOCK_TREND_GIVEBACK_PCT, default 0.55).

  • No-chase after a win: Blocks re-entering the same direction at a worse price than a recent winning exit (within POST_WIN_COOLDOWN_MINUTES). Stops the "take profit, then immediately re-buy the top" pattern; a genuine pullback (better price than the exit) is still allowed.

  • Regime throttle: In a hostile regime (bearish or RSI-exhausted daily) the confidence bar is raised (REGIME_CAUTION_MIN_CONFIDENCE) and position size shrunk (REGIME_CAUTION_SIZE_FACTOR), so the bot trades less and more selectively instead of churning low-conviction bounce-scalps in a downtrend. (Sep 25 2026: analyze_market_context now reports the floor code will actually enforce for that symbol as summary.entryGate.minConfidence; the prompt and snapshot used to show only the base MIN_CONFIDENCE, so the raised floor was learned from a rejection.)

  • Conviction-scaled sizing: Position size scales with how far the entry's confidence clears the (regime-adjusted) floor — a trade that barely clears it gets CONVICTION_MIN_SIZE_FACTOR of full size, ramping linearly to full size at CONVICTION_FULL_CONFIDENCE. Targets the failure mode where the agent takes a full-size position on a setup it itself reads as "mixed / low-conviction" (the pattern behind the SOL drawdown); only ever shrinks, never enlarges.

  • Sizing coherence (soft factors combine by their worst signal, not by a product): The soft size multipliers — regime, conviction, relative-strength, loss-streak, and expectancy — used to compound multiplicatively, so five independent 0.5–0.6 "be a bit cautious" reads collapsed a 1–2% risk budget to ~0.03–0.05% and every position became fee-dust that couldn't clear round-trip costs even on a win (the account's small-wins / big-losses shape). They are now combined by taking the single worst signal (a min), floored at SIZE_QUALITY_FLOOR (default 0.5), so a genuine edge is still sized to matter. The hard dollar-risk caps (volatility, RISK_PER_TRADE_PCT budget, concentration, portfolio heat) are applied separately and still shrink from there.

  • Noise floor on stop distance (STOP_ATR_FLOOR_MULT, default 2.5): the single biggest measured loss driver. Across the 27 closed futures lifecycles of 20–27 Jul 2026 the median stop sat at 1.4× the 15m ATR — about 0.7× the 1h ATR, i.e. less than one hourly bar of ordinary movement. Median favourable excursion was +0.27R against targets planned at 2.3–2.7R gross, and not one trade in the sample reached its take-profit; trades with tighter-than-median stops averaged −0.40R vs −0.08R for wider ones. A stop inside the noise band is not an invalidation level, it is a coin-flip on microstructure — the trade dies before the thesis can resolve, however good the read was. The code now floors the stop at STOP_ATR_FLOOR_MULT × ATR(15m) (the ATR the daily-gate analysis already computed, so no extra market call). Replaying those same entries on real 1-minute paths turns −4.5R into +4.9R, and every floor from 1.5× to 5× ATR is positive — the geometry is what matters, not the constant, so 2.5 is chosen as ≈ one full 1h bar of noise (median ATR(1h)/ATR(15m) = 2.18) rather than as the replay's argmax. This is a survival guard, not an opportunity gate: it never vetoes a setup and never changes which symbol or direction is traded, it only widens the risk leg — and because position size is stop-defined, a wider stop buys proportionally fewer contracts, so dollar risk per trade is unchanged. It also self-tunes (STOP_ATR_FLOOR_ADAPTIVE): if the bot's own winners routinely survive ≥0.6R of adverse heat, the stop is inside the working range and the floor widens by 0.5 (capped at 4.0). Adaptation is deliberately widen-only — winners showing little heat is ambiguous (it can mean the stop already eliminated everything that breathed, which is exactly what this account's data looked like: winners averaged 0.17R of heat because the tight stop truncated the rest), and the consequences are asymmetric, since a stop inside the noise destroys the strategy while a generous one only costs some size.

  • The floor stays reward:risk-neutral (STOP_FLOOR_SCALES_TARGET, default ON): widening the risk leg while leaving the target where the model put it mechanically destroys RR, and the admission gate then rejects the trade. This showed up in the first 12.5 live hours after the fix deployed (2026-08-02): the median gross RR of rejected setups fell 1.79 → 1.23, exactly cancelling the cost-model fix — which had correctly cut friction drag 0.62 → 0.28 — so NET RR BLOCKs stayed flat at 0.19/run while the order rate dropped 72%. The floor was converting itself into rejections instead of into smaller size. The resolution is that the floor is a statement about the scale of movement, not about the thesis: if the model's stop sat inside the noise band, its target was drawn on the same too-tight scale, so the reward leg travels with the risk leg and the intended R-multiple is preserved exactly. Size still shrinks, so dollar risk per trade is unchanged. This is not moving the target to pass the gate — only the unit of R changes — and it costs nothing in exit quality (replay: +4.12R with the target scaled vs +4.09R unscaled; outcomes are flat across 0.8–2.0R of target distance). ~43% of the observed rejections clear again with it on.

  • Exit discipline — measuring the model's own closes against its own brackets (edge.exit_discipline_stats, memory.record_exit_probe / settle_exit_probes, surfaced as edgeReport.exitDiscipline): the bot measured whether its entries predict and never measured whether its exits helped — and on the 2026-09-02 data the exits turned out to be the dominant behaviour. 16 positions were closed by the agent against 2 by the profit-lock, at a 13-minute median hold on brackets whose targets need hours to reach. Correction (Sep 19 2026): this entry originally claimed that replaying those closes showed the brackets worth +4.05R against the +1.88R taken — "a leak larger than the entire net loss". That replay read KuCoin futures candles in spot column order ([ts,o,c,h,l] instead of [ts,o,h,l,c]), so it tested stops against the close instead of the low and missed most stop-outs. Re-run correctly over the same window the brackets return −1.01R against −0.17R taken: the early closes helped, by +0.84R, on a sample too small to settle it either way. The scoreboard below was not affected by that column-order bug — it scores live — and it is now the only arbiter (it did read the SPOT ticker until Sep 25 2026, which invented bracket outcomes on wide-basis coins; it now settles on the futures mark the brackets trigger on, at poll resolution, so a touch-and-retrace inside one poll is still missed); the agent prompt no longer asserts a direction. What survives is the timescale argument, which never depended on that number: PRICE_CHANGE_TRIGGER_PCT=0.5 woke the model to re-decide whenever a symbol moved 0.5%, while a typical open trade's stop stood ~2.2% away — so it was asked "what now?" on wiggles a quarter the size of the trade's own risk unit, roughly 1,500 times over the window, and predictably found reasons to act. Two changes, both in keeping with code enforces survival, the model owns opportunity: (1) orchestration — regime.held_position_noise_pct raises the wake threshold for a symbol we already hold to that trade's own noise band (risk% / stopAtrMult), so sub-noise drift on an open position is no longer treated as news. The position keeps its bracket, ProtectionManager still evaluates it every poll, and the agent still wakes on a move that clears the band, on any other symbol, and on schedule — this changes when the model is asked, never what it may decide. (2) evidence, not a veto — every close that reached neither the stop nor the target is recorded and later scored against what its bracket would have returned (since Sep 25 2026, for the model's own closes, against the replayed live exit stack instead — see An early close is scored against the system below), and the running tally is handed back to the model. It is symmetric by construction: in that same sample 6 of 15 closes beat their bracket, and by family the verdict splits (fade_extreme −2.53R, but range_edge +2.23R, where closing early ducked stops) — so if discretionary closes start paying, the same number says closes add value and endorses them.

  • Scheduled macro events — the one news input that needs no latency (regime.macro_event_window / macro_event_entry_block, family macro_event, tool log_macro_calendar): the bot already had WebSearchTool, fetch_kucoin_news and fetch_coindesk_news, and the Research Agent called them — but the output was compressed into a single 0–1 sentiment score that landed at 0.5–0.62 every time, was never scored against forward returns, and gated nothing (SENTIMENT_FILTER_ENABLED defaults to false). Reacting to breaking news is not winnable here: a 60-second poll plus an LLM read loses that race by construction. Scheduled releases are the opposite — CPI, FOMC and NFP dates are published a year ahead, so this is a calendar, not a feed, and it shares the property that makes funding carry attractive: you do not have to predict the event, only know it is coming. The timing also matches the book — through 2026 BTC's explanatory power against CPI surprises concentrated at the one-hour horizon (Block Scholes), with single prints moving it 4–8% intraday, and the bot's own worst window (Aug 28) was Jackson Hole. Split along the usual line: survival — code opens no NEW position in the hour before a high-impact print (a risk control in the same family as the drawdown breaker, making no directional claim; open positions and their brackets are untouched, and the model is told an event is coming so holding into one is a deliberate choice rather than an accident); opportunity — the model may declare setup_family='macro_event' in the hour after, scored in signalEdge.by_family like every other playbook and explore-sized until it earns its verdict. The calendar is fetched by the Research Agent with the web search it already has and stored via log_macro_calendar, so there is no API key and no runtime dependency that can fail mid-poll. The effect is decaying — BTC shrugged off the July 2026 print and Deribit's CPI-day premium fell from +25% to under 5% — which is exactly why only the risk half is enforced in code and the tradeable half must prove itself on the scoreboard. An empty or stale calendar degrades to "no events known", so a missed refresh can never block trading (MACRO_EVENTS_ENABLED=false disables it entirely).

  • Funding carry — the one MECHANICAL playbook (regime.funding_carry_setup, family funding_carry): every other family here is a prediction and needs the direction call to be right. Across 178 measured probes none of them cleared costs — continuation −0.15%, fade_extreme −0.67%, both settled to no edge. Perpetual funding is different in kind: the exchange transfers value between longs and shorts at every settlement regardless of which way price moves — and each contract settles on its own clock (on 2026-09-24, 428 of 690 KuCoin contracts every 4h, 257 every 8h, 3 every hour; KuCoin shortens it when funding turns extreme). The rate is quoted per settlement: funding_carry_setup prints the real interval plus an 8h equivalent, and analyze_market_context reports futures.fundingIntervalHours. Figures quoted as %/8h elsewhere in this README are the per-settlement rates as the bot printed them before Sep 25 2026, when every rate was mislabelled per 8h — ONE-USDT's hourly −0.196% read ~8x smaller than the carry actually on offer. Positive funding means longs pay shorts (the short is paid); negative reverses it. It is also the best-documented edge in the crypto literature — delta-neutral funding carry is reported around Sharpe 2–6, against an "intraday momentum, reversal, or both" picture for the short-horizon technical signals this bot has been fishing in. The directional version taken here is not the pure hedged arbitrage (price risk remains), but it stacks two effects pointing the same way: you are paid to hold, and an extreme rate marks crowded positioning on the other side. The trigger is derived from measured cost, never hardcoded — it fires when one funding payment covers at least half the round-trip, so as execution costs fall the bar falls with them. The data was already being fetched (fetch_funding_rate, fundingFeeRate, funding divergence); nothing consumed it. It is scored in signalEdge.by_family like every other playbook, explore-sized while unproven, and stood aside if it fails to pay. Sep 2 2026 — size the expectation honestly: at the rates actually seen live the carry is small next to the price risk. One payment at 0.05%/8h against a ~1.9% stop is a 1:36 ratio, so the directional call still dominates the outcome and the carry is a kicker, not the thesis. The payments_to_cover threshold measures the carry against round-trip cost, which is the right test for "is this worth executing" but not for "is this worth risking a stop over". Treat funding_carry as an extreme rate marking crowded positioning, with the transfer paying you to wait — and let signalEdge.by_family decide whether that is real.

    Sep 1 2026 — the carry could not actually accrue, and the prompt alone could not fix that. The playbook told the model to "hold across at least one 8h settlement", but every exit mechanic in the bot is tuned for trades that resolve in minutes: the measured median hold is 13 minutes against a 480-minute settlement cycle, so a carry position was routinely breakeven-ratcheted and shaken out having paid the full round-trip cost and collected none of the transfer it was opened for — leaving nothing but a directional punt, in the one family whose entire premise is that the direction call does not have to be right. Holding to settlement is execution machinery, which is code's job, so regime.carry_hold_deadline now returns the first funding settlement after the fill (derived from the UTC 8h grid, not a tuned constant) and decide_protection(hold_until_ts=...) suppresses the code's own early profit-taking until then. The exchange-side stop is never touched, so the loss stays capped at exactly the 1R the model set — the trade only gains time, which is the one thing its thesis needs. The deadline anchors on the fill, not on "now", so a position that has already carried through its settlement becomes normally managed instead of rolling its hold forward forever. Every other family is unaffected.

    Sep 25 2026 — held to the wrong settlement on most carry contracts. "The first funding settlement after the fill" was computed on an epoch 8h grid for every contract, but extreme funding — the carry universe — is exactly where KuCoin runs short clocks: 15 of the first 18 carry trades were on 1h or 4h contracts, so breakeven, trail and early-cut stayed off for up to 7h past the payment the hold was waiting for. The hold now uses the contract's own clock: regime.first_settlement_after walks the exchange's grid (fundingTime + k × granularity, k negative or positive, so offset grids such as TRUSTUSDTM's 03/07/11 UTC work and a stamp older or newer than the fill is fine). Sources, in order: the exchange's own settlement history since the fill (exact — see below), the live clock (main._FundingClock, fetched only for held funding_carry positions, cached ≤15 min), the clock stamped on the entry (entryContext.funding — every futures entry now records {rate, intervalSec, nextSettlementTs}), then the old 8h grid; a failed fetch never drops to "no hold". Two follow-up fixes (Sep 25 2026 review): (1) walking the live grid back to the fill assumed the contract settled on today's grid ever since — false exactly when KuCoin shortens the interval because funding turned extreme. A carry filled 09:30 on an 8h clock that switched to 1h at 11:00 read a 10:00 "settlement" that never happened and lost its hold before the real 12:00 payment; the reverse (4h→8h) re-engaged a hold after the carry had collected. The loop now reads the funding history since each carry's fill until it lists the first payment, and when the live interval differs from the stamped one without history, the deadline is the earlier of the stamped clock's settlement and the live next one — never a walk of the new grid into the past. (2) The clock used to be fetched inside ProtectionManager's pass, under order_lock, and retried every poll on failure — a hanging funding endpoint cost up to the 15s request timeout per carry position per poll while holding the lock. _FundingClock is now cache-only for lookups; the poll loop calls refresh() before the protection pass and outside the lock, with a 5-minute backoff after a failure (one warning per failure streak). The effect depends on the price path and is not non-negative by construction — ending the hold sooner hands the trade back to the normal stack sooner, and on a 1h contract it can give up later payments. Illustration, not a rate: the 1m futures replay of Sep 15–24 changes only two ONE-USDT trades on 09-21 (−1.00R → about −0.05R and −0.40R → about +0.33R), net of −0.12R funding given up; none of the 4h-contract trades changed. The per-payment carry threshold is deliberately unchanged: normalising it to 8h would widen the carry universe, which is an opportunity change, not a clock fix.

  • The dashboard refuses un-renderable closed positions: a closed position needs a side and a price to draw as a trade. Rows carrying only a pnl are fragments — a pre-close estimate, a review note, a partial record — and publishing them renders an empty duplicate card beside the real trade. Seen live on 2026-09-01: NEAR appeared twice in Recently closed, once complete (RANGE EDGE / SHORT / entry 1.99400 to exit 1.99100 / +0.02R) and once with no side, no prices and no family. MemoryStore._authoritative_realized_rows catches upstream duplicates by shape, but this bug class has now surfaced under four different action names (hold-close-only, close_reviewed, close_short, and a bare *_triggered fragment), so the presentation layer now drops rows it cannot render rather than waiting to learn the fifth. It logs a warning when it does, because a fragment usually means a close was recorded twice upstream.

  • The equity index cannot run away (_MAX_DAILY_INDEX_STEP, _INDEX_SANITY_FACTOR): the published index is a daily chain — indexClose_today = prevDayClose x (1 + intradayReturn) — over an Azure series that is durable and never rewritten, so one bad point is permanent and every later day multiplies it forward. On 2026-08-31, after a two-week outage, the dashboard read +72,546,760% (index 72,546,860 against a base of 100: a 725,468x blow-up). The way corruption enters is update_limits anchoring each new day's baseline to whatever equity reading arrives first — and this account's log is full of KuCoin 504s and "futures event history is incomplete", so a partial snapshot (spot only, futures timed out) anchors the day near zero and the next honest reading computes a five-figure "return". Three guards now: the baseline refuses an opening equity more than 10x from yesterday's (a bad read, not an overnight move); the step holds the index flat when a day's return exceeds ±50%; and the chain re-anchors to the index base when the previous close sits outside base/1000..base*1000. The series read also hides already-corrupt points, because healing today is not enough while the chart still renders the old spike.

  • Overlapping probes count once per window: probes are recorded minutes apart, so their forward windows overlap almost entirely — thirty probes on one symbol inside four hours are close to one observation, not thirty. Counting them independently inflates the sample and the verdict with it. On 2026-08-11 that produced a false positive on live data: the 240m horizon read +0.224% (t=+3.66, n=131) and the verdict flipped to edge, but decimating to one probe per symbol per horizon window gave +0.068% (t=+0.51, n=31) — below the cost hurdle, i.e. nothing. Since this verdict governs how much capital each family receives, an inflated sample can size the bot up on noise, which is the most expensive mistake this module could make. signal_edge_stats now keeps one observation per symbol per horizon window; different symbols moving at the same time still count separately, because they genuinely are separate observations.

  • Probes record the CALL, not the order (memory.record_signal_probe): signal quality is a property of the direction call, so measuring only placed orders biases the sample twice — it drops every setup too small to clear the exchange's contract minimum (a systematic subset, not a random one), and it couples the evidence supply to the very risk factor the evidence is supposed to govern. That coupling deadlocked live on 2026-08-10: continuation measured no edge → risk cut to the floor → the resulting $7.49 notional fell under the $10.24 contract minimum → order rejected → no probe recorded → with no new probes the family could never earn back the evidence that would restore its size. Probes are now written when the call is made, before any sizing or RR rejection can discard it, into a dedicated bucket retained by count. Risk can fall as low as the measurement warrants without ever starving the measurement.

  • Making the alternative playbook findable (entryMap.fadeSetup): unblocking the fade gates was necessary but not sufficient — over the first 39 measured signals 38 were continuation and one was a fade. The model was not withholding fades; nothing at the point of decision told it one was on the table. Every piece of guidance the analysis emits (entry_hint, the entryMap note, the ATR-extension framing) is written for arriving at a good price on a trend trade, and the regime label reads trending on ~93% of symbols so the mean-reversion hints keyed on ranging essentially never fire. analyze_market_context now reports a fadeSetup block whenever 15m RSI sits at an extreme, naming the direction and stating that setup_family='fade_extreme' must be passed or the alignment gates will reject it. It asserts nothing about whether fading works — signalEdge.by_family keeps that score and risk follows the measurement. It exists so the hypothesis can be tested at all, which at one probe in 39 it could not be. A test pins the hint's thresholds to the gate that admits it, so the analysis can never advertise a setup the gate would then refuse.

  • The family allocator is evidence, not caution — so it sits outside the soft floor: size_quality_floor exists to stop several independent "be a bit cautious" opinions (regime, conviction, loss-streak) from compounding a real edge into fee-dust. The family factor is not an opinion — it is the measured forward return of that playbook against the cost it must clear — so applying it inside that stack clamped it: measured live on 2026-08-10 the continuation family sat at −0.29% net against a 0.166% hurdle, yet no matter how bad the evidence got the combined soft factor could not fall below 0.50. The one signal grounded in data was capped by a guard built for the ones that are not. It is now a separate multiplier, and the penalty is proportional to the shortfall in units of the cost hurdle (1.75× there), so there is no tuned constant and a marginal miss is treated differently from a deep one. The floor is deliberately non-zero (0.25): size is stop-defined, so driving it to nil pushes notional under the exchange contract minimum, the order is rejected, no probe is recorded, and the family could never earn back the evidence that would let it recover — the same doom loop the memory-retention fix had to undo.

  • Setup families and measured allocation: every entry declares a setup_family — continuation, fade_extreme, breakout, range_edge — and signalEdge.by_family scores each playbook on its own forward return versus costs, with risk flowing toward whichever currently pays (edge.family_size_factor, which never enlarges risk). Nothing in the code decides that trend-following or fading is correct; that is the market's call and it changes. A family reading no edge over a real sample is shrunk; a family reading insufficient data keeps full risk precisely so it can earn the evidence that judges it. (Superseded: since Aug 12 an unproven family trades at the explore size, and since Sep 25 each bet is judged per side — see "Judged per side, and told the truth" above.) This replaced a structural blind spot: the bot could only ever express continuation, because its regime label read trending on 68 of 69 recorded entries, the daily gate blocks counter-daily entries, and the 1h gate blocks anything the 1h opposes — which a fade has against it by definition. Meanwhile continuation measured −0.017% gross over 3,408 samples on the live universe: flat, and flat does not cover a 0.10% round trip. allow_fade_extreme now lets a declared fade past the alignment gates at a textbook RSI 30/70 extreme against the entry direction, so the alternative can at least be measured — deliberately without any starting advantage, since the same test showed fading positive in both halves but only at t≈1.0–1.6 over 135 independent events. Suggestive, not established, and not something to hardcode.

  • Zero edge → zero stake (stand aside) (edge.family_stand_aside, Aug 12 2026): family_size_factor floored a no-edge playbook at quarter-size rather than nil — a nil stop-defined size fell under the contract minimum and, when probes were recorded only on placed orders, starved the family of the evidence that would restore it (the doom loop above). That loop no longer exists: probes are recorded at the direction call, so a skipped trade still feeds the measurement. With evidence decoupled from execution the floor is free to fall to its mathematically-correct value — a playbook whose mean forward return does not clear its round trip has non-positive expectancy on the signal itself, and the growth-optimal (Kelly) stake on a non-positive-edge bet is zero. So the bot now declines an entry whose family reads no edge over a real sample (n ≥ 20) instead of staking floor-size fee-dust on it, and re-opens it automatically the moment its forward return beats cost. This is bet-sizing (survival), not a directional veto (opportunity): it acts only on the bot's own measurement of its own calls, never on which coin or direction is right. Replayed over the last 20 live closes it skips exactly the 9 continuation entries (that family measured no edge, n=44), sparing +6.64R, while leaving fade_extreme (insufficient data) and every other playbook free to trade. Off-switch: EDGE_STAND_ASIDE_NO_EDGE_FAMILY=false.

  • Corruption must be skipped, not made a reason to forget the history behind it (_prev_day_point + the chain-gap guard in dashboard_publisher, Sep 10 2026): the equity index compounds — indexClose_today = prevDayClose × (1 + intradayReturn) — so one bad point is permanent. An earlier heal handled that by re-anchoring to index_base (100) whenever the previous close was outside the sanity band. It healed the arithmetic and destroyed the record. Live sequence: day 20680 closed at 84.17 (the account genuinely down ~15.8%), the server was down 13 days, the restart wrote ~7.27e7 on days 20693–20695 (before the 50% step guard existed), and the next publish re-anchored to 100 — so day 20696 opened at 100.55 and every later point compounded off it. The published curve read −0.27% since June on an account actually down ~16.5%. The last sane close was sitting one row back the whole time, so _prev_day_point now skips corrupt closes and resumes from the most recent good one, falling back to base only when no sane close exists; it also returns that row's day, so the chain-gap guard measures the real gap rather than a one-day hop off a corrupt row. Separately, a gap longer than _MAX_CHAIN_GAP_DAYS (1) holds the index flat instead of booking the whole outage as one day's return — days the system was not trading produced no trading performance. A single missed publish still compounds, with the 50% step guard as backstop.

  • A guard on one side of a two-sided pipe is not a guard (SandBox/TrAIde/azurestorage.py, Sep 10 2026): the publisher's corrupt-point filter lived only on the write path, so the blobs it produced were clean — while the dashboard's own /traide/equity route queried the durable Azure table directly and served the raw rows straight to the chart. The live page was pinned to a y-axis of [0, 8e7] by three corrupt points, flattening four months of real history into a line along zero with one spike, while live.json was provably healthy. Azure history is never rewritten, so rows that predate a guard survive it forever: the same sanity band now applies in the reader too.

  • The macro calendar refreshes on code's schedule, not the model's memory (regime.macro_calendar_refresh_reason, Sep 23 2026). The dashboard showed "Calendar last refreshed 67h ago", and the stored calendar held a single event — NFP on Oct 2 — while CPI (Oct 13) and the FOMC decision (Oct 28) were missing. The refresh had been left to the Research Agent, but research only runs when forced after several runs with no trade, and forcing is blocked while positions are open. In a winning trend the book is almost never flat, so the better the bot traded, the blinder its macro guard became — zero research runs, zero refreshes. Code now decides: the calendar needs refreshing when it was never fetched, is older than 24h (the schedule is published a year ahead, so daily is enough — a property of the data, not of any market regime), or has no upcoming events. That triggers a calendar-only handoff which, unlike the whole-market overhaul, is allowed with positions open: one lookup and one tool call, never touching the coin list, so neither harm the flat-only rule prevents applies. A failed attempt retries at most every 6h. The Research Agent now sees what is stored (age + upcoming events) without calling the overwriting tool, and is asked for six weeks ahead instead of two — FOMC meets roughly every six weeks and the rest are monthly, so that window always holds at least one of each.

  • A restart must not silently disarm the exit rules on an open winner (initRiskPx / peakFePx in position_context.trade_context — main._trade_context until Sep 25 — seeded in ProtectionManager.run, Sep 22 2026). Audit of what a deploy + restart can lose: the memory file, .env and log are git-ignored; memory writes are atomic (tmp + fsync + os.replace, so a kill mid-write cannot corrupt); the wake-trigger baselines, taker-flow state and Telegram conversation memory are on disk; open positions and brackets live on the exchange; pending-entry expiry is measured from the order's own timestamp; and an in-flight agent run is cut by gunicorn's 30-second graceful stop with the exchange reconciled from persisted fill/close ids on the next poll. The one thing that lived only in the process was ProtectionManager's per-position state — and one piece of it mattered: _init_risk, the original stop distance that keeps every R-based rule anchored after a stop has been ratcheted to breakeven. It was captured only when a stop below entry was seen, so a winner already at breakeven before a restart could never regain it: risk_override=None, and the trail, breakeven and early-cut rules went silently inert for the rest of that position (the exchange stop still capped the loss — survival held — but the run gave back). The per-poll trade-context lookup now also returns the recorded entry stop distance and the lifecycle's recorded peak; the manager seeds its anchor from the first and can only raise its peak from the second (keyed on openTime:side, so a previous lifecycle on the same symbol never leaks in). Live values win whenever a real below-entry stop is visible. Tested against the real run() — a breakeven winner regains its anchor and the trail ratchets again — and mutation-checked both ways. Correction (Sep 25 2026): the peak half never fired. Its lookup keyed on KuCoin's raw positionSide (BOTH in one-way mode) while the recorder stored long/short, so no key ever matched — and the value it would have seeded was peakPnl / contracts, i.e. price × contract multiplier (10× on H/WIF/ONDO, 100× on G/F, 520,000× on PEPE). Fixing only the key would have market-closed multiplier>1 winners (~35% of closes) on ordinary polls. Now the recorder keeps a price-space peak (peakFePx, from markPrice − avgEntryPrice) keyed on the same identity ProtectionManager resets on — (openTime, side from qty, |qty|, avgEntry) via memory.peak_fe_key — so an add-on, partial reduction or flip resets it too; the lookup (moved to position_context.trade_context) derives side from currentQty and returns the peak only on an exact key match; and the manager applies it only on the poll it first sees a lifecycle, never per poll. Realized value so far is 0R (the ratcheted lock already sits on the exchange across a restart); the value is that a restart is now invisible to the trail — tested by driving the real writer, lookup and run() over one price path with and without a restart and requiring identical decisions — and the landmine is gone. The early-cut age clock still restarts with the process; with the peak restored that can only delay a cut, never add one.

  • A flip must not stamp the closed side with the new side's lifecycle (tools.lifecycle_for_close, Sep 22 2026): when the model logs a realized close itself (close_long / close_short with a pnl), the lifecycle fields were read from the current position book. On a flip — close the long, open a short in the same run — the book already holds the new short, so the long's close row was stamped with the short's openTime and side. That row can never match its own fill: no entry/exit prices, an empty dashboard card, and no exit probe, so the model's own close was invisible to the very scoreboard built to judge it (H-USDT, 06:04). The lifecycle is now attached only when the live side is the side the action closed; otherwise the poll-loop reconciliation, keyed by the exchange's own lifecycle, owns the row.

  • The trailing stop's record is now split by the entry's regime (regime on exit probes, exitDiscipline.trailByRegime): the trail exited 30 winners early in the Sep 19–22 rally, worth −10.77R against their brackets — the same trend leak seen a week earlier. Every one of those exits is tagged trending; there is no chop-regime row yet, and in chop the same tight trail saved ~35R. A regime-adaptive trail is the obvious fix and is deliberately not shipped until both rows exist — a widening tuned on rally data alone is exactly the one-regime constant this project refuses to carry. The tag is stored on every probe from here, so the evidence accumulates on its own and the dashboard shows the split.

  • Market state is recorded on every entry and probe (Sep 25, analytics.market_state, marketState on entryContext, signal probes and exit probes). The per-symbol regime tag read trending/strong on 164 of 188 closes and the BTC daily bias bullish on 186 of 188, so no stored tag could tell the Sep 17–22 rally from the August chop — and the trail's chop row above could never fill. The block is raw numbers only: breadth24 (share of the liquid, mature crypto perps — the screener's own SCREENER_MIN_TURNOVER_USD_24H + listing-age universe, no hand-picked basket — that are up over 24h), basketMedian24h, btc24h, btc72h, and BTC's daily ADX next to its bias (ADX is the chop marker the bias lacks). The poll loop refreshes it at most hourly (main._MarketStateClock, ~3 public calls an hour, after the profit-lock and probe settlement); the order path, the agent and the exit recorder only read the cache. It feeds two report-only splits keyed on rolling breadth terciles (the market's own recent distribution, not fixed edges): exitDiscipline.trailByMarketState (dashboard — with the same trail exits also split by the entry's BTC daily ADX terciles, nested as byBtcDailyAdx: breadth measures direction, so an August-style high-breadth chop spreads across every breadth tercile, and only the ADX split can isolate the chop row the adaptive-trail decision waits for) and signalEdge.by_market_state (Supervisor's get_edge_scoreboard, "insufficient data" below the family sample bar). The model sees only the current reading (marketState in its state, with a legend) — no per-state record and no size factor: outside that one rally, high-breadth trades lost like low-breadth ones, so a "your R in broad tapes" line would have trained the model on one episode. The EMA-50 breadth over a 29-coin basket from the original finding was dropped: both reviews showed breadth24 separates the same episodes at zero extra calls and without a coin list or EMA constant.

  • An early close is scored against the system, not the bare bracket (protection.replay_protection_stack, main._score_exit_probe_stacks, exitDiscipline.stackR, Sep 25 2026). exitDiscipline compared each of the model's early closes with the bare bracket — original stop and target, or an 8h mark — but the system never leaves a position on its bare bracket: breakeven, the noise-band trail and the carry hold manage it every poll. On the 5 attributed agent closes at the time the scoreboard read −6.23R vs the bracket; replaying the real stack gives about −3.44R. All of the gap was two trades (DASH, INJ) whose trail would have exited near +0.3..+0.5R long before the TP the bracket credited. The bias is regime-signed, not a fixed overstatement: in a trend the trail leaks against the bracket, so bracket-only makes closes look costly; in chop a run that armed the trail and then reversed scores −1R on the bracket but ~0R on the stack, so bracket-only flatters closes. Now each exit probe records what a replay needs (fillTs, initRiskPx, noiseBandR, holdUntilTs), and once a close's 8h horizon has passed the poll loop (at most 2 per poll, after the profit-lock and the settle step) fetches 1m futures bars from the fill, checks every bar's column order, and replays the exchange legs plus decide_protection with ProtectionManager's effective config (its breakeven fee raised to the round-trip cost, not the raw .env value). Replaying from the fill also rebuilds any ratchet the stop had made before the close, which the bracket probe ignored; a replay exit before the model's actual close is skipped (the live stack demonstrably had not exited), with the stop/peak state carried on. Agent closes are scored against stackR only — one comparator in the verdict: a close still waiting for its replay is left out (stackPending) so the verdict never flips between benchmarks, one whose replay was given up after a day is left out too (stackUnavailable), and older rows with only a bracket result are shown as audit (legacyBracketScored / legacyDeltaR) and not counted (Sep 25 review: falling back to the bracket for them blended two benchmarks — 5 legacy + 3 stack rows read 'closes destroy value' where the 3 stack rows alone were ~0 — and would bake the bracket's regime-signed bias into the next model's first verdict). The trail's own exits (otherExits) stay on the bracket — trail vs bracket is exactly the question there — now split by how the bracket resolved (take_profit / stop / expired), because the sign of that gap depends on the path. Known bias: 1m last-trade prices vs the live mark, ~0.05R median. Measurement only: no gate, no constant beyond the existing 8h horizon.

  • "Fresh" now means changed since entry (entryThesis on each futuresPositions row via position_context.entry_thesis; STEP 1b; counterAtEntry / htfAligned on exit probes, Sep 25 2026). STEP 1b said "if fresh 15m AND 1h biases both oppose the position, treat that as a confirmed intraday reversal" — without saying fresh compared with what, while the model's position view was raw exchange data. So it fired on the entry premise of counter-intraday playbooks: of 31 fires since Aug 10, 27 were already true at the fill (DASH and KCS were funding_carry longs entered with 15m/1h bearish and closed on that same condition). Each open futures position now carries the entry's own record — playbook, entry biases {15m, 1h, 4h, 1D}, planned stop/target, currentR on the original risk, noiseBandR, and for a carry the settlement the code is holding for (on the contract's own clock), minutesToSettlement, fundingRateAtEntry vs fundingRateNow and carryPerSettlementR (what one settlement pays, in R — it shows how small the transfer is against the stop). STEP 1b now says a reversal is a change since entry, that an opposition already present at entry is not new evidence by itself, and asks the model to name what changed before a reduce-only close. Deliberately neutral: the rule is kept for genuine flips (the 4–5 genuine post-entry flips weakly favoured closing), and nothing nudges toward holding carry trades — the only other instances of the pattern (G-USDT, XMR shorts) were closes that helped. The model's own record decides: exitDiscipline.byCounterAtEntry and byHtfAligned split its closes by those entry tags, shown with n before any verdict.

  • Importing the app must not start live trading (src/wsgi.py, Sep 20 2026): wsgi.py starts the trading loop and the Telegram supervisor at module import, deliberately — gunicorn imports the module and never calls application until traffic arrives, so the loop has to start without an HTTP hit. But "on import" means any import, and tests/test_module_imports.py imports every module in src to catch partial commits. Verified by counting threads: importing wsgi under pytest spawned Thread-1 (_start_background_loop) and supervisor-bot — a live trading loop and the Telegram bot against the real account, on every test run. A daemon thread cannot finish an agent run inside a short test, but ProtectionManager places and moves real orders well inside that window. The auto-start now skips when a test runner is detected (pytest in sys.modules, or PYTEST_CURRENT_TEST set); TRAIDE_WSGI_AUTOSTART=0|1 forces it either way. Production is untouched — and tests/test_wsgi_autostart.py checks both directions in a subprocess with the loop stubbed, because a guard that also stops gunicorn would silently end trading after a deploy.

  • Dependency refresh + the LangSmith deprecation (Sep 20 2026): LangSmith flagged calls to an endpoint it retires on 2027-01-31. That endpoint is POST /runs/batch; current clients post to /runs/multipart and fall back to the old one only if multipart 404s (and then permanently, for the life of the process). requirements.txt pinned no versions, so the server was running whatever was current at its last install. All direct dependencies are now floor-pinned to versions verified together — openai 3.16.2 (major bump from 2.x), openai-agents 0.22.3, langsmith 0.13.0, pandas 3.0.6 (major bump from 2.x), gunicorn 26.2.0, requests 2.34.2, azure-data-tables 12.7.0, azure-storage-blob 12.30.2. Verified by capturing the client's actual HTTP calls: 1× POST /runs/multipart, 0× /runs/batch. Also checked beyond the test suite, since tests do not exercise the live wiring: the Azure client still keeps its custom deployment-miss retry transport, OpenAIResponsesModel still wraps it, all 48 tools still compile to strict JSON schemas, every src module imports, the KuCoin client still returns 7-column futures candles, and every pinned package still supports the server's Python 3.12 (numpy is the tightest floor at >=3.12).

  • The exit scoreboard judges the model only on closes the model made (memory.note_agent_close, closedBy on exit probes, Sep 19 2026): the scoreboard used to score every exit that landed between the original stop and target — and a trailing-stop exit lands exactly there. So the live verdict read "closes destroy value, −12.41R over 30" when, cross-checked against the log, 16 of those exits were the code's own trailing stop (−8.99R) and one was the model (+1.54R — it helped). The model was being told its judgement was destroying value because of a mechanism it does not control, while the prompt told it to let that number decide how readily it closes. A reduce-only order accepted through the agent's own close tool now marks the symbol, the poll loop tags each exit probe closedBy: agent | protection, and only agent rows feed the verdict. The trail's exits are still scored — reported separately as otherExits on the scoreboard and the dashboard — because "the trailing stop left the bracket's target on the table" is worth knowing. In this rally it was: the brackets would have made +8.99R more. But that is a statement about the trail in a trend, where the tight trail is wrong, not in chop, where it saved ~35R; it argues for a trail that follows the regime, not a fixed widening. Exits recorded before attribution existed are reported as unattributed, so the model's own verdict restarts from zero.

  • A playbook graduates by degrees, not in one jump (edge.family_explore_factor, Sep 19 2026): a family used to trade at 0.40× while earning its verdict and then jump to 1.00× the instant it reached 20 probes, however weak the evidence. funding_carry graduated at n=24 on net +3.78% with a standard error of 3.20% — t = 1.18, barely distinguishable from noise — and its risk rose 2.5× in one step. The first full-size trade after that was a G-USDT short into a +28.6%-in-25-minutes spike (15:30→16:00, funding +0.21%/8h, every timeframe bullish): stop hit exactly as placed at 16:30, target never touched in the next four hours, −1.04R at 2.2× the size of the same trade an hour earlier. Size now ramps with the family's own confidence: the explore floor at t = 1 (where the stand-aside first releases it, so there is no step at graduation) rising linearly to full size at t = 2. funding_carry at t = 1.18 now sizes at 0.51× instead of 1.00×. The 2.0 is a statement about statistical confidence (two standard errors), not about the market; the family's net and SE — recomputed every run from live probes — do all the adapting, so a weakening edge shrinks the position automatically and a strengthening one grows it.

  • Data-integrity note — KuCoin futures and spot candles use different column orders. Spot klines are [time, open, close, high, low, volume, turnover]; futures klines are [time, open, high, low, close, volume, turnover]. The live analyzer converts correctly (tools.py, where futures rows are re-ordered before candles_to_dataframe), so indicators, ATR and stops were never affected. But several offline replays read raw futures rows in spot order — treating the low as the high and the close as the low — which silently tests stops against the close and misses most stop-outs. Re-checked: the Sep 2 XMR conclusion survives (stop hit first either way); the "agent's early closes cost ~2.6R" conclusion reverses (see the exit-discipline entry). The +13.05R marketable-entry replay below also came from 1-minute futures paths; re-run Sep 24 2026 on a new 124-limit sample with correct columns, its direction held but its RR-filter and lease-length sub-claims did not reproduce (see the marketable-entries entry). Verify with a one-liner: across real candles the HIGH column must be the per-row maximum.

  • Score each playbook at the horizon it is actually held, and let revealed trends breathe (edge.family_scoring_horizons, the runner branch of the trail, Sep 19 2026). Reconstructed against real candles for all 30 watchlist symbols across the Sep 17–19 rally (a CFTC-proposal-driven move: BTC +5.5% through $80K, HYPE +12%, SOL +10%, UNI +28%). The bot made +8.36R on 20 closes at an 85% win rate — but the rally was far bigger than that: ONE +130%, G +116%, F +67%, INJ +37%, ENA +33%, APT +32%, ARB +30%, OP +27%. Two distinct leaks, both measured rather than assumed:

    • The trend playbook was benched on the wrong clock. Every family's verdict was scored at one shared 60-minute horizon, but continuation is held a median 162 minutes. In a strong trend an entry is often mid-pullback at the 60m mark, so its verdict hovered at net −0.006% over ~110 probes and the stand-aside benched it for the entire second day of the move (100% of its entries refused from Sep 18 15:00). Those 143 blocked calls measured — by the same probe method the gate uses, from the market price at the call, no hindsight — +0.17% at 60m and +1.27% net at 240m with a 76% hit rate. Each family is now scored at the settled horizon nearest the median of its own most recent 20 realized holds — recent, so the horizon follows the market instead of averaging every regime the bot has seen (funding_carry's hold went 90m → 171m between chop and rally) — snapped on a log scale (ratios matter for time): continuation → 240m, fade_extreme and range_edge → 15m, funding_carry → 60m; a family with fewer than 6 realized closes falls back to the default. It is per-family on purpose: a blanket 240m would have released fade_extreme (+0.77% at 240m) on evidence from a holding period it never uses — at its own 15m it correctly stays benched (−0.30%).
    • Sep 26 2026 — snapping the median was itself a knife-edge, and the prompt could freeze it. Continuation's recent holds are bimodal (fast trail exits ~60m, runners ~4-14h), so its median sat at 118–141 minutes — right on 120m, the log midpoint between the 60m and 240m horizons. One fast +0.39R FET winner moved the median 125.5 → 118.5, the horizon 240m → 60m, and continuation longs from +1.1% net to −0.15%: zero stake, in the middle of a broad alt rotation (93 of the CoinDesk 100 up on Sep 25). The continuation longs the model then declined ran +0.05% at 60m but +1.44% at 240m (n=10) — the family was being judged on the wrong clock again, this time by a rounding step. Worse, a benched family makes no closes, so nothing could move the median back: the faster the trail banked winners, the more firmly the winning playbook stayed benched. Each family with ≥6 recent closes is now scored over its holding-time mix (edge.family_horizon_weights): every one of its last 20 holds votes for its nearest settled horizon, and the family's rows mix those horizons by vote share (edge._mixed_row) — each horizon keeps its own de-overlapped sample, the row's n is the thinnest leg and its standard error the weighted sum of the legs' errors (an upper bound, so mixing never overstates certainty); all weight on one horizon reproduces that horizon's row exactly. One close now moves the verdict by one twentieth, not all of it. Rows carry horizonWeights / nByHorizon. This does not manufacture an edge: over its real 55/45 mix, continuation's recorded calls read t≈0.7 — borderline — so the family re-opens only when new calls show the market paying. Which exposed the second half: once the stake was shown to the model before proposing (Sep 25), it stopped submitting the benched side and declined it instead (21 declines citing the stand-aside in a few hours) — and a decline records nothing, so the verdict could never see the rally. The prompt and the refusal hint now ask for the opposite: submit a genuine, honestly-labelled setup even on a stood-aside side (the call is recorded before the refusal; once per setup is enough), and never 'stand down until the regime turns' — for a benched side, new calls are the only way the bot can see the regime turn.
    • Investigated, deliberately NOT shipped: a wider trail for trend winners. After exit, price ran a further +11.5% on average, and the trend-adaptive is_runner flag turned out to be dead code while the trail is on — it only ever fed the give-back branch, which runs only when trailing is disabled. Widening the trail to a fixed 2× noise band once a trade reached 2R was non-negative on a 1-minute replay, but its whole benefit was one trade in twenty (+0.85R), and the 2× was a number chosen from a rally-only sample. A market-derived version — never trail inside a pullback the trade has already survived — scored identically to doing nothing, for a structural reason: it can only widen after a deep pullback, and the trade that benefited was stopped on its first. A fixed constant tuned on one regime is exactly what this system is meant to avoid, so it was reverted. The dead-code finding stands and is worth revisiting once there is evidence from more than one regime.
    • Not changed, and why: the take-profit cap is the larger remaining lever (most trades hit TP at 1.6–2.7R before becoming 2R runners), but target placement is the model's opportunity decision. The 25 RR-floor rejections were all near-misses (1.19–1.49 vs 1.50) that the model can clear by re-planning the bracket. Caveat: per-family retention had already evicted continuation's chop-era probes, so the horizon change could not be back-tested against a chop regime — it is justified by measuring each family at its real holding period, not by a chop replay.
  • funding_carry is measuring the POSITIONING half, not the transfer (Sep 15 2026): the family finally produced the book's best trade — CAP-USDT long, +1.81R / +$0.72, closed at target in 90 minutes — and funding_carry now reads n=18, net +1.04%, the only playbook of five that is not benched. But check why before believing the mechanism: 3 of the first 4 carry trades resolved at their bracket BEFORE any 8h settlement and collected exactly zero funding, the +1.81R winner included. Only 0G-USDT (+0.13R) ever crossed a settlement. So the evidence supports the playbook's secondary rationale — an extreme rate marks a crowded book prone to unwinding — and says nothing yet about being "paid to hold". The cost-derived threshold remains a sound filter for how extreme is extreme; the mechanism story does not. Both the funding_carry_setup note and the agent prompt now state the measured record instead of the intended one, so the model does not over-trust a transfer it is not collecting. carry_hold_deadline stays — it only suppresses the code's own early profit-taking, and a bracket resolving on its own terms is a legitimate exit, not a leak.

  • A directional gate must not judge a non-directional playbook (timeframe-conflict hatch in tools.py, Sep 14 2026): the fifth freeze of the same shape. 31 runs, zero orders. The only two qualifying carry setups on the entire watchlist were CVC (−0.32%/8h) and CAP (−0.19%/8h) — negative funding, so the long is paid — and the timeframe-conflict gate refused buy on each 11 times, because a 15-minute candle was bearish. But the 15m bias is a directional test and funding_carry explicitly does not claim to pass it: the transfer happens whichever way price moves. That branch was the one futures gate with no escape hatch at all, while the 1h gate directly above it has four. It now takes the same allow_mechanical_setup + mandatory verify_declared_setup carve-out the daily-exhaustion gate got, so a verified carry (funding clears the cost-derived threshold and pays the side being entered) passes, while a bare declaration still does not. Testing the pure function does not catch this — allow_mechanical_setup is correct either way; what breaks is the wiring, so tests/test_regime.py now asserts in source that every futures direction-gate branch offers the hatch and that every hatch is paired with verification.

  • One bad backend must not take the bot dark (agent._RetryDeploymentMissTransport, Sep 14 2026): the model endpoint is an APIM load balancer (.../abopenailb) fanning out over several Azure OpenAI backends. One backend lost the configured deployment and the tape shows the consequence precisely — calls alternated 200 / 404 (10 vs 8), and because a single 404 aborts an entire agent run, every run from 15:10 onward failed. The bot stopped trading completely on a fault affecting roughly half of one pool. The OpenAI SDK retries 408/409/429/5xx and connection errors but treats a 404 as permanent — correct for a single endpoint, wrong for a pool, where "this backend has no such deployment" is a per-backend condition. A narrow transport now retries only the Could not find an existing deployment 404, a bounded three attempts, re-dispatching through the balancer so the retry usually lands on a healthy backend. Any other 404 (bad route) passes straight through un-retried, and a deployment genuinely missing from every backend still surfaces the error and logs it as a configuration problem — that one must stay loud, not be papered over.

  • A benched playbook must not evict a trading one's verdict (memory._trim_probes_per_family + MAX_PROBES_PER_FAMILY, Sep 14 2026): a stood-aside family still records a probe on every direction call — deliberately, since that is the only way it can earn its way back (see record_signal_probe). The consequence nobody had costed: the benched family is usually the loudest. continuation, benched for weeks and never trading, held 283 of 479 retained probes (59%), and under a single global ring buffer it evicted the history of the families that were trading. fade_extreme's decimated sample fell from 27 to 17 — under the stand-aside's min_samples=20 — so its verdict flipped from no edge to insufficient data, and that releases a family to full size. It took two trades on the next two days and lost both (−0.59R, −0.46R). Reading the full retained set gave n=27, net −0.33%, no edge; reading the 400-row cap gave n=17, net +0.01%, insufficient data. The cap itself un-vetoed it. Two fixes: retention is now bounded per family (families are a small fixed set, so the file stays just as bounded) and the verdict readers ask for limit=0 — everything retained — because a second truncation at read time can only discard evidence the writer deliberately kept. The principle both serve: losing the evidence for a verdict must never be equivalent to never having had it. Note the union in signal_probes returns MORE rows than MAX_SIGNAL_PROBES (it folds in legacy trades-derived rows), so passing that constant as a read limit silently dropped the oldest of them.

  • The stand-aside needs hysteresis, and it needs all of the evidence (edge.family_stand_aside, Sep 4 2026): the family verdict is a sign test on a noisy mean, so a playbook parked near the cost hurdle flips between polls on the same evidence — and a flip to edge restores full risk (family x1.00) to the playbook with the longest adverse record. Live: continuation sat at net −0.03% against a standard error of ~0.30% over 130 samples; it was stood aside on four consecutive WIF-USDT attempts (family=0.71 / 0.26 / 0.25 / 0.44), then on the fifth poll read family=1.00 with no stand-aside, the long filled, and it lost a full −1.01R. Two fixes. (1) Hysteresis: entering the stand-aside still needs only the sign, but releasing it now requires the mean to clear cost by at least one standard error of the family's own sample — signal_edge_stats reports stderr_pct per family for this. The band is the sample's own dispersion, so it tightens automatically as evidence accumulates: a family that genuinely starts paying still escapes, it just has to do so by more than the measurement's own uncertainty. (2) Use the whole retained probe set: the entry path read signal_probes(limit=200) while MAX_SIGNAL_PROBES keeps 400 — throwing away half the evidence for free, and making the verdict swing with the window. On that day continuation read no edge at every window size except 200, which was the one the bot actually used. The dashboard now reads the same window, so the published verdict cannot disagree with the bot's own. (Correction, Sep 25 2026: what shipped is not hysteresis. There is no state — the rule is one stateless bar, stake zero while net < its own SE (t < 1), recomputed every run, which can flip back and forth near the line. The WIF case above is still caught by it; see "Judged per side, and told the truth".)

  • A refusal has to say where the edge went (edge.open_families, Sep 7 2026): the stand-aside handed back a hardcoded suggestion — "switch to a playbook that is currently paying, take a genuine fade_extreme at an RSI extreme". That advice went stale silently and then started pointing at a family that was itself stood aside: on Sep 7 fade_extreme sat at n=38, net −1.12%, so following the hint was guaranteed to be refused again. Live for the six hours after the probe outage below was fixed: 21 continuation stand-asides, the model re-proposing the one family it had just been refused, while range_edge sat unblocked at a measured +0.72% over n=22 and drew a single probe. The hint is now computed from the same scoreboard the refusal is: open_families names what is currently paying (best net first) and, separately, what is open but still gathering evidence — a family that is stood aside can never appear in either, and with nothing open it says so rather than inventing somewhere to go. This adds no gate and changes nothing about what may be risked; it is the "surface better wisdom instead of another veto" principle applied to the one message the model reads at the moment it is looking for an alternative.

  • A declared playbook with an objective trigger must actually show it (regime.verify_declared_setup, Sep 2 2026): the declaration carve-out below is right for breakout/range_edge — they are judgement calls with no computable trigger, so the model's word is the only evidence there is and rightly decides. funding_carry and macro_event are different in kind: both have a trigger the code evaluates itself. Accepting those on the label alone turned the carve-out into a universal gate-bypass token, and it cost a live trade. XMR-USDT, 2026-09-02: funding was 0.0523%/8h against a 0.0700% threshold derived from measured round-trip cost, so funding_carry_setup returned None and FUNDING CARRY AVAILABLE had never once been logged — yet a short was declared funding_carry, admitted past the 1h gate on the strength of the label with 15m/1h/4h all bullish against it, and closed at −0.6R. The carry it stood to earn was 0.0523% against a 1.89% stop — a 1:36 ratio; even a full day of holding earns 0.157%. It was a directional short wearing a carry label, and the funding_carry scoreboard was being scored on it. Now a declared funding_carry must have funding that clears the derived threshold and be on the side that actually receives the transfer (positive funding pays the short; a long declaring carry would be paying it away), and a declared macro_event must sit in the window after a scheduled release. This decides nothing about whether the trade is a good idea — that stays the model's call — it only says a mechanical playbook must have its mechanism present. The same trade may still be taken under the family it really is, where it meets that family's ordinary gates and is scored honestly.

  • Anti-FOMO refuses a bet on the trend; it should not refuse a bet on a mechanism (regime.allow_mechanical_setup, Sep 11 2026): the daily-exhaustion gate exists to reject one specific trade — "this has run to RSI 80 and I want to bet it runs further" — but it identifies that trade by side versus daily bias alone, so it cannot tell it apart from one whose payoff has nothing to do with the trend. It charged the full veto against a funding_carry, whose return is the 8h transfer, collected at settlement, whichever way price moves. RAY-USDT, Sep 10–11 2026: a ~96%-in-a-week rally on tokenized-stock volume pushed daily RSI to 83 while funding went to −0.1877%/8h paid to longs — against a 0.14% round trip, one payment covers more than the whole cost. The agent proposed the carry long ~54 times over two days; every one was refused "Daily exhaustion: bullish trend overextended — no continuation entry", and it correctly never retried or relabelled. The other side was no escape either, since fade_extreme is stood aside as a measured no-edge family. With a market-wide melt-up making daily bias bullish and overbought at the same time (BTC 14d RSI ~80, 56% of alts back over their 200DMA), that intersection shut both halves of the daily gate: the poll loop kept running ~190 times a day while orders went 15 → 2 → 0. Structurally the branch had no way out — the gate below it, for an entry opposing a healthy daily, has four escape hatches, while the exhaustion branch had exactly one (allow_trend_aligned_short, which needs a bearish daily and a sell), so a buy into an exhausted bullish daily had no path whatever it declared. Now funding_carry and macro_event may pass, and only those two: this carve-out is deliberately narrower than the declaration carve-out below, because there the thing being bypassed is a direction gate and direction is the model's call, whereas here it is an extension gate and "I promise this isn't FOMO" is exactly the claim a label cannot be trusted to make — accepting it would leave anti-FOMO one word away from disabled. breakout and range_edge stay blocked (a breakout long at RSI 83 is the continuation bet). Verification is mandatory, not optional: verify_declared_setup must find funding that clears the cost-derived threshold and pays the side being entered. Survival is unchanged and downstream — probe recorded at the call, explore-sized while unproven, stood aside at no edge. Widens what may be proposed, never what may be risked; disabling the family in DECLARABLE_SETUP_FAMILIES disables this too.

  • Every playbook can reach the book — measurement, not a trend gate, decides which gets capital (regime.allow_declared_setup + edge.family_explore_factor, Aug 12 2026): the alignment gates only admit trend-aligned entries freely, and a trend-aligned entry is tagged continuation — so continuation was the only family that could ever reach the book. It accumulated ~every probe, measured no edge, and the model kept taking it for want of anything else it was allowed to express. allow_fade_extreme opened one escape hatch keyed on an RSI extreme, but breakout and range_edge had none and are never auto-inferred, so a counter-trend one was rejected before it could log a single probe — samples: 0 forever, a chicken-and-egg that kept them permanently "untested". This generalises the fade-extreme carve-out: a deliberately-declared breakout/range_edge is admitted past the daily/1h gates on the model's declaration alone. Which playbook to run is an opportunity call (the model's job); a hardcoded trend gate deciding it is exactly the veto this codebase avoids. Survival stays fully code-governed, just downstream of the gate: the probe records at the call, family_explore_factor sizes an unproven playbook down to explore-size (0.4×) — cheap because a family's forward-return evidence records from the market price independent of our stake, so we learn a new playbook while risking little — family_size_factor shrinks it on any measured shortfall, and family_stand_aside skips it outright once it settles to no edge. A declared breakout that turns out not to pay still logs its evidence, is sized to a quarter, and eventually to zero — by measurement, not by a pre-trade veto. This widens what may be proposed, never what may be risked. Off-switches: DECLARED_SETUPS_ENABLED=false, DECLARABLE_SETUP_FAMILIES, EDGE_EXPLORE_UNPROVEN_FAMILY_FACTOR.

  • Signal-edge measurement (signalEdge in the edge report): the feedback loop the bot never had. Every other statistic measures an outcome, which conflates three different things — whether the direction call was right, whether the fill was any good, and whether the exit was managed well. That conflation is why several rounds of correct exit, cost and sizing fixes did not stop the bleeding. Each entry now stamps marketPriceAtSignal (the live price when the call was made, not the resting limit price), the poll loop settles 1h/4h forward prices from the ticker snapshot it already has, and edge.signal_edge_stats reports mean forward return signed by the traded direction, versus the round-trip cost. Measuring from the limit price instead scores the resting discount as prediction — on this account's data that produced a spurious +1.24%/15m at a 92% hit rate, where the correct measurement from market price was −0.007%. The verdict is deliberately blunt: edge / no edge / insufficient data. A signal whose forward return does not clear costs cannot be made profitable by any exit or sizing scheme — it can only lose more slowly.

  • Learning data ages out by being superseded, not by the clock: realized closes (pnl != None) and filled orders are exempt from the retention_days cutoff and retained purely by count (MAX_CLOSED_TRADES=200, MAX_TRADES=100). They are the training data for every adaptive guard — edge_stats, entry_quality_stats, expectancy sizing, the symbol bench, measured slippage and the adaptive stop floor all read them. Time-pruning them created a doom loop, measured live on 2026-08-06: as the trade rate fell the 7-day window emptied until only 8 closes remained, all recent losses. The controller then reported an 11% win rate and a 6-loss streak, halved position size, and the agent stood aside in 356 of 358 runs — which produced no new closes, so the window could only get staler and bleaker. A quiet spell must never be self-reinforcing. Declines and unfilled orders stay ephemeral.

  • Self-calibrating friction (SLIPPAGE_AUTOTUNE_ENABLED, default ON): every RR gate, net-profit check and fee-adjusted breakeven prices friction as fee_rate + ESTIMATED_SLIPPAGE_PCT per side, so a constant set once and never revisited silently becomes the strategy. The live config assumed 0.10%/side while measured entry slippage was 0.008% mean / 0.025% p90 — a ~12× overstatement. Round-tripped that is 0.32% of notional against a real ~0.08%, and at the account's median risk/notional (1.3%) it charged every setup a phantom 0.18R. To still clear a 1.5 net-RR floor the model had to plan ~2.7R gross targets — which the tape never reached, and which pulled the stop inward to keep the ratio. Overstating costs does not make a bot conservative; it makes it plan trades that cannot win. The estimate is now measured from the bot's own fills (conservative upper percentile, clamped to [0.01%, 3× prior], ESTIMATED_SLIPPAGE_PCT used as the prior until SLIPPAGE_AUTOTUNE_MIN_SAMPLES fills exist) and adapts in both directions if execution genuinely degrades.

  • Risk-targeted sizing (RISK_TARGETED_SIZING, default ON): position size is computed from the risk budget and the stop distance, not merely capped at it. bracket_risk_scale only ever shrinks, so the actual bet was whatever notional the model happened to name — and across the 35 recorded lifecycles realized dollar risk ranged over 18.9x ($0.06 to $1.17 on a ~$68 account), uncorrelated with conviction or outcome. The winners were systematically the small bets and the losers the large ones, so the last 9 trades closed +0.50R but −$0.12: a positive edge in R that still lost money. Conviction remains the size lever, but it acts through the regime/confidence/expectancy multipliers rather than an arbitrary number, so the unit of risk is constant while conviction still scales it. Pairs with the stop noise floor: a wider stop now buys proportionally fewer contracts instead of a bigger loss.

  • Coherent risk fraction: the effective risk per trade is min(RISK_PER_TRADE_PCT, CB_MAX_DAILY_DRAWDOWN_PCT / (CB_MAX_CONSECUTIVE_LOSSES + 1)). Per-trade risk and the daily drawdown stop are not independent settings — at 2%/trade against a 3% daily stop, two losers end the day on an account taking 5–9 trades daily. The inconsistency went unnoticed for months only because realized risk was never actually 2% (it averaged 0.52%), so the drawdown stop never had a chance to bite. Deriving one from the other keeps them coherent with no extra knob: with the shipped defaults that is 0.75%, which absorbs three full stop-outs inside the daily budget.

  • Marketable entries (fill-rate fix): a passive limit resting away from price fills far less often than one near it, and an unfilled limit captures nothing. The old 0.15%-band-plus-confidence rule admitted just 3 of 80 expired plans (against a median required cross of 0.82%), so MARKETABLE_ENTRY_MAX_DEV_PCT is now only an outer sanity bound; the post-cost RR gate still runs at the crossed price, and the atomic TP/SL bracket still attaches, so a marketable fill is never naked. Correction (Sep 25 2026): this entry used to present a Jul 30 replay of 82 expired limits as settled — crossing +13.05R over 80 (+0.163R), waiting longer −0.37R at 240min, and "still clears the RR floor after crossing" as a +0.388R filter vs +0.050R for confidence ≥ 0.80. That replay read futures candles in the wrong column order. On a new 124-limit sample (Sep 20–24, correct columns, replay within ~0.2R of real outcomes) the direction held (crossing ≈ +0.2R per placement) but the RR-filter and TTL sub-claims did not reproduce: crosses the gate passed did +0.16R (n=13) vs +0.22R for those it blocked, and extra fills after the lease were positive at every horizon to 240m. Even the crossing gain is one-regime — rally longs +0.32R (n=83), shorts +0.01R (n=41), fading below zero by Sep 24, and the live model's crosses averaged −0.03R. So none of these figures is in the prompt any more: the RR gate is described as what it is (a fee/payoff guard, not a quality filter) and the model reads its own live execution map instead (below). The 15-minute lease was not retuned on this evidence either way.

  • Entry-planning wisdom (not a gate): Chasing is an entry-timing mistake (opportunity), not a survival threat — the bracket and risk caps already bound the loss — so it is left to the agent's judgement rather than a hardcoded veto (which would fight the agentic design and need tuning). The analysis surfaces an entryMap: how far price sits beyond the 15m VWAP in ATR units (extensionAtrLong/Short) plus the nearest pullback/retest anchors (VWAP, band-mid, prior breakout shelf, fib of the last impulse). The prompt's OPTIMAL ENTRY PLANNING step then makes the agent reason about the highest-EV arrival price — after a vertical spike it rests a limit at the anticipated pullback (avoiding the chase and capturing the next leg) instead of buying the local peak. This improves automatically as the model improves; no threshold to maintain.

  • Strong-trend continuation (don't wait forever for a pullback): a pullback entry is only higher-EV if the pullback actually arrives. On a confirmed strong-trend leader that keeps running without retracing, "wait for the pullback" silently becomes "miss the whole move" — the bot churned unfilled ONDO limits while ONDO trended away. Unfilled entry-limit expiries are now recorded as agent-visible entryExpiries events (previously only logged), so the agent sees its own limits dying on a symbol and, per desk practice (scale-in), takes a reduced-size bracketed marketable continuation rather than re-placing the same never-filling limit — or stands down and rotates. Gated to an intact trend; a rolling-over trend gets no continuation.

  • Post-trade entry-quality feedback (compounding wisdom): On every close, edge.entry_quality_stats derives — purely from data already recorded — how well the entry was timed: avgMaeR (how far price went against the fill before the trade worked, in R), avgEntryExtensionAtr (how stretched entries were vs the 15m VWAP), and betterEntryRate. This is fed back in edgeReport.entryQuality so the agent can see its own pattern ("your last entries kept dipping 0.6R before working — rest limits at the pullback") and self-correct. It's a mirror, not a rule: no restriction is imposed, and the signal sharpens the model's entry timing over time. (Sep 25 2026: "rest deeper" is now checked against the execution map below — a deeper rest only helps on the fills it gets.)

  • Execution map — fill rate and fill quality by resting distance (edge.execution_map, edgeReport.executionMap, Sep 25 2026). The prompt used to steer entries with a fixed band — "target levels within 0.5–1.5 × ATR; farther than 1.5 ATR rarely fills" — plus the unreproduced Jul 30 figures above. Measured over 183 placements Sep 20–24, the fill rate inside the 15-minute lease by passive distance was 62% under 0.25 ATR15, 53% at 0.25–0.5, 33% at 0.5–1, 8% at 1–2 and 0% beyond 2: the band pointed into the 8–33% zone. Now the model reads its own recent limits bucketed by distance in 15m ATRs (marketable crossed / <0.25 / 0.25-0.5 / 0.5-1 / 1-2 / >=2 — units, not tuned thresholds): placed, filled within the lease, distinct ideas, median minutes to fill, the realized R of closed trades entered at that distance (from the closes' own entry context, one observation per entry order), and fillAdjustedR = fill rate × realized R, an unfilled limit counting zero. Placement and fill counts come from the same records as performanceSummary.limitFillRate (one predicate, memory.is_limit_entry_record), so its totals equal that rate by construction; realized R reads the longer closes list, because joining through the ~100 retained placements would leave the far buckets empty. Any mean under 10 samples is spelled insufficient, every mean carries its SE, the window span is shown, and mostReplaced lists the ideas re-placed most (one H-USDT short was re-placed 12 times in 9 hours on Sep 23, 0.9–2.3 ATR15 away, for 0 fills — while entryExpiries showed ×1 each run). Marketable is its own bucket on purpose: in the Aug–Sep chop it was the worst category, so the table never implies one answer. The futures-limit placement response echoes the matching bucket ("this limit rests 0.80 ATR15 from price (0.5-1); your last 22 limits there filled 41% within the lease; fills there averaged −0.16R (n=33, SE 0.11)") and, for a resting limit, the net RR the same bracket would have if crossed now. The prompt keeps only the mechanism — resting farther buys paper RR at the cost of fill probability, so compare fill-adjusted expectancy, not planned RR; the post-cost RR gate is a fee guard, not a quality filter; don't push the entry out to pass it — and no historical R figure. Counterfactual, every call at every depth: the table alone only sees the depths the model actually uses, so it would go blind on deep limits just when a regime (chop, range-edge fades) might make them pay. Every direction call's probe now stamps atr15Pct, the planned bracket, the lease and crossedNetRr (the post-cost net RR of the bracket if crossed at the live price — recorded, never acted on), and once the lease is over the poll loop stamps leaseLow / leaseHigh from 1m futures bars (column order asserted on every bar, one fetch per symbol per poll, written off as None — never back-filled — if the bars do not come within one lease length). executionMap.counterfactual then scores a limit resting 0 / 0.25 / 0.5 / 1 / 2 ATR15 away for every call: filled if the lease's adverse move reached it, outcome = the family-horizon settlement against that fill price, net of cost, in % and in the call's own R — labelled gross, before trade management (probe means ran ~0.1R above the validated managed replay). Validated offline on the 09-24 copy: the counterfactual fill agreed with the real filled flag on 82 of 82 resting limits at their own depth. crossBand buckets the same calls by crossedNetRr (<1.0 / 1.0–floor / ≥floor, per side), report-only: it is the zero-stake evidence a future sub-floor crossing lane would have to clear, replacing a rally-only replay (+17R over 40 crosses, ~23 independent, most from the retired model). Measurement only — no gate, no constant tuned to the rally.

  • Taker flow — closing the aggressor blind spot (analytics.taker_flow_summary, edge.taker_flow_edge_stats, Sep 4 2026): every playbook the bot runs reads closed candles at 15m and above, so it could see what price did and never who was pushing it — no kline carries the buy/sell taker split. KuCoin's public trade tape does: each row is a taker who crossed the spread, with a side. It is now sampled every poll (held positions first, then the watchlist) and the aggressor balance is stamped onto every direction call, because by settle time the tape is hours gone. This is deliberately measurement, not a new gate, and the prior is that it will not pay: order-flow imbalance is real (Cont, Kukanov & Stoikov 2014) but is measured in ten-second buckets and is largely contemporaneous — price moves because of the flow, with the forecastable part decaying inside a minute; the retail crypto version (CVD, taker ratio) has no rigorous public backtest behind it; KuCoin is not where these prices are set; and a ~0.12–0.21% round trip dwarfs a few basis points of short-horizon drift. So it goes through the same machinery every other claim here does — probe → settle → surface → let the model decide — reported as a spread (with-flow forward return minus against-flow), which cancels the book's own directional bias, at 5m/15m/60m/240m (the two short horizons added precisely because that is where flow information is documented to live). Two staged verdicts keep a real-but-unusable effect from being traded: informative clears its own standard error, tradable also clears the round trip. Nothing in the trading path reads it; the model sees edgeReport.takerFlow only once the sample can answer, and the dashboard publishes both the live tape and the running experiment. Two guards keep it from corrupting the existing scoreboard: report widely, act narrowly (the new horizons are reported, but the live signal-edge verdict and family sizing stay anchored to the horizons they were always computed on), and a forward stamp more than 20% of its horizon late is written off as missed rather than back-filled — without which adding a 5-minute horizon would have back-stamped every retained probe with the current price and labelled multi-day returns as five-minute returns. Off-switch: TAKER_FLOW_ENABLED=false.

  • Lifecycle risk + concentration caps: Every add-on shares the original position's stop-defined risk and projected same-symbol exposure budgets. Caps are reapplied after contract-lot rounding; an exchange minimum that exceeds either budget is rejected. Add-ons require a live fee-adjusted breakeven stop and may never loosen it or average down.

  • Correlation gate + relative-strength exception: Blocks ordinary non-major alt longs while BTC's daily regime is bearish. A rotating leader may pass only when its 1D/4H/1H/15M trends are all bullish, strength is high, and confidence clears RELATIVE_STRENGTH_MIN_CONFIDENCE; it is then reduced by RELATIVE_STRENGTH_SIZE_FACTOR. No symbol is hardcoded.

  • New-listing guard: Blocks futures entries on contracts younger than MIN_FUTURES_LISTING_AGE_DAYS (via the contract's first-open date). Freshly-listed perps are thin and ultra-volatile — RE-USDT had a ~100% intraday range on day one.

  • Minimum reward:risk (futures): Rejects any futures entry whose post-cost reward:risk is below MIN_FUTURES_RR, including estimated entry/exit fees and slippage. Dollar risk is controlled by position size, not by widening the stop or inventing a farther target.

  • Adaptive edge controller (src/edge.py): derives risk posture from rolling realized results and automatically relaxes when evidence recovers. Code-enforced actions surfaced in edgeReport: (1) direction/symbol risk scaling — a sufficiently sampled losing long/short direction or symbol trades at EDGE_NEGATIVE_EXPECTANCY_SIZE_FACTOR; (2) severity-scaled symbol bench — a repeatedly losing symbol is quarantined for an automatically scaled cooldown; (3) loss-streak throttle — consecutive losses reduce all new-entry size. Targets stay structural: weak results reduce capital at risk instead of moving take-profits farther away.

  • Give-back arming at 1R (PROFIT_LOCK_GIVEBACK_ARM_R): the give-back cap only acts once a run has reached this multiple of the trade's own initial risk (stop distance) — sub-1R wobble belongs to the original stop. Stops the cap from strangling winners into fee-scale scratch closes while losses ride to the full stop.

  • Trend-aligned shorts: In a confirmed downtrend the anti-FOMO gate would otherwise force the bot to only ever long oversold bounces. With TREND_ALIGNED_SHORTS_ENABLED, a short into an exhausted-bearish daily is permitted when 1h and 15m both confirm the downtrend is resuming and confidence clears a higher bar (TREND_SHORT_MIN_CONFIDENCE) — letting the bot trade with the trend, not only against bounces.

  • Reversal longs: The daily gate is a lagging signal — it reads bearish through the bottom of a move, so the bot is structurally forbidden from catching a reversal (in the Jul 2–5 2026 chop it sat out an +11% ETH bounce, blocked from every long). With REVERSAL_LONGS_ENABLED, a long against a bearish daily is permitted only when 1h and 15m have both turned bullish and confidence clears a high bar (REVERSAL_LONG_MIN_CONFIDENCE, default 0.80) — a confirmed turn, not knife-catching. Majors only (non-major alt longs stay blocked by the correlation gate), and the reward:risk floor still applies to whatever passes.

  • Reversal shorts (REVERSAL_SHORTS_ENABLED): exact mirror — when the regime flips bullish, the daily gate blocks every short even as intraday clearly rolls over (Jul 7–8 2026 pullback: "daily is bullish, shorts blocked" repeating while SOL fell 5%). A short against a bullish daily is permitted only when 1h and 15m have both turned bearish and confidence ≥ REVERSAL_SHORT_MIN_CONFIDENCE (0.80). Same discipline: confirmed turns only, R:R floor/bench/sizing still apply.

  • Anti-FOMO daily-exhaustion block: Refuses trend-continuation entries (long at bullish-overbought / short at bearish-oversold) when the 1D RSI is at an extreme (≥70 or ≤30). Counter-trend reversal setups remain allowed, a confirmed trend-aligned short can be re-permitted (see above), and a verified non-directional playbook passes (see below).

  • Anti-FOMO stacking: Refuses adds to an existing position — even a profitable one — when the daily is exhausted in the same direction. Stops doubling down at the top/bottom.

  • Volatility soft-gate: Above MAX_ATR_PCT_FOR_ENTRY (default 9%), position size is scaled down linearly (threshold/ATR, floor 50%) — a wider stop on a high-ATR name already sizes the position down via the risk budget, so the old quadratic penalty double-counted volatility and made the market's strongest movers untradeable at meaningful size exactly when they trended hardest. Above 1.5× the threshold the entry is hard-blocked as a data-quality / price-scale-discontinuity guard.

  • Squeeze-breakout signal: Structured squeeze_breakout field (long / short / None) surfaced in analyze_market_context. Fires only on the fresh transition out of a 1h Bollinger squeeze (BBW expanding ≥25% off the floor, ADX>20, price beyond BB band, RSI confirming). Takes the asymmetric upside after coiled-volatility periods; volume ≥1.5× 20-candle average is required confirmation. Anti-FOMO block still wins if daily is exhausted in the same direction.

  • 1h alignment requirement: Blocks new entries and add-ons when the 1h bias opposes the proposed side, regardless of what the daily trend says. The 1h timeframe captures the multi-hour trajectory — when daily EMAs are still bullish but 1h is bearish, the daily uptrend is in correction (not a healthy pullback) and buying bounces gets stopped out repeatedly. Catches the failure mode where 15m briefly turns bullish on a dead-cat bounce while the actual correction is still in progress.

  • Deadlock break: The daily gate (blocks the counter-trend direction) and the 1h-alignment gate (blocks the daily-aligned direction during a counter-bounce) can together strand the bot flat in both directions in a clean trend. With DEADLOCK_BREAK_ENABLED, the daily-aligned entry (short in a bearish daily, long in a bullish daily) is allowed past the 1h gate only when the counter-bounce is stalling — 15m no longer confirms it — and confidence clears DEADLOCK_MIN_CONFIDENCE. Takes the trend-continuation trade instead of standing aside, without knife-catching a live bounce. Disjoint from trend-aligned shorts (which covers the exhausted-daily case).

  • Timeframe-conflict gate: Secondary check on top of 1h alignment — blocks new entries when analyze_market_context reports timeframe_conflict=True AND the 15m bias opposes the proposed direction. Catches lower-TF disagreement that slips past 1h alignment (e.g., 15m bearish while 1h is neutral). Position management (manage/hold/protect) is unaffected by both gates.

  • Atomic entry bracket (ATOMIC_BRACKET_ENABLED): futures limit entries attach TP/SL via KuCoin's st-orders endpoint, so protection arms with the fill. If the atomic endpoint fails, the entry fails closed; there is no plain/unprotected-order fallback.

  • Live entry lease (src/safety.py): every background run receives revocable order authority. Incomplete account/fill truth, a 20-minute run timeout, or shutdown revokes it. Immediately before an entry, equity, positions, stops, and pending orders are fetched again under a serialized exchange-write lock; any structural change forces re-analysis.

  • Hard unrealized-loss cap: the polling protection loop closes a futures lifecycle when its unrealized loss exceeds RISK_PER_TRADE_PCT of current equity. This is a last-resort guard; gaps, fees, and slippage can still make realized loss larger.

  • Real OI sampling: OI/price quadrants use timestamped exchange open-interest observations with minimum age/change thresholds. The signal stays neutral without a valid baseline; 24h volume is never substituted for OI direction.

  • Unprotected-position safety net (EMERGENCY_SL_PCT, in src/protection.py): every poll, any open futures position found with no protective stop (after a short grace so an attached bracket can appear) gets an emergency SL at EMERGENCY_SL_PCT from entry + a TP at MIN_FUTURES_RR× that — within ~1 poll, not the next agent run. It's a floor, not the agent's considered bracket; the agent refines it next run. Guarantees no position is ever left naked.

  • Realized-vs-intended R:R reality check: each close records its MFE/MAE (peak/trough PnL), and the agent's edgeReport surfaces realizedRewardRisk (avg win ÷ avg loss actually achieved). When it's far below the intended floor, the agent is told the take-profits are set too far to reach and to pull them to the nearest realistic structural target with a tighter stop — so the RR floor passes at a reachable scale rather than aiming at targets that never fill.

  • Early invalidation cut (EARLY_CUT_*, in src/protection.py): a time/heat stop — a position that (after EARLY_CUT_GRACE_MIN) has never gone meaningfully green and has run EARLY_CUT_MAE_FRAC of the way to its stop is closed early rather than riding to the full SL, front-running an almost-certain stop. Per MAE research (and the bot's own data — winners here breathe ~0.57R of adverse excursion), the cut threshold must sit outside the MAE band of winning trades, or it stops you out when you're right; EARLY_CUT_MAE_FRAC was therefore raised 0.6 → 0.85 after a mid-coil trend long (NEAR) was cut at 0.65R moments before a +4% breakout. At 0.85 it only fires when the stop is nearly certain, leaving normal pre-breakout heat alone. Disjoint from the give-back/breakeven guards (which only act once a trade has gone green).

  • Mandatory TP/SL: Every position must have stop-loss and take-profit (no naked positions)

  • ATR-based stops: Stop distance computed from Average True Range for volatility-adaptive risk

  • Daily trade limits: Per-symbol and total daily trade caps

  • Fee-aware profit targets: Minimum net profit and ROI thresholds after accounting for fees and slippage

Order Execution

  • Futures-only new exposure: New positions use non-marketable place_futures_limit_order requests with atomic TP/SL. Spot and futures market orders are close/emergency-only; existing spot holdings can still be protected or closed.
  • Target-price limit entries: Futures entries wait at a technically derived level (EMA, Bollinger Band, swing high/low, VWAP, Fibonacci), preventing shorting into a dump or buying into a pump.
  • Pending order safety: Every run includes pending orders; code permits only one futures entry per symbol per run/book, tags bot-created GTC entries, and automatically cancels tagged entries older than ENTRY_LIMIT_EXPIRY_MINUTES. Manual and protective orders are never auto-cancelled.
  • Fee-aware entry gate: Atomic limit entries must clear configured net-profit/ROI floors after estimated fees and slippage.
  • Crossing vs resting follows the live execution map, not a frozen replay: a passive limit that never fills captures zero edge, and an unfilled winner is a real loss of opportunity. There is no confidence bar on crossing into the MARKETABLE_ENTRY_MAX_DEV_PCT band; the post-cost RR gate still runs on the crossed price, but it is a fee/payoff guard, not a quality filter — a pass does not mean the cross is good, and since exits come mostly from the trail, extra paper RR above the floor is not evidence of a better trade. The prompt used to tell the agent to "try the cross and let that gate answer", backed by the Jul 30 figures (+13.05R, −0.37R at 4h, +0.388R vs +0.050R) that did not reproduce on the Sep 20–24 sample (see Marketable entries). It now says: read edgeReport.executionMap, compare fill-adjusted expectancy by distance (marketable included, as its own bucket), and don't push the entry or the TP out just to clear the floor — for every playbook, since a range_edge or fade_extreme order resting at the level is the same passive limit. A wide cross is still paid out of reward and so fails the gate by construction — crossing is right when you are at the level, not chasing one that already ran.
  • Leverage control: Configurable max leverage (up to 125x) with automatic margin mode management
  • Fund transfers: Move USDT between spot, futures, and financial/Earn accounts

Memory & Learning

  • Trade memory: Records all trades, decisions, plans, sentiments, triggers, and fee snapshots
  • Persistent event inbox: Fill and close payloads survive restarts and model failures. They are acknowledged only after a successful agent run, eliminating repeated poll-triggered runs without silently dropping unprocessed events.
  • Two-tier decision retention: realized closed-trade outcomes (those with a PnL) are kept far longer (cap 200) than routine entry/decline decisions (cap 50), so win/loss history — and the exit prices the no-chase guard relies on — is never crowded out by no-trade decisions
  • One close, counted once: a close is reported twice — the bot's own pre-close estimate, then KuCoin's authoritative figure seconds later — and booking both corrupts realized PnL and every rolling stat built on it. _authoritative_realized_rows suppresses the estimate when it shadows a triggered close on the same symbol nearby. Rows that name no position at all (no positionId, no positionOpenTime, no lifecycle version) are the agent narrating a close rather than recording one, so their figure was never reconciled with anything and can be arbitrarily wrong: NEAR-USDT on 2026-09-01 had the bracket fire at +0.00335 and the agent log its own reduce-only close at +0.00666 five seconds later — 99% apart, so the "within half of each other" estimate band let it through and one trade was booked twice. For an anonymous row the test therefore widens from roughly equal to the same order of magnitude (<10×), which is the real claim — two reports of one close can disagree over fees, partial reductions and a few seconds of drift, but not by 10×, while two genuinely different trades on one symbol do differ by that much. Every closed-trade surface now goes through this one filter, including the dashboard's outcome bar chart, which previously filtered the raw decisions feed itself and so drew three bars beside a "recently closed" panel showing two trades.
  • Performance tracking: Win rate, PnL, trade counts split by venue (spot/futures) and mode (paper/live)
  • Position extremes: Tracks peak and trough unrealized PnL during position lifetime for post-trade analysis
  • Drawdown tracking: Per-venue daily drawdown percentage
  • Adaptive sizing: Kelly fraction adjusts position size based on actual win rate and profit/loss ratio
  • Automatic retention: Items older than configurable retention period are pruned

Coin Universe Management

  • Market screener (scan_futures_market): the Research and Trading agents can screen the entire KuCoin USDT-perp universe (~500 contracts) ranked by momentum / gainers / losers / volume, instead of only looking at coins already on the list. Every result pre-clears the liquidity floor (SCREENER_MIN_TURNOVER_USD_24H) and the minimum listing age, so discovery stays liquid and mature (no fresh micro-caps). This closes the gap where the scout could only evaluate symbols it already knew by name and never discovered the coin that was actually moving. Entry gates (daily/1h/R:R/correlation) still apply at trade time. Each row carries the contract's assetClass (CRYPTO / STOCK / METAL / COMMODITY).
  • Only what this bot can trade takes a screener slot (Sep 25): execution here is spot-anchored (add_coin, the universe and every order tool map BASE-USDT → BASEUSDTM), so a perp with no live spot pair — FHE, TAKE, 1000BONK (its spot pair is BONK-USDT), stock perps — can be analysed but never added or traded. About 16% of qualified perps were like that, they took 4 of 13 short-side slots in one snapshot, and 43 tool calls in 52h failed only because of it. One helper (agent._bot_can_trade over the cached live spot list, tools._SPOT_SYMBOLS, ~1h) now decides for both the screener and add_coin, so the two can never disagree: unreachable rows are dropped before the top-n cut and named in excluded, and add_coin says "no spot route" instead of the misleading "not found on KuCoin". If the spot list cannot be fetched it fails open (rows kept, tradeableCheck: "unavailable"). Trading perp-only contracts is deliberately not enabled: stock perps gap across their underlying's session close, a risk no survival code here models.
  • Scan rows say what the bot already knows (Sep 25, zero extra API calls): quarantined + remainingHours from the adaptive ATR quarantine, and lastAnalysisFailure + retryAfter / retryInHours from analyze_market_context's own data-quality refusals, which are now persisted (memory.analysis_failures, its own key — the coins list is capped at 50 and the model's remove_coin overwrites it). The retry time comes from the failure's evidence, not a TTL: a candle gap never heals, it ages out once the bar before it leaves the 50-bar analysis window (analytics.candle_quality_retry_after); anything else can at best change at the next bar close. A clean analysis clears the record. Rows are annotated, never dropped — the unchanged gate still decides. Live 09-23/24, TAKE's 1h gaps were re-analysed ~1.4 times per run for ~6 hours and ONE/4/POWER/MUBARAK were re-analysed while quarantined. list_coins returns the same analysisFailures.
  • Read-only futures tools accept any listed contract (Sep 25): the data tools (fetch_funding_rate, fetch_open_interest, fetch_futures_candles, fetch_futures_orderbook, fetch_futures_mark_price, fetch_contract_details) now map X-USDT → XUSDTM through the live contract list when the coin is outside the universe — 14 "Unknown symbol" errors in two days were spot-listed coins not yet added (KCS, BCH, SPX), minutes before the bot traded one. Order tools still resolve only through the allowed symbols. list_futures_stop_orders used to send the model's DASH-USDT straight to KuCoin, which accepts only DASHUSDTM — every filtered stop check failed with 100003 (9 times, each in the run that had just closed, entered or tightened that symbol). It now makes the same single call unfiltered and filters on the resolved contract, so stops on a coin the bot has since removed still list.
  • The decision feed no longer calls a failure a live order (Sep 25, agent._summarize_tool_output): any tool output carrying an orderId printed as live order, so lease-guard refusals and failed cancels read as live orders in the log, Telegram and the Supervisor's read_logs (10 cancel outcomes in two days; 2 printed correctly). rejected and error are now checked first, and a cancel whose exchange response is empty prints as cancel sent … (exchange returned no confirmation) — KuCoin can return null data on a successful cancel, so it is not called a failure either.
  • Seed with COINS env var; agent can dynamically add/remove coins with reasons and exit plans when FLEXIBLE_COINS_ENABLED=true
  • Auto-discovers unlisted holdings in spot account (worth >= $0.50) and adds them to the active list
  • Removes coins after 3 consecutive ticker fetch failures (flexible mode only)
  • Forced research handoff: after RESEARCH_HANDOFF_AFTER_NO_TRADE_RUNS consecutive no-trade runs (stuck / declining), the Trading Agent is forced to hand off to the Research Agent to overhaul the coin list and surface fresh opportunities — rate-limited by RESEARCH_HANDOFF_COOLDOWN_MIN so the costly web-research sweep can't fire every cycle

Setup

  1. Copy .env.example to .env and fill in credentials
  2. Install dependencies and run:
python -m venv .venv
source .venv/bin/activate  # or .venv\Scripts\activate on Windows
pip install -r requirements.txt
python -m src.main

The agent runs in a continuous loop: polls KuCoin, tracks price changes, performs web searches for market context, and invokes the AI agents when triggers fire. Keep PAPER_TRADING=true while testing.

Configuration

Required

Variable Description
AZURE_OPENAI_ENDPOINT Azure OpenAI resource endpoint
AZURE_OPENAI_API_KEY Azure OpenAI API key
KUCOIN_API_KEY KuCoin API key
KUCOIN_API_SECRET KuCoin API secret
KUCOIN_API_PASSPHRASE KuCoin API passphrase
COINS Comma-separated symbols (e.g., BTC-USDT,ETH-USDT,SOL-USDT)

Trading Controls

Variable Default Description
PAPER_TRADING true Simulate orders without real execution
MAX_POSITION_USD 500 Maximum spend per trade
RISK_PER_TRADE_PCT 0.02 Maximum stop-defined risk per trade as a fraction of equity (2%, aggressive); fees, slippage, and gaps can make realized loss higher
SIZE_QUALITY_FLOOR 0.5 Floor for the combined soft size factor (regime/conviction/relative-strength/streak/expectancy combined by their worst signal, not multiplied). 1.0 disables all soft shrink
MIN_ENTRY_NOTIONAL_USD 0 Fee-aware floor: bump a sub-floor sized entry up to this notional (still bounded by risk/concentration/heat caps). 0 = off (pure risk-budget sizing)
MIN_CONFIDENCE 0.65 Minimum confidence score (0-1) to place a trade
MAX_LEVERAGE 3 Maximum futures leverage (1-125)
MAX_TRADES_PER_SYMBOL_PER_DAY 6 Daily trade cap per symbol (curbs fee churn and loss streaks)
MIN_NET_PROFIT_USD 0.50 Minimum net profit target after fees
MIN_PROFIT_ROI_PCT 0.008 Minimum ROI target (0.8%) after fees
ESTIMATED_SLIPPAGE_PCT 0.001 Per-side slippage prior for RR/profit gates. Superseded by the bot's own measured fills once SLIPPAGE_AUTOTUNE_MIN_SAMPLES exist
SLIPPAGE_AUTOTUNE_ENABLED true Measure per-side slippage from real fills instead of assuming it (the static 0.1% was ~12x the measured 0.008% and taxed every setup a phantom 0.18R)
SLIPPAGE_AUTOTUNE_MIN_SAMPLES 8 Fills required before the measurement displaces the prior
RANGE_TRADING_ENABLED true Enable mean-reversion in ranging/sideways markets
SENTIMENT_FILTER_ENABLED false Require positive sentiment before trading
SENTIMENT_MIN_SCORE 0.55 Minimum sentiment score (0-1) when filter enabled

Advanced Trading Features

Variable Default Description
PARTIAL_TP_ENABLED false Opt in to 60%/40% staged take-profit; off by default because the current staging geometry reduces realized reward:risk
KELLY_SIZING_ENABLED true Use Kelly criterion for adaptive position sizing
KELLY_MIN_TRADES 30 Minimum trade history before Kelly sizing activates
PREFER_LIMIT_ORDERS true Legacy spot-entry preference; new spot exposure is currently disabled
LIMIT_ORDER_TIMEOUT_SEC 20 Timeout before falling back to market order (fee-saving path)
ENTRY_LIMIT_EXPIRY_MINUTES 15 Cancel unfilled target-price entry limit orders after this many minutes (recycles unfilled passive limits faster)
MIN_ENTRY_DEVIATION_PCT 0.0005 Minimum distance (0.05%) from current price for a low-conviction resting limit; high-conviction entries use the marketable band instead
STOP_ATR_FLOOR_MULT 2.5 Minimum stop distance as a multiple of the 15m ATR. The median live stop was 1.4x (< one 1h bar), which no trade could survive; size shrinks to hold dollar risk constant. 0 disables
STOP_ATR_FLOOR_ADAPTIVE true Widen the floor by 0.5 (max 4.0) when the bot's own winners survive >=0.6R of adverse heat. Widen-only by design
STOP_FLOOR_SCALES_TARGET true Scale the take-profit by the same factor the floor widened the stop, so the floor is RR-neutral and shows up as smaller size rather than as RR rejections
STOP_ATR_FLOOR_MAX_WIDEN 4.0 Never rewrite a stop by more than this factor — past it, it is a different trade
RISK_TARGETED_SIZING true Size each entry FROM the risk budget and stop distance instead of capping a model-guessed notional (realized risk varied 18.9x without it)
MARKETABLE_ENTRY_MAX_DEV_PCT 0.01 Outer bound on how far an entry may cross the spread to fill immediately; the post-cost RR gate at the crossed price is the binding test. 0 disables (passive-only)
MARKETABLE_ENTRY_MIN_CONFIDENCE 0 Optional extra conviction bar for crossing. Default off. (The Jul 30 comparison it was set on — confidence +0.050R vs RR-after-cross +0.388R — did not reproduce on the Sep 20–24 sample; neither filter showed selection value there. Crossing is read from edgeReport.executionMap.)
MAX_ATR_PCT_FOR_ENTRY 9 Soft volatility gate: above this daily ATR %, position size is scaled down linearly (threshold/ATR, floor 50%). Above 1.5× this value (13.5% default), entry is hard-blocked as a data-quality guard.
MAX_24H_VOLATILITY_PCT 30 Exclude/hard-block contracts whose absolute 24h price change exceeds this % (separate from ATR gate)
POST_LOSS_COOLDOWN_MINUTES 30 Block new entries on a symbol after a loss
MIN_TRADE_INTERVAL_MINUTES 10 Minimum interval between trades on the same symbol (anti-overtrading)

Circuit Breakers

Variable Default Description
CB_MAX_DAILY_DRAWDOWN_PCT 3.0 Restrict new exposure at a 3R daily drawdown with the default 1% lifecycle risk
CB_MAX_CONSECUTIVE_LOSSES 3 Restrict trading after N consecutive losses
CB_MAX_PORTFOLIO_HEAT_PCT 6.0 Maximum total capital at risk % across open positions (the "6% rule")
CB_COOLDOWN_MINUTES 120 Cooldown duration after consecutive loss trigger

When a circuit breaker fires, the agent enters close-only mode: it can adjust stops, close positions, and manage risk, but cannot open new positions. A Telegram notification is sent.

Profit Protection

Code-driven guards enforced outside the LLM (src/protection.py). They run every poll regardless of whether the agent runs, so a profitable trade can't quietly round-trip into a loss while the agent is idle.

Variable Default Description
PROFIT_LOCK_ENABLED true Enable the breakeven ratchet + give-back cap
PROFIT_LOCK_DRY_RUN false Log intended actions without placing orders (observe before arming on live funds)
PROFIT_LOCK_BREAKEVEN_TRIGGER_R 0.5 Move the stop to breakeven once favorable excursion reaches this multiple of initial risk. Lowered 1.0→0.5 on Aug 11 2026: the stop floor doubled since 1R was fit, so 1R now sits ~3× ATR and is almost never reached; 0.5R ≈ the ~1.4× ATR noise band the 1R value originally meant
PROFIT_LOCK_BREAKEVEN_FEE_PCT 0.0015 Round-trip cost buffer (KuCoin futures taker 0.06%×2 + slippage) so the breakeven stop nets ≥0
PROFIT_LOCK_GIVEBACK_PCT 0.35 Close after price retraces this fraction of the peak run; 0.35 retains ~65% of peak profit, mid-band of the 60–70% best-practice range (0 disables the give-back close)
PROFIT_LOCK_MIN_FE_PCT 0.005 Minimum run (fraction of entry) before the give-back cap can act — filters noise
PROFIT_LOCK_GIVEBACK_ARM_R 1.0 Also require the run to reach this multiple of the trade's own risk (stop distance) before give-back can act; 0 = pct-arming only
EARLY_CUT_ENABLED true Cut a trade that never went green and is failing toward its stop, before the full SL
EARLY_CUT_GRACE_MIN 20 Minutes to let a fresh entry work before early-cut can act
EARLY_CUT_MIN_FAVORABLE_PCT 0.003 Peak excursion (fraction of entry) below which a trade "never worked"
EARLY_CUT_MAE_FRAC 0.85 Fraction of the way to the stop that triggers the early cut. 0.85 (raised from 0.6) keeps the cut outside the MAE band of winning trades — a lower value stops out trades that are still right (per MAE research + the bot's own ~0.57R winner MAE)
PROFIT_LOCK_TREND_ADAPTIVE true Loosen the give-back cap for a revealed trend winner so it can run (chop keeps the tight defaults)
PROFIT_LOCK_TREND_RUNNER_R 2.0 Peak run (in R, the trade's own risk) at which a position is treated as a trend winner
PROFIT_LOCK_TREND_GIVEBACK_PCT 0.55 Once a runner, tolerate giving back this much of peak (vs PROFIT_LOCK_GIVEBACK_PCT in chop)
PROFIT_LOCK_TREND_GIVEBACK_ARM_R 2.5 Arm the loosened give-back only after this much favorable run
PROFIT_LOCK_TRAIL_ENABLED true Use the R-based trailing stop (lets winners run to TP) instead of the give-back market-close
PROFIT_LOCK_TRAIL_ARM_R 0.5 Arm the trailing ratchet once the peak favourable run clears the noise band the stop was drawn around (~1.4× ATR). This is a distance in stop units: the July 1m replay put it at 1.0R when the stop was 1.4× ATR, but the stop floor has since doubled (median 3.0× ATR), so 1.0R now sits ~3× ATR and never arms — leaving winners to round-trip to full stops (21 recent trades peaked +0.95R, kept +0.24R). 0.5R restores the validated arming distance under the wider stop
PROFIT_LOCK_TRAIL_LOCK_FRAC 0.33 Once armed, lock at least this fraction of the peak favorable run
PROFIT_LOCK_TRAIL_DISTANCE_R 1.0 ...or trail the stop this many R below the running peak, whichever locks MORE. Fallback only: when the position's entry recorded a stopAtrMult, the trail instead rides one measured noise band (1 / stopAtrMult in R) behind the peak, so the give-back is set by the symbol's own volatility and cannot go stale when the stop floor moves
NO_CHASE_ENABLED true Block same-direction re-entry at a worse price after a recent winning close
POST_WIN_COOLDOWN_MINUTES 45 Window after a winning close during which re-entry at a worse price is blocked
NO_CHASE_BUFFER_PCT 0.001 Tolerance band around the prior exit price
ATOMIC_BRACKET_ENABLED true Attach TP/SL to the futures entry order (KuCoin st-orders) so a limit fill is protected instantly, not on the next agent run
EMERGENCY_SL_PCT 0.02 Safety net: SL distance (fraction of entry) for an open position found with no stop; TP set at MIN_FUTURES_RR× it. 0 disables

Every automatic action (stop moved to breakeven, position closed, or a dry-run preview) is logged and sent as a Telegram alert.

Risk Guardrails

Blast-radius and selection guards added after the RE-USDT concentration blowup (one freshly-listed micro-cap alt at ~74% of equity, longed into a BTC downtrend).

Variable Default Description
MAX_POSITION_EQUITY_PCT 0.5 Cap a single position's notional at this fraction of total equity, regardless of leverage (0 = off)
MIN_FUTURES_LISTING_AGE_DAYS 7 Block futures entries on contracts younger than this many days — thin/volatile fresh listings (0 = off)
MIN_FUTURES_RR 1.5 Reject futures entries whose post-cost reward:risk (fees and estimated slippage included) is below this (0 = off)
SCREENER_MIN_TURNOVER_USD_24H 5000000 Market screener (scan_futures_market) liquidity floor: only surface perps with at least this 24h USDT turnover

Adaptive Edge Controller

Self-tuning risk (src/edge.py): posture derives from rolling realized closes, tightens while losing, and relaxes automatically when expectancy recovers. Once attributed history is sufficient it uses realized R (net PnL / planned maximum loss), so larger notionals cannot dominate the learning signal; legacy dollar PnL is only a migration fallback. The live RR gate remains structural and cost-aware, while the controller adapts risk rather than stretching targets.

Variable Default Description
ADAPTIVE_EDGE_ENABLED true Master switch for the adaptive edge controller
EDGE_LOOKBACK_TRADES 30 Rolling window of realized closes the stats are computed over
EDGE_MIN_TRADES 8 Minimum closes before adaptive actions kick in (below this, static behavior)
EDGE_DIRECTION_MIN_TRADES 5 Minimum closes before a long/short direction or symbol's expectancy can reduce its risk
EDGE_NEGATIVE_EXPECTANCY_SIZE_FACTOR 0.5 Entry-size multiplier for a sufficiently sampled losing direction/symbol; never enlarges risk
EDGE_RR_STEP, EDGE_RR_CAP, EDGE_RR_STALE_HOURS, EDGE_SYMBOL_RR_MIN_TRADES legacy Retained for configuration compatibility and offline comparisons; live admission no longer stretches targets after losses
EDGE_BENCH_LOOKBACK 5 Per-symbol recent closes examined for the bench
EDGE_BENCH_MIN_LOSSES 3 Losses within that window (with negative net) that bench the symbol
EDGE_BENCH_COOLDOWN_HOURS 12 Base bench rest; scaled by loss count so a persistent loser sits out longer
EDGE_BENCH_COOLDOWN_MAX_MULT 4 Cap on the bench-rest scaling (e.g. 12h × 4 = up to 48h)
EDGE_STREAK_THRESHOLD 2 Consecutive realized losses that trigger the size throttle
EDGE_STREAK_SIZE_FACTOR 0.5 Entry-size multiplier while on a losing streak
EDGE_STAND_ASIDE_NO_EDGE_FAMILY true Zero stake on a playbook — judged per side once that side has 20 probes of its own, else on the pooled family — while its net-of-cost forward return is not above its own standard error (t = net/SE < 1: no edge when net ≤ 0, unproven when positive but inside its noise). Stateless, recomputed every run; the refusal says which of the two fired. See edge.family_stake_status
EDGE_EXPLORE_UNPROVEN_FAMILY_FACTOR 0.4 Risk factor for a playbook with < 20 scored probes; once scored, size ramps from this at t = 1 to full at t = 2. A side still thin on its own stays at this factor whatever the pooled family earned
TAKER_FLOW_ENABLED true Sample KuCoin's public taker tape every poll and stamp the aggressor balance onto every direction call, so edge.taker_flow_edge_stats can measure whether order flow predicts on this venue at our horizons. Pure observation — no gate, no sizing input, no wake trigger reads it
TAKER_FLOW_MAX_SYMBOLS 12 Cap on tape samples per poll (one public REST call each). Held positions first, then the watchlist ranked by the move about to wake the agent, so the cap trims coins nothing is happening on rather than an open trade or the symbol about to be called. Readings expire after 10 polls
ALT_LONG_BLOCK_WHEN_BTC_BEARISH true Block longs on non-major alts while BTC's daily regime is bearish (alts are high-beta to BTC)
ALT_MAJORS BTC,ETH Symbols exempt from the alt-long gate (they have their own per-symbol daily gate)
RELATIVE_STRENGTH_LONGS_ENABLED true Allow a narrow all-timeframe bullish exception to the bearish-BTC alt veto
RELATIVE_STRENGTH_MIN_CONFIDENCE 0.82 Confidence required for that exception
RELATIVE_STRENGTH_SIZE_FACTOR 0.5 Reduced size applied to an exception trade
RESEARCH_HANDOFF_AFTER_NO_TRADE_RUNS 3 Force a Research handoff after this many consecutive no-trade runs to refresh the coin list (0 = off)
RESEARCH_HANDOFF_COOLDOWN_MIN 30 Minimum minutes between forced Research handoffs — rate-limits the costly web sweep (0 = off)

Regime-Aware Entries

Code-enforced entry adjustments that work alongside the daily gate (src/regime.py): be more selective in hostile regimes, size by conviction, trade with a confirmed downtrend instead of only longing bounces, and break the daily-vs-1h gate deadlock.

Variable Default Description
REGIME_THROTTLE_ENABLED true Raise the confidence bar + shrink size in a hostile (bearish / RSI-exhausted) daily
REGIME_CAUTION_MIN_CONFIDENCE 0.75 Elevated confidence floor in a hostile regime (base is MIN_CONFIDENCE)
REGIME_CAUTION_SIZE_FACTOR 0.6 Position-size multiplier applied in a hostile regime
TREND_ALIGNED_SHORTS_ENABLED true Permit a trend-aligned short past the anti-FOMO gate in an exhausted-bearish daily
TREND_SHORT_MIN_CONFIDENCE 0.78 Higher confidence bar specifically for a counter-bounce short
TREND_SHORT_REQUIRE_15M true Require 15m (not just 1h) bearish confirmation before allowing the short
REVERSAL_LONGS_ENABLED true Allow a long past a bearish daily gate when 1h+15m have both turned bullish (catch confirmed reversals)
REVERSAL_LONG_MIN_CONFIDENCE 0.80 High confidence bar for a counter-daily reversal long
REVERSAL_LONG_REQUIRE_15M true Require 15m (not just 1h) bullish confirmation before allowing the reversal long
REVERSAL_SHORTS_ENABLED true Mirror: allow a short past a bullish daily gate when 1h+15m have both turned bearish
REVERSAL_SHORT_MIN_CONFIDENCE 0.80 High confidence bar for a counter-daily reversal short
REVERSAL_SHORT_REQUIRE_15M true Require 15m (not just 1h) bearish confirmation before allowing the reversal short
CONVICTION_SIZING_ENABLED true Scale position size by how far confidence clears the floor (low-conviction → smaller)
CONVICTION_FULL_CONFIDENCE 0.85 Confidence at/above which full size is used (linear ramp from the floor)
CONVICTION_MIN_SIZE_FACTOR 0.5 Size multiplier at the confidence floor
DEADLOCK_BREAK_ENABLED true Allow the daily-aligned entry past the 1h gate when a 1h counter-bounce is stalling (15m no longer confirms it)
DEADLOCK_MIN_CONFIDENCE 0.72 Raised confidence bar to take the trend-continuation entry

Loop & Polling

Variable Default Description
POLL_INTERVAL_SEC 60 Seconds between polling cycles
PRICE_CHANGE_TRIGGER_PCT 0.5 Price move % that triggers an agent run
MAX_IDLE_POLLS 10 Force agent run after N idle polls
FLAT_AGENT_COOLDOWN_SEC 600 Quiet-market HUNT cadence while flat (~10min); triggered move magnitude/breadth automatically reduce it toward the active cadence
FLAT_BACKOFF_MAX_MULTIPLIER 1 Power-of-two backoff cap for repeated idle-only no-action runs; 1 disables backoff (frequent hunting), >1 opts in
ACTIVE_AGENT_COOLDOWN_SEC 300 Minimum interval between model runs with exposed capital or a recent lifecycle/trigger event
AGENT_MAX_TURNS 20 Max tool-call turns per run; a separate 20-minute wall-clock timeout revokes order authority

KuCoin

Variable Default Description
KUCOIN_BASE_URL https://api.kucoin.com Spot API endpoint
KUCOIN_FUTURES_ENABLED true Enable futures trading
KUCOIN_FUTURES_BASE_URL https://api-futures.kucoin.com Futures API endpoint
KUCOIN_FUTURES_MARGIN_MODE cross Futures margin mode (cross / isolated / auto); the cross-leverage call is only issued in cross mode
FLEXIBLE_COINS_ENABLED true Allow agent to add/remove coins dynamically

Azure APIM (Optional)

If AZURE_APIM_OPENAI_SUBSCRIPTION_KEY is set, the client uses APIM endpoint/deployment instead of direct Azure OpenAI (subscription key auth).

Variable Description
AZURE_APIM_OPENAI_ENDPOINT APIM gateway endpoint
AZURE_APIM_OPENAI_DEPLOYMENT Deployment name behind APIM
AZURE_APIM_OPENAI_API_VERSION API version (default: 2024-08-01-preview)
AZURE_APIM_OPENAI_SUBSCRIPTION_KEY APIM subscription key

Memory

Variable Default Description
MEMORY_FILE .agent_memory.json Path to agent memory store
RETENTION_DAYS 90 Auto-prune items older than this; both main loop and agent use the same horizon

Tracing (Optional)

Variable Default Description
ENABLE_TRACING false Enable OpenAI Agents SDK spans
ENABLE_CONSOLE_TRACING false Print spans to console (dev only)
OPENAI_TRACE_API_KEY — Export spans to OpenAI traces endpoint
LANGSMITH_ENABLED false Enable LangSmith tracing
LANGSMITH_API_KEY — LangSmith API key
LANGSMITH_PROJECT trAIde LangSmith project name
LANGSMITH_API_URL https://api.smith.langchain.com LangSmith API endpoint
LANGSMITH_TRACING true Send agent runs to LangSmith when enabled
LANGSMITH_SAMPLE_RATE 0.1 Head-sampling fraction of runs traced to LangSmith (avoids the monthly unique-trace cap)

OTLP export for Azure Monitor is supported via OTEL_EXPORTER_OTLP_ENDPOINT and OTEL_EXPORTER_OTLP_HEADERS.

Telegram Notifications

Get real-time updates on your phone for every trading decision, order execution, and error.

1. Create a Telegram bot

  1. Open Telegram and search for @BotFather.
  2. Send /newbot and follow the prompts to choose a name and username.
  3. BotFather replies with your bot token (e.g., 123456:ABC-DEF1234...). Save it.

2. Get your chat ID

  1. Start a conversation with your new bot (search its username and press Start).
  2. Send any message to the bot (e.g., "hello").
  3. Open this URL in your browser (replace <BOT_TOKEN> with your token):
    https://api.telegram.org/bot<BOT_TOKEN>/getUpdates
    
  4. In the JSON response, find "chat":{"id":123456789} -- that number is your chat ID.

3. Configure .env

TELEGRAM_ENABLED=true
TELEGRAM_BOT_TOKEN=123456:ABC-DEF1234ghIkl-zyx57W2v1u123ew11
TELEGRAM_CHAT_ID=123456789
TELEGRAM_SILENT=false
Variable Description
TELEGRAM_ENABLED true to activate notifications, false to disable (default: false)
TELEGRAM_BOT_TOKEN Bot token from @BotFather
TELEGRAM_CHAT_ID Your personal or group chat ID
TELEGRAM_SILENT true to send notifications without sound (default: false)

What you'll receive

  • Startup -- bot mode (live/paper), active coins, max position, leverage, futures status
  • Agent run summaries -- triggers that fired, every order placed (symbol, side, price, TP/SL, RR ratio, paper/live), declines with reason and confidence, and a narrative excerpt
  • Order details -- full breakdown of each executed order including stop-loss, take-profit, expected PnL for sells, and order ID
  • Circuit breaker alerts -- immediate notification when trading is restricted
  • Errors -- immediate alerts when the agent run or snapshot build fails

Messages are sent asynchronously via a background thread and never block the trading loop. If Telegram is unreachable, failures are logged and silently skipped.

Supervisor Agent (Interactive Telegram Bot)

Talk back to the bot. The Supervisor Agent listens for your Telegram messages, processes them through an AI agent with full read access to the system, and replies in the same chat.

What it can do

  • Query status -- ask about positions, balances, performance, win rate, recent trades, or recent decisions. It fetches live data from KuCoin and agent memory.
  • Check resting orders -- get_open_orders lists live unfilled limit/entry orders straight from the exchange. A pending limit entry is neither a position nor a fill, so this is the authoritative source for "is there still a pending order for X?" — it reconciles against a stale hold_pending decision that may predate the order's TTL expiry.
  • Read & search logs -- ask it to check logs for errors, search for a specific symbol, or show the last N lines.
  • Read source code -- inspect any file in the src/ directory.
  • View configuration -- see all non-secret config values (API keys are never exposed).
  • Fetch market data -- funding rates, open interest, mark price for futures symbols.
  • Read the gate scoreboard -- get_gate_scoreboard returns, per directional gate and side, what the gate blocks vs what it allows on the same day (from model-independent gate-state readings, hard refusals and hatch admissions), net of cost with a day-clustered SE. Report-only: the Supervisor is told not to pass a gate verdict to the trading agent unless you ask.
  • Read the edge scoreboard -- get_edge_scoreboard returns each playbook's measured edge per family and per side (n, net of cost, SE, t, verdict) with the standAside / stake the order path applies right now, plus per-model confidence informativeness — so "why did it refuse continuation?" and "does the new model's confidence mean anything?" have a direct answer.
  • Web search -- search the web for market context, news, or any other information.
  • Write notes for the trading agent -- influence the trading agent's behavior:
    • Temporary notes (one-time, highest priority): injected into the trading agent's system prompt on the next run only, then auto-deleted. These override any conflicting rules. Example: "Close all BTC positions immediately."
    • Permanent notes: added to the trading agent's system prompt on every run until manually deleted. Example: "Never trade DOGE-USDT."
  • Conversation memory -- the supervisor remembers the last 3 exchanges and maintains a rolling summary of older conversations, so you can have multi-turn dialogues without repeating context.

Enable it

SUPERVISOR_ENABLED=true
TELEGRAM_ENABLED=true
TELEGRAM_BOT_TOKEN=...
TELEGRAM_CHAT_ID=...
Variable Description
SUPERVISOR_ENABLED true to start the interactive bot (default: false)
LOG_FILE Log file path the supervisor reads (default: traide.log)
LOG_MAX_BYTES Max log file size before rotation (default: 5242880 / 5MB)
LOG_BACKUP_COUNT Number of rotated log backups (default: 3)

The supervisor runs as a daemon thread alongside the trading loop, using Telegram long-polling. Only messages from the configured TELEGRAM_CHAT_ID are processed; all others are silently ignored.

Example commands

  • "What's my current P&L?"
  • "Show me the last 5 trades"
  • "Search logs for ERROR"
  • "Add a temporary note: skip all trades this run, market is too volatile"
  • "Add a permanent note: always check BTC dominance before trading altcoins"
  • "List all notes"
  • "Delete permanent note 0"
  • "What's the current config?"
  • "Show me my KuCoin balances"
  • "What's the funding rate for XBTUSDTM?"

Backtesting

Run strategy backtests on historical data with parameter sweeps.

python -m src.backtest --symbol BTC-USDT --interval 1hour --lookback_hours 240 \
  --buy_rsi 55 --stop_atr_mult 1.5 --target_atr_mult 2.0 --fee 0.001

The backtester uses EMA crossover + RSI + MACD histogram for entries, ATR-based stops and targets, and computes total return %, win rate, profit factor, max drawdown, and best/worst trade. A parameter sweep mode scans ranges of buy_rsi, stop_atr_mult, target_atr_mult, and min_macd_hist to find optimal combinations.

How the Main Loop Works

Each polling cycle (POLL_INTERVAL_SEC seconds):

  1. Snapshot -- Fetches tickers, spot/futures/financial balances, open positions, stop orders, pending limit orders, recent fills, closed positions, and fee rates from KuCoin
  2. Reconciliation -- Sums USDT across all accounts, tracks daily drawdown per venue
  3. Price detection -- Compares prices with the last successful model-reviewed state. Each symbol learns an EWMA of ordinary poll noise and raises its trigger adaptively, bounded at PRICE_TRIGGER_MAX_MULTIPLIER× the configured floor (default 2× — safety-biased, so any move ≥ 2× the base trigger always earns a fresh model look even in the noisiest symbol; raise it to save more tokens), preventing oscillation from repeatedly calling the model
  4. Position extremes -- Updates peak/trough unrealized PnL for open positions
  5. Profit protection -- Ratchets stops to breakeven and caps give-back on live futures positions (code-driven, independent of the agent)
  6. Event tracking -- Logs triggered futures TP/SL closes as decisions (with exit price, for the no-chase guard)
  7. Circuit breakers -- Checks drawdown and consecutive losses against thresholds
  8. Agent run -- If triggers exist or the idle threshold is reached, starts one non-blocking Trading Agent run. Idle-only no-action cycles back off automatically up to FLAT_BACKOFF_MAX_MULTIPLIER; price/fill/risk events stay responsive. A pending atomic entry suppresses idle hunting and is managed by its deterministic lease/expiry instead of model babysitting
  9. Wait -- Sleeps until next cycle

Trigger types: initial:SYMBOL (new unreviewed symbol), price_move:SYMBOL:X.XX% (meaningful displacement), auto_trigger:SYMBOL:above|below:PRICE (one-shot, expiring explicit level), and idle_threshold (scheduled review). Crossed explicit triggers are persisted as pending events before consumption, so a restart cannot lose them. Cadence, productivity, reviewed prices, and learned price noise persist across restarts in agent_memory.json; deterministic protection still runs every poll.

Execution-quality metrics count a resting limit entry only when its recorded clientOid starts with traide-entry- (memory.is_limit_entry_record). Market and reduce-only close order IDs are excluded from limitFillRate. edgeReport.executionMap breaks the same records down by resting distance, so its totals always equal limitFillRate.

Project Structure

src/
  agent.py             Trading + Research agent assembly, system prompts, per-run context & helpers
  tools.py             All 47 agent tools (build_tools), organized by section: spot, futures, market data, screening, planning, news
  analytics.py         Technical indicators, regime detection, volume profile, multi-TF scoring, market state
  backtest.py          Strategy backtester with parameter sweeps
  config.py            Configuration dataclasses, env var loading, validation
  conversation_memory.py  Supervisor conversation memory (rolling summary + recent exchanges)
  kucoin.py            KuCoin spot + futures API client (HMAC auth, retries, error handling)
  main.py              Main trading loop, snapshot building, circuit breakers, trigger detection
  memory.py            Agent memory store (trades, decisions, plans, Kelly, cooldowns; signal, exit and gate probes)
  position_context.py  Facts about the trade behind an open position (carry hold on the contract's funding clock, noise band, original risk, restart peak) + the model's entryThesis + exit-probe replay inputs
  protection.py        Code-driven profit guards: breakeven ratchet, give-back cap, no-chase (runs every poll)
  safety.py            Revocable background-run authority and serialized exchange-write lock
  supervisor.py        Supervisor agent tools (read logs, memory, config, edge + gate scoreboards, write notes)
  telegram.py          Telegram notification sender (async, background thread)
  telegram_bot.py      Telegram long-polling bot for Supervisor Agent
  utils.py             Symbol normalization utilities
  wsgi.py              Gunicorn WSGI shim for service deployment
scripts/
  resettle_probes_futures.py  One-off: re-settle stored probes on 1m futures bars (+ funding) and backfill stackR on agent exit probes; bot STOPPED
tests/
  test_analytics.py    Analytics and indicator tests
  test_config.py       Configuration validation tests
  test_conversation_memory.py  Conversation memory tests
  test_memory.py       Memory store tests
  test_protection.py   Profit-lock decision + no-chase guard tests
  test_telegram.py     Telegram notification tests
  test_utils.py        Utility function tests

Running Tests

python -m pytest tests/ -v

Deployment

Direct

python -m src.main

Gunicorn (service-style on Linux)

gunicorn -w 1 -b 0.0.0.0:8000 'src.wsgi:application'

Keep -w 1 to avoid multiple loops. http://localhost:8000/ returns a health check while the background trading thread runs.

systemd

Create /etc/systemd/system/traide.service:

[Unit]
Description=trAIde Trading Agent (Gunicorn)
After=network.target
Wants=network-online.target

[Service]
Type=simple
User=traide
Group=traide
WorkingDirectory=/opt/traide
Environment="PATH=/opt/traide/.venv/bin"
ExecStart=/opt/traide/.venv/bin/gunicorn -w 1 -b 0.0.0.0:8000 'src.wsgi:application'
Restart=always
RestartSec=5

[Install]
WantedBy=multi-user.target
sudo systemctl daemon-reload
sudo systemctl enable traide.service
sudo systemctl start traide.service

Logs: journalctl -u traide.service -f

Quick setup script

sudo SERVICE_USER=$(whoami) ./setup_service.sh

or

sudo bash setup_service.sh

Environment overrides: SERVICE_NAME, SERVICE_USER, SERVICE_GROUP, WORKDIR, VENV_PATH, BIND_ADDR, REQUIREMENTS_FILE, FORCE_PIP_INSTALL, DEPS_ONLY.

The script syncs Python dependencies before it restarts anything. It creates the virtualenv if missing and re-installs requirements.txt whenever that file has changed since the last successful install (tracked by a SHA-256 stamp at $VENV_PATH/.requirements.sha256), then restarts the service. Without this, a deploy that changed requirements.txt restarted into a venv still holding the old packages — which is how the server sat for months on a LangSmith client calling an endpoint due for retirement.

Details that matter if you change it:

  • pip runs as SERVICE_USER, never as root. Running it under sudo as root leaves root-owned files inside the venv that the service user can no longer write, breaking every later install.
  • No --upgrade. requirements.txt pins floors (>=), so a plain install already lifts anything below the floor; --upgrade would pull the newest of everything and walk off the combination those floors were verified against.
  • Nothing restarts unless the install verifies. After installing, the script runs pip check and then imports every module the installed requirements provide (derived from package metadata, so the list cannot drift). If either fails it exits non-zero before the unit is rewritten, so the previously working service keeps running. The stamp is only written on success, so a failed deploy retries next time instead of believing it succeeded.
  • DEPS_ONLY=1 bash setup_service.sh syncs dependencies and exits without touching systemd — no root required. FORCE_PIP_INSTALL=1 re-installs even when requirements.txt is unchanged.
# sync dependencies only (no root, no service changes)
DEPS_ONLY=1 bash setup_service.sh

One-off: re-settle stored probes on futures

Probes recorded before the Sep 25 2026 settlement fix were settled on the spot ticker against a futures base (see Settle on the market that fills above). sudo bash setup_service.sh does this for you (since Sep 26 2026, after the manual step was skipped on two deploys): it counts the rows the script would migrate — by the script's own selection rules — and, only when there are some, stops the service, runs the script with --apply --bot-stopped as the service user (the script writes its own timestamped backup and refuses to write if the file changes under it), then starts the service as usual. A finished migration counts 0, so later deploys skip the step; a failure never blocks the deploy (the bot starts and the next run retries). The first run takes a few minutes, during which open positions keep their exchange TP/SL brackets but the code trail pauses. SKIP_RESETTLE=1 sudo bash setup_service.sh skips it. To run it by hand instead — with the bot stopped (it rewrites the memory file every poll):

sudo systemctl stop traide
cp .agent_memory.json .agent_memory.json.manual.bak
.venv/bin/python -m scripts.resettle_probes_futures --memory .agent_memory.json           # dry run: before/after report
.venv/bin/python -m scripts.resettle_probes_futures --memory .agent_memory.json --apply   # write
sudo systemctl start traide

For every probe with no priceSource stamp it re-stamps m5/m15/m60/m240 from the 1m futures kline that closed at the due minute (a thin contract's last close is carried forward at most 3 minutes, otherwise the horizon is None), adds the funding credit f{h} from the contract's funding history, and stamps priceSource=futures_mark_resettled. Horizons the live loop missed (no spot ticker for the symbol, or probes older than the 5m/15m horizons) are measured too — each at its own due minute, never back-stamped. Rows are re-settled, never dropped; a symbol whose data cannot be fetched is left untouched and listed, so a re-run retries it. Every kline row is checked (high = row max, low = row min) and one violation aborts the run — futures klines are [ts, open, HIGH, LOW, close], unlike spot. Exit probes that were resolved after their measurable life become unmeasured, and every other spot-settled exit-probe bracket (a resolved outcome with no priceSource — the live settle stamps futures_mark on every resolution now, so a re-run never touches a row twice) is re-resolved on 1m futures bar closes, the poll-like convention the live mark settle uses, with the old outcome kept under supersededOutcome (ONE-USDT 09-21 00:04 was stored as a +1.70R take-profit and was stopped on the contract). Those rows feed otherExits and the trail's regime / market-state splits — the evidence the adaptive-trail decision waits on; the dry run lists every flipped outcome and prints otherExits / trailByRegime before and after. It also backfills stack (stackR / resolvedBy) on every agent-closed exit probe whose 8h horizon has passed and has none — the same live-exit-stack replay the poll loop now runs for new closes (see An early close is scored against the system). Rows recorded before the replay existed lack its inputs, so each one's entry is found in the trades ledger (same symbol and side, filled before the close, same original stop) and the row gains fillTs / initRiskPx / noiseBandR / holdUntilTs and the counterAtEntry / htfAligned tags as well; the entry comes from the trades ledger, or — once the 100-row ledger has pruned it (~2.5 days) — from the realized close record, which carries the same entry context; one whose entry or bars cannot be found stays bracket-only (excluded from the verdict) and is listed. Run it soon after deploy: the trades ledger holds only about 2.5 days of placements. The dry run prints taken / bracket / stack per close and exitDiscipline before and after (on the Sep 24 copy: −6.23R vs the bracket → −3.44R vs the stack, n=5). The replay reads the bot's profit-protection settings from the VM's config so it replays the stack the live loop runs; --skip-stack leaves exit probes alone. --apply writes its own timestamped backup first, refuses a file modified in the last 3 minutes unless --bot-stopped is passed, and refuses to write if the file changed while it ran. Public market-data endpoints only — no API keys are used.


Disclaimer

This software is for educational purposes only. USE THE SOFTWARE AT YOUR OWN RISK. THE AUTHORS AND ALL AFFILIATES ASSUME NO RESPONSIBILITY FOR YOUR TRADING RESULTS. Do not risk money that you are afraid to lose. There might be bugs in the code - this software DOES NOT come with ANY warranty.

About

GenAI Trader

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages