# Analysing Good and Bad Trades The whole point of this system is that every trade produces data, and the data tells us what to change. This doc explains what we measure and — more importantly — what each number means when it's bad. The central design decision: **every trade stores the full `SignalSnapshot` from the moment of entry** (RSI, distance from MAs, relative volume, ATR, market regime). A trade without its entry context is just a P&L number; a trade *with* it is a labelled training example. ## The core scoreboard **Expectancy (% per trade)** — the single most important number. Mean return across all trades. Positive and stable = real edge. Everything else exists to explain *why* this number is what it is. A system with a 40% win rate can have great expectancy (small losses, big wins) and a 70% win-rate system can lose money (small wins, huge losses) — win rate alone is a vanity metric. **Profit factor** — gross profits ÷ gross losses. Below 1.0 the system loses money. 1.5 is workable, 2+ is good. More robust to outliers than expectancy because it's ratio-based. **Win rate + Wilson 95% lower bound** — we always report the Wilson lower bound next to the raw win rate. 6 wins from 8 trades reads as 75%, but the lower bound is ~41% — meaning we genuinely can't rule out a coin flip. This is the guard against the classic failure mode of trading-system development: celebrating noise. Rule of thumb: don't act on any bucket with fewer than ~20 trades, and treat 20–50 as suggestive only. **Payoff ratio** — average win ÷ average loss. Combined with win rate this fully determines expectancy. If expectancy degrades, this tells you which lever broke: are we winning less often, or winning less per win? **Max drawdown** — worst peak-to-trough fall of the cumulative P&L curve. This is the number that determines whether you'd actually keep running the system live. A strategy with great expectancy and a 40% drawdown gets abandoned by every human who trades it. ## The diagnostic pair: MAE and MFE These two are the real "good trade vs bad trade" analysis — they look *inside* each trade rather than at its outcome. **MAE (Maximum Adverse Excursion)** — the worst the trade looked while it was open. Diagnoses **entries**: - Winners with high MAE (e.g. finished +5% but was down 4% first) — the entry was early. We got paid, but by luck; a slightly tighter stop would have turned these into losses. Persistent pattern = enter later / demand more confirmation. - Losers with low MAE that hit the stop — the stop is tighter than the natural noise of the name. Compare the stop distance against the entry snapshot's `atrPct`: if we keep stopping out on names with high ATR%, stops should be ATR-scaled, not fixed-%. **MFE (Maximum Favorable Excursion)** — the best the trade looked while open. Diagnoses **exits**: - Losers with high MFE (was up 3%, exited −2%) — the trade *worked* and we gave it back. That's not an entry problem; the signal was right. It's a take-profit problem. Persistent pattern = tighten targets or add a trailing exit. - Winners whose exit ≈ MFE — targets are well-placed; we're capturing what the move offers. The summary stat we track: average MAE of winners vs losers, and average MFE of losers. When `avgMfeLosersPct` is materially positive, the fix is in exits, not signals. ## Attribution: which conditions make money? `bucketBy()` groups trades by any property of their entry snapshot and computes the full scoreboard per group. The standard cuts: - **By strategy** — the basic "which strategy earns its place" view. - **By RSI band at entry** — does "buy RSI 20–30" actually outperform "buy RSI 30–40"? This directly tunes entry thresholds with evidence instead of vibes. - **By distance from MA50** — is the sweet spot −5..−10% below MA50, or are the deepest dips (−10%+) actually falling knives? - **By relative volume** — do high-volume-day entries (capitulation) work better than quiet-drift entries? - **By market regime** (SPY above/below MA200) — most long strategies have their entire edge concentrated in bull regimes. If bear-regime expectancy is negative, the fix is a regime gate, not better entries. - **By symbol** — some names just don't suit the strategy. Persistent per-symbol losses → drop from universe. - **By strategy + symbol** (added 2026-07-19) — the direct "is this strategy working, and for which stocks" view. "By strategy" and "By symbol" are independent cuts, so neither alone shows the intersection (a strategy could look fine overall while quietly losing money on one specific name, or vice versa) — this cut pools trades by the combination instead. - **By exit reason** — what fraction of exits are stops vs targets vs time-stops? A rising stop-rate is an early warning before expectancy visibly degrades. - **By holding time** — if all the profit comes from trades that resolve in ≤3 days and the 10-day holds are flat, a time-stop adds expectancy for free. ## Execution quality **Slippage** — `entryPrice` vs `signalPrice` is stored on every trade. If the fill is consistently worse than the signal, the backtest is lying to us by that amount. This matters more as strategies get faster. ## The AI review (Phase 4, built 2026-07-19 — `pnpm learn`) The journal being plain JSONL is deliberate: a scheduled job hands the recent trades + bucket metrics to Claude, which writes a review answering, in order: (1) is overall expectancy holding, (2) which bucket is the biggest drag, (3) is the drag an entry problem (MAE pattern), an exit problem (MFE pattern), or a regime problem, and (4) one concrete, small parameter change to propose. One change at a time, human-approved, then measured for another few weeks. That's the learning loop. Runs **daily** after close, not weekly as originally sketched here — dan's call (2026-07-19): the ~20-trade-per-bucket minimum below already gates whether a run is actionable, so there's no cost to starting the cadence immediately. See CLAUDE.md §11 for the implementation (`src/scripts/learningReview.ts`) — it calls the Anthropic API directly via raw fetch, not a full Claude Code session, so it only ever sees what's in the prompt (this doc, the metrics report, current live params). ## Anti-self-deception rules These are the rules that keep the learning loop honest, written down so we don't quietly break them later: 1. Minimum ~20 trades before a bucket influences any decision; report Wilson bounds everywhere. 2. One parameter change at a time — otherwise attribution is impossible. 3. Never delete losing trades from the journal; the journal is append-only. 4. Judge changes on trades *after* the change was made, never re-scored history. 5. A bucket that looks great but has 5 trades is a hypothesis, not a result. 6. **Check the noise floor BEFORE running the study** (`pnpm power-check`). If the effect you are hunting is smaller than the smallest one your sample can resolve, the result is uninterpretable whichever way it lands — a null teaches nothing and a hit is almost certainly noise. This rule was added 2026-08-27 after it had already been broken. ## Why the signal studies kept coming up empty (2026-08-27) A long research thread produced six confident-looking findings about which entry conditions predict QSR's returns. Every one failed re-examination. That is not bad luck and the signals are not necessarily dead — **the measuring instrument was about 10x too coarse for the effects being hunted**, and nobody checked first. Measured on the real journal (`pnpm power-check` reproduces all of this): | | | |---|---| | SD of a symbol-day's 5-day forward return | **5.66%** | | Independent symbol-days | **81** (from 334 rows — see `journal/clustering.ts`) | | Smallest band-vs-band gap detectable at 95% | **2.48pp** | | Smallest correlation detectable at 95% | **r ≥ 0.218** | So "RSI 35–40 is the sweet spot" could only register if that band beat the rest by 2.48 percentage points over five days. Published technical-indicator correlations with forward returns run r ≈ 0.02–0.05; we can only resolve r ≥ 0.218. Those searches were unwinnable before they started. **The cost of simply waiting** is quadratic in the edge — halving the edge you want to detect quadruples the data needed: | True edge | Symbol-days needed | ≈ sessions | |---|---|---| | 1.00% | 124 | 28 | | 0.50% | 493 | **110** | | 0.25% | 1,971 | **438** | Realistic edges are the bottom rows. That is years at this trade rate. **A secondary cause, real but smaller: range restriction.** QSR only buys inside its own grade-A box, so the gates have already removed much of each metric's spread before we ask whether it predicts anything. Against the screener's own 301-name scan (2026-08-09), grade-A names keep this fraction of the universe's SD: days-since-low **0.24**, median up-move length **0.52**, RSI **0.66**, RelVol **0.71**, MA50 distance **0.76**. Narrowing a predictor's range mechanically attenuates its correlation (`power.ts#attenuatedCorrelation`). Do not oversell this — RSI at 0.66 is not collapsed, and two rhythm gates are one-sided minimums that leave grade-A names *more* spread than the universe (ratios 1.56 and 1.48), where the correction does not apply at all. Sample size is the dominant problem; this is a multiplier on top of it. ### What to do instead — pair, don't out-sample QSR's anchor and runner legs are the same stock, the same instant and the same fill price with two different exit rules: a natural paired experiment where stock, timing and regime all cancel. On the 61 pairs where both legs have closed, the SD of the *difference* is **2.62%** against **5.91%** unpaired, taking the noise floor from 2.48pp to **0.657pp** — roughly **3.8x the resolution from the same data**, bought purely with study design. (The pairing does not currently show a winner: runner −anchor is **+0.40pp**, 95% CI **−0.26 to +1.06**, runner ahead in **27 of 61** pairs. Not significant. An earlier qty-weighted version of this comparison suggested the runner was clearly better; it was comparing different things.) The practical rule: **stop band-splitting entries** (needs ~110+ sessions), and prefer questions that are either paired A/Bs on the same fill (~15–25 pairs) or deterministic and need no statistics at all — costs, slippage, capital turnover, whether stops actually fire.