# Learning-loop review — 2026-09-16 744 closed trades reviewed, 290 currently open. Proposal only — nothing here is applied automatically (CLAUDE.md §4: the AI review proposes, dan disposes). # AI Review — Trading Journal (per docs/METRICS.md) ## 1. Is overall expectancy holding? Pooled across all accounts: **expectancy +0.42% per trade, n=744, profit factor 1.61**. That raw count looks comfortably above the ~20-trade minimum, but two caveats matter before treating it as "holding": - **Clustering**: 598 of the 744 trades are `qsr`, and the power-check section shows those collapse to only **81 independent symbol-days**, not hundreds of independent observations — the effective sample is much smaller than 744. - **Account mix**: the positive pooled P&L is driven mostly by the paper accounts (`paper-main` +$3635 on 544 trades). The two live accounts are flat-to-negative: `live-1` **-$64.86 net on 104 closed trades**, `t212-isa` **-$2.99 on only 11 closed trades** (below the per-account 20-trade minimum, so not actionable on its own). Paper fills are optimistic (no slippage/liquidity limits), so the strongest evidence we have is closer to breakeven than the pooled number suggests. **Conclusion**: expectancy is nominally positive but not yet confirmed on live fills at adequate sample size. Treat the pooled +0.42% as encouraging, not proven. ## 2. Biggest drag bucket (≥20 trades) Several buckets clear 20 trades. The most negative expectancy among entry-signal buckets is **RSI 30–40, n=191, expectancy -1.20%**, versus the weighted expectancy of everything else (n=553) at ≈ **+0.98%** — a gap of **≈2.18 percentage points**. That gap is **below the 2.48pp noise floor** established in the statistical-power section. So even though n=191 comfortably clears the 20-trade minimum, **this bucket's underperformance cannot be distinguished from noise** given the measurement resolution the journal currently supports. It clears the count bar and still fails the resolution bar — flagging this explicitly rather than treating it as a finding. (Separately, `orb-v1`, n=68, expectancy ≈0.00% vs `qsr`'s 0.47%, is only a 0.47pp gap — nowhere near resolvable either. `swing-dip-v1` has **zero closed trades** in this journal at all, worth noting but not a "drag" since there's no data.) ## 3. Entry problem, exit problem, or regime problem? No bucket clears the noise floor, so a firm attribution isn't supportable. For general context, the book-wide MAE/MFE numbers don't show a clear exit-giveback pattern: **MAE winners -1.85%, MAE losers -5.23%, MFE losers avg +1.42%**. An MFE-losers value that high would need to be materially larger (multiple percentage points) to indicate we're routinely giving back winners before they turn into losses; 1.42% is modest. Losers' MAE (-5.23%) is in the same neighborhood as the live stop configuration (`stopAtrMult: 2`, `maxStopPct: 8`), consistent with stops doing roughly what they're sized to do, not obviously too tight. Nothing here points decisively at an entry, exit, or regime problem — the data is simply too noisy at the bucket level to diagnose which lever broke. ## 4. Proposed parameter change The candidate drag (RSI 30–40) is an **entry-signal/bucket-threshold** question, and its gap (2.18pp) is smaller than the noise floor (2.48pp). Per the anti-self-deception rules: **NO CHANGE to any RSI threshold.** Using the revisit-date table, an effect this size sits closest to the 1.00%-edge row (**124 symbol-days needed, projected revisit 2026-09-30**) — we should wait for that, or better, for a paired same-fill comparison rather than more band-splitting. Instead, proposing a **non-predictive, mechanical** change, which is fair game at current sample size: `qsr.maxHoldDays` is currently **`null`** (no calendar time-stop). The open-positions table shows a meaningful number of QSR legs sitting **30–37+ days** underwater (e.g. AEP -2.13% at 37.3d ×4, BTI -1.88% at 36.2d, CVS -4.28% at 30.3d, HON -0.69% at 35.3d) with no mechanism to force resolution. This is a capital-turnover/risk-sizing issue, not a claim about predicting returns. **Proposed change:** `qsr.maxHoldDays`: `null` → `30`. This is a single, small, mechanical change (bounding worst-case holding time), independently measurable after the change on new trades only, and doesn't touch any entry threshold. ## 5. Notable open positions - QSR has a very large open book (159 open in `paper-main` alone, plus dozens more across the other four accounts) — many positions are duplicated single-share/fractional-share entries across accounts for the same tickers on the same days. - Several positions have been open **30+ days with no time-stop and are underwater**: AEP (37.3d, -2.13%, ×4 legs), BTI (36–37d, -1.88%, multiple legs), CVS (30.3d, -4.28%, multiple legs), HON (35.3d, -0.69%). This directly motivated point 4. - A few longer-dated positions are healthy: EBAY (+5.50% at 36.3d), ISRG (+9.12% at ~8.2d) — so aging alone isn't uniformly bad, but the underwater aging cluster is worth watching. ## 6. QSR shadow comparison note The shadow window (2026-08-09 → 2026-08-23) is closed. Qualitatively: the legacy isTriggered+isBuy method fired only **48 times across 18 tickers**, versus the live `newlyA` method's **413 closed + 288 open across 117 tickers** in the same era — an order of magnitude more frequent and covering a far broader universe, with only **14 tickers overlapping** between the two approaches. This suggests the two selection methods are drawing from substantially different candidate pools, not just re-timing the same picks. This is observational only (the shadow trades never executed and can't be scored); a fuller outcome backtest comparing hypothetical shadow fills to real newlyA outcomes can now be requested. ## 7. t212-isa vs live-1 signal timing Not enough outcome data to say which gives the "better" signal — `t212-isa` only began real trading 2026-09-11 and has just 11 closed trades total, below any usable minimum. What can be said qualitatively: across 19 currently-open tickers held by both accounts, entry-basis divergence ranges from **0.0% (PLD, MS, BCS, BA — identical)** up to **~4.0% (BSX)** and **~3.9% (DASH)**, consistent with the documented mechanism that `t212-isa`'s pre-market-anchored entries can resolve to the prior day's frozen close while `live-1` waits for its post-open buffer. This is a timing/basis artifact, not evidence either strategy instance is better — no closed-trade outcome comparison exists yet to make that call.