# Learning-loop review — 2026-09-08 481 closed trades reviewed, 230 currently open. Proposal only — nothing here is applied automatically (CLAUDE.md §4: the AI review proposes, dan disposes). # 1. Is overall expectancy holding? Pooled expectancy is **+1.05% per trade** across **481 closed trades**, profit factor **1.70**, Wilson-lower-bound win rate **59.9%** (raw 64.2%). At the pooled level the sample is large enough that this isn't pure noise — but this number blends four very different strategies (qsr = 378 of the 481 trades) and four accounts with different risk configs and fill quality (paper vs live). The live account (`live-1`, 70 closed trades) is currently net **-$65.10**, which is the only real-money read we have and it's roughly flat/slightly negative — a useful caution against reading the pooled +1.05% as "the system is working," since most of the positive P&L is paper-simulated. # 2. Which bucket is the biggest drag? Several buckets clear n≥20 (orb-v1 n=68, qsr n=378, vwap-mr-v1 n=33, RSI bands, MA50-distance bands, relvol<0.8 n=343, regime SPY>MA200 n=396, several exit-reason buckets, qsr·NVO n=25, qsr·VST n=27). Of the true strategy-level cut, **orb-v1** is the weakest: n=68, expectancy **-0.00%**, PF **0.92** — the only strategy bucket at/near breakeven-or-negative with a real sample size. But: the gap between orb-v1 (-0.00%) and the pooled average (+1.05%) is only about **~1.0 percentage point**, which is **below the 2.48pp noise floor** established for bucket-vs-bucket comparisons. So even though orb-v1 clears the 20-trade minimum, this difference is **not statistically resolvable** from the rest of the book — it clears the n-bar but not the noise floor, and should be treated as unresolved, not as "the answer." (Exit-reason buckets like `stop_loss`, n=149, expectancy -5.21%, are much larger in magnitude, but that split is close to tautological — a stop-loss exit is a loss by construction — so it isn't a genuine "condition" drag in the entry/exit-diagnosis sense; it's discussed under point 3 instead.) # 3. Entry problem, exit problem, or regime problem? We don't have MAE/MFE broken out per strategy or per bucket in this report — only the overall figures: **MAE winners -1.82%, MAE losers -5.07%, MFE losers +1.14%**. So we can't cleanly attribute orb-v1's (unresolved) softness to entries vs exits specifically. At the whole-book level, though, the pattern doesn't look like a classic exit/give-back problem: `avgMfeLosersPct` (+1.14%) is modest, not the "was up 3%, exited down 2%" signature METRICS.md flags as a target/trailing-exit issue. It's closer to consistent with the stop-loss bucket's own numbers — stop_loss exits average -5.21% expectancy against MAE losers of -5.07%, i.e. losing trades are mostly running down close to their eventual stop-out level rather than being given back after looking good. That reads more like a stop-placement / risk question than a take-profit question, but this is a whole-book inference, not orb-v1-specific, and shouldn't be oversold given point 2's conclusion. # 4. Proposed parameter change No bucket both clears n≥20 **and** clears the noise floor, so per the anti-self-deception rules I am **not** proposing any entry-signal or bucket-threshold change (RSI band, relvol cutoff, MA50-distance gate, time-of-day gate, take-profit level chosen from bucket returns, etc.). The smallest edge our current sample (81 independent symbol-days) can resolve is ~1.0–1.5pp; a credible real edge (0.5% or less) needs ~493 symbol-days — projected revisit **2027-01-15** at the current pace (~2026-09-22 if the true edge were as large as 1.0%, which is itself not very credible). Instead, the one change I'll propose is a **non-predictive, exit-mechanics/risk-cap** tweak, which the rules allow at any sample size: > **qsr.maxHoldDays: `null` → `20`** Reasoning: this doesn't touch any entry signal or bucket threshold, so it isn't exposed to the noise-floor problem. Looking only at *closed* qsr trades, the strategy's own winning-exit mechanisms resolve well inside 20 days — trailing_stop exits average **10.0 days**, take_profit exits average **5.2 days**, and even the strongest large-winning symbols (qsr·TRGP 18.6d, qsr·AZN 20.2d) close out near or under that mark. A 20-day cap therefore would rarely have bound on the historical winners in this dataset, while providing a mechanical backstop against a position drifting indefinitely with no calendar exit (qsr currently has none). This is a single, small, reversible change; it should be judged only on trades entered after it's applied (rule #4), not on the currently-open aging positions noted below. # 5. Notable currently-open positions (context only, not evidence for #4) A large number of qsr open positions are aging well past the strategy's typical closed-trade hold times with `no calendar time-stop`: HDB (28/54 shares, **48.3 days**, -7.07% unrealized), several ETR/CI/AEP/TRP/BTI legs sitting **28–39 days**, and BTI in particular showing -3.19% unrealized across many aged legs. This is exactly the pattern that motivates the maxHoldDays proposal above, but it is *not* being used as the justification for it — the justification in #4 comes only from closed-trade hold-time distributions. Worth flagging separately: `live-1` currently has 42 open positions against only 70 closed, and its ~$480 equity means it's a price-biased subsample (per the account caveats above) — any read of its P&L should account for that. # 6. QSR shadow comparison note (observational only) The shadow window (2026-08-09 → 2026-08-23) is now **closed**, so a fuller backtest comparing hypothetical shadow-log outcomes to real newlyA outcomes can be requested. Qualitatively, the two selection methods look quite different in both volume and universe: the legacy isTriggered+isBuy method fired only **48** would-buy events across 18 tickers in that window, while the live newlyA method produced **207 closed + 214 open** entries across **63** tickers over the same-ish period — roughly 4x the ticker breadth and far higher frequency. Ticker overlap between the two methods is only 12 names (AEP, DDOG, IBKR, SU, COHR, TRP, BTI, ENB, NKE, ITUB, EBAY, VALE) out of 63+18 combined — the two logics are picking largely different names. This is worth a qualitative flag but, per the setup, is not scoreable evidence and is not used to support the parameter change in #4.