# Learning-loop review — 2026-09-04 461 closed trades reviewed, 176 currently open. Proposal only — nothing here is applied automatically (CLAUDE.md §4: the AI review proposes, dan disposes). ## 1. Is overall expectancy holding? Yes — pooled expectancy is **0.93% per trade** across **461 closed trades**, profit factor 1.65, max drawdown $1,675. At n=461 this comfortably clears the ~20-trade minimum for a headline read. But this number is a blend of very different strategies and accounts (qsr contributes 363 of the 461 trades and dominates the P&L), so "expectancy is holding" really means "qsr is holding" — see below. ## 2. Which bucket is the biggest drag? Restricting to buckets with **≥20 trades**: - By strategy: `orb-v1` (n=68, expectancy **-0.00%**, PF 0.92) and `vwap-mr-v1` (n=29, expectancy 0.25%, PF 1.35) lag `qsr` (n=363, expectancy 1.15%, PF 1.80). - By RSI band: `rsi 30-40` (n=103, expectancy **-0.41%**, PF 0.67) is the worst-performing band with a large sample, and it also carries the largest total dollar loss ($-438) of any attribution-cut bucket. `rsi 30-40` is the biggest drag by both n and $ impact. But check it against the noise floor: its gap vs. the overall pool (0.93% − (−0.41%) = 1.34pp) and vs. the next band `rsi 40-50` (1.65% − (−0.41%) = 2.06pp) are **both below the 2.48pp noise floor** established for bucket-vs-bucket comparisons on this journal. So even though this bucket clears the 20-trade bar, **the gap is not statistically resolvable** — it looks like a real problem but the instrument isn't fine enough to say so. Same caveat applies to `orb-v1`'s ~0.93pp gap vs. the pool. ## 3. Entry problem, exit problem, or regime problem? Because the RSI-band gap can't be resolved, I can't attribute *that specific bucket's* underperformance to entries or exits with confidence. Falling back to the portfolio-wide MAE/MFE numbers (which are just descriptive, not a new bucket claim): - MAE winners avg **-1.82%**, MAE losers avg **-5.13%** — losers go much deeper underwater before failing than winners ever visit, roughly matching the ~8% maxStopPct / 2×ATR stop distance, i.e. no obvious sign the stop is *tighter* than the noise (that would show up as low-MAE stop-outs). - MFE losers avg **1.14%** — losers were, on average, barely in profit before turning into losses. This is a **low** give-back number, which argues *against* a "the trade worked and we handed it back" exit problem. Put together, the pattern reads as mildly entry-side (trades that don't work tend not to have worked at all, rather than working-then-reversing), but this is a portfolio-wide description, not a confirmed diagnosis of the RSI 30-40 bucket specifically — that would need the resolvable-gap bar this bucket didn't clear. ## 4. Proposed parameter change **No entry-signal or threshold change.** The RSI 30-40 band, the strongest-looking "drag" candidate, has a gap below the 2.48pp noise floor. Per the anti-self-deception rules, acting on it would be indistinguishable from acting on noise. At the observed pace (~4.5 symbol-days/session), a 1.0% true edge becomes resolvable around **2026-09-18**; anything smaller (0.5% or below, which is realistic for RSI-type signals) isn't projected to be resolvable until **2027-01-11** or later. Recommendation: wait, don't touch `maxRsi14` or any RSI gate now. Instead, a **risk-sizing** change that doesn't depend on out-predicting the market: the QSR risk ladder is inconsistent with itself. `swing-dip-v1` halves size for its riskiest tier (`tier3RiskPerTradePct: 0.2` vs base `riskPerTradePct: 0.4` — a 50% cut), but **`QSR_RISK` gives tier-3 trades the exact same size as the base tier** (`tier3RiskPerTradePct: 0.12` = `riskPerTradePct: 0.12`, no reduction at all), even though QSR's own risk ladder (`riskLadderTier1MaxStopPct: 6`, `Tier2: 4`) explicitly recognizes that some entries carry wider stops / more risk than others. **Proposed change:** `QSR_RISK.perTrade.tier3RiskPerTradePct`: **0.12 → 0.06** (matching the 50% tier-3 haircut already used in swing-dip-v1). This is a pure position-sizing consistency fix, not a new signal claim — it doesn't require any additional statistical power, applies uniformly, and is trivially reversible/measurable (compare tier-3 trade $ P&L before/after on the same entries). One parameter, human-approved before applying. ## 5. Notable open positions - **QSR has no `maxHoldDays`** (`null`), and it shows: several positions have been open **20–45 days** with no resolution — HDB (2 legs, 44.3d, -3.72% each), ETR (6 legs, 30–35d, -0.62%), BTI (multiple legs, 23–25d, -3.03%), TRP (25d, -0.93%), ENB (22d, -1.45%). None of these are catastrophic individually, but a cluster of stale, underwater, uncapped-duration positions is exactly the kind of thing a time-stop is meant to catch — worth watching even though this is not evidence for point 4. - **crypto-trend-v1** has three positions open 15–45 days (BTC, ETH, DOGE) against a strategy with only **1 closed trade in the entire journal** — essentially unvalidated live exposure with no calendar stop. - Nothing currently open is catastrophically underwater (worst is BTI/HDB around -3–4%), so this is a duration/hygiene observation, not a P&L emergency. ## 6. QSR shadow comparison note The shadow-log window (2026-08-09 → 2026-08-23) is now **closed**, so a fuller outcome backtest comparing legacy shadow "would-buys" to real newlyA outcomes can be requested. Qualitatively, two things stand out: - **Frequency**: real newlyA entries (192 closed + 159 open = 351 over the comparison period) vastly outnumber legacy shadow would-buys (48) — newlyA fires roughly **4x more often** than the old isTriggered+isBuy logic would have. - **Coverage**: newlyA traded 57 distinct tickers vs. legacy's 18, with only **12 tickers overlapping** between the two methods. The two selection methods are picking largely different names, not just timing the same names differently. This is observational only — the shadow trades never executed and per METRICS.md cannot be scored or used to justify any parameter change.