1 · Why test variants at all
E8's apparent edge turned out to be a VWAP computed with future information; E4's turned out to be an entry price that predates its own signal. Both bugs happened to point the same direction — toward inflating a mean-reversion story. That leaves a real question unanswered: with VWAP/ATR computed honestly, is there any tradeable structure at all around these levels, or does the whole ATR-band-around-VWAP idea just not hold up at 5-minute granularity? Rather than declare the idea dead from one exit scheme, this page runs the idea's main alternatives.
2 · Five core variants (1-hour hold, top-100 universe)
| Variant | Trades/symbol | Gross win% | Net win% | Gross PF | Net PF | Net avg R |
|---|---|---|---|---|---|---|
| Fade to VWAP baseline, from E8 analysis | 2,208 | 33.0% | 32.3% | 1.04 | 0.57 | −0.79 |
| Fade to 1×ATR partial pullback, not full touch | 1,876 | 32.6% | 28.0% | 1.06 | 0.60 | −0.74 |
| Momentum @ 2×ATR, fixed 2R target | 2,248 | 35.6% | 32.2% | 1.02 | 0.55 | −0.78 |
| Momentum @ 3×ATR, fixed 2R target | 1,829 | 35.1% | 31.7% | 0.99 | 0.54 | −0.82 |
| Momentum @ 2×ATR, structural target next ATR band | 2,404 | 53.2% | 30.8% | 0.83 | 0.26 | −0.88 |
| Momentum @ 3×ATR, structural target | 1,957 | 53.1% | 31.2% | 0.82 | 0.26 | −0.90 |
A closer, more "achievable" fade target doesn't help (1.04→1.06 gross) — the problem was never target distance. Momentum is a coin flip at best (gross PF ~1.0-1.02) and actively bad with a structural target — and this is where win rate tells the real story: the structural-target variants win 53% of the time gross, comfortably "directionally correct" by the usual read of a win rate above half. But PF is still below 1.0 (0.82-0.83) — at that win rate, breakeven needs the average win to be at least as large as the average loss, and here it's only about 0.73x as large. Win rate alone was hiding a bad payoff ratio. And every variant's win rate drops several points from gross to net purely from the round-trip cost reclassifying marginal winners as losers, which is the fee-erosion effect directly.
3 · Does a longer hold help?
Extending the time stop from 60 minutes out to 24 hours, same entries, same 1×ATR stop:
| Horizon | Fade gross win% | Fade net win% | Fade net PF | Mom. gross win% | Mom. net win% | Mom. net PF |
|---|---|---|---|---|---|---|
| 1 hour | 34.8% | 30.8% | 0.575 | 35.6% | 32.2% | 0.548 |
| 4 hours | 30.9% | 29.8% | 0.613 | 33.7% | 32.3% | 0.569 |
| 8 hours | 30.7% | 29.7% | 0.616 | 33.7% | 32.3% | 0.571 |
| 24 hours | 30.7% | 29.7% | 0.616 | 33.7% | 32.3% | 0.572 |
Win rate barely moves past 4 hours either, in step with PF — confirming there's no slow-building edge being cut short by an early time stop. Performance saturates almost immediately after 4 hours — nearly every trade resolves via stop or target long before a 24-hour time stop would ever bind, so extending further changes almost nothing. The small gain from 1h→4h is just fewer trades forced out at an unfavorable time-stop price, not a stronger underlying signal. This isn't a "too fast" problem; waiting longer doesn't reveal an edge that was there all along.
4 · Regime conditioning — a real, if modest, surprise
Splitting entries by a causal (no-lookahead, trailing 4-hour window) Kaufman efficiency ratio — trending: ER≥0.30, ranging: ER<0.15 — to test the intuitive pairing of fade-in-chop / momentum-in-trend:
| Variant · regime | Trades | Gross win% | Net win% | Gross PF | Net PF |
|---|---|---|---|---|---|
| Fade, all regimes (baseline) | 178,164 | 30.7% | 29.7% | 1.05 | 0.62 |
| Fade, ranging the intuitive pairing | 115,120 | 30.9% | 29.9% | 1.05 | 0.60 |
| Fade, trending the opposite pairing | 8,031 | 30.5% | 29.6% | 1.12 | 0.73 |
| Momentum, all regimes (baseline) | 220,543 | 33.7% | 32.3% | 1.02 | 0.57 |
| Momentum, trending the intuitive pairing | 9,075 | 33.1% | 31.6% | 0.99 | 0.62 |
| Momentum, ranging | 144,141 | 33.9% | 32.4% | 1.04 | 0.57 |
Win rate is nearly flat across every regime split (30-34%) — the trend-fade result's PF lift comes from slightly larger wins relative to losses in that bucket, not from winning more often. That's worth flagging on its own: a PF improvement with an unchanged win rate is a thinner, less intuitive kind of signal, and part of why it didn't survive the out-of-sample check below.
Fading a deviation during a trend outperforms fading it during a range — backwards from the standard "fade the chop, ride the trend" intuition. Reading: in a genuine trend, a 2×ATR VWAP deviation is more likely a real, exhausted pullback opportunity; in a directionless market the same deviation is more likely just noise wandering around a VWAP that isn't anchoring anything. Momentum's apparent lift in the trending bucket is a false positive — its gross PF there is actually below 1.0 (0.99); the net-PF improvement is just fewer trades meaning less cost drag, not a real edge. Only the trend-fade result (gross PF 1.12) looked like a genuine signal worth chasing further.
5 · Chasing the trend-fade lead — and why it doesn't hold up
A 144-combo grid search around the trend-fade idea (stop 1.0–2.5×ATR, target 0–1.5×ATR pullback, trigger 2.0–3.0×ATR, trend threshold ER 0.30–0.50) found 22 combos with net PF above 1.0. The best: trigger 2.5×ATR, target 1.5×ATR, stop 2.0×ATR, ER≥0.5 → net PF 1.39, gross PF 1.60. That sounds like a discovery. It had 327 total trades across 100 symbols — about 3 per symbol over 8 months.
A 144-way search finding a handful of thin, extreme-parameter combos that clear PF 1.0 is exactly what a multiple-comparisons problem looks like, not evidence of an edge. The only honest test is: pick parameters using only the first part of the data, then check them, unchanged, against data the selection never saw.
| Config selected on Jan–Apr (train) | Train PF net | Test PF net | Test net win% | Test avg R, 95% CI |
|---|---|---|---|---|
| trig 3.0 · tgt 1.0 · stop 2.5 · ER≥0.5 | 0.96 | 0.88 | 37.2% | −0.46 [−0.84, −0.09] |
| trig 3.0 · tgt 1.5 · stop 2.5 · ER≥0.5 | 0.88 | 0.98 | 36.7% | −0.40 [−0.78, −0.02] |
| trig 3.0 · tgt 1.5 · stop 2.5 · ER≥0.4 | 0.87 | 1.09 | 36.8% | −0.24 [−0.35, −0.13] |
| trig 3.0 · tgt 1.0 · stop 2.0 · ER≥0.5 | 0.86 | 0.97 | 34.0% | −0.54 [−1.00, −0.08] |
| trig 2.5 · tgt 1.5 · stop 2.5 · ER≥0.4 | 0.86 | 1.17 | 38.4% | −0.29 [−0.42, −0.17] |
Win rate sits around 34-38% out of sample for all five — actually the highest win rates seen anywhere on this page, since a wide stop (2-2.5×ATR) against a modest target (1-1.5×ATR) mechanically raises hit rate. It still isn't enough: the payoff ratio needed to break even at a 37% win rate is roughly 1.7-to-1 (avg win to avg loss), and none of these clear it out of sample. Higher win rate without a matching payoff ratio is the same trap as the structural-momentum result in §2, just from the opposite direction. The best in-sample config (min. 100 trades to even qualify) only reached PF 0.96 — the 1.39 result from the unrestricted search doesn't reappear once thin samples are excluded. And of the top 5 configs the training window selected, every single one has a 95% confidence interval on out-of-sample average R that sits entirely below zero — including two whose test PF crept just over 1.0. That's not "inconclusive," that's the data actively saying no. For comparison, the untuned baseline (2.0×ATR / full VWAP target / 1×ATR stop / ER≥0.30, no cherry-picking at all) scores PF 0.70 on train and 0.72 on test — a boringly stable, consistently unprofitable number, which is exactly what an honest measurement should look like.
Grey tick = breakeven (PF 1.0). The "winner" only looks good because it's thin and unvalidated; every validated number sits near or below breakeven.
6 · Re-validated at 1-minute execution precision
Everything above runs on 5-minute bars — the same resolution as every entry check and every SL/TP walk-forward. Two things that fixes deliberately leave on the table: intrabar sequencing (when a bar's high and low both cross stop and target, which happened first?) and a fixed target getting gapped past within a single wide bar — the exact mechanism behind the E4 entry bug and the structural-momentum artifact in §2. So the whole study was rebuilt at 1-minute execution precision to check whether either problem was inflating or deflating the picture above.
Not a naive "redo everything at 1-min": ATR computed directly on 1-minute bars came out ~2.5x smaller than the 5-minute version for the identical 75-minute window — average bar range doesn't scale linearly with duration, so finer bars intrinsically undersize it. Using that would make every stop ~2.5x too tight, a different and noisier strategy, not a more precise version of the same one. The fix: VWAP, ATR, and the deviation threshold are still computed on 5-minute resampled bars (the actual strategy definition), and those values become available on the 1-minute timeline only once their 5-minute bar has fully closed (no lookahead) — entries and the stop/target walk-forward then run minute-by-minute against them. The same scaling trap resurfaced independently in the trend/range efficiency ratio (computed straight from 1-minute closes, it collapsed the "trending" bucket from 8,031 trades to 152) and was fixed the same way. Both are exactly the kind of subtle, easy-to-miss error this whole exercise is meant to catch — worth naming rather than quietly correcting off-page.
| Result | 5-min net PF | 1-min net PF |
|---|---|---|
| Fade to VWAP | 0.575 | 0.577 |
| Fade to 1×ATR | 0.595 | 0.600 |
| Momentum, fixed 2R | 0.548 | 0.532 |
| Momentum, structural target TP-mislabel 10.5%→6-7% | 0.26 | 0.31 |
| Fade, 24h horizon (all regimes) | 0.616 | 0.608 |
| Fade, trending regime | 0.729 | 0.739 |
The core, horizon, and regime results replicate almost exactly — this is the reassuring part. The structural-target variants improved somewhat, as expected once a fixed target is harder to gap past (net PF 0.26→0.31), but stayed clearly unprofitable; the TP-mislabel rate dropped by roughly a third but didn't disappear. None of this changes the verdict from §1-4.
The trend-fade OOS result: worse the second time you check it
Re-running the §5 grid search and OOS split at 1-minute precision initially looked like a real improvement — out-of-sample net PF up to 1.35 on the top training-selected configs. Checking why turned up the same tokenized-instrument distortion that inflated results early in this whole research line: XAUUSDT, XAUTUSDT, and PAXGUSDT trades with risk_pct as low as 0.0014%, which turn an ordinary 0.15% loss into an R-multiple past −100. After excluding those 8 symbols and computing proper 95% confidence intervals on the actual percentage return (not just the risk-normalized R), the picture is genuinely mixed: raw average return per trade is small and positive (+0.08% to +0.27%) but the CI includes zero in 4 of 5 configs; the risk-normalized average R is small and negative, with a CI entirely below zero for the largest sample (1,049 trades). Two honest metrics disagreeing in sign is what a result sitting at the noise floor looks like — this is reported as unresolved, not as a discovery in either direction.
7 · What this means
This isn't a failure to find the right parameters — it's ~15 structurally different ways of using the same idea (fade, momentum, partial targets, structural targets, four hold lengths, regime conditioning) plus a 144-point parameter sweep, cross-checked at two different bar resolutions, all converging on the same answer. The honest read is that a 2×ATR deviation from a correctly-computed VWAP just isn't informative on Bybit perps in this period, in either direction, at 5-minute or 1-minute granularity. If there's a real version of this idea, it likely needs either a different anchor entirely (a shorter or differently-weighted VWAP, or a structural level rather than a volume-weighted average), a much coarser timescale where intraday noise stops dominating, or a filter this study didn't try (order-flow/CVD conditioning, cross-asset regime, or funding-rate-driven positioning skew) — not a further tune of stop and target distance on this same signal.