ATR-VWAP Variant Study Research

Once E8 and E4 both turned out to be lookahead artifacts rather than real edges, the open question was whether any way of trading off ATR distance from VWAP works — fading it, riding it as momentum, at wider or tighter thresholds, over longer holds, conditioned on trend regime. This page tests that directly: ~15 variants and a 144-combo parameter grid, all on the corrected (non-lookahead) running VWAP/ATR, all cost-adjusted, with the headline result validated out-of-sample rather than just reported in-sample.

1 · Why test variants at all

E8's apparent edge turned out to be a VWAP computed with future information; E4's turned out to be an entry price that predates its own signal. Both bugs happened to point the same direction — toward inflating a mean-reversion story. That leaves a real question unanswered: with VWAP/ATR computed honestly, is there any tradeable structure at all around these levels, or does the whole ATR-band-around-VWAP idea just not hold up at 5-minute granularity? Rather than declare the idea dead from one exit scheme, this page runs the idea's main alternatives.

2 · Five core variants (1-hour hold, top-100 universe)

VariantTrades/symbolGross win%Net win%Gross PFNet PFNet avg R
Fade to VWAP baseline, from E8 analysis2,20833.0%32.3%1.040.57−0.79
Fade to 1×ATR partial pullback, not full touch1,87632.6%28.0%1.060.60−0.74
Momentum @ 2×ATR, fixed 2R target2,24835.6%32.2%1.020.55−0.78
Momentum @ 3×ATR, fixed 2R target1,82935.1%31.7%0.990.54−0.82
Momentum @ 2×ATR, structural target next ATR band2,40453.2%30.8%0.830.26−0.88
Momentum @ 3×ATR, structural target1,95753.1%31.2%0.820.26−0.90

A closer, more "achievable" fade target doesn't help (1.04→1.06 gross) — the problem was never target distance. Momentum is a coin flip at best (gross PF ~1.0-1.02) and actively bad with a structural target — and this is where win rate tells the real story: the structural-target variants win 53% of the time gross, comfortably "directionally correct" by the usual read of a win rate above half. But PF is still below 1.0 (0.82-0.83) — at that win rate, breakeven needs the average win to be at least as large as the average loss, and here it's only about 0.73x as large. Win rate alone was hiding a bad payoff ratio. And every variant's win rate drops several points from gross to net purely from the round-trip cost reclassifying marginal winners as losers, which is the fee-erosion effect directly.

3 · Does a longer hold help?

Extending the time stop from 60 minutes out to 24 hours, same entries, same 1×ATR stop:

HorizonFade gross win%Fade net win%Fade net PFMom. gross win%Mom. net win%Mom. net PF
1 hour34.8%30.8%0.57535.6%32.2%0.548
4 hours30.9%29.8%0.61333.7%32.3%0.569
8 hours30.7%29.7%0.61633.7%32.3%0.571
24 hours30.7%29.7%0.61633.7%32.3%0.572

Win rate barely moves past 4 hours either, in step with PF — confirming there's no slow-building edge being cut short by an early time stop. Performance saturates almost immediately after 4 hours — nearly every trade resolves via stop or target long before a 24-hour time stop would ever bind, so extending further changes almost nothing. The small gain from 1h→4h is just fewer trades forced out at an unfavorable time-stop price, not a stronger underlying signal. This isn't a "too fast" problem; waiting longer doesn't reveal an edge that was there all along.

4 · Regime conditioning — a real, if modest, surprise

Splitting entries by a causal (no-lookahead, trailing 4-hour window) Kaufman efficiency ratio — trending: ER≥0.30, ranging: ER<0.15 — to test the intuitive pairing of fade-in-chop / momentum-in-trend:

Variant · regimeTradesGross win%Net win%Gross PFNet PF
Fade, all regimes (baseline)178,16430.7%29.7%1.050.62
Fade, ranging the intuitive pairing115,12030.9%29.9%1.050.60
Fade, trending the opposite pairing8,03130.5%29.6%1.120.73
Momentum, all regimes (baseline)220,54333.7%32.3%1.020.57
Momentum, trending the intuitive pairing9,07533.1%31.6%0.990.62
Momentum, ranging144,14133.9%32.4%1.040.57

Win rate is nearly flat across every regime split (30-34%) — the trend-fade result's PF lift comes from slightly larger wins relative to losses in that bucket, not from winning more often. That's worth flagging on its own: a PF improvement with an unchanged win rate is a thinner, less intuitive kind of signal, and part of why it didn't survive the out-of-sample check below.

Fading a deviation during a trend outperforms fading it during a range — backwards from the standard "fade the chop, ride the trend" intuition. Reading: in a genuine trend, a 2×ATR VWAP deviation is more likely a real, exhausted pullback opportunity; in a directionless market the same deviation is more likely just noise wandering around a VWAP that isn't anchoring anything. Momentum's apparent lift in the trending bucket is a false positive — its gross PF there is actually below 1.0 (0.99); the net-PF improvement is just fewer trades meaning less cost drag, not a real edge. Only the trend-fade result (gross PF 1.12) looked like a genuine signal worth chasing further.

5 · Chasing the trend-fade lead — and why it doesn't hold up

A 144-combo grid search around the trend-fade idea (stop 1.0–2.5×ATR, target 0–1.5×ATR pullback, trigger 2.0–3.0×ATR, trend threshold ER 0.30–0.50) found 22 combos with net PF above 1.0. The best: trigger 2.5×ATR, target 1.5×ATR, stop 2.0×ATR, ER≥0.5 → net PF 1.39, gross PF 1.60. That sounds like a discovery. It had 327 total trades across 100 symbols — about 3 per symbol over 8 months.

A 144-way search finding a handful of thin, extreme-parameter combos that clear PF 1.0 is exactly what a multiple-comparisons problem looks like, not evidence of an edge. The only honest test is: pick parameters using only the first part of the data, then check them, unchanged, against data the selection never saw.

Config selected on Jan–Apr (train)Train PF netTest PF netTest net win%Test avg R, 95% CI
trig 3.0 · tgt 1.0 · stop 2.5 · ER≥0.50.960.8837.2%−0.46  [−0.84, −0.09]
trig 3.0 · tgt 1.5 · stop 2.5 · ER≥0.50.880.9836.7%−0.40  [−0.78, −0.02]
trig 3.0 · tgt 1.5 · stop 2.5 · ER≥0.40.871.0936.8%−0.24  [−0.35, −0.13]
trig 3.0 · tgt 1.0 · stop 2.0 · ER≥0.50.860.9734.0%−0.54  [−1.00, −0.08]
trig 2.5 · tgt 1.5 · stop 2.5 · ER≥0.40.861.1738.4%−0.29  [−0.42, −0.17]

Win rate sits around 34-38% out of sample for all five — actually the highest win rates seen anywhere on this page, since a wide stop (2-2.5×ATR) against a modest target (1-1.5×ATR) mechanically raises hit rate. It still isn't enough: the payoff ratio needed to break even at a 37% win rate is roughly 1.7-to-1 (avg win to avg loss), and none of these clear it out of sample. Higher win rate without a matching payoff ratio is the same trap as the structural-momentum result in §2, just from the opposite direction. The best in-sample config (min. 100 trades to even qualify) only reached PF 0.96 — the 1.39 result from the unrestricted search doesn't reappear once thin samples are excluded. And of the top 5 configs the training window selected, every single one has a 95% confidence interval on out-of-sample average R that sits entirely below zero — including two whose test PF crept just over 1.0. That's not "inconclusive," that's the data actively saying no. For comparison, the untuned baseline (2.0×ATR / full VWAP target / 1×ATR stop / ER≥0.30, no cherry-picking at all) scores PF 0.70 on train and 0.72 on test — a boringly stable, consistently unprofitable number, which is exactly what an honest measurement should look like.

Full-period search "winner" n=327, overfit
1.39
Best OOS-validated config n=1264
1.09
Untuned baseline, train
0.70
Untuned baseline, test
0.72

Grey tick = breakeven (PF 1.0). The "winner" only looks good because it's thin and unvalidated; every validated number sits near or below breakeven.

6 · Re-validated at 1-minute execution precision

Everything above runs on 5-minute bars — the same resolution as every entry check and every SL/TP walk-forward. Two things that fixes deliberately leave on the table: intrabar sequencing (when a bar's high and low both cross stop and target, which happened first?) and a fixed target getting gapped past within a single wide bar — the exact mechanism behind the E4 entry bug and the structural-momentum artifact in §2. So the whole study was rebuilt at 1-minute execution precision to check whether either problem was inflating or deflating the picture above.

Not a naive "redo everything at 1-min": ATR computed directly on 1-minute bars came out ~2.5x smaller than the 5-minute version for the identical 75-minute window — average bar range doesn't scale linearly with duration, so finer bars intrinsically undersize it. Using that would make every stop ~2.5x too tight, a different and noisier strategy, not a more precise version of the same one. The fix: VWAP, ATR, and the deviation threshold are still computed on 5-minute resampled bars (the actual strategy definition), and those values become available on the 1-minute timeline only once their 5-minute bar has fully closed (no lookahead) — entries and the stop/target walk-forward then run minute-by-minute against them. The same scaling trap resurfaced independently in the trend/range efficiency ratio (computed straight from 1-minute closes, it collapsed the "trending" bucket from 8,031 trades to 152) and was fixed the same way. Both are exactly the kind of subtle, easy-to-miss error this whole exercise is meant to catch — worth naming rather than quietly correcting off-page.

Result5-min net PF1-min net PF
Fade to VWAP0.5750.577
Fade to 1×ATR0.5950.600
Momentum, fixed 2R0.5480.532
Momentum, structural target TP-mislabel 10.5%→6-7%0.260.31
Fade, 24h horizon (all regimes)0.6160.608
Fade, trending regime0.7290.739

The core, horizon, and regime results replicate almost exactly — this is the reassuring part. The structural-target variants improved somewhat, as expected once a fixed target is harder to gap past (net PF 0.26→0.31), but stayed clearly unprofitable; the TP-mislabel rate dropped by roughly a third but didn't disappear. None of this changes the verdict from §1-4.

The trend-fade OOS result: worse the second time you check it

Re-running the §5 grid search and OOS split at 1-minute precision initially looked like a real improvement — out-of-sample net PF up to 1.35 on the top training-selected configs. Checking why turned up the same tokenized-instrument distortion that inflated results early in this whole research line: XAUUSDT, XAUTUSDT, and PAXGUSDT trades with risk_pct as low as 0.0014%, which turn an ordinary 0.15% loss into an R-multiple past −100. After excluding those 8 symbols and computing proper 95% confidence intervals on the actual percentage return (not just the risk-normalized R), the picture is genuinely mixed: raw average return per trade is small and positive (+0.08% to +0.27%) but the CI includes zero in 4 of 5 configs; the risk-normalized average R is small and negative, with a CI entirely below zero for the largest sample (1,049 trades). Two honest metrics disagreeing in sign is what a result sitting at the noise floor looks like — this is reported as unresolved, not as a discovery in either direction.

7 · What this means

This isn't a failure to find the right parameters — it's ~15 structurally different ways of using the same idea (fade, momentum, partial targets, structural targets, four hold lengths, regime conditioning) plus a 144-point parameter sweep, cross-checked at two different bar resolutions, all converging on the same answer. The honest read is that a 2×ATR deviation from a correctly-computed VWAP just isn't informative on Bybit perps in this period, in either direction, at 5-minute or 1-minute granularity. If there's a real version of this idea, it likely needs either a different anchor entirely (a shorter or differently-weighted VWAP, or a structural level rather than a volume-weighted average), a much coarser timescale where intraday noise stops dominating, or a filter this study didn't try (order-flow/CVD conditioning, cross-asset regime, or funding-rate-driven positioning skew) — not a further tune of stop and target distance on this same signal.

Universe: Bybit top-100 perps by 24h turnover, Jan 1 – Aug 25 2026. §1-5 run on 5-minute bars throughout (entries, VWAP/ATR, and the SL/TP walk-forward all at 5-min resolution); §6 reruns the same signal definitions with entries and SL/TP execution at 1-minute precision while keeping VWAP/ATR/efficiency-ratio computed on 5-minute resampled bars (see §6 for why). VWAP/ATR: session-cumulative running VWAP and Wilder ATR(15), the corrected (non-lookahead) versions validated in the E8 analysis. Entries: fresh crossing of the deviation threshold only (no re-chasing mid-excursion), one position at a time per symbol. Costs: 0.055% taker + 0.02% slippage per side, round trip, applied to every result on this page. Efficiency ratio: Kaufman ER over a trailing, causal 48-bar/4h window (5-min bars) or 240-minute/4h window (1-min bars) - a purely backward-looking regime label, no future information. Train/test split: Jan 1 – Apr 30 2026 (train) vs. May 1 – Aug 25 2026 (test), roughly even halves of the sample.