← How We Test  /  The Graveyard
The Graveyard

You've seen half of these sold as winning systems.

The "edges" in your ads and your favorite YouTuber's course — order-block fades, VWAP bounces, liquidity-sweep reversals, opening-range breakouts. We ran them all on 10+ years of real tick data. 95% died. This is where they're buried, and exactly what killed each one. We ship the 5% that lived.

5%
Of everything tested survives to release
7
Adversarial trials each must clear
Built twice
Clean-room rebuild before release
0
Unverified results shipped

Why publish the failures?

A test that cannot fail tells you nothing. Show only survivors and you can't tell a high bar from a lucky one. The kill rate is what shows the process has teeth — and the causes of death below are more useful than most of what gets sold as a strategy.

Causes of Death

Seven ways an edge turns out to be fake.

Grouped by what actually killed them — because the pattern is far more instructive than the individual strategy.

01

It was look-ahead all along

The most dangerous
Break-and-retest continuation
The rule armed and fired on the same bar — entering at a price the market had already left. It cleared the whole gauntlet, because every stage ran through the same leaky generator. The clean-room rebuild returned a coin flip.
KILLED · causal rebuild ≈ 49% · read the full case file →
Zone-confirmation screener
Entry was selected using the extreme of a window that hadn't finished forming yet. A tidy positive expectancy evaporated the moment entries were made strictly causal.
KILLED · edge was the hindsight, not the zone
Multi-timeframe confluence
A one-hour signal lit up when we combined it with a higher-timeframe filter — t-stat +6.18. Then we closed the higher-timeframe bar properly instead of reading it mid-formation, and it fell to t−0.57. One unclosed bar leaks the future, and the higher the timeframe the harder it leaks — here about 3.4×. Sealed up, 1,139 interaction tests came back null.
KILLED · an unclosed higher-TF bar is look-ahead
02

It only existed in the window we found it in

Overfitting
Opening-candle volume-at-price
Strong and stable — inside the search window. Extending the history to data the search had never touched killed it outright.
KILLED · out-of-sample
TICK-extreme washout reversal
Excellent through a single bull regime. Rolled forward onto a bear and a chop regime, it died. A t-stat from one regime says nothing about the next.
KILLED · regime-dependent
High/low-volume node rejection
Held up until the test period was extended by four years, at which point the effect disappeared. What looked like structure was a single-regime artifact.
KILLED · out-of-sample
Afternoon continuation
The morning version of this effect is real. We assumed the afternoon would behave similarly. It doesn't — the PM version tested significantly negative.
KILLED · assumption, not evidence
Opening-range volatility forecaster
An initial-balance reading that forecast the day's range — strong and genuinely useful, in 2020–21. That turned out to be a COVID-volatility artifact. On 2022-onward data the R² collapsed from +0.42 to +0.07 and it forecast worse than a plain 20-day median. The two sibling indices even disagreed within the same era — the tell that there was never any structure.
KILLED · a pandemic-vol artifact, not an edge
03

It failed the cross-market acid test

Data-mining
Expiry-week morning short
Statistically significant on one index. Re-run unchanged on its closest sibling market, the effect wasn't merely weaker — it wasn't there. A real structure shows up in both.
KILLED · single-market artifact
Auction ledge-versus-node read
A plausible, well-argued market-structure story. Tested across two markets and thousands of events, it left no residual edge in either. We also found a trap: a fixed profile bin size manufactured a phantom effect until the bins were scaled to the instrument.
KILLED · no residual, and the bin size was lying
04

The "signal" was a microstructure artifact

Plumbing, not alpha
Overnight reversal
Almost the entire effect came from the bid/ask bounce on the shared opening print. Delay the entry slightly to decontaminate it and the edge disappears.
KILLED · opening-print bounce
Cross-session momentum
Eight cells survived multiple-testing correction. All eight were the same opening-print artifact wearing different labels. Decontaminate before you correct for multiple testing, or you'll certify noise with great rigour.
KILLED · same artifact, eight disguises
Order-flow toxicity (VPIN)
The measure tracked venue fragmentation rather than informed trading. Its incremental explanatory power over a naive baseline was effectively nil.
KILLED · measuring the wrong thing
05

Real phenomenon — zero directional content

True but untradeable
Price-jump continuation and fade
The jump detector was sound; the jumps are unambiguously real. They simply carry no information about direction. Both the continuation and the fade versions failed.
KILLED · real event, no directional edge
The London session effect
London is genuinely real — in volatility. In direction it is absent. A positive control confirmed our test could detect the effect where it does exist, which is what let us trust the null here.
KILLED · real in vol, absent in direction
Initial-balance breakouts
Four separate attacks — level, exit, width, and catalyst-timing — all hit the same wall. An IB break is a volatility event on an efficient level. Narrow ranges were worse, because compression raises the cost hurdle.
KILLED · book closed after 4 attempts
Short-horizon order-flow features
Across a wide feature scan at 30–120 second horizons, information coefficients clustered around 0.003–0.008. That's noise. This one result explains most of our short-horizon graveyard: it was an inputs problem, not a rules problem.
KILLED · the inputs carry no signal
The "S&P leads the Nasdaq" cross-market signal
Everyone knows the S&P leads the Nasdaq. At one-minute bars the relationship is contemporaneous — correlation 0.94 — and the measurable lead is about 0.01. The real lead is milliseconds long and arbitraged away long before any retail bar prints. True, and completely untradeable at our resolution.
KILLED · real at microseconds, gone by the 1-min bar
Options-gamma regime filter
Dealer-gamma pinning near expiry is real — we could measure it. As a switch to turn strategies on and off it added nothing, and the pin itself isn't capturable at retail size and speed. The only usable residual was avoidance: below the gamma flip, ranges bleed — a volatility-regime tell, not an entry.
KILLED · real pin, non-tradeable
06

It never beat a fair baseline

Efficient market
The entire level-and-limit fade family
Order blocks, VWAP bands, value-area edges, prior-day and prior-week highs and lows, session extremes, fair-value gaps. Scored against the correct control — real bars shuffled, not a coin flip — every variant sat on the baseline. The market is efficient at these levels, and the pool type made no difference whatsoever.
KILLED · a dozen variants, one verdict
Barrier races scored against 50%
A family of setups looked to hit ~67% success. Against the correct random-walk baseline — stop distance over target plus stop — every cell scored at or below zero. The apparent edge was pure barrier geometry.
KILLED · the baseline was wrong, not the market
Volatility-scaled profit targets
Scaling targets to volatility tested profitable — and it would test profitable under any process with no edge at all. Matching a target's distance is not matching its probability (Jensen's inequality), so the method manufactures a positive from nothing. Re-scored on the barrier's actual hit-probability, the edge was zero.
KILLED · the metric was the edge, not the market
The trend-day / chop-day classifier
Know in advance whether a day will trend or chop and you'd know which playbook to run. We tried three independent measures of path shape, across every era and both indices. The one-day autocorrelation was ≈ 0 in every single cell. Path shape is noise — not just its direction, the whole shape.
KILLED · tomorrow's shape is unpredictable
07

It was a bug in our own code

The humbling one
A currency-futures port with six positive years
Six out of six profitable years on a new market. It traced to a stop placed relative to a moving average instead of relative to the entry — an implementation slip, not an edge. Fixing the bug erased the result entirely.
KILLED · our mistake, caught before release
A "not decaying" edge that was a clock error
A deployed book appeared stable across years. The stability was an artifact of a daylight-saving offset applied uniformly to dates where it didn't hold. Corrected, the edge was confined to a single regime and is now parked.
PARKED · certified only against the data it ran on
What The Graveyard Taught Us

The rules that came out of the failures.

Every one of these was bought with a dead strategy. They're now permanent parts of the process.

01

A validation pipeline catches luck and data-mining, but it is structurally blind to look-ahead — every stage inherits the same leak. Only an independent rebuild from zero shared code can catch it.

02

Score against the right baseline, never against 50%. Barrier races must be judged against stop-distance geometry; limit strategies against real bars shuffled. The wrong control manufactures edges out of nothing.

03

A significant coefficient is not a directional edge. Plenty of effects are real, replicate cleanly, and still tell you nothing about which way to trade.

04

Decontaminate before you correct. Multiple-testing correction applied to a contaminated result just certifies the artifact with more authority.

05

Check for unused history first. Before believing anything, ask whether there's data the search never touched — and then go run it there.

06

A "certified" number is only certified against the data it ran on. Nothing is certain in general; it holds only on a period, on an instrument, under a cost assumption.

07

A null needs a positive control. Before accepting "no effect here," prove your test can detect the effect somewhere it genuinely exists.

We buried what everyone else sells. Two models wouldn't die.

Everything we ship survived the process that killed the strategies above. See the survivors — and the record behind them.

Entries are summarised from our internal research log. Figures describe tests that failed and are provided to document methodology — they are not performance claims for any product. Trading futures involves substantial risk of loss.