Independent backtest audits

Your backtest says it works. Find out before your account does.

We try to break your backtest before the market does. Your strategy is re-simulated trade-for-trade on an independent, gap-verified 10-year OANDA dataset, and you get a plain answer on whether the edge survives out of sample. Most don't. Knowing that is worth more than the strategy.

Request an auditRead the case study

Why backtests lie

Not because you're careless. Because the failure modes are invisible from inside your own code. Look-ahead in an indicator warm-up. A stop and target that both sit inside one bar, resolved optimistically. A "best config" chosen from thousands, then scored on the same data that chose it. Spread charged once instead of on both sides. Each looks fine in review. Each inflates results.

The only defense is an independent re-simulation — a second engine, written separately, fed your rules and clean data, compared to yours trade by trade. If a single trade differs, one of them is wrong, and it's usually the flattering one. We've caught real bugs this way in our own engines, the ones we trusted most.

CASE STUDY · MACD CROSS STRATEGIES ON OANDA FX · SEP 2026

Three rounds of "improve the strategy." Same answer every time.

A retail FX trader had a daily MACD stop-and-reverse system with a genuine, independently validated edge. The question: could the same signal be made to work on faster timeframes, or improved with standard indicator filters? We built the test rig, then ran three escalating studies against a 323-million-row warehouse of bid/ask candles (68 pairs, M1–D, 2016–2026).

Every study used the same discipline: a fast vectorized simulator and a naive bar-by-bar simulator written separately, compared trade-for-trade before any number was reported. 60/40 time-split walk-forward. Configs selected on the training window only, then scored on the test window. Multiple-comparison rate disclosed every time.

StudyConfigsTest PF > 1Chance
1 · Every cross, spread-multiple TP/SL11,3404.5%~50%
2 · + RSI / ADX / EMA / Bollinger filters253,44024.4%~50%
3 · Strength-regime entry & exit54,11719.9%~50%

Fraction of all configurations profitable on the held-out test window. A strategy with no edge lands near 50% by luck. Below 50% means the approach is systematically paying more in spread than it earns.

FINDING 01

The best in-sample config collapsed out of sample. Every time.

Study 2, EUR/USD H1: the config selected on training data had a profit factor of 1.91. On the unseen test window: 0.39. Selecting the best of 11,520 candidates guarantees a flattering number; it does not guarantee a strategy.

FINDING 02

Verification caught a real bug that inflated trade counts ~20%.

The fast simulator treated any subsequent surviving signal as a reversal. With filters applied, consecutive signals can share a direction — which cannot reverse a position. The naive re-simulation disagreed on the first config compared; the fix changed results materially. No eyeball review would have found it.

FINDING 03

The "profitable" cells were single macro trends, not edge.

USD/JPY H4 showed a rolling out-of-sample profit factor of 3.86. Inspection: one 6-trade year during the 2023–24 yen collapse contributed more profit than the other eight years combined. Remove it and PF falls to ~1.1 on 440 trades. A headline number that depends on one event is not a system.

THE VERDICT

Stop searching this space. Fund the strategy that already works.

The client's existing daily system was the only design with an edge — and the 320,000 failed variants strengthened that conclusion by showing exactly which ingredient (the strength filter) was load-bearing. The report told him where not to spend the next year. That's the deliverable.

What an audit actually does

  1. Read your code and your claim.
    • Entry, exit, sizing, and fill assumptions
    • Data source and the numbers you believe
    • Code only — never broker keys, never account access
  2. Re-simulate independently.
    • Your rules, re-implemented as a naive bar-by-bar loop
    • Run on our gap-verified OANDA bid/ask data: 68 pairs, M1–D, 10 years, zero null quotes
    • Spread paid inside the fills; pessimistic intrabar ordering
  3. Diff trade-for-trade.
    • Entry bar, exit bar, exit reason, P&L — every trade
    • Mismatches located to the line
    • Equal totals with different trades is a fail, not a pass
  4. Walk it forward.
    • Time-split train/test
    • Your configuration selected on train, scored on test
    • If you tested N variants, we report how many would look this good by chance
  5. Verdict first, then evidence.
    • Supported, rejected, or conditional — one line
    • Then the trades and tables behind it
    • A negative verdict is a result. It is the cheapest risk decision you'll make.

Pricing

Sanity Check
$195
One strategy, one instrument, one timeframe. 48 hours.
  • Independent re-simulation on our data
  • Trade-for-trade diff with first mismatch located
  • Look-ahead, fill-timing and spread audit
  • Two-page written verdict
Start here
Research Engagement
From $1,500
Grid studies, engine builds, ongoing validation. Scoped quote.
  • Design and run a full parameter study with the discipline above
  • Build or harden your simulator against an independent one
  • Monthly re-validation of a live system
Discuss scope

Founding offer: the first three Full Verifications are $225 in exchange for an anonymized public case study. No profitability guarantees — we sell the truth about your numbers, not returns.

Questions worth asking first

What if my strategy passes?

Then you have something rare: an independent, reproducible confirmation with the trade log to prove it. A clean result is a real result, and it gets the same full report as a failing one.

What data do you use?

OANDA v20 bid/ask candles for all 68 currency instruments, M1 through D, September 2016 to present — 323M rows, gap-verified across every instrument × timeframe cell, zero null quotes. FX is our home market. Stocks, futures and crypto on request using your data or a purchased source, with provenance stated in the report.

Do you need my exchange keys or account?

Never. Code and data only. Materials are deleted 14 days after delivery and never published without written consent.

What formats can you audit?

Anything we can read as source code. Python (backtrader, vectorbt, backtesting.py, Freqtrade, plain pandas), Java, MQL4/MQL5 source (.mq4/.mq5), TradingView Pine Script, C# (NinjaTrader, cTrader), EasyLanguage, C++, Rust, JavaScript, or Excel models. We read the logic, re-implement it in our own engine, and run it on our data; where we can, we also run your code as-is and compare the two.

Compiled-only files (.ex4/.ex5 or any binary with no source) can't be read, so we audit the results instead: send the strategy tester report or a full trade list and we test it for overfitting, cost sensitivity and dependence on a single lucky period. A plain-English rule description works too, but then we're testing our build of your idea, not your backtest, and the report says so.

What do I send, and how?

Your source code, your results (a per-trade CSV is best: time, instrument, side, size, entry, exit, fees), where your price data came from, and roughly how many parameter combinations you tried before settling on this one. Email files up to 25 MB to [email protected]; for anything larger we'll send you a private upload link. Never send broker or exchange API keys or passwords. If one arrives by mistake we'll tell you, delete it, and ask you to rotate it.

How is this different from an automated backtest checker?

Automated tools run their own tests on your numbers. We read your code, rebuild the strategy independently, and compare the two trade by trade, which is how look-ahead, optimistic fills and position-handling bugs are found. A person signs the verdict and walks you through it.

Will you tell me how to fix it?

Yes — every defect comes with the fix. What we won't do is redesign your strategy into something profitable. That's research, not verification, and it's a separate engagement.

Is this investment advice?

No. We verify whether a backtest is an accurate simulation of its rules. Whether to trade those rules is your decision.

Request an audit

Tell us what you built and what it claims. We'll reply within one business day with a fixed quote and what to send.

Or email us directly: [email protected]

Opens your email app with the details filled in. Nothing is sent until you press send there.