Coldshim is a red team for trading strategies. We replay yours across a broad universe of markets and timeframes, with a real broker's costs, against thousands of random draws. What survives is rare — and what doesn't, you'll know before you put money on it.
Have a strategy checked See a sample report170 public strategies already tested — See the results →
The strategy tester answers "what would this script have made on this chart?" Coldshim answers "does this edge exist, at your broker, once the costs are paid?"
| Question | TradingView backtest | Coldshim report |
|---|---|---|
| Data | Chart feed, often different from the broker's | ✓M1 bars from your broker's MT5 server |
| Spread | Fixed or zero, entered by hand | ✓Real spread minute by minute, buy at ask, sell at bid |
| Commission & swap | Approximate or absent | ✓Commission measured from history, swap per direction, per night |
| Intrabar execution | Next bar's opening price | ✓Trade-list replay (Path 2): every order placed at the minute its price is actually touched |
| Chance | Not tested | ✓200 placebo draws: same durations, same directions, random dates |
| Stability over time | A single curve | ✓Three chronological thirds, each one must be positive |
| Generalisation | One symbol, one timeframe | ✓several symbols × several neighbouring timeframes |
| Lookahead | Not detected | ✓Future prices perturbed: past signals must stay identical |
| Verdict | A curve and a percentage | ✓ROBUST, REJECTED or INCONCLUSIVE, with every calculation |
A beautiful backtest is most often a portrait of past chance, not of an edge. The phenomenon has a name: overfitting.
Ask 50 people to flip a coin 10 times. One of them will probably get 9 heads out of 10 — no talent, just mechanics. Test 50 strategies, keep the best, and you get the same thing: a champion of chance that looks like an edge.
You try EMA 9/21, then 8/23, then 10/50, and keep the best. With each setting, the strategy clings a little more to the noise of a past that won't repeat.
You test 50 ideas and keep only the one that worked. None was optimised — and yet the result is just as false.
It isn't a trader's flaw. A tester that instantly recalculates with every shifted setting is an overfitting machine that feels like research. Many scripts and courses sold with perfect curves are its direct product — sometimes without their authors knowing.
| The trap | Coldshim control |
|---|---|
| The key filed for a single lock | ✓several symbols × several neighbouring timeframes |
| What only works on one period | ✓Three chronological thirds, each one positive |
| What chance alone would have produced | ✓200 placebos to beat |
| A protocol chosen after seeing the result | ✓Rules of engagement fixed before the test |
| A forgotten cost that manufactures a false edge | ✓Real broker spread, commission and swap |
"Too short" when it's short, "too old" when it's long — the question is badly posed. What matters is enough trades (at least ~300), at least one downturn and one high-volatility episode, on a market that's still the same today. For scalping or intraday: 3 to 5 years of broker data. And the protocol is decided before looking at the result.
Coldshim applies to your strategy what SECURIX applies to your systems: an attack run for you, not against you. We don't look for the strategy that wins; we attack yours to find out whether it holds.
| SECURIX red team (Scrytide, Nightprill) | Coldshim | |
|---|---|---|
| What we attack | Your perimeter, your brand | ✓Your trading strategy |
| Rules of engagement | Perimeter signed before the audit | ✓Symbols, timeframes, broker and history fixed before the test |
| Tools verified | Detection of known flaws | ✓Calibration: chance rejected, a future-reading signal caught, a planted edge found |
| False positives discarded | Every finding verified | ✓200 placebos: an edge must beat chance |
| Evidence | A reproducible finding | ✓A reproducible report, SHA-256 fingerprints |
| Independence | The auditor doesn't fix what it audits | ✓We don't optimise your parameters |
| Recurrence | The threat evolves | ✓Market regimes and spreads evolve |
Whoever commissions a red team isn't the one who wrote the code — it's the one who carries the risk. For a strategy, that's whoever allocates the capital.
A Coldshim validation replaces a quant analyst's work on one strategy — roughly 2 to 4 working days — delivered in 2 business days.
And you're not paying for a curve: you're paying to know whether to stop — before you put real money behind it.
Each one is applied to every strategy. A result on a single symbol and a single timeframe is not a result.
M1 bars pulled from the MT5 terminal connected to the broker's real server, up to 3 years of history.
For each symbol and timeframe, the spread / candle-range ratio: green where the cost is negligible, red where scalping is mathematically lost before it starts. You know before testing.
Every minute's spread, buy at the ask, sell at the bid; a deliberately conservative percentile is the reference cost.
Computed from the account's actual trade history, not copied from a rate card.
Per direction, per night, converted whatever the broker's mode; an unsupported mode is flagged, never silently counted as zero.
The Pine Script is translated with its exact semantics — down to the edge cases where a careless port and TradingView quietly disagree. When a faithful translation can't be guaranteed, we say so rather than guess.
We test whether any signal depends on data it couldn't have had at the time — if it does, it's reading the future.
Several markets across a ladder of timeframes — an edge that lives on one timeframe alone is noise.
200 ghost strategies with the same durations and directions, at random dates. The real strategy must beat the large majority of them.
History is cut into three periods; each one must be positive. An edge carried by a single lucky stretch fails here.
The gross edge must clear the cost by a comfortable margin, over a meaningful number of trades.
ROBUST only if all of this holds across several neighbouring timeframes and several markets. A single box is not a result.
Before every report it runs on three control markets: a random signal must be rejected, a signal that reads the future must be caught, and a deliberately planted edge must be found. If one of the three fails, the report is marked "not certified."
Clearing the grid is not enough. Test hundreds of strategies and some pass by luck — only unseen dates can tell them apart.
ROBUST requires the edge to hold across several neighbouring timeframes and several markets — not one lucky box.
One single replay on a period no earlier step has seen, against a success rule set in advance, before the calculation.
A rule set in advance stops the result being retouched after the fact. Without the vault, those four strategies could have been presented as winners. See the register →
Your code stays with you if you want it to: a protected strategy is validated from its trade list alone.
| Path 1 — Open code | Path 2 — Protected strategy | |
|---|---|---|
| What you provide | Your strategy's rules — a Pine or MQL script, or the logic written out | A trade list with timestamps — your TradingView "List of trades" export, or the equivalent from MT5 or your platform (CSV or Excel) |
| What we do | Port, lookahead detector, full symbol × timeframe grid | Each entry and exit replayed on the broker's M1 bars, at the minute the price is touched |
| Time offset | Not applicable | Auto-detected and aligned to the broker's time |
| Integrity check | Lookahead detector on every signal; port fidelity vs TradingView stated in the report | Alert if replayed gross differs from TradingView by more than 2% on the same feed |
| Possible verdict | ROBUST, REJECTED, INCONCLUSIVE | The same, on the exports you provide — the more coverage you send, the stronger the verdict |
| What stays invisible | — | Repaint and rule detail, since we only see the trades |
On a demo account, a strategy export was replayed to the cent: −81.50% on TradingView, −81.50% gross at the broker — then −91.18% once real costs were deducted.
A verdict on page one — then a full, multi-page document with every chart, table and figure behind it. Each report contains:
We took the TradingView strategies published under open licence and ran them through Coldshim, one by one, without changing their settings. Across 24,657 backtests, 4 cleared the grid. Replayed on dates they had never seen, all failed.
Popularity is not proof. A strategy loved by thousands of traders may have no edge — we checked, on 170 of them. See the full register →
We state the boundary of every verdict, in each report.
Slippage beyond the real spread isn't modelled — so live results can only be worse than what we measure. ROBUST means "the edge survived every test," not "you will win." The past is not a promise.
Not modelled beyond the real spread; the verdict is therefore an upper bound.
Invisible, since we only see the exported trades.
Not yet measured automatically against TradingView; a residual gap is possible and is stated in the report.
ROBUST means the edge survived every test — not that you'll make money.
A Coldshim verdict is reproducible: same data, same scripts, same fingerprints, same result.
| You are | Your question | What Coldshim brings |
|---|---|---|
| Prop firm, copy-trading platform | Does this strategy deserve allocated capital? | ✓An independent entry filter, before allocation, applied to every candidate by the same rule |
| Broker | Are our clients losing to scripts that don't survive our costs? | ✓A service to offer your clients, on your own server data |
| Allocator, family office | Does the presented track record hold after costs? | ✓Independent replay of the trade list, without access to the code |
| Honest script vendor | How do I prove my tool isn't a flattering backtest? | ✓A verifiable report to attach to your product page |
Your trading data is processed in our EU perimeter and erased after the report — we keep only the SHA-256 fingerprints, so a verdict stays verifiable. Prefer not to send files? We can pull the history over a read-only MT5 connection.
Being clear about this saves us both a wasted meeting.
A validation is one strategy, one parameter set — tested across a representative universe for your instrument, not a token sample, and delivered with a certified report.
Every plan runs the complete method — the tiers differ by volume and turnaround, never by depth.
One TradingView export — 1 symbol, 1 timeframe — replayed on your broker's M1 bars: TradingView result, broker gross, net after real costs. A one-page verdict, a 30-minute call to read it together. Under NDA, within 5 days. One per client.
A "no" costs exactly what a "yes" costs. We're never paid more for a robust result than a rejected one — anything else would pay us to tell you what you want to hear. And we don't optimise your parameters until they pass: a validator that tunes is no longer independent.
For an indicator or a coded strategy (Path 1), no: only the data of the broker connected to MT5 counts. For a trade list (Path 2), the same broker gives an exact integrity check; a different broker on the same instrument still works, and the replay refuses beyond a tight price-deviation tolerance.
No. Path 2 works from the trade list alone.
Any broker offering MetaTrader 5, and MT5 only. Costs are measured at the one you actually use.
The representative universe for your instrument — the markets your strategy is meant for, each across the timeframe ladder that fits it. Indices today; forex, metals and commodities on request. We don't test M1 by default: at one-minute bars the spread and commission usually dwarf the move, so the result is mostly cost noise — but if your strategy is genuinely built for M1, we include it.
MT5 is what lets us measure your own broker's real costs. If you're not on MT5, we can still test the strategy's edge on a representative MT5 broker's data for the same instrument — the overfitting and robustness findings carry over — but we can't certify your exact broker's costs. In short: the edge, yes; your-broker cost precision, only on MT5.
Within 2 business days for a validation (1 business day on Catalogue); the free check within 5 days.
Processed on our own servers in the EU, deleted 30 days after the report is delivered, never shared. Only the SHA-256 fingerprints are kept, to verify a report.
We test your strategy exactly as you define it — re-entries, SL and TP included if they're part of your rules. What we don't do is search for better parameters: trying dozens of SL/TP percentages and keeping the best is precisely what manufactures an overfitted backtest. We validate one configuration; we don't tune it.
No, and that's deliberate: a validator that optimises is no longer independent.
Test enough strategies — or one strategy on enough markets and timeframes — and some will look profitable by pure chance. We measure that directly: every result is put against hundreds of random draws with the same shape, checked on separate periods, and confirmed on neighbouring timeframes — with a final replay on dates no earlier step has seen. An edge that shows up in only one place, or that random data reproduces, is luck. We say so.
No — and that's the point. Hunting through markets and timeframes for the one combination where a strategy happens to look profitable is exactly the overfitting we test against: search long enough and you'll always find a lucky cell. We test whether the edge you claim holds up on the instrument it's meant for. A verdict is always scoped to what was tested — a "rejected" strategy didn't show a robust edge on the markets and timeframes listed in its report; it is not a claim that no edge could exist anywhere.
For the same reason you re-run a red team: the ground shifts. Market regimes, volatility and broker spreads change; an edge that was robust a year ago may be gone.
We'll replay it at your broker, with its real costs, against 200 placebos — and give you a one-page verdict. Free, under NDA.
Request a free check