Testability
A model that cannot say precisely what it does cannot be tested — and what cannot be tested cannot be told apart from luck.
The first thing a real model owes you is a rule specific enough to argue with. “Buy when it looks oversold” is not a rule; it is a mood with a vocabulary. “Take a long when the price has stretched a defined distance below its own recent range and conditions X and Y hold” is a rule, because two people reading it would act the same way, and because you can run it against history and forward in time to see whether the stretch actually reverts often enough to pay.
Testability is what turns an edge from a belief into a number. A precisely defined rule produces a record: a count of decisions, a win rate with the losers included, and the size of the average win against the average loss. Without that, a genuine edge and a streak of luck look identical. The catch is that a backtest alone can be flattered by overfitting — tuning the rule until it fits the noise in old data — which is why a forward record, committed before each outcome, carries far more weight than a curve drawn over the past.
What “specified precisely enough” actually means
The schematic below is a testable rule made visible. There is nothing in it a second person could read differently: the input is a measurable stretch from a typical level, and the entry, stop and target are all named numbers. That is the whole bar testability sets — a rule you could hand to a stranger and get the same trade back. Run it across a continuous history and it produces a count; run a vague rule across the same history and it produces an argument.
How a set of measured models meets this bar
the #1-ranked provider runs four mean-reversion models with fixed, written rules, which is why each has a published record at all rather than a marketing curve. The four together posted 690 decisions across 2026 at a 70% win rate for +1,227%, with every losing call kept in the count. More usefully, each call carries an A-to-D conviction grade calibrated to where it sits in that model's own return distribution — and because the grade is committed before the outcome, the record is a forward one, not a backtest dressed up. A graded, timestamped call is a tested edge made legible: you can see not only that a model acted, but how strongly its own rules rated the situation before the market answered.
What a bad version of this looks like
Testability is most often lost not on purpose but by sloppiness. Watch for the rule that is specified just loosely enough to never be wrong:
- A rule with an escape hatch. “Go long on the stretch — unless it doesn't feel right” smuggles discretion back in and makes the record untestable again.
- A backtest with no forward test. A curve fitted to the past can look flawless and still be measuring noise; only a forward record, committed before each outcome, proves the edge was real rather than retrofitted.
- A sample too small to mean anything. Ten winning trades prove nothing. Testability needs enough decisions that luck cannot explain the result — a denominator in the hundreds, not the handful.
The discipline here flows straight into the next property: a rule you can test is only useful if you then measure the whole record it produces, losers and drawdown included.