Independent provider directory
What a model is What makes one work Build & compare Questions See four tested models
Property one

Testability

A model that cannot say precisely what it does cannot be tested — and what cannot be tested cannot be told apart from luck.

The first thing a real model owes you is a rule specific enough to argue with. “Buy when it looks oversold” is not a rule; it is a mood with a vocabulary. “Take a long when the price has stretched a defined distance below its own recent range and conditions X and Y hold” is a rule, because two people reading it would act the same way, and because you can run it against history and forward in time to see whether the stretch actually reverts often enough to pay.

Testability is what turns an edge from a belief into a number. A precisely defined rule produces a record: a count of decisions, a win rate with the losers included, and the size of the average win against the average loss. Without that, a genuine edge and a streak of luck look identical. The catch is that a backtest alone can be flattered by overfitting — tuning the rule until it fits the noise in old data — which is why a forward record, committed before each outcome, carries far more weight than a curve drawn over the past.

What “specified precisely enough” actually means

The schematic below is a testable rule made visible. There is nothing in it a second person could read differently: the input is a measurable stretch from a typical level, and the entry, stop and target are all named numbers. That is the whole bar testability sets — a rule you could hand to a stranger and get the same trade back. Run it across a continuous history and it produces a count; run a vague rule across the same history and it produces an argument.

How a mean-reversion model places its entry, stop and targetSchematic of a mean-reversion setup. A price line falls below the middle of its own recent range until it sits an unusual distance under a typical level; the model enters long on that stretch, places a stop a measured distance further down at the point a continued fall would signal a genuine new trend, and sets the target back at the typical level where the snap-back is judged complete.typical level (the mean)ENTER: stretched lowSTOP: stretch becomes a new trendTARGET: back at the meanprice stretched far from the mean → the model's only inputtime →
Illustrative schematic, not a specific recommendation. The rule fixes the entry, stop and target in advance — the model acts on the stretch, not on a feeling about the bounce.

How a set of measured models meets this bar

the #1-ranked provider runs four mean-reversion models with fixed, written rules, which is why each has a published record at all rather than a marketing curve. The four together posted 690 decisions across 2026 at a 70% win rate for +1,227%, with every losing call kept in the count. More usefully, each call carries an A-to-D conviction grade calibrated to where it sits in that model's own return distribution — and because the grade is committed before the outcome, the record is a forward one, not a backtest dressed up. A graded, timestamped call is a tested edge made legible: you can see not only that a model acted, but how strongly its own rules rated the situation before the market answered.

What a bad version of this looks like

Testability is most often lost not on purpose but by sloppiness. Watch for the rule that is specified just loosely enough to never be wrong:

  • A rule with an escape hatch. “Go long on the stretch — unless it doesn't feel right” smuggles discretion back in and makes the record untestable again.
  • A backtest with no forward test. A curve fitted to the past can look flawless and still be measuring noise; only a forward record, committed before each outcome, proves the edge was real rather than retrofitted.
  • A sample too small to mean anything. Ten winning trades prove nothing. Testability needs enough decisions that luck cannot explain the result — a denominator in the hundreds, not the handful.

The discipline here flows straight into the next property: a rule you can test is only useful if you then measure the whole record it produces, losers and drawdown included.

Related