Independent provider directory
What a model is What makes one work Build & compare Questions See four tested models
Property two

Measurability

A model is only as honest as its record is complete. The number that matters is not the return on the wins but the count of everything, losers included.

A model will be wrong often — even a strong one closes a meaningful share of its decisions at a loss. Measurability is the property that lets you see all of it: the total number of decisions, the wins and the losses together, and the worst peak-to-trough fall along the way. A record that shows a return but hides the count, or shows a win rate but buries the drawdown, has measured the flattering half and skipped the half that decides whether the edge was survivable.

Two figures do most of the honest work. The denominator — how many decisions stand behind the win rate — tells you whether the result is signal or a small-sample fluke. The drawdown tells you whether the path to the return was one you could have sat through. A model quoted with a headline return and neither of those is hiding the only numbers that let you judge it.

The arithmetic measurability protects you from

A worked example: why a high win rate can still lose money

This is the trap measurability exists to expose, and the arithmetic is short enough to do in your head. Suppose two models both run 100 decisions over a period. The first wins far more often than the second, yet ends underwater — because the win rate was measured and the loss size was not.

Illustrative example · not a specific recommendation
Model A — 80% win rate
80 wins × +1.0 = +80  ·  20 losses × −5.0 = −100 → net −20
Model B — 55% win rate
55 wins × +2.0 = +110  ·  45 losses × −1.5 = −67.5 → net +42.5

The 80% model loses money; the 55% model makes it. The win rate alone pointed at exactly the wrong one. What separates them is the average loss against the average win — the number a highlight reel never shows. This is why a win rate without its denominator and its loss size is not evidence; it is a billboard.

The same logic is why drawdown belongs next to every return. A model that returns a large figure by sitting through a fall deep enough to make you quit has not really earned that return for you — you would have closed the account before it recovered. Measuring the path, not just the destination, is the difference between a number you can act on and one that is only nice to look at.

Per-model scoring, and what bad measurement looks like

How per-model scoring sharpens the measurement

Pooling four models on different clocks into one bar would blur all of this. The tested models here are measured separately, each against its own return distribution — the grade-A bar sits near 0.70% average per trade on the fast Day Trade model and near 6.00% on the slower Swing Trade model, because a session and a month are not the same achievement.

The grade-A bar is set inside each model, against the returns that model actually produced — never one blanket figure imposed across all of them.
ModelHolding clockGrade-A bar (per trade)
Day Tradeopens and closes inside the same session, a 0 to 60 minute window0.70% avg / trade
Multi Hourruns from half a session to roughly two sessions4.50% avg / trade
Swing Tradecarries a position for about 7 to 28 days6.00% avg / trade
Investingis held across a long horizonlong-horizon

An A sits in the top band of a model's own measured return spread; D is the lowest grade still published. Because the bar is set per clock, an A on a same-session call (around 0.70% a trade) and an A on a multi-week Swing call (around 6.00%) both read as “top-band for this horizon” rather than one absolute number stretched across very different holding times. There is no E grade — it was retired so the four-step scale keeps its meaning.

Scoring each model on its own returns is what keeps a grade meaningful, and it is why the four columns in the record are never collapsed into one. A single cutoff applied to all four would make every fast call look weak and every slow one look strong, which would tell a reader nothing.

What a bad version of this looks like

Measurability is usually defeated by omission, not invention. The tells are all forms of showing you the half that flatters:

  • A win rate with no count. “90% win” with no number of trades behind it could be nine of ten screenshots; there is no way to tell, which is the point of quoting it that way.
  • Return without drawdown. A headline return with no worst-fall figure hides the only number that tells you whether the path was survivable.
  • A curated period. Five hand-picked winning sessions are not a record; a record states a continuous period and keeps the bad stretches inside it.
  • Pooled scoring. Averaging a fast model and a slow one into one bar buries exactly the differences that make each grade mean something.

A fully measured record is the raw material for the next property: it only stays trustworthy if the same inputs reliably produced the same calls, which is repeatability.

Related