We tested our own confidence score against 465 trades. It does not separate winners from losers. We are publishing that.
Ours does not. Across 465 trades, the score ranked winners above losers almost exactly as often as a coin flip would, at p = 0.944.
That is a plain statement that one of our own headline numbers does not do what a reader would assume it does.
Because the alternative is worse. A confidence score that does not discriminate is not merely useless — it is actively misleading, because it invites you to size positions by it.
Every confidence band and every threshold we tested came out net negative, and the gradient was not even monotone: higher confidence did not reliably mean better outcomes. No threshold can rescue a score that carries no information. That is arithmetic, not opinion, and it closed off a whole family of “just raise the bar” fixes that would otherwise have looked sensible.
We stopped treating the score as a ranking of quality. Where a threshold still exists, it exists as an explicit brake on volume with a measured justification — not as a claim that the trades above it are better.
If you are evaluating any research product, the useful question is not what its best number looks like. It is whether it has ever published a number that embarrassed it.
This finding is marked fallen in our register and shown on purpose. The measurement that kills an idea is the same measurement that would have confirmed it; publishing only one direction would make both worthless.
All numbers on this page were published in Apex’s own record at the time; they describe the past under stated conditions and promise nothing about the future. All articles · The Apex day