Method note

How we decide whether a signal is real — and what happened when we checked our own headline finding and found it was an artifact of the way we had defined a win.

Educational and entertainment statistics only. Not investment advice. This is a research model showing its work, for learning and paper trading. It does not know you or your situation, and nothing here is a recommendation to buy or sell.

What the model does

It watches a short list of stocks on a fixed schedule and asks one question at each bar: are the trend measures lined up in the same direction, and is that alignment getting stronger? When they are, it records a setup and scores it from 0 to 1. Then it looks at a longer timeframe. If the bigger trend points the other way, the setup is vetoed rather than taken.

That is the whole idea. It is a trend-following model, not a prediction engine, and the interesting part is not the setups it takes — it is the ones it refuses.

How we decide whether it works

Five commitments, all of which predate the finding described below.

Pre-registration
Hypotheses are written down, with their falsifiers, before the data that would test them exists. The register is append-only: entries are never edited, only amended, and amendments are dated and say what they change.
Append-only forward evidence
Every decision the live model makes is logged as it happens, before the outcome is known. That record cannot be edited after the fact, which is what makes it evidence rather than a recollection.
Honest uncertainty
Setups on the same stock are not independent events, and overlapping windows share outcomes. Our intervals resample whole stocks rather than individual setups, which widens them considerably. Where we have too few distinct stocks for that to be reliable, we say so and decline to publish a number.
Stated power limits
Every result carries the smallest effect that sample could have detected. A finding of “no difference” from a sample that could only ever have seen a large one is not evidence of anything, and we label it that way.
A threshold set in advance
What counts as publishable is fixed before the numbers are looked at. Results that miss it are reported as missing it, rather than reframed.

The finding we withdrew

For several months our most striking result was that the model’s medium-confidence setups outperformed its high-confidence ones. It was counterintuitive, it was memorable, and we were preparing to build our public case around it.

It was an artifact of our own measurement, and here is the mechanism. A setup counts as a win if the price travels a set distance before a deadline. That distance is not fixed — it scales with how wide the trend is. And the score also rises with how wide the trend is. So the higher the model scored a setup, the further that trade had to travel before we would call it a success. We were handicapping our own best calls, then recording that they won less often.

In our data the required move rises from about 3.8% in the middle band to about 8.9% in the top band. The observed hit rate falls across those same bands — 62.9% to 46.8%. Those two facts are the same fact.

Correct for the distance and the effect does not reverse. It disappears. What is left is indistinguishable from no relationship at all.

What we can and cannot say now

We have not established that a higher score means a better trade. Nor the opposite — the corrected relationship is flat, not inverted, and anyone quoting our old numbers as evidence that low scores are better would be making our original mistake with the sign changed.

Against our own publication threshold, every timeframe we measure currently fails: the strongest one because its interval covers zero, the others because six or nine distinct stocks is not enough to support a claim about a whole market. We are publishing that failure rather than the subset of numbers that would have looked better.

This is also why the signal board is ordered alphabetically instead of best-first. A ranking asserts something we cannot currently support.

How to check us

What would change our mind

The measurements that would move us, stated before we have them: a distance-corrected relationship between score and outcome that holds across enough distinct stocks to be worth reporting; evidence that the setups the model refuses would in fact have lost; and results that survive on stocks we did not choose in advance. Our current evidence comes from a curated research library, which is a reasonable place to develop a model and a poor place to prove one.

Not advice

Everything here is educational and entertainment statistics. It is not investment advice, not a recommendation, and not a forecast. The model does not know you or your circumstances. Nothing on this site is a record of real trading; the paper terminal fills with play money, at the same prices it shows you.

New here? Start with the 2-minute guide