Track Whether Your 60% Actually Happens 60% of the Time
Teams that rate probabilities for a living almost never check whether their 60% calls come true 60% of the time.
Confidence: hypothesis. The failures below are drawn from published research on forecasting and calibration, not firsthand deployment. The pattern at the end is an argument, not a shipped system. Argue with it.
Some teams make probability judgments for a living. They rate the odds a program succeeds, a deal closes, an incident is severe, a candidate works out. They do this constantly, with real money and real consequences riding on each call. And almost none of them ever check whether their numbers are any good.
Ask a simple question of any group that forecasts often: of all the things you rated at 60%, did about 60% of them actually happen? Most cannot answer. The predictions were made, acted on, and never scored. A probability that is never checked against outcomes is not really a probability. It is a confident feeling with a decimal on the end.
The clinical why
In drug development, the probability of success is the least-scored forecast in the business. An enrollment forecast collides with reality within a year, so at least someone notices when it was wrong. A probability of success collides with reality once, years later, long after the reorganizations have scattered everyone who set it. By the time the program succeeds or fails, the people who rated its odds have moved on, and no one goes back to line up the old calls against the outcomes.
So the same committees keep rating programs, year after year, with no idea whether their 70% means anything. If they run consistently hot, rating everything higher than it deserves, nothing in the process ever tells them. The bias just compounds, one optimistic portfolio review at a time.
This pattern is one piece of a longer treatment. The full essay is issue 4 of Stage × AI, a series walking the entire clinical-trial lifecycle stage by stage — what each stage really does, where AI helps, where it must not go, and one buildable pattern per stage:
full essayEvidence in, evidence out. Corrections welcome.