Publish the Weights: Make Optimism a Number You Can DiscountPublic
Open access · through Aug 15, 2026
This piece is free to read, for now.
After Aug 15, 2026, it moves behind club sign-in. Join now — free, no card — and keep this piece, plus everything else, after the window closes.
A private calibration adjustment fixes the next forecast. A published one fixes the next submission.

Confidence: hypothesis. The transparency practice below is drawn from published incident-review and forecasting-calibration literature, not firsthand deployment. The internal-publication discipline is an argument I have not shipped at scale. Argue with it.
Every team that tracks forecast accuracy eventually learns the same uncomfortable fact: some of their inputs are systematically optimistic. A site's self-reported enrollment capacity, a sales team's pipeline estimate, a vendor's delivery date — the same source, missing in the same direction, forecast after forecast. The rare team that notices this quietly discounts that source in the next model. The much rarer team writes the discount down where the person who submitted the optimistic number can see it.
That second move — publishing the weight, not just applying it — is the difference between a calibration practice that improves the next forecast and one that also improves the next submission.
Here's why the private version fails quietly. A planning team notices, over several cycles, that one vendor's delivery estimates run about 30% optimistic. They start applying a 0.7x haircut in their internal model. The forecast gets better. The vendor's next estimate does not — because the vendor never saw the correction, has no reason to believe their estimating process needs fixing, and will submit the same kind of number next quarter. The team has built a permanent patch around a problem instead of fixing the problem, and the patch has to be maintained forever, quietly, by whoever remembers it exists.
The clinical why
A feasibility deck for a clinical trial runs on inputs from parties who all benefit from a feasible-looking answer: the site wants the study, the CRO wants the award, the internal team wants the program to advance. None of them are lying. The optimism is structural, not personal — which is exactly why quietly correcting for it in a spreadsheet doesn't fix anything. The site never learns their number was discounted, so they have no reason to submit a more honest one next time. The correction lives and dies inside one team's private model.
Publishing the calibration weight changes the incentive directly: “site self-reports have historically delivered 0.4× their promise; this projection discounts them accordingly” is a sentence a site coordinator can read, argue with, or try to improve on. It can't be ignored the way a silent internal adjustment can. The room can push back on the number — that's a feature, not a bug, because the pushback forces the disagreement into the open where it can actually be resolved, instead of staying encoded invisibly in someone's private discount factor.
The generalizable move
This is the same discipline behind blameless postmortem culture in software engineering, formalized at organizations like Google's Site Reliability Engineering practice: an incident review's value comes not from privately fixing the one system that broke, but from writing the finding down and distributing it to “the widest possible audience that would benefit from the knowledge” — often through a standing internal newsletter that shares well-written postmortems across the whole organization, not just the team that lived through the incident.¹ The same logic applies to a forecast's track record: a calibration weight kept private only fixes the next forecast one team builds. A calibration weight published where its source can see it changes behavior at the source. Google's own postmortem template makes this structural, not just cultural: a standard postmortem document carries a summary, a timeline, a root-cause analysis, an impact assessment, and named corrective actions with owners and due dates — every one of those fields is written to be read by people outside the team that lived through the incident, not filed away for the author's own reference. A calibration report for a forecast source can follow the identical shape: what the source predicted, what actually happened, the specific factor by which they diverged, and what the source (or the process consuming their number) should change next cycle.
Three pieces make this work in practice:
1. The weight is a number, not a vibe. “We've learned to discount site estimates a bit” is not actionable — it can't be argued with because it isn't precise enough to disagree with. “0.4× historically” is a specific, falsifiable claim a site can contest with its own data, or try to beat next cycle.
2. The weight is visible to the party it describes. A discount applied silently inside a model protects no one and improves nothing at the source; a discount printed on the same page the source's own submission appears on creates a reason to submit a better number next time. Concretely, this can be as simple as a line item the source sees before they submit their next estimate — “your last four projections averaged 1.6x actuals” — rather than a number buried three tabs deep in a planning team's internal model that the source never has occasion to open.
3. The weight updates on the same schedule the forecast does. A stale calibration factor from three years ago is worse than no calibration factor — it's precision borrowed from a context that may no longer apply, and a source that fixed the underlying problem two cycles ago deserves to have that improvement reflected, not to keep paying for a mistake they already corrected. The published number needs the same living-document treatment as the forecast it corrects, not a one-time adjustment set once and forgotten.
Where AI changes the economics: maintaining and republishing an accurate, current weight per source across dozens of sites or vendors is exactly the kind of continuous bookkeeping a model can sustain that a busy planning team usually can't keep current by hand — the capability gain is in currency and coverage, not in the underlying idea, which the forecasting and incident-review literature already established.
What to measure first
Look at one recurring forecast input your team currently corrects for informally — “we usually pad the vendor's timeline,” “we discount that team's estimate a bit.” Turn the informal correction into a number, and check whether the party being discounted has ever actually seen that number. If they haven't, you have a private adjustment, not a published one — and the source has no reason to improve.
Count, across your last several forecast cycles, how many informal corrections like this exist versus how many have ever been written down and shared with the source they describe. That ratio — private adjustments versus published ones — is a more honest measure of whether your organization actually has a calibration practice, or just a collection of tribal knowledge that lives in one planner's head and dies with their next role change.
Where this breaks in practice: publishing a weight can read as an accusation if it's framed as blame rather than as a shared, correctable fact about a noisy process — the postmortem-culture parallel holds only if the publication stays blameless, focused on the pattern rather than the person. And a published weight that never gets revisited becomes its own kind of stale authority, cited long after the underlying process has actually improved.
References
- Google, “Postmortem Culture: Learning from Failure”, Site Reliability Engineering (2016) — blameless postmortems distributed to the widest audience that benefits from the learning, including a standing internal newsletter sharing well-written reviews organization-wide.
This pattern is one piece of a longer treatment. The full essay is issue 3 of Stage × AI, a series walking the entire clinical-trial lifecycle stage by stage — what each stage really does, where AI helps, where it must not go, and one buildable pattern per stage:
full essayEvidence in, evidence out. Corrections welcome.