Why Averaging Your Forecasts Can Hide the Risk That Matters
Combining several forecasts usually beats trusting any one of them. But when the forecasts disagree about risk, the average can make the one warning that mattered disappear.


There is good evidence that combining forecasts works. Averages of several reasonable models or analysts tend to be more accurate, on average, than most of the individual forecasts that went into them. Errors in different directions cancel out. That is why ensembles, consensus estimates and committee views are everywhere in finance, insurance, engineering and planning.
The catch is in the words “on average”. When the decision is about risk — how much capital to hold, how much contingency to build in, how much protection to buy — the average forecast is often not the number you need.
Combining several forecasts usually beats trusting any one of them. But when the forecasts disagree about risk, the average can make the one warning that mattered disappear.
§ 02When the Average Is the Wrong Summary
Imagine three sources assessing the same risk. Two say it is small. One says it is large, perhaps because it has information the others lack, perhaps because it is simply wrong. The average lands somewhere comfortably small. If the dissenting source was right, the decision built on that average is badly under-protected, and nothing in the averaged number tells you that a warning was ever raised.
This is not a rare edge case. Catastrophe models for the same peril routinely disagree. Risk assessments for the same infrastructure project can differ by large margins. Analysts covering the same stock regularly take opposite views. Disagreement is normal; what matters is what the decision process does with it.
§ 03Five Agreeing Votes May Be One Vote
A second problem is quieter. Sources that look independent often aren't: they use the same data vendor, the same base model, or simply read each other's reports. Averaging them gives the shared view far more weight than it deserves, and makes the consensus look more certain than it is.
§ 04“Low Confidence” Is Not a Decision
Many systems respond to disagreement with a confidence score and then output a number anyway. A low-confidence number still gets used. For decisions that carry real downside, it is often better to say plainly that the evidence does not support a decision yet, and why, than to publish a figure that nobody should rely on.
§ 05What a Better Process Keeps
Three habits help. Keep every source's view on record rather than only the blended result, so that a dissent can be found later. Size risk decisions against the range of credible views, not only the middle of it. And allow the process to decline to decide when the evidence is too thin or too dependent on a single source.
These are the principles behind ARGUS Council, our approach to certifying decisions when models, analysts or AI agents disagree: every view retained, a decision that accounts for all of them, and a formal refusal when the evidence isn't there.


