Forecast accuracy: measuring error so it changes something
What this answers
How should forecast error be measured so that it identifies a cause somebody can act on?
Forecast accuracy is measured almost everywhere and used almost nowhere. The usual reasons are that the measure is taken at a level or a horizon that no decision depends on, that bias is hidden inside an averaged error figure, and that the review produces a score rather than a change to how the next forecast is made. Fixing those three things turns a reporting ritual into a feedback loop.
Written for: demand planners and planning managers, commercial teams contributing forecast input, supply chain analysts designing measurement.
Bias and dispersion are different diseases
Bias is a persistent tendency to forecast above or below actual demand; dispersion is how far individual periods scatter around the forecast regardless of direction. An absolute error measure treats an over-forecast and an under-forecast identically and therefore conceals bias entirely. Reporting a signed measure alongside an absolute one is essential, because bias is usually caused by an incentive or an assumption and can be corrected, whereas dispersion is largely inherent to the demand itself.
Measure at the level and lag the decision uses
Accuracy improves as you aggregate and as you shorten the horizon, so a headline figure quoted at total-company level one period ahead can look excellent while the item-location forecast that drives ordering is unusable. The measurement must be taken at the granularity of the decision it informs and at the lag at which that decision is committed — typically the point where the replenishment or production commitment is made, not the period immediately before delivery.
Weight by what matters
Averaging percentage error across a catalogue gives equal weight to a trivial slow-moving item and a major line, and percentage measures behave badly when actual demand is very small or zero. Weighting by volume or by value, or using a measure that divides total error by total demand, produces a figure that moves when something economically significant has gone wrong. Segmenting the result by item class prevents a stable core from masking a deteriorating tail.
Compare against a naive benchmark
An accuracy figure means little in isolation. The useful comparison is against a simple alternative — repeating last period's actual, or a seasonal repeat of the equivalent period last year. If the forecasting process cannot beat that, the effort invested in it is not returning value, and the honest conclusion is to simplify the method and redirect the attention to items where judgement genuinely adds information.
Close the loop on causes, not scores
The review should ask which adjustments improved the baseline and which degraded it, whether particular sources of input are consistently optimistic, and which items exhibit structural bias. Recording each override at the time it is made, with its reason, is what makes this analysis possible later. Without that record the review can only observe that the forecast was wrong, which nobody disputes and nobody can act on.
Frequently asked questions
- What accuracy level should we be aiming for?
- There is no universal target, because achievable accuracy is set by the demand pattern. A stable, high-volume line and a sporadic specialist item have entirely different ceilings, so the meaningful goal is improvement against your own baseline and against a naive benchmark rather than an external figure.
- Should planners be incentivised on forecast accuracy?
- Cautiously. Targets on accuracy alone encourage forecasting what is easy to hit and discourage flagging uncertainty. Pairing the measure with bias, and reviewing the quality of the assumptions rather than only the outcome, keeps the behaviour honest.
- Why does accuracy get worse after a system implementation?
- Often because measurement changed rather than performance. New systems commonly alter the level, the lag or the formula, so the earlier figures are not comparable. Restating the historic series on the new definition before drawing conclusions avoids a false alarm.
Data limitations
- Logistics figures are operator-supplied inputs, not market data. GeoBusinessIQ holds no freight rates, transit times, capacity, or throughput data and does not estimate them — every result reflects only the figures you enter.
Explore the graph
Related logistics topics
- Demand planning: turning a forecast into a usable number
- Demand variability: classifying how demand actually behaves
- Safety stock: buying availability with working capital
- Supply chain KPIs: a measurement set that survives scrutiny
- Bullwhip effect: why order swings grow upstream
- ABC analysis: directing attention across an uneven catalogue
- Business continuity planning for supply operations
- Capacity planning: sizing the ability to supply
Calculators
Sources
- United Nations Conference on Trade and Development — UNCTAD (accessed )Covers: Trade and development analysis, maritime transport review, and trade facilitation research.Does not cover: Real-time freight rates, company-level data, or operational carrier information.Why it matters: United Nations body producing long-running analysis of maritime transport and trade logistics; used for structural context rather than point figures.Review cadence: as published
Educational and operational information only — not legal, customs, tax, insurance, or financial advice. Requirements vary by jurisdiction, commodity, and contract; confirm with the relevant authority or a qualified adviser before acting.
Last updated: