Skip to content

How to Correct Forecast Bias in Ecommerce

Jannik SemmelhaackCEO & Founder, VOIDS

How to Correct Forecast Bias in Ecommerce Before It Distorts Inventory

What is forecast bias in ecommerce? Definition and impact

Forecast bias in ecommerce is a persistent tendency to forecast demand too high or too low for a SKU, product group, or planning horizon. Correcting forecast bias means measuring that directional pattern, isolating exceptional sales conditions, testing a revised rule on past periods, and using only validated changes in purchasing decisions.

Key Takeaways:
  • Forecast bias is repeated directional error rather than one missed forecast.
  • SKU-level review prevents offsetting errors from hiding behind an accurate category total.
  • Promotions, stockouts, and supply conditions require separate treatment before a baseline changes.
  • Backtesting tests a revised rule against completed comparable periods before it affects purchasing.
  • An owner, change record, and review date keep the correction controlled.

As of 2026, the operational consequence is clear: a forecast correction belongs in the purchase-decision record, not in an untracked spreadsheet override. Preserve the original forecast, retain the sign of the error, identify exceptions, and compare the old and revised rule against completed comparable periods before changing expected demand.

One published 2026 VOIDS customer case illustrates the inventory context in which this discipline matters. HEY HOLY, an eight-figure brand with around 100 SKUs and 1.5 FTE in purchasing, worked toward reducing target inventory coverage from 30 to 17.5 days while maintaining SKU availability above 99%; the case is described by Jannik Semmelhaack on his LinkedIn activity page.

TL;DR

  • Forecast bias is repeated directional error, not one missed weekly forecast.
  • SKU-level bias can be hidden by an apparently accurate category total.
  • Promotions, stockouts, and supply conditions need separate review before a correction.
  • Backtesting tests a proposed correction before it changes a purchase order.
  • A named owner and review cadence keep bias visible after the first fix.

I treat bias as a planning signal, not a verdict on a planner or a model. A positive pattern means the forecast repeatedly exceeds observed demand; a negative pattern means it repeatedly falls short. Both patterns can create operational pressure, because excess units and unavailable units are managed SKU by SKU, not through a reassuring total.

An aggregate forecast can conceal offsetting SKU errors: if a brand forecasts 120 units for SKU A and sells 80, then forecasts 80 for SKU B and sells 120, the combined forecast and combined sales both equal 200. The total looks exact while SKU A accumulates stock and SKU B faces a shortage. This illusion of accuracy is why I start with the level that drives a replenishment decision.

Direction matters alongside absolute error. A forecast that alternates above and below demand may need a different response from one that drifts low for six comparable review periods. Persistent directional drift can remain hidden in average accuracy reporting, so the report should show signed error as well as the size of error. A near-zero bias is a useful diagnostic target, but it does not prove that inventory is economically appropriate.

For a practical definition, compare the forecast available when a decision was made with actual demand from the same period, retain the sign of the difference, and inspect repeated observations. Over-forecasting can contribute to excess inventory and under-forecasting can contribute to stockouts; historical backtesting provides the discipline for testing whether a proposed correction would have behaved better before it governs future orders.

How to measure forecast bias before changing a forecast

I measure ecommerce forecast bias by preserving the original forecast, comparing it with actual demand over comparable periods, and checking whether error repeatedly points in the same direction at the SKU or coherent-product-group level. The planning horizon sets the review window: use periods that match the replenishment or purchase decision being evaluated. Change a forecast only after the proposed adjustment has been backtested on the comparable historical observations available for that horizon.

Start with a frozen forecast snapshot. Once a forecast is overwritten after sales arrive, the team can no longer distinguish the original prediction from a later manual adjustment. Record the forecast, actual demand, planning horizon, relevant availability constraints, and any exception that could affect the comparison.

  1. Set the decision horizon. Review demand in periods that match the replenishment or purchase decision. Do not combine horizons when assessing whether a directional pattern persists.
  2. Calculate signed error consistently. Use the same direction for every comparison so the team can see whether forecasts tend to run high or low; absolute error alone cannot show that direction.
  3. Compare like with like. Assess an SKU where the available observations are useful for the decision. If grouping is necessary, use a coherent product group with similar replenishment logic.
  4. Document exceptions and assess persistence. Keep unusual events, such as a promotion, a launch, or constrained availability, separate from the baseline comparison. Investigate a direction only when it recurs in the comparable observations available for the selected horizon.
  5. Backtest the proposed adjustment. Apply the old and revised rule to completed comparable periods and review what each would have implied for the earlier purchasing decision.

Intermittent or low-volume demand can make a small number of observations look more conclusive than they are. A total can also appear accurate while SKU-level over- and under-forecasting offset each other, as the discussion of intermittent demand and offsetting errors explains. I would treat a directional signal in these cases as a reason to inspect the available comparable history, not as proof of a correction.

Give the review an operational owner: purchasing can own the action, while demand planning or finance can challenge the assumptions. Bias measurement should inform planning decisions, rather than remain a dashboard metric. Backtesting keeps the question practical: would the revised approach have provided a more credible input for the earlier order? That is the historical simulation described in VOIDS’ demand-forecasting guidance.

What the review showsHow to interpret itNext action
An unusual period with a documented exceptionThe period is not a clean comparison for the selected planning horizon.Keep the exception visible and do not use it alone to revise the baseline.
Directional error recurs in the comparable observations availableThe forecast may be persistently high or low for that SKU or coherent group.Investigate the baseline input and backtest a proposed adjustment.
Category totals look accurate while item results differOffsetting SKU errors may be masking directional error.Review the SKU or coherent-group results rather than applying a category-wide change.
Sales occurred while stock was unavailable or constrainedObserved sales may not represent demand without further assessment.Keep the period separate while the team determines how to treat the demand signal.

Operational workflow: which signals should you inspect before correcting bias?

An ecommerce team should inspect baseline sales velocity, promotional periods, stock availability, lead-time conditions, product lifecycle status, and the forecast horizon before correcting bias. These signals identify whether the compared periods are suitable for a baseline adjustment; they do not, by themselves, prove why a forecast missed demand.

I separate the diagnostic conversation from the correction decision. A team may observe that a SKU sold less during an out-of-stock period, for example. That observation means recorded sales were constrained; it does not establish the demand that would have occurred with full availability. Treating constrained sales as ordinary demand can pull a later baseline down for the wrong reason.

Start with baseline velocity: sales from periods that resemble the period being planned. Then tag demand-shaping conditions. Promotions deserve their own field because campaign demand should not quietly become the everyday baseline. Lead-time and delivery changes belong in the same review because a forecast may be operationally sound yet unusable if the inbound date or order cutoff changes. Separating baseline velocity, promotions, lead-time variance, and stockout risk on SKU history is a sensible planning structure described here.

Next, test the granularity. A colour or size variant with thin sales history may be too noisy for independent correction, while a coherent group may provide a more stable view. The trade-off is clear: grouping creates a usable signal but can hide variant-specific demand; SKU-level review preserves specificity but can overreact to limited data. Use the level at which the purchasing decision can actually be changed.

Spreadsheets can support this work at a small scale, but manual updates create a control problem when forecasts, exceptions, and purchase orders live in separate versions. Teams should define one current planning view, a change log, and a routine for reviewing seasonality and demand shifts. Those operational practices are consistent with the ecommerce forecasting considerations outlined by Optiply. They are workflow choices, not proof that any given signal caused the bias.

Finally, revisit the SKU-level report after each annotation. A category can still look balanced while individual product lines pull in opposite directions, which is the exact masking problem described in the SKU-level bias discussion. I want the planner to state what was observed, what was excluded from the baseline, and what change will be backtested. That creates an auditable handoff to purchasing.

How do you turn a bias correction into an operating routine?

A validated forecast-bias correction can inform a purchasing decision by changing the expected demand view for the relevant review horizon. I then consider the available stock, confirmed inbound supply, timing, quantity, and the decision that needs revisiting next.

Start with the observed pattern rather than a single aggregate accuracy result. Forecast bias can reveal that forecasts repeatedly drift too high or too low even when average accuracy appears reasonable, and it should prompt a decision review, as this planning guidance explains.

For example, a planner finds repeated under-forecasting in one coherent product group. They separate periods with documented exceptions from the recurring signal, then backtest a revised forecast input against completed comparable periods. Next, they review available stock and confirmed inbound supply for the purchasing horizon. The planner selects a split order: it limits commitment to the uncertain portion while retaining a nearer replenishment option. They set the next review date to check whether the revised input still fits the completed periods and the current supply position.

The corrected forecast informs the order decision by updating expected demand; it does not settle quantity or timing on its own. A team may bring replenishment forward, raise quantity, or split an order when under-forecasting persists. When over-forecasting persists, it may reduce quantity, delay an order, or make no order. Each choice leaves a trade-off: a split order can reduce commitment but add operational complexity, while a delayed order reduces inventory exposure but leaves less time to respond if demand rises.

Backtesting tests whether a revised input would have improved forecast quality in completed comparable periods. Backtesting historical periods provides a way to challenge a proposed change before applying it to the purchasing view. Accounting for demand shifts and seasonality can also keep that view current, as discussed by Optiply.

For a next step on workflow setup, review Setup.

"No overstocks. No stockouts. Cash unlocked."

— Jannik Semmelhaack, Founder & CEO, VOIDS – AI-driven Demand Planning (2026-08-20) · Quelle

Examples from a first-party search data study

This first-party search data study records 14,247 query/page rows in the organisation’s Google Search Console performance data between 2026-06-05 and 2026-09-02. It describes the study scope and counting method only; it does not measure ecommerce forecast bias, establish causation, or support claims beyond that defined window.

The unit of observation is a query/page row, not a customer, order, SKU, forecast, or conversion. That boundary matters. Search-performance rows can document what was counted in this publication dataset, but they cannot validate a demand-planning correction. I keep this evidence separate from the operational records used in the workflow above.

Study fieldRecorded valueInterpretation boundary
DatasetGoogle-Search-Console-Leistungsdaten dieser Organisation (gsc_performance)Organisation first-party search-performance dataset.
Window2026-06-05/2026-09-02No statement extends outside this period.
MethodAusgezählt wurden alle 14247 Query-/Seiten-Zeilen mit Datum zwischen 2026-06-05 und 2026-09-02. Keine Gewichtung, keine Hochrechnung über das Fenster hinaus.Count only; no weighting or extrapolation.
Sample14247Query/page rows, not ecommerce demand observations.

The practical example is a separation-of-evidence rule: use search data to describe search data, and use frozen forecasts, availability records, actual demand, and purchase-order records to evaluate forecast bias. Combining unlike datasets without a defined mechanism can make a report look richer while making the decision less defensible.

Risks and limits of forecast bias correction FAQ

Forecast bias correction has limits: a near-zero bias can coexist with large individual errors, constrained sales can obscure demand, sparse SKU history can produce unstable signals, and a backtest cannot guarantee future purchasing outcomes. Teams should use bias as one controlled input alongside availability, supply constraints, and documented commercial exceptions.

A hard exclusion condition is missing or altered forecast history. If the team cannot recover the forecast that existed before actual sales, it cannot measure bias cleanly for that period. Reconstructing the number from memory or a current spreadsheet version is unsuitable evidence for changing a purchasing rule. Start capturing snapshots, then wait for enough comparable observations.

Intermittent demand is another boundary. Sporadic sales make errors harder to identify and correct, while an accurate-looking aggregate can still conceal SKU outcomes. The limits described for intermittent demand and aggregate accuracy support a cautious review cadence. Prefer a grouped view or a conservative manual exception where individual history is too thin, and document what is gained and lost by aggregation.

Bias also does not answer every inventory question. It cannot by itself determine the right service level, safety stock, supplier reliability, cash allocation, or promotional plan. A forecast can be directionally balanced and still produce an unsuitable order because its uncertainty, lead time, or commercial constraint was ignored. A near-zero bias is a diagnostic goal, not a certificate of cost-effective inventory management.

My final limit is governance. A correction without an owner, an effective date, and a review point becomes an undocumented override. Persistent drift deserves investigation because it can remain unobvious in average reporting, as this explanation of directional bias notes. The appropriate response is controlled learning: preserve the prior rule, state the hypothesis, backtest the alternative, and revisit the resulting purchase decisions.

How much history is needed to correct forecast bias?

Use enough comparable periods to distinguish a repeated directional pattern from a one-off exception. The needed amount varies with sales frequency, lifecycle stage, and planning horizon; low-volume SKUs require more caution than stable, high-velocity products.

Should a team correct bias at category or SKU level?

Correct at the level that drives the purchase decision. Start with SKU-level analysis where history is sufficient, then use a coherent product group when individual observations are too sparse. Do not let a balanced category total override opposing SKU signals.

Can a stockout period be used as actual demand?

Recorded sales during a stockout may be constrained by availability. Keep that period visible and assess it separately before using it as a baseline-demand observation.

What should be recorded after a forecast correction?

Record the original and revised forecast, the affected horizon, exceptions considered, backtest result, inventory position, purchasing action, owner, approver, and next review date.

Can forecast bias alone determine an order quantity?

No. A corrected forecast updates the expected-demand view, but the purchasing decision also requires available stock, confirmed inbound supply, timing, quantity, and the next review point.

HYROX is scaling merchandising to 9 figures with VOIDS: online and offline, across Europe, the US, and the rest of the world. 2,000 SKUs, specialized event demand forecasting, transfer logic, and global reorder quantities for Puma and other suppliers.

Jochen MollerCCO, HYROX

Read customer story
HYROX customer story
HYROXVOIDS customer