Retail demand forecasting,
scored as an inventory decision

A forecasting pipeline on 3,000,888 rows of real grocery sales, built around a question the competition metric does not answer: does this forecast lead to a better purchase order? Measured here, the two questions pick different models — the best RMSLE leaves a stockout on 41% of store-days.

Data: Corporación Favorita (Ecuador), 54 stores × 33 product families, 2013-01-01 to 2017-08-15, downloaded 2026-10-02 via the Kaggle API. Competition rules apply; the raw data is not redistributed. Granular retail demand is not public anywhere, so the choice was real demand from a comparable South American grocery chain or invented demand with a local label — this project takes the real data and says where it is from.

1.A quarter of the “zero demand” is not demand

The panel is 31.3% zeros, which reads as severe intermittency and sends you looking for Croston or a zero-inflated model. That number is wrong, and the reason is not in the schema.

Decomposition of the panel's zero rows into genuine zeros and stores that had not opened
Eight of the 54 stores open after the panel begins — the last on 2017-04-20, four months before the data ends. The panel is padded, not truncated: until a store opens, all 33 of its families report zero every day.

222,057 rows — 7.4% of the panel and 23.6% of every zero in the dataset — record the absence of a store rather than the absence of demand. The real zero rate is 25.8%, not 31.3%, and the count of severely intermittent series drops from 173 to 135.

A model trained on those rows learns “this store does not sell” about a store that had not opened, and a metric computed over them collects credit for predicting zero where there was nothing to predict. The pipeline flags them rather than dropping them silently, so both versions of each number stay visible.

Histogram of the share of zero-sales days across the 1,782 series
Even after the correction, not one of the 1,782 series sells every day. The distribution is bimodal — which is also why MAPE is absent from this project: it divides by the actual value, zero on a quarter of the observations.

2.The forecast that wins the metric loses the decision

The competition scores RMSLE on a point forecast. But nobody orders the mean. A buyer orders a quantity, and being short (lost sale) does not cost the same as being long (excess stock, markdown, tied-up capital). That is a newsvendor problem, and its answer is not the mean but a quantile set by the cost ratio, Cu / (Cu + Co).

Three models compared on RMSLE, service level and total cost
The same three forecasts, scored three ways. “Best” moves between panels.
ModelRMSLE ↓MASE ↓Service ↑Fill rate ↑Total cost ↓
Seasonal naive0.66661.76260.3%87.8%9.14M
LightGBM (mean)0.41480.81159.0%94.3%4.25M
LightGBM (quantile 80%)0.53131.17784.0%97.3%3.92M

The mean model’s service level is 59.0% — below the seasonal naive’s 60.3%, despite being dramatically more accurate. Ordering the conditional mean leaves you short roughly half the time by construction, however good the mean estimate is. Accuracy improved; the decision did not.

3.And it does not hold at every cost ratio

The critical ratio is an assumption about the business. If the conclusion flips under a plausible alternative, quoting only the favourable column is how a portfolio project becomes misleading.

Cost saving versus the mean forecast across cost ratios, and quantile calibration
Left: below roughly 3:1 the extra service costs more in excess stock than it saves in lost sales, and ordering the mean wins. Right: delivered service tracks the requested quantile at every level.
Cost ratioCritical quantileQuantile costMean costSaving
1:150%1.93M1.95M+1.0%
2:167%2.88M2.72M−5.9%
3:175%3.50M3.48M−0.4%
4:180%3.92M4.25M+7.7%
6:186%4.71M5.77M+18.5%
9:190%5.68M8.07M+29.5%

4.Method, in brief

Reproduce

pip install -r requirements.txt
kaggle competitions download -c store-sales-time-series-forecasting -p data/raw --unzip
python -m pytest -q          # 19 tests
python -m src.pipeline       # full run -> outputs/results.json
python -m src.make_figures   # redraws the figures above

5.Honest limitations