Do retailers need deep learning to forecast demand?
Not in this benchmark. XGBoost matched or beat five linear and neural-network alternatives on 913,000 historical sales records, while training faster than every neural model. The study does not prove the same choice will work in a live retailer.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
At a glance
- 1Across 500 store-item series and a 153-day holdout, XGBoost recorded the lowest average error, but its accuracy advantage over GRU was not statistically significant after correction.
- 2XGBoost trained in 3.2 minutes in the reported environment, compared with 20 to 95 minutes for the neural models, and produced the lowest simulated inventory cost under the study's assumed policy.
- 3The cost figures are modelled rather than observed savings, the main data end in 2017, and the larger Rossmann check tests only XGBoost—not all six alternatives.
Research topic
Whether conventional machine learning or deep learning offers the stronger practical trade-off for structured retail demand forecasting and simulated inventory decisions
The answer: start with the simpler benchmark, then test locally
Retailers do not automatically need deep learning to forecast daily demand. In this peer-reviewed benchmark, XGBoost—a gradient-boosted tree method—produced the lowest average error across 913,000 historical sales observations and trained much faster than the four neural-network designs. The result supports a practical rule: compare a well-engineered tree model before paying for a more complex architecture.
That is not the same as proving that XGBoost is universally best. Its root mean squared error was 7.95 units, against 8.02 for the strongest neural alternative, a gated recurrent unit model. The difference between those two was not statistically significant after the study's multiple-comparison correction. CNN performance was also close enough that its corrected comparison with XGBoost was not significant. Accuracy alone therefore does not justify a sweeping claim that trees beat neural networks everywhere.
The more useful finding is operational. The XGBoost model reached its reported result with 3.2 minutes of training and 1.1 GB of memory. Neural training took 20 minutes for CNN, 45 for GRU, 65 for LSTM and 95 for the CNN–LSTM hybrid, with memory use from 2.1 to 3.5 GB. For a retailer that has structured sales, calendar and lag features, the cheaper model may be the better first deployment candidate even when accuracy is effectively tied.[1]
The Impact Brief · Free
Follow the evidence in finance & business.
Get a five-minute weekday briefing on what changed, why it matters and where the evidence comes from. Choose the topics you care about.
What the researchers actually compared
Suleiman Ali Alsaif of Imam Abdulrahman Bin Faisal University in Saudi Arabia evaluated six approaches inside one pipeline: linear regression, XGBoost, a convolutional neural network, long short-term memory, a gated recurrent unit and a CNN–LSTM hybrid. The goal was benchmarking, not inventing a new model. Each approach received the same underlying engineered features and a time-aware validation process.
The main Kaggle dataset contains daily sales for 50 items across 10 stores from January 2013 to December 2017. The first 836,500 observations, through July 2017, formed the training period. The last 76,500 observations—500 store-item series over 153 days from August through December—were held out for testing. Features included calendar signals, one- to 28-day lags, rolling averages and variation, exponentially weighted averages, and store- and item-level summaries. The paper says rolling calculations and seasonal decomposition were fitted within each training fold to reduce leakage into validation and test data.
XGBoost achieved RMSE of 7.95, mean absolute error of 6.11, mean absolute percentage error of 12.5% and an R-squared value of 0.931. Linear regression's RMSE was 9.33. GRU's RMSE was 8.02 with R-squared of 0.930, while CNN's R-squared was 0.929. The statistical tests used errors from 500 store-item series across the test period and applied Holm–Bonferroni correction to seven planned comparisons. XGBoost clearly improved on linear regression, LSTM and CNN–LSTM; its edge over GRU and CNN was too small to treat as decisive.[1][2]
Why the inventory numbers need careful reading
The paper moves beyond forecast accuracy by feeding each model's predictions into a simulated periodic-review inventory policy. Across the 153-day test period and 500 store-item combinations, it assumes a daily holding cost of 50 cents per unit, a shortage cost of $5 per unit, a one-day review period and a three-day lead time. Under those settings, XGBoost produced the lowest total modelled cost: $656,709, compared with $662,831 for GRU and $769,672 for CNN–LSTM.
Those figures are not money saved by an operating retailer. They are outputs of a simulation whose price assumptions, replenishment rules and demand history are fixed by the study. The authors' annualised $270,000 difference between XGBoost and the worst model extrapolates the five-month simulation. It should not be copied into a business case without rerunning the policy using the retailer's actual margins, stockout consequences, lead times, perishability and service targets.
Uncertainty is also simplified. None of the six models produces a full predictive distribution, so the simulation uses each model's RMSE as a proxy for forecast uncertainty when setting safety stock. The paper calls this crude: average error does not capture changing uncertainty across items or time. A retailer would want calibrated prediction intervals or quantile forecasts, followed by a backtest of the ordering policy itself—not only point-forecast rankings.[1]
The larger dataset tests scale, not general superiority
A second experiment applies XGBoost to the Rossmann Store Sales benchmark, covering as many as 1,115 stores and adding promotions, competition distance and school-holiday variables. The reported R-squared value was 0.946, and training took 18.5 minutes. This suggests that the selected model can process a much larger and richer structured dataset in the paper's environment.
But the study does not train the other five models on Rossmann. It therefore cannot show that XGBoost retains its relative lead when every alternative receives those external variables. The authors state this explicitly. The comparison establishes cross-dataset performance for XGBoost, while the six-way ranking remains tied to the smaller Store Item dataset.
The benchmark also omits several tests that could change the conclusion. Transformer forecasting models and newer architectures such as PatchTST, N-HiTS, TimesNet, TiDE and TimeMixer were excluded. Hyperparameter searches used a fixed practical budget rather than exhaustive optimisation across repeated random seeds, which may disadvantage neural models. No ablation study shows how much of the performance comes from lags, rolling statistics, seasonal decomposition or calendar features. The paper is reproducible enough to guide a shortlist, not comprehensive enough to end model selection.[1][3]
Impact on workers, customers and decisions
For inventory planners and store teams, the immediate lesson is that model governance matters more than architectural prestige. A smaller, faster model can be easier to retrain, explain and monitor. That can reduce engineering workload and make it easier for planners to inspect why a forecast changed. The study does not measure staff time, usability or adoption, so those benefits remain reasonable hypotheses rather than observed outcomes.
Customers could benefit if more accurate and less biased forecasts reduce stockouts without creating excess inventory. They could also be harmed if a historical benchmark hides local demand shifts, if automated orders amplify a bad promotion signal, or if a single error metric causes planners to overlook essential but low-volume products. Human override rules, item-level monitoring and differentiated service targets remain necessary, especially for medicines, food and other consequential goods.
Procurement teams should ask vendors for a time-ordered comparison against a transparent baseline using the organisation's own data. The report should include every item and site in the denominator, confidence or prediction intervals, error by product group, forecast bias, compute and maintenance cost, and the effect of forecasts on a realistic ordering policy. A complex model should earn its place through measurable operational value, not by being labelled deep learning.[1]
Evidence limits and what would change the assessment
This is a peer-reviewed analysis of public historical benchmarks, not a randomised trial or live deployment. The main sales data are from 2013–2017, before pandemic disruption, current e-commerce patterns and many contemporary forecasting systems. Store and item identifiers are anonymised, limiting assessment of category-specific error. The second dataset is also a competition benchmark, and only the winning first-stage model is evaluated there.
The author reports no financial support and no commercial or financial conflict of interest. Generative AI tools were disclosed for language editing and proofreading; the author states that the scientific content was reviewed and verified. The transparent disclosure is relevant, but it does not independently validate the analysis. Replication by another group using the released pipeline and additional retail datasets would carry more weight.
The assessment would strengthen if all six models and modern forecasting baselines were tuned under declared, comparable budgets across several current operational datasets; if probabilistic forecasts were calibrated; and if inventory policies were evaluated prospectively with actual costs and service levels. Evidence that a simpler model sustained lower stockouts, waste, planner workload and total cost after deployment would turn this useful benchmark into a much stronger business decision case.[1][2][3]
What this means for people
- Inventory planners may get a faster, easier-to-maintain baseline before investing in neural forecasting infrastructure.
- Store workers and customers could see fewer shortages or less excess stock only if the benchmark holds under local costs, products and demand shifts.
- Leaders should keep override, monitoring and escalation controls because this study tests forecasts and simulated orders, not autonomous live replenishment.
Global context
The study comes from a Saudi university and tests widely reused public retail datasets, including a German pharmacy-chain competition benchmark. Its conclusion is globally relevant as a warning against assuming that more complex AI is automatically better. Generalisation remains limited: retail demand, promotion calendars, supply constraints, data quality and the consequences of shortages differ substantially across countries, firms and product categories.
What the evidence does not yet show
- The principal comparison uses public sales data from 2013–2017 rather than a current live retailer.
- XGBoost's accuracy advantage over GRU and CNN was not statistically significant after correction, despite having the lowest point estimates.
- Inventory costs are simulated under assumed holding, shortage, review and lead-time parameters; they are not observed savings.
- RMSE is used as a crude uncertainty proxy because the tested models produce point forecasts rather than predictive distributions.
- Only XGBoost is tested on the larger Rossmann dataset, so that experiment does not replicate the six-model ranking.
- Transformer and newer forecasting architectures, exhaustive tuning, repeated seeds and feature-ablation tests are absent.
What to watch next
- Independent replication of the full six-model comparison on Rossmann and other current, operational retail datasets.
- Probabilistic or quantile forecasting with calibrated intervals feeding a realistic inventory policy.
- Prospective deployment evidence reporting stockouts, waste, service level, planner workload and total cost.
- Comparisons with modern Transformer and MLP forecasting models under equal, declared compute and tuning budgets.
- Item-, store- and category-level results showing where a global winner fails important local groups.
Living evidence record
Impact record IAI-14LDTR0
Evidence stage
Studied
Confidence
Supported
Reporting basis
Source analysis
Independent or research support
Present
Record status
Monitoring
Last checked
10 October 2026
Source trail
3 direct sources across 1 source type.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
Related-source reporting disclosure
This record analyses 3 linked source records around the same underlying development. The extra records add method, date or context, but they do not by themselves constitute independent replication of every performance claim or predicted outcome.
Evidence trail
Sources used for this report
Links checked 10 October 2026
This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
Finance & Business
Do prompt embeddings improve demand forecasts?
Not in the strongest result from one peer-reviewed inventory benchmark. Structured XGBoost beat its prompt-embedding hybrid on 14,700 held-out records, while the hybrid helped only the weaker Random Forest family.
9 min · 1 source
Finance & Business
Should lenders turn borrower data into images for AI?
A 1.34-million-loan benchmark found a convolutional-neural-network and random-forest hybrid improved default discrimination, but much of its power came from LendingClub's existing risk grades and it was never tested in live lending.
7 min · 3 sources
Finance & Business
Is AI financing becoming circular?
A Bank for International Settlements analysis maps 1,246 AI firms and 972 investment relationships, finding that commercial links accompanied 16.1% of AI-to-AI deals by count but 46.4% by disclosed value. The pattern is consequential, yet it does not show that the transactions were improper, unprofitable or certain to spread financial stress.
10 min · 5 sources
Reader commentary
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Explore commentary across the portal →Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.