Abstract. Artificial intelligence is widely expected to transform supply chain management, yet many deployments stall between pilot and scale. This contribution argues that the recurring obstacle is not predictive accuracy but the translation of predictions into decisions. Drawing on operations management and behavioral research, it examines four sources of the gap: the conflation of model performance with decision performance, the constructed nature of supply chain data, the difference between predicting and deciding under uncertainty, and the social dynamics of adoption, and outlines implications for practice and for research.
Keywords: supply chain management; artificial intelligence; decision-making under uncertainty; data quality; algorithm aversion; Industry 5.0.
1. A promise that has become a rhetorical given
It is hard today to attend a supply chain conference without hearing that artificial intelligence will « change everything »: demand forecasting honed to the finest grain, real-time inventory optimization, early disruption detection, and autonomous planning. The narrative is appealing, and it is not wrong. Machine learning can capture weak signals that classical statistical methods miss, and the falling cost of computing makes these approaches accessible well beyond large principal investigators.
Yet in real deployments, a persistent gap appears between the value demonstrated in a pilot and the value captured at scale. Many projects stall at the proof-of-concept stage or produce technically impressive models that struggle to change day-to-day decisions. This gap is not a mere execution detail: it points to something structural about the supply chain as an object of management. This contribution examines that gap by crossing what research shows with what practitioners’ experience.
2. A model’s performance is not a decision’s performance
The first confusion to dispel is widespread: a good model does not mechanically produce a good decision. A forecasting algorithm can reduce mean error by several points without improving service level or inventory position, because the final decision depends on parameters that the model does not arbitrate, such as the cost of stockout versus the cost of holding, supplier capacity constraints, minimum lot sizes, and reliability of upstream transport.
Operations management has long shown that optimizing one link locally can degrade the performance of the whole system. The « bullwhip » effect is the classic illustration: improving a downstream forecast is useless if information does not flow and if actors keep reacting to distorted signals, with order variance amplifying as one move upstream (Lee, Padmanabhan, & Whang, 1997). In other words, predictive quality is necessary but largely insufficient. The real question is not « does the model predict well? » but « which decision does it change, and who owns it? ».
The practical consequence is strong. Before funding an AI project, it is more useful to map the targeted decision chain that decides, on what information, how often, under which constraints, than to compare model architectures. A project that cannot name the decision it intends to transform has little chance of producing value, however sophisticated the algorithm.
The invisible problem: data as a construct
We often hear that « data is the fuel of AI. » The phrase masks a difficulty every operational manager knows: supply chain data is not a raw fact; it is a construct. A « confirmed » delivery date can mean different things depending on the site, supplier, or system that records it. An order status may be entered in advance to satisfy a metric. Item master records merge with each acquisition, units of measure coexist, and exceptions live in parallel files that no one consolidates.
Models learn these imperfections without questioning them and may amplify them. A system trained on a history marked by chronic stockouts may « learn » that demand is low where it was simply not served the problem of censored demand, well documented in inventory theory: because a retailer cannot sell more than it holds, sales systematically understate true demand, and naïve estimation is biased downward (Tong, Feiler, & Larrick, 2018). Here, behavioral research converges with field common sense: a model is never better than its designers’ understanding of how the data was produced.
This calls for rehabilitating work often deemed unglamorous: fine-grained knowledge of data-entry processes, dialogue with those who populate the systems, and documentation of implicit business rules. Low in visibility and hard to « scale, » this work nonetheless conditions the reliability of the whole analytical edifice.
3. Deciding under uncertainty, not just predicting
A third tension deserves attention. Most effort concentrates on prediction, whereas value is created in decision-making under uncertainty. Predicting demand produces an estimate; deciding means choosing an action whose consequences are asymmetric. Being wrong by overstocking a perishable product does not carry the same cost as being wrong by understocking a critical component on an assembly line.
Approaches that settle for a point forecast obscure this asymmetry. More mature methods reason in distributions, scenarios, and robustness, not « what is the most likely demand? » but « which decision remains acceptable across most plausible futures? » Robust optimization formalizes exactly this logic, deriving ordering policies that perform well across an uncertainty set without committing to a single assumed demand distribution, while letting managers tune their level of conservatism (Bertsimas & Thiele, 2006). The relevance of this stance grows in a context of repeated shocks, geopolitical tensions, climate hazards, and volatile logistics costs.
For organizations, the implication is clear: resilience is not decreed by adding a predictive model, but achieved by embedding uncertainty into the decision itself, keeping room for maneuver, and arbitrating explicitly between efficiency and robustness.
4. Adoption: the real bottleneck
If so many projects stall, it is rarely for technical reasons. It is because the algorithmic recommendation collides with the organization. A planner who sees the tool propose an order contrary to their intuition can follow it, ignore it, or work around it. If they do not understand the recommendation, if they have been « caught out » by a tool before, or if their performance is judged on indicators the model does not optimize, they will ignore it legitimately, from their standpoint.
Behavioral research documents this precisely: people lose confidence in an algorithm faster than in a human after seeing it err, even when the algorithm outperforms the human a phenomenon termed algorithm aversion (Dietvorst, Simmons, & Massey, 2015). Crucially, the same research shows the aversion can be reduced: people are far more willing to rely on an imperfect algorithm when they can even slightly modify its output, because they regain a sense of control (Dietvorst, Simmons, & Massey, 2018). This argues for designs that involve operational staff from the outset, expose the reasoning behind recommendations, leave room for human adjustment, and clarify how responsibility is shared between human and machine. The « Industry 5.0 » perspective — refocused on human-machine complementarity, sustainability, and resilience rather than on automation alone (European Commission [Breque, De Nul, & Petrides], 2021; Xu, Lu, Vogel-Heuser, & Wang, 2021; Lu et al., 2022) offers a fruitful framing.
5. Implications for practice and research
A few convictions emerge, offered as an invitation to debate rather than as prescriptions. For practitioners, the useful starting point is not the technology but the decision: precisely naming what one seeks to improve, understanding how the corresponding data is produced, and designing the tool as support for an identified actor rather than as an anonymous substitute. A narrow scope that genuinely transforms one decision is worth more than an ambitious project that alters no routine.
For researchers, the gap between model performance and decision performance is a research agenda. It calls for work that does not stop at the predictive metric but measures decisional and organizational effects, and that takes data seriously as a constructed object and adoption seriously as a social phenomenon. It is precisely at the intersection of engineering, management, and the social sciences that the maturity of the AI-augmented supply chain will be decided.
Artificial intelligence does not exempt us from thinking about the decision: it raises the bar. Crossing the perspectives of science and the field is not a nice-to-have; it is the condition for the promise to become value.
References
- Bertsimas, D., & Thiele, A. (2006). A robust optimization approach to inventory theory. Operations Research, 54(1), 150–168. https://doi.org/10.1287/opre.1050.0238
- Dietvorst, B. J., Simmons, J. P., & Massey, C. (2015). Algorithm aversion: People erroneously avoid algorithms after seeing them err. Journal of Experimental Psychology: General, 144(1), 114–126. https://doi.org/10.1037/xge0000033
- Dietvorst, B. J., Simmons, J. P., & Massey, C. (2018). Overcoming algorithm aversion: People will use imperfect algorithms if they can (even slightly) modify them. Management Science, 64(3), 1155–1170. https://doi.org/10.1287/mnsc.2016.2643
- European Commission, Directorate-General for Research and Innovation (Breque, M., De Nul, L., & Petrides, A.). (2021). Industry 5.0: Towards a sustainable, human-centric and resilient European industry. Publications Office of the European Union. https://doi.org/10.2777/308407
- Lee, H. L., Padmanabhan, V., & Whang, S. (1997). Information distortion in a supply chain: The bullwhip effect. Management Science, 43(4), 546–558. https://doi.org/10.1287/mnsc.43.4.546
- Lu, Y., Zheng, H., Chand, S., Xia, W., Liu, Z., Xu, X., Wang, L., Qin, Z., & Bao, J. (2022). Outlook on human-centric manufacturing towards Industry 5.0. Journal of Manufacturing Systems, 62, 612–627. https://doi.org/10.1016/j.jmsy.2022.02.001
- Tong, J., Feiler, D., & Larrick, R. (2018). A behavioral remedy for the censorship bias. Production and Operations Management, 27(4), 624–643. https://doi.org/10.1111/poms.12823
- Xu, X., Lu, Y., Vogel-Heuser, B., & Wang, L. (2021). Industry 4.0 and Industry 5.0— Inception, conception and perception. Journal of Manufacturing Systems, 61, 530–535. https://doi.org/10.1016/j.jmsy.2021.10.006



