Better Forecasts. Same Bad Decisions.
A decade of planning-tech investment. Only about 1 in 5 leaders saw real value — just 7% from agentic or generative AI. The constraint isn’t accuracy; it’s the architecture.
A decade in, the results aren’t there: only about 1 in 5 leaders report real value from advanced planning automation, optimization engines, or AI — just 7% from agentic or generative AI tools. Each disappointing result gets the same defense: not enough clean data, not enough process discipline, not enough adoption. This piece argues the real constraint is architectural — not a job half-finished.
There is a quiet contradiction running through enterprise planning right now. Organizations have spent a decade and a fortune on advanced planning systems and AI, and the people who run them will tell you, honestly, that the decisions haven’t gotten much better. The forecasts improved. The dashboards multiplied. The decisions — the actual trade-offs that move service, cost, and capital — mostly didn’t.
This isn’t a hunch. BCG’s Supply Chain Planning 2026 report, drawing on a survey of more than 180 planning leaders, observes that many organizations “struggle to translate better forecasts into better decisions that generate business value.” In the same survey, only about one in five leaders say more advanced capabilities — planning automation, optimization engines, or AI — have delivered meaningful value so far, and just seven percent report meaningful value from agentic or generative AI tools. Demand forecasting is, by BCG’s account, the most commonly scaled AI use case in planning — and yet the value leaders report from advanced AI remains, in their word, rudimentary.
That finding deserves more than a shrug. If better forecasts aren’t producing better decisions, the problem isn’t forecast accuracy — it’s the whole pipeline that runs from forecast to decision, and an assumption baked into every stage of it that we stopped questioning a long time ago.
It is worth being precise about that assumption, because the way past the current disappointment with AI in planning runs straight through it — not around it with more tooling, but through the one idea every layer of the stack quietly takes for granted.

The assumption nobody questions
Almost every planning system in the enterprise runs on deterministic master data. A lead time is 14 days. Capacity is a number. Demand is a single best-guess figure. Teams spend enormous effort keeping those numbers clean, governed, and owned — and that effort is real and worthwhile.
But consider the world those numbers are trying to describe. In the same survey, the top external challenges leaders name are demand volatility (70%), global trade and geopolitics (64%), and supply volatility (62%). Volatility, not technology, is the defining condition of the modern supply chain. And here is the thing a cleaner number can never fix in a volatile world: a supplier lead time is not 14 days. It is a distribution — often a lumpy, bimodal one, where most orders land in a week and a stubborn tail stretches to five. The “14 days” in your master data is, at best, an average that describes almost none of the actual orders. And you do not plan against an average. You plan against the spread.

This is why governance, by itself, hits a ceiling. You can make the number perfectly accurate, perfectly owned, perfectly consistent — and it is still a single number, still incapable of carrying the variability that actually determines whether your plan survives contact with reality. Planning off deterministic master data in a volatile world isn’t a discipline problem. It’s a paradigm that has quietly become legacy.
And it explains the contradiction we started with — because determinism doesn’t enter the pipeline once. It enters twice.
First, the forecast itself is almost always a point. “We’ll sell 1,200 units” — maybe with a token confidence band stapled on — when the truth is a full demand distribution with real probabilities across a wide range of outcomes. The richest signal about one of the two main drivers of what to make, buy, and move is flattened to a single number before it is ever handed to the planning engine. Then it happens again on the supply side: that already-impoverished number is fed into a deterministic solve that collapses lead times, capacity, and yield — each its own distribution — into one more single plan. Uncertainty is deleted twice, at both ends of the very calculation that is supposed to manage it.

To be fair to modern planning systems, this is not the 1990s: good APS already tries to cope with variability — safety-stock math, multi-echelon inventory optimization, a library of pre-built scenarios, some Monte Carlo around the edges. But notice what all of that has in common. It bolts variability onto a deterministic core after the fact — buffers added on top of a single plan, a handful of scenarios run beside it. The uncertainty is treated as a correction to a point answer, not as the thing the decision is optimized over in the first place.
That distinction is the whole game, and it has two halves that deserve equal weight. The first is stochastic demand forecasting — the full range and probabilities of what customers might do, not a single number with a confidence band stapled on. The second, and just as important, is stochastic supply planning — a solve in which lead times, capacity, and yield enter as the distributions they actually are, and the optimizer produces a multi-stage policy that stays good across that whole supply-side spread, re-deciding as the future narrows. Not buffers bolted onto a deterministic plan — uncertainty carried inside the optimization itself.
Neither half is sufficient alone. A resilient supply solve fed a single-point forecast is still half-blind — it has thrown away the demand-side uncertainty that drives most of what you make, buy, and move. And a stochastic forecast handed to a deterministic supply solve is collapsed right back into a point the moment it arrives. What the decision actually requires is both at once: the demand distribution and the supply distributions solved jointly, so the policy you land on is good across the combined space of what customers might do and what your network might do. That is what it means to carry the distribution into the decision — end to end, on both sides, never collapsed to a point at either end and then reconstructed as a scenario after the fact.

It is fair to ask whether this is actually buildable, or just elegant. Three things have changed that make it practical now: compute is cheap enough that solving across many sampled futures is a budget line, not a research project; the methods for stochastic and multi-stage optimization are mature rather than experimental; and most enterprises are already sitting on years of transactional history — the raw material for learning the distributions — without realizing it. None of that makes it trivial. Long-tail SKUs are data-sparse, and you handle that with hierarchical pooling and regularization rather than pretending each one has a clean distribution. Demand patterns drift, so the distributions have to be re-estimated on a rolling basis rather than learned once. These are real engineering problems with known approaches — which is a very different thing from the structural impossibility of asking a deterministic solve to produce a resilient policy.
BCG’s report includes a case that is worth reading closely — partly because of what it shows, and partly because of how BCG reads it. A global medical technology company’s leading planning system was producing inflated, unreliable supply plans and low organizational confidence in its outputs. BCG attributes the root causes to foundational gaps: key planning parameters such as lead times, capacity constraints, and inventory netting that, in their words, “did not reflect actual operations,” together with incomplete inventory visibility. The detail I’d point to: even when the demand inputs were accurate, the system still overstated production, forcing planners into constant manual overrides. The company then corrected those parameters and strengthened its planning process, and BCG reports this reduced the demand signal by roughly $100 million while protecting service levels.
BCG traces this to a fix that worked: correct the parameters and governance, and the APS delivers. That's right, and worth taking seriously. What I'd add is a reading of why those parameters drifted in the first place, and this reading is mine, not BCG's. Lead times, capacity, and netting "did not reflect actual operations" for a structural reason: they were modeled as fixed numbers, and fixed numbers can't represent operations that are inherently a spread. Correcting them by hand is the organization re-injecting, once and up front, the variability the model had flattened, the same thing planners do every day in spreadsheets when they override the plan. That this works isn't in dispute. The question I'd press is what happens next: operations move again, the spread you hard-coded drifts, and the overrides return. The manual correction is the tell. Which points to a conclusion BCG doesn't draw, and I do: the limiting factor isn't dirty data or weak process. It's that a single number was the wrong object to plan against in the first place.
And the upside of treating it as a different object is not abstract. Setting BCG’s cases aside and speaking from our own work, the pattern is consistent even where confidentiality keeps the names off the page: carrying the distribution into the decision — illustratively — tends to hold the same service level at meaningfully lower inventory, or cut expedite and write-off cost at the same service, because the policy stops being optimized for one future and starts being robust to the range. The exact deltas vary by network and category; the direction does not.
Intelligence isn’t a layer on top. It’s the decision itself.
The common mental model treats the planning engine as the backbone and AI as an intelligence layer bolted on top — predicting, explaining, automating around the edges of a deterministic core. It is a comforting picture because it changes nothing structural. It is also the reason most AI pilots stall: intelligence that lives beside the decision, rather than inside it, is easy to ignore and impossible to scale. (BCG itself names “inadequate integration” — AI that runs alongside APS rather than within the planning workflow — as a core reason pilots fail to scale.)

Invert the stack — and be precise about what each piece is actually for, because this is where most of the industry quietly gives up on the value.
The decision engine’s job is to decide under uncertainty: to produce a policy that holds across hundreds of plausible futures, not a single plan that is optimal for one and fragile to the rest. Beneath it, the ERP and system of record do what they are genuinely built for — holding transactions and master data, executing and remembering what happened. That layer is real and it stays.
The deterministic planning solve also stays — but it is worth being honest about what it can and cannot do, because the whole architecture turns on this. Give a deterministic engine one fully specified future — one set of point values sampled from across the spreads — and it is genuinely useful: it will tell you what that world looks like, where the plan tears, what is infeasible, what the exposure is. That is risk diagnostic work, and it is real value. A deterministic solve, called many times across many sampled futures, is a perfectly good diagnostic engine.
What it structurally cannot do is resilient planning. It cannot tell you what to buy, make, store, and ship — where, and when — such that the decision stays good across most of the futures you might land in. It cannot check whether a decision holds across progressively elaborated trajectories, or re-solve over a rolling horizon as the future narrows. That is not a missing feature you could add with better adoption or cleaner data. It is a different mathematical object: a multi-stage policy optimized over a distribution, not a single plan optimized for a point. The deterministic solve was never built to produce it, and no amount of tuning makes it.
So the layers sort themselves honestly. The decision-intelligence layer holds the distributions, and it does two distinct jobs with them. For diagnosis, it samples futures and calls the deterministic solve — here APS earns its keep, as the fast engine that scores each world. For the resilient decision, it does the thing the deterministic solve cannot: it optimizes a policy across the whole spread and tests that policy forward across many trajectories, re-deciding as reality unfolds. APS isn’t displaced and it isn’t preserved as a backbone. It becomes the diagnostic solver-on-call inside a probabilistic engine that does the actual deciding.
This is exactly where the industry leaves the value on the table. The popular move is to keep the decision inside the deterministic engine and hand AI a smaller job: explain the plan, narrate the change, summarize the exception in fluent language. That is the band-aid — it asks AI to be a commentator on a decision it was never allowed to improve. This reading is mine, not BCG’s: where the survey records that only seven percent of leaders find real value in agentic or generative AI, I’d argue a structural cause sits underneath the adoption explanations — the AI was pointed at the explanation, not the decision. You cannot narrate your way to resilience.
Put uncertainty where it belongs instead. The distribution drives the decision directly — which futures get diagnosed, and which policy survives them — rather than being predicted early and discarded before the solve. The engine produces a resilient decision, not a deterministic plan with a chatbot bolted to the side.

It also changes the human job, for the better. When the engine absorbs the day-to-day variability, planners stop firefighting overrides and start doing the high-value work: setting the terms the optimizer solves under — risk appetite, service constraints, the decision policy itself — and calibrating that policy alongside the executives who own the strategy. Less reconciling numbers. More designing the rules the enterprise runs on. (This is consonant with where BCG lands too: planners shifting from data processors to decision orchestrators.)
Two systems are converging into one
This is also what finally ends the tired “AI versus APS” debate. The question was never which one wins; they do different jobs, and only one of them can carry the decision. And you can already see the boundary dissolving in real time. Planning vendors are embedding learning, optimization, and agents into their platforms. AI-native tools are reaching down into constraints and the solve. Each is racing to do what the other does. BCG observes the same convergence from the platform side — the line between classic planning and decision intelligence, in their phrasing, is blurring.
That is not two stable layers settling into a comfortable coexistence — it is convergence. The deterministic engine underneath isn’t exactly thriving on its own terms in the meantime: in the same research, planning-system programs frequently run over budget and behind schedule, only a fraction of intended planners use the systems consistently, and many quietly revert to spreadsheets. The lesson isn’t that the platform is worthless. It is that asking a deterministic point-planning paradigm to carry the whole decision, in a volatile world, has a structural limit. The platform is fine. The job we gave it was wrong.
So the two systems resolve into one: a single decision intelligence that decides under uncertainty, sitting over a system of record that executes, with the deterministic solve doing diagnosis inside it. The teams that see this early stop trying to perfect the boundary and start building for the world on the other side of it — one decision intelligence, one substrate, one flow from sensing to decision to execution.
The decision needs a currency
There is one more piece, and the same research points right at it. Asked where the largest untapped upside in planning lies, leaders name deeper integration of financial and operational decision making — and yet financial integration ranks among the least-mature capabilities in the entire study, second only to transformation capability itself. Almost everyone names the destination. Almost no one has the vehicle.
Here is the vehicle. Operations and finance integrate a decision when every trade-off is priced in the same unit — the one finance already answers to. Expected service against expected margin against the cost of the capital the supply chain quietly consumes, all expressed in one currency: economic profit. EVA.

Give the enterprise a shared decision currency and integration stops being an aspiration on a maturity curve and becomes, in principle, computable. The trade-off that used to require a cross-functional summit — do we expedite, do we hold buffer, do we re-source — becomes a single number every function already understands. There is real work underneath it, and it is worth naming honestly: you need a costing and capital-charge model good enough to price the trade-offs, and cost-to-serve allocation is itself a discipline. But that is a tractable prerequisite, not a reason the goal stays out of reach. Done, it turns service, margin, and capital from three scorecards in tension into one computation — decision intelligence as an enterprise-performance engine rather than a reporting upgrade.
Beyond the hype: a value-centric paradigm
Step back from the architecture for a moment and look at where the value actually moves. For three years the planning conversation has been organized around the technology that is loudest — generative AI, then agentic AI — and the survey numbers are the predictable result: a few points of value here, a successful pilot there, and a pervasive sense that the spend has outrun the return. That is what happens when the tool leads and the decision follows. The needle moves when you invert that: start from the decision you are trying to make well, and let each technology earn its place as a component in service of it.
Seen that way, none of the new capabilities go away — they get a job that fits them. Generative AI is genuinely good at explanation, so it explains: why the policy chose what it chose, which futures drove it, what changed since last week. Agentic AI is genuinely good at acting within guardrails, so it runs the event-driven replanning — a supplier slips, a node goes down, and the system re-decides within its impact radius without waiting for the monthly cycle. Both are valuable. Neither is the headline. The headline is the resilient decision itself — a stochastic view of demand and supply, solved jointly into a policy that holds across many futures and priced in one currency the whole enterprise shares. GenAI and agents are how that decision gets explained and executed, not what produces it. The cart stops leading the horse.

This is the shift from a hype-centric paradigm to a value-centric one. It is also why the same capabilities that have underwhelmed so far start to pay: not because the AI got better, but because it was finally pointed at the decision instead of the explanation of a decision that was never improved. The investments already on the books — the planning systems, the data, the pilots — begin to compound, because they are working in service of an engine that actually decides under uncertainty rather than around the edges of one that cannot.
Plans break. Policies hold.
A single plan, however clean and well-governed, is optimized for exactly one future; the moment reality diverges it breaks, and the organization spends its days in manual override. A policy is built to hold across the range — and when it is expressed in economic terms the whole enterprise shares, service, margin, and capital stop being three scorecards in tension and start being one decision. That is the difference between better forecasts and better decisions. Better forecasts were never going to be enough. The decision was always the point.

A note on sources. The survey statistics (including the 1-in-5 and 7% value figures and the volatility percentages), the medical technology company case, and the financial-integration maturity gap cited above are drawn from BCG's Supply Chain Planning 2026: Why AI Alone Isn't Enough (Boston Consulting Group, February 2026); other BCG observations are attributed inline. Where this piece advances a conclusion BCG does not — most notably that the limiting factor is a deterministic paradigm rather than an operating-model gap — it is flagged in the text as the author's interpretation. BCG's own thesis is that organizational readiness, not technology, drives planning excellence; this piece agrees on the symptom and diverges on the cure.