The replenishment pilot hit its service target, then stalled when the rest of the network could not trust its data.
AI replenishment creates daily value only when retailers turn a promising forecast into a managed operating discipline.
That shift is now visible in large retail operations. Macy’s is adding an AI forecast overlay to its replenishment process, moving the capability from a pilot toward a broader rollout. The goal is practical: improve in stock levels while creating inventory efficiencies that can be measured in stores, distribution centers, and financial results. The work sits inside a three point transformation plan launched in 2024, expected to drive $235 million in savings by 2026. The plan includes closing unproductive supply chain centers and opening an automated distribution facility. Supply chain efficiencies were expected to appear in the second half of the year and support gross margin. Meanwhile, inventory was up 2.5% in the quarter, in line with sales growth, a useful signal that growth did not require an outsized inventory build.
Other retailers are tackling adjacent parts of the same operating problem. Target built a digital twin of its middle mile inventory positioning system to improve availability. Lowe’s expanded a partnership with Relex Solutions to apply AI to in stock levels and demand trend analysis. Kohl’s adjusted inventory depth and allocation, then reported a smoother spring receipt transition. These examples differ in technology and scope, but they point to the same progression: AI is becoming less of a laboratory exercise and more of a decision layer inside normal allocation, ordering, and receipt routines.

Consider a planner responsible for several hundred items across stores and an online channel. The day rarely begins with a clean list of replenishment recommendations. It starts with exceptions: a promotion that pulled demand forward, a supplier shipment that will arrive late, a new item with no history, or a store count that changed after the forecast was published. The planner may spend hours cleaning data, checking whether a demand spike is real, and explaining why a system recommendation should be accepted or rejected.
An AI overlay can shorten that work, but it does not remove the judgment. The planner still needs to see the forecast, the assumptions behind it, the expected lead time, and the consequences of changing an order. A useful system brings the highest risk exceptions to the front, proposes an action, and preserves a clear human override path. The result is not a planner replaced by a machine. It is a planner with more time to resolve supplier, promotion, and store execution problems before they become empty shelves or aged stock.
The challenge is that a successful pilot usually starts in a friendly environment. Demand may be stable, the category may use one channel, item attributes may be mostly complete, and a small team may provide unusually close oversight. At scale, the operating conditions change. Messy item master data can link the wrong pack size or location. Promotions can cannibalize nearby products and make history misleading. Supplier lead times vary, so an accurate demand signal can still produce a late receipt. Planners may distrust a recommendation that cannot explain itself, especially after one visible miss.
Forecast error also compounds through the network. A 20% forecast error at the top of the chain propagates downstream, affecting purchase quantities, distribution center positioning, store allocations, and the inventory available for digital orders. The response cannot be to add another layer of automation without fixing the underlying process. Most replenishment automation fails because data quality and process ownership are weak, not because the algorithm lacks sophistication.

Moving from pilot to practice therefore requires a deliberate operating sequence. First, retailers need a common definition of availability, excess, and acceptable service risk. Next, they need ownership for item attributes, lead times, pack configurations, and promotional calendars. The pilot should expand by category and channel only after the team can explain its wins and misses. Metrics should connect forecast quality to outcomes that operators and finance both recognize: on shelf availability, lost sales, excess inventory, working capital, receipt timing, and gross margin.
The economics become visible when better decisions accumulate. Fewer avoidable stockouts can raise on shelf availability without simply raising safety stock everywhere. More precise ordering can reduce excess inventory and markdown exposure. Better positioning can keep working capital from being trapped in the wrong node while another store waits for product. Those benefits may arrive gradually because months of tuning are normal. A forecast overlay is not a switch flip, and the first version will need adjustment as planners discover new exceptions and suppliers expose new constraints.
Retailers should resist evaluating AI replenishment as a standalone software purchase. The more important question is whether the organization can make the recommendation reliable enough for daily use. Before buying another forecast tool, audit the item master, promotion history, lead time data, ownership model, and override workflow. Choose a narrow category where the financial and service effects can be measured, then expand only when the process works under real operating pressure. The retailers that gain lasting value will not be those with the most impressive pilot. They will be those that make accurate, explainable replenishment part of every ordinary day.