Redesigning Procurement Workflows for AI-Native Execution
Agents need stateful processes, not handoff chains built for humans.

Procurement workflows built for human teams fall apart when an AI agent, rather than a buyer, is the one carrying a request from start to finish. The failure isn't a shortfall in the technology; it's a mismatch between how the process is shaped and how an agent actually needs to work.
Why procurement workflows designed for human coordination fail under agentic execution
Human-run procurement is built around a chain of handoffs. Each link in that chain assumes a person is present to pick up the thread and decide what happens next. An AI agent doesn't work in handoffs. It holds state across the whole lifecycle of a task: it remembers the budget, the stakeholders involved, the evaluation criteria, and the surrounding context across weeks of a sourcing event, while a process built around handoffs is actively organized to break that continuity up into pieces.
Many tools sold as agents don't actually hold that kind of state. They take a prompt, produce an output, and forget everything the moment the task ends. That's fine for drafting an email. It falls apart for a supplier qualification that runs across weeks and dozens of separate decision points, because nothing is carrying the thread between them.
The mismatch appears the moment agents get layered onto a handoff-built process. The agent can't classify what it's looking at, can't route it to the right workflow, and ends up waiting for a human to step in at the exact point it was supposed to take over. This structural mismatch, handoff chains built for humans actively blocking agents that need stateful execution, is what Terminal Use looks for in its process audits before any redesign begins. The firm rebuilds the process from the ground up with the agent's need for persistent state treated as a starting requirement, instead of placing agents on top of an existing workflow; spotting the incompatibility has to come before anything else gets built. None of this reflects poorly on the people running procurement today. Broken intake didn't become a problem the moment agents entered the picture. It became a problem because agents need structured input to function at all, and the process was never built with that requirement in mind.
What the adoption data reveals about where organizations are stuck
The space between companies experimenting with AI in procurement and companies actually running it day to day is a gap in process design and in how the underlying data is organized, an evidence pattern now consistent across cases.
Across large organizations, individual employees have picked up AI tools on their own, while the organizations around them have changed very little at the operational level. Using a tool and transforming how a department runs are two separate outcomes, and most companies have only managed the first. A 2025 study from the MIT NANDA initiative found that most enterprise AI pilots produced no measurable effect on profit and loss. The cause wasn't weak models. The same research found that AI tools built through outside vendor partnerships succeeded roughly twice as often as tools built internally, which suggests that process knowledge and domain expertise mattered more to the outcome than how sophisticated the underlying model was.
Construction makes the pattern concrete. The intelligence is available. There's no redesigned process for it to run through, so it sits unused next to the problems it could otherwise help solve.
The structural logic that changes when an agent holds the thread end to end
Handing the thread of a procurement process to an agent instead of a buyer changes how data gets collected, how decisions get authorized, when a human needs to step in, and how one part of the process passes context to the next.
Start with data. Human-run procurement tolerates information that's incomplete, late, or scattered across formats, because a person is there to interpret it and fill the gaps from memory or judgment. An agent can't do that. Data completeness and structure stop being a nice-to-have and become a requirement the process has to enforce, because nothing downstream fills in what's missing. Decision rights change next. But the thresholds themselves have to be built into the process before the agent goes live. Working them out after something has already gone wrong is too late, and by then the damage is done.
A large multinational's deployment makes the shift concrete. The redesign wasn't a byproduct of the deployment. It was the precondition for it, and the original framing of the project, automating individual tasks, had to give way to a framing centered on orchestrating a process.
That kind of handoff between specialized agents, sourcing to legal to risk to negotiation, only works if context is explicitly structured and passed from one to the next. It can't rely on organizational memory or a buyer's accumulated knowledge of how things usually go, because none of that exists inside an agent unless someone has built it in. Escalation has to follow the same discipline. Setting those thresholds is a process design problem, not a setting to tune after the fact. At a recent BCG procurement event, practitioners described these organizational and process barriers as more serious obstacles than the technology or the algorithms themselves.
Where existing source-to-pay workflows break first under agentic pressure
Intake is where agentic procurement breaks first, and getting it right is the condition everything downstream depends on. When intake looks like that, an agent can't classify what's arrived, can't route it to the correct workflow, and can't apply the relevant policy, so a human ends up doing the very work the agent was deployed to take off their plate.
The problem compounds as the process moves downstream, because procurement data often gets confirmed only after the fact. An agent working against data like that is operating on a picture of reality that's already stale, and it can't monitor or respond to a problem it has no way of seeing in real time. The procurement data feeding that decision carries the same failure mode that drives cost overruns on construction projects.
Contract milestone monitoring and invoice validation are tasks agents are well suited to handle, and they're also where data quality problems pile up most visibly. An agent assigned to monitor a contract it can't actually read, because the contract sits as a PDF in a shared drive somewhere, functions as a bottleneck wearing an agent's name. Some organizations respond to this by cleaning data as they go, layering data-cleansing agents on top of the existing flow. BCG's own view allows for this as a partial answer: data-cleansing agents can catch data that's clearly wrong or implausible given historical benchmarks, and BCG recommends pairing that kind of cleansing with module-by-module upgrades to the underlying tech stack, treating neither one alone as sufficient.
What must be re-architected before any agent is deployed
Three things have to exist before an agent can be trusted to run a process: a unified data layer, decision thresholds that have been explicitly designed rather than left implicit, and a governance structure that spells out what the system decides on its own versus what it has to escalate. None of the three can be bolted on after the agent is already live.
Unifying procurement data is a process design decision. Data that lives scattered across an ERP system, email threads, supplier portals, and spreadsheets doesn't get unified by wiring those systems together. AI-native procurement platforms build intelligence directly into the data model itself, which only works if the data model was built with intention from the start. Skipping that step means whatever gets built on top of it won't hold once the volume scales up.
Decision thresholds, what an agent can execute on its own, what it has to flag, what it has to hand to a person, need to be designed in advance, because they define how the whole system is governed. Governance is part of the process architecture itself: which spend categories get proactive risk monitoring, which supplier decisions require a human signature, which contract deviations trigger an escalation, these are all decisions about how the process is built, and they establish whether the agent can be trusted to operate safely once it's running.
BCG's five-step framework for this work follows in order: move past copilots quickly, prioritize the workflows and decision points with the highest impact, redesign around the workflow rather than around a specific tool, build the data and governance foundation early, and only then scale toward a hybrid model where humans and agents share the operating load. In practice, that foundation gets built module by module: legacy systems get decoupled one piece at a time, and the order in which modules get replaced follows which agents are being deployed first, which in turn follows business value and readiness.
All of this depends on an audit that happens before any tool gets selected. The right starting point is one specific, painful process, examined in detail for how it actually runs today and where exactly it causes pain, before anyone chooses an agent or a platform to fix it. Organizations that skip the audit end up selecting tools for a process they don't fully understand, and the mismatch becomes visible later, after the money and the time have already been spent.
How to identify the right first process to rebuild
The right first process to rebuild is the one where today's failure is easiest to see, where the underlying data can most readily be structured, and where the cost of leaving things as they are is most concrete, not the process that merely looks easiest to automate.
The pull toward spend classification or invoice matching is understandable: these tasks produce visible, measurable results fast. The trouble is that quick wins on low-complexity tasks don't build the foundation an end-to-end agentic process needs, and they can quietly reinforce the instinct to patch. BCG is direct about this risk: don't get pulled in by an isolated use case just because it shows a return, because the real value of AI agents comes from connecting across systems, sourcing linked to supplier risk linked to compliance linked to contract, not from automating one task in isolation.
Field evidence keeps pointing to sourcing, negotiation, and supplier risk management as the highest-value places to start, because these are where decisions carry the most commercial weight and where an agent's ability to hold state and see across systems gives it the biggest edge over a process run by humans alone.
Two cases make the argument concrete. Bristol Myers Squibb redesigned its source-to-contract process, insourcing the work with roughly half the resources it had used before, running many times its previous RFP volume, and processing more than a billion dollars in spend in a fraction of the original timeline, against an original target of a hundred million. The redesign of the process is what made that possible, not a better tool layered on top of the old one. The speed was a result of the design, not a property of the model running underneath it.
Redesigning for agentic execution means treating decision rights, data structures, and escalation triggers as one connected system rather than separate fixes, and that kind of end-to-end rebuild starts with an honest look at how the process runs today before anything gets reconstructed around how agents actually need to work. That's architecture-first work, not tool selection, and the shape of the process is what determines what the agent can achieve, not the other way around. The process audit makes the choice of where to start legible: watching how a process runs now, where it waits, where data never gets recorded, where a human has to step in because the system can't, surfaces the right first candidate for a rebuild more reliably than any assessment of what a tool is theoretically capable of.
Why end-to-end redesign requires executive ownership, not procurement leadership alone
End-to-end redesign reaches into finance, operations, and enterprise strategy in ways procurement leadership has no authority to approve on its own. Without an executive sponsor, the rebuild stalls the moment it hits its first cross-functional dependency.
BCG states the requirement: this kind of transformation can't be driven by procurement teams working alone. The CEO or the COO has to take the lead, so the procurement architecture that gets built serves the competitive position of the whole enterprise. None of these sit inside procurement's authority alone. A deployment led only by procurement doesn't typically fail from a lack of ambition. It fails at the first cross-functional boundary, where the agent reaches the edge of procurement's own data jurisdiction and stops, because nobody with the authority to approve the integration ever signed off on it.
The cultural side of this carries its own cost: adoption stalls when the people running the process don't trust the agent sitting inside it. The speed advantage of agentic procurement comes from building the workflow to be agent-native from the outset, a principle that extends across industrial operations generally: Terminal Use deploys its first live process on a timeline of weeks by starting with one specific, painful workflow.
The payoff for getting the sponsorship right is concrete. When agents take over the execution layer, buyer time shifts toward supplier strategy, resilience planning, AI governance, and the negotiations complicated enough to still need a person in the room. That shift in where human effort goes is a decision about enterprise strategy, and it's the kind of outcome that justifies the executive attention the rebuild requires.


