智能体自动化如何构建超越“节省工时”的商业论证
Beyond hours saved: Building the business case for agentic automation
AWS 提出面向智能体自动化的 Agentic Value Model 商业论证框架,指出传统 RPA 时代的“节省工时×人力成本−构建成本”ROI 模型漏掉了大部分价值。
Agentic automation, software that reasons and adapts to complete tasks, is showing up on AI center of excellence (AI CoE) roadmaps. But the standard way companies justify automation investments, hours saved times labor cost minus build cost, was designed for rule-based tools like robotic process automation (RPA). That model misses most of the value agents create.
In this post, we introduce a framework AI CoE leaders can use to build a business case that captures the full value of agentic automation. You will learn why the RPA-era ROI model falls short, how to measure the value it misses, and which workflows are worth automating with agents.
Why the traditional business case falls short
The classic return on investment (ROI) model was built for stable, high-volume, rule-based work: count the transactions, measure the minutes, multiply by a loaded rate, subtract the build cost. RPA earned its place on those terms.
The model assumes the world it was built for. It assumes processes are stable, so it ignores the cost of maintaining automation as they change. It assumes tasks are rule-based, so it has no line item for exceptions. It assumes the work is the whole job, so it never counts the cost of human oversight. And it treats a saved hour as banked value, when freed capacity often refills with backlog and never reaches the profit and loss statement (P&L). Underfunding the work that turns saved hours into results is the mistake McKinsey says companies keep making: successful AI transformations follow a “1:3:5 pattern,” where “for every dollar invested in agentic technology, organizations spend three on process redesign and five on capability building and adoption. However, most companies invert this formula entirely” (McKinsey, “Agentic AI change management: Closing the adoption gap,” 2026).
We assume the value of automation lives in the task, but in agentic automation most of it lives around the task: the judgment, the exceptions, the coordination across systems. The real gains come from redesigning the workflow around agents, not from dropping an agent into an unchanged process.
The dimensions of value the old model misses
A better business case measures four dimensions of value, plus the condition that determines whether they reach the P&L. We call it the Agentic Value Model.
Time savings. This still counts, but agents extend it to work that fixed-rule RPA handles poorly or only through extensive exception logic. The measure is the same. Agents apply it to a larger base of work.
Exception handling. Exceptions carry much of the cost. AWS guidance gives planning ranges: correcting an error can cost 1.5–4 times the original transaction, and human error can account for 2–15 percent of operational cost (AWS Prescriptive Guidance, “Assessing your current human-process costs”). The rework multiplier prices the exceptions your team catches. The error percentage prices the ones that slip through. Labor-only cases omit both.
Decision quality. Agents can apply a common policy at scale with logged rationale, though consistency and error rates need continuous measurement. As AWS notes, “low-volume, high-value decisions might justify agentic assistance for improved decision quality rather than cost reduction” (AWS Prescriptive Guidance, “Understanding agentic AI economics”). A single better decision in credit, pricing, or risk can dwarf a year of saved minutes.
Change resilience and maintenance economics. This one goes both ways. Scripts are brittle: when a screen or upstream system changes, they break and someone rebuilds them. Agents absorb some variation without a rewrite, but they shift maintenance into evaluations, prompts, monitoring, and model operations rather than removing it, and running them carries its own cost. The case nets brittle-script maintenance avoided against ongoing agent operating cost, so in processes that change often the math may favor agents, and in frozen processes it might not.
The condition that governs all four is value realization: every benefit needs a defined mechanism that converts an operational improvement into economic value, plus an accountable owner. For released labor, that means reducing spend or redirecting capacity to a named, measurable outcome.
Consider a claims-triage process, with illustrative figures. It handles 200,000 claims a year at about 12 minutes per claim (before rework) and $45 an hour fully loaded, a base cost of roughly $1.8 million. Automate the routine 70 percent and you release about 28,000 hours, or roughly $1.26 million of capacity, before the oversight time those claims still need. That’s where most cases go wrong, because freeing hours isn’t the same as saving money. The company only spends less if it employs fewer people or cuts contractor, overtime, or outsourcing spend. Otherwise, the same people stay on the payroll and the P&L never sees the $1.26 million. If attrition and lower overtime capture half of it, the case should carry about $630,000, not the full figure.
Redeploying the people is the more common outcome, and then the value is what they now produce. For each released hour, count redeployed value or a cash cost reduction, not both. Then add what the labor-only case cannot see. Of the 8 percent needing correction (16,000 claims), a correction cost of 3.5 times the roughly $9 handling cost (12 minutes at $45 an hour) gives about $504,000 in annual correction exposure. That is a baseline, not agentic value. The value is the share the agent removes: a 40 percent reduction, adjusted by a 75 percent realization factor, is a modeled benefit near $151,000. That’s before you price a single improved fraud-flag decision. Because the 12-minute base excludes rework, this pool doesn’t overlap with released capacity. If your baseline includes rework time, count those savings in one pool only. That version of the case produces a number finance can defend.
To keep the math consistent, size every value pool the same way and net the costs, as summarized in the following table.
| Value pool | Baseline | Expected delta | Realization factor | Owner |
| Released capacity | Hours × loaded rate | % automated | Only if redeployed to a named outcome or spend falls | Ops lead |
| Exception and error cost | Correction cost + error losses | Expected reduction | Fraction captured | Quality lead |
| Decision quality | Value of better decisions | Uplift per decision | Attributable share | Domain owner |
| Change resilience and maintenance economics | Break-fix cost of scripts + downtime | Avoided rebuilds, faster change accommodation | Share of rebuilds avoided, with agent costs netted once in the formula | Engineering lead |
State the whole thing as one line, assigning each benefit to one pool so double-counting can’t slip through: Net annual value = realized capacity value + avoided correction and error costs + decision-outcome uplift + avoided maintenance and downtime, minus annualized implementation cost, minus agent runtime, integration, evaluation, oversight, governance, and change-management costs, minus losses from any new errors the agent introduces. Errors the agent does not remove already sit outside the avoided-cost term, so do not subtract them a second time. For a multiyear program, book implementation cost in the year you spend it rather than annualizing it, ramp benefits as adoption grows, and discount the yearly cash flows to net present value (NPV).
Three early deployments of Amazon Quick Automate, the automation service within Amazon Quick, show the dimensions at work. Kitsa, a clinical-trial site-selection company, automated extraction of more than 50 data points across hundreds of thousands of websites. It reported 91 percent cost savings and 96 percent faster data acquisition at 96 percent coverage, routing low-confidence cases to reviewers (time savings and exception handling). As Kitsa co-founder and CTO Rohit Banga noted in the AWS Machine Learning Blog post, unifying high-quality site data at scale broke a core bottleneck. That let its Site Finder Agent weigh more sites with greater precision.
dLocal, a cross-border payments provider, automated up to 75 percent of merchant-compliance reviews in controlled evaluations. That freed specialists for complex, high-risk cases (decision quality) and added continuous checks for policy drift that a periodic manual review would miss (change resilience). dLocal reports that specialists who used to spend hours on routine checks now spend that time on cases that need regulatory judgment.
And Genpact, automating supply-chain risk across multiple SAP systems, reports cutting disruption-impact analysis from 2–3 days to minutes. That’s the kind of cross-system coordination that no one system of record owns (time savings and decision quality). None of these deployments proves all four pools, which is why the case has to size each one separately.
A framework for prioritization
Score every candidate workflow on two axes. The first is task complexity, how much reasoning, context, and adaptation it requires. The second is decision risk, the cost of getting it wrong, which sets how much autonomy you can grant. Both axes come from the AWS economics guidance, and together they form a two-by-two matrix a steering committee can use.
Low complexity, low risk: keep it on RPA, since agents add cost here without adding value. High complexity, low risk is the capacity play, so automate for throughput and exception absorption. High complexity, high risk is the decision-quality play, so keep a human in the loop and justify on better decisions. Low complexity, high risk is the guardrails play, so harden controls around the step rather than adding reasoning it doesn’t need.
A useful shortcut: if a person moves across systems and interprets ambiguous context at each step, evaluate an agentic approach. If the path stays deterministic, RPA may still be cheaper. Use the Agentic Value Model to decide what to measure and this matrix to decide what to prioritize.
Figure 1: Agentic automation prioritization matrix, with axes adapted from AWS Prescriptive Guidance and illustrative classifications
Building the case for leadership
Present the investment as a portfolio of workflows rather than a single project. McKinsey’s research shows why isolated pilots stall: nearly two-thirds of enterprises have experimented with agents, but fewer than 10 percent have scaled them to tangible value (McKinsey, “Scaling agentic AI with data transformations,” 2026). Value comes from defining the outcomes, embedding agents deep in core workflows, and redesigning operating models around them. Bring leadership three to five prioritized domains, each with its four-dimension case.
The one thing that wins over a skeptical CFO is a stop rule: stage the investment with explicit break-even targets and predefined points to stop funding agents that underperform. Connect each domain to a key performance indicator (KPI) leadership already tracks, whether cost-to-serve, cycle time, Net Promoter Score, or compliance exposure. And treat governance as a way to move faster: bounded autonomy, oversight where risk is high, and auditability by design let you widen agent autonomy as the evidence builds.
Getting started with Amazon Quick
Amazon Quick brings research, business intelligence, and automation into one agentic experience. With Quick Automate, its automation service, teams orchestrate UI actions, API calls, and human review across enterprise workflows. This produces case-level execution data a CoE can pair with operational and financial KPIs to measure value as it’s realized.
That measurement turns a first project into a portfolio. After you look for the four value pools, you tend to find them wherever a person moves between systems, interprets context, and clears exceptions by hand. Put one workflow into production, baseline its value pools and operating cost, name an owner, and prove the number. Then the portfolio case accrues as each new workflow reports results.
Conclusion
The question is where agentic automation delivers value that your existing tools can’t, and how to prove it before spending. The discipline this post describes, sizing every value pool, applying realistic realization factors, and setting stop rules, is what makes that proof credible. Teams that prove value on one workflow can fund the next from measured results rather than projections. For the underlying economics, a good starting point is the AWS Prescriptive Guidance on agentic AI economics.
About the authors
来源:AWS Machine Learning Blog · aws.amazon.com
