Skip to main content
Workflow Automation

The Messy Truth About Layering AI Agents on Top of RPA

We've wired agents into RPA for claims processing, AP, and support triage. Here's what actually works, what breaks, and why we never let an agent click a button.

1.96%.

That's the score Claude 2 got on SWE-bench when it came out. Real software issues, real codebases. It solved about two out of every hundred. If a top model can't reliably patch a GitHub issue, why would anyone hand it the keys to their accounts payable ledger?

That number gets thrown around as a reason to avoid agents. Wrong conclusion. The right one: use agents for thinking, robots for clicking. We've been building this way for about 18 months now. Here's the unvarnished version.

Step 1: Sit with the people doing the work. Label everything.

We don't start with a whiteboard. We start with a chair next to the person who actually processes the invoices. Watch them for an hour. Write down every step. Then label each one.

Deterministic: copy this field, click that button, rename this file. Variable: decide if this invoice matches the PO, classify this email, route this exception. RPA handles the deterministic stuff well—high volume, repetitive, rule-based. Agents handle the variable stuff—judgment calls, exceptions, weird formats.

Skip this step and you'll buy the wrong tool for half your process. We've seen teams spend six figures on agent platforms only to realize 80% of their workflow was deterministic data entry.

Step 2: RPA where the UI is. Agents where the thinking is.

RPA is brittle. Change a button label and your bot breaks. Agents are unpredictable. Ask one to reconcile an account and it might hallucinate a transaction. So we don't ask agents to click through a legacy green screen. We ask them to decide, then hand off to a bot.

UiPath calls this "agentic automation." Fancy term. What it means: the agent calls a tool—often via MCP—and that tool triggers an RPA robot. MCP is an open standard for connecting AI apps to external systems. The docs call it a USB-C port for AI. That's the integration layer we standardize on.

One caveat: MCP is still young. We've hit edge cases where the tool call times out and the agent retries three times, triggering the same bot three times. Put idempotency keys on your RPA triggers. Trust us.

Step 3: Pick one orchestration pattern. Just one.

We see five patterns that don't fall apart: sequential chains, parallel swarms, human-in-the-loop checkpoints, event-triggered agents, and self-correcting loops. For a first project, pick one. We usually start with a sequential chain plus a human checkpoint.

IBM's legal research example is instructive: a multi-agent assistant routed queries through a low-cost classifier first, escalated only complex cases, and cut contract review from 90 to 45 minutes. That's a classifier plus a human escalation path. Not a swarm of 12 agents.

We tried a self-correcting loop on a customer support triage bot. It worked, but only after we added a circuit breaker that kicked in after two failed retries. Without it, the agent would loop forever, burning tokens and annoying customers. Now we ship with circuit breakers by default.

Step 4: Compare build options before you commit.

We evaluate four options. The table below is what we use in workshops. Scores are our judgment, not vendor claims.

OptionBest forSpeed to first valueMain riskOur recommendation
RPA onlyStable, rule-based, high-volume tasksWeeksBreaks when UI or rules changeGood starting point, but cap it
Agentic AI onlyJudgment-heavy, low-volume tasksWeeks to monthsUnpredictable, hard to auditRarely production-ready alone
Hybrid (agents + RPA)End-to-end processes with both types of stepsMonthsIntegration and orchestration debtOur default for serious workflows
Intelligent automation (AI + BPM + RPA)Regulated, cross-department processesMonths to a yearHeavy governance overheadRight for compliance-heavy work

If you're in banking, insurance, or healthcare, the fourth option is often mandatory. Intelligent automation combines AI, BPM, and RPA, and it helps prove consistent compliance. But don't buy the full stack on day one.

Step 5: Instrument, log, and set a human checkpoint.

We treat every agent action as a potential audit item. RPA already enforces process consistency and provides audit trails, so we extend that logging to agent decisions. We also define a human-in-the-loop checkpoint for any action that moves money, changes a customer record, or triggers a regulatory report.

AI agents are autonomous in decision-making but still need goals and predefined rules from humans. That's not a weakness; it's the control surface.

One concrete example: we built an AP automation flow where the agent decides if an invoice matches the PO. If confidence is below 0.85, it routes to a human. If above, it hands off to an RPA bot that enters the invoice into the ERP. The bot logs every field it touches. The agent logs its reasoning. Auditors love it.

What can go wrong

The biggest failure mode isn't model accuracy. It's organizational. Organizational change and human oversight, not technology, are the biggest barriers to AI automation adoption. We've seen teams buy agent licenses, skip the process mapping, and then blame the model when the bot clicks the wrong button.

The second failure mode is over-trusting benchmarks. SWE-bench's 1.96% result is a reminder that even strong models fail on real code. Your workflow is harder than a benchmark. Budget for exceptions, and staff a triage queue.

Third: integration debt. Every new agent you add creates another API call, another failure point, another thing to monitor. We've seen hybrid setups where the orchestration layer became so complex that debugging a simple invoice processing failure took two days. Keep it simple.

Our recommendation

Start with RPA for the deterministic 70% of your process. Add one agent for the variable 30%. Connect them with MCP tools. Put a human checkpoint on anything that touches money or compliance. Then measure.

If you need a single rule: don't let an agent click a button it can't explain, and don't let an RPA bot make a decision it wasn't programmed to make.

The market is moving fast—the global RPA market was estimated at $4.68 billion in 2025 and is projected to reach $35.84 billion by 2033—but the teams that win are the ones that keep the execution layer boring and the reasoning layer supervised.

Sources

  • Robotic process automation (Wikipedia) - https://en.wikipedia.org/wiki/Robotic_process_automation
  • UiPath (RPA) - https://www.uipath.com/rpa/robotic-process-automation
  • IBM (AI agents) - https://www.ibm.com/think/topics/ai-agents
  • Model Context Protocol (official docs) - https://modelcontextprotocol.io/introduction
  • SWE-bench paper (arXiv, ICLR) - https://arxiv.org/abs/2310.06770
  • Grand View Research (RPA market) - https://www.grandviewresearch.com/industry-analysis/robotic-process-automation-rpa-market

Share this article:

Comments (0)

No comments yet. Be the first to comment!