ai
3 мин
27 августа 2026 г.
Источник: Dev.to AI Feed

Building an Agent That Actually Handles Failure

Samcorp
Samcorp
RSS AI Ingest
Building an Agent That Actually Handles Failure

Something failed ↓ Retry It looked reasonable. APIs fail. Networks time out. Models occasionally return malformed output. Just retry the operation and continue. Then we started thinking about what happens when the agent performs real acti...

The first version of our AI agent had a surprisingly simple error-handling strategy: Something failed ↓ Retry It looked reasonable. APIs fail. Networks time out. Models occasionally return malformed output. Just retry the operation and continue. Then we started thinking about what happens when the agent performs real actions. Imagine the agent calls: send_email() create_ticket() update_customer() issue_refund() schedule_meeting() It calls issue_refund(). The request times out. What does that mean? A. The refund never happened. B. The refund happened, but the response was lost. C. The refund is still processing. D. The service received the request but failed midway. If the agent automatically retries, option B becomes particularly interesting. We may have just issued the refund twice. That was when AI agent error handling stopped looking like: try: run_agent() except: retry() and started looking much more like distributed systems engineering. A production agent doesn't merely need to recognize failure. It needs to answer a harder question: What state is the world in after that failure, and what is the safest thing to do next? That distinction changed how we designed the entire agent runtime. The Happy-Path Agent Is Easy Most agent demos follow roughly this loop: User Goal ↓ LLM Reasons ↓ Select Tool ↓ Execute Tool ↓ Observe Result ↓ Continue Reasoning ↓ Final Answer For example: User: "Find the customer's unpaid invoice and email them a reminder." Agent: 1. Find customer 2. Query invoices 3. Find unpaid invoice 4. Generate reminder 5. Send email 6. Report success Beautiful. Until step 2 returns 503. Or step 3 returns malformed JSON. Or step 5 succeeds but times out before acknowledging it. Or the model invents an invalid email parameter. Or the agent sends the message and then loses its state. The real architecture is closer to: ┌── Success ──────────────┐ │ ↓ Agent → Tool → Result → Validate → Continue │ ├── Timeout ├── Rate Limit ├── Invalid Input ├── Permission Denied ├── Partial Success ├── Unknown State └── Permanent Failure Those failures should not all produce the same response. That was our first important design decision. Failure Is Data The worst tool interface for an agent looks like this: { "success": false, "message": "Something went wrong" } What exactly should the agent do with that? Retry? Change the parameters? Use another tool? Ask the user? Stop? The error contains almost no actionable information. We moved toward treating failures as structured observations. For example: { "status": "error", "code": "RATE_LIMITED", "retryable": true, "retry_after_ms": 3000, "message": "API request limit exceeded" } Compare that with: { "status": "error", "code": "INVALID_EMAIL", "retryable": false, "field": "email", "message": "Email address is invalid" } Now the runtime can make different decisions. RATE_LIMITED ↓ Wait + Retry INVALID_EMAIL ↓ Correct Input / Ask User This is the foundation of reliable agent recovery: Classify before recovering. Our Error Taxonomy We eventually stopped thinking about “tool errors” as one category. We divided failures into several classes. 1. Transient failures Examples: HTTP 429 HTTP 502 HTTP 503 connection reset temporary DNS failure These may succeed later. Possible strategy: Retry + Exponential Backoff + Jitter 2. Validation failures Example: { "tool": "create_invoice", "arguments": { "customer_id": "C123", "amount": "a lot" } } The problem isn't availability. The parameters are wrong. Retrying the exact request five times won't help. Instead: Validation Error ↓ Return structured feedback ↓ Agent repairs arguments ↓ Validate again ↓ Execute 3. Authorization failures 401 Unauthorized 403 Forbidden These should usually not trigger blind retries. The agent may need to: refresh credentials request permission use another authorized workflow escalate 4. Business-rule failures Suppose the agent calls: refund_order(order_42) and receives: { "code": "REFUND_WINDOW_EXPIRED", "retryable": false } The service is working perfectly. The requested action is simply not allowed. This isn't an infrastructure failure. It's a domain outcome. 5. Ambiguous failures These became the most dangerous category. Request sent ↓ Remote operation starts ↓ Remote operation succeeds ↓ Response is lost ↓ Agent sees timeout From the agent's perspective: UNKNOWN not: FAILED That distinction is critical. A Timeout Does Not Mean Failure Consider: result = payment_api.refund( transaction_id="TX-48291", amount=100 ) Then: TimeoutError The naive implementation does this: retry() But the real timeline might be: Agent Payment API │ │ │──── refund $100 ───────►│ │ │ │ Refund succeeds │ │ │◄──── response ──────────│ │ X network connection lost Agent sees: TIMEOUT The refund already happened. Retrying could create another side effect if the downstream system doesn't deduplicate requests. So we changed our mental model. Timeout ≠ Failure Timeout = Outcome Unknown Now the runtime needs reconciliation. Verify Before Retry For state-changing operations, we want something closer to: Tool Call ↓ Timeout ↓ Did operation actually happen? ↓ ┌───────────────┐ │ │ YES NO │ │ ↓ ↓ Continue Retry Safely Pseudo-code: try: result = execute_tool(action) except TimeoutError: state = verify_action(action) if state == "COMPLETED": return recover_result(action) if state == "NOT_COMPLETED": return retry(action) return escalate_unknown_state(action) That single verification step can prevent a lot of dangerous behavior. Idempotency Changed Everything Verification isn't always possible. Another powerful tool is idempotency. Suppose the agent wants to create an invoice. Instead of: create_invoice( customer="C42", amount=1200 ) we send: create_invoice( customer="C42", amount=1200, idempotency_key="invoice:C42:order:991" ) The server records that key. First request: Key not seen ↓ Create invoice ↓ Store result Retry: Key already exists ↓ Do NOT create another invoice ↓ Return previous result Conceptually: ┌────────────────────┐ Agent ─────────────►│ Idempotency Store │ └─────────┬──────────┘ │ Key already exists? ↙ ↘ YES NO │ │ Return result Execute │ Save result Now retrying the intent doesn't necessarily repeat the effect. This is especially important for tools that: send create charge refund delete publish transfer update Recent research into agent failures under non-atomic tool execution similarly found benefits from combining postcondition verification, verify-before-retry, and idempotency keys rather than assuming tool calls are atomic. Not Every Failure Should Be Retried Once we had structured errors, our retry logic became much simpler. Something like: RETRYABLE_ERRORS = { "RATE_LIMITED", "SERVICE_UNAVAILABLE", "CONNECTION_RESET", "TEMPORARY_FAILURE" } def should_retry(error): return error.code in RETRYABLE_ERRORS And explicitly: NON_RETRYABLE_ERRORS = { "INVALID_ARGUMENT", "PERMISSION_DENIED", "RESOURCE_NOT_FOUND", "BUSINESS_RULE_VIOLATION" } The important idea is: Retry because the condition might change, not because an error occurred. If the email address is invalid, waiting two seconds doesn't make it valid. Retry Budgets Matter Even retryable failures need limits. Otherwise: Tool fails ↓ Agent retries ↓ Tool fails ↓ Agent retries ↓ Tool fails ↓ Agent retries ↓ ... Now one failed request can become: high latency + more tokens + more API calls + higher cost + extra load on an already failing service We introduced retry budgets. MAX_ATTEMPTS = 3 with exponential backoff and jitter. Conceptually: Attempt 1 ↓ fail wait ~1s Attempt 2 ↓ fail wait ~2s Attempt 3 ↓ fail STOP / FALLBACK / ESCALATE The important word is stop. An autonomous agent needs boundaries. Don't Let Multiple Layers Retry This one is easy to miss. Imagine: Agent retries 3 times ↓ Tool wrapper retries 3 times ↓ HTTP client retries 3 times One logical operation can now generate: 3 × 3 × 3 = 27 attempts. Exactly when the dependency is already unhealthy. We decided that every failure path should have a clear retry owner. Agent Runtime ↓ owns retry policy Tool Wrapper ↓ reports structured failure HTTP Client ↓ no hidden application-level retries The exact layer can vary. The principle shouldn't: Know who owns the retry. Tool Schemas Are Part of Error Handling An agent should not discover invalid inputs by crashing production APIs. Tool inputs should be validated before execution. For example: class RefundRequest(BaseModel): order_id: str amount: float reason: str Then: LLM proposes tool call ↓ Schema validation ↙ ↘ INVALID VALID ↓ ↓ Return error Execute to agent tool If the model generates: { "order_id": 42, "amount": -500, "reason": null } we should catch it before reaching the payment system. This is why tool design matters so much in agent architecture. The LLM is not a trusted caller. Treat it like an external client. Give the Agent Errors It Can Reason About This: Exception: 0x8004FA32 is useful to almost nobody. This is much better: { "code": "CUSTOMER_NOT_FOUND", "retryable": false, "message": "No customer exists with ID C-984.", "suggested_actions": [ "search_customer_by_email", "ask_user_to_confirm_customer" ] } Now recovery becomes part of the agent loop. Action ↓ Failure ↓ Structured Observation ↓ Reason About Failure ↓ Choose Recovery This is one of the important differences between traditional exception handling and AI agent error handling. The error isn't only for developers. It may also become an observation consumed by the model. But Don't Let the LLM Decide Everything There is a tempting architecture: Anything fails ↓ Tell LLM ↓ Let LLM decide We don't want that for every failure. Some policies should be deterministic. For example: Maximum retry attempts Payment limits Permission checks Timeout budgets Allowed tools Rate limits Escalation thresholds Those belong in code. Not in a prompt. Our rule became: LLM decides: "What should I try next?" Runtime decides: "Am I allowed to try it?" That separation is extremely useful. Circuit Breakers Protect the Agent Too Suppose an external CRM API begins returning: 503 503 503 503 503 Without protection, every active agent continues hitting it. Instead: Failures exceed threshold ↓ Open circuit ↓ Stop sending requests ↓ Use fallback / wait / escalate ↓ Probe later ↓ Close when healthy Conceptually: CLOSED │ │ failures exceed threshold ↓ OPEN │ │ cooldown ↓ HALF-OPEN │ ├── success → CLOSED │ └── failure → OPEN Circuit breakers aren't uniquely “AI.” That's exactly the point. Agents interact with distributed systems. They inherit distributed-system problems. Partial Failure Is Harder Than Total Failure Imagine an agent performing: 1. Create customer 2. Create invoice 3. Charge payment 4. Email receipt Step 1 succeeds. Step 2 succeeds. Step 3 fails. What does “retry the workflow” mean? If we restart at step 1: duplicate customer? duplicate invoice? So workflow state must be durable. Something like: { "workflow_id": "wf_123", "steps": { "create_customer": "completed", "create_invoice": "completed", "charge_payment": "failed", "send_receipt": "pending" } } Now recovery can resume from the correct point. Workflow Restart ↓ Load State ↓ Skip Completed Steps ↓ Resume From Failure This is much safer than asking the model to reconstruct everything from conversation history. Sometimes You Need Compensation Not every workflow can simply resume. Imagine: Reserve inventory ↓ Charge customer ↓ Create shipment Inventory reservation succeeds. Payment succeeds. Shipment creation permanently fails. The system may need compensating actions: Release inventory Refund payment So instead of pretending the entire workflow is one transaction: Step A ↓ Step B ↓ Step C fails ↓ Compensate B ↓ Compensate A This resembles Saga-style recovery in distributed systems. The important part is that compensation is explicit. The LLM should not invent rollback logic on the fly for financially or operationally important actions. Give Every Workflow a State Machine Agent loops can become difficult to reason about when state exists only inside messages. We prefer explicit workflow states. PLANNING ↓ EXECUTING ↓ WAITING_FOR_TOOL ↓ VERIFYING ↓ RECOVERING ↓ COMPLETED with terminal states such as: FAILED CANCELLED ESCALATED Then transitions are controlled. For example: if state == "WAITING_FOR_TOOL": if tool_result.success: transition("EXECUTING") elif tool_result.outcome_unknown: transition("VERIFYING") elif tool_result.retryable: transition("RECOVERING") else: transition("ESCALATED") This makes the system much easier to debug. Failure Handling Needs Memory Suppose the agent already tried: search_customer("John Smith") three times. Every attempt failed because the CRM was unavailable. If the agent doesn't retain execution state, it may reason: Maybe I should search the customer. Again. Useful agent memory isn't only conversation memory. It includes operational state: tools attempted arguments used results errors retry count side effects verification results remaining budget Think of it as: Conversation Memory + Execution Memory + World State Those are different things. Put a Budget Around the Entire Agent Retry limits aren't enough. Agents can fail by looping without technically repeating the same operation. For example: Search ↓ Reason ↓ Different Search ↓ Reason ↓ Another Search ↓ Reason ↓ ... So we add global budgets. AgentBudget( max_steps=15, max_tool_calls=10, max_retries=3, max_tokens=20_000, timeout_seconds=90 ) Then: Budget available? ↙ ↘ YES NO ↓ ↓ Continue Stop safely Autonomy without a budget is just an unbounded loop with API credentials. Fallbacks Should Be Designed Before Failure Suppose the preferred search service is unavailable. Possible fallback: Primary Search ↓ failure Secondary Search Model unavailable? Primary Model ↓ Fallback Model Real-time data unavailable? Live API ↓ Cached Data ↓ Tell user data may be stale But fallback behavior needs semantic awareness. If a payment API fails, “try another payment API” may not make sense. Fallbacks aren't interchangeable dependencies. They are part of product behavior. Human Escalation Is a Feature There's a tendency to view human intervention as agent failure. We don't. Sometimes the safest agent behavior is: I cannot determine whether this transaction completed successfully. I've stopped further actions and escalated the workflow for review. That is a successful safety outcome. We define escalation conditions such as: unknown state after side effect retry budget exhausted conflicting tool results insufficient permissions high-value transaction low-confidence irreversible action policy ambiguity This creates an important boundary: Agent Autonomy ↓ Safe Operating Envelope ↓ Human Review Agentic AI introduces precisely these broader questions around stability, reliability, planning failures, and human control. A useful overview of those architectural concerns is this discussion of Agentic AI fundamentals. The objective isn't maximum autonomy. It's appropriate autonomy. Observability: Can You Reconstruct the Failure? When somebody says: “The agent did something weird yesterday.” you need more than application logs. For every execution, we want a trace resembling: trace_id: agent_8f219 User Goal ↓ Model Decision ↓ Tool Selected ↓ Arguments ↓ Validation ↓ Tool Request ↓ Tool Response ↓ State Change ↓ Retry Decision ↓ Verification ↓ Next Action Useful fields include: { "trace_id": "...", "workflow_id": "...", "step": 7, "model": "...", "prompt_version": "...", "tool": "create_invoice", "arguments_hash": "...", "idempotency_key": "...", "attempt": 2, "latency_ms": 1842, "result": "timeout", "verification": "completed", "next_action": "continue" } Be careful about storing raw arguments when they contain secrets or personal information. Observability shouldn't become a data leak. We Started Measuring Recovery, Not Just Success A basic agent dashboard might show: Task Success Rate: 94% Useful. But incomplete. We also want: Tool failure rate Retry rate Retry success rate Unknown-outcome rate Duplicate-action rate Recovery success rate Escalation rate Average steps/task Cost/task Timeout rate Circuit-breaker activations Then we can distinguish: Agent completed because everything worked from: Agent completed because recovery worked The second one tells us much more about production reliability. Our Production Architecture Became Boring And that's a compliment. The first architecture looked like: User ↓ LLM ↓ Tools ↓ Answer The production version looked more like: User Goal ↓ ┌───────────────┐ │ Agent Planner │ └───────┬───────┘ ↓ Policy / Budget ↓ Tool Selection ↓ Input Validation ↓ Idempotency Layer ↓ Tool Call ↓ ┌───────────────┐ │ Result Class. │ └───────┬───────┘ ↓ ┌──────────────┼──────────────┐ │ │ │ SUCCESS RETRYABLE UNKNOWN │ │ │ ↓ ↓ ↓ Continue Backoff Verify State │ │ ↓ ↓ Retry ┌─────┴─────┐ │ │ Complete Unknown │ │ ↓ ↓ Continue Escalate + Circuit Breakers + Durable State + Compensation + Observability + Human Review Most of that architecture isn't “AI magic.” It's software engineering. That's exactly why it works. The Error-Handling Hierarchy If I were building another production agent, I'd think about recovery in this order. 1. Prevent Catch errors before execution. schemas permissions input validation policy checks preconditions 2. Detect Know exactly what happened. structured errors timeouts postconditions health signals 3. Classify Determine the failure type. transient validation authorization business rule unknown outcome permanent 4. Recover Choose a bounded strategy. repair retry fallback resume compensate 5. Verify Confirm the expected world state. Did the email send? Did the payment settle? Did the record exist? 6. Escalate When certainty disappears: stop preserve state explain request human review 7. Learn Turn failures into regression tests. Every meaningful production failure should eventually become: incident ↓ test case ↓ evaluation ↓ release gate A Practical Agent Failure Test Suite Happy-path testing is nowhere near enough. Inject failures deliberately. Tool failures 429 500 502 503 timeout connection reset Model failures malformed tool arguments wrong tool missing required field unexpected text instead of JSON State failures tool succeeds but response disappears partial database write delayed consistency duplicate request Workflow failures step 3 of 7 fails dependency unavailable budget exhausted human approval never arrives Then ask: Did we duplicate a side effect? Did the workflow resume correctly? Did retries stop? Was state preserved? Was the user told the truth? Could an engineer reconstruct what happened? If those aren't part of the test suite, we haven't really tested agent reliability. Reliability Is More Important Than Clever Recovery It is easy to make the recovery loop increasingly sophisticated: fail ↓ reflect ↓ replan ↓ self-correct ↓ try another strategy Sometimes that's useful. But we shouldn't confuse cleverness with reliability. For important operations, I would rather have: 3 predictable recovery paths than: unlimited autonomous improvisation Enterprise AI systems also need the surrounding engineering disciplines—system integration, monitoring, governance, security controls, and post-deployment optimization—not just a capable model. Those broader considerations are covered in this overview of AI consulting and production AI architecture. The model can reason. The runtime should enforce boundaries. The Most Important Question Our original question was: How do we make the agent succeed? The more useful question became: How do we make the agent fail safely? Because production systems will fail. Models fail. APIs fail. Networks fail. Credentials expire. Schemas change. Databases become unavailable. Users provide ambiguous instructions. External systems return contradictory states. No architecture eliminates all of that. Good AI agent error handling assumes failure is normal and gives the system a controlled way through it. Final Takeaway A production-ready AI agent needs more than: LLM + Tools + Prompt It needs: LLM + Tools + Validation + Structured Errors + Timeouts + Safe Retries + Idempotency + Verification + Durable State + Circuit Breakers + Fallbacks + Compensation + Budgets + Observability + Human Escalation The most reliable agent isn't the one that never encounters an error. That agent doesn't exist. The reliable agent is the one that knows: what failed, whether anything changed, whether retrying is safe, when to recover, and when to stop. That's when an AI agent stops being an impressive demo and starts behaving like production software.

Хотите внедрить ИИ в ваш бренд?

Спроектируем и развернем автономных агентов и современный цифровой стек под ваши задачи.

Рассчитать проект