# Decision Points in AI Agent Development
# Self-Correction Retry
🎯 The Hook
When your LLM output breaks, are you just blindly resending the same prompt? Self-correction retry feeds back *what went wrong* so the model can fix itself. But here's the catch: beyond 3 retries, improvement almost never happens. Knowing when to stop is as important as knowing when to retry.
📋 Overview
Self-correction retry controls how many times you re-prompt an LLM after its output violates a schema or falls short on quality, injecting error details into context each time. This is fundamentally different from network retries (resending the identical request). By providing specific feedback about what went wrong, you give the model a real chance to produce correct output on the next attempt. LLM outputs are probabilistic -- missing JSON brackets, out-of-range values, and incomplete responses happen routinely. Self-correction retry is one of the most practical strategies for handling these errors.
🔍 Decision Points
The primary driver is failure_cost: how much damage does a bad output cause downstream? Applying the same retry count to all errors is wasteful, so differentiate by error type.
Syntax-level errors (malformed JSON, type mismatches) almost always resolve with a single feedback round. Semantic-level errors (out-of-range values, nonexistent ID references) may improve with feedback, but if the second attempt fails, it's a structural problem. Quality-level errors (incomplete answers, missing information) are subjective and rarely improve through retries -- invest in prompt engineering instead.
💡 Key Details
Reference values to keep in mind 📊
- General case: 1-3 retries. If no improvement after 2, likely a structural issue
- Structured output (JSON Schema): 1-2 retries. Using response_format yields high first-attempt success rates
- High failure_cost domains (financial, legal, medical): 2-3 retries. Escalate to humans if quality plateaus
- Low failure_cost domains: 0-1 retries. If fallbacks exist, fail fast for efficiency
Error messages should be short and specific 🎯 Not "output is invalid" but "the delivery_date field is a past date; please specify a future date." Spell out what's wrong and what's expected. Keep error messages under ~200 tokens since they consume context budget.
⚖️ Trade-offs
Too few retries means discarding fixable errors. Throwing away output that's only missing a closing bracket? That's wasteful.
Too many retries and costs explode ⚡ If one LLM call takes 30 seconds, 5 retries means 2.5+ minutes of waiting. Each retry appends error messages to the context, causing cumulative token growth. Worst case, you hit context length limits and trigger an entirely different failure mode.
If the same error type appears twice in a row, question the prompt before attempting a third retry. Repeated identical errors signal that the LLM cannot produce correct output with the current prompt-schema combination.
🛠️ Use Cases
JSON schema violations: Provide specific error feedback; typically fixed in 1 retry. Using Structured Outputs (response_format) eliminates most syntax-level retries entirely.
Business rule violations (e.g., delivery date in the past): Include concrete constraints and current state in feedback. If 2 retries don't help, pivot to prompt redesign.
Quality shortfalls ("analyze from 5 perspectives" returns only 3): Try once; if no improvement, accept partial results or split the prompt into smaller generation tasks.
Persistent identical errors: Consider lowering temperature, simplifying the prompt, trying a different model, or escalating to a human operator 🔄
#
AIAgents# #
SoftwareArchitecture#