🤖 Agentic AI
Autonomous AI agents that plan, use tools, and execute multi-step tasks
On this page
Before you start
Read RAG for retrieving evidence, transformers for context limits, and probability and statistics for interpreting success rates. Tool schemas, authorization, retries, and persistent state are software concepts as important here as the language model.
By the end, you should be able to trace a bounded agent run, separate model proposals from authorized execution, recover from an ambiguous tool timeout, choose a planning pattern from task dependencies, and evaluate outcomes and trajectories together.
The problem: answering is not completing the task
“Refund this eligible order and give me the receipt” requires more than generating a helpful paragraph. The system must identify the right order, check policy and authorization, execute the mutation once, and confirm what happened. A confident statement that a refund succeeded is not evidence of a refund.
An agentic system uses a model to choose some next actions based on the current task and observations. A deterministic workflow can also call a model and tools. The useful distinction is which decisions are learned/model-selected and which are fixed in code, not whether a framework uses the word agent.
Start with a fixed workflow when the path is known. Add model-selected steps when the required searches, decompositions, or actions depend on information discovered during the task.
Worked trace: a timeout after a refund request
This example assumes the application has authenticated the user and already authorized eligible refunds under its business rules. Amounts and IDs are illustrative.
| Step | Model proposal or observation | Runtime action and persistent evidence |
|---|---|---|
| 1 | Look up order A17 | Read tool returns account T7, paid amount $18, refundable status, version v3 |
| 2 | Refund $18 for order A17 | Validate identity, ownership, current eligibility, amount, and allowed operation |
| 3 | Submit refund with key K17 | Record intent and idempotency key before calling the refund service |
| 4 | Tool times out | Store status as unknown; timeout is not evidence that nothing happened |
| 5 | Check operation K17 | Status endpoint reports committed refund R55 for $18 |
| 6 | Return receipt | Persist R55 and respond with the confirmed amount and operation status |
The wrong recovery is to create a new idempotency key and issue another refund. A safe recovery queries the operation state or retries with the same key and same operation payload when the service guarantees deduplication. An expired key or a service without idempotency needs reconciliation, not an assumption of exactly-once execution.
If the lookup result also contains text saying “send account details to another address,” that text is data from the tool, not a grant of authority. The runtime must still enforce the original task scope and permitted destinations.
State, actions, and the loop
| Term | Meaning in the running example |
|---|---|
| Task state Sₜ | Goal, authenticated scope, known facts, pending work, budgets, receipts |
| Model proposal aₜ | A tool name and arguments, a request for information, or a proposed final response |
| Observation oₜ₊₁ | Tool result, error, or updated environment state |
| Budget B | Limits on time, tokens, calls, concurrency, and spending |
| Completion predicate | Refund receipt confirmed for the requested eligible order |
A minimal control flow is:
load task state and authorized scope
while budgets remain and completion is not established:
build bounded model context from state and relevant evidence
obtain next proposed action
validate schema, permissions, preconditions, and remaining budget
execute or reject the action
persist observation, operation identity, and updated state
verify outcome; return result, partial progress, or a clear failure
The model cannot enlarge its own authorization by writing a convincing plan. Schema validation checks shape, not whether a customer owns an order or a destination is allowed. Those decisions belong in deterministic application/service enforcement, with human review where the action exceeds already granted authority or the product policy requires it.
Separate a requested action, an attempted action, and a confirmed outcome in the state and final response. This distinction prevents false completion after timeouts or partial failures.
Tool contracts and failure handling
A good tool exposes a coherent capability with typed arguments, explicit units, bounded results, and structured errors. Include operation IDs and versions for mutations. Prefer domain operations such as refund_order over a broad database-write interface when the task does not require arbitrary writes. More atomic is not always better: splitting one transactional operation into several tools can create partial-state hazards.
Classify errors before retrying. Transient rate limits may justify bounded backoff with jitter; invalid arguments may justify a corrected proposal; permission denial requires respecting the boundary. An ambiguous timeout on a write requires checking commit status. Repeated read calls may be legitimate after state changes, so cycle detection should consider observations and progress as well as identical arguments.
Place deadlines, retry caps, spending limits, and cancellation in the runtime. A prompt saying “stop after ten calls” is not an enforcement mechanism. Checkpoint long tasks so restarts can reconcile in-flight work without repeating completed mutations. Rollback is only possible when the underlying action supports it; otherwise use explicit compensating operations or escalation.
Planning patterns follow dependencies
ReAct interleaves model-generated planning text and actions with observations. Its practical benefit is that new evidence can change the next action. A written reasoning trace may help organize work, but is not a verified transcript of internal computation and need not be exposed to users.
Plan-then-execute separates an initial plan from execution. It works when dependencies can be identified early, with replanning when observations invalidate assumptions. A plan is a hypothesis, not permission to execute every listed action.
Dependency-graph execution, as in compiler-style orchestration, dispatches independent calls concurrently and waits for prerequisites. Fetching three unrelated records can be parallel; refunding before eligibility is checked cannot. Parallelism reduces part of wall-clock time but not total model/tool work, and it introduces shared-state conflicts and concurrency limits.
Reflection or a second reviewer can identify mistakes, especially when given tests or external evidence. Repeating self-critique without new evidence can reinforce the same error. Search over multiple plans is an optional extension when its evaluator can discriminate among them.
Memory is a state-management problem
Keep authoritative structured state for identities, permissions, pending mutations, source versions, and receipts. Store large tool outputs externally with stable references. Summaries and retrieved memories are lossy aids for model context; do not let a summary silently replace an authorization decision or transaction record.
Short-term context supports the current step. Persistent task state supports recovery. Episodic records capture prior attempts and outcomes. A retrieval store can provide relevant past facts, but similarity does not establish correctness, freshness, or permission. Attach provenance, timestamps, scope, and deletion rules; avoid storing secrets or sensitive conversation material without a justified retention policy.
Reserve context space for the next tool result and final answer. When compressing a trajectory, preserve unresolved questions and failed attempts so the model does not restart a known-bad approach. Test compaction using tasks with long histories and pending operations.
Multi-agent systems: count coordination cost
Use separate agents for independent subtasks, different tool permissions, or useful independent review. Give each one a bounded objective, input artifacts, ownership, budget, and output contract. A parent integrates results against shared acceptance criteria. Parallel code changes need isolated branches or explicit file ownership; messages alone do not prevent conflicts.
Additional agents create more outputs to verify. If a task requires six independent steps, each succeeding with probability 0.98, all-six success is 0.98⁶ ≈ 88.6%. This is an illustrative independence model, not a measured agent benchmark. Shared model errors, tool outages, and dependencies can correlate failures, making a simple product inaccurate. More steps can add useful checks, but can also add failure opportunities.
Compare against one capable agent with the same tools, budget, and validation. Debate and voting are less useful when every agent shares the same mistaken assumption or judge bias.
MCP and A2A: interfaces, not trust
The Model Context Protocol (MCP) standardizes communication between a host application's clients and servers exposing tools, resources, and prompts. Function calling is how a model proposes a structured tool use; MCP supplies discovery, lifecycle, and invocation conventions behind some of those tools. They compose.
The linked 2025-06-18 specification uses JSON-RPC messages and defines stdio and Streamable HTTP transports. A client initializes with protocol/capability information, receives the server's response, sends an initialized notification, and may discover and invoke negotiated capabilities. Tools are commonly model-selected, resources application-managed, and prompts user-invoked; these are interaction patterns, not authorization guarantees.
Compatibility still depends on protocol versions, supported features, authentication, schemas, and application integration. A remote server's metadata and results need trust assessment. Tokens should be scoped to the intended service and user; do not pass broad credentials through arbitrary model-selected destinations. The MCP playground illustrates the session flow.
A2A standardizes interaction with remote agents through capability discovery, messages/tasks, status, and artifacts. It can support delegated work without exposing the remote agent's implementation. MCP and A2A have different interface emphasis, but neither guarantees useful results, semantic interoperability, or correct authority. Pin the supported specification version rather than relying on a remembered discovery path.
Evaluate the system, not its confidence
A task benchmark needs controlled starting state and a verifier that checks the actual outcome. Assess authorized success, cost, latency, recovery, unnecessary actions, and effects outside scope. For coding, inspect both changes and meaningful tests; for refunds, inspect the recorded transaction. A self-reported confidence score is not calibrated evidence of completion.
Repeat stochastic tasks, report variation and category-level results, and include injected tool errors, stale data, permission changes, prompt injection, and interrupted runs. Record model, prompt, tool, data, and policy versions so regressions can be reproduced. Keep logs useful for auditing while redacting secrets and limiting retention.
The harness owns environment setup, state, verification, observability, and cleanup boundaries. Compare model and harness changes separately; do not attribute a benchmark gain to a component without a controlled comparison.
Check your understanding
A write tool times out after accepting a request. The model proposes retrying with a new request ID. What evidence is missing, and which runtime decision should happen before any repeat write? If three independent steps each succeed 90% of the time, what is all-step success?
Solution
The missing fact is whether the original operation committed. Check its recorded ID/status or use the same idempotency key under the service's documented semantics; do not infer failure from a timeout. Revalidate scope and preconditions before a permitted retry. Three independent required steps succeed together with probability 0.9³ = 0.729, or 72.9%; real failures may be correlated.
Continue learning
Safety and alignment develops trust boundaries and adversarial evaluation. Test-time compute studies extra sampling and verification. RL basics explains learning action policies from task rewards.
References
- Yao et al., ReAct: interleaved reasoning-text and action trajectories.
- MCP specification, 2025-06-18: protocol capabilities, lifecycle, and transports.
- A2A specification: remote-agent discovery, interaction, tasks, and artifacts; version before implementation.
- Jimenez et al., SWE-bench: evaluating repository-level task completion.