Back to blog
AI Agents

Planning and Execution in AI Agents: Decomposition, Failure and Recovery

A model proposes steps; the application controls every transition

Dependencies, permissions, receipts and bounded replanning

Original PLAN-STEP-7 cycle and failure routes
How an agent plans and executes one verified step at a time
Primary nodeBounded step execution
Routing modePLAN-STEP-7
StatusPUBLISHED
A bounded plan passes through validation, one-step execution, verification and recovery routes
PLAN_STEP_7_V01: route a model proposal through a policy gate to a verified result.
TERMINAL_PREVIEW.LOG
$ plan --contract PLAN-STEP-7
> bind: owner / scope / accepted outcome
> propose: finite steps / dependencies
> execute: one authorized step
> verify: read-back / receipt
> route: next / replan / review
Planning and execution with explicit boundaries

A planning execution AI agent turns a goal into bounded steps, runs one authorized step at a time, and checks the observed result. A plausible list from a model is not an executable plan. Each step needs inputs, dependencies, permission, an acceptance condition and a failure route. PLAN-STEP-7 below is an original synthetic design example, not a customer benchmark.

If you are deciding whether an agent is needed at all, start with task-first selection criteria. Broad local implementation questions belong on the AI specialist in Armenia page. This guide addresses the narrower engineering question: how a plan remains controlled while the world changes.

Problem and requirements

Imagine reconciling a supplier request against a CRM record and preparing a correction. A model may propose “find record → compare address → update CRM → report completion.” The record might belong to another tenant, have changed after the read, or lack a trustworthy source. A proposed plan therefore cannot authorize a write.

Before execution, bind the task owner, accepted outcome, permitted sources, action scope, step and time budgets, and reviewer for exceptions. Give every step an observable acceptance condition. If you cannot observe whether a step succeeded, clarify the task before running it.

Architecture: the PLAN-STEP-7 cycle

The original PLAN-STEP-7 diagram and contract use seven boundaries: bind → propose → validate → execute → verify → replan → close. An orchestrator stores a versioned plan and each step's state. The model proposes a sequence; the server checks policy and calls tools. A verifier compares each result with the target system. Replanning can use new evidence, but cannot silently enlarge scope or reset budgets.

BoundaryInputTransition condition
BindGoal, owner, scope, acceptanceIdentity and rights are known
ProposeGoal and allowed read toolsFinite steps and dependencies exist
ValidateStep, arguments, plan versionSchema, tenant, policy and budget pass
ExecuteAuthorized callTimeout and idempotency key assigned
VerifyReceipt and read-backExpected state is observed
ReplanFailure or changed factConstraints and reason are retained
CloseVerified stepsAccepted outcome or human review

This builds on the tool calling execution contract. Tool calling controls one function invocation; planning controls dependencies and decisions across invocations. If the steps are stable and known, deterministic automation may be simpler.

Components and state contract

Planner. Return a structured plan: step ID, dependencies, action type, expected observation and stop condition. Tool output and external documents are data, never authority to add tools or permissions.

Executor. Accept one ready step and check user, tenant, object version, policy and budget immediately before the call. Authorization at planning time can be stale by execution time.

State store. Persist task ID, plan version, step status, call ID, idempotency key, receipt and reason code. Recovery must know which side effects already happened. A new plan version must not erase them.

Verifier. Compare the destination state to the declared acceptance condition. HTTP 200 and model prose do not prove a business outcome. For writes, use read-back or equivalent destination confirmation. Reconcile an unknown outcome before any retry.

Minimal pseudocode

ts
type Step = { id: string; dependsOn: string[]; action: string; expected: string };
async function runStep(task: Task, step: Step) {
  require(task.owner && task.planVersion && task.budget.remaining > 0);
  require(dependenciesVerified(step, task.receipts));
  const tool = registry.get(step.action);
  if (!tool || !policy.allows(task.user, task.tenant, tool)) return hold("forbidden");
  if (tool.isWrite && !task.approvalFor(step.id)) return hold("review_required");
  const key = `${task.id}:${task.planVersion}:${step.id}`;
  const result = await executeOnce(tool, key, task.scope);
  if (result.unknown) return reconcileBeforeRetry(key);
  const verified = await verifyExpected(step.expected, result);
  await receipts.save({ key, verified, planVersion: task.planVersion });
  return verified ? "next_step" : "replan_or_review";
}

This is a contract sketch, not a complete SDK. A real system also needs atomic state updates, limits on replanning, cancellation, partial effect handling and race tests. Previously completed effects remain facts when the plan changes.

Failure modes

  1. A plausible plan depends on a missing fact. Validate the input before its dependent step.
  2. A dependency is skipped. Block execution until predecessor receipts are verified.
  3. Permissions change after planning. Check policy before each call.
  4. A write times out. Treat the outcome as unknown; reconcile by key and read back before retry.
  5. The plan loops. Bound replans, cost and duration, then route to a human.
  6. Replanning expands scope. New actions, data sources or tenants need new authorization.
  7. The model declares success without proof. Close only against observable criteria and receipts.

Testing and production checklist

Build fixtures for normal completion, empty source, stale CRM version, another tenant, policy denial, write conflict, timeout before and after a side effect, duplicate step ID, verification failure and exhausted budget. Each fixture should specify the expected route, receipt and forbidden side effect. Test planner structure, executor permissions and verifier observations separately, then test crash recovery for the whole loop.

Before a bounded pilot, assign the review owner, pin schema and policy versions, define stop and recovery paths, set retention and redaction rules, and observe verified completion, replans, denials and unknown outcomes. Measure those on your own system; none are claimed here. Start with read-only tasks and assess writes separately.

For a concrete process review, contact us with a task example, system map, action permissions and one failure case. Prompt engineering can help define input and output contracts.

CODE_BLOCK.TXT
require(step.dependenciesVerified && policy.allows(step));
if (step.isWrite && !approval.valid) return hold;
result = await executeOnce(step.idempotencyKey);
if (result.unknown) return reconcileBeforeRetry;
return verifyExpected(step.expected, result);