Back to blog
AI Agents

State and Memory in AI Agents: What to Keep Between Steps

Checkpoints record progress; memory supplies scoped evidence

Freshness, permissions, receipts and safe resumption

Original STATE-LEDGER-7 cycle and failure routes
What an agent should retain between steps and sessions
Primary nodeVersioned agent state
Routing modeSTATE-LEDGER-7
StatusPUBLISHED
Agent runtime connects versioned task state, scoped memory and execution receipts
STATE_LEDGER_7_V01: route checkpoint and scoped memory through policy to a verified result.
TERMINAL_PREVIEW.LOG
$ resume --contract STATE-LEDGER-7
> load: versioned checkpoint / owner
> retrieve: scoped and fresh memory
> gate: policy / current source
> act: idempotent call / receipt
> route: verified / reconcile / review
State and memory with explicit lifetimes

State memory AI agent design is a question of which facts must survive a step, a retry, or a new session. A model's context window is neither a transaction log nor an authority for future actions. This guide uses an original synthetic contract, STATE-LEDGER-7, to separate task state, durable memory and audit receipts. It does not report client results or benchmark scores.

For the broader choice of an AI implementation partner in Armenia, see the AI specialist page. This article addresses the narrower engineering question of storage and retrieval between agent steps.

Problem and requirements

Imagine an agent checking a supplier request against a CRM record, preparing a proposed update, then waiting for review. It must resume after a process restart without repeating a write or mistaking an old CRM value for current evidence. It also should not carry the whole conversation, private identifiers or obsolete instructions into every future task.

Define the task owner, tenant, accepted outcome, permitted tools, retention policy and review boundary first. Decide which facts need a short lifetime, which can be reused across tasks, and which must be retained solely as an inspectable receipt. A summary generated by the model is useful context, but its claims require provenance and expiry.

Architecture: STATE-LEDGER-7

The original STATE-LEDGER-7 contract has seven gates: bind → checkpoint → classify → retrieve → validate → act → reconcile. It is a design example, not a deployed system. A versioned checkpoint records the current task and step. A selective memory index offers prior facts as untrusted evidence. An append-only receipt records tool calls and verified outcomes. The application, not retrieved text, determines permissions.

GateStored inputRequired check
BindTask ID, owner, tenant, scopeIdentity and allowed outcome are explicit
CheckpointStep, plan version, object versionAtomic version comparison prevents stale resume
ClassifyCandidate fact and sensitivityRetention, purpose and deletion rule are assigned
RetrieveScoped query and candidate referencesTenant, provenance, freshness and permission are checked
ValidateEvidence plus current sourceContradictions and stale values are surfaced
ActAuthorized tool callIdempotency key and current policy are checked
ReconcileReceipt and destination read-backUnknown effects are resolved before retry

The checkpoint belongs to the current run; durable memory contains deliberately reusable facts; the receipt supports recovery and audit. Do not collapse the three into a single chat transcript. The planning and execution guide explains how one accepted step advances. The tool calling guide explains the execution gate.

Key components and contracts

Working state stores taskId, planVersion, stepId, status, input references, current object version and remaining budget. Update it atomically with a compare-and-swap or equivalent transaction. On restart, load the latest committed checkpoint; never infer progress from the assistant's last sentence.

Durable memory stores a minimal reusable fact with owner, tenant, source locator, observed time, expiry, classification and deletion path. Retrieval must filter by scope before ranking. A retrieved note can inform a proposal, but cannot add a tool, grant a permission or override a current system record.

Execution receipts bind a call ID and idempotency key to arguments hash, policy version, result status and destination read-back. Keep sensitive payloads out of default logs. A timeout means the effect may have happened; reconcile against the destination before repeating it.

Memory curator decides whether a fact merits retention. Prefer an explicit rule and human review for cross-task memory. A stale preference, an unverified model summary or a revoked instruction should be removed or quarantined. Keep an owner and retention window for every class.

Minimal pseudocode

ts
async function resume(taskId: string) {
  const checkpoint = await state.loadVersioned(taskId);
  const candidates = await memory.search({
    tenant: checkpoint.tenant,
    purpose: checkpoint.purpose,
    query: checkpoint.nextStep.query,
  });
  const evidence = candidates.filter(item =>
    item.source && item.expiresAt > now() && policy.canRead(checkpoint.actor, item)
  );
  const current = await source.read(checkpoint.nextStep.objectId);
  if (current.version !== checkpoint.objectVersion) return review("stale_state");
  const proposal = await model.propose({ checkpoint, evidence, current });
  if (!policy.allows(checkpoint.actor, proposal.action)) return review("forbidden");
  const key = `${taskId}:${checkpoint.planVersion}:${checkpoint.nextStep.id}`;
  const result = await tools.executeOnce(proposal, key);
  if (result.unknown) return reconcileBeforeRetry(key);
  return state.commitIfVersion(checkpoint.version, await verify(result));
}

This sketch omits database transactions, encryption, redaction and failure handling. Those need concrete implementation and tests. The model does not write directly into durable memory or advance a checkpoint without application validation.

Failure modes

  1. A transcript becomes the source of truth. Persist structured checkpoints and verify against the destination system.
  2. Old memory outranks current data. Attach timestamps and source locators; check freshness and current object version.
  3. A cross-tenant note is retrieved. Enforce access filters before ranking and test isolation with fixtures.
  4. The model stores secrets or personal data for convenience. Apply classification, redaction and retention before storage.
  5. A write is repeated after timeout. Use a stable idempotency key and reconcile the destination before retry.
  6. A summary invents a prior approval. Approval is an application record with scope and expiry, never a memory claim.
  7. Deletion misses derived copies. Track origin and downstream indexes so removal propagates.

Testing and production checklist

Test restart before a tool call, after a call but before receipt commit, stale CRM versions, concurrent workers, expired memory, revoked access, another tenant, contradictory source data, deletion and an unknown write outcome. Assert both the route taken and the absence of forbidden side effects. Measure resume correctness, stale retrievals, denied retrievals, duplicate writes and memory deletion latency on your own fixtures; no performance figure is asserted here.

Before a pilot, define storage ownership, access policy, retention, encryption, redaction, backup, deletion, incident review and observable stop conditions. Start with read-only retrieval and controlled checkpoints. For a process-specific design, request an architecture review with a sample task, data classes and one failure case. AI automation and prompt engineering provide related implementation context.

CODE_BLOCK.TXT
require(checkpoint.version && policy.canRead(memory));
if (source.version !== checkpoint.objectVersion) return review;
result = await executeOnce(step.idempotencyKey);
if (result.unknown) return reconcileBeforeRetry;
return commitIfVersion(checkpoint.version, verify(result));