State and Memory in AI Agents: What to Keep Between Steps
Checkpoints record progress; memory supplies scoped evidence
Freshness, permissions, receipts and safe resumption
Original STATE-LEDGER-7 cycle and failure routes
What an agent should retain between steps and sessions

$ resume --contract STATE-LEDGER-7
> load: versioned checkpoint / owner
> retrieve: scoped and fresh memory
> gate: policy / current source
> act: idempotent call / receipt
> route: verified / reconcile / reviewState memory AI agent design is a question of which facts must survive a step, a retry, or a new session. A model's context window is neither a transaction log nor an authority for future actions. This guide uses an original synthetic contract, STATE-LEDGER-7, to separate task state, durable memory and audit receipts. It does not report client results or benchmark scores.
For the broader choice of an AI implementation partner in Armenia, see the AI specialist page. This article addresses the narrower engineering question of storage and retrieval between agent steps.
Problem and requirements
Imagine an agent checking a supplier request against a CRM record, preparing a proposed update, then waiting for review. It must resume after a process restart without repeating a write or mistaking an old CRM value for current evidence. It also should not carry the whole conversation, private identifiers or obsolete instructions into every future task.
Define the task owner, tenant, accepted outcome, permitted tools, retention policy and review boundary first. Decide which facts need a short lifetime, which can be reused across tasks, and which must be retained solely as an inspectable receipt. A summary generated by the model is useful context, but its claims require provenance and expiry.
Architecture: STATE-LEDGER-7
The original STATE-LEDGER-7 contract has seven gates: bind → checkpoint → classify → retrieve → validate → act → reconcile. It is a design example, not a deployed system. A versioned checkpoint records the current task and step. A selective memory index offers prior facts as untrusted evidence. An append-only receipt records tool calls and verified outcomes. The application, not retrieved text, determines permissions.
| Gate | Stored input | Required check |
|---|---|---|
| Bind | Task ID, owner, tenant, scope | Identity and allowed outcome are explicit |
| Checkpoint | Step, plan version, object version | Atomic version comparison prevents stale resume |
| Classify | Candidate fact and sensitivity | Retention, purpose and deletion rule are assigned |
| Retrieve | Scoped query and candidate references | Tenant, provenance, freshness and permission are checked |
| Validate | Evidence plus current source | Contradictions and stale values are surfaced |
| Act | Authorized tool call | Idempotency key and current policy are checked |
| Reconcile | Receipt and destination read-back | Unknown effects are resolved before retry |
The checkpoint belongs to the current run; durable memory contains deliberately reusable facts; the receipt supports recovery and audit. Do not collapse the three into a single chat transcript. The planning and execution guide explains how one accepted step advances. The tool calling guide explains the execution gate.
Key components and contracts
Working state stores taskId, planVersion, stepId, status, input references, current object version and remaining budget. Update it atomically with a compare-and-swap or equivalent transaction. On restart, load the latest committed checkpoint; never infer progress from the assistant's last sentence.
Durable memory stores a minimal reusable fact with owner, tenant, source locator, observed time, expiry, classification and deletion path. Retrieval must filter by scope before ranking. A retrieved note can inform a proposal, but cannot add a tool, grant a permission or override a current system record.
Execution receipts bind a call ID and idempotency key to arguments hash, policy version, result status and destination read-back. Keep sensitive payloads out of default logs. A timeout means the effect may have happened; reconcile against the destination before repeating it.
Memory curator decides whether a fact merits retention. Prefer an explicit rule and human review for cross-task memory. A stale preference, an unverified model summary or a revoked instruction should be removed or quarantined. Keep an owner and retention window for every class.
Minimal pseudocode
async function resume(taskId: string) {
const checkpoint = await state.loadVersioned(taskId);
const candidates = await memory.search({
tenant: checkpoint.tenant,
purpose: checkpoint.purpose,
query: checkpoint.nextStep.query,
});
const evidence = candidates.filter(item =>
item.source && item.expiresAt > now() && policy.canRead(checkpoint.actor, item)
);
const current = await source.read(checkpoint.nextStep.objectId);
if (current.version !== checkpoint.objectVersion) return review("stale_state");
const proposal = await model.propose({ checkpoint, evidence, current });
if (!policy.allows(checkpoint.actor, proposal.action)) return review("forbidden");
const key = `${taskId}:${checkpoint.planVersion}:${checkpoint.nextStep.id}`;
const result = await tools.executeOnce(proposal, key);
if (result.unknown) return reconcileBeforeRetry(key);
return state.commitIfVersion(checkpoint.version, await verify(result));
}This sketch omits database transactions, encryption, redaction and failure handling. Those need concrete implementation and tests. The model does not write directly into durable memory or advance a checkpoint without application validation.
Failure modes
- A transcript becomes the source of truth. Persist structured checkpoints and verify against the destination system.
- Old memory outranks current data. Attach timestamps and source locators; check freshness and current object version.
- A cross-tenant note is retrieved. Enforce access filters before ranking and test isolation with fixtures.
- The model stores secrets or personal data for convenience. Apply classification, redaction and retention before storage.
- A write is repeated after timeout. Use a stable idempotency key and reconcile the destination before retry.
- A summary invents a prior approval. Approval is an application record with scope and expiry, never a memory claim.
- Deletion misses derived copies. Track origin and downstream indexes so removal propagates.
Testing and production checklist
Test restart before a tool call, after a call but before receipt commit, stale CRM versions, concurrent workers, expired memory, revoked access, another tenant, contradictory source data, deletion and an unknown write outcome. Assert both the route taken and the absence of forbidden side effects. Measure resume correctness, stale retrievals, denied retrievals, duplicate writes and memory deletion latency on your own fixtures; no performance figure is asserted here.
Before a pilot, define storage ownership, access policy, retention, encryption, redaction, backup, deletion, incident review and observable stop conditions. Start with read-only retrieval and controlled checkpoints. For a process-specific design, request an architecture review with a sample task, data classes and one failure case. AI automation and prompt engineering provide related implementation context.
require(checkpoint.version && policy.canRead(memory));
if (source.version !== checkpoint.objectVersion) return review;
result = await executeOnce(step.idempotencyKey);
if (result.unknown) return reconcileBeforeRetry;
return commitIfVersion(checkpoint.version, verify(result));