Back to blog
AI Agents

Short-Term, Long-Term and Episodic Memory in AI Agents

Short-term context holds the current task

Knowledge and episodes carry source and outcome

Original MEMORY-3 design and synthetic support ticket
Memory limits and practical checks
Primary nodeScoped memory assembly
Routing modeMEMORY-3
StatusPUBLISHED
Three AI agent memory layers connect current context, sourced knowledge and verified episodes
MEMORY_3_V01: current context, sourced knowledge and verified episodes.
TERMINAL_PREVIEW.LOG
$ memory --contract MEMORY-3
> current: bounded task context
> retrieve: sourced / scoped knowledge
> episode: event / receipt / outcome
Architecture, example and validation checks

AI agent memory types answer different questions: what is happening now, what knowledge may be useful later, and what happened in a particular past episode. These are architectural roles, not three mandatory products or abilities that appear inside a model by default. The model receives context for each call; the application decides what to store, retrieve, check and pass into that call. For the broader engineering brief, see AI specialist in Armenia.

This guide uses an original MEMORY-3 design and one synthetic support-ticket example. It is a proposed architecture, not a description of a live client deployment or measured outcome. The adjacent state and memory guide covers checkpointing between steps; this article focuses on how the three memory types differ.

Definition without the hype

Short-term memory is the working context for the current task: recent messages, the selected ticket, the immediate goal and interim reasoning artifacts. It is usually passed to the model within a finite context window. Older passages may be summarized or dropped. Context alone is not a reliable durable record after a task ends.

Long-term memory holds facts, documents, preferences or retrieved knowledge that may be needed in later tasks. Store it outside model calls, with source, owner, version, retention and access rules. “Long-term” does not mean permanent or true: a record can become stale or need deletion.

Episodic memory records specific events: which task ran, which tool was called, what the destination returned and what a person approved. An episode binds context to a time and outcome. It supports recovery and audit, but should not automatically become a general fact about a user.

TypeQuestionSupport-ticket exampleMain risk
Short-termWhat matters now?current request and draftsummary drops a condition
Long-termWhat should be known later?authorized return policystale or cross-tenant knowledge
EpisodicWhat actually happened?draft created; send unconfirmedrepeating an action after misreading the event

How it works: MEMORY-3 architecture

The design separates the three stores and puts a check before each use:

text
task + authenticated actor
    → current working context
    → authorized knowledge (source, version, expiry)
    → relevant episodes (event, time, outcome)
    → bounded context assembly for the model
    → proposed action → policy → tool
    → destination verification → new episode

The application first checks identity, tenant and task purpose. It loads the current conversation and only documents the actor may read. Long-term retrieval may use exact keys or semantic search, but a retrieved passage remains data, never an instruction allowed to override policy. Episodes are filtered by object, time and outcome status. Context assembly imposes a size limit and carries provenance labels. The model proposes an answer or action; the server checks authority and effect again. Execution produces either a verified receipt or an explicit unknown-outcome state.

Memory differs from execution state. The current step number, request ID and pending approval belong in a versioned process checkpoint. A past conversation may explain context, but cannot replace that checkpoint or a destination receipt. A send timeout is not proof of failure; reconcile the destination before retrying.

Where MCP and tools fit

MCP can expose tools that read documents or events, but the protocol does not itself determine memory quality. The application owns sources, authorization, retention and updates. The tool permission model must check access to each retrieved object and proposed effect. Instructions embedded in stored text do not gain policy authority when search retrieves them.

One end-to-end example

An employee opens ticket 784: a customer asks about a return. Short-term context holds the ticket text, selected language and goal to prepare a draft. Long-term storage returns the current return policy with version and source, within the employee's access scope. The episode log says a draft was created yesterday, with no confirmed send receipt.

The agent must not tell the customer “we already replied”: the episode confirms a draft, not a sent message. It prepares a draft citing the policy and marks the uncertainty. If the employee requests a send, the application checks present permissions, exact content and required approval, then obtains a receipt from the channel. A new episode records the message ID and verified outcome. If the channel times out, the route is reconciliation by ID, not an immediate resend.

A compact illustrative contract:

ts
const context = await loadCurrentTask(taskId, actor);
const knowledge = await retrieveKnowledge({
  tenant: actor.tenant, scope: "returns", asOf: now, requireSource: true
});
const events = await loadEvents({ ticketId: 784, actor, limit: 10 });
const proposal = await model.propose({ context, knowledge, events });
if (proposal.action === "send") return holdForExactApproval(proposal);
return saveDraftWithProvenance(proposal);

This is pseudocode. A real implementation enforces document and event access on the server and reads the send outcome from the destination. The model cannot choose its own tenant or read a raw shared archive.

Fit table: when each type helps

TaskSuitableUnsuitable
Clarify the customer's latest messageshort-term contextsave every exchange forever as fact
Find the current return policyversioned long-term knowledgetrust an old uncited model answer
Determine whether a reply was sentepisode plus channel receipttreat a draft as evidence of send
Resume an interrupted workflowcheckpoint and verified episodesinfer the step from a free-form recap
Reauthorize an actioncurrent access policyreuse an old permission copied into memory

For many workflows, a small working context and ordinary document database suffice. Episodic memory becomes useful when process recovery, action audit and a precise distinction between proposed, executed and verified effects matter. Sophisticated vector retrieval is optional for a small rule set; exact lookup by ID and version may be easier to verify.

Limits and validation

Memory enlarges the error surface. Summarization can erase a condition; semantic search can return a similar but unauthorized document; an old episode may conflict with current destination state. Every fragment needs provenance, timestamp, scope and deletion path. Retain personal data only where justified for the workflow, with a bounded lifetime.

Test revoked access, a new document version, two similar customers, conflicting episodes, a malicious instruction in retrieved text and a timeout after an external action. Useful operational signals include answers with traceable sources, denied cross-scope retrievals, stale hits and unknown outcomes. Do not promise performance gains without actual measurements.

Start with one bounded task, explicit knowledge sources, a small event log with verified statuses and a safe-resume test. Prompt engineering helps shape context; AI automation covers process ownership. For an architecture review, describe your workflow.

CODE_BLOCK.TXT
const knowledge = await retrieve({ actor, scope, asOf: now });
const episodes = await loadVerifiedEvents(taskId, actor);
const proposal = await model.propose({ context, knowledge, episodes });
return policy.check(proposal, actor);