Short-Term, Long-Term and Episodic Memory in AI Agents
Short-term context holds the current task
Knowledge and episodes carry source and outcome
Original MEMORY-3 design and synthetic support ticket
Memory limits and practical checks

$ memory --contract MEMORY-3
> current: bounded task context
> retrieve: sourced / scoped knowledge
> episode: event / receipt / outcomeAI agent memory types answer different questions: what is happening now, what knowledge may be useful later, and what happened in a particular past episode. These are architectural roles, not three mandatory products or abilities that appear inside a model by default. The model receives context for each call; the application decides what to store, retrieve, check and pass into that call. For the broader engineering brief, see AI specialist in Armenia.
This guide uses an original MEMORY-3 design and one synthetic support-ticket example. It is a proposed architecture, not a description of a live client deployment or measured outcome. The adjacent state and memory guide covers checkpointing between steps; this article focuses on how the three memory types differ.
Definition without the hype
Short-term memory is the working context for the current task: recent messages, the selected ticket, the immediate goal and interim reasoning artifacts. It is usually passed to the model within a finite context window. Older passages may be summarized or dropped. Context alone is not a reliable durable record after a task ends.
Long-term memory holds facts, documents, preferences or retrieved knowledge that may be needed in later tasks. Store it outside model calls, with source, owner, version, retention and access rules. “Long-term” does not mean permanent or true: a record can become stale or need deletion.
Episodic memory records specific events: which task ran, which tool was called, what the destination returned and what a person approved. An episode binds context to a time and outcome. It supports recovery and audit, but should not automatically become a general fact about a user.
| Type | Question | Support-ticket example | Main risk |
|---|---|---|---|
| Short-term | What matters now? | current request and draft | summary drops a condition |
| Long-term | What should be known later? | authorized return policy | stale or cross-tenant knowledge |
| Episodic | What actually happened? | draft created; send unconfirmed | repeating an action after misreading the event |
How it works: MEMORY-3 architecture
The design separates the three stores and puts a check before each use:
task + authenticated actor
→ current working context
→ authorized knowledge (source, version, expiry)
→ relevant episodes (event, time, outcome)
→ bounded context assembly for the model
→ proposed action → policy → tool
→ destination verification → new episodeThe application first checks identity, tenant and task purpose. It loads the current conversation and only documents the actor may read. Long-term retrieval may use exact keys or semantic search, but a retrieved passage remains data, never an instruction allowed to override policy. Episodes are filtered by object, time and outcome status. Context assembly imposes a size limit and carries provenance labels. The model proposes an answer or action; the server checks authority and effect again. Execution produces either a verified receipt or an explicit unknown-outcome state.
Memory differs from execution state. The current step number, request ID and pending approval belong in a versioned process checkpoint. A past conversation may explain context, but cannot replace that checkpoint or a destination receipt. A send timeout is not proof of failure; reconcile the destination before retrying.
Where MCP and tools fit
MCP can expose tools that read documents or events, but the protocol does not itself determine memory quality. The application owns sources, authorization, retention and updates. The tool permission model must check access to each retrieved object and proposed effect. Instructions embedded in stored text do not gain policy authority when search retrieves them.
One end-to-end example
An employee opens ticket 784: a customer asks about a return. Short-term context holds the ticket text, selected language and goal to prepare a draft. Long-term storage returns the current return policy with version and source, within the employee's access scope. The episode log says a draft was created yesterday, with no confirmed send receipt.
The agent must not tell the customer “we already replied”: the episode confirms a draft, not a sent message. It prepares a draft citing the policy and marks the uncertainty. If the employee requests a send, the application checks present permissions, exact content and required approval, then obtains a receipt from the channel. A new episode records the message ID and verified outcome. If the channel times out, the route is reconciliation by ID, not an immediate resend.
A compact illustrative contract:
const context = await loadCurrentTask(taskId, actor);
const knowledge = await retrieveKnowledge({
tenant: actor.tenant, scope: "returns", asOf: now, requireSource: true
});
const events = await loadEvents({ ticketId: 784, actor, limit: 10 });
const proposal = await model.propose({ context, knowledge, events });
if (proposal.action === "send") return holdForExactApproval(proposal);
return saveDraftWithProvenance(proposal);This is pseudocode. A real implementation enforces document and event access on the server and reads the send outcome from the destination. The model cannot choose its own tenant or read a raw shared archive.
Fit table: when each type helps
| Task | Suitable | Unsuitable |
|---|---|---|
| Clarify the customer's latest message | short-term context | save every exchange forever as fact |
| Find the current return policy | versioned long-term knowledge | trust an old uncited model answer |
| Determine whether a reply was sent | episode plus channel receipt | treat a draft as evidence of send |
| Resume an interrupted workflow | checkpoint and verified episodes | infer the step from a free-form recap |
| Reauthorize an action | current access policy | reuse an old permission copied into memory |
For many workflows, a small working context and ordinary document database suffice. Episodic memory becomes useful when process recovery, action audit and a precise distinction between proposed, executed and verified effects matter. Sophisticated vector retrieval is optional for a small rule set; exact lookup by ID and version may be easier to verify.
Limits and validation
Memory enlarges the error surface. Summarization can erase a condition; semantic search can return a similar but unauthorized document; an old episode may conflict with current destination state. Every fragment needs provenance, timestamp, scope and deletion path. Retain personal data only where justified for the workflow, with a bounded lifetime.
Test revoked access, a new document version, two similar customers, conflicting episodes, a malicious instruction in retrieved text and a timeout after an external action. Useful operational signals include answers with traceable sources, denied cross-scope retrievals, stale hits and unknown outcomes. Do not promise performance gains without actual measurements.
Start with one bounded task, explicit knowledge sources, a small event log with verified statuses and a safe-resume test. Prompt engineering helps shape context; AI automation covers process ownership. For an architecture review, describe your workflow.
const knowledge = await retrieve({ actor, scope, asOf: now });
const episodes = await loadVerifiedEvents(taskId, actor);
const proposal = await model.propose({ context, knowledge, episodes });
return policy.check(proposal, actor);