Skip to content
All guidesAI systems

The layers of AI memory, explained

Working context, session state, episodic, semantic and procedural memory. We trace a weekly report through five useful categories, showing what gets retrieved, what survives a restart and how to update a memory safely.

6 min read

An agent writes a weekly report. You ask for shorter summaries, then start a new conversation next week. For that preference to affect the next report, the application has to save it, find it again and include it in the model's input. Each of those operations can fail independently.

We can make the design easier to inspect by separating five categories of memory. Working context and session state describe where information is available. Episodic, semantic and procedural memory describe what information means. These categories overlap. They are a useful architectural model, not five compulsory databases or a universal industry standard.

1. Working context

Working context is the information supplied to the model for its current response. It can contain instructions, recent messages, tool results and selected memories. Tokens are the pieces of text a model reads and generates, and the context window limits how many fit into one interaction.

For the report, this could include the current request, this week's figures and the saved preference for short summaries. The application assembles those inputs. A fact stored elsewhere does not automatically become available to the model. Anthropic's context engineering guidance describes retrieving information when needed and keeping useful notes outside the context window.

2. Session state

Session state records where the current task has reached. Our reporting agent might hold the reporting period, completed data queries, draft location and approval status. A checkpoint is a saved snapshot that lets the application resume that task. LangGraph, an agent orchestration framework, associates these checkpoints with a thread identifier.

A checkpoint can survive a process restart if it is written to durable storage. An in-memory checkpointer keeps its contents in the running process and loses them when that process restarts. Resuming the same thread also differs from starting a new thread, which needs its own state or an explicit transfer.

3. Episodic memory

Episodic memory records what happened in a particular attempt. For example, last week's report missed refunds, failed the totals check and passed after the agent included the refund records. That episode can carry a timestamp, the action taken, the observed outcome and a link to the evidence.

We would retrieve that episode when a similar reporting task makes it relevant. It provides an example of a failure and recovery. It does not prove that every future report will miss refunds. Preserve the circumstances so the agent can judge whether the lesson applies.

4. Semantic memory

Semantic memory holds facts and preferences, such as a user's preferred report length. A structured record might store user_id, report_style, source_message_id, updated_at and version. The original message gives the preference a traceable source. Its version lets the application distinguish the current value from an older one.

Exact preferences can be fetched by a known key. Larger collections may use keyword or similarity search. An embedding is a numerical representation used to compare meaning, and a vector index helps search those representations. Semantic memory describes the content; semantic search describes one way of finding content. Neither requires the other.

5. Procedural memory

Procedural memory describes how to do the work. In this example, a versioned reporting skill might instruct the agent to fetch current figures, check dates and refunds, calculate changes, draft the summary and request approval before sending it. The skill is a reusable instruction file loaded when the task needs it.

A successful episode can suggest a better procedure, but we would review that change before promoting it into a shared skill. One lucky result is weak evidence for a rule every future report must follow. These application memory operations do not update the model's trained weights.

Follow one report through the layers

Retrieve → assemble context → act → check → write back
  1. RetrieveAgent

    The agent's retrieval tools fetch the authenticated user's report preference, the approved reporting skill and a relevant previous episode.

    SemanticProceduralEpisodic

  2. Assemble contextAgent

    The application combines selected information with this week's request and the current task state for the model call.

    Working contextSession state

  3. ActAgent

    The agent fetches this week's figures, drafts a short report and checkpoints its progress.

    Working contextSession state

  4. CheckYou

    The workflow checks dates and calculations, then presents the draft for your approval before sending.

    EvidenceApproval

  5. Write backAgent

    The application records the checked outcome and saves any explicit preference change under the correct user.

    Session stateEpisodicSemantic

This is an illustrative workflow. Failed checks return the task to revision. Shared procedure changes take a separate review; they are not written automatically after every run.

Persistence and access are separate decisions

Starting a new conversation does not make all stored information available. The application must retrieve authorised records into the new context. Restarting a process preserves only information saved to storage that survives that restart. Changing users should change the accessible scope, even if both users share the same agent.

LangGraph separates thread checkpoints from stores that can serve multiple threads. Its stores organise records using namespaces and keys. A namespace is a grouping such as an organisation and user identifier. The application still needs to enforce who may read and write that grouping.

This distinction is current engineering work. LangChain's Managed Deep Agents 0.8 announcement on 24 September 2026 introduced memory scoped to the authenticated user alongside shared agent memory. Those are access scopes, which can each contain several of the information categories above.

Build the update and forgetting paths too

  • Outdated preference. If the user requests detailed reports again, replace or supersede the short-report preference. Filter obsolete versions out of retrieval.
  • Conflicting records. Keep the source and effective date, then apply an explicit precedence rule. Ask the user when authoritative records still disagree.
  • Wrong user. Derive retrieval scope from authenticated identity and enforce it in the data access layer. Similarity ranking is not an access-control boundary.
  • Unwanted memory. Support deletion from active records and search indexes, and document what happens to retained logs and backups. Removing a message from the current prompt does not delete every stored copy.

Sources