Tutorials › Agentic AI › Memory and Context

Agentic AI · Part 5 of 7

Memory and Context

What an agent carries between steps and between sessions, and how to keep it within the context window.

The transcript is short-term memory

An agent's short-term memory is the transcript the loop sends on every trip: the goal, the model's replies, its tool calls, and their results. Frameworks keep it as part of the agent's state, scoped to one conversation. LangGraph, for example, persists it as thread-scoped checkpoints. Each trip appends to it.

Why the transcript needs managing

Everything the model sees has to fit in its context window (see Foundation Models and How They're Served). LangGraph's documentation notes that a full history may not fit, which results in an error. A transcript that fits can still hurt: Anthropic's guide on context engineering describes context as a finite resource with diminishing marginal returns, and cites needle-in-a-haystack benchmarks showing that the model's ability to accurately recall information from the context decreases as it fills.

Keeping the transcript small

Prompt Engineering as Application Logic covers context management for prompts in general.

Long-term memory

Memory across sessions lives outside the transcript. LangGraph stores long-term memories as JSON documents organized by namespace and key, available across conversation threads. Anthropic offers a memory tool that stores and retrieves information across conversations in files you control. In both, the agent reads and writes the store through tools, so it follows the mechanics in Tools and Tool Calling. Looking a memory up is a retrieval step, the mechanism covered in Retrieval-Augmented Generation.

Memory and prompt caching

Each trip repeats the same beginning of the transcript, which is the case prompt caching is built for (see Prompt Engineering as Application Logic). Anthropic's automatic caching moves its breakpoint forward as a conversation grows. Dropping or summarizing earlier turns changes that beginning, and a cache hit requires an identical prefix, so the cached entry stops matching.