Three layers: (1) short-term / working memory: the LLM context window holding recent observations, scratchpad, and the current task plan; volatile; resets per session; managed by compaction/summarization as it fills. (2) long-term / semantic memory: a structured store of facts about the user, prior tasks, world knowledge, indexed by entity and topic; often a vector store + structured (SQL/JSON). (3) episodic memory: timestamped records of past episodes the agent can recall to learn from experience (auto-summarized). Memory also implies write policies: what the agent is allowed to remember about the user (privacy / consent), and conflict resolution when new info contradicts old. Senior nuance: long-term memory is the differentiator between toy demos and persistent assistants, and the architecture is typically a vector store + metadata filter + periodic compaction. MemGPT and Letta (formerly MemGPT) formalized these layers in 2024-2025.