AI 智能体记忆系统设计:从短期记忆到长期记忆的完整架构

AI Agent 没有记忆,就像失忆的医生——每次看病都从零开始。本文讲 AI Agent 的记忆系统设计:从短期记忆(短期上下文)到长期记忆(向量数据库)的完整架构。
读完保证:能设计 4 层记忆架构 + 5 大记忆策略 + 6 个月落地。

一、为什么 AI Agent 需要记忆?

1.1 当前 AI Agent 的"失忆症"

场景:AI 客服 Agent

第 1 次对话:
用户:我姓王
AI:好的王先生,请问您的问题?
用户:我的订单 #123 物流很慢
AI:正在查询您的订单...物流正常

第 2 次对话(5 分钟后):
用户:订单怎么样了?
AI:您好,请提供您的订单号?  ← 失忆了

问题:每次对话 = 从零开始。用户重复信息,体验差。

1.2 真实场景需要"记忆"

场景没有记忆有记忆
AI 客服重复问订单号记住订单状态
AI 编程助手每次重新理解项目记住项目结构
AI 销售不知道客户历史知道客户购买历史
AI 陪伴每次都是陌生人了解你的性格
AI 研究助手不知道你研究过什么累积你的研究

1.3 2026 年记忆市场

2024 年:实验阶段(LangChain Memory 雏形)
2025 年:$500M 商业化
2026 年:$5B 爆发
2028 年:$50B+(成为 AI 基础设施)

→ 2026 年记忆 = AI 时代"数据库"

二、4 层记忆架构

┌────────────────────────────────────────────────────┐
│              AI Agent 4 层记忆架构                    │
│                                                      │
│   ┌────────────────────────────────────┐            │
│   │  L1:短期记忆(Working Memory)      │            │
│   │  存储:LLM 上下文窗口                │            │
│   │  容量:4K-200K token                │            │
│   │  持续:单次会话                     │            │
│   └──────────┬─────────────────────────┘            │
│              ↓                                       │
│   ┌────────────────────────────────────┐            │
│   │  L2:工作记忆(Scratchpad)         │            │
│   │  存储:Agent 内部状态               │            │
│   │  容量:1K-10K token                │            │
│   │  持续:单次任务                     │            │
│   └──────────┬─────────────────────────┘            │
│              ↓                                       │
│   ┌────────────────────────────────────┐            │
│   │  L3:长期记忆(Long-term Memory)   │            │
│   │  存储:向量数据库 + SQL            │            │
│   │  容量:100K-10M 条记录              │            │
│   │  持续:永久                         │            │
│   └──────────┬─────────────────────────┘            │
│              ↓                                       │
│   ┌────────────────────────────────────┐            │
│   │  L4:用户/全局记忆(Profile)       │            │
│   │  存储:PostgreSQL + 向量           │            │
│   │  容量:每用户 1K-100K 条           │            │
│   │  持续:永久 + 跨设备                │            │
│   └────────────────────────────────────┘            │
└────────────────────────────────────────────────────┘

三、L1:短期记忆(Working Memory)

3.1 设计原理

短期记忆 = LLM 上下文窗口

数据:当前对话的所有消息
存储:在 LLM API 调用时直接传
持续:单次会话(用户离开就清空)
容量:4K-200K token(取决于模型)

GPT-4o:128K
Claude 3.5:200K
Gemini 1.5:2M
Qwen2.5:128K
DeepSeek:64K

3.2 优化短期记忆

策略 1:滑动窗口(保留最近 N 条)

class ShortTermMemory:
    def __init__(self, max_messages=20):
        self.messages = []
        self.max = max_messages

    def add(self, role, content):
        self.messages.append({"role": role, "content": content})
        # 只保留最近 N 条
        if len(self.messages) > self.max:
            self.messages = self.messages[-self.max:]

    def get(self):
        return self.messages

策略 2:摘要压缩(超长时压缩)

def compress_if_needed(messages, model, max_tokens=8000):
    """如果消息超长,压缩早期消息"""
    if count_tokens(messages) > max_tokens:
        # 保留最近 5 条原样
        recent = messages[-5:]
        # 早期消息压缩成 1 条
        old = messages[:-5]
        summary = model.invoke([
            {"role": "system", "content": "请用 200 字总结以下对话"},
            {"role": "user", "content": format_messages(old)}
        ])
        return [{"role": "assistant", "content": f"[历史摘要] {summary}"}] + recent
    return messages

四、L2:工作记忆(Scratchpad)

4.1 设计原理

工作记忆 = Agent 内部状态

数据:当前任务的中间结果
  - 思考链(Chain of Thought)
  - 工具调用历史
  - 子任务状态
  - 推理步骤

存储:Agent 内存(Python 对象)
持续:单次任务(任务结束清空)
容量:1K-10K token

4.2 实战代码

class Scratchpad:
    """工作记忆 - 任务执行中"""
    def __init__(self):
        self.thoughts = []       # 思考链
        self.actions = []        # 行动历史
        self.observations = []   # 观察结果
        self.sub_tasks = []      # 子任务状态

    def think(self, thought):
        """Agent 思考"""
        self.thoughts.append(thought)

    def act(self, action, args):
        """Agent 行动"""
        self.actions.append({"action": action, "args": args})

    def observe(self, result):
        """观察结果"""
        self.observations.append(result)

    def get_context(self):
        """获取完整上下文(给 LLM 看)"""
        return {
            "thoughts": self.thoughts,
            "actions": self.actions,
            "observations": self.observations,
            "sub_tasks": self.sub_tasks,
        }

# 使用
sp = Scratchpad()
sp.think("用户问订单状态,需要先查订单号")
sp.act("query_order", {"order_id": 123})
result = query_database("123")
sp.observe(result)
sp.think("订单已发货,预计明天到")

五、L3:长期记忆(Long-term Memory)

5.1 设计原理

长期记忆 = 向量数据库 + 结构化数据库

数据:所有历史对话 + Agent 学到的知识
存储:
  - 向量数据:Qdrant / Milvus / pgvector
  - 结构化:PostgreSQL
  - 文本:S3 / OSS
持续:永久(除非主动删除)
容量:100K-10M 条

5.2 实战架构

from qdrant_client import QdrantClient
from qdrant_client.http import models
import openai
import uuid

class LongTermMemory:
    def __init__(self):
        self.qdrant = QdrantClient(host="localhost", port=6333)
        self.openai_client = openai.Client()
        self.collection = "agent_memory"

    def store(self, content, metadata={}):
        """存储记忆"""
        # 1. 向量化
        response = self.openai_client.embeddings.create(
            model="text-embedding-3-small",
            input=content
        )
        vector = response.data[0].embedding

        # 2. 存储到 Qdrant
        self.qdrant.upsert(
            collection_name=self.collection,
            points=[models.PointStruct(
                id=str(uuid.uuid4()),
                vector=vector,
                payload={"content": content, **metadata}
            )]
        )

    def retrieve(self, query, top_k=5):
        """检索相关记忆"""
        # 1. 向量化查询
        response = self.openai_client.embeddings.create(
            model="text-embedding-3-small",
            input=query
        )
        query_vector = response.data[0].embedding

        # 2. 检索
        results = self.qdrant.search(
            collection_name=self.collection,
            query_vector=query_vector,
            limit=top_k
        )

        return [
            {"content": r.payload["content"], "score": r.score}
            for r in results
        ]

5.3 长期记忆 5 大策略

策略 1:何时存储

def should_store(self, message):
    """判断是否需要存储"""
    # 存储条件
    if message["role"] == "user":
        # 重要信息:偏好、事实、问题
        keywords = ["我", "名字", "喜欢", "需要", "想要", "my", "I like", "I need"]
        if any(kw in message["content"] for kw in keywords):
            return True
    return False

策略 2:如何索引

def index_memory(self, content):
    """索引记忆(用于快速检索)"""
    # 关键词索引
    keywords = extract_keywords(content)
    for kw in keywords:
        self.keyword_index[kw].append(content)

    # 标签索引
    tags = extract_tags(content)
    for tag in tags:
        self.tag_index[tag].append(content)

策略 3:去重 + 压缩

def deduplicate(self, memories):
    """去重相似记忆"""
    unique = []
    for m in memories:
        # 相似度 > 0.95 = 重复
        is_dup = False
        for u in unique:
            sim = self.cosine_similarity(m["vector"], u["vector"])
            if sim > 0.95:
                # 合并(保留最新的)
                u["updated_at"] = m["created_at"]
                is_dup = True
                break
        if not is_dup:
            unique.append(m)
    return unique

策略 4:遗忘机制

def should_forget(self, memory):
    """判断是否应该遗忘"""
    # Ebbinghaus 遗忘曲线
    age_days = (now() - memory["created_at"]).days
    importance = memory["importance"]  # 0-1

    # 重要性 × 时间 = 保留概率
    keep_prob = importance * (1 - age_days / 365)
    return random() > keep_prob  # 30% 概率遗忘

策略 5:分层存储

# 热数据:最近 7 天
# 温数据:7-30 天
# 冷数据:> 30 天

# 存储分层
- 热:Redis(快但贵)
- 温:PostgreSQL(中等)
- 冷:S3(慢但便宜)

六、L4:用户 / 全局记忆(Profile)

6.1 设计原理

用户记忆 = 关于这个用户的专属记忆

数据:
  - 基本信息(姓名 / 偏好 / 习惯)
  - 行为历史(过去 N 天的对话)
  - 关系网络(与其他人 / 事务的关系)
  - 长期偏好

存储:
  - 结构化:PostgreSQL(user_id, key, value)
  - 向量:Qdrant(语义检索)
持续:永久 + 跨设备 + 跨平台

6.2 实战代码

class UserMemory:
    """用户专属记忆"""
    def __init__(self, user_id):
        self.user_id = user_id
        self.db = Database()
        self.vector_db = QdrantClient()

    def get_profile(self):
        """获取用户画像"""
        # 1. 从数据库查
        rows = self.db.query(
            "SELECT key, value FROM user_profile WHERE user_id = %s",
            (self.user_id,)
        )
        profile = {row["key"]: row["value"] for row in rows}

        # 2. 从向量库查最近对话
        memories = self.vector_db.search(
            collection_name=f"user_{self.user_id}_memory",
            query_vector=...,
            limit=10
        )

        return {"profile": profile, "memories": memories}

    def update(self, key, value):
        """更新用户记忆"""
        self.db.execute(
            "INSERT INTO user_profile (user_id, key, value, updated_at) "
            "VALUES (%s, %s, %s, NOW()) "
            "ON CONFLICT (user_id, key) DO UPDATE SET value = %s",
            (self.user_id, key, value, value)
        )

6.3 用户记忆 3 大类

类型 1:静态信息(很少变)
- 姓名、性别、年龄
- 职业、所在地
- 语言偏好

类型 2:动态信息(常变)
- 最近话题
- 最近购买
- 最近反馈

类型 3:派生信息(AI 推断)
- 兴趣偏好
- 行为模式
- 情感倾向

七、5 大记忆策略实战

策略 1:上下文记忆(Contextual Memory)

# 短期 + 长期结合
def get_full_context(user_id, current_message):
    """获取完整上下文"""
    # 1. 短期:当前对话最近 10 条
    short_term = chat_history[user_id][-10:]

    # 2. 长期:相关历史对话
    long_term = vector_db.search(
        query=current_message,
        filter={"user_id": user_id},
        limit=5
    )

    # 3. 用户画像
    user_profile = user_db.get_profile(user_id)

    # 4. 合并
    return {
        "short_term": short_term,
        "long_term": long_term,
        "user_profile": user_profile,
        "current_message": current_message
    }

策略 2:任务记忆(Task Memory)

class TaskMemory:
    """单次任务的记忆"""
    def __init__(self, task_id):
        self.task_id = task_id
        self.redis = Redis()
        self.key = f"task:{task_id}"

    def save(self, step, result):
        self.redis.hset(self.key, step, result)
        self.redis.expire(self.key, 3600)  # 1 小时

    def get_all(self):
        return self.redis.hgetall(self.key)

策略 3:共享记忆(Shared Memory)

# 多 Agent 协作时,共享记忆
shared_memory = SharedMemory()

agent1 = Agent(shared_memory)
agent2 = Agent(shared_memory)

agent1.write("plan", "已完成需求分析")
agent2.read("plan")  # 读取 agent1 的记忆

策略 4:链式记忆(Chain Memory)

# 把多个 Agent 的记忆串联
chain = MemoryChain()
chain.add(memory1)
chain.add(memory2)
chain.add(memory3)

# 查询
context = chain.query("用户的核心需求是什么")

策略 5:语义记忆(Semantic Memory)

# 把记忆按语义分块存储
semantic_memory = SemanticMemory()

# 存储
semantic_memory.add("用户喜欢喝咖啡")
semantic_memory.add("用户每天早上 8 点起床")

# 检索
results = semantic_memory.search("用户的生活习惯")
# 返回:喝咖啡、早上 8 点起床 等

八、6 大常见错误

错误 1:记忆不分类

错:所有记忆混在一起
- 短期 + 长期 + 用户 + 任务 → 全部存向量库

对:分类存储
- 短期:内存
- 长期:向量库 + 关系数据库
- 用户:PostgreSQL
- 任务:Redis

错误 2:记忆不遗忘

错:所有记忆永久保留 → 性能下降 + 成本增加

对:记忆有"半衰期"
- 重要记忆:永久
- 普通记忆:30 天
- 临时记忆:7 天

错误 3:记忆不可解释

错:AI 凭直觉决定记什么

对:记忆决策可解释
- 显式规则(关键词 / 重要性评分)
- 人工审核
- 用户控制("忘记这个")

错误 4:记忆被攻击

风险:恶意用户通过输入污染记忆

防护:
- 记忆写入前过滤(XSS / 注入)
- 重要性评分(防止垃圾记忆)
- 定期清理(垃圾记忆清除)

错误 5:记忆太详细

错:记录"用户说了一句话的所有 50 个字"

对:记录关键信息
- "用户喜欢咖啡" ✅
- "用户说:'我每天早上 8 点起床,喝一杯美式咖啡,加 2 颗糖...'"  ❌

错误 6:记忆不一致

问题:用户的偏好在不同对话中不同
- 3 月:用户说喜欢红色
- 5 月:用户说喜欢蓝色
- 8 月:记忆说喜欢红色

解决:记忆版本控制
- 保留最新版本
- 记录变化历史
- 必要时人工确认

九、6 个月落地路线图

Month 1-2:基础架构
□ 选 L1 短期 + L3 长期
□ 用 LangChain Memory 实现
□ 测试基本流程

Month 3-4:优化检索
□ 升级到 Qdrant / Milvus
□ 优化向量索引
□ 实现 RAG

Month 5:用户画像
□ 实现 L4 用户记忆
□ 跨设备同步
□ 用户控制(删除 / 编辑)

Month 6:遗忘机制
□ 实现 Ebbinghaus 遗忘
□ 性能优化
□ 监控告警

十、给不同角色的建议

Agent 开发者

推荐:LangChain + Qdrant + Redis
时间:3-6 月
复杂度:中
→ 先做 L1 + L3
→ 6 个月加 L4

企业 AI

推荐:PostgreSQL + Milvus + Redis
时间:6-12 月
复杂度:高
→ 自建 L1 + L3 + L4
→ 安全合规

AI 创业者

推荐:Mem0 / LangGraph / Letta
时间:3-6 月
复杂度:低
→ 使用现成框架
→ 快速上线

反思:记忆是 AI 时代的"数据库"

我之前讲 Agent 基础时提过"短期记忆"。这次写完整版时想说: AI 时代记忆 = 互联网时代的数据库。

  • 没有数据库 → 没有 Web 应用
  • 没有记忆 → 没有 AI Agent

未来 5 年:

掌握记忆系统 = 掌握 AI 时代的话语权。

立即开始:从 LangChain Memory 开始,6 个月后你的 Agent 会有真正的"记忆"。