AI 智能体记忆系统设计:从短期记忆到长期记忆的完整架构
# AI 智能体记忆系统设计:从短期记忆到长期记忆的完整架构 > AI Agent 没有记忆,**就像失忆的医生**——每次看病都从零开始。本文讲 AI Agent 的**记忆系统设计**:从短期记忆(短期上下文)到长期记忆(向量数据库)的完整架构。 > **读完保证**:能设计 4 层记忆架构 + 5 大记忆策略 + 6 个月落地。 ## 一、为什么
AI 智能体记忆系统设计:从短期记忆到长期记忆的完整架构
AI Agent 没有记忆,就像失忆的医生——每次看病都从零开始。本文讲 AI Agent 的记忆系统设计:从短期记忆(短期上下文)到长期记忆(向量数据库)的完整架构。
读完保证:能设计 4 层记忆架构 + 5 大记忆策略 + 6 个月落地。
一、为什么 AI Agent 需要记忆?
1.1 当前 AI Agent 的"失忆症"
场景:AI 客服 Agent
第 1 次对话:
用户:我姓王
AI:好的王先生,请问您的问题?
用户:我的订单 #123 物流很慢
AI:正在查询您的订单...物流正常
第 2 次对话(5 分钟后):
用户:订单怎么样了?
AI:您好,请提供您的订单号? ← 失忆了
问题:每次对话 = 从零开始。用户重复信息,体验差。
1.2 真实场景需要"记忆"
| 场景 | 没有记忆 | 有记忆 |
|---|---|---|
| AI 客服 | 重复问订单号 | 记住订单状态 |
| AI 编程助手 | 每次重新理解项目 | 记住项目结构 |
| AI 销售 | 不知道客户历史 | 知道客户购买历史 |
| AI 陪伴 | 每次都是陌生人 | 了解你的性格 |
| AI 研究助手 | 不知道你研究过什么 | 累积你的研究 |
1.3 2026 年记忆市场
2024 年:实验阶段(LangChain Memory 雏形)
2025 年:$500M 商业化
2026 年:$5B 爆发
2028 年:$50B+(成为 AI 基础设施)
→ 2026 年记忆 = AI 时代"数据库"
二、4 层记忆架构
┌────────────────────────────────────────────────────┐
│ AI Agent 4 层记忆架构 │
│ │
│ ┌────────────────────────────────────┐ │
│ │ L1:短期记忆(Working Memory) │ │
│ │ 存储:LLM 上下文窗口 │ │
│ │ 容量:4K-200K token │ │
│ │ 持续:单次会话 │ │
│ └──────────┬─────────────────────────┘ │
│ ↓ │
│ ┌────────────────────────────────────┐ │
│ │ L2:工作记忆(Scratchpad) │ │
│ │ 存储:Agent 内部状态 │ │
│ │ 容量:1K-10K token │ │
│ │ 持续:单次任务 │ │
│ └──────────┬─────────────────────────┘ │
│ ↓ │
│ ┌────────────────────────────────────┐ │
│ │ L3:长期记忆(Long-term Memory) │ │
│ │ 存储:向量数据库 + SQL │ │
│ │ 容量:100K-10M 条记录 │ │
│ │ 持续:永久 │ │
│ └──────────┬─────────────────────────┘ │
│ ↓ │
│ ┌────────────────────────────────────┐ │
│ │ L4:用户/全局记忆(Profile) │ │
│ │ 存储:PostgreSQL + 向量 │ │
│ │ 容量:每用户 1K-100K 条 │ │
│ │ 持续:永久 + 跨设备 │ │
│ └────────────────────────────────────┘ │
└────────────────────────────────────────────────────┘
三、L1:短期记忆(Working Memory)
3.1 设计原理
短期记忆 = LLM 上下文窗口
数据:当前对话的所有消息
存储:在 LLM API 调用时直接传
持续:单次会话(用户离开就清空)
容量:4K-200K token(取决于模型)
GPT-4o:128K
Claude 3.5:200K
Gemini 1.5:2M
Qwen2.5:128K
DeepSeek:64K
3.2 优化短期记忆
策略 1:滑动窗口(保留最近 N 条)
class ShortTermMemory:
def __init__(self, max_messages=20):
self.messages = []
self.max = max_messages
def add(self, role, content):
self.messages.append({"role": role, "content": content})
# 只保留最近 N 条
if len(self.messages) > self.max:
self.messages = self.messages[-self.max:]
def get(self):
return self.messages
策略 2:摘要压缩(超长时压缩)
def compress_if_needed(messages, model, max_tokens=8000):
"""如果消息超长,压缩早期消息"""
if count_tokens(messages) > max_tokens:
# 保留最近 5 条原样
recent = messages[-5:]
# 早期消息压缩成 1 条
old = messages[:-5]
summary = model.invoke([
{"role": "system", "content": "请用 200 字总结以下对话"},
{"role": "user", "content": format_messages(old)}
])
return [{"role": "assistant", "content": f"[历史摘要] {summary}"}] + recent
return messages
四、L2:工作记忆(Scratchpad)
4.1 设计原理
工作记忆 = Agent 内部状态
数据:当前任务的中间结果
- 思考链(Chain of Thought)
- 工具调用历史
- 子任务状态
- 推理步骤
存储:Agent 内存(Python 对象)
持续:单次任务(任务结束清空)
容量:1K-10K token
4.2 实战代码
class Scratchpad:
"""工作记忆 - 任务执行中"""
def __init__(self):
self.thoughts = [] # 思考链
self.actions = [] # 行动历史
self.observations = [] # 观察结果
self.sub_tasks = [] # 子任务状态
def think(self, thought):
"""Agent 思考"""
self.thoughts.append(thought)
def act(self, action, args):
"""Agent 行动"""
self.actions.append({"action": action, "args": args})
def observe(self, result):
"""观察结果"""
self.observations.append(result)
def get_context(self):
"""获取完整上下文(给 LLM 看)"""
return {
"thoughts": self.thoughts,
"actions": self.actions,
"observations": self.observations,
"sub_tasks": self.sub_tasks,
}
# 使用
sp = Scratchpad()
sp.think("用户问订单状态,需要先查订单号")
sp.act("query_order", {"order_id": 123})
result = query_database("123")
sp.observe(result)
sp.think("订单已发货,预计明天到")
五、L3:长期记忆(Long-term Memory)
5.1 设计原理
长期记忆 = 向量数据库 + 结构化数据库
数据:所有历史对话 + Agent 学到的知识
存储:
- 向量数据:Qdrant / Milvus / pgvector
- 结构化:PostgreSQL
- 文本:S3 / OSS
持续:永久(除非主动删除)
容量:100K-10M 条
5.2 实战架构
from qdrant_client import QdrantClient
from qdrant_client.http import models
import openai
import uuid
class LongTermMemory:
def __init__(self):
self.qdrant = QdrantClient(host="localhost", port=6333)
self.openai_client = openai.Client()
self.collection = "agent_memory"
def store(self, content, metadata={}):
"""存储记忆"""
# 1. 向量化
response = self.openai_client.embeddings.create(
model="text-embedding-3-small",
input=content
)
vector = response.data[0].embedding
# 2. 存储到 Qdrant
self.qdrant.upsert(
collection_name=self.collection,
points=[models.PointStruct(
id=str(uuid.uuid4()),
vector=vector,
payload={"content": content, **metadata}
)]
)
def retrieve(self, query, top_k=5):
"""检索相关记忆"""
# 1. 向量化查询
response = self.openai_client.embeddings.create(
model="text-embedding-3-small",
input=query
)
query_vector = response.data[0].embedding
# 2. 检索
results = self.qdrant.search(
collection_name=self.collection,
query_vector=query_vector,
limit=top_k
)
return [
{"content": r.payload["content"], "score": r.score}
for r in results
]
5.3 长期记忆 5 大策略
策略 1:何时存储
def should_store(self, message):
"""判断是否需要存储"""
# 存储条件
if message["role"] == "user":
# 重要信息:偏好、事实、问题
keywords = ["我", "名字", "喜欢", "需要", "想要", "my", "I like", "I need"]
if any(kw in message["content"] for kw in keywords):
return True
return False
策略 2:如何索引
def index_memory(self, content):
"""索引记忆(用于快速检索)"""
# 关键词索引
keywords = extract_keywords(content)
for kw in keywords:
self.keyword_index[kw].append(content)
# 标签索引
tags = extract_tags(content)
for tag in tags:
self.tag_index[tag].append(content)
策略 3:去重 + 压缩
def deduplicate(self, memories):
"""去重相似记忆"""
unique = []
for m in memories:
# 相似度 > 0.95 = 重复
is_dup = False
for u in unique:
sim = self.cosine_similarity(m["vector"], u["vector"])
if sim > 0.95:
# 合并(保留最新的)
u["updated_at"] = m["created_at"]
is_dup = True
break
if not is_dup:
unique.append(m)
return unique
策略 4:遗忘机制
def should_forget(self, memory):
"""判断是否应该遗忘"""
# Ebbinghaus 遗忘曲线
age_days = (now() - memory["created_at"]).days
importance = memory["importance"] # 0-1
# 重要性 × 时间 = 保留概率
keep_prob = importance * (1 - age_days / 365)
return random() > keep_prob # 30% 概率遗忘
策略 5:分层存储
# 热数据:最近 7 天
# 温数据:7-30 天
# 冷数据:> 30 天
# 存储分层
- 热:Redis(快但贵)
- 温:PostgreSQL(中等)
- 冷:S3(慢但便宜)
六、L4:用户 / 全局记忆(Profile)
6.1 设计原理
用户记忆 = 关于这个用户的专属记忆
数据:
- 基本信息(姓名 / 偏好 / 习惯)
- 行为历史(过去 N 天的对话)
- 关系网络(与其他人 / 事务的关系)
- 长期偏好
存储:
- 结构化:PostgreSQL(user_id, key, value)
- 向量:Qdrant(语义检索)
持续:永久 + 跨设备 + 跨平台
6.2 实战代码
class UserMemory:
"""用户专属记忆"""
def __init__(self, user_id):
self.user_id = user_id
self.db = Database()
self.vector_db = QdrantClient()
def get_profile(self):
"""获取用户画像"""
# 1. 从数据库查
rows = self.db.query(
"SELECT key, value FROM user_profile WHERE user_id = %s",
(self.user_id,)
)
profile = {row["key"]: row["value"] for row in rows}
# 2. 从向量库查最近对话
memories = self.vector_db.search(
collection_name=f"user_{self.user_id}_memory",
query_vector=...,
limit=10
)
return {"profile": profile, "memories": memories}
def update(self, key, value):
"""更新用户记忆"""
self.db.execute(
"INSERT INTO user_profile (user_id, key, value, updated_at) "
"VALUES (%s, %s, %s, NOW()) "
"ON CONFLICT (user_id, key) DO UPDATE SET value = %s",
(self.user_id, key, value, value)
)
6.3 用户记忆 3 大类
类型 1:静态信息(很少变)
- 姓名、性别、年龄
- 职业、所在地
- 语言偏好
类型 2:动态信息(常变)
- 最近话题
- 最近购买
- 最近反馈
类型 3:派生信息(AI 推断)
- 兴趣偏好
- 行为模式
- 情感倾向
七、5 大记忆策略实战
策略 1:上下文记忆(Contextual Memory)
# 短期 + 长期结合
def get_full_context(user_id, current_message):
"""获取完整上下文"""
# 1. 短期:当前对话最近 10 条
short_term = chat_history[user_id][-10:]
# 2. 长期:相关历史对话
long_term = vector_db.search(
query=current_message,
filter={"user_id": user_id},
limit=5
)
# 3. 用户画像
user_profile = user_db.get_profile(user_id)
# 4. 合并
return {
"short_term": short_term,
"long_term": long_term,
"user_profile": user_profile,
"current_message": current_message
}
策略 2:任务记忆(Task Memory)
class TaskMemory:
"""单次任务的记忆"""
def __init__(self, task_id):
self.task_id = task_id
self.redis = Redis()
self.key = f"task:{task_id}"
def save(self, step, result):
self.redis.hset(self.key, step, result)
self.redis.expire(self.key, 3600) # 1 小时
def get_all(self):
return self.redis.hgetall(self.key)
策略 3:共享记忆(Shared Memory)
# 多 Agent 协作时,共享记忆
shared_memory = SharedMemory()
agent1 = Agent(shared_memory)
agent2 = Agent(shared_memory)
agent1.write("plan", "已完成需求分析")
agent2.read("plan") # 读取 agent1 的记忆
策略 4:链式记忆(Chain Memory)
# 把多个 Agent 的记忆串联
chain = MemoryChain()
chain.add(memory1)
chain.add(memory2)
chain.add(memory3)
# 查询
context = chain.query("用户的核心需求是什么")
策略 5:语义记忆(Semantic Memory)
# 把记忆按语义分块存储
semantic_memory = SemanticMemory()
# 存储
semantic_memory.add("用户喜欢喝咖啡")
semantic_memory.add("用户每天早上 8 点起床")
# 检索
results = semantic_memory.search("用户的生活习惯")
# 返回:喝咖啡、早上 8 点起床 等
八、6 大常见错误
错误 1:记忆不分类
错:所有记忆混在一起
- 短期 + 长期 + 用户 + 任务 → 全部存向量库
对:分类存储
- 短期:内存
- 长期:向量库 + 关系数据库
- 用户:PostgreSQL
- 任务:Redis
错误 2:记忆不遗忘
错:所有记忆永久保留 → 性能下降 + 成本增加
对:记忆有"半衰期"
- 重要记忆:永久
- 普通记忆:30 天
- 临时记忆:7 天
错误 3:记忆不可解释
错:AI 凭直觉决定记什么
对:记忆决策可解释
- 显式规则(关键词 / 重要性评分)
- 人工审核
- 用户控制("忘记这个")
错误 4:记忆被攻击
风险:恶意用户通过输入污染记忆
防护:
- 记忆写入前过滤(XSS / 注入)
- 重要性评分(防止垃圾记忆)
- 定期清理(垃圾记忆清除)
错误 5:记忆太详细
错:记录"用户说了一句话的所有 50 个字"
对:记录关键信息
- "用户喜欢咖啡" ✅
- "用户说:'我每天早上 8 点起床,喝一杯美式咖啡,加 2 颗糖...'" ❌
错误 6:记忆不一致
问题:用户的偏好在不同对话中不同
- 3 月:用户说喜欢红色
- 5 月:用户说喜欢蓝色
- 8 月:记忆说喜欢红色
解决:记忆版本控制
- 保留最新版本
- 记录变化历史
- 必要时人工确认
九、6 个月落地路线图
Month 1-2:基础架构
□ 选 L1 短期 + L3 长期
□ 用 LangChain Memory 实现
□ 测试基本流程
Month 3-4:优化检索
□ 升级到 Qdrant / Milvus
□ 优化向量索引
□ 实现 RAG
Month 5:用户画像
□ 实现 L4 用户记忆
□ 跨设备同步
□ 用户控制(删除 / 编辑)
Month 6:遗忘机制
□ 实现 Ebbinghaus 遗忘
□ 性能优化
□ 监控告警
十、给不同角色的建议
Agent 开发者
推荐:LangChain + Qdrant + Redis
时间:3-6 月
复杂度:中
→ 先做 L1 + L3
→ 6 个月加 L4
企业 AI
推荐:PostgreSQL + Milvus + Redis
时间:6-12 月
复杂度:高
→ 自建 L1 + L3 + L4
→ 安全合规
AI 创业者
推荐:Mem0 / LangGraph / Letta
时间:3-6 月
复杂度:低
→ 使用现成框架
→ 快速上线
反思:记忆是 AI 时代的"数据库"
我之前讲 Agent 基础时提过"短期记忆"。这次写完整版时想说: AI 时代记忆 = 互联网时代的数据库。
- 没有数据库 → 没有 Web 应用
- 没有记忆 → 没有 AI Agent
未来 5 年:
掌握记忆系统 = 掌握 AI 时代的话语权。
立即开始:从 LangChain Memory 开始,6 个月后你的 Agent 会有真正的"记忆"。
💬 评论 25 条