|
Download agentframe_four_layer_blueprint.md from ljsysfurry/AgentFrame: direct link, hf CLI and curl.
- Browser
- Download file 15 kB
-
https://huggingface.co/ljsysfurry/AgentFrame/resolve/main/agentframe_four_layer_blueprint.md
- Command line
-
hf download hf://ljsysfurry/AgentFrame/agentframe_four_layer_blueprint.md
-
curl -L -o agentframe_four_layer_blueprint.md https://huggingface.co/ljsysfurry/AgentFrame/resolve/main/agentframe_four_layer_blueprint.md
15 kB
AgentFrame 四层融合蓝图
长上下文 Agent 终极方案:认知 × 路由 × 压缩 × 分页
版本: v1.0
日期: 2026-08-09
状态: 设计蓝图(待实施)
许可证: GPL-3.0
1. 设计动机
1.1 长上下文 Agent 的四个瓶颈
| 瓶颈 | 问题 | 现有方案 | 局限 |
|---|---|---|---|
| 认知 | 不知道什么该记 | 全量塞上下文 | 上下文爆炸 |
| 路由 | 不知道什么该读 | 全注意力扫描 | O(N²) 计算 |
| 存储 | KV 存不下 | 展开存储 | 显存线性爆炸 |
| 物理 | 显存放不下 | 全留显存 | 容量封顶 |
1.2 核心洞察
四个瓶颈是正交的——各自独立解决,可以全量叠加:
认知层 (Cognitive Workspace 思路) → 决定"记什么"
路由层 (HiLS-Attention 思路) → 决定"读什么"
存储层 (吸收式 MLA + 量化, 自研) → 决定"存多小"
物理层 (ICPE 思路) → 决定"放哪"
叠加效果(理论估算):
| 层 | 贡献 | 效果 |
|---|---|---|
| 存储层 | 270KB → 7.6KB | 35.6× 存储减 |
| 路由层 | 只读 top-K 块 | ~13× 计算减 |
| 物理层 | 冷 KV 换出 | 显存占用再降 10-100× |
| 认知层 | 只放必要信息 | 上下文需求减 3-10× |
综合:单卡 L40S 从 138 万 token → 理论 1 亿+ token 等效上下文
2. 四层架构总览
┌─────────────────────────────────────────────────────────────┐
│ Agent Loop (编排层) │
│ AgentFrame Orchestrator (现有, 增强) │
└──────┬──────────────────────┬──────────────────────┬────────┘
│ │ │
▼ ▼ ▼
┌─────────────┐ ┌──────────────────┐ ┌─────────────────┐
│ 认知层 │ │ 路由层 │ │ 存储层 │
│ MetaCog │ │ LandmarkRouter │ │ AbsorbedMLA │
│ 控制器 │ │ 块索引 + 分层软max│ │ 缓存 + 量化 │
└──────┬──────┘ └────────┬─────────┘ └────────┬────────┘
│ │ │
└────────────┬───────┴──────────┬───────────┘
▼ ▼
┌──────────────────────────────────┐
│ 物理层 │
│ KV Pager (热/温/冷分层) │
│ 显存 ↔ 内存 ↔ 磁盘 (mmap) │
└──────────────────────────────────┘
3. 层 1:认知层(MetaCog Controller)
3.1 来源
Cognitive Workspace (tao-hpu/cognitive-workspace, arXiv:2508.13171)
3.2 职责
决定哪些信息需要进入 KV 缓存,模仿人类认知的元认知控制。
3.3 组件设计
class MetaCogController:
"""
认知层控制器:任务分解 + 置信度追踪 + 信息缺口分析
"""
def __init__(self, confidence_threshold=0.7):
self.confidence_threshold = confidence_threshold
self.task_stack = [] # 任务分解栈
self.info_gaps = set() # 已知信息缺口
self.buffer_promotions = [] # 记忆提升记录
def decompose_task(self, task: str) -> list[SubTask]:
"""
任务分解:把复杂任务拆成子任务
每个子任务标注需要的信息类型
"""
...
def track_confidence(self, response) -> float:
"""
置信度追踪:模型输出时估算置信度
低置信度 → 触发信息检索
"""
...
def analyze_gap(self, task: SubTask, context) -> list[Query]:
"""
信息缺口分析:识别完成任务还缺什么信息
返回定向检索 query(比盲目 top-k 聪明)
"""
...
def promote_to_longterm(self, buffer_item):
"""
记忆提升:工作缓冲 → 情景缓冲 → 长期记忆
按 频率 + 重要性 + 关联度 打分
"""
...
3.4 与存储层的接口
# 认知层输出 → 存储层输入
class RetrievalDirective:
"""认知层告诉存储层:这次推理需要哪些 KV"""
required_chunks: list[str] # 需要的块 ID
priority: Literal["hot", "warm", "cold"]
reason: str # 为什么需要(可审计)
confidence: float # 认知层置信度
4. 层 2:路由层(LandmarkRouter)
4.1 来源
HiLS-Attention (Tencent-Hunyuan/HiLS-Attention, arXiv:2607.02980)
4.2 职责
决定读哪些 KV 块——用 landmark 摘要做稀疏检索,避免全量扫描。
4.3 组件设计
class LandmarkRouter:
"""
路由层:landmark 块索引 + 分层软max + 低秩校准
"""
def __init__(self, chunk_size=64, top_k=32, num_layers=27):
self.chunk_size = chunk_size
self.top_k = top_k
self.num_layers = num_layers
# landmark token 的可学习 query (每层一个)
self.landmark_queries = nn.Parameter(...) # [L, d_model]
# 低秩校准模块 (Q-Cal, 仅 0.6% 参数)
self.q_cal = LowRankCalibration(r=16)
def build_chunk_summary(self, chunk_hidden_states):
"""
块摘要:k'c = Attn(q'c, Kc, Kc) 的加权和
+ 熵偏置 b'c (捕捉块内质量分布)
"""
# 一阶泰勒展开的 LogSumExp 线性化:
# log Σ exp(q·k/√d) ≈ (q·k'c)/√d + b'c
...
def route(self, query, chunk_summaries) -> list[ChunkSelection]:
"""
路由:打分 top-K 块
分数 = (q̂ᵢ·k'c)/√d + b'c
q̂ᵢ = qᵢ + W_up W_down hᵢ (Q-Cal 校准)
"""
...
def hierarchical_softmax(self, query, selected_chunks):
"""
分层软max:块内归一化 × 块间质量(可学习)
w_ij = [exp(s_ij)/Z_intra] × [Ẑ_chunk/Ẑ_total]
使块选择端到端可训练 (LM loss 梯度直达)
"""
...
4.4 GQA 兼容
def group_select(self, group_scores):
"""
GQA 组内取 max 聚合:
组内各头归一化权重 → 取 max → top-K
保证任一头觉得重要的块都不漏
"""
group_max = np.max(group_scores, axis=0) # [num_chunks]
selected = np.argsort(group_max)[-self.top_k:]
return selected
4.5 与存储层的接口
class ChunkSelection:
"""路由层输出:本次推理要读的块"""
chunk_ids: list[int]
retrieval_scores: list[float]
selected_tokens: list[int] # 展开后的 token 索引
5. 层 3:存储层(AbsorbedMLA + 分层量化)
5.1 来源
自研(L40S 实测:270KB → 7.6KB/token, 35.6×)
5.2 职责
决定KV 存多小——吸收式 MLA 缓存潜在向量 + 分层量化。
5.3 组件设计(已有实现,增强)
class AbsorbedMLAEncoder:
"""
存储层:吸收式 MLA + 分层量化(现有 agentframe_core.py 增强版)
"""
def __init__(self, kv_lora_rank=512, k_rope=64, layers=27):
self.kv_lora_rank = kv_lora_rank
self.k_rope = k_rope
self.layers = layers
self.quantizer = LayeredQuantizer()
def encode(self, hidden_states) -> CompressedKV:
"""
压缩:只缓存潜在向量 (576维) 而非展开 KV
[seq, 512+64] → 30.4KB/token (8.9×)
"""
...
def quantize(self, kv, kind: Literal["thought", "tool"]):
"""
分层量化:
思考链 → INT8 (误差 0.011)
工具结果 → INT4 (误差 0.079)
per-channel 非对称 (32块, 27×精度提升)
"""
...
def fuse_router(self, router: LandmarkRouter):
"""
融合点 1:路由层直接作用于压缩表示
landmark 摘要基于潜在向量构建(不展开)
→ 路由成本 O(N/S) 而非 O(N)
"""
...
5.4 与物理层的接口
class CompressedKV:
"""存储层输出:压缩后的 KV 块"""
latent: np.ndarray # [chunk, 576] 潜在向量
k_pe: np.ndarray # [chunk, 64] 位置编码
quant_bits: int # 16/8/4
chunk_ids: list[int]
heat: float # 热度(物理层用)
last_access: float # 最后访问时间
6. 层 4:物理层(KV Pager)
6.1 来源
ICPE (matheusdelgado/infinite-context) 思想 + 自研轻量实现
6.2 职责
决定 KV 放哪——热留显存、温放内存、冷换磁盘(mmap)。
6.3 组件设计
class KVPager:
"""
物理层:KV 换页引擎(零拷贝, mmap)
显存 → 内存 → 磁盘 三级分层
"""
def __init__(self, vram_limit_gb=10.0, ram_limit_gb=32.0):
self.vram = LRUCache(limit=vram_limit_gb) # 热 KV
self.ram = LRUCache(limit=ram_limit_gb) # 温 KV
self.disk = mmap_store() # 冷 KV (mmap)
self.predictor = PredictiveEvictor() # 预测性驱逐
def place(self, kv: CompressedKV):
"""
放置:按热度决定层级
heat > 0.8 → VRAM
0.4 < heat < 0.8 → RAM
heat < 0.4 → DISK
"""
...
def prefetch(self, directive: RetrievalDirective):
"""
预取:推理前纳秒级把需要的冷 KV 换回
用注意力信号预测(非简单 LRU)
"""
...
def evict(self, policy="predictive"):
"""
驱逐:预测性驱逐(注意力信号驱动)
优于 LRU/FIFO
"""
...
6.4 预测性驱逐算法
class PredictiveEvictor:
"""
预测性驱逐:用注意力分数预测未来访问
不是"最久未用",而是"最不可能再用"
"""
def score_eviction(self, kv: CompressedKV) -> float:
# 特征: 注意力分数衰减率 + 访问频率 + 语义相似度
attn_decay = self.attention_decay(kv) # 注意力分数随时间衰减
access_freq = self.access_frequency(kv) # 历史访问频率
semantic = self.semantic_relevance(kv) # 与当前任务的语义距离
return 0.5*attn_decay + 0.3*access_freq + 0.2*semantic
7. 四层集成流程
7.1 推理前(认知 → 路由 → 物理)
Agent 收到任务
│
▼
[MetaCog] 分解任务 → 识别信息缺口 → 生成 RetrievalDirective
│
▼
[KVPager] 按 directive 预取冷 KV → 换入显存/内存
│
▼
[LandmarkRouter] 对显存中的 KV 块构建 landmark 摘要 → 打分 top-K
│
▼
[AbsorbedMLA] 展开选中的压缩块 → 计算注意力
7.2 推理后(存储 → 物理 → 认知)
注意力计算完成
│
▼
[AbsorbedMLA] 新 KV 压缩为潜在向量 → 分层量化
│
▼
[KVPager] 按热度放置 (新 KV = 热 → VRAM)
│
▼
[MetaCog] 置信度追踪 → 低置信度触发补检索
重要信息 → 提升到长期记忆
7.3 数据流伪代码
def agent_inference(agent, task):
# 认知层
directive = agent.metacog.decompose(task)
# 物理层预取
agent.pager.prefetch(directive)
# 路由层
selections = agent.router.route(
query=agent.current_query,
chunk_summaries=agent.store.build_summaries()
)
# 存储层
kv = agent.store.gather(selections)
output = agent.model.forward(agent.current_query, kv)
# 后处理
agent.store.encode_new_kv(output)
agent.pager.place(output.new_kv)
agent.metacog.track_confidence(output)
return output
8. 实施路线图
Phase 1:存储层强化(已有基础,2 周)
- 吸收式 MLA 缓存(已完成, 8.9×)
- INT8/INT4 分层量化(已完成, 35.6×)
- 增加块级组织(chunk 化,对齐路由层)
Phase 2:物理层(4 周)
- KV Pager 实现(mmap 冷存储)
- 热度追踪 + LRU
- 预测性驱逐 v1(注意力衰减)
Phase 3:路由层(6 周)
- Landmark 摘要构建(基于压缩潜在向量)
- 分层软max 前向实现
- Q-Cal 低秩校准
- GQA 组内 max 聚合
Phase 4:认知层(4 周)
- MetaCog 控制器(任务分解 + 置信度)
- 信息缺口分析
- 记忆提升策略
Phase 5:集成验证(4 周)
- 四层端到端集成
- 真实 DeepSeek-V2-Lite 加载
- 长上下文基准(RULER / LongBench)
- 与 HiLS 对比实验
总计:约 20 周
9. 风险与对策
| 风险 | 等级 | 对策 |
|---|---|---|
| 四层叠加延迟累积 | 中 | 每层接口零拷贝;路由层直接在压缩表示上操作 |
| 路由层稀疏性导致召回下降 | 中 | GQA max 聚合 + top-K 可调 + 兜底滑窗 |
| 认知层误判(该记的没记) | 高 | 置信度阈值可调 + 审计日志 + 失败回退全量 |
| 预测性驱逐误判 | 中 | 双写(驱逐前保留磁盘副本) |
| 物理层 mmap 延迟 | 低 | 预取提前量可调(纳秒级 → 毫秒级) |
10. 预期指标
| 指标 | 当前 | 四层融合后 |
|---|---|---|
| KV/token | 7.6 KB | 7.6 KB(存储不变, 但访问量降) |
| 单卡上下文 | 138 万 token | 等效 1 亿+ token(认知减负×路由稀疏) |
| 推理延迟 | 线性 | 亚线性(top-K 常数访问) |
| 显存占用 | 10 GB | 2-4 GB(冷 KV 换出) |
| 质量 | 无损 | 无损(分层软max 可端到端调优) |
11. 附录:参考项目
| 项目 | 来源 | 在本蓝图的角色 |
|---|---|---|
| HiLS-Attention | arXiv:2607.02980 | 路由层(分层软max) |
| Cognitive Workspace | arXiv:2508.13171 | 认知层(元认知控制) |
| InfiniteVL | arXiv:2512.08829 | 参考(线性架构, 多模态扩展) |
| ICPE | github.com/matheusdelgado | 物理层(KV 换页思想) |
| 自研 (AgentFrame) | 本仓库 | 存储层(吸收式 MLA + 量化) |
Cloud LTE Studio · 2026-08-09 · GPL-3.0