Back返回

GLOVE: Global Verifier for LLM Memory-Environment Realignment

Published:

NeurIPS 2026

GLOVE

GLOVE:面向大模型记忆—环境重对齐的全局验证器GLOVE: Global Verifier for LLM Memory-Environment Realignment

Xingkun Yin, Hongyang Du

The Problem问题所在

Memory-augmented LLM agents keep an external bank of past experience and retrieve from it to decide what to do next. That works well as long as the world stays the way it was when the memory was written. Real deployments do not.带记忆的大模型智能体会维护一个外部经验库,从中检索来决定下一步做什么。只要世界还停留在记忆写下时的样子,这套做法就没问题。但真实部署不是这样。

Implicit semantic drift in WebShop: the definition of a warm colour shifts from yellow to red
Figure 1. The drift that leaves no trace. In WebShop the attribute "warm colour" is silently redefined from yellow to red: the page still looks like a product page, yet every stored step that trusted the old definition is now wrong, and nothing in the observation says so.图 1。不留痕迹的漂移。WebShop 里「暖色」这个属性的定义被悄悄从黄色改成红色:页面看上去仍是商品页,但每一条信任旧定义的记忆步骤都已经错了 —— 而观测里没有任何东西提示这一点。

Existing methods decide whether a memory entry is still valid in one of two ways: they wait for an external evaluator to hand them a task-specific success signal, or they ask the model to reflect on its own memory. Both assumptions break down once the environment drifts. A web shop changes its search semantics, a maze's topology is rebuilt, a reward is reversed, a target object is swapped — and the agent keeps retrieving and reusing steps that are no longer true.现有方法判断一条记忆是否仍然有效,无非两条路:要么等外部评估器给出针对该任务的成功信号,要么让模型对自己的记忆做反思。一旦环境发生漂移,这两个前提都会失效。网店的搜索语义变了、迷宫的结构重建了、奖励被反转了、目标物体被换掉了 —— 而智能体仍在反复检索、复用那些已经不成立的步骤。

The failure is not gradual. Memory systems that look excellent before drift can collapse afterwards: in the authors' experiments, Voyager drops from 85% to 0% under WebShop semantic drift, and a Generative Agent drops from 80% to 0% under a FrozenLake topology change. Time-based forgetting helps only slowly and can hold on to high-confidence invalid memories. Worse, some drifts are invisible: the observation space looks unchanged while the consequence of an action flips, so nothing in the immediate state tells the agent its memory has gone stale.这种失效不是渐进的。漂移前表现优异的记忆系统,漂移后可能直接崩掉:论文实验中,Voyager 在 WebShop 语义漂移下从 85% 掉到 0%,Generative Agent 在 FrozenLake 拓扑变化下从 80% 掉到 0%。基于时间的遗忘只能缓慢起效,还可能把高置信度的错误记忆一直留着。更麻烦的是,有些漂移是看不见的:观测空间看起来没变,但一个动作的后果已经反转,于是当前状态里没有任何线索告诉智能体「你的记忆过期了」。

Key Idea核心思路

The paper's move is to stop treating retrieved memory as ground truth and start treating it as a hypothesis about how the environment behaves — one that has to be tested by acting.论文的做法是:不再把检索到的记忆当作事实,而是把它当作关于「环境如何运作」的假设 —— 一个必须靠行动去检验的假设。

The insight核心洞见

Memory validity is not something you can read off a memory entry or obtain from a label. It has to be constructed from fresh interaction evidence — a relative notion of truth.记忆是否有效,既读不出、也拿不到标签;它只能由新鲜的交互证据构建出来 —— 这就是「相对真理」。

GLOVE therefore watches what actually happens. After an action, it compares the observed outcome against the outcomes stored for the same state-action precondition. When the two disagree, the memory is not silently trusted: it becomes a claim to be verified.因此 GLOVE 盯着实际发生了什么。每次动作之后,它把观测到的结果与同一「状态–动作」前提下记录的历史结果做比较。两者不一致时,这条记忆不会被默默信任,而是变成一条待验证的主张。

Because the comparison is distributional rather than a single mismatch, one surprising outcome in a stochastic environment does not trigger a false alarm.由于比较是分布层面的、而不是单次不一致,随机环境里的偶发异常不会触发误报。

Three memory validation paradigms: internal cognition, external ground truth, and GLOVE's relative truth
Figure 2. The design space GLOVE opens. Memory has so far been validated either against the model's own cognition (MemoryBank, Generative Agents, MemGPT, G-Memory) or against an external ground-truth evaluator (Voyager, Memory-R1). Both assumptions break under drift; checking memory against the environment itself is a third dimension.图 2。GLOVE 打开的设计空间。以往判断记忆是否有效,要么靠模型自己的认知(MemoryBank、Generative Agents、MemGPT、G-Memory),要么靠外部真值评估器(Voyager、Memory-R1)。这两个前提在漂移下都会失效;把记忆拿去和环境本身对照,是第三个维度。

How It Works方法

GLOVE inside the agent–environment loop
Figure 3. GLOVE sits inside the ordinary agent–environment loop. Retrieved experience guides action selection as usual; after each action, the observed outcome is checked against the stored outcomes for the same precondition, and a discrepancy triggers verification rather than reuse.图 3。GLOVE 嵌入在常规的「智能体–环境」循环中。检索到的经验照常指导动作选择;每次动作后,实际结果会与同一前提下的历史结果比对,一旦出现不一致,触发的是验证而不是复用。
  1. Detect cognitive dissonance. For the current state-action precondition, GLOVE retrieves the counterpart experiences recorded earlier and compares the fresh outcome with their outcome distribution. A discrepancy means the response distribution that produced the stored experience no longer matches the current environment.检测认知失调。针对当前「状态–动作」前提,GLOVE 取出此前记录的对应经验,把新观测结果与它们的结果分布做比较。出现不一致,说明当初产生这条经验的环境响应分布,已经与当前环境不符。
  2. Probe actively. On a discrepancy, the agent re-executes the queried action under the same state context to collect fresh outcomes immediately. This is synchronous verification, and in deterministic settings a single probe settles the question.主动探测。一旦发现不一致,智能体在相同状态下重新执行该动作,立刻收集新鲜结果。这是同步验证;在确定性环境中,一次探测就足以定论。
  3. Or accumulate evidence asynchronously. When exact replay is unavailable, GLOVE suppresses the contradicted memory from future retrieval, marks the precondition as pending verification, and gathers evidence whenever a matching state is naturally re-encountered.或者异步积累证据。当无法精确重放时,GLOVE 把被证伪的记忆从后续检索中屏蔽,将该前提标记为「待验证」,并在自然再次遇到相同状态时持续收集证据。
  4. Realign the bank. Obsolete entries are replaced by the verified ones, so the agent's memory tracks the environment rather than the environment's history.重对齐记忆库。过期的条目被验证后的条目替换,于是智能体的记忆跟随的是环境本身,而不是环境的历史。

Crucially, none of this needs task-specific ground-truth labels, and none of it relies on the model introspecting about its own memory — the verification signal is observable interaction evidence.关键在于:整个过程既不需要任务级的真值标签,也不依赖模型对自身记忆的内省 —— 验证信号完全来自可观测的交互证据。

Results实验结果

GLOVE is evaluated as a drop-in module on top of existing memory designs, across three benchmarks and a real robot: WebShop for web navigation, FrozenLake for discrete planning, MountainCar for continuous control, and a HiWonder JetRover navigating a physical maze.GLOVE 作为即插即用模块叠加在既有记忆架构之上评测,覆盖三个基准与一台真实机器人:WebShop(网页导航)、FrozenLake(离散规划)、MountainCar(连续控制),以及一台在实体迷宫中导航的 HiWonder JetRover。

Every setting is run twice — once as a source environment, then again after a controlled drift. Explicit drift changes something observable (layout, transition dynamics, semantic mapping); implicit drift changes hidden task logic such as reward reversals, leaving immediate observations looking the same. Baselines cover No Memory, a plain RAG agent, and the Voyager, MemoryBank and Generative Agents memory architectures. Backbones span Llama-3.1-8B, Llama-3.3-70B, Qwen2.5-7B, Qwen3-30B, DeepSeek-V3.2, GPT-4o and Grok-3.每个场景跑两遍 —— 先作为源环境,再在受控漂移之后重跑。显式漂移改变可观测的东西(布局、转移动态、语义映射);隐式漂移改变隐藏的任务逻辑,例如奖励反转,而即时观测看起来毫无变化。基线包括 No Memory、普通 RAG 智能体,以及 Voyager、MemoryBank、Generative Agents 三种记忆架构。骨干模型覆盖 Llama-3.1-8B、Llama-3.3-70B、Qwen2.5-7B、Qwen3-30B、DeepSeek-V3.2、GPT-4o 与 Grok-3。

85% → 0%the collapse GLOVE fixes: Voyager under WebShop semantic driftGLOVE 修复的崩塌:Voyager 在 WebShop 语义漂移下
~90%WebShop success recovered after semantic drift, against 20% for MemoryBank语义漂移后 WebShop 成功率恢复水平,MemoryBank 仅为 20%
47.5 → 97.5Vanilla agent under FrozenLake hidden drift, once GLOVE is added加入 GLOVE 后,Vanilla 智能体在 FrozenLake 隐藏漂移下的表现
+65 ptaverage gain over non-augmented agents on FrozenLake Drift II在 FrozenLake Drift II 上相对未增强智能体的平均提升
Demo (3:23, silent). The real-robot experiment end to end: the JetRover runs on stale memory and keeps heading for the old target; once the target is swapped, GLOVE detects the mismatch, realigns the entry and recovers. The explanatory captions are burned into the video.演示(3 分 23 秒,无声)。实体机器人实验的完整过程:JetRover 带着过期记忆反复走向旧目标;目标被调换后,GLOVE 察觉到不一致、重对齐记忆条目并恢复。说明字幕已烧录在画面中。
Success rate and memory conflicts across the source environment and two drift phases
Figure 4. Verification is spent where it matters. Success rate (left axis) and memory conflicts (right axis) across the source environment and two drift phases: conflicts spike at each drift and GLOVE triggers verification only around them, so MemoryBank recovers within a few episodes instead of staying collapsed.图 4。验证只花在需要的地方。源环境与两个漂移阶段的成功率(左轴)与记忆冲突(右轴):冲突在每次漂移时集中爆发,GLOVE 只在冲突附近触发验证,于是 MemoryBank 在几个回合内就恢复,而不是一直塌着。

Against adaptive-memory baselines与自适应记忆基线的对比

Adding GLOVE to a memory architecture is one thing; the harder question is whether it beats simply forgetting more cleverly. Table 1 compares GLOVE directly against adaptive-memory policies on Qwen3-30B under FrozenLake explicit drift. Every method reaches 20/20 in the source environment, so the post-drift columns are the whole story.给记忆架构加上 GLOVE 是一回事;更难的问题是:它能否胜过「更聪明地遗忘」?表 1 在 Qwen3-30B、FrozenLake 显式漂移下,把 GLOVE 与各类自适应记忆策略直接对比。所有方法在源环境中都达到 20/20,所以漂移后的两列才是全部信息。

Table 1: Direct comparison against adaptive-memory policies (Qwen3-30B, FrozenLake explicit drift). Natural adaptation uses 20 episodes per drift phase; the last column reports completed successes under the same post-drift interaction budget of B = 310.表 1:与自适应记忆策略的直接对比(Qwen3-30B,FrozenLake 显式漂移)。自然适应为每个漂移阶段 20 个回合;最后一列为在相同漂移后交互预算 B = 310 下完成的成功次数。

MethodDrift IDrift IISuccesses @ B = 310
GLOVE16/2015/2031
LatestOverwrite16/200/2016
RecentK-K316/200/2016
SlidingWindow-W514/200/2014
RecencyWeighted-H512/200/2012
PageHinkleyReset14/200/2014
PeriodicRefresh-S305/207/2018

What the numbers say这些数字说明了什么

  1. Forgetting on a schedule is not the same as verifying. Under Drift II every adaptive-memory baseline scores 0/20 — discarding memory by recency or change detection throws away valid entries along with invalid ones. GLOVE is the only method that recovers, at 15/20, because it replaces the specific entries that fresh evidence contradicts rather than resetting the bank.按时序遗忘 ≠ 验证。在 Drift II 下,所有自适应记忆基线都是 0/20 —— 按时间新旧或变化检测丢记忆,会把有效条目和无效条目一起扔掉。GLOVE 是唯一恢复的方法(15/20),因为它替换的是被新证据证伪的具体条目,而不是把整个记忆库重置。
  2. It works when the drift cannot be seen. Implicit drift leaves observations superficially consistent, which is exactly when memory becomes harmful: MemoryBank loses 18.8% on WebShop and its forgetting mechanism does not help. GLOVE recovers by checking the observable consequences of actions rather than the appearance of states.漂移看不见时它也能工作。隐式漂移让观测表面上保持一致,而这恰恰是记忆变得有害的时候:MemoryBank 在 WebShop 上掉了 18.8%,它的遗忘机制帮不上忙。GLOVE 之所以能恢复,是因为它检查的是动作的可观测后果,而不是状态长什么样。
  3. The recovery is large, not marginal. On FrozenLake hidden reward reversal, adding GLOVE lifts the Vanilla agent from 47.5% to 97.5%. On Grok-3, Voyager with GLOVE reaches 100% success against 62.5% for Voyager alone.恢复幅度很大,不是边际改善。在 FrozenLake 隐藏奖励反转下,加入 GLOVE 把 Vanilla 智能体从 47.5% 拉到 97.5%。在 Grok-3 上,带 GLOVE 的 Voyager 成功率达到 100%,而单独 Voyager 只有 62.5%。
  4. It transfers across designs and backbones. The gains hold across plain retrieval, passive decay, reflective and iterative-refinement memory backends, and across seven open-weight and proprietary LLM backbones — the module is architecture-agnostic rather than tuned to one memory system.跨设计与跨骨干都能迁移。增益在普通检索、被动衰减、反思式与迭代精化式记忆后端上都成立,也横跨七种开源与闭源大模型骨干 —— 该模块与架构无关,而不是针对某一种记忆系统调的。
  5. It survives contact with hardware. On a physical JetRover navigating a maze with irregular walls, displaced targets and added obstacles, static memory keeps returning to the old target location after the target is swapped; GLOVE detects the mismatch, updates the stale entry, and recovers — separating real semantic drift from incidental physical noise.它也扛得住真实硬件。在墙体不规则、目标有位移、还临时加了障碍物的实体迷宫上,目标被调换后静态记忆会一直跑回旧位置;GLOVE 察觉到不匹配、更新过期条目并恢复 —— 把真正的语义漂移与偶然的物理噪声区分开来。

Citation引用

If you find this work useful, please consider citing it:如果这项工作对您有帮助,欢迎引用:

@article{yin2026glove,
  title   = {GLOVE: Global Verifier for LLM Memory-Environment Realignment},
  author  = {Yin, Xingkun and Du, Hongyang},
  journal = {arXiv preprint arXiv:2601.19249},
  year    = {2026}
}

The full manuscript is available as arXiv:2601.19249 (PDF), with code and experiment settings at github.com/NICE-HKU/GLOVE.完整手稿见 arXiv:2601.19249(PDF),代码与实验设置见 github.com/NICE-HKU/GLOVE。