Back返回

Experience Scaling: Post-Deployment Evolution For Large Language Models

Published:

arXiv preprintarXiv 预印本 · 2025

Experience Scaling

Experience Scaling:大模型的部署后演化Experience Scaling: Post-Deployment Evolution For Large Language Models

Xingkun Yin, Kaibin Huang, Dong In Kim, Hongyang Du

The Problem问题所在

Progress in large language models has come from scaling three things: model size, training data, and compute. All three are running into the same wall. Human-generated text is finite and increasingly exhausted, and the gains from adding more of it are diminishing — which is why attention has turned to synthetic data, with its own risks of reduced diversity.大模型的进展来自对三件事的规模化:模型规模、训练数据与算力。而这三者正撞上同一堵墙。人类生成的文本总量有限、且已接近枯竭,继续加量的边际收益在递减 —— 这正是注意力转向合成数据的原因,而合成数据又带来多样性下降的风险。

There is a second, quieter problem. A deployed model is frozen. Whatever it learned at training time is all it will ever know, no matter how many real interactions it handles afterwards. Every query it answers, every mistake it makes and corrects, every piece of feedback it receives — all of that experience is simply discarded. The model that has served a million users knows no more than the model that served none.还有第二个更安静的问题:部署后的模型是冻结的。训练时学到多少,之后无论处理多少次真实交互,它知道的也只有这些。它回答的每个问题、犯下并纠正的每个错误、收到的每一条反馈 —— 这些经验全都直接丢弃了。服务过一百万用户的模型,和服务过零个用户的模型,知道的东西一样多。

Key Idea核心思路

Experience Scaling adds a fourth axis to the usual three. Alongside model size, training data and compute, it proposes experience as a scaling dimension — with deployment no longer the end of learning but the beginning of a new source of it.Experience Scaling 在常见三条轴之外加了第四条。除模型规模、训练数据与算力之外,它把经验也作为一种可规模化的维度 —— 部署不再是学习的终点,而是一类新学习来源的起点。

The insight核心洞见

Interaction is data. The model's own post-deployment experience is a stream of information that no static corpus contains, because it is generated by the environment the model is actually operating in.交互本身就是数据。模型部署后产生的经验,是任何静态语料都不包含的信息流 —— 因为它由模型真实所处的环境生成。

The key is not to log that stream but to distil it. Raw interaction traces are noisy and unbounded; experience scaling captures them, condenses them into compact reusable knowledge, and periodically refines the store so it stays relevant and efficient rather than growing without limit.关键不在于把这股流「记下来」,而在于把它蒸馏掉。原始交互轨迹既嘈杂又无上界;experience scaling 先收集,再压缩成紧凑可复用的知识,并定期精炼,让经验库保持相关且高效,而不是无限膨胀。

Sharing distilled experience across a network lets the gains compound: what one deployment learns becomes available to related scenarios elsewhere.把蒸馏后的经验在网络内共享,收益就能叠加:一个部署学到的东西,可以被其他地方的相似场景取用。

How It Works方法

Experience scaling framework

Figure 1: The experience scaling framework. Interactions are captured as experience, distilled into reusable knowledge, and refined periodically so that relevance and efficiency are preserved as the store grows.图 1:experience scaling 框架。交互被采集为经验,蒸馏成可复用的知识,并定期精炼 —— 以保证经验库在增长过程中始终保持相关与高效。

  1. Capture the interaction. The framework records raw interactions between the model and its environment — the queries it receives and the responses it produces — instead of treating them as disposable.采集交互。框架记录模型与环境之间的原始交互 —— 收到的查询与产出的回答 —— 而不是把它们当作一次性消耗品丢掉。
  2. Distil them into experience. Raw traces are condensed into compact, reusable knowledge that captures patterns rather than replaying individual episodes, which is what makes the store usable and searchable.蒸馏成经验。原始轨迹被压缩为紧凑、可复用的知识,提取的是模式而非重放单次经历 —— 这正是让经验库可用、可检索的前提。
  3. Refine the store periodically. A server-side refinement cycle distils and compresses accumulated experience so it keeps its relevance and efficiency; refinement is what prevents over-saturation as the store grows.定期精炼经验库。服务端精炼周期对积累的经验再做蒸馏与压缩,使其保持相关与高效;精炼正是防止经验库在增长中过饱和的手段。
  4. Share across the network. Experience collected on one frontend can be retrieved by others, so learned patterns transfer to related but previously unseen scenarios rather than staying local.在网络内共享。某个前端采集到的经验可被其他前端检索,于是学到的模式能迁移到相关但此前未见过的新场景,而不是停留在本地。

Results实验结果

The framework is validated in three simulated real-world scenarios chosen to stress different parts of it: generalisation to previously unseen but related tasks, repetitive queries from many users on the same topic, and knowledge stores that have become over-saturated.框架在三个模拟真实场景中得到验证,分别针对它的不同环节施加压力:迁移到此前未见但相关的任务;大量用户在同一话题上的重复查询;以及已经过饱和的知识库。

The generalisation study uses the High School Biology subset of MMLU, split 80/20 into a train set and a test set, with the experience database built from the train set. Accuracy is then measured on the held-out test set under three conditions: plain inference, inference with memory, and inference with experience.泛化研究使用 MMLU 的高中生物子集,按 80/20 划分为训练集与测试集,经验库由训练集构建。随后在留出的测试集上测量三种条件下的准确率:纯推理、带记忆推理、带经验推理。

0.554average score with experience, against 0.508 for plain inference and 0.514 with memory带经验时的平均分,纯推理为 0.508、带记忆为 0.514
+4.6%accuracy over plain inference on the held-out test set在留出测试集上相对纯推理的准确率提升
+25%best gain from refining an over-saturated store (MMLU High School Geography)过饱和经验库经精炼后的最佳增益(MMLU 高中地理)
0additional human-written data — the experience comes from the model's own interactions额外增加的人工撰写数据 —— 经验全部来自模型自身的交互

Table 1: Average score by inference condition. Plain inference uses the model alone, memory supplies retrieved context, and experience supplies distilled and refined knowledge. Values are read from the paper's comparison of the three conditions.表 1:各推理条件下的平均分。纯推理只用模型本身;记忆提供检索到的上下文;经验提供经蒸馏与精炼的知识。数值取自论文对三种条件的对比。

ConditionAverage score
Plain inference0.508
Inference with memory0.514
Inference with experience0.554

What the numbers say这些数字说明了什么

  1. Experience is not just memory. Supplying retrieved context raises the score from 0.508 to 0.514 — a small gain. Supplying distilled experience raises it to 0.554. The difference is the distillation step: experience is organised in a form native to the model rather than being raw text dropped into the prompt.经验不等于记忆。提供检索到的上下文只把分数从 0.508 抬到 0.514 —— 增益很小。提供蒸馏后的经验则把它抬到 0.554。差别就在蒸馏这一步:经验是按模型「母语」组织起来的,而不是把原始文本塞进提示词。
  2. Post-deployment learning keeps paying. Accuracy rises steadily as interactions accumulate, and the model's self-reported confidence becomes better calibrated too — its confidence scores are consistently higher for correct predictions than for incorrect ones, so it increasingly knows what it knows.部署后的学习持续有回报。随着交互累积,准确率稳步上升;模型自报的置信度也变得更准 —— 答对时的自评置信度持续高于答错时,也就是说它越来越「知道自己知道什么」。
  3. Curation is not optional. An over-saturated store degrades retrieval, and refinement is what restores it: a BM25 threshold of 100 consistently produced gains, peaking at a 25% increase on High School Geography. Compressing too aggressively undoes the benefit — a threshold of 40 over-condenses the experiential knowledge and hurts retrieval quality.整理不是可选项。过饱和的经验库会拖垮检索,而精炼是把它救回来的手段:BM25 阈值取 100 时持续带来增益,在高中地理上最高达到 25% 的提升。压缩过猛则适得其反 —— 阈值取 40 会把经验知识过度浓缩,反而损害检索质量。
  4. There are domains where it backfires, and the paper says so. On College Biology the refined database performs worse than the full one. The explanation is specificity: those queries concentrate on far fewer high-similarity entities (26.28% under a BM25 score of 2, against a 17.81% average across other datasets), so condensation removes detail the queries actually depend on.有些领域它会帮倒忙,论文也如实写了。在大学生物上,精炼后的经验库反而不如完整库。原因是特异性:该数据集的查询高度集中在少数高相似实体上(BM25 分数为 2 时占 26.28%,而其他数据集平均为 17.81%),于是压缩正好删掉了这些查询真正依赖的细节。

Citation引用

If you find this work useful, please consider citing it:如果这项工作对您有帮助,欢迎引用:

@article{yin2025experience,
  title   = {Experience Scaling: Post-Deployment Evolution for Large Language Models},
  author  = {Yin, Xingkun and Huang, Kaibin and Kim, Dong In and Du, Hongyang},
  journal = {arXiv preprint arXiv:2509.18771},
  year    = {2025}
}

The full manuscript is available as arXiv:2509.18771 (PDF).完整手稿见 arXiv:2509.18771(PDF)。