Back返回

Carnator: Fast Text-to-Video Generation with Generation-Native Compatibility-Guided Cross-Request Reuse

Published:

arXiv preprintarXiv 预印本 · 2026

Carnator

Carnator:基于生成原生兼容性引导的跨请求复用加速文本到视频生成Carnator: Fast Text-to-Video Generation with Generation-Native Compatibility-Guided Cross-Request Reuse

Xingkun Yin, Xuebin Tang, Mingkun Xu, Hongyang Du

The Problem问题所在

Generating a video with a diffusion model is expensive. The model starts from noise and denoises the clip step by step, over a long sequence of spatiotemporal tokens, so a single request can keep a GPU busy for seconds to minutes. Serving video generation at scale means paying that cost over and over.用扩散模型生成一段视频很贵。模型从噪声出发、沿着一长串时空 token 逐步去噪,一个请求就能让 GPU 忙上数秒到数分钟。而要对外提供服务,这份开销要一遍又一遍地付。

There is an obvious source of savings hiding in the traffic. Real workloads are repetitive: many requests ask for broadly similar things — the same kind of scene, the same objects, the same layout — so the computation behind one finished request could help answer a later one. Reusing work across requests is called cross-request reuse.而真实流量里藏着一个显而易见的省算力机会:请求是重复的。大量请求要的是大体相似的东西 —— 同类场景、同样的物体、相似的构图 —— 那么一个已完成请求背后的计算,本可以用来回答之后的请求。这种跨请求搬运计算的做法,叫做跨请求复用。

The hard part is knowing when reuse is safe. Existing methods decide it from prompt similarity: if the text of the new request is close enough to a past one, reuse its cache. But similar prompts do not guarantee similar videos. Two prompts that read almost the same can drift into completely different visual structure as denoising proceeds, so a cache accepted on prompt similarity alone can quietly carry the wrong scene into the result — and the final quality score may not even reveal it.难的是判断什么时候复用安全。现有方法靠提示词相似度来决定:只要新请求的文本和历史上某个请求足够接近,就复用它的缓存。但提示词相似并不等于视频相似。两句读起来几乎一样的话,在去噪过程中可能演化出完全不同的画面结构 —— 于是仅凭提示词相似度接受的缓存,会把错误的场景悄无声息地带进结果,而最终的质量指标甚至看不出来。

A textually similar cached request still carries the wrong scene
Figure 1. Similar text, incompatible generation. The target asks for a peach on a snowy mountain; the recalled request — an apple on a mountain — is textually close enough to pass a similarity test, and its CLIP-T and VBench-Q look competitive. Reusing it keeps the grass and the wrong fruit. Carnator refuses the cache and falls back to full generation.图 1。文本相似,生成却不兼容。目标请求要的是「雪山上的桃子」,被召回的请求是「山上的苹果」—— 文本足够接近,CLIP-T 与 VBench-Q 看着也不差。可一旦复用,草地和错误的果实都被带了进来。Carnator 拒绝这个缓存,退回完整生成。

That leaves two questions, and Carnator is built around answering both:由此留下两个问题,Carnator 正是围绕它们设计的:

  1. Reuse validity. Is this historical computation actually compatible with the request we are serving right now?复用有效性。这段历史计算,与当前正在服务的请求真的兼容吗?
  2. Reuse scope. If it is, which parts of the computation can we skip — and which parts must still be computed for the new request?复用范围。如果兼容,哪些计算可以省掉?哪些又必须为新请求重新算?

Key Idea核心思路

Carnator stops guessing from the text and looks inside the model instead.Carnator 不再从文本去猜,而是往模型内部看。

The insight核心洞见

Prompt similarity is only a proxy for compatibility. The generation process itself is the real signal.提示词相似度只是兼容性的代理指标,真正的信号在生成过程本身。

Early in denoising, the partially formed state already reveals how the clip is taking shape — where objects are appearing, how large they are, what they look like. Carnator runs a lightweight early probe, summarises that evidence into a compact Early Signature, and compares signatures instead of prompts.在去噪早期,尚未成形的中间状态就已经透露出这段视频将长成什么样 —— 物体出现在哪里、有多大、是什么样子。Carnator 执行一次轻量的早期探测,把这部分证据压缩成一个紧凑的早期签名(Early Signature),然后比较签名,而不是比较提示词。

Because the same evidence also marks where the target-specific content sits, it answers the second question too: reuse everywhere the generations agree, and recompute only the regions that genuinely differ.同一份证据还标出了目标特有内容位于何处,于是第二个问题也一并有了答案:凡两次生成一致的地方就复用,只重算真正不同的区域。

Early denoising states separate compatible from incompatible request pairs
Figure 2. The signal is there early. Similar request pairs line up in their opening denoising representations, while easy and hard negatives drift apart in prompt-conditioned residuals and latent dynamics — which is what makes a short, cheap probe enough to tell them apart.图 2。信号在早期就已经出现。相似的请求对在最初几步的去噪表示上高度对齐,而「易负例」与「难负例」则在提示词条件下的残差与隐动态上依次偏离 —— 这就是一次短而廉价的探测足以区分它们的原因。

How It Works方法

Overview of Carnator
Figure 3. The pipeline end to end. Historical requests are recalled by text, filtered with generation-native Early Signatures, and — for an accepted cache — reused jointly across latent trajectories and sparse attention connectivity, while target-specific computation is recomputed.图 3。整体流程。历史请求先按文本召回,再用生成原生的早期签名过滤;对通过筛选的缓存,同时复用其隐变量轨迹与稀疏注意力连接,同时重算目标特有的部分。
  1. Probe early, not at the end. A short and cheap probe of the target's early denoising states builds an Early Signature. For each object named in the prompt it records a spatiotemporal support mask, the object's relative size, and its mean CIELAB colour.在早期探测,而不是等到最后。对目标请求的早期去噪状态做一次短而廉价的探测,构建早期签名:为提示词里每个物体记录时空支撑掩码、相对尺寸与平均 CIELAB 颜色。
  2. Filter under a risk budget. The target signature is compared with recalled candidates on four discrepancies — context appearance, object size, spatiotemporal shape, and object appearance. The thresholds are calibrated to maximise reuse coverage subject to a bound on unsafe reuses, then certified on a held-out set with a one-sided Clopper–Pearson bound. If nothing passes, the request simply falls back to full generation.在风险预算下筛选。目标签名与召回候选在四个差异维度上比较:上下文外观、物体尺寸、时空形状、物体外观。阈值在校准集上以「最大化复用覆盖率、同时约束不安全复用比例」为目标选取,再在一个独立认证集上用单侧 Clopper–Pearson 上界做认证。若没有候选通过,请求直接退回完整生成。
  3. Reuse along two axes at once. Once a cache is accepted, region-guided latent reuse recomputes tokens inside the evidence-derived active region and inherits the rest from the historical trajectory, while historical sparse-connectivity reuse decides where attention is evaluated — transferring the historical attention structure, not the historical attention values. The two are combined in the middle Transformer layers.沿两个维度同时复用。缓存被接受后,区域引导的隐变量复用只对证据推导出的活跃区域重算 token,其余从历史轨迹继承;历史稀疏连接复用则决定注意力在哪里计算 —— 迁移的是历史的注意力结构,而不是历史的注意力数值。两者在中间层 Transformer 中结合使用。

Results实验结果

Evaluated on three text-to-video backbones — Wan2.2-TI2V-5B, Wan2.1-T2V-1.3B and LTX-Video-13B — over a 1,000-request warm cache and a disjoint 300-request stream from VidProM, on NVIDIA H100 GPUs. Every online cost is counted, including retrieval, probing, signature construction, filtering and fallback.在三个文生视频骨干网络上评测 —— Wan2.2-TI2V-5B、Wan2.1-T2V-1.3B 与 LTX-Video-13B —— 使用 VidProM 的 1000 条历史请求作为热缓存、另取互不重叠的 300 条请求流,硬件为 NVIDIA H100。所有在线开销都计入,包括召回、探测、签名构建、过滤与回退。

2.17×end-to-end speedup on cache hits, Wan2.1-T2V-1.3B缓存命中时的端到端加速比,Wan2.1-T2V-1.3B
1.57×end-to-end speedup on cache hits, Wan2.2-TI2V-5B缓存命中时的端到端加速比,Wan2.2-TI2V-5B
0.6382VBench-Q, against 0.6331 for dense generationVBench-Q,稠密生成为 0.6331
4.50%certified upper bound on unsafe reuse, inside the 5% budget不安全复用的认证上界,在 5% 预算之内

Table 1: End-to-end performance on the VidProM workload across three video diffusion models. Within each model, all methods process the same request sequence under the same generation configuration. For cross-request methods, cache misses fall back to dense generation.表 1:三个视频扩散模型在 VidProM 工作负载上的端到端表现。同一模型内,所有方法处理相同的请求序列、使用相同的生成配置。对跨请求方法,未命中缓存的请求退回稠密生成。

MethodCLIP-T ↑VBench-Q ↑DiT Hit Latency (s) ↓E2E Hit Latency (s) ↓DiT Speedup (Hit) ↑E2E Speedup (Hit) ↑E2E Speedup ↑Hit Rate ↑
Wan2.2-TI2V-5B        
Dense0.33720.6331265.23293.881.00×1.00×1.00×N/A
NIRVANA-ViD0.29550.6457204.21256.091.30×1.15×1.12×80.67%
Chorus0.30050.6392175.48264.471.51×1.11×1.09×80.67%
Chorus II0.29860.6400201.58222.441.32×1.32×1.24×80.67%
Carnator0.29130.6382159.65187.301.66×1.57×1.36×72.67%
Wan2.1-T2V-1.3B        
Dense0.29840.6379974.21138.61.00×1.00×1.00×N/A
NIRVANA-ViD0.29440.6210884.51035.91.10×1.10×1.08×80.67%
Chorus0.29150.6089582.4719.81.67×1.57×1.42×80.67%
Chorus II0.28910.6445556.9577.11.75×1.97×1.66×80.67%
Carnator0.28100.6361496.0525.581.96×2.17×1.72×77.67%
LTX-Video-13B        
Dense0.26080.589679.3781.421.00×1.00×1.00×N/A
NIRVANA-ViD0.25860.606070.7872.821.12×1.12×1.09×80.67%
Chorus0.25740.599470.78136.241.12×0.60×0.65×80.67%
Chorus II0.26390.624475.2677.691.05×1.05×1.04×80.67%
Carnator0.25830.602262.5365.881.27×1.24×1.18×79.33%

What the numbers say这些数字说明了什么

  1. It wins by being pickier, not looser. Adding the Early Signature lowers the hit rate from 77.33% to 72.67%: roughly one in twenty textually similar caches is rejected as generation-incompatible. Latency barely moves (185.6 s to 187.30 s) while CLIP-T rises from 0.2725 to 0.2913 and VBench-Q from 0.6301 to 0.6382 — the filter removes caches that were hurting alignment, and costs almost nothing.它是靠更挑剔取胜,而不是更宽松。加入早期签名后命中率从 77.33% 降到 72.67%:大约每二十个文本相似的缓存里就有一个被判定为「生成不兼容」而拒绝。而延迟几乎没变(185.6 s → 187.30 s),CLIP-T 反而从 0.2725 升到 0.2913、VBench-Q 从 0.6301 升到 0.6382 —— 这个过滤器剔除的正是原本在损害文本对齐的缓存,代价几乎为零。
  2. Fewer accepted caches, larger savings each. On Wan2.1-T2V-1.3B, Carnator accepts fewer caches than the baselines yet still delivers the fastest cache-hit latency, dropping it from 1138.6 s (dense) to 525.58 s. Joint reuse extracts more computation out of each cache it accepts.接受的缓存更少,但每个缓存的收益更大。在 Wan2.1-T2V-1.3B 上,Carnator 接受的缓存比基线更少,缓存命中延迟却是最低的:从稠密生成的 1138.6 s 降到 525.58 s。联合复用的意义就在于,从每一个被接受的缓存里榨出更多可省的计算。
  3. Quality is preserved, not traded away. On the primary model, VBench-Q is 0.6382 against 0.6331 for dense generation, and CLIP-T stays in the same range as the baselines.质量是被保住的,不是拿来交换的。在主要评测模型上,VBench-Q 为 0.6382,稠密生成是 0.6331;CLIP-T 也保持在基线的同一区间。
  4. The two reuse axes are complementary. On a fixed 100-request subset, latent reuse alone reaches 206.06 s and attention reuse alone 160.91 s, but combining them reaches 155.85 s (1.71×) — better than either on its own.两个复用维度是互补的。在固定的 100 条请求子集上,只用隐变量复用是 206.06 s,只用注意力复用是 160.91 s,两者结合则是 155.85 s(1.71×)—— 优于任何单独一项。

On a disjoint certification set, the risk-aware rule accepts 50 of 70 request-cache pairs with no unsafe cases, giving a one-sided Clopper–Pearson upper bound of 4.50% and satisfying the prescribed 5% risk constraint.在独立的认证集上,风险感知规则在 70 对「请求–缓存」中接受了 50 对,且没有出现不安全案例,单侧 Clopper–Pearson 置信上界为 4.50%,满足预设的 5% 风险约束。

Qualitative comparison: dense source, dense target, three reuse baselines and Carnator
Figure 4. What reuse actually produces. Each row pairs a source request (dog / bear / airplane) with a target that changes the object (cat / panda / helicopter) on the same scene. The reuse baselines often keep the source object; the caches Carnator accepts reproduce the target.图 4。复用到底生成了什么。每一行是「源请求 → 目标请求」:场景不变,物体改变(狗 / 熊 / 飞机 → 猫 / 熊猫 / 直升机)。复用基线往往把源物体留在了画面里,而 Carnator 接受的缓存产出的是目标物体。

Citation引用

If you find this work useful, please consider citing it:如果这项工作对您有帮助,欢迎引用:

@article{yin2026carnator,
  title   = {Carnator: Fast Text-to-Video Generation with Generation-Native
             Compatibility-Guided Cross-Request Reuse},
  author  = {Yin, Xingkun and Tang, Xuebin and Xu, Mingkun and Du, Hongyang},
  journal = {arXiv preprint arXiv:2609.32420},
  year    = {2026}
}

The full manuscript — the complete method, related work, appendices and additional analysis — is available as arXiv:2609.32420 (PDF).完整手稿 —— 完整方法、相关工作、附录与补充分析 —— 见 arXiv:2609.32420(PDF)。