Papers
arxiv:2610.08777

CtrlCache: Accelerating Interactive Video World Models with Control-Aware Caching

Published on Oct 6
· Submitted by
shangyesong
on Oct 7
Authors:
,
,
,

Abstract

Interactive video world models need to generate each video chunk efficiently while responding faithfully to user controls. Many systems use chunk-wise autoregressive generation with few-step denoising, but each chunk still requires several costly denoising iterations. Training-free caching can reduce this cost, yet existing policies make reuse decisions primarily from model-internal denoising dynamics and do not explicitly account for control transitions. Actually, interactive generation explicitly exposes a signal they do not use: the controls for a chunk arrive before it is denoised, so a schedule derived from them costs no forward pass. To this end, we analyze adjacent chunks under different control regimes and find that structural similarity drops around action changes, while low-frequency structure remains more persistent than high-frequency detail. Motivated by these observations, we propose CtrlCache, a training-free control-aware caching framework that adapts computation to the current control sequence. Specifically, the action-aware scheduling and refresh policy detects action changes across and within chunks, and labels each chunk as initial, transition, turning, or steady state. At one selected interior denoising step, initial and transition chunks retain full computation, while turning and steady chunks reuse the transformer residual from the most recent fully computed step in the same chunk. To exploit the persistence of low-frequency structure during steady interaction, we further introduce a frequency-mixed history prior guidance that incorporates complementary information from the preceding clean latent without an additional DiT forward pass. Evaluated on Matrix-Game 2.0 and LingBot-World v1/v2, CtrlCache achieves 1.21x to 1.41x DiT-backbone speedups without model retraining while improving WBench Overall scores over original inference across all three models.

Community

Paper author Paper submitter

Hi everyone! We present CtrlCache, a training-free caching method for interactive video world models.

Key idea: in interactive generation, the user's controls for a chunk are known before the chunk is denoised. CtrlCache uses them to decide when cached transformer computation can be reused (steady motion / turning) and when it must be refreshed (action changes), and adds a frequency-mixed history prior during steady interaction.

On Matrix-Game 2.0 and LingBot-World v1/v2, it gives 1.21×–1.41× DiT-backbone speedups while improving WBench Overall over the original models.

Project page with side-by-side videos: https://wrecklong.github.io/CtrlCache/
Code: https://github.com/wrecklong/CtrlCache (coming soon)

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2610.08777
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2610.08777 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2610.08777 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2610.08777 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.