NeurIPS 2026

TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining

Shih-Ying Yeh♠♡† Daniel Z. Kaplan◇ Xuehai Wang♣△ Fu-En Yang★ Min-Hung Chen★ Shang-Hong Lai♠
♠National Tsing Hua University ♡Comfy Org Research ◇realiz.ai ♣Karolinska Institutet △Stockholm University ★NVIDIA
† Corresponding author: kohaku@kblueleaf.net
Paper (arXiv) Code BibTeX
Three encoder designs: spatial-only per-frame encoders, global 3D mixing, and TT-VidT with an explicit temporal axis that passes motion tokens across frames.
Spatial-only encoders never exchange information across frames, and global 3D mixing entangles appearance with motion throughout the network. TT-VidT keeps a wide per-frame spatial path for appearance and routes frame-to-frame change through a compact Temporal Transfer path that emits one set of motion tokens per frame.

TL;DR. Under a matched recipe of 24 architecture-objective combinations, only the pairing of a compact Temporal Transfer path with Diff Compression learns a representation that reads motion rather than appearance. TT-VidT leads Jester, Something-Something V2, ARID and Diving48 at the same time, at about half the encoder FLOPs of VideoMAE, V-JEPA 2 and DisMo.

Abstract

Comparisons in video self-supervised learning often evaluate complete training recipes rather than isolating the method itself: architecture, objective, data exposure, schedule, scale, and decoder capacity can all vary at once. This makes it hard to identify which choices yield motion-prioritized representations, whose gains concentrate on frame-to-frame change while retaining useful appearance. We address this with a matched 4×6 = 24 architecture-objective study at roughly 170M∼190M encoder scale on ~1.7M OpenVid and Moments-in-Time v2 clips for 8 epochs, and propose TT-VidT. TT-VidT combines a DINOv3-initialized ViT-B/16 per-frame spatial path with a compact Temporal Transfer Layer, trained by Diff Compression to reconstruct target frames from a first-frame appearance anchor and frame-specific motion tokens. The sweep shows that TT3D with Diff Compression, not either component alone, enters the strongest motion-sensitive regime, and decoder ablations favor a compact video-pretrained decoder. In final comparison, TT-VidT leads Jester, Something-Something V2, ARID, and Diving48 fine-tuning simultaneously, improving over the strongest non-TT row by 54%∼121%, while using 48% fewer encoder FLOPs than DisMo and 55% fewer than VideoMAE or V-JEPA 2. HMDB51, IARD, and EPIC-Kitchens bound the claim.

Method

The encoder interleaves DINOv3 ViT-B/16 layers, applied to each frame independently, with a Temporal Transfer layer. Eight learnable motion tokens join each frame's spatial tokens in the ViT self-attention, and the Temporal Transfer layer then lets the motion tokens, together with a 4×-downsampled view of the spatial tokens, attend across frames under a block-causal mask. The cross-frame sequence has 192 tokens instead of 2,048.

Diff Compression trains the encoder: a small diffusion decoder reconstructs every later frame from the first frame's full spatial features and that frame's motion tokens alone. Appearance is always available from the wide anchor, so the only useful content of a motion token is what the anchor cannot supply, which is how the frame has changed.

TT-VidT pretraining: the encoder interleaves 2D ViT layers with Temporal Transfer; the Diff Compression decoder denoises each target frame conditioned on the first-frame anchor and that frame's motion tokens.
TT-VidT pretraining. Left: the encoder. Right: the Diff Compression decoder, conditioned on the first-frame anchor and one frame's motion tokens.

Results

Setup. Every model in this study is trained from the same matched recipe, a deliberately small one: about 1.7M OpenVid and Moments-in-Time v2 clips, 8 epochs (about 13.6M samples seen), and encoders of roughly 170M∼190M parameters. This budget matches DisMo's published setting. It is far below the native recipes of VideoMAE or V-JEPA 2, so absolute numbers are lower than their reported ones, but every difference below comes from the architecture and objective rather than from data or compute. We evaluate frozen features with an attentive probe unless a table says finetuning.

73.25Jester, frozen attentive probe (strongest baseline 46.95)
25.92Something-Something V2, frozen (strongest baseline 14.95)
≤1%of clips keep their old label when the motion is inverted (baselines 9∼43%)
456 GFencoder FLOPs, about half of VideoMAE, V-JEPA 2 and DisMo

The only design to lead every motion-heavy benchmark

All methods share one pretraining recipe: the same data, schedule and encoder scale. Frozen attentive probing on six datasets, plus Diving48 full finetuning and EPIC-Kitchens verb anticipation.

MethodHMDBARIDIARDJesterSSv2EK-VD48 FTEK-V Antic.
V-JEPA 222.7016.0179.3427.4110.5932.618.0223.07
VideoMAE27.7324.3780.4439.4714.9535.628.4322.81
DisMo (dual aug.)22.5722.4589.8946.9513.3331.978.1222.72
TT-VidT (ours)25.1537.4774.8473.2525.9232.5418.6321.33

HMDB51, IARD and EPIC-Kitchens reward appearance, scene context or anticipation, and bound the claim: TT-VidT targets motion-sensitive recognition.

The representation reads motion, not appearance

Something-Something V2 contains class pairs that are exact mirrors under a horizontal flip, a vertical flip, or time reversal. We invert the motion of correctly classified clips and ask whether a frozen probe follows. A flip to the mirror class means the encoder re-read the motion, and a stay means appearance decided. TT-VidT follows the inverted motion on nearly every clip, while every baseline keeps part of its answers, most of all under time reversal, where no pixel changes.

ModelH-flip (flip / stay)V-flipTime reversalShuffle drop
TT-VidT (ours)81.2 / 0.659.5 / 0.971.4 / 1.050.9%
VideoMAE66.9 / 9.755.1 / 9.138.6 / 20.950.0%
DisMo34.5 / 35.08.4 / 43.234.5 / 16.431.9%
DINOv3, single frame37.1 / 24.511.8 / 22.70.0 / 100.03.1%

The gain belongs to the method

Four panels: the motion-inversion probe, frozen versus finetuned accuracy on Jester, three-seed means with error bars, and SSv2 gain over a 15-epoch continuation.
From left: the motion-inversion probe; the frozen probe against 30 epochs of end-to-end finetuning on Jester; the canonical configurations over three pretraining and three probe seeds; and the SSv2 gain over a 15-epoch continuation (dotted line: the 8-epoch budget).

End-to-end finetuning saturates the baselines near 51.5 on Jester, well below TT-VidT's frozen probe. A ViT3D initialized from DINOv3 moves only appearance-heavy IARD, so the image substrate does not create the motion regime. A VTok-style explicit feature difference is weaker than the learned motion token on all five datasets. The ordering holds across seeds and at a near-doubled training budget.

BibTeX

@inproceedings{
yeh2026ttvidt,
title={{TTV}idT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining},
author={Shih-Ying Yeh and Daniel Z Kaplan and Xuehai Wang and Fu-En Yang and Min-Hung Chen and Shang-Hong Lai},
booktitle={The Fortieth Annual Conference on Neural Information Processing Systems},
year={2026},
url={https://openreview.net/forum?id=ELiMecKQps}
}