NeurIPS 2026

Rethinking Causal Action Tokenization with Conditional Annealing in Flow Matching

CATok grounds discrete action tokens in the continuous flow-matching trajectory — a causally ordered, coarse-to-fine action representation for autoregressive VLA models.

Chenyu Zhang1,2*, Yuhang Cao1*, Daru Du1,2, Yingxi Lu1,2, Jing Shao1, Ruoqu Chen1,2, Jiajun Liu1,2, Liu Cao1,2, Yicheng Liu1, Hang Zhao1,2, Mengdi Xu1,2†

1IIIS, Tsinghua University  2Shanghai Qizhi Institute

*Equal Contribution  †Corresponding Author

Flow-Matching to Action Tokens
Flow-Matching to Action Tokens. Each discrete token encodes the residual reconstruction signal at a specific flow-matching stage, establishing a coarse-to-fine causal token space aligned with autoregressive modeling.

Abstract

Tokenization as a causally structured generative process

Autoregressive Vision-Language-Action (VLA) models offer a scalable path to robot learning, yet existing action tokenizers treat tokenization as a compression problem, producing representations that are semantically misaligned with the autoregressive backbone. We propose CATok, a causal action tokenizer that reframes tokenization as a causally structured generative process.

CATok introduces a conditional annealing mechanism that extracts action tokens by progressively annealing a flow-matching process: each token is conditioned on all preceding tokens and encodes the residual reconstruction signal at a specific noise level, establishing a coarse-to-fine causal token space whose generative semantics are structurally aligned with autoregressive modeling. A token-conditioned flow-matching decoder built on Multimodal Diffusion Transformer (MMDiT) reconstructs continuous action chunks from these discrete tokens with the precision of hybrid diffusion-head architectures. This discrete bottleneck enforces knowledge insulation by design, cleanly separating high-level semantic reasoning from low-level motor execution without requiring explicit attention masking.

Extensive evaluations across three simulation benchmarks demonstrate that CATok consistently surpasses existing tokenization methods in reconstruction fidelity–compression tradeoff and inference efficiency, while improving VLA task success rate, establishing a high-performance, scalable foundation for purely autoregressive VLA systems.

Method

Three components, one causal schedule

01

Dual-Stream Action Encoder

Maps raw action chunks into a compact set of latent query tokens via co-attention with an MMDiT-based transformer.

02

Bottleneck Vector Quantizer

Discretizes latent tokens into a finite codebook using factorized codes.

03

Flow-Matching Decoder

Reconstructs high-fidelity action trajectories with conditional annealing, enforcing a coarse-to-fine causal structure.

The key design is the conditional annealing mechanism: at flow-matching timestep t, the first ⌊tK⌋ token embeddings are progressively masked, forcing each token to specialize in a distinct generative stage—from coarse global guidance at high-noise stages to fine-grained residual corrections as the trajectory matures. This induces a causally ordered hierarchy naturally aligned with autoregressive action generation.

CATok Architecture Overview
Figure 1 · CATok Pipeline. CATok encodes an action chunk with a Dual-Stream Encoder into latent queries, discretizes them via a Bottleneck VQ, and reconstructs actions using a Flow-Matching Decoder with Conditional Annealing. The annealing schedule assigns each token to a distinct stage of the flow trajectory, establishing a coarse-to-fine causal token space.
VLA-CATok Overview
Figure 2 · VLA-CATok. VLA-CATok feeds visual, language, and proprioceptive tokens into a VLM backbone to autoregressively predict causal action tokens. A frozen MMDiT flow-matching decoder reconstructs the final continuous action. Discrete tokens serve as the sole interface, naturally insulating the VLM's pretrained knowledge from action-specific gradients.

Main Results

Better success rate, faster inference, fewer steps

95.9%

Avg. success on LIBERO

1.7×

Faster VLA inference than FAST

3.6×

VRR×CR of FAST at equal tokens

~50%

Steps to match the best baseline

We evaluate CATok on three simulation benchmarks—LIBERO, SimplerEnv (Simpler-Bridge), and RoboTwin 2.0—against three representative tokenizer baselines: BIN (uniform per-dimension discretization), FAST (DCT-based compression with BPE tokenization), and OAT (learned fixed-length tokenizer with prefix-based decoding).

  • Higher task success rate. CATok achieves the highest success rates across all benchmarks (95.9% avg. on LIBERO), with the largest gains on long-horizon tasks (LIBERO-Long: 91.0% vs. 90.1% for FAST).
  • Faster inference. 1.7× faster VLA inference than FAST and 3.9× faster than BIN, enabling more responsive real-time deployment.
  • Better fidelity–compression balance. 3.6× stronger VRR×CR score than FAST at the same token count.
  • Higher training efficiency. Higher success rates at the same training step count, attributable to stage-wise token factorization.
Method LIBERO Simpler-Bridge RoboTwin 2.0
SpatialObjectGoalLongAvg. SpoonCarrotStackEggplantAvg. CleanRand.
BIN 0.5860.8780.6800.6040.687 0.5420.3330.2080.7080.448 0.2130.221
FAST 0.9600.9980.9620.9010.955 0.5000.3750.3750.6250.469 0.4780.478
OAT 0.4280.8760.7040.2760.571 0.4170.2080.1670.6670.365 0.2290.233
CATok (Ours) 0.9780.9940.9540.9100.959 0.4580.4170.5420.5420.490 0.4890.531

Table 1 · Simulation benchmarking results. Accent bold marks the best result per column.

Training Efficiency
Figure 3 · Training efficiency on Simpler-Bridge. CATok achieves higher success rates with fewer training steps.
Inference Latency
Figure 4 · VLA inference latency. CATok (453 ms) is 1.7× faster than FAST (778 ms) and 3.9× faster than Binning (1747 ms).
Intrinsic Tokenizer Properties
Figure 5 · Intrinsic properties of action tokenizers. (a–c) CATok8 achieves the highest VRR×CR score (16.85), outperforming all baselines at the same compression level. (d–e) Encode and decode latencies on log scale—CATok achieves lower latency than FAST with comparable or better downstream performance.

Analysis

Causal tokens with generative semantics

Causal Predictive Ordering. Per-token Shannon entropy decreases monotonically in the forward token order (Figure 6a), while reversing the order flips the trend—confirming the learned ordering is directional and causal. FAST and BIN show no such structure.

Stage-wise Generative Semantics. t-SNE of VQ embeddings (Figure 6b) reveals clear slot-dependent geometry: early tokens form smooth elongated regions, late tokens form compact separated clusters. Prefix reconstructions (Figure 6c) show coherent coarse-to-fine trajectory updates, confirming each token contributes a distinct generative increment.

Causal Token Analysis
Figure 6 · CATok produces causal tokens with generative semantics. (a) Causal ordering: entropy decreases in the forward token order and increases in reverse. (b) Token embedding geometry: slot-dependent organization from smooth early-token regions to compact late-token clusters. (c) Prefix reconstruction: increasing prefixes produce structured coarse-to-fine trajectory updates.

Causality Validation

Coarse-to-fine token roles, validated by intervention

We use a native-decoder intervention—a donor-swap at the velocity level—to characterize the generative roles of individual token positions. Under conditional annealing, token k is bound to flow stage k via suffix masking, so as flow matching refines the trajectory from high- to low-noise stages, each token is trained to contribute at a particular refinement stage. We examine whether the learned tokens acquire distinct, flow-stage-aligned effects and a coarse-to-fine functional progression in action space; here, “coarse-to-fine” denotes the granularity of learned control information along the refinement process, not a prescribed frequency or fixed-basis decomposition.

Donor-swap influence matrix

We use the trained CATok tokenizer checkpoint (K=32, codebook 4096, horizon 20), its MMDiT velocity network vθ, and its ODE decoder D—the same modules used at inference, not a separately trained prefix decoder. We encode 128 held-out Franka action chunks (7-DoF + gripper) as Q = (q1, …, qK). For each token k, we replace its embedding with the same-position embedding from a donor chunk—keeping the intervention within the slot’s empirical distribution—while holding xt, t, and the noise realization fixed, so the resulting change isolates token k’s effect on the velocity prediction. Because conditional annealing masks token k outside its active range, we normalize its effect within that range.

Token x flow-time influence matrix
Figure 7 · Token × flow-time influence matrix Δ(t). Top: 32 individual tokens, raw (a) and per-token-normalized (b). Bottom: 8 groups of 4 tokens, raw (c) and per-group-normalized (d); cyan “+” marks the theoretical peak stage. The matrix has triangular support—token k is visible only for t < (k+1)/K, confirming the schedule κ(t)=tK—and each token’s influence concentrates near its own on-switch boundary (corr(g, t*) = 0.995 grouped).
  • Schedule-consistent support. Token k is visible only for t < (k+1)/K, consistent with the suffix-masking schedule κ(t) = tK.
  • Stage-specific peaks. Each token’s peak influence occurs near its assigned stage t* ≈ k/K, with corr(k, t*) = 0.985 (0.995 after grouping). Early tokens peak at high-noise (coarse) stages and late tokens at low-noise (fine) stages.

Real-World Deployment

From simulation to the physical world

To verify that CATok transfers to real-robot deployment under camera noise, lighting variation, and hardware differences, we design a suite of real-robot evaluation tasks. We collect 362 trajectories of “pick up the yellow cup and place it on the plate” and 220 trajectories of “stack the cups”. Using Qwen3.5-VL as the backbone, we train two VLAs that differ only in their tokenizer—QwenFAST and QwenCATok. CATok is pretrained on the collected real-robot data together with InternA1 joint data, so that its action reconstruction accuracy matches that of FAST.

Pick-Spatial scene layout
Pick-Spatial10 fixed cup-plate positions × 2 rollouts. Tests spatial generalization.
Pick-Color scene layout
Pick-ColorUnseen red cup, fixed position × 10 rollouts. Tests color generalization.
Stack-Long scene layout
Stack-LongStack two cups in sequence × 40 rollouts. Tests long-horizon completion.

Rollout comparison

Pick Pick (Red) Stack QwenCATokours
QwenFAST

Figure 8 · Real-world rollout comparison. Top row: QwenCATok (ours); bottom row: QwenFAST. Each cell is a single, unedited end-to-end rollout.

MethodPick-SpatialPick-ColorStack-Long
QwenFAST0.450.300.275
QwenCATok (Ours)0.650.500.475

Table 2 · Real-world success rates. QwenCATok improves over QwenFAST by +20pp absolute on all three tasks.

These real-world gains corroborate the mechanisms identified in simulation. CATok’s causally structured token space aligns the discrete bottleneck with the autoregressive backbone, and this semantic alignment generalizes to real-world perturbations: spatial repositioning (Pick-Spatial, 0.45→0.65, +44% relative) and out-of-distribution color (Pick-Color, 0.30→0.50, +67% relative). The largest gain appears on Stack-Long (0.275→0.475, +73% relative), consistent with the stage-wise causal factorization that drives training efficiency in simulation—by making each token prediction more informative, causal token ordering lets the policy better capture the long-range temporal dependencies required for compositional, long-horizon tasks.

Citation

BibTeX

@article{zhang2026catok,
  title   = {Rethinking Causal Action Tokenization with Conditional Annealing in Flow Matching},
  author  = {Zhang, Chenyu and Cao, Yuhang and Du, Daru and Lu, Yingxi and Shao, Jing and Chen, Ruoqu and Liu, Jiajun and Cao, Liu and Liu, Yicheng and Zhao, Hang and Xu, Mengdi},
  journal = {arXiv preprint arXiv:2609.35469},
  year    = {2026}
}