CATok
NeurIPS 2026 · arXiv:2609.35469

Causality Validation

Coarse-to-fine token roles, validated by intervention

We use a native-decoder intervention—a donor-swap at the velocity level—to characterize the generative roles of individual token positions. Under conditional annealing, token k is bound to flow stage k via suffix masking, so as flow matching refines the trajectory from high- to low-noise stages, each token is trained to contribute at a particular refinement stage. We examine whether the learned tokens acquire distinct, flow-stage-aligned effects and a coarse-to-fine functional progression in action space; here, “coarse-to-fine” denotes the granularity of learned control information along the refinement process, not a prescribed frequency or fixed-basis decomposition.

We use the trained CATok tokenizer checkpoint (K=32, codebook 4096, horizon 20), its MMDiT velocity network vθ, and its ODE decoder D—the same modules used at inference, not a separately trained prefix decoder. We encode 128 held-out Franka action chunks (7-DoF + gripper) as Q = (q1, …, qK). For each token k, we replace its embedding with the same-position embedding from a donor chunk—keeping the intervention within the slot’s empirical distribution—while holding xt, t, and the noise realization fixed, so the resulting change isolates token k’s effect on the velocity prediction. Because conditional annealing masks token k outside its active range, we normalize its effect within that range.

Token x flow-time influence matrix
Figure 7 · Token × flow-time influence matrix Δ(t). Top: 32 individual tokens, raw (a) and per-token-normalized (b). Bottom: 8 groups of 4 tokens, raw (c) and per-group-normalized (d); cyan “+” marks the theoretical peak stage. The matrix has triangular support—token k is visible only for t < (k+1)/K, confirming the schedule κ(t)=tK—and each token’s influence concentrates near its own on-switch boundary (corr(g, t*) = 0.995 grouped).
  • Schedule-consistent support. Token k is visible only for t < (k+1)/K, consistent with the suffix-masking schedule κ(t) = tK.
  • Stage-specific peaks. Each token’s peak influence occurs near its assigned stage t* ≈ k/K, with corr(k, t*) = 0.985 (0.995 after grouping). Early tokens peak at high-noise (coarse) stages and late tokens at low-noise (fine) stages, so token effects are distinct and aligned with their assigned flow stages rather than interchangeable.