Causality Validation
We use a native-decoder intervention—a donor-swap at the velocity level—to characterize the generative roles of individual token positions. Under conditional annealing, token k is bound to flow stage k via suffix masking, so as flow matching refines the trajectory from high- to low-noise stages, each token is trained to contribute at a particular refinement stage. We examine whether the learned tokens acquire distinct, flow-stage-aligned effects and a coarse-to-fine functional progression in action space; here, “coarse-to-fine” denotes the granularity of learned control information along the refinement process, not a prescribed frequency or fixed-basis decomposition.
We use the trained CATok tokenizer checkpoint (K=32, codebook 4096, horizon 20), its MMDiT velocity network vθ, and its ODE decoder D—the same modules used at inference, not a separately trained prefix decoder. We encode 128 held-out Franka action chunks (7-DoF + gripper) as Q = (q1, …, qK). For each token k, we replace its embedding with the same-position embedding from a donor chunk—keeping the intervention within the slot’s empirical distribution—while holding xt, t, and the noise realization fixed, so the resulting change isolates token k’s effect on the velocity prediction. Because conditional annealing masks token k outside its active range, we normalize its effect within that range.