CATok
NeurIPS 2026 · arXiv:2609.35469

Main Results

Better success rate, faster inference, fewer steps

95.9%

Avg. success on LIBERO

1.7×

Faster VLA inference than FAST

3.6×

VRR×CR of FAST at equal tokens

~50%

Steps to match the best baseline

We evaluate CATok on three simulation benchmarks—LIBERO, SimplerEnv (Simpler-Bridge), and RoboTwin 2.0—against three representative tokenizer baselines: BIN (uniform per-dimension discretization), FAST (DCT-based compression with BPE tokenization), and OAT (learned fixed-length tokenizer with prefix-based decoding).

Method LIBERO Simpler-Bridge RoboTwin 2.0
SpatialObjectGoalLongAvg. SpoonCarrotStackEggplantAvg. CleanRand.
BIN 0.5860.8780.6800.6040.687 0.5420.3330.2080.7080.448 0.2130.221
FAST 0.9600.9980.9620.9010.955 0.5000.3750.3750.6250.469 0.4780.478
OAT 0.4280.8760.7040.2760.571 0.4170.2080.1670.6670.365 0.2290.233
CATok (Ours) 0.9780.9940.9540.9100.959 0.4580.4170.5420.5420.490 0.4890.531

Table 1 · Simulation benchmarking results. Accent bold marks the best result per column.

Training Efficiency
Figure 3 · Training efficiency on Simpler-Bridge. CATok achieves higher success rates with fewer training steps.
Inference Latency
Figure 4 · VLA inference latency. CATok (453 ms) is 1.7× faster than FAST (778 ms) and 3.9× faster than Binning (1747 ms).
Intrinsic Tokenizer Properties
Figure 5 · Intrinsic properties of action tokenizers. (a–c) CATok8 achieves the highest VRR×CR score (16.85), outperforming all baselines at the same compression level. (d–e) Encode and decode latencies on log scale—CATok achieves lower latency than FAST with comparable or better downstream performance.