CATok
NeurIPS 2026 · arXiv:2609.35469

Real-World Deployment

From simulation to the physical world

We collect 362 trajectories of “pick up the yellow cup and place it on the plate” and 220 trajectories of “stack the cups”. Using Qwen3.5-VL as the backbone, we train two VLAs that differ only in their tokenizer—QwenFAST and QwenCATok. CATok is pretrained on the collected real-robot data together with InternA1 joint data, so that its action reconstruction accuracy matches that of FAST.

Pick-Spatial scene layout
Pick-Spatial10 fixed cup-plate positions × 2 rollouts. Tests spatial generalization.
Pick-Color scene layout
Pick-ColorUnseen red cup, fixed position × 10 rollouts. Tests color generalization.
Stack-Long scene layout
Stack-LongStack two cups in sequence × 40 rollouts. Tests long-horizon completion.

Rollout comparison

Pick Pick (Red) Stack QwenCATokours
QwenFAST

Figure 8 · Real-world rollout comparison. Top row: QwenCATok (ours); bottom row: QwenFAST. Each cell is a single, unedited end-to-end rollout.

MethodPick-SpatialPick-ColorStack-Long
QwenFAST0.450.300.275
QwenCATok (Ours)0.650.500.475

Table 2 · Real-world success rates. QwenCATok improves over QwenFAST by +20pp absolute on all three tasks.