How quantization, sampler tuning, kernel work, and rollout topology accelerate NVIDIA's 16B Cosmos 3 robot policies on 24 GB GPUs — from single-request serving to full 1,200-episode evaluations.
Vec-LUT turns repetitive scalar table lookups into contiguous vector reads, accelerating parallel ternary LLM inference on x86 and ARM CPUs by up to 4.2×.