A Unified KV Cache for Multi-Task Inference on Robots
OxyGen treats the KV cache as one shared resource across tasks and control frames, letting Mixture-of-Transformers VLAs keep full action frequency while generating language concurrently, up to 3.7× faster on an RTX 4090 and Jetson AGX Thor.