Reka AI's omni-model Rho-1 handles text, images, video, and robot control in a single model

Reka AI has introduced Rho-1, a 19-billion-parameter omni-model designed to handle text, images, video, and robot control actions within a single unified neural network. The core architectural idea is straightforward: instead of building separate specialized models for each modality and coordinating between them, Rho-1 treats every input and output - whether a word, an image patch, a video frame, or a robot action - as tokens processed in one shared context window.
This unified approach stands in contrast to the more common practice of assembling multimodal systems from distinct components, such as a vision encoder paired with a language model, or a separate policy network for robotics. By collapsing everything into a single token stream, Rho-1 can in principle draw on relationships across modalities that would otherwise be lost at the handoff points between specialized modules.
On the compute side, Reka reports that Rho-1 was trained on 320 NVIDIA H100 GPUs over approximately three months. That is a meaningful data point in context - many frontier models today are trained on thousands of GPUs for longer periods. Whether the efficiency gains come from architectural choices, data curation, or training methodology, the relatively modest resource footprint suggests that omni-model approaches may not necessarily require the largest compute budgets to cover a broad range of tasks.
The inclusion of robot control as a native output modality is notable. Most robotics work still relies on separate perception and planning stacks, often with language models bolted on for instruction following. A model that generates control actions in the same framework it uses to reason about images and text could simplify deployment pipelines and potentially improve coherence between perception and action. Rho-1 represents an early indication of where this line of research is heading, though real-world robotics performance will depend heavily on how well the model generalizes beyond its training distribution.

