Reka Releases Rho-1: A 19B Omni-Reasoning Model That Understands, Generates Video and Outputs Robot Actions in One
Reka has unveiled Rho-1, a 19-billion-parameter omni-reasoning model built from scratch that collapses several distinct AI capabilities into one network. Rather than stitching together separate specialist models, Rho-1 handles text, images, video comprehension, video generation, and robot action prediction through a single set of weights. The model processes all of these modalities over a shared key-value cache, which allows information from one domain to inform reasoning in another without the overhead of passing data between separate systems.
The inclusion of robot action outputs alongside generative video is an notable design choice. Most multimodal models treat perception and generation as the ceiling of ambition, but Rho-1 extends into embodied AI territory by producing low-level control signals that could drive robotic systems. This positions the model as a potential backbone for research that sits at the intersection of generative media and physical-world automation.
On the video generation side, Reka has also released a distilled variant of Rho-1 optimized for speed. That version can return a 5.3-second video clip in approximately one second - a significant reduction in inference time that makes the capability more practical for interactive or real-time applications. Distillation techniques typically trade some output quality for speed, though the degree of that tradeoff in Rho-1's case will be clearer once the research community can test the model more broadly.
For now, Rho-1 is available as a research preview only, and Reka has not released public weights. That limits immediate hands-on evaluation, but the architectural decisions - particularly the unified KV cache and the breadth of supported output types - give researchers and developers a concrete look at where all-in-one multimodal reasoning is heading. A full release with accessible weights would be needed before the model's real-world performance across each modality can be thoroughly benchmarked.
