gen‑ai.news
← Back
Multimodal

Reka AI's omni-model Rho-1 handles text, images, video, and robot control in a single model

Reka AI's omni-model Rho-1 handles text, images, video, and robot control in a single model

Reka AI has introduced Rho-1, a 19-billion-parameter omni-model designed to handle text, images, video, and robot control actions within a single unified neural network. The core architectural idea is straightforward: instead of building separate specialized models for each modality and coordinating between them, Rho-1 treats every input and output - whether a word, an image patch, a video frame, or a robot action - as tokens processed in one shared context window.

This unified approach stands in contrast to the more common practice of assembling multimodal systems from distinct components, such as a vision encoder paired with a language model, or a separate policy network for robotics. By collapsing everything into a single token stream, Rho-1 can in principle draw on relationships across modalities that would otherwise be lost at the handoff points between specialized modules.

On the compute side, Reka reports that Rho-1 was trained on 320 NVIDIA H100 GPUs over approximately three months. That is a meaningful data point in context - many frontier models today are trained on thousands of GPUs for longer periods. Whether the efficiency gains come from architectural choices, data curation, or training methodology, the relatively modest resource footprint suggests that omni-model approaches may not necessarily require the largest compute budgets to cover a broad range of tasks.

The inclusion of robot control as a native output modality is notable. Most robotics work still relies on separate perception and planning stacks, often with language models bolted on for instruction following. A model that generates control actions in the same framework it uses to reason about images and text could simplify deployment pipelines and potentially improve coherence between perception and action. Rho-1 represents an early indication of where this line of research is heading, though real-world robotics performance will depend heavily on how well the model generalizes beyond its training distribution.

Enjoy this story? Get the next one in your inbox.

Twice a week: the most important stories in generative image and video AI, distilled into a 2-minute read.

Free. Unsubscribe any time. No spam, ever.

Your next read

No image
Multimodal

Google’s new SynthID website can identify AI-generated media

Google has launched a public-facing website for SynthID, its AI watermarking technology, allowing anyone to upload and check whether an image, video, or audio clip was generated by AI. The tool extends SynthID beyond its previous developer and enterprise integrations, putting detection capabilities directly in the hands of everyday users. This marks a notable step toward accessible, practical tools for identifying synthetic media in the wild.

Uniting LED Volumes and Live Visual Chaos: Cinematographer Bradford Lipson on 'Rolling Loud'
Multimodal

Uniting LED Volumes and Live Visual Chaos: Cinematographer Bradford Lipson on 'Rolling Loud'

Cinematographer Bradford Lipson explains how he and director Jeremy Garelick used LED volume technology to blend real footage from the Rolling Loud music festival with controlled stage work, creating the illusion that actors were surrounded by tens of thousands of live concertgoers. The production shot on location at the festival for just three days before rebuilding that world on the Lux Stage at Trilith Studios. Lipson walks through the technical and artistic decisions behind matching concert

AMD Acquires World Labs for $8.2 Billion, Bringing Fei-Fei Li Onboard
Multimodal

AMD Acquires World Labs for $8.2 Billion, Bringing Fei-Fei Li Onboard

AMD has agreed to acquire World Labs, the spatial intelligence startup founded by AI pioneer Fei-Fei Li, for $8.2 billion. Li will join AMD as Executive Vice President and Chief Scientist, reporting to CEO Lisa Su. The deal gives AMD an in-house AI research unit built around world models that generate and simulate 3D environments.