One model, four jobs
Reka AI has unveiled Rho-1, a 19-billion-parameter omni-model. It can both process and generate text, images, video, and robot control actions inside a single neural network.
No task routing
Today's systems typically route different tasks to specialized models. Rho-1 instead runs all modalities as tokens in one shared context window. This approach lets the model learn a common representation across different data types.
Lower compute cost
According to Reka AI, Rho-1 was trained on 320 H100 GPUs in about three months. The company says this is only a fraction of the compute today's top models require.
Why it matters
Combining robot control with language and vision models under one roof could reduce the need to manage separate components in autonomous systems. Rho-1's low training cost also strengthens the case that such multimodal models may become more accessible.



