- Reka AI announced Rho-1 as a research preview on 5 October 2026, not as a generally available robot controller.
- The 19-billion-parameter scale, 320 H100 training run and video-speed figures are Reka’s claims; independent performance testing has not been published.
- The launch materials show simulated robot tasks. They do not establish real-world deployment reliability.
Source and verification note (6 October 2026): This report is anchored in Reka AI’s dated launch announcement. We also checked separately published accounts by The Decoder, RuntimeWire and AICoder. Those publications independently reviewed the release and distinguish the company’s demonstration from a deployed physical robot. They do not constitute independent benchmark tests of Rho-1. All performance, compute and model-size numbers below should therefore be read as company-reported, not independently validated.
Reka AI Rho-1: what is verified and what remains open
The useful question for readers is not whether a demonstration looks seamless, but what the preview actually proves. Reka AI’s October 5 release establishes that the company is offering a research preview and describing a single architecture for text, visual media and action. The published material does not provide a standardized robot-task success-rate table or a third-party latency comparison. For a deployment team in India, that separates an architectural research signal from a procurement-ready system.
For context, Universal Robots’ Gen 7 cobot platform illustrates the integration issues a factory must solve around hardware, sensing and software. Our NVIDIA Cosmos 3 Edge coverage explains another physical-AI approach. Neither is a direct benchmark against Rho-1; the links help distinguish a model research preview from a deployable robotics stack.
Reka AI has released a research preview of Rho-1, a 19-billion-parameter “omni-reasoning” model designed to understand and generate text, images and video while also producing robotic actions from the same underlying neural network. Unlike conventional multimodal systems that connect separate models for perception, generation and action, Reka says Rho-1 places these modalities inside one shared architecture and context.
The release represents a significant shift in Reka’s research direction from multimodal language models toward what it calls omni-world models: systems intended not just to describe the physical world but to simulate what happens next and act within it. The company says Rho-1 can continuously generate video, respond to new instructions while a simulation is running and translate its internal representation of an environment into robot-control actions.
Key takeaways
- Rho-1 is a 19-billion-parameter research-preview model.
- It processes and generates text, images, video and robotic actions within one neural network.
- Reka says the system does not rely on external model calls or conventional tool-based orchestration for these capabilities.
- The model can continuously generate video and accept new instructions without restarting the generation.
- Reka trained Rho-1 from scratch using 320 NVIDIA H100 GPUs for about three months.
- The company uses an inverse dynamics model to infer robot-control signals from ordinary video, addressing the shortage of action-labeled robotics data.
- Reka reports a median base-model video generation speed of 0.79× real time, while a distilled version reduces denoising from 99 steps to eight.
- The model remains a research preview, with important limitations including structural drift, weak temporal grounding, brittle editing and relatively low native video resolution.
- Reka has not yet published independent benchmark results demonstrating that Rho-1 outperforms specialized robotics or video-generation systems.
What is Reka Rho-1?
Rho-1 is Reka’s attempt to build a single AI system that can understand, simulate and act across multiple forms of information.
The model has 19 billion parameters and was trained from scratch, according to Reka. Rather than treating text, images, video and robot actions as separate workloads handled by different specialist models, the company represents them inside a shared architecture.
That distinction is important.
A conventional AI application might use a vision model to interpret a camera image, a language model to reason about the image, a video model to simulate what could happen next and a robotics policy to convert that plan into motor commands.
Each handoff introduces another interface between systems.
Reka’s argument is that a unified model can keep the information in one shared state.
Rho-1 therefore attempts to make the same model responsible for several stages of an interaction:
See → understand → imagine → decide → act
That is much closer to the way Reka believes future physical AI systems should operate.
How Rho-1 differs from conventional multimodal AI
Multimodal AI is already common.
Modern vision-language models can accept text and images and respond with text. Video models can generate or edit moving images. Robotics models can combine visual observations with language instructions to produce actions.
But these systems often have asymmetric input and output capabilities.
A vision-language model may understand an image but ultimately produce text. A video-generation model can create pixels but cannot necessarily control a robot. A robot policy may produce motor commands but have limited ability to generate a rich simulation of what will happen next.
Reka calls this the multimodal stack.
Rho-1 attempts to collapse that stack.
The company’s architecture treats text and symbolic reasoning as discrete tokens, while image representations, video frames, robot actions and proprioceptive information are represented through continuous tokens. Both types of information can participate in the model’s shared attention mechanism.
The result is intended to be a model in which understanding and generation are not separate applications connected by APIs.
Instead, they are different operations performed within the same model state.
One conversation can move from image to video to reasoning
Reka demonstrates this concept through a single multi-turn interaction.
In the company’s example, Rho-1 first creates an image of a lighthouse on a rocky coastline.
The user then asks the model to place a box around the lighthouse.
Next, the model animates the scene into video.
The user subsequently asks what changed between two videos, and the model explains the visual differences.
The important part is not simply that Rho-1 can perform each task.
The company says all five turns are handled by the same model and that the state accumulates in a shared key-value cache rather than being passed from one independent model to another.
This could reduce the need to repeatedly encode and transfer context between specialized systems.
It also creates a different approach to AI memory.
Instead of saying:
Model A → Model B → Model C
the architecture aims for:
One model → one shared state → multiple outputs
That is the central idea behind Reka’s omni-model research.
Rho-1 can generate and steer video continuously
One of Rho-1’s more unusual capabilities is continuous video generation.
Reka says the base model generates video at a median rate of 0.79 times real time, with the first watchable stream beginning in roughly six seconds.
The company has also created a distilled version.
The base model’s video generation process uses 99 denoising steps. Reka says the distilled version reduces that process to just eight steps while maintaining minimal quality loss in its internal evaluation.
This matters because conventional high-quality video generation can be computationally expensive.
If a model must generate an entirely new video clip every time a user changes an instruction, interactive use becomes difficult.
Rho-1 instead attempts to maintain a continuously evolving simulation.
A user could start with a scene and then change the instructions while the scene is unfolding.
For example, Reka demonstrates a virtual environment containing two balls on a countertop. A robot arm enters the scene, responds to an instruction to grab one ball and subsequently responds to another instruction telling it to put the ball down.
The model is therefore not just generating a sequence of disconnected clips.
It is attempting to maintain a persistent state and evolve that state over time.
From video prediction to robot control
This is where Rho-1 becomes particularly interesting for physical AI.
Reka’s objective is not merely to create better video.
The company wants the model to use video prediction as part of a robot’s ability to plan.
A robot needs to answer questions such as:
- What am I looking at?
- What will happen if I move this object?
- Where will the object be after I move my arm?
- What should I do next?
- What action should my motors execute?
A conventional robotics pipeline may assign these questions to different components.
Rho-1 attempts to place them inside a single model.
Reka describes this as a World-Language-Action Model, or WLAM.
The idea is that the same network predicting future camera observations can also produce the robot trajectory required to bring about those observations.
In simple terms, the model tries to imagine the next few seconds and act accordingly.
That is different from a system that simply maps an image directly to a robot command.
Why robot data is such a difficult problem
There is a major obstacle to this approach: robotics data is scarce.
The internet contains enormous amounts of video, but most online video does not include the underlying robot-control information.
A video of someone picking up a mug might show the movement of the hand, but it does not normally provide the exact joint angles, motor commands or torque values used to produce the movement.
Traditional robotics training therefore depends heavily on expensive data collection, simulation or teleoperation.
Reka’s solution is an Inverse Dynamics Model, or IDM.
The IDM attempts to infer the actions responsible for observed movement in ordinary video.
That potentially turns unlabeled video into a source of approximate action data.
The workflow is essentially:
Internet video → infer hidden actions → create action-labeled training data → train omni model
Reka says its inverse dynamics system can extract low-level motor commands from video, allowing Rho-1 to learn from a much larger pool of visual data than would be possible using only conventional teleoperated robot demonstrations.
This could prove to be one of the most important parts of the research.
The bigger challenge for physical AI may not be model size alone.
It may be finding enough high-quality data describing how the physical world changes when an agent takes an action.
Reka trained Rho-1 on 320 H100 GPUs
Rho-1 was trained from scratch using 320 NVIDIA H100 GPUs over approximately three months, according to Reka.
The company characterizes this as a relatively modest amount of compute compared with the resources used for some frontier AI systems.
That point is important because Reka is presenting Rho-1 as an architectural experiment rather than a finished frontier product.
The model’s current limitations suggest that scaling remains necessary.
Reka says long video rollouts can experience structural drift. A scene may remain visually convincing while the physical layout gradually becomes inconsistent.
The model also has problems maintaining object grounding throughout video, while targeted visual editing remains brittle.
Its native video output is currently limited to 672 × 384 pixels.
These limitations mean Rho-1 should not yet be interpreted as a production-ready general-purpose physical AI system.
Instead, the release demonstrates what Reka believes could happen if the unified architecture is scaled.
Rho-1’s architecture has two processing streams
Although Reka describes Rho-1 as a single model, its internal design is not simply one undifferentiated computation pathway.
The model contains two expert weight streams.
One is an understanding stream responsible for language and visual parsing.
The other is a generation stream that handles the transformation of latent representations into images and video.
The two streams share attention and state.
This is an important distinction.
Rho-1 is unified at the architectural level, but the company still uses specialized computational pathways within the model.
The difference is that these pathways are trained to work together inside a shared representation rather than existing as completely independent models connected through external APIs.
Reka says the architecture is trained using two objectives:
- Next-token prediction for discrete sequences such as language.
- Flow matching for continuous signals such as images, video and actions.
The company’s hypothesis is that improvements in one capability can reinforce other capabilities because they share the same representation and attention system.
Why Reka believes one model could beat a chain of models
Today’s AI systems increasingly rely on orchestration.
An agent might ask one model to interpret an image, another to generate a picture, another to create a video and another to operate a tool.
This approach is flexible, but it creates overhead.
Every model handoff can add latency, consume additional compute and require context to be transferred or reformatted.
Reka argues that an omni architecture could eliminate much of this overhead.
Its September research paper described the concept as a system with one shared representation that can process language, images, video and actions rather than chaining separate specialists.
The company argues that this could eventually reduce:
- Model-to-model communication
- Context switching
- Duplicate computation
- Inference latency
- Engineering complexity
But those are architectural arguments, not yet proven production advantages.
Rho-1 remains a research preview, and independent testing will be necessary to determine whether the theoretical benefits translate into better cost, latency, reliability and capability.
The robotics opportunity could be bigger than video generation
Video generation is an obvious application for Rho-1.
The model could potentially create interactive simulations for games, entertainment, training and virtual environments.
But the larger strategic opportunity may be robotics.
A robot operating in the real world cannot rely only on recognizing objects.
It needs an internal representation of how those objects behave.
A cup can be picked up.
A cloth can be folded.
A sponge can deform.
A box can be pushed.
A robot arm can collide with an obstacle.
These physical relationships are difficult to describe entirely through language.
Video provides a natural source of information about them.
Reka’s approach therefore combines three capabilities that are usually separated:
Perception + simulation + action
The model sees an environment, predicts how it could evolve and produces actions that can influence that evolution.
That is much closer to a world model than a traditional chatbot.
Reka is targeting the broader physical AI market
Rho-1 is part of a larger strategic direction at Reka.
The company says it is building models and infrastructure for the physical AI era, with applications spanning robotics, autonomous systems, wearables and real-time media.
Its September research described potential applications including autonomous robots, smart glasses, augmented reality, edge computing, interactive media and future operating systems.
The logic is straightforward.
If a model can understand images, language and physical movement within the same architecture, it could potentially be adapted to many devices that interact with the physical world.
A robot could use it to manipulate objects.
Smart glasses could use it to interpret surroundings.
An autonomous vehicle could use similar world modeling to anticipate how a scene will evolve.
A video application could use it to create an interactive environment rather than a fixed clip.
The same underlying technology could theoretically serve all of these applications.
Reka’s earlier work laid the groundwork
Rho-1 did not appear in isolation.
Reka has been working on multimodal models for years.
The company previously developed models capable of processing visual and language information, and its research has increasingly moved toward models that combine understanding, generation and physical-world reasoning.
In August 2026, Reka also published research on real-time video generation.
That system could generate a continuous video stream at 720p and 24 frames per second, according to the company, using a 30-billion-parameter model and NVIDIA H100 hardware. Reka reported an 11.8× end-to-end speedup through distillation and related optimization.
That work appears to have provided part of the technical foundation for the company’s broader omni-model strategy.
Reka’s research page now describes its direction as building world language action models that combine reasoning, visual understanding, generation and action.
The biggest challenge is reliability
The promise of Rho-1 is substantial, but physical AI has a much higher reliability requirement than ordinary content generation.
A video model can make a visual mistake and simply generate another frame.
A robot cannot always do that.
If an AI model incorrectly predicts the position of a physical object, the resulting action could damage equipment, drop an object or cause a collision.
That makes the current limitations particularly important.
Reka itself acknowledges that Rho-1 experiences long-horizon drift and limitations in temporal grounding and editing stability.
The difference between an impressive demonstration and a dependable robot controller is therefore enormous.
A model must remain accurate not only for one action but across hundreds or thousands of sequential decisions.
That is why independent robotics evaluations will matter more than visual demos as this technology develops.
What has not been proven yet
There is an important gap between Reka’s demonstrations and independent evidence.
Rho-1 is currently a research preview.
Reka has not published enough independent benchmark evidence to establish that it outperforms specialized robotics policies, video models or multimodal systems across standardized evaluations.
Its reported speed measurements are also company-internal results.
That does not make the results meaningless, but they should be treated as vendor claims rather than independent benchmarks.
The same applies to the model’s physical reasoning capabilities.
Reka’s demonstrations show promising behavior, but a handful of successful simulated robot tasks cannot establish that the system is ready for uncontrolled real-world environments.
The research preview should therefore be viewed as evidence that the architecture is technically plausible, rather than proof that the architecture has already solved physical AI.
Why Rho-1 matters for the AI industry
The most important part of Rho-1 may ultimately have little to do with the model’s 19 billion parameters.
It is the idea of one model serving as perception, simulation and action engine at the same time.
The AI industry has spent years building increasingly sophisticated pipelines.
One model understands language.
Another sees images.
Another generates pictures.
Another creates video.
Another controls a robot.
Another evaluates the result.
Rho-1 represents the opposite philosophy.
Instead of making the pipeline longer, Reka wants to make the foundation broader.
If that approach scales, the implications could be substantial.
Future AI systems may not need to think of text, images, video and actions as separate capabilities. They could become different expressions of the same underlying representation.
That would represent a significant change from today’s predominantly language-centric AI architecture.
The Bigger Picture
Rho-1 is an early but notable example of the AI industry’s movement from language models toward world models.
Large language models are extremely capable in symbolic environments, but physical intelligence requires more than predicting the next word. An agent operating in the real world needs to understand geometry, motion, time, causality and the consequences of its actions.
Reka’s omni architecture attempts to combine those requirements in one system.
The most compelling aspect is therefore not that Rho-1 can generate video or control a simulated robot individually. It is that the company is trying to make video prediction and robot control two sides of the same prediction problem.
If a model can predict what the camera will see next, it may also be able to predict which action will produce that future.
That is the foundation of Reka’s broader vision for physical AI.
Looking Ahead
The next important milestone will be independent evaluation. Rho-1 will need to demonstrate reliable long-horizon planning, physical grounding, low-latency control and consistent behavior outside carefully selected demonstrations. Improvements in training data, video resolution, temporal consistency and model scale could determine whether the architecture moves beyond research into practical robotics and simulation.
If Reka’s approach works at larger scales, the distinction between a multimodal AI assistant, a world simulator and a robot controller could become increasingly blurred. Instead of connecting separate models for seeing, reasoning, imagining and acting, future systems could use a single foundation that maintains a persistent representation of the environment and continuously converts that understanding into new predictions and physical actions.
Get the day’s top stories in your inbox
One concise email. No spam, unsubscribe anytime.



