Chinese technology giant Xiaomi has unveiled Xiaomi-Robotics-1, a new vision-language-action (VLA) foundation model that challenges a widely held assumption in robotics AI: that larger models inevitably deliver better performance. Instead, Xiaomi’s research shows that scaling high-quality training data has a much greater impact on robot capabilities than simply increasing model size, highlighting the growing importance of data collection in the race to build more capable autonomous robots.
The model was trained on more than 100,000 hours of real-world manipulation data collected using handheld, camera-equipped grippers operated by humans rather than expensive robotic systems. Xiaomi says this approach enabled it to build one of the largest robotics datasets to date while significantly lowering data collection costs. The company also developed an automated labeling pipeline that used AI to generate natural-language descriptions for the recorded motions, completing annotation of the entire dataset in about two weeks.
More Data Delivers Bigger Gains Than Larger Models
A central finding of Xiaomi’s research is that increasing the amount of training data consistently produced greater performance improvements than increasing the model’s parameter count.
During testing, researchers observed that:
- Larger models improved prediction accuracy.
- Increasing training data produced substantially larger performance gains.
- Performance continued improving as more data was added, with no clear plateau observed.
The findings suggest that robotics AI may benefit more from scaling diverse real-world datasets than from continually building larger neural networks, at least under current training approaches.
Key Research Highlights
| Metric | Result |
|---|---|
| Training data | 100,000+ hours |
| Data collection method | Human-operated handheld grippers |
| Data labeling | AI-generated automated annotations |
| Primary finding | More training data outperformed larger model size |
Human-Collected Motion Data Powers the Model
Instead of relying exclusively on physical robots to generate demonstrations, Xiaomi equipped handheld grippers with cameras and sensors, allowing people to naturally perform manipulation tasks.
This approach enabled researchers to:
- Collect data more quickly.
- Capture a wider variety of environments.
- Reduce the cost of robot demonstrations.
- Generate richer manipulation datasets.
The resulting trajectories were automatically converted into language-action pairs using an AI-based labeling system, creating a scalable pipeline for robotics training.

Strong Performance on New Tasks
Xiaomi reports that the foundation model can generalize to unfamiliar environments with relatively little additional training.
In evaluation tests:
- Less than 10 hours of task-specific training data were required for each new task.
- The model achieved an average success rate of about 75% across four unseen manipulation tasks.
- A competing model from Physical Intelligence reportedly achieved approximately 40% on the same benchmark.
Researchers also demonstrated the model performing complex activities such as suitcase packing, organizing household objects, and manipulating soft materials like paper.
Benchmark Performance
| Benchmark | Xiaomi-Robotics-1 |
|---|---|
| RoboCasa365 success rate | 57.6% |
| Average success on unseen tasks | ~75% |
| Task-specific fine-tuning | Less than 10 hours |
The accompanying research paper states that Xiaomi-Robotics-1 established a new state of the art on the RoboCasa365 benchmark, surpassing previous leading systems.
Foundation Model Designed for Multiple Robot Types
Following pretraining, Xiaomi adapted the model to operate different robotic platforms, including:
- Mobile manipulators.
- Dual-arm robots.
- Wheeled robotic systems.
The post-training process aligned the model with different robot embodiments and natural-language instructions while preserving the general manipulation capabilities learned during large-scale pretraining.
Implications for Robotics AI
Xiaomi’s findings reinforce a growing trend in robotics research: data quality and diversity may become the primary bottleneck for future robot intelligence.
Rather than focusing exclusively on larger neural networks, researchers are increasingly investing in:
- Large-scale real-world robot datasets.
- Automated data labeling.
- Cross-robot learning.
- Data-efficient adaptation techniques.
- Vision-language-action foundation models.
The company says it plans to release the model, code, and checkpoints publicly, potentially accelerating research across the robotics community.
Why the Research Matters
| Area | Potential Impact |
|---|---|
| Large-scale data collection | Better robot generalization |
| Automated labeling | Faster dataset creation |
| Foundation models | Broader task adaptability |
| Public release | Accelerated robotics research |
Looking Ahead
Xiaomi-Robotics-1 highlights an important shift in robotics AI development: collecting larger, more diverse real-world datasets may deliver greater improvements than simply building larger models. By training on more than 100,000 hours of human-collected manipulation data and demonstrating strong performance on unfamiliar tasks with minimal fine-tuning, Xiaomi has provided evidence that data scaling is becoming a critical driver of robot intelligence.
Looking ahead, the research could influence how robotics companies allocate resources, placing greater emphasis on scalable data collection, automated annotation, and efficient transfer learning rather than model size alone. If Xiaomi follows through on plans to release its model and training resources, the work could contribute to broader advances in autonomous robotics and accelerate the development of general-purpose robot foundation models across the industry.
Get the day’s top stories in your inbox
One concise email. No spam, unsubscribe anytime.



