Chinese technology giant Xiaomi has unveiled Xiaomi-Robotics-1, a new vision-language-action (VLA) foundation model that challenges a widely held assumption in robotics AI: that larger models inevitably deliver better performance. Instead, Xiaomi’s research shows that scaling high-quality training data has a much greater impact on robot capabilities than simply increasing model size, highlighting the growing importance of data collection in the race to build more capable autonomous robots.

The model was trained on more than 100,000 hours of real-world manipulation data collected using handheld, camera-equipped grippers operated by humans rather than expensive robotic systems. Xiaomi says this approach enabled it to build one of the largest robotics datasets to date while significantly lowering data collection costs. The company also developed an automated labeling pipeline that used AI to generate natural-language descriptions for the recorded motions, completing annotation of the entire dataset in about two weeks.

More Data Delivers Bigger Gains Than Larger Models

A central finding of Xiaomi’s research is that increasing the amount of training data consistently produced greater performance improvements than increasing the model’s parameter count.

During testing, researchers observed that:

  • Larger models improved prediction accuracy.
  • Increasing training data produced substantially larger performance gains.
  • Performance continued improving as more data was added, with no clear plateau observed.

The findings suggest that robotics AI may benefit more from scaling diverse real-world datasets than from continually building larger neural networks, at least under current training approaches.

Key Research Highlights

MetricResult
Training data100,000+ hours
Data collection methodHuman-operated handheld grippers
Data labelingAI-generated automated annotations
Primary findingMore training data outperformed larger model size

Human-Collected Motion Data Powers the Model

Instead of relying exclusively on physical robots to generate demonstrations, Xiaomi equipped handheld grippers with cameras and sensors, allowing people to naturally perform manipulation tasks.

This approach enabled researchers to:

  • Collect data more quickly.
  • Capture a wider variety of environments.
  • Reduce the cost of robot demonstrations.
  • Generate richer manipulation datasets.

The resulting trajectories were automatically converted into language-action pairs using an AI-based labeling system, creating a scalable pipeline for robotics training.

Image 57

Strong Performance on New Tasks

Xiaomi reports that the foundation model can generalize to unfamiliar environments with relatively little additional training.

In evaluation tests:

  • Less than 10 hours of task-specific training data were required for each new task.
  • The model achieved an average success rate of about 75% across four unseen manipulation tasks.
  • A competing model from Physical Intelligence reportedly achieved approximately 40% on the same benchmark.

Researchers also demonstrated the model performing complex activities such as suitcase packing, organizing household objects, and manipulating soft materials like paper.

Benchmark Performance

BenchmarkXiaomi-Robotics-1
RoboCasa365 success rate57.6%
Average success on unseen tasks~75%
Task-specific fine-tuningLess than 10 hours

The accompanying research paper states that Xiaomi-Robotics-1 established a new state of the art on the RoboCasa365 benchmark, surpassing previous leading systems.

Foundation Model Designed for Multiple Robot Types

Following pretraining, Xiaomi adapted the model to operate different robotic platforms, including:

  • Mobile manipulators.
  • Dual-arm robots.
  • Wheeled robotic systems.

The post-training process aligned the model with different robot embodiments and natural-language instructions while preserving the general manipulation capabilities learned during large-scale pretraining.

Implications for Robotics AI

Xiaomi’s findings reinforce a growing trend in robotics research: data quality and diversity may become the primary bottleneck for future robot intelligence.

Rather than focusing exclusively on larger neural networks, researchers are increasingly investing in:

  • Large-scale real-world robot datasets.
  • Automated data labeling.
  • Cross-robot learning.
  • Data-efficient adaptation techniques.
  • Vision-language-action foundation models.

The company says it plans to release the model, code, and checkpoints publicly, potentially accelerating research across the robotics community.

Why the Research Matters

AreaPotential Impact
Large-scale data collectionBetter robot generalization
Automated labelingFaster dataset creation
Foundation modelsBroader task adaptability
Public releaseAccelerated robotics research

Looking Ahead

Xiaomi-Robotics-1 highlights an important shift in robotics AI development: collecting larger, more diverse real-world datasets may deliver greater improvements than simply building larger models. By training on more than 100,000 hours of human-collected manipulation data and demonstrating strong performance on unfamiliar tasks with minimal fine-tuning, Xiaomi has provided evidence that data scaling is becoming a critical driver of robot intelligence.

Looking ahead, the research could influence how robotics companies allocate resources, placing greater emphasis on scalable data collection, automated annotation, and efficient transfer learning rather than model size alone. If Xiaomi follows through on plans to release its model and training resources, the work could contribute to broader advances in autonomous robotics and accelerate the development of general-purpose robot foundation models across the industry.

Get the day’s top stories in your inbox

One concise email. No spam, unsubscribe anytime.