THE SIGNAL IN ONE SENTENCE

Spirit AI, a Beijing embodied-AI startup, says robot software could reach a major capability milestone by mid-2027: a person would give a natural-language instruction and a robot would attempt a reasonable sequence of physical actions. That is a company forecast, not a demonstrated general-purpose system. The evidence available today is narrower and more interesting. Reuters visited Spirit AI's training center, where workers wearing sensors repeatedly opened refrigerators, unlocked safes, cut vegetables, and demonstrated other motions. The company says about 1,000 contractors collect this kind of real-world data in homes and factories across China. It reports a 90 percent success rate for simple tasks in structured living-room environments and says tens of its wheeled Moz1 humanoids are working on production lines at CATL and JD.com. Fine motor work, unseen tasks, unpredictable homes, and independent validation remain unfinished. The plain signal is that the robot brain race is not only a contest between clever models. It is an industrial system made from human demonstrations, task definitions, sensor records, factory trials, safety limits, and thousands of failures that the public rarely sees.

01

WHAT ACTUALLY CHANGED

Spirit AI co-founder and chief scientist Gao Yang told Reuters at the company's Beijing offices on September 17 that he expects what he calls a GPT-3.0 milestone for embodied intelligence by mid-2027. His definition is practical but still broad: a person speaks a task, and the robot attempts a sensible sequence of physical actions. The phrase is Gao's forecast and analogy. It is not a standardized robotics benchmark or a capability demonstrated in the article.

Reuters observed dozens of workers at Spirit AI's Beijing data center wearing sensors and repeating actions including opening refrigerators, unlocking safes, and cutting vegetables. Gao said the company uses around 1,000 contractors nationwide to gather movement data in homes and factories. That makes human demonstration labor part of the product, not a footnote hidden behind the model.

Spirit AI says its robots achieve a 90 percent success rate on simple tasks in structured living-room environments. The company did not publish the task list, number of trials, failure categories, tolerance for partial completion, model version, comparison systems, video record, or an outside replication. A tidy test room and a real home are different planets wearing the same sofa.

The company says tens of Moz1 wheeled humanoids are deployed on production lines at battery maker CATL and retailer JD.com, which is also an investor. That is stronger evidence than a stage demonstration because a factory can expose reliability, integration, and maintenance problems. It remains a company-reported deployment without disclosed work hours, task completion, human interventions, incidents, savings, or customer evaluation.

Gao said fine motor actions such as unscrewing a bottle cap and tasks the system has not seen remain difficult. He expects industrial uses to arrive first, simpler commercial-service work after roughly two years, and household deployment much later. This sequence matters because factories can constrain objects, routes, lighting, people, and acceptable behavior in ways a home cannot.

Spirit AI says it relies overwhelmingly on real-world data rather than simulation. Gao argued that simulators handle rigid objects better than flexible ones such as deformable electrical cables. He also said varied, imperfect demonstrations can improve learning faster than collecting only repeated clean motions. The claim suggests a valuable research direction, but the company has not published an ablation showing how much the broader data improves performance.

The startup says whole-body force control triggers emergency braking when interaction force becomes excessive. That is one useful protective layer, not a complete safety case. Force thresholds do not by themselves address perception errors, sharp tools, dropped objects, trapped fingers, navigation near children, software updates, remote access, worker override, or the long tail of unfamiliar rooms.

02

WHY THIS MATTERS

Physical AI makes the training-data supply chain visible. Language models learn from text that already exists. General-purpose robots need records of hands, arms, tools, surfaces, timing, resistance, mistakes, and recovery. The people creating those records deserve clear contracts, safe work, meaningful consent, limits on home data, and credit in the economics of the system.

A success rate without a denominator is decorative math. Ninety percent can mean nine safe completions out of ten identical trials or nine hundred completions across a thousand varied tasks. Buyers and regulators need the task distribution, operating conditions, severity of failure, intervention rate, time per task, and performance on novel objects before the number can support a decision.

Factories are likely to be the first honest proving ground. A production line can define the workspace, fixtures, materials, stopping zones, and fallback procedure. It can also measure throughput and downtime. If robots cannot produce durable value there, the dream of a machine wandering through a cluttered kitchen while holding a knife should stay in the dream department.

Human demonstration data can carry human variation and human surveillance risk at the same time. A system benefits from different heights, speeds, grips, disabilities, cultural routines, tools, and room layouts. Collecting those differences inside homes can also capture faces, voices, documents, medications, family members, and daily patterns unrelated to the task. Data minimization has to happen before the sensor suit turns on.

Safety is a system property, not a braking feature. Mechanical force limits, perception confidence, task permissions, tool restrictions, geofences, independent emergency stops, update controls, logs, incident review, and human authority must work together. The more a foundation model chooses its own action sequence, the less adequate one preset force threshold becomes.

China's advantage may come from operational scale as much as model architecture. Spirit AI has raised more than $670 million since 2024, employs about 300 people, and has organized a large demonstration workforce and factory relationships. Capital and data volume can accelerate iteration. They can also produce fast-moving systems whose labor conditions, evaluation methods, and incident records remain opaque.

FIG. 168TURN A HUMAN MOTION INTO A ROBOT CAPABILITY
1A WORKER DEMONSTRATES A TASK WITH SENSORS→
2DATA IS CLEANED, LABELED AND JOINED TO OUTCOMES→
3A MODEL LEARNS AN ACTION SEQUENCE→
4THE ROBOT TRIES IT IN A STRUCTURED TEST ROOM→
5NOVEL OBJECTS AND FAILURES EXPAND THE TEST→
6FACTORY EVIDENCE DECIDES WHETHER DEPLOYMENT GROWS
A demonstration is the beginning of the evidence chain. A safe, repeatable result in an unfamiliar place is the part that earns the word capability.

03

WHERE IT COULD HELP

  • Create a public task card for every robot capability that names the environment, objects, starting state, success rule, timeout, number of trials, intervention policy, model and hardware version, failure severity, and performance on familiar and unfamiliar variants
  • Treat demonstration workers as skilled data producers with training, ergonomic rotation, knife and tool safety, injury reporting, fair pay, privacy controls, access to their records, and a route to challenge unsafe or inappropriate collection assignments
  • Separate demonstration data from incidental household data by processing sensors locally when possible, masking bystanders and documents, collecting only required channels, setting deletion dates, recording consent, and auditing every reuse beyond the original task
  • Stage deployment from fenced factory cells to shared industrial spaces, then constrained commercial settings, with independent stop criteria, logged human interventions, near-miss reporting, software-version control, and customer evidence before expanding the operating domain
  • Design physical safeguards around the complete action chain, including permissioned tools, speed and force envelopes, redundant emergency stops, safe poses after uncertainty, human override outside the model, network isolation, update rollback, and investigation-ready logs
  • Measure value per completed task using throughput, quality, downtime, maintenance, energy, human supervision, worker injury, rejected output, and recovery time rather than counting robots shipped or demonstrations collected

KEEP A HAND ON THE WHEEL

Reuters directly visited Spirit AI's Beijing offices, observed the training center, and interviewed Gao Yang. The 2027 milestone, 90 percent structured-room result, contractor count, factory deployments, funding, valuation, training strategy, and safety descriptions remain company statements unless Reuters independently observed the specific fact. Spirit AI's product page verifies that it markets Moz1 as a full-force-controlled humanoid, but the page does not publish the evaluation behind the 90 percent figure. No independent benchmark, complete task set, raw trial data, customer performance report, workplace injury record, data-worker contract, privacy assessment, safety case, incident log, or home-deployment evidence was identified before publication. The comparison with GPT-3 is an analogy, not a shared technical scale. A robot attempting a sequence is not the same as completing it reliably, safely, or economically. Watch for a reproducible evaluation, novel-task performance, intervention rates, factory uptime, customer-authored results, worker protections, data-governance details, third-party safety testing, and a public account of failures severe enough to stop expansion.

04

TERMS WORTH KEEPING

SOURCES AND VERIFICATION STATUS

This article was written from the materials below. Product claims and dates were checked against those sources on September 18, 2026.

PUBLICATION RECEIPT: Revision 1. Published September 18, 2026.

THE PUBLICATION ENGINE

WANT A SIGNAL OF YOUR OWN?

We build source-grounded publications, private briefings, and editorial systems for organizations with something useful to say.

WORK WITH US