Tech Stack
On 17 September 2026, Figure published results for Helix 2.5: a humanoid performing three household jobs across 30 Bay Area homes, with zero data collected in any of them.
Not thirty demos in a lab that looks like a home. Thirty real homes the robot had never been inside, with furniture it had never seen, in layouts nobody trained it on.

## What actually happened
Figure adapted a single pretrained base model into three behaviours:
- -Tidying living rooms
- -Folding towels
- -Making beds
Together those span locomotion, rigid manipulation, deformable manipulation (a towel has no fixed shape), two-handed coordination, and active perception — looking around to decide what to do next. That combination is the point. Each one alone has been demonstrated before.
## Why "unseen" is the whole story
Here is the thing people outside robotics consistently underrate: making a robot do a task is not the hard part. Making it do that task somewhere new is.
A robot trained in one kitchen learns that kitchen. Different light, a different counter height, a bedsheet with a different drape, and the policy falls apart. The standard fix is to go collect data in the new place and fine-tune. That does not scale to houses, because there are a billion houses and they are all different.
Zero-shot means none of that happened. No data collection in those 30 homes, no fine-tuning, no adaptation.
## The number that makes it real
| | | |---|---| | Zero-shot success before Index pretraining | 9% | | Zero-shot success after | 56% |
A roughly 6× improvement from pretraining alone — the task-specific training did not change, the base model did.
Two details underneath that I find more convincing than the headline:
Helix 2.5 used half the task-specific data of a comparable Helix 02 behaviour. Better pretraining bought them a cheaper fine-tune. That is the same curve language models walked, arriving in robotics.
No single evaluation task makes up more than 1.90% of the pretraining dataset. This is the detail that kills the obvious objection. If "making beds" were 30% of training, 56% would mean memorisation. At under 2%, the behaviour is coming from general pretraining, not from having seen the answer.
## Where the data comes from
The base model is pretrained on Index, which Figure describes as a global-scale dataset of human behaviour currently generating roughly 35 minutes of new human experience every second.
Sit with that rate. It is the robotics answer to the internet-scale text corpus — the bet that you cannot hand-engineer household competence, you have to absorb it, and whoever absorbs the most wins. Figure announced Index on 25 August, three weeks before these results.
## The honest limit
56% means it fails nearly half the time.
Figure says so themselves, plainly: *"general humanoid robotics is not solved."* Their evaluation also aborts a trial whenever a human has to step in for safety — so that 56% is measured under supervision, not in an empty house.
And in a home, the failure rate is not the only thing that matters — the failure mode is. A robot that makes a bed badly is a nuisance. A robot that pulls a shelf over is a different category of problem. Figure publishes no failure-mode breakdown, and the write-up does not specify which robot model ran the evaluations.
## What this means in India
The Indian read on a home humanoid is different from the American one, and the difference is economic, not technical.
Household work here is already done, affordably, by people. A machine that succeeds 56% of the time and needs a human standing by for safety does not compete with that on cost or reliability — not now, and not at the price a humanoid will carry for years.
So the near-term Indian opportunity is not the home. It is everywhere the same generalization advance shows up first: warehouses, factory lines, and inspection work — structured, repetitive, supervised environments where a supervised 56% is a starting point rather than a disqualification, and where the labour is dangerous rather than cheap.
The thing to watch is not when a humanoid reaches your house. It is when this generalization curve meets a job that pays enough to absorb the failure rate.
## What I could not verify
While researching this I found several sites confidently quoting a $25 per robot-operating-hour price for Figure 03 at BMW, plus a specific 40-unit fleet size. Neither figure appears anywhere on Figure's own news page. I am leaving both out, and I would be careful with any article that states them as fact.
Figure's own BMW post is dated 30 June 2026 and does not carry commercial pricing terms.
## The question I cannot answer
What success rate makes a home robot actually tolerable?
It is clearly not 56%. But I do not think it is 99% either, because humans do not fold towels at 99% and we live with that fine. The real threshold probably has less to do with the success rate and more to do with what happens on the 5% — whether failure means a crooked duvet or a broken television.
Nobody is publishing that number. I would like to see it before I believe any home-robot timeline.
Want to build something like this?
I architect and deploy end-to-end AI systems — from MVP to revenue.
Let's TalkOr ask Angelina — my AI twin in the bottom-right corner. She knows my full build history, live GitHub, and how I'd approach your project.