Skip to main content
← All posts
Robotics9 min readSep 2026

Figure's Helix 2.5 Worked in 30 Homes It Had Never Seen. The Number That Matters Is 56%.

On 17 September 2026 Figure showed a humanoid tidying rooms, folding towels and making beds in 30 homes it had never seen, with no data collected in any of them. Pretraining took zero-shot success from 9% to 56%. Here is what that number actually means, straight from Figure's own write-up.

Figure AIHelix 2.5Humanoid RobotsZero-Shot LearningEmbodied AIVLA ModelsRobot Foundation Models
Dhruv Tomar

Dhruv Tomar

AI Solutions Architect

Tech Stack

Helix 2.5Index datasetvision-language-action modelswhole-body controlzero-shot generalization
30 unseen Bay Area homes — zero data collected in any of them
Zero-shot success 9% → 56% from Index pretraining alone (6x)
Three whole-body behaviours: tidying living rooms, folding towels, making beds
Half the task-specific data of a representative Helix 02 behaviour
Index generates roughly 35 minutes of new human experience every second

On 17 September 2026, Figure published results for Helix 2.5: a humanoid performing three household jobs across 30 Bay Area homes, with zero data collected in any of them.

Not thirty demos in a lab that looks like a home. Thirty real homes the robot had never been inside, with furniture it had never seen, in layouts nobody trained it on.

A half-made bed in morning light, with an unnaturally precise stack of folded towels on the chair beside it.
A half-made bed in morning light, with an unnaturally precise stack of folded towels on the chair beside it.

## What actually happened

Figure adapted a single pretrained base model into three behaviours:

  • -Tidying living rooms
  • -Folding towels
  • -Making beds

Together those span locomotion, rigid manipulation, deformable manipulation (a towel has no fixed shape), two-handed coordination, and active perception — looking around to decide what to do next. That combination is the point. Each one alone has been demonstrated before.

## Why "unseen" is the whole story

Here is the thing people outside robotics consistently underrate: making a robot do a task is not the hard part. Making it do that task somewhere new is.

A robot trained in one kitchen learns that kitchen. Different light, a different counter height, a bedsheet with a different drape, and the policy falls apart. The standard fix is to go collect data in the new place and fine-tune. That does not scale to houses, because there are a billion houses and they are all different.

Zero-shot means none of that happened. No data collection in those 30 homes, no fine-tuning, no adaptation.

## The number that makes it real

| | | |---|---| | Zero-shot success before Index pretraining | 9% | | Zero-shot success after | 56% |

A roughly 6× improvement from pretraining alone — the task-specific training did not change, the base model did.

Two details underneath that I find more convincing than the headline:

Helix 2.5 used half the task-specific data of a comparable Helix 02 behaviour. Better pretraining bought them a cheaper fine-tune. That is the same curve language models walked, arriving in robotics.

No single evaluation task makes up more than 1.90% of the pretraining dataset. This is the detail that kills the obvious objection. If "making beds" were 30% of training, 56% would mean memorisation. At under 2%, the behaviour is coming from general pretraining, not from having seen the answer.

## Where the data comes from

The base model is pretrained on Index, which Figure describes as a global-scale dataset of human behaviour currently generating roughly 35 minutes of new human experience every second.

Sit with that rate. It is the robotics answer to the internet-scale text corpus — the bet that you cannot hand-engineer household competence, you have to absorb it, and whoever absorbs the most wins. Figure announced Index on 25 August, three weeks before these results.

## The honest limit

56% means it fails nearly half the time.

Figure says so themselves, plainly: *"general humanoid robotics is not solved."* Their evaluation also aborts a trial whenever a human has to step in for safety — so that 56% is measured under supervision, not in an empty house.

And in a home, the failure rate is not the only thing that matters — the failure mode is. A robot that makes a bed badly is a nuisance. A robot that pulls a shelf over is a different category of problem. Figure publishes no failure-mode breakdown, and the write-up does not specify which robot model ran the evaluations.

## What this means in India

The Indian read on a home humanoid is different from the American one, and the difference is economic, not technical.

Household work here is already done, affordably, by people. A machine that succeeds 56% of the time and needs a human standing by for safety does not compete with that on cost or reliability — not now, and not at the price a humanoid will carry for years.

So the near-term Indian opportunity is not the home. It is everywhere the same generalization advance shows up first: warehouses, factory lines, and inspection work — structured, repetitive, supervised environments where a supervised 56% is a starting point rather than a disqualification, and where the labour is dangerous rather than cheap.

The thing to watch is not when a humanoid reaches your house. It is when this generalization curve meets a job that pays enough to absorb the failure rate.

## What I could not verify

While researching this I found several sites confidently quoting a $25 per robot-operating-hour price for Figure 03 at BMW, plus a specific 40-unit fleet size. Neither figure appears anywhere on Figure's own news page. I am leaving both out, and I would be careful with any article that states them as fact.

Figure's own BMW post is dated 30 June 2026 and does not carry commercial pricing terms.

## The question I cannot answer

What success rate makes a home robot actually tolerable?

It is clearly not 56%. But I do not think it is 99% either, because humans do not fold towels at 99% and we live with that fine. The real threshold probably has less to do with the success rate and more to do with what happens on the 5% — whether failure means a crooked duvet or a broken television.

Nobody is publishing that number. I would like to see it before I believe any home-robot timeline.

Want to build something like this?

I architect and deploy end-to-end AI systems — from MVP to revenue.

Let's Talk

Or ask Angelina — my AI twin in the bottom-right corner. She knows my full build history, live GitHub, and how I'd approach your project.