Technology
Figure AI Helix-02: Two Robots Make a Bed Without Talking to Each Other
Figure AI released a demo on May 8, 2026 showing two F.03 humanoid robots autonomously resetting a bedroom in under two minutes — including collaborative bed-making — using a single neural network with no communication link between the robots.
On May 8, 2026, Figure AI posted a video that most robotics engineers will spend the next several months thinking about: two F.03 humanoid robots, running a single learned neural network called Helix-02, reset a bedroom in under two minutes. No scripts. No teleoperation. No shared planner between them.
They opened doors, hung clothes, closed a book, operated a trash can foot pedal while balancing on one leg, and then worked together to spread and smooth a comforter across a bed, with each robot reading its partner's position from visual cues alone. The company called it "the first demonstration of a single learned neural network performing multi-humanoid collaborative locomanipulation, directly from pixels to actions." That framing is not hype. It is a precise technical claim, and it appears to hold up.
Key Stats
<2 min
Full Bedroom Reset
2
Robots, Zero Coordination Link
$39B
Figure AI Valuation
350+
F.03 Units Produced
How Helix-02 Works
Helix-02 is a Vision-Language-Action (VLA) policy — a type of neural network that takes raw camera pixels and proprioceptive sensor data as input and outputs full-body motor commands directly. There is no intermediate perception layer, no planner, no behavior tree. The network processes what the robot's cameras see and what its joints feel, then decides how to move every degree of freedom in its body at once.
The predecessor, Helix-01, focused on upper-body dexterity while the robot remained stationary. Helix-02 extends that to the whole body: walking, balance, stance shifts, and manipulation are all handled by the same unified policy. When a robot in the May 8 demo balanced on one leg to press a trash can foot pedal, that was not a separate balance controller handing off to a manipulation controller. It was a single network making both decisions at the same time, at policy rate.
🧠 Single Neural Network
One learned VLA policy handles locomotion, manipulation, balance, and multi-robot coordination simultaneously. No modular handoffs between subtasks.
📷 Pixels to Actions
Raw camera pixels feed directly into motor commands. No intermediate object detection, scene graph, or 3D reconstruction required for the core policy.
🤝 Implicit Coordination
Each robot infers its partner's intent from visual observation of motion alone, the same way two people read each other when folding a sheet. No messages pass between them.
📊 Data-Driven Expansion
New capabilities are added by collecting and training on new data. The core algorithm does not change, which means the cost of adding a skill is a data collection problem, not a software architecture problem.
The February 2025 version of Helix showed two robots coordinating to put away groceries. The May 2026 version handles a significantly more complex sequence involving deformable objects (bedding, clothing), articulated objects (doors, books), and dynamic balance demands (one-leg stance). The jump in complexity over 15 months is measurable and substantial.
Under the Hood: Why This Demo Is Technically Difficult
Three distinct technical challenges combine in the bedroom demo, each of which would be considered a research-grade problem on its own.
Multi-agent implicit coordination is the first. When two robots share a task, every action one takes changes the problem the other is solving in real time. Classic multi-robot systems handle this with explicit communication: robot A broadcasts its state, robot B factors it into its plan. Helix-02 does not do that. Each robot watches the other through its cameras and infers intent from motion, then acts. This requires the policy to model the behavior of another agent whose exact actions are unknown — a hard problem in reinforcement learning and game theory.
Deformable object manipulation is the second. A comforter has no fixed geometry. Its shape changes with every contact, and there is no canonical "grasp point" that works across different bed configurations. The robot must predict how the fabric will respond to a pull, which depends partly on what the other robot is simultaneously doing on the opposite side. Deformable object manipulation has historically been one of the hardest open problems in robotics — most manipulation research deliberately avoids it.
Whole-body locomanipulation across task types is the third. The bedroom sequence requires the robot to walk to multiple locations, switch between rigid object grasping (door handles, headphones), articulated object manipulation (book cover), deformable manipulation (comforter, clothing), and dynamic balance (foot pedal operation) without any scripted transition between modes. The full sequence runs in under two minutes, which means each subtask receives only seconds of execution time before the policy must transition to the next.
AI-generated image
The Helix-02 VLA architecture processes raw sensor data and camera pixels simultaneously to produce whole-body motor commands. Source: AI-generated editorial illustration.
Key Insight
The most significant claim in the Figure AI announcement is not the bedroom demo itself — it is that no algorithm changes were required to add these new capabilities. New skills came from new training data. That architectural property, if it generalizes, means the system can scale to new environments and tasks without rebuilding the underlying model each time.
Leading Humanoid AI Systems: Capability Comparison
| System | Figure AI Helix-02 | Tesla Optimus | Boston Dynamics Atlas | 1X NEO |
|---|---|---|---|---|
| Policy Type | Unified VLA (pixels to actions) | End-to-end neural (internal) | Model-based + learned | Teleoperation-to-autonomy |
| Multi-Robot Coordination | Yes (implicit, no comms) | Not demonstrated | Not demonstrated | Not demonstrated |
| Deformable Objects | Demonstrated (bedding) | Partial | Limited | Partial |
| Whole-Body Locomanipulation | Full integration | Partial | Strong locomotion | Household focus |
| Robot Height / Weight | 1.72 m / 61 kg | 1.73 m / 57 kg | 1.9 m / 90 kg | 1.67 m / 30 kg |
| Deployment Status | Industrial (BMW, scaling) | Internal (Tesla factories) | Commercial (Hyundai fleet) | Full prod. May 2026 |
| Company Valuation | $39 billion | Part of Tesla (~$1T) | Part of Hyundai group | Not disclosed |
Figure AI: The Full Picture
Figure AI, founded in 2022 and headquartered in Sunnyvale, California, has moved faster than most analysts expected. The company closed a funding round valuing it at $39 billion after raising over $1 billion from investors including Microsoft, OpenAI, and Nvidia. CEO Brett Adcock has publicly stated that the goal is general-purpose humanoid robots for household use, with industrial deployment as the near-term revenue path.
The F.03 platform is their current production robot. Over 350 units have been manufactured as of May 2026, with an active commercial deployment at BMW's manufacturing facilities. Earlier Figure models handled structured pick-and-place tasks in controlled factory environments. Helix-02 represents the company's push toward unstructured environments, where the layout, objects, and conditions are not predetermined.
• Founded: 2022, Sunnyvale, California
• CEO: Brett Adcock (previously Archer Aviation, Vettery)
• Valuation: $39 billion (post-latest round, 2025-2026)
• Investors: Microsoft, OpenAI, Nvidia, Intel Capital, Parkway Venture Capital
• Active deployment: BMW automotive manufacturing, expanding
• Units produced: 350+ F.03 robots as of May 2026
• AI partner: OpenAI (large language model integration for task understanding)
• Target market: Industrial in near-term, household in 3-5 year horizon
The Helix project has moved in distinct public milestones. In early 2024, Figure released the first video of OpenAI language model integration for task instruction. In February 2025, Helix-01 showed two-robot grocery collaboration. The May 2026 Helix-02 demo adds whole-body locomotion integration, deformable object handling, and a significantly harder multi-robot coordination task. Each release is meaningfully more capable than the last, which is not always true in a field that sometimes stages demos ahead of real capability.
What This Means for Physical AI
AI-generated image
Industrial automation at scale: the near-term deployment environment for systems like Helix-02, before household applications become viable. Source: AI-generated editorial illustration.
The household robot market has been declared imminent for decades and has never arrived at scale. What makes 2026 different from prior cycles is not a single company's demo — it is the convergence of three things that were not simultaneously true before: capable hardware at declining cost, large-scale training data from human demonstrations, and neural architectures (VLA models) that can generalize across task types without per-task programming.
The Helix-02 demo matters most as a signal about the skill acquisition cost curve. If new capabilities genuinely require only new data rather than new code, then the marginal cost of adding a household skill to a robot approaches the cost of collecting a few hundred hours of human demonstration footage. That is still significant but is categorically cheaper than hiring engineers to write and debug new controllers for every new task. It shifts the scaling problem from software complexity to data collection logistics.
For industrial buyers, the immediate implication is that robots deployed today on factory floor tasks can potentially be retrained for adjacent tasks without hardware replacement. A Figure robot doing parts sorting in 2026 could potentially learn to do quality inspection in 2027 with a data collection campaign rather than a new procurement cycle. That changes the ROI calculus significantly.
Competitive Signals to Watch
• Tesla Optimus: Internal deployment at Fremont Gigafactory continues; no external commercial sales announced. Tesla's data advantage from its vehicle fleet may translate to robot training at scale when the company chooses to accelerate.
• 1X Technologies NEO: Full production launched May 2026 in California at ~$20,000 per unit, targeting household use directly. Lighter at 30 kg, focused on teleoperation-to-autonomy transfer.
• Unitree Robotics: Airport trials at Tokyo Haneda with Japan Airlines launched the same week as the Figure demo — a coincidence that underscores the pace of real-world deployment across the sector.
• Samsung and Meta: Both companies announced accelerated humanoid investment in May 2026, entering a field that had previously been dominated by pure-play robotics startups.
The economic pressure driving deployment is also clarifying. Japan's trial at Haneda airport involves robots handling baggage amid documented labor shortages from an aging population. The U.S. warehouse industry has a structural gap in physically demanding roles that has widened through the 2020s. The demand side is real, which means companies reaching commercial-grade capability now have customers ready to buy.
The 12-Month Outlook
Figure AI has not announced a specific timeline for household deployment, but the company's trajectory points toward continued expansion in industrial settings through 2026, with household pilots likely in 2027. The BMW deployment provides real operational data at volume — 350+ units in active use generates far more training signal than any lab environment could. That data loop, proprietary to Figure, is a compounding competitive asset.
The technical milestones to watch over the next 12 months: whether multi-robot coordination scales to three or more units in a shared space, whether Helix can handle fully novel environments with no prior mapping, and how error recovery performs when the robot makes a mistake mid-task in an unsupervised setting. The bedroom demo was staged and photographed — production deployment means unpredictable environments, unexpected objects, and users who do not cooperate with optimal robot conditions.
Commercially, watch for the first Figure consumer product announcement, any new industrial partnerships beyond BMW, and whether competitors match the multi-robot coordination capability in 2026. Tesla has the manufacturing scale to accelerate quickly when it chooses to. Boston Dynamics has the locomotion depth. Neither has yet matched the VLA approach Figure demonstrated on May 8.
Frequently Asked Questions
What exactly is a Vision-Language-Action (VLA) policy?
A VLA policy is a neural network that takes visual input (camera images), language context (task instructions or scene descriptions), and sensor data as input, then outputs motor commands directly. The key difference from older robot control systems is that there is no separate perception layer, planner, or controller — the network handles all of it in a single forward pass. In Helix-02, the model outputs joint position targets and body force commands at the rate the robot needs to stay balanced and on-task, which is many times per second.
How do the two robots coordinate without talking to each other?
Each robot's VLA policy was trained on data that includes scenes with another robot present. Over training, the policy learned to predict what a partner robot will do next based on its observed motion, the same way people coordinate on tasks like folding laundry without verbal instructions. The robots use head orientation and body positioning as social signals — when one robot moves to one side of the bed, the other moves to the opposite side. There is no communication channel between them; all coordination happens through visual inference from each robot's own cameras.
When will Figure AI robots be available for home use?
Figure AI has not announced a consumer product or a home deployment timeline. The company's current commercial focus is industrial: the BMW partnership is the anchor deployment, and subsequent industrial clients are the expected near-term expansion. Industry observers generally estimate household humanoid robots from leading companies will reach limited pilot availability in 2027-2028, with broader commercial rollout in the 2029-2031 range, contingent on both capability maturation and cost reduction to under $30,000 per unit.
How does Figure AI's approach differ from Tesla's Optimus program?
The key structural difference is data source and deployment path. Tesla's approach leverages data from its automobile fleet (cameras, sensors, real-world driving scenarios) and applies analogous techniques to robot training, with internal deployment at Tesla factories providing more data. Figure AI's approach relies on human demonstration data collected specifically for robot tasks, with external commercial partnerships generating operational data. Tesla has not publicly demonstrated multi-robot coordination or deformable object handling at the level Figure showed on May 8. Both companies are pursuing VLA-style unified policies, but Figure is ahead in publicly demonstrated complexity as of May 2026.
Is the bedroom demo fully autonomous or is any human assistance involved?
According to Figure AI's announcement, the demo runs fully autonomously with no teleoperation, no scripted sequences, and no human intervention during the task. The robot receives a high-level task instruction (reset the bedroom) and executes the full sequence independently. The environment was staged and the bedroom was set up specifically for the demo, which is standard practice — Figure did not claim the robots can handle arbitrary unknown rooms. The distinction between "staged but autonomous" and "scripted" matters technically: the robots genuinely decide how to execute the task, they are just doing so in a prepared environment.
What This Demo Actually Signals
The Figure AI Helix-02 bedroom demo is not proof that household robots are ready for mass deployment. The environment was staged, the task sequence was chosen to be demonstrable, and the gap between a tidy bedroom demo and a messy real home full of novel objects is large. What the demo does prove is that the technical barriers that have blocked multi-robot coordination and deformable object handling for decades have started to fall, and they are falling faster than the prior decade of research suggested they would.
The more important number in the Figure AI story is not the two minutes to reset a bedroom. It is the 350 F.03 robots already in industrial deployment, generating real operational data at BMW and other facilities. That data compounds. Every robot-hour of real deployment makes the next version of Helix more capable, which justifies deploying more robots, which generates more data. Companies that reach that loop early hold a structural advantage that becomes harder to close over time.
The Bottom Line: Figure AI's Helix-02 is the clearest evidence yet that multi-robot physical AI is transitioning from research curiosity to commercial capability — and the company's data flywheel from active BMW deployment gives it a compounding advantage that money alone cannot quickly replicate.
The next 12 months will test whether the capability gains hold outside staged environments, whether competitors can close the gap, and whether the industrial deployment model generates enough revenue and data to fund the eventual consumer push. The robotics race of 2026 is no longer about who can build the most impressive demo. It is about who is generating real operational data at scale.