Skip to content

Physical AI

Humanoid Says Real-World RL Can Push Robot Manipulation Toward 99.9% Reliability

Humanoid introduced KinetIQ Ascend, a reinforcement learning approach designed to push robot manipulation toward 99.9% reliability at human speed or faster. The company reported task-level gains on real bimanual hardware, including 412 bearing rings per hour, 98% handover success, and 98.9% tote handling success.

By Cara Voss · July 7, 2026

Humanoid Says Real-World RL Can Push Robot Manipulation Toward 99.9% Reliability

Humanoid, the London robotics company behind the HMND 01 platform, has introduced KinetIQ Ascend, a reinforcement learning approach aimed at reaching 99.9% manipulation reliability at human speed or faster on industrial tasks.

The claim is not another choreographed warehouse video. Humanoid published task-level results from real bimanual hardware, including a 42% throughput gain in machine feeding, an 85% throughput gain in item handover, and a jump from 77.6% to 98.9% success in tote handling. The company says the runs took days of robot time, not months of simulation-only work.

Industrial training cell with bearing rings, conveyors, and machine vision sensors AI-generated image

A factory training cell is the right mental model for KinetIQ Ascend. The system learns from repeated real-world task attempts rather than only copying teleoperated demonstrations.

Key Stats

99.9%

Reliability Target

412/hr

Bearing Ring Throughput

98%

Item Handover Success

98.9%

Tote Handling Success

Why This Matters

Humanoid robotics is moving through a credibility filter. It is no longer enough to show a robot walking across a stage, waving at a crowd, or picking one object on camera. Industrial customers care about a narrower question: can the robot repeat boring work for thousands of cycles without constant rescue?

That question turns manipulation into the main bottleneck. Walking is visible, but hands and arms decide whether a humanoid earns floor space in a warehouse, plant, service counter, or logistics cell. A mobile robot that can navigate but drops one tote in five is still a demo machine. A robot that can move, pick, hand off, and recover under messy conditions starts to look like automation.

Humanoid frames KinetIQ Ascend as a way to close the last reliability gap. The company says behavior cloning, the imitation-learning method used across much of robotics, can teach a base policy what a task should look like. It cannot reliably push that policy past the speed and quality of the human demonstrator. It also does not teach the policy the cost of failed actions.

Reinforcement learning changes the training loop. Instead of only copying human demonstrations, the robot practices the task, receives a success or failure signal, and changes the policy based on what actually works. That sounds simple in a slide deck. It is much harder on real hardware, where force limits, gripper timing, lighting drift, object variation, and mechanical wear all show up.

The signal

The important part is not the 99.9% target by itself. The important part is Humanoid showing task-level gains on physical bimanual hardware, with explicit baselines, robot-time durations, and failure-rate reductions. That is closer to an industrial evidence package than the average humanoid launch post.

What KinetIQ Ascend Actually Does

KinetIQ is Humanoid's AI framework for controlling its robots. KinetIQ Ascend extends that stack with end-to-end, vision-guided reinforcement learning for manipulation. Humanoid says the method runs on production vision-language-action models, or VLAs, trained on real two-arm humanoid hardware under deployment-like conditions.

In practical terms, Ascend starts from a behavior-cloned policy trained on teleoperation data. The robot already knows the rough shape of the task. Then RL pushes on speed, precision, recovery, and failure avoidance. The company says it uses both simulation and real hardware, but the emphasis is on real-world training because the deployed policy has to face the same dynamics, contact physics, and edge-compute constraints that training sees.

Humanoid's post describes a disaggregated training setup. The trainer runs updates on server hardware while rollout workers collect experience on robots or simulators. The robots run inference on edge compute, with Humanoid naming NVIDIA Jetson Thor as the current platform. That matters because industrial fleets cannot stream hundreds of live camera feeds out to a datacenter every time they need an action.

The method also includes safeguards for physical exploration. RL needs the robot to try actions that differ from its demonstrations, but real arms can break real parts. Humanoid says it uses active compliance, force-torque monitoring at the wrist, and auto-termination when force crosses a safe threshold. Episodes that trip safety stops become negative training signals, so the policy learns to avoid the behavior that caused the stop.

Base Policy

Behavior cloning teaches the initial manipulation behavior from teleoperated examples.

Real RL

The robot practices on hardware and optimizes against sparse success and failure signals.

Fleet Loop

The long-term pitch is deployed robots turning interventions into fresh training data.

Close-up of machine vision cameras and sensor hardware above an industrial conveyor AI-generated image

Vision-guided manipulation depends on stable perception, consistent action timing, and feedback from force and success signals.

The Three Tests

Humanoid reported three task families: machine feeding, object picking and handover, and bimanual tote handling. They are useful because they stress different parts of the stack. Machine feeding is mostly a throughput problem. Handover is a reliability and human-interaction problem. Tote handling is a two-arm coordination problem.

Task Baseline RL Result Why It Matters
Machine feeding 291 bearing rings per hour, 27.2-second cycle, 0.60 pick success 412 rings per hour, 19.5-second cycle, 0.67 pick success Shows RL turning higher execution speed into real throughput instead of more failed grasps.
Item handover 80% success on bottle handovers 98% success, 85% higher throughput, 35% shorter mean episode duration Shows failure reduction on a task where the endpoint depends on a person receiving the object.
Tote handling 122 totes per hour, 22.9-second episode, 77.6% success 279 totes per hour, 12.8-second episode, 98.9% success Shows the same training recipe working on a coordinated two-arm lift without task-specific changes.

The machine-feeding result is the most industrially legible. The robot picked 110 mm steel bearing rings from a cluttered bin and placed them on a conveyor. Humanoid says a behavior-cloned baseline moved about 291 rings per hour at 60 frames per second. Pushing the policy faster without RL made grasp quality worse. After roughly five days of continuous hardware training and a speed curriculum that stepped from 60 to 75 to 90 frames per second, the RL policy reached 412 rings per hour.

The handover result is more interesting than it first looks because the model trained only on water bottles during RL, then improved on objects it did not see in the RL phase. Humanoid says throughput improved by 40% on cylindrical snack cans and 13% on deformable bags. That does not prove broad general-purpose dexterity, but it does suggest the training improved the underlying picking skill rather than memorizing one object.

The tote result is the strongest headline number. Throughput more than doubled, from 122 to 279 totes per hour, and success moved from 77.6% to 98.9%. For a warehouse buyer, the difference between those two success rates is not cosmetic. At 77.6%, a robot is still a supervised experiment. At 98.9%, it is not finished, but the remaining gap becomes an operations and safety-engineering problem that can be measured.

The Speed Problem

A robot policy cannot simply be played faster like a video. Grippers have closure times. Joints have acceleration limits. Objects settle, slide, collide, and deform on physical timescales. If a policy retracts before the gripper has closed, the robot gets faster only at failing.

Humanoid's answer is a speed curriculum. The company describes raising the policy's execution frequency in steps, allowing quality to drop, then using RL to recover quality at the new speed before stepping up again. In one example, the company says a baseline policy at 60 frames per second dropped in quality when run at 75 frames per second, then recovered through RL. The same process then repeated at 90 frames per second.

That is the core commercial argument for KinetIQ Ascend. Human demonstrations can seed behavior, but they do not automatically define the fastest safe way to do the task. RL can search the timing and correction space inside the robot's own dynamics. If it works reliably across more tasks, speed becomes trainable rather than hand-tuned.

What still needs proof

  • Whether the same gains hold in customer sites with different lighting, fixtures, objects, and worker behavior.
  • Whether the 99.9% target is reached in long-running production evaluation, not just approached in task trials.
  • How much human supervision is required to reset scenes, label failures, handle exceptions, and maintain the hardware.
  • Whether the economics work once robot time, engineering setup, maintenance, and safety validation are included.

Why Behavior Cloning Is Not Enough

The industry's current training stack leans heavily on imitation. A human teleoperator performs a task, the model learns from the demonstration, and the robot attempts to reproduce the behavior. That approach is valuable because it gives the model a starting point in a huge action space. It also creates a ceiling.

If the demonstration is slow, the robot learns slow. If the human avoids rare failure states, the model may never learn how those failures begin. If the camera view misses a subtle cause of the human's action, the policy can learn a visible correlation that breaks the moment the setup changes. Humanoid calls this one of the reasons passive imitation falls short of industrial reliability.

RL attacks that problem by making failure visible to the learner. A misgrasp, time-out, unsafe force event, or incomplete pick becomes part of the optimization target. The robot does not just learn what the human did when everything went right. It learns which actions lead to regrasp loops, force stops, time-outs, and dropped objects.

The hard part is making that practical. Dense task-specific reward shaping can turn into a custom engineering project for every new station. Humanoid says it uses sparse and generic rewards: task success, recoverable error penalties, and implicit safety penalties through episode termination. That design is important because a humanoid company cannot hand-build a fragile reward function for every bin, tote, box, and conveyor a customer uses.

Dark data center and industrial edge-compute servers used for robot training AI-generated image

The commercial version of RL for robots is as much infrastructure as algorithm: rollout workers, edge inference, training servers, safety stops, and evaluation baselines.

The Deployment Read

This is not a customer production rollout. It is a technical release that uses production-style tasks and real hardware. That distinction matters. The reported gains are meaningful, but Humanoid is still showing its own evaluations on its own platform, not an audited third-party plant report.

The company has been lining up industrial credibility through partnerships and proof-of-concept work. Earlier this year it announced Bosch as a contract manufacturing partner for HMND 01 after a logistics proof of concept, and Schaeffler as a strategic partner tied to actuator supply and future robot deployment. KinetIQ Ascend fits that roadmap because reliability, not demo fluency, is what those partners will eventually need.

There is also a wider market implication. Apptronik is building a Robot Park data facility in Austin. Agility is pitching investors on operating robotics data from real warehouse deployments. Figure has been publishing long-run autonomy tests. The common thread is that humanoid robotics is becoming a data operation. The companies that can turn real robot time into reliable policies may compound faster than companies that only ship impressive hardware.

Humanoid's own argument is explicit: once robots deploy, supervisor interventions can become reward signals. A fleet can keep learning at the exact sites where it works. Every rollout can feed the next model generation. That is a powerful idea, but it brings buyer questions with it. Who approves policy updates? How are safety cases maintained? How are site-specific behaviors validated before they change on a working line?

Deployment Reality Check

Stage: Demo and internal technical validation
Robot count: Undisclosed
Task: Machine feeding, item handover, bimanual tote lifting
Supervision: Unspecified, with operator resets and safety stops implied
Evidence: Company technical post and trade coverage
Customer: No customer named for these specific trials

What To Watch Next

The next useful milestone is not a better video. It is a long-run deployment metric at a named customer site. For machine feeding, that might mean rings per hour across shifts, intervention rate, safety stops per thousand cycles, and maintenance events. For tote handling, it might mean successful lifts per hour across varied tote weights, orientations, and surface conditions.

The second milestone is update governance. If a deployed robot keeps learning from supervisor interventions, the buyer needs a policy-management layer as serious as the model-training layer. Industrial automation buyers will want version control, rollback, site validation, and evidence that a local improvement does not create a new safety risk elsewhere.

The third milestone is cost. Days of robot time per task may be practical for high-value stations, but the economics depend on setup labor, hardware durability, data infrastructure, and the number of stations that can reuse a trained capability. The best case is that one bottleneck-trained skill generalizes across adjacent tasks, as Humanoid's object-picking result suggests. The harder case is a world where each customer site needs careful tuning.

For now, KinetIQ Ascend is a real signal because it speaks the language industrial buyers use: throughput, cycle time, success rate, interventions, safety stops, and task transfer. The humanoid market still has too much theater. This is closer to the evidence stack that will decide who gets past pilots.

FAQ

What is KinetIQ Ascend?

KinetIQ Ascend is Humanoid's reinforcement learning approach for improving manipulation policies on real robot hardware. It extends the company's KinetIQ framework beyond imitation learning so robots can improve through trial and error.

Did Humanoid reach 99.9% reliability?

Humanoid describes 99.9% task success at human or superhuman speed as its target. In the reported tests, item handover reached 98% success and tote handling reached 98.9% success. Those results are close, but they are not the same as verified 99.9% production reliability.

Why is reinforcement learning useful for robot manipulation?

Behavior cloning teaches a robot to copy demonstrations. Reinforcement learning lets it optimize task performance from its own successes and failures, which can improve speed, reduce rare failures, and teach the policy to avoid unsafe or unproductive actions.

Is this a production deployment?

No. This is best read as a technical demonstration using production-style tasks on real bimanual hardware. Humanoid did not name a customer deployment for these specific KinetIQ Ascend results.

Sources