From Demonstration to Deployment: What Figure 03 at BMW Reveals About Humanoid Robots at Work
A robot video can prove that a motion is possible. A production deployment must prove that the same work can be repeated safely, reliably and economically. BMW and Figure now provide one of the clearest cases for examining that difference.
Figure says its Figure 02 humanoid operated at BMW Group Plant Spartanburg for 1,250 hours, loaded more than 90,000 sheet-metal parts and contributed to production of more than 30,000 BMW X3 vehicles. BMW has separately confirmed the production figure and ten-hour weekday shifts. Those numbers make the project more informative than a staged demonstration, but they do not establish profitable, unsupervised or fleet-scale autonomy. Figure has not published complete data for uptime, human interventions, operating cost, safety events or the number of robots used. In June 2026, Figure 03 began a more complex sequencing workflow at the same plant, combining perception, manipulation and locomotion. The next decisive evidence will not be another polished video. It will be transparent performance over long periods: cycle time, first-attempt success, recovery, availability, maintenance and cost per completed task.
A demonstration answers the smallest question
Humanoid robotics is unusually easy to overinterpret. A short video may show a robot walking to a cart, selecting a component and placing it correctly. That proves the machine completed the recorded sequence under the recorded conditions. It does not reveal how often the task failed, how much preparation was required, whether an operator intervened, how long the machine ran before maintenance or whether the process cost less than an alternative.
A factory deployment asks a larger set of questions. Can the system repeat the task throughout a shift? Can it handle parts that arrive in slightly different positions? Does it recover safely after a poor grasp? Can technicians diagnose a failure without waiting for the robotics company? Most importantly, does the entire system—including supervision, charging, integration and maintenance—create enough value to justify its cost?
The Figure–BMW project matters because it has moved beyond the smallest question. It offers operating hours, production context and task-level metrics. The evidence remains incomplete, but it gives the humanoid industry a more serious benchmark than a laboratory performance.
What Figure 02 reportedly did at BMW
Figure describes an eleven-month deployment process at BMW Group Plant Spartanburg in South Carolina. The first application was sheet-metal loading. Parts were removed from racks or bins and placed into a welding fixture, after which conventional industrial robots performed the welding and continued the production flow.
According to Figure, the robot eventually operated ten-hour shifts from Monday to Friday, loaded more than 90,000 parts and accumulated more than 1,250 operating hours. The company says the work contributed to production of more than 30,000 BMW X3 vehicles. BMW has also reported the 30,000-vehicle figure and the ten-hour weekday schedule, providing customer-side confirmation of the central production claim.
This was not general factory labour. It was one constrained application inside a highly engineered production environment. That distinction does not make the result unimportant. Industrial automation succeeds by solving defined work reliably. It does mean that the project should not be described as evidence that a humanoid can move freely between unrelated jobs without new integration, training or validation.
“A useful deployment is not a robot doing something once. It is a complete operating system for repeating the work, detecting failure and returning safely to production.”
NV · NTS Editorial
The published metrics are meaningful—but easy to misuse
Figure identified three important key performance indicators for the Figure 02 application: cycle time, placement accuracy and human interventions. The stated production requirement was an 84-second total cycle, including a 37-second loading phase. The target for placing all three sheet-metal pieces correctly was above 99 percent per shift, and the intervention target was zero per shift.
These are the right categories because they translate a robotics demonstration into an operating process. Cycle time determines whether automation keeps pace with the line. Placement accuracy connects perception and manipulation to quality. Interventions reveal how often a person must rescue the system.
However, Figure’s public account describes these values as requirements or goals; it does not publish a complete time series demonstrating that every target was consistently achieved. The distinction is essential. A target of zero interventions is not the same as independently audited evidence of zero interventions. The 90,000-parts figure is substantial, but without a denominator for attempted cycles it cannot establish the full success rate.
Operating hours are not the same as availability
Figure reports more than 1,250 hours of runtime. Runtime shows that the robot accumulated real experience, exposed components to wear and generated data outside a laboratory. Yet manufacturers usually need a wider measure: how much of the scheduled production period was the system ready and able to work?
Availability includes unplanned stops, preventive maintenance, calibration, software updates, battery management and time waiting for technical support. A machine can collect many operating hours while still requiring significant downtime between them. Public disclosures do not provide enough information to calculate availability for Figure 02.
The same limitation applies to fleet size. Figure and BMW have not clearly disclosed in these updates how many units produced the reported totals. The result therefore demonstrates a real deployment, not a normalized per-robot productivity rate. Future reporting would be more useful if it included robot count, scheduled hours, productive hours and reasons for every significant stop.
The most valuable result may be the failure data
A long deployment does more than produce parts. It reveals which components fail first. Figure says the Figure 02 forearm became its leading hardware failure point at BMW. The compact assembly had to combine dexterity, electronics, communication and thermal management. For Figure 03, the company redesigned the wrist electronics to remove a distribution board and dynamic cabling, reducing complexity.
This is the less glamorous path from prototype to product: identify the subsystem that interrupts work, simplify it and test again. A robot designed for a video can contain delicate components and labour-intensive calibration. A robot designed for production must tolerate repetitive movement, temperature changes, vibration and inevitable contact with its surroundings.
Figure’s disclosure is useful because it links field experience to a specific hardware redesign. It does not yet show the resulting failure rate of Figure 03. The next generation must accumulate its own operating history before improved reliability can be treated as demonstrated rather than designed.
Figure 03 is attempting a harder class of work
In June 2026, Figure announced that Figure 03 had arrived at BMW’s Spartanburg plant for a sequencing application. Sequencing means selecting and arranging the correct parts for later assembly. Unlike a fixed stream of identical components, parts may arrive rotated, shifted, partly hidden or imperfectly organized.
The published demonstration combines precise handling of thin-walled components with whole-body movement, including stepping and repositioning while pulling a wheeled cart. Figure says its Helix 02 vision-language-action system coordinates the hands, arms, torso and feet using visual input and high-frequency motor control.
This is more demanding than repeating a fixed pick-and-place path because manipulation and balance influence each other. The robot must adjust its body to maintain reach, select the right object and correct small spatial errors. It is a credible test for learned robotics. It is still, as of this verification date, an early demonstration of the new workflow rather than a published multi-month production record comparable to the Figure 02 data.
Why a humanoid form might be justified
Factories already contain excellent robots. Six-axis arms can be fast, precise and reliable. Conveyors, automated guided vehicles and specialized grippers often outperform a humanoid when the object, path and workstation remain fixed. The business case for two legs and human-like arms cannot simply be that the machine looks versatile.
The strongest argument is environmental compatibility. A humanoid can potentially use aisles, carts, tools and workstations designed for people. BMW’s sequencing case involves variation and movement between positions that may be awkward for a fixed arm. If the same platform can later address several such processes, the cost of integration may be distributed across more work.
That advantage must be measured, not assumed. A mobile base with an arm may provide equal flexibility with greater stability. A small change to a rack may enable a simpler automated system. The correct comparison is not humanoid versus human in isolation. It is humanoid versus every practical combination of conventional automation, workstation redesign and human work.
The software is not general intelligence
Figure describes Helix as a generalist vision-language-action system. The term indicates a model that connects visual observations and language-level objectives to robot actions. It can reduce dependence on manually scripted trajectories and help a robot adapt to variations that traditional automation would need to constrain.
Generalist does not mean universal. A model may generalize within a family of objects and conditions while failing outside them. Industrial deployment therefore needs a controlled boundary around learned behavior: permitted work zones, safe states, validated payloads, speed limits and escalation rules when confidence is low.
The important engineering achievement is not unrestricted autonomy. It is useful autonomy inside a process that can be monitored and stopped. The more learned control replaces fixed programming, the more operators need logging and evaluation tools that explain when performance changes after a software or model update.
Safety is a system property
A humanoid working near people combines mass, moving joints, pinch points, batteries and software-driven behavior. Safety cannot depend on the model making the right decision every time. It requires layered protection: mechanical limits, force control, emergency stops, monitored zones, safe speeds, tested fallback behavior and clear responsibility for restarting the process.
Public material from Figure and BMW emphasizes collaboration with workers and deployment in human-designed environments. It does not provide a complete safety case, incident history or certification record for the Spartanburg application. That information may be confidential, but its absence limits what outside observers can conclude.
Workforce claims also require care. Automating a repetitive task can reduce ergonomic strain, while changing the number and type of jobs around it. A responsible analysis should track both effects: which tasks disappear, which technical roles are created, how workers are trained and whether productivity gains are shared.
Economics remains the largest missing layer
Neither operating hours nor production volume establishes a positive return on investment. A complete calculation would include the robot, integration engineering, computing, maintenance, replacement parts, supervision, charging infrastructure, safety changes and the cost of downtime. It would then compare those expenses with labour savings, increased throughput, lower injury risk, improved quality and the flexibility to reuse the system.
Figure has discussed designing Figure 03 for high-volume manufacturing and lower component cost. Its BotQ facility is intended to support production at a scale far beyond prototype assembly. Manufacturing capacity, however, is not the same as customer demand or profitable operation. Lower unit cost helps only if reliability and useful task coverage rise at the same time.
The factory is therefore the correct place to test humanoids, but not because factories guarantee success. They expose the machine to measurable economics. A deployment either meets the line’s requirements often enough and cheaply enough, or it does not.
The broader market is moving in the same direction
Figure is not alone in choosing automotive manufacturing as an early proving ground. Apptronik and Mercedes-Benz announced a pilot for the Apollo humanoid, including logistics tasks such as delivering assembly kits and inspecting components. Hyundai Motor Group and Boston Dynamics have positioned the electric Atlas for industrial applications, with deployments committed to Hyundai facilities and suppliers.
The convergence is logical. Automotive plants contain repetitive material movement, detailed process data, engineering support and high-value production lines. They also contain enough variation to test whether a humanoid provides an advantage over fixed automation.
These projects are at different stages, and announced plans should not be counted as completed deployments. Comparing them requires common measures: productive hours per robot, first-attempt task success, interventions, recovery time, maintenance hours, safety events and total cost per successful cycle.
What independent proof would look like
The strongest next step would be customer-reported data using clearly defined periods and denominators. BMW or another manufacturer could publish the number of robots, scheduled hours, productive hours, attempted cycles, completed cycles, interventions and safety-related stops. The robotics supplier could publish hardware replacements, software rollbacks and recovery performance without revealing trade secrets.
Independent assessment does not require access to proprietary model weights. It requires enough operational evidence to distinguish an isolated success from a repeatable system. Audited production totals, standardized definitions and long-duration testing would allow buyers to compare products rather than presentations.
Until that evidence exists, the fairest description of Figure at BMW is neither “just a demo” nor “humanoid labour at scale.” Figure 02 achieved a meaningful production deployment in a narrow task. Figure 03 is now demonstrating a more complex workflow whose sustained performance has not yet been publicly established.
An NTS checklist for humanoid deployment claims
When a company announces that a humanoid is “working” in a factory, ask seven questions. First, what exact task is completed from start to finish? Second, how many robots are involved? Third, how many scheduled and productive hours are recorded? Fourth, what percentage of attempts succeeds without help? Fifth, how are errors detected and recovered? Sixth, what maintenance and supervision are required? Seventh, what is the cost per completed unit of work compared with the best alternative?
Video quality is not part of this checklist. Neither are funding rounds, valuation or projected factory capacity. Those facts may matter to the company’s ability to continue developing the system, but they do not measure the robot’s performance.
A credible deployment report makes failure visible. It defines intervention, separates autonomous time from teleoperation and explains whether the task or environment was modified for the robot. The industry will mature when this information becomes routine.
The NTS View
The Figure–BMW project is one of the strongest public signals that humanoid robotics is entering a serious industrial evaluation phase. Figure 02 accumulated enough work to produce useful reliability lessons, and the shift to Figure 03 targets a workflow where perception, manipulation and locomotion must operate together. These are real advances.
The case does not prove that general-purpose humanoids are economically ready for broad factory adoption. Missing data on availability, interventions, fleet size, maintenance and cost prevents that conclusion. Figure 03’s sequencing work also needs a long operating record before it can be evaluated like the earlier deployment.
The important change is methodological. Humanoids are beginning to generate the kind of evidence that can eventually replace speculation: hours, cycles, parts, failures and redesigns. The winner of the humanoid race will not be the company with the most human-looking demonstration. It will be the one that can publish repeatable work, transparent limits and a credible cost for every task completed.