Imagine a robot that looks a little bit like a person: two arms, two legs, and a head with cameras instead of eyes. This is a humanoid robot.
Humanoid robots are being developed for many different environments, like warehouses, factories, hospitals and offices. But making a robot that can walk and perform a task is only one part of the challenge. Before it can be trusted to work around real people in busy environments, it has to show that it can handle situations that don’t go as planned. This is where testing becomes especially important.
From a QA perspective, testing a humanoid robot is similar to testing other complex systems. Controlled scenarios are tested first, followed by edge cases and more complicated environments. The main difference is that a problem is not always limited to software. Issues with perception, navigation or decision-making can cause the robot to stop, fall, damage equipment, or hurt someone.
The hardest part starts when a robot leaves the controlled environment of a lab and has to deal with the real world.
TL;DR
30-second summary
Why isn't lab testing enough for humanoid robots, and what does testing them in real-world environments actually involve?
- A robot that passes every lab test can still fail in the real world. A lab controls the variables that matter most — flat floors, stable lighting, a fixed number of people — which makes it easy to isolate and test one function at a time. A real warehouse is dusty, cluttered, and unpredictable, and Agility Robotics' Digit has logged over 65,000 hours across nine customer sites specifically because that scale of real-world data surfaces defects a lab never reproduces.
- The right testing approach depends heavily on where the robot will actually work. Warehouses demand consistency across thousands of repetitions, as shown by Digit moving over 100,000 objects at a single GXO facility. Factories demand long-term stability, demonstrated by Figure 02's eleven-month deployment at BMW helping produce over 30,000 cars with 99% accuracy. Hospitals and offices each introduce their own safety and navigation challenges no single test scenario can cover.
- Environmental variables like lighting, noise, and floor conditions directly affect how a robot perceives and behaves. A camera that recognizes an object under one lighting setup can struggle with the same object under different shadows. Background warehouse noise can affect voice command recognition. And uneven or slippery floors are a genuine balance risk for two-legged robots like Digit and Unitree's H1, which is why testing needs to deliberately reproduce these conditions rather than assume ideal ones.
- Human unpredictability is the hardest variable to test for, which is why supervision remains standard. People can change direction, drop objects, or walk too fast in ways no test case can fully anticipate. Testing with real employees performing their regular work, rather than in an empty facility, surfaces issues that a fully controlled environment never reveals — and it's part of why current deployments favor a small number of supervised robots over large, independent fleets.
- Gradual, task-specific deployment consistently outperforms broad, ambitious rollouts. Giving a robot one specific task, like moving objects in a warehouse or delivering hospital supplies, is easier to test properly than assigning multiple tasks at once. Figure's eleven-month, single-task pilot at BMW gave engineers enough time to demonstrate consistent performance rather than a one-off demonstration.
Bottom line: It's not enough to test that a humanoid robot can do something. It's just as necessary to test what happens when it can't — lost balance, sensor failure, a sudden lighting change, an unpredictable human. As more humanoid robots move from labs into busy real-world environments, this kind of testing is what determines whether they're not just functional, but safe and reliable for the people working around them.
Lab vs. real-world testing
A lab is the easiest place to start testing a robot because most of the variables can be controlled. The floor is usually flat and clean. Lighting can remain stable. Test objects can be placed in specific locations. The number of people around the robot can also be controlled. This makes it possible to test one function at a time, such as walking in a straight line, picking up an object, or moving from one point to another.
If something fails, the test can be stopped. The data is collected and the scenario is repeated once the problem is fixed.
The real world is far less predictable. A warehouse floor can be dusty, cluttered with small objects, uneven, or wet. Lighting can change during the day. People can walk into the robot’s path without warning, and objects may not be where the robot expects them to be. A robot can pass a test in the lab, but still fail it in a real warehouse.
That is why lab testing is not enough. From a QA point of view, passing controlled tests is more like the beginning of the next testing stage rather than the end of testing.
Agility Robotics’ Digit is a good example. The robot was trained and tested in controlled environments before being deployed at real Amazon and GXO warehouse sites. According to the company, Digit robots have now been used at nine different customer sites and have accumulated more than 65,000 hours of real working time.
This amount of real-world data is extremely useful. It allows engineers to find defects that can be difficult to reproduce in a lab. The robot can behave differently on a dusty floor, near many people that are working around it at the same time. These are exactly the kinds of conditions that can reveal problems before a robot is deployed on a large scale.
Warehouses, factories, hospitals, and offices

The testing approach also depends heavily on where a robot is going to work.
Warehouses
Warehouses are usually full of shelves, boxes, forklifts, and people. There can be a lot happening at once.
Digit has been tested at the GXO facility in Flowery Branch, Georgia, where it has moved over 100,000 objects. A task like this may look simple, but repeating it thousands of times generates a lot of useful information. It allows engineers to check whether a robot performs consistently, how often it fails, and how its behavior changes in a different environment.
Factories
Factories create another type of challenge. There may be heavy objects, moving equipment, vehicles, and other machines operating nearby. A robot also has to maintain a predictable speed and a safe distance from people and equipment.
Figure AI’s Figure 02 spent eleven months working at BMW’s plant in Spartanburg, South Carolina. During this period, it helped with the production of more than 30,000 cars and loaded around 90,000 parts with 99% accuracy.
For QA, a deployment lasting eleven months provides much more information than a short demonstration. It gives engineers enough time to see whether the system remains stable over a long period of time. It can also reveal issues that only appear after repeated use.
Hospitals
In hospitals, a robot has to work around patients, nurses, doctors, beds, medical equipment, and visitors. Safety is a major concern.
Moxi, developed by Diligent Robotics, is used in Medical City Dallas and Medical City Denton hospitals. It delivers lab samples and supplies and collects dirty linens. These are tasks that take nurses away from direct patient care. Moving such tasks to a robot can give staff more time to focus on patients. Clinical trial reports have suggested that these types of tasks can take close to 30% of a nurse’s working day.
Offices
Offices may look like an easier environment, but they still have many things that can affect a robot. Glass doors can be difficult to detect, hallways can be narrow, and chairs, bags, or other objects may be left somewhere unexpected. Someone can also walk around a corner without seeing the robot.
Each environment creates different test cases. There is no single test scenario that can prove a humanoid robot is ready for every possible environment.
Environmental variables: lighting, noise, and slippery floors
One of the things that makes humanoid robot testing different from normal software testing is the number of physical variables involved.
Robots use cameras and other sensors to understand their surroundings. Because of this, changes in the environment can directly affect their behavior.
Lighting
A hallway can have completely different lighting in the morning and in the evening. Shadows can also change depending on where the light is coming from. A camera may recognize an object correctly in one setting but struggle with the same object in another.
This is why lighting should be part of the test plan. QA teams can test the robot under different brightness levels, with shadows, and sudden changes in lighting. The goal is to check whether a robot can still navigate and identify objects correctly.
Noise
Noise can cause another set of problems. Warehouses and factories are usually loud, with forklifts, alarms, machines, conversations, and other sounds happening simultaneously. If the robot accepts voice commands, background noise can affect how well it understands them.
One way to test this is to reproduce real workplace noise during testing. For example, recordings of a warehouse environment can be played while voice commands are given to the robot. This provides a better indication of how the system may behave in an actual workplace.
Floors
For a human, a slightly wet or uneven floor may not be a major problem. For a two-legged robot, it can be. Robots such as Digit and Unitree’s H1 need to maintain their balance while walking on different surfaces.
We described testing humanoid robots with slippery surfaces, sudden lighting changes, and crowded areas in one of our previous blogs. These conditions are intentionally difficult because they can happen in real life.
A fall can damage the robot, but that is not the only concern. If someone is standing close to it, a robot could also injure a person. This is why safety standards for humanoid robots are becoming more important. ASTM International is working on testing approaches related to balance, falls, and recovery.
Older industrial robot standards were mostly created for robots that stay in a fixed position, such as robotic arms. A humanoid robot is different because it can move through the same space as people.
Human unpredictability

One of the hardest factors to test during humanoid robot testing is the human element, i.e. people. Test cases can be created for a person walking in front of the robot or for someone standing too close to it. But it is impossible to create test cases for every possible thing a person might do.
People can unexpectedly change direction, drop an object, stop in the middle of the robot’s path, or walk too fast. That is why human behavior is a difficult variable. That is also a reason why humanoid robots are usually deployed with human supervision. A report by TechCrunch about Agility Robotics described the current approach as a small number of robots performing a specific task under supervision, rather than hundreds of robots working independently for an entire shift.
From a testing perspective, this is reasonable. When a system can physically interact with people, it is important to understand how it behaves when something unexpected happens. Having a person nearby also means that a test can be stopped immediately if necessary. This is why testing with real employees is so important. Instead of testing in an empty warehouse or factory, QA teams can observe the robot while people are doing their regular work. What happens if someone crosses its path? Or what if two people approach it from different directions? What if someone stands too close to the robot while it is carrying a heavy object?
Such situations can reveal issues that never appear in a completely controlled test environment. Figure’s newer robots at BMW were designed to work alongside employees, instead of being separated in a safety zone by a fence. This makes the robot more useful, but it also creates additional testing requirements. The robot has to detect people, understand their position, how far they are, and react safely when they change direction or move unexpectedly.
In this case, homes and public spaces are even more complicated. -Children and pets do not follow predictable paths. A child could suddenly run toward the robot, or a dog could move directly into its walking path. These situations are much harder to predict than anything seen in a controlled workplace.
This is one more reason why many companies start with structured work environments before moving to more unpredictable public or home environments.
Testing a robot alongside real people takes more than a checklist
We've tested humanoid robots in real-world conditions, including how they behave around unpredictable human movement. Let's talk about yours.
Lessons learned
After looking at different environments and test conditions, a few things stand out.
First of all, real-world testing provides information that lab testing cannot. A robot completing a five-minute demonstration is not the same as a robot working for thousands of hours. A long deployment gives QA and engineering teams much more information about reliability, failures, environmental conditions, and unexpected behavior.
Another important point is that gradual deployment is usually a better approach. Instead of giving a robot multiple different tasks at once, companies can start with one specific task. Moving specific types of objects in a warehouse or delivering supplies in a hospital may not seem impressive, but these tasks are easier to test properly. Figure’s eleven-month pilot at BMW is a good example. The robot had a relatively specific job and enough time to demonstrate whether it could perform that task consistently.
Safety standards are still developing. Organizations such as ASTM and IEEE are working on new ways to test humanoid robots for balance, falls, recovery, and interactions with humans. Humanoid robots are still a relatively new area, so the industry is also learning how these systems should be tested.
For QA teams, this means that testing cannot be limited to checking whether the robot completed its main task. It is just as important to ask what happens when something goes wrong - lost balance, sensor failure, lighting change, interactions with humans, object detection.
In the end, a lab can teach a robot how to walk, pick up objects, or follow a route. But the real world shows whether it can continue doing those things when conditions are not perfect. And that is probably the biggest lesson from real-world humanoid robot testing: it is not enough to test that a robot can do something. It is also necessary to test what happens when it cannot.
As more humanoid robots move from labs into real, busy environments, this type of testing will be essential to ensure the robots are not only functional, but also safe and reliable for the people working around them.

Co-financed by the European Union. Project "Next Generation Micro-factories". Contract No. 5.1.1.2.i.0/4/24/A/CFLA/003. Funded by the European Union and the European Union – NextGenerationEU. However, the views and opinions expressed are solely those of the author(s) and do not necessarily reflect the views and opinions of the European Union or the European Commission. Neither the European Union nor the European Commission bears responsibility for them.
FAQ
Most common questions
Why isn't lab testing enough for humanoid robots?
A lab controls the variables that matter most for testing, flat floors, stable lighting, and a fixed number of people, which makes it possible to isolate and test one function at a time, like walking in a straight line or picking up an object. A real warehouse or factory floor can be dusty, cluttered, uneven, or wet, with lighting that changes and people who walk into the robot's path without warning. Agility Robotics' Digit has accumulated more than 65,000 hours of real working time across nine customer sites specifically because that volume of real-world exposure surfaces defects that a controlled lab environment simply doesn't reproduce.
How does the testing approach for humanoid robots differ across warehouses, factories, hospitals, and offices?
Each environment introduces distinct risks that shape the test plan. Warehouses are tested for consistency across high-volume repetition, as shown by Digit moving over 100,000 objects at a single GXO facility. Factories require long-duration stability testing, demonstrated by Figure 02's eleven-month deployment at BMW. Hospitals demand careful navigation around patients, staff, and medical equipment, as seen with Diligent Robotics' Moxi delivering supplies at Medical City hospitals. Offices introduce their own hazards, like glass doors and unpredictable pedestrian movement around corners. No single test scenario can prove a humanoid robot is ready for every environment.
How do lighting, noise, and floor conditions affect humanoid robot testing?
Robots rely on cameras and sensors to understand their surroundings, so changes in physical conditions directly affect their behavior. Different lighting and shadows can cause a camera to recognize an object correctly in one setting and struggle with the same object in another, which is why lighting variation should be part of any test plan. Background noise, like a warehouse full of forklifts and alarms, can affect how well a robot understands voice commands and is often tested using recordings of real workplace noise. Uneven or slippery floors pose a genuine balance risk for two-legged robots, which is why organizations like ASTM International are developing standards specifically for balance, falls, and recovery.
Why is human unpredictability one of the hardest factors to test for in humanoid robotics?
Test cases can be built for a person walking in front of a robot or standing close to it, but it's impossible to create test cases for every possible thing a person might do, changing direction suddenly, dropping an object, or walking too fast. This is one reason humanoid robots are typically deployed with human supervision rather than operating independently, as described in reporting on Agility Robotics' approach of using a small number of supervised robots for specific tasks. Testing with real employees performing their regular work, rather than in an empty facility, is important because it reveals issues that never surface in a fully controlled test environment.
Why do companies deploy humanoid robots gradually rather than assigning multiple tasks at once?
Giving a robot one specific, well-defined task is significantly easier to test properly than assigning several tasks simultaneously. Figure's eleven-month deployment at BMW is a clear example — the robot had a relatively specific job and enough time to demonstrate whether it could perform that task consistently across thousands of repetitions. This kind of long-duration, narrow-scope deployment gives QA and engineering teams far more information about reliability, failure patterns, and unexpected behavior than a short demonstration ever could, which is why most companies start with structured environments before expanding to more unpredictable public or home settings.
Functional isn't the same as field-ready
From environmental variability to human unpredictability, we help robotics teams build testing programs that go beyond the lab and hold up in the real-world conditions robots will actually face.





