A robot performing the task is no longer the whole test
A humanoid robot can learn a task and perform it successfully, yet still be unready for industrial work. Demonstrations at Cambridge Consultants’ Cambridge Science Park labs in Cambridge, England, showed where that gap now sits: commercial usefulness depends on the robot working quickly, reliably and safely around people whose behavior cannot be fully predicted. Task capability remains important, but deployment readiness puts much more of the system under test.
That broader test was visible when Cambridge Consultants, the deep-tech arm of French IT consulting firm Capgemini, opened the labs to media this week. Monday’s demonstrations included two humanoid robots modified with Cambridge Consultants’ own hardware and software. Those changes showed the next engineering problem for physical AI: capable models have to operate through hardware, compute and safety systems built for the conditions where the robots will work.
Those conditions raise the standard for success because a demonstration can show that a robot has learned an action without showing that the whole system can repeat it at useful speed around people. Industrial deployment also demands reliability and safety. Cambridge Consultants’ work presents a physical-AI stack in which model intelligence, human interaction, real-time computation, learning and verification work together.
The demonstrations show why intelligence extends beyond motion
The two modified humanoids make that wider stack concrete. Cambridge Consultants fitted one robot with a new head and backpack containing extra computing resources, connectivity and a safety kill switch. It redesigned the other robot’s hands to improve how it picks up and puts down boxes for factory and warehouse uses. Better reasoning has limited practical value when the machine cannot execute the resulting action effectively or be stopped safely when required.
Those physical changes were paired with an unnamed robotic awareness platform based on foundation models from Nvidia and Physical Intelligence, a company backed by Jeff Bezos and OpenAI. Foundation models are AI models trained broadly enough to support many inputs or tasks rather than one narrowly programmed behavior. Here, they help a robot understand the people and surroundings that give an action its context.
That awareness changes how a worker can instruct the robot. Gesture recognition allows a person to point toward the object they want the robot to move, while speech allows an instruction such as “move this here.” Because neither “this” nor “here” identifies an object or destination by itself, the system has to connect those words with what the person is doing and what exists around them.
For Tim Ensor, head of AI at Cambridge Consultants, combining those signals makes the interaction useful. “Matching together speech, gesture, pose, spatial understanding and object recognition allows a much more human-like interaction between a human and robot,” Ensor told TechTarget AI and Emerging Tech. “You can refer to objects in the room through pointing and through ambiguous references.” The robot has to resolve human intent from several signals when a command leaves parts of the action implicit.
Resolving that intent is part of useful understanding because mechanical ability does not tell the robot which object a person means. A robot may have the learned policy and physical ability to move a particular box, while the instruction still depends on identifying which box, where it should go and how the person’s gesture changes the meaning of their words. Human interaction therefore turns perception and context into inputs for the robot’s decisions.
A project in mind?
Schedule a 30-minute meeting with us.
Senior experts helping you move faster across product, engineering, cloud & AI.
Factories are structured; the humans inside them are not
The need to interpret people becomes harder to avoid when robots leave controlled demonstrations. Factories offer a comparatively structured environment for early deployment, especially compared with domestic settings, because equipment and workflows can be organized around defined operations. Workers still behave in ways the robot cannot assume will follow one predefined sequence.
That uncertainty limits practical usefulness. Ensor said, “The utility of your robot is limited if it can’t interact with humans who are unpredictable.” Human interaction is therefore an operating requirement: a robot that works correctly only when people behave exactly as its demonstration expects has a narrower range of useful conditions.
The same requirement appears earlier, when a general-purpose robot is learning what work to perform. A machine expected to handle multiple jobs has to acquire novel tasks instead of remaining tied to one fixed workflow, so people need an effective way to provide instruction and training. “In the course of getting a robot to a general-purpose level, somebody’s got to show them how to do those tasks,” Ensor said.
Teaching also matters after a robot can perform a workflow without constant supervision. As Ensor put it, “Even though it may not have to interact with humans constantly during a workflow, there still has to be a way for these robots to interact effectively with humans.” The engineering problem covers two forms of interaction: people communicate intent during operation, and they help robots acquire or improve behaviors that cannot all be fixed in advance.
Those forms of interaction change how deployment readiness should be judged. A controlled test can remove ambiguity from instructions, surrounding objects and people’s behavior, while industrial systems have to handle such ambiguity when it occurs. Speech, gesture, pose and spatial understanding therefore become parts of the deployment architecture once robots have to work with people.
Learning starts a larger deployment process
The need to acquire new behaviors also makes the model technology behind adaptable robots important. Vision-language-action models, or VLA models, connect what a robot sees and understands from language with the physical movements it performs. They have become a leading approach for robots that need to adapt rather than execute a single predetermined routine.
Cambridge Consultants has been experimenting with systems that learn from video and combine what they extract with language. When a movement is described verbally, the robot can form an expectation of what that movement should look like before attempting the action itself. Video examples and human language can therefore lead to physical behavior without requiring every task to be specified as a fixed sequence of movements.
Once a robot has learned the movement, deployment introduces measurable performance demands. “You can do that for a robot, and it will be able to perform the task at a certain speed and with a certain percentage reliability,” Ensor said. Speed and reliability matter to an industrial operator because they determine whether the machine completes work quickly enough and succeeds consistently enough for its intended environment.
Those measures expose a gap in the systems Cambridge Consultants has tested. “But we tend to find that they’re not actually fast or reliable enough to go straight into an industrial deployment,” Ensor continued. VLA systems can establish adaptable task performance, while deployment engineering still has to make that performance fast and dependable enough for industrial use.
Closing that gap has led Cambridge Consultants to work on what Ensor calls “non-core functional components.” One requirement is running large models efficiently on smaller computing platforms because the robot needs enough local computing capability to use its model under real-time operating constraints. The extra compute and connectivity fitted to the demonstrated humanoid show how model demands become hardware architecture decisions when AI has to operate inside a physical machine.
Performance can also improve through experience after the initial learning stage. Cambridge Consultants is developing self-learning loops in which repeated attempts, combined with some human intervention, allow the robot to improve. Those loops connect the earlier human-training problem with deployment engineering because people can contribute to improvement while the robot accumulates experience through repeated execution.
Improved performance then has to be tested for dependable and safe use. Cambridge Consultants is developing validation and verification processes designed to demonstrate reliability and safety. Verification checks whether a system satisfies specified requirements, while validation establishes whether its behavior suits the intended use. For a physical system working around people, both processes become part of the case for allowing learned behavior into production.
Those processes sit within the broader full-stack architecture Cambridge Consultants advocates. The stack includes computing architecture capable of executing the models fast enough for real-time use, a knowledge layer that gives the robot an understanding of its surroundings, and an intelligence layer that can choose useful actions from that information. Each layer affects deployment because a correct decision can still produce unusable behavior when it arrives too slowly, relies on incomplete environmental knowledge or becomes an unsuitable physical action.
The hardware visible in Cambridge shows how those layers reach the physical machine. Added compute and connectivity support processing and communication demands, the kill switch provides a direct safety mechanism, redesigned hands improve execution of jobs such as moving boxes, and multimodal awareness helps interpret people and the environment. VLA learning, self-learning loops, validation and verification address other parts of the same deployment problem. Cambridge Consultants has a commercial stake in this framing because, as Capgemini’s deep-tech arm, it sells engineering and consulting expertise to organizations working through such technology and deployment problems.
Near-term deployments can prove practical value
The remaining engineering gap still leads Ensor to a near-term commercial forecast. He expects commercial deployments of real robots doing “useful work” “within the next year.” That expectation can coexist with today’s limitations because a useful commercial deployment can operate within particular tasks and environments while the broader general-purpose goal remains under development.
That distinction fits Ensor’s characterization of physical AI as being at an early stage. “This is an exciting area that we’re really only at the beginning of,” he said. Near-term deployments would establish that particular robots can deliver practical value under particular operating conditions, while the larger challenge remains making increasingly general systems fast, reliable and safe around unpredictable people.
The longer-term ambition extends beyond automating another fixed industrial workflow. “The real promise of physical AI is a transformation of manual labor to a similar extent that we‘re seeing digital work being transformed by agentic AI. We’re really only going to see growth towards this goal from here,” Ensor said. For learned intelligence to reach that ambition, it has to survive real-time operation, repeated physical execution, human instruction and environmental uncertainty while meeting safety requirements that can be demonstrated.
Key highlights
- Judge the full system: Industrial readiness depends on speed, reliability, safety, compute and effective hardware alongside learned task performance. Buyers evaluating humanoid robots can assess these factors together under realistic operating conditions.
- Design interaction around human behavior: Speech, gesture, pose and spatial awareness help robots interpret ambiguous instructions and work around unpredictable people. Deployment teams can test these capabilities with real workers rather than tightly scripted interactions.
- Build human interaction into deployment: Factories provide structured environments, but workers still introduce uncertainty and play a role in teaching new tasks. Operators can evaluate how easily employees instruct, correct and work alongside robots as requirements change.
- Engineer learned tasks for production: VLA models can teach adaptable behaviors, while industrial use requires further work on real-time compute, repeated learning, validation and verification. Technology teams can set measurable targets for task speed, reliability and safety before production.
- Focus near-term investment on defined uses: Cambridge Consultants expects commercially useful deployments within a year while general-purpose physical AI remains early. Buyers can prioritize bounded tasks and environments where current humanoid systems can demonstrate measurable operational value.
A project in mind?
Schedule a 30-minute meeting with us.
Senior experts helping you move faster across product, engineering, cloud & AI.


