On August 31, 2026, Ricoh announced the development of Physical AI technology aimed at teaching robots actions associated with human work through image data, rather than linking the model directly to a specific robot architecture or control mechanism. The company said it had verified the technology’s effectiveness within a simulation environment constructed on the basis of physical conditions, in a step intended to transfer training from simulation to real robots.
The move is part of Ricoh’s efforts to expand its artificial intelligence business, alongside its participation in the AWS Japan program for accelerating the development of Physical AI, which provides computing resources and cloud services to support model training and testing.
Learning Actions Instead of Memorizing Robot Commands
The technology is based on the Latent Action Model (LAM). The model learns from changes in images or videos recorded during an operation, inferring actions such as grasping an object, moving it, pushing it away, or moving a target element. According to Ricoh’s description, LAM focuses on learning the actual behavior occurring in a scene, even when all movement details or commands are not explicitly described in the observation data.
The company says this approach can reduce the dependence of learning on the specifications of a particular robot, because the model learns a general representation of the action rather than relying solely on instructions tied to specific joints or control mechanisms. Human and robot action data can also be used to train the model, after which its outputs can support the development of Vision Language Action models that combine vision, language, and action.
What Changes in Practice?
The problem Ricoh is targeting is the cost of collecting training data. Differences in the architecture of each robot and its control mechanism usually require data to be collected separately for each device, which becomes more complicated for humanoid-like robots. The company believes that separating the representation of an action from the robot’s architecture could enable greater reuse of data and support the adaptation of models to multiple tasks and robots.
This does not mean that the technology has eliminated the need for robot-specific data. The material explains that Ricoh uses simulation environments for learning and then relies on Sim2Real techniques to transfer the results to real robots, while continuing to verify the suitability of the solutions for different robots and working environments.
AWS’s Role in Training and Simulation
Ricoh uses AWS cloud services, including Amazon EC2 and Amazon S3, to build a scalable learning environment, store data, and run model training. The company says that combining simulation with cloud services helps reduce the cost of collecting physical data and enables training on large datasets to be performed more efficiently.
The initiative falls within Ricoh’s announced vision for transformation through artificial intelligence, which includes layers beginning with AI documentation, generative AI, and AI agents, extending to Physical AI agents capable of interacting with the physical world. According to the material, the company is working on potential applications in offices, factories, logistics, and care.
Why Does This News Matter?
The significance of the development lies in its attempt to address a practical obstacle in AI-based robotics: how to reduce the need to recollect data and retrain the model for every robot or task. LAM offers a path to learning actions from more general visual data, but it does not by itself establish the technology’s readiness for large-scale production.
The announced results remain tied to simulation environments and verification tests, while Ricoh indicates that it has begun proof-of-concept applications for humanoid-like robots in factories and is working to progress gradually to more advanced verification stages. Therefore, open questions remain regarding the success of transferring results from simulation to reality, the amount of data required, and the stability of performance when robots, tasks, and operating environments differ.