Artificial Physical Intelligence

KAIST Exoskeleton Lab

Artificial Physical Intelligence (API)

Physical AI That Walks With a Person Inside

We build the intelligence of a robot that a person wears. Our team at KAIST EXO LAB develops WalkON Suit F1, a wearable robot with 12 actuated joints for people with complete paraplegia. F1 walks to its user unaided and is put on from the front without leaving the wheelchair. We are rebuilding its walking control with reinforcement learning (RL) in a simulation that we check against the actual hardware.

RL in large-scale simulation has taught legged and humanoid robots to walk. A robot with a person inside is harder and more valuable. Its dynamics include the wearer, a fall is not an option, and it measures the physical interaction first-hand. Our work rests on four pillars: learning to walk, state estimation, balance and reading the wearer's intention.

Concept illustration of a person walking in a lower-limb wearable robot, above photos of WalkON Suit F1 walking to its user by itself, walking with a person inside, and the KAIST team winning gold at Cybathlon 2024.

Learning by Wearing, Not Only by Watching

Watching a person walk is not the same as walking with one. Physical AI acts through a robot body. Much of it learns from videos and motion capture, which miss the force and intent behind motion.

A wearable robot, strapped to its wearer, measures posture, joint torque and ground reaction force (GRF) on one clock. That is a necessity: to walk safely, it must treat the human as part of its dynamics. It is also an opportunity: these measurements test our simulation against the actual robot. We read the lab's vision this way: Understanding Humans, Creating Real Technology.

As-is: learning by watching alone. Video and motion capture from a third-person view see motion without force, and intent is guessed afterwards.
As-is: Learning by watching alone. Sees the motion, not the forces behind it.
To-be: learning by wearing. Sensors on the body record posture, torque and ground reaction force on a shared clock, with a person inside every step.
To-be: Learning by wearing. Measures motion, force and response together.

Our Research

  • Research 1

    Learning to Walk: Planner-Conditioned Reinforcement Learning

    A pendulum model proposes each step; a learned policy executes it. Each wearer changes the dynamics, so our controller has two layers, like human motor control. An intention layer plans for an ideal body; a 1 kHz physical action layer makes the real body comply.

    A 3D linear inverted pendulum (LIP) planner proposes where to step. Conditioned on that step, the policy outputs residual joint targets around a nominal posture.

    A model-normalization torque makes every wearer look like one nominal human-robot model to the policy. For now, the wearer is an unknown payload, estimated from foot forces. In simulation, the suit model walks under random pushes of 1,000 to 2,000 N. With normalization, it walks with an unknown 50 kg payload, estimated within 1 mm and 0.01 kg.

  • Research 2

    Knowing Where the Body Is: State Estimation

    We feed the policy an estimated state, not raw sensor signals. Our estimator is a contact-aided invariant extended Kalman filter (InEKF; Hartley et al., IJRR 2020) on SE2+2N(3), covering base pose, velocity and N contact points on each foot. Its error dynamics are state-independent, so observability can be read from the sensor setup.

    IMU and leg kinematics leave global position and yaw unobservable. Visual odometry adds a velocity that survives foot slip; line and surface contacts constrain foot rotation.

    With a hand-held RGB-D camera in slow motion, visual odometry matched motion capture within 1.3 cm and 0.4° RMSE. Our ROS 2 InEKF package was tested offline on a public biped dataset. Related estimators from our group appear in Youn et al. (IFAC-PapersOnLine 2025) and Park et al. (IJCAS 2025).

  • Research 3

    Balance Begins at the Feet

    Balance starts with knowing where the weight rests under each foot. Each foot carries three 3-axis force sensors. From them we compute vertical ground reaction force and the center of pressure (CoP). On the actual hardware, the CoP matched a laboratory force plate within 3.7 mm fore-aft and 1.6 mm lateral (mean RMSE, 12 weight-shift trials). Both lie inside the 15 mm and 5 mm tolerance derived from the sensors' rated error.

    Balance is built up as a ladder of four capabilities, from push recovery while standing to walking at a changing commanded speed. During walking, balance comes mainly from foot placement, the idea behind the capture point (Pratt et al., 2006). The planner chooses where to step, and the policy realizes it.

  • Research 4

    Reading the Wearer: Intention and Human-in-the-Loop Learning

    Today the wearer selects motions with a hand controller; we want the suit to read intention instead. We plan to measure the wearer with inertial, muscle-activity and interaction-force sensors and a camera. Their signals mix three things: intention, active motion and passive motion caused by the robot. Only intention should reach the controller, as its high-level command.

    Our group has long kept the human in the learning loop. Park, Choi and Kong (IEEE T-RO 2022) adjusted gait patterns by iteratively learning the wearer's behavior. Kim et al. (Mechatronics 2025) optimized crutch-free walking with the wearer's adaptation in the loop. Park et al. (IEEE/ASME T-Mech 2026) introduced embodied teaching through variance-decomposed impedance learning. Our lab director co-authored a Nature Perspective on human-in-the-loop optimization (Slade et al., 2024).

  • Engineering

    Sim-to-Real as an Engineering Discipline

    A learned policy transfers only as far as its simulation matches the hardware. We measure friction, backlash and time delay on the actual hardware first. During gait, the ankle's linear actuators run in the boundary and mixed lubrication regime of the Stribeck curve. Their friction model is therefore load- and direction-dependent.

    Time delay appeared as a phase lag against simulation. We trace it through the communication chain and cut it where it arises. We judge the chain by jitter and deadline misses, not mean delay.

    A policy first runs in shadow mode: computers, buses and actuators on a bench, driving a simulated body. On the robot, a risk manager can fall back to high-impedance position control. Our group also published an ankle backdrivability compensator (An et al., AIM 2026).

Where We Are Heading

We work toward a robot that a person can wear and trust. Such a robot would keep its balance by itself, follow what its wearer means to do and adapt to each body. We begin with ground walking learned by reinforcement learning. Next, a nominal human model with state and intention replaces the payload in our training simulation. Then come motions beyond walking, imitated from recorded human movement.

To us, Artificial Physical Intelligence means intelligence that is learned in simulation, has to prove itself on real hardware and is built around a person. If you want to work where learning meets real hardware and a real person, join us.

Roadmap from ground walking learned by reinforcement learning (now), to control that knows its wearer (next), to a wearable humanoid for everyday life (vision). Topics you will work on: reinforcement learning, sim-to-real, state estimation, legged balance and dynamics, real-time control systems, and human modeling and intention.
How to apply →