Kinematics-Aware Diffusion Policy
with Consistent 3D Observation and Action Space
for Whole-Arm Robotic Manipulation

Anonymous Author(s)

The proposed KADP uses a set of 3D nodes on the arm body as both robot state and action representation for whole-arm manipulation, which is consistent with the 3D point cloud observation space and task space. Compared with using end-effector poses or joint angles, our method achieves higher spatial generalizability and sample efficiency while ensuring kinematic feasibility.

Abstract

Whole-body control of robotic manipulators with awareness of full-arm kinematics is crucial for many manipulation scenarios involving body collision avoidance or body-object interactions, which makes it insufficient to consider only the end-effector poses in policy learning. The typical approach for whole-arm manipulation is to learn actions in the robot's joint space. However, the unalignment between the joint space and actual task space (i.e., 3D space) increases the complexity of policy learning, as generalization in task space requires the policy to intrinsically understand the non-linear arm kinematics, which is difficult to learn from limited demonstrations. To address this issue, this letter proposes a kinematics-aware imitation learning framework with consistent task, observation, and action spaces, all represented in the same 3D space. Specifically, we represent both robot states and actions using a set of 3D points on the arm body, naturally aligned with the 3D point cloud observations. This spatially consistent representation improves the policy's sample efficiency and spatial generalizability while enabling full-body control. Built upon the diffusion policy, we further incorporate kinematics priors into the diffusion processes to guarantee the kinematic feasibility of output actions. The joint angle commands are finally calculated through an optimization-based whole-body inverse kinematics solver for execution. Simulation and real-world experimental results demonstrate higher success rates and stronger spatial generalizability of our approach compared to existing methods in body-aware manipulation learning.

Video

Method Overview

Overview of Kinematics-Aware Diffusion Policy (KADP). Taking the encoded 3D visual representations, the 3D robot nodes and time embeddings as input, diffusion model predicts the denoised 3D node trajectory iteratively. For execution, the joint angle commands are computed through an optimization-based whole-body inverse kinematics solver.

Real-World Task1: Pick up Cube

The robot only needs to grasp the object and lift it, which is designed to specifically analyze the spatial generalizability and sample efficiency of KADP. Since accurately controlling the end-effector pose is sufficient, DP3-EE is expected to perform well due to the alignment between observation and action space.

DP3-EE

DP3-Joint

KADP (Ours)

DP3-EE

DP3-Joint

KADP (Ours)

Real-World Task2: Open Door

The robot should first grasp the handle and then follow a circular trajectory to open the door. Due to the narrow width of the handle, even small grasping positional error from the handle’s center will cause the gripper to lose contact with it in the subsequent motion. Controlling only the end-effector pose is also sufficient but this task is obviously more challenging compared to the pick up cube above.

DP3-EE

DP3-Joint

KADP (Ours)

DP3-EE

DP3-Joint

KADP (Ours)

Real-World Task3: Put Cube in Cabinet

The robot should first grasp a cube and then put it into a deep and narrow cabinet, which is designed to evaluate whole-body collision avoidance performance. The primary difficulties arise from two factors: 1) The robot must reach near the cabinet's deepest position, which requires the entire robot to remain nearly horizontal to avoid collision with the top surface; 2) The cabinet is only 4cm wider than the gripper, making the successful insertion highly sensitive to even slight positional inaccuracies.

DP3-EE

DP3-Joint

KADP (Ours)

DP3-EE

DP3-Joint

KADP (Ours)

Real-World Task4: Push Button Elbow

The robot is required to press a button using its elbow instead of the gripper, making it meaningless to control only the end-effector pose. Learning directly from joint space is expected to yield good performance as only the angles of the first 3 joints change during the manipulation process.

Kinematics-Aware Diffusion Policy with Consistent 3D Observation and Action Space for Whole-Arm Robotic Manipulation

Abstract

Video

Method Overview

Real-World Task1: Pick up Cube

DP3-EE

DP3-Joint

KADP (Ours)

DP3-EE

DP3-Joint

KADP (Ours)

Real-World Task2: Open Door

DP3-EE

DP3-Joint

KADP (Ours)

DP3-EE

DP3-Joint

KADP (Ours)

Real-World Task3: Put Cube in Cabinet

DP3-EE

DP3-Joint

KADP (Ours)

DP3-EE

DP3-Joint

KADP (Ours)

Real-World Task4: Push Button Elbow

DP3-EE

DP3-Joint

KADP (Ours)

DP3-EE

DP3-Joint

KADP (Ours)

Kinematics-Aware Diffusion Policy
with Consistent 3D Observation and Action Space
for Whole-Arm Robotic Manipulation