Real-time 3D Human Mesh Recovery for Physical AI

HMR-Lite on a Unitree G1: metric 3D bodies from a single RGB-D frame.

For robots working around people, accurate 3D human geometry is only useful if it arrives in time to act on it. At Seeing Machines, I led the development of HMR, a real-time Human Mesh Recovery system and one of the leading features of Seeing Machines' Human-Centred Physical AI Platform.

We built HMR as a human-body expert. It combines an understanding of human semantics with 3D geometry to recover a person’s articulated pose, body shape and metric camera-frame position from a single RGB or RGB-D image. It runs entirely on the embedded device.

Highlights

On NVIDIA Jetson Thor, using the Harmony4D benchmark:

  • HMR reaches 31.8 mm MPJPE at 15.8 ms with 386M parameters. It is more accurate than Meta’s SAM-3D-Body (33.9 mm) with a ~33× speed-up and ~70% fewer parameters.
  • HMR-Lite runs at 5.5 ms (up to 180 FPS) with 146M parameters. It matches the accuracy of Fast-SAM-3D-Body with a ~20× speed-up and ~88% fewer parameters.
  • Both models support RGB and RGB-D input without increasing their parameter counts.
Accuracy, latency and size compared on NVIDIA Jetson Thor (RGB input, Harmony4D MPJPE). Bubble area is proportional to parameter count.

Demo: Unitree G1 humanoid

Real-time teleoperation of a Unitree G1 using HMR on NVIDIA Jetson Thor.
HMR-Lite running on a Unitree G1 outdoors, in difficult lighting.

How it works

HMR combines a geometry-aware foundation model with a Transformer-based Human Mesh Decoder. The foundation model can take monocular images, multi-view images and depth, and encodes both visual appearance and spatial structure. The decoder maps these features to body shape, articulated pose and camera-frame placement. The full 3D mesh and anatomical joints are then recovered from those parameters.

HMR design overview.

A 3D human mesh gives robots much more than a bounding box or a skeleton. It provides anatomical landmarks, articulation, body shape, body-part geometry, spatial occupancy and, with tracking, movement over time. Robots can use this for safer shared-space operation, human-aware navigation, coordinated handover, teleoperation and robot learning.

A 3D human mesh outputs rich geometric and semantic information.

Read more in the technical paper and the Seeing Machines blog post.

Bo Chen
Bo Chen
Senior Staff Machine Learning Scientist

Machine learning scientist specializing in human-centered 3D perception for Physical AI, robotics, and safety-critical systems.

Related