Mixed Reality · Generative Motion · Multi-Person XR

Mixed-Reality Platforms for Capturing Human Demonstrations and Learning Expressive Robot Behavior

We use mixed and extended reality to capture and retarget human expressive behavior, learn autonomous robot motion, and record synchronized multi-person demonstrations in shared XR—connecting individual expression, generative behavior, and social-physical interaction across one research direction.

Capture & Retarget Individual demonstration

Capture and Retarget Human Expression in Mixed Reality

An individual expert steps into the robot's point of view and performs the behavior directly. Facial activity, gaze, head motion, and upper-body movement are captured together and retargeted to how the robot should look, orient, and respond.

Mapping facial expression, gaze, head pose, and hand pose from an XR user to a virtual robot
Human-to-robot retargeting. Selected facial signals animate the screen eyes and ears; gaze positions the eyes; head and hand motion drive the robot's neck and upper body while the operator observes the interaction from a first-person view.
01

Face and gaze

Blendshapes and gaze direction control the robot's eye shape, eye position, and ear attitude.

02

Head and posture

Head pose transfers orientation and expressive timing while keeping the robot rooted in its scene.

03

Hands and arms

Tracked controllers or hands provide targets for broad, legible upper-body gestures.

Generate Autonomous expression

Generate Autonomous Expressive Behavior

The captured and retargeted demonstrations train a flow-matching policy rather than a library of fixed gestures. Conditioned on an emotion label, recent robot state, and a moving target, it produces coherent and varied action sequences in real time.

Demonstration capture, model training, and real-time robot inference pipeline
Demonstration data turns a high-level emotional intention into synchronized facial, head, and arm motion that can be learned and replayed by the robot.
Flow-matching architecture conditioned on robot history, target pose, and emotion
A history window, target pose, and emotion condition guide the model toward a continuous future action sequence for the robot.

Flow matching

Variation without losing intent

The model captures multiple valid ways to express the same emotion. This gives the robot motion that is recognizable yet not mechanically identical from one rollout to the next.

  • Explicit emotional conditioning
  • Responsive behavior around moving objects
  • Continuous joint targets for eyes, ears, neck, and arms

Demonstration to inference

See the complete learning loop

The same XR interface supports expert data collection and observation of the learned policy.

Visual summary. An expert performs expressive robot behavior in XR; the captured demonstrations train a policy that then generates emotion-conditioned motion around a target. Audio is not required to follow the research content.

Generated behavior spans different valence and arousal levels while combining facial, head, and arm cues. Open the image for the full-resolution comparison.

Preliminary qualitative observations

What we learned

Shorter context worked better

Two-to-four-frame histories were more effective than a 16-frame window in the current architecture.

Longer horizons felt more complete

Thirty-two-frame predictions produced fuller gestures; shorter horizons introduced occasional jumps.

Targeted data still matters

Six emotions transferred convincingly, while the distinctive curious “poke” remained underrepresented.

Generated expression

Emotion-conditioned motion in context

Visual summary. The robot performs distinct face, head, and arm behaviors around the same tabletop target under different emotion conditions. Audio is not required to interpret the comparison.

Multi-Person Capture Shared XR demonstration

Capture Multi-Person Demonstrations in Shared XR

XR3 extends demonstration capture from one expert to two people sharing the same physical space. A participant and a hidden operator perform a synchronized human–human interaction: the participant meets an expressive virtual robot in immersive VR while the operator uses passthrough MR to embody it, respond, and deliver physical contact.

A shared spatial anchor aligns both headsets. The system streams the enacted robot motion between views and records synchronized pose, face, participant gaze, pedal-triggered utterance events, and virtual contact events as one multi-person demonstration.
01

Two people, two perspectives

The operator works in passthrough MR while the participant experiences the enacted robot in immersive VR.

02

Shared spatial alignment

A common anchor and runtime adjustment align the virtual robot with the real interaction space.

03

Enacted social-physical cues

The operator coordinates robot head, gaze, face, hands, utterance events, and physical contact in real time.

04

Synchronized demonstrations

Time-aligned streams connect the enacted behavior with participant gaze, responses, and contact events.

Multi-person physical and social retargeting

Align Sight, Motion, and Touch

The hidden operator—not a physical robot—delivers the touch. Arm, palm, thumb, and index-finger inverse kinematics align that felt contact with the virtual robot hand seen by the participant.

Visual summary. A hidden, co-located operator animates the virtual robot and delivers touch while the participant experiences the aligned robot encounter in immersive VR. Audio is not required to understand the demonstrated interaction.

XR3 in action

Visible robot contact, physically felt

The operator controls head, gaze, face, arms, and hands in real time. Passive fingertip covers can further match the geometry and feel of the virtual robot's fingertips during light taps and presses.

Face and gaze animate the robot's screen face and ears; head and hand poses drive the head, arms, palm, thumb, and index finger. The interaction sequence moves from introduction to drawing on the participant's palm, acknowledging the answer, and closing the encounter.

Research outputs

Publications

Two publications document autonomous expression learned from individual demonstrations and synchronized multi-person demonstration capture in shared XR.

  1. 2025

    Generation of Real-time Robotic Emotional Expressions Learning from Human Demonstration in Mixed Reality

    Chao Wang, Michael Gienger, and Fan Zhang

  2. 2026

    XR3: An Extended Reality Platform for Social-Physical Human-Robot Interaction

    Chao Wang, Anna Belardinelli, and Michael Gienger