Face and gaze
Blendshapes and gaze direction control the robot's eye shape, eye position, and ear attitude.
Mixed Reality · Generative Motion · Multi-Person XR
We use mixed and extended reality to capture and retarget human expressive behavior, learn autonomous robot motion, and record synchronized multi-person demonstrations in shared XR—connecting individual expression, generative behavior, and social-physical interaction across one research direction.
Capture & Retarget Individual demonstration
An individual expert steps into the robot's point of view and performs the behavior directly. Facial activity, gaze, head motion, and upper-body movement are captured together and retargeted to how the robot should look, orient, and respond.
Blendshapes and gaze direction control the robot's eye shape, eye position, and ear attitude.
Head pose transfers orientation and expressive timing while keeping the robot rooted in its scene.
Tracked controllers or hands provide targets for broad, legible upper-body gestures.
Generate Autonomous expression
The captured and retargeted demonstrations train a flow-matching policy rather than a library of fixed gestures. Conditioned on an emotion label, recent robot state, and a moving target, it produces coherent and varied action sequences in real time.
Flow matching
The model captures multiple valid ways to express the same emotion. This gives the robot motion that is recognizable yet not mechanically identical from one rollout to the next.
Demonstration to inference
The same XR interface supports expert data collection and observation of the learned policy.
Visual summary. An expert performs expressive robot behavior in XR; the captured demonstrations train a policy that then generates emotion-conditioned motion around a target. Audio is not required to follow the research content.
Preliminary qualitative observations
Two-to-four-frame histories were more effective than a 16-frame window in the current architecture.
Thirty-two-frame predictions produced fuller gestures; shorter horizons introduced occasional jumps.
Six emotions transferred convincingly, while the distinctive curious “poke” remained underrepresented.
Generated expression
Visual summary. The robot performs distinct face, head, and arm behaviors around the same tabletop target under different emotion conditions. Audio is not required to interpret the comparison.
Multi-Person Capture Shared XR demonstration
XR3 extends demonstration capture from one expert to two people sharing the same physical space. A participant and a hidden operator perform a synchronized human–human interaction: the participant meets an expressive virtual robot in immersive VR while the operator uses passthrough MR to embody it, respond, and deliver physical contact.
The operator works in passthrough MR while the participant experiences the enacted robot in immersive VR.
A common anchor and runtime adjustment align the virtual robot with the real interaction space.
The operator coordinates robot head, gaze, face, hands, utterance events, and physical contact in real time.
Time-aligned streams connect the enacted behavior with participant gaze, responses, and contact events.
Multi-person physical and social retargeting
The hidden operator—not a physical robot—delivers the touch. Arm, palm, thumb, and index-finger inverse kinematics align that felt contact with the virtual robot hand seen by the participant.
Visual summary. A hidden, co-located operator animates the virtual robot and delivers touch while the participant experiences the aligned robot encounter in immersive VR. Audio is not required to understand the demonstrated interaction.
XR3 in action
The operator controls head, gaze, face, arms, and hands in real time. Passive fingertip covers can further match the geometry and feel of the virtual robot's fingertips during light taps and presses.
Research outputs
Two publications document autonomous expression learned from individual demonstrations and synchronized multi-person demonstration capture in shared XR.
Chao Wang, Anna Belardinelli, and Michael Gienger