Flow Generative Models for Robot Manipulation
To achieve robust general manipulation, we tackle the following key challenges: first, we perform parameter-efficient affordance grounding by prompt-tuning frozen foundation models to capture spatial and semantic scene cues; second, we execute visuomotor action generation via flow matching, transforming random waypoints into target trajectories with fast inference and stable training; and third, we integrate a multimodal predictive world model using generative latent flow matching to anticipate future sensory states—such as temporal audio and visual dynamics—enabling long-horizon reasoning and resolving physical ambiguities during complex tasks.
Highlights
Affordance-based VLA with Flow Matching
Affordance-based VLA with flow matching has been tested on tasks across Activities of Daily Living, and leads to consistently better performance than alternative behavior cloning methods. (Videos are 4x speed)
End-to-End Manipulation with Flow Matching
Multimodal World Action Model with Flow Matching
Paper
Frontiers in Robotics and AI - Robot Learning and Evolution, 2026.
Affordance-based Robot Manipulation with Flow Matching
Fan Zhang, Michael Gienger
arXiv:2512.08405 [cs.RO].
Learning Robot Manipulation from Audio World Models
Fan Zhang, Michael Gienger
Code is here: https://github.com/HRI-EU/flow_matching
Supervised Works
Reinforcement Learning Conference (RLC), 2026.
Assistax: A Hardware-Accelerated Reinforcement Learning Benchmark for Assistive Robotics
Leonard Hinckeldey, Elliot Fosong, Rimvydas Rubavicius, Elle Miller, Trevor McInroe, Fan Zhang, Patricia Wollstadt, Stefano V. Albrecht, Subramanian Ramamoorthy
★ RLC 2026 Outstanding Paper Awards - Tooling, Environments, and Evaluation for Reinforcement Learning ★
IROS, 2025.
CCDP: Composition of Conditional Diffusion Policies with Guided Sampling
Amirreza Razmjoo, Sylvain Calinon, Michael Gienger, Fan Zhang
★ Outstanding Paper Award at IROS 2025 Workshop on The Art of Robustness: Surviving Failures in Robotics ★
IROS, 2025, Foundation Models for Robotic Design workshop.
Generation of Real-time Robotic Emotional Expressions Learning from Human Demonstration in Mixed Reality
Chao Wang, Michael Gienger, Fan Zhang