Flow Generative Models for Robot Manipulation

Fan Zhang,   Michael Gienger
Honda Research Institute Europe

To achieve robust general manipulation, we tackle the following key challenges: first, we perform parameter-efficient affordance grounding by prompt-tuning frozen foundation models to capture spatial and semantic scene cues; second, we execute visuomotor action generation via flow matching, transforming random waypoints into target trajectories with fast inference and stable training; and third, we integrate a multimodal predictive world model using generative latent flow matching to anticipate future sensory states—such as temporal audio and visual dynamics—enabling long-horizon reasoning and resolving physical ambiguities during complex tasks.


Highlights

Flow matching represents a robot visuomotor policy as a conditional process of flowing random waypoints to desired robot actions.
Prompt tuning for vision-language-model to predict manipulation affordances in multi-task scenarios.
Flow matching exhibits more stable training and evaluation, and noticeably faster inference than diffusion policy.
A demonstration of robotic sorting toys while keeping storage bins open.

Affordance-based VLA with Flow Matching

Affordance-based VLA with flow matching has been tested on tasks across Activities of Daily Living, and leads to consistently better performance than alternative behavior cloning methods. (Videos are 4x speed)

Comb the hair
Sweep the trash
Hang the towel
Pass the water

End-to-End Manipulation with Flow Matching


Multimodal World Action Model with Flow Matching


Paper

Frontiers in Robotics and AI - Robot Learning and Evolution, 2026.
Affordance-based Robot Manipulation with Flow Matching
Fan Zhang, Michael Gienger

arXiv:2512.08405 [cs.RO].
Learning Robot Manipulation from Audio World Models
Fan Zhang, Michael Gienger

Code is here: https://github.com/HRI-EU/flow_matching


Supervised Works

Reinforcement Learning Conference (RLC), 2026.
Assistax: A Hardware-Accelerated Reinforcement Learning Benchmark for Assistive Robotics
Leonard Hinckeldey, Elliot Fosong, Rimvydas Rubavicius, Elle Miller, Trevor McInroe, Fan Zhang, Patricia Wollstadt, Stefano V. Albrecht, Subramanian Ramamoorthy
★ RLC 2026 Outstanding Paper Awards - Tooling, Environments, and Evaluation for Reinforcement Learning ★

IROS, 2025.
CCDP: Composition of Conditional Diffusion Policies with Guided Sampling
Amirreza Razmjoo, Sylvain Calinon, Michael Gienger, Fan Zhang
★ Outstanding Paper Award at IROS 2025 Workshop on The Art of Robustness: Surviving Failures in Robotics ★

IROS, 2025, Foundation Models for Robotic Design workshop.
Generation of Real-time Robotic Emotional Expressions Learning from Human Demonstration in Mixed Reality
Chao Wang, Michael Gienger, Fan Zhang


Team