HardFlow, published this week in IEEE Transactions on Pattern Analysis and Machine Intelligence, wraps pretrained diffusion and flow matching models — a newer class of generative AI — so their final output respects robotics, control, and vision
Today's pretrained image and trajectory generators produce convincing outputs that nonetheless bend or break the rules of the physical world: a robot path that crosses a joint limit, a control policy that drives a process outside a safe envelope, a synthesized image that ignores a labeled region. MIT researchers say their new method, HardFlow, lets those models satisfy such hard constraints at deployment time, without retraining.
The move, published this week in IEEE Transactions on Pattern Analysis and Machine Intelligence, is to relax constraints during generation and enforce them only on the final sample. Earlier approaches forced the rules at every denoising step, which trapped the model in low-quality solutions. HardFlow exploits a property its authors observe in modern diffusion and flow-matching models: the unconstrained trajectory already lands near a feasible region, so projecting onto the constraint at the end preserves both quality and feasibility.
"Generative models have shown remarkable power, but applying them to the real world requires respecting physical and operational constraints," senior author Navid Azizan of MIT's Laboratory for Information and Decision Systems said in the announcement.
The team, which includes lead author Zeyang Li and Kaveh Alim, reports evaluations across robotics, control of physical processes, and computer vision. The paper and release describe the technique as plug-and-play on existing pretrained models and name Stable Diffusion and FLUX as example targets. The current source bundle does not include independent validation of adoption or deployment; the IEEE TPAMI publication is the peer-reviewed record, and the arXiv preprint remains available for replication.