At the European Conference on Computer Vision (ECCV) 2026 in Malmö, Sweden, Stanford's Jiajun Wu argued multimodal AI (combining images, sound, and touch) breaks on new objects when sound and touch data vanish, and physics is the shared fix.
A robot that walks into a room it has never seen hears nothing and feels nothing through its sensors. On September 9 at ECCV 2026 in Malmö, Sweden, Stanford assistant professor Jiajun Wu called that vanishing data the real bottleneck of today's multimodal AI, the subfield of artificial intelligence that combines more than one kind of sensor input such as images, sound, and touch. The way out, he said, runs through a shared coordinate system built from physical properties.
ECCV, the European Conference on Computer Vision, is one of the field's two annual main events, alongside CVPR. Wu's talk, titled "Physics-Grounded Multimodal Perception: From Vision, Sound, to Touch," was the third time in roughly two months he has laid out the case in public. The argument, drawn from a translated talk recap of his ECCV talk, names a specific failure mode in the dominant recipe.
That recipe feeds images, audio, and touch into a single Transformer, the neural network architecture behind today's largest language and image models, and lets scale do the work. Wu's critique is direct: in any new environment with unseen objects, the audio and tactile channels collapse. The system has no images of how a teacup will sound when struck, and no recordings of how its surface will feel to a gripper. Stacking more layers does not solve that. Adding sensors does not solve that either. The problem is the data, and the absence of a prior that lets one modality borrow from another.
Wu's proposed prior is a set of physical attributes. Geometry, material, and stiffness (a property engineers call Young's modulus) describe the same teacup regardless of whether a camera, microphone, or skin sensor is reading it. Treat those properties as the common coordinate system, and a vision model can reason about how a teacup will sound; a touch model can predict what it looks like. The senses become projections of one underlying physical object.
His group has spent the last two years turning that idea into working systems. PhysDreamer distills physical properties from video diffusion models and uses them to predict how objects will move under 3D forces. It was an oral paper at ECCV 2024, one of the conference's highest distinctions. WonderPlay uses the same properties as an intermediate interface to a generative model, and was presented as a poster at ICCV 2025. DiffImpact brings the same logic to sound by making impact audio differentiable for learning. DexSkin, a soft electronic skin for tactile sensing, gives the touch channel a real sensor rather than a synthetic one. Together, the four papers form a portfolio that Wu's group page presents as one claim: physical properties, not raw sensor streams, are the right unit of fusion.
The Sept. 9 talk was Wu's third major lecture in roughly two months. The day before, at the conference's "Robotic Manipulation Research" workshop, he presented a complementary thesis: the right level of abstraction for robot skills is the atomic skill, a small self-contained action with explicit preconditions and postconditions that can be composed into long-horizon tasks. Pair the two, and a robot should be able to walk into a new room, identify the physics of what it sees and touches, and chain the right atomic skills without retraining.
The argument lands as the multimodal field is openly debating a scaling plateau. Larger Transformers have driven most of the gains since GPT-4, but their appetite for labeled multimodal data hits a wall wherever the environment is novel. Wu argues that researchers will get further by treating physics as a learned prior than by scaling the same architecture. The PhysDreamer and WonderPlay recipes, still preprint-stage demos, will face their first independent benchmarks in the workshop reviews and replication attempts expected before NeurIPS 2026's December deadline.