Back to all publications...

Vision-Language Models Fail to Generalize Across Modalities

Vision-language models exhibit many surprisingly simple failures, but why these failures occur remains unclear. We conjecture that their source is representational misalignment in the backbone’s vision and language representations. We demonstrate a new generalization failure that would not occur if the representations were easily alignable, followed by a set of theory-grounded experiments further showing that the representations cannot be aligned using any linear transform. The representations are not expected to be better aligned with sufficient scale due to each modality containing inherently different information. Modern paradigms, such as reasoning or in-context learning, do not alleviate existing failures either. These results suggest that existing paradigms are incapable of preventing these failures and falsify a strong version of the Platonic Representation Hypothesis – that sufficiently powerful models trained in different modalities should converge to equivalent representations.


Yonatan Gideoni, Yoav Gelberg, Tim G. J. Rudner, Yarin Gal
OpenReview
[paper]

Are you looking to do a PhD in machine learning? Did you do a PhD in another field and want to do a postdoc in machine learning? Would you like to visit the group?

How to apply


Contact

We are located at
Department of Computer Science, University of Oxford
Wolfson Building
Parks Road
OXFORD
OX1 3QD
UK
Twitter: @OATML_Oxford
Github: OATML
Email: oatml@cs.ox.ac.uk