Dominik Schnaus and collaborators say they aligned the embedding spaces of an image model and a text model without training on paired image-caption examples. In Schnaus’s summary, DINOv2 had never seen a caption and Qwen3 had never seen an image. He said the method also worked when the images and captions came from different datasets.
Computer-vision researcher Phillip Isola framed the result as evidence about converging representational geometry. He wrote that a single global map aligned the image and text embeddings well enough for reasonable text-to-image translations. The learned map was orthogonal, alongside centering and normalization.
Isola also stressed what the work does not settle. He said it remains open which structures actually align, and pointed to competing accounts that characterize convergence as local, coarse-grained or relational. The authors’ posts describe a research result, not an independent validation of general multimodal equivalence.
Other researchers responded with a mix of praise and caution. Julian Togelius called it surprising and praised the emphasis on new ideas rather than scale, while Sander Dieleman connected it to the broader semantic gap between language and perception. Yann LeCun noted that older unsupervised-translation work used a similar idea, tempering claims that the approach is entirely new.