• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
AI
Announcement

Unpaired translation between images and text draws praise

One post says Dominik's result looked more convincing than expected; another praises novel ideas rather than just scaling.

Yann LeCunYL
Sander DielemanSD
Phillip IsolaPI
14 Sources, 8h ago, first seen 8h ago

TLDR

One post calls unpaired translation between images and text a result its author had dreamed about for years, saying Dominik showed it was more possible than expected. Another says the result works and praises its skillful use of novel ideas rather than just scaling.

Combined views

50.2K

14 Sources, first seen 8h ago

1K likes28 comments647 saves92 reposts

Combined views

50.2K

14 Sources, first seen 8h ago

1K likes28 comments647 saves92 reposts
Video from X

Dominik Schnaus and collaborators say they aligned the embedding spaces of an image model and a text model without training on paired image-caption examples. In Schnaus’s summary, DINOv2 had never seen a caption and Qwen3 had never seen an image. He said the method also worked when the images and captions came from different datasets.

Featured Source

Computer-vision researcher Phillip Isola framed the result as evidence about converging representational geometry. He wrote that a single global map aligned the image and text embeddings well enough for reasonable text-to-image translations. The learned map was orthogonal, alongside centering and normalization.

Isola also stressed what the work does not settle. He said it remains open which structures actually align, and pointed to competing accounts that characterize convergence as local, coarse-grained or relational. The authors’ posts describe a research result, not an independent validation of general multimodal equivalence.

Other researchers responded with a mix of praise and caution. Julian Togelius called it surprising and praised the emphasis on new ideas rather than scale, while Sander Dieleman connected it to the broader semantic gap between language and perception. Yann LeCun noted that older unsupervised-translation work used a similar idea, tempering claims that the approach is entirely new.

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

Related Videos

  • It's a (Blind) Match! Towards Vision-Language Correspondence without Parallel DataVoxel51 · YouTube

Useful links

mcml.ai

MCML - Blind Matching – Aligning Images and Text Without Training or Labels

Voxel51 · YouTube

It's a (Blind) Match! Towards Vision-Language Correspondence without Parallel Data

scholarfeed.org

SOTAlign: Semi-Supervised Alignment of Unimodal Vision and Language Models via Optimal Transport

abhik.ai

DINOv2: Learning Robust Visual Features without Supervision

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

Useful Links

mcml.ai

MCML - Blind Matching – Aligning Images and Text Without Training or Labels

scholarfeed.org

SOTAlign: Semi-Supervised Alignment of Unimodal Vision and Language Models via Optimal Transport

abhik.ai

DINOv2: Learning Robust Visual Features without Supervision

Related Videos

  • It's a (Blind) Match! Towards Vision-Language Correspondence without Parallel DataVoxel51 · YouTube

Useful Links

mcml.ai

MCML - Blind Matching – Aligning Images and Text Without Training or Labels

scholarfeed.org

SOTAlign: Semi-Supervised Alignment of Unimodal Vision and Language Models via Optimal Transport

abhik.ai

DINOv2: Learning Robust Visual Features without Supervision

Related

Should 'AI welfare' be called 'morally obscene'?

One post calls 'AI welfare' morally obscene given human suffering and predicts it will add to a growing anti-AI backlash. A reply argues that calling either view morally obscene isn't useful.

Does an AI PhD still make sense as the field accelerates?

One post argues it’s a great time to pursue an AI PhD. Another says it’s an excellent time for many people to research advanced AI’s societal impacts, but is less certain the traditional PhD format suits many today.

14 Sources

Dominik Schnaus@dominik_schnausDINOv2 has never seen a caption, and Qwen3 has never seen an image. We still aligned their embedding spaces without a single image-caption pair. It even works when the images and the captions come from different datasets. Project page: https://dominik-schnaus.github.io/unpaired-rosetta ⬇️8h
Phillip Isola@phillip_isolaThis is a result I’ve dreamt about for many years: *unpaired translation between images and text* I thought it might be only slightly possible, the kind of thing you have to really squint at. But Dominik proved this wrong. You don’t have to squint. Worth looking for yourself:7h
Julian Togelius@togeliusCrazy that this works. But very cool. (Also, a bit of a palate cleanser to see people posting AI results that come from skillful application of novel ideas rather than just scaling.)7h
Benno Krojer@benno_krojerProbably the paper I am most surprised/excited by this year. Platonic Representation Hypothesis is back?!7h
Shao-Hua Sun@shaohua0116RT @dominik_schnaus: DINOv2 has never seen a caption, and Qwen3 has never seen an image. We still aligned their embedding spaces without a…6h
Djamé..@zehavocRT @benno_krojer: Probably the paper I am most surprised/excited by this year. Platonic Representation Hypothesis is back?!4h
Sander Dieleman@sedielemI've been doing some thinking about multimodality lately, specifically the semantic gap between common representations of language and perceptual signals. In that context, this is an exciting result! (These thoughts are likely to land on my blog in written form at some point✍️)3h
Yann LeCun@ylecun@togelius Old work on unsupervised translation using a similar idea.3h
Pete Skomoroch@peteskomorochRT @phillip_isola: This is a result I’ve dreamt about for many years: *unpaired translation between images and text* I thought it might be…2h
Kosta Derpanis@CSProfKGDRT @phillip_isola: In PRH, we argued that representational geometry is converging, but it's still an open question exactly in what sense.…2h
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    Phillip Isola

    14 Sources

    Dominik Schnaus@dominik_schnausDINOv2 has never seen a caption, and Qwen3 has never seen an image. We still aligned their embedding spaces without a single image-caption pair. It even works when the images and the captions come from different datasets. Project page: https://dominik-schnaus.github.io/unpaired-rosetta ⬇️8h
    Phillip Isola@phillip_isolaThis is a result I’ve dreamt about for many years: *unpaired translation between images and text* I thought it might be only slightly possible, the kind of thing you have to really squint at. But Dominik proved this wrong. You don’t have to squint. Worth looking for yourself:7h
    Julian Togelius@togeliusCrazy that this works. But very cool. (Also, a bit of a palate cleanser to see people posting AI results that come from skillful application of novel ideas rather than just scaling.)7h
    Benno Krojer@benno_krojerProbably the paper I am most surprised/excited by this year. Platonic Representation Hypothesis is back?!7h
    Shao-Hua Sun@shaohua0116RT @dominik_schnaus: DINOv2 has never seen a caption, and Qwen3 has never seen an image. We still aligned their embedding spaces without a…6h
    Djamé..@zehavocRT @benno_krojer: Probably the paper I am most surprised/excited by this year. Platonic Representation Hypothesis is back?!4h
    Sander Dieleman@sedielemI've been doing some thinking about multimodality lately, specifically the semantic gap between common representations of language and perceptual signals. In that context, this is an exciting result! (These thoughts are likely to land on my blog in written form at some point✍️)3h
    Yann LeCun@ylecun@togelius Old work on unsupervised translation using a similar idea.3h
    Pete Skomoroch@peteskomorochRT @phillip_isola: This is a result I’ve dreamt about for many years: *unpaired translation between images and text* I thought it might be…2h
    Kosta Derpanis@CSProfKGDRT @phillip_isola: In PRH, we argued that representational geometry is converging, but it's still an open question exactly in what sense.…2h
    Today's Rank

    #8

    Today's Rank

    #8