Are generative video models the path towards solving visual intelligence? | Robert Geirhos
heidelberg.ai
0:00 / 0:00
Are generative video models the path towards solving visual intelligence? | Robert Geirhos
160 просмотров · 3 недели назад
heidelberg.ai
2,67 тыс. подписчиков
160 просмотров · 3 недели назад
Are video models quietly turning into general-purpose vision models, the same way LLMs became general-purpose language models?
Think about how much language models changed things. A few years back you needed a separate model for every job. One for translation, a different one for summarizing, another for answering questions. Then LLMs showed up and you could suddenly do all of it by just asking. And the recipe behind that shift was pretty plain once you saw it: take a big generative model, train it on a huge pile of web data.
Robert's argument is that the same thing might be starting to happen in vision, and video is where to look.
His team has been poking at Veo 3, and it keeps doing things nobody trained it to do. It'll pick objects out of a scene, find edges, edit images, reason about how physical stuff behaves, work out what you can actually do with an object, even solve mazes and symmetry puzzles. None of that was the training goal. It just emerged, the same way a lot of LLM abilities did.
If that pattern holds up, it's a big deal. It hints at a future where one vision model handles most perception tasks, instead of a new specialized model every single time.