AI Needs to Learn to See
Gradient Flow
0:00 / 0:00
AI Needs to Learn to See
1 980 просмотров · 4 дня назад
Gradient Flow
6,61 тыс. подписчиков
1 980 просмотров · 4 дня назад
Episode notes: https://thedataexchange.media/andrew-...
Ben Lorica speaks with Andrew Dai, co-founder and CEO of Elorian AI, about why today’s frontier models still struggle with complex visual reasoning and why simply scaling language-centric architectures may not solve the problem.
Sections
What Visual Reasoning Means—and Why Frontier Models Still Struggle - 00:00:52
Post-Training vs. Pre-Training for Complex Visual Tasks - 00:05:15
Why Image Generation Is Not the Same as Visual Understanding - 00:07:22
Video Understanding, Embeddings, and Multimodal Benchmarks - 00:09:31
Elorian’s Approach: Data, Architecture, RL, and Real-World Use Cases - 00:13:16
Building a General Multimodal Foundation Model Instead of Custom Models - 00:18:30
Robotics, Physical AI, and the Vision Bottleneck - 00:20:30
Why Scaling Language Models Hasn’t Solved Vision - 00:22:35
Can a Visual-AI Startup Stay Ahead of the Frontier Labs? - 00:25:59
Why Visual Reasoning May Require a Different Architecture - 00:28:43
Reasoning Doesn’t Start With Language: Visual Chain of Thought - 00:31:46
How Enterprises Should Evaluate Visual Reasoning Models - 00:34:28
World Models, Multimodal Costs, and Ranking Frontier Models on Vision - 00:36:39
CAD, Vibe Cadding, and the Breakout Use Case for Visual AI - 00:41:37
Support our work:
subscribe to our newsletter 📩 https://gradientflow.substack.com/
leave us a tip 💰 https://buymeacoffee.com/gradientflow