Перейти к содержимому

Can this VLA Work with no Dataset? I Put MolmoAct2 to Test

Greg's Tech

0:00 / 0:00

Can this VLA Work with no Dataset? I Put MolmoAct2 to Test

3 930 просмотров · 1 месяц назад
Greg's Tech
1,01 тыс. подписчиков
3 930 просмотров · 1 месяц назад
I test MolmoAct2 — a zero-shot VLA (Vision Language Action Model) that claims to work without any dataset or pre-training. Can this embodied AI model really generalize to new objects and tasks out of the box? I put it through three levels: zero-shot robot arm grasping, VLA fine-tuning for a specific task, and transferring to a completely different embodiment — a drone. 🛠 Hardware • SO-101 robot arm • Tello drone 🤖 Models used in the video • MolmoAct2 (zero-shot + fine-tuned) • X-VLA (fine-tuned VLA) 📋 Chapters 0:00 — Intro 1:10 — Set up zero-shot inference 3:07 — Zero-shot test: picking up a rubber duck 3:59 — The CAPTCHA challenge 4:47 — How to fine-tune MolmoAct2? 8:53 — Comparison with XVLA 10:36 — MolmoAct on drone: changing the embodiment 11:57 — Done approaches fruits 13:09 — Out-of-dataset generalization test 14:00 — Landing test 14:58 — Outdoor test: chasing a moving rover 17:06 — Final verdict: should you use MolmoAct2? #MolmoAct2 #VLA #EmbodiedAI #Robotics