Can this VLA Work with no Dataset? I Put MolmoAct2 to Test
Greg's Tech
0:00 / 0:00
Can this VLA Work with no Dataset? I Put MolmoAct2 to Test
3 930 просмотров · 1 месяц назад
Greg's Tech
1,01 тыс. подписчиков
3 930 просмотров · 1 месяц назад
I test MolmoAct2 — a zero-shot VLA (Vision Language Action Model) that claims to work without any dataset or pre-training. Can this embodied AI model really generalize to new objects and tasks out of the box? I put it through three levels: zero-shot robot arm grasping, VLA fine-tuning for a specific task, and transferring to a completely different embodiment — a drone.
🛠 Hardware
• SO-101 robot arm
• Tello drone
🤖 Models used in the video
• MolmoAct2 (zero-shot + fine-tuned)
• X-VLA (fine-tuned VLA)
📋 Chapters
0:00 — Intro
1:10 — Set up zero-shot inference
3:07 — Zero-shot test: picking up a rubber duck
3:59 — The CAPTCHA challenge
4:47 — How to fine-tune MolmoAct2?
8:53 — Comparison with XVLA
10:36 — MolmoAct on drone: changing the embodiment
11:57 — Done approaches fruits
13:09 — Out-of-dataset generalization test
14:00 — Landing test
14:58 — Outdoor test: chasing a moving rover
17:06 — Final verdict: should you use MolmoAct2?
#MolmoAct2 #VLA #EmbodiedAI #Robotics