Перейти к содержимому

Guest lecture of EN.601.768 Language Model Agents @ JHU, Fall 2026

HAIVLab

0:00 / 0:00

Guest lecture of EN.601.768 Language Model Agents @ JHU, Fall 2026

20 просмотров · 9 дней назад
HAIVLab
12 подписчиков
20 просмотров · 9 дней назад
How much information does a machine actually need to understand a video? Long videos are full of pixel redundancy, but downstream tasks — finding an event, answering a question, retrieving a moment — only care about a small slice of the timeline. In this talk, I'll argue that video compression for AI should shift its unit of representation from pixels and generic visual tokens to persistent semantic entities: objects, interactions, and events. I'll introduce the O-I-E framework, a lightweight pipeline that builds structured, queryable video representations without running a large vision-language model on every frame, and show how a hybrid memory of structured facts, retrieval embeddings, and visual residuals supports coarse-to-fine question answering. I'll close with an experimental framework for separating semantic abstraction loss from compression loss, and a roadmap — using surgical video and robotics as proving grounds — toward codecs that preserve task utility at a fraction of the bitrate of pixel-based approaches.