Building AI Agents That Can See and Hear: Complete LiveKit Tutorial
App Vitals
0:00 / 0:00
Building AI Agents That Can See and Hear: Complete LiveKit Tutorial
8 762 просмотра · 1 год назад
App Vitals
463 подписчика
8 762 просмотра · 1 год назад
🤖 Learn how to build AI agents with vision and hearing capabilities using LiveKit! In this comprehensive tutorial, we show you how to create a multi-modal AI customer support agent that can see your screen, hear your voice, and provide intelligent assistance in real-time.
Watch as we build a complete system where the AI agent can:
Join video calls and see shared screens
Automatically detect and diagnose issues visually
Respond with voice using text-to-speech
Access knowledge bases for accurate support
Handle multiple participants in real-time
🔧 What You'll Learn:
LiveKit server setup and integration
Building React frontends with LiveKit SDK
Implementing speech-to-text and text-to-speech
Screen sharing and video frame capture
Multi-modal LLM integration with OpenAI
LangFuse tracing for production monitoring
Horizontal scaling with LiveKit workers
🔗 Resources:
Demo Repository: https://github.com/app-vitals/livekit...
LiveKit Documentation: https://docs.livekit.io/agents/
👋🏻 About Us:
Hi! We're Dan and Dave, the founders of App Vitals. On this channel, we share practical tutorials that teach developers how to build production-ready AI systems that actually work in the real world.
🗓 Need help implementing AI in your business? Book a free consultation:
https://cal.com/dan-mcaulay/discovery...
💬 COMMENT BELOW: What kind of multi-modal AI applications are you planning to build? Have you worked with LiveKit before?