Перейти к содержимому

Uber's GenAI Leap: Batch Predictions Using Ray and vLLM | Ray Summit 2024

Anyscale

0:00 / 0:00

Uber's GenAI Leap: Batch Predictions Using Ray and vLLM | Ray Summit 2024

742 просмотра · 1 год назад
Anyscale
15,5 тыс. подписчиков
742 просмотра · 1 год назад
At Ray Summit 2024, Anant Vyas, Baojun Liu, and Bo Ling from Uber present their innovative approach to large-scale Generative AI batch prediction. The talk focuses on Uber's integration of Ray and vLLM within their Michelangelo machine learning platform to enhance GenAI application development. The speakers discuss how this new approach addresses limitations in traditional Spark-based methods, particularly for GPU-intensive tasks. They explain how the combination of Ray's parallel processing capabilities and vLLM's advanced NLP performance has enabled Uber to create a more efficient and scalable batch prediction workflow. The presentation covers the architecture of Uber's new system, its integration with Kubernetes and Michelangelo's LLM evaluation workflow, and its application to various Uber services including Eats carousel, rider search, and customer obsession. The team shares benchmarking results and insights gained from developing and implementing this solution. This session offers valuable insights for organizations looking to scale their Generative AI capabilities, demonstrating how Ray and vLLM can be leveraged to improve prediction tasks, reduce latency, and enhance overall GenAI performance. -- Interested in more? Watch the full Day 1 Keynote:    • Ray Summit 2024 Keynote Day 1 | Where Buil...   Watch the full Day 2 Keynote    • Ray Summit 2024 Keynote Day 2 | Where Buil...   -- 🔗 Connect with us: Subscribe to our YouTube channel:    / @anyscale   Twitter: https://x.com/anyscalecompute LinkedIn:   / joinanyscale   Website: https://www.anyscale.com