Перейти к содержимому

This 1.7B Model Clones Your Voice Locally... With a Catch (KittenTTS 2)

Better Stack

0:00 / 0:00

This 1.7B Model Clones Your Voice Locally... With a Catch (KittenTTS 2)

25 568 просмотров · 3 дн. назад
Better Stack
222 тыс. подписчиков
25 568 просмотров · 3 дн. назад
KittenTTS 2 is a new 1.7B parameter text-to-speech model that can clone a voice from just a few seconds of audio, but it’s very different from the tiny Apache 2.0 KittenTTS models we originally loved. In this video, I test KittenTTS 2 locally, try its voice cloning and emotion features, look at CPU performance on a Mac, break down its model sizes and dependencies, and compare the experience with ElevenLabs. 🔗 Relevant Links KittenTTS 2 Hugging Face - https://huggingface.co/KittenML/kitte... Kitten Repo - https://github.com/KittenML/KittenTTS ❤️ More about us Radically better observability stack: https://betterstack.com/ Written tutorials: https://betterstack.com/community/ Example projects: https://github.com/BetterStackHQ 📱 Socials Twitter:   / betterstackhq   Instagram:   / betterstackhq   TikTok:   / betterstack   LinkedIn:   / betterstack   📌 Chapters: 0:00 KittenTTS 2 Changes Everything 0:38 How KittenTTS 2 Voice Cloning Works 1:18 Model Sizes, Ternary Weights & Built-In Voices 2:10 KittenTTS 2 Voice Cloning Demo 3:00 Mac Performance: CPU vs GPU 3:35 KittenTTS 2 vs ElevenLabs 4:03 What Changed From the Original KittenTTS 5:00 The KittenTTS 2 License Explained 6:00 Languages, Quality & Long-Text Limitations 6:29 Is KittenTTS 2 Really Real-Time on CPU? 7:24 Should You Actually Use KittenTTS 2?