This 1.7B Model Clones Your Voice Locally... With a Catch (KittenTTS 2)
Better Stack
0:00 / 0:00
This 1.7B Model Clones Your Voice Locally... With a Catch (KittenTTS 2)
25 568 просмотров · 3 дн. назад
Better Stack
222 тыс. подписчиков
25 568 просмотров · 3 дн. назад
KittenTTS 2 is a new 1.7B parameter text-to-speech model that can clone a voice from just a few seconds of audio, but it’s very different from the tiny Apache 2.0 KittenTTS models we originally loved.
In this video, I test KittenTTS 2 locally, try its voice cloning and emotion features, look at CPU performance on a Mac, break down its model sizes and dependencies, and compare the experience with ElevenLabs.
🔗 Relevant Links
KittenTTS 2 Hugging Face - https://huggingface.co/KittenML/kitte...
Kitten Repo - https://github.com/KittenML/KittenTTS
❤️ More about us
Radically better observability stack: https://betterstack.com/
Written tutorials: https://betterstack.com/community/
Example projects: https://github.com/BetterStackHQ
📱 Socials
Twitter: / betterstackhq
Instagram: / betterstackhq
TikTok: / betterstack
LinkedIn: / betterstack
📌 Chapters:
0:00 KittenTTS 2 Changes Everything
0:38 How KittenTTS 2 Voice Cloning Works
1:18 Model Sizes, Ternary Weights & Built-In Voices
2:10 KittenTTS 2 Voice Cloning Demo
3:00 Mac Performance: CPU vs GPU
3:35 KittenTTS 2 vs ElevenLabs
4:03 What Changed From the Original KittenTTS
5:00 The KittenTTS 2 License Explained
6:00 Languages, Quality & Long-Text Limitations
6:29 Is KittenTTS 2 Really Real-Time on CPU?
7:24 Should You Actually Use KittenTTS 2?