r/LocalLLaMA • u/Cyborg-2077 • 18h ago
I Built A Thing Local text to speech with Breeze is truly incredible
Been playing around with local TTS with Breeze combined with STT, and the results are amazing. Using Opus 5.5, I can hear the first sound after 500ms if there is no thinking involved, and with thinking on low mode, can be 1-1.5s.
I'm using a BLE remote (the kind that are used for taking pics with phones) combined with a wireless microphone. So I can just sit on the couch, and just talk to her.
She watches for any claude session that finishes, and sends me the results in a very short, spoken style summary, and tells me if there is anything waiting for my decision, then forwards my decisions.
Also impressed how consistent Opus 5.5 is in the communication. Even after more than 500k in context, he still remembers that he's in a live session with me, and has to keep messages short. Used to be an issue in the past.
The future is here guys.
Edit: For those who wanna give it a try, you can find free avatars such as this one: https://www.live2d.com/en/learn/sample/niziiro-mao/ or you can buy one from a marketplace.
Edit2: Might open-source that later next week with a free avatar. Let me know if anyone would like to contribute to the project.
95
u/Big-Farmer-2192 18h ago
1-1.5 goon/s on low thinking mode.
20
u/Constant_Art_20 17h ago
the bleeding edge performance on that benchmark at the moment. a feet of engineering
10
u/MentholMafia 17h ago
Will you open source this setup?
5
u/Cyborg-2077 15h ago
I can't really open-source it as it is, because the avatar is a paid model I bought, and cannot re-distribute it
17
5
7
u/Acceptable-Cycle4645 12h ago edited 11h ago
audio.cpp supports Breeze TTS 2 and makes it easy to integrate into existing pipelines, such as AIRI. https://github.com/0xShug0/audio.cpp. A single binary for running 100+ audio models (TTS, ASR, music generation, etc.) locally, so you can swap TTS/ASR components to find your favorite combo.
PS audio.cpp has become one of the audio engines in Unsloth.
1
u/vamsammy 10h ago
I've started playing with Breeze in audio.cpp. Very nice! Is there any chance I can approach realtime streaming on a Mac M1 Max 64, say using the 4 bit quant?
3
u/Acceptable-Cycle4645 9h ago edited 1h ago
Thanks for trying audio.cpp! For your question, maybe...GGML’s Metal kernels are less optimized than those of other backends, and the model has received fewer Metal-specific optimizations from the community (this is the major one https://github.com/0xShug0/audio.cpp/pull/554). Let me know if 4-bit works!
Update: just checked on M4 pro it can do realtime streaming with Q8.
Update: add q4_0 GGUF. The quality may not be as stable as Q8. Need to download directly https://huggingface.co/audio-cpp/audio.cpp-gguf/tree/main/Breeze-TTS-2-GGUF The built-in download tools won’t list it as supported yet, as we need more feedback on output quality.
Format Size Warm RTF Peak VRAM BF16 6.84 GiB 0.375 8046 MiB Q4_K 4.15 GiB 0.289 5412 MiB Q4_0 (pick) 3.53 GiB 0.291 4780 MiB
4
u/letsgoiowa 12h ago
I'm using Breeze to narrate some fanfic (sue me) and it does an incredible job. I'd give it a B+. 99% perfect, but that 1% is so standout bad it's like wtf.
For example, SLY there in instead of Slither in. And it only did this once! But man
3
u/Cyborg-2077 12h ago
I also tried it for voice-overs for ads, and my approach was to tell opus to generate it for me given a script, and he would do a great job at re-creating the parts where something is misspelled. Even for things like brand names that are not obvious, he'd be able to identify and fix it on its own. Either opus is really good at analyzing the audio, or I got lucky
6
u/MomentJolly3535 18h ago
How does Breeze TTS compare to Omni voice TTS for those who tested both?
5
u/Name835 16h ago
I find Breeze to be a bit better than Omni, albeit slower/bigger ofc.
I've gotten used to 11labs tts from my ai app using days (kindroid phone calls) and even if what I'm about to say is harsh, I think that compared to the proprietary tts models, breeze and omnivoice suck tbh. :D I still use Breeze atm cause it is the best of the bad options, however this is just my spoiled ass talking.
I tolerate Breeze and other open tts models, but am bummed that they are so much behind the competition (almost completely lacking wide sfx tag support for example), like about a year behind proprietary tts models, even 11V3 and ofc not even mentioning v4 which is even better.
I am still local tts only these days, and Im still super excited and really cannot wait for open tts models to get a lot better, hopefully in the near future! :)
1
3
u/KissMyShinyArse 14h ago
Omni supports more languages I think. For English, I strongly prefer Breeze.
2
2
u/FinBenton 12h ago
I tested breeze when it came out but I went back to Omni after a bit. Its not like omni is better, I think its slightly worse but its faster and uses slightly less VRAM so it fits me better. Breeze has text guidance but the generations became pretty unreliable with that so I didnt find that very useful so for me breeze without guidance wasnt really worth it but I can see it can be better for many.
5
u/Extreme-Way4665 11h ago
Maybe this is a stupid question, you said Local text to speech using Opus 5.5 - isn't Opus not local?
7
u/Cyborg-2077 11h ago
Opus is not local, I send the request to Claude and they return the response, with token streaming. As soon as I receive the first few words, I pass it into breeze, which generates the audio at 4x realtime speed (meaning it can generate 4 seconds of audio in 1 second), and play it immediately.
So the LLM is not local, but the TTS is.
2
2
1
1
1
u/rorowhat 15h ago
What software are you using for the avatar?
2
u/Cyborg-2077 15h ago
Vibecoded with html css
1
u/Any_Door7384 14h ago
and the assets?
1
u/Cyborg-2077 14h ago
Bought online from a website selling vtuber models. Forgot the name tho. You can also find some free ones online
2
u/Vivid-Snow-2089 14h ago
you can just ask opus 5.5 to draw the assets with math now
1
u/Cyborg-2077 14h ago
Hard to believe it would be able to achieve this level of quality. But happy to be proven wrong
1
u/Any_Door7384 14h ago
And then you got the vtuber model and made ai make a html css page using the vtuber model? pretty cool
1
1
u/geldonyetich 14h ago
When I put down the dosh for my AI box a year ago, I was thinking it'd be cool to have a hologram JARVIS agent. Backends like Hermes have the logic covered. Solutions like yours have successfully integrated the front end. Still waiting for that consumer spatial holographic projector. I'd settle for AR glasses but even those are slow to launch.
1
1
u/KissMyShinyArse 14h ago
It's 95% there. I recently converted a ~125k-word book into an audiobook. BreezeTTS 2 occasionally omits a word or phrase, reads something out of sequence, or pronounces the same word differently in different places. It doesn't happen often, though. I think that with a good AI pipeline, you can produce production-quality audiobooks, even with multiple voices.
I also tried BreezeTTS 2 VoiceDesign, and it's great.
1
1
u/aersel24 3h ago
What stt do you use? I feel most don’t understand what I say. Maybe I don’t speak English properly.
1
1
u/youcloudsofdoom 16h ago
I'm using breeze (and love it) for voice cloning and direction for a military sim I'm making, but this post has reminded me that I am in the non-goon minority for voice clone users
1
u/devdevgoat 18h ago
Nice, what are you running it on gpu wise for that speed? I like the idea for the button!
1
u/Cyborg-2077 16h ago
Rtx 4080 super, but I think could fit even on smaller ones. The tts model takes about 4 gigs of VRAM
1
u/p_tttttri 16h ago
Is there any possibility of open sourcing this setup
1
u/Cyborg-2077 15h ago
I can't really open-source it as it is, because the avatar is a paid model I bought, and cannot re-distribute it
1
u/Open-Adhesiveness-86 12h ago
the push-to-talk remote is probably doing a lot of the work on that 500ms number, since you skip the 500-800ms of VAD silence-wait that eats most "realtime" voice stacks. watch out though, those camera shutter BLE buttons tend to sleep after a couple minutes idle, so the first press just wakes the link and gets swallowed.
1
u/Cyborg-2077 12h ago
Yep, exactly that. I find push to talk to be perfect for this kind of "ai assistant".
I tried live models from both elevenlabs, and openai, and it's not great. Higher delay, you pay by the minute, they are stupid, and as soon as you stop talking, they start replying even if you're not done. So you have to add "hmmms" every time you need to take some time to think.
2
u/Open-Adhesiveness-86 11h ago
The latency floor on cloud voice APIs just kills the conversational feel, especially when pricing punishes pauses. Local PTT with a quantized model on-device feels way closer to natural walkie-talkie flow.
1
u/Cyborg-2077 11h ago
I might try with a local model to see how far I can bring the latency down. Need something that fits in my remaining ~4gigs of ram. Any suggestion?
2
u/Open-Adhesiveness-86 11h ago
Llama-3.2-3B or Qwen2.5-3B at Q4_K_M easily squeeze into under 2.5GB RAM with a 2k context window, leaving headroom for system and audio buffers. If you want even lower TTFT for conversational turns, SmolLM2-1.7B or Qwen2.5-1.5B respond almost instantly on CPU.
0
-5
u/Trustest-S 18h ago
this local voice setup beats most chatbot apps for companions, the quick response makes roleplay feel way more real than the laggy ones ive tried.
51
u/youcloudsofdoom 16h ago
I'm using breeze (and love it) for voice cloning and direction for a military sim I'm making, but this post has reminded me that I am in the non-goon minority for voice clone users