r/LocalLLaMA • • 18h ago

I Built A Thing Local text to speech with Breeze is truly incredible

Been playing around with local TTS with Breeze combined with STT, and the results are amazing. Using Opus 5.5, I can hear the first sound after 500ms if there is no thinking involved, and with thinking on low mode, can be 1-1.5s.

I'm using a BLE remote (the kind that are used for taking pics with phones) combined with a wireless microphone. So I can just sit on the couch, and just talk to her.

She watches for any claude session that finishes, and sends me the results in a very short, spoken style summary, and tells me if there is anything waiting for my decision, then forwards my decisions.

Also impressed how consistent Opus 5.5 is in the communication. Even after more than 500k in context, he still remembers that he's in a live session with me, and has to keep messages short. Used to be an issue in the past.

The future is here guys.

Edit: For those who wanna give it a try, you can find free avatars such as this one: https://www.live2d.com/en/learn/sample/niziiro-mao/ or you can buy one from a marketplace.

Edit2: Might open-source that later next week with a free avatar. Let me know if anyone would like to contribute to the project.

152 Upvotes

64 comments sorted by

51

u/youcloudsofdoom 16h ago

I'm using breeze (and love it) for voice cloning and direction for a military sim I'm making, but this post has reminded me that I am in the non-goon minority for voice clone users

6

u/Cyborg-2077 15h ago

I actually tried with serious looking 3d avatars such as meta human, but it's a pain in the ass. I figured I'd try vtubers models and it ended up looking so much better, also performance is much better

-1

u/Homberger 12h ago

You've built, what everyone is dreaming of, what personal ai should be like! At least, that's what I am dreaming of.

I'd really like to build something like you did, without propetiary/closed models. 

Are there any models, avatars, or creators you’d recommend checking out as a newbie to animation and tts/stt? Or maybe even a few specific models that you find very good?

2

u/Cyborg-2077 12h ago

1

u/Homberger 10h ago

VERY interesting! I would never ever have found this by myself. This looks like a good starting point to dive in! Thank you!

0

u/Imaginary-Unit-3267 5h ago

Meanwhile I'm like "I can barely speak out loud to human beings, why do you expect me to communicate over voice with a machine any more effectively?" It's like people don't know how to type.

95

u/Big-Farmer-2192 18h ago

1-1.5 goon/s on low thinking mode.

20

u/Constant_Art_20 17h ago

the bleeding edge performance on that benchmark at the moment. a feet of engineering

8

u/Sacra34 16h ago

State of the art performance in goon leaderboards

10

u/MentholMafia 17h ago

Will you open source this setup?

5

u/Cyborg-2077 15h ago

I can't really open-source it as it is, because the avatar is a paid model I bought, and cannot re-distribute it

17

u/MechwolfMachina 13h ago

The one with the 3 arms…?

5

u/SnowGrayMan 13h ago

I think he meant the actual setup, not your 3 handed avatar.

7

u/Acceptable-Cycle4645 12h ago edited 11h ago

audio.cpp supports Breeze TTS 2 and makes it easy to integrate into existing pipelines, such as AIRI. https://github.com/0xShug0/audio.cpp. A single binary for running 100+ audio models (TTS, ASR, music generation, etc.) locally, so you can swap TTS/ASR components to find your favorite combo.

PS audio.cpp has become one of the audio engines in Unsloth.

1

u/vamsammy 10h ago

I've started playing with Breeze in audio.cpp. Very nice! Is there any chance I can approach realtime streaming on a Mac M1 Max 64, say using the 4 bit quant?

3

u/Acceptable-Cycle4645 9h ago edited 1h ago

Thanks for trying audio.cpp! For your question, maybe...GGML’s Metal kernels are less optimized than those of other backends, and the model has received fewer Metal-specific optimizations from the community (this is the major one https://github.com/0xShug0/audio.cpp/pull/554). Let me know if 4-bit works!

Update: just checked on M4 pro it can do realtime streaming with Q8.

Update: add q4_0 GGUF. The quality may not be as stable as Q8. Need to download directly https://huggingface.co/audio-cpp/audio.cpp-gguf/tree/main/Breeze-TTS-2-GGUF The built-in download tools won’t list it as supported yet, as we need more feedback on output quality.

Format Size Warm RTF Peak VRAM
BF16 6.84 GiB 0.375 8046 MiB
Q4_K 4.15 GiB 0.289 5412 MiB
Q4_0 (pick) 3.53 GiB 0.291 4780 MiB

4

u/letsgoiowa 12h ago

I'm using Breeze to narrate some fanfic (sue me) and it does an incredible job. I'd give it a B+. 99% perfect, but that 1% is so standout bad it's like wtf.

For example, SLY there in instead of Slither in. And it only did this once! But man

3

u/Cyborg-2077 12h ago

I also tried it for voice-overs for ads, and my approach was to tell opus to generate it for me given a script, and he would do a great job at re-creating the parts where something is misspelled. Even for things like brand names that are not obvious, he'd be able to identify and fix it on its own. Either opus is really good at analyzing the audio, or I got lucky

6

u/MomentJolly3535 18h ago

How does Breeze TTS compare to Omni voice TTS for those who tested both?

5

u/Name835 16h ago

I find Breeze to be a bit better than Omni, albeit slower/bigger ofc.

I've gotten used to 11labs tts from my ai app using days (kindroid phone calls) and even if what I'm about to say is harsh, I think that compared to the proprietary tts models, breeze and omnivoice suck tbh. :D I still use Breeze atm cause it is the best of the bad options, however this is just my spoiled ass talking.

I tolerate Breeze and other open tts models, but am bummed that they are so much behind the competition (almost completely lacking wide sfx tag support for example), like about a year behind proprietary tts models, even 11V3 and ofc not even mentioning v4 which is even better.

I am still local tts only these days, and Im still super excited and really cannot wait for open tts models to get a lot better, hopefully in the near future! :)

1

u/MomentJolly3535 13h ago

Thanks for your answer, i gonna give it a try then !

3

u/KissMyShinyArse 14h ago

Omni supports more languages I think. For English, I strongly prefer Breeze.

2

u/MomentJolly3535 13h ago

Thanks for the answer, i will give it a try for english

2

u/FinBenton 12h ago

I tested breeze when it came out but I went back to Omni after a bit. Its not like omni is better, I think its slightly worse but its faster and uses slightly less VRAM so it fits me better. Breeze has text guidance but the generations became pretty unreliable with that so I didnt find that very useful so for me breeze without guidance wasnt really worth it but I can see it can be better for many.

5

u/Extreme-Way4665 11h ago

Maybe this is a stupid question, you said Local text to speech using Opus 5.5 - isn't Opus not local?

7

u/Cyborg-2077 11h ago

Opus is not local, I send the request to Claude and they return the response, with token streaming. As soon as I receive the first few words, I pass it into breeze, which generates the audio at 4x realtime speed (meaning it can generate 4 seconds of audio in 1 second), and play it immediately.

So the LLM is not local, but the TTS is.

2

u/Raredisarray 14h ago

Nice setup I need to build this out asap

2

u/stoppableDissolution 17h ago

What are you using for the avatar? Some vtuber tech?

1

u/Usual-Buffalo-1791 16h ago

Is it better than dots.tts and Qwen3-tts?

1

u/hiepxanh 15h ago

Can you share tech spec? Which repo or library to use and how you combine it?

1

u/rorowhat 15h ago

What software are you using for the avatar?

2

u/Cyborg-2077 15h ago

Vibecoded with html css

1

u/Any_Door7384 14h ago

and the assets?

1

u/Cyborg-2077 14h ago

Bought online from a website selling vtuber models. Forgot the name tho. You can also find some free ones online

2

u/Vivid-Snow-2089 14h ago

you can just ask opus 5.5 to draw the assets with math now

1

u/Cyborg-2077 14h ago

Hard to believe it would be able to achieve this level of quality. But happy to be proven wrong

1

u/Any_Door7384 14h ago

And then you got the vtuber model and made ai make a html css page using the vtuber model? pretty cool

1

u/Cyborg-2077 13h ago

Yep, pretty straightforward

1

u/Maligx 15h ago

set up guide?

1

u/geldonyetich 14h ago

When I put down the dosh for my AI box a year ago, I was thinking it'd be cool to have a hologram JARVIS agent. Backends like Hermes have the logic covered. Solutions like yours have successfully integrated the front end. Still waiting for that consumer spatial holographic projector. I'd settle for AR glasses but even those are slow to launch.

1

u/Any_Door7384 14h ago

Looks great. how did you do the avatar animations and art though?

1

u/Cyborg-2077 14h ago

Bought a vtuber avatar online

1

u/k8-bit 14h ago

I was using breeze for a while for DJ personas, but found it would too often blether absolute nonsense even with clean audio refs and transcriptions. Fish S2 pro was more consistent, but even it would sometimes laugh maniacally for like a minute.

1

u/KissMyShinyArse 14h ago

It's 95% there. I recently converted a ~125k-word book into an audiobook. BreezeTTS 2 occasionally omits a word or phrase, reads something out of sequence, or pronounces the same word differently in different places. It doesn't happen often, though. I think that with a good AI pipeline, you can produce production-quality audiobooks, even with multiple voices.

I also tried BreezeTTS 2 VoiceDesign, and it's great.

1

u/VoiceApprehensive893 transformers 12h ago

how did you make the avatar?

0

u/Cyborg-2077 12h ago

Bought it from a creator, it's an avatar rigged for vtubers

1

u/aersel24 3h ago

What stt do you use? I feel most don’t understand what I say. Maybe I don’t speak English properly.

1

u/caetydid llama.cpp 1h ago

cute!

1

u/youcloudsofdoom 16h ago

I'm using breeze (and love it) for voice cloning and direction for a military sim I'm making, but this post has reminded me that I am in the non-goon minority for voice clone users

1

u/devdevgoat 18h ago

Nice, what are you running it on gpu wise for that speed? I like the idea for the button!

1

u/Cyborg-2077 16h ago

Rtx 4080 super, but I think could fit even on smaller ones. The tts model takes about 4 gigs of VRAM

1

u/p_tttttri 16h ago

Is there any possibility of open sourcing this setup

1

u/Cyborg-2077 15h ago

I can't really open-source it as it is, because the avatar is a paid model I bought, and cannot re-distribute it

1

u/Open-Adhesiveness-86 12h ago

the push-to-talk remote is probably doing a lot of the work on that 500ms number, since you skip the 500-800ms of VAD silence-wait that eats most "realtime" voice stacks. watch out though, those camera shutter BLE buttons tend to sleep after a couple minutes idle, so the first press just wakes the link and gets swallowed.

1

u/Cyborg-2077 12h ago

Yep, exactly that. I find push to talk to be perfect for this kind of "ai assistant".

I tried live models from both elevenlabs, and openai, and it's not great. Higher delay, you pay by the minute, they are stupid, and as soon as you stop talking, they start replying even if you're not done. So you have to add "hmmms" every time you need to take some time to think.

2

u/Open-Adhesiveness-86 11h ago

The latency floor on cloud voice APIs just kills the conversational feel, especially when pricing punishes pauses. Local PTT with a quantized model on-device feels way closer to natural walkie-talkie flow.

1

u/Cyborg-2077 11h ago

I might try with a local model to see how far I can bring the latency down. Need something that fits in my remaining ~4gigs of ram. Any suggestion?

2

u/Open-Adhesiveness-86 11h ago

Llama-3.2-3B or Qwen2.5-3B at Q4_K_M easily squeeze into under 2.5GB RAM with a 2k context window, leaving headroom for system and audio buffers. If you want even lower TTFT for conversational turns, SmolLM2-1.7B or Qwen2.5-1.5B respond almost instantly on CPU.

0

u/brock_bowers 13h ago

Degeneracy

-5

u/Trustest-S 18h ago

this local voice setup beats most chatbot apps for companions, the quick response makes roleplay feel way more real than the laggy ones ive tried.