Indeed. Right now I think our open choices are Piper, Kokoro and Orpheus.

GaggiX · 2025-03-20T18:06:16 1742493976

He was talking about STT models, not TTS. Whisper is open source and a good solution in many cases (in particular finetuned ones).

pzo · 2025-03-20T19:50:07 1742500207

regarding STT we got also today 2 new models from Nvidia:

https://huggingface.co/nvidia/canary-180m-flash

https://huggingface.co/nvidia/canary-1b-flash

second in Open ASR leaderboard https://huggingface.co/spaces/hf-audio/open_asr_leaderboard

Sadly only supports 4 languages (english, german, spanish, french)

DrPhish · 2025-03-20T18:05:10 1742493910

In my opinion GPT-SoVITS is the best if you can put in the effort. I'm still using v2 since the output is so good. Its also the best multilingual one in my testing on Japanese inputs.

nickthegreek · 2025-03-20T20:03:46 1742501026

hadnt messed with that one before. my needs are more real time for voice assistant but was neat to play with on hugginface.

https://huggingface.co/spaces/lj1995/GPT-SoVITS-v2

pzo · 2025-03-20T19:57:48 1742500668

can it support more languages rather than only English, Chinese, Japanese, Korean?