Yes. Specifically, the pipeline is text -> phonemizer -> phonemized text -> TTS model -> audio You just have to modify the phonemizer's dictionary.
HN user
ZDisket
3 karma
https://zdisket.github.io/page.html
Contact: nika109021@gmail.com
Posts1
Comments6
Show HN: Real-time local TTS (31M params, 5.6x CPU, voice cloning, ONNX) 4 months ago
No multilingual capabilities yet, although that is planned for next iteration.
Ask HN: What Are You Working On? (March 2026) 5 months ago
I'll explain in detail once I've got the big release, but everything's been thoroughly modernized. Transformer, HiFi-GAN (now iSTFTNet w/Snake) vocoder, et al, plus a few additions.
Ask HN: What Are You Working On? (March 2026) 5 months ago
Multilingual and local? Try out Supertonic 2.
Ask HN: What Are You Working On? (March 2026) 5 months ago
I'm working on a voice cloning version of my TTS model, a highly upgraded VITS:
https://x.com/ZDi____/status/2013655958027669958
Right now, I only have single speaker checkpoints (as per the old video). That will change soon.
Upwork has candidates buy "connects" with real money that are spent when applying to jobs. Ultimately it seems some form of payment is a proven gate.