great idea, we'll do that too. we just decided to launch an onnx first and get some feedback. we'll be simplifying the process of running it everywhere including a command line executable.
HN user
rohan_joshi
got it, this is very useful. thanks a lot.
yes, indeed. we are working on adding mit licensed phonemizers too by this weekend, so you'll be able to use these models as you like :)
hey sorry for this issue, i think its a bug in our preprocessing. let me look into it and help fix it. i think you posted this in our discord so lets carry the conversation there.
for that, a 100KB model could be enough ;)
glad you liked it, thank you so much for the kind words. our team is really good at squeezing performance out of small models. we are working on a new launch and hope to release a technical report along with that which includes details. fyi, our current 14M model is better than our previous 80M model. and we expect this trend to continue.
i can confirm that we did not.
got it, this helps a lot. thanks!
thanks a lot for sharing this, its v helpful for fixing the env issues. we'll fix all of them by the weekend.
thanks a lot for the feedback. glad you liked it. we're gonna be launching more tiny models across use-cases.
thanks for pointing these errors out. we're looking into this and will help fix this.
yeah let me add uv and conda support to make it easier.
damnn, really sorry for the inconv, looks like some folks are having bad env issues. we're working on fixing this.
our next model(eta 3ish weeks) will support Japanese. would love to get your feedback then on how the quality is. can you share what usecase you want? would love to support it.
thanks a lot for helping w this. yes i'll fix this asap.
this is some env issue sorry for the inconvenience, lemme fix it. can you dm me w your env? discord / github / mail / anywhere works.
How is the Bruno voice for this one? there will also be another release in ~15-20 days where we have more professional voices. if you'd like to get early access and give feedback lmk, or dm me.
french, spanish and german models will also be out v soon. these are languages we are working on already. some lower resource languages will take longer.
spanish model will be out in a matter of weeks.
not v long. until then you start running tts on phones, wearables and r pis. at the model level, we'll have a model for this kind of mcu's later this year.
thanks a lot for trying it and giving feedback. custom preprocessing will fix this for 95% of use-cases. and as i mentioned, this will be fixed at the model level in the next release.
yeah we're fixing this at the model level too. but in the meantime, there is a way to add text preprocessing for you, and if you have a special use-cased, claude code should be able to one-shot custom preprocessing. its the way that most existing tts models (including sota cloud ones) deal w numbers and units, they just convert it into string.
damnn, lemme fix it, sorry for that. we may have forgotten to remove the redundant dependencies. i'll comment here once i push the change. thanks a lot for trying it and giving feedback.
thank you so much. glad you liked the model.
it can support chunk streaming, i'm working on adding it to the repo. should be up by tomorrow.
yeah we tried to include those voices in this release to showcase the expressivity. but we've already started adding more professional sounding voices for prod use-cases.
Yes, we've started working on it and will have a range of stt models v soon. lmk if you have a prod use-case in mind?
thanks a lot. yeah these models are way better than our previous launch. our 15M model now is better than our previous 80M model and we expect to continue seeing this rate of improvement.
yes, we just started working on this yesterday haha, great that you mentioned it. once we have it working it'll be out soon.
thanks, glad you liked it