Yes, the SOTA is currently much more advanced.
HN user
narrationbox
We do generative ML, specializing in speech technology.
https://narrationbox.com
Cost effective voiceovers and narrations at scale. Voices for agentic AI and audiobooks.
Check out our blog here: https://narrationbox.com/blog
Follow us at: https://twitter.com/narrationbox
Subscribe to our newsletter for generative text to speech, voice cloning, and accent conversion research.
Give us a try, I think we are what you are looking for
A lot of our customers use us [0] for that, it works pretty well if executed properly. The voiceovers work best as inserts into an existing podcast. If you see the articles of major news orgs like NYT, they often have a (usually) machine narrated voiceover.
Yeah, neural codecs are pretty amazing. The most incredible part is that they can do compression well across the temporal domain, something which has been non-trivial.
Their sentence segmentation heuristics were not configured correctly. It's not an inherent limitation of the technology itself.
The newer transformer based generators are a bit better in this regard (since they can maintain a longer context window, not just in short tiny snippets).
Mel + multispeaker vocoder is very much a classic (tacotron era) TTS approach
Since it does the signal processing in the Fourier domain, does this suffer from audio artefacts e.g. hissing in the output? Torch's inverse STFT uses Griffin-Lim which is probabilistic and if you don't train it sufficiently, you may sometimes get noise in the output.
https://pytorch.org/docs/stable/generated/torch.istft.html#t...
An alternative would be to use a vocoder network (or just target a neural speech codec like SoundStream).
Recognizes "human" and recognizes "desk". I sit on desk. Does AI mark it as a desk or as a chair?
Not an issue if the image segmentation is advanced enough. You can train the model to understand "human sitting". It may not generalize to other animals sitting but human action recognition is perfectly possible right now.
Your average mobile processor doesn't have anywhere near enough processing power to run a state of the art text to speech network in real-time. Most text to speech on mobile hardware are stream from the cloud.
For the high end stuff no, but many of the lower tier jobs are under threat.
We used to be in this field too (https://kloudtrader.com/narwhal). It is a very crowded market and monetisation is tricky.
You can throw this together pretty quickly using one of the AutoML APIs.
So if I am launching a Spotify/Audible-style music/audiobook streaming platform, how will the pricing work out? Do we pay you instead of the original author for any user uploaded content? If the original author chooses to upload their content onto our platform, do we scan it with your API and explicitly whitelist it and pay them directly?
Your company looks very cool btw. What's your email?
What about TransferWise ewallets?
Do you have any public pricing or startup plans?
Does the accounting system support Canada and other commonwealth countries? Or is this US only?
Some ETFs also allow easy investment in stocks of foreign countries without requiring additional brokerages on behalf of the end user. I presume this does not offer that functionality?
I will look into it, Wiki2SSML looks very handy.
It looks great, have you considered adding a visual editor?
We have one for our systems: https://narrationbox.com
Are there any free plans or discounts for HN users?
We never expected audio to be this popular either when we built Narration Box :)
Lua is a great language, but the fragmentation between versions makes it difficult to be adopted outside of niche embedded areas. On a side note, we maintain a Lua newsletter but haven't had time to update it. This article looks perfect for a new edition.
You can try locking down the app, it's not ideal but it is better than nothing:
A while back I wrote about this
https://medium.com/@kloudtrader/reducing-whatsapp-digital-fo...
Not sure if it still applies to the latest version of Android and WhatsApp but it might help. However it only mitigates certain real-time tracking and contact discovery, not to mention switching profiles is somewhat of a hassle.
This looks great! Congrats on launching. We used to be in this space too. The most tricky part is performance, especially if you are backtesting against an algorithm instead of manual trading. Having to wait for the results on 30 years of trading data can get rather annoying at times. If you do implement support for algorithmic trading, it might be helpful to rewrite the core in WebAssembly.
Interesting, it seems their products are geo-locked. The Canadian site still shows "Request invite".
Is this like Stripe Issuing where it is only available to a small number of companies? A lot of newer Stripe products seems to be available to large enterprises only. We applied for that waitlist multiple times but never heard back. What are the revenue or scale requirements for your targeted customers? Is it suitable for fintech e-wallets/banks?
https://rete.js.org is also a good choice.
When evaluating end-to-end human narration services, the most important factor is audio mastering. For services that target the podcast industry, good quality audio output that can be used as-is is a must.
Looks incredible! Will definitely try it out.