nw_wrld is an event-driven sequencer for triggering visuals using web technologies. It enables users to scale up audiovisual compositions for prototyping, demos, exhibitions, and live performances. Users code their own visual modules, then orchestrate them using the project's native UI composer.
SantaBench, a fun benchmark with a serious methodology. The task: play a cheeky Santa agent who researches users online and roasts them based on their social media.
Wasn't meant to imply the opposite. The video even has a watermark clearly saying it's generated. I genuinely found the video useful, so decided to share.
XTTS model release (Text-to-Speech and voice cloning)
# From the release notes:
This model is trained on top of XTTS v1, using output masking. We mask the part of the output that is used as the audio prompt while training and don't compute loss for that segment. This helps us to resolve the hallucination issue that V1 experienced.
- Add Japanese
- Resolve the hallucination issue (repeating the audio prompt)
- Increased expressivity
- Added ne_hifigan that was trained without denoising that brought some EQ and compression profile that might be unwanted for some use-cases
Prompt-to-voice creates a new, unique voice given a text prompt. Similar to stable diffusion, but for voices. Then those voices can generate speech via TTS.
imo this kind of tech is useful to supplement the artist, not replace them. You listen to a singer because you know their voice more than anything. Bob Dylan was a pretty bad singer, but I listen to him because of some emotional connection.
same company, yes, but major updates since our demo:
1) Quality: the underlying model is much, much better, and so is the quality of the Voice Clone.
2) Productization: previously we just had a stand-alone demo, now we've launched the product (user accounts, multiple voices, etc.)