I'm actually working on a voice controlled, tldraw canvas based UI – and I'm a designer. So I feel quite seen by this article.
For my app, I'm trying to visualise and express the 'context' between the user and the AI assistant. The context can be quite complex! We've got quite a challenge to help humans keep up with reasoning and realtime models speed/accuracy.
Having a voice input and output (in the form of an optional text to speech) ups the throughput on understanding and updating the context. The canvas is useful for the user to apply spatial understanding, given that users can screen share with the assistant, you can even transfer understanding that way too.
I'm not reaching for the future, I'm solving a real pain point of a user now.
You can see a demo of it in action here -> https://x.com/ojschwa/status/1901581761827713134