nice, are you using webrtc aec3 or sth custom?
to your question i did a thorough benchmarking (hence late reply) used 18gb m3 pro macpro
full on-device with gemma 4 e2b is ~1.5s since last utterance. vad silence th is set to 600ms
now, when I connect local hermes agent that uses deepseek v4 flash via their api, latency jumps to 3.5s (this of course includes 600s of silence to even trigger pipeline)
btw, suggestions are throttled to one every 20s by default, it's configurable from the app
afaik cluely is waaaaay over 5s, but it makes sense because their users need lots of ai reasoning to answer interview qs :D different target audience then stagewhisper