https://openai.com/index/mixpanel-incident/
openai is impacted too
HN user
https://openai.com/index/mixpanel-incident/
openai is impacted too
Deep Learning Researcher - Vancouver - Onsite
https://picovoice.ai/careers/deep-learning-researcher/
Picovoice is the only all-in-one on-device voice AI platform with offerings including wake word, speech-to-text, LLM, text-to-speech, and more. All on-device. Runs across mobile, web, desktop, and embedded.
thank you!
prompt engineering is a thing, and it's not a thing that you get on social media posts with emojis or multiple images.
is it a public-facing product?
the technology you're looking for is "on-device" or "local" speech-to-text
https://cloud.google.com/blog/products/ai-machine-learning/s... https://picovoice.ai/platform/cheetah/ https://github.com/mozilla/DeepSpeech
picovoice processes on the device and you can fine-tune the models https://picovoice.ai/platform/cat/
if you're running it locally, they don't and cannot.
if you're using the hosted whisper, they can. however, they don't specifically talk about it.
https://picovoice.ai/careers/applied-speech-scientist/
Picovoice is a deep tech startup founded by engineers and driven by engineers. We are accelerating the transition of voice AI from the cloud to the edge. Why? Privacy, reliability, and the environment. Numerous enterprises, including NASA and Stanford University, are leveraging Picovoice technology in their products.
In this role, you get to work on
Applied research to improve the accuracy and runtime efficiency of on-device voice recognition (https://picovoice.ai/blog/local-speech-to-text-with-cloud-le..., https://picovoice.ai/blog/end-to-end-intent-inference-from-s..., https://picovoice.ai/blog/direct-speech-indexing/)
Build Voice AI models from scratch for upcoming products
Improve existing Voice AI models: Leopard & Cheetah Speech-to-Text, Octopus Speech-to-Index, Rhino Speech-to-Intent, Porcupine Wake Word Detection, and Cobra Voice Activity Detection
Develop algorithms for automated data gathering and generation
Canada - Vancouver | On-site
I didnt disclose as I was not gonna promote anything, but I work for a startup specializing in on-device voice recognition. I am 100% biased towards on-device voice processing :)
i just wanted to share my 2 cents, as it's not unique to Apple and the cloud can be costly even if you own it. Big tech has been investing in on-device for a while. besides voice commands, apple and google do transcription locally too. because now you can have local speech to text with cloud level accuracy and of all the reasons you shared - cost, privacy, latency etc. (but again, i'm biased)
https://www.androidauthority.com/voice-typing-opinion-322134... https://techcrunch.com/2022/05/17/apple-adds-live-captions-t...
Amazon has been working on on-device voice for a while. Actually everyone is trying to do that. Running large speech models in the cloud is expensive, considering the number of devices, they probably need more than "surplus" :)
https://www.amazon.science/blog/on-device-speech-processing-...
Tutorial on adding subtitles to videos using the Picovoice Leopard Speech-to-Text Python SDK.
1. Install the software 2. Run the code to convert video to text 3. Create an SRT (SubRip subtitle) file
Example: Adding subtitles to a video on YouTube with pytube.
This tutorial requires Picovoice Console account- free to create, allows free transcription for up to 100 hours/month. Processes voice data locally on the device.
here's the accuracy and cost comparisons: https://picovoice.ai/platform/cat/
Used Express.js, a popular framework for Node.js web apps, and Picovoice Leopard Speech-to-Text engine, offers fully on-device audio transcription