What scraping challenges? From your pricing I can say you are just using other APIs and you are a layer on top of them. Also your docs mix insta and TikTok results
HN user
BetterWhisper
Better Whisper API - https://www.betterwhisperapi.com/
Hey, indeed Whisper can do the transcription of Japanese and even the translation (but only to English). For the best results you need to use the largest model which depending on your hardware might be slow or fast.
Another option is to use something like VideoToTextAI which allows you to transcribe it fast and then translate it into 100+ languages which you can then export the subtitle (SRT) file for
Do you support speaker recognition?
Not a 8 min read as stated in the beginning but nevertheless interesting.
In "The proof as we know" section he states that the dot is a NAND operation
Quote: "the · dot here can be thought of as representing the Nand operation"
Are you running whisper on that same $7 Server?
https://www.videototextai.com/ - an AI transcription, translation, chat with your video/audio platform. We are very close to releasing an update where it is possible to caption any video in any language - perfect for making social media content.
We're currently developing https://www.videototextai.com/ – ChatGPT for video and audio. The idea is to get to an all-in-one video and audio editing/insights platform. We’re actively building out new features to fully realise our vision, and we'd love to get any feedback from HackerNews!
You are allowed to delete any transcription you make and with that we do not keep any copy of the transcripts :) . The cookie banner is there to comply with the EU laws.
If you are looking for something automatic that also allows you to interact with your transcripts chatgpt style then I would recommend https://www.videototextai.com/
https://github.com/Emerge-Lab/gpudrive - the repo
Conclusion references a tweet by Elon to explain the most obvious solution is the most entertaining... Make of it what you want
Does it do speaker recognition/ diarization? Can't see it from the repo readme
Well, do you have the hardware to self-host? People nowadays usually are not okay with 1 week of downloading a torrent like it was in 2005. That is probably how long the embedding will take on your cpu for all the video/picture files?
Reading the notes aloud is a really good solution without having to spend a ton of time on trying to OCR handwriting.
I can recommend https://www.videototextai.com/ for transcribing huge amounts of audio. (Disclaimer, I am the founder of VideoToTextAI)
Developed https://www.videototextai.com/ exactly for this reason as it was quite impossible to search videos otherwise. Also you can copy the transcript into a LLM and ask questions from video content like that.
All of it, I have tested it and want to automate it currently.
Edit: the biggest hurdle previously was no midjourney API. Now that DALLE-3 is released it is good enough and has an API
And what is this unique opportunity to lead in AI?
Still no API...
Literally spent the last hour trying to figure out why my app was failing. Only thing I was seeing from my side is "permission denied".
why can't Google be bothered to put outages in Firebase console... Or allow us to see better logs?
Pretty sure this affected Hacker News login as well
Literally spent the last hour trying to figure out why my app was failing when Google can not even be bothered to put outages in Firebase console...
Pretty sure this affected Hacker News login as well
While it seems YouTube's auto-generated are hit or miss, I wonder if feeding them through an LLM can fix the mistakes and still get the video's idea out of them
Wow, why are they so expensive? Like even the regular whisperAPI by OpenAI is less expensive.
This is also why I decided to create https://www.betterwhisperapi.com/ . I believe most of the companies are charging pretty insane amounts for transcriptions...