Just made https://feycher.com thats similar, but has realtime lip syncing as well. Let me know if you are interested and we can chat
HN user
userhacker
A new age of empires game or any top down real-time strategy game.
I'm the creator of Revoldiv.com, We do speaker diarization and transcription at the same time. Give it a try.
Try to upload it on https://revoldiv.com/ we pre-process the file to make it a little Intelligible and you can supply your context when uploading.
Good point but the problem with local hosting is that if you want to use the larger models it will take a long time to transcribe a file. We use multiple gpus and we do speaker detection, sound detection and it is has a rich audio editor.
If you want a quick and free web transcription and editor tool, We've built https://revoldiv.com/ with speaker detection and timestamps. Takes less than a minute to transcribe 1 hour long video/audio
On the end user side Revoldiv.com lets you pick any podcast you want and transcribe it
Nice product! Any integration planned for jetbrain ides?
contact us at team@revoldiv.com and we are offering an API on a case by case basis
I suggest you give revoldiv.com a try, We use whisper and other models together. You can upload very large files and get an hour long file transcription in less than 30 seconds. We use intelligent chunking so that the model doesn't lose context. We are looking to increase the limit even more in the coming weeks. It's also free to transcribe any video/audio with word level timestamps.
Can you send me the audio that caused it, you can email me at team AT revoldiv .com. If there is going to be a lot of interest, yes we can provide it as an api service. Our service has some niceties like word level timestamp, paragraph separation, sound detection etc... for now it is a free service you can use as much as you want
For revoldiv.com we have profiled, many gpus, the best one is 4090. We do a lot of intelligent chunking and detect word boundaries and run the model in parallel in multiple gpus and we get about 40 to 50 seconds for an hour long audio but without expect 7 minutes for an hour long audio on tesla t4
on tesla-t4-30gb-memory-8vcpu google cloud
on tiny and tiny.en
for 10 minute = 30 seconds
on medium
for 10 minute = 1m 30s
for 60 minute = 7m
on large
for 60 miutes = 13m
on NVIDIA GeForce RTX 4090
on tiny
for 10-minute = 5.5 seconds
for 60-minute = 35 seconds
on base
for 10-minute = 7 seconds
for 60-minute = 50 seconds
on small
for 10-minute = 14 seconds
for 60-minute = 1 min 35 sec
on medium
for 10-minute = 26 seconds
for 60-minute = 3 mins
on large
for 10-minute = 40 seconds
for 60-minute = 3 min 54 secIt's not a model you can run on your own server but a free service on revoldiv.com. You can expect 40 to 50 second wait time to transcribe an hour long video/audio. We combine whisper with our model to get word level timestamps, paragraph separation and sound detections like laughter, music etc... We recently added very basic podcast search and transcription.
https://modal-labs-whisper-pod-transcriber-fastapi-app.modal...
Interesting, which model are you using? We use the medium model which is the sweet spot between time/performance ratio. We also chunk, We try to detect words and silences to do better chunking at word boundaries but if you do more chunking and you don't get the word boundaries right it seems like whisper loses some context and the accuracy suffers. We will soon support longer hours. We just want to make sure the wait time for transcription doesn't suffer for most users. But great demo, reach out to me if you want to collaborate
I recently swapped out the AI model for voice transcription on revoldiv.com and replaced it with Whisper. The results have been truly impressive - even the smaller models outperform and generalize better than any other options on the market. If you want to give it a try, our model is capable of faster transcription by utilizing multiple GPUs and some other enhancements, and it is all free
I created revoldiv.com. It's privacy focused and login is not required to transcribe. You can record your meeting and upload the video or audio to transcribe it
You can use revoldiv.com. If you go to export and choose audiogram, it will convert the audio/video you uploaded to text and create an audiogram.
Thanks yes it is, we implemented all the ai models in house, that cuts our cost.
You can use revoldiv.com to cut out filler words or any words of your choosing, after you upload your file and it finishes sound detection, you can click on the search box to bring up the toolbar to delete sounds
Hey IndySun, We are in the process of writing one, but we store the file to transcribe and it gets deleted automatically after a certain amount of time has elapsed (this varies on the server configuration, but in a short amount of time). No human accesses the audio or text transcription you have uploaded
The reason for capping it at an hour is, since the service is free we want to make the experience fast for everybody. We are gauging how our service is going to be used and allocating resources accordingly. We will allow longer transcriptions in the near future.
I'll hangout here to answer any questions
Pretty cool application, we are also working on a similar tool at https://revoldiv.com/
https://chrome.google.com/webstore/detail/text-to-speech/bkj... has a nice text to speech software, something like what the edge browser has
revoldiv.com has a similar feature set
If you are on a mac, I created a simple web wrapper for the google voice website and puts it on the menubar, you can SMS on the desktop that way. Link to the app voicenotifies.com
unless you change your holdings into fiat thats not the case. You can rebalance with exchanging among your coins
as a hack, you can use a regular gmail address and forward all your emails to that address. you can also respond from a regular gmail address to your abc@yourname.com by changing the from field (you have to verify you own the email address). currently using a similar setup, custom email address but responding from a regular gmail address
yup Atlanta is slow for me too
I second Readability, it works great for article heavy webpages. I used it to build a reading time estimator for chrome https://chrome.google.com/webstore/detail/read-time/nccohhim... and its open source https://github.com/usergit/read-time bonus, you can click on the extension to show only the main content of the page