That’s a great question! We partner with a number of different transcription providers that use AI to identify different speakers based on the sound of their voice. This prevents all the speakers from a conference room from being bundled together as the same person. We’re also going to be looking to add this functionality to our own transcription service in the coming months.
HN user
davidgu
Co-founder at https://www.recall.ai/
The postcards were part of a dev-focused campaign to get people curious enough to check us out. We kept it minimal to stand out amongst other mail.
For internal use cases like recording your own meetings into Google Drive, the native tools work fine.
Where we come in is for companies building products that need to support all of their customers across Zoom, Meet, Teams, Webex, etc. Most enterprises don’t want five different integrations, and native APIs often come with restrictions (like only the organizer being able to access the file, or recordings not being available until after the call).
We already support diarization in the Desktop Recording SDK by capturing the meeting platform’s speaker-change events, so you get a diarized transcript plus precise “speaker started talking” timestamps out of the box. We also support voice-signature diarization via third-party STT providers for participants calling in from the same room
For in-person meetings and audio uploads, this is on our roadmap and in development. More to come on this!
Just to clarify, we’re the infra layer that reliably captures and normalizes meeting data across platforms. The real value for users is what developers build on top: automated analysis, enrichment, and workflows (not the capture itself)
Modern LLMs can power sales coaching, medical scribing, legal review, support QA, and compliance reporting but they need consistent inputs to process. We handle capture/formatting/edge cases so developers can focus on models and UX
Thanks and love to hear this!
Amanda says thank you so much!
I actually agree that it’s become incredibly easy to transcribe conversations using open-source models, and that’s not where Recall adds the most value. The hard part is building the infrastructure that allows you to get real-time access to the raw audio, video, and transcript data directly from the meeting platforms. We abstract all of that away and provide you with a clean interface to access that data. Once you get the data, you could use any of the models that you mentioned to do your own transcription, or transcribe using Recall’s transcription models.
Thanks! Really appreciate the kind words
Did you get one? :) This was a part of our Series B raise to help get our name out
Enabling transcription/recordings per platform and remembering to record creates user-dependent setup. Also the host often needs to install apps which adds security friction, and you still have to build/maintain separate implementations for Zoom/Meet/Teams which is often a cost that devs don't want to deal with
Instead, we built a single API that can get the same results without the issues mentioned above so you can focus on building the features your users care about
Usage includes silent time too as we are still processing the media streams
$0.70/hr is our starter rate for low-volume testing. In production, developers will see higher usage and choose to commit to volume and longer-term usage. Because of this, we've seen most teams don’t pay the starter price once they scale beyond early pilots
You're right, and I agree that participants should be aware when they’re being recorded
Because consent laws are complex and vary by region and industry, we leave the consent flow to the developer and we provide the tools and guidance to do it correctly. As with our Meeting Bot API, we also urge teams to follow local laws and make recording clearly visible to users
You might be looking for Recall.ai which is an API to record meetings.
We offer an API to record via desktop app like Granola/Krisp/Amie, and also an API for meeting bots like Grain/Fathom/Fireflies.
Here are docs so you can get a better sense on if this is what you're looking for: https://docs.recall.ai/
At Recall.ai, we built a 10,000-node cluster, processing over 1TB/sec of raw video in real-time.
After a major migration, we faced a strange audio issue that led us on a deep dive through our infrastructure.
The culprit? Not the audio code—but a hidden interaction with AWS’s virtual serial ports.
We wrote about our journey discovering the artifacts and finding a clean fix!
If you're looking to build an AI notetaker for medical use-cases (or others), https://www.recall.ai is a HIPAA compliant API to capture conversations from video conferences.
P.S. I'm the founder, so obviously biased :)
I'm obviously biased as the founder, but check out https://recall.ai/ for a meeting bot API.
We serve over 300 companies, process millions of hours of recording a year, and are officially partnered with Zoom.
Writing action items during a meeting is super distracting (plus mechanical keyboards are insanely loud). A tool like this is definitely something I need!
Would Taro be a good fit for a startup CTO looking to improve their management skills, or is it mostly targeted towards ICs at larger companies?
Really enjoyed the post Tom. I really like the structure of introducing a potential solution then pointing out the problems with it.
P.S. Svix is great, super happy customer here :)
Super cool product, congrats on the launch Maitham :) Good to hear that all the data is encrypted & anonymised, since that would be my #1 concern with health data. (Evervault also looks interesting, I didn't know there was a plug-and-play way to do this!)
I think it heavily depends on the kind of work you're doing.
I've used i3wm for the last 4 years, and it's very helpful for when you have a bunch of short-lived windows (like terminals) or you frequently switch between two different sets of windows (like docs/code, or code/web-app).
Perfect Recall | Full Stack Engineer (first hire) | Full Time | Onsite Waterloo, Canada | https://www.perfectrecall.app
We let you share key highlights from your Zoom calls, replacing written notes. Our software records and transcribes Zoom calls. Users can then highlight transcript text to create short, consumable clips.
We’re making our first engineer hire. You'll have a major role in developing the core product and creating the culture of our engineering team. You'll be working closely with the founders. You'll have full ownership over major features - frontend and backend, from prototype to production.
Our frontend is written in React, with Typescript. Our backend is Django and Celery deployed with Terraform on AWS. We use Postgres and Redis. Experience with our stack is not necessary.
Engineering problems we’re solving: Recording Zoom calls at scale, producing dynamic media at near-real time performance.
Video calls today are an inferior substitute to in-person interactions; low-resolution, high-latency, and without additional capabilities. We believe that the potentials of video as a communications medium have just barely been explored, just as the first movies were essentially recorded stage-plays. We see a future where video calls are more productive than face-to-face, because of the value the intervening software delivers.
In terms of our company culture, we believe that raw hours make a difference. Working too many hours doesn't guarantee success, but working too few leads to failure. We expect everyone to put in their best effort every day, but we generally don't expect you to work late evenings or on weekends. We are careful to avoid burnout, but we also don't want to sugarcoat the fact that raw hours can make a huge difference in a startup.
If what we’re doing sounds interesting, apply below or email david@perfectrecall.app
https://www.ycombinator.com/companies/perfect-recall/jobs/kc...
and maintains really good performance even with large documents
Out of curiosity how large were these documents, and were they completely text or did they contain a fair number of images?
I'm asking since I once ran into a situation with a long and image-filled Word document (150 ish pages, maybe 50 high res images) where there was a noticeable lag when typing and saving would take 30s.
I am very seldom in a situation where I must download my IDE while traveling in a rural area, run it with Photoshop simultaneously, or use it for the entirety of a 3-hour plane ride.
While multi-hundred-megabyte text editors consuming double digit CPU to render some text are definitely a sign of inefficiencies _somewhere_, I value any marginal productivity benefits from these additional features over (possibly significant!) usability in very resource constrained situations.
Hi HN,
After our last post (https://news.ycombinator.com/item?id=22415714), we listened to your feedback and made a ton of changes to the UI to make it more intuitive.
Copying the description of the project from our last post:
We found that it was difficult for people who wanted to contribute to open source to actually start contributing since searching and finding relevant existing documentation related to the code itself was tough (which is important for new contributors, since they don't really know the codebase).
We built Hyperdoc, a documentation tool to help solve this problem and used it to write a code walkthrough guide for the popular open source software Socket.io to show how it works.
You can use the links in Hyperdoc to step through the guide.
If you want to contribute to Socket.io and you found this guide useful/insightful, we'd love to hear about it. Also, if you want to see a guide like this for one of your favourite open source projects or have any feedback, please shoot us an email at hello(at)gethyperdoc(dot)com.
Gmail is using their influence to modify email standards[0]. This is exactly how it works in reality. If a certain service provider holds most of the user base, they can implement breaking changes without fracturing the community. In some cases it's a net positive, helping protocols evolve, and in some cases it's a net negative.
[0]https://www.theverge.com/2018/2/13/17007100/google-amp-gmail...
I question the methodology behind this figure. The nature of the decentralized network makes it difficult to get good estimates. All I have are anecdotes. Yes, the overwhelming majority of users are on Mastodon. However, I personally interact with several users on several different Pleroma instances and a couple on GNUSocial instances, and I did not seek them out for this purpose.
However they are the exception. Because the overwhelming majority of users are on Mastodon, the Mastodon developers have the power to modify protocol implementations. If it was closer to an even split, any breaking changes would fracture the community, while in this case some smaller groups may be lost.
But not federation. No one can talk to people on other servers. Such forks already exist, none of them are successful.
They are not successful because Signal has not yet done anything particularly egregious. If they do, I believe that a fork would quickly gain popularity.
I know of many people on both of the platforms I mentioned. Collaboration between all of these platforms is frequent and they drive improvements in each other.
We've both given anecdotes; user counts would be more conclusive. I've found +1 million [0] for Mastodon. Do you know how many users GNU Social and Pleroma have? I can't find these numbers from a quick Google search, but they might be somewhere.
[https://en.wikipedia.org/wiki/Mastodon_(software)#Adoption]
Moxie toes the line so that Signal's strengths always outweigh its faults for most people, which prevents an alternative from reaching a critical mass.
Signal is released under the GPL, meaning that if something did happen to shake user faith (data breach, data mining) a fork pointing at different servers could be instantly created with complete feature parity.
Every traditional barrier to user migration is removed. Features are the same. UX is the same. Compatibility is the same.
That's a precarious line to toe.
Also, Mastodon is built on open standards that have several competing and compatible implementations (GNUSocial and Pleorma are the main two), which are run by their own maintainers.
Yes, Mastodon is inter-operable with OStatus and ActivityPub. However I do not know of any significant user base that interacts with Mastodon through an alternative platform. Given the will, Mastodon could implement its own extensions to ActivityPub and break compatibility, with most users unaffected due to the strength behind the core development team.