HN user

shade

278 karma
Posts0
Comments129
View on HN
No posts found.
Transcribe.cpp 4 days ago

I have not formally published it, but it's open source: https://github.com/edmistond/larmindon and https://github.com/edmistond/larmindon-core - you'll want to clone them into the same root directory.

Right now it only supports languages supported by parakeet-rs and Nemotron (so... English only as far as I'm aware) and you'll need the ONNX version of Nemotron: https://huggingface.co/altunenes/parakeet-rs/tree/main/nemot...

The first run experience isn't great, you'll need to download all the files from the model, start the app, and then go to settings and configure the model directory. It runs well on Mac and Windows; I haven't tested it on Linux in a couple of months since my Linux install is out of commission currently.

Transcribe.cpp 4 days ago

Nice - I'm definitely going to take a look at this. I've built my own cross-platform (Mac/Win/Linux) live captioning app on top of Nemotron, and it works well but dealing with ONNX is kind of annoying. With this having Rust support (I built it on Rust/Tauri) it should be a pretty solid candidate; I'll have to see if I can find a Silero VAD implementation that doesn't depend on ONNX, or maybe I'll see if the clankers can migrate it for me.

I have a 2023 Crosstrek, my wife has a '21 Ascent. I have the same habit you do - edging away from large trucks slightly - and both of them do the same thing you described to me.

It's essentially that Subaru's lane system actually has two levels: it has lane keeping where it's just trying to keep you inside the lines, and then on top of that it also has lane centering which is pretty much what it says.

Just a note for you or anyone reading who has a recent Subaru and doesn't know already: if you find the centering really bothersome, you should be able to be able to go into the settings on the instrument cluster display (up/down arrows at the lower left behind the wheel, toggle it until you get to the "hold for settings" option), find the Eyesight settings, and turn off lane centering. It will still try to keep you inside the lane markers but won't try to park you right in the center of the lane. In that mode, it's more like the Honda Sensing system I had on my 2016 Civic.

I go back and forth a bit on it but mostly keep it in lane centering mode now - I've gotten used to how it positions the car in the lane, and it lets me focus more on what's going on around me than micromanaging lane position and such.

I'm deaf, so I test a lot of speech to text and transcription apps from an accessibility point of view.

My answer to "why have a monthly subscription" would be that you need capabilities that Whisper doesn't handle well, like real-time transcription in noisy environments.

That's not the niche you're targeting here, though. :)

My experience is that Whisper - not being built for real time speech to text - isn't as good at it as other tools are. You can hack something together by stacking together progressively more audio frames to feed to Whisper to give it context, but IME, you're going to get better results from a model that's designed for real-time STT in the first place, or by using a service like Azure Speech to Text which has excellent noise resilience... but which is also an ongoing cost which would justify a subscription. Real-time Whisper also devours your battery quickly.

That said - while I've had very good experiences with Parakeet in MacWhisper, I'm curious if you evaluated Apple's SpeechAnalyzer APIs at all. It's unfortunately limited macOS/iOS/iPadOS 26+ since it's a new API, but it's on device, has comparable quality of results to Whisper Large v3 Turbo and Parakeet, and seems to be better on battery usage.

Yep, I'm also deaf (since age 6), went through a lot of speech therapy, and have a very pronounced deaf accent. I live in the midwestern US (specifically, Ohio) and at least once a year I get asked where I'm from - England being the most common guess, but I've also had folks ask if I'm Scottish or Australian.

AI struggles massively with my accent. I've gotten the best results out of Whisper Large v2 and even that is only perhaps 60% accurate. It's been on my todo list to experiment with using LLMs to try to clean it up further - mostly so I can do things like dictate blog post outlines to my phone on long car rides - but I haven't had as much time as I'd like to mess around with it.

Meta Ray-Ban Display 10 months ago

Yeah, I've been deaf for over 40 years now and captioning glasses are something that I've wanted ever since I was a kid. I'm not a particularly big fan of Meta and I have some serious reservations around privacy that need to be satisfied, but at the same time it's really exciting to see this going from "pie in the sky thing I dreamed about having when I was ten" to "actual existing product."

There's a few other companies/startups working on this too, but a lot of the glasses they're producing are very ugly. There's a couple that didn't look bad, but from what I'm seeing Meta's are a combination of the best-looking ones and best display so far, and I'll be very curious to see the reviews.

One of my weird hobbies is radar chasing storms, and all of that stuff is completely normal. NEXRAD is very sensitive, especially when it's in clear air mode (it has different modes depending on if it's raining in the area) and can pick up things like dust, birds, bats, and insects. There's also ground clutter from things like buildings, wind farms, and even cars.

The National Weather Service has a good brief explainer: https://www.weather.gov/iwx/wsr_88d

They also have an interesting PDF covering some of the more unique signatures you might see, though it's not exhaustive: https://www.weather.gov/media/btv/research/Radar%20Artifacts...

In the last 10 years has technology actually made my life better?

In my case? Yes, absolutely. Automatic speech to text is now cheap or free, ubiquitous across most platforms (even Linux!), and generally very effective. Total game changer to my ability to participate in meetings at work and in society generally.

I would say it's cool in the sense that building anything is cool, but I find myself mostly in agreement with your take, although with a caveat.

I can't find the quote now, but someone (I think simonw?) said that they feel a bit of an obligation to spend at least as much time working on writing something as it would take to read it, and I agree with that... if you want me to spend time reading your post, I'd like to know you actually made an effort on it.

For me, writing is thinking, and helps me refine my thinking, so I don't use AI to assist writing process. I agree with the comments that AI writing tends to have a specific voice, and I don't care for that voice and don't want my writing to come across that way.

Where I do find it useful in writing, however, is as an editing pass in an advisory role. I don't ask it to rewrite anything for me, but I will ask it to double-check for excessive passive voice, tone, does it raise unanswered points, etc. I typically write my draft posts in Zed, and use Zed's AI chat panel to throw a request at Claude. The big thing though is not blindly accepting every suggestion the AI makes - I read them, think about it, and sometimes adjust the post based on that feedback. It's a useful sanity checking step and while a real human editor would be preferable, I can't justify the cost to hire an editor for my little blog that probably gets zero hits most days. :)

Yup. I grew up in Hancock Co and used to ride occasionally with the bike club there, had a couple of days when the ride out was brutal because of persistent, endless wind, but then the ride back was awesome for the same reason. :)

Yep, I'm in the exact same situation as you.

The tools for in-person are getting better, but aren't frictionless to set up and sometimes require you to spend time futzing with getting your iPad or iPhone to actually see an external microphone. I don't know if Android is better about this or not, unfortunately. I would _hope_ that interviewers would extend people a bit of grace about this, but who knows.

As an aside - I saw your post on Apple Live Captions, and completely agree with you. I've been slowly adding to a collection of reviews of various captioning tools, and was _very_ critical of some of the choices Apple made there.

M4 MacBook Pro 2 years ago

I have the OG 13" MBP M1, and it's been great; I only have two real reasons I'm considering jumping to the 14" MBP M4 Pro finally:

- More RAM, primarily for local LLM usage through Ollama (a bit more overhead for bigger models would be nice)

- A bit niche, but I often run multiple external displays. DisplayLink works fine for this, but I also use live captions heavily and Apple's live captions don't work when any form of screen sharing/recording is enabled... which is how Displaylink works. :(

Not quite sold yet, but definitely thinking about it.

Best retro I ever had, on a small team of seniors: we all sat down, looked at each other, agreed that a sprint happened and we couldn't think of anything that was good or bad about it. Then we called in our manager, who also acted as scrum master, so we could do planning for our next sprint. I thought this was reasonable enough - ostensibly we'd do retro and then planning back to back, and none of us minded the chance to take a few minutes and reflect if we had anything we should discuss.

By contrast, I've worked with scrum masters who were strict about the process and insisted _every_ retro needed to have at least one improvmenet or action item out of it, preferably more. I found this pointless and I've rarely seen them actually followed up on.

Yup - I've done double-ended safety razors, I've tried the Gillette and Harry's disposable ones, but in 30+ years of shaving what I keep coming back to is Braun electric razors. So far, they're the only option I've found that doesn't leave me with razor burn.

Currently using a ten-year-old Braun Series 7. I can get away with infrequent shaves (~3x a week) since my beard doesn't grow very quickly, so I replace the foil and cutter heads every 2-3 years. It's not as cheap as a double-ended, but for me, it's a better shave, worth the cost, and cheaper/less waste than disposables.

If I upgrade anytime soon, it will probably be to something that's designed to be used in the shower - I do miss a nice shaving cream sometimes, and unless I have a bottle of 'Lectric Shave handy, I don't usually like to shave right after showering.

Yes, I was on an internal project recently that wanted to use LLMs in a way that was appropriate to evaluate if changes between two versions of a text were semantically meaningful, and limited to that scope, it would've been a really valuable tool.

We had a directive from management to, for political reasons, use AI in the tool as much as possible to show how innovative and forward-thinking the company is. This led to a bunch of poorly-thought-out choices and while the project is in production and has internal users... I don't think it was particularly successful.

Not all of that is due to the "use AI" directive; there were also poor technology and deployment stack choices that made things overly complicated and cost us a bunch of time.

I did something similar recently with my daughter's school calendar and ChatGPT with gpt-4o; her school has a ton of closures/teacher work days, I fed the PDF of the calendar in and asked it for all the dates that impacted 2nd grade, then asked it to create an icalendar file for them.

Oddly, it didn't want to create the actual file, but gave me a Python script for doing so; you just need to be sure to tell it what time zone you are working in and that you want it to be an all-day event, or you'll get suboptimal results.

Speaking as someone who's deaf and uses these services a lot: for speech to text, the AI stuff is getting rather good.

I'm not saying it's perfect for every situation, but I have a very high success rate using InnoCaption[0] for captioned phone calls, including to places like restaurants with a lot of noise going on in the background. InnoCaption does both live person and AI-based captioning; since they started offering the AI-based option I've left that on, and I've never had to switch to human operators to continue a conversation.

That said - I'm not deaf from birth (lost my hearing in elementary school), so I voice for myself and that does simplify the process. I have used the old school text-only relay services and that was always such a miserable experience for me that I would crawl over broken glass to avoid making phone calls, especially going through phone trees. That's one area that relay operators still have a major advantage on. IIRC, Google's Pixel phones are supposed to be able to navigate phone trees for you, but since I use iOS I have no personal experience there.

[0] https://www.innocaption.com/

As I recall, this was more or less the concept behind Brilliant Pebbles [0], except Starship makes it cost-effective to launch.

I'm not going to argue whether building it is a good idea, but it also seems like Starship has the potential to make launching a kinetic bombardment system [1] possible given the large payload capacity.

[0] https://en.wikipedia.org/wiki/Brilliant_Pebbles#Brilliant_Pe... [1] https://en.wikipedia.org/wiki/Kinetic_bombardment#2003_Unite...

Yeah, I agree with this take. For web and console apps, C#/dotnet is a great choice and should continue to be. Blazor, I think, is also fine (for some use cases, it's situational) and I think it's at the point it'll achieve liftoff.

I've been playing around with MAUI the past few months and it... isn't great. Desktop feels like (and honestly, kinda is) an afterthought, and the documentation is sparse. I spent several hours fighting with a couple of native layout controls trying to get them to work, before giving up and implementing everything I wanted to do in half an hour in a web view.

NSubstitute is good, I used it at a previous job.

I've favored Moq in the past because I think there are a couple of things it makes a bit easier or is a bit less opinionated about, but NSub is perfectly cromulent as well.

Someone posted a quick guide to migrating a bunch of it easily in one of the issues in the Moq repo discussing this whole mess: https://github.com/moq/moq/issues/1374#issuecomment-16712411...

That's a good point, and I had a similar reaction. I got laid off from my first job out of college in late 2006. Unemployed for 3 months, found a role with a new company 2 hours away, moved, and stayed there over 12 years because I didn't want to take the risk of moving. That, and I let myself be too intimidated by the interview process.

I eventually got laid off from there and found a new job after a ~2 month stint of unemployment; stayed with my new company for about 18 months and then changed to my current job of my own volition to find something I was happier with rather than sitting around being miserable.

Also built into Windows 11 as of the 22H2 release, just for the record.

That said, I may have to try this out - I've wanted to give Linux another go on my desktop but since I use captions rather heavily, that's been a disincentive. I'll have to see how this stacks up against other options; Apple's live captioning doesn't work as well as I'd like, Google's live captioning on the Pixel is great; Microsoft's live captions on Windows are pretty fantastic.

Unfortunately I'm tied to iOS for the longer-term since my hearing aid (a Resound model) integrates well with iPhones but not so much with Android, unless I want to buy an additional $400+ accessory.

I definitely had to experiment as well - I bounced around between Claritin and Allegra for a couple of years. I did try Zyrtec and experienced significant fatigue with it, which led me to stop using it pretty quickly.

Earlier this year I discovered Xyzal (expensive, but does have generic versions, too), which appears to be a close relative to Zyrtec. Interestingly though, they recommend taking it at bedtime rather that in the morning, the idea being you'll get a better night's sleep from not being congested AND you'll sleep through the worst part of the fatigue from it. This has worked out well for me - my allergies have bothered me less this year than they have in the last ten years or so, and I sleep better at night (I've had problems with insomnia off and on for ~20 years).

RE: the OP, I'm in my mid 40s and have always had some issues with following through and not procrastinating, and my sister was recently diagnosed with ADHD, so it's probably something I should at least discuss with my doctor at my next physical.

If the audio was actually sent to MS, I'd feel no worse than I would about Google getting it for live transcribe; the privacy implications annoy me but I'd rather have the captions.

That said, Win11 live captions work on-device and are not server-bound, and I've confirmed this by playing audio files and having captions still work while entirely off the internet.

For that matter, while I'm not an Android user - Google has moved most of their captions processing on-device and that is the same approach Apple is taking with theirs.

I've been testing 22H2 via Insiders Beta channel for a couple of months now; the Live Captions feature is great and has been almost life-changing for me. Really glad it's rolling out to wider availability now.

Pilot's ultra-fine pens are great, though for the money, I think the Juice and Juice Up 0.3mm pens are nicer all-around than the Hi-Tec C. A bit less scratchy, more comfortable to hold. Just in case you needed a rabbit hole to fall down, again. :)