HN user

nharada

3,342 karma

ML in real life

Posts40
Comments484
View on HN
www.youtube.com 22d ago

Free Electron Lasers (2017) [video]

nharada
2pts0
www.pi.website 7mo ago

Emergence of Human to Robot Transfer in VLAs

nharada
1pts0
techcrunch.com 8mo ago

Waymo robotaxis are now giving rides on freeways in LA, SF and Phoenix

nharada
341pts437
www.bloomberg.com 9mo ago

America's Tech Right Is Obsessed with Building Giant Statues

nharada
6pts1
thelifeelectric.us 10mo ago

My EV roadtrip experience after upgrading from Chevy Bolt to Ioniq 6

nharada
2pts0
www.cli-agents.click 10mo ago

Show HN: Give Claude Code control of your browser (open-source)

nharada
10pts3
www.theverge.com 11mo ago

Microsoft is getting ready to return to the office

nharada
8pts1
www.bloomberg.com 1y ago

Mount Everest's Trash-Covered Slopes Are Being Cleaned by Drones

nharada
4pts3
dronexl.co 1y ago

Skydio's Tracking Mailers to Police Spark Privacy and Security Concerns

nharada
3pts1
twitter.com 1y ago

We CT scanned several popular water filters before and after use

nharada
3pts1
news.ycombinator.com 1y ago

Ask HN: Help finding a smart file organizer that uses a local LLM

nharada
3pts3
www.youtube.com 1y ago

Meteorologist sets up surge cam livestream for hurricane Milton [video]

nharada
1pts0
next.voxcreative.com 1y ago

How Do We Feel About Autonomous Vehicles?

nharada
1pts4
www.youtube.com 2y ago

Electra First ESTOL Flight May 2024 [video]

nharada
2pts0
twitter.com 2y ago

Reverse engineering a software crack

nharada
213pts87
en.wikipedia.org 2y ago

Republic XF-84H Thunderscreech

nharada
1pts0
www.avweb.com 2y ago

Unleaded aviation fuel to go on sale in California by summer

nharada
1pts0
usezeroshot.com 2y ago

Show HN: Build an open-source computer vision model in seconds using text

nharada
64pts14
www.youtube.com 3y ago

Nopia, a Tonal Harmony Instrument

nharada
4pts0
moonshineai.readthedocs.io 3y ago

Show HN: Moonshine – open-source, pretrained ML models for satellite

nharada
86pts15
www.theverge.com 4y ago

Labrador's robotic shelf for people with mobility issues

nharada
112pts28
www.youtube.com 4y ago

Art of the Long-haul: Pilot's video series of ferry flights during Covid

nharada
2pts1
www.theverge.com 4y ago

Labrador's robotic shelf for people with mobility issues

nharada
2pts0
news.ycombinator.com 4y ago

Ask HN: How to accommodate everyone's office preferences?

nharada
1pts1
news.ycombinator.com 4y ago

Ask HN: What Was the iPhone Release Like?

nharada
6pts9
www.politico.com 5y ago

Biden plan could save California high-speed rail if state leaders can ever unite

nharada
1pts0
authentic.sice.indiana.edu 5y ago

The Affective Growth of Computer Vision [pdf]

nharada
2pts0
news.ycombinator.com 5y ago

Ask HN: What technology are you excited or optimistic about?

nharada
16pts18
news.ycombinator.com 7y ago

Ask HN: Books for how to be a better tech lead?

nharada
1pts0
www.usatoday.com 8y ago

A permanent emergency: Trump becomes third president to renew post-9/11 powers

nharada
4pts0

Yeah if it has the amount of content people are expecting I'd say it'll be well worth 80 bucks

This is super cool and exactly what I've been looking for for personal projects I think. I wanna try it out, but the "agent" part could be more seamless. How does my coding agent know how to work this thing?

I'd suggest including a skill for this, or if there's already one linking to it on the blog!

Saying nothing about the actual performance of this model, it does strike me how .... minimal(?) this announcement is. Their safety section is like 2 paragraphs about bioweapons. Go look at the reports for OpenAI and Anthropic's model releases. It's like 50+ pages of tests, examples, reports, and benchmarks across a bunch of safety and wellfare metrics.

If Meta wants to be seen as a cutting edge massive lab they need to come across as one instead of looking like a school project version of a frontier model.

I just can't find myself summoning the energy to be mad about markdown. It's good enough for like 99% of the things I use it for. Sometimes I get annoyed at specific extension support or whatever when I realize I shouldn't be using markdown for that task.

The devices are either dangerous, or they're not

That's not actually how it works though, it's all a risk and percentages. Nobody says "driving is either safe or it's not" or "delivering a baby is either safe or it's not"

I’m curious if you think viewpoints have also gotten more extreme in this period. It feels like the gap in political ideologies has widened a lot since I was younger.

Yeah as long as the chatbot is empowered to fix a bunch of basic problems I'm okay with them as the first line of support. The way support is setup nowadays humans are basically forced to be robots anyway, given a set of canned responses for each scenario and almost no latitude of their own. At least the robot responds instantly.

Nice, I like the idea. It sounds like qualitatively you haven't had any performance regressions while doing this, but have you tested it at all on any sort of benchmark or similar eval? I'm curious how well the actual system performs with less context like this. I mean it's possible it actually improves...

Another interesting thing here is that the gap between "burned out but just producing subpar work" and "so crispy I literally cannot work" is even wider with AI. The bar for just firing off prompts is low, but the mental effort required to know the right prompts to ask and then validate is much higher so you just skip that part. You can work for months doing terrible work and then eventually the entire codebase collapses.

+1

First, I agree with most commentators that they should just offer 3 modes of visibility: "default", "high", "verbose" or whatever

But I'm with you that this mode of working where you watch the agent work in real-time seems like it will be outdated soon. Even if we're not quite there, we've all seen how quickly these models improve. Last year I was saying Cursor was better because it allowed me to better understand every single change. I'm not really saying that anymore.

This is awesome and I'm really happy to see this progress. Landing a new chemistry in a production car THIS YEAR is some crazy velocity, especially compared to where other Na-Ion batteries are in the development cycle elsewhere. Is anyone else even close to having a car on the road with their cells?

The reason this is so exciting for me personally is for stationary energy. Because the raw materials are so abundant and have good cold weather performance, both grid and home level energy storage costs should come down significantly as this is commercialized further.

Claude Opus 4.6 6 months ago

That's a massive jump, I'm curious if there's a materially different feeling in how it works or if we're starting to reach the point of benchmark saturation. If the benchmark is good then 10 points should be a big improvement in capability...

worry that the US will fall behind the curve

Man it's already over. It's hard to imagine the US autos EVER catching up at this point, even with state support.

I'm obviously biased, and I probably have more gripes than most about Waymo as a corporate entity, but the premise this article seems to be based on is "Waymo is a zombie company who will never release a real product" or something similar?

They seem to be scaling just fine. Here in SF they're ubiquitous and most people I know use them regularly (and usually prefer them to rideshare). Sure, it's not the type of growth possible with pure software, but they're doing 500k rides/week and are looking to be doing 1MM/week by the end of the year. What scale does this business need to be for the author to consider them a real company?

In that case working at a startup would be a thing someone would only do as a last resort, and the talent pool would consequently be extremely low quality. Sounds damaging to the scene to me.

It definitely feels like a jump in capability. I've found that the long term quality of the codebase doesn't take nosedive nearly as quickly as earlier agentic models. If anything it's about steady or maybe even increasing if you prompt it correctly and ask for "cleanup PRs"