HN user

bjackman

4,857 karma

[ my public key: https://keybase.io/bjackman; my proof: https://keybase.io/bjackman/sigs/VqJOkkjZ3NpAA_GO1V1yRLqqvoDbwJu0bLFsLxgdj8c ]

Posts21
Comments889
View on HN
yawn.io 11mo ago

`curl | sudo bash` isn't a security issue (on Linux)

bjackman
1pts0
news.ycombinator.com 1y ago

Ask HN: Anyone struggling to get value out of coding LLMs?

bjackman
345pts279
yawn.io 1y ago

The YX Problem (Reverse XY Problem)

bjackman
3pts1
github.com 1y ago

Show HN: Limmat – Local Immediate Automated Testing

bjackman
1pts0
lwn.net 2y ago

Extensible scheduler class to be merged for 6.11

bjackman
3pts0
news.ycombinator.com 2y ago

Ask HN: Any tool for managing large and variable command lines?

bjackman
49pts54
yawn.io 3y ago

Some Observations About German

bjackman
5pts0
www.theguardian.com 3y ago

Public invited to swear their allegiance as king is crowned

bjackman
2pts0
news.ycombinator.com 3y ago

Ask HN: Help me pick a front-end framework

bjackman
141pts176
yawn.io 6y ago

EBPF Is Turing Complete

bjackman
2pts0
yawn.io 6y ago

Pangrammatic Autograms

bjackman
3pts1
news.ycombinator.com 6y ago

Ask HN: What book to read to get a footing in CS theory?

bjackman
306pts82
news.ycombinator.com 9y ago

Ask HN: Does anyone record data about their lifestyle?

bjackman
2pts6
en.wikipedia.org 9y ago

Phantom Island

bjackman
1pts1
blog.cloudflare.com 9y ago

The Porcupine Attack: investigating millions of junk requests

bjackman
5pts1
cloudflare.com 9y ago

The Porcupine Attack: investigating millions of junk requests

bjackman
7pts2
github.com 10y ago

Intrade: Market Accuracy and Exploitable Miscalibrations in a Prediction Market

bjackman
5pts0
news.ycombinator.com 11y ago

Ask HN: Can we talk about FreeBSD vs. Linux?

bjackman
352pts214
news.ycombinator.com 12y ago

Ask HN: What does your CV/resumé look like?

bjackman
1pts1
sciencevsmagic.net 12y ago

Ancient Greek Geometry

bjackman
1pts0
github.com 12y ago

A first-principles implementation of arithemetic in Python

bjackman
3pts1

This is a very fast model.

I was already impressed by how fast 3.5 Flash was. But I've never compared it to other models in its class for coding.

Why? Coz models in that class are not very useful to me. Time saved waiting for responses usually just turns into time wasted replying to low quality responses.

Google need to release a Pro model ASAP. I am skeptical of the "maybe they don't have the compute to run it" thing. Anthropic were (probably) in that situation with Mythos and they announced it anyway - that's the obvious play for investor relations as well as hype for your product.

Coz you want to know what flooring, doors, cupboards, bath, sinks, railings, windows, etc they are putting in.

I went to view a flat this week, it was a building site. That's still the most important bit coz you get a feel for the size and shape of the space which is what really matters. But I'm glad to have the AI renders too.

Requiring disclosure seems obvious.

Using AI for these pics is also not inherently deceptive though.

I live in an extremely overheated housing market where properties are usually sold/rented long before they actually get completed. I'm fine with landlords using AI in their renders to make claims about how the place will eventually look.

You also see people using AI to put furniture into the image (I assume they are also taking out the furniture that's actually there, belonging to the previous tenant, but doesn't fit their desired aesthetic). Again, nothing _inherently_ deceptive about this.

Main thing is just whether tenants are empowered to back out of the contract if they don't get what they were promised.

Anyone who e.g. uses AI to expand rooms/windows... Jail please.

After a rebase, mybranch@{1} refers to the previous location of mybranch, so you don't need to manually track these before-rebase branches etc.

(In practice I find this syntax super annoying and usually end up typing `git reflog mybranch` and then copy-pasting the commit hash from the output).

Even if nobody is "cheating" your particular definition of cheating, the benchmarks are _somewhere_ in the super-structural gradient descent. Models are benchmark-maximising machines at some level, so I think the benchmarks are inherently a bit useless.

This is not really surprising, benchmarking _people_ doesn't work. You can only get a decent measure of someone's coding abilities by personally interacting with them. Given that models are basically person simulators it would be weird if benchmarks kept being useful as the simulation got more accurate.

I think what I've just said is basically just a more roundabout way of what you said: "Goodhart's law at work". It really is a law.

Great to see there are others doing the same.

It's a extremely cool that ESPHome is able, just by existing and being good, to create this little industry of no-bullshit products. What an awesome project that is!

Yeah I was thinking airtightness might be the difference. My flat seems to be bizarrely hermetic (when you turn on the kitchen extractor fan, it struggles if you don't have a window open somewhere).

So maybe a few leaky cracks are enough that when you open a window you get a bit of a through-draft.

That's interesting coz I found the opposite, at my place to keep the level below 1k I usually have to open a window in the room I'm in, or use a fan.

I live on a noisy street so I don't usually want to do that, if I open a window at the back and keep internal doors open it will stay reasonable but significantly elevated.

So yeah I think the lesson here is you probably need to buy a sensor, different homes are gonna differ.

My home is quite small (probably 80m²) and has literally zero ventilation built in (even in the bathroom!). I live in Switzerland where it's traditional to actively ventilate your home twice a day. But that doesn't do anything for CO2. Also it's such a fucking waste of time lol. Looking forward to moving into a modern building.

IMO it's something where an intervention is often cheap enough that it's worth it even without great evidence.

But also bear in mind that regardless of "are we operating at max effectiveness", OSHA sets a legal limit of 5000ppm in a workplace, and that's about _safety_.

This article is talking about keeping levels below 1000 which is a very high standard IMO (still arguably justified given the studies mentioned). But if you are in a poorly ventilated home office you could easily hit 3000. At that point you are closer to "illegal in the US" than "earth's atmosphere".

So yeah even if you are unconvinced about micro-optimising your CO2 levels there's a very long established argument in favour of at least paying _some_ attention to it.

As a middle ground I can also recommend this unit: https://apolloautomation.com/products/air-1

Looks like it's increased in price unfortunately but I like the idea, it's basically just what you would do as a DIY project but ready built. So you can either use it like a normal commercial product, or you can just fork the ESPHome config that's on GitHub and flash it exactly like any normal ESPHome project.

And I believe the accuracy is also not great on these cheap ones. The product in the OP's photo costs $200 where I live! And ISTR finding the sensor itself contributes a lot to this cost.

IIUC they also need fans. The one I have in my home has one that's actually integrated into the sensor unit.

Immich 3.0 19 days ago

Ah yeah I see. I guess the fallback there would be to split the service up into a remote encrypted storage layer that goes on the VPS and then host the actual service (with the decryption keys) locally?

But ISTR reading Immich kinda assumes the storage is on a plain local filesystem so you get perf issues if you do something clever under its feet. Could be out of date on that.

Immich 3.0 20 days ago

I really don't think you want E2EE for this. I host storage for family and friends, I haven't set Immich up yet (don't think I'd have space for everyone's photos) but the choice is between:

1. "Hey just so you know, I have access to everything you upload here".

2. "Do NOT lose your password or your data will be GONE FOREVER and I CANNOT get it back".

I definitely prefer 1 and I'm sure my users do too. They shouldn't upload it if they didn't trust me anyway.

In my case I follow it up with "and I might actually go digging around in your files if I need to debug something or you're wasting disk space". But I think you could also follow it up with "but I do promise not to look" and that would be valid too.

This whole thing only makes sense for people you're pretty close to.

(I do tell people not to back up their password managers on my system though).

I guess maybe for Immich specifically it would be nice to have a "vault" feature where people can upload nudes etc where they are willing to trade risk of loss for privacy on a per-photo basis.

No? Some subsets of the world has papyrus I guess but for most people during most of that time people were pressing text into clay and wax and stuff, it must have fucking sucked.

Then we got paper and pens and that was a pretty decent interim for a short period. Then about 100 years ago we got typing. Then about 20 years ago we reached a world where almost everyone is better at typing than they are at writing.

Obviously it's still important that people can write by hand, but expecting people to do it for more than a few hundred words at a time is idiotic. Would you like your clothes to be hand sewn too? That also worked for thousands of years (much longer than writing) but we stopped doing it for a very good reason.

Well, how many times in that 9 years have you written on paper for 2 hours straight? Even as a kid who did it regularly, it sucked.

Doing it now I really don't think I could deliver my intellectual best while worrying about if anyone can read my handwriting and whether I'm gonna cramp up by the end of the exam.

Pen and paper is just not a very good way to produce text.

As a $BIGCORP member I don't think this would be a great solution. I suspect there are plenty of vibe coding PR spammers that work for my company. And the admins of the GitHub org would not really care, making it easy for staff to contribute to third party projects is nowhere near their top priority (and policing the behaviour of their org members outside of org-owned repos is not in their mandate even if they wanted to).

Nothing major just a few little details:

- the /artifact thing is quite useful (don't think CC has it?)

- the /tasks is a bit better than CC's equivalent

- there are a few built-in skills that I haven't found CC equivalents for in the built in set (but the fact that I haven't sought out 3rd party versions shows you they aren't very important).

And more generally it does a better job of making the agent available. When Claude is debugging something complex and running a bunch of experiments it's often unavailable for like 20 minutes at a time, you only have /btw. Whereas AGY tends to more aggressively use timers and background jobs.

But now I wrote that out, I realised it's probably just as much of a system prompt thing as a harness design thing. Coz Claude _can_ operate that way too.

Anyway, like I said none of these come anywhere near balancing out the model quality gap.

I'm a Brit living abroad, when I visit the UK I use a Tailscale network with an exit node at my home, and yeah this always seems to work for me.

Going the other way around to try and watch British TV I used to find with a normal hosted VPN services could still figure out I wasn't in the country, but now I have a Tailscale exit node at my mum's place in the UK it always works fine.

So I suspect it all comes down to the IP source, probably a residential IP is the best possible case and with commercial VPNs it depends on how hard they work on isolating their IP blocks from known datacentres.

You don't save memory with code, you save it with architecture, platform decisions, and feedback loops.

Feedback loops are the important bit. If you want to reduce your service's memory footprint, don't at the code look at the memory profile and monitoring. You will find something like "oh shit 30% of our RAM is used by these buffers that we could basically eliminate if we tweak the flush frequency".

If you automate/regularise those investigations you will get an efficient service.

Same is true of every other performance metric, and reliability. It comes from your engineering processes (alerts, qualification, prod experimentation) not "write better code".

For my personal use, I really fucking want a humanoid robot, coz my home and all the bullshit in it was built for humanoids, I want a robot to do the bullshit for me. I don't want to move into a new, robot-oriented home.

I've never been to a factory but I bet there's a lot of the same bullshit. Ditto in a mine.

On the other hand, I've been in a datacentre. I don't see much need for a humanoid form in there, everything is flat and predictable. Why don't we have robot DC techs? This is probably an interesting clue re the next 10 years of robotics and maybe the reason Boston Dynamics is only valued at $1.1B.

Seems we might still be pretty limited on usecases. Maybe a dexterity bottleneck.

This is weird coz as a user of both Gemini and Claude I have the opposite feeling.

Antigravity CLI is quite decent, it's a huge step up from Gemini CLI (like, for example, it actually fucking works) and has some genuine advantages over Claude Code. Does Codex have something over both of them? I haven't tried it.

But the model just fucking sucks. Before I switched to Claude for personal stuff a few weeks ago, I was like "damn model capabilities are really slowing down" but no, it's just Gemini that's slowing down.

Will have to see if 3.5 Pro is any good when that comes out. But it feels like they would be attempting to catch up to Opus, not to Fable.

FWIW issue is never really about the code it writes it's about general intelligence. Gemini hallucinates like it's 2024, fails to follow instructions, and goes down wildly wrong debugging paths. Opus just gets the job done, first time, every time. With Gemini it feels like "I _am_ glad this intern is working for me but I'm tired of babysitting him" and with Claude it's like "this new PhD guy can replace me soon".

The "driving" tech I want in my car is:

- Cruise control.

- Camera for parking. I guess sensors too. These are just unbelievaly useful IMO, it makes parking trivial in cases that used to require quite intense focus. I see the appeal of fully automated parking, but with cameras and a car that you have lots of experience parking I think I am fine Austin-Powers-ing into any space that the car physically fits into.

- I guess, maybe, I kinda like the thing where it automatically watches your blindspot and has a little orange light to remind you that there's a car there.

I dunno, when did cars get all that stuff? (Cruise control was basically universal in the US before I was even born I think, but not sure when the others showed up).

But then there's some non-driving tech that I do want:

- Completely frictionless navigation and media control. Android Auto just seems to be fucking nonfunctional so I think maybe what I want here is actually just a Qi mount and a reliable bluetooth controller?

- I've never had it but I bet remote climate control is really nice (warm up the wheel 5 mins before you set off on a frozen morning / turn on the AC 2 mins before you get into a car that you couldn't park in the shade).

I have no plans to buy a car but I'm curious: what is the sensible choice for technical people with a reasonable amount of money?

I rent cars whenever I travel to the US and I've never not been pissed off by a car's software.

If you live in a country that makes it practica/affordable and you don't need too much range, I wonder if buying an old car with a broken engine and paying someone to do an electric conversion is a good choice?

Or maybe generally just buy a ~10 year old car, find a mechanic and say "I want this car to last a really long time, if we can build a trust relationship I will spend a lot of money in your business" and just budget for extensive proactive maintenance? Maybe with this approach you can still save money relative to a new car?

Or, is it possible to buy a newish car and then just rip out and completely replace the infotainment/climate control/etc while still keeping stuff like the parking cameras working?

Hosting DeepSeek Pro yourself is gonna be wildly expensive though?

You have a wide choice of providers available, so if you can find one you trust you can get inference without data harvesting and it's still very cheap. But dedicated HW is insanely inefficient.

Opus estimates you can do it for $13k/month if you get committed pricing on the HW.

Re being away from the HW: with Tailscale and llama-server it's now super easy to just run an inference server at home and use it from wherever you are.

It's a good technical artifact yeah but it would need to be forked and degoogled, today it is only really useful with Google services as a backend.

Also it's coupled to the device ecosystem which is organised by Google. This coupling with the HW is one of its major technical strengths though, including for the security things I'm yapping about.

So yeah I think the two options for a EuroOS are:

- Fork and degoogle ChromiumOS/AOSP

- Invest in a Silverblue/bootc/Flatpak style system and just keep filling the gaps there

Hard to say which would be the better option. Both require at least tens of millions in investment over 5+ years.