HN user

StevenWaterman

1,535 karma

https://github.com/stevenwaterman

HackerNews@StevenWaterman.uk

Posts26
Comments213
View on HN
stevenwaterman.uk 1y ago

Authenticating JavaScript WebSockets

StevenWaterman
3pts0
stevenwaterman.uk 1y ago

Nomerge Comments in CI/CD Pipelines

StevenWaterman
1pts0
stevenwaterman.uk 2y ago

Luckily I got hit by a car

StevenWaterman
21pts67
www.cloudflare.com 3y ago

The DNSSEC Root Signing Ceremony

StevenWaterman
8pts0
pubmed.ncbi.nlm.nih.gov 3y ago

A mathematical model for the of area under glucose tolerance curves (1994)

StevenWaterman
2pts1
talkjs.com 3y ago

Onboarding at a Radical Enterprise

StevenWaterman
2pts1
stevenwaterman.uk 4y ago

Don't Worry, It's Rocket Science

StevenWaterman
2pts0
stevenwaterman.uk 4y ago

Decorate Your Blog with AI

StevenWaterman
2pts0
stevenwaterman.uk 4y ago

Show HN: My Balance Box – A mood tracker that publicly shares how I feel, live

StevenWaterman
2pts1
stevenwaterman.uk 4y ago

Kind and True

StevenWaterman
2pts0
stevenwaterman.uk 4y ago

Opening Up about Burnout

StevenWaterman
2pts0
twitter.com 4y ago

The overlap between Software Engineering culture and a happy relationship

StevenWaterman
1pts1
philippe.bourgau.net 4y ago

The story about how we do Agile Technical Coaching (2020)

StevenWaterman
1pts0
lexoral.com 4y ago

We Need to Talk (if you want)

StevenWaterman
1pts0
lexoral.com 4y ago

You can learn to Speak Confidently

StevenWaterman
1pts0
lexoral.com 4y ago

Things you don't need JavaScript for

StevenWaterman
501pts243
lexoral.com 4y ago

Database sync like magic, with Svelte and Firestore

StevenWaterman
1pts0
lexoral.com 4y ago

Lexoral is open-source so you can punish us

StevenWaterman
3pts0
narration.studio 5y ago

Show HN: Narration.studio – Automatic in-browser audio editing for voiceovers

StevenWaterman
1pts1
blog.scottlogic.com 5y ago

Down the ergonomic keyboard rabbit hole

StevenWaterman
251pts184
www.youtube.com 6y ago

Prototype It with Sat [video]

StevenWaterman
2pts0
blog.scottlogic.com 6y ago

3D Rendering on a Children's Toy

StevenWaterman
1pts0
blog.scottlogic.com 6y ago

GitHub is a free CI/CD/Hosting solution

StevenWaterman
2pts0
blog.scottlogic.com 6y ago

Embrace Your Obsessions

StevenWaterman
3pts0
github.com 6y ago

Show HN: MuseTree – Human-AI interaction for music composition

StevenWaterman
8pts2
blog.scottlogic.com 6y ago

Planning 56 sprints per second with SAT4J

StevenWaterman
2pts0

The set of models that are pareto-optimal, IE for some set of variables, no other model strictly dominates them = no other model is better than them on every variable.

So like, on a cost-intelligence graph, the cheapest and most intelligent models are pareto optimal. Then in-between those if you have

- cost $3 intelligence 6

- cost $1 intelligence 5

- cost $2 intelligence 4

The 1st and 2nd are pareto optimal, the 3rd is not, because it's dominated by the 2nd (2nd is cheaper AND more intelligent at the same time)

I'm pretty sure that's the plan. Currently they're legally bound to keep it within 0.9s of solar noon (or something like that) but in 2035 it's changing to +-1 minute, which basically kicks the can down the road for another century or so

I say we let it reach 15 minutes then countries can solve it themselves by shifting timezone by 15 mins. Since making sure solar noon matches noon on the clock, is the point of timezones existing in the first place

I thought about this more and realised your question might have been "what's the difference between knowing and learning". IE, how can we say the model believes something without having been taught it.

I think you're right that they're basically the same thing. I'd argue they're very slightly different because what an AI model ends up knowing isn't perfectly predictable based on what they were taught (emergent intelligence), but the sentence you quoted is using believing and learning to mean the same thing, it's just trying to draw attention to the fact that the training process structurally enforces "cheat as much as possible without getting caught".

IE, the contrast in the original sentence wasn't "believe" vs "learn", it was "good" vs "permissible"

I think we'll get there. Right now it works for me, because I'm naturally pretty verbose in my prompts, and know the codebase well, so I know what it needs to look at. Plus subagents for anything exploratory.

I think deepseek v4 pro has 1m context and does pretty well up to around 600k. But if you have the hardware to run that locally, you already know

Even then if there's a smaller model with 1M context, you'll need a ton of RAM to actually run it at full 1M. I guess that's why you don't see it too much. Anyone that could run Qwen 3.6 27B with 1m context would be better off running a much bigger model with smaller context instead, in the same amount of VRAM.

In terms of optimizing further, huge context + KV quantization sounds like a terrible idea, but there's some decent innovation in sparse attention, KV cache rotation allowing Q8 to perform nearly as well as full 16-bit precision, plus some ideas around offloading KV cache to system RAM (but I'm skeptical)

Yep, I daily drive Qwen3.6-27B (including for work), have done pretty much since it came out. IMO it's the only (small-ish, local) model worth using, if you can run it. It might not be as good as Opus at "add X large feature" but I don't want that in a model. I want to do the thinking while it does the typing. And Qwen 3.6 27B is perfectly good at that (while in my experience models like the 35A3B and gemma are significant downgrades)

Plus, I never have to worry about rate limits, quotas, or sitting in a queue during peak time. And I can always see its full thoughts, don't have to worry about where my data is getting sent, and know it can't get secretly nerfed.

Running on 2x 3090, 500-1000tok/s prefill and 60tok/s output at Q6_K_XL with MTP on llama.cpp, 220k tokens context window (starts to get a bit dumb above 160k ish), no KV quantization

The problem is, people see "they're not profitable once you account for training" and equate that to "AI will go away soon"

But if all the AI companies stopped training new models, they would all instantly become profitable (and stick around)

The thing that makes them unprofitable, is having to compete (which means training models). If / when enough companies exit the market, the cost to compete goes down and you end up in an equilibrium

My first instinct was that the essay would just be "67" as a stupid and harmless but nonsensical response.

Somewhat amusingly, mine depends on the examiner knowing how advanced AIs are. In the 1960s mine would just look like a trickle AI. It feeling human demands we assume the ai would actually be competent

Yours is even more effective. Both hinge on the solution being "be as unexpected and out-of-distribution as possible"

I somehow imagine they wouldn't like your essay that is made of 100% slurs though, regardless of how effective it is at the stated task

This thread started because of "the cheapest bridge that just barely won't fail"

My point was that safety factors are a part of this. A safety factor of 1.0, designing bridges so that they can perfectly withstand the expectations of intended use, means that some unacceptable % of those bridges will fall down in practice.

In other words, it's true that you can explain safety factors as:

Assuming perfect construction, and no defects, under designed maximum load, make sure that this bridge really stays up by a wide margin

But that misses the point of why we use safety factors. Nobody is paying for a bridge to really stay up by a wide margin. Because there's no material difference between a bridge that stays up, and a bridge that really stays up, right up until the point that the weaker one falls down due to inevitable over-loading or defects in construction / materials.

Safety factors account for uncertainty. Uncertainty the quality of materials, of workmanship, of unaccounted-for sources of error. Uncertainty in whether the maximum load in the spec will actually be followed.

Without a safety factor, that uncertainty means that, some of the time, some of your bridge will fall down

"If the heat shield breaks then I will die" is the exact situation for the astronauts, and yet we still have astronauts.

In fact it's worse for the astronauts, because in this hypothetical only the heat shield failing will condemn the POs to death, whereas any critical part failing kills the astronauts

Yes, it's a much sexier job than project manager, but clearly there are some people, in some circumstances, that would accept it.

A covert signal is still beneficial even if the signal is secure. The existence of the signal is valuable metadata.

For a contrived example, imagine I'm in a warzone:

- Secure = Enemies can't read my messages. Good. But they can still triangulate my position.

- Covert = Enemies don't know I exist

This is almost textbook countersignalling. The same as:

- Signalling: I dress more formally than everyone else to make up for the fact I'm less professional in other ways

- No signalling: I dress like everyone else because I am like everyone else

- Countersignalling: I wear ratty old clothes with holes in them, and nobody will dare to question it because I'm the important one here

Kevo app shutdown 10 months ago

If you want to run it overnight, or while you're at work, so it finishes as you arrive and doesn't leave the clean clothes in a clump for hours (or so it runs during cheaper power hours)