HN user

robkop

399 karma

Rob Kopel robkopel.com

Posts12
Comments111
View on HN

Could you please elaborate a bit more for my understanding?

What in particular about this method breaks correct token boundaries?

On my first read I read your comment as there are special tokens that require multiple tokens to emit, hence you can't get certain tokens emitted alone - but I don't think that's what you're getting at on a second read?

Interesting that you've found similarities between "d" and the hidden tokens for opening an xml tag, pressing caps lock and the other hidden tokens of note. I haven't run into any trouble extracting "d" tokens, is it a particular model that you see create that pattern?

A lot of enterprises were doing that but now they hit the 150 user limit on Claude and are paying seat+api rates.

Codex is still going strong but it’s hard to imagine they won’t do similar eventually.

So now im honestly hearing a lot more folk stick it out with cursor while waiting for the dust to settle.

There’s a lot of tradeoffs to play with, those inference ASICs may not carry the gradient but they are still optimised for larger batches and to run any model. They need enough memory for the weights, wide batch inference, and ideally leftovers for kv cache efficiency.

For personal inference you’re given a lot more room to play in - much of it poorly explored today - enough to concern an argument of cost advantages evaporating

You can ablate surprisingly large chunks of a model with near to no effect, you can try this easily - download an open weight model in torch.

Obviously it’s not ideal but you could likely have single digit % of all weights affected and still have a useful model (many caveats here: e.g. locality of damaged weights matters, distribution of errors matters, fail high/low matters, …)

I can’t speak for the states, but in AU I clearly see a massive displacement of undergrad and junior roles (only in AI exposed domains).

I say this as both someone who works with many execs, hearing their musings, and someone who no longer can justify hiring junior roles themselves.

Irrespective of that; if we take this strategy of only taking action once it is visible to the layman - our scope of actions available will be invariably and significantly diminished.

Even if you are not convinced it is guaranteed and do not believe what myself and others see. I would ask you is your probability of it happening now really that close to 0? If not then would it not be prudent to take the risk seriously?

LLM=True 5 months ago

We’ve got a long way to go in optimising our environments for these models. Our perception of a terminal is much closer to feeding a video into Gemini than reading a textbook of logs. But we don’t make that ax affordance at the moment.

I wrote a small game for my dev team to experience what it’s like interacting through these painful interfaces over the summer www.youareanagent.app

Jump to the agentic coding level or the mcp level to experience true frustration (call it empathy). I also wrote up a lot more thinking here www.robkopel.me/field-notes/ax-agent-experience/

Rumours say you do something like:

  Download every github repo
    -> Classify if it could be used as an env, and what types
      -> Issues and PRs are great for coding rl envs
      -> If the software has a UI, awesome, UI env
      -> If the software is a game, awesome, game env
      -> If the software has xyz, awesome, ...
    -> Do more detailed run checks, 
      -> Can it build
      -> Is it complex and/or distinct enough
      -> Can you verify if it reached some generated goal
      -> Can generated goals even be achieved
      -> Maybe some human review - maybe not
    -> Generate goals
      -> For a coding env you can imagine you may have a LLM introduce a new bug and can see that test cases now fail. Goal for model is now to fix it
    ... Do the rest of the normal RL env stuff

I get this at least once a week. And then once you have to dig in and understand the full mental model it’s not really giving you any uplift anyway.

I will say that doing this for enough months has made my ability to pick up the mental model quickly and to scope how much need to absorb much quicker. It seems possible that with another year you’d become very rapid at this.

I added a "Human" LLM provider to my local OpenCode a few months ago as a joke, and it turns-out acting as a LLM is quite painful. But it massively improve my agent harnesses dev skills.

So I thought I wouldn't leave anyone out! I made a small oss game - You Are An Agent - youareanagent.app - to share in the (useful?) frustration

It's a bit ridiculous. To tell you about some entirely necessary features, we've got: - A full WASM arch-linux vm that runs in your browser for the agent coding level - A bad desktop simulation with a beautiful excel simulation for our computer use level - A lovely WebGL CRT simulation (I think the first one that supports proper DOM 2d barrel warp distortion on safari? honestly wanted to leverage/ not write my own but I couldn't find one I was happy with) - A MCP server simulator with full simulation of off-brand Jira/ Confluence/ ... connected - And of course, a full WebGL oscilloscope music simulator for the intro sequence

Let me know what you think!

Code (If you'd like to add a level): https://github.com/R0bk/you-are-an-agent (And if you want to waste 20 minutes - I spent way too long writing up my messy thinking about agent harness dev): http://robkopel.me/field-notes/ax-agent-experience/

Ax Not UX 6 months ago

It's a fair question - I think the fact that they hold abilities (read 200k tokens instantly, can clone themselves, ...) that we don't would suggest they will have quirks and differecnes.

What downstream implication that will have on a AX sense is certainly arguable, but I would put forward that we're already seeing it with effective harnesses such as Claude Code. The experience the agent has there is quite different to how you'd build an IDE for a human.

2025 Letter 7 months ago

Can you elaborate? I would have thought the main driver for the price of a service is the labor?

Occam's Razor - this complexity arises from the human nature to try and build consistent abstractions over complex situations. It's exactly what we do in software too. To an outsider it's going to look nonsensical.

I want to share a thought experiment with you - atop an ancient Roman legal case I recall from Gregory Aldrete - The Barbershop Murder.

Suppose a man sends his slave to a barbershop to get a shave. The barbershop is adjacent to an athletic field where two men are throwing a ball back and forth. One throws the ball badly, the other fails to catch it, and the ball flies into the barbershop, hits the barber's hand mid-shave, and cuts the slave's throat-killing him.

The legal question is posed: Who is liable under Roman law?

- Athlete 1 who threw the ball badly

- Athlete 2 who failed to catch it

- The barber who actually cut the throat

- The slave's owner for sending his slave to a barbershop next to a playing field

- The Roman state for zoning a barbershop adjacent to an athletic field

Q: What legal abstractions are required to apply consistent remedies to this case amongst others?

Opinion: You'd need a theory of negligence. A definition of proximate cause. Standards for foreseeability. Rules about contributory fault. A framework for when the state bears regulatory responsibility. Each of those needs edge cases handled, and those edge cases need to be consistent with rulings in other domains.

Now watch these edge cases compound, before long you've got something that looks absurdly complex. But it's actually just a hacky minimum viable solution to the problem space. That doesn't make it fair that citizens bear the burden of navigating it - but the alternative is inequal application of the law

For those curious about the "consistent principle of law" here - SCOTUS wrestled with nearly exactly this question in Free Speech Coalition v. Paxton earlier this year, and effectively emboldened more of these laws.

Previously the Fifth Circuit had relied heavily on Ginsberg v. New York (1968) to justify rational basis review. But Ginsberg was a narrow scope - it held that minors don't have the same First Amendment rights as adults to access "obscene as to minors" material. It wasn't about burdens on adults at all. Later precedent (Ashcroft, Sable, Reno, Playboy) consistently applied strict scrutiny when laws burdened adults' access to protected speech, even when aimed at protecting minors.

In Paxton the majority split the difference and applied intermediate scrutiny - a lower bar than strict - claiming the burden on adults is merely "incidental." Kagan had a dissent worth reading, arguing this departs from precedent even if the majority won't frame it that way. You could call it "overturning" or "distinguishing" depending on how charitable you're feeling.

The oral arguments are worth watching if you want to understand how to grapple with these questions: https://www.youtube.com/watch?v=ckoCJthJEqQ

On 1A: The core concern isn't that age-gating exists - it's that mandatory identification to access legal speech creates chilling effects and surveillance risks that don't exist when you flash an ID at a liquor store.

Note: IANAL but do enjoy reading many SC transcripts

Ahh, there's a bug with the z-index on the Turing one (I made it a "legendary card" for surviving 70 year), will fix shortly.

Here's the link for the moment: https://arxiv.org/pdf/2405.08007

Also if you want to read the original Turing paper (it was interesting to look back upon, I think the future of benchmarks may look a lot more like the Turing test): https://courses.cs.umbc.edu/471/papers/turing.pdf

For my year end I collected data on on how quickly AI benchmarks are becoming obsolete (https://r0bk.github.io/killedbyllm/). Some interesting findings:

2023: GPT-4 was truely something new - It didn't just beat SOTA scores, it completely saturated several benchmarks - First time humanity created something that can beat the turing test - Created a clear "before/after" divide

2024: Others caught up, progress in fits and spurts - O1/O3 used test-time compute to saturate math and reasoning benchmarks - Sonnet 3.5/ 4o incremented some benchmarks into saturation, and pushed new visual evals into saturation - Llama 3/ Qwen 2.5 brought Open Weight models to be competitive across the board

And yet with all these saturated benchmarks, I personally still can't trust a model to do the same work as a junior - our benchmarks aren't yet measuring real-world reliability.

Data & sources (if you'd like to contribute): https://github.com/R0bk/killedbyllm Interactive timeline: https://r0bk.github.io/killedbyllm/

P.S. I've had a hard time deciding what benchmarks are significant enough to include. If you know of other benchmarks (including those yet to be saturated) that help answer "can AI do X" questions then please let me know.