HN user

anp

4,881 karma
Posts14
Comments174
View on HN

It’s far from a perfect analogy but I would imagine that people were pretty hyped about the novelty of the first legitimately useful compiled programs where they didn’t have to allocate their own registers. I wonder how long it took for that novelty to wear off?

Or in other words I’d argue novelty is contextual and that these kinds of discoveries’ novelty will eventually wear off too but for right now it’s pretty cool that the “math discovery compiler” works well enough to do this (again imperfect analogy).

I feel similarly but IIUC I think that doesn’t strictly require an open source development model. I’ve benefited a huge amount from consuming and contributing to open source projects and I’m a bit worried that the “unit economics” changing might break some of the social dynamics upon which the ecosystem is built.

I read the Hyperion books during a particularly intense period of my life and found them quite powerful. I didn’t know anything about Simmons at the time, but I suppose I shouldn’t be surprised that like Tolkein these stories started with an oral format for children.

Keep Android Open 5 months ago

any [] property can be [taken by the state] from its [original] owners simply by [those owners becoming more powerful than the state wants]

When rephrased like the above, I think what you’re describing is pretty common in history. Many industries and assets have been nationalized when it serves the state’s interests.

IMO the moral justification is that there is no ownership or private property except that which is sanctioned by the state (or someone state-like) applying violence in its defense. In this framing, there’s little moral justification for the state letting private actors accrue outsized power that harms consumers/citizens.

I’m not sure I see that assumption in the statement above. The fact that no prompt or alignment work is a perfect safeguard doesn’t change who is responsible for the outcomes. LLMs can’t be held accountable, so it’s the human who deploys them towards a particular task who bears responsibility, including for things that the agent does that may disagree with the prompting. It’s part of the risk of using imperfect probabilistic systems.

FWIW I understood GP to mean that it suddenly makes sense to them, not that there’s been a sudden focus shift at google.

This has mostly been my experience as well although I don’t tend to run yolo mode outside of an isolated VM (I’m setting them up manually still, need to try vagrant for it). That said, it seems like some of the people who are more concerned about isolation are working with more untrusted inputs than I’ve been dealing with on my projects. It’s rare for me to ask an agent to e.g. read text from a random webpage that could bring its own prompt injection, but there are a lot of things one might ask an agent to do that risk exposure to “attack text”.

Anyone who finds this relatable (like me) might benefit from learning more about the last couple of decades of research on emotional regulation, trauma, and the nervous system. I have a great “trauma informed” therapist and over time this tendency of mine feels much less compulsive and more like a choice I can make because I know I’m good at something. At least for me having a calmer internal life has made it way easier to pick my battles and it usually means I end up feeding my desire to be useful on more satisfying and impactful things than I would have chased in more obsessive times in my life.

LearnixOS 7 months ago

Others think someone from the Rust (programming language, not video game) development community was responsible due to how critical René has been of that project, but those claims are entirely unsubstantiated.

I understand that I have a bias, which is why I was disclosing it. I think it strengthens my question since naively I'd expect a self-professed zealot to buy into the narrative in the blog post without questioning the data.

I find a lot of these points persuasive (and I’m a big Rust fan so I haven’t spent much time with Zig myself because of the memory safety point), but I’m a little skeptical about the bug report analysis. I could buy the argument that Zig is more likely to lead to crashy code, but the numbers presented don’t account for the possibility that the relative proportions of bug “flavors” might shift as a project matures. I’d be more persuaded on the reliability point if it were comparing the “crash density” of bug reports at comparable points in those project’s lifetimes.

For example, it would be interesting to compare how many Rust bugs mentioned crashes back when there were only 13k bugs reported, and the same for the JS VM comparison. Don’t get me wrong, as a Rust zealot I have my biases and still expect a memory safe implementation to be less crashy, but I’d be much happier concluding that based on stronger data and analysis.

I tend to agree but there are a few scenarios where I really want it to work. Debuggers in particular seem hard to get right for the current agents. I’ve not been able to get the various MCP servers I’ve tried to work, I’ve struck out using the debug adapter protocol from agent-authored python. The best results I’ve gotten are from prompting it to run the debugger under screen, but it takes many tool calls to iterate IME. I’m curious to see how gemini cli works for that use case with this feature.

Not GP but 2CB and psilocybin were never very visual for me compared with LSD in my tripping days. I have aphantasia and the only chemical to give me full eyes open visuals was DMT. Mescaline was a very distant second.

This matches my experience and I was quite surprised to find out other aphantasiacs have their “minds eye open” when tripping. For me psychedelics only ever produced a fractal overlay on top of what I was already seeing.

I wondered for a long time why everyone else experienced such strong visuals and eventually decided on my own it must be related to aphantasia. It’s nice to find out I might not have been a total crank with that hypothesis :).

(I work on a project that uses Chromium’s commit queue infrastructure)

I think there’s a big difference between Chromium’s approach and the “not rocket science” rule. AIUI Chromium’s model there are still postsubmits that must pass or a change will be reverted by a group monitoring the queue. This is a big difference in practice vs having a rotation or team that reorders the merge queue and rolls changes up to merge together. In the commit queue model you land faster at the expense of more likely reverts than in the merge queue model.

Comments so far seem to be focusing on the rejection without considering the stated reasons for rejection. AFAICT Alsup is saying that the problems are procedural (how do payouts happen, does the agreement indemnify Anthropic from civil “double jeopardy”, etc), not that he’s rejecting the negotiated payout. Definitely not a lawyer but it seems to me like the negotiators could address the rejection without changing any dollar numbers.

FWIW this closely matches my experience. I’m pretty late to the AI hype train but my opinion changed specifically because of using combinations of models & tools that released right before the cut off date for the data here. My impression from friends is that it’s taken even longer for many companies to decide they’re OK with these tools being used at all, so I would expect a lot of hysteresis on outputs from that kind of adoption.

That said I’ve had similar misgivings about the METR study and I’m eager for there to be more aggregate study of the productivity outcomes.

This was an interesting read since it's unlike any conversation I've had with the current bots, I haven't done a lot of exploratory probing of conversations I wouldn't have otherwise had with a person.

That said, I'm a bit surprised it agreed with you about the persuasiveness of the final approach (and maybe there's the agreeability to counter my previous point?). I agree a consequentialist argument could be compelling in the abstract but in my experience many bigoted people who care about things like the NAP will have emotional responses to social compromise so extreme that it wouldn't be a good idea to challenge them directly on the consequences of their actions. Without having any prior relationship with someone I would maybe expect that I'd achieve more influence with them if I learn to speak their own priorities back to them before I gently challenge them via contrast rather than argument.

you are a highly critical thinker and this is tempered by your self-doubt: you absolutely hate being wrong but you live in constant fear of it

Is there any work happening to model these kinds of emotional responses at a “lower level” than prompts?

I see work around like councils of experts and find myself wondering if some of those “experts” should actually be attempting to model things like survival pressure or the need to socially belong that many would normally consider to be non-rational behaviors.

Maybe I’m just falling victim to my own cognitive biases (and/or financial incentives as a Google employee), but I get the closest to that experience among the frontier labs’ chat interfaces when I chat with Gemini.

I chuckled and upvoted but I think it might be more subtle. It’s best to avoid negation with humans if possible too, but we are also way better at following negative examples than an LLM. I suspect it might have something to do with emotional responses and our general tendency towards loss aversion, traits these mind emulators currently seem to lack.

Cursor CLI 12 months ago

FWIW at least with Claude and Jules on a project I have a decent setup where I put all of the real content in an agents.md and then use “@agents.md” in CLAUDE.md. If all of the tools supported these kinds of context references in markdown it wouldn’t be that hard to have a single source of truth for memory files.

I have said a lot of times that I feel like people keep trying to reinvent a state and taxes to pay for shared infrastructure with open source maintenance. I don’t know how to use the state to solve the problem without severely degrading the quality of what gets built though.

I think there might be interesting time scales in between “now” and “my entire career” to which the bitter lesson may or may not apply. As an outsider to ML I have questions about the longevity of any given “context engineering” approach in light of the bitter lesson.

Not GP but as a 90s kid who didn’t learn programming until my 20s my impression is that script kiddies with cute hacks are a huge source of creativity. And maybe even playful tools, which I personally enjoy. Something about not focusing on deep programming or tech maybe makes it easier to just have fun and accomplish your goals.