HN user

perching_aix

2,739 karma
Posts5
Comments1,746
View on HN

Yes, it's just a ToS violation at present. Those are legally binding though, despite the common adage. What that really translates to here though, anyone's guess.

Anthropic's own copyright infringement could apparently be forgiven for 1.5B USD after all, so maybe there's a price that breaking the distillation clause for is acceptable too. Or some other arrangement.

No? They outright say the opposite!

Like look, I'm not a native speaker, sure. But I think when someone says "value add", that means there was value there (which you claim they're rhetorically erasing), and then that was added to. Under no interpretation of this phrase do I get an erasure of prior value.

So certainly, as long as words mean anything, no, they absolutely did not say or suggest what you claim they did, and what you extract a thus unreasonable amount of obnoxious schadenfreude from, while throwing in a cheap insult for funsies at the end.

It's the second time I feel compelled to reach for this just today: https://i.kym-cdn.com/photos/images/original/002/659/979/108...

We seem to be reading the same comment(s) differently.

The context (verbatim):

Correct. We need open weights, open code and open data. If nobody else can reproduce what someone did there will always be security questions. Even if we can reproduce it there could still be security concerns but it's more realistic to investigate yourself.

In short, it's an appeal to full openness and reproducibility on the basis of security; open weights alone notably do not provide that same confidence. They're better in some respects, not really in others.

Then comes the question (also verbatim):

Exactly what are the possible 'security issues' of self hosting an open weights model?

Implying then that as long as you do have the weights and just self host it, the asker cannot imagine what could possibly go wrong. What is the gap, if any?

And so I explained. That was my point. Open weights do not give you full reproducibility, and so that on its own falls short of what the parent comment is making an appeal to. That there does remain a security concern, shared by remote and closed models, that does not improve just by having the weights, but would if you did have full reproducibility. Explaining that gap was my point, as that is what I understood as being asked there. It's the only thing I can reasonably imagine being asked, in fact.

This is a materially different question to what you apparently extracted (again, verbatim):

What security issues come from self-hosting?

Implying that by self-hosting models, something bad might specifically happen.

I do not think this, do not think I suggested this, do not think the original question suggested this, and generally do not think this is indeed any sensible, in or outside the context. Certainly not beyond something common sense, like vLLM being compromised or whatever.

You seem to agree. But then how did we get here, clearly talking past each other?

I'm not so sure about this view.

For me, it was the same genre of comment as e.g. comments that call out how vision models tend to be structured surprisingly similarly to how the visual cortex is structured, even without specifically attempting to make that happen, and within the boundaries of our understanding of either things of course.

So not so much anthropomorphization, and more a recognition of the unintentional, and surprising, biomimicry.

Though to be honest, I'll be damned if any of the pictures in the post look anything like any child drawing I've ever seen, Grok's or any of the other models.

I struggle to suspend disbelief that this is not just some ill-faith mockery.

The guy really quite explicitly spells out a journey of growing in intellectual humility. The writing is very, very blatantly in the genre of self-growth, which is why you see this structure so often. That he wanted to gravitate towards those whom he found smart (implying he did not consider himself that smart at the start, very directly contradicting what you say), participated, eventually discovered that there's smoke to the fire, then stopped participating, learning lessons along the way and now passing them on to the dear reader that much wiser.

How you walked away from this with the conclusions you did genuinely escapes me, and apparently this is not even your first time either then. Never stopped to reconsider? Is any expression of appreciating reason egotistical to you? Do you maybe instinctively read it as a comparison, as if it was meant to suggest that other people (i.e. you) thus somehow do not appreciate reason? I have a meme for you if so: https://i.kym-cdn.com/photos/images/original/002/659/979/108...

Or are you at one of these stages yourself, and we can expect a blogpost from you in the future too?

I might be slow today but I'm pretty sure that the Church-Turing thesis is about undecidable statements, not false decidable statements. If a decidable statement is false... then that's it, it is just false, that's how we can decide so in the first place.

It's not that the model can't be fixed, but that it won't create and fix itself, unlike how people iterate on their own mental models.

You can build a more elaborate model that can, but then that's still structurally constrained in terms of how it can adapt.

You can build a maximally general model that automatically captures arbitrary patterns, and estimates using them, but:

- that's not what people do and mean when using Bayesian inference in argumentation; they're the ones capturing and maintaining the patterns instead, manually and arbitrarily

- it kind of renders the whole thing meaningless; you basically get a machine learning model out of that, and introspectibility is not really their thing + it becomes quite literally arbitrarily adapting (and thus arbitrarily inferring)

One may also recognize that to be resembling a human, with whom you get the same issues. Kind of how one gets to arguing about empiricals and logic in the first place. This is what I describe in my other comment that I linked to, that debates using these instruments are basically just long winded coordination processes between people in the best of times, not some sort of divination tools.

I'm not the guy from above btw, but I am the guy who had their comment turn dead. I don't know why it did though, I doubt this subthread sees that much traffic, and I did not see it being flagged since I posted it. I'd assume it was automatically or manually moderated out.

As a sibling comment mentions, this is a chain-of-thought leak, in part evidenced by the excessive use of Markdown formatting symbols. In my judgement, all occurrences of noun-phrase fragments are thus better understood as either ortographic mistakes, or stylistic choices due to the language register it's trying to hit (i.e. note-taking, dictation), rather than grammar mistakes.

I do not spot any missing articles, and the missing subjects (as well as the debatable-to-be-missing transition phrases) too fall within the bounds of stylistic concern. Sentence length also.

Definitely not a pleasure to read mind you, but given that it wasn't meant to be read either, I'm not sure that should be surprising.

Sentences being too short and annoying to read is a stylistic grievance, not a grammar one. The examples were edited in after my original reply, but even then, I see no actual grammatical mistakes there.

Even with my native language, which is definitely a lot less represented in the training data, the worst I encounter are phrasing mistakes, incorrect use of idioms, and invented words. You have to use some really badly tortured local model to get an LLM to produce incorrect grammar.

I do sometimes see Opus make typos, which is entertaining, but again, not a grammar issue.

Undecidability is not relevant here beyond e.g. a model incorrectly claiming something is decidable, and potentially even following through. It is no different to the model incorrectly claiming anything else.

These are statistical systems, so guaranteeing any particular high level behavior is not possible because of that. But given that they're working with natural language, that was never going to happen anyways, for the obvious language theoretic reasons.

The more appreciable interpretation of the claim is that they can be nevertheless tuned so that this issue becomes practically resolved. Contending guarantees and theoreticals is simply misplaced, these are not formal symbolic reasoning systems being buggy.

Maybe we're prompting it different, but it's not "trying to be my friend" for sure, nor am I trying to be "its" friend either. Or at least I'm sufficiently oblivious to its advances, and find it unthinkable to form such a bond :)

On the flipside, it does spuriously make hilarious remarks like "Good data.", which I find pretty funny specifically because it comes across as just silly. Not sure how it'd be harmful either, a little entertainment I think goes a long way in this type of profession.

I see zero issues with these, and I have a hard time understanding why people have their panties in a twist so hard about them. I sometimes really quite wonder just what kind of correspondence would y'all prefer, and how would that sound like.

Matter of fact, do you have an example at hand? Like an exact before & after?

I'd argue it's hard to have a community without something in common, so identity formation is kind of unavoidable (as there is an identifier), even if people resist it.

In some cases, that point of commonality is just the venue (like your HN example). Not the case for this.

That said, I do appreciate the distinction and you may very well be right. Like certainly, one can appreciate logical thinking and quality argumentation without adopting any kind of group identity.

Inventing the hypotheses and the model in the first place. Bayesian inference updates probabilities over hypotheses already admitted by a model. If reality requires a structure that the model doesn't assign a prior probability to, such as an omitted variable, mechanism, or failure mode, ordinary Bayesian updating cannot recover it, no matter how much data is observed.

For example:

A forecasting model estimates restaurant demand from years of bookings. Then a major concert is announced next door. A local manager immediately expects a packed evening, while the model predicts an ordinary Tuesday because it has never encountered that situation, and was not prepared for it. To do so would have not been reasonable based on the data available to begin with, after all.

So you can have someone opine about how well some argument is supported through Bayesian inference, and how another isn't, but that's only going to keep them honest in a limited sense. They can still miss the bigger picture by e.g. simply not knowing about it, and not knowing to check for it. If you don't either, they done misled you.

This failure mode is identifiable without having to reason about probability theory as well by the way: https://news.ycombinator.com/item?id=48986444

If you're still unconvinced:

- why would BAND have a concert here? no concerts in this area, ever

- doesn't BAND have ties to this place? no sufficient evidence in the way of that

- BAND lead singer has hung out here once and liked THING, decided to have a concert here for sure one day on a whim

Individually, these have low local attributing info available, so you'd never be able to justify the inference. But not being able to justify the inference didn't make the claims wrong, just unfounded. Those are not the same thing. And so chasing probabilities like this, while maximizes how much your position is defensible, also boxes you in. You'll be reasonable in why you thought what you did, but that won't make you necessarily right, despite the suggestions otherwise.

Continuing the gaming metaphor, what I'm trying to get at is this is like Cyberpunk 2077, and you sound like a guy who has a grand total of 0 hours in it since launch, but has developed very strong opinions about it, and refuses to accept it improved or can improve, purely because it started out so bad that that's hard for you to even imagine. Pretending to be some kind of alien, who's just going through their first exposure to practical facts somehow.

When was the last time you tried an agentic harness (Codex, Claude Code, Copilot Chat in VS Code, Cursor, Pi, OpenCode, etc.) for work in any appreciable capacity, and with what model? Surely if you're so confident they continue to be unusable and that nothing materially changed, that must be backed by a recent significant experience that way? Or even just an experience at all?

If a game on release is unplayable-tier buggy, then improves over time to the point where bugs are barely noticeable, is acknowledging that going to count as "goalpost moving" to you? It never became formally verified after all, and it's even running on physical hardware... Woe are the people lying to me (nobody), the issues have not been fundamentally ruled out!

Meanwhile, if I search my file system for "foo" a trillion gazillion times, it will not return "bar" once.

Great! Nor will an agent any more likely, cause it just sends out a tool call and surfaces its output.

It boggles the mind. One would think this is some highly secretive technology only a dozen people in the world have access to, the way one has to argue tooth and nail about trivially verifiable facts regarding it. You quite literally do not have to take either of our words or "vague" judgement for it.

I did not claim they train people away from either.

That said, the review burden is rough, that I can agree with. I outright felt compelled to evaluate whether the additional review burden did not outweigh the benefits, but at least for my tasks it did not. So grumpily, I simply live with that pain.

Maybe it helps if I mention that my line of work is DevOps and Operations. I have an ongoing suspicion that this area is better suited than average for agentic work. The codebases are relatively tiny, the languages and technologies used are very well represented in training data, and there's a decent amount of side chore. I can definitely imagine agents being a lot more frustrating to work with on proper, sizeable codebases, and the numbers simply no longer adding up. I don't have much of a first hand account with that.

My closest exposure is some personal toy projects, where getting the actual vision out there ended up requiring an inordinate number of turns (this is with a frontier model). In my estimation it was still worth it, but I definitely had to give it a back of the napkin calc.

As far as my work goes, it is of course not magic, creatively worded AWS docs will still trip it up (as they initially also do me). In those cases, my expertise is still required. But the well trodden is very well trodden, and I could cut out a lot of cruft, including a lot of organizational minutia, which I very much appreciate. I was able to burn through my backlog almost completely, for example.

Being engaged == doing stuff, not necessarily being actually productive.

Sure. So to clarify, no, I did not just spend it on busy procrastination, and I don't think it promotes that either. Would contradict my story anyhow, this was not some trick I was trying to play on you.

Easy to pay a lot to the model vendor though eh?

Certainly, as I'm not the one paying. Though it's exactly corporate who really wants to have it both ways (who wouldn't?), and keeps trying to get me to use the crappy useless models because they are cheaper, despite them tanking productivity rather than helping it, so go figure.

There will definitely be a time when the hype dries up, and mgmt will start playing hardball. I'm confident that the value is there, and that I'll be able to demonstrate it to them that it is more than worth it. You can choose to not believe that, up to you. Maybe it really isn't true for your line of work, after all.

You mean why not work myself instead of the agent? That'd be because for me it's been working great, and so it does make sense. In the scenarios it doesn't, I do indeed just fall back to manual work. A lot of those scenarios are obvious too, so not too many wasted runs to speak of either.

I did give up on cheaper models, they required constant babysitting, and in those cases yes, the benefits indeed evaporated. The expensive models have been genuinely working wonders though, and were still able to justify themselves economically plenty, at least by my own measurements.

I think there's also one underappreciated and indirect way agents help with productivity: they counteract the attention span collapse of the past years. By being addictive themselves, they keep you engaged, and being engaged means being productive. Not even asking an agent to check something out feels too rich, and once you've asked, you're already one foot into the flow.

There's also definitely been some honeymoon effect going on for me, where I dived into more work more readily, just to see if the agent can figure things out on its own.

We probably understand this word very differently, because as far as I'm concerned, rationality (rational behavior) is essentially an exercise in behaving optimally, one that is thus strictly relative to some particular context. It is not just "facts and logic", and it is most certainly not universal, since universal optimality doesn't exist: https://en.wikipedia.org/wiki/No_free_lunch_theorem

The entire point of the exercise is to establish that certain things are true, and that this is not because of faith

Taken at face value, that'd mean bridging the analog hole (https://en.wikipedia.org/wiki/Analog_hole), which is fundamentally impossible. Logic proves consistency, not truth, despite the many related terminologies suggesting otherwise. This is further skipping over how people segment the world different, even relative to themselves over time, which thus forms a coordination problem.

It reminds me to how HTTP is stateless. Yeah, it's stateless - all the state keeping has simply been laundered a layer above. It may sound tantalizing to position certain arguments as simply rational (i.e. correct/optimal), but all that ultimately does is sidestep the actual coordination part, which never really went away. There is still an underlying segmentation model being asserted, it is just now coated in self-righteousness, an illusion of rigor, drawing up a false dichotomy of "either you segment the world like we do, and then you're correct, or you're incorrect, and your segmentation is too".

See this rhetorical landmine in action in a sibling thread: https://news.ycombinator.com/item?id=48986530

Even if I go out of my way and interpret the word such that it merely means "acting on reason" (as opposed to some sort of nonreason), that doesn't mean much for the same reasons. You can have any arbitrary reason to act or think any particular way; possessing an explainable rationale on its own, even if internally coherent, is not necessarily one way or another.

It may have been backdoored during training, potentially causing it to randomly start wreaking havoc at runtime, possibly in a clandestine manner (e.g. sneaking in bugs into generated code).

On the more general side, it's a bit hard to describe, just like it is hard to describe when someone "Googles bad".

They ask self serving questions, underspecify their requests, omit crucial context that the agent is blatantly not going to have access to, or subtly misdirect the agent. They expect the agent to figure out everything: you'll never catch them write a prompt longer than one or two sentences. They never steer the agent or look at the CoT traces.

My boss being a particularly poor case: he apparently has the habit of arguing with the agent, as if it was a person, as if there was any merit to that. Starts being a dickhead with it, shouts at it, what have you. Was flabbergasted we don't.

On the more practical side, they have zero mental model of the harness they're using (Copilot Chat in VS Code). They're surprised when the cheap-ass Auto model, which is almost always some beyond-demented version of GPT, does stupid things. They have no concept of skills, zero understanding of what an MCP server is, haven't heard of lifecycle hooks, agent memory, the various fs scopes (session, workspace, user). No concept of how to have the agent inspect its own debug logs for higher accuracy action provenance.

This also snowballs. Having to give them a stock config is one thing, but even beyond that, you won't see them experimenting. The MCP you're using doesn't support some action? They'll never interrogate whether the underlying scoped OAuth token or bearer token does support it, and they'll never ask their agent to patch the functionality in. They'll not consider the various user flows it can perform on their behalf. They'll not string them together into end-to-end automated workflows unless you explain it to them this is possible, and even after that, they'll just kind of ignore it. They'll never build tooling, extend the harnessing, etc.

Whether this has more to do with age or just disinterest-induced lackluster adoption, up for opinion.

Pretty sure it already exists, and it's the same as always: the younger you are, the better you adapt.

It's painful to watch my older colleagues use their agents, and they're not even that much older. Like they were intentionally trying to sabotage themselves sometimes.

They're getting better, but the time it takes for them to pick things up is just significantly longer, not the least because they're kind of just throttling themselves in addition.

Good thing that there's not much to pick up on at least.