HN user

Grimblewald

1,361 karma
Posts3
Comments778
View on HN

And I have evidence that anthropic distills from openai, moonshot, alibaba etc. So unless anthropic holds itself to the standard they're implying should be followed, why should I care others did to them as they do unto others? Seems like a nothing burger. Has the same vibe as a bully crying foul because they got hit back. Also, if distillation is what got K3 to where it is, why is it better in many areas? Also, big fucking kudos to k3 team for putting this together do fast givne how new fable access is, if it is true and had a meaningful impact. The real story there for me becomes one of extreme competence.

I get resistsnce. GDPR is good in theory but implementation has been shocking. Dark patterns in banners, banners when not required (ie no relevant cookies) etc. the gdpr banners as a whole make the internet a shitter experience. I categorically refuse work on anything requiring gdpr specific compliance and recommend against any practices which would require it on work I accept, and walk if demanded anyway.

I'm lucky in that I mostly work in areas where it really has no value add, and so far always get my way, but fuck I wish GDPR banners had some form of standard that didn't wreck internet experience in general.

Was it too hard to require the already opt-in nature of the cookies to no fuck with average user experience? no one in their right mind clicks accept all.

Was it too hard to enforce some default stance one can set?

Was it too hard to not fuck everything up without fixing the issue?

apprently. so now i have to constantly read and interpret information I have zero interest in to consistently hit the exact same button, to access information I _might_ be interested in. all of which is an excercise in dark pattern gymanastics so draining im near ready to swear off the open web entierly.

Qwen 3.8 2 days ago

for usable local layperson applications, I think qwen models are king. However, at hosted/frontier scales I'd agree for code. However for visual comprehension etc. for me qwen is the top of the line, and it isnt even close.

Qwen VL models nail tasks frontier models dont even get close to acceptable on.

and claude will call itself chatgpt etc.

nothing new, all ai labs are immoral and not bound by any reasonable oversight or ethical constraints. All outlaws in their own rights on that front. Absolutely none of them have true rights on the matter of being distilled from given historic and continued behaviour. I'm not sure why this is a talking point at all? We know AI companies steal, the least interesting behaviour among this is them stealing from one another.

For me, a far more interesting and important point of conversation on this matter is anthropic buying rare or evwn unique books, processing them for training data, and then destroying the books for others cannot use it as well.

Permanemt destruction of priceless primary source materials is so many leagues beyond copying a copy that I cannot fathom it even registering as a discussion point.

nah, things are worse. Basic userfacing components in windows frequently shit the bed. File explorer in W11 is near unsuable for example. grim state of affairs. Problem extends to many other platforms at scales where its not a quirk here or there, but rather, I can not longer use system instability and unexpected behaviour as a signal that maybe something is wrong that needs fixing.

Bugs and errors were rare and spaced out enough in the past, that when they happened I'd wonder if hardware was on it's way out, or if I broke something. Now, I simply assume software is fucked and I am almost always right.

Nokia is in for a massive comeback through the graphene collab. I feel it in bones. Normal people at work are talking guillotines and getting these lovecraftian fucking ghouls out of our private affairs. If they nail the grapheneOS collab launch and anouncement, they'll be back in a big way.

I had a similar experience, only, I was my natural way of talking.

e.g. one pattern I had/have is

<scope problem> the good news is <solution by way of analogy> is aviliable. The <constraint/requirement> is loadbearing though, ...

a near automatic script I rattle off in discussions/consults. When you solve similar problems several times, you figure out what works for communicating things you stick with it and you recycle/polish.

problem is, when for whatever reason, that pattern ends up as part of the core statistical distribution a model uses. You could royally fuck ones life up if working for a "frontier" lab, by simply finding a person with an acceptable speech rythem and cloning it, making it synonymous with ai slop. you'd destroy that persons image every time they open their mouth without them realising it. I for one started randomly getting quite hostile reactions from software devs who would be exposed to more llm putput than others.

Imagine an AI lab steals your voice, and uses it to scam call folks all day. Now, every time you call anyone, you are met with an immediate hangup. you'd have to put on a fake voice just to get the call to stay connected.

I love ai, and what it promises excites me, but as usual, humanity has a way of taking a cool tool and fucking it up royally. maybe the solution is to simply pepper slurs into everything one writes to blacklist ones content from training.

I second this, I read too much AI slop already so when something triggers that part of my brain, at this stage I immediatley lose the capacity to engage outside of work, largly because it feels like work. Scrolling through, this article looks like it holds useful info. Info i'd likely love to engage with, but realistically I cannot force myself to spend my weekend reading more ai outout, even if human seeded.

people like him exist today. people who will do things the scientific community knows is dangerous for pay. Due to a lack of morals or intellect required to understand rammifications is unknown.

Midgley didnt invent most if any of what was packaged up and commodified, he just appropriated research at the time into products without a care in the world. its for that reason that me and many other scientists simply dont publish some results anymore.

The mind it takes to discover and formalize a thing, will understand its problems. The mind it takes to commodify something once formalized, can be many steps down, and so, is unlikely to be capable of seeing the problem.

Therefore, anything which can cause serious issues wont be released. why? because even if we as the inventors go against the profit narrative, we'll be destroyed. people only listen when there's money to be made, but when we say "ok, that's eboygh fossil fuels, we need to slow down" funding gets cut, reputation and character gets attacked, and you largly get deplatformed.

there's already captcha systems like this and are already easy to beat. You wouldnt buy much time, if any, with this. the reason is that all it takes to defeat is a human style temporal averaging vision system, take any 4 connected frames and take the average, text stands out and even quite bassic llm's can read. works not just for this but for a slew of other "ai" combating methods, so likely scrapers and the like wouldnt be slowed by this at all. Decreasing SNR to force longer frame windows to get avceptable snr for reading makes it harder for humans than it does bots where dynamic temporal averaging via tool call is trivial. run tool, it temporally averages until snr in the region of interest improves, measured by a plateu in change of importance of high frequency terms in fourier space, since if the moving background has been blurred to a smooth gray, the noise text stands out clearly in contrast. Not something an LLM could solve when asked, but now that I've made this comment publically for nothing other than a few internet points and ego stroking, it's a matter of time before some of the larger llms start suggesting this solution. He'll claude probably already will, since I got it to write code for this series of tests/experiments during early opus 4.x era. and I know other ideas LLM's were shit at that I discussed heavily with claude ended up being the go-to recommendation a generation later, despite no "public" discourse on the matter.

I feel like fable is simply several 4.5s strapped together with consensus voting on next token.

Outputs i've seen so far are on par with my tests for 4.5, where 4.6+ were consistently regressions on 4.5 and their predecessors. One notable improvement being significantly lower retries to good output (1.1 avg. Vs 1.7 prev. On harder tasks)

given all the smoke and mirrors and OAI style fear-hype, it wouldn't surprise me if they intentionally degraded opus 4 for a few iterations, so they can resell "coke classic" at a markup with a minor quality of life feature put in, but charging way more than just re-attempting a poor output would have been previously.

unless anthropic starts acting in the image they claim and starts contributing to research, we'll never know either. Ultimately, the secrecy in how and why things are done would mostly be beneficial to this kind of buisness practice, since as it has always been, the moat is the data not the tech, so I cannot imagine what they hope to gain from the recent uptick in paranoia, jealous guarding and secrecy other than trying to huck a previous peak performance model as an imorovement when really, it is simply coke classic.

Eh, let em. If the US economy wants to self-sabotage, let them. Its a dying empire, lets its fall be hastened. I'm ready for china to fill the US vacuum. At least china controls it's billionairs (see jack ma saga) rather than the inverse.

Leanstral 1.5 21 days ago

I did get a refund quite quickly, after I finally figured out how to contact their support, which in some stroke of cosmic humor took deepseek to figure out because using their website left my trapped in a dead-end loop. Asking mistral for help lead to links that all 404'd or the same useless help section / faq page.

Them strugglig to prevent others from doing what they themsleves do, is my problem to a degree where I must pay for the pleasure to getting sabotaged? Are you fucked in the head? In what world are we "being fair" by claiming that? It takes a problem anthripic has hallucinated, and makes its consequences mine for no reason. Anthropic chose to create both this problem and it's consequences, why am I wearing it?

If I decide I dislike words starting in S, despite myself being a prolific user of words starting in S, and smack some child who'd never even heard the rules, simply for saying "sorry", simply because some people say "shit" and it makes me mad, despite it being my most used word, are we "being fair" by saying "to be fair, managing curse words is a difficult problem"

no right? Insane take.

Changing well documented requirememts and functioning code to do subtly incorrect things, over multiple areas, for no reason. When confronted you get the usual, "youre right to push back, there was no reason to touch that and it did make everything worse"

Why do I suspect faulty trigger? The things wrecked weren't wrecked by even much less capable, including local models, while reverting everything to pre fucking up and asking claude again led to similar results. Once on whatever shitlist that was, claude also failed tasks it previously aced when given the exact same prompt and project files. I attributed it to opus 4.6 being a downgrade which people always pushed back on, claiming they thought it was better, but i had empirical proof, it couldnt do tasks 4.5- could do quite well. Now all of this? It's clicking all of a sudden that its quite plausible i got flagged somehow and ended up with an intentionally degraded service.

So, claude is off table for me these days, and deepseek gets very deep git commit read-throughs with every file even being read being carefully monitored. This obviously ruins the "agentic" promise, but the reality is we cannot trust these fucks (being the companies). The irony is deepseek now queries claude on problems its stuck on for input via openrouter, with mild success (the gap really isnt all that big, if its even there), before escalating to me for input on direction on solving a given problem. So now chinese models get more claude training data, not less. In fact, they get training data on frontier ML stacks and problems. Anthropic did that to themselves.

So block people, instead of having false positives be secretly fucked over, and having them pay for the pleasure?

Given the hidden model degradation of fable and now this, what makes you think this is where it stops? That's just what we know about and there's clearly a long-standing and deeply rooted malicious intent here.

I've had Claude fuck over clean well documented code-bases for no reason, and there's a good chance this is due to some faulty trigger. Luckily I don't trust these things one bit, and claude only ever runs in an isolated VM, however, I am pissed I am being made to pay for their errors in detection and waste my time fixing things I apparently paid to have fucked up.

That's unacceptable conduct. It's witch-hunting. Punishment and attacks on you for things without real proof. That isn't right.

Leanstral 1.5 22 days ago

Got curious, sign up, add money to account, try to use. Can't, it's a labs model. Fine, let's enable labs. Can't, unspecified error. Fine, lets contact customer support as instructed, can't no customer support, just a half-assed FAQ, that seems vibe-coded and searched poorly, totally irrelevant answers coming up for all queries tried. Then it hit me:

If AI makes good customer support, then why does no AI company use theirs to provide customer support?

I might not use all the same terms as you, but i largly agree. However llms live in a world almost entierly divorced from our own. Ours is physical, it's is human information corpi. For tasks suited to its world, and llm is as "human" as I am in my own domain. Same same but different. Intelligence is a fundamental trait of our universe, not a surprise.