HN user

gozucito

61 karma
Posts0
Comments38
View on HN
No posts found.

This is the post Nemus was replying to:

Perhaps even more importantly, the current frontier LLM models are self-admittedly the product of enormous quantities of copyright infringement and even less savory inputs, so calling them out for distilling the fruit of that tainted tree reads as highly hypocritical at best.

Context is important. And in this context, their argument only mentions creativity when it belongs to an AI lab. That omission is the blind spot I pointed out. Bottom line is whether or not Anthropic are being hypocritical and yes, they most definitely are, regardless of any attempted sophistry.

There is a reason courts want you to tell "The whole truth" and not just "the truth".

There is an even higher level of creativity in creating books, songs and all sorts of art used in model training though. That's your apparent blindspot.

There is no world in which me vacuuming the entirety of human knowledge to make a genai model is ok but hoovering my model answers is not. The hypocrisy is stunning and risible.

Now if you go and make a model based on purely synthetic data and not a single work made by humans, you would have a valid point.

Good news, and necessary to keep me on the MAX plan since Opus 4.8 is now 3rd best after Sol and Kimi 3 for SWE work.

The false cyber positives when doing normal coding tasks remain a big issue. They were not a big deal for weeks, as long as I stayed away from explicit hardening work or security audits, until recently.

They are a prime weakness for competitors to exploit.

I've reas the tweet and...Am I mistaken or is the plan "Do what Anthropic has been doing and advocating for this past year" ?

My problem with this plan is that it seems to have faith in mankind, despite the fact we've consistently failed to rise to the occasion for decades now. The last time we rose to that occasion was probably when we eradicated smallpox, many decades ago.

Ironically, nowadays, many people don't even trust vaccines. A dramatic regression.

Could you please give more detail about this for those of us less informed of EU things? What corrupt things has Ursula done? Or what has she been accused of, exactly? How long has she been cancelling court dates?

Who is the engineer of all this evil? Her? Someone else?

GPT-5.6 13 days ago

The meat of the report for SWEs:

SWE-Bench Pro Sol: 64.6% Fable: 80% Opus: 69.2% (!!!!)

So, it still trails Opus, significantly, and is not a next-gen coding model like Mythos/Fable 5.

Disappointing to say the least, but somewhat expected.

Grok 4.5 14 days ago

They get to compare their model to the old ones from the competition.

In this case, ChatGPT 5.6 Sol / Ultra releases tomorrow, so today is the last day Grok can compare Grok 4.5 to Codex 5.5. If they did it tomorrow people would point out they're comparing themselves against old models.

2 to 5 tabs in warp. Kinda want to figure out how to properly use iterm+tmux to have the boris cherny experience. already used both but there were issues with tmux with either scroll back, or copy paste or other similar things that get broken, even after mesing with settings.

i use git worktrees in different tabs as needed.

i have git push hooks that audit the code diffs for security issues by 2+ frontier models For code quality with a FAIL/CLOSED condition where both have to give the OK.

i have to do a pass and ask it to shorten the code, remove unnecessary comments and excessive exception gathering, etc. Generally cuts the code by half. The process is repeatable.

i just use claude code or codex with minimal plugins (HUD, frontend design). I would do even more if I had 10x or 100x the tokens and/or token/s available. I spend a lot of time waiting on 5.5 or Fable 5 to work, even when multi-tasking.

I spend the downtime writing detailed follow up or unrelated prompts.

After thinking about it, my advice is to start a new company with the rest of Mullvad folks who don't support this:

Markus Allard has voiced support for the idea of large scale remigration on multiple occasions. On one occasion, in a debate with a Liberal member of parliament, he asked why the Liberal party "does not wish to deport 100 000 social welfare-Somalis?"[19] In the same debate Allard also claimed that "Sweden belongs to the Swedes. We have to make sure that we take care of our own damn people and we must deport these damn parasites who sit and live at our expense."[20] Regarding deporting those born in Sweden, Allard said in a podcast that "They will also be forced to leave, even if they are born in Sweden, because they have no natural connection to Sweden. They are not Swedish."

I will be happy to continue giving you my money then. Keep in mind the situation also sucks for all of us who have been recommending Mullvad for year, only for our money to go towards that kind of hatred. It is a betrayal. Now I have to start advising against it and explaining the money goes to neo nazis.

You're in a a shitty no-win situation.

Your choices are to either end a 20 year friendship and enormously fruitful collaboration or defend someone who has demonstrated quite clearly that he holds repulsive, despicable beliefs.

Imagine you were of immigrant background and your closest friend gave millions of dollars to a party whose sole platform is to strip you of citizenship and kick you out of your country.

You and I are a trillion times more likely to choke on food than die of an AI-created nerve agent. Yet I doubt anybody wastes time worrying about that. Don't trust your faulty, misleading human gut. It is woefully outdated and dangerously obsolete in this modern world. Trust mathematics instead.

a worse experience for there paying customers

Actually, the legit buyers experience is better because this bypass is not a "proper" crack

1. They have to disable Windows security features before playing

2. Reboot their PC twice (before and after)

3. They're still running Denuvo code, same as legit buyer

Legit buyers experience is thus significantly better than pirates.

You're right. I didn't scroll down. I wonder why they didn't update the top cards that everyone see. They do it for claude Cowork but not claude code? That is not very transparent. How does it make sense? It's not like claude code is too niche to be included, it's in the main app and I know multiple non-techie people who use it.

The reality of losing TSMC is no joke either. I remember Covid times when many G20 leaders went to Taiwan begging for some chips so that they could keep exporting cars and other things that need computer chips.

But what would you rather have? 2000 Shahed/Lucas drones or a single F35? Same cost for both.

The saying "Quantity has a quality all of its own" is not obsolete in 2026.

I believe the lack of quick evident profit increases are partly a failure of imagination or a failure of understanding that AI agents are different from people. More impressive or faster in some ways, but much much less reliable in others.

The evolution of harnesses like claude code or open cause, and metaharnesses like Ralph loops, gas town, claws, etc. Will progressively allow for gradually better results and abilities even if models stopped evolving, and if the Mythos eval numbers are to be believed, there is still no hard ceiling to be felt yet.

At the same time, small models that can run on PCs VRAM/UNIFIED RAM have like Qwen are becoming more useful.

I predict that having more and more loops within loops within loops and layers of cloud/local models of different capabilities will solve a great many limitations of LLMS today...at the cost of speed and token count.

We've never had a tool that is at the same time so unreliable and complicated as GenAI before. It will take us a minute to figure out how to use it properly.

I think the suspicion regarding skills and plugins is fair and logical. And it is absolutely the case that some use significantly more tokens.

with that said, on my 5x plan, I could have multiple sessions working and the limit was far away. Around when you introduced the whole more tokens during off-peak hours and fewer tokens during working US hours, Even with a single session, using no plugins at all (I uninstalled OMC) I run into limits very often.

I have not performed any rigorous tests but it feels like I have about 25% of what I used to have or less. This is all without using teams of agents, or ralph loops or anything like that. Just /plan and execute in a single session. I have restored the /clear context before executing plan to try and mitigate things. I will also try the 400k context since, in my experience, the 1M tokens have not made Opus 4.6 noticeably smarter for my small webapp use-case.

Best of luck to you!

ps: whenever you introduce a change, please make it optional AND ask the user about it at first. Don't just yank things suddenly (like the /clear context and apply plan option.) as I spent hours trying to figure out how I broke it before I saw your note and how to re-enable it.

Waymo Safety Impact 4 months ago

Is it a bullshit stat though? it's not like you or I can go to a different dimension where all drivers are healthy, fully awake, undistracted, sober, competent, etc.

Simplicity brings us closer to truth — Occam's razor has underpinned the development of our species for centuries.

I keep thinking of emergent complexity. Even starting with very simple rules and components, the amount of complexity that arises as a consequence of ever rising interactions can boggle the mind and seems to validate our current predilection for elegant and succinct laws of physics to be enough to model the universe.

Coincidentally, LLMs being so good at coding that it became the #1 source of income for Anthropic is one such example of emerging complexity from deceptively simple ingredients:

A giant pile of matrix multiplies, next-token prediction, and enough data somehow climbs the ladder from autocomplete to writing code well enough that people will pay $20-200/month per seat for it. It is completely bonkers.