Maybe true, I'm not sure, but this isn't the conversation where the counterexample was found.
HN user
furyofantares
Depends on what you mean by determinism.
I find 5.6 Sol will pick a direction and aggressively pursue it in long horizon tasks. I've got it porting an older game from Pascal to my own game framework. I gave it some instructions on doing a full 1:1 port. I had already ported the game rules and multiplayer support to a very different system than the original, but all of the UI and features and such needed doing, and needed to be integrated into this very different system.
The first attempt it had files tracking both hashes and semantic hashes of every individual line of Pascal code, mapping to what code in the port is responsible for that line of pascal. It had written tooling to parse Pascal in service of this for some reason as well. I asked why it was doing this, it said it was because the reference code is .gitignore'd so it needs to thoroughly maintain the mapping in case someone working on it does not have the reference code, or in case the reference code changes.
I started over with Claude 5 Fable, and with better instructions about focusing on UI. I got a long ways with that before I hit my weekly limits, and switched back to 5.6 Sol. It picked up and did a great job for a while, although it interpreted my desire for a 1:1 port to mean every pixel must be perfect. I let it go on and it did some good work in that regard, but then it decided it must perfectly reproduce a hash of the game state in various replays & etc. It had clearly lost track that I didn't need game rules ported, and it found that the original code produces a hash of the gamestate for various purposes, so it ended up reproducing this in a game that represents its state totally differently. It also rolled its own version of Pascal's RNG source in order do this. I've burned through 3 weekly limit resets on this to see if it's actually going anywhere, and it has found some bugs, but man it is going hard in a direction I didn't even ask for.
Getting the most out of out agent. Knowing how to give it context that will cause it to succeed. Knowing what things it will succeed at and what it won't. Knowing how to tell when it's failing.
Figuring out when it's appropriate to fully wrap your head around the details of a solution, and when it's appropriate to let the agent handle that and only know a bunch of properties of the solution. Knowing when it's appropriate to try 10 things, throw them all away, then go deep. Knowing how to learn from these things so you can make something better, not just faster.
It's true that there are new skills to acquire (and that those with talents for the old skills may be less talented with the new skills).
It's absolutely not true that it "just" moved the difficulty around. If that were true then I'd be just as well off continuing to use my decades of programming expertise just as I always have; but the reality is I can get more done, and get better work done (depending on level of vibing), than ever before.
Thing is, though, making programming easier doesn't mean programmers will work less hard. That's how competitive markets work. LLMs made programming easier. Capitalism prevents workers from capturing that value.
If you're only using the frontier providers, it's nearly trivial to support switching between the 2-3 of them that exist without using OpenRouter.
If you are ALSO routing to open-weights models on OpenRouter, then sure. But what we've come back to is that almost everyone using OpenRouter is someone who wants to use open-weights models, and some of them also want to use closed models, so of course open models take all the top spots.
There very little reason to use OpenRouter to access closed models, instead of just using the provider directly.
Once you get to the LLM judging them, you've given up, and might as well just prompt the LLM to make up funny item co-occurrences.
Nothing wrong with giving up though, it's a hard problem for all the same reasons recommendation engines are.
no
Are you also calling it literature if it's not written down?
The idea that it only reveals missing abstractions and that this is bad, is pretty off the mark in a lot of ways.
First off, there's plenty of cases where it make no sense for me to spend all my time abstracting everything possible in my pipeline. Maybe I'm a game developer who doesn't want to spend all my time abstracting away all the initial scaffolding for game prototype and just make a dang game prototype.
Or maybe the abstraction would needs to much confirmation to not really be any different than code (again, games.)
Or maybe the LLM is indeed a perfectly good abstraction for such a task, reliable and customizable.
Just keep writing abstractions all the way up the chain until you're so abstracted that you take english input and can product whatever is asked for, and, oh right.
No, I typically want to target either mobile or PC first - but I strongly want stuff to run well in the browser too, for early playtesting and for demos at the very least.
I released 3 games my friend designed, and built a game framework in WebAssembly for building them. Some 9 or 10 months ago I asked my friend (legendary game designer Mike Elliott) if he wanted to work through his backlog of designs he has that he'd love to see brought to life. I'm very protective of my family life and dedicate a lot of my time to it, and have a normal full time job, and so LLM-agents really enabled being able to make stuff like this happen in a reasonable timeframe just working on and off in my free time.
I started them with ebitengine (Golang) but got somewhat frustrated with its web builds, and so built my own thing for small games that I want to work great on mobile or native PC, but also on web. I call it NanoGame, the host is written in Rust and the games are AssemblyScript. I've ported a number of other small games I had written to it as well, but haven't released any.
Two of the games I released a couple days ago were actually the ebitengine versions, but have partial ports to my framework, and the third I released the version using my stuff.
https://scramblequest.app - ebitengine, word search game where you slay monsters with the words, has a long campaign as well as a daily challenge and unlimited play
https://wordpeek.app - ebitengine, another word search game, this one reveals pieces of a picture and your goal is to guess the picture
https://playsilhouette.app - my own framework, this is a simple matching/hidden object(ish) game, more for kids
I also made a little umbrella site for them at https://playthese.wtf
I have a personal game framework that I have LLMs write games in, which is in AssemblyScript. AssemblyScript is certainly closer to TypeScript than it is to WebAssembly, but it's still this thing where the host shares some big chunk of memory with the script and you pick some memory locations to read and write as your means of exposing APIs, and there's not a lot of training data on games written in AssemblyScript and even less in my game framework specifically (none) - and the LLM does an excellent job.
Also I want to highly recommend to anyone that has game designer friends (no matter the kind of game, tabletop games too) ask your friends if they have a folder with designs they're exited about that never got implemented. Most of them have this. And it's suddenly quite cheap to get something out there, I bet they'll be excited to see some of their ideas in action.
That's not true at all, human tends to come through (in a way I didn't notice pre-LLM) with varied tone, with opinions injected, and with varying degrees of weight to different components, in all but the most egregious of examples (linkedin motivation pieces, apple marketing speak).
LLMs have a bunch of tricks for dressing up their infodumps, but they are almost purely infodumps, and no real opinion comes through. There's no sense of more importance to one statement or the other, it's all monotone (and usually over the top.)
A problem with timers in word games is players vary A LOT, like a lot a lot, -- no, more than you're thinking now, even after I said that -- in terms of how fast they are at word games.
So a timer needs to either accept that a lot of interested parties will be turned off by it - or must be designed in a more accommodating way.
Sure, maybe - but just because 123,000 is cheap doesn't mean it's OK to make your headline "You can buy a house from the government for $3,000" if the reality is that it's 40x as expensive and you don't actually get a house, you get a demolition project you have to complete in order to use the land.
I'm not sure exactly how you're misreading it, but you are.
The Nothing isn't executing all the taps, some are blocked by the animation. It is responding visually and haptically to all of the taps, but some are blocked from doing any work by the animation.
You also said the Nothing was 6 taps but I'm not seeing anywhere the article says that. I believe it was 8 taps on both.
The person with the vehicle is who should ultimately be held responsible in the case of an accident, but I also find it absolutely wild when I venture out into the city and see people on their phone with headphones on crossing the street when the walk sign comes on without so much as glancing in the direction of traffic.
Sure, although personally the comparison would be to delivery from other stores.
But they don't avoid solving it, they offer it by partnering with instacart.
This is about 80% spot on, but the last 20% fails to mention that you can avoid the in store experience if it isn't for you, and in fact get the stuff you want delivered to your door in a short period of time, using services like instacart. Costco even partners directly with instacart for same day delivery. You can use your membership to get same day delivery shopping on costco's website and they will use instacart to fulfill it for you. Or you can use instacart directly, in which case you don't even need a membership yourself.
"in general" in mathematics means "in all cases, without exception", rather different than the normal usage where it means "usually, but not always".
If a mathematician using the mathematical sense of the word general says "it is not in general possible to tell if a program will halt by inspecting it", they're talking about the halting problem, even if you've looked at lots of programs where you can tell if they'll halt or not, and your experience might be correctly described using the normal usage of the word with "in general I can tell if a program will halt or not".
It's June 30 and the ban has been lifted, I'm gonna say Polymarket's 32% chance of the ban being lifted in June was far and away a better guess than bemoaning that non-Americans will never access an American model better than Opus 4.8.
Virtual, not Very
Just another data point about this, not an argument for (or against) blocking a TLD, but personally I wouldn't register an .xyz again, due to what I presume is related to them having to fight abuse. I still have one site on there and have migrated another off.
My domain was flagged for abuse (it's a static site with a daily word game, no ads or anything else) and the TLD took it down. Not my registrar or host, the TLD itself. There was no communication on this, it took some effort to work out what even happened, and appealing was a pretty blind process of claiming to have fixed the issue and issuing proof (which felt a bit strange to fabricate proof that it was fixed, since no issue existed to begin with - I sent a screenshot of the page or something, I can't recall) and hoping they'd unblock it, with no communication at all beyond a place to send such a claim.
They did unblock it, and while I am sympathetic to them having to fight abuse, I still moved away from them.
Why would AirPod prices increase?
Having the LLM rewrite your writing always, in my experience, loses any sense of what the author cared about, everything is given monotone weight instead of somethings more important than others.
But does your writing have to become blunted after a red-team _review_? You're running a technical blog and want to detect technical errors, that's great - a red-team review sounds like something that would tell you the technical errors and you'd fix those up without rewriting your words.
Is there a contradiction here? If resold Opus tokens are sold at a 93% discount, you can be a lot cheaper than Sonnet while also a lot more expensive than resold Opus tokens.