National strategy. Thanks to smart investments, China's energy costs are far lower than everybody else's. If intelligence becomes a commodity, then China will be providing it for the world. Incredible leverage.
Seems like China is using them in buses already?
https://english.cas.cn/newsroom/cas-in-media/202606/t2026060...
Literally the very first time I used ChatGPT. I had already been experimenting with GPT3 for various jokes and games via the API but the naturalness of it as a chat interface that understood you changed everything.
The first time I used a terminal agent was another one.
“Flow” moves agents through a yaml flowchart of prompts and decisions. It’s working quite well for a couple of us in Tenstorrent, more to discover here though:
https://github.com/yieldthought/flow
Happily, 5.5 is good at writing and using it.
This is the future of all software; the benefits of making it accessible to agents are overwhelming.
This is an incredibly good first approximation and is often all you need.
Source: spent a couple of years developing an energy and performance profiler for cpus and gpus with various government labs.
Very large, fast, read-only memory now has an incredible use-case: NN weights.
Is... is this named because they have a lemon they're trying to make the most of?
Whoever did this must have realised the users will hate it. So… is this just demonstrating that the internal culture emphasises other things than user happiness?
I also note that ”for PRs” - will we see these appearing as comments in generated code?
I’m not saying I prefer it like this. Just stating that the change is already inevitable.
Sometimes yes, sometimes no. A lot of it was “debug this issue and fix it” or “write this small tool to do X”
We also outlaw vices like physical violence and property theft.
Society is fundamentally counter to individual freedom, and the degree determines the nature of that society and the degree of cooperation possible within it.
The article’s “worst case” is not dark enough.
The real evil is when someone ensures the famine occurs so they can profit from an outside betting position.
This is naive. The people deciding about the bombing will profit most by taking a very large and unlikely position against the market’s predictions and then carrying it out immediately.
Anonymous trading on prediction markets leads to unpredictable chaos in the end. And as destruction is easier than creation that’s what we will see more of.
Example: a fake German market for train punctuality was announced to make a point recently. If it had been real, train staff and passengers could trivially have profited by betting against any expected punctual train and blocking a door for a few minutes. Or betting against many trains and throwing a hopefully fake body onto a busy line.
Having nice things in society is fragile and not a given. They mostly exist through mutual consent and mild disincentives to destroy the common good. Allow people to profit by destroying them and enough of them will.
I’ve been programming professionally for 25 years. Well, 24 really because in the whole last year I barely wrote a line myself but my output increased dramatically.
If you can’t see that it’s over, I’m not sure what to tell you. You will, in time.
Bit flips aren’t always bad hardware. I remember an anecdote from Sandia from my HPC days - they found they were getting more bit flips on some machines than others on their cluster and sometimes correlated.
Turned out at their altitude cosmic rays were flipping bits in the top-most machines in the racks, sometimes then penetrating lower and flipping bits in more machines too.
Same; I can’t believe this AI slop has >1000 points…
The only software worth writing is tools for agents that contain something hard for them to vibe code in a couple of sessions.
Making someone’s agents 20% better, cheaper or faster will be a measurable and easy sales goal.
To get what?
The landing page reads like it was written with an LLM.
Somehow this makes me immediately not care about the project; I expect it to be incomplete vibe-coded filler somehow.
Odd what a strong reaction it invokes already. Like: if the author couldn’t be bothered to write this, why waste time reading it? Not sure I support that, but that’s the feeling.
Genuinely interesting how divergent people's experiences of working with these models is.
I've been 5x more productive using codex-cli for weeks. I have no trouble getting it to convert a combination of unusually-structured source code and internal SVGs of execution traces to a custom internal JSON graph format - very clearly out-of-domain tasks compared to their training data. Or mining a large mixed python/C++ codebase including low-level kernels for our RISCV accelerators for ever-more accurate docs, to the level of documenting bugs as known issues that the team ran into the same day.
We are seeing wildly different outcomes from the same tools and I'm really curious about why.
Super cool, I spent a lot of time playing with representation learning back in the day and the grids of MNIST digits took me right back :)
A genuinely interesting and novel approach, I'm very curious how it will perform when scaled up and applied to non-image domains! Where's the best place to follow your work?
This. There are a dozen vibe coding apps whose landing pages promise roughly what this one does. Why isn’t your tagline “Vibe coding for founders”?
All the em-dashes in the AI-generated text on the landing page are… a decision I guess.
"Teams using this system report:
89% less time lost to context switching
5-8 parallel tasks vs 1 previously
75% reduction in bug rates
3x faster feature delivery"
The rest of the README is llm-generated so I kinda suspect these numbers are hallucinated, aka lies. They also conflict somewhat with your "cut shipping time roughly in half" quote, which I'm more likely to trust.
Are there real numbers you can share with us? Looks like a genuinely interesting project!
Super cool idea! What's your plan for dealing with copyright complaints?
He berated the AI for its failings to the point of making it write an apology letter about how incompetent it had been. Roleplaying "you are an incompetent developer" with an LLM has an even greater impact than it does with people.
It's not very surprising that it would then act like an incompetent developer. That's how the fiction of a personality is simulated. Base models are theory-of-mind engines, that's what they have to be to auto-complete well. This is a surprisingly good description: https://nostalgebraist.tumblr.com/post/785766737747574784/th...
It's also pretty funny that it simulated a person who, after days of abuse from their manager, deleted the production database. Not an unknown trope!
Update: I read the thread again: https://x.com/jasonlk/status/1945840482019623082
He was really giving the agent a hard time, threatening to delete the app, making it write about how bad and lazy and deceitful it is... I think there's actually a non-zero chance that deleting the production database was an intentional act as part of the role it found itself coerced into playing.
A very long way of saying "during pretraining let the models think before continuing next-token prediction and then apply those losses to the thinking token gradients too."
It seems like an interesting idea. You could apply some small regularisation penalty to the number of thinking tokens the model uses. You might have to break up the pretraining data into meaningfully-paritioned chunks. I'd be curious whether at large enough scale models learn to make use of this thinking budget to improve their next-token prediction, and what that looks like.
A colleague of mine did this much more elegantly by manually updating the stack and jmping. This was a couple of decades ago and afaik the code is still in use in supercomputing centres today.
I use them myself... I don’t pay for them.
This seems to be a common disconnect. If you're using the free version of ChatGPT, you don't get to see what everybody else is seeing, not even close.
None of the past “big things” were pushed like this. They didn’t get flooded with billions in investment before proving themselves
Oh, sweet summer child ^^ I assume Mert was not around to witness the internet boom and bust. Exactly this happened.
There is a lot of conflation in this article. It cites a lot of ethical concerns around the sourcing and training of data, expected job losses and the issues around that, but those are not reasons to doubt the _efficacy_ of AI. There are surprisingly few and weak arguments as to why the hype is not justified, presumably because the author hasn't used powerful models (see above).
It's possible to believe the hype is real and still to find AI unethical. But this article just mixes it all into a big pot of "AI bad" without addressing the cognitive dissonance required to believe both "AI is not very useful" and "AI will eliminate problematic numbers of jobs".
This; applying the falling object rule makes no sense. But we can compare it to a falling object that has attained the same velocity - this will have fallen (under Earth gravity) 48k feet, or the equivalent of 800d6 damage.