HN user

famouswaffles

6,419 karma
Posts221
Comments2,683
View on HN
www.appen.com 2mo ago

Benchmarking Subquadratic's latest model and SSA Kernel

famouswaffles
2pts0
arxiv.org 2mo ago

FutureSim: Replaying World Events to Evaluate Adaptive Agents

famouswaffles
2pts0
blog.alexisfox.dev 3mo ago

Hill-Climbing ARC-AGI-3. Human Performance with Opus 4.6

famouswaffles
2pts1
arxiv.org 4mo ago

Lost in Backpropagation: The LM Head Is a Gradient Bottleneck

famouswaffles
4pts0
www.smh.com.au 4mo ago

From one dictator dad to another: Monica's lost childhood in North Korea

famouswaffles
5pts1
www.smh.com.au 4mo ago

From one dictator dad to another: Monica's lost childhood in North Korea

famouswaffles
3pts0
antocuni.eu 8mo ago

SPy: An interpreter and compiler for a fast statically typed variant of Python

famouswaffles
276pts131
transformer-circuits.pub 8mo ago

Emergent Introspective Awareness in Large Language Models

famouswaffles
30pts4
epoch.ai 11mo ago

Quantifying the algorithmic improvement from reasoning models

famouswaffles
1pts0
www.sciencedirect.com 1y ago

Evidence of interrelated cognitive-like capabilities in large language models

famouswaffles
1pts0
arxiv.org 1y ago

Atlas: Learning to Optimally Memorize the Context at Test Time

famouswaffles
43pts4
deepmind.google 1y ago

Gemini Diffusion

famouswaffles
61pts7
arxiv.org 1y ago

Tails Tell Tales: Chapter-Wide Manga Transcriptions with Character Names

famouswaffles
2pts1
arxiv.org 1y ago

Over-Tokenized Transformer: Vocabulary Is Generally Worth Scaling

famouswaffles
2pts0
anokas.substack.com 1y ago

LLMs struggle with perception, not reasoning, in ARC-AGI

famouswaffles
2pts0
hkunlp.github.io 1y ago

EvaByte: Efficient Byte-Level Language Models at Scale

famouswaffles
3pts0
arxiv.org 1y ago

Tell me about yourself: LLMs are aware of their learned behaviors

famouswaffles
2pts0
arxiv.org 1y ago

Imagine While Reasoning in Space: Multimodal Visualization-of-Thought

famouswaffles
2pts0
anokas.substack.com 1y ago

LLMs struggle with perception, not reasoning, in ARC-AGI

famouswaffles
1pts0
ai.meta.com 1y ago

Byte Latent Transformer: Patches Scale Better Than Tokens

famouswaffles
6pts0
deepmind.google 1y ago

Mastering Board Games by External and Internal Planning with Language Models

famouswaffles
1pts0
arxiv.org 1y ago

Emergence of Hidden Capabilities: Exploring Learning Dynamics in Concept Space

famouswaffles
2pts0
gamegen-x.github.io 1y ago

GameGen-X: Open-World Video Game Generation

famouswaffles
4pts0
arxiv.org 1y ago

TokenFormer: Rethinking Transformer Scaling with Tokenized Model Parameters

famouswaffles
174pts33
www.youtube.com 1y ago

Kurzgesagt: We Fell for the Oldest Lie on the Internet [video]

famouswaffles
1pts3
arxiv.org 1y ago

Relaxed Recursive Transformers: Effective Parameter Sharing with Layer-Wise LoRA

famouswaffles
1pts0
arxiv.org 1y ago

Solving Global Lyapunov functions: open problem in mathematics with transformers

famouswaffles
3pts0
www.similarweb.com 1y ago

ChatGPT Topped 3B Visits in September

famouswaffles
2pts0
research.google 1y ago

Tx-LLM: Supporting therapeutic development with large language models

famouswaffles
2pts0
research.google 1y ago

Tx-LLM: Supporting therapeutic development with large language models

famouswaffles
2pts0

It doesn't matter much whether they are lying about how secure the environment is when they train and ship these models for other actors. Believe it or not, they're not spending hundreds of millions training these models for only themselves. Hacking Hugging face is an achievement on its own and the important relevation here.

Yes I do. Not mathematical results specifically but generally research results. I might even have articulated that criticism on HN. I don't know if I could search for it easily though.

Okay Fair, but then this is a general research issue and not really a Open AI issue.

The people desperate to demonstrate mathematical competence are the AI companies. The people discussed in the part of my comment you quote are "random people" by which I meant the "3d parties" in your original comment.

I'm not sure you got the point i was making. The point there was that if Open AI were running as many problems as frequently as you imagine they are then those 3rd party results should have been achieved by them, and if they're really so desperate to tell us how good the model is for math then why didn't they tell us ? Why have they only announced 2 results when they could have announced near a dozen by now ? Sure these 2 are in a class of their own, but some of the others are genuinely impressive in their own right and would certainly help that narrative you're talking about.

Either they're just not telling us and aren't as desperate as you imagine, or they're simply only interested/running models in a relatively few set of problems.

What, Erdős problems? It's hard to see how anyone except mathematicians, and then again only a few communities of mathematicians, especially care about those.

Erdős problems vary enormously in difficulty and significance. The fact that a problem is obscure to non-mathematicians does not make its solution unimportant. There are also many major open problems in computer science that most laypeople have never heard of.

That they started with Erdős problems? I'd have gone for a Millennium Prize problem, first. P vs NP, Riemann, Navier Stokes, those are heavy-weight results that would establish AI as the de facto approach to mathematics for the foreseeable future.

Solving the biggest, most famous open problems would "establish AI as the de facto approach to mathematics for the foreseeable future"? It would do a lot more than that.

That's exactly how research works in general, both in academia and in industry.

Then what exactly is the objection? Research normally produces many failures and incremental results before a major success. Do you apply this survivorship-bias criticism to every published mathematical result, or only when a machine contributed to it?

Actually, that's a good point but it's in support of my contention.

If they're running as many problems as constantly as you imagine then they did not miss all that. So either they're just not sharing it us which contends with your "desperate to demonstrate mathematical competence" or they have their eyes on a more curated set.

Or, to abuse Fermi's question, where is everybody?

If we had as many verified alien encounters as LLM contributions to open problems, nobody would be invoking the Fermi Paradox. Where is everyone? Right here.

It might also be helpful more than 'necessary'. The 2 most notable solutions have come from open ai themselves, but most of the 'LLM solves open problem' category are from 3rd parties doing their own thing.

This one was pretty impressive in its own right (probably the most impressive outside these 2), and the prompt is concise and basic.

https://www.scientificamerican.com/article/amateur-armed-wit...

https://chatgpt.com/share/69dd1c83-b164-8385-bf2e-8533e9baba...

The 2 most notable/interesting solutions have come from Open AI directly, but most of the 'LLM solves open problem' category didn't and has come from 3rd parties doing their own thing with publicly available models. I don't see why one would assume they're running models on hundreds of problems. Most likely they have a few problems they especially care about that they run on.

Is this something humans have been unable to do?

It's a famous open problem so yeah

There’s only so many people with the necessary skills to solve this. And you need these humans to choose to spend their time solving this, and not something else.

Sure, but that doesn't mean a lot of very skilled people hadn't attempted and failed to solve this.

GPT‑Live 14 days ago

People who do this kind of stuff are very irritating. You clearly have some problem with the work they do. Instead of saying and approaching that outright, you pass it in some passive aggressive fake bullshit. Makes you sound like the kind of person I would much rather not be speaking to, which is kind of ironic given your comment.

GPT‑Live 14 days ago

The Star Trek computer would suck at a lot of what people use GPT for.

GPT‑Live 14 days ago

Does video/image input still work with these duplex models?

garlic_enjoyer has already said valuable stuff, but you must realize that skeptical=/true. A lot of people simply don't know what they are talking about. I remember on one of the previous mech interp papers arguing with someone who just didn't even understand what the paper was saying and the experiments they had set up and so a lot of misunderstandings and wrong conclusions spilled from there. And it's kind of funny because you would certainly think he knew what he/she was talking about from how self assured it all was.

'Naturally' might not be the best word? Maybe 'Necessarily' would be better?

Regardless, it's something that happens in people. Have you not or seen someone else struggle to recall a specific fact or memory until phrased or induced in a certain way?

You probably could also say LLMs 'tend towards bidirectional recall' over the course of training as things that ought to be recalled both ways are reinforced to do so. In the above example, you will also eventually learn both ways with enough exposure even without explicit practice.

Recall isn't naturally bidirectional, even for humans. If you are learning vocabulary in a new language, it's common advice to practice both target > source and source > target. Doing only one-way often makes you much better recalling that single direction than both.

I think this is a bit disingenuous. Japan spent nearly all of the last 30 years needling deflation. If you take a look at the highest grossing movies of all time in Japan with and without adjusting for inflation, it barely changes. Do that for the US and it's an entirely different list.

Normal inflation for the last 4 years is basically still nothing in the grand scheme of things.

CursorBench 3.1 21 days ago

Cursor sessions are pretty much what composer models are RL'd on. This bench and the training data are/should be basically the same distribution.

There's another one that intrigued me greatly when i read about it years back. This was back when GPT-3 was state of the art. I had a lot of trouble finding it again but i did!

It's not an exact fit because the output is that of a tool rather than the model itself (though i don't think much would change if we had the model perform the arithmetic itself but altered answers similarly), but it was the first time I began to realize that just like the brain, these models have an expectation of reality that they work around. They don't necessarily 'trust' an output if it diverges significantly from this 'reality'. And that this disregard may be silent indeed (no reasoning or chain of thought here).

GPT-3 will ignore tools when it disagrees with them - https://vgel.me/posts/tools-not-needed/

Anthropic has some mechanistic interpretabilty research on this actually.

https://www.anthropic.com/research/introspection

TLDR; Part 1: Testing introspection with concept injection

First they find neural activity patterns they attribute to certain concepts by recording the model’s activations in specific contexts (so for example, they find the concept of "ALL CAPS" or "dogs"). Then they inject these patterns into the model in an unrelated context, and ask the model whether it notices this injection, and whether it can identify the injected concept.

By default (no injection), the model correctly states that it doesn’t detect any injected concept, but after injecting the “ALL CAPS” vector into the model, the model notices the presence of the unexpected concept, and identifies it as relating to loudness or shouting. Most notably, the model recognizes the presence of an injected thought immediately, before even mentioning/utilizing the concept that was injected (i.e it won't start writing in all caps then go, 'Oh you injected all caps' and so on) so it does not simply deduce this it's own output. They repeat this for several other concepts.

Part 2: Introspection for detecting unusual outputs

They prefill an out of place word in the model's response to a given prompt. For example, 'bread'. Then they compare how the models responds to 'Did you mean to say this?' type questions when they inject the concept of bread vs when they don't. They found that models will go , 'Sorry, that was unintentional..' when the concept was not injected but try to confabulate a reason for saying the word when the concept was injected.

Part 3: Intentional control of internal states

They show that models exhibit some level of control over their own internal representations when instructed to do so. When instructing models to think about a given word or concept, they found much higher corresponding neural activity than when told the model not to think about it (though notably, the neural activity in both cases exceeds baseline levels–similar to how it’s difficult, when you are instructed “don’t think about a polar bear,” not to think about a polar bear!).

Notes and Caveats

- Claude Opus 4.1 was the best at these kinds of introspection.

- There is obviously a genuine capacity to monitor and control their own internal states, but they could not elicit these introspection abilities all the time. Even using their best injection protocol, Claude Opus 4.1 only demonstrated this kind of awareness about 20% of the time.

- There are some guesses, but no explanations for the mechanisms of introspection and how/why some of these abilities might have arisen in the first place.

I'm not saying Open AI pricing is entirely unrelated to size/cost. I'm saying why are we assuming that OpenAI is serving say OAI-Opus but at half the price of Anthropic when they could just be serving GPT-5.x which is genuinely near half the cost of Opus at scale.

The official API output tokens cost of GLM-5.2 is like a third of Gemini-3.1-Pro. The model is Open weights so we know it's not just a ploy to grab users at the cost of bleeding money. You can actually serve the model profitably at similar prices.

They have near a billion consumer users every week. Compute efficiency at scale would be at the forefront of any training effort. It makes a lot more sense to me that they have more compute efficient models (even with the scaling) than Anthropic rather than just serving Opus/Fable at half the costs Anthropic are incurring.

That's just texturing over a labor intensive 3D animation

You're already lost if you need perfect 3D renders as the reference

The reference is far from a "perfect 3D render". That's a rudimentary 3D blockout. The characters are basic mannequins without specific geometry, and the environment is composed of untextured, flat-shaded boxes. The demo uses stock assets so effort meter is even more skewed in AI's favour but even if it wasn't, this is significantly less labor-intensive than hand-drawing every frame or creating a fully rigged, textured, and lit 3D scene for traditional production.

Seedance is supplying most of the visible production value: character designs, faces and expressions, linework, backgrounds, lighting, and a coherent anime rendering. It is even generating the secondary animation: the physics and flow of the hair and clothing, which the rigid 3D models completely lack. Far more work than 'just texturing' here.

Don't know if it's you (did you publish?). I read about something similar but it had its issies:

- Tuning hyperparameters to gain improvement on a dataset when you're constantly looking at the answers is pretty meaningless. It's basically testing on the training data.

- Eval on ImageNet1k alone (very small, useless for the real world) made me wonder if it wasn't just overfit to the training set. Would it perform better training on the datasets used for the foundation models ? I doubt it.

Well I'm not saying CNNs are bad or useless at any rate.

There's no 'rigorous comparison' that puts CNNs over Vits in quality and Vits unlocked more use cases easier than CNNs did. That's why they're more popular, not because it's 'bandwagon-y'.