HN user

ertgbnm

870 karma
Posts0
Comments98
View on HN
No posts found.

I've had the feeling that labs aren't pelicanmaxxing specifically but that they do have some sort of RL environment for SVGs that they are letting the AIs overcook in. Specifically I'm thinking of the gemini 3.1 pro annoucnement that seemed to have a huge leap in animated SVG performance but not much else impressive about it.

So they aren't pelicanmaxxing but they are benchmaxxing in a way. The benefit of the pelican was originally that uplift on the pelican signaled an overall uplift on intelligence. I don't believe that is the case anymore and it is just another jagged edge of model intelligence.

MAI-Thinking-1 2 months ago

with AI-generated content excluded from pre-training.

without distillation from third-party models

sounds like zero unless they are lying.

Claude Opus 4.8 2 months ago

My point is that if I made someone "smarter" they wouldn't suddenly know "What day, month, and year was Carrie Underwood’s album “CryPretty” certified Gold by the RIAA?" which is an example of a question in the SimpleQA benchmark.

So (in my opinion) knowledge benchmarks stagnating for small models is not evidence that small model agentic coding performance improvement will stagnate soon. Small models do not struggle with syntax, the barrier is not knowledge. The barrier is long context coherence and problem solving, which I don't see a bottleneck on improvements for small models in the near horizon as we get more and more high quality reasoning traces to train upon.

Claude Opus 4.8 2 months ago

Knowledge benchmarks can't really be improved upon via distillation or RL. It requires those facts be added to the training corpus and for the model to memorize them better. Neither distillation or RL really do that and thus we shouldn't expect improvements on SimpleQA unless some other interventions are being made.

Model intelligence and knowledge aren't necessarily directly related. If we can pack greater intelligence and agency at the cost of it forgetting factoids, that would actually be a good thing. We don't need LLMs to memorize facts, we need them to learn how to interact with the world such that they can find the facts that are necessary and surface them to the user.

If we could distill all of the knowledge out of an LLM and just be left with a very agentic model that only knows facts in it's context, I think some very interesting stuff would happen.

My threshold for asking for help might be a little higher than the median, but, I like this operating style personally. Maybe it's just how I was raised, but the thought of not trying to figure something out for myself first is unthinkable. I fill like I get plenty of human connection at work and collaborate with peers and other disciplines plenty.

I'm not in tech, I'm in civil engineering so maybe it's just a difference in the types of problems we have in different industries.

I do find it very frustrating when an EIT asks me how to do something and it's clear that they haven't even read the instructions page to the excel sheet that they literally have open. I have time to mentor peers and subordinates but I want them to treat my time with the same respect that they treat their own.

The null hypothesis isn't just the opposite of whatever your opposition believes.

For LLMs the null hypothesis would be that there is no relationship between the input and output tokens. Something that is so obviously not true that it's not even worth calculating the number of sigmas away from the null hypothesis that LLMs are.

So clearly we discarded the null hypothesis sometime in 2017. Now we have a system that is really really good at pattern matching and seems to understand consequences. Is that "seeming" just a ruse or does it really understand stuff? A proper scientists would look at that evidence and put forward the hypothesis that maybe it really does understand stuff and begin working on experiments that would disprove that alternative hypothesis, moving forward with the assumption that the hypothesis is true until disproven or a better hypothesis is proposed that explains previous evidence more accurately. Naysayers saying "you haven't proven that pattern matching becomes understanding to my satisfaction" is not a rebuttal. They need an alternative hypothesis that can make predications that better fit the model and can be tested.

The only rebuttals I've heard are "AI can't actually understand stuff and therefore can't do X" which is a testable hypothesis at least. But Invariably AI eventually does X, just in a different way than anyone really expected.

Sending an AI response to a question that someone asks you is insulting because it's a bit like sending them a link to letmegooglethat where it just animates typing the question you have into google.

I think it's only appropriate when you are trying to insult the asker. Like if an employee asks a really dumb question that indicates that they didn't even bother googling the question or asking AI first, then sending them back an AI response is appropriate specifically because it's a bit insulting to do.

In fact it does exist for gpts: https://letmegpt.com/

Personally, If I'm asking for help it's because I've surely exhausted other avenues of approach like googling it or asking chatGPT. I've come to the person because I need their input specifically. The people I work with are professional enough and I've developed such a relationship with them that I don't have the problem the OP is discussing very much.

It makes scams like that scalable. Once you discover one vector of scamming an AI bookkeeper, you can scam all of the users of that AI, using your own AI to scale it for you.

GPT-5.5 3 months ago

can't wait for "our worst and dumbest model yet"

Instagram follows is not a good way to hire football players but it's probably a good way to hire instagram influencers. The football analogy is a little unfair because VCs are investing in more than just a company's ability to "play football" they are investing in the brand, the marketing, and the vision. GitHub stars are at least an indication of a startup having a promising brand or some ability to market themselves.

Nevertheless, VCs are in fact pretty dumb sometimes and it'd be stupid to invest soley based on stars.

Agreed. I think the starting comparison actually works here. It's a bit like the automobile. The advice of "just don't" doesn't work for cars. It takes a deliberate effort on every scale of society to accomplish, it's not something an individual can just do and succeed at. An American can't just not have a car the same way someone from the netherlands might be able to.

Most breakthroughs that are published are for efficiency because most breakthroughs that are published are for open source.'

All the foundation model breakthroughs are hoarded by the labs doing the pretraining. That being said, RL reasoning training is the obvious and largest breakthrough for intelligence in recent years.

Does the data not support a 2X increase in packages?

Pre-ChatGPT, in ~2020, there were about 5,000 new packages per month. Starting in 2025 (the actual year agents took off), there is a clear uptick in packages that is consistently about 10,000 or 2X the pre-ChatGPT era.

In general, the rate of increase is on a clear exponential. So while we might not see a step change in productivity, there comes a point where the average developer is in fact 10X productive than before. It just doesn't feel so crazy because it can about in discrete 5% boosts.

I also disagree with the dataset being a good indicator of productivity. I wouldn't actually suspect the number of packages or the frequency of updates to track closely with productivity. My first order guess would that AI would actually be deflationary. Why spend the time to open source something that AI can gen up for anyone on a case by case basis specific to the project. it takes a certain level of dedication and passion for a person to open source a project and if the AI just made it for them, then they haven't actually made the investment of their time and effort to make them feel justified in publishing the package.

The metrics I would expect to go up are actually the size of codebases, the number of forks of projects that create hyper customized versions of tools and libraries, and other metrics like that.

Overall, I'd predict AI is deflationary on the number of products that exist. If AI removes the friction involved with just making a custom solution, then the amount of demand for middleman software should actually fall as products vertically integrate and reduce dependencies.

That would be great if journals bothered publishing replication studies. But since they don't, researchers can't get adequate funding to perform them, and since they can't perform them, they don't exist.

We can't look for failed replication experiments if none exist.

I am reminded by the perhaps revisionist history but still applicable belief that slavery was really ended by industrialization making abolition economically advantageous and not actually a socially driven movement. (In reality it was certainly a convoluted mixture of the two I'm sure.)

I hope we are in a similar era with regards to climate change. Surely there's a lot of money to be made in harnessing effectively unlimited renewable energy that literally falls from the sky like manna. With a bit of social pressure we should be able to extinct the fossil fuel industry in my opinion.

Gemini 3.1 Pro 5 months ago

Animated SVGs are one of the example in the press release. Which is fine, I just think the weird SVG benchmark is now dead. Gemini has beat the benchmark and now differences are just coming down to taste.

I don't know if it got these abilities through generalization or if google gave it a dedicated animated SVG RL suite that got it to improve so much between models.

Regardless we need a new vibe check benchmark ala bicycle pelican.

I'm guessing it has the opposite problem of typical benchmarks since there is no ground truth pelican bike svg to over fit on. Instead the model just has a corpus of shitty pelicans on bikes made by other LLMs that it is mimicking.

So we might have an outer alignment failure.

Explanation I've heard in popscience books:

Healthy grandparents that are around to support their children and take care of grandchildren increase the fitness of the entire lineage by helping their children have more children and those grandchildren to be healthier/safer.

Stoicism has always struck me as cognitive behavioural therapy (specifically the cognitive triangle) but for boys who think therapy is for women and is rife for misuse from people who don't understand it.

I understand stoicism is deeply entwined with modern CBT and the roots can be traced back basically, but why misuse the ancient form when we have decades of evolution and study on CBT?