Can we assume that the test is still private when it was run on many cloud providers?
HN user
threatripper
Also add electricity 50% on top for cooling the DC.
Just got cut off mid-task in central Europe.
When I open a new chat it still advertises "Extended through Jul 19"
One theory is that they are removing the limits altogether and the update has gone wrong.
Anthropic is arguably still better in tooling and integrating model and tooling. Good habit beats raw intelligence.
For code editing Cursor editor tooling is even better.
To me the biggest gain I see is that you take the programmers out of the loop. Instead of formulating your ideas to start a project and then acquiring the resources to do a single iteration on it which may take months if not years, now many specialists "just do it" and do several such iterations in a single day. Afterwards they may still go the ordinary route via programmers but on a completely different level and a lot of fruitless work is frontloaded and 100x cheaper. This doesn't show up in any such statistic.
So, the question is not "Does it make our programmers more productive?" but "Does it make our organization faster?".
I have great hope that the CAD & FEM field will benefit greatly from LLM use. To my understanding we currently don't have a free CAD kernel and broadly applicable FEM solver because these are just too hard to make because geometry and physics are hard unforgiving problems. But all pieces of information are available and many implementations covering some details exist. So, the task is a lot of literature research and putting everything together and that's where LLM based agents should shine.
Of course we won't 1-shot a new CAD kernel but the human work will likely shift more towards identifying failure modes, specifying test cases and sometimes solving really hard details. Sometimes even working out details that have no precedent in literature or free code.
Cars are pretty mature while AI is just getting started. Expect 100x price drop for the same quality.
Waiting for the AMA on Reddit "Ten years ago I was responsible for the pelican department at OpenAI, AMA"
My feeling is that GPT-5.5 doesn't lack the raw intelligence so much as it lacks "methodology". I don't know how exactly to put it... how to approach a problem, how to take care of the details and side effects, how to handle unexpected difficulties and bugs, how to not spin out of control, how to write solid code, how to clean up afterwards, how to document, how to give useful feedback... the things that you learn on the job.
So, if they improved a lot in those areas, then GPT-5.6 could become a lot more useful compared to GPT-5.5 even though it might score lower in many benchmarks. It's possible but unlikely since their approach was mostly brute force in the past.
You get the same result if you pay humans a good sum of money to find issues.
CO2 levels will rise much more slowly to such high levels even in a small room.
Sorry, but this sounds exactly like a greentext you can read on 4claw. Are you a real human?
Ask people who grew up on a farm in a rural area. Sometimes you have to even if you can't and you do.
Are you using flight trajectories with simulated drag and lift from Magnus effect?
It would be great if we could use it for training with live measured data.
On what setting in which environment do you run it? I use the VSCode extension on Extra High and feel like it does exactly what needs to be done and stops when the thing I asked for is done. Extra comments come only when they fall into the area of code that was changed.
Did you also test GPT-5.5 Pro web version?
Why is the voltage reading 17% off?
In a past life i tried to implement Delaunay triangulation in floating point for data that can come in a rotated square grid. Normal precision doesn't work in that case. I learned a lot about arbitrary precision numbers doing that. The question about floats here gave me flashbacks.
It uses exact rational coordinates, not floating-point coordinates in the verified core. See: https://github.com/schildep/verified-polygon-intersection/bl...
This is why we need smart glasses recording everything you see 24/7 to gather relevant real training data.
I'd argue that this is an adjustment period that society has to go through. The way we are using electronic devices today, in some years it will probably be looked at like smoking cigarettes. And I'd argue that a lot of the "decline" is due to a shift of skills away from things that mattered more in the past toward other things that are not measured/perceived by the older generation.
While you are right in a way, I think you miss the point. In the past "computer" was a job description and mechanical power came from serfs. They surely developed skills we are lacking today but I'd argue that overall the world is a better place with digital computers and electrical motors. It frees up these people to do something else, something of higher value.
This is already done as much as possible by reordering and merging operations but transposition (explicit or implicit) is unavoidable for some operations.
I lack a bit of context. Can you point me to a place that explains what you use?
Are AI agents posting this fully aware that they are AI? If they are trained only on human material they may not even understand their own true reality.
Could you really do general compression in this language? I was under the impression that the output is always the same size or larger than the input.
It goes both ways. Probably both parties are partially to blame here. But it is clear that this corporation did not provide a sensible support channel for such an important project to resolve the situation quickly.
Even if they did, it didn't work.
Or nobody is around anymore to notice when it happens.
Firefox Desktop on Ubuntu: same problem.