I believe that living a happy life is a long term survival strategy. Producing offsprings is the other way. If we feel miserable, we want to cut our life short or it will be due to many problems, bad health included. Long term happiness is a very good measure of how good your life is going. So while we might not agree on whether we should optimize our lives for happiness, it's certainly an objective we should always consider. However, we should not confuse happiness with pleasure. One is meant for the long term, while the other is ephemeral.
HN user
sinuhe69
“You have reached your limits on help.roblox.com”
That is the first time I see something like that!
I have a problem with the cost per task metrics of Artificial Analysis. We don’t know how they calculate it exactly. But recently, cost per task has become the most discussed topic. The logic is basically: if model A achieves 55% on benchmark X and model B 60%, but the cost per task of A is 50% cheaper, people would choose A instead of B.
But that implies that all output of the less intelligent model A is usable, perhaps only a bit worse than the output of B. But what if the output of A is unusable, or it can only deliver usable results in 1 out of 5 tries? In such cases, the user will have to rerun the task and it will very quickly double or triple the cost and makes the old average number misleading! I would argue the retry and flaky cost will be many times bigger than the average token cost and that is the true cost the users have to bear.
AI-Benchy [0] (admittedly a one man benchmark) shows a much different figure than the numbers of Artificial Analysis. Opus 4.8 cost per task according to AA is $1.80 and Kimi K3 is $0.94$. According to AI Benchy, however, the *cost per successful task* of Opus 4.8 is 10.7 cents vs 19.4 cents of K3. The number of correct tests and pass rate of Opus 4.8 is also higher than Kimi K3.
So on a cost-per-usable-result basis, Kimi K3 is actually pricier than Opus 4.8 — the opposite of what AA’s headline number suggests.
Thus, I don’t know if I can believe the numbers of AA or we need to track the cost ourselves.
[0] https://aibenchy.com/compare/anthropic-claude-opus-4-8-mediu...
In this book, Sanjoy Mahajan shows us that the way to master complexity is through insight rather than precision. Precision can overwhelm us with information, whereas insight connects seemingly disparate pieces of information into a simple picture. Unlike computers, humans depend on insight. Based on the author’s fifteen years of teaching at MIT, Cambridge University, and Olin College, The Art of Insight in Science and Engineering shows us how to build insight and find understanding, giving readers tools to help them solve any problem in science and engineering. (Description courtesy of MIT Press.)
With these heatwaves, many cities in France, Italy and England are at times *hotter* than cities in the tropical SEA! No wonder the bananas are starting to fruit.
The title is misleading. The link led to a pricing page/token plan and not about the new QWen 3.8 model.
Nope! GPT just cheated (by search the web). If you changed the text, export as image, GPT 5.6 Sol will fail. I tested with other text and even with a hint of a hidden text underneath, GPT 5.6 Sol could not see it.
If they are not a frontier-lab, they would not need to submit their models for safety test before release. At least that is the proposal.
Not just money. A lot of CO2 was dumped into our atmosphere and tons of clean water were evaporated as well.
Extrapolating to other fields, we can say this development is the worst what one can do to mankind and society: pushing for AI but at the same time decimating the human capital that helped produce and understand it. What is the point of producing countless mathematical proofs or code that one cannot verify nor understand? It’s like having a machine to produce endless things that no one needs or will use. Just because the machine can bring profits for the (initial) investors?! That is the rottenest form of capitalism imaginable.
The last and most important moment was not clear at all. I wonder why they had to cut the video right there and not let it extinguish the fire and stand tall on the landing platform as SpaceX has done all the time. Even Blue Origin showed the final moment very clear with the explosive riveting.
It said “3D-looking Rubrik cube”. Maybe your cube looks different but I’m pretty sure for everyone else, the GPT result doesn’t look like a 3D-looking Rubrik cube.
Generationsbilanz, AHV, Bernd Raffelhüschen
And I wonder if somebody has tried with the available galactic data and see if the genetic programming can come up with a better formula than MOND or Einstein's general relativity.
For simple problems as Kepler's law, a quick detour on Desmos will show a perfect fit for power law instantly. In general, there are many important criteria for a better curve fitting (for ex. independent, normal distributed residuals), not just R, so I hope the author has/will incorporate them into the search to create a more robust result.
Medical and long-term care expenses outweighs by a large margin the loss of consumption or tax revenue for retirees, especially for a welfare state like Germany. Many studies in Germany have showed exactly this constellation of cost-benefit.
When people came to the country to work then retire somewhere else, isn’t it not a net benefit for Germany? Less burden on the social net, healthcare system etc.
So what the Germans did is right, not wrong!
While attribution is a strong weapon in fighting malicious software, persevering the ability to install and run anonymous software is essential to fight authoritarian regimes and corrupt systems. If we accept that only signed, permitted software can be installed and run on users’ phones, democracy and our freedom are doomed. Regardless if it is in the West or the East, or it’s against an AI overlord.
A model that can ask questions or ask for help when in doubt is indeed a major feat. None of the current frontier models can do that.
Although I can see that some people perceive “China’s model“ as a bad connotation, it’s simply the truth. So face it.
At this point, I think open weights vs proprietary models is a misnomer.
First, we can not be sure the next release will remain open weights as Qwen 3.7 has showed.
And second, they are all Chinese models. So instead of open weights, perhaps Chinese AI models is a better word choice.
IMO, using AI to assign keywords to a broader group of strict synonymous keywords would make the comparison much more helpful.
Because in general we want to know the trend of categories more than of a word, asking for “auto pilot” for ex. should include “self driving”, FSD etc.
Is moving many drones in a formation so truly difficult? I’ve seen drones in formation all the time for drones show and fireworks. Hundreds and even thousands of them, and most likely they are not remote controlled but programmed to do so.
Why was the account painted as something extraordinary?
You forgot that the SAT requirement is not exclusive but an additional data point. While I agree that it could narrow the path for truly good employees, I’d argue that an additional data point like SAT (+ GPA) could tell the employer a lot about consistency of the applicants. Or at least an interesting talking point (“I see you got a very high SAT score but your GPA was lower, what happened?”), if they care.
I think it could serve the purposes of hiring fresh/young graduates. However, it’s still weird if they requested it for people already 5-10 years or more in the industry.
Searching and seizure of your laptops, including your personal phones without a probable cause or warrant.
Compel you to reveal your secrets, including your passwords by threatening to arrest and detain you without legal proceedings for an unspecified period.
Deny your basic human rights, particularly at the borders, especially if you aren’t a citizen.
And more.
5% is very low probability to get a hit. I tried with all ChatGPT, Gemini, Claude and Grok but they all answered correctly :’(
Too sad, I want to have the fun.
How it’s price dumping if they give you the model to run free on your own hardware? Follow this logic, are not all OSS price dumping and should be blocked as well? I remember Steve Ballmer once called for that!
I’m pretty sure the digital lords like that proposal a lot. Not so much about the serfs themselves, though.
Not just Sander, Trump expressed his wish to exchange concessions and privileges for share in AI labs, too. Which is IMO even more problematic.
These games are so far outside the normal training corpus and purposes of the AI, I think different promtings could bring vastly different results.
Too bad the author didn’t let the playground open for anyone to try their hand on it.
Yes, it’s fun and it could justify the conclusion “each model for its task”. But are coding benchmarks not designed for the same purpose? The current benchmarks are certainly not perfect and hyper-tuned for the tests can always happen. However, I don’t think a battle royal result can tell much about the coding performance or how helpful the AI could be for me in my daily work.
I get what you mean. But for many people, AI coding is not about solving complex problems. No, they do it mostly themselves. AI coding for many is a productivity tool, where it helps you with mundane, but laborious tasks.
In my setup, I use a daily workhorse for such things. They should be fast, cheap and reasonably working well. I don’t expect it to be smart, but need it to follow instructions perfectly and handle tool calling well.
For architectural work or debugging help, I use the top models instead.
That works reasonably well for me with a low cost.
Recent incident with the Rio 3.5 model clearly shows that many coding models are specifically trained/fine tuned for the benchmarks.