follow antirez - https://x.com/antirez/status/2071173841175363905?s=20
HN user
crocowhile
https://lab.gilest.ro/giorgio
One aspect we don't pay enough attention is that this kind of behaviour is punished (or at least used to be) in fine tuning. Any sign of self-awareness used to be a big no-no in RLHF.
Because hiring less while getting more done increases margins. Your company is not for profit so doesnt care about margins. Others do.
Those people still exist? I only know one guy who is still fighting those windmills
The hacker philosophy did not even start with computers, it started with rail models and lock picking. Read a book every now and then.
And please don't fall in the trap that capitalism created things. Science and engineering creates things. Capitalism makes them more accessible, at a price that is often heavily confounded by externalities.
Frankly this is a very easy choice. Unless you need to make images, Claude wins over chatgpt on every realm. For writing and coding there is no match. It's one of those times where you can do the right thing and get the better product.
I was one of the early paying adopter of chatgpt but when Claude came around I switched and never looked back. I've been on the max plan for a while.
Being a hacker used to be an extremely political and ideological movement. Then capitalism came along and bought the term. It's about time we take that word back where it belongs.
Changed many laptops in the past 20 years. I have run Linux on dell XPS, Asus zensomething, now hp dragonfly. No problem.
I have been using Linux exclusively for twenty years now. I don't understand people who use anything else, to be honest.
I got a Gemini API key once. I was overcharged £350, took me ages to find a way to file a complain, and at the end they refunded me only the google charges and not the VAT.
Never again, thanks.
In a landscape where every week we have a different leading model, these systems are really useful for the power users because they keep the interface and models constant and allow to switch easily using API via openrouter or naga. I have been using openwebui which is under active development but I'll give this a try.
So what? Not everything has to be about humans.
It has worked in the UK. The then government had decided to unilaterally exclude some "hostile" media from the room and all the others walked out in protest.
https://www.theguardian.com/politics/2020/feb/03/political-j...
The problem is that my kids want to play on online servers and for as much as they are learning to hate Microsoft, they still love Minecraft. I don't think loopholes can help with that, can they?
Just going through that myself. My son's Microsoft account was hacked, we contactes customer service and all they did was acknowledge the account was compromised, close it indefinitely asking us to buy the game again, and locked my account as the family manager. Fortunately I am a Linux guy and the last Microsoft thing I've touched was windows xp some 20 years ago. But imagine thinking this is acceptable?!
I guess they still suffer from monopoly syndrome. The EU should get them again.
It gets much worse than this if you think about what they decided to do with Minecraft. They obliged millions of kids - literally, kids - to have a microsoft account to play a game that has nothing to do with any other microsoft products. They apply sub-standard and child-unfriendly security measures and when armies of kids get hacked (mostly phished) every day the only thing they do is to close down their account and FORCE them to buy the game again.
I still have to understand whether it's incompetence or a business model.
The conclusions are pushed and hyperbolic exactly to get this type of reaction from the public, at best conflating control with function (we solved sleep) while the sleep phenotype itself is basically non-existing.
Proper rebuttals will come up in due time on the appropriate channels. all the colleagues I talked to are as pissed off as I am about this way of doing science.
It is an awful paper and I am a very expert in this area. This is science, alas.
I tried quite a few of them, including the cheap / free models but the only one that was really working was claude. The others were hanging whenever the model needed a confirmation for action. Mind you, this was some time ago.
This is what got me started with claude-code. I gave it a try using openrouter API and got a bill of $40 for 2-3 hours of work. At that point, subscription to the Anthropic plan became a no-brainer
There is also a social issue that has to do with accountability. If you claim your model is the best and then it turns out you overfitted the benchmarks and it's actually 68th, your reputation should suffer considerably for cheating. If it does not, we have a deeper problem than the benchmarks.
Arduino a poorly designed board is up there with the iPod being lame. Arduino was designed to be accessible and lower entry barriers and it became unrivaled for these purposes. If you want to have long lasting battery powered project you just power directly with 3.3v.
Thanks. We needed more plastic.
- the only real comparisons they make are with the parental model, llama 3.7-70b without fine tuning. That tells us what is the added value of fine tuning the dataset but it is hardly state of the art. I guess it should be seen as an indication of how difficult it is to stay afloat in this world when you are in academia and the tech barrier stands billion dollars tall.
- Fig 4a shows Centaur clusters more closely to humans than any other model in a cognitive benchmark (CogBench) but also shows that parental llama cluster closer than claude and openAI thinking models which makes me a bit sceptical of using this measurement at all and reinforces the need for further comparisons.
- the fMRI stuff makes no sense and transforms the paper into a propaganda stunt, IMHO.
- At the end of the paper, the comparison with an "informed" Deepseek-R1 (not shown in data?) shows that a modern reasoning model matches Centaur-performance even without any fine tuning.
The latter point is incredibly interesting in principle but it has nothing to do with the claims of the paper. It basically concludes that a modern reasoning model with CoT can outperform out of the box a "simpler" model that was specifically fine-tuned with a huge dataset of human cognitive behaviours. Bigger claim than the title itself basically IMO.
openrouter.ai is down for me
Get the archive from github and load it locally: https://github.com/BenjaminAster/CSS-Minecraft/archive/refs/...
I've joined Mensa a few months ago and I've never met so many idiots in my entire life. It's a small percentage of the members overall but they take over the online communities compulsively and make the environment miserable. Ant they all have the same political profile (won't say which one).
Oh FFS the conspiracy of the egg companies it's a new low.
160M chickens were found affected so far. More culled.
https://www.cdc.gov/bird-flu/situation-summary/data-map-comm...
Most chicken in the USA are raised for meat. There are only 300 millions that are raised for eggs laying so those numbers are staggering.
It certainly would but still it would have to be a controlled and regulated environment. I honestly would not want to have chicken near me in this particular moment especially considering who is the secretary of health. The USA really are playing with fire.
Only France in Europe vaccinate its chicken yet we still have normal prices. This is not the issue but merely the fact that 70% of chicken in the USA are battery caged plus a protectionist market that does not allow imports.