xAI is renting out compute
https://techcrunch.com/2026/05/20/anthropic-will-pay-xai-1-2...
HN user
xAI is renting out compute
https://techcrunch.com/2026/05/20/anthropic-will-pay-xai-1-2...
Spent some time playing on it just now on Linux. I couldn't easily get the automatic ingestion or gardening working, but everything else is working okay. For my workflows I think I actually prefer wiki updates being manual and deliberate.
This isn't as easy as it sounds. Every ML model is struggling to balance between generalization and test performance.
Taking a good model like GLM5.2 and just fine tuning it on coding can decrease real world performance due to mechanics like catastrophic forgetting. There is also other interesting behaviors were training on a broad training set can improve coding performance because there is positive transfer.
There is 100% an effort to make solid coding focused models, but it is very hard to do that without including capabilities across a broad set of adjacent tasks.
Help me understand this viewpoint that AGI being possible in the near-ish future is a myth, I see it repeated quite a lot.
I've been in NLP since the LSTM days and it's hard for me to look at LLMs and not just think they are incredible. It's truly a different level of expressiveness. So much of capabilities research is pointing to LLMs effectively learning a world model.
RLVR is also proving really effective. It is hard for me to imagine a world in the future where LLMs aren't at human level performance across a wide variety of tasks.
I fully acknowledge that current LLM labs have a financial interest in people believing AGI is very near, but from what I'm reading in the literature and seeing myself experimenting with the SOTA models it doesn't seem totally unreasonable.
What evidence are you seeing that makes you confident that AGI in the soon-ish future is a complete myth?
There doesn't really seem to be anything of substance in the actual executive order.
Section 1 doesn't say anything
Section 2 seems to boil down to: "improve cyber security and maybe use AI if we can find funding for it"
Section 3 proposes building a benchmark for evaluating cyber security performance of models that developers can choose to benchmark against. This seems like a good idea, I know Jack Clark has been a huge advocate for government's getting in with benchmarking.
Section 4 says to prioritize prosecuting cyber crimes. Not sure why they wouldn't already be prosecuted.
Section 5 doesn't say anything
Are we looking at the same data? On that site I see that opus 4.7's and gpt 5.5's g scores are within each others confidence intervals, and both significantly ahead of the number 3 model.
Your comment makes it sound like they are miles apart, which the benchmark doesn't seem to support.
Edit: I looked at the data more and the two models are only basically equal when looking at the mean of all the tests. Gpt 5.5 significantly outperforms opus 4.7 in coding, while opus 4.7 significantly outperforms in "decision making." I'm not seeing details on what decision making explicitly means.
An alternative explanation is that cold places with long winters are depressing, and because they are depressing fewer people want to live there.
Alaskan winters are hard regardless of how many friends you have.
I had a similar experience recently, where I logged in to Facebook after not using it for years and was shocked by how much garbage was there. My spouse does use Facebook somewhat regularly so I looked at her feed and it was much more reasonable.
I wonder if for those of us that haven't used Facebook in years the recommendation algorithm is essentially default. Which much like the default youtube algorithm, is completely garbage. But if we did use it (which I have no intention of doing), it would start being more reasonable.
In practice you can use 2d generation on spheres with simple UV mapping techniques. Your pixel height becomes distance from the sphere origin.
I worked on something very similar for my master's degree.
The problem I could never solve was the speed, and from reading the paper it doesn't seem like they managed to solve that either.
In the end, for my work, and I expect for this work, it is only usable for pre generated terrains and in that case you are up against very mature ecosystems with a lot of tooling to manipulate and control terrain generation.
It'll be interesting to see of the authors follow up this paper with research into even stronger ability to condition and control terrain outputs.
The main issues I see are immich identifying things like statues or paintings as people, and not dealing with people (especially kids) aging.
Google photos isn't perfect either but I never saw these kind of issues when I was still using it.
There are still some features that a miss from Google photos. There isn't any way (that I know of) to auto add pictures to an album based on the face. I used to have dedicated albums for family members, and it was nice to have the auto updated.
Face recognition in general just isn't as good as Google Photos.
It's still an amazing piece of software and I'd never go back, but it isn't perfect yet.
Diffusion LMs do seem to be able to get more out of the same data. In a world where we are already training transformer based LLMs on all text available, diffusion LMs ability to continue learning on a fixed set of data may be able to outperform transformers
DETR model have outperformed YOLO models for a while, but they have been much slower making them impractical for real time detection.
Well, the US publishes numbers for a lot of its programs so we can see exactly how much is spent on the bureaucratic nightmare.
Medicaid
FY 2023 Budget: $900.3b ($620b federal, $280b state) [1]
FY 2023 Budget not spent on benefits (admin overhead): 5% ($45b) [2]
SNAP
FY 2024 Budget: $100.3b [3]
FY 2024 Budget not spent on benefits (admin overhead): $6.5b [3]
TANF
FY2024 Budget: $31.5b ($16.5b federal, $15b state) [4]
FY2023 Budget spent on program overhead: %10.1 ($3.2b) [5]
Total Admin Spending $54.7b -> $169 per person in the US
So not totally negligible but also not exactly a basic income
[1] (https://www.macpac.gov/topic/spending)
[2] (https://www.congress.gov/crs-product/R42640) See figure 4
[3] (https://usafacts.org/answers/how-much-does-the-federal-gover...)
[4] (https://www.gao.gov/assets/880/872093.pdf)
[5] (https://acf.gov/sites/default/files/documents/ofa/fy2023_tan...)
This exactly. For parents it is not a choice, you absolutely must have a parent sitting by a young child. The effect of not automatically putting parent and children next to each other would just be making tickets more expensive for parents.
You're completely right, it is very similar to the western style credit score, and is often either accidentally or deliberately misrepresented. That being said, it covers behaviors not covered by western credit scores that does have elements of tracking "how good of a citizen are you".
I think this article from Beijing University does a great job of highlighting some of the issues.
Web page:
https://fzzfyjy.cupl.edu.cn/info/1035/11343.htm
Pdf: https://www.law.pku.edu.cn/docs/20210927092933450384.pdf
I agree that the American understanding of the Social Credit System is flawed, but to suggest it doesn't exist is an extreme overreach.
Furthermore, it clearly is a system that is important to the CCP and has an effect.
From the Baidu page:
建立社会信用体系是保持国民经济持续、稳定增长的需要
Establishing a social credit system is necessary to maintain sustained and stable growth of the national economy
Certainly Mainland China does need a credit system, and undoubtedly the Social Credit System will and has helped in that regard, but it does have legitimate flaws with regards to privacy.
Its goals extend beyond ensuring creditworthiness to
社会信用体系具有揭示功能,能够扬善惩恶,提高经济效率;
The social credit system has a revealing function, can promote good and punish evil, and improve economic efficiency;
And its integration with the National Healthcare Security Administration, and other government and private entities extend its reach far beyond what the Western credit systems do.
What are you referring to?
gov.cn page on social credit system plan https://www.ndrc.gov.cn/xxgk/zcfb/tz/202406/P020240604321155...
gov.cn page on social credit system suggested changes/opinions https://www.gov.cn/zhengce/202503/content_7016535.htm
baidu page on social credit system https://baike.baidu.com/item/%E7%A4%BE%E4%BC%9A%E4%BF%A1%E7%...
Are you making the distinction that its name is actually the social credit system and not social score? Or that the system isn't fully in place yet.
Or perhaps are you suggesting that gov.cn and baidu are part of the deep state's propaganda plan against China.
Here is a list of items that matter for an American. It is much worse if you are from Hong Kong or Taiwan.
1. Instagram, Facebook, Youtube, Twitter (basically all major American social media platforms) are not accessible without a VPN. Some major VPN providers have also been banned as well. The counter argument is that many Huawei products are also banned. But I actively use Harmony OS on my Huawei smart watch. I can view Bilibili content, XiaoHongShu, QQ, and other platforms on any device without issue.
2. State media propaganda. My first time seeing state movies was shocking how in your face the propaganda was. It is completely blatant. From talking with locals about it, they recognize it as propaganda (although the Chinese word for propaganda doesn't have the same negative connotation as in English). I haven't watched state sponsored news to get a feel for their bias, but from the amount I have seen they are certainly selectively with their content. Everything focuses and the achievements Xi JinPing has recently achieved. The idea of a media outlet reporting on something silly the president has done is absurd.
3. The surveillance infrastructure in China surpasses even the UK. The number of cameras almost looks like a gag. Once again, the nature of how opaque the government is means it is difficult to say with any degree of certainty what they use the data for, but they certainly collect a lot.
Plane tickets to China are relatively cheap right now. LAX to PEK round trip for $750. Take a trip and see for yourself.
The nature of Chinese censorship makes it difficult to provide hard numbers, but it really is worse. America's handling of censorship is certainly not the best, and it has gotten worse recently, but it is not on the level of China.
How would you respond to the critique that it makes that tax associated with a property dependent on the improvements to adjacent properties? I could see a situation where a single family home owner would deliberately oppose improvements (i.e. parks/bike lanes), because their derived utility from those improvements is less than the potential increase in taxes.
I feel like most people that say WeChat is a super app haven't actually used it for any period of time. WeChat achieves their "able to do everything" by embedding sub apps within the app. Switching between them is jarring, and is sometimes less smooth than just opening a different app. Saying WeChat is a super app is like saying an app store is a super app.
While I strongly doubt this fully disables tracking, you can at least disable your watch history on youtube which will have the effect of the recommendation algorithm not adjusting to your preferences.
You can change it from Google account > Data & Privacy > History Settings > youtube History
If you have youtube premium + a general purpose ad blocker + disable watch history its really hard to tell if you are being tracked.
If you do decide to disable watch history, be prepared for just how terrible the median youtube interest is. All recommendations become beyond worthless.
To me, the point the friend is making is, just like you said, that you don't need to review every line of code in a package, just the interface. The author misses the point that there truly is code that you trust without seeing it. At the moment AI code isn't as trustworthy as a well tested package but that isn't intrinsic to the technology, just a byproduct of the current state. As AI code becomes more reliable, it will likely become the case that you only need to read the subset of the interface you import/link and use.
The truth that may be shocking to some is that open source contributions submitted by users do not really save me time either, because I also feel I have to do a rigorous review of them.
This truly is shocking. If you are reviewing every single line of every package you intend to use how do you ever write any code?
At least with traditional Chinese, reading isn't as bad as people make it out to be. A lot of characters are pictophonetic characters(形聲), where one element describes the sound and the other meaning. While not perfect they allow a reader to guess with decent accuracy the meaning and pronunciation of a character they have never seen before.
Measure words in Chinese are great. They provide so much descriptive capacity in such a short simple way. 一棍棒, 一把棒, 一根盪, 一條棒, all would translate to English as "a stick", but they convey different perspectives about what that stick is. I can appreciate the frustration with learning words that only have one specific measure word that only really describes it, but even then you can honestly get away with 個.
html