Pretty clear they are defending against (a) AI scrapers (b) impending age verification requirements. Honestly, requiring login is probably the least they could do with old Reddit to keep it running with those two constraints.
HN user
zmmmmm
We have information ...
Hard to think of a weaker way to express this. Strongly suggests veracity of said information is poor.
It's probably a mix of things but I do think they are viewing "edge AI" as their strategic play: on-device, small efficient models (Android / iOS) and instant AI summaries in google search etc. So all of their focus is on delivering strong performance in a compute constrained environment.
I do think it's still also simultaneously true that they have an actual problem with competing with current frontier progress. It's just that has gone from an existential threat to something they are willing to defer addressing because they see the long game for them sitting at the smaller end.
I think the subscriptions pay them back in spades because the same dev who maxes out their subscription on their personal account transfers that exact behaviour over to their enterprise work and - guess what - it's all billed per token there. This is a large reason why corporates are reeling from the cost right now, I think.
considering that Llama, the mother of all open-weight models, has led to anything but success for Meta.
To me, Llama was the ONLY successful thing they did it. it was when they stopped that they fell off the radar as an interesting AI company. They had a genuine chance to be the "substrate" that Ben talks about here. I can't actually figure out why they threw it away.
yes... I think for startups there's a strong desire to demonstrate technical independence - there's a strong smell with looking like a wrapper on Anthropic or OpenAI. With an open model you can weave a story where you are exercising differentiating knowhow by deploying or tuning models yourself.
it's because they are being heavily subsidized to do the research activity
Is it actually true? This seems pivotal because currently most theories rest on the idea that individual Chinese companies are acting in China's overall economic or strategic interest. It's a tough sell to believe they all just do that through implicit desire to align with the CCP's direction. I would believe it much more easily if there were concrete incentives involved.
I'm sort of baffled by what the entities that train the open-weights models get out of it though
I agree. The thesis in the article is interesting insomuch as I had not heard it expressed this way before: US restrictions on GPU exports have made it feasible to train models in China but not serve them. Therefore open model is a hack to get around the export restrictions, since models can be trained internally but shipped out of the country to be served elsewhere under the banner of open weights. I don't really buy this argument - inference is much cheaper than training and they are hosting their models anyway.
I think it is more likely (a) they have the money to do it and they need it for internal reasons - these are huge companies (b) there is a lot of prestige in China associated with besting American technology (c) people are still basing logic on outdated ideas of Chinese capability which are no longer true.
So it is easier than people think for Chinese labs to do this, they need to do it anyway and there is a lot of prestige from opening the weights. It is honestly not that different to why American companies themselves have released open weight models.
Train your own base model - but tune it off Claude output to make it perform more in line with Claude
Is that actually genuine distillation though? Distillation suggests the core model is being pre-trained using output from another model. For the above to work, you have to already have all the core intelligence trained into your base model.
If distillation just comes down to post-training then it's tantamount to admitting that the Chinese base models are just as good as frontier US lab models. Because you can't post-train frontier intelligence into a model. It has to be there in the base. Then you can change how that intelligence is expressed through post-training.
this for me was one of the truly liberating parts of LLMs for me: simply being able to ask a question and get a straight answer without having to run the gauntlet of why I shouldn't be doing what I am trying to do in the first place. The friction of trying to explain why I still wanted to do it that way deterred me from asking so many questions.
One more tiny piece of the global system of international order falling apart.
There was a time when people would have felt safe enough to rely on the multiplicity of strongly allied nations with steel production capability. Now that is not considered safe. Now steel production has to be protected because nations that were previously considered reliable strategic partners no longer are behaving that way.
It will happen slowly but piece by piece things will move and we will all pay a cost for it.
The obvious burning question is how performance looks over different network conditions on some standard models. Have you done much benchmarking? Is it mainly latency affected or is overall throughput less than the capacity of the GPUs due to being distributed?
I'll throw a shout out to the new Google Translate practice feature - it generates sentences around a theme you specify and speaks them to you at varying speeds in your native and learning languages.
Yes it has completely turned me around - was all in on Anthropic but now it just looks too risky. Better off leaning into open models. Even if I found a way to work with the restrictions as they are, who is to say they won't suddenly change tomorrow. It's not worth it.
It envisions delivery of “20 mission-ready aircraft” by 2031.
Hmm, I'm not sure they fully address the problem if that is what is being proposed. The world will be an entirely different place by 2031 and 20 drones is .... meaningless? Surely they should be talking in the thousands or tens of thousands.
Good to see Meta finally back to releasing something at least worth evaluating. And it sounds like they did at least a bit skate to where the puck is going by focusing on tool and computer use.
Yes and Zuck effectively disbanded the entire team that did that. Not saying we shouldn't cast a critical eye on it, but it probably does warrant a second chance.
I don't care about the politics, I wouldn't trust anything made by Musk.
Hard not to reflexively reference the XKCD here. I think the authors of most charting languages would say they look good by default. It seems more likely this is achieving a subjectively different presentation than something objectively better. Higher level implies information loss. So it can only be better if it is doing so by assuming some better defaults. But then you have to ask if it lost expressiveness.
Like a lot of the rest of the world they would probably rather take the alternative option and accelerate the transition to clean energy. Has the upside of not handing more power to an authoritarian state on the other side of the Atlantic that clearly hates them and routinely threatens them.
And yet the will badger you endlessly to the point their photos app is near unusable to turn on auto sync which slurps up every photo and makes it very awkward to then delete them after. To me, this makes Google a liable party even if real CSAM is stored.
I would love to know that inside story. The whole saga is starting to look like one of the biggest own goals in history - Meta went from being widely respected and considered a peer with leading frontier labs to having no competitive technology. How a company seemingly willfully threw away a leading position in the most valuable tech race of all time should be a business case study, apart from a technology one.
I do have a theory : Llama3.1 marks the point where Zuck got seriously interested and took over the reigns in driving the work. From the minute he started directing things instead of considering the AI work as a quirky side project, things went downhill. He tried to force a huge scale up in Llama4 which didn't work. Then as we know he disbanded the whole team and brought in a new crowd of mercenaries who may or may not have had the technical skills but they came into an organisation in disarray and still driven by Zuck himself who is continually forcing decisions that are not well founded in the science.
All the above is an entirely evidence free fan fiction version of things, but I would be completely unsurprised if it is true.
To me a lot of the anti-short leash sentiment is reflective of the low accountability SWE have always had for their output. Software devs seem to strongly reject the concept that it isnt ok to ship defective products and fix later. It will be interesting to see if it persists as incidents start to occur due to fully automated code.
I can't decide if this will make science better or be the death of it. The potential wave of slop about to hit journals is frightening. Essentially what happened with GitHub code reviews is about to hit academic peer reviews and it isn't going to be pretty.
It's honestly quite baffling that the EU would want to put any more power in the hands of any US controlled company at this point. The US is a borderline hostile state, only recently threatening to invade Greenland among numerous other examples. The situation with Anthropic has illustrated that the US government will not hesitate to leverage power over US companies when it feels its interests are advantaged by doing so. If anything, the EU should be banning use of Google or Apple dependent architectures, not pseudo mandating them.
That's a great write up.
The one thing I feel it seems to under estimate is the likelihood of improvement. Even the authors acknowledge it's not even worth comparing local models from a year ago to what we have now. In fact, people widely see Opus 4.5 in November last year - 8 months ago - as the first time agentic coding became viable broadly viable even with frontier hosted models.
So why would we lock in hard on any concept at this point of what a local model is and isn't good for? Whatever it is right now, it probably won't be that in a year. It might be naive optimism to think we'll ever get to long horizon tasks with models that run on consumer / pro grade hardware. But so far the naive optimists are winning.
Engineering for the sake of engineering has no value to the economy
I think that's the adventure we're on now. If recreating something is low cost, what is the value in investing in designing it well in the first place? We can empirically discover issues and the the AI to address them.
I certainly routinely find in supervising what the LLM is writing that it's making terrible internal design choices and correct them. Usually things one level up from code. "This will cache every image on the client and cause a huge amount of bloat. Change it to pull the image in real time from the server" kind of stuff. You do slowly build that up in the project documentation - "Never store unnecessary data on the client: we assume they are using low powered devices without substantial storage". But it takes time and the road to discovering that empirically is through a lot of unhappy users.
So I think there is still a lot of room for genuine engineering - that is, at the technical design level. Levels up from that - code structure etc - are much less clear. I am guessing that over time we will heavily optimise code written by AI for maintenance by AI. Which may be mostly about matching the context window to the code module size. Factoring something to 5 modules may be less of a good idea if it means the context window has to hold all of them for the LLM to work. But that is the path of discovery we are on which history tells us is a 20 year journey.
Now you get not just the 5 LoC to review but a 5 page essay to read in the form an auto-generated review as well. Which makes the submitter even more indignant when you start nit picking things about how it's implemented.
I've tried so many of these and paid for a lot of them and I still can't find what I want. It sounds like this is closer than most:
- record and separate two sides of the conversation
- save meetings in a simple transcription format in a local folder
- connect with my calendar (Outlook, Google Calendar) and name meeting transcripts accordingly
- for recurring meetings, append rather than create a new transcript
- let me label speaker voices and recognise those voices across different meetings
A tool that did all this and then ALSO built a knowledge base to let me RAG query my meetings would be the holy grail for me.
they would have started verifying the identity of their customers.
Very good point. Yes i think this part goes to hubris. Amodei probably didn't think the ban would cut along those lines if it happened. And in fact it wouldn't surprise me if the government specifically made it that way (singling out foreign nationals) as a way of punishing Anthropic for putting them in this position. It's clear they absolutely hate being dictated to by anybody, but especially Amodei and they probably thought through what would hurt them a lot to implement and deliberately made it that way.