The market is more unpredictable than it’s been in a long, long time so I hesitate to make a firm prediction but to me the odds that SpaceX will be a successful IPO over a 3-6 month window are significantly lower now. S&P inclusion basically requires funds to hold a position by default, and per their own estimates $20tn of assets are indexed/benchmarked to the S&P.
HN user
louiereederson
Anthropic's annualized run rate is >$40b according to outside reporting. AWS hit that by Q4 2019. There were still debates on public cloud vs on prem at that time, but by late 2019 public cloud had facilitated the creation or adoption of entire categories of software within SaaS and PaaS, not to mention consumer internet businesses like Uber and Airbnb. The net impact of AI coding tools is far more ambiguous in comparison.
The profitability comparison is fraught but worth noting that by then AWS was already extremely profitable.
I wonder if/when the US limits market entry of Deepseek and other Chinese model vendors like they have done with Huawei
I'm wondering if companies are 'diverting' engineering resources from core products to AI products with the view that the former are legacy. Kind of two sides of the same coin though.
Per sanguinem ad astra
Late last year I tried asking ChatGPT to summarize a collection of 10 researchers' views/findings on a topic and provide representative quotes. It initially looked plausible but when I checked the links, the quotes were from clearly AI generated summaries of actual interviews. The paraphrasing was also plausible but subtly and profoundly incorrect.
I haven't tested this again on the latest models though, so not sure if there's been an improvement.
It’s true, this distracts from the real atrocities
The extension of Full Self Destruct mode
I think as someone pointed out earlier, this is more likely about margin preservation as their gross margins are deteriorating really quickly.
Yikes, so incremental margins are in the 50s. I think this says it all.
Tesla FSD didn't crash your car, you did
This article is about Uber, not Meta
This article seems to fundamentally misunderstand what 'enterprise IT' is all about (enterprise IT being different from IT for a tech-native).
IT is a highly dynamic system, and enterprises optimize for a minimal set of capabilities at the maximum level of abstraction under high levels of uncertainty and different inherited states.
This results in decisions that may not appear technically optimal but which are still an optimal outcome under the extreme uncertainty that an 'enterprise' operates in vis a vis technology paradigms.
Add to this that there is no one technology operating model. everyone has a different starting point, different inherited technical debt. They are optimizing to their own starting point, not a clean slate.
This is what people don't get about what Microsoft actually does - it abstracts both at the technical level and the operational (contracting) level. This is valuable for an organization whose core competency is not technology, even if it does not lead to the most optimal outcomes from a pure technology perspective.
It is possible to question the sustainability of the AI buildout and not have a dogmatic position on AI development.
There are still major unanswered questions here. For instance, all of the incremental data capacity build out is going to businesses that have totally unknown LT unit economics and that today are burning obscene amounts of cash.
More like he wants to ban accelerator chip sales to China, which may be about “national security” or self preservation against a different model for AI development which also happens to be an existential threat to Anthropic. Maybe those alternatives are actually one and the same to him.
Are they war or defense products when they are used against your own citizens?
Chip costs strongly impact the economics of model serving.
It is entirely plausible to me that Opus 4.7 is designed to consume more tokens in order to artificially reduce the API cost/token, thereby obscuring the true operating cost of the model.
I agree though, I chose poor phrasing originally. Better to say that GB200 vs Tranium could contribute to the efficiency differential.
For a 56.7 score on the Artificial Intelligence Index, GPT 5.5 used 22m output tokens. For a score of 57, Opus 4.7 used 111m output tokens.
The efficiency gap is enormous. Maybe it's the difference between GB200 NVL72 and an Amazon Tranium chip?
Surveillance for all
The people in this image remind me of early this person does not exist, in the best way
The image of the messy desktop with the ASCII art is so impressive - the text renders, the date is consistent, it actually generated ASCII art in "ChatGPT", etc. I was skeptical that it was cherry-picked but was able to generate something very similar and then edit particular parts on the desktop (i.e. fixing content in the browser window and making the ASCII dog "more dog like"). It's honestly astounding, to me at least.
That Reddit account is only 3 days old and this is their only post so probably not credible. This would have strategic merit as noted already, but seems difficult to justify for Anthropic as a private company given they are horribly computer constrained and have more pressing needs for cash. It could be possible if/when they IPO though.
I think it's difficult to say agentic and human developer labor are fungible in the real world at this point. Agents may succeed in discrete tasks, like those in a benchmark assessment, but those requiring a larger context window (i.e. working in brownfield systems, which is arguably the bulk of development work) favor developers for now. Not to mention that at this point a lot of necessary context is not encoded in an enterprise system, but lives in people's heads.
I'd also flip your framing on its head. One of the advantages of human labor over agents is accountability. Someone needs to own the work at the end of the day, and the incentive alignment is stronger for humans given that there is a real cost to being fired.
I meant reference Toby Ord's work here. I think his framing of the performance/cost frontier hasn't gotten enough attention https://www.tobyord.com/writing/hourly-costs-for-ai-agents
LLMs exist on a logaritmhic performance/cost frontier. It's not really clear whether Opus 4.5+ represent a level shift on this frontier or just inhabits place on that curve which delivers higher performance, but at rapidly diminishing returns to inference cost.
To me, it is hard to reject this hypothesis today. The fact that Anthropic is rapidly trying to increase price may betray the fact that their recent lead is at the cost of dramatically higher operating costs. Their gross margins in this past quarter will be an important data point on this.
I think the tendency for graphs of model assessment to display the log of cost/tokens on the x axis (i.e. Artificial Analysis' site) has obscured this dynamic.
I mean there is a runtime layer that needs to be developed, and some of it may live in CC/Codex and some might live in the various enterprise systems. Someworkflow automations and some amount of the semantic layer may for instance exist in your CRM/ERP/data platform. Yes the front-end would be owned by the chat interface, but part of the solution may exist in the various enterprise systems. This would be closer to a distributed system than a monolith. The demos and marketing language point to this as the direction of travel (i.e. the reference to Atlassian Rovo, etc.).
Maybe but the product category is not necessarily a monolith in the same way that Claude Code is. These general purpose tools will have to action across a heterogeneous set of enterprise systems/tools. A runtime environment must be developed to do that but where that of the agent ends and that of the enterprise systems begins is a totally open question.
I don't doubt it, I just mean the decision to release/not release generally may also be informed by the commercial/economic viability of the model for general usage patterns versus extremely high value patterns like vulnerability assessment
I don't think you can say this with confidence, outside-in. It's not just about safety. The additional unknown is cost - I don't just mean API cost, but fully loaded cost for a given task. Is the model cost effective for tasks such that it has product market fit?
We don't yet know if Mythos was a level shift in the capability/cost frontier, or a continued extension of the same logarithmic capability/cost curve.
Successful over what time frame? It's way too early to declare victory.