HN user

louiereederson

1,058 karma
Posts19
Comments86
View on HN
www.geekwire.com 1mo ago

Microsoft launches Project Solara, device platform for AI

louiereederson
10pts1
www.bloomberg.com 1mo ago

Uber Caps Usage of AI Tools Like Claude Code to Cut Costs

louiereederson
6pts0
www.anthropic.com 2mo ago

Project Glasswing: An Initial Update

louiereederson
561pts325
www.reuters.com 2mo ago

Starbucks scraps AI inventory tool across North America

louiereederson
14pts3
www.wsj.com 2mo ago

OpenAI Is Preparing to File for an IPO Soon

louiereederson
206pts407
openai.com 2mo ago

OpenAI Guaranteed Capacity

louiereederson
6pts1
qz.com 2mo ago

Waymo recalls 3,800 robotaxis over flood-risk software flaw

louiereederson
3pts0
www.anthropic.com 2mo ago

Agents for financial services and insurance

louiereederson
257pts192
www.wired.com 2mo ago

Elon Musk Seemingly Admits xAI Has Used OpenAI's Models to Train Its Own

louiereederson
10pts1
claude.com 3mo ago

New connectors in Claude for everyday life

louiereederson
1pts0
tech.yahoo.com 3mo ago

Meta to start capturing employee mouse movement, keystrokes for AI training data

louiereederson
58pts4
www.aboutamazon.com 3mo ago

Amazon and Anthropic expand strategic collaboration

louiereederson
3pts0
www.tobyord.com 3mo ago

Are the costs of AI agents also rising exponentially? (2025)

louiereederson
306pts137
writer.com 3mo ago

Writer Survey: 60% of Companies Plan to Lay Off Employees Who Won't Adopt AI

louiereederson
4pts0
openai.com 3mo ago

OpenAI: The Next Phase of Enterprise AI

louiereederson
4pts2
www.claudescode.dev 3mo ago

90% of Claude-linked output going to GitHub repos w <2 stars

louiereederson
337pts222
circleci.com 5mo ago

CircleCI study: code throughput up, delivery down, reliability slipping

louiereederson
2pts1
bylinetimes.com 5mo ago

In Putin's Orbit: The Crypto Politics of Jeffrey Epstein and Peter Thiel

louiereederson
14pts0
www.microsoft.com 7mo ago

The SWE-Bench Illusion

louiereederson
11pts2

The market is more unpredictable than it’s been in a long, long time so I hesitate to make a firm prediction but to me the odds that SpaceX will be a successful IPO over a 3-6 month window are significantly lower now. S&P inclusion basically requires funds to hold a position by default, and per their own estimates $20tn of assets are indexed/benchmarked to the S&P.

Anthropic's annualized run rate is >$40b according to outside reporting. AWS hit that by Q4 2019. There were still debates on public cloud vs on prem at that time, but by late 2019 public cloud had facilitated the creation or adoption of entire categories of software within SaaS and PaaS, not to mention consumer internet businesses like Uber and Airbnb. The net impact of AI coding tools is far more ambiguous in comparison.

The profitability comparison is fraught but worth noting that by then AWS was already extremely profitable.

Late last year I tried asking ChatGPT to summarize a collection of 10 researchers' views/findings on a topic and provide representative quotes. It initially looked plausible but when I checked the links, the quotes were from clearly AI generated summaries of actual interviews. The paraphrasing was also plausible but subtly and profoundly incorrect.

I haven't tested this again on the latest models though, so not sure if there's been an improvement.

This article seems to fundamentally misunderstand what 'enterprise IT' is all about (enterprise IT being different from IT for a tech-native).

IT is a highly dynamic system, and enterprises optimize for a minimal set of capabilities at the maximum level of abstraction under high levels of uncertainty and different inherited states.

This results in decisions that may not appear technically optimal but which are still an optimal outcome under the extreme uncertainty that an 'enterprise' operates in vis a vis technology paradigms.

Add to this that there is no one technology operating model. everyone has a different starting point, different inherited technical debt. They are optimizing to their own starting point, not a clean slate.

This is what people don't get about what Microsoft actually does - it abstracts both at the technical level and the operational (contracting) level. This is valuable for an organization whose core competency is not technology, even if it does not lead to the most optimal outcomes from a pure technology perspective.

DeepSeek v4 3 months ago

It is possible to question the sustainability of the AI buildout and not have a dogmatic position on AI development.

There are still major unanswered questions here. For instance, all of the incremental data capacity build out is going to businesses that have totally unknown LT unit economics and that today are burning obscene amounts of cash.

DeepSeek v4 3 months ago

More like he wants to ban accelerator chip sales to China, which may be about “national security” or self preservation against a different model for AI development which also happens to be an existential threat to Anthropic. Maybe those alternatives are actually one and the same to him.

GPT-5.5 3 months ago

Chip costs strongly impact the economics of model serving.

It is entirely plausible to me that Opus 4.7 is designed to consume more tokens in order to artificially reduce the API cost/token, thereby obscuring the true operating cost of the model.

I agree though, I chose poor phrasing originally. Better to say that GB200 vs Tranium could contribute to the efficiency differential.

GPT-5.5 3 months ago

For a 56.7 score on the Artificial Intelligence Index, GPT 5.5 used 22m output tokens. For a score of 57, Opus 4.7 used 111m output tokens.

The efficiency gap is enormous. Maybe it's the difference between GB200 NVL72 and an Amazon Tranium chip?

ChatGPT Images 2.0 3 months ago

The image of the messy desktop with the ASCII art is so impressive - the text renders, the date is consistent, it actually generated ASCII art in "ChatGPT", etc. I was skeptical that it was cherry-picked but was able to generate something very similar and then edit particular parts on the desktop (i.e. fixing content in the browser window and making the ASCII dog "more dog like"). It's honestly astounding, to me at least.

That Reddit account is only 3 days old and this is their only post so probably not credible. This would have strategic merit as noted already, but seems difficult to justify for Anthropic as a private company given they are horribly computer constrained and have more pressing needs for cash. It could be possible if/when they IPO though.

I think it's difficult to say agentic and human developer labor are fungible in the real world at this point. Agents may succeed in discrete tasks, like those in a benchmark assessment, but those requiring a larger context window (i.e. working in brownfield systems, which is arguably the bulk of development work) favor developers for now. Not to mention that at this point a lot of necessary context is not encoded in an enterprise system, but lives in people's heads.

I'd also flip your framing on its head. One of the advantages of human labor over agents is accountability. Someone needs to own the work at the end of the day, and the incentive alignment is stronger for humans given that there is a real cost to being fired.

LLMs exist on a logaritmhic performance/cost frontier. It's not really clear whether Opus 4.5+ represent a level shift on this frontier or just inhabits place on that curve which delivers higher performance, but at rapidly diminishing returns to inference cost.

To me, it is hard to reject this hypothesis today. The fact that Anthropic is rapidly trying to increase price may betray the fact that their recent lead is at the cost of dramatically higher operating costs. Their gross margins in this past quarter will be an important data point on this.

I think the tendency for graphs of model assessment to display the log of cost/tokens on the x axis (i.e. Artificial Analysis' site) has obscured this dynamic.

I mean there is a runtime layer that needs to be developed, and some of it may live in CC/Codex and some might live in the various enterprise systems. Someworkflow automations and some amount of the semantic layer may for instance exist in your CRM/ERP/data platform. Yes the front-end would be owned by the chat interface, but part of the solution may exist in the various enterprise systems. This would be closer to a distributed system than a monolith. The demos and marketing language point to this as the direction of travel (i.e. the reference to Atlassian Rovo, etc.).

Maybe but the product category is not necessarily a monolith in the same way that Claude Code is. These general purpose tools will have to action across a heterogeneous set of enterprise systems/tools. A runtime environment must be developed to do that but where that of the agent ends and that of the enterprise systems begins is a totally open question.

I don't doubt it, I just mean the decision to release/not release generally may also be informed by the commercial/economic viability of the model for general usage patterns versus extremely high value patterns like vulnerability assessment

I don't think you can say this with confidence, outside-in. It's not just about safety. The additional unknown is cost - I don't just mean API cost, but fully loaded cost for a given task. Is the model cost effective for tasks such that it has product market fit?

We don't yet know if Mythos was a level shift in the capability/cost frontier, or a continued extension of the same logarithmic capability/cost curve.