Definitely not my experience. Fable is better but I'd prefer K3 to Opus based my experience with both.
HN user
neosat
Good observations. There's definitely a trend in pricing increasing but also balanced by innovations and availability of other models (both open and closed) emerging as alternatives. It's natural for the labs to explore how much they can push pricing, and for competitors to explore how they can treat that margin as their opportunity to grow their business.
Eventually the pricing should be more stable.
I've been using GLM 5.2 recently (company hosted, for non-coding tasks) and it's been strong and reliable. There are areas where GPT 5.5 and Opus 4.x still feel marginally better but only marginally. For most tasks if GLM 5.2 is the only model I have to use I'm productive and happy. This was not true before GLM 5.2. No doubt in my mind that the gap is closing quickly and for most tasks that are not very specialized open models will be usably on par on flagship closed models and have an edge factoring in cost.
For coding I still use 5.5 w/ Codex and prefer that to other models + harness combinations.
Anthropic is really on a tricky path here. When you have had runaway success due to a hit it is easy to believe that it is the natural way of things. However, that happened due to unique convergence of tech paradigm shift, the competitive landscape, and how they were positioned to capture that value through claude code.
They somehow conflate their value with 'safety'. While it's an admirable internal quality for the company to have, their treatment of their user base (developers, users) has been bordering on indifference and their stance bordering on arrogance.
As competition heats up, there is a very real chance of them shooting themselves in the foot with friction such as this (to be fair not completely in their control but also they had their share of responsibility that led to this)
Agree. Audio has strongly temporal so there is almost certainly some positional encoding one way or another.
You need to see the response in light of the original discussion. Referencing here for clarity since I should have included it in the first place: "We used the claude code and codex harness and I implemented some prs they needed with gpt5.5 and opus4.7 and asked them to identify which came from which only from the code."
So the same person, was using similarly competitive tools, and showing that the output was hard to discern (indirectly the implication was also that implementation was fairly trivial in both of those). A better analogy would not be different process and widely different tools but for example two power drills. Sure, folks could still prefer one over the other, but that's a different claim that saying X is objectively better than Y when both are directly competing on very similar dimensions.
Assuming you meant Claude code: I'd love to learn more about "Codex and Claude are very different" because maybe I'm assuming just based on my use case where I use both of them interchangeably for the same thing (coding web and mobile apps)
That's a fair callout and I agree my statement was too general in just mentioning 'output', as you correctly pointed out. To define 'better' you would indeed need to agree on the dimensions you would evaluate candidates against.
I think a more appropriate rephrasing would be 'You cannot simply make a claim that (model + harness) X is better than Y, but then have no discernible difference on dimensions you care about'. In the case of latest of claude code vs codex with gpt 5.5) both are similar enough in the dimensions people will care about in evaluating (vs. differing wildly in cost or time taken).
Your argument is fine but different from the claim the OP is making. You cannot simply make a claim that (model + harness) X is better than Y, but then have no discernible difference in the output. Subjectively, people might still prefer one over due to anything from design to marketing, but that's very different from the claim that X is better than Y for coding (see: "A colleague was convinced Claude is better"). Basically, I prefer Claude is a different claim than Claude is better and the latter has a higher bar of proof.
Exactly, I was confused too. The authors clearly mention what the parent comment talks about, albeit towards the end of the article, that the 'J' bundle meant that these firms were not set up for success once they 'caught up' and were required to innovate not just process but from the ground up to envision new categories (e.g. iPhone).
Revenue is not the right metric when you compare space trips to trips inside a city. The more relevant numbers are EBITDA, Operating cash flow, Profits.
has anyone done the math on: 1. cost to build out and run the data centers 2. cost of compute (hardware and energy) 3. depreciation of legacy GPU and thus value at the end of 3 years.
And then compare the $45B revenue from Anthropic to see if it's mostly break even or if one of Anthropic/SpaceX came out ahead on the contract.
That's true, I should have mentioned active. Actual params are closer to 12B-14B likely, given the 40GB VRAM usage.
Do you find the video understanding work there also to be 'silly little slop', or did you only look at the gifs on the page and not read about the understanding work in a 3B model?
This is not ground-breaking by any means, but achieving this in a 3B model and sharing the approach + weights advances engineering and certainly more contribution that 'silly little slop videos' imo.
If that's the case, a way to test the theory and understanding (assuming some parts of reservoir and signal channel can be reliably identified) would be to prune the high-confidence reservoir significantly reducing the model size while still getting good predictions. I don't believe the authors mention this (though I skimmed and didn't read the full paper in detail so I may be wrong)
"What slows down a team where agents do the implementation is the production of specifications precise enough for an agent to pick up and run. Roadmap, written down. Acceptance criteria, written down. The “what we actually want” forced into precision, be it via a test suite, a ticket, or a written design."
This is merely speed of development and not the velocity of a company towards higher value. There are many PMs confidently (using the same AI tools), without a clear deep understanding of the user problems or why the requirements will be adopted by their target users (or even who the target users really are), writing these done elaborately.
So yes this will lead to faster end-end execution. But if the product is used or if it sits unused will depend on things beyond the above.
Agree with your points on the primary two questions and the circular argument in the original article. However, re: " How is it that atoms/electrons/photons suddenly start experiencing pain? What is it, in terms of atoms/forces, that's experiencing the pain?" that's an interesting question but not necessarily fundamentally refuting of #1. If you start with #1 "Consciousness is an unknown physical something (force/particle/quantum whatever)" then it has 'perceivable' properties of it's own different from those of it's constituent atoms or electrons. A toy example is the 'wetness' of water. If you only look at atoms and molecules with no way to 'experience' water then it's hard to conceive how water can have properties (though in the case of water it is tractable)
Consciousness *may* be something similar. If it is (e.g. the purest form of energy) then it is not inconceivable that it has some properties that not not tractable if we only look at more granular manifestations of it.
Apart from a cool project, this evolved my perspective on what an MCP is, along with some cool architecture insights and inspiring ideas. Thank you!
"We are investigating an issue preventing users from reaching Claude.ai, and will provide an update as soon as possible."
Who is We? I thought software engineers were going to be redundant and AI could do it all itself? (not to take anything away from Claude code + Claude both of which I love)
Just refreshed and see 5.5 now - yay! Love the speedy resolution ;) Thanks folks, I'll complain faster next time....
Enterprise user here and still seeing only 5.4. Yesterday's announcement said that it will take a few hours to roll out to everybody. OpenAI needs better GTM to set the right expectations.
As a player myself, and having seen much higher level player than me, reading the spin from the ball rotation (and in fact trajectory) of the ball is a common (if advanced) skill. Sometimes the movement of the bat can be deceptive (since with the same movement, where it contact on the bat, the finger pressure can affect the spin).
For example, backspin/underspin balls will move slower after the first bounce and feel 'damper' while topspin will jump. So it's def. possible (and in fact reliable) to read the spin from the spin and trajectory of the ball.
Great work on the feature and sure I'll do that. :)
Tried it to automate something that was on my to do list for the day. I had blocked off a few hours for this and managed to get the agent working reasonably well (85%) of the way there in < 15 mins.
The main remaining part is the poor docx / pdf / final output but will create a skill/workflow to get around that.
Worked really well end-end!
The juxtaposition of MCP vs Skills in the article is very strange. These are not competing ways to achieve something. Rather skills is often a way to enable an optimization on top of MCPs.
A simplified but clarifying way to think about it is that MCP exposes all the things that can be done, and Skills encode a workflow/expertise/perspective on how something should be done given all the capabilities.
So I'm not sure why the article portrays one to be conflicting with the other (e.g. "the narrative that “MCP is dead” and “Skills are the new standard” has been hammered into my brain. Everywhere I look, someone is celebrating the death of the Model Context Protocol in favor of dropping a SKILL.md into their repository.").
You can just not choose to use a skill if it's not useful. But if it's useful a skill can add to what an MCP alone can do.
The sign of a company in absolute decline is when the worst possible place to get info about it, is their official announcement. Whoever wrote the announcement went out of their way to obfuscate the crux of the announcement (compare the clear heading on hackernews to their own heading)
If your (well paid) job is to write and communicate clearly, and for a major announcement you come up with this...not much left to say.
I don't think Walter is implying anything about how common or uncommon this is. His core insight seems fairly objective and plausible to me: "...your chances for happiness are increased if you wind up doing something that is a reflection of what you loved most when you were somewhere between nine and eleven years old". I.e. if you do end up being lucky and wise to do something as a profession closely related to what you *loved* doing when you were ~11 , because you end up spending time doing what you love (and equally importantly not spend that time doing something that sucks up energy) you increase your chances of being happier.
Kudos on the beautifully and thoughtfully designed landing page - which is becoming a rarity these days. Most product landing page highlight adjectives and abstract value propositions with links to join waiting lists for 'priority' access - providing little insight into the value proposition of the product itself. Not to mention the gratuitous visual effects.
So its a pleasure to see a thoughtfully done product landing page (which strongly signals that the same care will have gone into the product). The page is performant, no gratuitous visual effects. It clearly highlights the core product value propositions in the context of product visuals. Addresses key hesitations clearly and upfront (e.g. no cc required, pricing information), and a simple, obvious call to action.
Hopefully more people follow this template than the slop generated by auto generators.
Delightful explanation! A great example of how deep concepts can be made accessible and fun.
President Trump signs off on TikTok deal that puts US app's value at $14 billion
This metaphor may be misleading. For a compelling alternate view, read the excellent: "Is the cell really a machine?" https://www.sciencedirect.com/science/article/abs/pii/S00225...
From the article:
"It has become customary to conceptualize the living cell as an intricate piece of machinery, different to a man-made machine only in terms of its superior complexity ......"
" ..... However, the recent introduction of novel experimental techniques capable of tracking individual molecules within cells in real time is leading to the rapid accumulation of data that are inconsistent with an engineering view of the cell ...... which emphasizes the dynamic, self-organizing nature of its constitution, the fluidity and plasticity of its components, and the stochasticity and non-linearity of its underlying processes."