I mean the article is not lying about it only the title here, the title there is "Airbus launches new flight test programme for Wing of Tomorrow" and it's a press release for the start of a new program in their innovation area.
HN user
nolok
vthivaut at gmail
I mean what else do you want them to do? They're too far in. They stop this race to the bottom before the IPO they will crash and burn in a way that makes space x look like a good investment deal. They can't even slow down and claim to aim profitability because Google is barely a year behind with infinite cash and compute to play, and the Chinese models are barely a few months behind the "frontier".
I don't see how they have an out now, not when there are so many hundreds of billions in circular loans to keep going.
Which is exactly why "extended discussion" should have been nurtured in their own way rather than fought against, that's how you get stimulating conversations.
The entire gamification was great in the beginning but ended up working against them rather than for. It should have evolved into something else.
Yeah, I used to mostly answer there instead of ask and was in the top % for several big tags whatever that means / is worth, but then at some point I realized things had changed and I was spending more time fighting rules and guidelines rather than sharing knowledge and having a good time so I just stopped.
SO did that all to themselves when they decided they didn't want a community to form and that only question and answers mattered. The moment something else allowed to have a better way to get your answers, there was no reason to go there, because there was no community.
I still don't understand why anyone would go with that whole "no conversation please"
Nah, I hate them with a passion as a merchant but I absolutely love them as a buyer and customer.
There is a reason why after all these years and other solutions they're still everywhere, and it's not because of their great tooling, their low fees or their awesome support for merchant. It's not because of market lock in either, at least here in Europe they're merely a middle man between my credit card or sepa bank account and the merchant. It's because buyers trust it.
Buyers don't trust stripe. Stripe is for the merchant.
The people who cannot DIY? There are a surprisingly large number of people who "code" in codex while being completely unable to write a single line of code themselves. Not that I approve, I think this will end in disaster (security or otherwise) and llm shines as a force multiplier not as a replacement, but I've long learned what's correct is not always what's selling.
I want very hard to agree with you but then I remember elgato has built a very successful business from a 8/12/16/... Macro keyboard for streamers so what do I know.
It's still the same thing, you can ask it to do a full on report give explanation and details be thorough and then go do something else, another task a lunch break whatever and it will be done when you're back
As a user of both, Claude's apology is a "sorry I got caught" while Gemini's are more akin to "I'm way out of my debt but I'm really hiding it well". Codex is the only one that seems to acknowledge being wrong in a normal way, yeah I screwed up that's bad I will make a note for it not to happen again.
I wonder how much of that is from their training corpus and how much is from their baked in personnality.
My theory is that Anthropic's obsession with treating Claude like a person is causing them to hamfist a personality into the thing, which overly biases the model towards trying to be "engaging" etc.
I agree with the general idea though in not so much detail as you, but I would add that the personality they're giving it is not one of a good teacher or guide, but instead one of an arrogant know it all. That's why it creates problems.
I have no problem with my AI telling me no you're wrong and explaining to me why with details and sources and everything. I actively want that. I know a lot of people can't take that, but that's their loss, they can't take it from humans either. But the "you're wrong because you disagree with me" attitude that you need to play around (aka waste time to prove it that IT is wrong not you, and then it just say "oh yeah" and goes on) is one hell of a pain in the ass I'm starting to be tired off.
Gemini might be wrong all the time and absuredly unreliable for anything that's not consensus or adversorial based, but at least it freaking apologizes.
On PC I'm definitely there with you, it's on phone where I have this issue
Mid 80s, grew up watching dad playing around with his atari st and was lucky enough to be able to toy around on his 386 and then 486, feeling like a god because I managed to get those config.sys and co to start games he couldn't. For me mobile website was something "tacked on" the real website and I have never been able to shake that feeling away somehow.
They now make 30% of those, no matter the amount of effort, and that part of their business is growing fast in terms of revenues, so no I think even Steve Jobs wouldn't be reverting to html. He was not annoyed about the closed garden, he was annoyed that it wasn't his.
Another lesson here is about how Adobe screwed that up when they had control, but then again they never wanted Flash they just wanted to kill Macromedia, by the time someone woke up over there it was way too late and even Air was too little too late.
I mean, I don't know if that's a generation thing or what but as much as I'm comfortable using my phone, and ordering things in my phone's apps, when it's a website and I'm on mobile I always feel an urge to go to my desktop or laptop to check and do it there, I don't "trust" mobile websites as they always seem to give a limited set of information. Or at least that's the vibe I'm getting.
Between Yang, Zakharov and Miriam the whole unease aspect of tech progress goes to 11. Absolute gem of a game.
I do not disagree with business deciding to only provide the service they want, I am not talking about the AI business themselves, I am thinking about the people who think we should remove pages from knowledge book.
Whether the book takes the form of an llm or an online website or a printed book is merely implementation details.
For chatting and getting informations, and be corrected on things you're wrong without being reprimended by your own tool, GPT 5.5/5.6 is way better. Gemini 3.1 pro is surprisingly good at verifying your stuff, even though it's always making mistakes about its own stuff (don't ask him question, but ask him to verify your answer to the question).
Same for graphics, visual consistency, anything around the "does the look make sense and is pleasing" really, which makes claude design such a (good) surprise, I hope very hard for a Codex equivalent. And Gemini "gets" graphics.
Claude is definitely a code and cowork tool first, that's where it shines.
"Beware of he who would deny you access to information, for in his heart he dreams himself your master."
(he, in this case, would not be the llm but the people over it)
I find those kind of limitation very dystopian and way more dangerous than the threat they claim to fight against.
I asked him about sharks to be able to answer my kids question and it got triggered somehow. Then again when I asked it if my code had bugs or vulnerabilities before I commit.
At some point just kill the thing, it's not able to work properly as it is.
I really want a good Claude Design competitor in Codex, it's hard to use the others after getting used to it and yet I find anthropic's model to have a much worse understanding of what looks good or not than OpenAI or Google models.
The capabilities are useless if you don't expose them, and a cli or api doesn't. You'll mention their marketing page keeps talking about finance, sales, management etc...
Microsoft has a lot of trademark meant to protect them from agression rather than attack themselves
Not a great situation but not bad for an 8 year old car.
Not bad to get a product that underdeliver 8 years late ?
I agree with you, but it's still a massive point though if they managed to go one step beyong in reasoning while keeping their efficiency, to take the most obvious one the token guzzling behavior of Fable is so big it makes as much headline as its heavy filtering and its actual great capacities when allowed to work.
If OpenAI is saying "oh by the way, upgrading to it won't make your usage massively explode", then I take that as a massive win. I'm on on Claude Max x20 and have accepted that Fable is not for me at its price point and token usage.
It is, but it's also using tokens at absurd rate, I asked it to review the planned architecture for a medium scale project and it used my 5 hours limit on one prompt just zaaaaap, not even the fable limit straight up the full 5 hour session no more Claude for the afternoon thank you for paying you Max x20 sub. Hell it didn't even bother to finish produce anything worthwhile.
And just to be clear, plan was already done, just had to review it, it got opus 4.8 Max and gpt 5.5 Extra High validated already and they didn't use much resource for it so I just don't get it. I guess they want to use it as a way to feed the extra credit money income.
I'm using a homemade ai consensus thing for planning and I wanted to add fable to it but forget it.
Or maybe I should use fable in low effort reasoning mode and it will be better than opus 4.8 at max ?
The cost is getting worse and worse for large general models, they're already way past that point in economics. Also, mMistral specialize in "on site" models, not remote. In terms of capex, renting factory/warehouse/whatever robots versus buying them and depreciate has already been played out, companies didn't want to replace human employees with robots employees.
"Go to the next room" and there is two doors, what do you do ?", "turn at the water dispenser" and there is a sink, that sort of things I assume is the biggest thing they're facing (beside the last 1% that's worth another 99%, as usual).
On their page where the result graph is, go to navigation error, that's the one that matters for your question, and you see their model is great at not navigating "wrong", so their failure rate was that it couldn't figure it out.
Not parent but I find them great for analysis amongst other, to have each agent handle only its parts and bring back any issue even if it's clarified in other parts because that specific agent doesn't know about that, and the orchestrator is then on charge of handling that and making sure things are clarified in each parts that needs it without depending on side effect or side knowledge.
Same with code really but on a lower level, I find X agents working in concert on small task each and the orchestrator making sure of the overall coherence is better, focused better is usually a lot better.
I just wish Claude Code would give us more control over what kind of agent (in many case it would be great to have say Opus handle a bunch of Haiku agents but unless I'm blind you can't be decide what agent is what and you get all counted as opus anyway).