I very much enjoyed Grok 4.5's rendition of the Mona Lisa as Elon Musk with tentacles.
HN user
maxall4
https://maxc.codes
Especially an MRI which is a 3D medium —something current LLMs are very bad at.
Or, alternatively, it may suggest that the NSA’s classified systems are not very secure, which seems at least as possible: they may rely on requiring physical access to these systems to even attempt to penetrate them.
It doesn’t seem like it? Unless I am misunderstanding these Nasdaq insider trading reports: https://www.nasdaq.com/market-activity/stocks/cbrs/insider-a...
This is rather reminiscent of the Bogdanov affair: https://en.wikipedia.org/wiki/Bogdanov_affair
We have reviewed the report and validated that the level of capability displayed there is widely available from other models (including OpenAI’s GPT-5.5), and is used every day by the defenders who keep systems safe. We will share more details over the next 24 hours.
So much for all of the rhetoric about Mythos supposedly far surpassing GPT 5.5 (edit: in cybersecurity, in particular). Of course, the AISI benchmarks also showed this, but it is amusing that Anthropic is saying it now that it is to their advantage.
Indeed, according to METR, Mythos only achieved an 80% success rate with 3 hour tasks. https://metr.org/time-horizons/
I’m very aware of this as well.
As someone who works in bioinformatics, and, as such, does a great deal of machine learning, this makes Fable unusable for me as well.
These safeguards are ridiculously sensitive: a prompt as simple as “ Why is an infinitely slow process reversible?” gets flagged as a ToS violation.
“The strait of Hormuz is open so long as Iran does not fire missiles at ships.”
Only an LLM could liken a first-grader to a scholar: "In a stratified society where the imperial family sat at the symbolic center, that gesture mattered. The randoseru moved, almost overnight, from battlefield to classroom, from soldier’s kit to scholar’s gear." This is an interesting topic, but this kind of AI writing gets very, very grating. Additionally, though this is somewhat unrelated, I feel like LLMs tend to argue points through gaslighting, rather than actual argumentation; they prefer to stack a bunch of tangential, or parallel, evidence and then assert that it proves their point when, in reality, it does not have any logical coherence—unless, perhaps, one reads it at 2am, in which case it might make sense.
Well, what else are we going to do while waiting for the bench scientists to finish collecting data?
Is OpenBSD actually more secure than Linux? I have not been able to find any data to support this—only some vague opinions.
It’s fairly overt: „The film was just starting. It was science fiction, of all things; called Dark Star.“ (pg. 111 in the mass market paperback)
I stumbled across a reference to Dark Star while reading the novella The State of The Art by Iain M. Banks.
At this scale, that kind of thing is not really a problem; you just dump all of the data you can find into the model (pre-training)1. Of course, the pre-training data influences the model, but the reinforcement learning is really what determines the model’s writing style and, in general, how it “thinks” (post-training).
1 This data is still heavily filtered/cleaned
on ~2% of new prosumer signups.
I, and everyone else I have asked, see this new updated sales UI; sounds like more than 2%.
I have a Claude Pro tier subscription; Claude Code, as of right now, is still functional for me. If Anthropic does boot Pro-tier users off Claude Code, I will be cancelling my subscription.
I was part of a team researching MS at a university a while ago. It truly is an endlessly fascinating disease. Most evidence currently points to MS being caused by a combination of Epstein-Barr infection and genetic factors [0,1]. It is hypothesized that Epstein-Barr triggers autoimmunity which results in the prototypical demyelination [2].
[0] https://www.science.org/doi/10.1126/science.abj8222
I smell bad data. This sounds too good to be true and most studies of this kind have turned out to be false a few years down the line.
Edit: one of many examples: https://www.science.org/content/article/journal-retracts-inf...
Theoretically, you can’t benchmaxx ARC-AGI, but I too am suspect of such a large improvement, especially since the improvement on other benchmarks is not of the same order.
I think you could have discovered this bug more easily by looking at the commit(s) that were made when the problem started.
In this article we'll tell you why we decided to put Claude Code into RollerCoaster Tycoon, and what lessons it taught us about B2B SaaS.
What is this? A LinkedIn post?
Starlink receivers are actually very complicated. They make use of a bunch of high-end FPGAs and a bunch of other expensive and uncommon components. See this teardown: https://youtu.be/h6MfM8EFkGg?si=m-sN6UW4nh8_HzPR.
If I read one more article/press release/whatever with such clumsy use of antithesis, I’m going to go insane. I have no problem with using AI to write if it is done well, but this…
Sure, but I’m not convinced that producing a blacklist and filtering system is that difficult. More importantly, it’s little things like this that slowly and insidiously degrade the user experience. Sure it starts with one 300ms API call, maybe most people won’t notice. But when you reach for solutions like this to every minor technical problem, the next thing you know it takes 5 seconds to sign-up.
I can’t tell if this is some complex joke or a real product. This is literally string.contains() as a service.
Edit: 300ms?!
The problem with HTML and CSS is there are encapsulation boundaries where there shouldn’t be. Tailwind, by contrast, does not separate the layout from the styling; creating a more cohesive developer experience. Anyone making a point like this does not understand why Tailwind—and similar libraries—are superior to classical encapsulated HTML/CSS.