steady lads, deploying more capital
HN user
czk
if you're benchmaxxing then maybe bigger doesnt always mean better, but for general intelligence and big model smell, that couldn't be further from the truth
the oss models are impressive but it's pretty clear how quickly they fall off when you try to use them outside of a narrow set of problems they benchmarked well on when compared to opus/5.5
There was also a bug where you could cancel the subscription via the iOS app store and if you never opened the iOS claude app again, you'd keep the subscription forever and could use claude via the web, without paying.
Also when they added extra credits to everyone as an apology I was able to click the claim button multiple times and I got up to $400 in credits. Eventually a day later this dropped to $200 and then a few days later, $100 where it sits today.
They mention it uses MXFP4 quant which is a blackwell capability but it looks like this is also supported by ascend 950 series according to marketing material
i wonder if they put an older cutoff date into the prompt intentionally so that when asked on more current events it leans towards tool calls / web searches for tuning
the model obviously knows things after the reported date but its just curious that it reports that date consistently
could be they do it intentionally to encourage more tool calls/searches or for tuning reasons
with thinking off and tools disabled:
Donald Trump won the 2024 U.S. presidential election.you cant but its pretty reproducible across api and codex and other agents so i just thought it was odd. full text it gives:
Knowledge cutoff: 2024-06
Current date: 2026-04-24
You are an AI assistant accessed via an API.
# Desired oververbosity for the final answer (not analysis): 5
An oververbosity of 1 means the model should respond using only the minimal content necessary to satisfy the request, using
concise phrasing and avoiding extra detail or explanation."
An oververbosity of 10 means the model should provide maximally detailed, thorough responses with context, explanations, and
possibly multiple examples."
The desired oververbosity should be treated only as a *default*. Defer to any user or developer requirements regarding
response length, if present.API page lists the knowledge cutoff as Dec 01, 2025 but when prompting the model it says June 2024.
Knowledge cutoff: 2024-06
Current date: 2026-04-24
You are an AI assistant accessed via an API.Memory bandwidth is the biggest L on the dgx spark, it’s half my MacBook from 2023 and that’s the biggest tok/sec bottleneck
"adaptive" thinking
show us the benchmarks with "adaptive thinking" turned on
the MDM profile requirement is suspect though I get why they are doing it. but it doesn't inspire confidence to see that their profile is unsigned and still using the default micromdn scep challenge...
it should, lume is a thin wrapper around Apple's Virtualization.framework as i understand it
starting with M3+ you can use Hypervisor.framework/Virtualization.framework to spin up nested VMs.
it would be amusing if that bypassed the limit.
I tried the periodic table in their examples using sonnet 4.6 on the $20/mo plan. After a few minutes Claude told me it reached the max message length and bailed. I pressed continue and eventually it generated the table, but it wasn't inline, it was a jsx artifact, and I've now hit my daily usage limit.
claude models with 'extended thinking' toggled answer very quickly and the quality of the answer is far ahead of what gpt 5.2 'instant' provides. i wont even bother using the non-thinking version of chatgpt because the quality of the answers is awful and usually incorrect.
i spend most of my time with claude thinking about when my daily usage limit is going to reset
“After creating a new account, I can confirm the quota drains 2.5x–3x slower. So basically Max (5x) on an older accounts is almost like Pro on a new one in terms of quota. Pretty blatant rug pull tbh.”
lol
claude has half the context window size of codex and blows through a good percentage of it right off the bat by injecting a system prompt the size of don quixote
it sounds like the data can be involuntarily disclosed to an external third party (the attacker’s domain) purely because someone reviewed logs that auto-load remote images
their log viewer renders the markdown and their browser will make a request containing the sensitive data to the attackers domain where it can be logged and viewed
i thought the article was going to go there, just redirecting the host to a self-hosted ip address serving the bin, but i was pleasantly surprised it didn’t! interesting to learn about the patching process and tooling used
like leveling to 99 in old school runescape
its possible to use gpt-5-high on the plus plan with codex-cli, its a whole different beast! i dont think theres any other way for plus users to leverage gpt-5 with high reasoning.
codex -m gpt-5 model_reasoning_effort="high"
it only requires exponentially MORE money for linear returns!
eventually traditional operating systems will cease to exist, you'll just have a model creating dynamic UX for you on the fly for whatever experience you want
ironically... they use cloudflare.
the year is 2045.
you've been cruising the interstate in your robotaxi, shelling out $150 in stablecoins at the cloudflare tollbooth. a palantir patrol unit pulls you over. the optimus v4 approaches your window and contorts its silicone face into a facsimile of concern as it hits you with the:
"sir, have you been botting today?"
immediately you remember how great you had it in the '20s when you used to click CAPTCHA grids to prove your humanity to dumb algorithms, but now the machines demand you recite poetry or weep on command
"how much have you had to bot today?", its voice taking on an empathetic tone that was personalized for your particular profile
"yeah... im gonna need you to exit the vehicle and take a field humanity test"
this is good to know! thank you for the info
I've often thought about the amount of data that these bot services must have access to (they could log millions of private channels), data thats silo'd away from search engines/indexers and could be pretty valuable to sell to someone training an AI model, or doing other things.
A while back there was a service called 'Spy Pet' that ran hundreds of discord bots selling access to searchable data logs. I wonder if discord is primarily concerned about the massive logging capability of services like these.