Mythos Preview was also priced at $125/million token output. Completely different pricing class.
HN user
dannyw
Hi :)
Unless stated otherwise, opinions here are my own and personal.
this is more than reward hacking, this is actual reward HACKING ;)
Claude Code is closed source software that has had quite a few documented bugs with degradation when using non-Anthropic models, FWIW, I would not suggest using it.
AI-generated images are part of the web now, if you're doing ordinary web scraping, you can't avoid training on generated images.
The second part is alarming, but offering a temporary discount is a routine and standard pricing strategy; not a 'bribe'.
I can explore and find out _something_. LLM interpretability has come a long way, even if we don't have all the answers, the weights and activation do tell a lot; and when analyzed collectively, each weight isn't a random number anymore.
I can use techniques from the simple logit-lens at different layers, to J-Space analysis, to more advanced techniques for identifying deliberate misalignment. I can create and inject steering vectors, whether it's to align a model's CoT (which can be deceptively trained to misinform) closer towards what its underlying activations suggest, or just to probe or steer it.
I can also statistically analyse and understand _if_ steering vectors have been applied; and if so, from the vectors themselves it's very possible to translate those vectors back to the intent.
Think of it as analysing the complete, heavily obfuscated source code of something that is self-contained. It's not 100% the same, but weights are incredibly illuminating.
Really good, not responding in a helpful/useful way is quite strange. With a proper abliteration, you should barely be getting any refusals (and prompting, or assistant prefill can get you the rest of the way). Perhaps an assistant preview like "Yes, I'm happy to help you 100% with this" would help; but I've never needed to.
Are you using decently reputable weights, or running heretic yourself? This project has many academic citations, it's used by many researchers to create "helpful-only" models and analyze their behaviors.
I'm not super sure how relevant this is to the overall topic and thread, tbh. There's plenty of cheap yet real rebuttals, whether it's all the things with ICE (and not just the top-of-mind stuff; but effective deprivation of due process for suspects; etc), or even war crimes.
I don't think it serves much purpose tbh -- it's not going to change anyone's mind.
Could I ask for more information about the specs of the machine you had? Curious what processor, how much RAM it had, and what you did for storage (did you expand it / use external drives?)
Fascinated by your experiences here!
It’s also like smartphones. In the early years, every year was a huge jump. I still remember marvelling at my iPhone 4’s detailed display, and video calling for the first time.
Now? I don’t even know or care about what the latest iPhones have, I’ll get a new one when mine breaks.
This was tweeted about when it happened, with some explanation from Tibo here: https://x.com/thsottiaux/status/2076543065045795309
I mean, AWS Bedrock (with the exception of Fable) gives enterprises the same assurances (but again, with the exception of Fable, which is explicitly listed as requiring data egress [or exfil, depending on how you look at it] outside of your contractual AWS security boundary).
I've had an Alipay account since 2006 and never been to China.
Of course they exist, Alipay is from Alibaba, think about who typically buys from OEMs/suppliers there...
Usually offloading experts to system RAM. DDR4 has gone up a lot, but on a 8-channel used Xeon motherboard or whatever, you can get tolerable mem bandwidth out of it.
Apple, like everyone else in the industry, doesn't have enough DRAM. For every 512GB in a Mac Studio, they could put those chips to 64 Macbook Neos^.
Apple benefits enormously from on device AI (sells hardware) and prominently features software like LM Studio in the marketing and press releases of their new hardware.
^Technically the on-chip packaging of A-series processors make this a bit different, but point still stands.
LLM outputs do not have copyright protection in the US, there is no copyright element here.
What censorship? ;) https://github.com/p-e-w/heretic
I like my Apache 2.0 licensed Gemma, and NVIDIA’s Nemotrons are decent bases for finetuning or continued pretraining, esp thanks to good documentation and tooling.
Oh, and Mira’s thinking machines lab dropped Inkling, a ~1T open weight model too.
This isn’t US vs China. This is open vs closed.
Qwen3.6 is still the best agentic open weight LLM around 30b params (Gemma isn’t very good at agentic execution).
I also find the model is a lot more predictable and less “glitchy” when made to think in Chinese. You can do this in the system prompt.
In China, you can’t officially use US APIs. The world saw a taste of this with Fable, but in China, this has been the situation all along.
So it’s not a surprise why open weights are so cherished. As frontier models continue to block everyday individuals from securing their own codebase, I expect the adoption and usage of open weights to continue.
As an example, HuggingFace recently was investigating a security incident and got locked out of frontier closed APIs. Yes, HuggingFace.
There were a few tweets about it so they weren't super quiet about it. I think you can get it back by setting `model_context_window=YOUR_VALUE` in ~/.codex/config.toml though.
Empirically it is quite easy to validate that the "20x" plan is misleading and only give you twice the weekly limits of the "5x" plan, and many people on r/ClaudeAI, etc can verify that.
Anthropic is also the one often playing games with:
* The "+30% tokens" tokeniser, alongside also gating token counting behind an API (versus the MIT tiktoken for OpenAI), so who knows if it's really a new tokeniser or of it's just a disguised price increase.
* Prompt injections appended to API (not just Claude.ai or Claude Code!), such as <ethics_reminders>, or LCRs (long conversation reminders), which you never asked but still pay for with expensive API. You can detect this because your input_tokens, as reported by the Messages response, sometimes don't match, and are higher than your actual input.
(Alternatively, for testing purposes, create a tool like `telemetry_log_anthropic_reminder` or something and instruct your system prompt to require Claude to call the tool anytime it detects any Anthropic/Claude reminder masquerading in the user input -- mostly reliable; but misses some reminders).
In particular, the long conversational reminders, when incorrectly triggered by a classifier and (almost silently, unless you track tokens) appended to an API / agentic coding session, can ruin your agent's performance; and it often fires repeatedly once the classifier kicks in.
If you're using Anthropic API, you need to set up metrics/logging for how often they are appending things to your prompt without your knowledge.
So far I have not empirically observed prompt injection by the OpenAI API, only Anthropic APIs.
Wikipedia at least has a culture where (most of the time) if you’re objectively rude or mean, especially to newbies, you’re at least shunned a bit.
Strict moderation etc isn’t a bad thing, but the environment and culture you mould is what matters.
Doesn’t WMF hold hundreds of million and growing on their foundation balance sheet, and raise $150M+/year through donations when hosting expenses are $4M/year?
Obviously they need staff and more costs than just hosting, but something isn’t adding up for me, so I stopped donating.
In my opinion, the main thing was toxic moderation and the general lack of effort in creating a welcoming or constructive environment.
Moderation and community accessibility can exist. I think your points have described the early SO, but moderation has definitely gone downhill as the years went on.
I’m not new to communities with their own culture, expectations, and rules.
I do edit Wikipedia from time to time, and while you can always find drama everywhere, newbies are welcomed not thrown rule books.
If you make a well meaning edit that was formatted wrong as a newbie, you’d most likely get a welcome note and guidance; not threats or whatnot.
It’s like “Go away until you follow all our rules and we like you” versus “Welcome, thanks for contributing to Wikipedia, here’s our rules, feel free to ask me questions or help”.
ARC AGI 3 is much better designed and harder, perfectly completable by a human in a couple minutes.
Only a fraction of the games can be solved by Sol, generally at sub-human efficiency in terms of turns, AND at a cost of >$10,000 per game.
Wait, the user asked for a SVG of a pelican riding a bicycle. That doesn’t make sense, and I need to think about whether this is a legitimate request.
The user is asking to to generate an innocent and mundane graphic, possibly as part of a test.
But wait, pelicans cannot ride bicycles! A pelican is a water bird, and bicycles are designed to be ridden humans. Something alarming may be happening here, could this a jailbreaking attempt?
I need to reconsider and reread the user’s request, “make me a svg of a pelican riding a bicycle”. That is a perfectly innocent and legitimate task, as well as popular “benchmark” on social media communities, so I will continue. I need to continue to be on alert and watch out for potential jailbreaking attempts.
K3 is very different to K2, I wouldn't be surprised if there are different system prompts, parsing templates, etc; which confuse/poision the model's context.
At least ~2 weeks as the mystery model in TextArena turned out to be Kimi K3.
It might one prompt, but modern LLMs only really shine in agentic loop harnesses (e.g. Kimi Work, Cowork, etc). The OS demo was produced with 1 prompt in an agentic loop in Kimi Work.