HN user

gck1

738 karma

saggy.eskimo294@passmail.net

Posts8
Comments332
View on HN

It's always slightly amusing to me how SE has the most restrictive cloudflare turnstile config set.

On the rare occasion that I DO want to go there, they greet me with an impossible gate that takes 15+ seconds to pass on Brave and gets invalidated quickly.

If anyone from SE is reading this: you already failed to protect your data from the LLM crawlers, SE is no longer that much interesting to them. The only visitors you're gating today are not bots, they're humans.

And they had actual community! Never have I been frequent on any manufacture's discussion forum, but with OnePlus, I was. It was wild to see a manufacturer helping you with rooting their device.

Then I saw the announcement that they'll be merging the bloatware Chinese version of their OS and that was the last day I held my OnePlus phone.

Had they not taken that path, I'd likely have bought at least 3 of their phones since then.

GPT 5.6* throw fits on anything even remotely related to reverse engineering, and I'm not ever paying anything more than $20 to Anthropic anymore.

How's Kimi in this area?

And recently, since GPT 5.6, OpenAI basically doesn't show anything but a single line, 5 word titles of reasoning traces - titles of summaries of reasoning i presume.

It's effectively just a completely hidden thing now.

You can't drive prosumers anywhere near API prices. I would guess that the maximum you can extract from vast majority of prosumers is maybe $500/mo, and even that is a big stretch.

Once you cross that threshold, prosumers will simply fall back to using Chinese models and/or self-hosting smaller models, with more efficient and tight workflows.

You'd be killing your consumer line completely.

Both are verbose in their own way, and both - terrible. Claude models love to throw huge blobs of text in architecture planning / interview conversations, but in not a mentally draining language. OpenAI models are more compact, but very dense & formal - they will speak in RFC language for a button that clicks and submits a form.

So claude: 10 paragraphs of prose

codex: 1 paragraph of jargon over jargon.

Forget 5.5-pro. Why isn't everyone talking about the fact that there's no 1M context window model in codex?

Yes, I know attention degrades above ~200k, but it's still useful in many applications.

What does it matter which tool I use when I hit the limit?

A third party harness may have a misconfiguration of prompt caching, leading to more load on Anthropic's servers, they could also have wildly different usage patterns (Hermes, openclaw etc), also making it hard to predict the load. If you constrain everyone to use your own harness, at least you're in control of that.

The first point is sort of funny though, because their own harness had a misconfigured prompt caching for several months during the times where they were crying about third party harnesses (/resume was busting cache all the time).

I'm never going back to claude from codex, including for the reasons you mentioned, but it must be said that web chat inference on ChatGPT is magic incantation, and I'm almost certain they're not serving the same models there as in Codex.

Claude web definitely feels like it's the same models behind as in API, with much less extra behavior/layers that make it behave differently.

This seems to be conflicting with the adoption of Web Bot Auth, which is still in infancy stage.

I do have some bots, they're nice and predominantly used for grounding AI harnesses which I use interactively. Knowing that most operators will whitelist maybe 5 well know bots and route the rest to the micropayments, what's the incentive for me to have my bots identify as bots with Web Both Auth when it's easier to make them mascarade as humans?

Again, my bots are nice. They're making roughly the same number of requests I would make manually via browser if I was manually working on something.

Since Anthropic has eroded all the trust it could possibly have, I'm going to allow myself to be cynical and say that this move is just another pillar of their shady marketing practices.

I know a few real persons who will praise Fable solely because it's scarce and unavailable to them. Heck, they've already been doing that in the past month, as if Fable allowed them to do unimaginable things.

And once this play has done its job, Anthropic is going to come as a savior and put it back into subscription, driving even more hysteria and visibility.

There's probably a name for this tactic in some marketing playbook which I'm unaware of.

I've been working on my own private harness for the past 8 months, and I've been collecting ideas from such repos I've stumbled upon.

pi-tmux is one such example (seems to be archived now) which inspired me to use tmux as communication layer and provide visibility of subagents of multiple models in their native harnesses [1].

There's also herdr, which is not 0-stars, but is super interesting but lesser known project [2]. This also has interesting substrates to allow agent coordination.

None of these are harnesses per se, but they're pointing towards clear gaps in existing harnesses. For example, we've known for a while now that compounding knowledge of different class of models achieves better performance. Why is there no harness where this is a native functionality? And there's no harness where subagents are first class citizens both in terms of capabilities and UX.

[1] https://github.com/offline-ant/pi-tmux

[2] https://github.com/ogulcancelik/herdr

It's sad to see that the teams that have the most resources that can contribute to development of next-gen harnesses are essentially copying the same exact thing from each other, with no meaningful changes.

And most of the advancement and experimentation happens in some random 0-star github repos.

It's very likely that OAI models will have even more restrictions. Firstly because now they know what feds will do if you don't tune the safety classifiers towards more false positives and secondly, OAI models were always more restrictive than ANT.

This has no mention of what happens to the prompt cache, including the "learn more" link.

Knowing Anthropic, it wouldn't surprise me if it will result in a full cache miss/rewrite at fallback, with potentially up to 1M tokens in the context window.

Location: Europe

Remote: Yes. 12+ yrs, comfortable fully async across US/EU hours

Willing to relocate: No

Technologies: Rust, Python, Go · LLM/agent systems, multi-agent orchestration, LLM eval · browser automation + anti-detection (CDP) · data pipelines · dev tooling

Senior engineer & founder, 12+ years. Last ~2 years deep on Rust + LLM agents:

- Built a Rust CDP / browser-automation driver from scratch (anti-detection, agentic pipelines).

- Multi-agent orchestration + an LLM judge/eval layer - claim-level extraction and verification against ground truth for my own AI SaaS.

- Built AI harness orchestrating different agents in their own native harnesses (agent-to-agent comms, clean room reviews, sandboxing)

- Wrote an architectural linter enforcing pure/effectful separation in large Rust codebases.

Before founding: Software Engineering Team Lead on high-throughput messaging infrastructure. Earlier: national-scale public data platforms (company registry, procurement monitoring, government transparency).

Strongest at taking an ambiguous problem to a shipped system solo, hard systems work in Rust, and LLM evaluation/verification/orchestration. Most interested in agent infrastructure, dev tooling, LLM eval, and automation.

Open to: senior/staff IC or founding-engineer roles, or focused AI contract work.

Résumé/CV: Upon request

Email: In profile

Based only on the third quote (you're literally in the thread discussing second iteration of it), and your username, you can't possibly be acting in good faith here, so I'm not going to waste my time providing references to the events that were widely discussed even here on HN in the past 6 months.

This started a few months ago when anthropic started beating openai.

From where I'm standing, this started a few months ago when Anthropic decided to gaslight users, sabotage their projects, ship malware and attempt a regulatory capture.

If there's an anti-ANT propaganda, it is solely of ANT's own making.

The gap between Chinese models and American frontier models is estimated at 10 months by Anthropic themselves, and it's growing.

#1 I've had use cases where it was clearly obvious the Chinese models were behind.

#2 I've also had use cases where I couldn't tell a difference at 1/20th of the price.

The problem is - the #1 is the use case where American frontier is gated behind saboteur classifiers and is tiny minority anyway. Vast majority of work is #2.

The gap doesn't matter anymore.

The trigger is ANTHROPIC_BASE_URL, Claude Code's API base URL override

I had a use case where I had to MITM CC's traffic to strip credentials that could have accidentally made it into the harness.

I'm happy my paranoid self told me "You don't really know what they're doing with that flag or if they're honoring it for all requests", so made a decision to proxy it at the network extension level.

Also, does anyone remember Anthropic quite literally sabotaging your project if the classifier in front of fable thought you were working in the AI industry? After backlash, they pulled it back, now they did this. Anthropic is on a weird tangent to ship malware. If someone doesn't stop them, one day, this will backfire catastrophically.

It's not really a 90% discount (I went into the rabbit hole) and none of the sites from this list are what people use (looks like some labs and random sites). It's more closer to 30% specifically for Claude models, and it's constantly changing.

It's also a discount relative to API prices. It would still be much more expensive than a Claude subscription, because that's what these providers are actually doing - pooling subscriptions.

Eh, pretty much everyone that spent some time tweaking their harness already had a homemade 'ultracode' long before Anthropic did it.

OpenAI is just way more careful with what features they add or enable by default in their harness. Anthropic's harness is a junk drawer of random features, with a new feature added every few hours. It feels like they're in panic mode, dropping random things to see what sticks when models are eventually commoditized.

I prefer OpenAI way - slow and steady.