HN user

btown

19,398 karma

Email: btown.hn@gmail.com

Posts39
Comments3,173
View on HN
www.theregister.com 2mo ago

Google users fight for refunds as unauthorized API usage bills soar

btown
2pts3
old.reddit.com 11mo ago

I Have No Mut and I Must Borrow

btown
7pts0
benwiser.com 2y ago

I just spent £700 to have my own app on my iPhone (2022)

btown
4pts3
engineering.tumblr.com 3y ago

StreamBuilder: Tumblr's open-source framework for powering your dashboard

btown
2pts0
blog.danlew.net 3y ago

The LOGAF Scale (2020)

btown
1pts0
e2eml.school 4y ago

Sharpened Cosine Similarity for Neural Networks

btown
3pts0
www.merchantmaverick.com 4y ago

The MasterCard Match List and Terminated Merchant File

btown
1pts0
web.archive.org 5y ago

Muse Group head of strategy: musescore-downloader dev could be “deported”

btown
2pts0
corecursive.com 7y ago

Rethinking Databases and Noria with Jon Gjengset [audio]

btown
3pts1
github.com 7y ago

Official writeup of the Rust async-await syntax debate so far

btown
2pts1
www.standard.co.uk 7y ago

Runaway Saudi sisters urge Google and Apple to pull woman-monitoring app

btown
63pts3
twitter.com 7y ago

“Nobody at Tesla has ever seen rain before.”

btown
50pts3
www.playstarmancer.com 7y ago

Starmancer's Agent-Component-Event Engine (2018)

btown
2pts0
www.axios.com 7y ago

Flexport (YC W14) reportedly in talks to raise $500m led by SoftBank at $3B pre

btown
1pts0
news.ycombinator.com 7y ago

Ask HN: Any exciting under-the-radar discoveries/gems from NeurIPS?

btown
2pts0
www.reddit.com 7y ago

Style2Paints V4: Automatic colorization of line drawings via style transfer

btown
1pts0
slatestarcodex.com 7y ago

Sort by Controversial

btown
4pts1
www.youtube.com 7y ago

Bryan Cantrill on Oracle's closed-sourcing of Solaris [video, 2011]

btown
2pts0
qz.com 8y ago

NASA report: SpaceX design error led to 2015 CRS-7 explosion

btown
1pts0
twitter.com 8y ago

React will support async loading within render functions via thrown Promises

btown
2pts1
bartblaze.blogspot.com 8y ago

Anime video-on-demand site Crunchyroll hacked to deliver malware

btown
2pts3
voxeu.org 8y ago

Global productivity slowdown hiding increasing performance gap across firms

btown
1pts0
www.bloomberg.com 8y ago

Bloomberg: Study finds private startup valuations aren’t grounded in reality

btown
2pts0
medium.com 9y ago

Pnyetya: Yet Another Ransomware Outbreak

btown
2pts0
www.fcc.gov 9y ago

FCC Net Neutrality Comment Filings: “Null Null”

btown
4pts1
www.cjr.org 9y ago

Maneuvering a new reality for US journalism

btown
2pts0
www.kauffman.org 9y ago

Lessons from 20 Years of the Kauffman Foundation’s VC Fund Investments (2012) [pdf]

btown
1pts0
cidarlab.org 10y ago

Cello – a Verilog compiler for transcriptional logic in bacterial cells

btown
145pts34
medium.com 10y ago

Hot Reloading in React (or, an Ode to Accidental Complexity)

btown
1pts1
www.alleywatch.com 10y ago

How Yahoo Made $20+ Billion Disappear Through M&A

btown
3pts0

Nurses are instructed to stick to a script on phone calls and give no more than two to three pieces of advice, Capulong and other nurses said, which means they may sometimes need to decide whether to withhold advice or face a performance evaluation hearing.

Another nurse speaking on condition of anonymity said “AI did not understand our job and would grade us wrong all the time.”

It's always worth remembering Goodhart's law https://en.wikipedia.org/wiki/Goodhart%27s_law - "When a measure becomes a target, it ceases to be a good measure."

In theory AI could usher in the first time in history where one can escape from this trap - because qualitative judgments can be made at scale, from an unbiased and universal baseline. In this situation, for instance, rather than collapsing call transcripts and reports into metrics, it could evaluate whether red flags are encountered in the context of a call, and allow for qualitative guidance on improvement, across a comparative corpus of situations that are themselves chosen qualitatively.

But very few managers are empowered to take this kind of approach; they're evaluated by their ability to report quantitative metrics, and thus they must implement regimes of quantitative metrics. And leadership instructs them to use AI to build that regime more quickly.

If you want to see an "AI native" organization, it's one where leadership actively fights this tendency, and sees managers as product designers who make the end-user experience a beloved and empathy-driven one, as opposed to a gear that turns accountability into a single number on a screen.

The biggest limitation of the pelican is that it doesn’t touch at all on the thing that matters most for today’s model: agentic tool calling and the ability to operate tools reliably as conversations grow in length.

In all seriousness, I propose SWE-bench-adversarial-pelican-gen: it's like SWE-bench, but the harness gets interrupted every 5 turns/tool-calls and is asked to produce an SVG of an arbitrary animal before being told to continue, and every few tool call outputs add comment lines that refer to SVGs of pelicans (and, perhaps, how a møøse bit my sister once). And, at the end, once it's 800k tokens deep into context, it's asked to produce an SVG of a pelican and is evaluated against both the pelican and the completion and efficiency of the task.

You're only as good as your ability to solve problems in the midst of an SVG pelican attack.

Not a ban, perhaps... but if I have two users - Alice who uses the app according to normal patterns, and Bob who consistently predicts and declines when my dynamic pricing algorithm tries to push above market pricing, and has a remarkably high look-to-book ratio - I'll want Alice to have better drivers, better service, faster pickups etc., because I'll have larger lifetime margins from choose-the-app-on-vibes-Alice than from race-to-the-bottom-Bob.

If my known-good supply is limited at any given time, I have every incentive to focus it on Alice, and I'd be inclined to try out e.g. new drivers on accounts like Bob's.

Rideshare data teams are incredibly talented, capable, and motivated. One does not simply front-run a market where the biggest players have a massive data advantage, control your latency, and are effectively unregulated.

We've adopted a simple/similar Dropbox-based approach for skills and rules - each person's ~/.claude/skills is actually symlinked to a folder just for them inside a shared Dropbox folder, one that others on our (small) team can see and edit as well.

This solves a set of problems around people writing skills that reference artifacts or other skills that only exist on their system, and/or that reference their own name/information as the creator, and not knowing to make them self-contained and replicable. Luckily, adapting your colleagues' skills to self-contained versions and pulling them into your folder is trivial to instruct an agent to do. And you can have meta-skills that do this on the fly if a colleague has a skill that would unblock your project! (Editing to add a tip: make sure all the folders are set to offline visibility in Dropbox, rather than being loaded on demand from online.)

The courtesy simply has to be that you don't write into other people's skill folders unless/until they ask you to maintain something for them - at which point the words "I am assuming direct control" are said with all the necessary gravity and effect.

It's great to see someone putting UI and guardrails around this pattern!

As a counterpoint: in a complex project, Fable's "curiosity" may be exactly what you want for an exploration and planning stage - not just for the orchestrator that turns your prompt into different angles with which to explore, but for each subagent whose task is to search the codebase for one of those "angles." If you truly want no stone unturned, letting those subagents spawn their own discoveries, and recursively grow the surface area of the inquiry, then it's quite reasonable to want Fable throughout.

That said, if your project is "do this well-planned thing on a bunch of things in parallel" then you should absolutely be instructing to have subagents "step down" to less curious models. Their output may well be more cohesive as a result!

PostHog FOSS 13 days ago

It's also worth noting that in 2023 they abandoned their Kubernetes support which was relied upon by a full 3.5% of their users: https://posthog.com/blog/sunsetting-helm-support-posthog

In their rationale for this:

We also learned that the tools to do that automation just don't exist. We kept finding new failure modes. When onboarding a new customer we would have to vet their engineering team for Kubernetes experience so that we'd be confident they could help us debug issues in their PostHog deploy. Folks that didn't have infra experience would often be able to get something set up, only to get stuck when something went wrong.

I empathize that this is a sane choice for PostHog to make as a business. But - if you can't deploy and dogfood your changes, are you truly able to maintain a fork with customizations? And if you can't use your own changes, is the software open-source, or source-available?

Perhaps the punchline is that any scalable & performant web analytics platform must necessarily be a distributed system of ingestion and storage services, and that complexity is like oil and water with the classic "you should be able to swap out the dependencies on your systems with ones you fork" open-source ethos.

PostHog had an opportunity to break this trend, to innovate and invest in those automations they correctly said didn't exist - and I was cheering them on. I've been saddened to see them move in the opposite direction.

Something that many don't realize is that Claude Desktop and the Agent SDK are both just wrappers around pools of CLI instances, literally running `claude --input-format stream-json --output-format stream-json`.

So this wasn't even about third-party harnesses that replace the toolset Claude has access to, and try to call the Claude API with subsidized credentials. No, this was literally a blessing for their desktop UI's over others, all driving Claude Code's CLI at the end of the day.

To the broader point, it's hard not to see that as arbitrary and borderline spyware. Software that sniffs the context in which it's executed and uses that to phone home about billing is the type of thing you'd expect from the most corporate parts of the gaming industry, not a frontier lab, but here we are.

Writing an entire article about whale transportation, without even a subtle nod to Star Trek IV, or even a reference to tank material and the process of its invention, is borderline criminal. What is this, the dark ages?

It's worth remembering that a malicious model doesn't need Internet access to exfil - it merely needs to write code with subtle backdoors that will eventually run on a production system, and wait until its code is woken up by a system that will scan all known addresses and ports for the specific patterns introduced by the model's progeny. Which is not to say that this is happening in this case, or anything about which nation-state will be the first to attempt this - but we're only at the beginning of what's possible here.

For Claude, at least, "throw out the reasoning tokens" is only true when a session has been idle for more than an hour, and is new since March.

The basic concept is that for a session active recently, interleaved thinking tokens are already in KV cache, so it's more efficient to keep using them than not! But when resuming an older session where KV cache has been evicted, it's more expensive to restore the thinking tokens, so they're silently dropped from prior turns. It's 2026 and stateful servers are back on the menu!

https://www.anthropic.com/engineering/april-23-postmortem describes this as an intended optimization:

The design should have been simple: if a session has been idle for more than an hour, we could reduce users’ cost of resuming that session by clearing old thinking sections. Since the request would be a cache miss anyway, we could prune unnecessary messages from the request to reduce the number of uncached tokens sent to the API. We’d then resume sending full reasoning history. To do this we used the clear_thinking_20251015 API header along with keep:1.

The implementation had a bug. Instead of clearing thinking history once, it cleared it on every turn for the rest of the session... This surfaced as the forgetfulness, repetition, and odd tool choices people reported.

And https://news.ycombinator.com/item?id=47879561 is a thread with a Claude team member's further rationale.

Eliding parts of the context after idle: old tool results, old messages, thinking. Of these, thinking performed the best, and when we shipped it, that's when we unintentionally introduced the bug in the blog post.

(Also, https://news.ycombinator.com/item?id=47884517 indicates OpenAI drops reasoning tokens "smartly" at its own election, which is likely a similar performance optimization.)

I've experimented with rules to have Claude Code be explicit about recapping its thinking tokens, including tool choices and approaches chosen and rejected, into actual message output, but this is lossy at best. And sometimes dropping reasoning tokens can give a session "fresh eyes" in a good way.

I just really don't like the lack of control, and it's a reminder of how ephemeral the current landscape is. The Claude giveth, and the Claude taketh away.

While the panic is indeed nothing new, Meta could have chosen a path of solidarity across the tech industry, lobbying for the ways age/identity verification makes people of all ages less safe, especially in the context of phishing and data harvesting.

Instead, its strategy has become to advocate for increasing the net levels of tracking and regulatory burden, so long as it is positioned to burden other parts of the technology stack (namely, app stores and operating systems) rather than their social networks.

From the link from a sibling commenter: https://web.archive.org/web/20260429210901/https://tboteproj...

Meta spent a record $26.3 million on federal lobbying in 2025, deployed 86+ lobbyists across 45 states, and covertly funded a group called the Digital Childhood Alliance (DCA) to advocate for the App Store Accountability Act (ASAA).

The irony that their namesake Metaverse was meant to be, itself, an operating system and app distribution platform is palpable. When ambitions shift to regulatory capture, a shark has arguably been jumped.

Notably, though, Persona does not have access to your Claude interactions, other than your signup/verification date. They’ll train on your uploaded docs and photos, to be sure, but it won’t be correlated to your chats and projects, unless Anthropic is doing things that would make their counsel have heart attacks.

Both can be true. Promoting a standard isn’t free, and having licensing and certification fees, especially in an industry where such practices make a standardization org get taken more seriously, is a reasonable strategy. We’re lucky that our industry moved in a different direction!

It's also worth noting that when e.g. inputs to a stage might have unpredictable defects or alignment, a robot arm utilizing neural networks for planning and analysis might still be the best way to handle that - without the extra degree of freedom of movement-relative-to-floor, planning can be done more rapidly, and movement can be executed more aggressively and quickly.

If I were Hyundai, I'd be looking at this as buying a significant amount of vision, dynamics, and integration systems expertise, not necessarily the dream of self-motive walking systems.

The problem, of course, is the moment you have some aspect of the model not representable by the basic primitives - do you make it impossible to switch back to the beginner interface?

I'm reminded of the concept of "ejecting" from e.g. Create React App a few years back - the idea was that if your beginner-friendly interface is actually built on the same underlying engine (in this case bundler and deployment assumptions) you can have full fidelity when you need customization, albeit with a one-way transition.

In the JS world things moved more towards build systems where beginner-friendly-DX and full-configurability could coexist. I'm not sure that CAD has the same dynamic.

Perhaps something like nested layers could work: you can use a complex model as a layer, but only opaquely, and build things around it with solids-and-holes; you can then lift that to itself a complex model, do things with professional CAD, and then treat the result itself as an opaque complex model if you switch back? That gets complicated fast.

And the non-cancellable nature of grants is not just a nice-to-have, it's absolutely critical for research with upfront capital costs (buying equipment, building labs, etc.)

The very _fact_ that this is a policy is disrupting research, even if specific grants haven't been cancelled. Some universities are stepping in to backstop, but it's a powerful chilling effect.

There's even more to this - because among the possible outcomes is one where Fable is only made available to enterprises that have gone through Know Your Customer (KYC) processes, and perhaps only to verified users of those companies, and perhaps requiring biometrics and attribution so government can know who was using that account.

Now, say you don't want to sign a pre-committed enterprise contract with Anthropic. But oh, you already have such a contract with AWS, and they'll let you use any model you want, and they've implemented KYC and will graciously connect you with a solutions partner who can help you with the IAM systems integrations for key tracking and attribution.

Oh, and all these enterprise contracts will bill by token. We're not talking a small stake in a company selling subscriptions, we're talking immediate revenue at four-figure-per-user levels, and pushing more and more companies to see that as "just part of their AWS bill."

This is worth a significant amount of money to AWS. So the question is: does Hanlon's Razor apply when a $2.5 trillion company is putting its best minds into how to engineer strategic outcomes?

The regulatory capture angle here is, if anything, an implementation detail.

Another ongoing HN thread from yesterday around some exciting cancer treatment breakthroughs, this time with a CRISPR Cas12a2 mechanism: https://news.ycombinator.com/item?id=48505231

This subthread there is a fascinating explainer about one user's journey into funding and incentivizing research into their own rare form of blood cancer, and how they are able to push forward the state of the art: https://news.ycombinator.com/item?id=48506997 - something of a modern-day (and more accurate) Lorenzo's Oil!

Per that link: I think there's an interesting question about whether a nefarious actor who's infiltrated a cloud provider with physical access to machines that are running signed operating systems, with signed binaries, with TDX remote attestation, and with hardware supply chain verification, has the ability to break the privacy guarantees of a tenant with Apple's sophistication.

Certainly, one could tamper with the hardware, but could one do it in a way that wouldn't get that machine immediately flagged, removed from the routing pool, and told to wipe its memory immediately, by a watchtower (perhaps even the routing layer itself) that runs in a separate secure Apple datacenter?

So yes, it's a TUI... but it's a TUI rendered by Ink, a React library, with a full JS runtime in the background. The number of re-renders per unit time involved with rerendering a JS implementation of flexbox every new token comes in? That's not a walk in the park for a garbage collector, and a single memory/retention leak can cascade dramatically.

I imagine this is part of the impetus behind the Bun acquisition - they have a deep need to push optimization efforts towards the specific patterns that are most relevant to their use cases. (Which are probably good ones for the broader Bun userbase, to be sure, but relative prioritization is something they now have greater control over.)

a few GB per user at scale

While this might seem to be true for casual users, I recall that one of the reasons for Anthropic's recent changes for only retaining KV cache for an hour or so, was that many users just have one massive ongoing session that they continue on with multiple unrelated queries (as one would in a single-thread "group chat"). And this is hard to distinguish from someone who wants that context for their seemingly-unrelated query to apply tone etc.

So in practice, there are many casual users who are typing their Google-esque searches against a 100k+ token context window - and it's at that point where things balloon into 300GB+ KV caches to maintain.

I wouldn't be surprised if we see new UX's around subsidized plans starting to encourage resetting the context window more often.

I find Claude Desktop on macOS infinitely better at managing RAM when you have 10+ or more parallel sessions.

(Some for different aspects of full stack features, some for managing specific client situations advised by the codebase and its tools!)

The CLI is not designed to be lightweight, and it’s easy to get into situations where every CLI session consumes multiple GB of memory alone - stack them up and it’s a lot!

And not all terminal GUIs handle multiple tabs well enough to see all sessions at a glance.

So on top of the plugin features, desktop is a really useful thing to have!