HN user

cjbarber

2,484 karma

chris.barber@alumni.stanford.edu

https://x.com/chrisbarber

Posts184
Comments237
View on HN
slashquiz.org 3mo ago

Show HN: I made a quiz to help people learn Claude Code features (with macOS UI)

cjbarber
3pts3
majorcontext.com 3mo ago

Moat: Run AI agents in isolated containers

cjbarber
3pts1
livenearfriends.com 3mo ago

The Friend Compound Field Guide

cjbarber
3pts0
huggingface.co 4mo ago

Synthetic Data Playbook

cjbarber
1pts0
www.conference-board.org 5mo ago

The Leading Economic Index for the US Continued to Decline in December

cjbarber
1pts0
arxiv.org 5mo ago

Soft Contamination Means Benchmarks Test Shallow Generalization

cjbarber
2pts1
epoch.ai 5mo ago

How close is AI to taking my job?

cjbarber
2pts1
news.ycombinator.com 6mo ago

Show HN Guidelines

cjbarber
2pts0
www.noahpinion.blog 6mo ago

Zero-sum economics keeps failing

cjbarber
3pts0
twitter.com 6mo ago

There are broadly two ways people think about AGI and labour

cjbarber
2pts0
www.vibekanban.com 6mo ago

Tips and best practices for working with AI coding agents

cjbarber
2pts0
sankalp.bearblog.dev 6mo ago

A Guide to Claude Code 2.0 and getting better at using coding agents

cjbarber
5pts0
chrisbarber.co 6mo ago

I asked AI researchers and economists about SWE career strategies given AI

cjbarber
13pts18
secondthoughts.ai 7mo ago

Presenting the Case That the Future Will Be Unrecognizable

cjbarber
2pts1
vitalik.eth.limo 7mo ago

Let a Thousand Societies Bloom

cjbarber
3pts1
www.smithsonianmag.com 7mo ago

How to Watch the Radiant Geminid Meteor Shower Tonight

cjbarber
7pts3
walkingtheworld.substack.com 7mo ago

Why Are Americans Unhappy?

cjbarber
18pts23
chrisbarber.co 7mo ago

I asked AI researchers and economists about SWE career strategies given AI

cjbarber
5pts2
www.erdosproblems.com 7mo ago

Aristotle from Harmonic has solved this Erdos problem

cjbarber
2pts0
www.youtube.com 7mo ago

Ilya on Dwarkesh [video]

cjbarber
4pts0
www.vals.ai 8mo ago

Vibe Code Bench

cjbarber
5pts1
twitter.com 8mo ago

Software 2.0 easily automates what you can verify

cjbarber
1pts0
asteriskmag.substack.com 8mo ago

Common Ground Between AI 2027 and AI as Normal Technology

cjbarber
1pts0
www.youtube.com 8mo ago

Satya Nadella – How Microsoft is preparing for AGI [video]

cjbarber
4pts0
www.writingruxandrabio.com 8mo ago

What will it take for AI to change drug discovery?

cjbarber
2pts0
medium.com 8mo ago

How to give (and get) writing feedback

cjbarber
3pts1
tecunningham.github.io 8mo ago

Forecasts of AI and Economic Growth

cjbarber
1pts0
epoch.ai 8mo ago

What does OSWorld tell us about AI's ability to use computers?

cjbarber
2pts0
epoch.ai 8mo ago

Open database of large AI data centers, using satellite and permit data

cjbarber
1pts1
secondthoughts.ai 8mo ago

A Project Is Not a Bundle of Tasks

cjbarber
2pts0
Precursor 9 days ago

Bright Data ranks #1 on Foil's leaderboard [1], but is still detected. ScrapFly is #4 and ZenRows #7. And I guess Capsolver isn't really a scraping thing but is more just for the captcha component.

I think it'd be good if there were more products that did a better job of making an actually undetectable agent, but doesn't seem like any exist yet.

[1] https://usefoil.com/research/stealth-browser-leaderboard

Precursor 9 days ago

This (agent detection) is now a kind of emerging space. Obviously it'll get much more important, too.

Other products in the space:

- Foil (https://usefoil.com/), I'm biased, a friend is building this

- Kasada https://www.kasada.io/

- DataDome (https://datadome.co/)

- Castle (https://castle.io/)

- Fingerprint (https://fingerprint.com/)

- HUMAN (http://humansecurity.com/)

- Google Cloud Fraud Defense, which is basically the updated reCaptcha (https://cloud.google.com/security/products/fraud-defense?hl=...)

- this, Cloudflare Precursor

It seems like some of the main reasons people care so far are:

- Preventing automated credential stuffing

- Preventing bots from creating a bunch of fake accounts (eg free trial abuse, which can also lead to high twilio SMS bills!)

- Reducing payment fraud

- Blocking LLM scraping

- Blocking automated scalpers (!) eg for tickers or sneakers

I'm curious to see which use cases end up dominating as the reason companies care about this. And I'm hopeful that my agents will still have good ways for me to browse and do things on the web on my behalf - eg detect agents and route them to an agent path, rather than blocking them.

(I'm interested in tools for detecting AI agents and seeing how this shifts as bot traffic goes way up.)

My view is different. Agent products have access to tools and to write and run code. This makes them much more useful than raw models.

The history of both knowledge work and software engineering seems to be increasing in both volume and complexity, feels reasonable to me to bet on both of those trendlines increasing?

Maybe but the product category is not necessarily a monolith in the same way that Claude Code is. These general purpose tools will have to action across a heterogeneous set of enterprise systems/tools.

What would make it not be a monolith? To me it seems like there'll be a big advantage (e.g. in distribution, user understanding) for most people to be using the same product / similar interface. And then the agent and the developer of that interface figure out all the integrations under that, invisible to the user.

For all the benefits that agents offer, they can be asymmetrically harmful. This is not a solved issue.

Strongly agreed.

I saw a few people running these things with looser permissions than I do. e.g. one non-technical friend using claude cli, no sandbox, so I set them up with a sandbox etc.

And the people who were using Cowork already were mostly blind approving all requests without reading what it was asking.

The more powerful, the more dangerous, and vice versa.

There are businesses that want bespoke AI tools and don't have the discipline to deploy them in-house. I don't know if it is ever possible for OAI & friends to develop a "hyper" agent that can produce good outcomes here automatically. There are often people problems that make connecting the data sources tricky. Having a human consultant come in and make a case for why they need access to everything is probably more persuasive and likely to succeed.

Sort of agreed, though I wonder if ai-deployed software eats most use cases, and human consultants for integration/deployment are more for the more niche or hard to reach ones.

Non-technical users expect a CEO's secretary from TV/movies: you do a vague request, the secretary does everything for you. LLMs cannot give you that by their own nature.

What are you using today? In my experience LLMs are already pretty good at this.

Please for the love of god actually go outside and talk to people outside of the tech bubble.

In the past week I've taught a few non-technical friends, who are well outside the tech bubble, don't live in the SF Bay Area, etc, how to use Cowork. I did this for fun and for curiosity. One takeaway is that people at startups working on these products would benefit from spending more time sitting with and onboarding users - they're very powerful and helpful once people get up and running, but people struggle to get up and running.

People don't want "personalized interfaces that change every second based on the whims of an unknowable black box". They have plenty of that already.

I obviously agree with this, I think where our view differs is I expect that models will be able to get good at making custom interfaces, and then help the user personalize it to their tasks. I agree that users don't want something that changes all the time. But they do want something that fits them and fits their task. Artifacts on Claude and Canvas on ChatGPT are early versions of this.

Yes, and the same thing will happen in non-coding knowledge work too. Making knowledge work cheaper will cause complexity to increase, more knowledge work.

My current expectation is that the Cowork/Codex set of "professional agents" for non-technical users will be one of the most important and fastest growing product categories of all time, so far.

i.e. agents for knowledge workers who are not software engineers

A few thoughts and questions:

1. I expect that this set of products will be extremely disruptive to many software businesses. It's like when a new VP joins a company, they often rip and replace some of the software vendors with their personal favorites. Well, most software was designed for human users. Now, peoples' agents will use software for them. Agents have different needs for software than humans do. Some they'll need more of, much they'll no longer need at all. What will this result in? It feels like a much swifter and more significant version of Google taking excerpts/summaries from webpages and putting it at the top of search results and taking away visits and ad revenue from sites.

2. I've tried dozens of products in this space. For most, onboarding is confusing, then the user gets dropped into a blank space, usage limits are uncompetitive compared to the subsidized tokens offered by OpenAI/Anthropic, etc. It's a tough space to compete in, but also clearly going to be a massive market. I'm expecting big investment from Microsoft, Google etc in this segment.

3. How will startups in this space compete against labs who can train models to fit their products?

4. Eventually will the UI/interface be generated/personalized for the user, by the model? Presumably. Harnesses get eaten by model-generated harnesses?

A few more thoughts collected here: https://chrisbarber.co/professional-agents/

Products I've tried: ai browsers like dia, comet, claude for chrome, atlas, and dex; claw products like openclaw, kimi claw, klaus, viktor, duet, atris; automation things like tasklet and lindy; code agents like devin, claude code, cursor, codex; desktop automation tools like vercept, nox, liminary, logical, and raycast; and email products like shortwave, cora and jace. And of course, Claude Cowork, Codex cli and app, and Claude Code cli and app.

Edit: Notes on trying the new Codex update

1. The permissions workflow is very slick

2. Background browser testing is nice and the shadow cursor is an interesting UI element. It did do some things in the foreground for me / take control of focus, a few times, though.

3. It would be nice if the apps had quick ways to demo their new features. My workflow was to ask an LLM to read the update page and ask it what new things I could test, and then to take those things and ask Codex to demo them to me, but it doesn't quite understand it's own new features well enough to invoke them (without quite a bit of steering)

4. I cannot get it to show me the in app browser

5. Generating image mockups of websites and then building them is nice

I think this is smart and very interesting. I see it like an aggregator marketplace. A powerful position to be in.

Cloudflare, GitHub (if they shipped more), Anthropic and OpenAI are also in decent positions to do this.

I wrote notes on this previously [1]. If you believe agents are going to be big consumers, it's helpful to make things that today allow users of agents to easily discover and purchase services via apis.

[1] https://x.com/chrisbarber/status/2026331038994321898

I've tried a few computer use and browser use tools and they feel relatively tok/s bottlenecked.

And in some sense, all of my claude code usage feels tok/s bottlenecked. There's never really a time where I'm glad to wait for the tokens, I'd always prefer faster.

It could be interesting to do the metric of intelligence per second.

ie intelligence per token, and then tokens per second

My current feel is that if Sonnet 4.6 was 5x faster than Opus 4.6, I'd be primarily using Sonnet 4.6. But that wasn't true for me with prior model generations, in those generations the Sonnet class models didn't feel good enough compared to the Opus class models. And it might shift again when I'm doing things that feel more intelligence bottlenecked.

But fast responses have an advantage of their own, they give you faster iteration. Kind of like how I used to like OpenAI Deep Research, but then switched to o3-thinking with web search enabled after that came out because it was 80% of the thoroughness with 20% of the time, which tended to be better overall.

Claude Sonnet 4.6 5 months ago

I wonder if it's actually from CC harness updates that make it much more inclined to use subagents, rather than from the model update.

From the author on X (https://x.com/g_leech_/status/2023384135201349633), below is all me quoting the tweet thread:

New paper on a long-shot I've been obsessed with for a year:

How much are AI reasoning gains confounded by expanding the training corpus 10000x? How much LLM performance is down to "local" generalisation (pattern-matching to hard-to-detect semantically equivalent training data)?

tl;dr

- OLMo 3 training corpus contains exact duplicates of 50% of the ZebraLogic test set.

- We embed the corpus to find semantic duplicates of test data in the wild. 78% of the CodeForces test set had >=1 semantic duplicate

- The semantic duplicate rate is maybe >4 in 10000

* at least 50% and at least 78% that is

arxiv.org/pdf/2602.12413

Imagine you're head of training at at OpenAI, and you want your benchmark scores to be meaningful (: to estimate OOD performance)

You have a hard task ahead of you! Your models have seen so much, memorisation is so easy - as is local generalisation (noisy pattern-matching).

What can you do? Well, obviously you take every benchmark you're going to test on and try to "decontaminate" your training corpus (remove test data from the training data).

By default this is just one level above string matching ("n-gram matching" - if sentences overlap in (say) a 13-token window, remove them from the training corpus).

But you're actually trying, so you also translate the test sets and delete translations of test from train.

But! every piece of test data has an arbitrary number of logical equivalents and neighbours (like how `x + y = 10` is the same problem as `2x + 2y = 20`). And LLMs are amazing at semantic search, so maybe this inflates benchmark scores.

The cutting-edge tech for detecting these "semantic" duplicates is... an LLM. But you simply can't do 100T x 1M calls. There's not enough compute in the world (yet).

So you do what you can - maybe you

- categorise the entire corpus & do intense search inside relevant partitions (e.g. maths > number theory > ...)

- embed the whole corpus & look for things really close to test data

- train a wee 300M filter model & do what you can with that

How much does this process catch? How many semantic duplicates of test data slip through? And what's the impact on final benchmark scores?

We don't know, This (finally) is where our paper comes in:

We experiment on OLMo 3, one of the only really good models with open training data. Since we have its entire training corpus, we can exhaustively check for real "natural" duplicates and finetune it to estimate their impact. We embed the entire Dolma Instruct corpus.

Firstly: we were surprised by how ineffective n-gram decontamination was at catching exact duplicates - 70% of harder tasks had a match. But the spurious performance gain wasn't so large, at most +4pp.

Secondly, every single MBPP test example and 78% of CodeForces have semantic duplicates

Thirdly we generated 10k synthetic duplicates for MuSR, Zebralogic, and MBPP problems and finetuned on them.

- MuSR +22pp. Semantic duplicates as strong as exact

- ZebraLogic +12pp. Exact much stronger

- MBPP +17pp. Exact stronger

Fourthly we guess that 4 in 10,000 training datapoints are a strong semantic duplicate for a given benchmark datapoint (where strong means just "obvious to Gemini")

So: n-gram decontamination is not enough even for the easy (exact) stuff, semantic duplicates are at least a moderately big deal, and this probably transfers to frontier models to some degree. The above are probably underestimates too (since our detection pipeline was cheapo).

Data contamination is a huge field. Here's how we're new

This is preliminary work on a shoestring - we didn't get at the big questions yet ("what share of benchmark gains come from interpolation over a hidden training corpus?", "does this even matter?")

And local generalisation across very different strings is anyway pretty miraculous

The grand aim of this research programme is to decompose benchmark gains / apparent AI progress into 4 estimates:

1. benchmaxxing (memorising exact duplicates)

2. usemaxxing (RLing narrow capabilities)

3. hidden interpolation / local generalisation

4. OOD generalisation

We have a lot of ideas! If you're interested in funding this, grab me at gavin@arbresearch.com

Nearly all of the real work done by Ari Spiesberger, Juan_VaGu, Nicky Pochinkov, Tomas Gavenciak, peligrietzer and NandiSchoots

And ofc this work wouldn't be possible without allen_ai and natolambert working in public and enabling actually scientific evals.

It'll be nice when there's smarter routing between models, or easier routing, so some things get sent to the fast model, some get sent to the cheap model, some get sent to the smart model, etc.