HN user

trjordan

4,868 karma

Working on https://tern.sh

tr at tern dot sh

Posts37
Comments564
View on HN
tern.sh 13d ago

A compiler that never says No

trjordan
2pts0
tern.sh 4mo ago

Volume, Ambition, Clarity

trjordan
2pts0
tern.sh 1y ago

You Have to Decide

trjordan
3pts0
blog.turbinelabs.io 8y ago

Coworkers Will Become Customers – It’s a Good Thing

trjordan
1pts0
www.saastr.com 9y ago

Reasons I Won’t Fund You

trjordan
93pts73
news.ycombinator.com 10y ago

Ask HN: What's the state of Python web frameworks in 2015?

trjordan
24pts19
mydailyjava.blogspot.com 11y ago

Make agents, not frameworks

trjordan
2pts0
www.theguardian.com 11y ago

Delivering Continuous Delivery, Continuously

trjordan
1pts0
computationallyendowed.com 11y ago

MVC in a Reactive World

trjordan
14pts0
mumbledrantings.blogspot.com 12y ago

Marketing Must Build a Product

trjordan
1pts0
www.appneta.com 12y ago

How to Easily Capture TCP Conversation Streams

trjordan
6pts0
www.appneta.com 12y ago

Automated Testing of Hardware Appliances with Docker

trjordan
18pts0
www.appneta.com 12y ago

Profiling Python Using lineprof, statprof, and cProfile

trjordan
4pts0
www.appneta.com 12y ago

How to Save 90% on your S3 Bill

trjordan
368pts98
www.appneta.com 12y ago

The Right Stuff: Breaking the PageSpeed Barrier with Bootstrap

trjordan
118pts33
www.appneta.com 12y ago

The FTC is more responsive than NASA

trjordan
9pts3
www.appneta.com 13y ago

How to make a telepresence robot with node.js and a Raspberry PI

trjordan
6pts0
techcrunch.com 13y ago

CA Sues New Relic for Patent Infringement, Seeking Injunction

trjordan
4pts0
www.tracelytics.com 14y ago

Page Guide: Providing contextual help for feature-packed webapps

trjordan
4pts0
www.tracelytics.com 14y ago

Performing under pressure, pt. 1: Load-testing with multi-mechanize

trjordan
6pts0
betaspring.com 14y ago

Startups Find Providence 31.42% More Wicked Funner

trjordan
17pts9
www.tracelytics.com 14y ago

Tracelytics: The Structure of your Success [Analyzing for Performance]

trjordan
4pts0
steve-yegge.blogspot.com 14y ago

Emacs Tips and Tricks from Steve Yegge [2006]

trjordan
1pts0
techcrunch.com 14y ago

Web App Performance Solution Tracelytics Raises $600K

trjordan
15pts1
doodler.melodylu.com 14y ago

A creative application for Google Doodler

trjordan
1pts0
mumbledrantings.blogspot.com 15y ago

5 Minute Intro to Cassandra

trjordan
2pts0
www.wired.com 15y ago

Proving It: Wind-Powered Cart Goes Faster Than the Wind

trjordan
1pts0
www.lala.com 16y ago

Apple shuts down cloud-music site Lala.com

trjordan
1pts0
googleenterprise.blogspot.com 16y ago

Collaborative mapping for major disasters

trjordan
1pts0
mumbledrantings.blogspot.com 16y ago

Javascript's dynamic scoping

trjordan
1pts0

Claude wasn’t just a compiler here. I never handed off a task and let an agent make a bunch of decisions in order to reduce it to practice.

I’d say that, in all the ways that matter, I understand the code.

I think the dissonance here is really important, and not a bad thing at all. A lot of the decision _were_ handed off the the AI, but they weren't the decisions the author cared about. This is a big selling point of AI! If something is doable with a computer, it’ll figure it out. 30 minutes and 200m tokens later, it’ll take any idea and declare “the feature is fully implemented.”

The hard part is figuring out where to inject that friction, so you can see where it's making decisions for you that matter. The author approached this by incrementally building the thing, reviewing and poking and prodding at every step. A week of attention following a bunch of design discussions is fast, but that's still not trivially cheap.

I want to see us talk more about the decision exhaust of agents, because the better the models get, the more decisions we'll want them to make.

I wrote a bit more here: https://tern.sh/blog/compiler-never-says-no/

Pushinka 6 days ago

Breed: mixed

That's a corgi.

Pushinka subsequently became irascible, and "a little nippy" according to Caroline Kennedy, which she attributed to her upbringing in a scientific laboratory.

No it's because she's a corgi.

So, we tried feeding the logs back to the LLM, and it mostly produced slop. Lots of decisions nobody cared about. The biggest things that moved the needle were:

- Baseline it. We mine previous logs, github comments, etc. for "what you care about." That helps pull out decisions that you actually care to read.

- Anchor to code. "The code enshrines this decision" is more interesting than "the agent self-talked this." Agents don't always self-talk decisions, and the thing that ultimately matters is the behavior in code.

to your edit (and all totally fair):

- Yes, closed source and signup required. A lot of what we're driving towards is easy team sharing, so we're taking the bath early instead of building an OSS thing and rug-pulling later.

- Code doesn't leave your machine. There's an agent that runs locally. I know "trust me" isn't the strongest stance, but this comes from the multi-player future.

- Honestly, Goose AI is a remnant of a previous product. It's inert and we'll clean it up once we've gotten the last couple folks off the previous iteration.

100% important. But what decisions do you care about seeing?

The whole point of the agent is to make decisions for you. If you want to make every little detailed decision, just write the code.

The whole art of this problem is figuring out which decisions matter to you, and how to surface them.

(Disclosure: we're working on this too. https://tern.sh)

The agent will always fill in the gaps in your understanding. It's not a compiler. It's categorically different from any of the other ways we've built software.

I'm not sure reading code is coming back. The ritual of reading code must come back, because that's the only way to build products that don't collapse under their own incoherence, both technically and visibly.

"just ask Claude" is fine, but it's not the end state

I am no fan of Zuck. But this is his whole deal.

Instagram was a purchase. Facebook wasn't his idea. Threads is a copy. The 1 thing that Zuck understands better than anybody is that engagement is the only thing that matters to social networks, and he's willing to throw the entire company at the problem. He has been for 20 years.

He's good at addiction. He knows how to build an org that's world-class at addiction. It's entirely reasonable that the EU regulate it, and Zuck is exactly the person to point the regulation at.

lmao hi Matt

I agree, though maybe the middle ground is something more like: the constraints of our environments shape us. It's easy to say that big companies are a weird and unique cave that produces weird and unique outcomes, but other companies are somehow constraint-free. Smart, talented founders do weird and constrained things all the time because they don't have capital or customer bases or brands, and those are also constraints that bind just as hard.

"Race for MVP to learn what your bottleneck is" is a handcuff, just like "you can't deploy more than 3x / year because our customer base hates change."

Most startups fail. Most big company projects are kind of worthless. These are two sides of the same coin.

Producing something novel and valuable is HARD. Unbelievably hard. The idea is hard. The building is harder. The scaling and steering and feedback is ego-crushingly hard.

When it's valuable, it's frequently enormously valuable. That funds the experimentation, the incremental expansion, the waste. It's hard to really internalize how valuable localization, admin controls, FedRAMP, and onboarding tweaks are, truly, because they all compound. You can't just have the idea and the MVP, you also have to have all the other stuff, and it's hard to come up with new ideas while you're trying to keep a million users happy.

I vehemently disagree that people working at big companies are stupid, or making themselves stupid. There are VPs and SVPs at Adobe and Salesforce that are smarter, more knowledgable, and more productive than any startup employee. It's just structurally hard to move the needle there, and their successes aren't written about in TechCrunch. They're also paid a million dollars a year, and are unbothered by the lack of external recognition.

I'm off founding a startup now, and it's good for the soul, but I don't delude myself into thinking everybody else is blind.

AI is so miserable for this. It's so focused on doing what you ask, it forgets that there's stuff worth doing that you didn't ask for, like defining reasonable abstractions.

Getting away from stuff like this is exactly why I want to use AI. When I say "implement this for idle but active users," I _want_it to define isUserActiveIdle() and stuff these 4 conditionals in it. Having to check the generated code for stuff like this undoes, like .... all the benefit of using AI.

AI makes all these little decisions for us. I can about some of these decisions. I just want to notice when it's doing this without having to make my eyes bleed reading 10k lines of generated code a day.

98% isn't much 16 days ago

I was heading to dinner with a friend who worked in infra. Google maps said we could bike across town in 20 minutes. He suggested we leave 40 minutes ahead of time and grab a drink at the bar if we got there early. When I raised an eyebrow, he goes:

"What, do you not live your life based on 99th percentiles?"

I tend to think of work as upside-based on downside-based. Most feature work is upside. 10% lift on conversions is great, 40% adoption is winning, and you're playing for the moonshot of 10x. Infra work is downside-based. 98% secure, 98% available, 98% acceptable performance -- that'll all failure. Winning means the thing works as expected and nobody notices.

Not everything sorts cleanly into upside vs. downside, but a lot does. Allocate your risk accordingly.

It's because it mostly doesn't matter what you are trying to get the code to do. What matters is what the code does.

Session logs can absolutely be useful, but not when building further. It's just that that the place they slot in is during validation. You know, that place between the markdown plan and CI passing, where there's 800 new lines of code and it all seems sort of fine when you click around?

Session logs can show you what sort of manual validation happened. CI will run the tests you had, and the code will show you what new unit tests were added, but session logs can show you that the agent drove the app with Playwright, or that the agent read and considered the prod config as well as the dev config.

Nothing bulletproof, but not every piece of validation work merits a test in the repo that lives forever. We've gotten a lot of mileage out of re-analyzing the sessions, figuring out where the agent made decisions without asking, and forcing the agent to consider validation for those decisions. That's the sort of thing that's hard to dictate up front but easy to highlight with the session logs.

100%. The problem with them isn't making sure they're doing the right thing, it's making sure they're not making bad assumptions.

IMHO this is where code review goes until we fix the individualized model thing: you need to review the decisions the agent made, where you didn't steer. Most will be right. A few will be disastrously wrong. But decision-by-decision is a lot less to review than line-by-line of code.

This is RL, right? Like, this is exactly why models have mostly converged around obvious style, because we train them literally on thumbs-up/thumbs-down data of what good behavior and good code looks like.

And that's why it's so hard to get a model to reproduce the specific taste of a person or an organization. My taste is different than yours, so if we dump our aggregate preferences into RL, in averages out to nothing interesting.

For the code-writing case, this means you end up reviewing every line of code, looking for places where you'd thumbs-down the code. Not every line of code contains a real decision, though, so it feels like a waste of time.

You can't unit test for taste if you haven't written down what you mean by taste. If you can externalize it, then you can.

Follow this line of thinking, and the AI-friendly answer is easy: we just have to externalize everything we know, so Claude can implement what I want.

Except that I can't fully externalize myself. Debugging a system takes more resources than running the system. If I could write down everything I know and hand it to a machine, I'd do that, but it impossible.

People aren't books or hashmaps. If you want to build something, you need to use the tools, not teach the tools to use you.

[edit: I'm trying to figure out if there's something to be done about this. Email me if you want to chat -- tr at tern dot sh]

If you didn't take the time to write it, why should I take the time to read it?

This is a band-aid. Maybe even a good band-aid, because it'll keep individual contributors from flooring the zone. But the core problem is Github's model that assumes code is worth reading.

I'm much rather see the agent logs stapled to PRs. Make it easy to understand if there's a brain behind the suggested changes before engaging.

The Coming Loop 30 days ago

I think there's 2 important, but separate, ideas in this post:

- Models are not good at or getting better at creating strong invariants, which his fundamental to good software

- It is unclear how to keep tabs on what the agent is doing, so you, a human, can intervene.

These are related, obviously: one of the highest-leverage things you can do is force you agent to use a strong, minimal set of types or data invariants or other constraints. They get much better when your codebase broadly supports this!

I do suspect they're separable, though.

If you had the right levers and visibility, you should be able to get the model to produce code that doesn't feel like slop. But every time I've had a model try to keep me in the loop, it inundates me with irrelevant decisions and busywork. Its inability to see what's structurally important still shows up, just differently.

[If the models get better at defining and respecting invariants, maybe there's a new flavor of slop, that's less obvious today.]

The worst thing that can happen at an early company is that it sort of works.

I like the deal where I roll the dice and don't have to work again if I win. I'm fine with the deal where I take a barely-passable salary and do something wacky for a year.

The worst deal I can imagine is that the startup slowly grinds its way to profitability over 3 years, can't raise, and grows 15% / year.

Every company I've seen do that never fixes the salary issue. Everybody's still making their seed-stage base, or maybe +25%, which is still a 30% pay cut from the last job they had. Their equity is worth nothing. There's no career advancement, because there's 2 staff jobs and 3 EM jobs and 1 VP job.

There's lots of ink spilled about how founders expect early employees to work hard, perhaps too hard for what they're paid. It goes the other direction, too: early employees should expect founders to succeed, because there's always another startup to join.

The core of the problem is that there are a million tools that make AI better, and no ways to measure whether AI is working better.

Big companies with popular products have it. They do something between normal product analytics and chatbot evals to figure out if users are being successful in their sessions. That's the job.

But any given dev, with between 3 and 50 sessions a day? Like, I have no idea what makes the LLM better. It's all vibes.

My company has a whole stack here. Preferred harnesses, preferred models, skills, the shape of our code, everything. There's gotta be a way to measure whether this setup is working for us, at 1 / 1-million-th the scale of a Claude Code.

Those are not code problems. They are evaluation problems.

Code becomes precious when it is the only place knowledge lives.

Reading AI code all day is _agonizing_. Just, a horrible way to live, and it melts people's brains at the moment you need them to be the most capable.

Manual programming has this really productive and gratifying feedback loop, where you read the code, write the code, and fix it until it compiles/runs/does what you want. AI code not only does half that for you, but it makes the "click" at the end uninspiring because you're never sure if it's cheated a bit to get to that moment.

Trying to operate with AI-generated code as the only durable artifact of programming is a dead end for the industry. Charity points to (and correct discards) architecture diagrams/specs as an interesting space to work in. My suspicion is that it's closer to the thing that's hand-written: prompts, markdown plans, and other nudges. Focus on the thing that you, as a human, produce, and that's the basis for both the core loop of "did the AI follow my instructions" and it's higher-leverage when you go to code review.

By the time you get to the PR, you've probably typed enough to Claude that you can regenerate the code, but the current industry default is to just throw away all those sessions and ship the code. That's backwards!

It used to be hard 1 month ago

There's a neat / weird ladder that I keep seeing friends go through as they work through this.

- Volume. Kill the backlog! 8 agents in terminals, frantically!

- Ambition. Do the things you always want to do! You have the power!

- Clarity. Oh god I have to figure out what to do next.

That last one is honestly super-hard, but it's also the most valuable. Like, do you want to wake up every day and find new work, because you understand the machine better than everybody else? I know a bunch of people that love that stuff, but also a bunch that don't. I totally get that the transition is hard.

https://tern.sh/blog/volume-ambition-clarity/

Totally. Every "we're losing our craft" article has the same gloomy shape. That's enough of a bummer, but they also argue against themselves halfway through.

This one, for instance:

But exactly which details are deemed “unimportant” is a very consequential and sometimes subjective decision. And eventually, the details always leak through.

Right, so you're saying this new technology will still reward deep technical understanding, because there's no way around it. I agree. Why is the whole tone of this thing "AI is making my craft a cheap commodity?"

Websites are largely better, technically, than they were 10 years ago. They're more full-featured, they're faster, SSL/a11y/responsiveness are stronger defaults. Content mills / SEO / news sites are a separate, terrible failure mode of ads and corporate incentives. That's not React's fault!

It feels like we're far past the point of where having AI do more faster is helpful.

It's telling that they used "rewrite Bun in Rust" as the proof point here. It's cool! But the vast majority of software engineering doesn't start with tens of thousands of tests, where making them pass is the whole job.

In my experience, AI still drifts from what I meant it to do on anything bigger than building a widget. My time is spent suspiciously reviewing output for changes the agent snuck in, or invariants it broke. I talked with a friend recently where the agent broke the test harness badly enough that none of the tests mattered for 3 weeks. They did pass, though, so CI never complained.

There's something at the intersection of context engineering, managing that sloppy pile of markdown plans, and good old fashioning system understanding that's the real bottleneck.

They've got, ballpark, $5t to $10t to make back in the next 5 years, or the hardware buildouts will start getting written down.

This means we're going to need $1t+ per year in spending, per year, on tokens. 200m knowledge workers in the world, 30m developers. We're talking about a world where you need 5% of every knowledge workers salary to go into tokens. 20% if you're a developer.

That's a _huge_ shift. Most people I know cite +20%-40% velocity with these tools, against the actual work their company cares about doing. +20% speed for +20% spend isn't going to motivate a trillion dollars a year in spending.

We're not there yet. This is still the upswing of the hype cycle, and unless we figure out how to make developers 2x, 5x, 10x as productive on stuff that matters, this isn't going to play out well.

This is ... not new at all?

App-bundling apps existed. Apple rejected them.

Low-code apps existed. Terminals existed. Apple rejected them.

LLM apps exist. Apple allows them, because they render text, pictures, and video, but they don't run arbitrary code.

Running arbitrary code is flatly forbidden, because users can't reason about them. I see absolutely no evidence that software is moving away from versions, any more than it was when apps could first search the internet, render recommendations, or deliver messages.

Yes, but: writing code always teaches you something.

I've worked at founder-sized startups and $xxb dollar public companies. I've never read a product spec, a pitch deck, or a PRD that describes a solution that, if implemented in the way described, would solve the problem. Building the thing teaches you how it should behave.

Software is a complex, interactive medium. Iterating in the code, with people who understand the problem and care to see it solved, is the only way I've seen valuable products get created. Meetings and diagrams help, but it's not until you write some working software that you know whether you have something.