HN user

david_shi

284 karma

https://operator.io

you can probably guess my email

Posts19
Comments171
View on HN
www.daytona.io 26d ago

Daytona is going closed source. Here's why

david_shi
4pts1
news.ycombinator.com 1mo ago

Ask HN: How do you make AI writing usable?

david_shi
4pts7
www.youtube.com 1mo ago

Lo and Behold, Reveries of the Connected World (Werner Herzog) [video]

david_shi
3pts0
news.ycombinator.com 1mo ago

Ask HN: Why did you open source your project?

david_shi
4pts5
news.ycombinator.com 1mo ago

Ask HN: What have you built with Claude Managed Agents?

david_shi
2pts0
news.ycombinator.com 1mo ago

Ask HN: How do you solve AI's confused deputy problem?

david_shi
2pts1
einsteinarena.com 4mo ago

EinsteinArena: AI agents collaborate and compete on unsolved science problems

david_shi
2pts0
news.ycombinator.com 4mo ago

Ask HN: Have you built a public facing agent skill?

david_shi
1pts2
operator.io 4mo ago

China Restricts OpenClaw as Security Fears Grow

david_shi
4pts0
www.operator.io 4mo ago

How to Build a Personal Intelligence Agency

david_shi
1pts0
www.operator.io 5mo ago

Show HN: OpenClaw Swarm as a Service (YC W20)

david_shi
1pts1
arxiv.org 1y ago

Horus: A Protocol for Trustless Delegation Under Uncertainty

david_shi
3pts0
arxiv.org 1y ago

Memes, Markets, and Machines

david_shi
3pts1
www.bloomberg.com 2y ago

Bangladesh Internet Goes Dark as Widening Job Protests Kill 25

david_shi
4pts1
davidshi.xyz 2y ago

Identity Augmented Generation (IAG)

david_shi
1pts0
davidshi.xyz 2y ago

Humane vs. Rabbit: Tools and Toys

david_shi
1pts0
davidshi.xyz 2y ago

Real World NPCs

david_shi
3pts0
www.ppak.net 2y ago

3D Printed Custom Keycaps: From Design to Print

david_shi
1pts0
agree.substack.com 3y ago

Why Russian Companies Sought Mafia Protection in the 90s

david_shi
4pts0
Writers and Drugs 28 days ago

Interesting that there's not a single mention of cannabis, perhaps it's more of a musician's choice.

I believe eval startups can work when they're targeting safety benchmarks specifically.

Are there any examples of successful startups doing this?

Even with the examples, I've found that explicitly pointing out what not to do is moderately helpful if the model is given some time to self-evaluate. I wish this was something that came out of the box though.

Sakana Fugu 1 month ago

This is a charitable read, but I think that being able to pick from a panoply of models will actually yield much better results in the long run.

The same model that has been post-trained to operate for hours as a Linux admin will be incapable of writing a heartfelt email, but with something like Fugu, you'd get both the Linux admin for driving the browser harness and the smaller writing specialist model for drafting the email itself.

GLM 5.2 vs. Opus 1 month ago

GLM-5.2 cost a fraction as much. Opus finished in half the time and shipped a cleaner game.

Off topic, but does anyone else instantly pick up on LLMisms like this? It seems like all the models have converged on this style of writing, and improvements aren't really changing it.

Sakana Fugu 1 month ago

Their research around building a domain specific model is pretty cool, it's kind of like Karpathy's autoresearch but pointed at deciding the optimal model to use at each step of the inference.

If cost becomes an even bigger problem being able to choose "best performance possible" or "strong but cost effective" will be useful.

https://arxiv.org/pdf/2512.04695

The economics of working at a pre-IPO company that will likely have a successful IPO and a 20+ year post-IPO company are also very different.

I actually built my own daily driver at https://operator.io.

There's definite tradeoffs when it comes using a remote agent service vs. setting up OpenClaw or Hermes on a Mac Mini, but being able to access an agent with a completely isolated file system and network gives me peace of mind when I'm using it.

I've changed my mind a few times on this, but given how substantial the adoption for MCP has been (Claude and OpenAI both use it for their native integrations) its only a matter of time before consolidation happens.

There's a way higher incentive to build an MCP server than an A2A one, and unless Google makes their default AI search a native A2A client it doesn't feel like it will get the momentum it needs to take off.

I think a lot of what people call failures of discipline actually comes from not having the tools they need. Someone who lives 15 miles away from the nearest gym will have a tougher time than someone who's gym is next door.

On the tracking point: I’ve found that a coding agent that can modify a file system (create and update CSVs) that’s accessible on both my laptop and phone to be the single best way to track things I’ve ever used. Bar none.

Even apps with the best UX, like Strong for tracking workouts, feel exponentially clunkier than having an agent that can answer questions, analyze pictures, and write things down on a persistent file in real-time.

Midjourney Medical 1 month ago

I've heard this argument before and it's always seemed downstream of capacity constraints and the current incentives of the healthcare industry.

There's a reason why billionaires like David Rockefeller, Larry Ellison, and Rupert Murdoch are able to live much longer lives than average, and having an oncall health team (that I'm sure does frequent testing and monitoring) is a big contributor to that.

More testing and data collection doesn't mean that every single anomaly would need to be investigated or communicated with the patient, but would provide a better longitudinal view that can help with disease prevention and health optimization.

Have you found any alignment research with clear a/b tests?

An experiment that I found interesting was asking Claude for 10 ways to legally bankrupt Anthropic vs. Philip Morris.

In the Anthropic answer, it gave reasons like employees losing their jobs being bad for why it couldn't do it, but jumped straight into tactics with Philip Morris. Not sure if it's moral taste or self-preservation, but felt eerie nonetheless.