HN user

petesergeant

5,239 karma

vivid.art0944@fastmail.com

Posts31
Comments2,672
View on HN
github.com 1d ago

Show HN: Byre; a free/free agent sandbox with a focus on comfort

petesergeant
3pts0
sgnt.ai 1d ago

LLM spambots liked my Show HN post more than real people did

petesergeant
13pts3
pleasedonotescape.com 2d ago

Show HN: A comprehensive, filterable list of AI agent jails

petesergeant
8pts3
sgnt.ai 20d ago

The prompt is not the work; describing AI contributions

petesergeant
1pts0
news.ycombinator.com 2mo ago

Ask HN: What are you doing during inference?

petesergeant
8pts3
github.com 6mo ago

Claude Code superpowers: core skills library

petesergeant
4pts0
www.wsj.com 11mo ago

Amazon to Pay New York Times at Least $20M a Year in AI Deal

petesergeant
3pts0
www.technologyreview.com 11mo ago

Quidnet wants to use the Earth as a giant battery

petesergeant
3pts1
sgnt.ai 1y ago

RAG chunking isn't one problem, it's three

petesergeant
4pts0
bsky.app 1y ago

Quitting programming 'cuz of LLMs is like quitting carpentry 'cuz of table saws

petesergeant
3pts0
sgnt.ai 1y ago

Understanding Modern AI Is Understanding Embeddings: A Guide with Lots of Dogs

petesergeant
1pts0
sgnt.ai 1y ago

Understanding modern AI is understanding embeddings: a guide with almost no math

petesergeant
3pts0
sgnt.ai 1y ago

When Users Won't Wait: Engineering Killable LLM Responses

petesergeant
4pts0
sgnt.ai 1y ago

In-memory free-text search is a super-power for LLMs

petesergeant
3pts0
sgnt.ai 1y ago

Don’t let an LLM make decisions or execute business logic

petesergeant
325pts169
sgnt.ai 1y ago

Four bad definitions of "Agentic AI"

petesergeant
2pts0
www.sgnt.ai 1y ago

Street-fighting RAG: chain-of-thought prompting

petesergeant
1pts0
clarotesting.wordpress.com 1y ago

Y2K – why I know it was a real problem (2015)

petesergeant
50pts28
www.bbc.com 2y ago

Julian Assange to plead guilty in deal with US, go free for time served

petesergeant
22pts3
www.macrumors.com 2y ago

Meta CEO Mark Zuckerberg Says Quest 3 Is Better Than Apple Vision Pro

petesergeant
3pts0
www.washingtonpost.com 2y ago

China's Xi, in need of a win, appears ready to engage with Biden

petesergeant
1pts0
www.washingtonpost.com 2y ago

FDA approves Mounjaro for weight loss

petesergeant
2pts0
www.washingtonpost.com 2y ago

Opinion: The Chinese economy is doing better than you might think

petesergeant
1pts2
news.ycombinator.com 3y ago

Ask HN: Does anyone have a contact at Y Combinator who can help with research?

petesergeant
7pts0
pastebin.com 3y ago

ChatGPT got 6 of my 50 Xmas Trivia questions spectularly wrong

petesergeant
8pts12
www.washingtonpost.com 3y ago

Animated Guide to the Offside Rule

petesergeant
1pts0
www.livescience.com 3y ago

Dolphins have names they choose for themselves in infancy (2006)

petesergeant
6pts0
mojojs.org 3y ago

Mojo.js is a port of Perl's Mojolicious to TypeScript

petesergeant
73pts22
blog.jooq.org 3y ago

Creating the tables dee and dum in Postgres (2017)

petesergeant
1pts0
news.ycombinator.com 4y ago

Ask HN: Digital nomad developers, how much are you charging?

petesergeant
19pts13

and is now doing rather well

I was going to write a shitty reply to this, but the more research I did, the more it seems actually this went pretty well for them. UK and NL governments that lost deposits in the Icelandic banks mostly got their cash back (eventually), there was recession and unemployment but not that much worse than other countries, and the country is in good standing again with the markets.

I would note that this is much much easier to do if your investors are foreign, rather than domestic pension funds, so it doesn’t bring down the government, though.

LLMs are inherently adversarial (read, "relentlessly proactive") [and] you still tend to end up with something you care about on the same side of the trust boundary as the LLM

I mean this is a problem with many coworkers too, so you deal with it in the same way: limit what they can do to creating pull requests.

This is not how it’s going right now, and FWIW I also don’t think it’s going to go that way, but it’s certainly a plausible investment thesis.

The valuation is based on one lab getting a decisive first advantage, and turning that into a durable self-improving advantage that can never be caught up to. If any can pull it off (a gigantic if), they will effectively own most AI value, and the people who own their shares will live happily ever after. Divide your investment between the labs that could plausibly do this, and your EV may not be dreadful.

Habibi, I don't think anyone is really blaming you personally. We are piling on to the fact that this piece of critical infrastructure many of us depend on day in day out is being built in such a way that a single developer can wake up one morning, ship a change they thought might be nice, and then walk it back the next day.

Yah. And it's not like they can't afford the talent to do this right either. I've said elsewhere, I think it's an attribution error. Claude Code is massively popular, but arguably because of the model/subscription, but I think the brass reads this as "great success throughout"

https://github.com/pjlsergeant/byre can add small but very convenient layer around whichever agent you want, initially trapping it in a specific folder, but you can easily mount more, firewall it, bring in MCPs and skills, or whatever. It's unambiguously AI-assisted, but a very very great deal of thought and work has gone in to the design of it. Ask your favourite agent to review the code if you're skeptical.

closing at $135.27

ffs, wake me up when it's at least 10% below what it IPO'd for. The idiotic tulip mania that followed in the few days after it floated was noise, but as of today, it seems the IPO price was pretty much right. However, endless headlines about the price crashing etc.

From a fundamentals perspective, it's an insane price, obviously. But the narrative that it's all coming crashing down is obviously not correct (today).

almost definitely requires an LLM to tackle at all

Conveniently I have some of those… first day of trying to script Grok Build I think I sent in 6 bugs of slightly weird behaviour I discovered, it will be much more useful to (have an agent) check the source and see if stuff looks deliberate or like a bug, etc

Genuinely curious about whether comments like this consider all AI generated codebases to be slop? Are you just knee-jerking or is this one an example of actual trash? I have been building a product[0] where I’ve not written a single line of code; is it also definitely “just tokenmaxxed slop” or is any consideration going into comments like this?

0: https://github.com/pjlsergeant/byre

I have found Grok Build to be decent, and the harness to be competitive with similar harnesses. What will Cursor add if I check it out?

Neat, trying to reverse engineer some specifics of how it does stuff has been a pain in the ass, and this will make it easier.

Classification of types of user frustrations and sentiment analysis, content trends, engagement and gap analysis, as well as then looking at changes from the previous week. We also look at how certain queries turn into actions in the system (eg: which users take actions we offer them). We run it once a week, rather than every day, and it provides an exec-facing overview, as well as areas for support to dig further in to. While it's some good work, as far as I'm aware it's almost all just a text prompt and a connection into Langfuse.

In my experience they're pretty well aware of this failure mode, to the point where I learned the word accretive from adversarial agent code reviews this week. Less good at fixing it, but once you know it'll happen you can ask it not to.