HN user

CuriouslyC

7,955 karma

Chief scientist/ceo, Sibylline Software. nathan @ sibylline.dev

Posts49
Comments4,425
View on HN
sibylline.dev 4mo ago

I Changed My Mind About MCP

CuriouslyC
3pts0
sibylline.dev 4mo ago

You Don't Need to Detect Prompt Injection to Stop It

CuriouslyC
2pts0
github.com 4mo ago

Smith: The Secure Open Source Multi-User AI Assistant Framework

CuriouslyC
4pts2
sibylline.dev 4mo ago

Defeating Prompt Injection with Protocol Firewalls

CuriouslyC
2pts0
sibylline.dev 5mo ago

The Pillars of Agent Security

CuriouslyC
2pts0
sibylline.dev 5mo ago

The Pillars of Agentic Security

CuriouslyC
1pts0
github.com 5mo ago

Clean: High Performance Prompt Injection Detection and Mitigation

CuriouslyC
1pts1
z.ai 5mo ago

GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

CuriouslyC
484pts520
github.com 5mo ago

Scurl: Agent First Curl Wrapper with Markdown Extraction and Secret Blocking

CuriouslyC
1pts1
sibylline.dev 5mo ago

Pi Is the Linux of Agent Harnesses

CuriouslyC
2pts0
sibylline.dev 5mo ago

Averting the Code Quality Apocalypse

CuriouslyC
3pts0
sibylline.dev 5mo ago

Stop screwing around with agent orchestration, your bottleneck is validation

CuriouslyC
2pts1
sibylline.dev 5mo ago

The Problems with Spec Driven Development

CuriouslyC
1pts0
sibylline.dev 5mo ago

Agents with Scribe Solve Hard Problems 57% More Often

CuriouslyC
1pts0
sibylline.dev 5mo ago

Stop screwing around with agent orchestration, your bottleneck is validation

CuriouslyC
2pts1
sibylline.dev 5mo ago

Scribe Increases Agent Success Rate on Hard Problems by 57%

CuriouslyC
1pts0
sibylline.dev 6mo ago

Scribe reduces agent token usage by 30% with no loss of accuracy

CuriouslyC
2pts0
sibylline.dev 6mo ago

Scribe reduces SWE-bench token usage by 30% with no loss of accuracy

CuriouslyC
1pts0
codehealth.sibylline.dev 6mo ago

Is your codebase holding back your AI tools?

CuriouslyC
1pts1
sibylline.dev 6mo ago

Digital Alchemy: Turning Slop into Gold with Ralph and Valknut

CuriouslyC
1pts0
codehealth.sibylline.dev 6mo ago

The Code Quality Leaderboard

CuriouslyC
1pts0
sibylline.dev 6mo ago

Model Evals Can Fix Education

CuriouslyC
2pts0
sibylline.dev 7mo ago

Valknut: Code Intelligence for Agents

CuriouslyC
2pts0
sibylline.dev 7mo ago

Valknut: Advanced Code Intelligence

CuriouslyC
1pts0
github.com 7mo ago

Show HN: Valknut – static analysis to tame agent tech debt

CuriouslyC
2pts0
sibylline.dev 8mo ago

Calligraphers and Storytellers

CuriouslyC
2pts0
www.ft.com 9mo ago

OpenAI Should Make a Phone

CuriouslyC
3pts2
sibylline.dev 9mo ago

Claude Skills Considered Harmful

CuriouslyC
4pts0
sibylline.dev 9mo ago

AI Is Too Big to Fail

CuriouslyC
8pts3
www.cl.cam.ac.uk 9mo ago

The Space and Motion of Communicating Agents [pdf]

CuriouslyC
3pts0

Random note, the biggest single thing you can do to improve your aim is to crank down your sensitivity and learn to strafe with the target rather than try to track with the mouse. The second biggest thing you can do is stop trying to exactly track them, and instead try to predict their strafe patterns based on their surroundings so you can place shots where they're likely to be, at least for any weapon with a travel time.

For code you can do better than humans on a lot of metrics pretty easily bro. This is because you can RL on objective verified outputs. Kind of like how RL was used to make agents wipe the floor with humans in Dota2, chess and go.

RLing taste and discernment are harder, but don't doubt that researchers can get a model that produces a more performant solution than you, while using less code, and being more secure/robust/etc, in a fraction of the time. Your moat for the moment is alignment with stakeholders and high level taste, enjoy it while you can.

We had O1 and Gemini 2.5 Pro last year, which were both very good models, just not for long horizon agentic tasks. They were surprisingly capable, just not without a lot of hand holding.

Current models are very good for long horizon agentic tasks, but they lack taste and higher level organizational principles. They can get a tremendous amount done with limited hand holding, but they still tend to build messy slop unsupervised/without intervention.

By this time next year hand writing code will be dead outside the rarest domains, and agents will be able to build larger projects coherently, and in two years human software taste and architectural guidance will be basically redundant. I predicted the complete automation (3+ 9s) of software engineering in 3 years back in mid 2025, if anything we're slightly ahead of schedule.

If houses are capital, renting can never be just as good without de-incentivizing renting via legislation, which in reality will result in a massive rental shortfall, since nobody is going to maintain rental property when they could just invest their money in another asset.

The biggest problem with that saying is that it's completely ironic. "If you couldn't be bothered to think for yourself, don't bother speaking to me" is what people wish it meant, but what it signals to people outside of the haters club is "if you don't copy from the sources I like, don't communicate with me."

I think virtual reality and AI are a prelude to humanity living in Hong Kong style coffin apartments, and things will get progressively worse until we're in Matrix pods.

Criticizing use of agents for skill atrophy is valid, it definitely atrophies blank slate coding ability, though I don't think it atrophies engineering abilities unless you just YOLO all decisions to the agent. The data center/oligarchy complaints are also valid.

Saying agents produce shitty code is a bad argument though. They produce shitty codebase organization, but at a micro level their code is solid if not elegant. If you let them turn your codebase into a spaghetti mess, that's on you.

The things it loses are all the things that google models are historically excellent at, so that's a reasonable performance. I think the take home here is that the 1 bit models are probably better, but it's not a slam dunk given advanced quantization techniques.

A programmer's job is to deliver business value to their employer. If you're slowing down PR turnaround by mindlessly auto-rejecting on stuff that the suits don't care about, you better have a rock solid case for why that is going to deliver business value down the line, otherwise you're actively sabotaging your employer to bikeshed your personal preferences, which is the hallmark of a bad employee.

Rejecting an obviously bad PR after scanning the code quickly is one thing, burning business cycles on PR turnaround/latency to bikeshed bookkeeping without spending any time on the actual value producing portion of the PR is just bad. At the minimum you wasted an opportunity to give feedback on the proposed solution, thus probably necessitating another round of reviews, with the associated org latency.

A professional with standards who wasn't also unpleasant would put the time in to review the content of the commits with a request to clean up the history. Someone who looks at the history, thinks to themselves "not how I like it" and just auto-rejects the entire PR without any further thought is just a bad coworker.

People have been making games in a weekend using asset packs/store resources and engines for a while now. Indie gaming speed ran a lot of the marketing bullshit that every kind of app is suffering through now, and there's already a pipeline in place that works pretty well: itch.io demo -> vlogging/streaming/discord -> steam demo -> early access -> streamer promotion.

Small stacked PRs are a NIGHTMARE unless your org is consistently turning around PRs in low single digit hours, and team members are working on decoupled code so you don't have any stack weaving. You end up in a situation where the engineering director and some tech leads really like tools like Graphite, while the entire team working under them mutters irately under their breath daily.

Making small atomic commits as you go in the age of AI tends not to go great because it forces too much human in the loop in a lot of cases, and the percentage of AI code rework is significantly higher than manual code, so the history tends to be harder to keep clean.

It's ironically easier to create a messy agent work branch then have the agent cherry pick independent PRs from it into atomic commits post-work.

You sound pleasant to work with. I bet your coworkers route around you when they can, and when they can't they cherry pick from their working branch to deliver monolithic commits while rolling their eyes.

Squashing everything has a significant downside compared to functional atomic commits: git-bisect gets WAY less useful.

Games are complex systems, you actually have to playtest them to observe whether the changes you made increase or decrease the fun, despite your best guess. A lot of times things that seem like they should be fun actually aren't for reasons you couldn't predict. This is what makes them immune to being completely automated.

Indy games markets have been flooded with human slop since long before AI was a thing, the marketing playbook hasn't changed at all, if you just fire games into the void you never had a shot.

The answer to this question after a lot of reflection: games.

AI can slop fork or clone existing software well, but a clone of an existing game is pointless, it's basically guaranteed to be derivative and worse than the original game, and games aren't so expensive that you can't just buy the original. AI can't know if new mechanics or angles to an existing genre will feel good to play, or if a new genre is fun, that requires a human to experience the game in its totality.

Games are also very resilient to sloppy AI coding, and if an indy game crashes nobody is getting paged.

Holding cash is a terrible strategy, the US dollar is going to get devalued by the end of the Orange Troll's term due to money printing/bond market repression/trade hijinks.

We're in the pump of a pump and dump by the richest people in the world, who are trying to engineer exit liquidity in a way that doesn't immediately crash the market. They're going to cycle their holdings several times before things go to shit to avoid the brunt of it, we'll be able to see the wave on the horizon. I don't think just the SpaceX dumping will trigger a broader loss of confidence, but Anthropic + OpenAI dumping post IPO will ripple through to the neoclouds, which will ripple through to the semis, and probably tank things. In the short term (pre AI IPO) the market is volatile but pretty safe IMO, so just trade techically.

In the medium-long term I'd hold a mix of mining/industrial equipment/etc stocks in Euros, and Chinese AI/renewables stocks in Renminbi. The Euro stocks are safer and a significant portion of their increased valuation will come from currency shifts, and the Chinese stocks are riskier but have a big upside and are likely to benefit from Renminbi revaluation as China onshores "finishing"/final assembly of products (which much of the world is going to push them to do).

You act like humanity doesn't exist in a competitive environment. If you think AI codegen is a mistake? Just relax, keep writing code by hand and wait for the pendulum to prove you right while showering you in wealth. There are plenty of people making this bet, and I wish the best of luck to you because I'm 99% certain you're on the losing end of it.