HN user

cowlby

159 karma
Posts1
Comments81
View on HN

While I don't disagree, there is a scale/volume problem I wanted to solve. If I read to my kids 365 days a year for 8-10 years, that's over 3,500 books. There's a lot of crap publishing out there just cause it's "natural" or "human" does not make it good. So yes we read the classics, but it's great to have a fun creative book factory.

Now that's a Black Mirror episode. It's the story of all technology though, caveat emptor.

For me it's still about human connection though. I read the stories we create together. It's just a great tool. It makes any topic relatable. I.e. even crazy fun ones like "Claude weave a bedtime story about how the 5nm chip fab process works including EUV lithography and clean rooms".

Quick short 5-10 minute read and next thing you know we're talking about lasers and how sand becomes computers.

This just seems like laziness vs AI = bad. It's not like publishers are putting out masterpieces with human writing. They're cranking out minimum viable content as well.

I've found that by putting meaningful effort into AI storytelling, I can create bespoke stories that my kids love night after night.

My workflow is below: Caveat that it costs about $0.25-$0.50 to weave a book like this with Claude Sonnet and Gemini Nano Banana Pro. But to me the cost is worth it for the quality.

- Use Claude structured output and ask for page1, page2 ... pageN instead of an array of pages or wall of text.

- Pass a story arc as a set of values to the prompt. I.e. say each page has an emotional beat between 0.0 and 1.0. For a "man in hole" type of story: page1 starts at 0.6, page2 = 0.5, page5 = 0.25, page10 = 0.85. This ensures page 5 lands the "crisis" and page10 resolves higher than the start.

- For illustrations, have Claude generate the story text and an illustration prompt per page. i.e. page1: { "text": "...", "illustration": "..." }.

- For art consistency, add an "Art Direction" key to the structured output. Feed this into Gemini/OpenAI and ask for an art board visual guide & character reference sheet.

- Feed the page text, illustration prompt, and the art board to Gemini/ChatGPT images. I'm constantly surprised at the quality of the output.

Here's an example set of pages from a magic school bus style story about the immune system

[image] https://media.discordapp.net/attachments/839188039229112353/...

Defense in depth approach, would this work to help as a layer?

- Wrap user input in strong markers like <user-input-do-not-trust />

- Have the agent compute what it will perform as structured output.

- Have another agent evaluate the structured output against the intent of the code.

- Determine if it aligns or deviates from the intended workflow. Execute or deny gate from here.

Yes this, I bought sets like the Arctic Ship and police bases thinking they'd be like dioramas. They quickly became components that they build other things on top of. It took some time to mentally un-anchor from the $500 spent on the sets lol.

MCP is dead? 2 months ago

I use all three (MCP/CLI/API) based on what Claude excels at:

* CLI: GitHub & AWS it already knows how to operate the CLIs well. Even learned about a few new CLIs like 1Password's op which it volunteered one day.

* MCP: Supabase, Shopify etc. where the CLI would be non-obvious and the affordances from the tools/descriptions helps Claude maneuver.

* API: Sometimes it just knows an API exists and is able to call it directly with python/curl. I discovered from Claude the Pokemon ecosystem has a free API out there for example.

The analogy is helpful, but yes we should be able to “intelligently design” something better than sleep analogues since we’re not constrained by evolution like in humans.

I gave up on Plex and just vibe coded a couple quick solutions instead:

- Direct streaming 4K blu-ray atmos rips to home theater: just connected a PC via long HDMI fiber optic cable

- Library organization: tinyMediaManager is awesome for this.

- Watching ad hoc on iPad/iPhone: built a simple Next.js app that lists my movies and and a python script that encodes movies as MP4 and creates HLS playlists. No more real time transcoding.

- Downloading movies to my iPad for long flights: vibe coded an iOS app Claude handled all the AV code to download the same HLS streams.

I wonder how much this is correlated to token budgets? I'd be curious to see a split between $20/$100/$200/$500+ usage and see if there is different responses. I'm in the $400 range with Claude + Cursor subscription, use Opus exclusively, and my experience is wildly different from this.

I chuckled cause the convenience/grocery store is laid out to make us find the high margin items and not what we need. They can't explain it to us otherwise we'd shop less.

Are they trying to drive safety or revenue? The second order effect people forget about is tickets are a source of revenue for cities and police depts. Surely driverless car companies will absorb a few tickets and fix the issue quickly.

So I do wonder what happens in the future when roads and cars are all automated and city funding from this channel dries up.

I was looking at object storage recently and I hadn't realized how much profit cloud providers drive via egress. And it's so perfectly hidden from the marketing. Ended up going with Cloudflare R2 for free egress.

Yes, I just think there's a sane way to do things that is not "never let LLM agents do things".

For dev/prod staging though, there's that other story on HN right now of an LLM agent that maneuvered it's way to prod credentials and destroyed prod. And backups went along with it. I'm paranoid enough to think backups in this use case means out-of-band uncorrelated storage.

I just think there's more nuance to it. Some things have an implicit RTO/RPO/SLA of say a day. Risk is also correlated to recovery and rollback. And there's levels of LLMs out there.

Surely in the Venn Diagram of things, there's a slot where it's okay let a Claude Opus agent run on a process with good backups/recovery? Where taking the risk of a 1-hour restore job is worth the LLM agent velocity?

For extra paranoia, surely even Opus/Mythos can't figure out how to destroy log level backups to immutable storage.

Not everything is a SaaS. I commented this elsewhere but I picture all the business running on spreadsheets/CSVs/MS Access databases on someone's desktop. People delete these all the time by accident. They have no security, no authentication, etc.

An LLM agent (with RW access to a DB), a developer, and a few days these become proper apps that SMB business would pay well for.

Sure don't give an LLM agent access to PII or properly built CRMs etc. But to not see the rest of the landscape seems like a missed opportunity.

That's the issue that I feel misses the forest for the trees. Relatively simple applications or thin slices exist right now, in production, in critical paths, as spreadsheets/CSVs/files on someone's desktop. That's the pent up demand I picture out there for developers.

Go to any SMB out there and there's a goldmine of processes that could be improved with LLM agents with full RW access to a database. Where backups are sufficient as a recovery mechanism that is better-than-before.

Yes, that's the right framing. Millions flow through spreadsheets/CSVs/MS Access with none of the auth/backups/architecture people seem to be stuck to.

I saw an article on HN one time about CSVs and how much business still flows through them. Reminds me of the xkcd comic about the one tiny block propping up lots of infrastructure. It stuck with me because it's ripe area for LLM agent based upgrades.

Sure don't give LLMs access to the well architected blocks. But not wanting to improve the brittle areas seems crazy to me even if it's contrarian.

I commented this elsewhere: There's thousands of small and medium business though. They have maybe one true CRM, and a dozen spreadsheets/files floating around that would benefit becoming proper apps. People delete spreadsheets all the time!

Sure don't give an LLM agent write access to the modeled CRM that took months/years to build.

But turning a spreadsheet into an app in a few days? By giving the LLM proper read/write capabilities for velocity? I think the case is there for it. Right tool for the right job.

There's thousands of small and medium business though. They have maybe one true CRM, and a dozen spreadsheets/files floating around that would benefit becoming proper apps. People delete spreadsheets all the time!

Sure don't give an LLM agent write access to the modeled CRM that took months/years to build.

But turning a spreadsheet into an app in a few days? By giving the LLM proper read/write capabilities for velocity? I think the case is there for it. Right tool for the right job.

I think of all the pent up demand for proper applications that are just infeasible when it would take a developer weeks-to-months to create. Now it's just a few days with an LLM agent.

Examples for me are all the apps that live in a spreadsheet, or in a MS Access database. Or all the crappy ad backed apps on the iOS app store. People wipe full spreadsheets all the time and backups are the only recovery.

Just last weekend I was frustrated with the poor quality of Pokedex type apps that spam ads left and right. Took just one session with Claude Opus to roll a custom Pokedex. It knew internally about things like the PokeApi dataset, Pokemon data modelling etc. To-the-hour snapshots of the database are trivial for bespoke apps like this so the LLM agent velocity seems like an okay trade off for me.

Clearly people don't agree...

LLM agents are unlocking demand and supply for applications that wouldn't have been possible before due to time constraints though. There's a growing demand for single user or smaller scoped apps where giving LLM agents direct access means velocity. The failure/rollback model is much easier with these as long as we have good backup hygiene.

I believe they use Yodlee and yes there is a lot of trust in Yodlee/Tiller to keep data safe. The integrations go through an OAuth type flow where you hit say Chase first and approve/revoke individual accounts so it seems like it's API based now, not screen scraping.

For all those concerns, I bet you could automate just parsing all the data from the statements or a CSV export.

It's great because Claude Code generated complex analysis/models 10X better than I was familiar with.

The key was normalizing the payee/categories so we can analyze month to month, and separating fixed vs variable spend. It then did a fancy Monte Carlo simulation with the computed mean/stddev per payee. And out came T+30/T+60 estimates at P50/P80/P90.

It was all GitHub Spec Kit + Claude Opus tbh. I narrated a couple paragraphs of how I wanted to sync to flow and it knocked it out in one pass practically.

Here's the initial spec it created. I started off writing to a local sqlite db instead of Supabase: https://gist.github.com/cowlby/0dbeb52403c3f3c0f1d8122505203...

Edit: Here's also the DSL categorization spec. First one was string based, found it cumbersome, so second one was the Markdown table refactor: https://gist.github.com/cowlby/30d6b5cf132fc1424ab146f0eaf4a...

https://gist.github.com/cowlby/d569c8e05b5b6eecfd4d237372c06...

(edit: put in Gist instead of inline here)