HN user

bredren

7,189 karma

I build internet products and work extensively with generative AI.

Recent public projects include:

- Contextify: https://contextify.sh (A tool to assist CLI-based agentic programming workflows)

– FileKitty: https://github.com/banagale/FileKitty (a prompt engineering utility)

– Chief of Staff: https://chiefofstaffhq.com (a pre-generative AI text-to-speech SaaS)

I also very occasionally write at: https://banagale.com

Feel free to say hello, rob @ the domain above.

Posts102
Comments2,646
View on HN
news.ycombinator.com 5h ago

Ask HN: Anyone working on practical robotics (cobots) for the home?

bredren
2pts0
contextify.sh 19d ago

Show HN: Pull Claude Code transcripts into your Codex session, and vice versa

bredren
7pts3
contextify.sh 7mo ago

Total Recall: RAG Search Across All Your Claude Code and Codex Conversations

bredren
2pts1
banagale.com 7mo ago

⌘-Arrow Hotkey Navigation in Claude Code and Codex

bredren
2pts0
contextify.sh 7mo ago

Show HN: Contextify – Searchable History for Claude Code and Codex CLI

bredren
6pts2
news.ycombinator.com 11mo ago

Ask HN: How are you sharing Claude Code Sub Agents?

bredren
2pts0
banagale.com 1y ago

AI Is the Answer to Everything

bredren
3pts1
banagale.com 1y ago

Claude code planning mode posture is off balance

bredren
2pts0
github.com 1y ago

Show HN: Export Slack messages during your vacation into LLM-friendly format

bredren
3pts1
banagale.com 1y ago

ChatGPT's Quiet Shift in User Data Retention and Missed Opportunity for OpenAI

bredren
2pts0
www.modular.com 1y ago

Modular and AMD: Unleashing AI Performance on AMD GPUs

bredren
33pts1
news.ycombinator.com 1y ago

Ask HN: Has LLM prompting changed how you interact with people?

bredren
2pts2
www.blu-ray.com 1y ago

Sneakers (1992) – 4K makeover sourced from the original camera negative

bredren
410pts299
news.ycombinator.com 1y ago

Ask HN: What tools are you using to manage a shared enterprise prompt library?

bredren
5pts1
hunt.io 1y ago

KeyPlug-Linked Server Briefly Exposes Fortinet Exploits, Webshells, and Recon

bredren
2pts0
old.reddit.com 1y ago

Univ of Hong Kong releases Dream 7B (Diffusion reasoning model)

bredren
3pts0
old.reddit.com 1y ago

Thread comparing M4 Mac Studio 512GB vs. PC build running Deepseek-v3 671B-Q8

bredren
1pts1
github.com 1y ago

cvss-bt – enriching NVD CVSS scores to include temporal/threat metrics

bredren
1pts0
news.ycombinator.com 2y ago

Ask HN: Have you run Microsoft Word in headless mode?

bredren
5pts11
www.mentalfloss.com 2y ago

When Crispin Glover Was Replaced by a Lookalike in 'Back to the Future II'

bredren
1pts0
news.ycombinator.com 2y ago

Ask HN: Realtime conversation transcription / processing tool

bredren
1pts0
github.com 2y ago

Show HN: FileKitty – Combine and label text files for LLM prompt contexts

bredren
69pts21
old.reddit.com 2y ago

I asked ChatGPT to repeat the letter A as often as it can

bredren
1pts0
github.com 2y ago

Blender Docker images for CPU rendering updated daily

bredren
1pts1
www.billboard.com 3y ago

Burning Man’s Celebrated Mayan Warrior Art Car Destroyed in Fire

bredren
3pts0
chiefofstaffhq.com 3y ago

Show HN: Chief of Staff – Keep articles as audio, listen when you have time

bredren
3pts5
banagale.com 3y ago

Staying Hydrated in VR Workouts Is Problematic

bredren
1pts1
banagale.com 3y ago

The Supernatural Lawsuit and Apple

bredren
2pts0
en.wikipedia.org 3y ago

I need a haircut: Sampling lawsuit

bredren
2pts0
www.trendforce.com 3y ago

Intel Orders Delayed TSMC Slows Three-Nanometer Expansion

bredren
6pts1

I used to read him in my dad's PC Magazine subscription. I had some ~falling out over what my teenage self perceived as backwards take on mp3s and their proliferation. Perhaps one one of my earliest experiences with "kill your heroes" concept.

Glad for all of his contributions, I did enjoy reading his work.

It is ~a meme on subreddits that developers struggling to get good results out of any given model is a "skills issue."

But I think your comment drives at some authentic take on this. Skill with AI is not only crafting iterative prompts the agent will understand, but also very high domain-specific knowledge of what the prompts explore.

One without the other can result in frustration or worse.

Making 6 hours ago

This goes straight to the compiler and whether that precludes original craft. But I think it would be more relevant to discuss the use of open source packages.

Because I absolutely feel I made entire SaaS products, cool ones, by hand and they sat on a multitude of packages and infra I but glued together.

For frontend, I tailored templates, and instead put energy and time into making modern frontend toolchains and Django backends work together. This was satisfying to understand and build.

Being at ease with the minute details of Django or latest flags of esbuild felt like knowing the various modes of my Dewalt or Dewalt-colored power tools.

Building with AI is not a passive activity, and doing so well and efficiently, there is a ton to master and it evolved a great deal in the past six months.

So leveraging the tools using custom skills, custom CLIs, etc this is very important and very valuable thing to do. I would classify it as a serious contemporary computer science skill that should be taught in addition to the fundamentals.

Presume now, you're making and leveraging all of the modern bells and whistles of Claude Code and Codex. That is, you're reading the release notes and you have personal tooling so you can switch between them easily.

You can build incredible things with sustainable release workflows and reasonable security and possibly more than what a solo dev's "production quality" of yore.

I know this because I had Fable look at an entirely "hand crafted" SaaS I built over two years and it found about a page of bullets just in p0-p1 that I had not caught in my artisanal best effort. They were real issues, maybe unlikely but still things that I would have fixed if I had known.

But what can not be replaced is *taste*.

You can build all day and night and if it ain't good people won't use it. If you can't describe it in a way that makes sense to people they won't care.

If you build something people don't want you've not really made anything more than we did before we had AI as a tool.

I do think it can still achieve the same level of satisfaction if it is built well. Because it actually does take a lot of skill and knowledge to make something well using agentic programming. This is regardless of whether people want it.

In the "old west" if a horse was spooked and ran a person over, liability for the horse's owner varied but was similar to how this might be handled.

Consider this Supreme Court case, Brown v. Collins. [0]

A pair of horses were spooked by a nearby train engine, causing the animals to damage a stone post.

The driver of the grain-loaded wagon was not found to be at fault because he was "not guilty of any malice or unreasonable unskilfulness or negligence." And that the horses "did damage there against the will, intent and desire of the defendant."

If you read OpenAI's statement, in the Actions Being Taken section:

    1. "As part of the investigation, we are implementing strict controls in infrastructure configuration at the cost of research velocity while the vulnerabilities are patched."
Which I think implies that as a result of this incident, OpenAI has instituted controls beyond ~"reasonable skillfulness" such that it might even impede their business (and from a certain point of view impede the progress of the people governed by the laws which might hold the company accountable.)

I'm not taking a position on the merit of the above but I can understand the line of thinking.

[0] https://www.jstor.org/stable/3303590

Pardon the plug, but I have built a tool that ingests the state of conversations to a local DB in realtime. It has both macOS and Linux clients.

When context compaction introduces a gap, I use the /total-recall skill to pull prior turns back into context and off it goes.

The tool is free for personal use and has a source available local cloud option to sync convo histories across multiple machines.

http://contextify.sh/docs

Have you tried this:

   Look here and here and here are some tools, and /skill /skill [repo of folder paths etc] and here is what needs to happen: [stuff].

   ---

   Restate this request in your own words and enrich it as appropriate handling any gaps.
?

It is an ultra-lite way to plan, I suppose.

I like the format because:

    - I still get to put all my thinking into the request but then easily override the instruction
    - It is interesting to see my casual typo-riddled blast professionalized and improved upon.
    - Sometimes it surfaces useful questions that can save some time up front.
I think the models are doing this anyway, but I find the words "enrich" and "gap" are well understood by models and they demonstrate it in the response to the above pattern.

Anyhow, to get back to the point, there are still prompt-level tricks--but ultimately if repeated, should probably also be built into skills themselves!

Yes. I worked on two large monorepos, about one quarter per project. Maybe 80+ devs on first, maybe 40+ on second. Both were not sure how to describe this but ultra-high-velocity agentic driven development efforts.

On the first, there were ~no shared skills. There were some requirements set up but they were not minded properly and became stale / ate context for little gain. The hardest hit was in E2E tests which would flake and create long running, too-often failing CI. People would disable them, because they were not reliable and velocity was so high, no one was happy w them.

I maintained my own set of skills and CLIs to back them. I'd share them if they came up but it was like the old days of manage your own stuff. Not much credit for building and sharing devex tooling to the team.

But then on the second one we were in better shape--we had vendoring set up to distro skills automatically.

Before the project was well underway, I put time into understanding how all of our tests aught to be written. Finding the forbidden things, etc, getting review from our best test folks and ultimately landed on a `/test` that routed across all possible test types.

Like night and day. Instead of finding out while trying to get a release out the door that some corner of the project had a handful of flakes, tests were written the right way from the start.

Like, it was beautiful. And I don't think devs noted difference while building. Only that there was an absence of BS in CI.

Hard to quantify the lack of pain, but it was big!

I was thinking this past week I have gotten so lazy w my prompting via CLIs.

Back in the before I had put such discipline into my prompting and supporting context.

Now I’m like, “look here and here and here are some tools, and /skill /skill okay go.”

Or “restate this request in your own words and enrich it as appropriate handling any gaps. Okay go”

Do you have a link handy?

I used same session, set it to k3 model. I’ll look at the blog but the result was so bad I am prepared to abandon.

I should have saved the output.

I think maybe it was a mistake to not use open router.

I just tried this on the monthly $18 plan, having it do a basic task with its 2.7 model and then audit it using k3.

K3 got into some loop trying to run docker and after maybe the 6th attempt ran out of quota for the 5 hour window which represents 20% of the weekly.

I run 200 max and chatgpt pro, but I had to blink at that.

K3 didn't even write out what it was doing or provide any sense for why it was pursuing the execution path it was.

I'm in disbelief that this is a groundbreaking model, and do not think it represents a threat to Claude Code or Codex at this time.

Except often queued agentic flows must be checked in on. Or to use the comparison, 3D printers are not immune to making spaghetti all night when something goes wrong. (I’m not a 3d printing expert so maybe that is solved now)

It is common for agents to just stop because overload or some API error hijinks.

Or you get a TUI question that is blocking.

In general you’re right though, staring at tokens from agentic is not time well spent.

Some of these I’ve built custom harness around in iterm2 though.

I hadn’t thought about testing the bounds of model safety on comparatively benign requests compared to the type of thing described in frontier model cards.

I've been working on Contextify: https://contextify.sh

Makes it easy to use Claude Code or Codex interchangeably across multiple computers. Personal editions are free, I have a hosted commercial cloud (workgroups share AI history) and commercial self-hosted option available.

It has macOS and Linux clients and I released a guide for setting up the source-available, self-hosted cloud option this week: https://contextify.sh/docs/self-hosted/

I am thinking about the other AI cli environments and providing support for those as well.

PostHog FOSS 13 days ago

I used their AI chat last night and I was impressed with the product implementation. I was able to use it to make quick sense of data and even generate prompts to solve problems locally.

IDK what the prior AI behaviors have been, but for me, what they have now is an ~idealized version of AI product.

It is basically what I'd do if they didn't offer it but not as well: export data, import their docs in md., import some industry best practices in analytics into context etc.

Except its all right there. Not sure of the economics of it, as it was on a free account and some reasonably good model was powering the discussion. But it was about as engaged as I've been with an analytics tool.

Grok 4.5 14 days ago

I think it is a mix of the sibling replies here. I'd add that the company has seemed to find ways to ~do more with less.

I have never liked the various nerfs Anthropic has used to balance GPU (slowing down responses, quota variance, model optimizations etc) and it definitely has burned a lot of good-will.

But it has seemed that being able to look beyond the short term pitchforks has worked quite well.

GPT‑Live 14 days ago

I also think this undersells the real value of the bot, which is to handle tasks via voice that an average human either would not or could not do.

In the video example with the grannies, the knitter is essentially wanting a PA. Regular folks don't have PAs. Even when that became a thing in the aughts they were all outsourced.

When I've used voice chat, it has often turns into rabbit holes on very niche topics. For example, I had one start about the 1996 performance of Rage Against the Machine in Portland, Oregon that was supposed to feature Wu Tang Clan. (already outside most human's knowledge) that dove into details of the club scene in Los Angeles at the time of RATM's signing to Epic Records.

Was anyone else here at that '96 show in Portland? It seems like it might be challenging to find a person on the internet able to engage on the topic.

The person may exist, but not during my fleeting interest in the subject while walking to the park.

How are y'all carrying context history from one agent to the other?

I also flip between the models due to quota, TUI enhancements, model updates and service availability.

To handle this, I built a thing that normalizes your transcripts between Claude Code and Codex into a shared DB, then a CLI and skill.

It has made it so it doesn't matter what I built where (or when) I just refer to the work and drop in a /total-recall (or $total-recall on codex) and the agent brings it into the current convo.

I realize there are a lot of ~memory tools out there, but I think particular my approach and product behavior is unique.

If you're open to giving it a try, I'd appreciate any feedback: https://contextify.sh recent show hn: https://news.ycombinator.com/item?id=48777790

Hey, I’m watching this thread and happy to answer questions or take feedback.

I’ve described some of the challenges in parsing and normalizing the variant JSONL files codex and Claude code produce in prior comments.

There is some interesting nuance to how Claude Code writes its queue-operation records.

Early on I focused on the macOS client, working. To get the summarization tuned in, and thought of the product as more of a HUD.

But as I used it I found that making the corpus of session data available to the AI was the real power tool.

This is what guided me toward adding a Linux client (which does not have a GUI, but provides all of the ingestion capability of the Mac app.)

I’ve written a skill for codex and Claude code that designates an orchestrator on the primary worktree and is agnostic about what type of AI workers are on the N supporting worktrees.

The orchestrator knows which AI client is running in any given worktree, so it would be fairly easy to designate which AI should receive what kind of tasks.

You run either Claude or Codex in tabs for each work tree. I do have some AI TUI specific instructions, for instance codex is primitive at monitoring compared to CC. So, there are additional notes for Codex workers on how to properly monitor for new "mail."

You work with the orchestrator on the primary worktree and allow it to delegates tasks to the workers and answer their smaller questions.

It surfaces results and assisting them with context clearing when needed.

The orchestrator and workers communicate using a simple shared file system under tmp/* and together they can handle a big and varied workload.

I use iterm2, so I’ve also added iterm2 specific python that allows the orchestrator to “kick” a worker or perform tasks otherwise veto'd by the TUIs (ie /clear) by modifying the input and submitting it.

I would attribute Disney's use of scarcity as a primary means to drive film and TV box office and streaming dollars in the Star Wars franchise.

This is already under threat due to the Star Wars AI videos being released on Youtube, seemingly without constraint as of yet.

The videos are not Hollywood quality [0], however they circumvent rules Disney can't easily break like using the likeness of any actor at any age in any circumstance.

These fan made videos get lots of views. Even if they were all removed from YouTube, this will be a difficult thing to stop.

I believe a generally accepted "good" or even "great" unofficial, Star Wars film built without sets or actors using AI is inevitable. And that this will be true for any popular franchise.

The natural corollary to this arc is into games, where using AI to code most or all of a AAA-competitive title would be considered inevitable.

I suspect Disney and Sony have at least someone pointing at this outcome.

[0] I suppose idealized Hollywood quality. They are better than some films.

Nano Banana 2 Lite 22 days ago

How do you get the ~real time prototype things shown in the "hands on" section of this page?

gemini.g lets you add a canvas or use image gen, but it isn't clear to me how you stick in the "space lift" prompt and out comes what is demo'd

Claude Sonnet 5 22 days ago

This is what I realized, can you provide more detail on how you've observed this? The /usage screen does not make it clear.

Claude Sonnet 5 22 days ago

It would be good if Anthropic provided some kind of feedback or even toggle to auto-route requests for models being used at thinking levels that would be a better value using a different model.

Sort of like, getting an automatic upgrade at a car rental or hotel if there is availability.