HN user

dbreunig

2,753 karma
Posts259
Comments252
View on HN
github.com 6h ago

Show HN: Brew doctor` for Skills and MCP loadouts

dbreunig
2pts0
www.dbreunig.com 18d ago

Understanding the Dynamics of the AI Ecosystem with Pace Layers

dbreunig
3pts1
www.dbreunig.com 29d ago

The Problem is Prompt Debt: You can't be model agnostic and hand-tune prompts

dbreunig
4pts0
www.dbreunig.com 1mo ago

When an agent can explain anything, what is the role of human-centric docs?

dbreunig
4pts0
www.dbreunig.com 2mo ago

Overfitting to First Party Harnesses

dbreunig
1pts0
www.dbreunig.com 3mo ago

Cybersecurity looks like proof of work now

dbreunig
562pts213
www.dbreunig.com 3mo ago

Claude Code builds a system prompt

dbreunig
4pts0
www.oreilly.com 3mo ago

The Cathedral, the Bazaar, and the Winchester Mystery House

dbreunig
2pts0
www.dbreunig.com 3mo ago

The 2nd phase of OSS in the agentic era: From clones to reimaginings

dbreunig
2pts0
www.dbreunig.com 3mo ago

The Cathedral, the Bazaar, and the Winchester Mystery House

dbreunig
190pts66
www.cmpnd.ai 4mo ago

Build a deep researcher and learn DSPy Signatures and Modules

dbreunig
2pts0
www.dbreunig.com 4mo ago

Can chat bots accommodate advertising?

dbreunig
2pts0
www.dbreunig.com 4mo ago

Learnings from a No-Code Lib: Keep the Spec Driven Development Triangle in Sync

dbreunig
4pts3
www.dbreunig.com 4mo ago

Claude and the Dow: AI is unlike other tech because AI has embedded judgment

dbreunig
1pts1
www.dbreunig.com 4mo ago

Two Beliefs About Coding Agents: Devs Don't Realize What They Bring

dbreunig
2pts0
www.dbreunig.com 5mo ago

Why is Claude an Electron app?

dbreunig
428pts458
www.dbreunig.com 5mo ago

Analyzing How System Prompts Define Agent Behavior

dbreunig
3pts0
www.dbreunig.com 5mo ago

The Potential of RLMs

dbreunig
3pts1
www.dbreunig.com 5mo ago

The Rise (and Limits) of Spec Driven Development

dbreunig
3pts0
www.dbreunig.com 6mo ago

A OSS Library with No Code, Only Specs

dbreunig
3pts2
www.dbreunig.com 6mo ago

2025 in Review: Jagged Intelligence Becomes a Fault Line

dbreunig
2pts0
www.dbreunig.com 6mo ago

Why I write (and you should too)

dbreunig
1pts0
www.dbreunig.com 7mo ago

Applied AI in 2025: From 'Naked' Model Calls to Tool Use Environment Calls

dbreunig
1pts0
www.dbreunig.com 7mo ago

Enterprise Agents Have a Reliability Problem

dbreunig
2pts0
www.dbreunig.com 8mo ago

Don't Fight the Weights: Learn to Spot Contexts That Go Against Training

dbreunig
3pts0
www.dbreunig.com 10mo ago

About That MIT Report: Enterprise AI Looks Bleak, but Employee AI Looks Bright

dbreunig
3pts0
www.dbreunig.com 10mo ago

AI Companies School Like Fish to the New Use Case

dbreunig
3pts0
arxiv.org 10mo ago

Monolingual speech recognition models beat multilingual models ~30x bigger

dbreunig
3pts0
www.dbreunig.com 10mo ago

Are Chatbots Incompatible with Advertising?

dbreunig
3pts0
www.dbreunig.com 11mo ago

The AI Job Title Decoder Ring

dbreunig
91pts71

Elizabeth Lopatto at The Verge makes a strong case we _do_ have proof that Musk is actively gathering and throwing fuel on the fire: https://www.theverge.com/ai-artificial-intelligence/929129/s...

But the thing is, Molo doesn’t actually have to be good at this job, because the point of this trial isn’t to win — though I’m sure Musk wouldn’t mind a win. The point is to punish Altman, Brockman, and OpenAI. Musk has done that pretty thoroughly — reinforcing in the public’s mind that Altman is a liar and a snake. This morning, I read an exclusive in The Wall Street Journal that assorted Republican AGs and the House Oversight committee wanted to look into Sam Altman’s investments. References to the trial are peppered throughout the article.

Model testing and swapping is one of the surprises people really appreciate DSPy for.

You're right: prompts are overfit to models. You can't just change the provider or target and know that you're giving it a fair shake. But if you have eval data and have been using a prompt optimizer with DSPy, you can try models with the one-line change followed by rerunning the prompt optimizer.

Dropbox just published a case study where they talk about this:

At the same time, this experiment reinforced another benefit of the approach: iteration speed. Although gemma-3-12b was ultimately too weak for our highest-quality production judge paths, DSPy allowed us to reach that conclusion quickly and with measurable evidence. Instead of prolonged debate or manual trial and error, we could test the model directly against our evaluation framework and make a confident decision.

https://dropbox.tech/machine-learning/optimizing-dropbox-das...

"Think step by step," was just a sentence you appended to your prompt.

It ended up kicking off reasoning training which enabled the massive gains in coding, tool use, and more over the last 18 months.

So yeah, it's "just using LLMs in a specific way."

Last year they pushed out an update stating if any “Meta AI” is left on, they can access image data for training,

I turned the AI off and used them as headphones and taking videos while biking. After a couple rides, I couldn’t bring myself to put them on because people started to recognize them and I realized I didn’t want to be associated with them (people are right to assume Meta has access to what they see).

Meta Ray Bans, if kept simple, could have been a great product. They ruined them.

Check out “Recursive Language Models”, or RLMs.

I believe this method works well because it turns a long context problem (hard for LLMs) into a coding and reasoning problem (much better!). You’re leveraging the last 18 months of coding RL by changing you scaffold.

Author of the post here.

I didn’t say AI was bad and I acknowledged the benefits of Electron and why it makes sense to choose it.

With 64gb of RAM on my Mac Studio, Claude desktop is still slow! Good Electron apps exist, it’s just an interesting note give recent spec driven development discussion.

Agree. I bucket things into three piles:

1. Batch/Pipeline: Processing a ton of things, with no oversight. Document parsing, content moderation, etc.

2. AI Features: An app calls out to an AI-powered function. Grammarly might pass out a document for a summary, a CMS might want to generate tags for a post, etc.

3. Agents: AI manages the control flow.

So much of discussion online is heavily focused towards agents so that skews the macro view, but these patterns are pretty distinct.

There was a good study on this a few years ago that ran the numbers on this and landed on white paint for residential homes as the best option, for a few reasons, if I remember correctly:

- Installation, maintenance and transmission costs are lower when solar is aggregated on farms - Solar offsets air conditioning, but that moves the heat outside. White roofs reduce the need for AC, which helps significantly with urban heat scenarios

A quick search yields a UCL study, which supports the lower claim: https://phys.org/news/2024-07-roofs-white-city.html

We're Joining OpenAI 11 months ago

Yeah, I agree. Almost mentioned in the post how I imagine an ad PM at OpenAI is jealous of an ad PM at Perplexity.

I also dislike the term. It feels concocted to evoke “tacticool” vibes.

Unless you’re pushing new firmware onto a drone in Ukraine, FDE is stolen valor.

I’m wondering that too. I think better routers will allow for more efficiency (a good thing!) at the cost of giving up control.

I think OpenAI attempted to mitigate this shift with the modes and tones they introduced, but there’s always going to be a slice that’s unaddressed. (For example, I’d still use dalle 2 if I could.)

I’m aware of LoRA, Civitai, etc. I don’t think they are “widely known” beyond AI imagery enthusiasts.

Krea wrote a great post, trained the opinions in during post-training (not during LoRA), and I’ve been noticing larger labs doing similar things without discussing it (the default ChatGPT comic strip is one example). So I figured I’d write it up for a more general audience and ask if this is the direction we’ll go for qualitative tasks beyond imagery.

Plus, fine-tuning is called out in the post.

Came here to say this: people present MCP’s verbosity as all the context the LLM needs. But almost always, this isn’t the case.

I wrote recently, “ Connecting your model to random MCPs and then giving it a task is like giving someone a drill and teaching them how it works, then asking them to fix your sink. Is the drill relevant in this scenario? If it’s not, why was it given to me? It’s a classic case of context confusion.”

https://www.dbreunig.com/2025/07/30/how-kimi-was-post-traine...

Sometimes buzzwords turn out to be mirages that disappear in a few weeks, but often they stick around.

I find they takeoff when someone crystallizes something many people are thinking about internally, and don’t realize everyone else is having similar thoughts. In this example, I think the way agent and app builders are wrestling with LLMs is fundamentally different than chatbots users (it’s closer to programming), and this phrase resonates with that crowd.

Here’s an earlier write up on buzzwords: https://www.dbreunig.com/2020/02/28/how-to-build-a-buzzword....