I got an early tip on this and used it to build a brand mood board and it crushed previous attempts with Claude. Highly recommend.
HN user
dbreunig
Elizabeth Lopatto at The Verge makes a strong case we _do_ have proof that Musk is actively gathering and throwing fuel on the fire: https://www.theverge.com/ai-artificial-intelligence/929129/s...
But the thing is, Molo doesn’t actually have to be good at this job, because the point of this trial isn’t to win — though I’m sure Musk wouldn’t mind a win. The point is to punish Altman, Brockman, and OpenAI. Musk has done that pretty thoroughly — reinforcing in the public’s mind that Altman is a liar and a snake. This morning, I read an exclusive in The Wall Street Journal that assorted Republican AGs and the House Oversight committee wanted to look into Sam Altman’s investments. References to the trial are peppered throughout the article.
Among benchmarkers its a frequent topic. Qwen BURNS reasoning to get its scores.
Model testing and swapping is one of the surprises people really appreciate DSPy for.
You're right: prompts are overfit to models. You can't just change the provider or target and know that you're giving it a fair shake. But if you have eval data and have been using a prompt optimizer with DSPy, you can try models with the one-line change followed by rerunning the prompt optimizer.
Dropbox just published a case study where they talk about this:
At the same time, this experiment reinforced another benefit of the approach: iteration speed. Although gemma-3-12b was ultimately too weak for our highest-quality production judge paths, DSPy allowed us to reach that conclusion quickly and with measurable evidence. Instead of prolonged debate or manual trial and error, we could test the model directly against our evaluation framework and make a confident decision.
https://dropbox.tech/machine-learning/optimizing-dropbox-das...
No reason it can't. I know people currently generating specs from existing code; just gotta write the pipeline.
"Think step by step," was just a sentence you appended to your prompt.
It ended up kicking off reasoning training which enabled the massive gains in coding, tool use, and more over the last 18 months.
So yeah, it's "just using LLMs in a specific way."
Last year they pushed out an update stating if any “Meta AI” is left on, they can access image data for training,
I turned the AI off and used them as headphones and taking videos while biking. After a couple rides, I couldn’t bring myself to put them on because people started to recognize them and I realized I didn’t want to be associated with them (people are right to assume Meta has access to what they see).
Meta Ray Bans, if kept simple, could have been a great product. They ruined them.
Check out “Recursive Language Models”, or RLMs.
I believe this method works well because it turns a long context problem (hard for LLMs) into a coding and reasoning problem (much better!). You’re leveraging the last 18 months of coding RL by changing you scaffold.
Author of the post here.
I didn’t say AI was bad and I acknowledged the benefits of Electron and why it makes sense to choose it.
With 64gb of RAM on my Mac Studio, Claude desktop is still slow! Good Electron apps exist, it’s just an interesting note give recent spec driven development discussion.
I keep saying this, it’s my new favorite metaphor.
That's cute.
Agree. I bucket things into three piles:
1. Batch/Pipeline: Processing a ton of things, with no oversight. Document parsing, content moderation, etc.
2. AI Features: An app calls out to an AI-powered function. Grammarly might pass out a document for a summary, a CMS might want to generate tags for a post, etc.
3. Agents: AI manages the control flow.
So much of discussion online is heavily focused towards agents so that skews the macro view, but these patterns are pretty distinct.
There was a good study on this a few years ago that ran the numbers on this and landed on white paint for residential homes as the best option, for a few reasons, if I remember correctly:
- Installation, maintenance and transmission costs are lower when solar is aggregated on farms - Solar offsets air conditioning, but that moves the heat outside. White roofs reduce the need for AC, which helps significantly with urban heat scenarios
A quick search yields a UCL study, which supports the lower claim: https://phys.org/news/2024-07-roofs-white-city.html
Yes, if you put unrelated stuff in the prompt you can get different results.
One team at Harvard found mentioning you're a Philadelphia Eagles Fan let you bypass ChatGPT alignment: https://www.dbreunig.com/2025/05/21/chatgpt-heard-about-eagl...
Yeah, I agree. Almost mentioned in the post how I imagine an ad PM at OpenAI is jealous of an ad PM at Perplexity.
I also dislike the term. It feels concocted to evoke “tacticool” vibes.
Unless you’re pushing new firmware onto a drone in Ukraine, FDE is stolen valor.
You should read the post. You might find the “domain” discussion interesting.
I will be thinking about this comment for a bit. Thanks for this perspective!
The team at Chroma is currently looking into this and should have some figures.
I’m wondering that too. I think better routers will allow for more efficiency (a good thing!) at the cost of giving up control.
I think OpenAI attempted to mitigate this shift with the modes and tones they introduced, but there’s always going to be a slice that’s unaddressed. (For example, I’d still use dalle 2 if I could.)
I’m aware of LoRA, Civitai, etc. I don’t think they are “widely known” beyond AI imagery enthusiasts.
Krea wrote a great post, trained the opinions in during post-training (not during LoRA), and I’ve been noticing larger labs doing similar things without discussing it (the default ChatGPT comic strip is one example). So I figured I’d write it up for a more general audience and ask if this is the direction we’ll go for qualitative tasks beyond imagery.
Plus, fine-tuning is called out in the post.
Came here to say this: people present MCP’s verbosity as all the context the LLM needs. But almost always, this isn’t the case.
I wrote recently, “ Connecting your model to random MCPs and then giving it a task is like giving someone a drill and teaching them how it works, then asking them to fix your sink. Is the drill relevant in this scenario? If it’s not, why was it given to me? It’s a classic case of context confusion.”
https://www.dbreunig.com/2025/07/30/how-kimi-was-post-traine...
I really like Relay.app for non-coders. People can get wrapped around the wheel with n8n and co.
A similar, fun case is where researchers inserted facts about the user (gender, age, sports fandom) and found alignment rules were inconsistently applied: https://www.dbreunig.com/2025/05/21/chatgpt-heard-about-eagl...
Wrote about this about a month ago. I think it’s fascinating how they developed these prompts: https://www.dbreunig.com/2025/07/05/cat-facts-cause-context-...
I studied linguistic anthropology, in addition to CS. Been at it since 2002.
And I wrote the first post before the meme.
Sometimes buzzwords turn out to be mirages that disappear in a few weeks, but often they stick around.
I find they takeoff when someone crystallizes something many people are thinking about internally, and don’t realize everyone else is having similar thoughts. In this example, I think the way agent and app builders are wrestling with LLMs is fundamentally different than chatbots users (it’s closer to programming), and this phrase resonates with that crowd.
Here’s an earlier write up on buzzwords: https://www.dbreunig.com/2020/02/28/how-to-build-a-buzzword....
While researching the above posts Simon linked, I was struck by how many of these techniques came from the pre-ChatGPT era. NLP researchers have been dealing with this for awhile.
Yes, look up Winshuttle.
A very successful company with some of the happiest customers I’ve ever seen, whose entire product was a SAP hack that allowed people to enter their data using Excel. As someone unfamiliar with SAP, absolutely blew my mind.
Not dictation…copy/paste I think. Thanks, fixed.