And here I've been using Maybelline for nothing!
HN user
mergesort
Indie developer building plinky.app — the easiest way to save links for later. Also teaching people how to use AI @ build.ms/ai.
I love writing: professionally at build.ms and personally at fabisevi.ch.
Sorry, when I say I think I have a balanced take on AI what I mean is that I do my best to weigh both the pros and cons of this technology as opposed to a more extreme behavior like spending all day chatting with LLMs or posting all day on X about how AI is already better than me at everything and that jobs are over.
If I had to assign a confidence score for whether agents will change the way we all work and many aspects of how we live, I would put it at a 7/10, maybe 8/10. I felt about the same about the smartphone. While many things we do look the same way they did in 2005 (we still drive on roads, kids still go to school), at the same time it's undeniable that much of our lives are intermediated through a small screen and many societal dynamics have shifted due to that technology's existence.
I will concede that you should read my post with that context and draw your own conclusions about the veracity of my perspective — but I think it is more well-reasoned than what people generally attribute to "LLM hype". (Of course it's a bit tautological that I believe that, but I try to surround myself with people of all kinds technical and non-technical and like to think I stay reasonably grounded.)
All that said, I think the code from a leading company being bad and yet delivering good results is more a sign of the technology's jagged frontier[^1]. Calculators can't write sonnets the same way that LLMs are bad at math, but that doesn't make them useless — it just makes them a tool. This is a tool in our tool belt and I find is surprisingly useful as a general purpose technology despite it's limitations. (Which is related to the main argument I make in the post that bad code leading to good results may imply that we're under and overweighting certain aspects of what is important in software development, and that our expectations of code may may need to be recalibrated often as we gather more evidence.)
[^1]: https://www.oneusefulthing.org/p/centaurs-and-cyborgs-on-the...
Hey there! I'm not sure I have a universally applicable answer, but I can do my best to map out some things about my process and flow that hopefully help a bit and answer your question.
- I've had an iPhone for half my life (I'm 36 and got one when I was 19), so I've gotten pretty acclimated to typing on the go. I try switching to dictation every couple of months but the iPhone's dictation trips up over enough words that I find it more frustrating than typing as I walk.
- I don't do this but if you're worried about the thoughts disappearing I would absolutely recommend recording a voice note. As I'll touch on in a moment — do not let those thoughts disappear! Even the act of codifying them into something tangible allows you to process them more deeply.
- I live in NYC but I start most mornings by taking a walk along a relatively quiet street, so I rarely end up having to worry about bumping into someone. That is definitely not universally applicable advice. (:
- I look up as I'm typing and let autocorrect take the wheel. That works at least 95% of the time, so if I make the occasional typo it doesn't really matter, I'll just fix it in post.
- It helps to have an app with a great text editing experience. I've found that there are very few out there that are fluid, many have incredibly subtle hitches that make it hard to quickly jot down thoughts onto a canvas. I really love Craft (https://craft.do) and have been using it for years, so at this point it feels more like an extension of me than an app.
- This is surely unique to everyone but my writing tends to start from a few keystone thoughts. Once I have one written down, I let myself almost free associate, writing down whatever comes to mind from that initial thought to make sure I do not forget. I can always edit after the fact, and often the editing process leads to more interesting insights as well. But the main thing I want to avoid is losing those sparks, in the same way that you're mention your thoughts evaporating. Don't let those go, just get 'em on paper and sort through 'em afterwards.
- That's all a lot easier to do on my phone than if I approached the problem as "type an essay on my phone", so I'll almost always edit a post on my computer before publishing. Yesterday was more of an exception than the rule though because I was bouncing around between doctors all day, so I wrote all of this on my phone [not expecting it to blow up or get a ton of scrutiny].
Not sure if anything's missing but I'm happy to share anything that may be helpful! Clearly this post wasn't perfect, but I've been much happier since I started letting myself write out long-form thoughts on my phone and sharing them as blog post rather than firing them off as pithy tweets that decay into the ether once the algorithm says it's time for them to go.
Hey there, post author here. I think if you read deeper into my blog post history you’ll see that I have a reasonably balanced take on AI.
I generally think this will be a very important technology so I teach the subject to make sure people understand how to use it as leverage in their lives. (Yes as paid workshops, but I also volunteer weekly for 3-4 hour sessions at a non-profit where I get nothing more than the joy of helping people learn a valuable skill.)
At the same time just last week I wrote a post decrying the slop people are hoisting on their coworkers[^1], because I want people to use this technology in a positive way to create the lives they want, not to create downstream consequences for others. Ultimately I think agentic systems are incredibly powerful but also a technology that lends itself to anti-social behavior because of how independently empowering it can be. And so I hope that with the right exposure, discussion, and teaching we can take advantage of its democratizing nature, while reinforcing that what makes us special as humans is that we care and coordinate to do greater things. Value in this world — not just in the financial sense that we often boil it down to when we talk about this subject.
Hope that context helps provide a better lens into the piece, and that I still do care a lot about code and everything else that got me here, but that you are also reading personal reflections of who I am in a time of change, which is making me question (or reinforcing) some of the fundamental things I believed about software and sometimes the world more widely.
Heya, post author here. I just want to say that I actually agree with your surprise. In the piece what I’m trying to articulate is not that this can go on forever, but that I’m genuinely surprised we’ve made it a year and they continue to ship at a rapid clip and the product is still reasonably good.
I’ve had to question the value of code a lot over the last couple of years, and this leak continues to reinforce the notion that I’ve vastly overestimated it my entire career.
Now we could be moments away from hitting any of the rules described on https://how.complexsystems.fail, but if you’d asked me a year ago how long it would take to get there with people working this way I would have definitely taken the under. That difference in what I believed and what I see with my own two eyes is what has me questioning my priors, because my calibration seems to need readjustment (maybe large or maybe small) for the world of software we’re in right now.
Heya, post author here. I would say I’m not trying to be apologetic for the idea that software quality is diminishing, especially cause I became an indie developer because I care about quality.
When I say “it doesn’t matter” I mean more in an existential sense, and that people don’t seem to care. On the other hand people should do things because they care, which is why I personally still review the code that goes into my apps and spend the time to refactor and improve the stability and foundation rather than slopping like there’s no tomorrow.
Maybe I’m growing cynical but I understand why a business doesn’t care (at least until it comes back to bite them — which may take longer than some have assumed). And most of what you read about the subject is ultimately being driven by business needs of the desire of businesses.
Heya, post author here. I think I was just wrong about this assertion. I got into a discussion with a copyright lawyer over on Bluesky[^1] after I wrote this and came away reasonably convinced that this wouldn’t be a valid example of a clean room implementation.
[^1]: https://bsky.app/profile/mergesort.me/post/3mihhaliils2y
Hey there, author of the post here. I actually agree with this! That is in fact why I used the word maybe — my comment really was meant to be more speculative than definitive.
Hey there, author of the post here! I actually wrote this piece myself on my phone while I was out for a walk this morning. It was initially meant to be a quick note more than a full blog post —- whereas Coding As A Creative Expression took me a couple of days to write.
I made a commitment to write more this year and put my thoughts out quicker than I used to, so that’s likely the primary reason it’s not as deep of a piece of writing as the post you’re referencing. But I do want to note that this wasn’t written using AI, it just wasn’t intended to be as rich of a post.
The reason it came out longer is that I’ve honestly been thinking about these ideas for a while, and there is so much to say about this subject. I didn’t have any particular intention of hopping on a news cycle, but once I started writing the juices were flowing and I found myself coming up with five separate but interrelated thoughts around this story that I thought were worth sharing.
To be honest that's most of my pitch for Codex in the blog post. Codex works great without any configuration, and amazingly with. If you want to spend less time configuring then maybe Codex is the right agentic system for you.
I don't want to restate my thesis too much — but I really do believe it's worth experimenting with these tools every couple of months to see if the latest updates better match your preferences.
I've only skimmed it since I'm between Christmas and a longer vacation that starts in 24 hours, but this actually looks really neat! I'll definitely take a closer look to at these skills in depth — but this is exactly the kind of thing I've been telling people to take the time to invest in for their agentic environments. :)
You raise a couple of really good questions!
1. I find Claude Code's handling of the context window to be pretty poor, and one of the reasons why I use it for smaller things versus multi-hour coding sessions. I'm not sure what dark magic OpenAI has done to make their context window feel infinite, but Codex has become a better choice for that at the moment.
2. A small note on subagents but Claude Code did this right. Subagents are granted their own context window, so they don't spill over into your context window until they're done doing their own work — and the added context is relatively minimal. I'd love to see OpenAI adopt this pattern as well, especially in combination with something like Skills rather than leaning into MCP.
3. When I suggested adding skills, I mean ones that are far more complicated than your example, and can drive a chunk of work autonomously. The skill I use for writing in-app copy (which I'm bad at because you can see I'm never short for words) is about 100 lines long. It includes my style guide as an accessible resource, and a mostly complete history of my Bluesky posts to help achieve the authentic tone I when discussing Plinky. (I write all of my posts, so this really is my voice.)
These kinds of skills save me a lot of time as an indie developer! As I mentioned I have ones for data insights, fact-checking, and of course for code. My main suggestion would be to think through every step of your work and see if they can be automated, and then turn small pieces of that into skills.
—--
It's hard to assign a specific percentage to how much my effectiveness has improved, but it's a lot. The reason I don't want to put a number on it is that what I've gotten is a far broader set of skills (no pun intended) that allows me to execute in parallel. The metaphor I'd use to describe all this is to say that I'm no longer single-threaded.
I am a big believer that right now models work best for people who are effectively running small businesses — or teams that operate lean. The work of 10 can be done by 4-5 motivated and well-armed people, or an indie like me can do every facet of the work involved and do it well. I sit down and focus on explaining the big picture with great detail, and then set things off so I can do every part of the work involved in a round-robin style.
While an engineering task is going I'm off writing my newsletter with my words but with a skill that does meaningful research for me. While I'm running some research I'm in Figma working on social media assets. While I'm doing code review for my app's code I've got the server side building in the background.
Last week I had Codex finding a domain for me, with specific requirements. (Here's a simplified version of the prompt.)
I need a domain to represent this concept [+ 200 words], based on the code in this repository. [Code included so Codex really knows what the heck I'm building and talking about.] Don't show me any domains over $50/year at this registrar. Make sure it's a real word with no fun typos like tumblr.com is short for tumbler, and no compound words like "thisisfun.com". You can start with this list of tlds, but if you think there are any other ones that could be a good match then you can make a suggestion.
And after about 10 messages back and forth Codex found something that would have taken me far longer to research on my own — in parallel.
This all means that I'm able to write code, do marketing, design, support (which is always me and not AI), and run my business. If I plan well what I get is an extra set of hands to hand things off to, and most of the time (honestly) it does the work perfectly. But even for the times it doesn't, if it gets me 80-90% of the way there, that's a huge head start over where I would have been previously.
So the reason that I'm hesitant to answer this with a specific percentage is that your experience across organizations will vary. But I've seen in my work (solo engineering work, teaching, and consulting) is that the gains are pretty prosperous. That's true for roles where you're singularly focused on writing code — but the key is to lean into the strengths of this system and be creative about how you use it.
As I said — incapable of keeping my writing short so I hope that helps!
Thank you so much! That’s very kind of you to say. :)
Hey there, article author here! I spent a lot of time writing up a comment that talks through my process so I'll just lazy link it here if that's ok. [^1]
But the truth is I do build real production features all the time! That just wasn't the focus of this article. :)
That does sound like a good article! Sadly I've never written it and probably won't because I think a lot of that stuff doesn't provide as much value as people assume — which is why my personal conclusion in the post is to just lean into Codex.
This isn't a value judgment, it's just a question of where my priorities and tradeoffs lie. That said, I think Skills are the killer feature because they are a very composable tool — which I'll get to in a bit.
- Your CLAUDE.md should be a good high-level description with relevant details that you add to over time. Think of it as describing the lay of the land the way you would to a new coworker, including the little warts they need to know about before they waste hours on some known but confusing behavior.
- MCP has it's purposes, but it's not really a great tool for software development. It's best served for interfacing with a remote service (because it provides a discovery layer to LLMs on top of an API), but if you use them the way developers are told to, you're almost always better off using an equivalent CLI.
- I'll skip over agents, because an agent is basically a skill + a separate context window, and the main selling point is the context window bit. I think over time we'll see a separation of concerns where you can just spawn a skill with a context window and everyone will forget about the idea of agents in your codebase.
So now Skills. I wrote a well-received post [^1] a few months ago about Claude Skills, and why I think they are probably the most important of these tools. A skill is basically a plain-text description of an application.
The app can be something like I describe where Claude Code converts a YouTube video to an mp3 based on natural language, or you can have a code review skill, a linter skill, a security reviewer skill, and so on. This is what I meant when I said skills are composable.
You can imagine a team having lots of skills in their repo. One may guide an agentic system to build iOS projects well (away from an LLM's bad defaults when building with Xcode), skills that are very contextually relevant to the team, or even skills that enforce copy in your app to conform to your marketing team's language.
Skills are just markdown so they're very portable — and now available in Codex and many other places.[^2] (I had been using OpenSkills to great effect since the way Skills work is just through prompts). I now have a bunch of skills that do lots of things, for coding, marketing, data analysis, fact checking, copy-editing, and more. As a nice benefit they run in Claude — not just Claude Code. If you have ideas for processes you need to improve, I would invest my time and energy into building up Skills more than anything else.
[^1]: https://build.ms/2025/10/17/your-first-claude-skill [^2]: https://agentskills.io/specification
So I will state upfront that my current experience is not the most common team dynamic because I'm an indie developer [^1]. But I've worked at many companies — as small as 2 and as large as Twitter — so I am very familiar with the variety of engineering processes.
I can share how I work with agentic systems, because I (and now others) have found it to be very effective. I still have the engineering-like experience of thinking deeply — I've gotten great results across codebases small and large — and almost everyone who I've run a workshop with has come back to me and said that this was a missing piece for them when they work with agentic systems.
I'm the kind of person I alluded to at the end of my blog post when I wrote "Some people couldn’t start coding until they had a checklist of everything they needed to do to solve a problem.", so this description will be representative of that.
1. I start a document in Craft [^2] whenever I think of a great feature, and keep adding to that doc over the next few months whenever I have a new idea. I try to turn that document into something cohesive — imagine something like a PRD without the formality.
2. Then when it comes time to build the feature, I will just sit and write out a prompt (with lots of pointers to source code and relevant screenshots) that considers everything that needs to be built. I'll write out our goals for the feature, how the client should work, how the server should behave, the expected user experience, and anything else that's relevant. That process is really clarifying because it unearths a whole bunch of meaningful context — and context is exactly what a large language model needs!
3. Last but not least I'll simply add something like "Please ask any clarifying questions you may have, or for any additional details that you may find helpful". That leads to questions which I spend anywhere from another 5 to 30 minutes on, which fills in the gaps that I hadn't even considered to consider. And sure that may take time, but now the model has *so many useful details* that most people never add to their context window.
4. Once you have that, the model can act much more surgically than the experience most people have with agentic systems. Since it's so surgical I can go do something else like work on my newsletter, my AI workshops, or even go for a walk. This is why I much prefer to work this way, as opposed to the hands-on process I described Claude Code users [often] preferring in the blog post. (Which as I mentioned there is perfectly fine, just not my cup of tea anymore.)
---
I'd still like to touch on working with people though. I do quite a bit of open source work and there I still follow what people would consider standard processes and best practices. If I'm doing a week's worth of work I still don't want to dump a whole ton of code in one commit, so I'll break everything down into very atomic commits that spell out exactly what I'm doing. I also write lots of documentation, update references, and add tests like a person should.
But there's also nothing to say you have to generate a week's worth of code in one go. It's important to remember that you're in control of how you work. It may be more fitting to define smaller tasks (which will take less time for each independent step) and work on them serially, which you can then hand off to your coworkers one by one.
Ultimately my message is that people still need to exercise their best judgment and think for themselves. AI doesn't change what we've come to accept as best practices, it automates and accelerates them. In fact, the models keep getting better the more they are trained on our best practices, so my assertion is that success using AI seems to correlate well with autonomy, creativity, and critical thinking skills.
Anyhow, long answer for a short question — but I hope it helps! And if there's anything unclear: please ask any clarifying questions you may have, or for any additional details that you may find helpful.
[^1]: https://plinky.app [^2]: https://craft.do
Hey there, post author here. I do actually generate days worth of code in minutes and weeks worth of code in hours (not minutes) — but didn't really cover them because this post was more conceptual than specifically covering a tactic or technique as other posts do.
But in case you're interested a talk that I gave in Spain this year just went live yesterday, and discusses not only real world use cases but also discusses a lot of the fundamentals of how these AI systems work to make that possible.
https://www.youtube.com/playlist?list=PLztE34GS_piKKQ6y1dkku...
So I definitely understand where you're coming from, but let me provide a little bit of context.
The workshops are 3-4 hours and we do spend a lot of time discussing how things work in reality vs. how they work in the context of the workshop. It's worth noting that these workshops span the gamut of non-technical people in sales to seasoned developers, so a lot of people simply won't learn much (or have the excitement to learn on their own) if we spend the first 2-3 hours setting things up.
In my experience the heaviest lift for teaching practically any technical subject is getting someone interested by showing them how to accomplish something they care about, and then leaving them with lots of information and resources so they can continue experimenting and growing even once we're done. The way I do that is to make sure they leave the workshop having built their own idea — without taking shortcuts!
Being able to use Codex to accomplish something because you spent an hour crafting a good prompt isn't cheating, it's learning the skill of becoming a better technical communicator — in a short period of time — and taking advantage of the skill you've just learned. I don't consider that magic, it's actually the core tenant of building with AI, and is very much how I work with AI every day.
I'm late for dinner so I should probably stop here, so I'll leave just one final note. After every workshop I send each student a list of personalized resources that will help them continue on their journey by demystifying things that we may have glossed over or weren't clear in the workshop — so they should be armed with the tools to take their next steps away from any magical thinking.
It's a bit hard to boil down exactly what I do and how I try to design for best hands-on pedagogical practices in an HN post I'm writing on the go — but I am absolutely open to your thoughts! :)
Hey Jose, author here! That's a great call out. I write predominantly in Swift and for a long time Claude was the only usable option. But sometime around GPT-5 OpenAI's models got much better at Swift, so the choice started becoming more about aesthetics (as a descriptor of preferences). So you're right — if the model can't write coherent code then it doesn't matter what kind of flow you feel as you're working with the tools — but I do imagine this will continue to improve for all languages including Elixir.
Heya, author of the post here. That's a good call out because it's probably a lot!
And now that you mention it, that's also one failure case for why some people look at AI and go "this just isn't very good at coding". I'm not saying it has to be that way nor will it be that way forever, but there are absolutely a lot of people who just download Claude Code or Cursor or Codex and dive right in without any additional set up.
That's partially why I suggest people use Codex for the workshops I offer, because it provides the best results with no set up. All of these tools have a nearly unending amount of progressive disclosure because there's so much invisible configuration and best practices are changing so fast. I'm still trying not to imply that one tool is "better" than another (even if I have my preference), but more so hit on the fact that which AI tools people like is mostly about your preferred set of tradeoffs.
No problem at all! I read it as a bit pithy, but I didn’t think it was particularly mean spirited.
If you check out my writing on build.ms and fabisevi.ch you’ll see that the majority of it is meant to be evergreen observations of a concept or a moment in time. My goal is to make people think and to think about thinking, more than it is to tell people what exactly to think.
If I had to summarize my style in one sentence, it would walking people to and around an idea, and leaving the rest as an exercise to the reader. Naturally, this means I have less control over how people interpret my writing so I do try and cover my bases with fact and experience, but that still means sometimes I won’t deliver a complete picture to everyone.
In that case, sometimes I come to a place like HN or Bluesky or Mastodon where my post is being discussed and try add some perspective and clarity through constructive conversation. :)
If I’m being honest, I think we’re too early in the state of generative AI as a coding tool to draw very strong factual conclusions for many of our experiences using AI to code that will hold up well. I’m not implying it’s all vibes, but I think it would be pretty hard to wrap up my post in a bow the way you’re suggesting. On the other hand I’m always open to well-considered feedback — and would love to know more about your experience if you’re interested in sharing!
That’s a long way of saying happy holidays to you as well!
Heya, author here! I do agree with you that this is a big downside, but I don’t know if this is the primary reason.
In my experience teaching people, most people don’t actually know much at the time they make this decision. They’ve heard about Cursor, they’ve heard of Claude Code, and they may have heard about Codex. But what they’ve heard is anecdotes and marketing — they don’t yet have hands-on experience.
They make a big choice and then assume that this is how all AI works, because they don’t have a full breadth of context yet. And that’s to be expected! That’s how most things work.
That is why I teach the workshops I do to make AI accessible, so people can walk through the tradeoffs and make the best educated choices for them.
A couple of comments here have said that the post is subtly pro-Codex, but I tried to make my point very explicit: people should try a lot of things and see what works best for them. But it’s very hard to do that without investing a lot of time because the market is so nascent and moving so fast. This post exists to try and nudge people into exploring more of the tools they haven’t tried yet, so they can make their own informed decisions like you have. :)
All that’s to say, people definitely hit limits with Claude Code (as I have done myself) — especially if they’re hesitant to upgrade to Claude Max because they haven’t gotten enough out of Claude Pro. But I think the real reason people make the choices they do starts earlier in the process, even before they get a lot of hands on experience with Claude Code or Codex.
Heya, I’m the author of the post! To be clear I have AI write probably 95% of my code these days, but I review every line of code that AI writes to make sure it meets my high standards. The same rules I’ve always had still apply — to quote @simonw “your job is to deliver code you have proven to work”.
So while I’m enthusiastic about AI writing my code in the literal sense, it’s still my code to understand and maintain. If I can’t do that then I work with AI to understand what was written — and if I can’t then I’ll often give it another go with another approach altogether so I can generate something I can understand. (Most of the time working together to understand the code works better, because I love to learn and am always open to pushing my boundaries to grow — and this process can tuned well to self-directed learning.)
And to quote a recent audit: “this is probably one of the cleanest codebases I’ve ever audited.” I say that emphasize the fact that I care a lot about the code that goes into my codebase, and I’m not interested in building layers of unchecked AI slop for code that goes into my apps.
As the author of the post I think it was a nice quick post to share my perspective of a behavior I’ve been seeing across many (but not all) developers recently, but I’m always open to feedback for how to improve my writing!
And as I mentioned here (https://news.ycombinator.com/item?id=46392900) I have no affiliation with any of the organizations, nor care to evangelize any of them. Nobody pays me to write, I’m just a guy on the internet sharing his thoughts, building software, and teaching people how to use AI better with any tool people want to use. :)
Heya, author of the post here! I think you're right in everything you've said, but I want to note that the programming language comparison was meant to be metaphorical more than literal. Everything is changing so fast (as I mention in the post a few times), but I have seen some (far from all) people get locked into Claude Code or Codex in a way where they won't even consider alternatives the same way people they chose Ruby to start their career and now identify as Ruby developers.
My goal was to open people's minds just a little bit by saying exactly what you're getting at — everything is moving fast and we should be reassessing often. A meaningful difference is that you can start a codebase with Claude Code and then switch to Codex with almost no friction, while you can't just migrate a TypeScript app to Python in 15 minutes.
All that's to say, we agree!
Heya, author here! I completely agree with you — and why the post is titled Codex vs. Claude Code (Today). I also have this very specific disclaimer in the second paragraph to note that this post is a reflection of a moment in time. :D
Before we continue, I need to make a disclaimer: This post is about the Claude Code and Codex, on December 22, 2025. Everything in AI changes so fast that I have almost no expectations about the validity of these statements in a year, or probably even 3-6 months from now.
That said I do what you do and try different models when I want to see if things have changed. I run my own private little benchmarks with a few complex real world tasks, and I really love seeing how things are progressing — both in terms of quality but also the novel quirks that are introduced, changed, or removed. :)
Heya, author here!
I'll try to answer these one by one, but I will just note that a lot of my prompts are domain specific so it's hard to share those.
- I don't use any plans — my writing is the plan. The Plan Mode in Claude Code is excellent, but as I've switched to Codex (which doesn't have one) I will simply write up a nice long prompt and then add "Please ask any clarifying questions you may have, or for any additional details that you need" — and it works great! I may go back and forth for anywhere from 5-30 minutes depending on what else is needed, but that's basically the experience of using Plan Mode in Claude Code too.
- I've built quite a few recent features for my app Plinky [^1]. I've made a few meaningful contributions to my open source project Boutique [^2] (and have been having AI asynchronously sketch out a large new database relationships feature). I built my new blog and my workshops pages [^3] with Codex as well. Truth is I do practically everything in Codex and Claude Code these days, so I'd have more trouble listing what I haven't built lately.
- Plinky's upcoming Reader Mode is a good example of a prompt that took me two hours, but the feature isn't yet in the app so I'd prefer not to share the prompt. But I can share the first draft of the prompt for Boutique's relationships feature sine that's open source. [^4] I've been experimenting with using ChatGPT Pulse to make progress on it every day (simply by asking it to!), and much to my surprise it's been designing a new API day by day in a way that's far from perfect but certaintly has been very interesting.
The honest truth is that this one did not take two hours and I wrote it on the bus so it's probably not perfect, but the descriptive process is effectively the same. For a feature like Reader Mode you would have to capture more details to scale up to the additional complexity of a domain-specific feature with client and server components, a new download queueing pipeline, amongst other abstractions.
Hope that answers your questions!
[^1]: https://plinky.app [^2]: https://github.com/mergesort/Boutique [^3]: https://build.ms [^4]: https://gist.github.com/mergesort/04a77c47ea4cb6433aa9ade4e1...
Heya, author here! That's a great question! I fully understand the vendor lock-in concern, but I'll just quickly note that when it comes to a first workshop I do whatever makes the person most comfortable. I let the attendee choose the tool they want — with a slight nudge towards Codex or Claude Code for reasons I'll mention below. But if they want to do the workshop in Cursor, VS Code, or heck MS Paint — I'll try to find a way to make it work as long as it means they're learning.
I actually started teaching these workshops by using Cursor, but found that it fell short for a few reasons.
Note: The way that my workshops work is that you have three hours to build something real. It may be scoped down like a single feature or a small app or a high quality prototype, but you'll walk away with what you wanted to build. More importantly you'll have learned the fundamentals of working with AI in the process, so you can continue this on your own and see meaningful results. We go through various exercises to really understand good prompting (since everyone thinks they're good but they rarely are), how to build context for models, and explore the landscape of tools that you can use to get better results. A lot of that time is actually spent in a Google Doc that I've prepped with resources — and the work we do there makes the code practically write itself by the time we're done.
Here's a short list of why I don't default to Cursor:
1. As I noted in another comment, the model performance is just so much better [^1] when accessed directly through Codex and Claude Code, which means more promising results more quickly. Previously the workshops were 3-4 hours just to finish, now it's a solid 3 with time to ask questions afterwards. You can't beat this experience, because it gives the student more time to pause and ask questions, seep in what they've done, and not spend time trying to understand the tools just to see results. 1a. The amount of time it took someone to set up Cursor was pretty long. The process for getting a good set up is pretty long — especially for someone non-technical. This may not be as big of a deal for developers using Cursor — but even they don't know a lot of the settings and tweaks to make to get Cursor to be great out the box.
2. The user experience of dropping a prompt into Codex/Claude Code and watch it start solving a problem is pretty amazing. I love GUIs — I spend my days building one [^3], but the TUI melting away everything to just being chat is an advantage when you have no mental model for how this stuff works.
3. As I said in #1, the results are just better. That's really the main reason! I
Not to toot my own horn, but the process works. These are all testimonials in the words of people who have attended a workshop, and I'm very proud of how people not only learn during the workshop but how it sets them off on a good path afterwards. [^2]. I have people messaging me 24 hours later telling me that they built an app their partner has wanted for years, to tell me that they've completed the app we started and it does everything they dreamed of, and hear more process over the weeks and months after because I urge them to keep sending me their AI wins. (It's truly amazing how much they grow, and I now have attendees teaching ME things — the ultimate dream of being a teacher knowing you gave them the nudge they needed.)
Hope that helps and isn't too much of an ad — I really just want to make it clear that I try to do what works best and if the best way to help people learn changes I will gladly change how I work. :)
[^1] https://news.ycombinator.com/item?id=46393001 [^2]: https://build.ms/ai#testimonials [^3]: https://plinky.app
Heya, I'm the author of the post and I just wanted to say I do appreciate the configurability! As I mentioned in the post, I have been that kind of developer in the past.
This is a perfect match for engineers who love configuring their environments. I can’t tell you how many full days of my life I’ve lost trying out new Xcode features or researching VS Code extensions that in practice make me 0.05% more productive.
And I tried to be pretty explicit about the idea that this is a very personal choice.
Personally — and I do emphasize this is a personal decision — I‘d rather write a well-spec’d plan and go do something else for 15 minutes. Claude’s Plan Mode is exceptional, and that‘s why so many people fall in love with Claude once they try it.2
For every person who feels like me today, there's someone who feels like you out there. And for every person who feels like you, there's someone like me (today) who finds it not as valuable to their workflow. That's the reason my conclusion was all about getting folks to try out both to see what works for them — because people change and it's worth finding out who you really at this moment in time.
Anyhow, I do think that Codex is also very configurable — I was just trying to emphasize that it's really great out the box while Claude Code requires more tuning. But that tuning makes it more personal, which as you mention is a huge plus! As I've touched on in a few posts [^1] [^2] Skills are to me a big deal, because they allow people to achieve high levels of customization without having to be the kind of developer that devotes a lot of time to creating their perfect set up. (Now supported in both Claude Code and Codex.)
I don't want this to turn into a bit of a ramble so I'll just say that I agree with you — but also there's a lot of nuance here because we're all having very personal coding experiences with AI — so it may not entirely sound like I agree with you. :)
Would love to hear more about your specific customizations, to make sure that I'm not missing out on anything valuable. :D
[1] https://build.ms/2025/10/17/your-first-claude-skill/ [2]: https://build.ms/2025/12/1/scribblenauts-for-software/
Heya, I'm the author of the post! This was probably unintentional but I think you're making a really valuable observation that will be helpful to others.
The models Cursor provides to use in their product are intermediated versions of models that companies like OpenAI and Anthropic offer. They are technically using Codex, but not in the way that they would be if you were in a tool like Codex (CLI) or Claude Code.
If you ask Cursor to solve a tough problem, Cursor will break down the problem into a different problem before sending that request to OpenAI so they can use Codex. They do this because: 1. To save money. By restructuring the prompt they can use less tokens, saving them money for running Cursor since they are the ones paying for the tokens with your subscription cost. 2. [Based on things the Cursor team has said] They believe they can construct a better intermediate prompt that is more representative of the problem you want to solve.
This extra level of abstraction means that you are not getting the best results when you use a tool like Cursor. OpenAI and Anthropic are running their harnesses Codex CLI and Claude Code at a loss (because VC), but providing better results. This is not the best way to make money, but it's a great way to build mindshare and hopefully get customers for life. (People are fickle and cheap though so I doubt this is a customers for life strategy the way people buy the same brand of deodorant once they start buying Dove.)
Happy to answer any questions you may have, but mostly I would highly suggest trying out Codex CLI and Claude Code to get a better feel for what I'm saying — and to also to get more out of your AI tools. :)