HN user

ianberdin

519 karma

https://x.com/ianberdin

Founder of https://playcode.io - ai app builder.

Posts43
Comments181
View on HN
playcode.io 4d ago

Pelican on a bicycle is a good benchmark, right?

ianberdin
2pts0
playcode.io 9d ago

The pelican benchmark is doubtful, let's draw MacBook Pro in SVG. Is Fable best?

ianberdin
3pts0
playcode.io 9d ago

The real prices of frontier models

ianberdin
157pts83
playcode.io 9d ago

Fable 5 on Playcode. As well as Sol, Grok 4.5 and GLM 5.2

ianberdin
3pts0
playcode.io 9d ago

Playcode Cloud – Firecracker backed and Neon inspired

ianberdin
3pts0
playcode.io 4mo ago

I've built a better Lovable clone alone

ianberdin
2pts4
getpartner.ai 6mo ago

Show HN: Partner – An AI co-founder that remembers you

ianberdin
1pts0
playcode.io 6mo ago

Show HN: I've Built a Python Playground

ianberdin
2pts0
playcode.io 7mo ago

Show HN: I've built the best online Python Compiler

ianberdin
1pts0
playcode.io 7mo ago

Show HN: I've built the best Python Playground

ianberdin
2pts0
playcode.io 7mo ago

Show HN: I've built a best JSON formatter app

ianberdin
1pts1
playcode.io 7mo ago

Show HN: 3 yrs later, my JS sandbox has 11M users and an AI agent

ianberdin
1pts1
boing.playcode.io 7mo ago

Opus 4.5 is 2x cheaper and 2x better relative to Sonnet in reality. A quick demo

ianberdin
2pts6
boing.playcode.io 7mo ago

Show HN: Boing #2

ianberdin
4pts2
playcode.io 7mo ago

Show HN: I built a browser-based Cursor alternative as a solo dev

ianberdin
1pts0
playcode.io 7mo ago

Show HN: I've built a Cursor alternative in browser. AI Coding Agent.

ianberdin
6pts0
news.ycombinator.com 8mo ago

Zig is so cool, C is cooler, but JavaScript everywhere

ianberdin
4pts1
twitter.com 8mo ago

I spent $1,800/month on Cursor for a year to rewrite playcode.io from scratch

ianberdin
1pts0
twitter.com 9mo ago

Vite + Rolldown. 78 seconds to 3.5 seconds build time reduction!

ianberdin
1pts0
news.ycombinator.com 10mo ago

Flow state is the best pleasure in life. How to keep it sustainable?

ianberdin
1pts0
twitter.com 10mo ago

I need two addictions, not one to fight burnout

ianberdin
1pts1
news.ycombinator.com 11mo ago

I've spent 4 months and $800/mo AI bill on Cursor, Claude Code. Later is better?

ianberdin
3pts0
twitter.com 11mo ago

I've spent 4 months with $800/mo AI bill on Cursor, Claude Code

ianberdin
2pts0
twitter.com 11mo ago

Never update to the beta version of iOS. It's a damn nightmare

ianberdin
2pts0
twitter.com 12mo ago

Why I Refer to Myself as "We"

ianberdin
6pts0
news.ycombinator.com 1y ago

I've been fighting with burnout for 18 years

ianberdin
7pts8
twitter.com 1y ago

Sounds weird, but how to stop working?

ianberdin
2pts0
twitter.com 1y ago

Pathological Demand Avoidance (PDA): When Every Task Feels Like a Threat

ianberdin
5pts4
nypost.com 1y ago

Nvidia stocks drops 16% because of DeepSeek R1 breakthrough

ianberdin
4pts0
github.com 2y ago

Express-zod-API – Zod based Node.js framework

ianberdin
2pts0

Thank you for the criticism. I heard you. I added a TLDR. I cleaned up many AI constructions. By the way, I tweaked it a bit, compressed it.

One way or another, I want to note that yes, this text was made in collaboration with AI. My English is non-native. It helps me translate, helps me structure better. Yes, there is a downside, it can bloat the text with unnecessary words. But that, unfortunately, is the price.

But the key thing is that I tried very hard to share my many years of experience, or rather a part of it, which I acquired, with all of you. And I am very glad that this information turned out to be useful to you.

The key here is: * The information that is written in the article. * Not how it is written, but what I was trying to convey to you.

Thank you very much for reading and responding.

Well, criticizing is, of course, great. But the reality is that English is not my native language and I dictated most of it with my voice, then processed it with the help of AI, translated, added, corrected, and converted.

It is actually a big result of work, a lot of research and attempts. And to just say that "oh, this is AI-slop," I consider unfair, but that is your choice.

There is a difference: - There are people who do, - And there are those who criticize.

We've switched the default model in playcode.io among Opus 4.8, Opus 4.6, Sonnet 4.6, and Sonnet 5. I must admit, Opus 4.8 is quite expensive, and the costs accumulate quickly. Opus 4.6 is about 50% cheaper, while Sonnet 5 is significantly more affordable. According to the data, Sonnet 5 is about 2-3 times cheaper. Fable 5 is unaffordable at all...

Today, I tested Sol 5.6 on various tasks. It performs similarly to Opus 4.8 but is still noticeably more expensive than Sonnet 5. Although Sonnet 5 isn't the top model, it's quite effective for creating typical websites for small and medium businesses. However, they will increase the price starting September 1, as their free offer is ending.

I'm also actively testing Grok 4.5. There's something promising about it. The design is mediocre, in my opinion, but it operates quickly and reliably without any deadloops. Usually, Grok models would fail or loop, but this one is stable.

Overall, I really want a benchmark based on real tasks.

Well, I both agree and disagree with you.

On one hand, the price is just astronomical for Fable, well, not exactly astronomical, but I would say unaffordable. That is to say, so expensive that it is impossible to use.

But on the other hand, Fable is simply incomparable to anything else. I mean, it is just amazing. There is nothing even close to being equal to it.

Well, in my view, it's just the most ordinary manipulation to avoid creating unrest. There is most likely no improvement inside.

Of course, these are my guesses, but did anyone feel the difference in the transition from Opus 4.5 to 4.6? In my opinion, no. And it's unlikely to be a matter of the tokenizer.

I have a large monorepo that includes about 15 TypeScript services and many Rust services. Everything is well-documented and organized, with standardized and structured custom code.

When an issue arises, I often test the systems by providing a minimal prompt, like: "this user, this is their email, this isn't working, figure it out in production." I send this to both Opus and ChatGPT, but it doesn't help. I've set up Agents.md and Quote.md identically, with the same access and linkers, so the Harness is consistent.

ChatGPT rarely succeeds. If the task is complex and requires a multi-step process to identify the true cause, ChatGPT usually stops after a few initial ideas and wrongly claims it has found the solution.

- For simple tasks, like identifying a missing item in a to-do list, ChatGPT performs well. - However, for issues like memory leaks or file system corruption, it struggles.

On the other hand, Opus 4.8 always finds the solution, albeit slowly. I can rely on it without worrying about whether it will succeed. It just gets the job done.

Recently, Fable 5 has emerged, which resolves issues without needing any prompts. It operates even faster than Opus.

When I ask ChatGPT or Opus to create a new feature: - ChatGPT often produces superficial results, ignoring existing code and building unnecessary independent code. - Interestingly, the outcome from ChatGPT appears functional, but it's usually incorrect, focusing on a superficial "aha!" moment.

Opus, however, plans thoroughly, executes, and cleans up, ensuring everything works correctly. If needed, I can provide more realistic examples, though it's challenging due to the monorepo's size and complexity, with hundreds of thousands of lines of code.

Everyone probably has the same question: what about Fable? Fable 5 is sick.

It is simply the best model in the world out of everything we have ever tried. It's absolutely fantastic. It solves almost any task from start to finish, the way it should be done — no errors, perfect code. It's a miracle.

If there's any way to make it a little more affordable, that would be incredible.

As for GPT-5.6 Sol — it doesn't even come close. I honestly don't understand why people even try to compare them. It feels like Sam's attempt to hold onto his audience with those endless daily limit resets. A clever trick, nothing more.

We at Playcode.io - a company similar to Ploy are still using Opus 4.6. "Why?" you might ask.

Because GPT 5.6 Sol, while fast and pleasant to use, is essentially the same model as 5.5 wrapped in new marketing packaging, just to avoid losing ground to Anthropic. In practice, it's the same quality: it generates the same garbage, tons of code, and can never solve even a single complex task. We simply don't trust it to write code for clients that they'll end up throwing away anyway.

"Then why not Opus 4.8?" you might ask.

Well, because Opus 4.8 and 4.7 are just another lie, a price hike with no actual quality improvement.

That's why at Playcode, we give our clients the best possible quality/price - which is Opus 4.6. Regardless of what people write in articles like this.

I also want to point out that all of this sounds fun and great — until the load kicks in and usage starts to grow.

That's when you start seeing:

- *Rate limits* from object storage - *Dropped packets* - *Hanging S3 requests* - *Overloaded NVMe drives* — because it turns out they're nowhere near as fast as they seem

For example, we recently discovered that the read speed is 5 GB/s, but the *average write speed is only 400 MB/s* — not several gigabytes as expected. Surprise! Who would have thought that Bare Metal could ship such underwhelming drives?

And then there's the CPU — which is also easy to kill, for instance, if you're compressing chunks. And so on, and so on.

A lot of things surface once you're in *production usage*. On paper, of course, everything looked much simpler.

Surprisingly, I'm already encountering a second solution that involves storing data chunks on S3 — and this is all within the same week.

This is becoming popular. At Playcode, we built what we believe is a revolutionary file system for our Playcode Cloud (https://playcode.io/cloud), which enables the creation of full-stack web software. The FS built completely from scratch using Rust. We thought we were the smartest ones around and that nobody else had figured this out. But it turns out Databricks, Neon, and several others have as well.

The idea behind a *Bottomless File System* is really cool, and it works very well for us. Essentially, as described here:

- There is a *page server* - A *Linux file system* split into chunks (let's call them chunks instead of pages) - A *cache on NVMe* - And of course, *object storage*, where everything is asynchronously synchronized

It works quite well, though it has its downsides.

One clear advantage is that NVMe drives have become expensive lately, while object storage remains cheap — so the benefits are undeniable. That said, latency is also a factor.

On top of that, uplink costs are rising. To run an object storage-backed file system, you need a very strong uplink with consistent speed — 1 Gbps is simply not enough. Ideally, you want *5 to 10 Gbps*, depending on the load.

We spend a lot of time optimizing and experimenting with different hosting providers — specifically bare metal hardware. The main challenges are:

- *Slow disks* - *Slow uplink* - And as it turns out, *object storage can be unreliable* — unless you're using S3

But AWS hardware is expensive, so nothing in life is ever that simple.

Claude Sonnet 5 22 days ago

Anthropic outsmarted everyone again.

They released Sonnet 5 with a temporary price reduction until August. Everyone was excited, but in reality, they increased the tokenizer size by 50%. As a result, the actual cost went up by 50%, they shifted everyone's attention to decrease.

Thus, Anthropic is raising prices but not telling anyone about it. Nobody is really aware of it. You go to the pricing page, the price looks the same. Yet people are actually paying 50% more.

Very shady marketing.

And of course they lie about 35% again. In reality with coding it is 50%.

UPD: I run playcode.io, so it’s my job test all models, their pricing, quality in order to provide best price/quality/speedy/reliability to non-techy.

I say hello to every happy HN and twitter post “we moved to hetzner and saved 10X”.

I told ya about silent happiness…

Thanks for the read. It is a bit more complicated than you think. I completely rebuilt this sync engine + orm with relations, lazy loading etc, using Vue + Pinia in https://playcode.io. Google Linear’s videos, they explained in detail their architecture.

Yes, I spent a few months. But it worth it. Every new field, model I need to add, it is so straightforward. I do love frameworks and foundations. They make live easier later by a lot.

When playing busy Dota 2 (realtime game), it was crashing sometimes. I asked Claude Code any advice (without any hope) and it debugged somehow that I have unstable IP address and a rented VPS server will improve my connection. I could not believe, it worked…

The real problem is this: no cheap model right now produces a genuinely beautiful, usable UI when it comes to website building. Not one.

And here’s the core tension. The models keep getting better. GPT 5.5 improved. But it also got more expensive. Opus 4.7 to 4.8 has become outrageously priced too, up 50%, and 4.6 was already brutally expensive to begin with. API pricing is a real pain.

What’s missing is any meaningful supply of affordable, democratically priced models you can actually embed into your own service. For me that’s playcode.io, whether it’s the website builder or the app builder. The moment we give users access to these models, the cost becomes a serious blocker. There’s no way around it.

The same dynamic explains Cursor. Why did they go build their own Composer 2.5 model? Because relying on third-party models is simply too expensive for users unless they’re carrying a Claude Code or Codex subscription. So Cursor had to roll their own. It’s a real mess, honestly.

And Chinese models don’t close the gap either. They’ve improved, the free-tier ones especially, which is great to see. But the limitations are significant:

• No multimodality. They don’t accept image input. • You can’t attach a screenshot, show a UI, or hand it a PDF. • They feel heavily stripped down overall. • They’re just not polished. Not even close.

Opus, by contrast, feels like a finished, deeply refined product. Everything else is still rough around the edges. And that’s exactly why Anthropic can charge what they charge: because they actually deliver. That’s the whole problem in a sentence.

Don’t get me wrong, but 7K LoCs means it is still an early attempt to make a coding agent. It starts easy “ah it can edit and read files!”, but it requires a lot of extra effort to make properly for many edge cases, especially caching, price optimizations, etc.

I’ve been implementing custom coding agent in https://playcode.io for 3 years already. Far beyond of 7K LoCs.

So when you compare to “shitty slow” Claude code - I don’t agree.