HN user

bluelightning2k

1,592 karma
Posts29
Comments467
View on HN
news.ycombinator.com 13d ago

Tell HN: GPT5.6 Is Imminent?

bluelightning2k
1pts2
news.ycombinator.com 7mo ago

Ask HN: How much more Antigravity do you get on paid plans?

bluelightning2k
2pts0
news.ycombinator.com 8mo ago

Ask HN: Why is it OK for Cursor and Windsurf not to credit their model

bluelightning2k
1pts4
news.ycombinator.com 1y ago

Ask HN: Why do Cursor, Windsurf and Claude Code dominate the conversation?

bluelightning2k
28pts38
news.ycombinator.com 1y ago

Tell HN: Windsurf edits are failing and burning infinite credits for most users

bluelightning2k
13pts5
moddable.app 2y ago

Show HN: mod-support for extendable business software

bluelightning2k
1pts0
areweturboyet.com 2y ago

Turbopack looks like it's about to go stable?

bluelightning2k
1pts0
news.ycombinator.com 2y ago

Ask HN: How do you debug prod data in dev?

bluelightning2k
1pts3
docs.getrerun.com 2y ago

Speedrunning a startup in 3 days for YC's challenge

bluelightning2k
1pts0
news.ycombinator.com 2y ago

Ask HN: Does YC read all applications?

bluelightning2k
2pts3
sellsitself.substack.com 2y ago

Winning a big new Blue Ocean market

bluelightning2k
14pts3
news.ycombinator.com 3y ago

Ask HN: Co founders getting part time engineering job?

bluelightning2k
1pts2
chatspot.ai 3y ago

ChatSpot – HubSpot's GPT Chatbot

bluelightning2k
4pts0
demotime.com 3y ago

Algorithm edits every sales demo into a highlight-reel video

bluelightning2k
59pts48
sellsitself.substack.com 3y ago

The golden rule of SaaS: Don't Be Boring

bluelightning2k
4pts0
news.ycombinator.com 3y ago

Ask HN: How did you meet your VC

bluelightning2k
5pts3
sellsitself.substack.com 3y ago

Making sofware user-extendable (adding features on the fly to popular tools)

bluelightning2k
2pts1
news.ycombinator.com 3y ago

Ask HN: Can I edit your sales demos into highlight-reel videos (for free)?

bluelightning2k
2pts0
news.ycombinator.com 3y ago

Ask HN: Edge-Cases in Software Demos

bluelightning2k
1pts1
news.ycombinator.com 3y ago

Ask HN: Building a Hard Product

bluelightning2k
2pts1
news.ycombinator.com 3y ago

Tell HN: Russian antivirus flags NPM package as malicious for logged message

bluelightning2k
2pts1
news.ycombinator.com 3y ago

Ask HN: Better Storytelling for Tech Demos?

bluelightning2k
12pts9
timetravel.dev 3y ago

Show HN:A time-travel debugger for JavaScript/TypeScript

bluelightning2k
2pts0
news.ycombinator.com 3y ago

Ask HN: Advice after a decade trying to get into YC

bluelightning2k
35pts30
news.ycombinator.com 4y ago

Ask HN: What to do with non-commercialized tech?

bluelightning2k
2pts1
news.ycombinator.com 4y ago

Ask HN: How to create a webpage like Remix?

bluelightning2k
1pts1
news.ycombinator.com 4y ago

GitHub Copilot is a brilliant piece of **product management

bluelightning2k
1pts0
medium.com 4y ago

Going solo to a conference. The kindness of strangers

bluelightning2k
2pts0
demotime.com 4y ago

Show HN: Best SaaS demo follow-up (automatic “highlight reel” summary)

bluelightning2k
1pts0

Why is it hard? Ultimately you take whatever your signal is and send it to some relatively cheap LLM.

How is it easier to sign up and manage a different service, implement a different API, etc.

And from the company side the fatal flaw is that these types of tools rely upon 1% of their users having huge spend. Nobody is going to be a huge spender here because it's easier to hand roll than navigate procurement on this (not to mention impossible to justify the spend, additional security/privacy risk, etc.)

It feels approximately impossible for this company to have large accounts.

Must admit, for this particular case I don't see the appeal in using a wrapper.

Why would I not just use Codex directly?

The we wrote a bunch of prompts argument is kind of meh. That sort of thing has not only diminishing value with subsequent model releases but I actually believe will turn negative. The model will know better by default.

For example, initially giving the model some advice on code best practice was helpful. But now it's unhelpful because the model already knows best.

Agreed!

Example: for large Eloqua/Marketo/HubSpot emails we would previously make a planner which delegated the sections to their own call.

GPT5.6 can do the whole thing. The planner is unhelpful.

My suggestion: feature flag your complex implementations so you can rapidly contrast with and without it. (Or a formal eval suite if you have one).

If you prefer the simpler path, delete the old path.

Note: the challenge is making things compatible with these tools. Obviously generating html directly has been simple for ages.

(Source: mopsy.ai)

Good observation.

I actually started typing the same point that the chances are actually high because of train/eval overlap then realised you answered your own question with that same observation.

It is interesting though!

Perhaps in some way this means we should decide which eval set aligns best with our taste?

Back to the blog post. This is an excellent write up of an excellent technical achievement.

I have a lot of respect for the Cognition/Devin (always "Windsurf" to me) and Cursor teams.

I found it interesting - but justified - that they referred to themselves as a foundation lab rather than a dev tools company.

No. It makes you faster. You can prove it by looking at the lines of code.

Source: Gary Tan, who can write more code than Jeff Dean and John Carmack.

If you want to go really fast, you put a Ralph Wiggum dungeon. It's where you orchestrate a team of Ralph Wiggum loops together using subagents to win the economic Darwin Award which companies are handing out for who can burn the most tokens.

Claude Fable 5 1 month ago

Congratulations to Anthropic for solving safety on Mythos exactly when the SpaceX compute came online. Nice how that lined up for them.

Claude Fable 5 1 month ago

To hide the severity of the price increase, the plan is to move everyone right one model.

Haiku = essentially phased out Sonnet = the Haiku use cases Opus = the new Sonnet class Fable = the new Opus class

If I am right, the other "5.0" models will be conspicuously absent, possibly even for a couple of months. (If Opus 5 follows soon and is even modestly better than 4.8 then I was wrong.)

Perhaps "potentially billions" is justifiable? In that it's a "change the scenario entirely" level of change.

Certainly if you compare it to another likely scenario where Vercel buys them and fast forward 2 years, it's plausible that a huge number of projects went one way or the other because of what the AIs defaulted to.

The reason this is worth it to CloudFlare is it will cause AI to recommend them more.

The agents already reach for Vite. When they reach for Vite it's very logical they will default to CloudFlare after. (Much like they will guide users to setup Vercel for NextJS).

This could be a $20m acquisition which will generate $billions from the increase in the agent equivalent of SEO.

Your linked post shows they raised 12.5 million?

But it's also possible they haven't spent much of that money.

The investors don't need to be happy. They just need to be made whole (assuming they have a minority control).

It could literally be that only $2m ever got spent and that's been paid back.

It could also be that when literally nobody said they would pay for Vite+ the investors and team in general lost confidence and were actually very happy just to get their money back and pivot into this acquisition.

Vite is great. Vite+ seemed to offer very little added value and was paid for (or was going to be one day?).

The article didn't mention what happens to paying Vite+ users. Is that because there basically aren't any?

DaVinci Resolve 21 2 months ago

Respectfully no, thank you.

I know a lot of people are/will build this. I would be specifically interested in Black Magic doing it first party.

DaVinci Resolve 21 2 months ago

So much respect for Black Magic. They are absolutely World Class and their business model is extremely generous.

Having said that, for all the AI features, the big one would be setting key frames etc. with an agent, driving the general editing workflow with text,etc. I realize this is non trivial but it's certainly viable for a team of this calibre.

I think if BM added a paid for agent which helped execute their traditional video editing tools (even if it "only" supported a subset) then that's a subscription a lot of people would be willing to pay for, especially as their core tool is so generous.

This is an absolute joke.

Anthropic capitalized upon a brief window of being more code-focused, which turned into enterprize contracts.

Then on renewal rug-pulled those same enterprises - going from your seat includes all the usage a user would reasonably need, to being you pay for the seat + all tokens at API pricing. (Which they raised by how many times in a year? I don't know the actual number.)

Revenue spikes like crazy through basically hostage taking made possible by Sonnet 3.5 era sentiment + enterprise purchasing lag.

Parlay the revenue spike into the valuation.

Crazy. Those same enterprises will get sticker shock and leave. Absurd short-term thinking.

OpenAI is the better company (transparency, open sourcing things, how they handle things in general e.g. OpenClaw, how they compete, etc.) and they have the vastly better brand, the better consumer presence, and (for me and many others) they have the better coding app + models.

Anthropic doing deeply customer hostile stuff - again and again - to produce a short term revenue spike does NOT make for a long-term sustainable business.

For such a young business to have such a long history of bait-and-switch is absolutely crazy. (Raising prices repeatedly, lowering rate-limits repeatedly, changing the terms, banning calls which contain "OpenClaw", turning on their IDE partners, turning on their enterprise partners.)

AFAICT anyone who's ever shown faith in Anthropic has been immediately exploited by them to some degree. They will quickly get the reputation of being "the Oracle of AI companies".

I wouldn't even value them at half of OpenAI.

Cloudflare Flagship 2 months ago

Thanks, good reply. I can see the argument for sure.

I guess I like boring software too much to reach for a dependency but I do see how the tooling matters here.

Cloudflare Flagship 2 months ago

I've never understood feature flags. How are they fundamentally different to a Boolean in a database?

They note that Mythos "found a way to inject code into a config file that would run with elevated privileges, and designed the exploit to delete itself after running".

This is more impressive than what the benchmark was supposed to be measuring. The Kobiachi Maru.

Sometimes outcomes and achievements and work product are useful beyond just... stack ranking yourself against your peers. Seems so odd to me that this is your mentality unless you're earlier in your career.

Great write up. Thank you. Really great!

I was reminded of the factorio blog. That game's such a huge optimization challenge even by today's standards and I believe works with the design.

One interesting thing I remember is if you have a long conveyor belt of 10,000 copper coils, you can basically simplify it to just be only the entry and exit tile are actually active. All the others don't actually have to move because nothing changes... As long as the belts are fully or uniformly saturated. So you avoid mechanics which would stop that.

Curious why Dynamic Workers instead of Workers for Platforms?

Seems there's a lot of conceptual overlap between your WFP, Dynamic Workers, and Sandbox products.

(I guess there's an expectation of at least some permanence with WFP?)

Reading these comments aren't we missing the obvious?

Claude Code is a lock in, where Anthropic takes all the value.

If the frontend and API are decoupled, they are one benchmark away from losing half their users.

Some other motivations: they want to capture the value. Even if it's unprofitable they can expect it to become vastly profitable as inference cost drops, efficiency improves, competitors die out etc. Or worst case build the dominant brand then reduce the quotas.

Then there's brand - when people talk about OpenCode they will occasionally specify "OpenCode (with Claude)" but frequently won't.

Then platform - at any point they can push any other service.

Look at the Apple comparison. Yes, the hardware and software are tuned and tested together. The analogy here is training the specific harness,caching the system prompt, switching models, etc.

But Apple also gets to charge Google $billions for being the default search engine. They get to sell apps. They get to sell cloud storage, and even somehow a TV. That's all super profitable.

At some point Claude Code will become an ecosystem with preferred cloud and database vendors, observability, code review agents, etc.