HN user

sauwan

477 karma
Posts1
Comments162
View on HN
GPT‑Live 14 days ago

I think you're missing the counterfactual. I now have two choices: 1) sit behind my desk and work, or 2) walk and talk something out.

Before this model, the voice models were pretty dumb and annoying to work with. We'll see if this changes that.

My wife and I had a really complicated spreadsheet for tracking how much we owed our babysitter – it was just complex enough to not really fit into a spreadsheet easily. I vibecoded a command line tool that's made it a lot easier.

Ok, please help me understand. Or is this more of a nanny?

Well, the benefit here is the application of fluoride, which isn't something most people are actively thinking about maximizing.

For pure text responses, agree 100%. Gemini falls way short on tool/function calling, and it's not very token-efficient for those of us using the API. But if they can fix those two things or even just get them in the same ballpark like they did with flash and flash-lite, it would easily become my primary model.

Hah, I quit Prime and they just gave it back to me. No charge. I can't cancel it. I can't figure it out, but perhaps they realized that their margin on me before I canceled was well over the cost of Prime? I'm not sure, but I still only use it a fraction of what I used to...

This is perfect timing for me - was just thinking about how to do this. But pricing is a bit steep for a startup currently looking to prove the market. Would you consider a cheaper option (e.g. 1 free inbox, or maybe $20/mo for 5 agent inboxes and a more limited storage level)? I'm building something that I might consider this for, but I don't know how long my runway is before I get sustainable client revenue, so $100/month is a deep hole being burned in my personal pocket before I can prove my MVP out.

Would be great to see how our previous months usage stacked up and when, if at all, we would have been rate limited.

I'd be pretty surprised if I were to get rate limited, but I do use it a fair amount and really have no feel for where I stand relative to other users. Am I in the top 5%? How should I know?

It's not cheaper than Deepseek V3.1, though, and Deepseek outperforms on nearly everything. And only between 1- 3x the throughput based on the openrouter metrics (near equivalent throughput if you use an FP8 quant). Wish I could be a little more excited about this one.

I'd be surprised if this was a new base model. It sounds like they just did some post-training RL tuning to make this version specifically stronger for coding, at the expense of other priorities.

Gemini 2.5 1 year ago

I'm assuming this is true of all experimental models? That's not true with their models if you're on a paid tier though, correct?

You might consider reading up about Patrick Lencioni's Working Genius models. He has a book that, in my opinion, is a blog post of content turned into a book, so you can get the information distilled in better places. He has a podcast about it which is much better than the book. But sounds like you don't get energy from the type of work you get from a leadership/management position. That's totally fine and normal!