HN user

pstorm

135 karma
Posts0
Comments84
View on HN
No posts found.
GPT-5.6 13 days ago

Cost and intelligence aren't the only axes. Terra has better latency and output speed than Sol for example.

I’ve been building this out too, and your comment made me realize the missing piece for me. I’ve given the agents tools to validate its own work, but I haven’t improved the experience of humans verifying the agents’ work.

I played again and found issues. After level 40ish, basically nothing can hurt you. You heal too fast. And once I got past level 80ish, I was running so fast, I ran through walls. And the saddest part, I got to level 200+ and couldn’t die to lock in my high score! It was a fun time though, thanks for sharing this game!

At a minimum, you increase top-k to cast a wider net, then after reranking, take the N you really want. You have to play around with it a bit, but that’s the idea.

I’m very surprised this isn’t getting more attention. Am I missing something?

It seems at or above SOTA on the given benchmarks, doesn’t have context rot, is orders of magnitude faster, and uses less compute that current transformer models. I suppose it’s just an announcement and we can’t test it ourselves yet.

I built one of the top 3 results on Google when you search “compound interest calculator” and a dozen other similarly popular calculator pages.

The value isn’t the interface, it’s the trust that its calculations are accurate. I can’t tell you how many meetings I had with accountants and finance people to validate all the calculations.

This resonates a lot. For years, I would dive in for a week or two, sometimes I would make it to a few months. Lately, I've just been telling myself no, to keep focus on my start up, but it sometimes feels like I'm squandering this spark of energy you described. I absolutely love your approach of doing the planning in that initial burst rather than just getting some part of the idea working in a day or two.

I think those ultra-productive days earlier in my career account for some of the best learning, but now, over a decade into software engineering, focusing on the planning aspect seems much more prudent.

I've recently come to believe the differences in hypertrophy are negligible between eccentric and concentric focused movements, but can't find any recent, compelling research that says so. There was a 2017 meta analysis by Brad Schoenfeld (basically THE hypertrophy researcher) that showed a pretty significant different in hypertrophy: an average of 10% vs 6.8%.

I know Greg Nuckols from StrongerByScience believes this is mostly caused by lifters, especially untrained lifters which most of the research is on, having spent less time in eccentric phases so there is more opportunity for growth there, but it will eventually plateau.

I'm trying to understand this approach. Maybe I am expecting too much out of this basic approach, but how does this create a similarity between words with indices close to each other? Wouldn't it just be a popularity contest - the more common words have higher indices and vice versa? For instance, "king" and "prince" wouldn't necessarily have similar indices, but they are semantically very similar.

I’ve never really thought about this before, but it is obvious in hindsight. When I finally made enough money to move into a newer apartment, the noise from neighbors and outside was almost nonexistent. It was a stark contrast, but I never thought about how that affected my sleep/productivity, let alone how it would affect all apartment dwellers.

While I haven't read 3BP yet, you might be interested in the Zones of Thought series by Vernor Vinge. Extremely grand and unique universe, some politics, diverse cast, lots of interesting tech. I use it as a benchmark for grand space operas.

You raise a good point - AI + human review might end up being more time than just a human doing everything. I can see a certain subset of issues could be simple enough to done by AI and a quick review - like changing a button color or fixing clearly defined bugs. Time will tell how much work gets shifted over to AI + human review, but I'm betting on most of it.

What percent does an average junior engineer solve? If it is even close, these models can be run all day and night for cheaper than one yearly SWE salary.

I've been planning on building some of this for an internal tool, but now it looks like I don't have to. I'm impressed by the demo, it looks really polished.

I'm particularly surprised by the speed considering all of the pre and post processing. I am doing some similar things and that is one bottlenecks. I'll dig in, but I'm curious what models you are using for each of these steps.

Sorry late reply. I believe the goal of separating by age is to have a more apples-to-apples comparison and to see trajectories. For instance, it highlights that at a young age, Gen Z is actually ahead of previous generations on homeownership, but also the upward slope looks like it may trend flatter. It isn't drastically different, in this case, than just looking at the averages.