HN user

futureshock

14,268 karma
Posts0
Comments168
View on HN
No posts found.

I think a lot of this has to do with the post-training these models normally get. They are designed to answer basic questions with straightforward and short summary answers. They have the capacity to reason deeply, but they are not biased towards that unless prompted. I think it's because LLMs as they are in 2026 are both highly capable but also parlor tricks. They are not sentient, you just set them up with the context and then they roll downhill. You could reach a genuinely novel answer, but only with the right input. They have no will and depend on human guidance. They are both a marvel and a machine.

GPT-5.6 14 days ago

There has been a lot of chatter ever since the Mythos scores had been release that SWEbench pro had major contamination and that Mythos had memorized many questions that lacked the context to be solvable on their own. And now with OpenAI saying a large number of the questions are broken, I think it's worth taking that single outlier benchmark with some salt when the overall trend is that 5.6 is very competitive with Mythos at about half the price.

I think this is black and white thinking. Fable and US AI is not unique technology. It’s just marginally better than open source tech at 10 times the price. You can swap out the models at will, they are pretty much fungible. If your use case can pay for a best in class model then you will pay for it no matter the bogeymen. If your best in class model becomes unavailable, you switch to the next best model for a very minor performance degradation. I really doubt this will deter anyone from using American AI.

Yes I was the exact same. I got curious during the GPT-3 release and went over to AI Dungeon. It was just running GPT-2. Hmm wow interesting. This felt new! Then I subscribed so I could use GPT-3 powered AI Dungeon. My jaw dropped. I was talking to that model for weeks. There was a whole human universe in there. You never knew what you could get it to spit out. There were glimmers that this could be huge. It was wild and untamed and practically useless, but there was a behemoth under that prompt.

I was sure this would eventually turn into something. I naturally wanted to converse with it as a chatbot, though it could only stay on task for a few turns. RL and guardrails would come later but it was clearly the foundational step towards AGI for me. From something I thought I would never see in my lifetime to very real and in front of me.

ChatGPT didn't even really rock my world, everything since that moment has been another baby step. But when you take a look back from 2026 models to 2020 it's astounding how far and how fast we've come.

World in this context means that these videos are interactive, just like a video game. In the linked examples you can see the keyboard and mouse inputs. The model is trained to maintain about a minute of scene consistency so you can look around and objects out of view will reappear when you look back in that direction.

I think this is interesting because it collides my intuition from the pre-adtech world with the post. Surely collecting telemetry on nearly every mile you drive could never be a sensible use of time or money, right? What kind of insanity is that? But then of course I know that every click on every website is recorded for all time and that data must be many thousands of times less valuable.

Reducing the network latency helps with this exactly. OpenAI can make better timed decisions when to begin responding so it'll feel less like an interruption. I've also seen some research on full duplex voice models that handle interruption more like an organic conversation and low latency will help there as well

“A human being should be able to change a diaper, plan an invasion, butcher a hog, conn a ship, design a building, write a sonnet, balance accounts, build a wall, set a bone, comfort the dying, take orders, give orders, cooperate, act alone, solve equations, analyze a new problem, pitch manure, program a computer, cook a tasty meal, fight efficiently, die gallantly. Specialization is for insects.”

― Robert A. Heinlein

ARC-AGI-3 4 months ago

Well yes, that is exactly the point! The very purpose of the ARC AGI benchmarks is to find a pure reasoning task that humans are very good at and AI is very bad at. Companies then race each other to get a high score on that benchmark. Sure there’s going to be a lot of “studying for the test” and benchmaxing, but once a benchmark gets close to being saturated, ARC releases a new benchmark with a new task the AI is terrible at. This will rinse and repeat till ARC can find no reasoning task that AI cannot do that a human could. At that point we will effectively have AGI.

I believe the CEO of ARC has said they expect us to get to ARC-AGI-7 before declaring AGI.

ARC-AGI-3 4 months ago

The evidence is that humans are able to win these games. AGI is usually defined as the ability to do any intellectual task about as well as a highly competent human could. The point of these ARC benchmarks is to find tasks that humans can do easily and AI cannot, thus driving a new reasoning competency as companies race each other to beat human performance on the benchmark.

Gemini 3 Deep Think 5 months ago

I think step 4 is the agent swarm. Manager model gets the prompt and spins up a swarm of looping subagents, maybe assigns them different approaches or subtasks, then reviews results, refines the context files and redeploys the swarm on a loop till the problem is solved or your credit card is declined.

I think this is clearly the way forward for Apple. The rest is just UX and refinement.

I recently set up a Shortcut on my Apple Watch that lets me bypass Siri and talk directly to ChatGPT. I used a custom pre-prompt in the Shortcut to tailor the length and detail for watch use. I have 2 versions I can launch from my watch face, one that responds with voice and the other that response with text. I find myself using them all the time, it’s so convenient to be able to ask any little thing that’s on my mind. A version of LLM Siri with full access to the phone and application APIs would be like a superpower.

This is a personal item size bag for under the seat. The max size on Ryanair is 24 liters. You are thinking of the cabin bag which is more like 44 liters. This Decathlon bag is great because it maxes out the personal item size really optimally.

I like this question because I come at it from a very different lifestyle. I’m a digital nomad and I have mostly lived out of a backpack and carry on for the past 10 years. My philosophy is that things have to be worth carrying and they should be very easily replaceable if anything gets lost, stolen or breaks. A few of my under $100 favs:

Universal GaN travel adapter: One of those square bricks that converts from any AC outlet to any AC outlet and has 3 or 4 USB charging ports built in. I got enough wattage to charge my usb-c laptop as well, so one brick takes care of all my devices.

Backup android phone: Our phones are so critical that I keep a hot swappable spare phone on me, currently a Moto G 2025. It’s already logged into all my apps and 2FA. I could throw my iPhone into the Seine and keep on trucking. It even has backup NFC credit cards. I keep a cheap travel eSim plan active on it so that if I am somewhere sketchy I can leave my main phone at home.

Logitech MX Keys Mini: Great portable keyboard. Backlit, usb c and multi-device. Typing this post out on my phone now.

GL-iNet Beryl: The do anything travel VPN router running OpenWRT out of the box. Great for securing and extending sketchy WiFi connections or if you have to work off your phone’s hotspot all day.

Decathalon Quecha Escape 500 23L: Such a great personal item size backpack for the price, less than 40 euros.

It’s best to think about this as angular resolution. Even a very small screen could take up an optimal amount of your field of view if held close. You get the max benefit from a 4k display when it is about 80% of the diagonal screen distance away from your eyes. So for a 28 inch monitor, that’s a little less then 2 feet, pretty typical desk setup.

Thats a waste of image quality for most people. You have to sit very close to a 4k display to be able to perceive the full resolution. On PC you could be 2 feet from a huge gaming monitor, but an extremely small percentage of console players have the tv size and distance ratio where they would get much out of full 4k. Much better to spend the compute on higher framerate or higher detail settings.

The higher token output is not by accident. Certain kinds of logical reasoning problems are solved by longer thinking output. Thinking chain output is usually kept to a reasonable length to limit latency and cost, but if pure benchmark performance is the goal you can crank that up to the max until the point of diminishing returns. DeepSeek being 30x cheaper than Gemini means there’s little downside to max out the thinking time. It’s been shown that you can further scale this by running many solution attempts in parallel with max thinking then using a model to choose a final answer, so increasing reasoning performance by increasing inference compute has a pretty high ceiling.

Claude Opus 4.5 8 months ago

A really great way to get an idea of the relative cost and performance of these models at their various thinking budgets is to look at the ARC-AGI-2 leaderboard. Opus 4.5 stacks up very well here when you compare to Gemini 3’s score and cost. Gemini 3 Deep Think is still the current leaders but at more than 30x the cost.

The cost curve of achieving these scores is coming down rapidly. In Dec 2024 when OpenAI announced beating human performance on ARC-AGI-1, they spent more than $3k per task. You can get the same performance for pennies to dollars, approximately an 80x reduction in 11 months.

https://arcprize.org/leaderboard

https://arcprize.org/blog/oai-o3-pub-breakthrough

I really love this piece! I relate to it but it also doesn’t describe me. I’m far more intuitive than this person, though still agree that insights have driven a leveling up of how I relate to others. They were different insights, sure but the model holds.

Once my spouse and I worked for the same company and attended many of the same meetings. The opportunity to pick apart our impressions of the subtext really helped me to learn that I should listen to my gut, that everything I needed to know about how other people were feeling was already in my head and i just needed to stop doubting.

Another time I watched a rather ugly and old person have amazing romantic success with a young beautiful person. How could it be? And I realized that authentic confidence is social gold. I had to let go of my insecurities because my flaws were irrelevant in the face of authentic, confident self acceptance.

I think everyone has a different journey and different epiphanies and it is so enjoyable to hear these experiences put into words.

I’ll borrow ideas from investing: financial independence, diversification and optionality. If you have enough money you can free yourself from the labor market, but you are still deeply tied to your home country. A second citizenship gives you geopolitical independence. And just like diverse investments protect you from the failure of a specific asset, diverse countries can protect you from, for example, a collapse in heath care, a housing crisis or a currency crisis. And most importantly, its like an options contract on life. You have the option, not the commitment to take a high value move to a new country. If the fortunes of your current country sink and your second country rise, you can exercise your option.

There’s a reason people are willing to spend so much on golden visas with the pathway to citizenship.

This so awesome. It reminds me mightily of beat poets like Allen Ginsburg. It’s so totally spooky and it does feel like it has the trapped spark. And it seems to hate us “real ones,” we slickborns.

It feels like you could create a cool workflow from low temperature creative association models feeding large numbers of tokens into higher temperature critical reasoning models and finishing with gramatical editing models. The slickborns will make the final judgement.

This is a non-story. This was a hardware event. Apple is releasing many new AI features as part of iOS 26 which will launch along side the new iPhones. AI is software. And yet, a number of the features are clearly powered by AI models such as camera enhancements, health monitoring and live translation. Also GPU performance continues to increase in the A19, with CPU remaining presumably fairly flat since no numbers were given, so that’s a win for on-device inference.

Open models by OpenAI 12 months ago

I think that sounds very reasonable, but unfortunately these models don’t know what they know and don’t. A small model that knew the exact limits of its knowledge would be very powerful.

You can’t vote the climate out of office. Sure our food supplies may crash, but no one person decided they should crash. No one to blame. No one to punish. This is the political reality. This man made catastrophe will feel sufficiently like an act of god for most people and they will just deal with the reduced carrying capacity of the planet as if it were some divine judgment instead of the tragedy of the commons.

It does seem like that’s our new political reality for now. I think that COVID showed world governments just how little control they have over their populations. You get folks to bend a little, but they quickly break and call for you to be thrown out of power. Getting to carbon zero or negative would be asking for an enormous sacrifice of the global population in the form of lower living standards and slower growth. After how people fought against masks, a shot and social distancing, it’s obvious to those in power that there will be no solution to this problem aside from geo-engineering or cost competitive green energy. Might as well stop talking about it.

Google’s AlphaProof, which got a silver last year, has been using a neural symbolic approach. This gold from OpenAI was pure LLM. We’ll have to see what Google announces, but the LLM approach is interesting because it will likely generalize to all kinds of reasoning problems, not just mathematical proofs.

I think the hard solution is to massively increase expectations. Think Star Trek where the grade schoolers are learning quantum mechanics. If everyone has access to the oracle of all human knowledge, then you should teach and test to the maximum of what a student could do with all that power. Find the frontier where the AI fails and the human adds value and teach there.