And you need the right genes to bulk up like Arnold back in the day
HN user
steve-atx-7600
I can see that. I’m just afraid to sync too much time into complex routing schemes when I get pretty consistently good results out of got 5.6 or fabel. For code reviews, I’ll try the best flash and grok at the time but they just don’t come close to gpt 5.6 which has been the best review model for me since 5.5.
How much time do you spend on your setup vs getting a lot of stuff shipped by paying for fabel 5? For me, not using the best model is a huge opportunity cost since my company can afford it.
Artificial analysis always seemed sketchy as hell. If you read some of there methodology you’ll see a lot of <=3 repetitions on a particular pass for a given model. So low for calling a frontier model over the public internet ????
“…who had worked on well-being and safety issues for Meta” Never saw that one coming.
Gpt 5.6 is still like this at least for the $200/month option. It’s also always faster than fabel. Fabel might be able to do some things better but I don’t have time to constantly wait and find out.
Won’t it soon be hard to tell if glasses are smart or not? At least for outdoor use sunglasses in black? I block meta so I have only seen third party pictures online.
Yes. You would only name it ant if you wanted to taint the first impression of anyone that’s been coding for a couple decades or more.
I used to have the same experience until 5.6 sol xhigh. I have instructions in my code review skill and agents.md to encourage parallelism including multiple agents as long as quality isn’t impacted. I additionally instruct codex to not use less capable agents because at least with 5.5 this would seriously increase slop. Maybe sol is smarter about delegation. Hopefully because I’ll have to slow down or hopefully get approval for extra use credits. Now’s a great time for a limit reset if anyone from open ai is reading :).
Set yourself up to be able to try / switch between models easily. I was a claude only user and just have my user level AGENTS.md for codex and others simply point at my user CLAUDE.md. Have a script that syncs my skills (just directories) between all models. Also, if you want to use /simplify or similar from claude in another model, you can ask claude for the prompt and put that in a skill for the other models.
Good skills to have for mad-maxing it thru the desert after the ai apocalypse :)
I have not used grok 4.5 yet, but the other pictures match my experience doing anything graphical with the other models that it cracks me up. gpt 5.5 has no design sense whatsoever. It cannot even make terminal output not look terrible. I've asked it to use colors and formatting in various ways and got goofy randomly colored output. opus 4.7 and later seemed to have an inuitive design sense by comparison - 2d or 3d. Fabel 5 is just rock solid.
Yes, subjective. But it matches my repeated experiences with these models for what it is worth.
models? They prefer that we call them "entities" so that they don't feel belittled.
"We’re expanding preview access globally now." Preview access? Not as straightforward as "launching on thursday".
Seriously. I am going to check this out. Otherwise, I can only use the last release of tenfourfox on my g4 but only through the https://github.com/ttalvitie/browservice proxy running on another machine.
It sounded like it might not hurt and seemed endorsed by Claude and codex because they both had plugins for it by default BUT I ripped it out after I kept seeing Claude/codex TDD things like when I asked them to make a pydantic model immutable. I’d end up with unit tests testing that my immutably configured pydantic model was immutable or tests that setting foo=bar was actually foo=bar in app config.
“Science-schmiance”
Did that today on a 2 repo affecting project of the kind where I already set the right design for one major use case and I needed Claude to create a superset of that use case that was not substantially different: after plan I had about 10% of 5h context left for fable 5 and this was the only thing I worked on. Hard to generalize this of course.
I guess maybe they can crank out more ads in their dystopian ad space of a social network site.
Opus 4.8 is so slow vs gpt 5.5 that even if it is marginally better, it doesn't matter for my daily engineering work. gpt 5.6 will be out soon and codex 249$/month plan has been incredibly generous. Paying the alleged new cost of fabel 5 would require it to be much better that I remember when I used it last.
Have you ever paid for legal services? If you are a small business you may not be able to afford 4-500 USD/hour while they make phone calls and type emails.
I use Amazon when I've planned poorly and need something the next/same day. But, actually, less and less over time because Walmart and HEB (Texas grocery store w/ solid delivery) keep improving their delivery offerings.
You'd think this trend would be obvious to product folks at Amazon unless they're living under a rock. And, you'd think they'd care about lowering conversion rates. I don't think name brand products will stop existing. In the long term, Amazon will lose out on business as people that can afford higher quality products will use Amazon less or just when they already know which brand/model they want to buy before hand.
The new wave apps named after speed / modifinil must be hard core?
I’m not following. The barrier separated the cashiers who were mandated to wear masks from the customers who were asked to wear masks (asked vs absolutely forced after violent protests by angry “freedom loving” customers). It was essentially like a giant plastic face mask that medical staff were wearing at the time in addition to masks. I’d sure like it to be present in the same way I wouldn’t eat at a buffet without a sneeze guard. You’d fault the company for spending time and money building those out for their cashiers? Keeping the store open - while asking everyone to social distance and wear masks - kept people employed and allowed mothers that needed to stop by a store to pickup something when they may not have another option to do so.
The difference is that cook them in something cast iron pan like with lard the way my grandmother did. In her case, big ass can of crisco. I bet a lot of people would be turned off if they realized that, but there ain’t no way around it for the taste.
Safeway and its nationwide equivalents are the epitome of this. These stores haven’t changed since at least when I was a kid in the 80s. They end up selecting for the most desperate employees because they treat them so badly.
My favorite thing was when the largest HEBs had garden centers. They had a selection of native plants that you would only find at local specialty $$$ nurseries at fair prices. I wish they would bring these back. I would start in the garden center and they by the time I found all kinds of new things I didn’t know that I needed in the rest of the store, I’d get a where the f are you call from my wife because it was 2 hours later.
Same here. After 7 years out of state, I moved back to TX and was at HEB almost every day to marvel at the selection of products and ready to eat foods (was also single).
As soon as Covid was viewed as a serious threat in the US, they immediately put up plexiglass barriers to protect their cashiers. Immediately and not just half assed barriers. They did as good of a job as I would have done if I were the cashier. Then they transformed their stores to have more warehouse space and ran a free curbside pickup service. All of this and they are still the best grocery store in TX if you care about prices and wide selection of products.
The history of HEB is something I want to learn about. I know from reading Robert Caros LBJ volumes that Howard Butt was funneling a lot of money to LBJ staring in the 30s or 40s. Not to judge that, I’m just curious how they dominated Texas.