Agreed, it's becoming a bit of an echo chamber anytime Chinese models are discussed here. This is really par for the course for China though, in physical manufacturing they've been doing this kind of thing for decades, I'm not surprised the mentality has persisted into the AI race. Looks like rather than jumping ahead, they will remain persistently 6 months behind.
HN user
jrflo
festudio.net
I'm not sure if you're taking the piss or genuinely don't know who he is
I'd be doing the same thing if I were them as a marketing move
Doesn't disqualify them, but it may call into question their results seeing as they have a potential conflict of interest.
Hmmm, a company that hosts open models is telling us how good open models are...
That's not what airgapped means. Airgapping means the model exists on a system where there is no ethernet cable plugged in to a router or wifi card installed, it is physically impossible for it to access the internet because the hardware connection does not exist. If it was able to get on the internet, it was not airgapped.
Right, this is the classic VC playbook and people should really know better. They use capital to undercut competitors on cost to gain market share, then slowly crank up the costs until they are (hopefully) profitable. It makes 0 sense long-term to spend tens of millions on training a frontier model and releasing it for free for someone else to host on their GPUs. It is naive to believe that when the capital starts to dry up and the VCs are looking for a profit that the models will continue to remain open-weight.
Chinese labs are a bit different since they are somewhat state-funded, so I'd expect them to shift to a model where Chinese models are hosted on Chinese infra.
LLM assisted education =/=> no empathy is taught human-to-human
Just like 99% of reddit threads these days unfortunately, lol
It is super annoying when you first set up a Mac and is really over the top. Definitely geared more towards the average user rather than developers. But, once you get through the barrage of approvals during initial use you're basically good to go for the lifetime of that machine. That said, I really wish there was a "I know what I'm doing" checkbox to avoid a lot of that BS...
Looks like the issue was closed?
Agreed. At someone who works primarily in hardware, it is definitely hard. A midi-to-bluetooth recorder is practically a software project in a fancy box. There is very little physical complexity to deal with here, which is what makes hardware difficult- the interaction between firmware and the "real world".
Congrats though, making a half-million a year business is a massive accomplishment! The headline just irked me a bit, haha (but, it got my attention, so I guess you win there)
I think it's mostly to spread hype for the new models. Sol can be ridiculously long running, even without /goal so it can run for 12hr+ on a problem with defined and verifiable output. So it's a good way for people to get hyped about the capabilities without worrying about usage limits.
Originally yes, everyone's usage limit went back to 100% at the same time regardless of if you were at 0% or 99%. Now they tend to give out "banked" resets where users can choose when to use it for up to a month (except the most recent one, which I think was non-banked)
We really need to stop using $/M tokens as the pricing benchmark. I've found that the number of tokens used tends to be a bigger factor than the listed per token price. The cost per task vs. intelligence curve is really what you care about, and in my estimation Chinese models are just not there. They are focused on benchmaxing and getting the highest raw score they can, rather than efficiency.
Agreed. It's motivational for some people, and that's great, but the "data driven" information it gives you fails to pass the "use your brain" baseline.
100% agreed. I wore an apple watch daily for almost 10 years and have switched back to automatic watches. The notifications caused me to be way too connected to my phone, and the "tracking" aspects of it were kind of useless to me. I know intuitively if I got enough exercise in a given day (did I go to the gym? Did I go on a long walk?), I don't need a computer to tell me that. Same thing with sleep tracking: do I feel like shit in the morning? I probably didn't sleep well!
The one thing it's good for is run tracking, I'll still put it on for that activity.
Tech != the entire economy
NASDAQ is famously overweighted in tech. It saw an 80% drop in the aftermath of the dotcom bubble, while the S&P500 only had a 40% drop. It's a double edged sword, with the AI boom it's benefiting, if that reverses it will fall proportionally to those gains.
That is true of all tech announcements. Marketers will do their thing regardless of reality.
Thank you! I've personally found that having it on iOS especially is a game changer, it's really improved my attention span across the board.
Scrolless, a Safari extension that keeps all the human parts of social media (search, DMs, stories, posts from friends) while removing all the algorithmic garbage designed to suck up your attention.
What destructive actions are you afraid of in particular? Honestly the models are pretty smart, I let the agents go --yolo and nothing bad has ever happened (yet) that couldn't be solved with git.
/goal tune 5.6 Luna parameters until performance is maximized across all benchmarks
My guess is that it's the same for Haiku/Sonnet/Opus: Biggest model for architecture and high level planning and technically challenging problems, medium model for simple implementation tasks, small model is for nothing
Agreed. GPT 5.5 will come up with more straightforward solutions with far fewer tokens than Claude. Also, the usage limits are much more generous for Codex than Claude Code for the same monthly plan.
Depends on which side of the pond (or Canadian border?) you're on ;)
Ah got it, that makes more sense. Thanks for the info!
I don't really understand the point of this, I feel like LLMs have been able to one-shot matplotlib since GPT 3.5. I have extensively used LLMs to do data viz and haven't run into any problems. What is a specific instance where an agent struggles to generate a visualization and Flint solves it?
That would explain a small number of true OS X users, but the graph reports 2x OS X as macOS, which is ridiculous.