My 2 cents, all these high level videos kind of suck, even the professional ones, when they don’t show the product. Even videos I produce for this purpose! I don’t think this is that bad. I’ve seen worse.
HN user
mchusma
Founder & CEO SignNow Founder & CEO tidy.com CTO HotMic
Incredible. This is definitely the launch of the day. Just crushing Google's releases.
The pricing here is incredible. This is the first US release that's competitive with DeepSeek V4 Flash. Very excited about this.
Wow, Laguna S 2.1 (released today) just destroys Flash-Lite underly and completely. What a weak and embarrasing release from Google.
Maybe, but they said they have “started” the Gemini 4 pretrain. So not having done any significant pretrain in a year or so seems odd to me.
3.6 Flash would be a great model at 3.0 flash pricing. At this pricing, its thoroughly trounced by about 10 models on cost/performance including Grok 4.5. 3.5 Flash-ite would be a great model at 2.5 flash-lite pricing, as is, its trounced by many models including Deepseek v4 Flash.
As is, they are thoroughly outclassed for most usecases. I will say the one area where i do see Gemini punching above its weight class is in tasks that are effectively "Google this for me" / knowledge stuff. So it does have a role, and I do use it. So while I think Google is still in a strong position overall, they are really stuck as a tier 2 AI player right now with text models. They are tier 1 in bio, images, and video.
The issue is not investment at all, we want more investment into housing.
But you are absolutely correct you cannot both want housing to be a good investment (increase in value faster than inflation) and an affordable (drop in price lower than inflation).
Don’t subsidize demand. Allow more supply. Private homeowners, big developers, all of the above. Build.
I am generally for a light government touch, but I think excessive light and sound are forms of pollution and should be generally regulated in a not-overly burdensome way (give people fix it tickets etc). Policing excessive light, and excessive blue light, would be pretty trivial. And I think cameras that automatically ticket loud vehicles would be amazing.
I counted the other day and there were at least 12 different providers with "better than Opus 4.5 performance" on Artificial Analysis, Opus 4.5 being Anthropic's December release that many say kicked off the latest acceleration. Which is totally insane competition, particularly given how low switching costs. I personally think that Opus 4.5 level performance is sufficient for most apps and usecases, as they get deployed.
I’m not categorically against bans of this kind but I’m pretty sure everyone knows these are fake. I am worried about freedom of speech slippery slope issues, and also San Francisco (a relatively small eccentric city) trying to dictate US policy.
I actually think a reasonable strategy of Google is to focus on biotech, where the demand for compute is infinite, and focus on cost effective solutions with really good search for other verticals (eg legal, accounting, etc), and offer loss leaders in coding to undercut Anthropic/openai, with “cheap just behind the frontier” models.
They are not really doing this, but it makes sense to me.
This paints everyone as right or left, which I don’t find accurate or helpful. I think other labels even if they meant the same category, to be more helpful.
I don’t know if it’s a great business model but it makes perfect sense to me. Open models when fine tuned are capable at better than frontier performance at a fraction of the price for many (probably most) domain specific tasks. If companies help make that easy to implement, there is value to capture. But I kind of like Unsloths model here which is to be really good at just layer, and not bothering with building their own models.
Non paywalled alternative: https://www.theverge.com/science/965849/spotify-founder-ek-s...
This looks similar in concept to the Midjourney health, but much more about multiple modalities (different surface scanners). Its interesting, I do think it could be better then dermatologists at picking up skin issues if widely deployed.
I do think this would be interesting if they made these easy to finetune, as I do think this level of intelligence is likely sufficient for many applications and could be extremely cheap to run.
Incentivizing usage during peak times makes total sense, but if price swings are this wild, how are grid scale batteries not highly economical? My rough ballpark math was that you need roughly 20 kilowatts of battery storage to make this issue basically nonexistent, and that would cost about 10 billion dollars, which doesn't seem that much for this.
I would say that overall there are pros and cons to this, I really want to be allowed to use agents on my behalf, and don't want to see sites prevent me from doing this. On the other hand, I do recognize there are cases when its good/ok to have only humans allowed to take some action. In my opinion, the line is likely when you are representing you are a human, its ok to prevent bots. otherwise, you can't.
I stayed in downtown LA recently and looked like the set from the walking dead. Literally blocks of people wandering in traffic. I guess you could argue you definitely don't need flock cameras to see the problem, but also I don't know how anyone would not do everything possible to stop it.
I agree with parent, the full quote is: "The whole thing relies on donations to keep it afloat, which is really what tax dollars are for."
I think this is a great site, love what they are doing, and support them (including a literal donation). But a government maintained website for this data is low on my list of things of what tax dollars are for. In fact, I think this is better done privately. To be clear, many of the things every US administration does including this one I also think is better done privately.
I like the linked Scott Alexander post, but I also genuinely wonder what is the rate of change on these tests? The linked test Prenuvo has competition from Ezra + Function and others. It this drops from $2k to $500 over time, it makes it look considerably better. The more we can use different testing modalities, we should be able to reduce false positives in each modality.
I will say, that for cancer specifically, tests like Galleri seem better, but as that cost comes down I could see in 5-10 years an annual $500 scan that offers a full body scan of some kind, plus comprehensive bloodwork including blood cancer screening, and the type of thing that could be done annually by many in the US.
I was similarly confused. Saying a MRI is the equivalent of stopping smoking for 1 year earlier, or driving a motorcycle 10,000km less seems actually really good! Go MRIs!
As another point, most of the negative costs of getting full body scans are actually poor reactions to the full body scans. The phrasing is "Hey, if you get more information, we are going to act badly on this information." I think the solution here should be just acting better on the information, not getting less information.
This was my favorite as well.
I will plug Willow for mac recording. IMO it's basically to me a "better than perfect transcription" as it cleans things up and is almost instant. I liked Superwhisper but switched to Willow as it was a big difference.
Its so good that I'm not sure that it's possible to get any better. Speech to text seems like basically a solved problem, if not now then definitely in 5 years. I don't know if any of these speech to text businesses will work in the long run, but for consumers they are great. My guess is the 2030 version of Apple's SpeechAnalyzer will be so good that nobody will need to use 3rd party software.
I think the content here is not controversial EXCEPT that it sounds way too doomerish. The only mention of positive effects of AI is "It could bring...major gains in living standards." It really should say "In the US, social security runs out in 2032. 65 million people die annually. AI progress is critical to solving these and other issues."
Full text of statement: "1. AI may become radically more powerful over the next 10 years.
2. This could drive an unprecedented transformation of our economy, larger than the Industrial Revolution, but unfolding over a vastly shorter time frame. It could bring risks, including large-scale job displacement, as well as opportunities such as major gains in living standards.
3. Economists, policymakers and technology leaders must act now to understand the economics of transformative AI and to build the incentives, guardrails, and institutions needed to steer AI in a direction that complements humans and benefits society."
After seeing studies like this, and how the shingles vaccine reduces dimensia, I have become increasingly convinced that it’s bad to get almost any disease, even transiently. I used to think that it was kind of good to train your immune system (kind of whatever doesn’t kill you makes you stronger). I no longer believe that. I believe that diseases often cause unknown effects and it’s better to avoid disease entirely, that vaccines are actually more beneficial than current studies show in this regard, and new universal vaccines to prevent the common cold and flu will likely have significant health span improvements over time beyond the acute prevention of symptoms.
This is great. Federally subsidized loans is directly (not solely) responsible for rapid inflation of college costs in the US. Anything to limit its use is a good thing. I’d argue that this test would be better expanded to actually having an ROI, not just do no harm, to encourage schools to not only provide value but also constrain costs (eg your school may make sense at $10k debt not $100k debt).
This and/or making loans dischargeable in bankruptcy.
I agree with parent, Meta has been at this a long time and its only because they have recently fallen off that they pushed this "oh give us credit its really a new org" thing. Basically, if you can't actually "win" then try to fake a restart and say we are the fastest.
Even given that, this is their second try (they had Spark 1.0). Spark 1.0 was uninteresting, this is potentially interesting, but we can't really try it yet it seems (at least not in Openrouter).
Ultimately, competition is now fierce in this broad level of intelligence/cost: Spark 1.1, Grok 4.5, GPT 5.6 Luna, GLM 5.2
Sonnet not in the same ballpark of pricing (more expensive than Opus in many cases). Haiku has been basically abandoned.
This not being available on Openrouter really makes it hard to test. I was going to compare vs Grok 4.5 and GPT-5.6 Luna, but I don't want to deal with signing up for Meta for it unless it checks out. Please Meta make this available.
Looks like a great set of models, but there are about 20 different thinking/model levels here in this family and they are very complex to pick the right one for the task
E.g. for GeneBench Pro, it looks like you would always use GPT-5.6 Sol over Terra/Luna, its pareto optimal.
For Agents Last Exam, you would maybe want Luna, then Terra, then Luna, then Sol as you increasingly budget for tasks.
I feel that there may need to be a new auto mode in many of these cases. It selects the best model and thinking given a particular problem.
Feels like it's going to have to go that way eventually, because here we have about 20 different model and thinking levels you could use, and they're not obvious which ones are right for the given use case.
Yeah, this is most directly comparable to xAI Grok 4.5. In both cases, directionally "opus level intelligence for haiku prices" which is a really big deal for application developers who want to include models like this in their applications. I have been testing switching out haiku and sonnet for Grok 4.5, and may give this a try too (it is quite a bit cheaper, particularly for cached).
Great model, very nice. Opus class performance at Haiku level pricing (or cheaper with the token efficiency). This seems like a GLM-5.2 killer and this is what Sonnet 5 should have been.
This is a model I could really see used inside applications, where Opus or Sonnet or GPT-5.5 are too expensive.
I would really like to see a strong Deepseek v4-Flash competitor, which ideally is something like Sonnet 4.6 performance at <$0.30 per token. This is missing from main US labs.