Surely this is realistic now (or soon) in the form of LLM curation. A few auto-librarians reading everything, looking at different versions side by side, making choices, etc.
HN user
ALittleLight
"AI, please cure cancer."
"Okay, all humans dead, technically a 100% cure."
I hate not being able to use the latest models. There needs to be a much faster resolution to whatever is happening with the federal government.
Hmm. I also recently got a subscription to the NYT. I noticed they sent me a lot of email so I just filtered it with the click of a couple buttons. It's obnoxious but in the minor way that having to click a couple buttons one time is.
The paper says the professors have a median of 200 comparisons each. It also says they only used 2 models because using more models would require more comparisons and they selected Google models because Google was branded/advertised as being education focused. When you see other models show up elsewhere, that's because they extended the main idea to other models but using LLMs to judge instead of human professors.
I don't understand the "deathbed" perspective. Are you going to wish you made more hackernews comments on your deathbed? Probably not. Does this mean you should stop using hackernews?
If you optimized for minimizing deathbed regret perhaps you'd regret that on your deathbed!
If I have cogent thoughts on my deathbed I expect they'll be along the lines of "I wish I wasn't dying" and not regretting the many ways I enjoyed my time on Earth (which includes vibe coding apps nobody uses).
This just says 60% of systems, but not the frequency for those systems. They were evaluating 20 systems, so for 12 systems there were mistakes in the prescriptions, but there isn't information about how common those mistakes were and it's hard to judge relative to a human system.
Now that there is proof of communication while dreaming, it is only a matter of time before someone manages to vibe code while sleeping.
Think of it as paying for tokens. The tokens you could buy 3 years ago are better and two orders of magnitude cheaper today. If that happens again over the next 3 years then the tokens you can buy today to do a job for 20k will cost 200.
This isn't optimistic in my opinion. It's not even fully realistic because Gemma 4, which you can run on local hardware, is even better and another few orders of magnitude cheaper. A 20k job today might a few dollars in a few years.
3 years ago the best model was DaVinci. It cost 3 cents per 1k tokens (in and out the same price). Today, GPT-5.4 Nano is much better than DaVinci was and it costs 0.02 cents in and .125 cents out per 1k tokens.
In other words, a significantly better model is also 1-2 orders of magnitude cheaper. You can cut it in half by doing batch. You could cut it another order of magnitude by running something like Gemma 4 on cloud hardware, or even more on local hardware.
If this trend continues another 3 years, what costs 20k today might cost $100.
For what reason would they attack a single school? Some strikes being well some doesn't mean others can't be mistaken.
Now, I guess. They aren't releasing this one generally. I assume they are using it internally.
The difference is you can appeal or ignore a game result. If Ukraine lost a strategy game tournament, would they give up their territory? Or fight to hold it still?
Europe is currently being invaded by Russia and your idea for a "hardliner" course is to start sending Russia money?
This seems like the kind of foolishness it takes a lot of money to believe. Anthropic blew up their contract with the Pentagon over concerns on lethal autonomous weapons and mass domestic surveillance. OpenAI rushes in to do what Anthropic wouldn't.
If you think that means your company isn't going to be involved in lethal autonomous weapons and mass domestic surveillance... I don't really know what to tell you. I doubt you really believe that. Obviously you will be involved in that and you are effectively working on those projects now.
"Uh, excuse me. This ad for a payday loan company says the vast majority of households can't afford this thing that most households with children do, in fact, afford and pay for. Which figures do you dispute?"
Concluding something farcical and then asking people to debate it is silly and a waste of time. 400k is an extremely high household income. Childcare is a common expense for families with children. The claim that only extremely high income people can afford a common expense is wrong on its face and needs no further analysis.
Median family income in the US is a little over 100k. 400k to afford childcare is a ridiculous lie/error.
If someone is dying of colorectal cancer, trying to get them to worry about their credit score is not only not helpful, it's actively harmful. Messing up your credit for a few years is in the category of "inconvenience" and it's not the kind of thing that you need to worry about when surviving cancer.
Nonsense. There are many ways to get free or reduced price medical care in the US, especially if you are poor. Your doctor will have resources to help you if needed.
You can also rack up huge medical debt and then not pay it. The hospital will sell your debt to bill collectors who will call you for a while, and eventually sue you. At that point you can offer to settle for pennies on the dollar, or you might lose the lawsuit and have to declare bankruptcy which would mean you have negative credit for a few years.
Obviously it will be a difficult time, and hopefully you have something else, but they won't just let you die because you can't afford it.
Aren't all of these things you can do with Claude Code? Granted, the chat app one is novel, but you could ask Claude Code to set that up.
I agree with you that someone who is good with a screen reader can efficiently move through web interfaces. A good screen reader user is faster than the typical user.
However, not all blind people are good with screen readers. For them, an AI assistant would be useful. Even for good screen reader users an AI could be useful.
An example: Yesterday, I needed to buy new valve caps for my car's tires. The screen reader path would be something like walmart -> jump to search field, type "valve cap car tire" and submit -> jump to results section -> iterate through a few results to make sure I'm getting the right thing at a good price -> go to the result I want -> checkout flow. Alternatively, the AI flow would be telling my AI assistant that I need new car tire valve caps. The assistant could then simultaneously search many provider options, select one based on criteria it inferred, and order it by itself.
The AI path, in other words, gets a better result (looking through more providers means it's likelier to find a better path, faster delivery, whatever) and also, much easier and faster. Of course, not only for screen reader users, but also just everyone.
The first email I ever wrote was to Scott Adams. He actually replied!
I was a child and had just read and enjoyed one of his older books, maybe the Dilbert Principle. I came from a religious household and I was surprised by something in the book that revealed him to be an atheist.
I looked up his email, or maybe it was in the back of the book, and wrote him a quick message about how and why he should convert. He replied to me (unconvinced) and I replied back, at which point he realized I was a child and the conversation ended.
When I heard he was dying of cancer I wrote him another email, again offering my own unsolicited thoughts, this time on cancer and experimental treatments. He did not reply, but I thought there was a kind of symmetry to it -- I wrote him towards the start of my life and again towards the end of his.
Interesting guy, I've enjoyed several of his books and the comics for many years. He had a big impact. Tough way to die.
Ah, yes. Trump and friends are in the White House because nobody called them racist. Excellent political analysis.
You can feel animosity towards someone without thinking they've met the elements of a crime.
The 20k model has no subscription? How does that work? Surely it's using a fair amount of compute, isn't it?
The only part of this article I believe is the legal and bureaucratic burdens part.
"Human radiologists spend a minority of their time on diagnostics and the majority on other activities, like talking to patients and fellow clinicians"
I've had the misfortune of dealing with a radiologist or two this year. They spent 10-20 minutes talking about the imaging and the results with me. What they said was very superficial and they didn't have answers to several of the questions I asked.
I went over the images and pathology reports with ChatGPT and it was much better informed, did have answers for my questions, and had additional questions I should have been asking. I've used ChatGPT's information on the rare occasions when doctors deign to speak with me and it's always been right. Me, repeating conclusions and observations ChatGPT made, to my doctors, has twice changed the course of my treatment this year, and the doctors have never said anything I've learned from ChatGPT is wrong. By contrast, my doctors are often wrong, forgetful, or mistaken. I trust ChatGPT way more than them.
Good image recognition models probably are much better than human radiologists already and certainly could be vastly better. One obstacle this post mentions - AI models "struggle to replicate this performance in hospital conditions", is purely a choice. If HMOs trained models on real data then this would no longer be the case, if it is now, which I doubt.
I think it's pretty clearly doctors, and their various bureaucratic and legal allies, defending their legal monopoly so they can provide worse and slower healthcare at higher prices, so they continue to make money, at the small cost of the sick getting worse and dying.
Talented individual(s) who want to do a startup.
First, just from a "danger" standpoint - more people in the EU die from heat than from guns in the US. And roughly 8 times more people die from cold than heat in Europe. So, I would say, that we live in an environment where our neighbors are armed the same way you live in an environment where you're often dangerously hot or cold - i.e. we get used to it.
Second, you can walk or drive on a street. Every passerby in a car could kill you if they wanted to by colliding with you. It rarely happens. Stand next to a tall ledge or overpass with crowds walking by and watch the teeming masses - you're unlikely to see any of the thousands of people walking by leap off to their end. Similarly, in life, even though basically anyone could kill you, it's very rare to encounter someone who is in the process of ending their own life, and killing you would basically end, or severely degrade, their own life. Almost nobody wants to do it.
Charlie Kirk is/was kind of an extreme example. He said many things that severely angered hostile people. He went into big crowds and said provocative things many times before being shot. I think in most situations you have to push pretty hard to get to the point where people are angry enough to shoot at you. If you can avoid dangerous neighborhoods and dangerous professions (drugs and gangs) and dangerous people (especially boyfriends/husbands) then you are pretty unlikely to be shot and you benefit from being able to carry guns or keep guns in your home to protect yourself and your family.
For one example, consider the "Grooming gangs" in the UK, where thousands of men raped thousands of girls for decades with the tacit knowledge/permission of authorities - and despite the pleas of the girls and parents for help. Such a thing could be handled quite differently in a society that was well armed. If the police wouldn't help you, you might settle the matter yourself.
This looks pretty trivial. Obviously modern gains in life expectancy were from removing things that killed us in early age. This says nothing about future gains in life expectancy which may come from biological/medical interventions that reduce senescence.
This is like the worst case of "Sales promises features that don't exist" ever.