Similar to the story of George Dantzig, who was late to class and solved two open problems in statistics because he mistook them for homework, I think the current batch of frontier LLMs are chained up by knowing which problems are supposed to be unsolved. If they're let free (probably via some targeted RLHF) we might get a flurry of solutions to open problems.
HN user
doctoboggan
Owner, Lulim Jewelry (https://lulimjewelry.com)
http://jack.minardi.org
contact me at:
python -c "print '{}'.join(['jack', 'minardi', 'org']).format('@', '.')"
[ my public key: https://keybase.io/jminardi; my proof: https://keybase.io/jminardi/sigs/UeFUbvxi2yZJWRicpP5aAtgs0AK0mzQytx2Zefu2RM8 ]Yeah agreed, from one standpoint I couldn't care less that they did a "distillation attack", but I am interested in knowing if China is able to develop open weight frontier models without the prior existence of a huge model to distill from.
Does the apple passwords app work this way (log in to a public machine by scanning a QR code?)
The legal argument is that we gave this data "voluntarily" to the data broker so the government no longer needs a warrant.
This is an area that desperately needs new laws to catch up to the reality of what is going on, but I don't really see much motion toward that goal in the near future.
I have a side business selling custom fingerprint jewelry and I use gemini nano banana to clean up customer submitted fingerprint images. This was a step I used to do by hand at 10 - 15 minutes per image and nano banana is the first model that is able to do the task (it is astonishingly good at it). I can't wait to see what the next nano banana can do, hopefully its released soon.
I feel bad for the engineers at Cursor who have to use Grok in these sorts of experiments.
I think the SotA is moving too fast for the production timelines of an ASIC, wouldn't you think? People are just now coming out with LLAMA ASICS but who would want to use LLAMA? Or I guess you are arguing that the models _now_ will be durably useful enough to commit the time to creating the ASIC?
What are the steep angle arrows indicating? Too steep for an escalator/stairs, but not 90 degrees like an elevator would need to be. Anyone know?
Is there anyone credible who thinks this is a plausible pathway for SpaceX to make huge amounts of profit?
Scott Manly (who I think is credible) has a video where he goes over the logistics of SpaceX's space based data centers. He seems to think its an idea worth pursuing, but its important to note that his expertise is space tech, and not business strategy.
Can you tell me more about the cache misses causing a hefty bill? I think I read somewhere that interacting with a CC instance that has been idle for over an hour can cause a cache miss. Is this what you are referring to? How hefty of a bill are we talking? (using Fable for instance)
Agreed, and the prevailing wisdom now seems to be that unless you can release a truly frontier model, you might as well release yours as open source to undercut your competition.
I do not mind when I am coding with Claude and it uses all the typical claudisms. I am much more bothered when I am reading a blog post, email, or other form of prose and I see those same claudisms.
I guess they are not annoying since I know I am talking to an LLM and expect the typical responses. When I am reading prose online that I previously would have expected a human to write, it can be quite jarring to realize its an LLM.
Very interesting, I wonder what happened in 2020 that causes the rotational speed to start drifting the other way?
Pandemic -> more people working from home -> less people in tall office buildings -> faster rotation (like a skater pulling in their arms).
Probably not remotely true but it would be funny.
I also don’t understand the reference to 2017.
My guess is that is when they last changed the offset, so the -37s has been in effect since then.
What causes the unpredictability in this? I would have guessed we have earth's rotation and orbit down to many decimals. Does geological activity, weather, or something else cause rotation speed differences that we just can't predict?
You can never ask why a model did a certain thing, or what it was "thinking" when it said something - just like you can't ask a human which neurons were firing when they had a certain thought. The information just isn't available at that level.
You absolutely can have deep nuanced discussions with LLMs however, you just need to better understand their strengths and weaknesses.
Yes, good science writing almost always gets an opinion from someone not involved in the research for the article. I would guess varying definitions of "not involved" depending on the repute of the publication.
Ordering a main, a side, and a drink isn't really "working to reach it". Your original post was insinuating that the OP or their friend lied about the cost and I was just demonstrating that it's quite plausible to reach it.
I don't think you can get to $68 for half as many people, even with drinks and tax.
A 5pc chicken tenders, Mac and cheese, and a large drink is $25 before tax. If there are three people who get a similar meal (but not exact so they don't share the family meals) then the total is $75 before tax. Seems like the original price quote of $68 is certainly plausible for a group of three. I am sure its possible to feed three people for less like you claim, but that doesn't mean the $68 is impossible to reach.
Interesting project and (possibly more) interesting explanation of the development process. I agree with the author that the primary difference between vibe slop and real engineering is just reading the lines of code. However it does feel like we are just on the cusp of only needing to read the tests and _not_ all the lines of code. Maybe a few more model generations and we will be there.
No, you are misunderstanding the graph. Draw a vertical line anywhere, that is a "constant cost" line. For any given cost, Opus 4.8 has a higher performance than Sonnet 5. Only where Sonnet 5 effort is at medium or low would it make any sense to use it, as there isn't even an equivalent Opus effort level to compare to.
Alternatively you can draw a horizontal "constant performance" line and see that Opus is cheaper for a given performance level.
You have to pay more for that, and/or go through some USG vetting process.
The cost per task chart is telling me that I should _never_ use Sonnet 5 above medium effort level - Opus always performs better for a given cost. So I guess the takeaway is that if Sonnet 5 medium isn't good enough for you, switch models, not effort levels.
If the Chinese government is as involved in LLM development strategy as many people claim, wouldn't you expect them to immediately cease releasing open weight models and restrict access as soon as they start producing the frontier models? I am assuming this is what the USG thinks and is why they are trying to cut off the flow to foreign nationals ASAP.
LLMs are an undeniably valuable tool, and governments like to control those.
Your blog post doesn't get found by anyone in Google until you've built up your SEO mojo, your LinkedIn post isn't read without the followers you need to accumulate and your content has to get engagement for people to see it even then, you don't start off line with a million followers on X, etc.
I hate that this is true. It's the worst part about selling stuff online IMO and I found that you have to spend so much time doing it. In many cases, selling something online can be optimized to the extreme such that spend on marketing should be as high as possible and spend on the product R&D, manufacturing, support, etc should be minimized as much as possible. This equation gives you the most profit, but also gives the customer the absolute worst product that is possible to sell.
Capitalism doesn't really have a solution to this problem that I've seen yet.
Mr. Robot had that kind of writing (at least in the first season which is the only one I watched).
Anthropic and Google have both accused China-based rivals including DeepSeek of using “distillation attacks” to train their models by siphoning knowledge from American companies’ AI.
“distillation attacks” is definitely an interesting way to phrase that.
Usually JavaScript is blocked when you load pages that way.
That is, I would say that creativity requires that the new things generated be Evaluated. Without evaluation, and retention of the best, there is nothing created. The novelty flickers into existence but, if its value is unrecognized, it flickers away and is lost.
I really like the way he frames this here. I think a lot of people in the twitter comments (and maybe a few here) aren't reading past the introduction. He isn't saying AI systems are incapable of creativity and discovery. He is claiming generative AI without a harness is not capable of creativity and discovery. There needs to be some other system that "recognizes the value" of the novel idea and remembers it. He gives examples of where this value recognition step is automated and thus by his definition achieve creativity and discovery in a fully automated system.
I guess we are well into the enshittification phase of starlink. Here's hoping Amazon Leo comes soon so we can have some competition in this market.