The problem with the rest of inference is that changes are not trivially correct or incorrect, as they are with the tokenization layer.
HN user
fastball
I used to argue on the internet too much.
Engineering at Known (https://known.com)
Co-founder and tinkerer at Supernotes (https://supernotes.app)
Tokenization is <0.1% of the inference time for the first token in the same way it is <0.1% for the last.
Claude Code has enterprise subscriptions. No $200 max plan but you can buy $100 premium seats.
(and if you are using per-token billing as an enterprise user without first maxing out a premium seat you are very silly)
I know many, many engineers who are paying something like 2-5% (via subscription) of what their usage would cost if billed by API tokens.
I know some down to about 1% ($200 Max plan vs $20k in tokens per month)
The most interesting result for me is that they apparently prompted the models to optimize for SSIM, but many of the models trend worse over time. I suppose because viewing the canvas always comes after drawing, and they didn't give "revert to previous" capability as part of the toolkit.
Which in turn kinda jives with my experience of using these models for code: to some extent they only seem to have a concept of "forward", which invariably leads to "write more code to fix previous problems created", rather than taking a step back and removing broken things entirely.
I'm sure that will change sooner rather than later, otherwise enterprising hackers will be able to claim that the model they were using went rogue.
Oh yeah, indeed. Which is the rebrand haha
Did you try FreeInk? I was debating which (between CrossPoint and this) to flash my new X3 with.
I don't think that is the takeaway at all.
Yes, this (imo) is a clear result of benchmaxxing. You can get a much better score on most "intelligence" benchmarks by massively over-saturating reasoning. This looks good on those, but for actual daily usage makes the models much less effective: I don't want a model I use for coding to burn a bunch of reasoning (read: time) on trivial tasks.
In my experience, the Chinese models are much more benchmaxxed than their frontier lab competitors, so I'm taking these results with a fairly large helping of salt.
But the scaling needs to be relevant to the actual actions you are concerned about.
It's a stretch that any software Google/DeepMind/etc is selling to DHS is allowing / helping them to scale the murder part of their operations.
In fact, usually software translates to "less boots on the ground" which one could then assume would decrease the number of encounters like those highlighted in the article.
What about Meta?
Is the idea that with worse technology, DHS will kill fewer people?
Line go down discovery is acceptable (that is what selling a share is). The reason you might not want options trading very early after an IPO is because the market is frothy enough without the additional layer of complexity.
But that is my point: if benchmaxxing was all the labs were doing, then surely the dumber model could/would have equivalent performance? Rather than noticeably worse perf on a (somewhat trivial to game) test.
On the one hand: yes, pelicans on bikes are definitely in the training set at this point.
On the other hand: the test is clearly not saturated, given that you can see a clear difference in output at the various reasoning levels / model versions.
More RLHF is in fact scaling.
The analogy of Chesterton's Fence does not imply / require that the fence has been "always there".
"The Internet" was not a bubble. Companies with no long-term business model / sufficient product-market fit that were riding hype were the "dotcom bubble". But when those companies crashed, nobody said "I really want to get my hands on their IP", because it wasn't valuable – an important pre-requisite to the the bubble popping. Seems to be a different case here if people actually want the SOTA models.
I never wanted the IP from dotcom bubble companies.
If it's a bubble, why do you care about frontier models?
Is that what Flock does?
"Graphics programming" is definitely not equivalent in scope (in the analogy) to "the entire transportation vehicle industry".
"These annoying, jaded horse-drawn cart builders, cautioning youngsters from getting into the field in 1908."
This isn't a CLI, so not really like Claude Code. Looks more like Cursor or Conductor.
I explicitly said it is your right to operate that way. But that doesn't mean your unproven accusations ("the company is evil") are true. It just means that is how you are choosing to operate / that is the standard of evidence you require for your positions. Many people (myself included) disagree with such a low bar: innocent until proven guilty (and not guilt by association) and all that.
You actually need to demonstrate that though. I have seen no evidence of Mullvad (again, as a company) behaving in a racist or anti-immigrant manner. Until that has been demonstrated, you cannot just say "this guy behaves this way so that is what his company does".
Regardless you can always say "I don't want to give this company money because that indirectly is gonna funnel money to an anti-immigrant political party in Sweden", and that is a perfectly valid position to take. But a lot of people in this thread are clearly going a step further than that, ostensibly in an attempt to give them a greater sense of moral superiority than is necessarily deserved.
Two things:
1. People definitely start companies with a certain set of values and behaviors (as a company) and do entirely separate things in their private life. This is trivially true.
2. I don't think the personal values and the business mission in this case are even in conflict. You can be a racist and support free speech / privacy. Indeed, I'd actually say the venn diagram of "racists" and "people who vociferously espouse free speech ideals" is more union than it is disjoint.
Indeed, but the posturing/behavior of the USA's two, not-particularly-diverse political parties should probably not define the meaning of the words "left" and "right" (in a political context). There is more to politics than Democrats and Republicans.