This was a big concern for earlier models, but with modern CoT trained models they should be able to come to the conclusion entirely in the thinking trace.
HN user
mattnewton
matthewnewton.com
I think a) the labs are releasing very fast and b) why would they implement the long tail of app features when they can effectively sell tokens to every user to write their own version of the app, which is what is currently happening?
It’s not that poor people can’t afford a stamp it’s that they aren’t going to spend money on an automated service or stamp if there are other places to apply to that don’t require this.
So I think it’s going to lower recall a a lot - it reduces volume of good and bad actors equally. Anecdotally, I’m not going to bother with such an application unless it’s my top choice or I’m desperate. For small companies that probably means throwing the babies out with the bathwater.
Why do your think Meta doesn’t make money? Their ad platforms are incredibly lucrative.
There are plenty of services to send mail form the internet for a small fee, so this will only discourage the most poor candidates and add friction for the best ones.
I’m saying there is basically no way to both make vlms able to understand the long tail of PDFs where the layout conveys information (like charts and tables) and to make it as token efficient as text formats. Current approaches have mostly chosen to work more often than not at the cost of token efficiency.
Because PDFs are a nightmare of a format and the only thing that’s is reasonably guaranteed about them is they will render to an image that people can read, the parsing of which will be much less token efficient than the equivalent text
But then I close my laptop and it’s not running on the headless host anymore right
I read point 19 as Palantir’s goal being to import Chinese and Russian style surveillance, and the comment saying effectively “it can’t happen here” and “taking them less seriously” because they are raising alarm. After rereading my comment, I still think it clearly responds to that.
How hot does the water need to be before you raise the alarm?
I think there is value in pointing out trend lines and voicing opposition even if there are other countries that have more authoritarian views on speech. This is not a competition, what matters is the experience of the people in the country today not the fact that if they moved to Russia it would be worse. What is important is that the US state has both gained capabilities to act that way, and has shown predilections for it.
Texas just gave a man 30 years for transporting zines because of the politics of those zines. The trend lines are potentially very bad. And it only gets harder to reverse if the concerned people are right; would you just say “I don’t think it can happen here” and have people wait until it does and delay talking about it until we are not allowed?
In the US, I think we are being intentionally DDOS-ed.
This strategy was laid out by Steve Bannon in the old frontline PBS interview where he called the media “the opposition party.”
“They’re dumb and they’re lazy, they can only focus on one thing at a time,” he said. “All we have to do is flood the zone. … Bang, bang, bang. These guys will never – will never be able to recover. But we’ve got to start with muzzle velocity.”
Also see Vance’s recent comments about how nobody would hold Nixon accountable for watergate if it happened today, it would be lost in the next news cycle.
Yeah people are overdoing it, and not everyone is being motivated to do something by feeling mad, but maybe instead of throwing up your hands that you can’t care about soybean tariffs, you could try to educate yourself on tariffs and choose your political representatives based on whether or not they are doing a good job. Or read more narrow news. Completely shutting down the channel seems to be throwing the baby out with the bathwater.
This kind of checking out / mass abdication and apathy seems really dangerous in a democracy.
_in a polling place_ no less
Definitely encourage you to test the models. We tried to optimize for realistic focus and not over-sharpening, which leads to a "hyper" AI-look. It's hard to benchmark because people generally prefer sharp, saturated orangish pictures all else equal, but I believe these are bad shortcuts for the model to learn realism.
Krea 2 Large (on the website and api) was trained with the FLUX 2 VAE, if you want to test it out and push realism. After working with both I think the flux VAE has a slight edge in learning realistic textures but it's smaller than you might think, the Qwen VAE was overall very good in ablations and good at learning to produce a diverse set of styles.
You can find some links and details in the GitHub readme for finetuning / LoRA support. Ostiris, musubi tuner, fal and hugging face diffusers are all day-0 supported :) https://github.com/krea-ai/krea-2
We recommend training off the undistilled, Raw checkpoint, and then applying the LoRA to the Turbo model for inference.
Hi HN, we're releasing weights for our latest text to image model and publishing this writeup on how we trained it in quite a bit of depth.
I hope there is something in the report for everyone, we included a fair bit on the actual training and data infrastructure usually not written about much, that I think will be interesting to people here. There's more that didn't fit, happy to answer questions!
The goal is to encroach on privacy, and those interests are using the children able to do it. There is no “solution” to the children using the internet anonymously problem that will satisfy them. We have to keep fighting for privacy.
I do not need to have the real secret of immortality to say to the emperor that swallowing mercury is bad.
The burden of proof is on those who put forward a solution.
In my view, it’s very simple. There are places like schools, or parents buying phone plans, to identify children. Will some children get access outside of that? Sure. But 100% enforcement isn’t possible even if you thought it was worth destroying privacy on the internet.
Because banning smartphones in schools doesn’t affect adults not in those schools, whereas age verification does?
How you implement these protections matter.
They have a few in models that powered the app, including one that applied edits. Recently they also started fine tuning Kimi models under the “composer” brand, and Composer 2.5 is a very cost effective coding model. But I suspect that the real value is in the distribution they have, which is what I primarily meant.
I recommend watching the video, he makes an IMO excellent case that #1 without a lawyer really is a bad idea. You can still help solve the case in way that protects you, the stakes here are often incredibly high.
Congrats to the Cursor team!
When I first saw the company built on top of vscode in such a crowded field way back at the end of 2022, I thought "forget having a moat, they are renting their castle from the invaders!" - I couldn't see how see how a single team could execute well enough to effectively muscle their way in between Microsoft and OpenAI, who at that point looked destined to control the developer ecosystem between GitHub, VsCode and the then-best coding models. I think it's easy to forget how insane this seemed even just a few years ago.
But every year since then they managed to simply ship a better product on the axis that mattered to the most users. And now they are sitting between a huge user base and a massive stream of valuable tokens, they can sell to SpaceX. Incredibly impressive.
they have had one for a while now. https://cursor.com/cli
I think it is less about competing for athletes and more about competing for national attention (in the form of sports viewership that turns into money and school programs).
I don’t disagree with your conclusion that this is likely ai rewritten, but I do find it strange that you say “normal people don’t write like this” when it is mimicking how people write, and using patterns I have seen people write. I think models are at the point where style is not really reliable as an indicator anymore.
this is almost certainly too recent to have been used for training data, no? Unless they optimistically included most repos somehow?
No insider info, but just wanted to mention that pricing signals things too. If Mythos is only servable at $X*Y dollars and isn’t Y times better than $X of compute at another provider, it’s quite possible that affects the IPO price negatively versus the halo of having the worlds most expensive model that is “too powerful to release” unpriced and unbenchmarked.
I think that most people at Anthropic are true believers from my interactions with them so I don’t believe this theory anecdotally. The simplest explanation is that it really is taking a while to gain confidence they won’t be used for a spree of bad cyber attacks. Knowing how long it takes institutions to fix security issues when filed by humans I would be more suprised if this wasn’t the case.
But I would forgive anyone who did think it was deliberately sandbagged; given the staggering sums at play, true believers might believe the ends justify the means to a little “marketing” like this.
I wouldn't be surprised if each of the frontier American labs and individually has compute access similar to the entire EU. Chinese firms are a more interesting comparison since there are a fair amount of great models there, and it's estimated about 15% of the ai relevant compute is in China versus maybe 5% in the EU under European companies (and 70% ish in the US is the most common ballpark I see)