Well, many EU countries like Italy and Germany officially freaked out about DeepSeek, ordering it to be banned from app stores etc.
HN user
montroser
This is why God gave us 1.58-bit ternary quants?
This was always where this was heading, but we got here much faster than expected.
Once western governments declare it to be a "national security" risk for citizens to have access to open-weight frontier models, and once they classify using these models as acts of terrorism, what will that world be like?
Will using Kimi K3 come to be like how napster was in the olden days? Everybody knew it was technically illegal, but come on -- any track at your fingertips? But surveillance is quite more evolved now.
Or it will be like cannabis, where a guy in the neighborhood will low key rent you metered access to the 8x5090 rig in his basement he cobbled together from parts on ebay? Or everyone will flock to VPNs?
Or will the oppressors actually succeed? The same way that napster is long gone, and everyone accepts that they must pay spotify for a homogenized collection, where artists must take only a minuscule cut (more than napster though)... We'll be stuck with nerfed Cohere or Mistral models for open-weight options, as if they need more lobotomizing. Or else we can pay through the nose for Anthropic/OpenAI for "American Frontier" models which will fall increasingly far behind China.
Or else, like how Kindle Fire was subsidized by ads, we'll have "Kindle AI" where influence is sold to the highest bidder, where the LLM will tell us that smoking is actually healthy if big tobacco can engineer its renaissance by turning its lobbying dollars to pay-to-play, pumping its propaganda into the training pipeline for Amazon's extra commercialized line of ultra budget LLMs.
One time I hired two different random guys on craigslist to assemble three IKEA bureaus. They didn't know each other beforehand, and I hired them separately to work together.
They started by jockeying to see which one was going to be in charge, arguing about how to proceed, throwing shade, making claims of superiority. One guy wanted to follow the instructions, and the other guy just wanted to go on intuition. They couldn't agree and asked me which one should be in charge, and I said, I don't know just work together, follow the instructions, and get this stuff built.
The instructions guy assumed the lead, and then they settled into a rhythm. Three hours later, they couldn't stop gushing about each other, and a most beautiful bromance was born. They exchanged telephone numbers and made plans to hang out on the weekend.
This is very cool! Slightly off-topic though, I miss technical people writing in their own voice about the awesome things they've built.
Welcome to the future. I think open weight models are our only hope for LLMs being net positive for society.
Well, it's not that hard: just give the LLM a user-scoped access token, same as if the user themselves were asking their own LLM to act on their behalf.
Basically, just like we don't show users information they shouldn't be able to see, and don't let them take actions they shouldn't be able to take -- we can use exactly those same explicit mechanisms (scopes, roles, permissions) to limit what the LLM can see and do.
The LLM could try to do more than what's allowed, but they get shot down with an access denied message just like anyone else.
The anti pattern is to think that you can reimplement access control with prompt engineering and give the LLM root access. That is doomed to fail every time.
I'm ready for my home helper robot that makes dinner and does the dishes and takes out the trash.
But I'm scared for when those home helpers get drafted to fight in wars, either for or against me...
This article only promises to get into "the coming AI margin collapse" in a yet to be published "part two". This part only makes the point that GLM 5.2 is pretty good (no shit).
I'm glad that your trivial migration was trivial.
All this was a big source of churn for plugins and frameworks: https://github.com/vitejs/vite/discussions/16358
Vite had five major version in the four years 2022-2026. Version 3 => 4 => 5 => 6 => 7 => 8. Each one of those had breaking changes and required devs to go through a migration. It's too much. And for what? It's not as if it is dramatically better now than it was in version 3.
I can't say I would really look forward to bringing this level of needless churn and constant disruption to the rest of my development toolchain. Anyway, Vite+ is really just wrapping existing tools into an abstracted command-line interface? And so I have more layers of indirection to wade through in order to get the thing to do what I want? So far I am not optimistic about this prospect...
That "thank you" at the end is particularly classy. Thank you for getting fucked and giving us your money.
Yeah, but they're talking about fine-tunes.
Sounds cool. How do agents know what else is going on in the doc? They have an embedded browser and they do like mutation observer type stuff? Or does the integration do polling?
Nice. Can we get `nub --compile` up in there like Bun has?
I've been driving glm-5.2 for a day or two now. It feels like a mature, seasoned colleague.
It could be luck, but I don't know -- it keeps one-shotting relatively hard stuff. And taking initiative to think about what potential regressions it should look out for, and choosing to do strategic refactoring when it should do. It is not confidently incorrect hardly at all, doesn't tell me that it's fresh risky pile of changes is ready for production without having exercised all the code paths and writing a bunch of tests, etc.
We might be reaching the next level here...
Reading the paper, by "Fertility" the author really means "rate of pregnancy" which seems kinda like a different thing.
Basically, access to the Internet is correlated with less teenage pregnancy. Idiocracy in motion...
Makes sense to split the foundation, as the communities have split (and withered).
From its inception Perl 6 was an incredible journey that resulted in a genuinely weird and interesting new programming language, and squandered a broad wealth of momentum and good will and enthusiasm from the Perl community at large. It was a dramatic slow death over the course of a decade, where people who had built their careers, and small and large companies who had built their economic engines on Perl got to come to the realization that the whole thing was over, killed somewhat inadvertently by its own creator...
Well, this is certainly not benchmaxxed, I'll give it that. And props for being honest about how far behind Qwen 3.6 MoE is this model.
But yeah, it's not the best look to have to stretch and say it's "competitive" with other models in it's weight class, when it offers not much else that's useful or novel.
deepseekv4 pro via opencode go is $10/mo and has very generous limits. I use pi for the harness and go just as a model provider. It goes a good long way...
This is a finetune of Qwen-3.5. Interesting to see this coming from the government of Brazil.
https://www.reddit.com/r/LocalLLaMA/comments/1u4fzg1/new_mod...
Well, this is the exact opposite of his point. Of course it should make sense when not animating! That is given. The entire crux of his point is that it should also make during an animation.
In an ideal world, it is hard to argue with. Yes, sure it should make sense. But also, please don't spend precious cycles on this unless all the other bugs are fixed, and this animation consistency is truly the most important remaining issue to address.
In Cellpond, I handpicked hexadecimal values for each channel so that the resultant colours would better fit my app's theme and needs.
Well, this is an admission that trying to balance "wiggle room" without too much "fussing" with 1000 colors didn't really work.
Evenly sampled in rgb space, a 1000 color palette yields neither enough flexibility (especially in the blacks, greys, whites), nor enough constraint to really make it dead simple.
For app development at least -- choose 20 gradations of blackish to whiteish; 8 gradations of an accent color and so too for a couple of secondary colors...and you're good. That's like 48 colors instead of 1000.
Result is ~12 tokens per second, as reported by OP down in these comments here.
An impressive effort, and better than I would have thought possible on this hardware -- but still pretty far short of what one needs for an satisfactory interactive session.
This is my daily driver laptop. It's pretty good for what it is. Runs Linux perfectly, not trying to be especially too fast, very nice pixel density, all metal case, sturdy build. Battery life is not the best. Beautifully compact.
In practice, my experience is that it's mostly a lose-lose proposition. You have to invest in learning a bunch of same-but-difterent framework apis to do what the language already does natively. And in return, the code is more complex, and harder to debug, and so it has more bugs.
We once hired a very smart fellow to build out a media processing pipeline. He did with rxjs, but it wasn't scaling well. We tried to get with the paradigm for a bit and help scale it up, but flame graphs in profiler output were all crazy, and it was a pain to wire in timing traces, etc. We built a POC imperative version just to prove that we could indeed achieve the throughput we thought we could, and then we just said, well hey, this is faster and simpler, so... let's just go with this instead. And so we did.
Come on, now. The human writes the plan up front, which includes guidance on testing strategy, classes of tests, particular test cases to cover, etc. And just like normal, of course you don't just ship the code without doing manual verification, code review, auditing the test cases, and all the rest.
It matters for at least a few reasons:
- Depending on the nature of your application, it may be very important to be able to audit the business logic and intended behavior. For compliance reasons, for operational reasons, for moral/ethical reasons -- you very well might want to affirm what the code is actually trying to do.
- A coding agent may get very creative in order to write code that passes a tightly-defined unit test. It may come up with approaches that technically pass, but work against the overall intention of the app in the first place. This becomes an arms race rather than a productive collaboration, where the agent's increasing creativity has to be matched by a sprawling test suite.
- Eventually, inevitably, business requirements will change, and the blob will need to evolve. It will be much easier for an agent or a human alike to understand how to safely make the change, if the existing implementation is transparent and understandable.
Well, my team does what we call vibe engineering.
You do ask hard questions up front, define boundaries, give lots of high level architectural guidance, declare interfaces, and bounds of abstraction... And then you ask the LLM to make it so, and it does. You give it the structure, and it fills in the implementation.
This is engineering, more or less.
Slate is just some renderings though, right? Is there anything actually real about it more than just marketing?