the line `ssh late.sh`?
HN user
kennywinker
Good thing you don't have to install anything to run `ssh late.sh` and check this out
https://en.wikipedia.org/wiki/Wells_Fargo_cross-selling_scan...
https://en.wikipedia.org/wiki/Subprime_mortgage_crisis
Turns out banks run on trust and honor too. I personally know people who got loans because of knowing the right person not their credit score.
But even if banks did run on transparency and accountability, which they don’t, governments aren’t banks. You can’t double-entry your way to equal application of the law. Transparency and accountability are good things, but they are bolsters for the honor system.
Democracy runs on the equal application of the rule of law.
And how exactly do we maintain equal application of the rule of law? We put honorable people in the position of applying the law. The honor system.
All democracies run on the honor system. Which is fine, as long as you’re aware of it and don’t let people who don’t respect that honor system get into positions of power
If it had more vram it could really cook. Qwen3.6-35b-a3b quantized to 3bits is genuinely usable for coding running on the ten year old pascal card with 16gb i picked up recently.
Idk, what happens when you do that? Does it apply traditional color filters and tweaks, or does it process the photo thru generative ai potentially altering it in ways beyond what you intended?
Well, if nothing else, posterity.
People have absolutely hated their loan applications being rejected since before ML was being used anywhere near it.
As evidence, let me cite the “computer says no.” sketch from 2004
Pot, meet kettle.
What do you call a reverse slippery-slope argument? “All images are edited, therefore ai editing is ok.”
Degrees of alteration matter, pretending ai images are the same as color retouching is dumb.
You’ve clearly not used photoshop recently, hey?
Generative features are all over Photoshop and other image editors. Removing a coffee cup off a table is a pretty small use of AI that nobody would really object to
he doesn't actually have any power to do anything here.
Landlords in nyc are doing business in nyc, which means the city can regulate them, does it not?
Analytic, but badly made.
Idk could you swap out an attachment and make them to something completely different?
I think of robots as general purpose, machines are specific purpose. When it works, we make it single purpose because that’s far far cheaper than general purpose.
At this point AI is a marketing term not an actual category
There is still a categorical difference between how they are being used. Specifically analytic vs generative. Generative AI (LLMs and image generators) are the ones people have issues with - pretty much nobody cares about ML processing for analysis.
I agree it’s too small a benefit to justify the investment, and I agree the bubble will pop. I just don’t think that means hardware prices become sane again for quite a while. I think if you half the price of a server GPU because demand from the big ai companies drops out, we’ll still have a shortage - it’ll just being going into commodity data centers to run open weight models.
I’m not assuming that at all, I’m responding to someone suggesting we’ll be able to run 1T models on phones in 2-ish years.
I absolutely agree that models are going to advance on to “edge” hardware over the next few years by becoming small + specialized.
* Software inference optimizations
Absolutely. I'd be surprised if they couldn't 2x performance in the next year. Still doesn't make a 1T model fit on your phone.
* Heavy quantization
I think this is a dead end if you're trying to fit a 1T model into a phone. Makes much more sense to train a model that's designed to be small, than train a model that's smart and then quantize it into stupidity.
* Chips with hardcoded transformer architecture
Totally, this will probably work great. Now good luck booking fab time any time in the next 2 years.
* Much cheaper HBM
Totally, this will probably work great. Now good luck booking fab time any time in the next two years.
* Much sparser models - 1T total with ~1-10B active params e.g.
Fewer active params helps with the speed of token generation, but if the whole model doesn't fit into ram it doesn't solve the issue of having to constantly stream portions of the model from disk to ram.
* Not to mention - 2 years of today's frontier models writing RTL and kernels at superhuman levels.
IMO this is a delusional myth-making idea being sold to us by ai companies. Machines that generate output based on statistical averages won't generate genuinely new ideas. They can help us try out ideas faster, but they're simply not capable of the kind of creativity and understanding required to push a field forward, except incrementally.
Even if the bubble pops and anthropic and openai et al implode - genie doesn’t go back in the bottle. The usefulness of LLMs for coding is proven, and a chip in a datacenter running 24/7 is always going to be more valuable than in a personal device running occasionally.
That doesn’t change until production capacity exceeds the datacenter demand. When that happens, they’ll start selling them down the market until it eventually reaches phones and toasters and whatever. But not in two years.
First off the math doesn’t math. Datacenters are willing to pay $50k for a single high end GPU. If you have unlimited capacity, yeah sell millions for $100 a pop or $10 a pop or whatever the bom cost of a phone GPU would be - but if you have limited capacity, you’re gonna sell all of that to the customer who is willing to pay the most PER UNIT.
Second off, this doesn’t work from a power consumption standpoint. When I run qwen3.6-35b, a far smaller model than op is suggesting, power usage spikes to 150-200W during inference. To fit a 1T model in the palm of my hand, the amount of processing required doesn’t fit the amount of power available.
Now I’m not saying this will never happen - there are some great leads, e.g. burning models directly on to a chip - but op’s scenario is definitely not happening in two years. Maybe 5, a lot more likely 10, unless of course local ai is made illegal
Unless there are major improvements to how much hardware it takes to run a 1T model, this is deeply unrealistic. First because why release hardware that puts your biggest customers (data centers) out of business. Second because as I understand it the data centers have bought up all the high end chip production capacity for at least the next year and unless the bubble pops that'll continue for a while.
It totally does follow the mold of social engineering, but LLMs aren’t part society, which is why it seems fundamentally different to me.
Anyway, agree with what you see saying - this is well worth a payout, embarassing they haven’t
I don’t think it counts as social engineering if it’s exploiting an llm, we might need a new word. Prompt injection doesn’t cover it, because it’s not about a malicious prompt.
I’m thinking some play on highjacking. AIjacking? Agent-jacking? Claudejacking?
…? I said why.
the load on your machine is gonna make doing other stuff while it’s running painful
Is your question about something else?
Gamers complaining about disc less games despite that problem pipeline and waste.
Tbf the issue is the user-hostile parameters of buying a diskless game. Most people would be happy to download their games if they could back them up to a thumb drive and never get locked out of them and sell the game when they’re done with it.
Seems deeply tangential, but it’s not. Blaming people for wanting a physical thing because the alternative is being further abused by a corporation - that’s a miss. Be mad at game platforms for not offering real ownership in whatever the most climate-friendly way possible. Be mad at governments for not forcing companies to cost in the negative externalities of their business.
Yes well my mouth was full of breakfast so the toast was censoring me.
Incidentally, taking down a domain used for short links doesn’t prevent speech or publication, since they have about 20 other domains that the same info is available at. Like how knocking over a newspaper box doesn’t censor the paper. So, by your own definition this isn’t censorship. Which is weird because it probably is censorship. Almost like your definition is bad.
Toss the rtx into a cheapo optiplex or thinkcenter, and run it headless - the load on your machine is gonna make doing other stuff while it’s running painful. Plus that frees up the rest of your vram.