HN user

drillsteps5

356 karma
Posts1
Comments167
View on HN

A decent gaming machine perfectly doubles as your friendly local inference server. Just start llama-server with the model of your choosing and start chatting with it through its Web interface or connect any chat completion-compatible client (agentic or not) which will use REST to send requests and receive responses. From any device on your network. Voila.

I honestly don't get the hostility against local models in this thread (and in some other threads recently).

I haven't seen anyone make an argument they are as good as SotA (OpenAI, Anthropic). It's just they are approaching state where they are "as good" for some _limited_ set of use cases. Which will allow us to resolve 2 primary issues with these SotA models: privacy and vendor lock-in. Plus, they're very useful for education purposes, you get to explore what things looks like under the hood, play with various models, tools, maybe put something simple together yourself.

You get Macbook - great. You got gaming rig with a decent GPU - great (set it up as a dedicated server that you connect to through simple REST).

What exactly is wrong with any of that?

As another commenter said a "model" is a file (or group of files, there's multiple formats available; GGUF format is all in one file for example). You download it to the hardware of your choice (ie your own desktop with NVIDIA GPU). You run the inference engine (llama-cpp, ollama,lm studio etc) and tell it where the downloaded model is and it runs inference (so you can start chatting with it, or run agents).

"Open weights model" means the developer made the model available for everyone for free. You can download it from huggingface.co for example and do whatever you want with it.

Why "open weights" and not "open source"? Because the "source code" for LLM would include things like training data, training methodologies and tools, so that you can do the training and produce the model (files) yourself. That would be like compiling from source code. Which is not done with these models, it's company's know-how, they only share the end result.

It's more analogous to "freeware" which is what we traditionally call freely distributed binary executable files. But people started calling them "open weights" instead and the term stuck.

Open weights models are cheap in the context of the article (when you run inference in the cloud) because they are free. When I pay for inference for running DeepSeek open weights model I only pay the inference service provider for compute/memory/storage/network throughput. The model itself is free, the developer isn't getting a dime.

Developing these things is NOT free, there's a lot of labor, hardware, compute/memory/storage/network that goes into that. Who's paying for all this? Chinese govt? Developers themselves? What's the revenue model here?

I absolutely LOVE ability to either run them locally or access inference providers on the cheap, but having a hard time understanding the financial side of this.

Aside from googling "how to download and run open weights model" check out localllama (yes 3Ls) subreddit. Huggingface.co is where many of them are published.

There's many providers that run open weights models and give you access. Many decent open weights models cannot be run on consumer-grade hardware (DeepSeek, GLM, many others).

I'm looking forward to the trial where Anthropic will have to disclose sources of their training data, and then explain why they are entitled to charging customers for using regurgitated training data but Alibaba which trains their models on Anthropic's models are not.

Should be fun.

Edit: clarification

Spying on people and charging different prices to different people for the same thing (i.e. surveillance pricing) is not.

They are not "spying", they are _legally_ using various data _legally_ collected on their current and prospective customers to set the pricing. Do you ever read T&C of any online services you're using? So how is this illegal?

Can you imagine the mayhem if companies just straight up knew salary information for all of their customers?

I filled in FAFSA and CSS Profile for multiple colleges my child applied for last year. So yes, that's exactly how some industries work. Sucks for me, the customer. Not illegal.

This does not sound right to me. It would be correct if all companies and their hiring managers had the same requirements/looking for candidates with the same qualifications - but they're not.

Obviously as a hiring manager you're looking for a hard working individual with a number of successfully completed projects and glowing referrals from multiple places of employment, but you're also looking for a person with expertise in particular technologies/industries/whatever other areas of expertise. To a large extent requirements for each role are unique, however many do have some overlap.

So being rejected from one position might simply mean there's a misalignment between what the company is looking for and what the individual has. Which might not be the case with other companies.

So if we're seeing increasing number of candidates being consistently rejected at multiple places the question "why" is a valid one.

I heard that "surge pricing" , or "surveillance pricing", is wrong, but don't understand why. Shouldn't the service provider be free to charge whatever they want for the service they provide? And the consumer should be free to choose between multiple service providers to find the optimal value (price/convenience/features/whatever)?

In other words, the root issue is not "surveillance pricing" but lack of competition.

If you think there's no way for Uber and Lyft to infer anything about your purchasing power/habits when you install an app running on your primary computing device with generous privileges, logged in with your unique phone number/email... You might be unpleasantly surprised

Not a good take. It wasn't work that was invented recently, but ability to sustain yourself by performing some repetitive (and not always meaningful or productive) actions at pre-defined time periods (like 5 days a week x 8 hours a day). Which does go back to Industrial Revolution and even more recently to Ford Motors and similar enterprises and business models. If you were to ask a hunter-gatherer or a nomad or a slave or even a trade laborer (ie a shoemaker) in pre-industrial times, they'd tell you it's a pretty sweet deal.

No worrying where the next meal will come from, if there's going to be enough crops for the next few months, or if you'll be able to find an animal to kill large enough to feed you but but not large enough to kill you, if you can protect yourself against predators, or aggressive neighboring tribes, if you will be able to find/maintain a shelter good enough to protect you from the elements, esp in extreme cold or hot climates. If you'll be able to make enough shoes to earn enough to sustain yourself and the family, while competing with other shoemakers for a limited demand and limited materials, and million other things.

In fact, quantitative studies revealed that the average adult hunter-gatherer spent about 20 hours a week at hunting and gathering, and a few hours more at other subsistence-related tasks such as making tools and preparing meals (for references, see Gray, 2009). Some of the rest of their waking time was spent resting, but most of it was spent at playful, enjoyable activities, such as making music, creating art, dancing, playing games, telling stories, chatting and joking with friends, and visiting friends and relatives in neighboring bands.

I'm surprised the author didn't add that they also didn't suffer from obesity or dental cavities or cancer (which is mostly because living past 30 wasn't invented until like 14th century).

Even if you were a "point" (an endpoint assigned to the node) you still had to set up the software and (in the mid-to-late 90s at least) set up a modem to call your node to upload/download. And sometimes you had to set up repeated dialing until you got through because the node could be busy (some nodes doubled as BBSs), or connection could be bad and it'd had to retry etc. Wasn't an easy task, so it served as a sort of a filter so that most people on there were geeks.

Later on of course some nodes started distributing over the Internet so setting up a node became much easier (and I think there was a way for the node to allow multiple users read/write without even setting up a node/point at all).

Direct quote:

We have covered this math before. The $725 billion that Amazon, Microsoft, Alphabet, and Meta are spending on AI infrastructure in 2026 has to come from somewhere. For many companies, the somewhere is headcount. Not because AI replaced the work. Because the budget line got moved to a different row on a spreadsheet.

So "headcount" (because that's what we call people) is being cut because of crazy spending on genAI infrastructure (part of which btw goes to Mr Hwang's company). Or, if you're not a hyperscaler, crazy spending on tools/tokens. But no, you should NOT tie reductions in "headcount" to genAI. That's "irresponsible" and "lazy".

Did I get that right?

I've lived through both 2000 and 2008. They do happen. And typically not when everybody says there will be a recession, but when almost everybody finally agrees there won't be one.

Not that us plebs can do anything about it anyway... :(

"Enshitification" is not a new concept. A business should always be willing to make their product cheaper, even at the cost of quality, until the customers start turning away. Of course you need to be able to catch that moment early enough so that you don't lose too much market share to competition. But that will give you increased profits. The same with increasing prices.

On a side note, I'm curious as to how "600% increase in AI usage" is measured. Are their agentic workflows' bills skyrocketed 600% in the last 3 months? That would be in line with what other people using agents are seeing (costs are way higher than they expect/used to be). In that case, that would mean that LLM/agents are no longer necessarily cheaper than human labor, no?

Labor market data this week came out stronger than expected, even as large layoffs in IT continue to happen and IT job market continues to be very slow.

Canvas is back up as of Friday US morning for me (HS student's parent). My kid got a few panicked emails yesterday from the teachers but it looks like Instructure got it resolved quickly.

Canvas does provide a lot of value (all courses, teachers', students', and parents' contact information, all learning plans, schedules, room numbers, all grades, a lot of tests and assignments themselves, all upcoming assignments and deadlines, a lot of other coursework is in there, as are the final grades) but it shows that with external SaaS you might be one attack away from not only losing all that convenience but also in a world of hurt 'cause you lost all the data and now have to figure out how to proceed without the data and the system.

US high schools are in the middle of the finals, and seniors are getting ready for college (the transcripts to be finalized and sent out in a few weeks) so that was a scary timing.

I wasn't saying that this is the optimal solution (it clearly is not). I was saying that it makes perfect sense for both sides - HR has their work automated and candidates have better chance to be noticed - and therefore became a common practice in many places.

The well has been already poisoned, to survive you have to get in on the action.

Don't want to play this game? Make connections, set up the network, and use it to get/stay employed.

Even if you define Gen X birth years starting at 1965 (and not 1961 as some do), their oldest are already 61. So anybody between 55 and 61 is a Gen X, not a boomer. And for the 55+ group employment participation rate has been decreasing. Which does not mean that "boomers are retiring", it's Gen X's turn now.

t the same time, we're still seeing 55+ leave the labor force rapidly (~4M Boomers continue to retire per year, ~330k/month)

Youngest Gen X are 46, oldest are 65. Can we stop with this "retiring boomers" thing?

Ufff.

Those readings were the strongest since February 2020 (63.3%), signaling that more people were re-entering the workforce and helping to ease labor shortages in the wake of the pandemic. Labor shortages. In the wake of the pandemic. That ended 4 years ago. Ssuuure.

For workers 55 and older, demographic factors are key, with more individuals retiring, particularly since the pandemic. Got it. "Retiring": giving up on finding a job because nobody hires 50+ yo (either in the trades or white collar)

These findings reinforce a clear post pandemic trend: young men remain the most likely to be on the sidelines of the labor force. This underscores the need for policymakers and businesses alike to develop strategies to draw more of them back into the labor market—efforts that could help ease workforce shortages by expanding the overall supply of labor.

"Nobody wants to work"

As labor force participation among young people declines, fewer teenagers are gaining access to those formative early work experiences, and the essential skills that come with them, which have lasting implications for both individual career trajectories and the broader economy.

Have 2 teenage sons, at 14, 15, and 16 (last 2 summers) they were looking for a summer job (US, suburbs), retail, fast food... Anything, really. I made them apply to as many places around us as possible. Nothing.

It also aligns with widespread anecdotal reports that finding employment today is more difficult than it was a year or two ago.

"anecdotal" Who wrote this stuff???

Oracle has also owned JD Edwards since early 2000s which is in many large legacy companies (I think a lot of them are still in mainframes).

Oracle the company has not been about Oracle the DB server for 20+ years.

Oracle the company specializes in acquiring software, integrating it in their ecosystem, selling the installations, and living off the recurring licensing fees (NetSuite is one example).

They might drop it for end consumers but I doubt it.

It's such a small niche right now they do not even care if they're in their cloud. Enterprise users are the absolute majority of their user base revenue-wise.

However dropping the requirement might force them to change some things. Like in Azure-related stuff such as OneDrive where you have to design/build/test it behave differently if the user is not constantly logged into the Azure account. This means that they might decide to continue to force the Azure account and if they lose more of the end consumers so be it.

Unless they decide to separate Home and higher versions of Windows even more and drop the requirement for the home version users. But it might be more trouble than it's worth.

Enterprise is where the money is.

I very much would like to know how much of this presumably ordered (and backordered) hardware (RAM/SSD/.../wafers) is going to end up being released back to the market when the dust settles. I haven't seen any estimations but in order to put all this hardware to work the hyperscalers need to be building data centers at ludicrous speed. That should be appearing in construction data, jobs data, and many other places. Are we actually seeing any of that? Or is it all just based on the back-of-the-napkin math by Mr Altman and Co and they put all the money they got towards the future projects?

I actually printed this one out to give it a thorough read, something I haven't done in a while :) I'll probably comment on the author's blog later.

My first reaction is that it's a bit of oversimplification, LLMs have been on the scene for what like 4 years now? I doubt the effects will be visible when you look at the data.

But on the topic of inequality as a result of changes how value add is distributed between business owners and the labor - there's a very relevant book by Thomas Pickety called "Capital in the 21st century" (I believe he used the title of a similar work by another economist in 18th century on the same topic). He collected as much data as he could on income and wealth of individuals and groups for Western nations (some going back to like 17th century) and did some analysis.

In a nutshell, the share of profits going towards the business owners (simplifying, inherited capital) in the West had been increasing in the 18th and the 19th century, likely due to Industrial Revolution. Which in the US culminated in extreme inequality personified in Robber Barons (Carnegie, Rockefeller and Co). It was followed by economic and financial collapse, mass unemployment, and arguably, 2 world wars. It also resulted in creating various mechanisms designed to prevent these things in the future, such as anti-trust regulations.

After WWII (especially in the US) the distribution of value add between the ownership/capital and the labor changed so that the labor started getting larger and larger share, to the point that the economists declared that the capitalism solved the issue of inequality. That trend reversed around 1970s/1980s, which coincided with the invention of complex electronic computing and communication devices (and therefore inventions of the new business models). From that point on the share of profits going towards the business owners started increasing again, and that speed has been actually accelerating since the 2000s.

imho at this point US is basically where we were in the end of 19th/early 20th century. The individuals' names are different, the issues (extreme monopolization/concentration of capital) are the same. LLMs and other GenAI stuff are simply part of that trend.

Hopefully people at the power learned something from what happened 100+ years ago. But I kinda sorta doubt it.

Some people strongly prefer to have an external battery that can be swapped (as in pull a tab, remove the battery, and plug in another, which is fully charged). Older Thinkpads (T480 and earlier) had a second, smaller, battery inside that would keep the laptop running while the main (external) battery is being replaced.

I've been using various Thinkpads for 10+ years and have yet to use this feature. But hey, to each his own :)