HN user

clusterhacks

710 karma
Posts0
Comments208
View on HN
No posts found.

On the one hand, sure, why not have a default install throw a bunch of bells+whistles via skills and extensions.

But I like pi precisely because it is so minimal. I want understand and work around the simplest possible agentic coding setup, find the sharp edges, maybe even improve my prompting ability. And doing all three with a locally hosted LLM.

At some point, if I don't understand the foundations, am I just punting on actually thinking about what I'm doing?

Of course, making individual choices about how to do agentic coding are precisely just making individual choices. People should do what makes them happy and productive.

This may be too naive, but I created a user on my linux box who doesn't have very many permissions. Then I sudo to that user, use firejail to start pi in a dev project directory, and let it have at it.

My projects are usually very limited with respect to external dependencies and that is part of prompts or markdown files describing various project goals, plans, and current state.

My operating theory is that this probably won't get my systems borked. I wasn't patient enough to dig deeper.

FTA's conclusion:

"Is this decline a distinct change from the recent behavior of the labor share in the U.S.? Along the two key dimensions we investigate, our answer is no. <later> ... and they provide little evidence that it will evolve differently from past episodes."

This conclusion seems to be against "this time is different" arguments. Should we be generally encouraged by similarity to past declines pre-2000 or bearish and think that there is more drop to come like the 2000-2007 and 2007-2019 periods they graph out?

I guess there is no way to predict other than check back in after time passes.

Gotcha. I'm past the point of having any confident thoughts about what happens to their share price at IPO.

What about the idea that there is a high likelihood that the potential share price for OpenAI and Anthropic are both going to be pretty divorced from a rational market price for either?

I used to agree with you but now do not. I now think the floor for this market is probably no worse than the annual revenue of cell phone plans in the US market. So say, $250 billion.

Now, that probably doesn't justify the valuations and hype being thrown around, but I think it gets at a real revenue number.

I also don't know how that number fits into the funding rounds already raised and VC dreams of IPOs for these two.

This isn't coming from deep analysis on a verifiable source, but I started asking people in my social circle (includes white-collar and blue-collar folks) about their LLM use. The biggest surprise in 2026 for me was that almost all of these people told me about regular (and sometimes sophisticated) use.

A more intriguing observation - I work on the side with high school students and have two college kids of my own. Their LLM usage (and their peers) is much, much lower than expected . . . that's a little counterintuitive given "popular" perceptions I read.

--what this means for the valuation of the AI companies

Probably nothing. Most users have no idea what an LLM is or how it runs. Anecdotally speaking, I see many LLM users default to whatever their day job provides to them. And even slightly more sophisticated users seem ok with paying for their openai or anthropic subscriptions.

Maybe we will see a small but dedicated group of open weight model users who prefer local llm, but everybody else will just consume from the big providers? The scenario might look something like OS choices today - a small, committed group of Linux users vs the vast majority of other users running Windows, MacOS, or Chrome?

Appreciate the anecdote and your other comments on HN. But I strongly suspect you are incredibly atypical based on your background and previous work experience in ways that would tremendously down weight the probability that any part of your experience with recruiters would apply to even above average engineers.

Check out the Taulbee survey results:

"In 2023–24, Bachelor’s degree production fell 5.5% compared to the previous year across CS, CE, and I departments. Among departments reporting both years, the decrease was 4.3%. Despite this drop, production remains well above pre-pandemic levels and reflects continued strength following the post-2020 rebound. CS saw a 7.4% decrease and CE a 13.3% decrease."

But it also looks like enrollment in CS programs increased in 2024/2025:

"U.S. CS departments reported an increase in new majors per department of 12.8%"

  https://datavisualization.cra.org/TaulbeeSurvey/CRA_Taulbee_Survey_Report_2024.html#Bachelor%E2%80%99s_Program_Production_and_Enrollments

"I personally dropped $20k on a high end desktop . . . "

This is where I think current hackers should be headed. I grew up with lots of family who were backyard mechanics, wrenching on cars and motorcycles. Their investment in tools made my occasional PC purchase look extremely affordable. Based on what I read, senior mechanics often have five-figure US dollar investments in tools. Of course, I guess high quality torque wrenches probably outlast current GPU chips? I'd hate to be stuck making a $10K investment every 24 months on a new GPU . . .

I have been renting GPU resources and running open weight models, but recently my preferred provider simply doesn't have hardware available. I'm now kicking myself a little for not simply making a big purchase last fall when prices were better.

--> I can spot a person's social media app of choice is in 5 minutes.

I find this sadly hilarious. What are the current tells you see? I'm similar in that I read a lot of HN and don't have other social media accounts. But I couldn't even guess at what a person's preferred social media is.

Very cool - thanks for the info.

That you are writing AI agents for a living is fascinating to hear. We aren't even really looking at how to use agents internally yet. I think local agents are incredibly off the radar at my org despite some really good additions as supplement resources for internal apps.

What's deployment look like for your agents? You're clearly exploring a lot of different approaches . . .

Good grief. I'm here cautiously telling my workplace to buy a couple of dgx sparks for dev/prototyping and you have better hardware in hand than my entire org.

What kind of experiments are you doing? Did you try out exo with a dgx doing prefill and the mac doing decode?

I'm also totally interested in hearing what you have learned working with all this gear. Did you buy all this stuff out of pocket to work with?

Wow, thanks for the link to Texerau. I had no idea a pdf was floating around and have wanted this book for some time. You video looks interesting, especially the part around Ronchi and Focault testing. I have 'Understanding Focault' but have to admit that reading it doesn't give me confidence.

One question I always think about is how much time and effort a "one-time" mirror maker should plan on making to exceed the quality of a generic 8" or 10" F/5-F/7 available from the Chinese mirror makers.

Zambuto seems to imply that whatever magic happens for his mirrors might be in very long, machine driven polishing to smooth out the final surface imperfections that cause scatter. With his retirement and with few mirror makers in the US, it seems like options for buying "high end" mirrors in the 6"- 10" size are very limited. I have been debating an 8" F/7 and would love to just purchase a relatively high quality mirror, but most of the mirror makers seem more taken with significantly larger mirrors.

Watch your local craigslist or facebook marketplace. With a little patience, you will probably find a good 8" or 10" dobsonian at a great price. I picked up a lovely 8" dob for less than $200. Most of the generic 8" F/6 dobsonians seem pretty decent.

Or check your local library. It may have a smaller Starblast table-top dobsonian you can check out - I did that when traveling once.

Whatever you do, do NOT buy a small cheap refractor on some flimsy mount. They are mostly awful.

Sorry, I don't much track or keep up with those specifics other than knowing I'm not spending much per week. My typical scenario is to spin up an instance that costs less than $2/hr for 2-4 hours. It's all just exploratory work really. Sometimes I'm running a script that is making a call to the LLM server api, other times I'm just noodling around in the web chat interface.

No, I don't blog. But I just followed the docs for starting an instance on lambda.ai and the llama.cpp build instructions. Both are pretty good resources. I had already setup an SSH key with lambda and the lambda OS images are linux pre-loaded with CUDA libraries on startup.

Here are my lazy notes + a snippet of the history file from the remote instance for a recent setup where I used the web chat interface built into llama.cpp.

I created an instance gpu_1x_gh200 (96 GB on ARM) at lambda.ai.

connected from terminal on my box at home and setup the ssh tunnel.

ssh -L 22434:127.0.0.1:11434 ubuntu@<ip address of rented machine - can see it on lambda.ai console or dashboard>

  Started building llama.cpp from source, history:    
     21  git clone   https://github.com/ggml-org/llama.cpp
     22  cd llama.cpp
     23  which cmake
     24  sudo apt list | grep libcurl
     25  sudo apt-get install libcurl4-openssl-dev
     26  cmake -B build -DGGML_CUDA=ON
     27  cmake --build build --config Release 
MISTAKE on 27, SINGLE-THREADED and slow to build see -j 16 below for faster build
     28  cmake --build build --config Release -j 16
     29  ls
     30  ls build
     31  find . -name "llama.server"
     32  find . -name "llama"
     33  ls build/bin/
     34  cd build/bin/
     35  ls
     36  ./llama-server -hf ggml-org/gpt-oss-120b-GGUF -c 0 --jinja
MISTAKE, didn't specify the port number for the llama-server
     37  clear;history
     38  ./llama-server -hf Qwen/Qwen3-VL-30B-A3B-Thinking -c 0 --jinja --port 11434
     39  ./llama-server -hf Qwen/Qwen3-VL-30B-A3B-Thinking.gguf -c 0 --jinja --port 11434
     40  ./llama-server -hf Qwen/Qwen3-VL-30B-A3B-Thinking-GGUF -c 0 --jinja --port 11434
     41  clear;history
I switched to qwen3 vl because I need a multimodal model for that day's experiment. Lines 38 and 39 show me not using the right name for the model. I like how llama.cpp can download and run models directly off of huggingface.

Then pointed my browser at http//:localhost:22434 on my local box and had the normal browser window where I could upload files and use the chat interface with the model. That also gives you an openai api-compatible endpoint. It was all I needed for what I was doing that day. I spent a grand total of $4 that day doing the setup and running some NLP-oriented prompts for a few hours.

I ran ollama first because it was easy, but now download source and build llama.cpp on the machine. I don't bother saving a file system between runs on the rented machine, I build llama.cpp every time I start up.

I am usually just running gpt-oss-120b or one of the qwen models. Sometimes gemma? These are mostly "medium" sized in terms of memory requirements - I'm usually trying unquantized models that will easily run on an single 80-ish gb gpu because those are cheap.

I tend to spend $10-$20 a week. But I am almost always prototyping or testing an idea for a specific project that doesn't require me to run 8 hrs/day. I don't use the paid APIs for several reasons but cost-effectiveness is not one of those reasons.

All those choices seem to have very different trade-offs? I hate $5,000 as a budget - not enough to launch you into higher-VRAM RTX Pro cards, too much (for me personally) to just spend on a "learning/experimental" system.

I've personally decided to just rent systems with GPUs from a cloud provider and setup SSH tunnels to my local system. I mean, if I was doing some more HPC/numerical programming (say, similarity search on GPUs :-) ), I could see just taking the hit and spending $15,000 on a workstation with an RTX Pro 6000.

For grins:

Max t/s for this and smaller models? RTX 5090 system. Barely squeezing in for $5,000 today and given ram prices, maybe not actually possible tomorrow.

Max CUDA compatibility, slower t/s? DGX Spark.

Ok with slower t/s, don't care so much about CUDA, and want to run larger models? Strix Halo system with 128gb unified memory, order a framework desktop.

Prefer Macs, might run larger models? M3 Ultra with memory maxed out. Better memory bandwidth speed, mac users seem to be quite happy running locally for just messing around.

You'll probably find better answers heading off to https://www.reddit.com/r/LocalLLaMA/ for actual benchmarks.

whole point of the time compression is to spread the grades out

I suspect that is true for standardized tests like the SAT, ACT, or GRE.

I suspect in classroom environments that there isn't any intent at all on test timing other than most kids will be able to attempt most problems in the test time window. As far as I can tell, nobody cares much about spreading grades out at any level these days.

Sure, but that answer doesn't address the questions of the value of time limits on assessment.

What if instead we are talking about a paper or project? Why isn't time-to-complete part of the grading rubric?

Do we penalize a student who takes 10 hours on a project vs the student who took 1 hour if the rubric gives a better grade to the student who took 10 hours?

Or assume teacher time isn't a factor - put two kids in a room with no devices to take an SAT test on paper. Both kids make perfect scores. You have no information on which student took longer. How are the two test takers different?

I share your paranoia.

My kids use personal computing devices for school, but their primary platform (just like their friends) is locked-down phones. Combining that usage pattern with business incentives to lock users into walled gardens, I kind of worry we are backing into the destruction of personal computing.

I started my career as a software performance engineer. We measured everything across different code implementations, multiple OS, hardware systems, and in various network configurations.

It was amazing how often people wanted to optimize stuff that wasn't a bottleneck in overall performance. Real bottlenecks were often easy to see when you measured and usually simple to fix.

But it was also tough work in the org. It was tedious, time-consuming, and involved a lot of experimental comp sci work. Plus, it was a cost center (teams had to give up some of their budget for perf engineering support) and even though we had racks and racks of gear for building and testing end-to-end systems, what most dev teams wanted from us was to give them all our scripts and measurement tools to "do it themselves" so they didn't have to give up the budget.

I was playing around with Qwen3-VL to parse PDFs - meaning, do some OCR data extraction from a reasonably well-formated PDF report. Failed miserably, although I was using the 30B-A3B model instead of the larger one.

I like the Qwen models and use them for other tasks successfully. It is so interesting how LLMs will do quite well in one situation and quite badly in another.

I never looked at short videos and couldn't understand how my friends and family would open their phones given even a minute of quiet time for these things. I viewed some Youtube shorts (maybe the least effective short video provider in terms of content and the recommendation algorithm?) and was shocked at how easy it was to burn time looking at crap. The experience really opened my eyes about how a person can be pulled into endless viewing.

I think the crack house comparison is entirely appropriate. The brain is weird . . .

Gemini 3 8 months ago

That's interesting. While I suspect the pricing will lean heavily into enterprise sales rather than personal licenses, I personally like the idea buying models that I then own and control. Any steps from companies that make that more possible is great.

Gemini 3 8 months ago

I wish I could just pay for the model and self-host on local/rented hardware. I'm incredibly suspicious of companies totally trying to capture us with these tools.

I was a light or social drinker for decades. Probably 3-5 drinks per week.

In November of 2024, I decided to avoid alcohol as a personal experiment - no GLP-1 medications involved. I have not consumed any alcohol since.

After 3-4 months, my interest in alcohol seemed to really fall off a cliff. I joked with friends that I was going "dry in 2025", but I am now more seriously considering taking 2026 off from alcohol as well before making a decision about whether to add alcohol back into my diet.