HN user

eurekin

1,020 karma

"thoughtful considerations about practical applications"

Self hosting, SBCs, AI/Vision/LLMs

Posts0
Comments622
View on HN
No posts found.
Qwen 3.8 4 days ago

It's a mcp, so connects quite easily to agents. With mcpo, I also connected it to open-webui (which has better support for OpenAI style tools/functions). Used it in claude code with that mcp plugin set-up too. Only ever used it for managing homelab information, but it met initial expectations. 27b is a great model, if grounded. The query about physical hosts and routing... I haven't found a single hallucination (altough Codex 5.6 as a reviewer mentioned something was wrong with some parts, and those were exactly the never properly documented ones. Codex/gpt had extra knowledge, because it was the conversation I used to set it up).

Qwen 3.8 4 days ago

With 3.6 27b, I just stopped changing local models and started tinkering with things on top (like mem0). Feels genuinely useful and more than a toy

Disclaimer, I only use it to grow the "knowledge hub".

It's a single git project at my $USER home, that is referenced in global memory. It contains as much information about work things, as possible, to be productive.

I found that, if I allowed Claude to create the notes, it actually became more and more useful, but without the guideline, I just could barely get through reading it manually.

I'd never publish anything with such origin.

That's where the annotation in plannotator helps.

I'm asking to scan projects on gitlab, go through some docs to find more grounding material, write a subarticle (in the same style), scan logs on the test env, issue some curls, etc.; until the whole article is digestible - in the "backing knowledge graph" department.

Oh, I know exactly what you mean. I use plannotator with claude a lot and have much better time, since I asked for a specific styleguide.

I used "CD era MSDN reference and Raymond Chen blogging style" as a starting prompt for the styleguide and my work ability to digest AI plans raised a lot.

Couldn't recommend it more. Humble, insightful and respecting the reader

A shift is happening among major AI labs, who are becoming increasingly skeptical of endless parameter count and training data scaling

I'm pretty sure it's mostly due to the training data quality. No idea, why this never gets mentioned in those discussions.

It was obvious right from the get go, that the scaling law just enabled some abilities, that were described by the underlying data and allowing the ANN to abstract it in the latent space.

Noticed few cliffs. Sometimes it was a spurious stop (had to write "go on" or "continue" to restart), othertimes it was randomly saying: "Oh the user wants [the thing we already resolved]" and goes back in history. Cleared all out on fp16

The model is running so hot, that it shoots past the goal and starts looping

later:

My latest experiment was setting up vLLM (the gold standard for production and concurrent serving) and even with an NVLink (175GBP) and tensor parallelism turned on, it was 3 tokens/second slower than llama.cpp during generation for an equivalent setup.

In all my tests, getting vllm to run is worth it. It was the single biggest thing, that helped for looping issues, agents going whack and losing focus on the task, long context being essentially useless.

FP8 model, unquantized cache in vllm an you have a league better overall experience, with any other stack I tested. Then, you can actually focus on using the model for other things and stop tinkering with settings.

CrankGPT 1 month ago

I'm still sour they had only one toast in, in a two slot toaster

I kept getting recipes with "that one ingredient", which was either a major PITA to source or produced too much waste, even from a real world dietician consultation. Example, use 1/4th of a pumpkin for something. Those were good recipes, in terms of macronutrient composition, but doesn't work long term due to logistics.

I'm years after that strict diet needs, but that itch of fixing or easing some parts of the process stayed.

I keep finding more and more usecases for Q3.6 27b (same league) and the best performance is, when answers to my question is already in the context.

The moment I'm trying something open-ended or ambitious, Claude/ChatGPT clearly take you to the goal quicker.

For things, where there's a way to build a knowledgebase though, the local llm definitely can be a true contender. Plus, having a big context and no worries about filling it over and over - you can get quite far.

I'm writing this, literally in between cooking a pasta, that the local llm ordered products for me online. I've built a grocery shopping skill, so that it roughly knows what I have in fridge (losely), my last 10 representative orders (general preferences plus rich info about shops and skus around me) and actual real-time in stock info. The last part has been my personal pet peeve for every product that promised cooking ingredient delivery (that is not packaged specifically for that).

This is what has been promised to us by every big tech company with an agent, and now a local llms actually solved that for me fully.

Yeah, it won't. SWE here with a sidegoal to tackle the deployment side through various means (homelabbing, grabbing sre/cloud/observability tasks at work).

The biggest observable improvement in my post and pre ai development is the ability to tackle two projects at once, if the agent is on track. If not, and I have to do a deep dive to debug, it basically regresses to plain old everything like before.

It feels now like an alternative timeline, one which performance optimisations were first and foremost still. Sometimes I fantasize, thinking how would our current development ecosystem look like, if we never abandoned the "be very vigilant with all resources you use" approach, that includes the whole webdev liftoff, where we ship a few hundred mb chromium engine for a dock app

PyInfra 3.8.0 3 months ago

On my homelab. It really feels like a dream come true for my usecase. No more puppet agents. No more declarative syntax, that you have to work around to do basic imperative ways. Or use a module, that stopped being maintained 3 years ago. Just plop a file here and there through ssh.

Never understood any appeal of a screen inside a car:

1. Reflections make you tilt, just to make some pesky highlights go away. Even if they are angled properly, there's always something (like a sun reflected by a watche's face) what causes nuissance at any angle

2. Car can go from a tunnel to a sunny valley in few seconds. That's 5 to 8 stops of dynamic range difference, that a human eye is easily designed to handle. Auto adjusting screen brigtness is never as bright as necessary in sunny conditions. Even if it were, it would be a significant battery drain and an element, that heats the cars interior already unnecessarily.

3. You don't have pure blacks in many of them, so that annoying halo at the corner of the eye is often present. You can solve it with an OLED, but those are even worse in bright daylight

4. All of the usually mentioned tactile feedback facts - you can reach with your hand to a AC knob, feel it's current set by finding the bulge with a finger and gently turn exactly how you want them. Zero lag, no eye contact necessary at all (keep that on the road!), instant feedback. Nothing that any screen can ever give.

5. Biggest gripe of all - modality. I think that there were some high ranking studies done early in design exactly against this type of input for high risk applications. Modality is the biggest enemy of discoverability and throws extra delays into otherwise instant input.

6. If you use a LCD variant, they interact with sunglasses polarity filter and, at some orientations, can be blocked altogeter. As you often use sunglasses exactly, because you want to see the road the best, it's contrary to the main objective of the control again.

7. Refocusing. If you can use a tactile control, with a good feedback, you're freeing your eyes from the need to adjust it's lens to focus from far to near to far again. Not many people are aware, that this is even happening, and can lead to overestimating your ability to keep engaged attention on the road.

I'd pay extra for a zero screen variant in a jiffy. Had I ever need to use a screen, I would've put my phone in a holder instead.