HN user

digdugdirk

2,667 karma
Posts7
Comments715
View on HN

What model? And what hardware do you run it on?

I find these style of models are great, but fail hard, and fail randomly. I'd be hesitant to use it for a daily driver, but I'm using dual 3060s, so it's not like I'm quantizing a frontier model here.

How do you find the overall experience? And do you have any special sauce or recommendations for going this route?

In regards to people with reasons to illuminate several sq km at once - I'd bet that major metro areas would see a massive savings in electricity/maintenance if these were deployed over a metro region. Whether that is more than the cost of a satellite? Who knows, it's still fiction until these people try it out. But it's at least an interesting use case.

Other fields of engineering usually have a regulated licensure, upon which they can call themselves a Professional Engineer. This gives them the ability to make final approval/sign-off on designs and technical reports. It's most common in civil engineering, where a PE license is required for all publicly funded projects (and most privately funded ones as well, due to local/regional/national regulations) to be approved.

This license requires the holder to uphold code of professional ethics, and makes the engineer themselves be personally responsible for the safety and viability of the design itself. Losing a PE license is rare, but it does happen. The industry board (usually a regional board) can also discipline/reprimand engineers who fail to meet the professional standard - rubber stamping projects, personal misconduct, etc. Losing a license is a huge deal, but even reprimands can have a serious negative impact on someone's career.

In the industry the previous commenter works in their hypothetical would absolutely meet the bar for discipline or reprimand.

Really? It's what he did after he spent nearly a third of a billion dollars to influence an election (which doesn't include the incalculable value of his influence via Twitter!) that makes him a Bond villain.

It's Chesterton's Fence[0] in action - Elon took a self-proclaimed woodchipper to the United States government without understanding what he was destroying and without concern for the repercussions to society. He used that influence to benefit himself and his companies directly. It's not sharks with laser beams, but it's far more impactful and the world will living with it for decades to come.

[0] https://en.wikipedia.org/wiki/G._K._Chesterton#Chesterton's_...

It's an inherently different thing. Fibre is infrastructure. It'll still function decades later. GPUs at this scale are consumables. I've heard 3-5 years lifespan before they fail out. This might be a low estimate, but even if you double it - they're 50% of the cost of datacenters. We're flushing entire countries worth of economic value down an Nvidia shaped toilet.

I think part of it is the feeling of false understanding that comes from using llms regularly. They let you operate at a higher conceptual level, and they paper over enough of the actual details that your conceptual model might not actually be correct.

I'm a mechanical engineer by training, and have similar vibes with the similarities I see between llm training and metallurgy. I could probably put together a formal concept for these vibes at this point, but is there actually a "there" there? I have no idea. And it would take me years to actually dive in and learn everything to gain the deep understanding that would be required to know if I'm just experiencing my own brand of AI psychosis or not.

It's a brave new world, that's for sure.

Ding ding ding! This is a huge reason. Being able to bootstrap a domestic weapons manufacturing base is a massive win for any country these days. South Korea is one of the few countries that are both willing and able to do so with high quality modern materiel.

I've always wanted to figure out how to implement a cooperative source license. Something like, you're allowed to do what you want with it, but any derivative work requires the same license, and X% of any income goes to the cooperative?

Not sure how it'd work, but there's absolutely a niche for a privacy focused data cooperative out there.

Claude Opus 4.8 2 months ago

Do you have a collection of these benchmark apps saved anywhere? I'd be particularly interested in seeing the relative cost differences between different models in a use case like this.

It's always been this way. America has just been able to coast on being the only remaining major economy after WW2, and exploited the rest of the world instead. That exploitation of the rest of the globe has been mostly optimized now, so those shareholder returns are now coming at the expense of the 90% of Americans who aren't sitting at the table.

Parts of these cities worse off than the third world? Have you been to a third world country? Or Seattle, for that matter?

The commonly scapegoated cities in the United States are not experiencing third world conditions. Appalachia is experiencing third world conditions. Hollowed out rust belt cities in the Midwest are experiencing third world conditions. These areas are not run by lefty politicians. The United States has a systemic problem, not a local one.

And yes, the systemic problem is that there are a tiny number of ultra wealthy people with wildly outsized influence on the government of the United States, doing everything they can to reduce the amount they need to pay in taxes while simultaneously ensuring they extract the maximum amount of profit from the US government's wildly excessive expenditures.

That looks like a really nice hackathon! That said, the fact that they probably had a majority of the best NixOS developers in the world under one roof and they weren't solely focused on NixOS error messages is borderline criminal...

It doesn't have to do any thing interesting - it's completely fascinating all on it's own. If you understand anything about the math and science behind LLMs, you'll understand that this is an achievement worthy of sharing to a community like HN.

That being said, small models like these have plenty of use cases. They allow for extra "slack" to be introduced into a programmatic workflow in a compute constrained environment. Something like this could help enable the "ever present" phone assistant, without scraping all your personal data and sending it off to Google/OpenAI/etc. Imagine if keywords in a chat would then trigger searches on your local data to bring up relevant notes/emails/documents into a cache, and then this cache directly powers your autocomplete (or just a sidebar that pops up with the most relevant information). Having flexible function calling in that loop is key for fault tolerance and adaptability to new content and contexts.

Its cool. Enjoy it.

Mojo 1.0 Beta 3 months ago

It does almost seem like they're trying to recreate the Nim programming language in this regard.

How does that get integrated into the scoring system? I'm imagining a scenario where a cheaper model may get close, but only needs a small follow up to get the desired result. How would this score in comparison to a larger model that got it right the first time - even if it may have been much more expensive overall?

Interesting! I've been thinking about how to create a similar type of evaluation system for myself. How do you handle tweaks to agentic tasks? Say that a model gets pretty close to what you want, so you just need a quick follow up prompt to the original response?

From someone who definitely doesn't fully understand what you made, this looks really cool!

I'm seeing some functionality that seems like it could replace some personal services I currently host via my tailscale network. Am I understanding this correctly? If so, do you have a feel for what the performance implications would be?

I don't understand why you're getting downvoted? Of course an LLM will return the answer to a widely known and commonly cited riddle that exists because of the far more rigid societal gender norms 50 years ago?

LLMs are just statistics based on vibes. Switching the gender of the character in the beginning of the story, but keeping all else identical is going to be a huge signal into the noise, and that response is going to be wildly likely to occur.