HN user

imiric

8,470 karma

hn@${username}.com

my public key: https://${username}.com/pgp

Posts10
Comments3,236
View on HN

Creativity is connecting ideas from different domains and see if something from one field applies to another.

That's true. The question is whether the produced pattern has any value. LLMs are incapable of determining this, and will still often hallucinate, and make random baseless claims that can convince anyone except human domain experts. And that's still a difficult challenge: a domain expert is still needed to verify the output, which in some fields is very labor intensive, especially if the subject is at the edge of human knowledge.

The second related issue is the lack of reproducibility. The same LLM given the same prompt and context can produce different results. This probability increases with more input and output tokens, and with more obscure subjects.

The tools are certainly improving, but these two issues are still a major hurdle that don't get nearly as much attention as "agents", "skills", and whatever adjacent trend influencers are pushing today.

And can we please stop calling pattern matching and generation "intelligence"? This farce has gone on long enough.

Shipping is just a milestone. We all know that "AI" can produce code much faster than any human.

Productivity should be measured over time and take into account the cost of maintenance, reliability, amount of issues, etc.

never ask a model for confirmation or encouragement; but you can absolutely ask it to critique something, and that's often of value.

What's the difference? The end result is equally unreliable.

In either case, the value is determined by a human domain expert who can judge whether the output is correct or not, in the right direction or not, if it's worth iterating upon or if it's going to be a giant waste of time, and so on. And the human must remain vigilant at every step of the way, since the tool can quickly derail.

People who are using these tools entirely autonomously, and give them access to sensitive data and services, scare the shit out of me. Not because the tool can wipe their database or whatnot, but because this behavior is being popularized, normalized, and even celebrated. It's only a matter of time until some moron lets it loose on highly critical systems and infrastructure, and we read something far worse than an angry tweet.

And yet you have certainly used and enjoyed software published by others free of charge, and your employer, company or favorite service has relied on it. Your career may even be entirely dependent on it.

If you demand remuneration for all your work, then it's only fair for you to also pay for every single piece of software you ever use. If OTOH you're willing to trade some of your time and effort for the time and effort someone else spent on the software you enjoy for free, then you might appreciate that a financial transaction is not required for value to be created in the world. What is required is fair collaboration.

There is currently no way to prevent this apart from not giving the LLM full control. It will not delete what it can not delete.

But deleting something is just one action you might not want it to take.

The recent "agentic" craze is fueled by the narrative pushed by companies and influencers alike that the more access given to an LLM, the more useful it becomes. I think this is ludicrous for the same reasons as you, but it is evident that most people agree with this.

We can blame users for misusing the tools, and suggest that sandboxing is the way to go, but at the end of the day most people will favor convenience over anything else a reasonable person might find important.

So at what point should we start blaming the tools, and forcing "AI" companies to fix them? I certainly hope this is done before something truly catastrophic happens.

This is missing the point.

The issue isn't with the amount of guardrails in place to perform an action. Yes, it is obvious that there should be some in place before doing any critical operation, such as deleting a database.

The issue is that the "agent" completely disregarded instructions, which in the age of "skills" and "superpowers" seems like an important issue that should be addressed.

Considering that these tools are given access to increasingly sensitive infrastructure, allowed to make decisions autonomously, and are able to find all sorts of loopholes in order to make "progress", this disaster could happen even with more guardrails in place. Shifting the blame on the human for this incident is sweeping the real issue under the rug, and is itself irresponsible.

There are far scarier scenarios that should concern us all than losing some data.

The major hurdle right now is actually pivoting LLMs from just generating code: integrating those tasks into workflows.

Funny, I thought that the major hurdle is improving accuracy and reliability, as it's always been. Engineering is necessary and useful, but it's a much simpler problem, which is why everyone is jumping on it.

I consider that to also be a wrongly held position, because you'd need proof either way.

Proof that something doesn't exist? Ever heard of Russel's teapot?

The burden of proof is on the claimer.

What exactly do we mean by God?

Absurd question. Pick up any holy book, or ask any believer. An atheist is simply a person who doesn't hold those beliefs.

A famous Dawkins quote is apt in this discussion:

We are all atheists about most of the gods that humanity has ever believed in. Some of us just go one god further.

Having certainty something that can be perceived as God by believers cannot exist in our universe is in the end a belief, with no proof.

Again, you're mistaking what atheism is. It's not being certain that a deity cannot exist—it's not having any reason to think that it does.

People who claim certainty in either direction are equally delusional. The problem is when a belief crosses into realms of reality, defines the identity and culture of people, and influences the rest of society. Based on history and personal experience, theists are far more prone to this than atheists.

your extreme atheists aren't much different from your extreme believers; they both have strong beliefs about things they can't prove, and for some reason want to go off on them.

You have a mistaken understanding of what atheism is. It is not a belief in anything, but an absence of belief in a deity.

there are a whole lot of things that we all go around everyday "not believing."

Sure, and yet theism is part of 75% of the world population and influences everything from education to politics. It's perfectly reasonable to talk about atheism within appropriate settings.

From what I understood, any color and material involved in high precision manufacturing requires careful design and thorough testing. They likely prioritize the brown color and material due to branding, so changing this to anything else requires redoing large parts of the pipeline.

Understand Anything 3 months ago

There's no such thing, but all content publishing web sites should at the very least provide tools for users to self-moderate, which this forum heavily relies on anyway.

Now that the internet is flooded by machine-generated content, which is often published and promoted autonomously as well, all content should be scanned and labeled with a value that indicates the likelihood of it being machine-generated and published.

I'm thinking of JSON fields like `machine_gen_probability` and `machine_pub_probability` returned by the API. Then the frontend should expose settings to show these labels next to each post and comment, and filtering rules to decide what should be done with content above a certain value (hide, adjust feed rank, etc.). Some people might even want to boost this content, for whatever reason, so making the system flexible would be smart.

The scoring system of course won't be perfect, but I figure that a company like YC should know a few talented individuals that could do a solid job of implementing this. They've certainly profited from investing in companies that cause this problem.

But... considering HN is merely a promotional tool for YC that runs on limited resources as it is, I wouldn't hold my breath that such a system would ever be implemented. So all we're going to get are changes to "guidelines", and hope that the system won't be abused. Which is laughably naive in this day and age. So this forum will most likely be overrun by the noise, and end up with minimal participation from reasonable humans, as is happening and will continue to happen on most online platforms.

Understand Anything 3 months ago

I'm exhausted by these shiny vibe coded projects that overpromise and underdeliver.

Knowledge comes from doing the hard work, not from being spoon fed information. All these fancy graphs represent a tentative mental model produced as a result of research and learning. Everyone's model is different based on their own experience and focus, so trying to present it as a unique map will more than likely not be conducive to understanding at all. Besides the fact that it will almost certainly miss important details or be hallucinated.

HN users: stop upvoting and promoting this garbage. HN mods: please give us tools to label and filter this content.

The idea that a tool intended to replace all human cognitive work is the next level of abstraction is so fundamentally flawed, that I'm not sure it's made in good faith anymore. The most charitable interpretation I can think of is that it's a coping mechanism for being made redundant.

Nevermind the fact that these tools are nowhere near as capable as their marketing suggests. Once companies and society start hitting the brick wall of inevitable consequences of the current hype cycle, there will be a great crash, followed by industry correction. Only then will actually useful applications of this technology surface, of which there are plenty. We've seen how this plays out a few times before already.

So this is proof of the models actually getting stronger (previous generations of LLMs were unable to solve this one).

No, it's not.

While I don't dispute that new models may perform better at certain tasks, the fact that someone was able to use them to solve a novel problem is not proof of this.

LLM output is nondeterministic. Given the same prompt, the same LLM will generate different output, especially when it involves a large number of output tokens, as in this case. One of those attempts might produce a correct output, but this is not certain, and is difficult if not impossible for a human not expert in the domain to determine this, as shown in this thread.

This resonates a lot with me.

Long breaks help. Take your mind off of things that bothered you. Do things you enjoy. Which may include tech work, but on your own terms.

I wouldn't be surprised if you decide to not go back. The status quo of most organizations is grim. But there are still people who care about the same things as you. You can seek them out and work together, much like you did 15 years ago. This is more difficult now among the noise, but you can tune that out. The industry will never recover altogether, but this current period is a blip of high insanity, which will subside in a few years.

Good luck!

As disturbing as the film "Midsommar" was, I found the concept of a human life being divided into 4 seasons of 18 years each pretty compelling. Not necessarily that life should end after Winter, but a person's contributions to society probably should. Having politicians in office pushing 80 is a disgrace.

For crying out loud, why are we discussing and paying attention to articles and claims about a product that doesn't even exist yet?!

If this isn't a sign of a bubble, where marketing is more important than the actual product, I don't know what is. This industry has completely lost the plot.

Incus wants to own the VM lifecycle the way libvirtd does, and once you have that you're back to "two sources of truth" if you ever shell out.

That's true. But I didn't want to reinvent what Incus or any hypervisor abstraction does. I simply wanted to add some sugar on top that allows me to easily declare infra using small abstractions, and to tie in the provisioning aspect along the way. I still use Incus directly, and can benefit from their work, as you say. State is also managed by Pulumi, so really, there are 3 places for it to exist. There are some challenges with this, of course, but I think the tradeoff is worth it.

Good luck with your project, I'll be keeping an eye on it. I'll probably make a Show HN post when I release mine. Cheers!

Ah, that's neat as well.

I took a slightly different approach in that I don't want to use YAML as the authoritative source. Many projects abuse it, and end up creating a DSL on top of it with all sorts of hacks to achieve the flexibility of a programming language. Pulumi and Pyinfra already provide user-friendly primitives and idempotent(ish) APIs that work much better than YAML. I simply want to expose some (opinionated) building blocks to make them easy to use, and allow users to customize them and add their own as needed. E.g. I definitely don't want to write any shell scripts inside YAML. :)

BTW, Pulumi already supports YAML[1], which can be used with any provider. But to me it's too verbose and generic, and of course, it lacks the provisioning primitives.

[1] https://www.pulumi.com/docs/iac/languages-sdks/yaml/

Very cool, thanks for sharing.

I built something similar recently on top of Incus via Pulumi. I also wanted to avoid libvirt's mountain of XML, and Incus is essentially a lightweight and friendlier interface to QEMU, with some nice QoL features. I'm quite happy with it, though the manifest format is not as fleshed out as what you have here.

What's nice about Pulumi is that I can use the Incus Terraform provider from a number of languages saner than HCL. I went with Python, since I also wanted to expose a unified approach to provisioning, which Pyinfra handles well. This allows me to keep the manifest simple, while having the flexibility to expose any underlying resource. I think it's a solid approach, though I still want to polish it a bit before making a public release.

That's ludicrous hair splitting.

If I have evidence that a crime has been committed based on my layperson understanding of the law, I will surely inform others before the case is even brought to courts. Journalists can and should do the same.

By your logic, reporting based on evidence provided by whistleblowers shouldn't exist. Things like Watergate would likely have never happened.

Journalists shouldn't accuse anyone of committing a crime, and goes without saying that facts shouldn't be fabricated, which is unfortunately common nowadays as well, but they should report events that happened based on the information they have, whether these happen to be related to crimes or not.

I can't say whether this was machine-generated or not, but the reason LLMs use these patterns is because they're often used by humans, which is what they're trained to mimic. LLM spam has now made it annoying, but there are many people who still write like this. Asking them to change their writing patterns because LLMs have ruined it for readers is not just unfair—it's offensive. (See what I did there? Double whammy!)

Claude Opus 4.7 3 months ago

A major reason for that is because there's no way to objectively evaluate the performance of LLMs. So the meme projects are equally as valid as the serious ones, since the merits of both are based entirely on anecdata.

It also doesn't help that projects and practices are promoted and adopted based on influencer clout. Karpathy's takes will drown out ones from "lesser" personas, whether they have any value or not.