HN user

jcgl

429 karma

sysadmin and low-key programmer

j+hn@cgl.sh

(my website and mail are hosted on my laptop; they may be offline at any given time.)

Posts0
Comments396
View on HN
No posts found.

Technical people, maybe. Even if we hypothetically grant that vibe-coded backup systems can be trusted, deploying, using, and maintaining are still hurdles. And there’s absolutely no way that’s all tractable for the vast, vast majority of the population.

Mind you, an average person probably doesn’t even really know what a server is. And these average people need backups just as much as (probably more than) people technical enough to vibe-and-deploy.

Not my area, but isn't it really only because glibc doesn't maintain stable interfaces across versions? If it did, you absolutely could use the same ld.so with different glibc versions. But it doesn't, so here we are.

First of all, techy nerdy people like things to be easy too. They're just somewhat more likely than other people (on average) to overcome not-easiness.

Second of all, how long is IndieWeb supposed to cook for before it's supposed to be ready for a broader audience? This is no shade on the IW folks if they like what they're building. But if what they're building is supposed to catch on somewhat, what's the path supposed to be? The project seems to be 15 years old already: https://indieweb.org/IndieWebCamps#

You mean a few customers? Yes, I think that’s perfectly reasonable to expect that changes made to a very expensive product are well-documented for those customers to whom that matters.

Huge fan of jet. It lets you just write SQL, but in your Go. I basically can’t imagine using anything else now.

Why is that better than writing plain SQL like in sqlc? My main reason was being able to dynamically construct queries and reuse different bits. Plain SQL statements simply don’t compose at all, and I don’t recall sqlc giving any solution to help with this.

Poor writing is not a new thing, of course. Most of the moderation mechanisms that it uses were perfected a quarter century ago when sites like Slashdot were popular as a defense mechanism against bad user behavior. Bad user behavior impacts commenting, article submissions, and moderating itself. While bad users now have AI to abuse, the problem of a large volume of low quality content is is largely the same.

I respectfully disagree on almost all counts (beyond poor writing not being new, of course!).

Moderation mechanisms have not been perfected. They're certainly not perfect, and, given that, I don't know what would make one call them perfected. Humans have probably gotten more accustomed to being moderated, but it'd take a lot to convince me that we have even reached a decent place for moderation at moderate scale, let alone something that is good-to-perfect.

Most importantly, the problem of low quality content is now not the same. Magnitudes matter, and a difference in degree eventually becomes a difference in kind. LLMs have escalated the problem of garbage content beyond what would've previously been conceivable.

To illustrate: How do you dispose of several trash bags at once? Take them to the trash can. How do you dispose of several tens of trash bags at once? Need to rent a dumpster.

Or a classic: If you owe the bank $100, that's your problem. If you owe the bank $100 million, that's the bank's problem.

Bringing it back to written content: doing human-driven moderation on hundreds of submissions a day is tractable with (idk) a couple of people. For thousands or tens of thousands? Intractable. And bear in mind that human-driven moderation is one of the things that keeps HN a better place on the net than many (most) others.

That's definitely the pragmatic choice when working with shell and what roughly everyone does. But it's is also a UX compromise that, if needed anywhere other than the Stockholm Syndrome-ridden world of unix, would be routinely derided.

Shell and plain text are too embedded in the unix-descendant world to get rid of, and I'm not advocating for anything like that. Just trying to push back against oft-repeated maxims about the power of plain text. Structured data has lots to offer for composability, power, and UX.

Tl;dr: Plain text's bad composability forces the dichotomy that you identify between --sort and sort. I agree with you that --sort often can be a sensible UX choice regardless, but losing out on composable middle ground is strictly a loss in terms of power and expressiveness.

Thanks for taking some time to think about it. Despite my pretty absolute wording in GP, I do think that there's nuance here. But what I want to drive home is that the tight coupling/brittleness inherent in plain text composition systematically limits composition.

What I'm not saying is that every --sort option is bad sign for composition. Like you point at, sometimes it's just a simpler UX to have such an option included with your command. As a matter of fact, you see that in PowerShell sometimes with the -Filter option on various cmdlets.

Plain text's brittleness limits composition by promoting exactly the extremes that you point at (the extremes being overloaded sort, or overloaded producer):

the producer understands image size semantics better than the sort command, you can either bake everything into sort, making it an all encompassing command or you can simplify the previous step with a sort flag.

I think it should be obvious that baking everything into sort is a bad idea in the general case. Like you say, the producer understands the size semantics better. Moreover, the ergonomics are lousy (see comments by users alloyed and i15e).

Baking everything into --sort is bad because it limits sorting to the producer's own predefined semantics. While arguably better than relying on sort's, the producer won't always have the semantics that the user cares about. E.g. maybe the user wants to analyze the disk space used only on a certain datastore. Maybe some datastores do transparent compression and sorting should be done by physical disk usage. And so on.

These two extremes are basically your only options in a plain text world, but structured data gives you more opportunities for composition. By moving to structured data and eliminating the need for ad-hoc parsing, users and their code can operate at a higher semantic level. In particular, the loose coupling introduced by this approach gives you access to things like lambdas. You're not alone in objecting to PowerShell's verbosity. Here's a terser version that would work if podman image output structured data rather than text, thereby kinda steelmanning this position:

  podman image ls --all |
      sort {podman size -h $_.size} -d
Basically all that has happened is that the plain text->structured parsing was dropped (again, to steelman the structured data vision) and naming conventions were made unix-y.

An alternate version if no specialized podman size command is needed and the sort cmdlet would by default look at an object's size property:

  podman image ls --all |
      sort -d
In many common cases, I think I agree with you that having --size on the producer is pragmatic and fast. A good UX choice. What's bad about plain text is that there is not/cannot be any middle ground between --sort and sort.

Now, tools shouldn't always oblige the user to compose. Composition is frequently not the best UX. But a shell that systematically limits composition is prima facie worse than one that promotes it. And plain text shells do indeed limit composition.

Sorry for the essay.

There is some cool stuff here.

I like using column to format the table. Appending it to alloyed's command fixes their header problem.

The stdbuf to multi-command block (term.?) is a neat trick. Although, one time when I ran this, I only got a couple lines of output. No idea why and I can't replicate it, but there could be some flakiness that results from the buffering somehow?

Question: how do the ? markers on the sort and column invocations work/what do they do?

Yes, the point was about doing it in a pipeline. The pipeline is the basis for composition of plain text in the unix shell. If something as basic as sorting a table is hard to do, it should make us question just how good the unix shell/plain text philosophy actually is.

Baking --sort flags into shell tools is a sign that the tools do not compose well.

Ooh, I was so hoping someone would take up the challenge! This is a far shorter answer than I had honestly thought was possible. More readable too, somehow? Great use of tee that I would never have come up with (though I hear what you say about there maybe not being ordering guarantees).

Unfortunately, it's not 100% correct, due to misaligned headers:

  REPOSITORY TAG IMAGE ID CREATED SIZE
  registry.fedoraproject.org/fedora-toolbox 44 5a36f433c691 2 months ago 2.14 GB
  quay.io/keycloak/keycloak latest 1361d6e49205 9 days ago 478 MB
  ...
I think that speaks to your final point, which is spot-on:

I'd probably just end up dropping the header and living with worse output in reality

This pretty much sums up plain text and unix shell imo. It's very much the pragmatic solution here, and it's what ~100% of shell scripters would choose to do. And it should make anyone question the orthodoxy around the "power" of plain text in shells.

Great example! How do you like using elvish? Even though I am a proponent of structured data and like PowerShell a lot (mentioned in a nearby comment of mine), I use fish as my regular shell. Big fan of fish's careful focus on user experience, but would be open to trying something structured.

I think you're mistaking text-with-structured data for structured data itself.

Because unix shell is irrevocably text-oriented, kludging in something like JSON is basically the best that can be done when you start to want to do structured operations on structured data. (I'm sympathetic to your point about the AWS CLI tools doing JSON by default though--that just sounds like bad design.)

Being text-oriented imposes drastic limits on composability. Because there is no structure, every element of a pipeline needs to do its own parsing of the input data. This leads to brittle pipelines where every element is tightly coupled to its input's textual representation.

As an exercise, try to write a pipeline that sorts podman images by size without removing the column headers[0]:

  $ podman image ls --all 
  REPOSITORY                                 TAG         IMAGE ID      CREATED       SIZE
  docker.io/prom/prometheus                  latest      937690d77350  2 months ago  367 MB
  quay.io/keycloak/keycloak                  latest      da9433c9fac3  2 months ago  466 MB
  registry.fedoraproject.org/fedora-toolbox  43          a32da54355ca  4 months ago  2.19 GB
  docker.io/powerdns/pdns-auth-49            latest      8c1385c9deed  4 months ago  208 MB
  docker.io/testcontainers/ryuk              0.13.0      b75bc7ce94c3  6 months ago  7.21 MB
As far as I can tell, there is no way to do this in a manner that's even remotely composable. Your best bet is to basically do everything from within awk. Whatever the result would be, it certainly won't be pretty!

Contrast that with what you can do in PowerShell. You can write a couple of standalone functions[0] that are readable and composable, resulting in this pipeline:

  podman image ls --all |
      Replace-SpacesWithTabs |
      ConvertFrom-Csv -Delimiter "`t" |
      Sort-Object -Property {Convert-HumanSizeToBytes -Size $_.size} -Descending

[0] Repurposing this from a blog post I wrote: https://www.cgl.sh/blog/posts/sh.html#this-should-be-basic

Sure, I take your point that the smell-test works reasonably well for domains related or adjacent to one's own.

But (not speaking about your use specifically here) many people (most, I'd wager) use LLMs for many things beyond their own expertise. And it's there that they're most likely to be ensnared without even knowing it.

I definitely agree with your notion that learning-by-doing is a helpful salve for LLM falsehoods. It's no panacea (working != correct (an incorrect solution can appear correct over a given interval)), but it's a good way of working in general that helps keep LLMs in check. And it's very natural to code or other things that can be immediately applied.

But learning-by-doing of course doesn't work with topics that aren't immediately applied. Which includes lots of topics that people use LLMs for (Wikipedia too, for that matter).

The set of unfamiliar-or-unapplied is practically a lot larger than the set of familiar-or-applied.

bad info becomes apparent, almost immediately

Can you elaborate on this? I suspect that you’re thinking mostly of cases in which you already have a fair bit of domain expertise. But in the general case, this seems to be very untrue. Which is why it’s so pernicious that LLMs can generate such quantities of syntactically-plausible-but-factually-untrue text.

This is covered by allowing for single-use credentials. IIRC the EU personal IDs will use this. Basically, the wallet requests a batch of single-use eIDs that all use different device key-pairs. Each credential is only used for one request and then deleted.

But this then means that the issuers and the verifiers can trivially collude to deanonymize holders/users.

Citation needed. These numbers are quite consistent with the growth pattern that started well before usable LLMs were even a thing.

First of all, multi-party democracy needn’t be slow. Parliamentary systems in multi-party countries often react faster than the US. This is due in part to legislation being systematically easier to pass.

Second of all, winner-take-all presidential mechanics don’t imply a four-year cycle of funding instability for research. That only happens if the president has sufficient control of funding. Which, through the administrative state (which is supposed to basically be a delegate of congressional authority), really is supposed to be insulated in large part from presidential, partisan politics. With increased centralization of power in the president (which, imo, is largely just an ~evolutionary response to Congress’ sclerosis), this insulation is lost, exposing research more to the four-year cycle of heavily partisan presidential politics.

I’m no expert, but as long as they’re represented by tokens in the end, they’re just tokens. Even if you train the transformer to treat them specially, a token is a token, and there’s no free lunch. At best, you’re going to be trading off between paying attention to this would-be security boundary and delivering high-quality results; the more you focus on one, the more you lose on the other.

Leaving Mozilla 1 month ago

Case-in-point: just started a new chat with a new person (we had a previous room in common)--My desktop client, NeoChat, shows "This message is encrypted and the sender has not shared the key with this device." for all of their messages. FluffyChat on my phone shows their messages correctly.

Welcome to Matrix. It's the best option there is, and it's not very good.

Leaving Mozilla 1 month ago

has good clients

So far, I've only found clients with different bugs. Calling them good would be a stretch. Passable, perhaps. But the scene as a whole is more of a choose-the-bugs-to-live-with situation than choose-a-good-client.