HN user

abathur

2,097 karma

Fixing Shell with Nix + https://github.com/abathur/resholve. Blogging @ t-ravis.com

Posts35
Comments918
View on HN
www.theguardian.com 2y ago

Hyperphantasia and the quest to understand vivid imaginations

abathur
4pts0
www.theguardian.com 2y ago

No solar glasses? How to tell if you damaged your eyes during the eclipse

abathur
1pts3
economicsfromthetopdown.com 2y ago

Nixing Technological Lock In

abathur
27pts9
fosstodon.org 2y ago

Accidental FUD around setup.py

abathur
5pts0
www.wiumlie.no 2y ago

PhD Thesis: Cascading Style Sheets (2005)

abathur
2pts0
news.mit.edu 2y ago

Desalination system could produce freshwater that is cheaper than tap water

abathur
344pts191
www.oilshell.org 3y ago

Oils Is Exterior-First (Code, Text, and Structured Data)

abathur
1pts0
discourse.nixos.org 3y ago

NixOS S3 (public binary cache) Short Term Resolution

abathur
11pts1
support.google.com 3y ago

Reports of missing contacts on Google Account Community forum

abathur
1pts1
support.google.com 3y ago

A: Have all your contacts disappeared? (Android)

abathur
1pts0
old.reddit.com 3y ago

Old searches keep showing up below my [Google Chrome] search bar?

abathur
2pts0
t-ravis.com 3y ago

modular bash profile scripting with shellswain

abathur
2pts0
imgur.com 3y ago

Visible mending is the new couture

abathur
1pts0
t-ravis.com 3y ago

What color is your markup?

abathur
1pts0
williamcolgan.net 3y ago

Melting ice reveals two-million-year old peat

abathur
62pts65
t-ravis.com 4y ago

No-look, no-leap Shell script dependencies

abathur
3pts0
t-ravis.com 4y ago

semantic: the 8 letter s-word

abathur
2pts0
www.t-ravis.com 4y ago

The Gizmo's Role in Markup

abathur
1pts0
t-ravis.com 4y ago

Nix-shell, but make it lovely

abathur
1pts0
www.t-ravis.com 4y ago

Neighborly Shell with bashup.events (part 3 of 3)

abathur
2pts0
www.t-ravis.com 4y ago

The missing comprehensive package manager for Shell (part 2 of 3)

abathur
1pts0
www.t-ravis.com 4y ago

Modularity in the age of antisocial Shell (part 1 of 3)

abathur
1pts0
t-ravis.com 4y ago

Advanced shell packaging: resholve YADM's nixpkg

abathur
1pts0
github.com 5y ago

Gandalf-/Bashkell: Functional bash scripting

abathur
2pts0
keith.github.io 5y ago

Cryptex filesystem hierarchy specification (Darwin manpage)

abathur
2pts1
github.com 5y ago

Porglet is a low cost and small form factor drone detection system

abathur
1pts0
www.dallasnews.com 5y ago

Texas takes Griddy to court over electric bills sent during storm

abathur
3pts0
reuters.com 5y ago

Texas electric firm files for bankruptcy citing $1.8B in claims

abathur
152pts268
adoptoposs.org 6y ago

Adoptoposs – Keep open source software maintained

abathur
1pts0
fivethirtyeight.com 6y ago

University of Nebraska Helping Student-Athletes Become Social Media Influencers

abathur
1pts0

I guess, but have you actually encountered a teacher grading an assignment solely based on word count?

I certainly wish more teachers encouraged parsimony and penalized fluff and bullshittery, but I'd be surprised to find them doing it outside of some narrow cases where the point is just to make you write something at all.

Tthey generally want to encourage their students to engage with the topic at a certain level and practice the thinking needed to research, structure, and implement an argument of a certain length. They want you to put at least 5 pounds of idea in the 5-10 pound idea bag.

If you're convinced you've hacked word economy and satisfied the assignment except for this goshdarnpeskyminimumwordcount, you're probably misunderstanding the lesson the instructor is willing to read through a bunch of bad writing to impart and cheating yourself.

I feel like we need a name for css-the-syntax (and maybe -the-semantics) as separate from css-the-body-of-rules/functions/units/etc-defined-by-csswg.

There's juice in it, but it's hard to talk about and survey other uses without just searching GH for code using css parsers and just see what kind of shenanigans people are up to.

I've been playing around with a weird thing that's kinda like a template engine, but driven by a mix of a lightweight node-based markup language, css selectors for expressing what goes into the template, and a css-alike for controlling exactly how all of these parts come together.

I agree with you that a face-to-face q&a is a reasonably good way to detect low-effort cheating, but I'll still quibble a bit:

- I don't think this lowers the cost of detection as much as you imagine. You still need to know the paper better than the student and have to sacrifice already tight instruction/planning/grading time to have all of these conversations. Even if you catch enough to successfully deter most, it likely means not covering something else. It won't be too hard to catch low-effort cheaters who can't be bothered to read the paper, but you're on the low-leverage side of an arms race with the remaining students. You have experience on your side and they can't know what you'll ask, but they outnumber you and can certainly read the paper and use LLMs to quiz them on it. You have to invest your effort without knowing how each student prepared, so you'll spend about as much effort on every low-effort cheat as you do on the highest-effort cheat you are prepared to catch.

- Not sure it is "from the wrong direction" since both approaches raise the cost of cheating and lower the cost of detecting it.

- While this does avoid encouraging students to dumb down their work, it does still raise the cost of not-cheating. Unless you surprise the students with these conversations, the ones that care most will still anxiously prepare.

There are many disciplines in which students work on effectively distinct projects.

For example, the life-changingly-well-designed newswriting course I took in college assigned every single student a different story to spend several weeks reporting out so that we wouldn't all be out harassing the same poor people for interviews.

Sure--yes--the student will learn something if they actually wrote a 20-page paper on some given topic. But how are you going to evaluate their ability to compose the 20-page argument?

I would prefer not to be confrontational here, but I am having a hard time imagining that you've deeply considered the pedagogy of how to teach and evaluate students on squishy skills like this.

Knowing a bunch of facts about something is a world apart from structuring a compelling in-depth argument about it.

Does crapping on the average school's deep well of expertise for evaluating how effectively AI software solutions address their problems somehow fix the underlying problem (that the cost of catching cheaters is significantly higher than the cost of cheating)?

(This is roughly the same problem as evaluating software that only does an approximation of what it claims to do.)

(Aside: AI-based variations on this theme are in the early stages of proliferating across our society. They're being developed by many people using this forum and being sold to our schools, businesses, governments, and other organizations with little regard to whether they actually do what they claim.)

I don't disagree with you that a reasonable way to cope with the current problems is to ensure everything that "counts" is done in a controlled environment, but pedagogy and its goals are vast.

There are things you learn from spending several days structuring a 20-page argument that you will not learn (and cannot assess) from oral examination or a 5-paragraph essay written in a blue book.

Granted, but this reads a bit like a headline from The Onion: "'Hard to imagine a more favourable situation than pressing nails into wood' said local man unimpressed with neighbour's new hammer".

Chuffed you picked this example to ~sneer about.

There's a near-infinite list of problems one can solve with a hammer, but there are vanishingly few things one can build with just a hammer.

You (or the person I was replying to) basically have to make the case that Simon Willison is ignorant about LLMs and programming, is desperate about something, or is deluding himself that the port worked when it actually didn't, to keep the original claim.

I don't have to do any such thing.

I said the experiments were both interesting and illuminating and I meant it. But that doesn't mean they will generalize to less-favorable problems. (Simon's doing great work to help stake out what does and doesn't work for him. I have seen every single one of the posts you're alluding to as they were posted, and I hesitated to reply here because I was leery someone would try to frame it as an attack on him or his work.)

Is it? I can't use an example where they weren't useful or failed.

  https://en.wiktionary.org/wiki/cherry-pick

  (idiomatic) To pick out the best or most desirable items
  from a list or group, especially to obtain some advantage
  or to present something in the best possible light. 

  (rhetoric, logic, by extension) To select only evidence which supports an argument, 
  and reject or ignore contradictory evidence. 
> any number of people failing at plumbing a bathroom sink don't prove that plumbing is impossible or not useful. One success at plumbing a bathroom sink is enough to demonstrate that it is possible and useful - it doesn't need dozens of examples - even if the task is narrowly scoped and well-trodden.

This smells like sleight of hand.

I'm happy to grant this (with a caveat^) if your point is that this success proves LLMs can build an HTML parser in a language with several popular source-available examples and thousands of tests (and probably many near-identical copies of the underlying HTML specs as they evolve) with months of human guidance^ and (with much less guidance) rapidly translate that parser into another language with many popular source-available answers and the same test suite. Yes--sure--one example of each is proof they can do both tasks.

But I take your GP to be suggesting something more like: this success at plumbing a sink inside the framework an existing house with plumbing provides is proof that these things can (or will) build average fully-plumbed houses.

^Simon, who you noted is not ignorant about LLMs and programming, was clear that the initial task of getting an LLM to write the first codebase that passed this test suite took Emil months of work.

If a Tesla humanoid robot could plumb in a bathroom sink, it might not be good value for money, but it would be a useful task. If it could do it for $30 it might be good value for money as well even if it couldn't do any other tasks at all, right?

The only part of this that appears to have been done for about $30 was the translation of the existing codebase. I wouldn't argue that accomplishing this task for $30 isn't impressive.

But, again, this smells like sleight of hand.

We have probably plumbed billions of sinks (and hopefully have billions or even trillions more to go), so any automation that can do one for $30 has clear value.

A world with a billion well-tested HTML parsers in need of translation is likely one kind of hell or another. Proof an LLM-based workflow can translate a well-tested HTML parser for $30 is interesting and illuminating (I'm particularly interested in whether it'll upend how hard some of us have to fight to justify the time and effort that goes into high-quality test suites), but translating them obviously isn't going to pay the bills by itself.

(If the success doesn't generalize to less favorable situations that do pay the bills, this clearly valuable capability may be repriced to better reflect how much labor and risk it saves relative to a human rewrite.)

I think both of those experiments do a good job of demonstrating utility on a certain kind of task.

But this is cherry-picking.

In the grand scheme of the work we all collectively do, very few programming projects entail something even vaguely like generating an Nth HTML parser in a language that already has several wildly popular HTML parsers--or porting that parser into another language that has several wildly popular HTML parsers.

Even fewer tasks come with a library of 9k+ tests to sharpen our solutions against. (Which itself wouldn't exist without experts trodding this ground thoroughly enough to accrue them.)

The experiments are incredibly interesting and illuminating, but I feel like it's verging on gaslighting to frame them as proof of how useful the technology is when it's hard to imagine a more favorable situation.

I've been tasked with doing a very superficial review of a codebase produced by an adult who purports to have decades of database/backend experience with the assistance of a well-known agent.

While skimming tests for the python backend, I spotted the following:

    @patch.dict(os.environ, {"ENVIRONMENT": "production"})
    def test_settings_environment_from_env(self) -> None:
        """Test environment setting from env var."""
        from importlib import reload

        import app.config

        reload(app.config)

        # Settings should use env var
        assert os.environ.get("ENVIRONMENT") == "production"
This isn't an outlier. There are smells everywhere.

This is the kind of dismissive sneer the HN guidelines advise against.

You can write dev docs for humans and still want machine readability (without caring about whether some LLM can make sense of the docs).

Machine readability is how you repurpose your own documentation in different contexts. If your documentation it isn't machine readable it might as well be in a .doc(x) file.

Agriculture feeds people, Simon.

It's fair to be critical of how the ag industry uses that water, but a significant fraction of that activity is effectively essential.

If you're going to minimize people's concern like this, at least compare it to discretionary uses we could ~live without.

The data's about 20 years old, but for example https://www.usga.org/content/dam/usga/pdf/Water%20Resource%2... suggests we were using over 2b gallons a day to water golf courses.

I'll cop to not reading the whole list before commenting, but I skimmed this and didn't really notice anything about speed or performance.

When using tools that can emit 0 to millions of lines of output, performance seems like table-stakes for a professional tool.

I'm happy to see people experiment with the form, but to be fit for purpose I suspect the features a shell or terminal can support should work backwards from benchmarks and human testing to understand how much headroom they have on the kind of hardware they'd like to support and which features fit inside it.

It seems reasonable to me. Sorry you got that reaction. Just from the mailing list thread size alone, it looks like you put quite a lot of work into it.

I wasn't readily able to find where the discussion broke down, but I see that there's a -p <path> flag in bash 5.3.

When you say library system, do you mean something more or less like a separate search path and tools for managing it?

I've written a little about how we can more or less accomplish something like meaningfully-reusable shell libraries in the nix ecosystem. I think https://www.t-ravis.com/post/shell/the_missing_comprehensive... lays out the general idea but you can also pick through how my bashrc integrates libraries/modules (https://github.com/abathur/bashrc.nix).

(I'm just dropping these in bin, since that works with existing search path for source. Not ideal, but the generally explicit nature of dependencies in nix minimizes how much things can leak where they aren't meant to go.)

Death by AI 1 year ago

It's an outdoor seating counter serve kind of place, so yeah :)

Death by AI 1 year ago

A popular local spot has a summary on google maps that says:

Vibrant watering hole with drinks & po' boys, as well as a jukebox, pool & electronic darts.

It doesn't serve po' boys, have a jukebox (though the playlists are impeccable), have pool, or have electronic darts. (It also doesn't really have drinks in the way this implies. It's got beer and a few canned options. No cocktails or mixed drinks.)

They got a catty one-star review a month ago for having a misleading description by someone who really wanted to play pool or darts.

I'm sure the owner reported it. I reported it. I imagine other visitors have as well. At least a month on, it's still there.

Sure! Something date-based is a simple way to handle perpetual storage while keeping the active history set from over-growing.

I did anything at all because the default shell profiles in macOS can cause history loss, and I'd found my last straw. I put them in sqlite because--if I was bothering to build something bespoke--I wanted to track more command-time context to build tooling around later. (Including to play with some ~curation ideas.)

I might have decided to just use Atuin if it existed at the time, but I did it three years and change before its first release. (There were maybe 5 or 6 barely-used public examples of this idea on GH at the time, but none tracked everything I wanted among other issues.)

Are you sure it's loading history in there? My .zsh_history isn't completely empty, and when I run the same with 'history' swapped for 'exit' it doesn't print anything. (But this might have something to do with macOS default shell profile stuff.)