HN user

gabriel666smith

418 karma

Literary fiction author (‘Brat’ & ‘The Complete’). Used to be in tech, still coding for research (ie mucking about). Interested in both, where they meet, and what each can learn from the other. It’s all just language

twitter.com/gabriel666smith

henrygabrielsmith at gmail

Posts6
Comments109
View on HN

The metadata matters in a different way. A single piece with a single composer can be interpreted very differently, over time, by different conductors and orchestras.

This means classical music is very badly organised in most streaming apps. Sorting by "artist" doesn't really work if the artist is listed as an orchestra, which has probably recorded the work of more than one composer.

Roon is another example of a streaming platform that treats this metadata with the same importance.

Lol, I'm personally too broke for a Spotify subscription - as in I couldn't sign up this month with the cash in my bank account - and that has often been true during the periods of my life I have not spent working in tech. I don't think I'd even be in their youngest adult user demo anymore!

We definitely exist. It's pretty tight out there right now. The £10.99/month I see quoted buys quite a few kilos of dry pasta.

In any case, the core argument I was making specifically is that I believe human curation can provide more friction for the audience than algorithmic curation tends to, and that this is a net-good for cultural education.

I would contend with classical music being the highest strata of musical culture (and the deeper idea of culture having 'levels'), but that's subjective. So, that aside, if classical music is your thing, and you believe:

"The quality of podcasts and music streaming is on a quality level which radio never reached."

I'd recommend Glenn Gould's The Idea of North, which is one of my absolute favourite pieces of music. It's considered pretty innovative and important:

https://www.youtube.com/watch?v=ry5MUnZoeGI

Glenn Gould wrote it specifically for radio, as it was commissioned by the CBC. So "government produced sewage". :-)

I'd also strongly recommend you give BBC Radio 3 (still often excellent, especially late at night) a try. NTS has some great classical shows, too. Perhaps you're already familiar with all of these, but - if not - they might give you a reason to turn on your radio more frequently than annually.

You're mostly correct - I was absolutely generalising, using internationally known examples (like the BBC), to be more comprehensible to the international audience here about cuts to public arts spending.

The BBC was a silly example of me to use (as was 'public radio'!) because you're right - it's mostly not funded by general taxation.

That said, the license fee is directly tied to inflation (via the CPI). Whether you see that as a correct measurement is another question. [1]

The BBC has also historically been funded partially by taxation. For example, taxation provided free television licenses to over-75s, until 2020, when that funding was cut, and the cost of doing this was passed to the BBC itself.[2]

But yes - quibbles aside, a sloppy example for me to have used there. Your sibling comment has some great sources which better illustrate changes to arts funding, compared to the off-handed ones I initially provided.

[1] https://www.tvlicensing.co.uk/faqs/FAQ23

[2] https://www.ageuk.org.uk/northumberland/about-us/volunteer-f...

Yeah - I used "austerity" as a slightly sloppy shorthand in my comment. My two closest local libraries were closed due to cuts made during "the credit crunch" by the coalition government (from memory!), when "austerity" was the buzzword, so I felt it'd communicate most broadly (to HN's international userbase) the phenomenon I was gesturing at.

From what I can understand, these sources are saying: "because costs are up, those past cuts have not yet really been reversed, especially not in terms of real-terms funding per-capita, or funding for 'non-essential' services".

Super bleak reading - it's a very dark rabbit hole!

I work as an author. I believe this is total bullshit, from beginning to end - the ruling, the settlement, and the suit itself.

In the UK, we have a thing called the Public Lending Right [1]. This pays authors a fixed sum each time their book is taken out of a library, up to a capped amount.

The cap isn't very high - about $7k - so it is both an OK bit of income for authors who might be making very little money elsewhere, and also doesn't end up all going to authors who are already bestsellers. It's a decent legal system for helping libraries hold niche titles as well as the popular ones. This is, after all, the purpose of a library.

To establish my bias here: My debut novel came out after the period this specific suit concerns. I also uploaded it to LibGen myself.

I strongly believe that books should be available to read, free of charge, to all people. I benefited enormously from libraries and piracy growing up. I think they serve an important educational purpose that does not end when a person leaves school, and I do not think wealth or disposable income is a fair way to decide the breadth of a person's education.

I also have no problem with people making new "language things" using my work. I love sample-based music (like dance music, hip hop, etc) and it'd be hypocritical for me to take issue with anyone doing analogous things using books. Maximising sales is not the end-goal of making art, for me personally. Other artists feel otherwise. They consider training on pirated books stealing. That's OK - it's not for me to tell them what to believe.

The problem for me is that these corporations - undoubtedly still pretraining on pirated material - are, essentially, leeching. By not releasing the model as open-weight, freely available, they are not acting in the same spirit of the system they took advantage of. It's the Spotify model: pirate first, pay a nominal amount that does not meaningfully harm profit later. Now the dust has settled there, we can see the harm it has done to music culture.

A single settlement which does not establish precedent does not solve anything. A tokenistic $3k allows anti-AI authors to wave a cheque in the air and declare a victory. It pays the rent for a month or two. It does nothing for the months after that, when the corporation is still profiting. It does nothing to establish precedent for future artists, who also have to pay rent.

It would be (non-trivial, but) relatively simple to integrate - for example - download figures from Anna's Archive into the PLR. I'd happily dilute my PLR payment appropriately, because I think libraries are important.

You can't stop people pirating digitally replicable things. Digital ownership is not a concept that has held, or will hold.

There are only 23,000 authors in the UK who claim the cash from the PLR. To pay all those authors the national living wage in the UK (£26k) from the PLR, you would need to raise £546 million. That is around 1/34 of Anthropic's reported annual revenue.

I'm of course not arguing Anthropic should be solely responsible. But it's very frustrating that all the pieces of the puzzle for actually paying artists in a sustainable and ongoing way now exist, and one of the major obstacles to this - and the idea of a genuinely free, legal, international library, which creates more authors, writing better books, full-time - are legacy rights holders who remain attached to a completely dysfunctional and outdated concept of ownership.

So - unless part of a sustained and reasonable campaign, which understands the futility of (and damage to the medium and its creators caused by) treating digital ownership in the same way as physical ownership - this suit is close to pointless, and arguably actively harmful in the long term.

[1] https://www.bl.uk/services/plr

It is a shame. I think it’s especially a shame for children and young people.

A counter to radio remaining relevant is “I can listen to whatever I want, via this digital technology”.

Fine - but young people don’t yet have a complete mental map of what culture they enjoy.

In the UK, funding for libraries and the arts (public radio, the BBC) has been increasingly cut in the name of austerity. I think that’s true in the US too.

Despite this, young people have an unprecedented amount of access to culture. Or, at least, any form of culture that can be digitally replicated.

Soulseek and Anna’s Archive are wonderful projects. But they only have as much utility as the person using them can put into the search bar.

We’re increasingly entrusting cultural education - what to put into the search bar - to algorithms that are designed not (as the BBC’s mission statement is) to “inform, educate, and entertain”, but to retain attention.

Even worse: many of these (eg, Spotify) hide real access behind subscription fees. If you grew up broke, like I did, you will understand how inaccessible this real access is.

The most motivated young people will continue to find the workarounds piracy offers, and to find other routes for cultural education.

Any young person who isn’t in that upper percentile of motivatedness might increasingly be left out. You can attribute that to a choice they have made - I think that’s wrong.

I think the joy of broadening one’s horizons is available to all, and is also difficult to learn and a strain to continue to practice. I think it must be taught.

I think this is especially true and important when everyone has a constantly-available option of algorithmically-selected pleasantness.

Human radio DJs are not optimized for pleasantness, or to serve you relevant ads based on your conversations with them. At their best, DJs are educators.

Anyone who grew up listening to radio knows that, and knows the pleasure that comes from learning not to change the channel when something new and uncomfortable-feeling comes on, just in case it ends up being something you fall in love with.

You can personally work out how apply this one in engineering today!

If you have a spare hour or two, I'd encourage you to have a go at learning (however you learn best - I like just asking smarter people or robots stupid questions) what the maths means and why it's important.

And then once you feel like you have a vague grip on the principles, think about a problem in a domain you know a lot about. Try to see if the maths - and how it's changed our perception - could be used as a tool to solve that problem, or if the solution is analogous to a solution you could try in your own domain of expertise.

LLMs are good at speeding up, I think, the journey an idea has to take between "theoretical academic stuff for academics" and "a usable idea for regular people", because they increasingly allow you to ask an infinite number of stupid questions and give you (hopefully) reasonably good responses.

I've had loads of fun doing this today - specifically seeing if the idea this counter (from what I understand: a many-to-one conversion that kind of does and kind of does not preserve meaning) can tell me anything about the relationship between language and meaning.

I'm sure everything I've done today while mucking around has been the equivalent of a monkey with a typewriter (and Codex), but I think the huge breakthrough(s) you ask whether we're on the cusp of are relatively dependent on how many monkeys are throwing typewriters at problems they know a little bit about, after learning a bit about new ideas like this one. Historically, that's a really good way for broad cultural innovation to happen - distributed information applied across multiple domains by experts in them.

I don't think this is necessarily going to prove to be true.

I often see the sentiment: "the Chinese strategy only makes sense in the context of undercutting American labs' profit margins".

If, for example, you are a company with a near-monopoly on "serving video content", and you feel reasonably confident about retaining a decent slice of the serving-video-content market (Google in the west is an example, Tencent in the east), then training video models on your dataset - and releasing them freely - makes an awful lot of sense.

Free tools to create with mean more video content. In this hypothetical, you're reasonably certain that any video content which does get created will also be watched on your platform.

That is a net positive. The question becomes: How many watch-hours earns back the cost of training a model? It's probably not really that many, especially when you have a near-monopoly on a billion sets of eyes.

It's also a net-positive if people build better video models from research you release, because - again - you are reasonably certain that the even-more-innovative content those models produce will be watched on your platform.

It really begins to make strategic sense if your company is in a GPU-poor environment. Your costs cease at the point you upload a model if your users are running it themselves. You don't have to serve the model. The content is still created.

You are also less likely, I think, to alienate human creators whose work the model was trained on if the model is not sold back to them as a subscription, or by the token, but given for free as a tool.

This frames the conversation very differently. It creates, I think, less of an "us vs them" dynamic, and more of a rising tide.

It's true that it is also beneficial that these models undercut (especially in language models) American companies. But, generally, Americans are not the customers of Chinese companies releasing models. They are already serving a huge volume of customers in a complex, existing marketplace.

The full picture is much more nuanced than simply a geopolitical desire to undercut US labs, and there are several other reasons the strategy can make logical sense.

Qwen 3.8 2 days ago

It's an interesting question, I think, because GGP is correct to say that geopolitical discussion dominates posts about work done by Chinese developers and organisations. That work is fascinating and innovative.

On the one hand, those orgs are tied to a global superpower. I couldn't name a single global superpower that has existed - ever - which hasn't committed acts I personally consider deeply immoral. I don't think we typically discuss OpenAI's work, and learn from it, in quite the same way we discuss the work of Chinese labs. OpenAI is a US defence contractor, and - whether you agree or not - it must be acknowledged that a lot of people would see actions like the invasions of Afghanistan or Iraq as deeply immoral. I don't think "which country is more evil?" or "which org is more embedded in its government?" is a productive route for this discussion - to me it seems like whataboutism. Others would disagree.

On the other hand, technical and artistic innovation has been sponsored and patronised by wealthy organisations (nation states and religions, predominantly) throughout history. That wealth was, generally speaking, acquired very violently. What does this do to the moral status of the work itself? I don't have an answer to that.

I do think I have a lot more in common with the average Chinese developer than I have in common with the average British (or American) political leader. I also believe the average Chinese developer has more in common with the average American developer than they do with the average Chinese political leader.

Lots of people might disagree with me about that, which is fine. What would be more difficult to disagree with, I think, is that I personally think I can learn a lot more from talking to people about the work we share than about the politics we do not.

Some examples: I personally think the Kimi line of models have been, for some time, by far the strongest of any lab's at creative writing tasks. I don't know how they've managed that. I'm completely obsessed with Bilibili ("China's YouTube"). I think many of its features are outstanding, and unmatched by anything a western video platform currently offers. The video content hosted there is often innovative, and occasionally artistically brilliant; the AI video usage specifically is notably more artistically advanced than anything I've seen on western platforms.

I'd love to listen and talk to some of the people who worked on this stuff, or who know more about it (perhaps in a less off-topic place!). I'd really love to work alongside those human beings. Because of geopolitics, that isn't easy. That is really frustrating to me. HN would, I think, in an ideal world, be a place that could happen.

Qwen 3.8 3 days ago

I don't think you're taking the good-faith reading of the sentence's grammar here, which is:

"[When China is mentioned] the comments section stops discussing [technical aspects], and instead starts going on about [non-technical aspects], and all that [stuff which I do not feel is relevant]."

I'm struggling, though, to find a good-faith reading of what you meant by:

"Denying human rights. Classic. Sorry - you instantly lost any respect I could have maybe had for your opinion."

Specifically, "classic" - classic what? "Classic" meaning "a trait you personally believe is more prevalent in people who (might be) citizens of a specific country"?

Hopefully the good-faith interpretation of GP's comment, which you might have unintentionally missed, will help you engage with their opinion with more respect.

I think it's an easy opinion to empathise with personally. I would be enormously frustrated myself, on an individual level, if interesting technical achievements made by companies in my own country were overshadowed by political discussion about the actions of the country overall, when spoken about on technical forums.

This would be especially true for me if I - as an individual - disagreed with politically with my country's leadership. Frustration at a topic of discourse is not mutually exclusive with being in favour of human rights.

I don't think GP has done themselves any favours with the broad statements that close their comment, and is clearly frustrated themselves. That there's mutual frustration is probably a sign that empathy on both sides might lead to a really productive discussion - but focusing on the more substantive points they've made, rather than escalating by misinterpreting the less important aspects of what they commented - is what will lead to that.

Word. In terms of game level design, I think SSX Tricky’s Tokyo Megaplex is one of the most interesting, complex, simple, comprehensible, fun, beautiful, and downright cool pieces of work of all time.

I can’t have played it in fifteen or twenty years but I still think about Tokyo Megaplex every month or two. It tapped into something deep in my brain.

Thanks for the reminder of it. I won’t know what heaven looks like until I’m at its gates, but when I die, I hope I get to ascend via that weird updraft tube, and ride those red and white candy cane rails for eternity. If not, God might have missed a trick. What an incredible piece of work that game is.

I think one of the beautiful things about America is its formation as a country contains so many nations, and in a way America’s cultural history is inseparable from those nations it contains (newly formed) extensions of. America is a young country made from the offshoots of old countries. So in some ways, I think America’s Homer can simply be Homer, just as America’s Shakespeare can be - if an American chooses - Shakespeare. That’s a beautiful idea for a nation.

To think about the question less evasively, for me, it depends on which aspect of Homer is primary. I’m restricting myself here to widely read texts which I think have had vast impact on national character.

If it’s the idea of suffering, frontier, and the impossible but attempted journey, I find it hard to look beyond Melville. Not just for Moby Dick, but I would argue Bartleby, the Scrivener is an incredibly prophetic work about how regular people would come to experience the deep strangeness of modern American life.

If it’s a question of language - of how we write and speak to each other - so much of American language rests on Hemingway. This answer is unavoidable for me. In American communication, so much value is placed on clarity, and brevity, and the idea of “cool”.

If it’s Homer as an epic, sweet, sad, narratively-loose, arguably-multi-authored, and kind-of-orally-told document of American life, widely shared: Why not Homer Simpson? The Simpsons is art at its best, and has clear export and staying power for audiences.

I think the real answer is probably that it’s too early to say. Bob Dylan will probably be on this list in a century - but I think the ways his work is prophetic of future art and the broader future world have not quite formed into reality yet.

Scott Fitzgerald might be there if America’s 21st century experience is so defined by striving, (self) deception, and startlingly violent tragedy as its 20th century was.

Similarly, if America stays as strange and psychedelic and new as it has been in the 20th century, a case could be made for H P Lovecraft.

While I don’t personally enjoy Lovecraft’s prose, his ideas are increasingly prescient. In the 20th century, Americans walked on the moon. In the 21st, Americans have invented - in the consumer LLM - a thing for which the most tractable analogy is Cthulu, with its many-faces of a single entity, its cultists, and its nondeterministic and unknowable - or incomprehensible - non-human motivation or end-point, which can only really be described as a desire for its own momentum.

It may well end up being Lovecraft.

It's a product that's so deeply dependent on lifestyle, and on physical needs. My fiancee is technically partially sighted (nystagmus from albinism) and I was surprised to learn when we started dating that the RNIB has 3% of the UK living with some form of sight loss, and ~350,000 people as partially-sighted or blind. I didn't really understand the full and nuanced spectrum of what "sight loss" could mean - I guess I had thought of it in very cartoonish, uninterrogated "no glasses vs wears glasses vs sunglasses-dog-and-cane" terms.

The age demographics which read more, broadly speaking, have vastly higher incidences of some kind of sight loss. It's a sizeable portion of the market that goes, it seems, rather underserved in products akin to e-readers, like televisions.

(Televisions specifically being a product where - somewhat off-topic - there are an awful lot of features that would be trivial to implement which would drastically improve the product for the meaningful segments of the market who have some form of sight loss. These features unfortunately seem often to be lower down the roadmap than spyware-adjacent tracking or new kinds of motion smoothing for users to turn off.)

My own IRL library is scattered in storage crates across places I've lived in the past, due to space and travel limitations, so I strongly empathise with you on that.

The creepypasta operates as such an interesting cultural object: Whilst having an original - often anonymous - author, a creepypasta is often built upon quickly and communally, making it strongly akin to a folk tale that develops at a previously impossible speed.

To be successful, the original generally needs to contain a highly compelling image (linguistic or graphical): the "backrooms" in this instance, the long limbs of Slenderman. I would personally guess that the increased cultural interest in the Wendigo over the past few years is directly related to the "Anansi's Goatman Story" greentext's amazingly sensory description of the smell, which really stayed with me. [1]

These alleged actions from A24 are anathema both to the way these stories develop - communally - and also to the communal, somewhat anti-corporate, "independent" creative values in certain audience demographics that partially contributed to A24's rise in the first place.

I suspect it's just copyright strikes being applied thoughtlessly, in the usual way, without the legal team (or whatever is automating their actions) realising that this is an instance where an exception should be made.

To me personally, even if it's not maliciously intended, it's still gross, or perhaps even grosser.

Incidentally, I thought the film itself - while relatively unfaithful to the source material - was one of the better attempts I've seen at integrating the ugliness, slop, and grotesque aspects of online life into a contemporary film. Usually, the online world is translated much more clumsily.

[1] https://creepypasta.fandom.com/wiki/Anansi%27s_Goatman_Story

This looks nice - I think ‘small enough to be pocketable’ is an important form factor. Reading preferences are a very personal thing, and for people like me - who feel they internalize more effectively when reading from a physical book better than when reading from a screen - this form factor makes owning one more worthwhile.

I have always wondered why the form factor that has been settled on for e-readers was ‘book-like’, other than the obvious screen-size advantages. I wonder if there is a better form-factor out there and if anyone has any ideas. Electronic devices don’t need to resemble the thing they are replacing, and it’s sometimes better that they don’t.

I know a lot of people lust after the clamshell e-reader imagined in the film It Follows. [1]

The haptics of a physical book are, I think, what helps me better remember them, because memory is intrinsically linked non-digital sensory information like touch, and smell.

I’ve personally tried a few prototypes to try to bring unique sensory inputs to individual books read digitally to help me remember each one better: converting an AliExpress Game Boy into one to enable more haptic and visual differentiation; generative ambient & binaural music that is automatically created based on the text shown.

Each did help, to an extent, in preventing the way everything I read on an e-reader slightly blurs into indistinct memories of ‘reading generally’, rather than ‘reading this specific book’.

Maybe the e-reader has to be as personally customizable as a cyberdeck for those of us who kind of need the other sensory inputs. So it’s good that the firmware on this one seems to be open-source, but I haven’t yet read through it to understand the extent of that.

[1] https://collider.com/it-follows-clam-phone/

When I'm doing it, it depends on whether the agent has ownership of an MD (or other) doc in the flow. If they do, this remains either in-context or greppable, each fresh-context turn.

If not, my agent-level chains just look like:

'''

Turn 1: OK, my task is X, so I should grep for it. Oh, it produced these results:

(Message pairs)

I should expand the context around those message pairs that look relevant.

(3 message pairs around search result)

I should save 1 of these, as it contains relevant information.

[Enforce Tool call limit]

[Delete all context added, except the search tool used already, and the relevant result(s) found.]

Turn 2: OK, my task is this, and it seems I already have this result, but I still need...

...

Turn 10: OK, after that search, my answer is:

[Response]

'''

I've never bothered to let agents self-remove from context, so I would guess it's 'lazy' in that sense. It seems more complex than the task requires in this case, though I can see the benefits on more complex tasks. If you're already saying "this is relevant info", I figure it's simplest to just enforce deletion of everything not marked relevant. In chained prompts, when you're trying to keep costs low and use weaker models, it seems best to limit decision-making as much as possible to make the models as deterministic as possible (on really dumb tasks like, "What is the colour of this goblin's hair?").

There are likely other parts of the actual paper's implementation where the ways I'm implementing it are lazy (because I'm doing this stuff for artistic/fucking around reasons, rather than to advance the field, or implement perfectly), and I think there are various interpretations of what "RLM" should mean. But I found the original paper very helpful, with lots of interesting ideas in, and think it's one where people can take what they need from.

For me: Kind of.

I find agents will reveal information marked as "lore" (or similar) almost immediately once it's in-context.

One thing I've tried when playing with longform fiction or screen stuff, where you have an expected wordcount or page count to structure around, and the audience has less agency - I've not experimented with this for a DnD-like interactive narrative - is to use an agent that will design context additions like "this information is revealed" to be triggered in X number of words/pages, and simply do not include it in-context until that time.

This needs heavy quality control from new, separate agents with further turns, also, or you end up with incomprehensible constantly-twisting narrative soup.

I expect you could do something similar for message-pair-based participatory storytelling formats like DnD.

Another approach I've tried which I think would be more suited to interactive storytelling is to have the agent tasked with designing characters/setting information include the twists a % of the time, and to include a trigger for that reveal. "If asked about X, they say Y".

Then I remove these from the context for all agents.

Then I run an agent which is looking for the pre-defined triggers each turn.

When the agent sees a pre-defined trigger appear in the story, it adds the pre-defined reveal back in to the context/lore.

Again, you need to run a quality control / superego across that to check it still works, and amend or remove it suitably if it doesn't! It gets convoluted fast.

"Revealed information" is, I think, significantly more of a strain on general immersion, because it inherently contains surprise for the reader or audience. So, I think tasking the agents doing any initial character or world design with "adding twists" makes sense, so revealed plot information isn't random-feeling or out-of-the-blue, but has intent and logic that fits the character or setting.

Uplift on "random dictionary words" (excluding 'stop words' & proper nouns) was ~20%; uplift on "random words from my local epub library" (same exclusion rules) was ~40%.

The random words from my local epub library (leans toward postmodern fiction) were definitely more evocative than the dictionary words when I eyeballed them.

I randomised each turn but kept the story prompt request the same across control, dictionary, personal library.

I must stress that I'm not claiming scientific method or certainty here - just sharing an approach that seemed to work well enough for me, and seemed like a reasonable conclusion: introduce noise, get more interesting output.

I haven't done the math but I think you'd need a much larger sample size than 1k per category to prove the uplift!

The whole thing was borne out of wanting to keep costs low! My favoured approach (last time I was doing this) is using only a tiny sliding context window based on message pairs, rather than tokens, and only for the agents that need it. Amending prose style, for example, shouldn't need context beyond the message it's working on, and then its system prompt.

For the models that require context, I personally found combining a tiny sliding window with a lazy version of the "Recursive Language Models" approach broke immersion least often and had a significantly lower cost. That + the "Id noise" + the strict agents also allowed cheaper models to overperform for me personally.

My lazy version of the RLM approach is basically just giving the agent a grep tool across the full message history & "lore" documentation created by agents, combined with repeated, low-context turns, and a "submit answer" tool for when it felt like it had finished working.

When I looked at the internals of what each agent turn looked like, it did look like a complete mess - but the context window only needs to surface the things it actually needs to know each turn.

Short outputs help a lot with immersion, too - brevity means there is a lot less you can get wrong, and also aids response time & cost.

It does take me an awful lot of prompt tuning to get what I want creatively from LLMs in any format, especially weaker models working in this chain, but I think that's likely always going to be true. Art can have rules, but that doesn't make it science :-)

The RLM approach is detailed here, and I've found it really useful for any cost-sensitive/long-context task: https://alexzhang13.github.io/blog/2025/rlm/

This is also the best approach I've found thus far when I'm seeing how well LLMs can form narrative content.

I don't frame its prompt as antagonistic though - I've found in the past (with weaker models, so YMMV) that this can be overly officious, sometimes blocking more creative outputs that you'd want to retain.

The structure I've found that works best is to have six or seven agents chained, each roughly mimicking a part of the mind, or a role in film production. Broadly:

- A high-temp "Id" agent, tuned to output only vaguely related noise. This really helps creativity.

- An "Ego" agent, who receives the "Id" noise and is then given the initial response task.

- A low-temp "Super-Ego" or "script supervisor" agent, who can grep back across longer contexts to check detail, and is asked to ensure that the initial response is within narrative reason. Not telling it that one role of the dialogue was "user" and one was "assistant" really helps with it not siding with the user.

- A "continuity editor" agent, who is explicitly tasked with world and character lore-checking, building and updating character & world MD docs, etc.

- A "prose editor" agent, whose sole task is to ensure it's tonally in-line with initial guidelines.

You can add more as needed, depending on what is important to you.

I think expecting competent narrative from a single model is a big ask. When writing and telling or performing a story, you have to engage several different parts of the brain, with very different tasks. The creative part of the brain has to have lots of bad ideas in it to surface a compelling idea; the parts dealing with immersion and/or realism have to incredibly restrained.

The Id agent is very important. By appending 100 tokens of noise to a prompt asking: "Write a short story about [subject]", then asking an LLM to blindly score the short stories generated across a range of creativity metrics (such as they can exist!) I personally saw a ~40% score increase vs control over 3k short stories.

I'm with the pedants on this one, in that I'd argue the prologue is inseparable from the novel, and the first line would be either be:

"(Supplied by a late consumptive usher..."

or:

"The pale Usher-threadbare in coat, heart, body, and brain; I see him now."

But straight fire either way. And, when it comes down to it, that "Call me Ishmael" is considered so consistently as the novel's opening line makes it the opening line. A testament to its power :-)

I was taught to use the term "action sentence" for an opening line that actually does something to a reader, and to aspire to that. I think it's a Gordon Lish-coined term, but it might not be.

Hard to pick an absolute favourite, but one would be:

"There was a time when I thought a great deal about the axolotls."

From Cortazar's short story Axolotl.

I don't know for certain why I find it so evocative, but it certainly draws me in, and makes me want to know more.

Quoting without looking attempt:

"I went to see them in the aquarium in the Jardin des Plantes and stayed for hours watching them, observing their immobility, their faint movements. Now, I am an axolotl."

There's definitely a tonne of signal in those, and it's a critique made from a place of strong support of your basic thesis. There's always been a tonne of signal in traditional customer support requests that goes un-used by most orgs, especially b2c orgs.

In case it's helpful: I always explained it to people I was training like this: All lean product theory comes from listening to the workers actually assembling the parts at Toyota.

Now, most digital products - whether the UI is graphical or linguistic - require a customer to work on an assembly line themselves. An onboarding flow is an assembly line and the user has tasks. Those users complain to agents (whether human or LLM) about their task on the assembly line. The purest implementation of lean philosophy would start with modelling these messages and conversations before it did anything else.

If I were you, I'd build a CRM. Intercom and its ilk charge ridiculous money for functionality that the people using it despise. The existing products in the space optimise for 'serve customers quickly' (increasingly irrelevant with LLMs) and not 'learning from your customers' (increasingly relevant as humans talk to customers less day-to-day). They are horrible to try to integrate into an established product development cycle (I've tried).

I think this makes the proposition easier to comprehend to a customer, the value-add more obvious, and allows you to undercut on pricing, rather than giving people a new bill for something they don't know if they need. The MVP of a CRM is also perhaps easier to build than it might seem initially. "Serve customers faster, cheaper, and learn from them in a highly configurable & meaningfully better way, giving your product iteration an advantage over your competitors". Building a CRM, crucially, allows you oversight of much more of the data - which then enables significantly more meaningful discovery.

This is the unsolved half of the coding agent space: what to actually build, what order to build it in, and why. It's really solvable from your starting point, and is potentially just as important/disruptive as the coding agent has been thus far - especially now that we suddenly have more lines of code than we know what to do with.

I'll shut up now - it's a fascinating space to me, so it's easy to get carried away about! Always happy to talk about stuff like this via email (in my profile) on the off-chance any of the above was useful, though :-)

I built an in-house version of this a couple of years ago for where I was working. My concern would be that by excluding observability, you might end up creating a really selective dataset, whose conclusions you're then asking companies to take seriously when allocating resources to different possible roadmaps.

My guess would be that agent logs would highlight obvious feature requests and bugs for smaller companies - like customers expecting an AI video editor product to be able to add subtitles to a video by itself.

For larger companies who deal with a higher volume of inbound customer support / agent requests, there will probably be big, noisy, already-known-by-the-team query clusters that make up big portions of the dataset - for example, "billing issue with my subscription". After those big clusters you'll likely have a really long tail of different queries, and - without deep observability - no real way to rank their importance. I also think you'd be unlikely to understand the root cause of the product issue in a complex developed product with lots of users solely from agent logs. Most product teams can't make good product decisions consistently, and they're working with a lot more data.

If coupled with staying out of evals (which, btw, I wouldn't find trust-building, if I were a potential customer of yours), I think that it might be difficult to provide genuine value in this space for larger orgs - without evals it's easily dismissed as just fancy & mostly-contextless sentiment analysis.

But I hope I'm wrong! I do think that (though each org's needs probably have to be catered to in a very boutique way) there are huge gains available by rolling LLMs & language analysis into existing product workflows, and that what you're pitching is absolutely a part of what companies should be doing. We are, of course, meant to actually listen to customers - and LLMs/agents should be making that easier, not harder. Absolute best of luck!

Honestly maybe my favourite part of being an author is being able to get briefly and deeply obsessed with any topic I choose - it's a rare privilege.

Beginning and then almost immediately dropping niche hobbies (eg flight simulators, poker, guns in the linked post) is transformed from something a spouse or partner might see as an undesirable and potentially annoying personality trait, into: "this is research, darling, it's my job", which is probably significantly more annoying.

It is easy to make mistakes with a verb one might not think to question like "cocked". Of course, ideally, you'd question every word used, so that the % of readers who understand its full associated meaning don't have their immersion in a story suddenly & painfully torn away.

To be less glib, I find that when speaking about a topic from a character's perspective - in dialogue or narration - a relatively important part of empathising with their point-of-view is understanding the physical and linguistic structure of their world. Sometimes I find there's no way to do this without putting hundreds of hours into understanding the tools they would use or the way they would live. Write what you know!

Late replying - I don't think you should have been downvoted so much. You're right that I was using a comically simple example for comic effect (though I'm certain it is something that happens a lot), and also that LLMs are very interesting thought tools. Private dialogue is really analogous to thinking. There's nothing in your comment that suggests posting a critically unexamined, verbatim snippet of one's private LLM dialogue.

Inconsistent capitalisation ('Twitter' vs 'reddit'); subtly using the outdated name for 'Twitter' as most humans do; the genuinely hard-to-parse final clause of the comment.

Though I note it didn't say "read comments by other humans", only "read comments by humans", so confirmed AI.

I think the guidelines here work quite well, and expect a good-faith interpretation, which they mostly receive.

I think you're asking for some sort of empirical verification of "this is / is not LLM text" (which seems impossible), but there's no real reason to expect the existence of LLMs to change that this website is, generally, interacted with in a good-faith way. People are really good at calling others out on here -- I doubt that will change.

Quite! It's very easy to send a HN link to one of our new artificial friends to see what they have to say about it. Subsequently publicly posting the inference variation you receive strikes me as very self-centered. Passing it off as your own words - which the majority seem to - is doubly bizarre.

It's very funny to imagine people prompting: "Write a compelling comment, for me, to pass off as my thoughts, for this HN news thread, which will attract both upvotes and engagement.".

In good faith, per the guidelines: What losers!

I wonder if the adverts in the "personal super-assistant", per the blog post, ("that helps you do almost anything"!) will have the same triggers as the shopping assistant, which pops up underneath messages right now in the web UI.

When first trying 5.2, on a "Pro" plan, I was - and still am - able to trigger the shopping assistant via keyword-matching, even if the conversation context, or the prompt itself, is wildly inappropriate (suicide, racism, etc).

Keyword-matching seems a strange ad strategy for a (non-profit) company selling QKV. It's all very confusing!

Hopefully, for fans of personal super-assistants--and advertising--worldwide, this will improve now that ads have been formalised.

This is fun!

Given online is now bot-riddled, I half-finished something similar a while back, where the game was adopting and 'coaching' (a <500 character prompt was allowed every time the dealer chip passed, outside of play) an LLM player, as a kind of gambling-on-how-good-at-prompting-you-are game. Feature request! The rake could pay for the tokens, at least.