HN user

benbreen

45,773 karma

benjaminpbreen.com

email: breen85 [at] gmail [dot] com

Posts2,188
Comments504
View on HN
resobscura.substack.com 8h ago

Quality non-fiction books are the antithesis of AI slop

benbreen
36pts16
en.wikipedia.org 5d ago

Pushinka

benbreen
50pts15
www.nytimes.com 15d ago

The revenge of the philosophy majors

benbreen
164pts270
www.historytoday.com 16d ago

The Victorian War on Rabies

benbreen
25pts21
www.pnas.org 21d ago

Natural history on canvas: Brueghel knew about bird-eating noctule bats

benbreen
13pts0
www.historytoday.com 22d ago

The Meadows of Medieval Summer

benbreen
3pts0
www.theatlantic.com 28d ago

Paradise Revisited: What Darwin Saw in the Galápagos

benbreen
48pts2
aeon.co 29d ago

Words, Words, Words

benbreen
35pts15
publicdomainreview.org 1mo ago

The Art of Kite Flying (1430–1929)

benbreen
29pts11
www.nytimes.com 1mo ago

Carlo Ginzburg, Who Told the History of the Obscure, Dies at 87

benbreen
18pts5
www.historytoday.com 1mo ago

Cypriot Graffiti in Ancient Egypt

benbreen
4pts0
bethmathews.substack.com 1mo ago

Vacuum-Form Signage

benbreen
117pts24
www.metropolitanreview.org 1mo ago

Easy Writer: On Ted Geltner's Biography of Denis Johnson

benbreen
3pts0
arstechnica.com 1mo ago

Blue Origin’s New Glenn rocket exploded during a static fire test

benbreen
209pts8
www.the-hinternet.com 1mo ago

The Science of Weather and the Nature of Science

benbreen
17pts2
en.wikipedia.org 1mo ago

R vs. Dudley and Stephens

benbreen
1pts0
www.diffuseai.pub 2mo ago

The Structural Barriers to AI Lawyers

benbreen
66pts85
resobscura.substack.com 2mo ago

What Is Happening to Publishing?

benbreen
83pts59
en.wikipedia.org 2mo ago

Russian–American Telegraph

benbreen
3pts0
publicdomainreview.org 2mo ago

Twilight of the Velocipede: Typesetting Races Before the Age of Linotype

benbreen
49pts6
resobscura.substack.com 2mo ago

Why they stopped building wooden stupas: on survivorship bias in history

benbreen
5pts0
spectrum.ieee.org 2mo ago

Archivists Turn to LLMs to Decipher Handwriting at Scale

benbreen
3pts0
michaelnotebook.com 2mo ago

Notes on Tanya M. Luhrmann's Book 'How God Becomes Real'

benbreen
1pts0
cookingthehashishcookbook.com 2mo ago

Cooking the Hashish Cookbook

benbreen
9pts3
blog.mozilla.ai 2mo ago

Sovereign AI: Control, Choice, and Why It Goes Beyond Geopolitics

benbreen
4pts0
www.sciencehistory.org 2mo ago

The Thinking Plant's Man (2025)

benbreen
59pts19
www.rightscon.org 2mo ago

A statement about why RightsCon 2026 will not take place in Zambia

benbreen
105pts48
resobscura.substack.com 2mo ago

Thoughts on Historical Language Models and Talkie-1930

benbreen
17pts3
porticoquarterly.com 2mo ago

The Last of the Lost Generation

benbreen
25pts6
www.newyorker.com 3mo ago

When your digital life vanishes

benbreen
75pts20

Author here - just wanted to clarify in case there is any confusion that those two (intentionally bad/weird) figures of speech about the echo and the trouts in the redwood roots were human written, by me! I wrote them as parody of an AI trying to do "literary" writing. The actual (probable) AI written excerpts are below that part.

I was thinking about the immortal Twin Peaks line "there's a FISH... in the PERCOLATOR" when I wrote the trout one.

Just wanted to flag that the works of Ian Hacking, especially The Emergence of Probability (1975) and The Taming of Chance (1990) are excellent on this. Dense and challenging at times but also well written and the product of a very original mind.

The latter book has a Wikipedia page with some more info - was surprised to see Hacking not mentioned here since the featured article is partly based on his work: https://en.wikipedia.org/wiki/The_Taming_of_Chance

Gemini 3.1 Pro 5 months ago

One underrated thing about the recent frontier models, IMO, is that they are obviating the need for image gen as a standalone thing. Opus 4.6 (and apparently 3.1 Pro as well) doesn't have the ability to generate images but it is so good at making SVG that it basically doesn't matter at this point. And the benefit of SVG is that it can be animated and interactive.

I find this fascinating because it literally just happened in the past few months. Up until ~summer of 2025, the SVG these models made was consistently buggy and crude. By December of 2026, I was able to get results like this from Opus 4.5 (Henry James: the RPG, made almost entirely with SVG): https://the-ambassadors.vercel.app

And now it looks like Gemini 3.1 Pro has vaulted past it.

Thank you, this sort of insight is exactly why I've felt such kinship with what software engineers like Karpathy and Simon Willison have been writing lately. It seems obvious to me that there is something special and irreplaceable about the thought processes that create good code.

However, I think there is also something qualitatively different about how work is done in these two domains.

Example: refactoring a codebase is not really analogous to revising a nonfiction book, even though they both involve rewriting of a sort. Even before AI, the former used far more tooling and automated processes. There is, e.g., no ESLint for prose which can tell you which sentences are going to fail to "compile" (i.e., fail to make sense to a reader).

The special taste or skillset of a programmer seems to me to involve systems thinking and tool use in a different way than the special taste of a writer, which is more about transmuting personal life experiences and tacit knowledge into words, even if tools (word processor) and systems (editors, informants, primary sources) are used along the way.

Sort of half formed ideas here but I find this a really rich vein of thought to work through. And one of the points of my post is that writing is about thinking in public and with a readership. Many thanks for helping me do that.

I don't have a good answer to your question, but I do think it might be comparable, yes. If you had good taste about what to get Opus 4.6 to write, and kept iterating on it in a way that exposes the results to public view, I think you'd definitely develop a more fine grained sense of the epistemological perspective of a writer. But you wouldn't be one any more than I'm a software developer just because I've had Claude Code make a lot of GitHub commits lately (if anyone's interested: https://github.com/benjaminbreen).

Love the faux Nature article: https://sw.vtom.net/hn35/pages/90098000.html

Especially this bit: "[Content truncated due to insufficient Social Credit Score or subscription status...]"

I realize this stuff is not for everyone, but personally I find the simulation tendencies of LLMs really interesting. It is just about the only truly novel thing about them. My mental model for LLMs is increasingly "improv comedy." They are good at riffing on things and making odd connections. Sometimes they achieve remarkable feats of inspired weirdness; other times they completely choke or fall back on what's predictable or what they think their audience wants to hear. And they are best if not taken entirely seriously.

Was going to say - it would be fascinating to go a step further and have Gemini simulate the actual articles. That would elevate this to level of something like an art piece. Really enjoyed this, thank you for posting it.

I'm going to go ask Claude Code to create a functional HyperCard stack version of HN from 1994 now...

Edit: just got a working version of HyperCardHackerNews, will deploy to Vercel and post shortly...

I read Ulysses Grant's memoirs awhile back, and loved his description of being in San Francisco in the 1850s. (Another tidbit I loved is that he imagined an alternate path for his life where he would have settled down in the Bay Area and become a math teacher):

"The immigrant, on arriving, found himself a stranger, in a strange land, far from friends. Time pressed, for the little means that could be realized from the sale of what was left of the outfit would not support a man long at California prices. Many became discouraged. Others would take off their coats and look for a job, no matter what it might be. These succeeded as a rule. There were many young men who had studied professions before they went to California, and who had never done a day's manual labor in their lives, who took in the situation at once and went to work to make a start at anything they could get to do. Some supplied carpenters and masons with material—carrying plank, brick, or mortar, as the case might be; others drove stages, drays, or baggage wagons, until they could do better. More became discouraged early and spent their time looking up people who would 'treat,' or lounging about restaurants and gambling houses where free lunches were furnished daily."

I think the median response is something between revulsion and mild dislike if it's spoken about in the context of the classroom. But there are also a pretty significant group of people who find it interesting as a potential research tool. (Also the question of what "it" is matters a lot here - if you asked people in history about ChatGPT, the response would be massively different than if you asked about machine learning tools for OCR and data mining historical documents, which there is a lot of support for).

Personally I think it absolutely will lead to major changes in historical research. The transcription and translation abilities of transformer models alone are already leading to significant changes and advances. For instance, I'm working on a post about new transformer based OCR tools like Leo that are geared specifically for historical research and led by historians (https://www.tryleo.ai - I'm not involved in the project, just an interested observer).

IMO AI tools will definitely still be used by a minority of historians in a 5-10 year horizon. Historical research is not like some STEM fields where there is a lab-base culture oriented around adopting new tech and finding applications quickly. It's a lot more of a solo, idiosyncratic process of personal research and that is partly why I like it, but it also means that uptake of new tools is much slower. That said, historians do use technology and digital tools all the time and are not inherently adverse to it. It's interesting reading history books from the 1970s, like the works of Lawrence Stone (https://en.wikipedia.org/wiki/Lawrence_Stone) and seeing the footnotes about how the data was encoded in punchcards and analyzed by mainframes. I expect we will be seeing history books by the end of the 2020s that use custom data analytics and tagging tools developed by the historians themselves using vibe coding.

Thanks for the question, will be writing more about this. Feel free to get in touch any time.

Agreed. Or at least the best non-Shakespeare play I've ever read, and among the best works of 20th century literature. I really can't recommend Arcadia highly enough. It's both deeply moving and extremely thought-provoking, clever, and intellectual interesting.

Currently working on an idea like this, but its a history simulator for educational use - I find that LLMs respond rather well to being grounded in a specific time/setting in real world history, as opposed to being told to roleplay a fictional setting. The latent space of any fictional world is close enough to other fictional worlds that they will rapidly slide off into other similar-sounding settings. Whereas if you return them to a super-specific historical context each go-around ("The time is now 3:13 pm. It is August 3, 1348. You are currently simulating the functioning of a small vineyard in Normandy. The farmer, [NPC name], is looking for helpers in the fields") they will be able to pull from a pretty solid baseline of background knowledge and do a decent job with it.

Some fun things I've been experimenting with is 1) injecting primary sources from a given time and place into the LLMs contex to further ground it in "reality" and 2) asking the LLM to try to simulate the actual historical language of the era - i.e. a toggle button to switch to medieval French. Gemini flash lite, the only economical model for this sort of thing, is not great at this yet but in a year or so I think it will be a fascinating history and language learning tool.

Have been meaning to write this project up for HN but if anyone wants to try a very early version of it, it's here - you can modify the url to pick a specific year and region or just do the base url for a fully random spawn, i.e. here is Europe in 1348: https://historysimulator.vercel.app/1348/europe

Exactly, this is the reason why I struggle with this sort of solution to the problems we are all facing in education currently. On the surface it seems to make sense, but the blue book exam is entirely artificial and has basically no relationship to real world skills (subconsciously, even the quality of student handwriting handwriting could influence how graders assess blue books). Even leaving AI aside, anyone writing anything nowadays is using Google and Wikipedia and word processors, so why constrain those?

Oxford and Cambridge have a "tutorial" system that is a lot closer to what I would choose in an ideal world. You write an essay at home, over the course of a week, but then you have to read it to your professor, one on one, and they interrupt you as you go, asking clarifying questions, giving suggestions, etc. (This at least is how it worked for history tutorials when I was a visiting student at an Oxford college back in 2004-5 - not sure if it's still like that). It was by far the best education I ever had because you could get realtime expert feedback on your writing in an iterative process. And it is basically AI proof, because the moment they start getting quizzed on their thinking behind a sentence or claim in an essay, anyone who used ChatGPT to write it for them will be outed.

I realize this isn't the same thing as your point about images as part of training data, but just flagging it in case anyone isn't aware: Claude Code lets you copy and paste images into terminal. I've been designing a "universal history simulator" game for use in my history classes lately, and it is really helpful to be able to make a mockup of a ui change I want and then paste it in, rather than trying to explain it verbally. Also good for debugging graphics issues.

[dead] 12 months ago

Sorry just realized I posted the wrong URL - resubmitting with correct url now.

Yes, I think this kind of combination is where higher ed is going to land. I've been talking to a colleague lately about how social skills and public speaking just got more important (and are things we need to focus on actually teaching). Likewise, I think self-directed, individualized humanistic research is currently not replicable by AI nor likely to be - for instance, generating an entirely new historical archive by conducting oral history interviews. Basically anything that involves operating in the physical world and deploying human emotional skills.

The unsolved issue is scale. 5-10 minute Q&As work well, but are not really doable in a 120 student class like the one I'll be teaching in the fall, let alone the 300-400 student classes some colleagues have.

The ability to seamlessly integrate generated images is fascinating. Although it currently takes too long to really work in a game or educational context.

As an experiment I just asked it to "recreate the early RPG game Pedit5 (https://en.wikipedia.org/wiki/Pedit5), but make it better, with a 1970s terminal aesthetic and use Imagen to dynamically generate relevant game artwork" and it did in fact make a playable, rogue-type RPG, but it has been stuck on "loading art" for the past minute as I try to do battle with a giant bat.

This kind of thing is going to be interesting for teaching. It will be a whole new category of assignment - "design a playable, interactive simulation of the 17th century spice trade, and explain your design choices in detail. Cite 6 relevant secondary sources" and that sort of thing. Ethan Mollick has been doing these types of experiments with LLMs for some time now and I think it's an underrated aspect of what they can be used for. I.e., no one is going to want to actually pay for or play a production version of my Gemini-made copy of Pedit5, but it opens up a new modality for student assignments, prototyping, and learning.

Doesn't do anything for the problem of AI-assisted cheating, which is still kind of a disaster for educators, but the possibilities for genuinely new types of assignments are at least now starting to come into focus.

Gwern and others who have dug into it this far might be interested by this footnote in the Crespo article: "I have tried to lay my hands on the original version of the conversation, as I am sure Simon did, too. I contacted Gabriel Zadunaisky, who, as the article explains, participated in the meeting. He is a professional translator. I asked him for the original version, and he replied on WhatsApp: 'Mr. Crespo: I am very sick. Unfortunately, I am unable to provide you with the information requested.' My hypothesis is that Zadunaisky translated the conversation directly from the recorded version and that this original version has been lost."

My read is that most likely, it was recorded on an old school reel-to-reel tape recorder. It's entirely possible that the tapes are still sitting on a shelf somewhere in Argentina, though the chances of actually tracking them down are pretty low. I worked with some reel-to-reel tapes that Alan Ginsberg made (now held at Stanford) in the mid-60s (including one where he is talking to Bob Dylan!) and they held up pretty well. Had to use audio editing software to remove tape hiss, but they were not as badly preserved as I expected.

PaperBench 1 year ago

I've been developing a more elaborate variation on the "chat with a pdf" idea for my own use as a researcher. It's mostly designed for a historian's workflow but it works pretty well for science and engineering papers too. Currently Flash 2.0 is the default but you can select other models to use to analyze pdfs and other text through various "lenses" ranging from a simple summary to text highlighting to extracting organized data as a .csv file:

https://source-lens.vercel.app

(Note: this is not at all a production ready app, it's just something I've been making for myself, though I'm also now sharing it with my students to see how they use it. If anyone reads this and is interested in collaborating, let me know).

Such a confusing comment, because when I enter the text from case study #1 into Deepl, it's very clearly much worse than what Claude or GPT4o can come up with (the first few lines from Deepl are: "With his expositions to all the Tables, particularly of the quality of the countries, et of the most notable things, to be found in them. Which Tables, can be, and t are taught to reduce' together" and so on).

Likewise with using Google translate on both case studies #1 and #2 - the results are self-evidently far worse. In both cases there were multiple errors in each line and in case study #2 it was entirely unable to transcribe or translate the title line. If you see this, please email me at bebreen [at] ucsc dot edu to share the better results you are seeing because I genuinely am interested and open to using alternative tools - I just am not seeing what you are seeing, apparently.

In terms of typos not changing the meaning, yes naturally a real human needs to double check absolutely everything if it's being used in research. We agree on that - the point is simply that this significantly speeds up the initial research process, not that it replaces the expertise necessary to, for instance, double check that a name or year is transcribed correctly. A huge amount of historical research is simply about skimming through documents looking for relevent info to zero in on - this is where LLMs can really help.

I agree, I probably should've gone into more detail on the actual case studies and implications. I may write this up as a more academic article at some point so I have space to do that.

To your point about OCR: I think you'll find that the existing OCR tools will not know where to begin with the 18th century Mexican medical text in the second case study. If you can find one that is able to transcribe that lettering, please do let me know because it would be incredibly useful.

Speaking entirely for myself here, a pretty significant part of what professional historians do is to take a ton of photos of hard-to-read archival documents, then slowly puzzle them out after the fact - not by using any OCR tool (because none of them that I'm aware of are good enough to deal with difficult paleography) but the old fashioned way, by printing them out, finding individual letters or words that are readable, and then going from there. It's tedious work and it requires at least a few days of training to get the hang of.

If anyone wants to get a sense of what this paleography actually looks like, this is something I wrote about back in 2013 when I was in grad school - https://resobscura.blogspot.com/2013/07/why-does-s-look-like...

For those looking for a specific example of an intermediate-difficulty level manuscript in English, that post shows a manuscript of the John Donne poem "A Triple Fool" which gives a sense of a typical 17th century paleography challenge that GPT-4o is able to transcribe (and which, as far as I know, OCR tools can't handle - though please correct me if I'm wrong). The "Sea surgeon" manuscript below it is what I would consider advanced-intermediate and is around the point where GPT-4o, and probably most PhD students in history, gets completely lost.

re: basically perfect, the errors I see are entirely typos which don't change the meaning (descritto instead of descritta, and the like). So yes, not perfect, but not anything which would impact a historical researcher. In terms of existing tools for translation, the state of the art that I was aware of before LLMs is Google Translate, and I think anyone who tries both on the same text can see which works better there.

re: "irrelevant books," there's really no way to make an objective statement about what's relevant and what's not until you actually read something rather than an AI summary. For that reason, in my own work, this is very much about augmenting rather than replacing human labor. The main work begins after this sort of LLM-augmented research. It isn't replaced by it in any way.

Thank you! Have been a big fan of your writing on LLMs over the past couple years. One thing I have been encouraged by over this period is that there are some interesting interdisciplinary conversations starting to happen. Ethan Mollick has been doing a good job as a bridge between people working in different academic fields, IMO.

Author here, I agree that my read may not be correct either. It’s tough to make out. Although keep in mind that “ph” is used in Latin and Greek (or at least transliterations of Greek into the Roman alphabet) so in an early modern medical context (I.e. one in which it is assumed the reader knows Latin, regardless of the language being used) “ph” is still a plausible start to a word. Early modern spelling in general is famously variable - common to see an author spell the same word two different ways in the same text.

Author here, I had the same question and looked into it. The author of that comment seems to be onto something because hygrine is indeed found in nightshades as well as in coca. Interesting.