Heh, this one brought a smile - well played.
HN user
thadt
Yes, we've been using Transkribus for this extensively. My wife is a historian who spends quite a bit of time sorting through old letters and diaries, and it has been a considerable quality of life improvement.
Even if you are able to read someone's scratches, having a model to do the bulk lifting saves your eyes a lot of squinting. One thing that makes Transkribus useful for research vs a chat interface is that it can line up its interpretation alongside the original image so you can examine its work directly.
Whereas if we're talking about lossy compression (as is the person to whom you replied) we certainly can compress arbitrary data - almost as much as we want.
The hard question, then, is how much the decompressed output looks like the original.
I do not want a "connection" with a business.
That's because we don't make connections with businesses - we make connections with people. That one nice hair dresser. The pharmacist that goes out of her way to make sure my mom's medicines are right. A cashier that's just pleasant to talk with.
Years ago I did most of my grocery shopping at Target. I cared almost nothing about Target the chain. Or Target the super store. I did care about Betty the checker and wanted to know how her grandkids were doing this week.
Right - and if I ever go to raise VC then that's a guy I'll want to talk to. Even if they themselves don't invest, their recommendation would be valuable.
Life's too short to screw people over - reputation is one of the few things that last after we're gone.
Agreed. I scanned a short book with my phone, and a dedicated scanner would have been nice to have.
But with page flattening and separation and automated capture, it went much faster than I would have thought. If I were going to do a lot more, I'd want something like a scan tent [1]. It's not as ergonomic as a dedicated solution, but in 2026 a phone and some light can get you a lot of the way there, pretty fast.
Maybe not years ago, but scanning documents with the phone in your pocket has become incredibly efficient. That combined with AI transcription and indexing for search makes such a project faster and cheaper in 2026 than at almost any other time in the past.
That’s not necessarily the case. It depends on how autonomous your drone is and what you need to guide it to do…
They're doing transcription, not translation - so, turning someones pages of scrawled script into typewritten text. They have around 20 people nationwide that are able to do this. Most of them are older volunteers who aren't all that interested in computer assistance, but about a third of them have started leveraging the newer AI tools and it has accelerated their throughput significantly.
Having a 'best guess' at the lettering is really handy - in some cases the writing is really rather difficult to make out at all. Even being able to run something as simple as frequency analysis on stroke patterns would be a massive benefit.
At this point they're becoming throughput bound on the scanning process. Diaries are digitized since the archive is in one place and their transcription experts are spread out over the country.
AI had been a super useful for processing historical data. Interviewed a volunteer last month from the diary archive in Germany, and they're using supervised AI for diary transcription. Going from (old) personalized hand script to text is a lot of work, even for experienced transcribers. Being able to automate the first pass of that has been a huge boon to their processing pipeline.
Ironically, the reason I used Google the most then was because it indexed Usenet while so many other parts of the Internet offered by the other engines were "slop". My, how the turn tables.
In general I agree and suspect that memory safety is a tool that will continue to pay dividends for some time.
But there are tradeoffs and more ways to write correct and 'safe' code than doing it in a "memory safe" language. If frontier models indeed are a step function in finding vulnerabilities, then they're also a step function in writing safer code. We've been able to write safety critical C code with comprehensive testing for a long time (with SQLite presenting a well known critique of the tradeoffs).
The rub has been that writing full coverage tests, fuzzing, auditing, etc. has been costly. If those costs have changed, then it's an interesting topic to try to undertand how.
So the intersting question: are we long term safer with "simpler" closer to hardware memory unsafe(ish) environments like Zig, or is the memory safe but more abstract feature set of languages like Rust still the winning direction?
If a hypothetical build step is "look over this program and carfully examine the bounds of safety using your deep knowledge of the OS, hardware, language and all the tools that come along with it", then a less abstract environment might be at an overall advantage. In a moment, I'll close this comment and go back to writing Rust. But if I had the time (or tooling) to build something in C and test it as thoroughly as say, SQLite [1], then I might think harder about the tradeoffs.
Bananas are like XML that way. If you're not getting the results you want, you're just not using enough of them.
They specifically call out Yingxin Li[1] in the acknowledgements section of the paper?
Clarification - in the past when I've written high performance data tools in JS, it was almost entirely to support the use case of needing it to run in a browser. Otherwise, there are indeed more suitable environments available.
To your question, I was about to point out Firefox[1], but realized you clarified 'mainstream'[2]...
As someone who spent hours playing Jedi Knight with friends and lots of mods, allow me to say - thank you :)
Browsers
Getting a broad overview of "world history" is useful for having basic context for large events, but, IMHO, history gets so much more interesting and educational when you're deep into individual people's lives and stories. I'm probably a bit biased, but tend to agree with the suggestions that you pick a time and place and dive deep into an individual or event that catches your fancy.
Oh man, have I gotten to read a lot of history recently.
And also fiction.
Frequently at the same time.
I used to love using em dashes.
I still do - but I used to, too.
In the 80's we had a way to deal with that kind of thing [1]. Just gotta practice to get the technique right.
A cursory search doesn't seem to turn up anything solid.
But it would be interesting to know. I'll be in the German diary archive in March and made a note to keep an eye out for it.
Zyklon-B wasn’t much of a secret - it was used all over the place as a pesticide. Most soldiers would have been about as familiar with it as we would with Raid spray or bug traps.
Nuclear measurements, where the speed of a gamma ray flying across a room vs a neutron is relevant. But that requires at least nanosecond time resolution, and you’re a long way from thinking about NTP.
What does a "digital education" look like, specifically?
Having spent several years teaching kids to code everything from games to lightbulbs on Chromebooks, I can confirm that there are certainly difficulties - but they're tradeoffs. I could spend my time coming up with a way to work through the platform restrictions, or I could spend my time maintaining a motley crew of devices and configurations. Having done it both ways, they both have different pain points.
Well, it depends on the granularity of the time scale right? When you're measuring milliseconds, then the cable length probably isn't a thing factoring into your latency calculation.
When we're measuring time on the scale of nanoseconds then, yes, cable length is definitely something we care about and will reliably show up in measurements. In some situations, we not only care about the cable length, but also its temperature.
When I see juries in American courts, for example when I've served on one, it seemed like a group of people who take their job quite seriously. You are correct in that what a jury gets is a very curated set of information. The intention being to keep the jury focused on the details of a very specific situation with evidence that is processed in such a way as to be as "reliable" as possible.
It is by no means an accurate or incorruptible system. When we design and prove out a better, more robust alternative, I'll be eager to learn about it.
Well, it's fascinating right? Can we cleanly separate out military, politics and sociology? A whole lot of military capability comes down to not just technology and tactics, but the entire culture and makeup of the people. When we think of famous examples such as the Spartans at Thermopylae, the whole Spartan culture is important to understanding the how and why.
Context is really important. As you correctly note, many of the people the Romans were conquering could be even more ruthless as well (by 2025 standards). My point was more that historians wear a lot of different hats, depending on what they're doing. When you're wearing your 'investigator hat' learning how and why things worked, your thoughts might be different than when you're wearing your 'builder hat' and thinking about the society you might want to live in today (and tomorrow). It isn't a contradiction to weigh the tradeoffs that various people in history have made when designing their culture (and politics, and military capabilities).
Weird right? My wife studied Nazi Germany, and doesn't really like the Nazis either.
History is messy - we can and should learn from those that came before, both the good and bad. One can both admire the things the Romans accomplished while simultaneously despising the way they went about accomplishing them. It isn't a contradiction.