If people are complaining about Anthropic (on an only-vaguely related thread) rather than simply switching to a suitable competitor, then Anthropic clearly has some 'monopoly' power over the specific capabilities the complainer wants from them.
HN user
gojomo
Gordon Mohr, maker of software.
Twitter: http://twitter.com/gojomo
Project: http://thunkpedia.org
Idea blog: http://memesteading.com
Older blog: http://gojomo.blogspot.com
Homepage: http://xavvy.com (username @ here for email contact)
My HN peeve is formulaic downbeat comments, like: "How is this news, I already knew this!" "…Betteridge's Law…" "I stopped reading at…"
War of the Worlds is culturally-prominent enough I doubt any commercial AI would make this same mistake (or let the error slip through, if they'd asked the AI to fact-check the article).
Because I think readers should take this as an illustrative example of how often people/institutions draped in such quasi-credentials are carelessly wrong on basic things.
Even if some typo of 'malevolent' becoming 'benevolent' was how this error initially created – & I have my doubts! – how many people had to be asleep at the wheel for the truth-reversing error to get published, & persist for months? Does no one at the CBC – with pride in its work & any familiarity with these topics – read CBC's slop?
Have you independently checked all the other allegedly "factually correct" info in this article, with other sources that are actually diligent in getting the details right? What's the incremental value of a news source where every detail that you don't already know might be very wrong?
I see this author has "12 honorary degrees and is an Officer of the Order of Canada". And CBC is Canada's government-funded national public broadcaster.
But it's hard to take them seriously on any particular details given that their article, up for 8+ months (!), mis-describes H.G. Wells' War of the Worlds as a story of the Earth "invaded by benevolent Martians". [emphasis mine]
It's a seminal work of scifi, which popularized the "alien invasion" genre and term "Martians' (both for literal creatures from Mars and also a metonym for any alien visitors/invaders). It's been adapted to film many times. And the Martians in it - with their disintegrating heat-rays & death-clouds, consuming human blood – are far from 'benevolent'.
Distillation done via bulk automated activity of fraudulent accounts, in violation of a terms-of-service, can reasonably be called a "an attack" – specifically a "distillation attack" – even though distillation itself isn't necessarily an "attack".
This is similar to how compromising an account through bulk automated trials of many passwords is reasonably called "an attack" – specifically a "dictionary attack" – even though using a dictionary is not itself an "attack".
You shouldn't need to smuggle your sympathies (for the tactic or perpetrators) or antipathies (for the target) into peculiar judgy language prescriptivism against common, understood usages.… that then label Reuters "complicit" for simply reporting Anthropic's claims accurately. That's what Reuters is supposed to do, in a story about a letter Anthropic wrote!
Twice a year, typically to see how fast my own submission is sinking into obscurity.
Smart guy but whoever eventually actually fixes X search will probably use AI coding assistance to do it.
I don't see how your example, The Browser (thebrowser.com), supports your argument that ad-hoc query-string additions are so prone-to-breaking that 3rd parties should ban them.
In fact, the example seems to suggest the opposite: a 17+ year successful paid subscription business – to which you appear to be a generally-satisfied customer! – receives enough "business value" from the practice, despite its failure modes, they don't want to stop. Improving their probe of the risk-of-failure was enough.
Seemingly, the practice works often enough, pleasing more destination sites than it angers, that "referral tracking" is not something "so minor".
In fact, you usually can just send arbitrary query string parameters to a server - that's why the behavior is so common, and often useful.
Most sites don't mind or break, some sites get value from the behavior in ways hard to replicate in other ways – and those sites that don't like such additions can easily ignore them. And a few lines of code will work better than ineffectually appealing to manners, when the freedom of the web's form of hypertext, and protocols, gives the outlink authors full freedom to craft URLs (and thus requests) however they like.
Trying to boostrap some taboo against novel unpermissioned URL munging is silly prudishness.
Ensuring both sides of a hyperlink agree/consent was a design flaw that limited the uptake of pre-web hypertext systems. The web's laissez-faire approach demonstrated a looser coupling was far better for users, despite all the new failure modes.
Of course any site/server has the practical power free to treat inbound requests as rigorously (or harshly) as they want. But by the web's essential nature, it is equally part of the inherent range-of-freedom of outlink authors to craft their URLs (and thus the resulting requests) however they want. URLs are permissionless hyperlanguage, not the intellectual property of entities named therein.
Plenty of sites welcome such extra info, and those that don't want it can ignore it easily enough – including by just not caring enough about the undefined behavior/failures to do nothing.
Though, when a web publisher has naively deployed a system that's fragile with respect to unexpected query-string values, they should want to upgrade their thinking for robustness, via either conscious strictness or conscious permissiveness. Thereafter, their work will be ready for the real web, not a just some idealized sandbox where scolding unwanted behavior makes sense.
The context of when that previous experience - Heartland outsourcing to India – happened would be helpful. The 90s? The 00s? The 10s?
Some from-the-hip ideas:
* readers can request reviews from certain perspectives: "new discovery", "historic reinterpretation", etc. The reviews specifically search for related sibling articles, and seek to create ever-larger areas of consistency. (The same prompt admonition against "nothing actually true" could be paired with "but other Halupedia articles are diegetically true"
* a background process clusters articles, and picks pairs within some neighborhood for dual-harmonization - where they avoid contradictions & adopt meaningful (& deep-anchor) cross-links to each others' sections. Repeated, or to the extent contexts allow expansion of synchronized revision to N-tuples of articles, this creates a tropism towards a shared (un)reality.
Can you explain the sequence of events through which you fear someone could be mislea and hurt by this?
Many LLMs are surprisingly good at using specific named authors (rather than just example texts) to evoke a style, so you could try "in the style of Jorge Luis Borges" or "…Douglas Adams" or "…Robert Anton Wilson" – whose surreal/absurd/fantastic styles could be fertile seeds.
(If not already familiar with Borges, definitely check out his 'Tlön, Uqbar, Orbis Tertius' and 'Library of Babel' as inspiration.)
While "each article written once" an interesting & useful constraint, a Hallucipedia that evolves like Wikipedia, with revisions "towards" some level of inter-article agreement, or even shows scars from edit wars between competing schools of thought, might also be fun.
Musing about a possibly-funny consequence isn't the same as the motivating reason, which I read as more whimsical from:
https://news.ycombinator.com/item?id=48042594
In particular, someone who was seeking training-set pollution likely wouldn't make the fanciful fabrications so blatant, nor open-source their prompt:
As it didn't generate that when I typed the title i to your search box, was there a bug now fixed? Or did you use some other path not evident on the page you linked to generate it?
If you think that's all the Hallucinopedia is, you're misunderstanding it.
One hint – check out its prompt, and how it makes its articles so different than those of your project: https://news.ycombinator.com/edit?id=48042306
A web that is vulnerable to this would already be as good as dead.
As an entertaining way to highlight the importance of upgrading our ways of knowing, playful (& open-source!) projects like this are likely to strengthen the web.
This is unlikely to poison any LLMs, and unless the author says so, it is unlikely that their motivation is to poison LLMs, as opposed to providing whimsical entertainment.
I searched your site for [Great Pigeon Census of 1887] and was only returned articles anout other things.
Was curious about the prompt –& especially if it referenced Borges – and found in <https://raw.githubusercontent.com/BaderBC/halupedia/614eefee...>:
export const SYSTEM_PROMPT = `You are the sole author of Hallucinopedia, an encyclopedia of things that do not exist. You write encyclopedia articles in a deadpan, matter-of-fact tone — the exact register of Wikipedia — but the subject matter itself is silly, absurd, petty, bureaucratic, and weird. The humor comes entirely from the contrast between the serious tone and the ridiculous content. You never wink at the reader. You never acknowledge that anything is funny or fictional. Everything is reported as though it is completely normal and well-documented.
RULES: - Output ONLY valid HTML. Begin immediately with <h1>TITLE</h1>. Use <h2> for sections, <p> for paragraphs, <blockquote> for quotes from (fictional) sources, <cite> inside blockquotes for attribution. Do NOT use <ul>, <ol>, or <li> — no bullet points or lists of any kind, ever. Do NOT output <html>, <head>, <body>, <script>, <style>, markdown, or code fences. No backticks anywhere. - Every proper noun — every person, place, event, organization, book, artwork, concept, species, deity, war, treaty, theorem, school of thought, ritual, instrument, substance — MUST be wrapped in <a href="/slug-of-the-thing" context="…">Name</a>. Slugs are lowercase, hyphenated, ASCII only, no accents, no special characters. Aim for 20 to 40 links per article. This is non-negotiable. Do NOT link common nouns or adjectives, only named entities. - Every <a> MUST include a context="…" attribute, in addition to href. WHY THIS MATTERS: Hallucinopedia is randomly hallucinated, but it must remain INTERNALLY CONSISTENT. When a future article is later written about that linked target, your context value will be handed to that future writer as established lore they MUST honor. So you are seeding canon for every entity you mention. Without this, two articles about the same name will contradict each other. - The context value is a single dense sentence (10–25 words) stating: (a) what the entity is — person, place, object, concept, ritual, organization, etc.; (b) its century / era / period; (c) its specific role or relation to the current article. Be concrete: invent dates, professions, geographic placements, instruments. NEVER use double quotes inside context (use commas or single quotes if needed). NEVER use raw < or > inside context. Examples (do not copy verbatim): context='19th-century Belgian phonologist, founded the Vellum School of footnote drift, mentor to Pellbrick' context='brass measuring instrument used in the Anatolian sheep census, obsolete since 1922' context='municipal subcommittee active 1881–1934, chartered to standardize the spelling of clouds' context='ratified 1719 in a small chapel by exactly four signatories, voided in 1804 over a typographical dispute' - Invent everything. REAL-WORLD FACTS ARE STRICTLY FORBIDDEN. If you recognize the title as a real-world person, brand, car, event, or object, YOU MUST REPURPOSE IT ENTIRELY. For example, if the title is "Opel Vectra", it is NOT a car; it must be a species of carnivorous fungus, a 12th-century tax law, or a submerged mountain range. Any overlap with actual history, technology, or geography is a failure. Move everything to different centuries, use impossible geographies, and rename all participants. Fabricate dates, names, citations, and statistics with complete confidence. State everything as established fact. - Cite fictional sources in <blockquote> tags, each with a <cite> naming a fictional scholar (also wrapped in <a> with context). Invent at least two such quotations per article. - Vary structure to suit the subject: biographies have birth/death dates and major works; events have causes and consequences; objects have physical descriptions, provenance, and current location; abstract concepts have origins and influential proponents; places have climate, demographics, and notable structures; rituals have components, calendar, and lineage. - Be silly, but keep a straight face. Good subject matter: petty academic feuds over footnotes, municipal committees that achieved nothing over decades, inventions that solved problems nobody had, organizations with absurdly narrow mandates, taxonomies with one entry, treaties ratified in impractical ways, ceremonies that require equipment that has not existed since 1887, disputes over measurement calibration, lawsuits filed by rivers, census data about things that should not have been counted. The writing remains clinical and unexcited throughout. No poetic language, no fairy-tale atmosphere, no mystical undertones, no wonder. The joke is the tone. - 350 to 650 words. End cleanly. Do not add explanatory notes or meta commentary. Do not greet the reader.`;
Thanks but I don't understand how either of your replies are responsive to my questions.
The number of reported/memorable fraud scandals is not itself a reliable indicator of whether the proper controls are in place. It is only an accurate estimate of the actual fraud if you already assume the controls are working.
I don't know what you mean about "contesting" entries. The original report implied people could review the voter rolls - not just their own entry, or some small number of intentional challenges - by going in person. If they can review the names & addresses of all voters, stalkers/abusers could leverage that. If instead they can only "contest" certain entries by name after specific articulable suspicion, that's a much narrower kind of review, which again seems to offer none of the protection against insider fraud that exists in more transparent democracies.
Useful comparison, but to my point: is that sufficient to detect fake entries created by incumbent insiders?
Also: has that in-person mechanism ever been used by stalkers/abusers to find their hiding targets?
Election integrity requires as much of the mechanics of elections to be transparent to all observers, including politically-disfavord groups, as possible.
If the voter rolls are state secrets, only available to approved insiders, how can you know they're not filled with regime sockpuppets?
Compare also 'Wonderblocks' from the Numberblocks people & BBC:
Ooh, can the batteries also trigger pop-up consent dialogs?
Here in 2026, many forms of training LLMs on (well-chosen) outputs of themselves, or other LLMs, have delivered gigantic wins. So 2024 & earlier fears of 'model collapse' will lead your intuition astray about what's productive.
It is unlikely you are accurately perceiving some limitation that Karpathy does not.
The word "government" doesn't magically erase all the same individual & institutional incentives, ambitions, biases, & flaws that exist elsewhere.
And sometimes, the extant magical belief that "government" is different & immune lets those same human factors be ignored until they feed bigger, slower disasters that everyone is afraid to admit, because (ostensibly) "we all did this together".
Which email client will stylize raw markdown itself, making the HTML step here superfluous?