HN user

photonthug

1,768 karma
Posts8
Comments709
View on HN

Adding complexity is just one aspect. Everywhere there is someone whose job is to ensure the bottom line never changes and status quo for the powerful is preserved. Insurance, taxes, rents.. in the absence of effective regulation, the average number of successful appeals will simply get factored in and average costs go up so that profit stays the same and grows at the same rate as before. Similar to how chains factor in losses due to spoilage or theft.. of course they don't actually take a profit loss, they just price it in.

I really don't get people who see this kind of thing as empowering because in the end your (now strictly necessary) appeal with lawyers or AI to get a more fair deal just becomes a new tax on your time/money; you are worse off than before. A good capitalist will notice these dynamics, and invest in AI once it's as required for life as healthcare is, and then work on driving up the costs of AI. Big win for someone but not the downtrodden.

Yes, unfortunately a phrase that's used in an attempt to lend gravitas and/or intimidate people. It sort of vaguely indicates "a complex process you wouldn't be interested in and couldn't possibly understand". At the same time it attempts to disarm any accusation of bias in advance by hinting at purely mechanistic procedures.

Could be the other way around, but I think marketing-speak is taking cues here from legal-ese and especially the US supreme court, where it's frequently used by the justices. They love to talk about "ethical calculus" and the "calculus of stare decisis" as if they were following any rigorous process or believed in precedent if it's not convenient. New translation from original Latin: "we do what we want and do not intend to explain". Calculus, huh? Show your work and point to a real procedure or STFU

As something of a technophile myself.. I see a lot more value in arguments that highlight totally ridiculous core assumptions rather than focusing on some kind of "humans first and only!" perspectives. Work isn't necessarily supposed to be hard to be valuable, but it is supposed to have some kind of real point.

In the dating scenario what's really absurd and disgusting isn't actually the artificiality of toys.. it's the ritualistic aspect of the unnecessary preamble, because you could skip straight to tea and talk if that is the point. We write messages from bullet points, ask AI to pad them out uselessly with "professional" sounding fluff, and then on the other side someone is summarizing them back to bullet points? That's insane even if it was lossless, just normalize and promote simple communications. Similarly if an AI review was any value-add for AI PR's, it can be bolted on to the code-gen phase. If editors/reviewers have value in book publishing, they should read the books and opine and do the gate-keeping we supposedly need them for instead of telling authors to bring their own audience, etc etc. I think maybe the focus on rituals, optics, and posturing is a big part of what really makes individual people or whole professions obsolete

fully AI generated and fully AI review

This reminds me of an awesome bit by Žižek where he describes an ultra-modern approach to dating. She brings the vibrator, he brings the synthetic sleeve, and after all the buzzing begins and the simulacra are getting on well, the humans sigh in relief. Now that this is out of the way they can just have a tea and a chat.

It's clearly ridiculous, yet at the point where papers or PRs are written by robots, reviewed by robots, for eventual usage/consumption/summary by yet more robots, it becomes very relevant. At some point one must ask, what is it all for, and should we maybe just skip some of these steps or revisit some assumptions about what we're trying to accomplish

You are how you act 9 months ago

In philosophy 101 the usual foil for Rousseau vs.. would be Hobbes, but that framing with a realist/pessimist would not be popular with the intended audience, where the goal is to lionize the nationalist, the inventors/owners, the 1%.

Despite his own moral lapses, Franklin saw himself as uniquely qualified to instruct Americans in morality. He tried to influence American moral life through the construction of a printing network based on a chain of partnerships from the Carolinas to New England. He thereby invented the first newspaper chain. https://en.wikipedia.org/wiki/Benjamin_Franklin#Newspaperman

To be clear Franklin's obviously a complicated historical figure, a pretty awesome guy overall, and I do like American pragmatism generally. But it matters a lot which part of the guy you'd like to hold up for admiration, and elevating a preachy hypocrite that was an early innovator in monopolies and methods of controlling the masses does seem pretty tactical and self-serving here.

A definition of AGI 9 months ago

Funny but the eyebrow-raising phrase 'recursive self-improvement' is mentioned in TFA in an example about "style adherence" that's completely unrelated to the concept. Pretty clearly a scam where authors are trying to hack searches.

Prerequisite for recursive self-improvement and far short of ASI, any conception of AGI really really needs to be expanded to include some kind of self-model. This is conspicuously missing from TFA. Related basic questions are: What's in the training set? What's the confidence on any given answer? How much of the network is actually required for answering any given question?

Partly this stuff is just hard and mechanistic interpretability as a field is still trying to get traction in many ways, but also, the whole thing is kind of fundamentally not aligned with corporate / commercial interests. Still, anything that you might want to call intelligent has a working self-model with some access to information about internal status. Things that are mentioned in TFA (like working memory) might be involved and necessary, but don't really seem sufficient

A definition of AGI 9 months ago

Yeah, I mean I hope there are not many people that still think it's a super meaningful test in the sense originally proposed. And yet it is testing something. Even supposing it were completely solved and further supposing the solution is theoretically worthless and only powers next-gen slop-creation, then people would move on to looking for a minimal solution, and perhaps that would start getting interesting. People just like moving towards concrete goals.

In the end though, it's probably about as good as any single kind of test could be, hence TFA looking to combine hundreds across several dozen categories. Language was a decent idea if you're looking for that exemplar of the "AGI-Complete" class for computational complexity, vision was at one point another guess. More than anything else I think we've figured out in recent years that it's going to be hard to find a problem-criteria that's clean and simple, much less a solution that is

A definition of AGI 9 months ago

Outcome would depend on the rest of the test, but I'd say the "human" version of this answer adds zero or negative value to chances of being human, on grounds of strict compliance, sycophancy, and/or omniscience. "No such thing" would probably be a very popular answer. Elaboration would probably take the form of "love it" or "hate it", instead of reaching for a comprehensive answer describing the inside and the outside.

Experimental design comes in here and the one TT paper mentioned in this thread has instructions for people like "persuade the interrogator [you] are human". Answering that a green eggplant is green feels like humans trying to answer questions correctly and quickly, being wary of a trap. We don't know participants background knowledge but anyone that's used ChatGPT would know that ignoring the question and maybe telling an eggplant-related anecdote was a better strategy

A definition of AGI 9 months ago

I think the TT has to be understood as explicitly adversarial, and increasingly related to security topics, like interactive proof and side channels. (Looking for guard-rails is just one kind of information leakage, but there's lots of information available in timing too.)

If you understand TT to be about tricking the unwary, in what's supposed to be a trusting and non-adversarial context, and without any open-ended interaction, then it's correct to point out homework-cheating as an example. But in that case TT was solved shortly after the invention of spam. No LLMs needed, just markov models are fine.

A definition of AGI 9 months ago

Hah, tools-or-no does make things interesting, since this opens up the robot tactic of "use this discord API to poll some humans about appropriate response". And yet if you're suspiciously good at cube roots, then you might out yourself as robot right away. Doing any math at all in fact is probably suspect. Outside of a classroom humans tend to answer questions like "multiply 34 x 91" with "go fuck yourself", and personally I usually start closing browser tabs when asked to identify motorcycles

A definition of AGI 9 months ago

Turing test isn't actually a good test of much, but even so, we're not there yet. Anyone that thinks we've passed it already should experiment a bit a with counter-factuals.

Ask your favorite SOTA model to assume something absurd and then draw the next logical conclusions based on that. "Green is yellow and yellow is green. What color is a banana?" They may get the first question(s) right, but will trip up within a few exchanges. Might be a new question, but often they are very happy to just completely contradict their own previous answers.

You could argue that this is hitting alignment and guard-rails against misinformation.. but whatever the cause, it's a clear sign it's a machine and look, no em-dashes. Ironically it's also a failure of the turing test that arises from a failure in reasoning at a really basic level, which I would not have expected. Makes you wonder about the secret sauce for winning IMO competitions. Anyway, unlike other linguistic puzzles that attempt to baffle with ambiguous reference or similar, simple counterfactuals with something like colors are particular interesting because they would NOT trip up most ESL students or 3-5 year olds.

Agree with all this; see my other comment in thread for more color, more math. I don't want to come across as embracing pseudo-science, misinformation, etc.

I do like to think about the distinctions and boundaries for hard/soft/squishy knowledge though, and try to challenge assumptions and misconceptions about it. Invariably people have weird ideas about how "hard" their pet area is and how "soft" that thing they love to hate really is, which is itself a kind of dogma or superstition. Plus I think it's a public service to try and interest gear-headed nerds in things like criticism and philosophy (or vice versa, pushing engineering and math at the literature nerds). Last time I waded into this kind of debate I was pointing out that Frege worked on semiotics.

Not OP, but I think the plot twist is, maybe we need to be able to entertain "obviously absurd" ideas to be able to land on a correct position if the culture we're inside of is not ready for those ideas yet. (No idea if the journal was really that early on this particular position though)

Crucially, entertaining ideas isn't the same as believing them, it's about giving them some time and space so you can work out whether it's consistent, rich, useful. Even in math this stuff is hard to get right, just look at the resistance and ridicule that Cantor had to go through, or look at the development of non-Euclidean geometry. And that's a space where proof is actually possible. Critical theory is a real thing but is always walking this fine line between being nonsense or being revolutionary.

There's some interesting stuff in here if you can tolerate the meandering and the way-back-when. Like you'd expect from po-mo wonks, everything's gotta be infinitely subtle and infinitely contextualized. So no big mea-culpa and no big defensive denial either. All of that's been hashed and rehashed many times already I guess. You'll find some self-deprecating humor, some spots with surprising self-awareness, some with a surprising lack of it. The main fresh thing is how they'd like to try and compare/contrast/contextualize it in this moment. For example:

Being a gatekeeper by maintaining high intellectual standards is not what public opinion would associate with Social Text, to say the least. Yet that is what the journal practiced, mainly. And it is a practice worth defending, however elitist it might look. All the more so because of how the Trump administration has weaponized both the idea of the hoax and the program of anti-elitism. [..] We know what has befallen intellectual standards. [..] Is this ChatGPT, or is it Orwell’s doublethink?

Well ok, there's a conversation to be had about these things! This is not the time to pontificate though, it's the time for sweet revenge. There's never been a better time for po-mo wonks to lean on AI slop and blast physics journals with fake stuff about gravity until someone understaffed falls for the trick. Then you can do a big scandalous reveal about how you can't believe you got away it ;)

Most people do this so that they can eat the rocks afterward. They are shiny and very nutritious, and it strengthens the teeth. It's normal for some teeth to break off during this phase, but a) you already have colorful rocks to replace the teeth with, and b) old broken teeth can now be placed inside the tumbler for smoothing. 9/10 geologists agree that unsmoothed teeth that aren't made of rock are the number one cause of oral hygiene problems

The capital was better fortified and the French wanted food and loot. So softer target sounds nice, especially if you think this crushes morale immediately and don't believe the opposition will go scorched earth (but they did).

If I would take S.P., I would hold Russia by the head. If I take Kiev, I will hold Russia by legs. If I take Moscow, I will reach right into its heart!"

https://history.stackexchange.com/questions/27588/why-did-na...

And temperature 0 makes outputs deterministic, not magically correct.

For reasons I don't claim to really understand, I don't think it even makes them deterministic. Floating point something something? I'm not sure temperature even has a static technical definition or implementation everywhere at this point. I've been ignoring temperature and using nucleus sampling anywhere that's exposed and it seems to work better.

Random but typical example.. pydantic-ai has a caveat that doesn't reference any particular model: "Note that even with temperature of 0.0, the results will not be fully deterministic". And of course this is just the very bottom layer of model-config and in a system of diverse agents using different frameworks and models, it's even worse.

It DOES fail more when the numbers are longer (because it results with more text in the context),

I tried to raise this question yesterday. https://news.ycombinator.com/item?id=45683113#45687769

Declaring victory on "reasoning" based on cherry-picking a correct result about arithmetic is, of course, very narrow and absurdly optimistic. Even if it correctly works for all NxM calculations. Moving on from arithmetic to any kind of problem that fundamentally reduces to model-checking behind the scenes.. we would be talking about exploring a state-space with potentially many thousands of state-transitions for simple stuff. If each one even has a small chance of crapping out due to hallucination, the chance of encountering errors at the macro-scale is going to be practically guaranteed.

Everyone will say, "but you want tool-use or code-gen for this anyway". Sure! But carry-digits or similar is just one version of "correct matters" and putting some non-local kinds of demands on attention, plus it's easier to check than code. So tool-use or code-gen is just pushing the same problem somewhere else to hide it.. there's still a lot of steps involved, and each one really has to be correct if the macro-layer is going to be correct and the whole thing is going to be hands-off / actually automated. Maybe that's why local-models can still barely handle nontrivial tool-calling.

Thanks for doing this. OpenAI is not in fact open, so referencing their claims as obviously true on anything else is just a non-starter. Counterpoint though, it's been a while since I've run this kind of experiment locally, so I started one too. For reasoning I only have qwen3:latest and I won't clutter the thread with the output, but it's complete junk.

To summarize, with large numbers it goes nuts trying to find a trick or shortcut. After I cut off dead-ends in several trials, it always eventually considers long form addition, then ultimately rejects it as "tedious" and starts looking for "patterns". Wait, let me use the standard multiplication algorithm step by step, oh that's a lot of steps, break it down into parts. Let me think. Over ~45 minutes of thinking (I'm on CPU), but it basically cannot follow one strategy long enough to complete the work even if landed on a sensible approach.

For multiplying two-digit numbers, it does better. Starts using the "manual way", messes up certain steps, then gets the right answer for sub-problems anyway because obviously those are memoized somewhere. But at least once, it got the correct answer with the correct approach.

I think this raises the question, if you were to double the size of your input numbers and let the more powerful local model answer, could it still perform the process? Does that stop working for any reason at some point before the context window overflows?

Agree, this stuff was trending up very fast before AI.

Could be my own changing perspective, but what I think is interesting is how the signal it sends keeps changing. At first, emoji-heavy was actually kind of positive: maybe the project doesn't need a webpage, but you took some time and interest in your README.md. Then it was negative: having emoji's became a strong indicator that the whole README was going to be very low information density, more emotive than referential[1] (which is fine for bloggery but not for technical writing).

Now there's no signal, but you also can't say it's exactly neutral. Emojis in docs will alienate some readers, maybe due to association with commercial stuff and marketing where it's pretty normalized. But skipping emojis alienates other readers, who might be smart and serious, but nevertheless are the type that would prefer WATCHME.youtube instead of README.md. There's probably something about all this that's related to "costly signaling"[2].

[1] https://en.wikipedia.org/wiki/Jakobson%27s_functions_of_lang... [2] https://en.wikipedia.org/wiki/Costly_signaling_theory_in_evo...

“Settling Defendants have agreed not to provide nonpublic data to RealPage for use in competitor pricing recommendations and to refrain from using RealPage’s RMS that relies on non-public competitor data to make pricing recommendations,” attorneys wrote in the settlement filing.

https://www.multifamilydive.com/news/realpage-class-action-l...

I agree that "nonpublic" is barely related to the problem so how it's related to a solution is unclear. But it seems like this is the only general aspect of the outcome. Otherwise the outcome is just to stop doing this specific bad thing this specific time, and fines that are less than the profit made from bad behaviour.

Seems like an optimistic read on things. This is the kind of common-sense approach you would expect in a world without lawyers, just observing that collusion is bad because the effects are bad, and digging into the details of the causes are completely irrelevant for the public/plaintiff because it's really just on the company to fix the undesirable result.

IANAL but if realpages outcomes were definitive or reasonably generalized results dealing with the core issue, then similar arguments against e.g. Amazon would be a slam dunk. AFAIK, actual case outcome just hinges on details about "nonpublic data" and similar. Not remotely on bad effects for consumers or anything like that. Since printing realpages database in the newspaper would not actually help apartment-hunters, then this just tells landlords and third party markets how to do price-fixing legally next time? Most likely algorithmic pricing, surveillance pricing, etc is still coming to your grocery store after the issue is "settled" for property rental, or at least settled for realpages, in certain jurisdictions, for now.

Modern price collusion is more apt to happen with A/B testing if prices at locations to see what the local market will bear.

One of my first thoughts as well. If you're big enough, you collect so much data and run so many experiments all the time that you know exactly what you'd do if/when there's any competitor on the scene. Not only is there no need to talk to them and make backroom deals, but barely any need to even observe them. You priced like they did/would/could at some point already anyway. At a certain scale and if you already know the price that the market can tolerate.. the most relevant hidden information you want to know is how much cash your competitor has access to. That tells you whether you can win the price-war to sell at a loss for long enough to ruin them, buy them, move on to integrating verticals etc.

Game theory is interesting but also a bad model to the extent that it assumes persistent players with changing strategies, whereas average case in late-stage capitalism is more likely to have players eating players, no new players can enter, players changing rules, etc. As a CS nerd I still like a game theoretical approach better than most econ, but at some point we need to give up on tidy formulas and closed-form answers, and go all in on messy simulations.

feels like the wrong question to me

I agree but had different questions. TFA mentions the consideration of whether failure cases are correlated, but of course if OpenAI wins big, there's a good chance this directly or indirectly creates much instability and uncertainty in many other loans/partners. What's the EV on whether that is net-positive considering this is a loan at 5% and not an investment?

On the other side, if OpenAI crashes hard, is it really such a sure thing that Microsoft will be the on the hook to pay off their debts? Setting aside whatever the lawyers could argue about in a post-mortem, are they even obligated to keep their current stake / can they not just divest / sell / otherwise cut their losses if the writing is on the wall?

Replacement.ai 9 months ago

What's to be gained from talking like this in public as a corporate figure?

Diving into the game theory of a 4-player setup with executives/investors/customers/workers is tempting here but I'll take a different approach.

People who actually face consequences have trouble understanding how the "it might help, it can't hurt!" corporate strategy can justify almost any kind of madness. Especially when the leaders are morons that somehow have zero ideas, yet almost infinite power. That's how/why Volkswagen was running slave plantations in Brazil as late as 1986, and yet it takes 40 years to even try to slap them on the wrist.[1] A manufacturing company that decided to run FARMS in the amazon?, with slaves??, for a decade??? One could easily ask, what is to be gained by doing crimes against humanity for a sketchy, illegal, and unethical business plan that's not even related to their core competency? Power has it's own logic, but it doesn't look like normal rationality because it has a different kind of relationship with cause-and-effect.

Overall it's just a really good time to re-evaluate whether corporations and leaders deserve our charitable assumptions about their intentions and ethics.

[1] https://reporterbrasil.org.br/2025/05/why-is-volkswagen-accu...

If we fill those abandoned buildings with people, air-conditioning the inside of the building for them will obviously add even more heat to the outside? Parking lots that are full of cars aren't going to be that much cooler than empty ones?

Basically the real story is just that trees make shade (yes, we know already) and "vacant or abandoned" isn't much involved (yes, but we want to discuss zoning/taxes/urbanism things)

Replacement.ai 9 months ago

It's actually kinda noteworthy that corporations don't talk like this (yet). Masks are off lately in political discourse, where we're all in on crass flexing on the powerless, the othering, cruelty, humiliation. How long before CEOs are openly talking about workers in the same ways that certain politicians talk about ${out_group}? If you're b2b with nothing consumer-facing to boycott, may as well say what you really think in a climate where it can't hurt and might help. The worst are filled with passionate intensity, something something rough beast etc.

Obviously platforms get advertiser dollars, but the question is what the business paying for that gets. The answer is.. almost nothing? Dedicated marketing/advertising resources at your business is probably just a fifth column, where shareholders and business owners are swindled into paying the salary and other maintenance for people who are actually working for google/facebook/amazon/whatever.

Source? Like 20+ years of ubiquitous surveillance, tracking and micro-targeting.. and yet non-pet owners still get ads for dog food, males get ads for feminine hygiene products, single people are offered deals on family vacations, and people who just bought a car get car advertisements for the next 3-5 years which only taper off when it might actually be time to buy another car.