HN user

aakresearch

133 karma
Posts2
Comments141
View on HN

What gives you basis to believe and accept that all of the above is exhaustive, complete and devoid of fakes? That "finding" finds ALL what exists and NONE of what does not exist? Ditto for categorising, summarising and labeling?

There is no fundamental guarantees provided by LLM architecture, so to assert anything of sorts would require manually executing the operation and comparing the results. And why not asserting the correctness is not negligence or fraud?

I am not saying it is impossible, but so far I have not found any explanations that would close this huge gap in logic for me. LLMs themselves "admit" the gap cannot be explained rationaly.

Ooooh, don't get me started how mad it makes me! I am paranoid (or just lucky) enough that I didn't yet had it happen to me, but my wife's phone had done it four times in the last year. Each time I check and double check that all "backups" are turned off, and each time it somehow pops back.

So, Google "backs up" a 128Gb worth of photos on the phone onto 15Gb free storage combined with Gmail and who knows what, completely clogs it (as if it couldn't be predicted) and then has audacity to suggest paying for "extra storage". There is no way in online UI to just delete the whole "backup". And the cherry on top: when you finally get to delete some there is a fine-print - "the selected photos will be deleted from all synced devices". Well, I guess I must be thankful that they at least show this warning. This is what passes as "backup" in Google's parlance these days.

I think we are in full agreement. Likewise, I am not in charge of any practical matter, and glad I don't have to answer to any moral or physical dilemma.

The link is fun :). My kill count was 75.

I was hinting at full impossibility to quantify those problems. Starting with un-define-ability (sorry, English is not my native tongue) of the very term "need" upon which the predicate "the needs of the many outweigh the needs of the few". Maybe that is why I am leaning towards "deontological" resolution of trolley problems, rather than "utilitarian".

As for LLMs, I use them and benefit from them. For fun, and hopefully for profit too. But if I had a lever I could pull to make this tech disappear - I would not hesitate one second, deontology be damned. Even better if smartphones suddenly follow too.

Wait, what? I thought everyone agrees that modern models post September 2025 (or whenever Opus or whatever 5.6789 was released) do not hallucinate, make things up, contradict themselves and can review their own output into perfection regardless of task, goal or context???? /s

If the life of everyone (all 8 billion people) is improved by tech by some margin (pick the margin - 10%, 1%, 0.1%) at the cost of x people (pick the x - a thousand, a million, a billion) being killed by the same tech - is it still positive or negative? How do you even reason about it?

No, government must not do any such thing. They must back away, as my literacy, curiosity and intelligence is not their business in the slightest. Nobody owes anything to government, especially trust. But if "government" insists on importance of things they communicate, it is absolutely on them to ensure it is calibrated to everyone's profile. And I still am not obliged to trust them. Once a liar always a liar.

That would end exactly at "Harrison Bergeron" world as described by Kurt Vonnegut, would it not? If every perceived advantage would require you to wear a "handicap".

We should shuffle the pile and throw half of them in the trash. I don’t want to hire unlucky employees

Modern filtering using "AI" and ATS of all kinds is such a loaded dice that I believe an honest randomizing like you described would be a significant improvement for both ends of the funnel. Beyond jokes.

like asbestos or lead

Perfect framing! I was leaning towards "AI Chernobyl" analogy in my predictions, but I think "asbestos" or "fossil fuels" captures its nature much better. May I borrow it? I definitely see the harmful consequences of "AI" exceeding both, however with sad realization that it usefulness was absolutely below par with any.

Impatiently waiting for the follow-up "Why Earn a Billion Dollars". I believe it is a few keystrokes from being finished. Any minute now...

Ghosting does much more harm to candidates than the burden to not-ghost does to companies. I'm not going to cry over all "talent acquisition experts" having to do some actual work. Uncertainty that businesses face does not compare to uncertainty that a candidate face after 18 months of unemployment. Ask me how I know.

In trading (of securities) posting an order without intent to execute is considered market manipulation, which is illegal and harshly prosecuted. There is a consideration that change of mind is possible, but you'll have a hell of a lot to prove in such case before authorities let you off.

I agree with many, pointing that companies will (try) find ways to fleece any regulation imposed. And I am not a fan of regulations myself, at all. But I think it is fair to hold businesses to some standard in many aspects, including hiring. It is already being done in regards to some, like discrimination and equality. Un- and under-employment is a matter, dealt with by society through institutions and funded by taxpayers. The "clearance rate" of job applications (from both "buy" and "sell" sides) is, therefore, a state concern. I do not think extending requirements of "business license" to demonstrate "genuine intent" would place insurmountable burden on HRs or CEOs. But of course, such extension must have some teeth.

To be clear, the current situation with excessive ghosting is not helped by decades-old push to "commoditize" jobs, particularly IT jobs. And the regulations we discuss will be a not very well-veiled recognition of its de-facto success. Which I am also not a fan of. But flip side seems worse, when companies are allowed to pretend they'd only settle for unicorn while not demonstrating a "unicorn-shaped sieve" at all.

How do you know that someone is about to lose their computing chops? They waste a paragraph explaining a convoluted roundabout way to answer question "how many months of 93% growth it takes for something to grow 500x" with useless precision. A kid just started with BASIC learns by heart in a week that 512=2^9 :)

And a follow-up to my comment above, as wanted to reply to "create an issue before filing a PR" remark too.

I find it actually extremely useful practice. We engineers tend to center our thoughts around "code" and with such code-centric mindset to accumulate knowledge as "code-adjacent" - in repo's commits, PRs, markdown files of all sorts. But in reality most if not all projects extends past the code, and it makes much more sense to have a "project-centered" mindset. As such, an external "issue" captures much more context and provides more useful insight about project impact than PR description alone. Love or hate JIRA, beyond microscopic solo-projects it makes full sense to use broader-scope external tools for project management.

As a nice cosmetic effect it also removes (some) bickering about whether to put "type" or "scope" first in the commit message :). Simply provide a reference to the ticket! (No, really, I hate JIRA, honestly!)

In my opinion review will always be a bottleneck, in OS as well as in commercial development.

To my understanding Code Review is first and foremost a trust-building exercise, seeking to establish common understanding behind the piece of code which is to be delivered. That it also may lead to improvements or "catch some bugs" is a distant tertiary side-effect, not the primary goal. At least this is the vantage point I reviewed any code from in the last two decades, and found that team morale and overall quality of teamwork - and delivered software, and customer satisfaction, as a result - responds very well to such interpretation. Regardless of the side you are on and competency level of your vis-a-vis. I would put "elevating competency level of both partners" as a secondary goal of Code Review, and very closely connected to the primary.

With current crop of LLMs there simply nothing on their side to which words "trust" and "understanding" can be meaningfully applied. Hence the "review" takes drastically different shape and implies very different goals. As both the primary and secondary goals described above cannot apply too. The only remaining goal of "improving and bug-catching", in absence of trust, understanding and learning, now requires much more work, which is also much more exhausting.

I grant we are - and also to facilitate understanding of one's thinking by other team members. But what I write "between commits" is not any more an evidence of "thinking" or conducive to "understanding" than the NSFW mouse gestures I make or rubbish I am humming to myself.

Indeed, I hear you 100%. It may be painful to watch "your" military engage in wars that are not righteous, from one's perspective. Are there even "righteous" conflicts these days? But short of demand to abolish all state military, what is appropriate way to express one's indignation? Hopefully all can agree that such demand would be insane; but why, then, those "pacifist" performances, which effectively are calls to deprive military of the best weapons, people, thoughts, strategies, are not considered insane?

Agree or disagree with particular foreign policy or military action, why do people forget that the bulk of military is staffed with their fellow citizens? Many of whom aren't terribly privileged to enjoy ample alternative choices to elevate themselves socially or financially. It is exactly this lot who benefits the most from DEI policies, cherished by "pacifists", is it not? It is them who are the first and most massive direct casualties, caused by not having access to the best, superior materiel, doctrine and training on and beyond the battlefield.

I'll be the first to point that military and paramilitary forces attract many with unchecked lust for violence. That "pride", "honor" and "patriotism" are often terribly misused, to uphold goals of those with impure, malicious ambitions. Who, I grant it, also disproportionately represented in the command echelons of military and beyond. But if we are honest, that scum won't be shaken or taught a lesson by SotA technology being withheld from their use or corporation refusing cooperation. It is their subordinates, who, maybe naively, subscribe to "ideal", unquoted interpretation of Pride, Honor and Patriotism, will bear the brunt of being crippled (by the consequences of the withholding and refusal) on the battlefield, and pay with their lives. Don't their lives matter?

- A little script that is best described as "symbolic submodule" workflow enabler for git: https://github.com/fontmaniac/symgit. As in "symbolic link" vs "hard link" - where hash tying host repo and pseudo-submodule is not git-generated, which makes "submodule" completely substitutable. For the description of creation process see the relevant section of Readme.md

- Companion script named Dr.Emmett - to retrospectively time-walk git commit history and rebuild the repo in a new place. It is not opened yet, and may not be ever as it doesn't satisfy LLM/human ratio worthy of publishing.

- Both scripts were made in service to my FNA-"game" project VibeSopwith: https://github.com/fontmaniac/VibeSopwith.Game. A more or less comprehensive description, and treatise on my stance about vibe- and agentic-coding can be found in the project's Readme.md

- VibeSopwith catalysed creation (accretion?) of Nage.Strata: https://github.com/fontmaniac/Nage.Strata - "Not-a-game-engine" collection of useful "primitives" for FNA-based game development.

- VibeSopwith Readme.md has notable mention of "Aether Bodybuilder" - a helper tool for visually coupling Aether.Physics2D "bodies" to specific sprite textures. Fully vibe-coded in a couple of hours, and as such won't ever be publicly opened, having LLM/human ration approaching infinity.

I fear that workflow is just chat, search, and being a rubber duck for my thoughts

This is exactly what I settled upon after my own trying really hard. It is liberating, I have no fear at all!

Oops... I am deeply sorry, thank you for the heads up! It seems I've myself committed a cardinal sin that I am usually quick to point in others - rushing to reply without comprehending the full message. (Meta-oops: I realized how LLM-ish it sounds. Quick, reboot before my cover is blown!)

I happen to believe that the flaw being discussed IS fundamental and inherent in the design and architecture of LLM - this is why I always put "AI" in scare quotes. I've spoken about it in some of my other comments, namely this https://news.ycombinator.com/item?id=47162553 and to some extent this https://news.ycombinator.com/item?id=48046333. And as you do, I, too, hope that I am wrong about the hype and its eventual clash with reality, but do not hold my breath.

>At the same time, do you really want every conversation you have with your doctor recorded

Yes. This is what medical records are.

No. Medical records are limited extracts from conversations, which is your doctor and only your doctor is qualified to make, using "semantic analysis applied to your unique situation", not "linguistic probabilistic inference applied to conversation about your situation using token weights averaged over billion unrelated samples"

It's not like the doctor is talking to you about which anime series are the best. You're talking about your health, your body, your disease, your treatment.

No jokes, no banter, no chit-chat, no complements to doctor's new Tesla?

It's important to keep track of that.

Same fallacy Meta fell into when started tracking employees' keystrokes and mouse gestures. 90% of my mouse movements are just fidgeting, with no relation to the task at hand - and it is not a crime! But if I knew my mouse fidgeting is being watched, I'll make sure that percentage goes up to 99% - for the LLM which is gonna be trained off it to self-immolate over its NSFW nature.

Knitting bullshit 3 months ago

I may have different perception facilities in the brain, but to me the images in the article are horrifying. Such bad, Frankenstein-level, criminal butchering of absolutely everything!

I am by no means an expert, but I'd like to offer my mental model - up to you to decide if it is solid or not, but it works for me.

I think the core intuition is that, like with any other "rasterized" system with finite memory that cannot encode an absence of anything - relation, concept, entity, LLM cannot encode an absence of something through its internal weights. Say, you can have "Product" or "Order" tables in you database, but you cannot have "NotAProduct" or "NotAnOrder" tables - for obvious reasons of such relations being infinite and uncountable. So, to establish an absence of Product or Order your application must execute a "search" operation through the relevant tables. But in LLM-space "search" operation does not exist. It is mathematically undefined. LLM arrives at output (or "what to do") through a sequence of projections of input token vector through its "latent space". It "moves toward" high-probability clusters, fundamentally unable to "move away". So, the success of any "negation" in the prompt ("don't touch this file", "draw me a ballot box without a flag on it") depends on how heavily such scenario represented in the training data/model space. And again, the absence-of-something may be hard-to-impossible to usefully encode, especially if "something" is not fixed. Therefore, to expect "don't touch this file" sentence to result in, well, not touching the file is pure gambling. Sometimes it may look like working, albeit for wrong reasons, and some other times LLM may do exactly the opposite - because its weight matrix statistically pushes it towards "touch this file", completely ignoring (nonexistent in its latent space) "don't".

There is no way to reliably know what will work, and no "skill" or "art" in this. Well, no more than in dice rolling or horoscope casting.

I'd like to add that for the above reason I find "agentic development" usefulness on par with avian remains reading. But when I explored it two practical advises seemed to be helpful in nudging LLM around negation problem:

- Omit the "don't" prompt completely, thus not creating a false "attractor" for LLM; and

- Provide an alternate positive directive ("what to DO", not "what to NOT DO") to act as "escape hatch" when LLM might "want" to touch the sacred file or drop the production DB.

While it looked like somewhat working, I think it is trivially obvious that trying to predict all the nonsense LLM might want to perform and coming up with possible "escape hatches" for everything very quickly becomes utterly impractical.

It seems that author unironically advises to write your commit messages like this: "Restructured Claude’s module architecture, rejected initial state management approach, rewrote error handling from scratch", to have a chance at defense in potential court hearing. I find it funny, if vindicating for my personal approach. If the expectation is to "restructure, reject, rewrite" what "AI" spits out, why use "AI" at all at this point???

To my understanding, LLM, by design, is unable to encode negation semantics. Neither negation "operation", nor any other "subtractive" operations are computable in LLM machinery. Thinking out loud, in your example the "Read code" and "Form hypothesis" seem to be useful instructions for what you want, while "Do not write any code" and "Not to fix the bug" might actually be misleading for the model. Intuitively (in human terms) one would imagine that, when given such "instruction", LLM would be repelled from latent-space region associated with "write any code" or "fix the bug". But in reality LLM cannot be "repelled", it is just attracted to the region associated with full, negated "DO NOT <xxxx>". And this region probably either has a significant overlap with the former ("DO <xxx>") or even includes it wholesale. This may explain why it sometimes seems to "work" as intended, albeit accidentally. My 2c.