HN user

wonnage

1,608 karma
Posts4
Comments805
View on HN

AI cannot tell you why it did something. If you ask it why it didn't do X even though your prompt said so, it can come up with some plausible sounding explanation, which might even be correct, but there's a fundamental impedance mismatch between giving you the most plausible sounding answer and actually grounding that in knowledge of how it generated that answer.

TBF having a 60k personal style guide is also a bad idea, people get borderline AI-psychosed about how their special prompts are really steering the AI when in reality there is no way to evaluate this stuff. You invariably end up trying to define your style in terms of what it is and it isn't, it's like trying to define "red". "follow the style of my 5 most recent PRs" works pretty well.

Not sure if this is still up to date (2023), but https://arxiv.org/abs/2307.03172 shows that performance degrades mostly in the middle of the context.

Anecdotally I've been stuck in that situation of being at 400-500k tokens and "just one more prompt bro" will get the task done, and I appreciate not having to wait through a compaction. If anything, keeping the bloated context helps with accuracy at the expense of speed in these cases.

Agree that the study design is flawed. Going with the possibly-hallucinated AI answer is rational as long as you know the hallucination rate isn't 100% (or whatever % you get after factoring in the monetary rewards they introduced later in the study)

But in reality when google gives you the wrong answer, you at least have some signals you can use to infer confidence. For example, the number of results, whether the sources are trustworthy, etc.

AI at best tucks that away in a footnote and discourages further critical thinking.

This is a useless post-hoc rationalization. "It worked out, so it doesn't matter". You're trying to galaxy brain yourself into ignoring the obvious conclusion.

The point is that if you were starting a new TUI LLM harness today, you would basically use CC's architectural decisions as a guide for what not to do.

The top marginal tax rate in the US was 90% for most of the post WWII 20th century and that didn’t seem to hurt anyone. Invented transistors, went to the moon, built interstate highway system, mass construction of nuclear power, and became the world leader in manufacturing

Today we have low tax rates and can’t make chips, still working on that moon landing, can’t build high speed rail, can’t build nuclear, and are trying to tariff our way back to a manufacturing sector

WhatsApp and Messenger have billions of users but don’t make much ad revenue. I’m sure there exists a way to show ads in ChatGPT that will be figured out eventually but so far nobody’s figured out how to monetize chat as a medium

Also it’s basically free for Google to show you a sponsored result but embedding one in a ChatGPT response actually costs money (assuming they’re part of the generated response).

Lastly I will bet you one Stargate datacenter that Meta has thought about LLM-based advertising, and if there’s any low hanging fruit there it’s already been tried

Dell has a few. I have the U2723QE. It reliably wakes up my Windows desktop but waking up my Macbook is a crapshoot. I'm pretty sure the issue is in Apple's software, not the docks.

All the more reason to focus on those service guarantees, integration, and lawyers while making the underlying model easily swappable to whoever’s winning the frontier model involution battle at the moment

This is much harder than it sounds. Most techniques I’ve seen end up using separate agents to do the planning, implementation, and judging.

The elaborate workarounds you have to build to help an agent which fundamentally doesn’t know what it’s doing reminds me of this old blog post about TDD: https://pindancing.blogspot.com/2009/09/sudoku-in-coders-at-...

IMO present technology is tailored for an experienced developer to give agents manageable tasks that can be one-shot. The marketing right now reminds me of the 90s when AskJeeves promised natural language search when the technology was fundamentally still stuck in keyword search, and learning to craft a search query for Google is today’s prompt engineering

Yeah, the neoclouds and hyperscalers are taking massive losses right now, self hosting is basically signing yourself up to do the same. There are philosophical reasons to do so but it’s a terrible economic decision

But those proofs are showing that the fundamental axioms (which are generally simple and elegant) are still enough to build a complex result.

I think of elegance as not having to add epicycles, not that everything in the system has to be simple.

Also, without a working theory the, the space of possible solutions is near infinite. LLMs manage to pluck out the space of comprehensible English strings from n-dimensional hell. Even if this is done with a black box of billions of parameters, it’s still elegance in the sense that such a space even exists and was found

Yeah, they could just inflate Bancor the same way a country can with their own currency today (i.e, printing dollars) to manage their debts

Tbf it was proposed in a time where globalization was a good thing and there was naive optimism about international organizations!