Reddit post so take with a grain of salt but this does seem possible. But unclear whether this is actually a viable architecture https://www.reddit.com/r/LocalLLaMA/comments/1t8s83r/nvidia_...
HN user
wonnage
I mean the questions are on random movie trivia, the alternative is just guessing. I think you're overthinking it.
AI cannot tell you why it did something. If you ask it why it didn't do X even though your prompt said so, it can come up with some plausible sounding explanation, which might even be correct, but there's a fundamental impedance mismatch between giving you the most plausible sounding answer and actually grounding that in knowledge of how it generated that answer.
TBF having a 60k personal style guide is also a bad idea, people get borderline AI-psychosed about how their special prompts are really steering the AI when in reality there is no way to evaluate this stuff. You invariably end up trying to define your style in terms of what it is and it isn't, it's like trying to define "red". "follow the style of my 5 most recent PRs" works pretty well.
Not sure if this is still up to date (2023), but https://arxiv.org/abs/2307.03172 shows that performance degrades mostly in the middle of the context.
Anecdotally I've been stuck in that situation of being at 400-500k tokens and "just one more prompt bro" will get the task done, and I appreciate not having to wait through a compaction. If anything, keeping the bloated context helps with accuracy at the expense of speed in these cases.
Agree that the study design is flawed. Going with the possibly-hallucinated AI answer is rational as long as you know the hallucination rate isn't 100% (or whatever % you get after factoring in the monetary rewards they introduced later in the study)
But in reality when google gives you the wrong answer, you at least have some signals you can use to infer confidence. For example, the number of results, whether the sources are trustworthy, etc.
AI at best tucks that away in a footnote and discourages further critical thinking.
They provide a sample of hallucinated answers from ChatGPT at the end of the study.
This is a useless post-hoc rationalization. "It worked out, so it doesn't matter". You're trying to galaxy brain yourself into ignoring the obvious conclusion.
The point is that if you were starting a new TUI LLM harness today, you would basically use CC's architectural decisions as a guide for what not to do.
The top marginal tax rate in the US was 90% for most of the post WWII 20th century and that didn’t seem to hurt anyone. Invented transistors, went to the moon, built interstate highway system, mass construction of nuclear power, and became the world leader in manufacturing
Today we have low tax rates and can’t make chips, still working on that moon landing, can’t build high speed rail, can’t build nuclear, and are trying to tariff our way back to a manufacturing sector
WhatsApp and Messenger have billions of users but don’t make much ad revenue. I’m sure there exists a way to show ads in ChatGPT that will be figured out eventually but so far nobody’s figured out how to monetize chat as a medium
Also it’s basically free for Google to show you a sponsored result but embedding one in a ChatGPT response actually costs money (assuming they’re part of the generated response).
Lastly I will bet you one Stargate datacenter that Meta has thought about LLM-based advertising, and if there’s any low hanging fruit there it’s already been tried
Dell has a few. I have the U2723QE. It reliably wakes up my Windows desktop but waking up my Macbook is a crapshoot. I'm pretty sure the issue is in Apple's software, not the docks.
Turns out that when you amass a significant portion of society’s resources then society will be interested in what you do with them.
It can also be tried for general anomaly detection https://dl.acm.org/doi/fullHtml/10.1145/3624062.3624121
Open weights != local models.
All the more reason to focus on those service guarantees, integration, and lawyers while making the underlying model easily swappable to whoever’s winning the frontier model involution battle at the moment
This is much harder than it sounds. Most techniques I’ve seen end up using separate agents to do the planning, implementation, and judging.
The elaborate workarounds you have to build to help an agent which fundamentally doesn’t know what it’s doing reminds me of this old blog post about TDD: https://pindancing.blogspot.com/2009/09/sudoku-in-coders-at-...
IMO present technology is tailored for an experienced developer to give agents manageable tasks that can be one-shot. The marketing right now reminds me of the 90s when AskJeeves promised natural language search when the technology was fundamentally still stuck in keyword search, and learning to craft a search query for Google is today’s prompt engineering
"Some of you may die, but that's a sacrifice I'm willing to make" was supposed to be a comically evil statement in Shrek
I'm gonna go ahead and assume you don't believe in driver licenses and speed limits either
They're legal to buy like an hour out of the city unfortunately
Also the 2A wingnuts think banning fireworks is akin to gun control https://www.tully-weiss.com/blog/fireworks-and-the-second-am...
It's telling that the proponents of vibe coding are the same people who think memory management is no longer a concern with Java
You created a throwaway to “not troll” and post the same three tired tropes every tokenmaxxing vibe coder trots out
Healing by reverting to seven year old pre-slop articles :)
It’s crazy that straightforward rules like this can’t be followed and yet they think they can gate Fable
Yeah, the neoclouds and hyperscalers are taking massive losses right now, self hosting is basically signing yourself up to do the same. There are philosophical reasons to do so but it’s a terrible economic decision
I mean based on all the "coding is solved" hype that's what these companies are aiming for
But those proofs are showing that the fundamental axioms (which are generally simple and elegant) are still enough to build a complex result.
I think of elegance as not having to add epicycles, not that everything in the system has to be simple.
Also, without a working theory the, the space of possible solutions is near infinite. LLMs manage to pluck out the space of comprehensible English strings from n-dimensional hell. Even if this is done with a black box of billions of parameters, it’s still elegance in the sense that such a space even exists and was found
Joke’s on you, Claude is writing that too
That’s the scenario where we’ll all be using Chinese models
This is the is/ought problem (https://en.wikipedia.org/wiki/Is%E2%80%93ought_problem) and it’s unclear whether an objective general solution to this even exists, especially constrained within the framework of language that LLMs are stuck in
Yeah, they could just inflate Bancor the same way a country can with their own currency today (i.e, printing dollars) to manage their debts
Tbf it was proposed in a time where globalization was a good thing and there was naive optimism about international organizations!
There is an idea floating around that trade imbalances create global inequality (Trade Wars are Class Wars by Klein & Pettis) and the original sin was adopting the dollar as the reserve currency instead of something like Bancor (https://en.wikipedia.org/wiki/Bancor)