Surprising they didn't clean that from the data before training. It's easy to identify, a simple search->replace gets most of it, and a cheap LLM can identify the edge cases (e.g. avoiding "Claude Shannon" -> "Kimi Shannon" or something).
HN user
ComplexSystems
How many tokens would it cost to write some library functions to fill in the gaps?
Who, exactly, should help people host a website?
So the post-introductory price is set such that Sonnet 5 will cost 100%-135% as much?
Sonnet 4.6 is ahead of Opus 4.7? Hm.
The guy is genuinely worried for his future. I see nothing shameful or narcissistic about that, and I don't think it makes sense for you to wish him any harm.
Yes, they certainly should have taken one of the many other jobs that are widely available right now.
Who can afford to use this damn thing though? They're pricing everyone out of the market with stuff like this.
Can you imagine ChatGPT terminating a conversation because it thought your question was "low effort?" The behavior wouldn't be viewed as helpful or aligned, and nobody would use it.
StackOverflow was, all too often, not helpful or aligned. It died because the staff were unable to get the moderators to be helpful.
That's because the purpose of this article is not to have an objective debate over their abilities at all. Most interesting research in this field isn't. Instead, it's to present a new technique to improve LLM performance, which is much more interesting than (once again) rehashing the philosophy of LLM personhood.
That may be true for now, but it seems clear enough that letting the model use Lean in its internal reasoning process would be a great idea
Sure, you can substantially reduce the information collected about you as long as you don't just give it away somewhere else instead.
The reason I think this is a bad idea is that it lulls you into a false sense of security. The article makes recommendations that seem thorough and sensible - keyword "seem" - but, as mentioned elsewhere here, there are other potential hidden sources of telemetry (in CarPlay and Android Auto), and who knows what else.
For this kind of thing to succeed as a general lifestyle, you would need to invest an enormous amount of time making potentially irreversible modifications to all kinds of electronic equipment - only to be virtually guaranteed to miss something.
Do this kind of thing if you want, but don't be fooled into thinking you're actually solving the problem for real.
You know, back when it was a noble democracy where all men were free, or something.
Surely there's room for the view that this is misaligned behavior for ChatGPT to have. I would guess this is during the "sycophantic" phase last year.
Yet another round of layoffs. Is there a fallback career? :-/
That makes sense. If you could magically just get the top d PCs in quadratic kernel space without having to compute the whole kernel matrix, and then just do top-d quadratic PCs -> ridge, would that be better than doing the PCA -> top-d -> quadratic kernel ridge as you are now?
Absolutely. Who cares if the LLM automates some of the grunt work? Mathematicians are artists, and they paint with ideas. The goal is map out more of the beautiful structure of how things work. The enjoyment in it derives entirely from the payoff of seeing the larger view of how things fit together. If part of their process involves bouncing things off of other people, or even LLMs, I don't think it matters much, nor does it take away from the enjoyment in getting things figured out.
What makes this different from just kernel PCA with the quadratic kernel?
People can and do see unidentified things and take plenty of photos of them.
While I am sure FreeBSD is more secure than your average Linux distro, I sure hope they are using these new AI models to harden everything.
Good article, but
"We take the exponential of each input and normalize by the sum of all exponentials. This transforms a vector of arbitrary real numbers into values between 0 and 1 that sum to 1, it technically this is a pseudo-probability distribution (they're not derived from a probability space), but it's close enough to a probability distribution and for practical purposes they work just fine."
Why is this a "pseudo-probability distribution?"
It seems like some kind of technique is needed that maximizes information transfer between huge LLM generated codebases and a human trying to make sense of them. Something beyond just deep diving into the codebase with no documentation.
Yes it does; you can build the absolute value as sqrt(x²), and sqrt(x) and x² are both constructible using eml.
This line of reasoning makes no sense when the AI can just be given access to a fuzzer. I would guess that it probably did have access to a fuzzer to put together some of these vulnerabilities.
So the data cannot possibly tell you anything about how likely is the observed outcome, because the observed outcome is the only outcome that you observe.
This could also be viewed as supporting the Bayesian perspective, where the observed data are not viewed as random variables - they are fixed. This is because, as you say, the observed outcome is the only outcome that you observe. It is the classical setting, in comparison, where we instead do our analysis by treating the sample as a random variable, placing the counterfactual on other non-observed values ("what if I had drawn a different sample?"), even though we didn't. Bayesian methods treat the data as gospel truth, and place the counterfactual on the different parameters ("what if the population were different?"), even though it isn't.
The other criticism you have is
The problem with this approach is that we can only observe ONE level of treatment effectiveness, i.e., the level of treatment effectiveness that the treatment actually possesses. All other possible levels of effectiveness are entirely hypothetical.
This is true of both Bayesian and classical methods. We build models that would explain how different hypothetical levels of effectiveness would affect what data we should expect to see - that is the whole point. Classical methods also involve exploring scenarios in which purely hypothetical values of the parameter may be potentially true, and characterizing counterfactual samples that could have been drawn from them, even though in real life they couldn't have been.
I found it surprising that this article persistently did not capitalize the word "Bayesian." Is this a new trend or something?
I certainly don't. If a software developer has found a way to use these tools that works well for them and produces good results, that's a good thing.
They are mathematical models of what human beings would say. That's it.
the code is not public, so we can't know.
I feel like you're making this statement in bad faith, rather than honestly believing the developers of the forum software here have built in a clause to pin simonw's comments to the top.