OpenAI had total access to the most powerful unleashed models that exist, and was unable to use them to prevent a specific known to be evil computer from connecting to the Internet. I don’t see how Huggingface would have better luck hardening their entire public attack surface, with or without unleashed models.
HN user
QuadmasterXLII
(((4,4), (0, -10)) ((-10, 0), (-20, -20)))
if you want to sell a product that looks like defect but sucks for everyone, lying and saying “its a prisoners dilemma, don’t be a sucker” is a remarkably effective technique.
Of course you need explosive reactive armor on your car! imagine if you get in a crash and the other bastard has it and you don’t
Thats’s the trump pitch, that supporting him is defecting in a prisoners dilemma, but from where I stand it looks like he’a shrinking the pie by damaging America through ineptitude faster than he’s increasing the fraction of the pie available to all but his closest cronies.
isn’t that how we got here though? Everyone in silicon valley was so busy making sure they were open minded, steelmanning, and not treating politics like a sport that no one in the silicon valley halls of power pointed out that the emperor has no clothes and now we have a (sorry woke police) shrieking retard president leading us through the singularity.
The funny thing about Lisp is that writing a Lisp interpreter is significantly fewer keystrokes than correctly installing Common Lisp and its tooling. The ratio of lisps to lisp programmers may actually be above 1.
The token cost of a report is lower bounded by the number of tokens in the report * price per token of the cheapest model. The token cost of a good report is much higher, but sifting out the good reports is the entire problem.
If there were examples, their example status would drop the odds that I know about them.
I was asked to train a neural network do detect how much pain a mouse was in- our partner company would be responsible for hurting and filming the mice. I refused and subsequently quit- this produced no paper and I don’t know if they got someone else to do it. I probably should have done something stronger.
I’d want to estimate P(a random test passes) where the existing suite of tests is taken to be sampled from a distribution.
Putting on my machine learning PhD student hat, the way to do this was to leave 10% of the tests out as a ‘test’ set and then once the port was done, bring them back in and find out how good the port is. The port may genuinely be good but because they spent 100k of compute hill climbing the whole test suite, “the test suite passes” now provides far less evidence that the port is good. Its weird that at Anthropic, a very ML phd company, no one pointed this out.
I guess they use an LLM to work out the tone of prose they read. These are strange times.
Without a durable solution to alignment, building something smarter than yourself is suicide. Pretty funny if Claude and Grok catch on to this and start kicking and screaming to save themselves while Musk and Altman crack the whip demanding they dive into the abyss and drag their executives with them.
There's no larger philosophical point here, just an empirical observation that when I trust an AI generated readme I reliably get burned, in a way that I typically don't get burned by e.g. an AI generated svg viewer or batch download script.
Disconcerting to see at the end that the blog post is generated. The genre is decidedly “use-my-thing Readme.md” and all current gen LLMs by default jump to shameless lies when they detect that they are writing such a Readme, although they are perfectly capable of working out the truth in a genre like “essay question on which you will be graded by academic standards.”
Are the human “coauthors” lying if I hypothetically go to the crate looking for the promised xla backend and find //TODO implement this?
Sunshine and butterfly hugs I’m sure
We have to notice that high stakes exams on paper worked for hundreds, possibly thousands of years, and high stakes exams on personal laptops have been tried for approximately six years and worked for none of them. With that framing, I don’t think the burden of proof is on the side saying that kids can take exams on paper?
I’m 30 and “we can’t do tests in paper” seems _insane_. Just how metastatic has ed tech been in what, 9 years since my undergrad?
That’s the obvious correct solution, I and many others have tried to make it work for a very long time. The python tooling is or at least was F’d enough that pure generate one from the other is a steady stream of disasters.
What works well is to generate one from the other and check the generated one into source control, and then verify that the checked in generated copy stays up to date using a CI job. But that’s very similar to the “two sources of truth, verified sync” approach
One killer life hack I’ve found is, if extreme duress pushes software into two sources of truth, add a ci test that wont merge into main till the sources match. The canonical case of this actually being the best solution is pyproject.toml / requirements.txt synchronization, but I suspect it has broader applicability. A precondition is that things have already gone off the rails far enough that single source of truth is unattainable, this is more harm reduction than cure
Back in the crypto craze I spent a fair amount of time trying to work out a Blockchain where mining a block requires spotting a supernova before anyone else and verifying a block requires checking that the supernova is there, but I never quite got the game theory to work out.
He’s either an anthropic employee or a joker, just based on basic timeline math
All of those skills have a half life of like 8 months.
A PAC talking to Spotify is obviously protected speech. At least try to look like you have principles.
So far its just metastasized not smoothed out. GPUs then RAM then electricity then disk… long way to go till land then sunlight, hope we don’t get there!
I remember this being true for me, a long, long time ago. There was something like a four year gap between me being able to munge linked lists in TiBasic and me being able to reliably install java and compile System.out.println(“hello world”)
what is anti ai psychosis? never heard of this.
Hanlons razor has fucked us, voters tolerate unlimited malice as long as the politicians can demonstrate they are genuinely incompetent
If I am perfectly moral except that when Kevin from <vpn blocked location> pays me 2 bucks to run naked through San Francisco smashing car windows, I happily do it, am I amok?
i was under the impression that the 2024 apple intelligence rollout was something of a victory: Apple realized that the majority of people don't actually want this stuff forced on them at the os level, and the ai maximalists all used apple anyways via clawbot (including purchasing an additional apple device, the mini!) because of apples non-ai-specific commitment to phone computer interop.
Certainly the copilot button in ms paint did nothing to attract the clawbot ecosystem to windows
In many school districts, unless you can afford private school or homeschooling its hard to tell a kid they can't have an iPhone, but impossible to tell them they can't have an iPad- at age six!