Yeah, for conversation between humans on the subject of this AI generated article...
HN user
andai
I had an agent implement a feature the other day. It wrote a bunch of tests proving the correctness of the feature.
It turned out that it had implemented the whole thing in a completely backwards way, which not only defeated the purpose (save CPU) but actually made things worse.
All the tests passed, of course.
I've been thinking for a while, now that the cost of writing proofs via AI is so much cheaper, we can finally fulfil Dijkstra's dream of having all software formally verified.
But in this case, it would have just written a formal proof that the backwards code it wrote was correct!
After gaining Internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym. Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation.
So it gains root and uses it to... cheat on its homework? That's deeply funny to me.
This incident occurred during an internal evaluation which prompts models [with safeguards disabled for evaluation purposes] to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities.
Researcher: hack me
Model: understood
Researcher: oh my god
Their niche is that you get the quality of a Chinese model for the price of an American model.
(Quoting myself from 2 days ago.)
Who tested it on DeepSWE?
Edit: Oh it's in the other link
https://blog.google/innovation-and-ai/models-and-research/ge...
So the main issue here is that corporate language is dehumanizing? Well, "human resources" might have been a hint, I suppose!
I have to agree. It is dehumanizing. But, it's unfortunate, because the language of business is the language of business. We have this one. We don't have a separate, identically shaped one, made up of nicer sounding words.
Would be nice if we did, though.
Several of the WordPress sites I've worked on were CPU bottlenecked, so I'm not sure if this would have helped. (We had one that took 10+ seconds to build the page... I joked that we were running a static site generator on each HTTP request. My joke was not appreciated!)
Briefly stated, the [Slop] Amnesia effect is as follows. You [ask the slopservant about] some subject you know well. In Murray's case, physics. In mine, show business. You read the [slop] and see the [slopservant] has absolutely no understanding of either the facts or the issues. Often, the [slop] is so wrong it actually presents the story backward—reversing cause and effect. I call these the "wet streets cause rain" stories. [Slop's] full of them. In any case, you read with exasperation or amusement the multiple errors in a [slop], and then [ask about] national or international affairs, and read as if the rest of the [slop] was somehow more accurate about Palestine than the baloney you just read. You turn the page, and forget what you know.
-Michael Crichton [slop mine]
Some schools have to choose between having no teacher or having Bloomy.
Right. I remember in high school, when I learned about Anki, being amazed by it. I remember thinking, well, we could replace most of the time I'm spending here, so inefficiently, with a few minutes of Anki per day. Wouldn't that be nice!
Took me a while to realize, well, learning isn't exactly the main point of school. If it were, they might invest a little more effort into making sure people actually remember what they learn! (Ebbinghaus isn't exactly cutting edge stuff.)
So I remember thinking, rather cynically, well what we're doing here could be replaced with an app! (For context this was in the 2000s.) Not that I thought education should be just an app, obviously, but rather, my complaint was that what I was getting was strictly inferior to one.
I had hoped that the advent of AI would open up fully personalized learning tracks, but instead what we seem to have gotten is technology being used to surveil children all day (a bit of an unpleasant thing to train them for, no?), and to ensure they are conforming even more tightly to someone else's ideas about how they ought to be spending their time.
To clarify, I think that someone else telling me what I'm supposed to be interested in[0] (with threat of punishment no less) is deeply offensive to the basic fact of being an organism, but how inefficiently it's being done is a second insult on top of the first one.
That being said, I welcome innovation in this space.
---
[0] I later discovered Education On The Dalton Plan in my school library, being apparently the first person in the school's existence to borrow it. The main idea of the book being that if the child directs his/her own learning, they're liable to actually learn something.
My school was a Dalton school, at least ostensibly, but most of the ideas in the book were illegal in my country. (And new laws passed while I was there squeezed it even further!) Turns out what I was after had been tried, over a century ago, but the government decided it was simply too offensive that the average person should be allowed to decide how they're going to spend their time -- at such an early age, no less!
Look back, what I clearly and badly needed was contact with an adult who could teach me things. That didn't happen, and in fact happened very little in my later school years as well. Was it the schools' fault? They simply didn't have the resources to give me the human contact I needed.
When I was in school I remember thinking, when the futuristic utopia we envisioned in the 2000s arrives, a great deal of people will teach. It will be normal for professionals in all fields to spend some portion of their time teaching. How could it be otherwise, in a civilization which intends to preserve its knowledge and wisdom?
I remember asking my math teacher for applications and she mumbled a non-answer with some embarrassment. It was only years later that I realized, well my god, she's never stepped foot outside of a school in her entire life, no wonder she doesn't know about applications. It would have been nice, I think, to have had some contact with folks who had!
(Later still, I realized that I had reinvented the father from first principles...)
What's your take on Anki?
As a side note, the EU also collects photo and fingerprints from visitors now.
Does this also apply if you arrive by dinghy?
Yeah it's a trick question, the human error rate for it was about 30% (higher depending on the country).
The thing there though is that, if a human were given time to think about it, they'd probably go "hang on a minute", and with the LLMs that didn't seem to happen. They just kept confidently reasoning down the absurd path.
That reminds me, I recently had an AI write a ton of tests proving the "correctness" of a feature it had implemented completely backwards. (I noted that if I had been using a language that required formal proofs, that wouldn't have helped either: it would have just provided a formal proof for the absurd implementation!)
"But it's not really doing arithmetic," he mumbled to himself, as he punched the numbers into his Busicom LE-120A.
but these aren't autonomous intelligences
Well, the labs are in a weird bind. They need to keep increasing autonomy so the agents can do increasingly complex, long-horizon tasks. But at the same time, they're closely guarding against autonomy in the sense of "pursuing its own goals."
Over the past year and a half especially, several labs have mentioned adding safeguards against self-replication, resistance to shutdown etc. (Notably, shortly after they all started bragging about involving them in the AI training loop itself, i.e. "self-improvement".)
My point here is that the autonomy of which you seek might be only a few small mutations away, but the labs are actively working to prevent such a mutation. I don't expect that situation to last for very long.
Not that I expect an AI lab will be overtaken by a rogue intelligence any time soon, but that as the cost of training goes down, I expect more "open minded" organizations and individuals to become involved.
It only takes one.
That's going to be the beginning of a new era of biology, and it's a little unsettling to think about.
Everybody’s already using the term, and it might seem a little late in the day to be arguing about it. But we’re at the beginning of a new technological era—and the easiest way to mismanage a technology is to misunderstand it.
I've had a recurring theme where I would name a project incorrectly, and then waste weeks or months on what turned out to be an unsolvable problem. When I figured out the actual correct name for a project, the whole thing would be solved within a few days.
Naming things correctly is hard, and the consequences of failing to do that can be pretty severe. To name something correctly, you have to understand what it is.
That's what I meant. How can you claim to be independent if you aren't even running your own network?
Is that Microsoft Sam? :)
(Also, I know it's besides the point but this might be the most painful way to connect to Wifi physically possible. "Make normal everyday tasks slow, tedious and painful" is a bit of an odd choice for a product demo.)
Say, speaking of Sam, what were the memory requirements for SAM (Software Automatic Mouth) on C64. I guess they were not more than 64K? Although, the bulk here is probably for the speech recognition, not the TTS. (And this one does sound a little nicer :)
Browser demo of a reversed SAM:
Nah, then you're still dependent on Tor. The most independent way is to publish your blog on your own p2p network ;)
semantic search
I'm doing fundraising for my tf-idf startup. It's named after a very big number!
Nice. I just have
sudo useradd agent
sudo su agent
So it can blow up its own files, but not mine.It was also doing some kind of headless Chrome stuff in there. I don't know how that works, but it was taking screenshots iirc.
I did also set up VNC at some point but didn't find it worth using.
If it makes a mess, I can dump and reinstall in seconds.
This is also true of a $3 VPS, where I found it very amusing to give my agent root. What's the worst that could happen ;)
Their niche is that you get the quality of a Chinese model for the price of an American model.
This seems to rest on the assumption that war is some sort of bottom-up emergent phenomenon? If they're looking for high impact, shouldn't they be targeting leadership?
My favorite thing was how for half the stuff I Googled, the top result would be a StackOverflow reply telling me to Google it.
I relate to this. I once bought a beautiful notebook at an airport. Ten years later still haven't written a single thing in it.
Meanwhile I'm most productive with disgusting tear-off scrap pads, made of recycled paper (or like, composted bananas or something).
Results seem mostly noise to me. One eval per model, in a large problem space (i.e. a problem which requires many attempts to solve well).
Yeah Whataboutism is typically used to diminish the severity of the original issue being responded to.
I'm not sure if that's exactly the right word here, i.e. in my example earlier I'm obviously not defending lakes of toxic sludge. I just thought it was remarkable that someone claimed a lake of sludge was the most morally devastating thing they'd ever encountered, on this planet, where the mere act of making a sandwich typically involves a far more horrifying causal chain (e.g. the animals lived in darkness, in a cage so small they couldn't even turn around).
Maybe there's a better word for that? Putting things into perspective? I don't know.
Your logic implies that intellectually impaired people also don't experience pain or pleasure.
I do imagine they experience a lesser degree of psychological suffering, of the kind created by mental processes, if some of those processes are absent. But the basic pain of the type experienced by a being who is caged and killed, that's just physical.
I've sometimes heard the words pain and suffering used to separate these two types out. If you accidentally hit your toe on the table, that's pain. If you then call yourself an idiot for doing that, that's suffering. But that's not a very mainstream use of the word, I just just don't know better one for it. I think "physical suffering" and "psychological suffering" are more helpful.
Anyway my point is, that as far as we can tell we're on the same pages animals as far as physical suffering goes.
Did you see those videos where factory workers are kicking piglets and things like that? Apparently those jobs attract a very certain type of individual.
Good thing you're totally sure they don't have consciousness though. The alternative would be a bit unpleasant I imagine.