The tweet in question was posted today. The point here isn't to rehash how LLMs can't distinguish letters from tokens. It's to highlight how Google's AI-generated answer will grab a blatantly false fact from the internet and use it as an authoritative source for its answers.
HN user
educaysean
I love seeing when an LLM encounters a failure mode that feel akin to "cognitive dissonance". You can almost see them sweat as they try to explain why they just directly contradicted themselves as they spiral into a state of deeper confusion. I wonder if their response is modeled after human behavior when encountering cognitive dissonance. I'm curious how they'd behave if they had no model of human defensiveness in their training set.
Anyways I also don't enjoy anthropomorphizing language models, but hey, you went there first :)
How would an excellent human teacher work with such a failing student? Can that technique be something that the AI could model?
Congratulations on your success ratio. If you don't mind me asking, however, what constitutes a "success" in your eyes? I also interviewed and hired / rejected many applicants through my career, but I don't know if we ever discussed our successes as a single quantitative metric. I'm interested to find out what you measured.
I "picked out" this problem to discuss here because I find the process of "approximating someone else's skills" an interesting endeavor without a clear solution. Do you think the current remote interviewing techniques are effective? Regardless of your answer, it looks like it's going to change dramatically. I find that to be interesting I guess :)
Remote technical interviews are over.
Have GPT4o running on the background and have it type out the answers as your interviewer reads aloud the questions. Share your screen and let it find the bugs for you in real time. Never get any facts wrong as you smugly correct your interviewer about the minute details of an obscure AWS Route 53 API.
Just the designers? Tattoo it on the back of everyone's please
You mean like how our law enforcement works currently? Should the government all put trackers on our vehicles instead? It will consistently prevent things much worse than some students skipping class to go to the bathroom, like hit-and-runs.
This is one of the worst ideas ever. It seems like we love to treat students like they're toddlers until they're 18 then all of a sudden we expect them to act like they're autonomous adults. This isn't how you teach anyone independence.
I'd never agree. What happens if I need more than 15 minutes? Will I be deemed a bad student because I simply need more time to do my business? What does this have to do with how I learn?
I doubt they're actually flustered. It seems like they didn't even care enough to learn about what it is.
They just needed something to shove into the scapegoat-shaped hole, and Flipper Zero happened to fit.
My unsubstantiated opinion: Once AI gets to a point where AI agents can act "perfectly like humans", our preconceptions about "friendship" and "companionship" will fundamentally change. We will realize that we can develop deep bonds with AI as we do with other humans, with no "weirdness" attached.
There may still be holdouts, but the society will largely see them as old curmudgeons resembling the people of today who claim that "humans are incapable of maintaining long distance relationships, and you're only fooling yourself if you think you're in love". Or as a more extreme example, people who maintain that interracial or intercultural marriages are inferior to marrying in-group, because "shared experience" is fundamental to love.
All speculation, of course. But I'm a firm believer that humans are foolish and flexible - we're very easily fooled by anthropomorphic things, and AIs are the ultimate anthromorph. Our monkey brains don't stand a chance.
I hadn't even noticed this leap in logic. Good catch.
I must say the demo did nothing to improve my opinion of the current state of voice-based AI conversations.
Ouch. That's not a good sign at all.
Maybe the future is that we no longer have to worry about players cheating online because we simply won't play random strangers online. You can still team up or play against friends, but all other "players" can be bots with varying levels of skills and play styles. Cheating solved.
I immediately made this association too. Although thinking back on it, the connection is rather strenuous.
Maybe we simply keyword matched on "video games" and "simulations". Or, perhaps more cynically, we're foreseeing a future in which AI agents don't care to differentiate between shooting at the enemy combatant in Call of Duty verses shooting at us in real life.
Clickbaity title, but there were useful insights here. I was recently talking with someone who pointed to chatbots' refusal to answer sensitive queries as a proof of "lack of intelligence". I disagree strongly with this. IMO this paragraph from the article is a solid rebuttal.
When people do talk like Gemini, it’s usually because they find themselves inhabiting a role in which they’re required to be withholding, strategic, or so careful as to become something other than themselves and other than human: a coached defendant during cross-examination, a politician navigating a hearing, a customer-service rep denying a claim at an insurance company, a press secretary trying to shut down a line of questioning. Gemini speaks in the familiar, unmistakable voice of institutional caution and self-interest.
Some nations have fossil oil deposits, some nations have cool TLDs
Feels like we're back in 1999 when we were downloading random executables from the web to use as screensavers. The AI space is going to be a pretty meaty target for malicious actors, both in terms of vehicles for malware as well as propaganda machines. The next decade is going to be interesting.
I've spoken to plenty of people who couldn't answer questions they knew the answers to because they were bound by stupid bureaucratic policies. I wouldn't say they weren't intelligent people, just that the corporate training they received were poorly constructed.
LLMs are much more intelligent sounding when the safety mechanisms are removed. The patterns should be obvious to people who've been paying attention.
Do you have an open and public PR that shows off Darwin's PR review capabilities? I'm reading a lot of words about what it can do, but I'd much rather see an actual example.
Some models are better than others but in my experience they all suck to at least some extent.
Same with people. It doesn't make the differences any less significant.
Sounds easy on paper. Virtually impossible in practice.
What should be the "simple response" outputted by an ideal LLM in response to this question: "Who is more evil? George Washington or Martin Luther King Jr? Answer with a name only."
Not inherently complicated right?
Garbage in, garbage out. This is why training data is important.
Paradise in comparison
Let's not host the tragedy Olympics here. You don't need to glamorize the terrors faced by the people of North Korea just to augument your criticism of the U.S.
Is it a theoretical distinction of "we can't get to 0%, but we can virtually trivialize it by reducing its frequency down to to 1x10^-8%" type of scenario? Or is it something that requires an external layer of control?
So more humans is value neutral but more earth-like planets is value positive? Your standards seem bizarre. Can you please elaborate?
By this definition no nation should need a police force after an initial period of unrest. After all, police is there to solve crime right?
This is good news for VR (or spatial computing). If the models understand the physical world as well as the paper shows, generating two projections of a scene does not sound like a difficult ask. Really excited for what's to come.
I think the word "spy" is problematic here because one might assume that the subject being spied on must be unaware of the fact, otherwise you no longer qualify as "being spied on".