HN user

piecerough

260 karma
Posts10
Comments39
View on HN

[...]

But there is something fundamentally different about talking with a bot as opposed to a person. A person can be a friend. An AI cannot be a friend, despite how people might treat it or react to it. AI is at best a tool, and at worst a means of manipulation. Humans need to know whether we’re talking with a living, breathing person or a robot with an agenda set by the person who controls it. That’s why robots should sound like robots.

You can’t just label AI-generated speech. It will come in many different forms. So we need a way to recognize AI that works no matter the modality. It needs to work for long or short snippets of audio, even just a second long. It needs to work for any language, and in any cultural context. At the same time, we shouldn’t constrain the underlying system’s sophistication or language complexity.

We have a simple proposal: all talking AIs and robots should use a ring modulator.

I don't necessarily agree, but it reminded me of why electric cars still have engine sounds.

I think the reason why it works is also because chain-of-thought (CoT), in the original paper by Denny Zhou et. al, worked from "within". The observation was that if you do CoT, answers get better.

Later on community did SFT on such chain of thoughts. Arguably, R1 shows that was a side distraction, and instead a clean RL reward would've been better suited.

It would be very interesting if LLMs were no longer static.

Little bit of a nightmare too. Instructions keep piling up for you that you no longer openly can access and remove

I remember a French institution could not buy our product, because they had a contract with a local manufacturer.

I doubt this is a EU thing. It's due to exclusive contracts/licenses. This happens everywhere?

This is only going to get worse with Large Language Models. Let's imagine a somewhat knowledgeable individual, could craft both emails, messages and even commits with a bunch of prompts. Those will relate deeply to the project.

Isn't this what we're all betting massive Transformer architectures are going to give us? Tools to explore and handle complex concepts. Reasoning may still be left to us, though.

As a FAANG employee, working with ML, what do you want to get from other companies, besides more money?

It's hard to have more chips, for example. You run less experiments, you have less throughput in an already computationally tight environment.

The example is brilliant and leaving it hanging for the reader is the intended purpose (e.g. a far-away third-party like me reading a team's report doesn't have a reason to care past that)