Have you tried 2.5 Flash Lite to cut costs further?
HN user
piecerough
"quantize enough"
though at what quality?
What's a decent european enterprise?
What US government attacks?
[...]
But there is something fundamentally different about talking with a bot as opposed to a person. A person can be a friend. An AI cannot be a friend, despite how people might treat it or react to it. AI is at best a tool, and at worst a means of manipulation. Humans need to know whether we’re talking with a living, breathing person or a robot with an agenda set by the person who controls it. That’s why robots should sound like robots.
You can’t just label AI-generated speech. It will come in many different forms. So we need a way to recognize AI that works no matter the modality. It needs to work for long or short snippets of audio, even just a second long. It needs to work for any language, and in any cultural context. At the same time, we shouldn’t constrain the underlying system’s sophistication or language complexity.
We have a simple proposal: all talking AIs and robots should use a ring modulator.
I don't necessarily agree, but it reminded me of why electric cars still have engine sounds.
SFT forces the model to output _that_ reasoning trace you have in data. RL allows whatever reasoning trace and only penalizes it if it does not reach the same answer
I think the reason why it works is also because chain-of-thought (CoT), in the original paper by Denny Zhou et. al, worked from "within". The observation was that if you do CoT, answers get better.
Later on community did SFT on such chain of thoughts. Arguably, R1 shows that was a side distraction, and instead a clean RL reward would've been better suited.
Who's Tavi?
Have you had a common theme for these projects you navigated?
It's very related to LLMs. Though instead of text tokens you are working with audio tokens (e.g. from SoundStream). Then you go to audio corpus, instead of text corpus.
It's great!
It would be very interesting if LLMs were no longer static.
Little bit of a nightmare too. Instructions keep piling up for you that you no longer openly can access and remove
I remember a French institution could not buy our product, because they had a contract with a local manufacturer.
I doubt this is a EU thing. It's due to exclusive contracts/licenses. This happens everywhere?
This is only going to get worse with Large Language Models. Let's imagine a somewhat knowledgeable individual, could craft both emails, messages and even commits with a bunch of prompts. Those will relate deeply to the project.
Isn't this what we're all betting massive Transformer architectures are going to give us? Tools to explore and handle complex concepts. Reasoning may still be left to us, though.
"We are also releasing three new datasets: Screen Annotation to evaluate the layout understanding capability of the model, as well as ScreenQA Short and Complex ScreenQA for a more comprehensive evaluation of its QA capability."
Looks useful to me for replicating some things. Good stuff!
That seems brutal, indeed. Why did you move there in the first place?
So what's next?
As a FAANG employee, working with ML, what do you want to get from other companies, besides more money?
It's hard to have more chips, for example. You run less experiments, you have less throughput in an already computationally tight environment.
In today's market, if you are available for the intro call, recruiters go hunt for the next hard-to-get candidate. Bigger likelihood it'll be an actual conversion.
Happened to me.
The company added that some Search features will be removed in Europe to comply with the DMA, including Google Flights
What?!
Tell us more about "I don't aee anything suspicious". How exactly do you know it's not a binary that hashes all your files using a key and asks for btc to revert?
Super hard question. I think it's a bit too early to put these models in hands of kids unsupervised.
Thanks for writing this question up ;)
It's not missed - pornpen.ai
What is this klima bonus?
Nice, one would expect so. Though if you think more deeply about it, I think not. The language itself is the manifestation of capabilities, but the process exhibiting them is the underlying system, eg. neural nets or human brain.
I came here for Tensorflow.
Too bad it's paywall'ed
The example is brilliant and leaving it hanging for the reader is the intended purpose (e.g. a far-away third-party like me reading a team's report doesn't have a reason to care past that)