Well making an actual sandbox before testing offensive abilities of their supposedly _really dangerous model_ would have been nice...
HN user
desterothx
interesting, I haven't played with any of them yet, but i thought the point of orthogonalizing the weights towards the restriction vector was that there is no loss in capability while removing guardrails. Does it affect other parts of the RL alignment too?
Almost like the entire worlds manufacturing was pushed to China the last 40 years (and is now getting moved to cheaper countries with the rising Chinese middle class). Is there a similar explanation for why the US number is so high?
Well implemented by the current US administration. Every accusation is a confession
IP laws work if the penalty is larger than the use case of the data. Pretty sure literally every US company is using the data knowing they won't be punished equivalently. We're living in the age of trillionaires, Facebook paid 5b$ for the cambridge analytica scandal, something Elmo could write off as a business expense at this point...
In all open weight model discussions there is at least some anti-Chinese sentiment
Probably better to enable it only for strikes/spares as it would drain too much time, but the idea is great.
According to Pat Toulme, the thinking traces and outputs that the Chinese researchers distilled are useful for getting initial trajectories, to prevent a cold start during RL. Once you get those initial correct trajectories (the model actually solving a task), you can generate you own new traces, so the distillation is already done, there is no need for further reasoning traces. Sure they could probably get better alignment with the frontier by distilling fruther, but in any case the damage is already done, at this point hiding reasoning is mostly just hurting users
What's your config, how does it compare to pi
the article is literally about the model upgrade not being a one liner
It doesn't seem correct at all though? the supposedly best one isn't even the best gemini flash output (the medium one looks better than the high one)
If you don't have an agent heavy workflow, you'd be surprised how far the 20$ subscription stretches. I used ~500$ of usage in the last month on a team 20$ plan.
Really almost all benchmarks I look at have a cost per task column, which is basically the code size metric if you take an extra step
Just tested on xiaomi hyperOS, and their copying of iOS has gone to such lengths that it works the same, it buffers the clicks properly
People really need to read Dijkstras Go to statement considered harmful letter [1]. If the obscurity of go to for static analysis of the code was too much, of course bringing in a literal ai black box is harmful for stable processes.
[1] https://homepages.cwi.nl/~storm/teaching/reader/Dijkstra68.p...
You're right, he should have said people became less horny
Someone hasn't been following Le chaton fat news, agi is coming
And yet the average intelligence continued to rise through the reduction in numeracy. What job do you have where your performance is graded by your "raw intellect" instead of capability? A student?
If you have a problem with the math notation, you should open a law book, and look at the terminology mess they have going on. I like math notation, because it can be simply converted into natural language, thus removing the notation from the equation and leaving pure logic. Once you have that, thinking about it and working out what the original notation actually meant should be on you, that's how you build an understanding of math.