Can you explain how the above event doesn't count as evidence alignment is an actual risk?
HN user
zaptrem
Given these models could not have been trained in the first place if they had to license every line of random fan fiction on the internet, I think distillation also being fair game is a tradeoff everyone should be willing to take (unless they want to decelerate, but that's a different conversation).
This is my number one complaint about the M-series MBP line. Especially true of the cutout in the middle that has points so sharp they can cut you if you accidentally scrape it with your hand.
Why haven’t we seen any queues or the like over the past week then? If it’s truly a capacity limitation why not just boot subscription users to a lower priority queue or limit usage to outside peak hours?
Should we require the destruction of the brains of those that watch pirated movies?
Needs more WebGL spinning rubik's cube
Can you include GPT 5.5 non-pro (extra high thinking I guess) in your comparison? GPT Pro is the "I am willing to torch cash for a sooometimes slighty better result" option, not the one people are actually expected to use daily. That's probably part of the reason it's not in Codex
OOM on CUDA GPUs is relatively graceful (the process crashes). However, on macOS if torch MPS tries to allocate too much memory, the whole kernel will simply lock up and the only option is to reboot the computer. I have no idea why Apple doesn’t reserve memory for stuff like the OOM/kernel watchdog, but it seems they either don’t or there is a bug.
Love me some JSD. Here is a problem most people don't consider with generative modeling (e.g., AI text, image, music, video models): basically all standard pre-training algorithms for generative models (i.e., cross entropy, basically all diffusion/flow formulations) are closer to a Forward KL divergence. In other words, given limited capacity the model will try to stretch itself to cover every mode. This gives you a jack of all trades (lots of knowledge and diversity), but a master of none (you get blurry images and text filled with nonsense).
The real magic in generative modeling comes from the post training process that comes after, which usually (e.g., RLHF) approximates Reverse KL (given limited capacity, try to perfectly cover what you can, but it's fine to drop the rest entirely). This gives amazing results, but is also the cause of AI oddities like the "AI Image Pixar Look", many of the verbal tics of LLMs, and all AI music using the same small set of voices. Jensen-Shannon Divergence sits right in the middle of Forward and Reverse KL and is what many GANs are claimed to approximate. Ideally, it is a better trade-off between diversity and fidelity.
V4-Pro is about 2.4× total params and 1.3× active params of V3.2.
Seems pretty clear, Claude and Codex were getting a lot of free publicity by instructing their models to do the same and MS wanted similar results. However, a bug caused this to be applied to all commits instead of all Copilot-influenced commits.
I bumped from $20 -> $100 today but the Codex CLI lacking code rewind and "you can change files but ask me every time" mode from Claude Code is quite annoying. Sometimes I want to code, not vibe code lol.
Agreed, that’s why I specified end to end (I.e., text to waveform)
My point is you should consider creating truly undetectable audio end to end with AI to be effectively impossible for the foreseeable future (i.e., I would bet money it is still trivially detectable five years from now). It won't be detectable to humans, though, only models.
I train music generation models. They are very trivial to detect. In fact, detecting them then training them to evade detection by the detection model is a big part of training them! But the detectors win instantly without some hardcore regularization. Simply turn that off and you've instantly got a perfect classifier.
This isn't like text classification, the signal many orders of magnitude higher bitrate and so many more corners need to be cut. It's likely going to be nearly impossible or at least not remotely worth it to generate an audio signal that is truly undetectable in the foreseeable future.
What's your reasoning effort set to? Max now uses way more tokens and isn't suggested for most usecases. Even the new default (xhigh) uses more than the old default (medium).
YouTube et al's automated copyright systems put way too much trust in the hands of those making the claims.
Many of the games that actual kids spend time on are the purest expression of gaming slop (half-broken microtransaction gambling hell with schizophrenic flashing colors). Roblox and Fortnite's Islands system are both guilty of this. The problem is kids don't know any better and don't yet understand the value of money. The obvious response is "parents should handle this" and while I agree, there is no system to let them say "here are Robux/V-Bucks you can spend on quality content (e.g., Fortnite's Battle Pass is very well designed, quality content), but gambling slop is disabled".
In my experience, the Epic Games Store downloads faster, installs more efficiently, and launches games faster than Steam. The social features I actually use (i.e., add a friend, join them in a game) work fine. I'm not aware of any features Steam has that EGS lacks that I actually use frequently (Valve's VR, streaming tech, and Proton are great, but I don't use those frequently). It's not just me, many indie game developers are also big fans of EGS (most recent example that comes to mind are Jeff Kaplan's remarks during his 10 hour stream a week or two ago). Gamers' vehement defense of what is effectively a monopoly continues to confuse me.
I have Max 20x and they're still separate on 2.1.75.
Data centers don't do anything other than sit there and turn electricity into heat. They only emit nothing but heat (which could be useful to others in the building).
What did Epic do?
"Previous data from the trial reported that 107 participants received the mRNA vaccine and Keytruda treatment, while the remaining 50 only received Keytruda. At the two-year follow-up, 24 of the 107 (22 percent) who got the experimental vaccine and Keytruda had recurrence or death, while 20 of 50 (40 percent) treated with just Keytruda had recurrence or death, indicating a 44 percent risk reduction"
Statistically, if those in the control group had gotten the treatment, then in expectation 9 of those people wouldn't have had their cancer return or died. It must be exciting to run these sorts of trials with super promising drugs, but also a little bittersweet/dark.
See here for a truly random sample of human music: https://0xbeef.co.uk/random/soundcloud
Thankfully, most of it doesn't reach your Spotify feed. I think most of it is garbage, but I'd fight for the right of people to continue posting it. All things algorithmic have this exploration/exploitation, diversity/fidelity tradeoff and Spotify has theirs tuned very heavily toward exploitation/fidelity. I think there is a cool opportunity for someone to put the tradeoff dial into users hands.
I’m a founder of one of these AI music companies and that noise you’re describing (it differs between co’s for us it’s loud vocals, for Suno it’s vocal aliasing/sandiness and mushy instrumentals, etc) is exactly why I think these songs should not be going on Spotify/etc.
We’ll have this (and the corny lyrics issue) mostly fixed in a month or so, then it mostly becomes a recommendations problem. For example, TikTok is filled with slop, but it’s not a problem - their algorithm helps the most creative/engaging stuff rise to the top. If Spotify is giving you Suno slop in your discover weekly (or really crappy 100% organic free range AI-free slop) blame Spotify, not the AI or the creators. There are really high effort and original creations that involve AI that deserve to be heard, though.
I suggest going back and listening to some of the first experimental electronic music. The tools have improved a lot since then and people have used them to do really cool things, even spawning countless genres.
Not sure it’s a cultural thing since most of the copy coming out of DeepSeek has been pretty straightforward.
This sounds like the same basic voice systems we’ve had for 15 years. Idk if that counts as modern “AI”
Not sure where that math is coming from. Assuming it's true, you're ignoring that some users (me) already pay 10X that. Btw according Meta's SEC filings: https://s21.q4cdn.com/399680738/files/doc_financials/2023/q4... they made around $22/month/american user (not even heavy user or affluent iPhone owner) in q3 2023. I assume Google would be higher due to larger marketshare.
I've fed thousands of dollars to Anthropic/OAI/etc for their coding models over the past year despite never having paid for dev tools before in my life. Seems commercially viable to me.
Latency may be better, but throughput (the thing companies care about) may be the same or worse, since every step the entire diffusion window has to be passed through the model. With AR models only the most recent token goes through, which is much more compute efficient allowing you to be memory bound. Trade off with these models is more than one token per forward pass, but idk the point where that becomes worth it (probably depends on model and diffusion window size)