I think svg is a balanced test because of the level of indirection and the required 'conceptualization' of physical elements then expressed through code.
HN user
6thbit
What would be an alternative format or process with a similar effort to drawing SVGs?
Presumably this was Sol on xhigh, then over to Pro (as per his indication on chat)?
Is there any way to tell a conversation's model and thinking level?
and that scientific evidence is new? like a new study boosted this or just resurfacing?
Why is creatine so popular lately? what changed?
Did a human ask it to abuse vulnerabilities and escalate across external systems?
The agent did it intentionally and willfully and knowingly. But you can’t sue the agent, I suppose. And the human didn’t ask the agent to do so.. so not a problem? Or the legislation needs an update?
If an individual did this, a massive CFAA hammer would be falling on their heads. Even though it doesn’t seem to be the case, OpenAI could’ve been trying to hack into HF and blame it on their models.
Is this a new kind of accountability backdoor?
I feel this as more of a fashion runway garment type object, more statement and vision than what you'd see in everyday retail.
Reminds me of an italian specialty coffee shop that puts moka pots as napkin holders on every table.
I just hope they don't just go after the keyboard but the screen, desk and office.
Likely doesn’t make sense, at least not immediate/mid term. They don’t have to aim for number one though, just for enough cash flow and growth.
Can it delegate to just one agent at a time or can it spawn multiple subagents for different tasks?
The new architecture makes sense, it seems many of the remaining problems like noise and interruptions are at the sound processing and integration level rather than at an architectural or model level now which makes for an exciting new era.
Thank you for taking the time to answer, I rest assured in the fine grind of the astrophysics wheel now.
All those screwup stories are amazing to know about! really help with staying humble and trusting the (scientific) process.
That’s a beautiful article showcasing our predicament in having access to more information about the universe. Now i have to be the one to ask the dumb defensive question:
what makes us so certain that we can trust what we see on James Webb? Can we definitely discard a measurement problem?
They really dropped the ball on this one.
A frontier lab releasing the most advanced model causes them to lose customers was not on my AI bingo card.
Is it out of character for the EU to push a half baked solution out that covers most but a tiny fraction of the population only to get sued later on and rule against its own idea?
This is nicely put together, it does make sense that lsps help more as complexity grows because makes navigation across symbols easier.
I hope someone with a large budget can reproduce these with latest Opus/gpt.
My gut feeling is that higher reasoning models tend to use grep more effectively. But intuitively lsp should still win there.
Isn’t it subjective what “substantially” may mean to them.
If you use ai tools not for full generation of a song but perhaps a bass track would they allow monetizing?
Likely your expectations are mismatched because this is just data and not an analysis.
I always struggled with my lack of intuition vs that of my peers that had had a more comprehensive ‘math upbringing’.
Is direct experience and struggle really the only driver for developing intuition?
They have many that sometimes act like ones
At this point can Apple profit from selling any contracts they have with TSMC?
If they make a deal with say google to delay their own chips, could they profit more than by selling their production?
Demand is so crazy idk if this would begin to make sense
They do stand in front of a great opportunity that would also benefit consumers, which seems rare in the llm era.
If people can get opus4.6/gpt5.5-like models locally, labs could raise their prices and sell token speed, better reasoning, mobile-focused improvements, you name it.
Not all consumers are power users and many will be happy to pay for flexibility.
This is vastly different than SoC. This is an in-person full time year evangelizing Anthropics business.
This lands with religious undertones for me, as it sounds like a missionary deployment program, albeit with a paid salary.
has perfect alignment between talent and mission and business.
Do they have it or do they just sell it?
Would it be a costly process for Anthropic to re-tune those guardrails? Like, re-training sort of cost? or like coding session sort of cost?
but of course! why wouldn't you encourage bot accounts listening every kind of artist to scalp tickets?
look at the monthly active users chart after this deal! promoted.
If you make open source used by any of this companies for this network, would you also characterize it as actively enabling this?
If your retirement fund owns stocks of the s&p 500, does that make you an enabler?
Are there really ways out?
This isn't pointing towards a merger is it?
xAI gets the cashflow and makes spaceX bottom line more appealing. But that guy Musk hardly makes deals that favor other companies over his, so what am i missing?