It does.
HN user
pylotlight
You didn't read the blog post at all did you? Clearly you have zero clue what you are talking about.
Does that not prove this test should be well burried by now?
The cost there is multiple rounds of review tokens making it both slow and expensive.
But isn't it running basically 1 request at a time? This would make agentic coding difficult right? Compared to running as many sequential tests as you want via api?
wait until you hear why, and that you can optionally use e2e but lose other functionality.
Because the two sets of skills aren't related. Library and design work are not the same thing. That's why
It's really that weak?
worked for me.
and yet 5 more popped up.
roughly ~50–56GB, although this is somewhat configurable with iogpu.wired_limit_mb. By default, macOS reserves ~25% of memory for the system.
There are none. People forget the other side wanted to shut down research entirely, not just release. No idea why people think the other side would have been any better, it would have been even worse. On top of that, anthropic got exactly what they wanted.
No anything but wasteful, weak, expensive, environmentally harmful solar. Nuclear is the only path forward for superior energy production, at least until we figure out fusion.
I'm pretty sure the vision/hwa reqs for cars is much less than an LLM/genai in general so that doesn't quite work out. But it would be nice to have an AI server with wheels :p
It's not tied at all, the AI integration is on top of. So OSS applies to the app itself.
You'll note that's the claude app if you actually look at the preview, which while confusing advertising what isn't even your app, does show how it works with the hand off.
Ye built in AI is the only way this makes sense to me. Or otherwise could just add a terminal where you can run any TUI mapped to that notes/vaults dir or something.
Do I dare mention the family guy way to achieve this? :P
I think everyones been asking for "what's next" around git for a while as well :P Theo certainly mentions it plenty.
SVG generation is a useless test, what's there more to know?
The only real essential item here is tool calling capability is it not? So I assume they tested a strong read/write/edit tool consistency?
As in, you learnt that a useless test that no one should be using was tested here, that's what you meant right?
Melbourne Australia feed in rates have also been significantly reduced since the start.
Brew installation? Not looking to use pip or load manually.
For single bins or otherwise, brew is definitely preferred.
Just plain incorrect.. please stop spouting this nonsense, this is not the reason whatsoever.
rein in? You don't like going to space? What do you have against progress?
If that was how it was phrased I think there would have been less push back, but that's not at all how it's been communicated. There is no assumption to rereview at a later date at all given the focus on the AI usage etc.
If they said we will rereview in 1-6 months or whatever the whole discussion would be mute.
They used to have good limits that lasted hours, now I wiped mine in a couple of minutes..
My general understanding of the concenus on most models these days is that people consider google models to be some of the worst at tool calling, so certainly an interesting choice. Did you do any evals on this?
Performance? Second only to rust and other lower level langs. Surely you don't need this spelled out for you...