My guess is for distillation, they need to forward the prompt to Anthropic to get the real Anthropic model's response so they can train their own models on it
HN user
andrewmunsell
Any content I post on Hacker News is strictly personal opinion and does not reflect the views of my employer or anyone other than myself.
- https://www.andrewmunsell.com/ - https://mastodon.munsell.io/@andrew
[ my public key: https://keybase.io/andrewmunsell; my proof: https://keybase.io/andrewmunsell/sigs/iTG3yiQopPB2Za_q0DZMOVBh_5rxp_aExEVcXVUyDv4 ]
Improvements in model performance aren't always strictly compute-constrained in a way that makes them reliant on Moore's Law. Open weight models-- in particular, from Chinese labs-- are optimizing model intelligence with less compute. They're "behind" frontier models by months, but as others have noted, it's possible to get Sonnet 4.5+ level performance at reduced cost, today, from open weight labs.
This has been a thing for years, and so much so, that there's an entire TV with a dedicated second screen that shows you ads underneath your main screen: https://www.telly.com
No one rich enough flying what the average person would consider a "private jet" or private plane would be flying VFR from uncontrolled airport to uncontrolled airport. The "ultra rich" are not puttering around in single-engine Cessnas
Yes, but I had to update the Codex CLI manually via NPM to see it. The VS Code extension auto-updated for me
incorrect, its an o3 finetune.
This is Open AI's fault (and literally every AI company is guilty of the same horrid naming schemes). Codex was an old model based on GPT-3, but then they reused the same name for both their Codex CLI and this Codex tool...
I mean, just look at the updates to their own blog post, I can see why people are confused.
https://openai.com/index/openai-codex/
Edit:
Google just did it too. "Gemini Ultra" is both a model (https://deepmind.google/models/gemini/ultra/) and their new top-tier subscription plan (a la Open AI's Pro plan). Why is this so difficult?
The reason I have never bothered with Claude Code (or even other agentic tools), is that I still code mostly by hand.
When I am using LLMs, I know exactly what the code should be and just am using it as a way to produce it faster (my Cursor rules are extremely extensive and focused on my personal architecture and code style, and I share them across all my personal projects), rather than producing a whole feature. When I try and use just the agent in Cursor, it always needs significant modifications and reorganization to meet my standards, even with the extensive rules I have set up.
Cursor appeals to me because those QOL features don't take away the actual code writing part, but instead augment it and get rid of some of the tedium.
Given that there's a dozen agentic coding IDEs, I only use Cursor because of the few features they have like auto-identification of the next cursor location (I find myself hitting tab-tab-tab-tab a lot, it speeds up repetitive edits). Are there any other IDEs that implement these QOL features, including Void (given it touts itself specifically as a Cursor alternative)?
On the Apple Edu store, it's $499 for the 16/256 and $1079 for the 32/512
Another data point:
17.6 tokens/s on an M4 Max 40 core GPU
Data privacy-- some stuff, like all my personal notes I use with a RAG system, just don't need to be sent to some cloud provider to be data mined and/or have AI trained on them
Sellers are actually supposed to mark items as AI generated/assisted, where applicable: https://techcrunch.com/2024/07/09/etsy-new-seller-policy-202...
Whether they actually do this (and whether there's any incentive to do so), is obviously not a given
If Apple's "Private Cloud Compute" is used even on the latest devices for some tasks that are too computationally complex to be done on-device, then is there some reason (other than money) that they can't launch Apple Intelligence on all iOS 18 devices but use the cloud for all "AI" requests that would be "too slow" because of the older chips?
From MacRumors:
Notes can record and transcribe audio. When your recording is finished, Apple Intelligence automatically generates a summary. Recording and summaries coming to phone calls too.
So the functionality exists, maybe just not in the Voice Memos app?
CarPlay is rendered by the phone itself, so it's not strictly a function of how powerful the car infotainment is. You've been able to talk to Siri since the beginning of CarPlay so additional voice control is really just an accessibility thing
My current assumption is that this has to do with whatever "AI" Apple is planning to launch at WWDC. If they launched a new iPad with an M3 that wasn't able to sufficiently run on-device LLMs or whatever new models they are going to announce in a month, it would be a bad move. The iPhones in the fall will certainly run some new chip capable of on-device models, but the iPads (being announced in the Spring just before WWDC) are slightly inconveniently timed since they have to announce the hardware before the software.
You appear to be right, as far as I can tell the standard itself has not been ratified
https://www.ieee802.org/11/Reports/tgbe_update.htm
https://en.wikipedia.org/wiki/IEEE_802.11be
Development of the 802.11be amendment is ongoing, with an initial draft in March 2021, and a final version expected by early 2024
I was very excited by this, but I found that some of my dumber IoT devices would refuse to connect to the network if it used PPSK. If I connect them to a separate SSID I use for IoT devices with a basic WPA2 PSK, they work totally fine, but I didn't dig too much so it could also be user error
There's always the option to turn off your phone entirely, leave it at home, or simply buy a dumb phone incapable of using these new satellite features.
There's always going to be that one person blasting music in the wilderness, just because some people want to be disconnected doesn't mean we should eschew progress towards tech that can save peoples' lives.
Honestly no, it's mostly due to inexperience with operators and not really understanding what the "best" way to find operators is. I did also look at the Crunch Data one (I was having some issues setting that one up), but didn't even find Zalando during my search.
OperatorHub is currently the main resource I use, but GitHub stars aren't exposed in the search so I have been looking at the "Capability Level" chart and checking for Github popularity when I find one with the feature support I want.
I'm facing this exact same issue now when trying to find an operator for Redis. I am not sure if I am just missing out on the "right" option by limiting myself to Googling and Operator Hub and looking for the one with the most Github stars, so I am open to tips.
My holiday project was doing another pass at my Homelab Kubernetes cluster, part of which involved switching to a proper operator to manage Postgres. Coincidentally, I setup cloudnative-pg (https://github.com/cloudnative-pg/cloudnative-pg) yesterday.
The article is essentially an ad for an app that does just this
Everyone made fun of me for having one, but it was the perfect city car. Great turning radius & it fit basically anywhere. And the plastic body meant that if someone opened their door into you, it just bounced off instead of leaving a dent.
It's a shame they never made any significant improvements on the range. After degradation, the range on the i3 ends up pretty abysmal. The REx was a fun idea too and even let me take the car on multi-hundred mile road trips, without recharging, after coding the REx to one of the buttons in the car.
Since I had the same question as everyone else, it seems like it must be using just the transcript. When asking about one of those "8k HDR" showcase videos (with no speech), Bard responds with:
I'm sorry, but I'm unable to access this YouTube content. This is possible for a number of reasons, but the most common are: the content isn't a valid YouTube link, potentially unsafe content, or the content does not have a captions file that I can read.
They've had a Leica partnership for a while in the form of the Insta360 One R (and RS) 1-inch lenses, plus with their 1-inch 360 module. There's still some fundamental physics limitations with gathering light on a small sensor, but I enjoy using the 1-inch lens on my One RS when I can over the "normal" action lens.
As with many things in life: Hanlon’s razor
Why do you think that blue represents E2EE and not simply iMessage? If data isn't available and the iPhone sends an SMS, like you mentioned, the bubble is green, but this doesn't necessarily have anything to do with encryption. For example, the satellite SOS messages are represented as gray. It seems more like the color represents the transport.
It's actually both, at least in the US.
Tesla has the "Magic Dock", which is the built-in CCS adapter that unlocks when you use the app with a non-Tesla car. This is rolling out in a few places.
The other thing is that "NACS" (North American Charging Standard, the same physical plug that Teslas use but with the CCS protocol, with some minor backwards compatible changes for optional higher voltages) is being adopted by almost every EV brand, beginning ~2025. Further, these same brands have committed to providing or selling NACS-CCS adapters to existing cars so they can use Tesla chargers, even without buying a new model of car.
And it absolutely worked on me.
Prior to having a Steam Deck, my overall video game time was fairly low since it took time to boot the PC and start everything up. With the SD, it's much easier to grab it and get a small session in, and I've purchased a number of games (and will even buy games on Steam at a higher price than elsewhere) because of the Deck. It's the price of convenience, but well worth it in my opinion.