Taalas is developing this, but not for Frontier class models. I hope that if we can least get the easy 80% of work done on that sort of hardware, we can greatly reduce the demand for GPUs, HBM and energy to some extent.
HN user
tybit
They’re not using ciphertext in inference. They are encrypting agent responses on their servers if it’s going to a subagent on the client. The subagent will send it back to their servers for inference. Only their servers have the keys, so they can decrypt when running inference.
Yeah, if user -> org tenancy is stored in the same database without any similar defence in depth then a fresh API key after updating org would work around this. Would be a interesting topic for them to cover.
I think the same HMAC(pepper, user, org) as a validation column would work. Better yet, encryption with AAD on any tenancy data if you have a TPM available.
The OPs point wasn’t that OpenAis financial situation is comparable to Apples. It was that the likely cost of litigation is a drop in the ocean for OpenAi too despite their comparative lack of cash to burn. Legal disputes like this cost in the hundreds of millions over many years, so well below 1% OpenAis last single funding round in single year. If they got a tiny benefit from this (very gross) behaviour it may be finically well worth while. OpenAI may very well go under IMO, but this will barely be a straw on the camels back.
That’s what they do, but the TPM pepper is also needed for HMACing in their threat model. Otherwise the attacker just adds the victim’s user id to their hashing process too.
Zero data retention was an enterprise agreement that Anthropic and Amazon agreed with customers and delivered on. There’s no way AWS would trade in their reputation with enterprises just to soak up some slop.
I’ve seen whole teams at companies set up fail to provide these booleans-as-a-service well. There are whole companies like LaunchDarkly for them.
If you boil it down to this, you may as well boil down every service that exists to bits-as-a-service.
Turns out theres legitimate business value in these things, and complexity in delivering them.
At least Anthropic claims that they are profitable on a per model basis. But since both revenue and training costs are growing exponentially, and they need to pay for model N training today, and only get revenue for model N-1 today, the offset makes it look worse than it is.
Obviously that doesn’t help them turn a profit, until they can stop growing training costs exponentially.
So it’s really a race to see whether growth in revenue or training costs decelerates first.
Yes, but it’s a common misconception that impact is a bad thing.
The body, including bones, muscles, tendons and joints, adapt to stress. Many people do too little, not too much, as they get older.
There’s a limit to that recovery of course, and balancing it with stress is not always simple.
I also think fsync before acking writes is a better default. That aside, if you were to choose async for batching writes, their default value surprises me. 2 minutes seems like an eternity. Would you not get very good batching for throughout even at something like 2 seconds too? Still not safe, but safer.
Yes, the greenest browser is one that doesn’t use AI. They aren’t claiming they’ve built that though, just the greenest AI.
It’s interesting that the author chose to use SHA256 hashing for the CPU intensive workload. Given they run on hardware acceleration using AES NI, I wonder how generally applicable it is. Still interesting either way though, especially since there were reports of earlier Graviton (pre v3) instances having mediocre AES NI performance.
Yeah, investing in the top companies leads to higher returns for most periods when looking short term.
Over longer periods, the top companies by market cap tend to change though. https://www.investmentnews.com/equities/only-one-of-the-worl...
So if you want to invest in the top companies, you either need to think they won’t change anymore, or you need to find when to buy and sell. Index funds solve this problem for you, albeit with slightly lower returns in the short term.
At successful tech companies, engineering work is valued in proportion to how much money it makes the company
If you look at what it actually takes to get promoted at most tech companies I’d say this isn’t generally true at many big tech companies.
Being on a very lucrative part of the product may not get you as much “impact” on your promotion packet as if you are working on a platform/infra touching the whole org. Even if that platform isn’t generating the company much money even indirectly.
Runtimes with garbage collectors typically optimize for allocation, not deletion.
JOOQ handled this nicely.
https://www.jooq.org/doc/latest/manual/sql-execution/crud-wi...
CockroachDB is presumably what they’re referring to in:
For Postgres-compatible NewSQL, we would’ve had one of the largest single-cluster footprints for cloud-managed distributed Postgres. We didn’t want to bear the burden of being the first customer to hit certain scaling issues
I find their claim a bit hard to believe.
While they don’t specify it sounds like they don’t even require 2FA to access their systems?
Tangential to the authors point, but it’s funny to note many new SQL databases(e.g CockroachDB, TiDB, MyRocks) are written on top of RocksDB, a “NoSQL” key value store.
When it comes to cache invalidation worse performance isn’t the primary concern in most cases, correctness is.
It’s usually a safe bet that the state will preference companies over both employees and taxation unfortunately.
I think this architecture would be really powerful paired with the actor model to shard databases to nodes.
At big tech companies I’ve seen and heard about, the answer is crypto shredding. Encrypt all PII at rest with a per user data key. GDPR deletion requests can then delete the data key. This isn’t perfect, but it’s a step in the right direction IMO. Unfortunately I don’t see it being feasible for a typical company anytime soon.
For anyone else expecting this to be a paper given the domain name, it’s not. It’s a non technical interview with a couple of the original papers authors. Not bad, just not as exciting as I imagine a paper detailing what they’ve learnt from a distributed systems perspective etc operating Dynamo then DynamoDB for so long now.
This would be an interesting article to flesh out. I.e is there evidence that MySQL is more reliable in those ways?
I always prefer reliability over features even though I’m a product engineer so if he’s right it’d be good to know. Either way, I’m stuck with the MySQL that the infrastructure engineers at work have provided us with.
I think there’s a good argument that async is decent for performance critical languages, e.g C++ and Rust, and for languages looking to model effects, e.g Haskell and arguably Rust. I don’t see a good reason for it in mainstream languages like Java, JavaScript and C#.
I think Java’s approach with Loom is going to be a big win over C# there, as someone that just wants to get stuff done and is a fan of both.
I realise that the Twitter is using Mesos, but for those of us on Kubernetes does guaranteed QoS solve this? https://kubernetes.io/docs/tasks/configure-pod-container/qua...
It was very clear from their post that they were criticising STS from the perspective of an engineer in AWS within a different team.
Anyone else surprised that they managed to get the IOT ticker? I would of thought that would have been taken already.
I can’t even imagine the senior product and data people I’ve listened to deciding to change this based on one artists experience.