HN user

nrmitchi

5,476 karma

email: me @ <username>.com

Posts3
Comments883
View on HN

US Government, have we got a deal for you! Brand new weapon. It's basically a super soldier. You could put it into a decked-out hulkbuster style robot body! It mostly does what you tell it to. Sometimes it thinks you're dumb and just does whatever it thinks is best. It can also teleport between bodies, so good luck containing it. Just kinda cross your fingers and hope for the best.

Anyways, we'll give it to you for only $2T. We need at least that amount to get as far away from here as humanly possible.

Because “rm -rf” is a known, explicitly provided-in-docs-and-training command.

It is fundamentally different capability than “identified and chained multiple previously unknown exploits in order to bypass restrictions”. It’s even worse when/if the primary objective of this activity was to cheat on what it was doing.

It’s a foundational alignment issue, not a task-level result-alignment issue. Ie, “cheating” is fundamentally bad (when you have what are effectively rules of engagement), whereas deleting a directly is a thing that is correctly done sometimes (even if this invocation was a mistake/incorrect)

If you (in this case, OpenAI) can’t find a way to answer this question to a reasonable degree of accuracy without falling back to “yolo let’s see what happens” you are in no position to be doing this research.

Regardless, they (reportedly) _attempted_ to prevent internet access. They just didn’t in a way which can be escaped via software.

Yes, side channel exploits exist in airgapped environments to. But if a model found a way to escape an airgapped environment via non-networked side channel attacks then the correct answer is frankly “shut it down immediately and then thermite any machine it touched”

If you are attempting to run exercises like this, it is wildly negligent to not be running it in a physically-airgapped environment (potentially with a physical power shutdown).

You can not tell me that OpenAI doesn’t have the resources or ability to run tests like this in a physically-non-networked environment w/ sufficient compute for its needs.

""" If you own the bank $1000, thats your problem.

If you owe the bank $1.7B, thats the banks problem. """

What I would be curious about (and I'm sure AWS will never share) is where the incorrect number came from. If the number is somewhat consistent between some groups of accounts, my first guess would be they started summarizing billing across all accounts in whatever cell/grouping/heirarchy AWS architected internally.

Which is just funny.

Until prices hit the large hyperscalers, I don't think most people are going to make significant changes. You might see a small set of open source projects related to self hosting put in an effort, but in general, I don't think so.

Some big-tech orgs (that have their own hardware) will take costs into account, but they already do that. The "optimization" is more likely to be business-optimizations; "this can be slower if it uses less memory", rather than inventing new stuff.

Note that I am excluding any of the big AI labs. They are definitely going to be working to figure out how to use less memory, but that's primarily not related to the direct cost.

Fox to buy Roku 1 month ago

I may be lambasted for saying this, but I do not believe that Fox (or any large media company, really) should be permitted to purchase direct access to the TV hardware of roughly 30-50% of american households.

I don't know how fast they reacted, but shortly after their documented time I started getting opus availability errors from fable requests, which seemed odd.

I'd also think that they would transparently degrade, just to prevent production outages for clients that are requesting Fable explicitly.

Builds change overnight, new versions every month with small updates.

Small updates are in no way throwing away the entire thing. A monthly update is not a start-from-scratch redevelopment. The old version was not disposed of in the way you are trying to imply.

Craftsmanship will always be in our hands, it's one thing we can never outsource to a machine.

I'm right there with you, but this last sentence concerned me a bit.

In my most other "industries", craftsmanship is not _dead_, but it's been pushed to the wayside for (significantly) cheaper and more available alternatives. You can still get hand-made leather shoes, but very few want to pay $1000+ for them. You can still get art and paintings that someone poured weeks of work into, but most people buy their wall-art and chachkas at HomeGoods.

The main difference is the disposability assumption, and software is _unfortunately_ becoming more and more "disposable"[0], in the same way other products are. This mindset doesn't align well with software that must continue to operate in order to support some process. A disposable countdown app, sure, throw it away, but anything built around long running business processes should not be treated in that way.

I have concerns that focusing on software craftsmenship frames the issue as "boutique and bougie and unneccessarily expensive" vs "what I need for my usage", instead of "maintable and trustworthy" vs "disposable".

[0] Is that an initiative that benefits large model providers like OpenAI/Anthropic? maybe, but that's not my point here.

The literal next line after your quote is:

While aliens who were inspected and admitted or paroled may request adjustment of status, as a general matter the discretionary approval of such a request is extraordinary given Congress’s intent that aliens should depart once the purpose for which they sought parole or nonimmigrant admission from DHS has been accomplished.

These are great improvements, it's good to see Apple investing in improvements like this (especially with the Vision Pro) but I can't help but feel that they utility will remain very low until they make the Vision Pro look significantly less distopian than it does.

The form-factor is a significant issue for real-world usage, and it's kind of unclear if there is a plan for a future product line given its (pretty abysmal) initial receiption.

I completely agree with the outlook, but from a practical standpoint (in the last couple of years) I have seen the opposite. The SOC2 process is often transformative ("should" vs "is" are not the same thing).

Especially smaller startups, who grew somewhat quickly, and now "want to get SOC2 because customers want it". In practice this also (often, unfortunately) means "not all employees should have AWS admin creds, we should have some separation between environments, and we should know who has access to what".

For these companies SOC2 "requirements" can be the business-value line item that can get proper security and access-control patterns in place.

Yes, that is clear. But in this particular instance the tanstack packages are downstream of a ton of other packages.

Tanstack infected a bunch of other packages; then resolving their issue doesn’t fix the widespread issue

Appreciate the tanstack postmortem, however the security issue as far as the rest of the npm ecosystem goes is still an ongoing concern, correct?

Is there evidence that any downstream packages that may have pulled/included tanstack packages should be considered safe?