HN user

avianlyric

7,334 karma
Posts7
Comments1,859
View on HN

Really, now? Yeah, they can ask, but do data center operators just go "oh grid needs our power! Let's switch to backup!"

Of course they do. The grid is going to load shed regardless, better to disconnect in an orderly fashion (especially as the grid will pay you to do so) than wait for your local substation to disconnect you.

That seems unlikely. Why would a data center (or any large consumer of electrical power) be forced into a backup/contingency plan, when no crisis exists at all?

Because they get paid to do it. This is pretty basic demand side management that’s been the bread and butter of grid management for decades now. Grids will offer discounts to large loads that allow themselves to be disconnected at short notice, and will then pay those loads for the periods they’re disconnected. There’s usually multiple tiers of discount and payment depending on how fast a load can disconnect, and how often the grid expects to call upon that facility.

This is where harness, and the fact that a machine can be endlessly prompted to try again comes in.

Even if an LLM starts by pursuing things that follow human bias, continuous failures and re-prompting to try something different will eventually force it to consider things outside of what ever biases it has.

You can do the same thing to a human. But most people would consider it unethical to lock someone in a box and force them to keep trying to solve the same problem over and over again until they figure it out.

I think the point is that an “open source project” is more than just the code and license, it’s also the community that builds and maintains it.

A fast rewrite of Bun in Rust has effectively alienated most of the people in the Bun community, so in that sense the “open source project” has died. It’s no longer a community project, it’s just become a personal project again.

There’s a big difference between processing multiple streams, and processing multiple streams simultaneously.

You can achieve the former, without the latter, by doing time slicing. Spending a small amount of time processing stream A, then dropping that and processing stream B for a moment, then swapping back. Just like how a single core CPU can process multiple threads.

Proving the brain is continuously processing and encoding multiple streams simultaneously is an interesting finding that helps us better understand how our brains handle multitasking. That’s absolutely something worth studying and understanding, even if the headline discovery “feels” obvious. It the precise mechanism that’s interesting, not the effect the mechanism produces.

You can only get grid connections if the grid thinks they can support your demand. That’s why grid connections take so long, they’re not just plugging in a massive cable, they’re reinforcing other parts of the grid to make sure your massive cable won’t cause a blackout if use all the power it can provide.

Building new power plants immediately isn’t really required to serve data centre, they’re huge amounts of spare capacity in most grids that only exists for extreme situations. The hard part is making sure any existing capacity can actually be delivered to the new data centre. New power plants will get built to rebuild spare capacity over time, or more likely, solar and wind will get built to reinforce that spare capacity over time.

Datacentres are useful grid loads because they all have backup power on site capable of powering the entire site as an island. Which means grid operators can ask them to disconnect quickly if they need to extra power for short periods of time to handle power spikes, or local grid faults. So it’s possible to connect quite a lot of extra data centre type load before generation capacity becomes a serious issue.

This argument would hold more weight if Anthropic and OpenAI main customers weren’t massive trillion dollar companies with legal teams capable of burying just about anyone, anywhere, for even the mildest contract violation. Something that OpenAI is getting some close up experience with at the moment.

It was also disruptive because it was open weight, meaning anyone and their dog could theoretically compete with the frontier labs for their inference revenue.

The frontier labs need to recoup a huge amount of cash to cover their model development costs, and justify their valuations. That’s plausible when they’re only ones capable of selling inference on these models, it a lot less plausible when models themselves become cheap commodities, and you’re just competing on your ability to provide compute. Anthropic and OpenAI can’t compete with people like AWS on that front.

Most data centres connect to the grid, they don’t connect to a single power station, except in scenarios where there’s a uniquely low cost power supply nearby, like a small Hydro Plant.

Utility scale power stations have outputs measured in GWs. Data centres are measured in MWs, although people are trying to build GW scale data centres at the moment. But even then a data centre will want a proper grid connection, otherwise they have a massive single point of failure in the form of the directly connected power station.

It’s also very unlikely that purpose built power station is capable of offering cheaper than grid power anyway, except in the very special situations like Hydro. So if you’re gonna build a datacentre, you will want a proper grid connection capable of providing all you needs. Even if you’re running on dirty gas turbines in car park initially while waiting the grid hardening happen. In the long term, that grid connection is always going to be the cheapest, most reliable source of power, ignoring it completely would be foolish.

I think I covered all that with

the value of human life drops to value a person can produce defending their society.

OP was talking about how US munitions software needs to go through a 6 week long integration test before it can be deployed. I seriously doubt that kind of testing would last very long if the US was engaged in peer level war, and their enemy had found a flew in their munitions guidance systems that made the munitions useless.

At the end of the day, when at war time isn’t just money, it’s also lives. When you have people dying on the frontlines, the risk of equipment failures from lack of testing will be substantially smaller that loss of lives from ineffective equipment, that can only be improved every other month.

I wonder if anyone is going to learn a lesson about overregulation.

Seems unlikely. Regulation and Health & Safety are both societal luxuries, which only happen once societies are stable and prosperous enough to start valuing human life beyond its ability to perform labour.

The moment the bombs start dropping, the time for luxuries also stops, and the value of human life drops to value a person can produce defending their society. There isn’t the money or resources for anything more than that.

The US (most developed democracies) places an extremely high value on the lives of soldiers, because dead soldiers in foreign wars does terrible things to politicians in power. Paying 1000X more for the same tech as Ukraine to minimise the number of service members killed using it, is a pretty small price to pay.

OpenPrinter 17 days ago

You can't make a commercial competitor of this printer using their design

Which means it's not open source. Open source means you have the right to distribute work however you want, including commercially, provided you also provide the source under the same license terms as the original.

The second you slap a non-commercial limitation on there, it ceases to be open source.

So while pedestrian deaths are climbing, the overall deaths are still trending downward, and I shouldn't have to defend that the overall count is more important than the pedestrian subset.

A better question is why is the US the only developed country that’s seen pedestrian deaths increase over the past 10-15 years. Every other developed country has seen both occupant and pedestrian deaths decrease over the same time period, and has seen a larger combined drop in deaths than the US. And to be clear, I’m talking about deaths per mile driven, not absolute counts, so the size of the US is already factored into the numbers.

but think how many people drive in America, and how useful it is

You should tell that to the families of the 8000 that are killed each year. I’m sure it’ll help them accept the accept the loss of their loved ones.

The US is the only developed country that has seen a steady increase in the number of pedestrian deaths per 1000km driven over the past 10-15 years. And 10-15 years ago it was one of the worst performing developed countries for pedestrian deaths. Every other developed country has seen decreases in pedestrian deaths over the same time period, which means the US is an extreme outlier when it comes to pedestrian safety, or lack there of.

disguise the underlying dogma, which serves as an unsupported conclusion: humans are assumed to be completely entirely unique in every way whatsoever

Is that the argument the paper is making? In my reading they seem to primarily be making the point that assigning anthropomorphic concepts to LLM is dangerously misleading, and more importantly, not needed to properly study and evaluate LLMs.

I don’t think you have to make the assumption that humans are unique for that argument to hold up. I would argue that really it’s a comment on how loose and poorly defined all anthropomorphic attributes are. At the end of the day we have to make the assumption that other humans feel and experience broadly the same mental activity as each other, because we’ll never directly experience anyone else conscience, we can only experience our own.

We can barely link our own mental experiences to concrete empirical measurements. The vast majority of the measurements we make are entirely self-reported, and we simply assume strong correlation between self-reported measurements and the individuals actual experiences. We also have to assume that somehow all of our self-reported measurements are “calibrated” to some reasonable degree. Even measuring anthropomorphic properties in humans is pretty fuzzy and inaccurate, the only reason accept such poor data is because it’s the best we’ve got, and there enough signal in there for us to develop useful tools like talking therapy, physiological profiles, mental health scores etc which have some level of predictive and healing power when applied to _humans_.

It’s honestly amazing that what we have works for measuring and predicting humans, and we only know that works through decades of empirical measurement and study. But to then try and directly apply that fuzzy mess to a completely different system, and just assume the same level of predictive power, strikes me as kinda crazy. It requires huge assumptions, which effectively can never be tested (because even the human mind is a total mystery to us), to be made, and if we can study these systems without making those assumptions, then why make the assumptions at all?

Isn’t the entire paper is trying to point out that the second you ask the question “Do LLM have <anthropomorphic property X>”, you have to assume that they do, even before you make any assessment?

Just because the person asking the question isn’t aware of they’re implicitly making that assumption, doesn’t change the fact that a logical assumption has been made. It just makes the questioner ignorant of the assumptions they’re making.

Personally don’t totally understand the argument being made in the paper. But I can understand the idea that I can ask a question, without properly understanding the assumptions I’m making when asking the questions. Indeed I can also understand that I might not even notice the assumptions I’ve made with my question, and why that would make my entire exploration and conclusion invalid, _after_ doing the investigation. Logical fallacies can be really difficult to spot and understand.

SoftBank hold huge positions in companies like OpenAI, funded using debt. The interest on those loans is killing them, and until OpenAI actually IPOs and SoftBank can sell their stake, they have to pay that interest using cash from somewhere else.

There are definitely technical issues with refinery capacity, but I don’t think they’re insurmountable

If they were cheap or easy to solve, don’t you think US refineries would have already converted to support domestic crude? Domestic crude is cheaper than imported crude, the only reason to import is because it so expensive to convert a refinery. My, admittedly very limited, understanding is that you generally don’t convert refineries, it’s cheaper and easier to just build a new one that targets a new type of crude. Building refineries takes a few years, they’re not something you throw together in a few months when oil markets go crazy.

Pricing on SToA models probably won’t fall, there’s no reason for the frontier labs to lower their prices.

But we’re seeing lots of open weight models that are either pretty close to SToA, or more importantly, perfectly capable of doing all the low level token insensitive grunt work when writing code. Pairing them with SToA models for long horizon task management, and you’ve got a very cost effective system.

The frontier labs have put little effort into cost efficient inference, they don’t need to, but folks like DeepSeek clearly are, and have achieved some impressive cost improvements. Given DeepSeeks models give you 70% of the capabilities for 30% of the cost, expect people to start moving lots of workloads to providers that provide cheap inference for open models, and huge competition to appear to provide that cheap inference. It’s truly commodity LLM inference.

In turn expect more companies to focus on building inferences efficient models, because someone that can build a model that provides 70% of SToA capabilities for 10% of the token cost, immediately eats up huge amounts of the available inference market.

Another factor in all this, is it’s becoming increasingly clear that building custom agents/workflows for LLM to operate in, is required to get the best out of these models. That means people are implicitly building the infra needed to use multiple model types and evaluate workflow performance end-to-end. Which in turn means they have everything they need to plugin in future, cheaper, inference providers and quickly evaluate if they can change their model provider.

Pace of data creation ignores the fact that the majority of the big gains in LLM “intelligence” has come from scraping and feeding in the huge amount of public data that already exists.

Unless we’re producing data on the order of an entire new internet every couple of years, then it’s hard to see how LLMs can achieve further huge leaps in capability compared to training on effectively 0% of the internet vs 100% of the internet.

Contractual obligation, external third party audits, and above all, AWS’s reputation.

AWS isn’t going to risk their reputation, and thus huge chunks of their business, just so a few AI labs can get some extra training data. That’s an insane risk with zero upside for AWS. AWS knows full well they will make insane quantities of cash without breaking legal contracts with companies who pay them billions each year for infra.

You’re making the assumption that customer product prototypes are the only prototypes produced by 3D printers.

There’s plenty of other more valuable things that are prototyped using 3D printers, such as high end commercial machines, or components that go into those machines.

I suspect that getting hold of STLs from US defence manufacturers would be extremely valuable. Why bother trying to capture a copy of your enemies technology, when they’ll happy just send you all the prototype STLs. Even if it’s not defence, don’t you think access to prototype components from EUV machines from ASML would be crazy valuable to Chinese companies trying to close the gap between Chinese and Western chip fabrication technologies?

Proper utility scale gas generators come with proper utility scale pollution controls to make sure nasties like fine particulate and NO is filtered or properly reduced into some much less harmful to human health.

CO2 is bad for us long term. But there are plenty of other nasty combustion products that are extremely bad for humans in the short term. Which is why we have pollution and air quality regulations.

Portable generators don’t meet any of the stronger requirements that utility scale systems have to meet, because it’s assumed they’re only operated in small numbers for short periods of time. They’re not designed to safe to operate in large numbers over long periods of time in the same place. For that you need proper pollution controls

They have done. The Three Mile Island accident happened when it was being operated by Navy vets [1]. Simple training isn’t enough.

During the investigation of the accident the Admiral that built and ran the Navy nuclear program was asked how the Navy had managed to operate accident free, and what others could learn. This was his response:

Over the years, many people have asked me how I run the Naval Reactors Program, so that they might find some benefit for their own work. I am always chagrined at the tendency of people to expect that I have a simple, easy gimmick that makes my program function. Any successful program functions as an integrated whole of many factors. Trying to select one aspect as the key one will not work. Each element depends on all the others.

So recreating that accident free operating environment requires a lot more than just training. It would require wholesale adoption of the Navy’s approach across the entire industry. Which probably doesn’t scale very well. Not to mention the Navy operates much smaller nuclear reactors compared to utility scale reactors, and has extremely easy access to lots of cooling water, which probably gives them a little more wiggle room when dealing unexpected reactor behaviour.

[1] https://jackdevanney.substack.com/p/tmi-lessons-what-was-lea...

It’s always been true. If you want to build a limited distribution app Apple has mechanisms for private distribution which is used by companies for internal apps etc

They don’t want the App Store filled with app that can’t be used by the vast majority of people that might see and download it.

Because it’s a pain in the arse to design, manufacture and build a specialist device just for use in your stores.

I’m sure Apple could do everything that box does and more. But why bother designing, building and manufacturing your own specialist device when someone else already sells a perfectly good tool that does the job.

Don’t forget this is for use in a retail store by people who will have been given 5mins training on how to use the device. You want something that just requires a person to plug two phones in and hit a big “go” button. And it needs to work 99% of the time with zero messing around.

You can implement either approach on iOS as well.

But if you have strong end-to-end encryption for messages, then you don’t have to care about the transport anymore, you assume they’re all compromised. At that point you might as well use the push notification system as your transport, given both OSs allow applications to intercept the push notification locally and decrypt it before it’s displayed to the user.