HN user

impure-aqua

82 karma
Posts0
Comments33
View on HN
No posts found.

Commoditizarion is a process, not a binary state.

If I run an oil refinery, my fractional distillation system needs to be reworked depending on the exact mixture of crude I'm taking as input. So there are still switching costs even in the textbook example of a commodity.

Crucially though the exact upstream I use has minimal impact on the downstream. Closer equivalents, say, another barrel of WTI grade crude from a nearby regional supplier, require extremely minimal reworking. Oil from a different region, that might require more retooling, so I might be willing to sustain a longer shock in market conditions before making that switch. The important thing is the output broadly remains the same, but even this is broad, e.g. a different mix of inputs yields different ratios of output.

LLMs are quite similar no? Maybe switching to another SOTA model has minimal reworking, as you can delegate at the same level of abstraction to the model, whereas switching to a slightly-behind-frontier model you need to do more hand holding. Switching costs being nonzero does not preclude them from being broadly an interchangeable input in the production process. Any non-frontier use (99% of SWE) will be delivered on pretty much the exact same timeline irrespective of which model was used, so my requisition process looks more like buying barrels of oil than e.g. shopping for a new phone.

Phosh 0.56.0 17 days ago

of course everyone, it's definitely the 200MB usage surplus that makes your phone feel heavy.

It actually is though? My Pixel 9a has a perfectly serviceable CPU but I am often frustrated by its 8GB of RAM. Switching to the Revolut app to generate a disposable card number consistently evicts my browser tab from RAM. ~100% of the time this happens and I get extremely frustrated by it losing my state in the checkout flow.

I don't know if that tracks, senior leadership was heavily influenced towards implementing the one child policy by the works of Song Jian, who came from a rocketry background and presented a model whereby the population would grow to an unsustainable level unless corrective control was applied.

I think it is unlikely philosophers would have suggested to treat population growth like tuning a PID controller.

WhatsApp performs dynamic code loading from memory, GrapheneOS detects it when you open the app, and blocking this causes the app to crash during startup. So we know that static analysis of the APK is not giving us the whole picture of what actually executes.

This DCL could be fetching some forward_to_NSA() function from a server and registering it to be called on every outgoing message. It would be trivial to hide in tcpdumps, best approach would be tracing with Frida and looking at syscalls to attempt to isolate what is actually being loaded, but it is also trivial for apps to detect they are being debugged and conditionally avoid loading the incriminating code in this instance. This code would only run in environments where the interested parties are sure there is no chance of detection, which is enough of the endpoints that even if you personally can set off the anti-tracing conditions without falling foul of whatever attestation Meta likely have going on, everyone you text will be participating unknowingly in the dragnet anyway.

On my Pixel 9a (also on GrapheneOS) the biggest limitation is it can't be set to higher than 1080p, and the upscaling algorithm with my 4K display (not sure where in the chain that happens, monitor or phone) was quite terrible to the point of text legibility being a concern.

The usage experience otherwise is quite good, it's perhaps my preferred way to sync data to and from my phone, I have it all stored on a NAS so I connect to my Type-C display (which has keyboard, mouse, and ethernet connected to its switch), fire up a terminal, type in my rsync commands, and my pictures & music are synced ~instantly at LAN speeds.

That is true of press, weld, and paint stages, which gives you a chassis and nothing else. It is absolutely not lights out for "final assembly" which despite the name is how massive amounts of the car comes together.

Robots are great at the bulk movement required for sticking sheet metal into huge stamps as well as repeatably welding the output of these stamps together. Early paint stages happens by dipping this whole chassis and later obviously benefits highly from environmental control (paint section is usually certain staff only to enter.)

But with this big painted chassis you still need to mount the engine/transmission, the brake and suspension assembly needs installing, lots of connectors need plugging in for ABS- and supporting all the connectors that will need plugging in is a lot of cabling that needs routing around this chassis. These tasks are very difficult for robots to do, so they tend to be people with mechanical assists, e.g. special hoisting system that takes the weight of engine/trans while the operators (usually two on a stage like this, this all happens on a rolling assembly line) drag it into place, and do the bolting.

Trim line is also huge, insert all these floppy roof liners, install the squishy plastic dashboard, the seats, carpets, door plastic trim, plug in all your speakers and infotainment stuff, again the output of the automated stages is literally the shell of a car, and robots are extremely bad at doing precise clipping together of soft touch plastics or connection of tiny cables. Windshield install happens here too, again these things are mechanically assisted for worker ergonomics but far from automated.

Each of these subassemblies also can be very complex and require lots of manual work too but that usually happens at OEM factories not at the assembly factory. Automation in these staffed areas mostly is the AGVs which follow lines on the floor to automatically deliver kanban boxes which are QR tagged (the origin of the QR code, fun fact) to ensure JIT delivery of the parts needed for each pitch.

It is far from lights out even in the most modern assembly plant and I think it will be a long time until that is true. The amount of poka-yoking that goes into things like connector design so there is an audible "click" when something is properly inserted for example- making a robot able to perform that task at anywhere near the quality of even a young child will take vast amounts of advancement in artificial intelligence and sensing. These are not particularly skilled jobs but the robotics skill required is an order of magnitude more than we can accomplish with today's technology.

I also am a programmer, and I care about all of these things on my laptop. I used my trackpad to click reply on this webpage, that's not a rare thing!

If you ever have a meeting where multiple people huddle around a laptop, that uses speakers, webcam, and microphone, and the MacBook does so much better in that scenario. We have interrupted meetings to swap from a Framework 16 (old CTO's laptop) to my MacBook Pro because participants simply couldn't hear those of us slightly further away from the laptop!

Zonal dimming is an advantage whenever you have black areas on the screen, and good fan tuning is an advantage if you want to compile some changes during a meeting without thinking "this task will turn my laptop into a jet engine and distract everyone else".

If you don't care about these things, then you can find way cheaper devices than the Framework that are cost competitive on core specs. Let's get some Framework pricing as a datum, Framework will sell me the AI 350 and 2.8K display for $1939CAD, it has no RAM, no SSD, no charger, no ports... if I add 16GB RAM, 512GB SSD, charger, and 2xUSB-C, 1xHDMI, 1xUSB-A, we're looking at $2403CAD.

If I don't care about the less measurable components, why would I not buy something like this $400USD (~$550CAD) laptop [1] another poster in this thread found which also has an AI 350, 16GB RAM, and a 512GB SSD? I can buy four of these laptops for the price of the Framework and still have some cash left over! If I need more RAM I'm sure I can find a similarly cheapo laptop with a SODIMM by actually googling myself.

I think the reality is both you and I do care about these other parts, just maybe with a different minimum acceptable quality. But even inside PC land Framework is not competitive. Higher-end X1 Carbons have haptic trackpads at the same price point as Framework is offering diving boards. Across the market there are OLEDs for less money than Framework is charging for LCDs.

Personally, I don't care about trackpad alone so much, merely that the pointing device situation be acceptable. When programming, I type a lot and then do a few small mouse actions (e.g. expand some segment on a docs webpage, or mouse around some GUI to test the feature I have been building out). With a haptic trackpad, I can move my thumb from the spacebar to the top of the trackpad which is just below it and do my mouse actions without significant hand movement. This is not possible with a diving board design as the top of the trackpad is not clickable. A pointing stick is absolutely an acceptable solution to this problem, but Framework also does not offer those, again despite price-competitive offerings from, say, Lenovo offering it.

Let's briefly look at Lenovo's website. I can spec out a ThinkPad P14s Gen 6 here in Canada from Lenovo's website [2] with a 120Hz OLED screen, trackpoint, Ryzen AI 350, 1x16GB SODIMM and 512GB NVMe for $1529CAD, that's a fully working computer for less than the barebones Framework, with a better display and pointing device situation! I can use the empty second SODIMM port with a single 48GB stick and get 64GB, and stick the NVMe in an external enclosure to use as an external SSD, and deck it out with whatever market-rate drive and RAM I can get.

The Framework is broadly uncompetitive even if you won't consider a MacBook.

[1] https://slickdeals.net/f/18984394-hp-omnibook-5-16-fhd-ips-r...

[2] https://www.lenovo.com/ca/en/configurator/cto/index.html?bun...

There are all sorts of other things that don't show up on a spec sheet so easily that Framework isn't competitive on.

It has a diving board trackpad, significantly worse speakers, no zonal dimming on the display (comparing to MacBook Pro, which higher end specs of the Framework cost as much as), general poor body rigidity, an aggressive fan curve that ramps up audibly on short loads (the Air doesn't even have a fan and the Pro can handle a couple mins of all-core 100% load without becoming audible), etc etc.

As much as I dislike Apple's business practices it's undeniable that other vendors are generally selling significantly cheaper feeling devices at the same price point. These are not niche things, you feel the cheapness on the Framework with every touchpad click, short bursty CPU task, HDR video, audio playback, heck even picking it up off the desk.

This is not really anything new, back in AIM and SMS messaging days, people would type "wuu2" or "whats up" to a friend, but to express the same idea in an email, you would probably be sending some variant of "What are you up to?"

There is massively different subtext between the two. Autocapitalization and autocorrect represents a limit on the subtextual bandwidth you can communicate along with a message. Restrictions on subtextual bandwith are not ideal when your generation relies on text-based communication for evermore intimate interactions - that "whats up" message might be the start of you asking someone out on a date, I don't want it formatted the same way as a message I would send my boss.

When I drive cars with old headlights, they are clearly inferior to the point of feeling nearly dangerous in some situations. I would also not call modern lights less reliable, although I am sure it is more expensive to repair modern lighting technology.

In a North American city where there is overhead lighting and the streets are a mile wide, sure, I could probably turn the lights off even and be totally fine.

In the middle of the British countryside on a single-track road that has hedges on either side, not enough space for cars in the oncoming direction to pass me, a 60mph speed limit, during a rainstorm? I want the nice lights.

I feel like this problem is better in the UK than in North America.

For starters, there is higher market penetration for better headlight technology, particularly ADB (adaptive driving beam). North American road safety regulations have made it very difficult to get this technology into cars, whereas in Europe it is reasonably widespread. Even rental cars I have had in the UK have this technology- most recently a Mazda3 which had a very good implementation of it, I could drive through the countryside with high-beams on constantly, and you could see the car quickly dim the beam facing towards oncoming traffic if any came around a bend. These are not high-end cars; I have rented cars with a manual transmission and cloth seats yet better headlights than the fanciest S-class in North America.

There is also less variation in vehicle size, and better emphasis on road safety testing. In Canada I often encounter lifted pickup trucks, which changes the alignment of factory lighting, not to mention the lights on these are often aftermarket anyway and usually installed without any thought for alignment. British pickup trucks are rarer, smaller, and would fail their yearly MOT for having improper headlamp aim.

There are all manner of health conditions that can occur that have little to do with your own healthy living, and don't incapacitate you or make life not worth living, but will cripple you financially.

You might end up with Crohn's or all manner of autoimmune conditions where patented biologics easily costs north of $100k US a year just in medication, but your quality of life if you find a medication that works is not particularly degraded from the average person.

CrossFit will not prevent you from getting into that situation, and I think it would be a vast overreaction to commit suicide in response to such a diagnosis.

The walled gardens got a lot more appealing.

When we moved to Canada from the UK in 2010 there was no real way to access BBC content in a timely manner. My dad learned how to use a VPN and Handbrake to rip BBC iPlayer content and encode it for use on an Apple TV.

You had to do this if you wanted to access the content. The market did not provide any alternative.

Nowadays BBC have a BritBox subscription service. As someone in this middle space, my dad promptly bought a subscription and probably has never fired up Handbrake since.

I don't see what advantage any company gets from choosing to build products that enable personal data ownership. I say this as someone working on a venture with these sorts of design aims, it feels like pushing a boulder uphill often.

The business model of cloud service providers makes a lot of sense- we have a system which stores and operates on your data, you pay some rental fee for us to store it and operate on it, easy peasy. The cost is related to both the utility of the operations the operator performs (to both the operator and the user) and the amount of data the user stores.

Fundamentally this is how everything from Dropbox to Facebook is governed- Dropbox does not devise much utility per GB and users store a lot, so you rent per GB, but at Facebook, they don't store lots of your stuff, and on the data side maybe you don't get much value from it as it's a cesspit, but the data is valuable to Facebook to sell ads, etc, so they can provide the service for free.

Importantly, you don't need to improve the product to continue extracting this rent, because the product you are selling is not Dropbox v4, Facebook v2.3, rather you are selling ongoing access to the rental.

As soon as you introduce even simply a federated system where a few corporate operators are involved, it becomes very hard to justify extracting rent there as the network designer, as the operators are taking on the cost of actually storing the data. You have to really be iterating on the core product to use a SaaS business model here. Some things simply don't need a v4, does Dropbox really need that much iteration?

Meanwhile as the system designer, life has become a lot more complex for you. Suddenly you cannot push unilateral sweeping changes to APIs, you need to version things in a way that is compatible between, say, one university updating their system but not the other. Since your users are a few large operators rather than millions of individuals, you lose the network effect advantage of being able to screw over a few users for the "greater good", since if you irritate one corporate client, you lose a lot of your install base. Why would you voluntarily choose this harder path as a company?

Things get even worse as you increase the level of decentralization. The reality is users expect the polished experience that the rental companies can give you; they want their data always accessible so that their friend can see the pic they shared without needing to keep their own computers running, they want the "like counter" to go up without their personal node subscribing to messages from other nodes, etc. The only users that will accept a worse experience are people who have are motivated by their philosophy re: personal data ownership, and this crowd will want a FOSS solution, so you can say goodbye to charging them for Dropbox v4, they are simply not interested if you're not giving them the source code for free. (I suspect this is where the author sits, but fundamentally I don't think it will get mass appeal, most people simply do not care about data ownership above something that "just works".)

So now you are dealing with problems like dynamic generation of redundant data and fault- and Byzantine-tolerant consensus algorithms so that your system can maintain function even when the user turns their computer off, and you have to deal with wrapped-key cryptography so that the redundant data can be split across all these user nodes without you worrying that an unauthorized user can read it, and then you have issues like how do you deal with nodes that are too slow to process updates (perhaps some user data needs to be stored in this conflict-free replicated datatype you devise), and eventually you go through all of this to... create a system that is less monetizable than the rental model, because you can't extract that rent for ongoing data storage, and we know users are not interested in actually paying for software.

I don't think Framework will be able to compete on efficiency with their design philosophy.

The NVMe disk is swappable, which means it has its own controller which manages power management itself. I did my research to pick an efficient SSD and ended up with a Lexar NM790. It tops Tom's efficiency charts and comes in third place for lowest idle power consumption [0]. This is still ~0.8W at idle. On a 60Wh battery an idling drive alone will kill the battery in 3 days.

Now technically there is the APST (Autonomous Power State Transition) feature in the NVMe specification. Is there some lower APST power state that can get the power draw down? Potentially, but that is a feature well beyond the purview of any SSD reviews I have seen, so I don't know- does this drive have reliable and well-implemented APST state support? How does this interact with the platform-specific sleep state implementation, which presumably wakes the disk sometimes to do some Modern Standby features- how often is it spending time in that 0.8W state versus lower? This can vary between board rev or BIOS version certainly. Beyond the actual drive configuration and ACPI interaction, there is also kernel interaction. Do certain drives behave poorly with Linux? Etc etc.

On the RAM side of things, they are using DDR5 and not LPDDR5. There is a lower voltage on LPPDR5 which is a constant inefficiency, but also LPDDR5 has dynamic voltage scaling and dynamic frequency scaling. There is also technically some voltage drop across the SODIMM connector which you don't need to contend with when you solder RAM, which would be a constant source of loss, but I am not sure how significant that is.

Beyond this you have different behaviour for every model of RAM. This post on the Framework forum shows the user could get 7.82 days of suspend time with the HMCG66MEBSA092N DDR5-4800MHz 16GB kit whereas only 2.25 days with the CT2K48G56C46S5 DDR5-5600MHz 96GB kit [1]. Consider that there are effectively infinite combinations of memory people can run, and even inside a model series, vendors can swap their chip providers, etc. Which kits give the best battery endurance? I can't tell you.

Now someone could certainly embark on a long adventure to test different drives, RAM kits, and measure their performance, recommend tunables for the Linux kernel you want to set for each particular set of hardware, etc. But this is effectively what Apple is doing for you with the MacBook. They are choosing their memory supplier, their flash supplier, and integrating as much as possible into their SoC with presumably an entire team focused on extracting the most efficient behaviour out of both.

Consider this same thing extends to display behaviour (beyond VRR support, which I believe Framework has now, you also have local dimming behaviour to tune on the MBP), wireless behaviour, all sorts of embedded controllers that Apple can wrap inside the SoC that I probably wouldn't think of... I don't see how a modular system like Framework can achieve anything close to the idle efficiency of a MacBook.

[0] https://www.tomshardware.com/reviews/lexar-nm790-ssd-review/...

[1] https://community.frame.work/t/impact-of-ram-density-on-susp...

iPhone Air 11 months ago

The case is part of the utilitarianism. I need to attach my phone to a bike mount. I use a Peak Design case that has a locking attachment point in the centre to securely fasten the phone onto the handlebars.

I would gladly ditch the case if Apple had a strong mounting system integrated into the phone (MagSafe has nowhere near the resistance to shear forces sufficient to hold a phone over bumps on a bike.)

I suppose I am looking for the phone equivalent of a camera thumbscrew mount. If Apple iterated on MagSafe to include an actual mechanical fixture as part of the attachment, I would buy that phone right away so I can avoid using these crappy pieces of rubber/plastic that degrade so much more quickly in appearance than the phone frame rails.

iPhone Air 11 months ago

We could potentially see one-time-purchase model checkpoints, where users pay to get a particular version for offline use, and future development is gated behind paying again- but certainly the issue of “some level of AI is good enough for most users” might hurt the infinite growth dreams of VCs

iPhone dumbphone 11 months ago

In the UK it has become very common to need to scan a QR code on your table to order at a restaurant, which takes you to a website.

Most certainly you can still order at the bar the old fashioned way, but since COVID, physical menus have been removed, so how is your group meant to decide what it wants to order before one of you goes up on its behalf? (You cannot all go up if you want to hold the table.)

I don't even particularly mind the experience of using the website; the interface enables the display of all ingredients & allows you to specify allergens they need to avoid. If the kitchen runs out of an item, they can mark it as unavailable in the webpage. Finally, fighting to order at a busy bar was never a fun experience to begin with (it is the norm in non-fine-dining experiences in the UK to not have your order taken at your table.) But, this does require you allow arbitrary internet access on your device, which complexifies the blocking situation.

Unless you need the GPIO pins it is likely a better choice to go with one of the many x86 mini PCs on Amazon, although you need to ensure it is not totally no-name.

I got a GMKtec G5 which is about ~3"x3"x2", has an Intel N97 CPU, 12GB of RAM, and a 512GB SSD (upgraded to a 2TB NVMe disk). No need to buy an additional power adapter (included, it's Type-C), HDMI adapters, or case/fan, either. I think it was about £110 with next-day shipping from Amazon.

It has remarkable performance; I tried GNOME on NixOS and it felt instantly responsive for all general purpose desktop use (web browsing, vscodium with my linting extensions, etc). The only area of my everyday workflow in which it clearly fell behind my M1 Max MacBook Pro was in Rust compilation which is obviously expected - I was just shocked how close it was for everything outside of that. This is in huge contrast to Raspberry Pis which suck to use graphically, even with the Pi 5.

It has happily been sitting on my desk running Forgejo, Mastodon, Vaultwarden, and acting as a personal storage server with that 2TB drive for the last ~6 months and I never even hear the fan. Sits at 0.1 load average, despite Mastodon with this many relays previously eating up the contabo VPS' CPU I was using quite handily.

I do not think determinism of behaviour is the only thing that matters for evaluating the value of an abstraction - exposure to the output is also a consideration.

The behaviour of the = operator in Python is certainly deterministic and well-documented, but depending on context it can result in either a copy (2x memory consumption) or a pointer (+64bit memory consumption). Values that were previously pointers can also suddenly become copies following later permutation. Do you think this through every time you use =? The consequences of this can be significant (e.g. operating on a large file in memory); I have seen SWEs make errors in FastAPI multipart upload pipelines that have increased memory consumption by 2x, 3x, in this manner.

Meanwhile I can ask an LLM to generate me Rust code, and it is clearly obvious what impact the generated code has on memory consumption. If it is a reassignment (b = a) it will be a move, and future attempts to access the value of a would refuse to compile and be highlighted immediately in an IDE linter. If the LLM does b = &a, it is clearly borrowing, which has the size of a pointer (+64bits). If the LLM did b = a.clone(), I would clearly be able to see that we are duplicating this data structure in memory (2x consumption).

The LLM code certainly is non-deterministic; it will be different depending on the questions I asked (unlike a compiler). However, in this particular example, the chosen output format/language (Rust) directly exposes me to the underlying behaviour in a way that is both lower-level than Python (what I might choose to write quick code myself) yet also much, much more interpretable as a human than, say, a binary that GCC produces. I think this has significant value.

Of the prescription options, estradiol is probably the most common and easily available, between hormonal birth control and HRT.

It is also likely the most easy to study, as we have 60-70 years of usage that is not correlated with prevalence of other diseases that might skew life expectancy (like metformin etc.), and quite high-quality medical records by virtue of it being vended on a prescription basis.

Despite this, there is not really any clear evidence that it increases life expectancy.

Documentation on features your SQL dialect supports and key requirements for your query are very important for incentivizing it to generate the output you want.

As a recent example, I am working on a Rust app with integrated DuckDB, and asked it to implement a scoring algorithm query (after chatting with it to generate a Markdown file "RFC" describing how the algorithm works.) It started the implementation with an absolute minimal SQL query that pulled all metrics for a given time window.

I questioned this rather than accepting the change, and it said its plan was to implement the more complex aggregation logic in Rust because 1) it's easier to interpret Rust branching logic than SQL statements (true) and 2) because not all SQL dialects include EXP(), STDDEV(), VAR() support which would be necessary to compute the metrics.

The former point actually seems like quite a reasonable bias to me, personally I find it harder to review complex aggregations in SQL than mentally traversing the path of data through a bunch of branches. But if you are familiar with DuckDB you know that 1) it does support these features and 2) the OLAP efficiency of DuckDB makes it a better choice for doing these aggregations in a performant way than iterating through the results in Rust, so the initial generated output is suboptimal.

I informed it of DuckDB's support for these operations and pointed out the performance consideration and it gladly generated the (long and certainly harder to interpret) SQL query, so it is clearly quite capable, just needs some prodding to go in the right direction.

Reversing the brain drain will certainly be difficult, but stemming the flow of graduates is perhaps a different story.

For Canadians considering working in the US, the recent politically-motivated detentions and deportations against green card holders - a group that has significantly stronger rights of abode than TN visa holders under USMCA - certainly factors into the calculus of whether a US job is worth it.

Further, the appeal of the American political environment is inversely correlated with the distance between one's personal characteristics and the feature vector of able-bodied, white, male, cisgender, and heterosexual. This is nothing new, but the importance of these characteristics, especially the last three or four, have dramatically increased with the rightward shift of the US over the past few years - and there are also more individuals in Canadian STEM (i.e. TN-eligible) degrees than ever before who don't have these characteristics.

Just to put some HN-relevant ballpark numbers to it, the University of Waterloo, a notorious Silicon Valley hiring pool, is reporting 38.6% women [0] in their 2024 engineering admissions. This was 21.2% in 2014 [1], the earliest year with statistics available.

I don't think it is much of a stretch to say the 2014 first years were more likely to aspire to intern at, say, Tesla or Meta than the 2024 first years (who will be entering their first co-op internships in a month's time). That is a function of both the American tech companies having cozied up to the political right, and an increased proportion of the students being both more directly affected [2] and morally repulsed [3] by this state of affairs.

Add an additional nationalism multiplier for Canadians being turned off by the annexation rhetoric coming out of the US, and I think we may see a change in this trend towards Canada retaining more of its local talent.

[0] https://uwaterloo.ca/engineering/about/faculty-engineering-s...

[1] https://web.archive.org/web/20140804070109/https://uwaterloo...

[2] https://www.theguardian.com/commentisfree/2022/jun/27/roe-wa...

[3] https://www.ft.com/content/29fd9b5c-2f35-41bf-9d4c-994db4e12...

Apple M3 Ultra 1 year ago

I think that is much too hand-wavy regarding the performance differences.

Both Passmark and Geekbench are aggregates of a variety of tasks. If you dig into the individual tests that constitute this aggregate score, you will find different platforms perform better, or worse, on certain tests than others. I would wager that, for many applications, only a subset of these tasks are relevant to the performance of the application, yet such benchmark suites distil out all nuance into a single value.

Here is a personal anecdote. I have tried running CASTEP (built from source), a density functional theory calculator, on both an M1 Max MacBook Pro [0], and on a Ryzen 7840HS Lenovo laptop [1]. A cursory glance at those Geekbench results linked might make you expect that the performance is roughly equivalent, but the Ryzen outperforms the Mac by about 4x, a huge difference.

What happens if we try and dig into any particular benchmark to explain this? If you click on any particular benchmark in the Geekbench search lists, you will see they test things like "File Compression", "HTML5 Browser", "Clang". Which of these maps most closely to the sorts of instructions used in CASTEP? Your guess is as good as mine.

If anything, I would say Passmark is quite a bit less abstract about this. Looking at the Mac [2] and Ryzen [3] Passmark results, you can see the Ryzen outperforms the Mac by about 2x on "extended instructions", which appear to involve some matrix math, and also about 2x on "integer math". The Mac, meanwhile, appears to be extremely good at finding prime numbers, at over 3x the speed of the Ryzen. Presumably the Ryzen's balance of instruction performance is more useful for DFT calculations than the Mac's, which perhaps is weaker in areas that might matter for this application, but stronger in areas that might matter for others.

Of course, optimization is likely a component of this. How much effort is put into the OpenBLAS, MPI, etc, implementations on aarch64 darwin vs. x86-64 linux? This is a good question. It is, however, mostly irrelevant to the end consumer, who wishes to consume this software for use in their further research, rather than dig into high-performance computing library optimization.

[0] https://browser.geekbench.com/search?q=7840hs

[1] https://browser.geekbench.com/search?q=m1+max

[2] https://www.cpubenchmark.net/cpu.php?cpu=Apple+M1+Max+10+Cor...

[3] https://www.cpubenchmark.net/cpu.php?cpu=AMD+Ryzen+7+PRO+784...