gcc and LLVM have also copied ideas from C2, and there's been cross-pollination between Java and C++ compilers going back decades. All three are among the most sophisticated optimising compilers right now.
HN user
pron
https://pron.github.io/
Working on OpenJDK at Oracle
(BTW, when I wrote "Not one of those projects ... changed their language", I meant after less than a decade, as a continuation to the previous comment)
First, technically respected products like the ones you describe either 1. plan or expect to switch in advance (e.g. they start with, say, Python/Ruby, expect that if they grow they'll switch to, say, Java) or 2. they improve their chosen language runtime (e.g. Facebook with PHP/Hack or Shopify with Ruby). Projects that switch a language without expecting to always show a pattern of bad decisions (clearly, if they thought their chosen language will carry them through growth and then they're convinced that it won't, that means that they don't know to judge languages' merits).
Second, this is clearly not the situation here, is it? There is absolutely no new information that Bun learnt in the past year that they didn't have five years ago, and certainly this has nothing to do with growing workloads on some service. They say they believe the language they have chosen lacks the features needed for the very core of the domain, which is dealing with JS objects. As someone working on the HotSpot JVM, I can tell you this is not true, but fine - that's what they believe. What could have taken them five years to come to that conclusion? Again, this happens to be a domain close to my own, only much simpler, and I seriously doubt they made some novel discoveries in the past year. And if it's taken them five years to acknowledge what they now think are fundamental limitations with the language they had chosen, how can they be so confident they've made a right choice now after a few weeks? It looks like they chose Zig on a whim and then chose Rust on a whim, and neither of these choices is the source of their problems and neither is the solution to them.
It's not necessarily the thing that matters most to executives, who are often those making decisions, but it's always been the thing that mattered most to programmers (at least those of them who have any emotions or strong preferences toward programming languages).
This is just so weird to me, because I would say the same about Zig.
Then why is it weird if you're saying the same thing? Different programming languages appeal to programmers with different tastes, and so it makes sense that some programmers would be drawn to language X and dislike language Y, while others would be the opposite.
I should add that in the 30 years I've been a professional software developer, I've worked on and advised many projects. They all ran into serious challenges at one point or another. Not of those projects that was held in high technical regard changed their language (except for things like JS -> TS or when the project planned to change languages, starting with one suitable for prototyping and expecting to switch if and when their workload grew).
All the ones that opted to switch language after less than a decade were those with serious shortcomings in their technical decision process, and those problems, unsurprisingly, persisted after the language change. After all, the very decision to switch so soon is an admittance that they'd made a very serious misjudgment, but these projects never properly debrief why they'd made such a big mistake and how they can avoid making one again.
Sounds to me like his choice of Zig was made in haste, as was his choice of Rust. If you find yourself changing a project's primary language more than once a decade (more like 15 years, but let's say a decade), the problem isn't the language but your technical decision process, and that's what you should look into first.
Some of the world's more important software - from browsers to the JVM - mix high-level languages with a GC and low-level languages, and it works not because of a style guide (even though one may exist). As someone working on the HotSpot JVM, I can say that it's done with a lot of thinking about constructing the right primitives that make this work well. Zig doesn't lack the features to construct the mechanisms required for getting good results in that domain, and Rust doesn't have features that could save you the thinking about such mechanisms.
The first solo-founder unicorn isn’t built by a genius doing the work of three hundred people. It’s built by one person sitting at the center of a coordination layer that does the work the three hundred people mostly used to do... What’s new is AI that scales a single person’s coordination capacity, attacking exactly the cost the gig economy couldn’t.
Two problems with that:
1. AI isn't free, and how cost-effective it is remains to be seen.
2. AI can't currently really do the full job of one person, let alone three hundred [1]. And when it is able to do the job of three hundred people, the very structure of the economy is likely to change so much that any transfer of details from the existing economy to that imagined one may well be irrelevant. In other words, at the point AI is able to do something so transformative, it's unreasonable to think that the structure of one company will be revolutionised without everything around it also being revolutionised.
It stands to reason that the economic value that one person can do with a relatively cheap tool (assuming that the AI that could do all that is cheap enough) will be similar to whatever one person could do with a relatively cheap tool at any other point in time. An increase in productivity in the presence of competition lowers the price of the product by about the same factor as the increase in productivity. People have more stuff, but not necessarily more money. A person with a laptop and a 3D printer might be a "unicorn" if they were transported back in time to 1526, but it doesn't make them a unicorn today because many other people can do that, too.
[1]: So much of the old grunt work in the knowledge economy is already automated (typing, copying, posting letters), and so three hundred people are probably doing some non-trivial work already, and replacing them means AI with much better capabilities than we have today.
We comprehend, but these "scaling laws" have been in effect for less than a decade (and they're more historical observations over a very short period than actual laws), and while some technologies progress exponentially for some amount of time, the complexity of some computational problems grows exponentially forever. For example, if computational resources double every year, it may still be five centuries before some computational problems can become practically computable. It is mathematically proven that no amount of intelligence can compensate for resources, and even if exponential growth could be sustained for a long time, and that's a big if, an exponential curve still grows slowly at the beginning.
For example, suppose AI helps us figure out a way to exploit much more of the sun's energy. Accomplishing that necessary preliminary can, on its own, take many decades, and if we split our efforts among multiple approaches, it may take longer. Intelligence can't break the actual laws of mathematics or of physics. And that's before we consider things like how resources will be allocated when AI tells people that man-made climate change and transgenderism are real.
This generation of AI hasn't even hit its first crisis. Assuming there will be none is like settlers assuming their town will never suffer a major earthquake because they haven't had one in five years.
Of course, improvements in problems that may not grow exponentially also matter a great deal, but there are too many unknowns (look at how many unknowns there are around quantum computing).
robots seem mainly limited by software
First, I don't think so. Second, some resources would take even robots decades or centuries to collect. Even something as fundamental as energy production takes a lot of time and money to build. The question of where we'll get the energy to run the robots is not a simple one, and it's one of the simplest questions involved.
I think that's what original comment was trying to argue, and I think it's possible within the limits you've laid out.
Possible isn't the same as likely, and the reason we don't extrapolate is that different extrapolations lead to very different results.
But regardless of what happens and when, I think that we, people educated in computer science, should remember that many questions are simply not answerable in a short amount of time, and we know with absolute certainty that answering them is not a question of intelligence but of computational power and time. We no that no human or machine, however intelligent, can predict or control the nonlinear systems that are all around us because they are computationally intractable.
Because the thing pursuing the goal would be able to improve it's ability to pursue the goal
But not necessarily without needing a huge amount of resources. It is a mathematical certainty that no intelligence can solve computationally intractable problems (including forecasting the weather, or the economy) without access to resources we simply don't have.
I suspect we just have different beliefs about how close we are to RSI.
I don't have any belief on the matter, but my scepticism isn't necessarily about RSI itself, but about how much it would matter even if it does happen soon. Too many things are limited by lack of resources that no brain in a jar can obtain. And if such a brain in a jar itself is very expensive to operate, it may not be easy for it to justify its existence. My point is that the technical aspect is uncertain, but it is also only a part of a larger system that's has many sources of uncertainty.
Because most systems in nature and society, including technological progress, aren't linear.
It's certainly plausible that improvements will continue, but the pace is completely unpredictable. It's also plausible that the training material is polluted and progress will not continue. I'm just saying that predicting the rate of technological progress is not easy, and historically, it's rarely been smooth in the long run.
There are many complicating technical factors, but also non-technical ones. Technological improvement in the short term will not necessarily yield commensurate economic gains, in which case, investment may not grow enough to sustain the progress.
As for the recursive part, currently it's either hypothetical or based on too few data points. Not saying it won't happen, but it's far from being the only plausible trajectory.
I just don't understand why people are incapable of extrapolating.
I guess it's a matter of education. People with a mathematics or computer science background know that unless the dynamics of a process are known to be linear, extrapolation is usually wrong. Since we don't know that the dynamics here are linear and have every reason to believe they aren't, extrpolation is unlikely to teach us anything valuable.
Sure, but because work is something we spend so much time on, when workers can quit when the work feels meaningless or absurd, that's a good thing. We should aspire for a society where all workers can do that.
Every time you compile a statically-typed programming language you are using formal verification
Yeah, this is not what we're talking about here. We're talking about proving properties with deep alternative quantifiers.
That isn't just hard. Proving software correct in complete generally is impossible. There are all kinds of practical and fundamental constraints that leave it to be impossible. Verification is only useful when you are acting within the scope of a compressed specification of a system's behaviour.
Nobody said anything about complete generality. We're talking about the practice of applying formal methods. It's not writing in Rust, and it's not a general program verifier, but a practice that's applied in some parts of the industry and not others, as the article says.
Put another way, the question is: for those programs and those properties that humans are able to prove with proof assistants, how expensive is it for LLMs to do that work.
I would think the cost multiplier in those cases is much lower for an LLM as compared to a human that doesn't have an inherit understanding and needs to give it thought. Wouldn't you?
No. I don't see why proving would require less relative effort for an LLM. In fact, years ago, long before LLMs, I wrote about why it is relatively easy to write sort-of-correct software yet hard to write provably correct software, and I don't see why it's any different for LLMs. Their power lies in inductive "intuition", while deduction requires effort, just as it does for humans: https://pron.github.io/posts/people-dont-write-programs
But there's no need to speculate. Those who think verification-by-LLM is feasible and cost-effective on an industrial scale, are welcome to try it and report what they find. So far I've seen only tiny examples, and even they don't show effortless (i.e. token-light) work by the agent.
Whatever the cost multiplier is, I see no reason why that same multiplier won't remain with AI.
Personally, I don't think that picture is quite accurate. Yes, there is a high cost multiplier for small programs, albeit perhaps not so prohibitive. But for large programs, that multiplier is, for most intents and purposes infinite, unless, perhaps, you have experts who know what's worth proving and what is not.
Anyway, I'd like to see that put to the test. Have an LLM write a 50-100KLOC program and prove all correctness properties - with the properties themselves approved by an expert human - and tell us what it cost. A colleague of mine stopped his AI proof experiment when he got an email from some functionary at the company to stop doing what he was doing with the model, because it was costing too much money.
It’s no longer just for safety-critical systems with the budget for specialized proof engineers. It’s for anyone who has a property worth proving
... and the budget to pay the AI to prove it.
I have quite a bit of experience with formal verification, but I don't understand the claim made in the article. As an aside, AI's ability to reliably prove the correctness of significantly large programs is still theoretical at this point, but let's assume it's possible. The claim in the article is that writing 10,000 lines of proof to prove a 100-line program was very expensive, and that's why it isn't done. But this increase in cost continues with AI! Whether you pay people to write the proofs or you pay an LLM to write the proof, you still have to pay for it. If I run a software company, saying that "verificaton is the AI's problem" isn't much different from saying, "it's the engineers' problem." Either way I'm not doing the work myself, but I am paying for it.
If the premise is that writing proofs was 100x more expensive than testing, I see nothing in this article to even suggest why it wouldn't still be 100x more expensive when an LLM is doing the work.
(BTW, the reason there aren't many specialised proof engineers is because they aren't in high demand; they're not being paid that much more than other engineers at a similar level)
Nominal. The inflation-adjusted price today is 2/3 of what it was then.
If we view Rust (including unsafe) as a memory-unsafe language, then it's the same as C++, since we can then view Rc/Arc as optional. But if we want to look at Rust as a memory-safe language, then it mandates the use of GC when an object may have multiple owners. In other words, Rust depends on GC to ensure the memory safety of common functionality. It is true that Java depends on GC for even more operations, but the fact remains that it's very hard to write many large Rust programs without the use of the GC in its runtime (unless you go unsafe, in which case it's like C++, where the GC is optional).
I think many people, especially those with insufficient experience with both low-level languages and modern garbage collectors incorrectly assume that the presence or reliance on GC necessarily implies some performance overhead. In actuality, some GCs (moving GCs in particular) were invented, among other reasons, to reduce the overhead imposed by malloc/free that causes significant performance problems in large programs written in low-level languages. Of course, refcounting GCs, as well as some tracing GCs (non-moving ones) also rely on malloc/free, so they may still suffer from the same issues.
Another misconception is that "a GC" is some necessarily large and sophisticated runtime mechanism compared to "no GC". The problem with that view is that modern malloc/free are also large and elaborate runtime mechanisms (in the range of 10KLOC), and they're elaborate because clever sophistication, as well as CPU/footprint tradeoffs, are required to get decent performance from such allocators (another fact that experienced low-level programmers know). Modern malloc/free allocators may be larger and more complex than simple moving collectors (although it is true that modern moving collectors are larger and more complex than modern malloc/free allocators, but they both require non-trivial runtimes).
https://www.taylorfrancis.com/books/mono/10.1201/97810035953...
Note that this is a pretty new technology. The first production-quality "pauseless" moving collector for commodity hardware was released in 2010 by Azul, but it was proprietary. The first open source implementation (that was also generational) was in JDK 21 (https://openjdk.org/jeps/439), i.e. it's less than three years old. The book I linked to is by one of the primary designers of that GC (and one of the world's foremost experts on memory management).
ZGC does no work in stop-the-world pauses; no marking, no compacting, not even root scanning.
Of course, if your language targets the JVM, it will automatically get to enjoy that amazing GC.
so "low footprint" also generally means "low latency"
Not anymore.
You're absolutely right that one of the reasons moving collectors were not used more widely was that, while their throughput was always very impressive, their latency wasn't that great, but that changed a few years ago.
E.g. Generational ZGC in OpenJDK (released in September '23) introduces hiccups or "pauses" that are not dependent on the size of the liveset and are no larger than latency hiccups introduced by the OS (assuming no realtime kernel), i.e. <1ms (and typically <<1ms) up to heaps of 16TB. In fact, the latency can be smoother than approaches that have an explicit free operation and require maintaining a free list, as freeing a large object graph can be quite slow and occur in surprising places.
So modern moving GCs no longer have a latency penalty, but this is newer than even ChatGPT.
But that doesn't mean any language that allows you to implement reference counting as a library, is a garbage-collected language.
The concept of "a garbage-collected language" is not well-defined. There are languages, like Java, Rust, and Python that depend on a garbage collection mechanism, and languages like C, C++ and Zig, which don't. C++ happens to offer a GC in its standard library, however.
That "working developers" use some other terminology is not what matters. What matters is whether the terminology they're using expresses important distinctions or not (and may, in fact, express misconceptions about distinctions). In the case of memory management (as in the case of "transpiles", although there the damage isn't as high), the colloquial terminology is misleading as it is used to hint at distinctions (such as about performance) which are simply not there. E.g. moving GCs are used to avoid the performance overheads of malloc/free, especially in large and/or concurrent programs. This performance overhead that C and C++ suffer from is well known to experienced low-level developers (which is partly why large programs that benefit from moving collectors are relatively rarely written in such languages anymore), but now the terminology is used as a cargo cult, which leads to conclusions that are sometimes the very opposite of what's really going on.
Right, and there are differences within tracing GCs that are just as big as between refcounting (and even manual malloc/free) and tracing. For example, Go uses tracing to determine when an object lifetime ends. But the moving collectors in Java, .NET, and V8 don't know and don't care when objects die, and they have no "free" operation at all. In many ways, the performance profile (of favouring smaller footprint or higher throughput) of memory management in C++, Rust, Python, and Go share more similarities among themselves than Java, .NET, V8, and Zig, which also share a more similar profile (arenas, like moving collectors, don't need or want to know when an object's lifetime ends).
Another distinction without a difference that is really just giving a name to a misconception is the notion of "a runtime". When I learnt C in the late 80s or early 90s, the book said something like, "C is not just the language, but a rich runtime". Indeed, modern malloc/free implementations mean that a C program ends up needing a larger and more elaborate runtime than a program in some educational language that uses a trivial implementation of a mark-and-sweep collector. Modern malloc/free allocators also sometimes come with an impressive set of tuning knobs. It's just that people who haven't had a lot of experience writing large programs in low-level languages don't know about them (or they just work to avoid allocations as much as possible, because that's what they've been told to do).
I think a lot of people just want to be able to discuss different areas of the automatic memory management design space separately, and maintaining the distinction between reference counting and garbage collection (meaning tracing GCs) lets them do that.
The problem is that there are many differences in memory management techniques that offer different tradeoffs, and the difference between refcounting and tracing is not necessarily the biggest of them.
For example, one of the most important distinctions in memory management is whether it optimises for footprint or speed (or some compromise), and the line isn't where people who don't understand memory management think it is. It can matter (often a great deal) whether you determine that an object is dead dynamically (say, by counting references) or statically (by manually writing free or by having the language track lifetimes), but it doesn't matter as much as whether or not the mechanism needs to know when objects are dead in the first place. So reference counting, manual free, static lifetimes, and even non-moving mark-and-sweep tracing collectors (like Go's) generally optimise for footprint at the expense of speed (although different allocators can have some control over that tradeoff), while arenas and tracing moving collectors optimise for speed at the expense of footprint (although here, too, they have some control over the tradeoff). So the line for this super-important tradeoff is between [manual, static, refcoutning] and [arenas, moving tracing]; non-moving tracing collectors are somewhere in between but may be closer to the first group.
People who don't understand memory management and may not have a lot of experience in low-level programming sometimes think that manual or statically-determined freeing must be fast because low-level languages, which inexperienced people think are fast, use them. In fact, low-level languages have some concerns that are much more important than speed and that preclude them from optimisations such as moving pointers. To get around that performance handicap, these languages try to avoid using their heap memory management as much as possible because they're using a rather slow technique because of their constraints.
It's not about restricting the language. It's that practising programmers often don't know a subject well enough, so they use different words to make distinctions that don't matter as much as they think (see "transpile"). "Dynamically typed" is actually not that big of an offence (because the distinction is real, it's just that the terminology is a bit muddled), and the people in PL theory who are bothered by this (most notably one person) are considered pedants even among their colleagues.
E.g. many practising programmers don't know that tracing moving collectors are used to avoid some of the high overheads associated with memory allocators (malloc/free), which are themselves big and complex beasts that make up substantial "runtimes" (another misused and misleading word).
What I didn’t like about this series of books was choosing “garbage collection” as umbrella term for both tracing GC and reference counting, without verifying if programming community would agree with that, which turned out they didn’t.
This has been the standard terminology in memory management research for many decades. The only programmers who don't like it are those who don't understand the principles of memory management.
By that definition, C++ code has garbage collection if it uses std::shared_ptr
That's right.
going against widespread common usage of the term “garbage collected programming language” which specifically contrasts manual languages like C++ or Rust against garbage collected ones.
Since this contrast mostly exists in the minds of people who don't understand memory management, going against this common misconception is good. That's not to say that there aren't some interesting tradeoffs that often align with the colloquial perception, "garbage collection" isn't the interesting part. As you said, both C++ and Rust use GC; in fact, they use a GC somewhat similar to the one used by CPython.
I'm saying I don't like this proposal
And I'm trying to tell you that this document is not what you think it is. It's a rough sketch of a building's foundations and your critique is about the roof. Even if this were to become a proposal, it's likely the matter of defaults of this feature will be covered by a different JEP, because Java features are usually broken down into multiple JEPs. What you're complaining about may be part of the feature, but it is not covered by this document. If this becomes a JEP and another JEP talks about nullabilty defaults, then you could criticise the selection of defaults, but that particular aspect of this feature is outside the scope of this particular document. So one, this is not a proposal, and two, you haven't seen the description of the part of the feature you want to criticise.
We are well aware that splitting features over multiple JEPs can invite such misunderstandings, but that doesn't change the fact that Java features are, at least currently, split over multiple JEPs. We are also well aware that if this does become a proposal, adding ! everywhere is not what we want, but we want to cover that aspect in a different document, as we usually do. Most Java users aren't confused by this because they don't read the JEPs at all, but such splits help focus the discussion on one aspect of a feature at a time. So your desire for a non-nullable default is very understandable, it's just not relevant as a critique of this document.
For example, the virtual threads JEP described a pinning limitation. We knew it reduced the applicability of that JEP and said as much. We just wanted to address it in a different JEP, and so we did (https://openjdk.org/jeps/491). Ever since the JDK switched to time-boxed, semiannual releases, this is how we've delivered features: in multiple pieces. The same applies to Valhalla.
it's still a bad choice
Except there's no choice here. It's a draft of an exploration of a portion of a feature.
Drafts are ideas that have not even been submitted for consideration for inclusion in the roadmap. You can find drafts over ten years old that have long since been superseded by other ones or abandoned altogether: https://openjdk.org/jeps/0#Draft-JEPs. Some JDK engineers write JEP drafts when they feel they get close to something they would like to propose, while others write drafts for pretty much any idea they have.
The problem is that it defaults to everything nullable and adds noise for non-nullable
It doesn't, though. Even if this draft ever becomes a JEP (and I don't know if it or anything like it will), in its present form or another, it would still, like most JEPs, describe only part of the feature. It's perfectly on point for one JEP to describe the explicit nullability annotations, while another describes the defaults. Smaller features than this have been split into two or three JEPs. This is just how Java features have been described and delivered for years now (see how many JEPs patterns were split into: https://openjdk.org/jeps/0).
It's perfectly okay to dislike some design, but I find it strange to assume, based on a draft of a portion of a feature that one of the most experienced and successful programming language design teams in the history of software is likely to get it wrong. Maybe wait until there's an actual proposal and a roadmap for the complete feature before critiquing it? It's like seeing a draft of some building's foundations by a prestigious architectural firm and saying, these idiots forgot to plan a roof.
Java is getting that power in a different, more orthogonal way, IMO, through Project Babylon: https://openjdk.org/projects/babylon/