HN user

jandrewrogers

24,440 karma

Designer of bespoke analytical database engines for applications with unusual performance and scale requirements. Knows a lot about geospatial and sensor data. Went to school for chemical engineering.

email: andrew at jarbox dot org

Posts109
Comments4,740
View on HN
www.quantamagazine.org 26d ago

A Dark Dimension Could Link Two of the Universe's Great Unknowns

jandrewrogers
5pts0
arxiv.org 1mo ago

From AGI to ASI

jandrewrogers
9pts0
www.beren.io 1mo ago

Capital Ownership Will Not Prevent Human Disempowerment

jandrewrogers
2pts0
www.quantamagazine.org 1mo ago

When Quiet Undersea Volcanoes Turn Disruptive

jandrewrogers
3pts0
quantumai.google 3mo ago

Securing Elliptic Curve Cryptocurrencies Against Quantum Vulnerabilities [pdf]

jandrewrogers
54pts32
www.washingtonpost.com 3mo ago

China Bars Executives at Meta-Owned AI Company from Leaving Country

jandrewrogers
9pts2
www.quantamagazine.org 5mo ago

Climate Physicists Face the Ghosts in Their Machines: Clouds

jandrewrogers
4pts0
www.quantamagazine.org 5mo ago

Are the Mysteries of Quantum Mechanics Beginning to Dissolve?

jandrewrogers
16pts0
sebastiangaliani.substack.com 5mo ago

Marginal Revolution

jandrewrogers
6pts3
www.reuters.com 5mo ago

US Accuses China of Secret Nuclear Testing

jandrewrogers
13pts4
www.quantamagazine.org 5mo ago

Long-Sought Proof Tames Some of Math's Unruliest Equations

jandrewrogers
5pts0
papers.ssrn.com 5mo ago

The Role of Unrealized Gains and Borrowing in the Taxation of the Rich

jandrewrogers
4pts0
cedardb.com 5mo ago

Efficient String Compression for Modern Database Systems

jandrewrogers
153pts47
www.oxfordeconomics.com 5mo ago

No US-Style AI Investment Boom to Drive EU Growth

jandrewrogers
1pts0
eml.berkeley.edu 5mo ago

Monopsony, Markdown, and Minimum Wages [pdf]

jandrewrogers
1pts0
www.quantamagazine.org 6mo ago

Monster Neutrino Could Be a Messenger of Ancient Black Holes

jandrewrogers
4pts0
www.quantamagazine.org 6mo ago

Why There's No Single Best Way to Store Information

jandrewrogers
4pts0
www.quantamagazine.org 6mo ago

String Theory Can Now Describe a Universe That Has Dark Energy

jandrewrogers
1pts0
wunkolo.github.io 6mo ago

vpternlog: Signed Saturation

jandrewrogers
2pts0
strathprints.strath.ac.uk 6mo ago

Meritocracy and Inherited Advantage in the United States [pdf]

jandrewrogers
2pts0
www.nber.org 6mo ago

O-Ring Automation

jandrewrogers
23pts10
www.nber.org 6mo ago

Why Care About Debt-to-GDP?

jandrewrogers
1pts0
academic.oup.com 6mo ago

Coexisting with Humans: Genomic and Behavioral Consequences in Bear Population

jandrewrogers
2pts0
www.nber.org 6mo ago

Delivering Higher Pay? Impacts of a Task-Level Pay Standard in the Gig Economy

jandrewrogers
1pts0
www.nber.org 7mo ago

Delivering Higher Pay? Impacts of a Task-Level Pay Standard in the Gig Economy

jandrewrogers
2pts0
www.nber.org 7mo ago

"Captain Gains" on Capitol Hill

jandrewrogers
1pts0
www.quantamagazine.org 8mo ago

Mixing Is the Heartbeat of Deep Lakes. At Crater Lake, It's Slowing Down

jandrewrogers
3pts0
www.sciencedirect.com 8mo ago

Possible Enhanced Meteor Impact Risk in 2032 and 2036 from Taurid Asteroids

jandrewrogers
1pts0
papers.ssrn.com 9mo ago

Evidence from Minimum Wage Increases and University Lab Employment

jandrewrogers
3pts0
jack-vanlightly.com 9mo ago

Beyond indexes: How open table formats optimize query performance

jandrewrogers
95pts3

Most scalar-to-SIMD conversion requires changing the design of data structures and algorithms to be effective. Compilers are required to exactly reproduce the specified data structures in a deterministic way for obvious reasons.

Even if compilers were clever enough to transform your data structures and algorithms for SIMD (they're not), the data structures are a contract that can't be unilaterally modified.

This has been discussed in strategic geopolitical circles for a long time. When surveying defense analysts and similar if they believe a "great powers conflict" will occur in the near future, the percentage that say yes has been rising rapidly (currently 40-50%). If you search for "probability of global war in $YEAR" you will find a lot of discussion of varying quality.

The major alignment is almost always assumed to be something like Russia/China vs Europe/US/Japan.

You can find circumstantial evidence for this in the change in military posture of many countries and rapid scaling of production capacity. Some random examples:

French General states that they are preparing for open war in "3-4 years" in 2025:

https://www.politico.eu/article/french-top-general-expects-s...

Sweden increases defense spending by 40% starting in 2021:

https://www.defensenews.com/global/europe/2020/12/15/sweden-...

A war with large geographic scope has been independently modeled as having a high probability of occurring circa 2030 by major governments for a decade or more. As time has passed, reality has not diverged from those models, so confidence in the conclusions have grown. It is worth noting that this pre-dates current political dramas, Russia's invasion of Ukraine in 2022, etc. The underlying strategic geopolitics driving it is much deeper than those issues.

It has been accompanied by a large spike in military spending globally. The backlog of foreign military sales of US weapons just waiting for approval is ~$1T.

Property tax pays for maintaining the utilities and services connected to the property. If you buy land that is completely unconnected from public services, you will find that the corresponding property taxes are extremely low.

I've had annual property taxes as low as $1/acre.

In the US, bricks and concrete are unsafe, not "higher quality". Much of the US has severe earthquakes which collapsed the European-style masonry buildings they used to build. There are many old photos of what happened to cities that were built with brick/concrete after major earthquakes. The main alternatives are steel and wood, same as in Japan.

My home is rated to survive earthquakes stronger than any in recorded European history without structural damage.

It depends on the country. The US has the ability to do the equivalent of deep "duck typing" with almost no documentation even if you've been off the grid in the developing world for a long time. This is a case where the intelligence apparatus works in your favor. They can know you are a US citizen with high probability absent obvious evidence of such.

Of course, if you fall into a crack that is beyond their reach you will almost certainly have a more difficult time. For a variety of historical reasons, there has been relatively little reliable documentation of American citizenship so the system adapted to that reality.

The direct cost of manufacturing the units is only a small percentage of the total costs of the retail product. Logistics is often a bigger cost. You can create substantial economic efficiencies elsewhere in the supply chain by allowing some manufacturing waste.

Production processes that require more production and supply chain customization for each order have significantly higher costs that need to be amortized. It is cheaper to pack and ship identical boxes at the factory than to customize the contents and logistics of each box for every retailer or customer. The more variation and complexity you allow into the supply chain, the more capital infrastructure, equipment, and people you need, all of which must be amortized into the retail unit cost.

The costs of increased supply chain variability and customizability can easily exceed the cost of wasting a few units. You may have wasted hundreds of t-shirts but you also didn't have to invest the millions of dollars in systems and equipment that would have prevented that waste. These are low-margin businesses, everyone is carefully tracking and attributing these costs.

Supply chains in most industries continuously and ruthlessly optimize to squeeze out waste while trying to increase flexibility. The number of items that are produced on demand -- and therefore produce little waste -- has grown dramatically over the last couple decades. However, many goods intrinsically have long, slow supply chains which makes waste all but unavoidable.

Unfortunately, there is an issue with food pantries where people who are not in need use them because free food. People can be shameless. It is a minority but still too common and doesn't come with the stigma it deserves in some places. In Seattle, I've even heard a few anecdotes of people trying to resell food from the food pantries.

This behavior does impact prices in the normal market at the margin, particularly if it becomes normalized.

There is no global definition of "less common size". It varies greatly from one locale to another. At the same time, production has relatively high fixed costs and is centralized.

It would be very expensive for the global factory to customize the distribution of sizes manufactured for a retail store in Des Moines, Iowa. The order is tiny and it would require customized logistics, all of which greatly increases cost and complexity.

Currently, unpopular sizes are over-produced because they are subsidized by popular sizes. If the unpopular sizes have to be paid for, the logistics and production processes would push producers to under-produce popular sizes.

A key insight is that what constitutes an "unpopular size" is a very local phenomenon. Every point of retail sells a different, semi-predictable distribution of sizes. It is much cheaper to ship sizes no one will buy than to manage the logistics of exactly matching local demand for a specific distribution of sizes.

I asked the same question to someone who works in this business and got an eye-opening detailed explanation that made it obvious in hindsight why things the work the way the do. The difference in product cost and logistics infrastructure was not small.

Who pays for the logistics cost of moving and stocking these products in discount stores and giveaway centers? That is a large percentage of the total cost of production and the reason disposal is cheaper.

If those costs are paid for by taxpayers then the consumers are in effect involuntarily buying products they would not have otherwise bought, just with more steps. We already see this with agricultural subsidies.

If those costs are charged back to the producer then it becomes economically optimal to under-produce, which will cause prices to rise and risk shortages but eliminate waste. One can make the argument that higher prices for basic goods to reduce waste is a social good but it also impoverishes consumers.

All of these scenarios have happened empirically countless times. That almost every producer over-produces to some extent at no profit to themselves when allowed has strong "Chesterton's Fence" characteristics.

I think you may be confused about what "throughput-optimized" means. HFT is not throughput-optimized by definition. The LMAX link literally says it is a latency-optimized system. An optimal throughput-optimized system has unbounded worst-case latency -- the opposite of "latency-optimized".

None of those links contradict anything I wrote, I am already familiar with all of them.

That "waste" is almost certainly excess production that can't find a market + spoilage. It is entirely normal for farmers to literally dump their crops in these cases. Everyone in the supply chain does it to some extent.

Lack of market and spoilage makes it unavoidable. When I lived in farming communities there were massive piles where people dumped excess crops.

Any effective approach to atmospheric CO2 reduction on a timescale that matters will require extremely large quantities clean energy, far beyond current generation capacity. This will be the biggest bottleneck so start there.

Given that this is known required input, we can start by expediting the building this energy infrastructure with all due haste. This would require ignoring various activists with a litany of reasons for why deploying solar/wind/nuclear/geothermal at scale is stupid/immoral/unethical.

TBH, physics limits how latency-sensitive weapons systems need to be and you can largely just disable the GC in these contexts. They use CPUs from the 1990s to do hypersonic terminal guidance. You don’t have to do any performance engineering for many latency-sensitive weapon systems. Could probably write it in Javascript.

For throughput-optimized systems, some of which are real-time, you never see a GC. That loss in performance is simply too large such that the computation becomes intractable. A lot of really poor systems admittedly exist but no one considers them “good”.

How so?

The ships sunk in the Falklands War were all less than half the displacement of a typical US Navy destroyer. The sole exception is the Belgrano, which was built in the 1930s!

The ship being tested here is ~25x the size of the largest British ship that was sunk. Generally speaking, there is a super-linear relationship between ship size and the amount of explosive required to sink it. There is mountains of empirical data on this that you are choosing to ignore.

Every military knows this. They are making a tradeoff between size, which makes the target more difficult to destroy and easier to defend, and the number of ships they can build which allows them more flexibility in force projection.

I would guess they want a large enough explosion to generate peak acceleration of the entire ship without a local enough explosion to actually damage it. Getting enough separation to make it non-local requires a lot of explosive thanks to the inverse cube law.

If you look at e.g. seismic damage models, peak acceleration is correlated with most of the worst outcomes.

The structure is engineered to survive a multitude of conventional threats intact. It is testing properties of the design rather than specific weapons per se. Also, these tests are intended to be non-destructive which impacts their design.

Exercises where the US military uses decommissioned aircraft carriers and other large ships as targets are illustrative. They are basically unsinkable. You can hit them with torpedoes, bombs, missiles, etc all day. At the end of the exercise they usually have to send over a specialized demolition crew to actually scuttle the ship. Astonishingly damage resistant.

A nuke would of course do the trick but now you are playing a different game.

People chronically underestimate how difficult it is to get enough conventional explosive on target to sink a major naval vessel, even ignoring the extensive active defenses.

If you are severely I/O bound it isn't intrinsic, it means your server is badly under-provisioned in the I/O department. Linux on a modern server can push 200 GB/s of I/O. Even if web services were engineered to a standard that could consume that much I/O, which they are not, you would have to be astonishingly wasteful to burn it all.

It is rare to be severely I/O bound because software engineered for I/O performance tends to run out of memory bandwidth first.

Games are not throughput-optimized systems in any conventional sense. They are a canonical example of latency-optimized systems.

I have nothing against GCs, I use them regularly even in performance-sensitive contexts. But too many people understate the adverse impact of GCs on performance contrary to evidence and theory.

Sophisticated throughput-optimized systems rely on deep latency-hiding. Schedulers see millions of atomic operations into the future, continuously rewriting the schedule globally to maximize locality and minimize resource contention based on real-time changes to workload, resource availability, and system behaviors.

In short, for each of the millions of in-flight operations (which might only map to a handful of user operations), it is trying to precisely optimize the concurrency, timing, and dependency sequencing such that when operations are executed every resource required is hot, uncontended, and available with high probability. When this works well it dramatically reduces the number of hidden stalls in execution. The schedule is constrained by tail latency requirements; a theoretically throughput-optimal schedule can defer execution indefinitely.

For an analytical database engine, an "atomic operation" is typically a query operation on a database page. A modern server can retire 100M ops/sec. While I am oversimplifying a bit, a 1 millisecond GC pause can blindly wreck the schedule for 100,000 operations in an unpredictable way. In these architectures we try to eliminate all context switches for the same reason which are 100x cheaper.

Practically, 1µs stall is a good heuristic for a noise floor. The schedulers have pretty wide concurrency on big systems, so the implied 100 operations are unlikely to have a dependency. Many stalls that are difficult to precisely control like cache line fills fit in here too.

If there was a GC that had a worst-case stall of 1µs then you could probably use it for these cases. Unfortunately, "low-latency" GCs tend to be more like 1000x that. I don't think there is any way of closing that gap short of putting a GC in hardware.

Sampling bias. Most of the people responding are probably those with a strong opinion because of what they work on. Everyone else is likely relatively indifferent to it.

It is a misconception that GCs only affect latency-sensitive systems. High-performance throughput-optimized systems are also sensitive at ~1µs granularity for different reasons, so GCs are not used there either.

That a GC is adverse to the performance both latency-oriented and throughput-oriented workloads doesn't leave many use cases in "high-performance" systems. Maybe systems that are severely I/O bound but is barely a thing these days.

This was a major issue in the early days of organic high-explosives generally. The explosives and their precursors are often not just toxic but also biologically active. We still use a number of explosives as medicine because the effects on the body are potent. Many can also be absorbed through the skin.

It took a long time for industrial production to become sterile enough that risk of exposure was nominal. The early days of mass producing these chemistries was horrifically unhealthy. But then, so was the frontline. It was a different time.

It's literally the most sophisticated scheduling engine in the world.

That seems unlikely regardless of how good it is. This is a domain where state-of-the-art research is not in the public literature. Scheduling is an AI-complete problem.

OSS only dominates for software that is commoditized and the published computer science research for that software domain is close to the frontier.

OSS struggles at being relevant when software is non-commodity e.g. office suites. In software domains like databases where the state-of-the-art computer science research is often unpublished, OSS struggles to be relevant at the higher end of the market on technical merits.

When deciding what should be OSS, it is useful to consider the preconditions that have made it successful.

No one implied the intent was self-harm. The observation is that was the manifest result.

I’ve lived in serious poverty in many parts of the US, more than most people imagine exists. I’m not hypothesizing, I have first-hand knowledge. The evidence is so overwhelming that I shouldn’t have to play that card. It is disappointing the extent to which people will deny this reality for ideological reasons against all evidence.

Poverty is a mixed bag but the policies there, as with everywhere else, are defined by the worst cases. A significant subset of people in poverty cannot be trusted with their own welfare. That is a fact.

The widespread existence of secondary markets for SNAP benefits, which convert subsidized food for poor people into cash, carries this implication. This is pretty normal in some poor communities and that cash is commonly diverted to various vices. Some malnutrition is a consequence of this. Adding friction to the conversion of welfare benefits to cash is a feature.

You can't force poor people to spend cash on proper nutrition and a minority of them don't. It isn't a moral judgement but an observable fact. A lot of policies around welfare are targeted at trying to prevent this minority from slowly killing themselves in public.