HN user

hydroreadsstuff

235 karma
Posts16
Comments65
View on HN
blog.automaton2000.com 11mo ago

The Next Step for AI – Full Personal Interaction Capture

hydroreadsstuff
1pts0
github.com 1y ago

Measuring Apple CPU Core-to-Core Latency Without Core-Pinning

hydroreadsstuff
4pts2
twitter.com 1y ago

Biased AI Code Completion Example: calculateWomanSalary

hydroreadsstuff
1pts0
www.lynalden.com 5y ago

The Case for a Longer-Term Oil and Gas Bull Market

hydroreadsstuff
2pts0
www.youtube.com 5y ago

When Virtual Conferences Become Video Games: It's the Future – Ian Cutress [Yt]

hydroreadsstuff
3pts0
www.masterworks.io 5y ago

Masterworks – Learn to Invest in Fine Art

hydroreadsstuff
1pts0
youtu.be 6y ago

Long-term Covid19 cases – Clinical Features

hydroreadsstuff
4pts0
www.ecb.europa.eu 6y ago

ECB announces package of temporary collateral easing measures

hydroreadsstuff
3pts0
www.dw.com 6y ago

Angela Merkel to quarantine after meeting infected doctor

hydroreadsstuff
1pts0
de.wikipedia.org 6y ago

Socially Acceptable Early Death

hydroreadsstuff
1pts0
en.wikipedia.org 6y ago

Helicopter Money

hydroreadsstuff
2pts0
en.wikipedia.org 6y ago

List of OECD countries by hospital beds

hydroreadsstuff
33pts11
www.youtube.com 6y ago

Stanford Seminar – Centaur Technology's Deep Learning Coprocessor

hydroreadsstuff
1pts0
www.youtube.com 6y ago

Designing Path of Exile to Be Played Forever. Chris Wilson at GDC 2019

hydroreadsstuff
1pts0
www.anandtech.com 6y ago

Next-Gen Nvidia Teslas Due This Summer; to Be Used in Big Red 200 Supercomputer

hydroreadsstuff
1pts0
pages.stern.nyu.edu 6y ago

Annual Returns on Stock, T.Bonds and T.Bills: 1928 – Current

hydroreadsstuff
74pts77

The 4x comes from the neural accelerators (tensor core in NVIDIA jargon). It's 4x fp16 over the vector path (And 8x compared to M1 because at some point they 2x'd the fp16 vector path). Therefore LLM prefill(context processing/TTFT), diffusion models (image gen), and e.g. video and photo effects that make use of them can be up to 4x faster. At fp16 that's the same speed at the same clock as NVIDIA. But NVIDIA still has 2xfp8 and 4xnvfp4.

Batch-1 token generation, that is often quoted, does not benefit from this. It's purely RAM bandwidth-limited.

Some companies like to stress the efficiency or performance of Arm SoCs, but really this is a hedge against more expensive x86 hardware. AMD has increased prices of mobile SoCs radically recently. I'm looking forward to having more affordable SoC options for laptops, handhelds and desktops, perhaps from Mediatek or other lower-cost vendors.

The history of the PC is one of commoditization. A fractured multi-polar landscape is detrimental to the ecosystem/productivity and should ultimately fail.

x86 emulation is an important puzzle piece, and I'm happy Valve recognizes this and sponsors it.

The Llama 4 herd 1 year ago

This means GPUs are dead for local enthusiast AI. And SoCs with big RAM are in.

Because 17B active parameters should reach enough performance on 256bit LPDDR5x.

I wish the author would have put some actual numbers on the cost of the different technologies. Arguing about tradeoffs including that information would be much more meaningful.

The system is set up to avoid such shenanigans. If desired, the question is will they find a way to work around the rules. I agree with you, but then the trust in the system would be greatly damaged. They might as well have two different interest rates / not pay IOR.

Afaik the needed difference in interest payment from the fed comes from the treasury. So it’s indirectly connected to the budget. In theory though you can still raise more debt to pay the interest. But I’m not sure about the longterm consequences of this.

Having inflation above the interest rate helps decreasing the debt/gdp ratio.

I hope we will see this on more devices. This is a huge boon to performance.

Might even forebode soldering RAM onto packages from here on out and forever.

Steamdeck will probably have a crazy 100gb/sec ram b/w. twice of current laptops and desktops.

How do they get 200/400GB per second RAM bandwidth? Isn't that like 4/8 channel DDR5. 4/8 times as fast as current Intel/AMD CPUs/APUs? (E.g. https://www.intel.com/content/www/us/en/products/sku/201837/... with 45.8GB/s)

Laptop/desktop have 2 channels. High-end desktop can have 4 channels. Servers have 8 channels.

How does Apple do that? I was always assuming that having that many channels is prohibitive in terms of either power consumption and/or chip size. But I guess I was wrong.

It can't be GDDR because chips with the required density don't exist, right?

As someone who worked in HPC Projects in the EU, I'd say the vast majority of the grants is wasted. E.g. on

- obscure software and programming paradigms/libraries

- one-off's

- projects that create only reports, papers and recommendations

- endless recreations of the same software over the years and decades

- plain dumb projects

- attempts at recreating other people's software (e.g. U.S.)

There is also a large amount of overlap in work/content between projects, and lots of time is appropriated for unrelated work e.g. employees working on their PhDs.

On the other hand the industry takes a lot more money and plays similar games.

Project partners do not take the right path, because it's not conducive to fulfilling a grant, getting the next grant or increasing your citation count. The incentives are wrong.