HN user

kevmo314

3,204 karma

haha, computer go brrr

kevmo314@gmail.com

blog.kevmo314.com

[ my public key: https://keybase.io/kevmo314; my proof: https://keybase.io/kevmo314/sigs/jKuWflZVNiTzhIe0ttNcauvsm_uNCKaCk5U7Ni6SuFk ]

Posts85
Comments625
View on HN
github.com 1mo ago

Lupine: A GPU-over-IP Bridge

kevmo314
4pts0
github.com 2mo ago

Resolving Neighborhood Info with HTTP Range Requests

kevmo314
2pts0
www.flightsinasia.com 3mo ago

United's Unique Hub in the Pacific

kevmo314
1pts0
pluralistic.net 4mo ago

The whole economy pays the Amazon tax

kevmo314
9pts1
pion.ly 8mo ago

WebRTC Survives When You Walk Out

kevmo314
4pts0
codelabs.developers.google.com 1y ago

The 10th most common commit message is "can't you see I'm updating the time?"

kevmo314
3pts2
ninetyfive.gg 1y ago

Solving Character Prefix Conditioning with Coverage Beam Search

kevmo314
2pts0
github.com 1y ago

Show HN: Vibe HTTP – AI generated HTTP responses

kevmo314
9pts0
www.dwbowen.com 1y ago

Plant Machete

kevmo314
1pts0
www.youtube.com 1y ago

In search of the perfect dynamic array growth factor [video]

kevmo314
1pts0
github.com 1y ago

The Qubit Factory

kevmo314
1pts0
github.com 1y ago

Snake Diffusion

kevmo314
2pts0
gist.github.com 1y ago

Llama Inference in 150 Lines

kevmo314
6pts0
www.youtube.com 1y ago

Beating every possible game of Pokemon Platinum at the same time [video]

kevmo314
6pts0
www.youtube.com 1y ago

Rory Sutherland – Are We Now Too Impatient to Be Intelligent? [video]

kevmo314
2pts0
github.com 1y ago

Scuda – Virtual GPU over IP

kevmo314
207pts40
gist.github.com 1y ago

Please Stop Reinventing JSX

kevmo314
43pts89
github.com 2y ago

USB Video Camera (UVC) Devices with Go

kevmo314
1pts0
trekhleb.dev 2y ago

Content-aware image resizing in JavaScript (2021)

kevmo314
1pts0
www.youtube.com 2y ago

Do Bad Reviews Kill Companies? [video]

kevmo314
5pts2
www.cutercounter.com 2y ago

Free Visitor Hit Counter

kevmo314
1pts0
sites.google.com 2y ago

Federal Income Tax Spreadsheet

kevmo314
1pts0
kevmo314.github.io 2y ago

Show HN: Appendable – Index JSONL data and query via CDN

kevmo314
8pts1
www.ling.upenn.edu 2y ago

The Origin of the Terms Big-Endian and Little-Endian

kevmo314
6pts2
www.bzero.se 2y ago

How the append-only btree works (2010)

kevmo314
209pts100
github.com 2y ago

Show HN: Appendable – A Statically Hosted Database

kevmo314
1pts0
kevmo314.github.io 2y ago

Querying time zone data with HTTP Range Requests

kevmo314
2pts0
github.com 3y ago

Show HN: Meta's Segment Anything Model in a Chrome Extension

kevmo314
1pts0
www.youtube.com 3y ago

IETF Celebrates the Standards

kevmo314
2pts0
vimium.github.io 3y ago

Vimium – A browser extension that provides Vim-style keyboard controls

kevmo314
441pts209

This is not the first time we can see Nvidia taking shortcuts to achieve maximum performance of their GPUs

Why is implementing it correctly not performant? For context I have no idea how rounding is typically implemented anyways.

It's surely possible but if it's, for example, 10% slower, that easily eats into execution time and that directly translates into a sense of "maybe it's just worth it to pay the license fee for this year" after just a few 20h place and route runs.

Of course, if it were faster, that would be a huge win for the open source implementation.

The difficult part is the place and route algorithm, not the bitstream. The proprietary ones already take quite a long time to solve: I regularly have 12-24h runs. Perhaps an open source one could do better? But it's not quite as straightforward as reverse engineering a proprietary bitstream.

Same, the Apple silicon chips have been huge.

I bought a 2019 Intel MBP and that was by far the worst laptop I've ever had. After just a year of use it was constantly overheating and running out of memory and disk space, barely able to open a terminal. It was so bad that I hesitated to buy the Apple silicon versions, but the good reviews convinced me and it has been going strong ever since.

Are you writing this from the future? The latest gen nvidia gpus sit at around 2-2.5 GHz and the latest gen amd cpus sit 4-5 GHz.

That matches my personal experience too, writing naive cuda code that doesn’t take advantage of parallelism is roughly half the speed of running it on cpu.

a block of Rust threads that are properly programmed to take advantage of the vector processing by avoiding divergence

Sure, if you have that then of course it would be fast. But that’s not what this library is proposing.