HN user

emcq

860 karma

http://emmettmcquinn.com/blog/

Posts5
Comments284
View on HN

If we are going to consider using prior runs of the program having the file loaded in RAM by the kernel fair, why stop there?

Let's say I create a "cache" where I store the min/mean/max output for each city, mmap it, and read it at least once to make sure it is in RAM. If the cache is available I simply write it to standard out. I use whatever method to compute the first run, and I persist it to disk and then mmap it. The first run could take 20 hours and gets discarded.

By technicality it might fit the rules of the original request but it isn't an interesting solution. Feel free to submit it :)

The 1.4s is _after_ having the file loaded into RAM by the kernel. Because this is mostly I/O bound, it's not a fair comparison to skip the read time. If you were running on a M3 mac you'd might get less than 100ms if the dataset was stored in RAM.

If you account for time loading from disk, the C implementation would be more like ~5s as reported in the blog post [1]. Speculating that their laptop's SSD may be in the 3GB/s range, perhaps there is another second or so of optimization left there (which would roughly work out to the 1.4s in-memory time).

Because you have a lot of variable width row reads this will be more difficult on a GPU than CPU.

[1] https://www.dannyvankooten.com/blog/2024/1brc/

This is even slightly more direct: access to WSJ data requires paying LDC for the download, and the pricing varies depending on what institution / license you're from. The cost may be a drop in the bucket compared to compute, but I don't know that these licenses are transferable to the end product. We might be a couple court cases away from finding out but I wouldn't want to be inviting one of those cases :)

Be wary of using this model - the licensing of this model seems sketchy. Several of the datasets used for training like WSJ and TED-LIUM have clear non-commercial clauses. I'm not a lawyer but releasing a model as "MIT" seems dubious, and hopefully OpenAI has paid for the appropriate licenses during training as they are no longer a research-only non profit.

Shockingly the human genome itself has not been fully sequenced, despite the human genome project completing years ago [0]. There are difficult to map regions of the genome, some of which are interesting. Only recent advances in long read sequencers have helped to solve some of these issues [1].

For the future to truly be amazing with one sequencing the lab prep, chemistry, and equipment required needs to advance. Oxford Nanopore has some advancements here [2] but it's still a ways to go before you could have a sample prepared as easily as an ultrasound or x-ray.

[0] https://www.statnews.com/2017/06/20/human-genome-not-fully-s... [1] https://www.ecseq.com/support/ngs/are-there-regions-in-the-g... [2] http://nanoporetech.com/products/voltrax

I used to work for a drone company. I made sure to go watch the mechanical testing for prop safety. One of the test objects was a chicken leg, and the prop cut through more than 2mm of bone. The props can do some serious damage!

I don't know what, if any, approaches are implemented but one solution is that you can design a safety feature in the motor control to turn off the motor if it detects blockage.

This article is quoting employees conversations from 2016. It wasn't until late 2016 that Google published research on wide and deep recommender systems. I'd be interested to see if people still think behavioral data is poor and ML has not provided advances.

That said some demographics are broadly correct from behavioral tracking (e.g. a child who watches baby shark on repeat), but good luck telling the difference between a rich child and a less rich child.

You can get all this functionality and more working beautifully with Android with Garmin watches. Mine lasts for many days on a charge (perhaps a week without activity tracking), has better sport tracking features, but the screen isn't quite as good. There are currently a few good deals on the Vivoactive 3, but also many higher end models available too.

Elon's was quoted saying "you have a forest of redwoods and the little trees cant grow." While pointing his finger at California, I can't help but think this really reflects back on him.

Elon Musk's success and several others began as the result of PayPal.

To my knowledge none of Elon's companies since have produced a group like the PayPal mafia. The PayPal folks are smart people, but the bay area is full to the gills of smart visionaries lacking capital to take on ambitious visions. Many have worked at Elon's companies. To really see little trees grow people like Elon would need to change their equity structures to let the next group of innovators thrive.

AirPods Max 6 years ago

It feels like we are seeing the result of diminishing returns with technology advancement.

To make a product that Apple believes is significantly better than the competition they had to design a very intricate solution that includes:

* High end look and feel not similar to the bulk of their products with lots of textile webbing and memory foam ear cups. These get way more wear and tear than normal electronics being exposed to sweat, sunscreen, etc. On top of that these must be safe for long skin exposure and comfortable across many head shapes and sizes. * High quality magnets with custom speaker design for low THD and large frequency range. * 2 custom ASICs built for sound processing and low power bluetooth and 10 audio cores each. * 10 microphone array for ANC and wind noise cancelation. * Multiple accelerometers + gyros head tracking with spatial audio.

It you remove one of those components, I'd be surprised if Apple still shipped this. At least 40 people worked on designing and engineering this new headphone. There are not many companies in the world with the right kind of talent for this and Apple happens to be one of them.

Google colab is a much better tool than a new computer for getting into ML. It's free, requires no dependency management, easy to share, and has tons of example notebooks you can reproduce as easily as duplicating a google doc.

Not to mention a Linux based workstation in the limit will have fewer headaches than mac these days. Package management doesn't require homebrew or dockerized everything, selinux is surprisingly easier to configure than the Mac security subsystems, etc.

You can make a model of any size with deep learning.

If your concerns are about over fitting there are lots of regularization techniques used in practice like dropout, weight decay, and data augmentation.

There's nothing preventing you from sharing weights across layers, and would be interesting to see some research about that.

A better rule of thumb is miles because hours don't reflect the amount of energy put into your drive train.

Typical recommendations are every 500-750miles for a mountain bike. For the average mountain biker I think that would be more like every 125 hours at 6 mph. Mountain bike chains see much more abuse than road, and you can go further too.

You can be more scientific by using a chain ware tool to measure how much stretch there is.

You can't really make all things be equal and get increased efficiency and longetivity for most cases.

The trade-off are explicitly made to have reduced longevity to have better efficiency.

Take tires for example. A low rolling resistance lightweight XC tire will have significant advantages over a dual casing 2.6" DH tire. The XC tire will smoke a DH tire but not have better longevity on dirt. The XC tire will be slower on some DH tracks and the DH will be slower on XC. If you ride a DH tire on a road it will ironically wear down more quickly than a typical XC tire, and you can shred an XC tire in one day on technical DH.

Years ago I worked at a small, like 5-7 person startup in soma where we were surprised to find Rupert himself personally come by, later send some chiefs to do due diligence, and cut us a check. We worked with a few of his brands after that.

It didn't feel different than any other VC experience I've seen as an employee, whether that was A16z, Sequoia, etc. If it looks like a duck...

Blaine Washington is right along the Canadian border and next to many islands. The islands are sparsely inhabited and infrequently visited by humans it makes me wonder if maybe there is still a colony out there. It seems like a difficult area to do containment so glad they tracked these down.

It looks like in 2019 Vancouver island (which is huge) had a colony get eradicated, but otherwise haven't seen much else reported.

https://www.ontario.ca/page/asian-giant-hornets

A common approach in rendering engines to convert screen space coordinates to objects is to render a second image with light and shadow disabled where the color uniquely maps to an id. You then can uniquely identify 24 bits worth of objects without needing to maintain a KD tree.

The most accurate description is probably the authors built a hearing aid from the 1960s with only $1 today [0]. Its worn around the neck, doesn't have any fitting, doesn't appear to have beamforming or feedback management, and low amplification.

It turns out these things are all important in clinical outcomes for listening comfort and intelligibility. I've built a modern hearing aid and done extensive patient testing - the details matter and it's not just electrical or algorithm problems but tough mechanical and UX challenges to get something more state of the art.

While I haven't worn many neck mounted devices, the Bose Headphones can get crazy feedback and the feedback path is similar. This would cause squealing discomfort for patients and everyone around them.

[0] https://en.wikipedia.org/wiki/History_of_hearing_aids

Unfortunately live listen is more like a replacement for remote microphones, not the hearing aids themselves. It adds around 70ms latency with a bluetooth headset. This is enough to be very uncomfortable. For comparison, most hearing aids on the market today have around 8ms of end to end latency.

Source: I've built a hearing aid and done extensive latency tests :)

I was surprised to see no discussion about environmental impacts.

One nasty consequence of non organic canola oil is that the industry largely uses glycophosate for pest control. The organic version of Oatly does not use canola oil, and Califa uses sunflower oil (and has less sugar!).

And yes, industrial cattle practices suck. But the cattle themselves aren't the inherent problem. In 1800 there were approximately 60 million wild Buffalo in North America. Today there are about 90 million cows, of them about 30 million used for milk. The environment has the capacity to support a similar order of magnitude of bovines sustainably. Sustainable cattle farming in some cases is even carbon negative [0]. Bovines historically played a big role in the grassland ecosystem but what industrial scale farms are doing look nothing like nature.

It's easy to vilify industrial scale cattle farming but it's not truthful to say that there are no ways to have environmentally friendly cattle farms.

[0] https://www.whiteoakpastures.com/meet-us/environmental-susta...

For small networks it's often a win to stay on chip at least on the power side. But if you do need to go off chip for memory it's hard to beat the memory bandwidth you have on a GPU.