HN user

lightcatcher

1,252 karma

http://ericmart.in (not much here)

https://github.com/lightcatcher

http://www.linkedin.com/pub/eric-martin/34/636/628/

My email is protected behind this captcha: https://www.google.com/recaptcha/mailhide/d?k=01d1UiHRvpkZRXt99xIP9iGw==&c=Klp_6m19BlnmignxY1w00mYErHe263koplO9E_NWDbo=

Interested in startups, machine learning, and high performance computing.

Posts22
Comments257
View on HN
www.wired.com 9y ago

Uber Hires Raquel Urtasun in the Quest to Rehab Its Future

lightcatcher
5pts0
www.wired.com 9y ago

Google Opens Montreal AI Lab to Snag Scarce Global Talent

lightcatcher
1pts0
www.wired.com 10y ago

How GM Beat Tesla to the First True Mass-Market Electric Car

lightcatcher
10pts2
www.wired.com 10y ago

Obviously Drivers Are Already Abusing Tesla's Autopilot

lightcatcher
10pts0
www.fastcompany.com 10y ago

Inside Pinterest's Plans to Fix Its Diversity Problem

lightcatcher
1pts1
www.wsj.com 11y ago

Mixpanel (YC S09) Raises $65M to Build Predictive Data Tech

lightcatcher
136pts27
recode.net 12y ago

Mixpanel: How addictive is your app?

lightcatcher
2pts0
www.wired.com 12y ago

Love, Actuarially

lightcatcher
6pts1
zinkov.com 12y ago

Vowpal Wabbit for the Uninitiated

lightcatcher
3pts0
blog.mixpanel.com 12y ago

ECommerce activity on Black Friday and Cyber Monday

lightcatcher
1pts0
blinkdb.org 12y ago

BlinkDB: Queries with Bounded Errors and Response Times on Very Large Data

lightcatcher
85pts4
www.bizjournals.com 12y ago

Jeff Dean on how neural networks are improving everything Google does

lightcatcher
1pts0
qz.com 13y ago

Four apps that would make Google Glass truly useful

lightcatcher
4pts0
blog.ionelmc.ro 13y ago

Python debugging tools

lightcatcher
13pts0
mitadmissions.org 13y ago

Meltdown, a description of life at MIT

lightcatcher
4pts1
jacobtracey.com 14y ago

Lessons from the porn industry on how to evolve content distribution

lightcatcher
6pts1
media.caltech.edu 14y ago

Neuroscientists Find That Status within Groups Can Affect IQ

lightcatcher
2pts0
media.caltech.edu 14y ago

Elon Musk to Deliver Caltech Commencement Address

lightcatcher
1pts0
www.economist.com 14y ago

Where Rats and Robots Play

lightcatcher
2pts0
www.newyorker.com 14y ago

Brain Gain: The Underground World of Neuroenhancing Drugs

lightcatcher
124pts109
www.quora.com 14y ago

Reflections on Interning at a Startup

lightcatcher
2pts0
techcrunch.com 14y ago

Mixpanel Now Funnels Into The Past

lightcatcher
34pts20

I have no experience in PNG encoding, but found https://github.com/brion/mtpng The author mentions "It takes about 1.25s to save a 7680×2160 desktop screenshot PNG on this machine; 0.75s on my faster laptop." which makes me think your slower performance on smaller images either comes using the max compression setting or using hardware with worse single threaded performance.

Although these don't directly solve the PNG encoding performance problem, maybe some of these ideas could help?

* if users will be using the app in an environment with plenty of bandwidth and you don't mind paying for server bandwidth, could you serve up PNGs with less compression? Max compression takes 15s and saves 35MB's. If the users have 50mbit internet, then it only takes 5.6s to transmit the extra 35MB, so you could come out 10s ahead by not compressing. (yes, I see your comment about "don't say to use lower compression", but no reason to be killed by compression CPU cost if the bandwidth is available).

* initially show the user a lossy image (could be a downsized png) that can be quickly generated. You could then upgrade to a full quality once you finish encoding the PNG, or if server bandwidth/CPU usage is an issue then you could only upgrade if the user clicks a "high-quality" button or something. If server CPU usage is an issue, the low then high quality approach could let you turn down the compression setting and save some CPU at the cost of bandwidth and user latency.

I went to high school in a science magnet program in Texas. Students from 3 high schools were eligible for the magnet program, and the magnet program was housed in one of the 3 high schools. Our math+science classes were in the magnet program, but our English/history/PE/art/all other classes were in the host school. The program made up about 10% of the host high scool, and students in the program were counted as part of the host school for purposes of university admission.

This understandably made people unhappy at the host school - ~7% of the class is academic high achievers from out of the school zone who take most of the admission spots reserved for the top N%.

I don't know of any cases of parents moving to avoid the extra competition, but I probably wouldn't have heard of that if it happened. I do know of some people set on going to UT who did not apply to the magnet program so they could have less competition.

The point here is you don't need to move to a rural area to decrease school competition. There are plenty of cases where you can move a mile to get into the zone of a less competitive school.

Parallel prefix sum is the most underappreciated parallel algorithm in my opinion, and this paper is the best explanation and visualization of the concept I've seen.

A few years ago I worked on a deep learning project using parallel prefix sum as a new way to accelerate recurrent neural nets on GPUs[0]. The paper in this post was the most important reference and source of inspiration. I'm happy to see this paper shared on HN in hopes that it also sparks ideas in others.

[0] https://arxiv.org/abs/1709.04057

How much wealthier would you be if you paid no taxes at all in tax year 2020?

An example with some made-up numbers that could apply to some HN users:

Let's say your net worth was $500K at the start of the year, you earned $200K income in the year and spent $50K. Without taxes, your end of year net worth would be $650K. However, you pay $40K+ in taxes, which makes your net worth <=$610K. So effectively you paid $40K/$650K = 6.1% of your wealth in taxes.

"Regular" people build wealth through income, while wealthy build wealth through appreciating assets. The point I take from headlines like these are not "US executives illegally avoid taxes", it is "tax rules favor the wealthy". Increasing tax on appreciated assets by raising capital gains rates, removing step-up basis, or (maybe) taxing unrealized gains could shift some of the tax drag on wealth from income earners towards asset holders. All of these would need to be done very carefully to not overly hurt small business owners, perhaps through something like a lifetime capital gains exemption (similar to gift exemption, apparently existed in Canada in the 80s[0]).

[0] https://static.twentyoverten.com/5b9280ab0420c067d6b36505/Wo...

I've found Walmart.com to be about as good as Amazon for my online shopping (in the US). I particularly find Walmart to be a lot better for some dry goods like cereal and Clif bars. They can mix delivery from their warehouses and from local stores.

This is not me shilling Walmart. I've been pleasantly surprised by it in the last year, and find it to be a real competitor to shopping at Amazon.

I'm a runner who never goes on an "exercise walk", but there are plenty of (non-injury) reasons someone might prefer walking:

- less strenuous

- much less fitness required to walk for an hour than to run for 30 minutes

- easier to avoid sweating while walking, which can be useful for some commutes

- easier to talk on the phone while walking

- easier to find someone to walk with you than it is to find someone who will run with you

The article was about making fitness easy and not about making fitness efficient. I 100% agree that running is more efficient, but I'd recommend walking to anyone who thinks "bleh/eww" at the thought of exertion.

I also find the "or" wording of the law interesting.

I do think it's racist as it grants the privilege of abandoning the Jewish religion while remaining a legally privileged class (Jew) to people with some ancestries (Jewish) but not with others.

Parler is the right starting their own company in response to Twitter censoring/banning some of their discussions. Parler's current deplatforming indicates that "start your own company so you're the governor" is a borderline unrealistic point of view, since the internet and especially mobile phones are run by mega-corporations that don't support an open platform.

If we look at Parler being dependent on other companies, there are 3 dependencies that bit them in the last few days:

(A) running servers on AWS

(B) distributing Android app through Google Play Store

(C) distributing iPhone app through Apple Store

Parler could and should have avoided the dependency on AWS by hosting with a company more aligned with their views or outside of the US.

However, I cannot in good faith say "Parler should have written their own mobile OS if they want people to use their platform from their phones". I believe Android supports app installation without Google Play. It's a little inconvenient, but I think an extra minute or two of clicking around is an ok price to pay for access to unpopular speech.

However, as far as I know there's no way to have an app on iPhones without Apple's approval. This is a strong point for "even if you start your own company, you're not the governor - 45% of America's phone users can only access your service at the mercy of Apple".

Fairness, Accountability, Transparency (aka FAT) is a real sub-field of machine learning, that is as technical as any other machine learning sub-field in my opinion. It is not a PR stunt. I've published at top ML conference a paper about some CUDA kernels I wrote to accelerate training a class of RNNs, so I feel have technical grounding to make these claims.

Here are some FAT papers I've enjoyed: https://arxiv.org/abs/1806.08010 https://arxiv.org/abs/1802.04023 https://arxiv.org/abs/1803.04383 (won best paper award at ICML 2018)

Many FAT papers are published at NeurIPS or ICML (generally considered top two machine learning conferences). There's also a conference just on the topic: https://facctconference.org/

2015 BS alum

Trolling, flicking - I never heard anyone say these, but had heard they were used in the past. Along those lines, flaming or flaming out was used for failing out or needing to take terms off.

A few people still did 3AM Tommy's. Negative time Tommy's was a significant event in at least a few houses.

The Ride tradition was definitely around, with the only sanctioned playing on mornings of finals

I never heard anyone say finesse, brute force, or ignorance

I believe there's work to be done for each new CPU architecture (Broadwell, Skylake (AVX-512), Cascade Lake, let alone ARM or other architectures). The code needs to be updated for things like L1 cache size, number of registers per core, and number of adders per core. So there will likely continue to be frequent work on BLAS implementations until there's some very smart optimizing and profiling compiler (which is related to what ATLAS does I think).

It feels as if the point that I'm trying to make is that mindful archiving is a better solution than to just 'keep all the things'

On the topic of plain text things (such as text messages) - how much data are you actually hoarding?

Let's say you type 100 words per minute for the next 40 years (and each word is 10 bytes). No sleep, no breaks, just 40 years of typing. Congratulations, you just produced 21GB of data. This fits on an SD card (<$30) or in the cheapest tier of cloud backup like Dropbox or Google Drive. You can search your 40 years of typing in well under a minute. If you remember the year you typed in, you can grep the data from that year in under a second.

I don't like the term "hoarding" for this. Hoarding has a negative connotation. Storage of plaintext is so incredibly cheap (and search so fast) that I feel that option value of retaining the text is almost always greater than the miniscule cost of storage and slower retrieval.

I don't think are any valid analogies between storing physical items and digital items, as digital storage and search is orders of magnitude cheaper. Consider the same experiment where one writes with pen and paper for 40 years, and then wishes to search for the name "George".

Making a decision of what to keep must be more expensive and time-consuming than just keeping everything.

Yes.

There is no way to move messages from one iOS device to another (such as a new phone). My girlfriend recently got a new iPhone and wanted to transfer our Signal message history from her old iPhone onto the new one. She said it wasn't possible, and then I spent an hour or two reading about it figuring there must be some hacky awful way to accomplish it. I couldn't find one. This has been an open issue for years [0][1].

Android has an inconvenient backup flow (that involves randomly generated 30 digit PIN and manual transfer of file), but that's infinitely better than the total lack of options on iOS. I do wish Android (and iOS) had a method to download all message history to decrypted plaintext (or JSON) for use with other apps. If I own my data, decrypting it should be my choice.

I regret recommending to my girlfriend that we use Signal, and won't recommend Signal to more people after this.

[0] https://github.com/signalapp/Signal-iOS/issues/2542 [1] https://whispersystems.discoursehosting.net/t/ios-backup-kee...

Agreed! I'm on an eastern edge of a timezone and DST ending moves sunset forward from 6PM to 5PM (and up to around 4:20PM during winter solstice). This feels like doubling down on how much winter sucks: not only is it cold, but it's now not possible for me to exercise outside in daylight after work

a company that is doing serious core refactoring/redevelopment on the back of an underpaid "intern" is probably exploiting that labor in an unfair way

Handing out throwaway projects like you posit is an easy trip to a class action suit.

So if an intern shouldn't do work on the core product but also shouldn't do throwaway projects, what do you suggest an intern should do? What do you consider an "underpaid" intern? Would you consider 75% salary of campus new hire as underpaid?

I think of it as a trade-off between the tool perfectly designed for the job and the tool that can be used for the job of which I'm already an expert user.

As an example, my current team's codebase is Python and C++. When we need to do some basic Linux scripting (check if this file exists, if not send an error email), the two prime candidates are Bash and Python. Bash might be exactly designed for this type of thing, but a lot of the team would need to Google stuff like "bash logical and of two booleans" for Bash where they already know the Python syntax. My current rule of thumb here is "<10 lines of code: use Bash. else use Python". More generally, I often try to err on the side of using the tool I know than the unknown tool that might be perfect for the problem.

I was an intern at Mixpanel at the time (but not the author of this article or involved in the discussed rewrite). For context, the company was then <10 people. (I'm surprised to see anything from my first professional software experience on top of HN!)

My recollection of events (of 8 years ago): The Erlang endpoint was written a year or two earlier as someone's first Erlang project, in a tiny org with no Erlang experience. The endpoint worked well enough and people moved onto other projects in the mostly Python codebase. Eventually, it became painful to debug or add features to this endpoint because no one was particularly proficient in Erlang. Ankur (author of the post, intern) then rewrote it in Python.

I wouldn't read this as a negative about Erlang or a positive article about Python. I'd instead read it as "use a language that makes sense for your organization and already existing codebase".

The typical neural net matrix multiplication is N_EXAMPLES X N_FEATURES_IN multiplied with N_FEATURES_IN X N_FEATURES_OUT.

The output feature count is completely independent of the data size, and input feature count is only dependent on the dimensionality of the data (not the number of points), and that's only in the first layer of the network. Even with datasets with huge number of examples, the net usually only trains on a small "minibatch" of examples at a time, typically somewhere between 16 and 1024. This minibatch size is the algorithmic N_EXAMPLES. Given these numbers, the typical neural net matrix multiplication is probably something like (32, 256) x (256, 128). This is not nearly large enough for non-N^3 tmatmul algorithms o accelerate things.

Deep learning mostly uses matrices with largest dimension <1000. The size of the matrix directly corresponds to the number of input and output units to to a fully connected layer.

Even a 2000x2000 matrix is relatively small as far as non O(N^3) matrix multiplication algorithms go. The faster asymptotic algorithms have larger fixed costs and do not map as nicely to current CPU/GPU architectures. Additionally, the people capable of writing high performance matrix multiplication mostly haven't spent time on the non-N^3 algorithms. I believe every common implementation (MKL, cuBLAS, Eigen, and various other BLAS's) all use the N^3 algorithm with optimizations more focused on computer architecture (cache and register blocking, limiting instruction dependencies, SIMD) than on algorithmic cleverness.

The only attempt that I know of to use a non-N^3 algorithm for practical matrix multiplication is this 2016 paper[0] called Strassens' reloaded. Figure 5 shows that their implementation of Strassen's outperforms MKL on a single core on about a (2K, 2K, 2K) matrix multiplication problem (look at upper right part of figure). However, you're rarely multiplying on a single core. Figure 6 shows that for a 10 core system, Strassen's becomes marginally faster with square matrices of size 4K, and only becomes significantly faster at size 8K. Although these results are super cool in my opinion, they're mostly not applicable to deep learning in it's current incarnation. Finally, Strassen's algorithm is one of the simplest non-N^3 matmul algorithms (known since 1969, complexity O(n^2.8)) and I believe it has much lower fixed costs than something more recent that has complexity more like O(n^2.37).

[0] Strassen's Reloaded paper: https://www.cs.utexas.edu/~jianyu/papers/sc16.pdf

Chicago definitely does not "absolutely require" a car. I've lived here for a few years without one, and 80-90% of my friends in the city don't own one.

The trains don't go everywhere, but train + bus + bike share + Lyft make it very easy (and probably more convenient) to not own a car.

Most of the time when people talk about "Houston" they're really talking about Houston + suburban sprawl. Pearland to Kingwood is 40 miles. The area is huge, and there aren't common trip starts and trip ends.

I'll compare to two cities I do know that have more public transit:

* SF/Bay area handles similar distances with BART and Caltrain. Try to imagine the bay area (and transit) if there was no bay, the western part of the peninsula was as heavily developed as what's along the bay, and the metro area went further inland. Although the bay causes many transportation woes, it does linearize transit which means more overlapping commute patterns

* Chicago. Buses and trains within the city supported by density and the huge fraction of jobs in the Loop area (small geographic area, huge fraction of jobs). Metra rail serves the suburbs, and taking Metra generally involves driving to a Metra station.

At the press conference, city officials shared 80 exabytes worth of heart ultrasound videos, according to one company that participated.

Do you know anything about this? 80 exabytes is an insane amount of storage and is orders of magnitude larger than any dataset I'm familiar with. I believe Common Crawl is in the petabytes range, and [0] describes training a neural net on 300 million images (which would be 300PB if each image was 1MB, but I suspect images are smaller). 80 exabytes is 80 million terabytes and would cost hundreds of millions to store.

My guess is that some journalistic error occurred here, or perhaps someone confused "80 exabytes of data are generated during heart ultrasounds" with "we've stored 80 exabytes of heart ultrasound data".

I'd argue that a neural net is "black-box" in the sense that nobody really can give a coherent answer to "what happens if I perturb/double/negate this parameter" where the parameter might be deep in some weight matrix. Maybe this isn't a useful question because of the distributed representations within neural nets, but it is at least an answerable question for other models.

Do you know of any work on interpreting neural nets that are being used for non-image tasks?