HN user

brianchu

2,341 karma

Looking for the next thing.

Worked on machine learning for medicine at Cardiogram. Studied CS at UC Berkeley. Former undergrad researcher and teaching assistant for Berkeley's ML class. Former intern at Twitter on Ads Analytics Infrastructure.

Email: bc@brianchu.com

github.com/bchu

twitter.com/brrrianchu

Posts124
Comments300
View on HN
www.bloomberg.com 8y ago

Coinbase Says It Was Wrong About SEC Approval of Acquisitions

brianchu
2pts0
medium.com 8y ago

The Vanity Fair 'Brotopia' party was way worse than it sounds

brianchu
27pts9
blog.google 8y ago

Earth to exoplanet: Hunting for planets with machine learning

brianchu
1pts0
techcrunch.com 8y ago

Palmer Luckey’s new company Anduril interested in AR and VR on the battlefield

brianchu
2pts0
techcrunch.com 8y ago

Activehours raises $39M for its new take on cash advances

brianchu
2pts0
www.cnbc.com 9y ago

Zenefits' controversial former CEO gets $7M for new startup

brianchu
3pts0
techcrunch.com 9y ago

Zenefits founder Parker Conrad takes another crack at HR onboarding

brianchu
1pts0
techcrunch.com 9y ago

Uber will apply for a self-driving test permit in California

brianchu
1pts0
www.propublica.org 9y ago

When Evidence Says No, but Doctors Say Yes

brianchu
3pts0
medium.com 9y ago

Toxicity and Tone Are Not the Same Thing: Analyzing the Google API on Toxicity

brianchu
2pts0
www.nytimes.com 9y ago

Snapchat Founders’ Grip Tightened After a Spat with an Early Investor

brianchu
111pts67
spectrum.ieee.org 9y ago

The 2,578 Problems with Self-Driving Cars

brianchu
2pts0
www.cnbc.com 9y ago

Waymo's self-driving cars improve performance in California tests

brianchu
1pts0
www.axios.com 9y ago

Cisco spikes AppDynamic's IPO

brianchu
1pts0
www.buzzfeed.com 9y ago

Introducing Pound: Process for Optimizing and Understanding Network Diffusion

brianchu
2pts0
www.bloomberg.com 9y ago

Food Delivery Startup Munchery Cuts Staff, Parts with Founders

brianchu
3pts0
www.nytimes.com 9y ago

F.B.I. Arrests Volkswagen Executive on Conspiracy Charges in Emissions Scandal

brianchu
2pts0
research.googleblog.com 9y ago

Equality of Opportunity in Machine Learning

brianchu
1pts0
research.google.com 9y ago

Attacking discrimination with smarter machine learning

brianchu
2pts0
fortune.com 9y ago

Bed Bath and Beyond Paid Just $12M for One Kings Lane

brianchu
14pts6
medium.com 9y ago

Introducing Drive.ai and a New Vision for Self-Driving

brianchu
3pts0
avc.com 9y ago

Checking in on Chat Bots

brianchu
4pts2
techcrunch.com 9y ago

GM expressed interest in buying Lyft, but Lyft declined

brianchu
1pts0
www.technologyreview.com 10y ago

Tougher Turing Test Exposes Chatbots’ Stupidity

brianchu
2pts0
www.youtube.com 10y ago

CVPR 2016 keynote: Autonomous Driving

brianchu
2pts0
thehill.com 10y ago

Comey: Hacker Guccifer never gained access to Clinton server

brianchu
2pts0
www.youtube.com 10y ago

My Four Months as a Private Prison Guard: Part One

brianchu
4pts0
xeniaschmalz.blogspot.com 10y ago

What happens when you try to publish a failure to replicate in 2015/2016

brianchu
332pts137
www.motherjones.com 10y ago

My Four Months as a Private Prison Guard: A Mother Jones Investigation

brianchu
40pts1
fortune.com 10y ago

Andreessen Horowitz Raises $1.5B for New Fund

brianchu
77pts35

I have the opposite assumption: the raw data is usually more reliable than an editorial. And I confirmed it: the raw table is more reliable since it is the same survey across different years. The article is inconsistently comparing numbers from 2 different surveys. The 2018 figure (2.63) is from the "American Community Survey" and the 2010 figure (2.58) is from "Census SF1 data".

But the 2010 "American Community Survey" says the average houshold size is 2.63 (https://data.census.gov/cedsci/table?q=b25010&tid=ACSDT1Y201...), so for this survey the trend is flat.

Your example is hogswash. Absolutely, the perceived long term value dropped 30% as people feared a million deaths (with lockdown) and dead bodies piling up outside hospitals across the country (with lockdown, and not just New York City). The stock market is rising now that people realize the pandemic, while still bad, isn't going to be as bad as those predictions. Our perception/understanding of the pandemic has rapidly changed.

Some models were predicting multiple hundreds of thousands of deaths, with lockdown. The Imperial College model was predicting 1 million deaths, with lockdown. I completely agree cases will rise as places begin reopening. But whatever the outcome it will be with reopening - still better than what the market expected in March.

Yes, NYC was the only place in America where the system was close to overrun and some hospitals actually were overrun, I'm not disputing that.

The explanation is very simple. The pandemic is not nearly as bad as people thought it would be in March.

Models were predicting hundreds of thousands of deaths in the USA over the next few months, with lockdown. Many people were predicting hospitals would be widely overrun in New York City, parts of California, etc (again, with lockdown). These models and predictions, of course, were wrong.

Printing money and stimulus should have been expected (given the government's response in 2008) and therefore priced in, at least in theory. If we actually had massive numbers of bodies piling up outside hospitals in all major US cities, no amount of money printing would have propped up the markets.

The WHO is no longer a trustworthy source. But this is still (mostly) correct, and it's not that hard to dive into the studies directly instead of appealing to authorities. All indications are that it is possible for SARS-Cov-2 to be airborne but it is rare.

The main study from Wuhan that people cite for airborne SARS-Cov-2 only found high levels of airborne SARS-Cov-2 in poorly ventilated areas of a hospital setting (where certain medical procedures like intubation are known to generate aerosols): https://www.nature.com/articles/s41586-020-2271-3. Well-ventilated areas had very low levels. A separate study in Singapore found no airborne samples in a hospital setting (https://jamanetwork.com/journals/jama/fullarticle/2762692). Documented spreading events are consistent with the disease not being aerosolized (one example: https://twitter.com/zeynep/status/1251556084424347649).

Yep. Obviously this is anecdotal and limited to my own social circles, but I live in the Bay Area, and almost everyone (90%+) I know in San Francisco does not own a car. The ones that do own cars only use them if they need to get out of the city (within the city they bike/scooter/walk/Uber/Lyft/transit).

When I did the math, if I lived in the city it would be cheaper to use rideshare everywhere and rent cars when needed, than to own a car.

I'm coming to the view that gig workers are neither employees nor independent contractors. Employees don't get to unilaterally set their own hours, and contractors don't get prices unilaterally dictated to them or barred from their profession if their rating falls too low.

All this regulatory squabbling is arguing over whether a square peg fits a round hole or fits a triangular hole.

We need a third classification for gig workers that affords them some protections while preserving the economic viability of ridesharing companies. Disregarding any problems we might have with specific companies, I think ridesharing companies are a benefit to consumers.

This is really interesting and thought-provoking, but I'm skeptical.

1. In my experience, each data source and each data format requires a lot of custom work. Each kind of prediction task requires additional custom work. I don't see this going away. Even if a company develops a solid core of reusable engineering infrastructure, it will always need to be adapted to the problem at hand. At this point, this company would seem more like a consultancy, with non-trivial marginal/variable costs. This reminds me of Palantir, which operates this way - core set of tools and infrastructure, consultants implement and apply these tools/infra at each client company with a lot of custom integration work.

2. Assuming this is not a problem and the shared infrastructure is able to generalize enough of the custom work to be feasible, this thesis actually seems like an argument for the big tech companies dominating all data companies. Google, Microsoft, and Amazon have the engineering talent and resouces to develop this hypothetical infrastructure. They also have the internal political will because they can then expose this as cloud APIs. Indeed, it appears they are already attempting this in certain domains.

3. Superior engineering infrastructure is indeed a competitive advantage, but isn't enough of a moat for a single company to dominate this space. Yes, great engineers are hard to find, but there are enough of them for more than one company to feasibly develop this infrastructure, with a lot of money. You can't say the same about trying to buy the social network of Facebook or Instagram.

The interesting thing about the $3000 "upgrade" is that if you actually want to get it, I think you're better off taking the $3000 and investing it in Tesla.

Assuming autonomy is make or break for Tesla in the long term: if Tesla fails at it your stock will still be worth something (versus a useless upgrade). If Tesla succeeds you'll make more than enough to cover a retrofit.

1. I wouldn't take much away from the LSTM benchmark. It's more a benchmark of Keras since Keras only supports CuDNN's LSTM via Tensorflow right now. AFAIK CNTK does supports CuDNN LSTM but not through Keras. Keras actually implements its own LSTM in terms of the base math operations (it doesn't call the Tensorflow or CNTK LSTM operations which are in some cases optimized in C++ etc.), so on the CPU you probably could get better performance if you were using the Tensorflow or CNTK functions directly.

2. Compiling Tensorflow from source on CPUs is a bit of a hassle but I have seen nice performance gains (10-20%) for LSTM tasks. I bet you would get even higher gains for CNNs since they're more parallelizable. (Note: I've never gotten the latest TF to work with Intel MKL).

3. I haven't fully tested this myself, but with the P100s you also have full support for half precision floats, which supposedly offer a huge speedup.

4. Also would have liked to see benchmarks of other frameworks like PyTorch, etc. I haven't used them myself but everything I've heard indicates that Tensorflow is often slower.

When possible, most DL frameworks take advantage of Nvidia's specialized CuDNN libraries, which often provide a 10x+ speedup (and are obviously not available via WebGL). So at least on the latest Nvidia cards, you will likely see a 10x slowdown, and probably even more.

Cardiogram | San Francisco, CA | ONSITE: FULLTIME

Our mission is to reinvent preventive medicine using consumer wearables. Our goal is for Cardiogram to be a “doctor” on your wrist that continuously screens your health based on your exercise and wearable data.

We have an Apple Watch app and an Android Wear app that help users track their heart rate and exercise, both built using React. We’re a small team funded by a16z and we’ve scientifically validated our algorithms with UCSF (https://techcrunch.com/2017/05/11/apples-watch-can-detect-an...).

We’re looking for engineers wth 2-4+ years of frontend, mobile, UI, product design, and/or product development experience to help us design and build out our mobile apps. Our stack is React, Node.js, and Postgres. You’d be working directly with our two founders and another engineer-designer. There is a ton of exciting design and frontend work we need to do around helping our users stay healthy, screening them for medical conditions, and providing them with options for diagnosis and treatment. We have hundreds of thousands of daily active users and our Apple Watch app has been featured by Apple in the past. All the time we hear anecdotes from our users about how their cardiologist recommended they use our app to track their heart rate.

Email me at brian@cardiogr.am.

Tensorflow sucks 9 years ago

This article is not that detailed, but it's a sentiment I agree with, so I'll add one major shortcoming of Tensorflow: its memory usage is really bad.

The default behavior of TF is to allocate as much GPU memory as possible for itself from the outset. There is an option (allow_growth) to only incrementally allocate memory but when I tried it recently it was broken. This means there aren't easy ways to figure out exactly how much memory TF is using (e.g. if you want to increase the batch size). I believe you can use their undocumented profiler, but I ended up just tweaking batch sizes until TF stopped crashing (yikes).

TF does not have in-place operation support for some common operations that could use it, like dropout (other operations do have this support, I believe). Even Caffe, which I used for my research in college, had this. This can double your GPU RAM usage depending on your model, and GPU RAM is absolutely a precious resource.

Finally, I've had issues where TF runs out of GPU RAM halfway through training, which should never happen - if there's enough memory for the first epoch, there should be enough memory for every epoch. The last thing I want to do is debug a memory leak / bad memory allocation ordering in TF.

Braess’ paradox 10 years ago

Sure, but dropout actually increases training error (makes you less likely to find the globally optimal training error), with (sometimes) a decrease in generalization error (test error). So any connection is very thin.

Braess’ paradox 10 years ago

I'm not sure dropout has anything to do with local optima or removing greedily optimal paths, since it is random.

The original dropout paper's handwavy justification for dropout is that it prevents co-adaptation. It prevents individual units (nodes/neurons) in the network from relying on specific units in the previous layer firing as well. This is a bad thing because it's fragile (if one unit is off). I say handwavy because this is just intuition; there is not really any proof that this is actually what is happening.

Another commonly cited motivation is that dropout is like learning an ensemble of multiple networks.

The only paper I've seen that theoretically analyzes dropout is: https://arxiv.org/pdf/1506.02142v6.pdf, which proves it's equivalent to approximating gaussian processes (this is beyond me).

The very method of using a word embedding space assumes the manifold is smooth, so the fact that vectors extracted from a method that assumes a smooth manifold, are in fact on a smooth manifold, is just circular and not evidence of anything.

All those sentences sound (mentally) a lot differently. Some of those sentences give you the impression the speaker is an idiot, for example.

I specifically and mindfully added those words because everything is really an open research question. Would you rather I dissembled a false sense of confidence? If anything, you're stating your vague case way over-confidently. Turing-completeness is broad and nonspecific. Doing "some computation" is an obvious statement that doesn't add any information. The human brain does not seem to have time limits when it comes to thinking about what to say, and further we don't understand enough about neuroscience to make statements like that. Like I said, these are all active areas of research; the jury is still out on whether any specific approach will be the breakthrough.

EDIT (reply to below): in general these statements are either vague and nonspecific, or perfectly correct and non-informative, comments that don't have much to do with my original point.

Sorry, I've edited my original comment to be clearer. What I really meant is that there is wide tolerance of noise in those domains. "How long does stars last" has a completely different meaning than "How long do stars last" - not tolerant of noise.