HN user

dthal

2,135 karma
Posts140
Comments86
View on HN
www.wired.com 2y ago

The Suspicion Machine

dthal
29pts31
longform.asmartbear.com 3y ago

Pricing determines your business model (2014)

dthal
4pts0
www.wired.com 3y ago

The Suspicion Machine

dthal
4pts1
publicdomainreview.org 4y ago

Harold Fisk’s Meander Maps of the Mississippi River (1944)

dthal
43pts7
www.biorxiv.org 4y ago

SARS-CoV-2 exposure in wild white-tailed deer

dthal
19pts8
www.technologyreview.com 5y ago

How to talk to conspiracy theorists–and still be kind

dthal
2pts0
www.theatlantic.com 6y ago

Internet Speech Will Never Go Back to Normal

dthal
4pts1
youtu.be 6y ago

Domino Chain Reaction

dthal
1pts0
www.bloomberg.com 6y ago

Portrait of an Inessential Government Worker

dthal
3pts0
cityobservatory.org 6y ago

Quantifying Jane Jacobs (2018)

dthal
24pts0
www.citylab.com 6y ago

Technology Sabotaged Public Safety

dthal
1pts0
www.azcentral.com 6y ago

Where will the West's next deadly wildfire strike? The risks are everywhere

dthal
1pts0
cityobservatory.org 7y ago

Why homeownership is frequently a bad bet

dthal
2pts0
www.theatlantic.com 7y ago

The Slackification of the American Home

dthal
2pts0
www.theatlantic.com 7y ago

Living Smaller (1991)

dthal
1pts0
cityobservatory.org 7y ago

Quantifying Jane Jacobs (2018)

dthal
2pts0
www.theatlantic.com 7y ago

The Lessons of the Seattle Plane Crash

dthal
3pts0
bunching.github.io 8y ago

Pittsburgh Bus Bunching (2016)

dthal
82pts60
www.newyorker.com 8y ago

Turn-of-the-Century Pigeons Photographed Earth from Above

dthal
88pts12
humantransit.org 8y ago

Apps Are Not Transforming the Urban Transport Business

dthal
1pts0
www.propublica.org 8y ago

New York City Moves to Create Accountability for Algorithms

dthal
79pts22
www.seattletimes.com 8y ago

Amazon’s Seattle hiring frenzy slows sharply

dthal
4pts0
fivethirtyeight.com 8y ago

The First FBI Crime Report Issued Under Trump Is Missing a Ton of Info

dthal
2pts0
qz.com 9y ago

Instead of hacking self-driving cars, researchers try to hack the world they see

dthal
2pts0
www.vox.com 9y ago

The little-known deal that saved Amazon from the dot-com crash

dthal
2pts0
www.bloomberg.com 9y ago

What's Really Driving the Trade Deficit with China

dthal
1pts0
www.seattletimes.com 9y ago

We’re losing the information war

dthal
221pts312
www.theatlantic.com 9y ago

Why Nothing Works Anymore

dthal
113pts92
www.bloomberg.com 9y ago

Battling the Tyranny of Big Data

dthal
1pts0
www.vox.com 9y ago

How “open source” seed producers are changing global food production

dthal
3pts0

This blog post went up on Saturday, September 12. It says:

> MUCH cleaner air will push in by 4 PM Sunday over the coastal zone and will just reach Seattle late in the afternoon

> By 1 AM Monday, air quality will be hugely better in western Washington

As of now (Tuesday) there is no sign of clearing. Here in Seattle we are still in the "very unhealthy" range for AQI.

From that CityObservatory article:

  Zumper has also made a name for itself through its “National Rent Reports”
  —more or less monthly press releases that claim to track median rental prices around the country. 
  These reports have received copious media coverage, from the Bay Area to Seattle to Nashville to Chicago to Boston to LA to Miami to Denver, and so on.

Note that Zumper is not a reliable source for median rent data. CityObservatory wrote an article about their data problems a few years ago [1]. Zumper's data is, of course, based on apartments that are for rent and doesn't include currently occupied units. That alone skews high, especially when there is a lot of higher-rent new construction hitting the market. Also, it looks like Zumper's data skews towards higher-end neighborhoods.

For a broader look at the rental market, including occupied units and rent-controlled units, you could just consult ACS data. That says that median rent for all occupied 1-bedrooms in San Fransisco was $1912 in 2017 [2].

[1] http://cityobservatory.org/journalists-should-be-wary-of-med...

[2] https://censusreporter.org/data/table/?table=B25031&geo_ids=...

The CO2 is also an input; it originally comes from the environment, so this is net-zero emissions in the same way that biofuels are. From the article:

   Although the bus emits CO2, Team Fast argues that the original CO2 used to create the hydrozine is 
   taken from existing sources, such as air or exhaust fumes, so that no additional CO2 is produced
    - it's a closed carbon cycle in the jargon.

Well, supposing this is correct...Congratulations to Anthony and the rest of the Kaggle team! Those guys do a great job. Hopefully they get rewarded for it.

R for Data Science 10 years ago

I use both R and python quite a bit. I prefer python as a programming language. Here's my take on 'Why learn R?': (1) R/ggplot is hands-down better for plotting than anything in python. I also think that R is better for EDA generally. (2) Many smart, knowledgable people use R and publish their code. To learn from it, you need to know enough R to read and modify it. (3) R has better package support than python in several common data analysis domains. For example, in forecasting and in graph analysis, the best R packages available are much better than the best python packages.

>“It’s a paradox,” said Valentin Bote, head of research in Spain at Randstad, a recruitment agency. “The unemployment rate is too high. Yet we’re seeing some tension in the labor market because unemployed people don’t have the skills employers demand.

There's no real paradox there. Employment of young people, and therefore normal career progression for that cohort, essentially shut down for 6-8 years. Now the pipeline is a little empty.

A lot of people are going to dump on this billion dollar valuation, but we don't have the terms. Presumably there is liquidation preference. According to Crunchbase [1], the total capital raised is about $160 million after this money, so the investors don't need this company to be worth $1 billion to come out fine. Sam Altman made the point a few days ago that a lot of these late-stage private financings are kind of debt-like [2, about paragraph 8-9 and footnote #2], with the result that the valuations aren't really meaningful.

[1] https://www.crunchbase.com/organization/udacity#/entity

[2] http://blog.samaltman.com/the-tech-bust-of-2015

I wouldn't be so down on vocational training for software...It is a field that has an unusually high requirement for ongoing learning. And when you have to learn some new tech, usually you have to do it now, and in-place, not next September, in the nearest college town. Given the high ongoing learning requirements that software has, I think there is a pretty clear need for training that's delivered where you are and when you need it, and traditional schools won't be able to do that.

>I'm sorry, continuing resolutions are now our ideal?

I did not say that continuing resolutions are an ideal. My point was that although budgets are complex, only a relative small part of them changes from one cycle to the next. People don't have to digest the whole budget, because they already know what was in it before. They only have to understand the changes.

The ACA wasn't negotiated in secret. Budgets are normally continuing resolutions; they only negotiate the delta from previous years, which is usually small. And again, budget negotiations don't involve years of secret negotiations.

The point about the late-stage investments being not-really-equity is a great point. So my question is: why do financial journalists just about always miss it when they write about tech unicorns?

In the post, PG states that First Round's study is evidence of gender bias in VC financing. But footnote [2] is important: Uber was excluded as an outlier. Now...excluding Uber is reasonable (it is sort of an outlier), but so is not excluding it (it was a company that First Round invested in). When the conclusion from a data analysis depends on which way you go on something like this - which of two reasonable alternatives you pick - then the results are fragile and they don't really support either conclusion very well.

I've played around with this some, and the recognition isn't perfect, but I'm very impressed with how well it can pick up a new pattern from even just one example.

Fair enough... but the issue here is about findings of an effect versus no effect. It's about statistical significance. In order for single-comparison p-values and the like to be valid, there has to be a (one) single comparison. There is a way to do 'any-of-k' testing, but the required effect sizes get larger.

It seems that the main issue here is that with pre-registration, study authors have to pick a single measure of primary benefit at the outset, whereas before, they might have made that choice after getting results back. The original study is at PLoSONe, and it is not a difficult read[1]. From that source:

>Prior to 2000, investigators had a greater opportunity to measure a range of variables and to select the most successful outcomes when reporting their results... Among the 25 preregistered trials published in 2000 or later, 12 reported significant, positive effects for cardiovascular-related variables other than the primary outcome.

That is, in most cases, there are large effects for some outcome, and if they get to choose the primary outcome after looking at some results, they could have been cherry-picking the outcome variables.

[1] http://journals.plos.org/plosone/article?id=10.1371/journal....

Scroll down to where it says:

   It even works for dresses that were not in the training set:   
Yes, that first dress is in the training data, but she did do reconstructions on dresses not used to build the PCA basis. They look decent, but they aren't as good as that first example. And, as she points out, they can't reproduce patterns that are not in the initial data very well, and can't reproduce accessories that were not in the initial data at all.

Nonlinear methods can result in a smaller (lower-dimensional) representation, but they are non-linear so they are harder to use and usually require more data. On the other hand, PCA is easy and having 50 or more principal components is often not a problem, unless you are doing visualization. With the additional representational capacity from just keeping extra dimensions you can still get good reconstructions.

> Was there image processing to distill the dress into major components, hand tagging of variables (like short/long, color, etc), or just raw pixel data processed?

Its pretty clear that this is PCA on raw pixels of the images, so no, there is no hand-tagging or anything else. The eigendresses are just the eigenvectors from PCA projected back into the original, image-scale, basis.

It does survive, in a certain sense, in scientific programming and data science. Both iPython notebooks and Rmarkdown are a sort of literate programming, although with the emphasis on the text more than the code. In that setting, the executable artifact is not really more important than the explanation of why the code does what it does, so the extra overhead is justifiable.

Rmarkdown example: http://kbroman.org/knitr_knutshell/pages/Rmarkdown.html

iPython notebook example: http://nbviewer.ipython.org/github/empet/Math/blob/master/Do...