The fed tracks this data: https://fred.stlouisfed.org/series/APU0000709112
HN user
tansey
wes.tansey@gmail.com
Asst. Prof in Computational Oncology at Memorial Sloan Kettering Cancer Center. PhD in CS from UT Austin.
Bio: http://wesleytansey.com
Github: https://github.com/tansey
Twitter: http://twitter.com/#!/wesleytansey
Happy to talk with anyone-- email me! :) http://hnofficehours.com/profile/tansey/
Somebody has to pay for the trial. Drugs are expensive and the amount needed to dose a single person is orders of magnitude more than mice. So who funds the study?
If it's the government or philanthropic fund, you have to put in a grant and show that it's competitive in terms of preliminary results etc.
If it's the drug companies, you have to deal with lots of complicated things with combination therapies. Drug companies don't like their drug paired with other drugs that aren't in their portfolio. They also need to see a way to make a profit on it, which means they need to evaluate whether this is the most likely successful trial, assess bang for buck etc.
If you want to go compassionate use, then you need to get the pharma to donate the drugs or insurance to cover it. This is spain, so I guess insurance is just the government since they have universal healthcare. I have no idea how that works but I am guessing it doesn't move fast.
Somewhere between $500K and $2M for an mRNA vax is my guesstimate. You don't have to hypothesize about what the rich might do, though. The co-founder of GitLab was diagnosed with osteosarcoma a few years ago. He relapsed, leaving him without any standard of care treatments. He's spent the last couple of years throwing loads of time, effort, and of course money at lots of experimental treatments: https://osteosarc.com/ -- even released all the data gathered on his tumor over timepoints.
For all the folks complaining about "it's only in mice! things never work in humans!" -- I work at MSK and we definitely have seen success treating PDAC in humans: https://www.nature.com/articles/s41586-023-06063-y
"Why don't I see these treatments hitting the general public?" Because trials like these are phase I/II. Then you need a phase III that takes a long time to recruit a large cohort and has overall survival as an end point so you need a long time to measure the actual outcome you care about. And most trials fail in phase III because the surrogate end points used in phase II studies, like progression free survival (ie how long did patients go before their disease advanced in screens), are not necessarily great predictors of improved overall survival.
Specifically for cancer vaccines, this paper was a driving force behind MSK establishing a cancer vaccine center to scale up these personalized neoantigen mRNA vaccines. It's very very difficult to do and extremely expensive right now.
PI at a cancer center here. These two ideas are not mutually exclusive. Cancer is indeed not 1 disease but many countless ones. At the same time, cancer diagnoses are based on site of origin and histology (how it looks under a microscope). But often what drives a cell of one tissue toward pathogenesis is the same mutation or other molecular malfunction as a cell of another tissue type. In those cases, we can develop drugs that target that specific component and it may work across both cancer types.
Unfortunately, there are countless ways things can go wrong in cells. There are also rarely drugs that truly span a large swath of cancer types effectively. That's because even though two cancers from different sites may have the same driver, how they respond to treatments can differ. The difference in cell state may allow one of the two cancers to adapt to the treatment, such as by activating an alternative growth pathway, whereas the other cancer type's cell of origin may not have such an easy road to therapeutic escape.
Sugar basically doesn’t exist in nature without fiber
Honey?
It's political news: https://news.ycombinator.com/newsguidelines.html
"Off-Topic: Most stories about politics, or crime, or sports, unless they're evidence of some interesting new phenomenon. Videos of pratfalls or disasters, or cute animal pictures. If they'd cover it on TV news, it's probably off-topic."
Here goes nothing:
- bio startups rise, ag tech startups rise, food startups rise-- everything to do with engineering life.
- second and third tier cities rise in the US while first tier cities ebb. Places like Austin see continued growth, while cities like Baltimore and Cincinnati begin to really revitalize as local market become more important than global markets.
- mobility increases as people are less tied to their jobs/families
- Google slows on the innovation front. FB is increasingly weak as social networking just isn't as profitable anymore. Amazon keeps churning away. Netflix wins best picture at the Oscars.
- AI progress slows. The ML community fragments again. Neurips is no longer where the best ML researchers publish.
- At least one climate protest with over 5M people involved nationally, calling on Congress to act now
- China, mired in internal political upheaval, faces a lost decade. They will either no longer be one of the two biggest economies or they will be on a clear trend down but third place is a long way to fall
- twice as many people consider themselves vegetarian or vegan. Foods for this demographic have gotten much better and more diverse. Consumption of meat is still high, but the trend in 2030 will be clear: the meat industry is shrinking rapidly in the US and Canada. Other nations will lag here, meaning almost all of the innovation will happen in North America
- NYC will have built only two new subway stations
- a third political party will gain at least 5 seats in the House
- there will be a recession. Likely due to housing again. As the boomers die out, their millennial children inherit their suburban homes. Unfortunately they don't want to live in the suburbs and selling the house would pay off their student loan debt. But who is going to buy all these houses?
- political divide in the US heals, but it's not pretty and it all feels chaotic. Political parties weaken in favor of some new form of factions.
- CS undergrad enrollment declines but CS course requirements pervade nearly every STEM field. The common wisdom will now be that interdisciplinary jobs, not vanilla software engineer, are the growth sector, especially in bio/medicine/food/agriculture.
- Medicare age lowered to 50 as a transition toward single payer that will take another decade
- the world has missed its chance to avoid global warming. The schism in the debate will now be about what to do. Radical positions (open borders for mass refugees, a trillion dollar climate change R&D bill) will become more mainstream.
- Cities all over the US will be calling for federal infrastructure to build new train and tran systems, obviating the market demand for autonomous taxis. Uber and Lyft go under. Long haul autonomous vehicles are at 10% of all interstate traffic
- a common political point will be about how the US needs to stop subsidizing corn. As crops becomes more important in this decade (ag tech, rise of vegetarianism, climate change) it will be clear that corn is an over investment but political inertia will not let things change this decade.
- A dominant Canadian tech company will arise that rivals Google/FB or is at least rising rapidly
- towards the end of the decade, robotics is starting to reach the early-success stage. This will have impacts on all sectors as co-working spaces pop up with access to reprogrammable devices that can prototype commercial products. Think: Roomba for X. 2030 is still early days for this, but there's buzz, articles about robo-spaces Wired, etc.
The type of quants talked about in this article probably has a PhD in math, stats, CS (with a focus on machine learning), or something similar.
Yes. Exactly this. Building models that assess risks and potential gains. A quant is a catch-all term though, so one person may be working on models that predict some sector of the market and another person could be looking at online allocation algorithms for maximizing risk-adjusted returns.
This is why people do twin studies. Orphaned twins that were separated at birth are as close to a natural instrument as one could ever hope to get.
You can if it's defective: https://en.wikipedia.org/wiki/Lemon_law
The problem with the idea of cheap screening tools is Bayes' theorem. If doctors go ordering this for most people since it's just a blood test, and if only 1% of people ever really have cancer when tested then the 1% false positive rate means there's only a 50/50 chance you have cancer given the test is positive.
Except it seems there is good case law to show that in fact suspected cannot be forced to open a combination lock, as it falls under fifth amendment protection. They can, however, be compelled to provide a key if it is a key-based lock. This applies similarly to biometric-based locks.
It's hard to believe that an encryption key is any different than a combination lock in this "encryption is like a safe" metaphor.
Relevant cases:
https://supreme.justia.com/cases/federal/us/425/391/case.htm...
https://supreme.justia.com/cases/federal/us/487/201/
https://supreme.justia.com/cases/federal/us/530/27/case.html
The usual way they reduce variance in these man vs. machine poker showdowns is to do "pairs" play. You have two humans playing simultaneously in isolated locations. The decks for both humans are the same, but player 1 and 2 are swapped for one human. That way, the bot strategy has to play both hands.
It does totally eliminate variance, but they also take that into account and correct for it when looking at final outcomes usually. Right now the bot is up by something like 800K over 60K (out of 120K) hands. If that rate continues, it will win by around 1.6M or 400K per human. The blinds are 50/100, so that would equate to roughly 33 millibets (thousandths of a big blind per hand). That isn't too far off from standard win rates in bot vs. bot tournaments [1].
I'd say it's likely that the results of this tournament will be a statistically significant win for the bot.
[1] http://www.computerpokercompetition.org/downloads/competitio...
Educate them to do... what, exactly? What could Uber do with two million mostly-uneducated truck drivers that are scattered all through out the country? Most do not want to move, do not have the patience/desire/grit to go through long retraining periods for a vastly different job, and are currently making something in the neighborhood of $55K/yr. How could any company possibly be expected to help that large a workforce not take a dip in its standard of living, when literally the only skill they have is about to become nearly worthless? And how could you do it while still upholding your fiduciary responsibilities to your shareholders?
Stan (http://mc-stan.org/documentation/) is arguably the most advanced language. It's especially pushing the bounds of doing automatic variational inference, for the scenario where your model does not have a nice conjugate form that would be amenable to Gibbs sampling. It's not quite reached what I would say is production-quality, but some of the best people in the world of computational Bayesian methods (e.g., Michael Betancourt, most of David Blei's lab, etc.) are working on it.
The data is useless unless it is annotated (e.g. a human labels where the lanes are, where the obstacles are, what the bicylist is doing, etc.) - that's the bottleneck, not collecting large amounts of raw sensor data + driver actions.
Except that is exactly what Nvidia did, and it worked out fine for them: https://arxiv.org/abs/1604.07316
Not sure what the source is on that chart, but I'll assume it's credible. If that's the case, then despite what the article says, it's hard to believe that most of this decline is not accounted for by the decline in smoking. Surely the decline in smoking accounts for the vast majority of the lung cancer decline, and the peaks for colon and prostate cancer are very close to the same time.
It's not that much data. You're storing about 2M voxel image every 2s or so. There are open repositories out there already: https://openfmri.org/
Most (good) encryption schemes have the property that knowing part of the message, or some other information like message length, will not help you decrypt the message.
I guess I don't understand what is really novel about this. Don't brands like Digiorno and Freschetta already have highly-optimized pizza assembly lines? Is it really that hard to make the topping/sauce dispensing dynamic instead of static?
I also don't understand the appeal of having the pizza be super hot-n-ready. If you have an efficient baking-to-delivery pipeline, and an insulated container for the pizza, then how much are you really improving the quality? Ten minutes of sitting still in an insulated bag is fine with me.
And how do they handle the hand-off issue with an AUV? You get a notification and then the car sits there waiting to deliver the pizza for several minutes while you walk outside? At that point, the pizza is just as cold as it would have been if you had a human bring one to your door-- and at least the latter doesn't make you put on shoes.
Just seems fraught with issues that really require what you might call "AUV-complete" or "robot-complete" tech. If you have a very smart, powerful drone that is capable of picking up a pizza, keeping it warm, and effectively delivering it to people's doors with 99.99% reliability, then this sounds obvious. Otherwise, this seems like an idea that's too soon to succeed.
The contrastive divergence paper that Hinton published in 2006 definitely set the field off again. I remember entering grad school in 2010 and everyone was still really excited about using unsupervised pretraining. However, nowadays no one uses it.
It just turns out that with GPUs and stochastic gradient descent, no one needs any of that stuff. There are some tricks out there to making it really work, though. In that sense, Hinton's dropout paper has probably had a longer lasting effect on the field.
But either way, I doubt what OP is saying will be true. None of the real advances in deep learning are coming from self-taught coders in the middle of nowhere. They're coming from big labs with lots of resources, both physically and intellectually. This stuff takes a lot of hard thinking by a lot of people who understand optimization and probability. It also takes a ton of compute power and massive datasets, which won't be available to a hobbyist.
Isn't the labeling really tricky, though?
In my limited experience, EHRs aren't usually setup to handle structured labeling of something like an image. There are lots of different fields for text entry that can be unstructured. Then the only label left is the billing code, which ends up being a poor choice of label since the hospital often bills for what it can get reimbursed for, not what you actually had.
Just read (skimmed) this paper yesterday actually.
Looks interesting-- but there are no timing graphs! It's kind of a strawman argument to say "We can't use Newton's method because it's too slow to calculate the Hessian," and then go and present all your performance graphs in terms of number of iterations.
This seems pretty likely to be placebo effect, at least in the human trials which are the only parts of the paper I read. The paper cited [0] pre-registered on clinicaltrials.gov [1] -- which is great and everyone in the field should do so. However, looking at what they pre-registered, we have:
- Two different diet treatment regimes: fasting-like for 3 months followed by Mediterranean diet (FMD), and ketogenic diet (KD).
- Control diet was simply telling people to eat the same way as usual. So there's no real accounting for how well a placebo would do here.
- Measurements included a 54-question survey, adverse event counts, and various lab measurements. These measurements were taken at start, 1-month, 3-month, and 6-month intervals.
The problem then is that what was report was:
- Results of the first half of the first treatment (fasting-like for 3 months) for a subset of the measurements. What if things only worked in the second half? Or if things worked only for KD? So many implicit comparisons here.
- Comparison against the control group at 3 months, with reported p-values. Even though one of the reported measurements was the overall survey results, all of the values reported are p-values without any mention, that I can find, of multiple hypothesis testing across all measurements. This comes despite the fact that for all of the mouse results, they explicitly state they used Bonferroni correction.
- Baseline performance which involved no placebo. How many people would have improved if simply given some bullshit diet? Or if they had simply been given a diet that was vegetarian, or something that gave them the impression it was a treatment? Especially in surveys, placebo effect is a huge thing to look out for. Their more hard metrics like lab results show a more mixed bag with WBC dropping for fasting subjects. Sure, it returns once the 3 months is over, but then the supposed quality of life scores drop; so you can't have it both ways, though their writing makes it sound that way.
I'm not saying it isn't a great result from a bio standpoint. I'm sure the Cell reviewers found the mouse model results compelling. I just don't see any way to conclude the broad sweeping title of the article from the actual content of the paper, and it's unethical to do so without strong evidence.
[0] http://www.cell.com/cell-reports/pdfExtended/S2211-1247(16)3...
[1] https://clinicaltrials.gov/ct2/show/NCT01538355?term=NCT0153...
Very nice write-up. Several years ago, I took a graduate class that covered NEAT and ended up doing a little project [0] to see if you could apply similar ideas (recurrent nets evolved via NEAT). My idea was to use it for multi-agent control problems, though, where you want agents to try and "teach" other agents whenever they learn something useful.
[0] https://github.com/tansey/social-learning (note: not as clean of a structure as OP's code!)
There's a wearables company: https://www.atlaswearables.com/
It seems like most of these complaints center around the idea of the inherent subjectivity of the prior. In cases like astronomy and other hard sciences, the prior reflects actual scientific knowledge and is not really subjective at all. In cases where we don't have that kind of evidence, empirical Bayes methods work very well by just peaking at some subset of the data and finding a good point estimate for the prior.
I'm also not sure why the OP thinks that calculating the normalizing constant is a huge issue. Most of the time you rarely need it since you're likely going to end up doing an MCMC or some other sampling method for the posterior, in which case you only need proportionality.
There are lots of problems with Bayesian methods in practice, but most of them revolve around the scalability of modern methods to massive data sets and very complicated models. Many Bayesians tend to think that it's absolutely crucial to quantify uncertainty and that the added computational cost and human effort is worthwhile. In practice, point estimate methods to find MAP or even just maximum likelihood values work really well for most problems. If you look at the trend in most machine learning, for instance, generally people find a cool way to solve some problem with good performance (e.g. SGD + Deep Nets), then some Bayesian lab spends a few years trying to interpret everything as a generative model and coming up with a clever way to sample everything (e.g. Lawernce Carin's lab at Duke has done a lot of this work in Deep Bayesian Nets). The end result is usually better, but by then most people have moved on to a newer problem and the appeal of getting a marginal boost in performance is harder for me to see. The Bayesian nonparametrics crowd has historically done a pretty good job of hitting a sweet spot of compromise on this by keeping a Bayesian view but still (usually) treating everything as an optimization problem first (e.g. variational inference methods).
Thanks for the write-up Chris. Now I understand why I couldn't follow the path of logic you were laying out in our original discussion in PG's article's comments.
The main problem I was having is that you are assuming our observation variable is the latent skill or potential value variable (which you're calling x here). However, the article by PG was talking solely about the average of returns (let's call it y).
So the reason I was confused is that, assuming that the outcome of a startup is dependent only on x, we are really observing y ~ f(x) = \int_0^1 g(x)h(x)dx, where h is your cut-off criteria for x, g(x) is some unknown payoff distribution for a given skill level, and I'm assuming our x is in [0,1] without loss of generality. So in essence, the real problem here, even if you could see all of the individual returns for a given portfolio, is that you have to perform a very, very difficult deconvolution problem. And I'm pretty sure it's non-identifiable without some other information or additional parametric assumptions.
Thinking out loud a bit, let's assume that y is actually log(return), where a return of 1 is breaking even and 0 is losing everything. Since log(0) is undefined, most startups return 0, and very few exit for less than 1, I would think we could model this as a point-inflated normal distribution: p(y) = c * \delta_0 + (1-c) * N(\mu, \sigma^2). Given this, we could then model our latent parameters (c, \mu, \sigma) as being functions of x. Since the model is separable, we can even just look at the zeros and non-zeros in isolation. Then we can come up with a test from there, but I'm not really sure what that test would be at this point. Anyway, that's a completely different line of thinking, but it seems much more tractable in practice.