HN user

entee

2,970 karma

Biochemist and data scientist. Lover of art, history, politics and a good argument.

Currently: Founder, Anagenex.com

Posts5
Comments618
View on HN

A lot of this post relies on the recent open ai result they call GDPval (link below). They note some limitations (lack of iteration in the tasks and others) which are key complaints and possibly fundamental limitations of current models.

But more interesting is the 50% win rate stat that represents expert human performance in the paper.

That seems absurdly low, most employees don’t have a 50% success rate on self contained tasks that take ~1 day of work. That means at least one of a few things could be true:

1. The tasks aren’t defined in a way that makes real world sense

2. The tasks require iteration, which wasn’t tested, for real world success (as many tasks do)

I think while interesting and a very worthy research avenue, this paper is only the first in a still early area of understanding how AI will affect with the real world, and it’s hard to project well from this one paper.

https://cdn.openai.com/pdf/d5eb7428-c4e9-4a33-bd86-86dd4bcf1...

I think what’s missing is what the software allows. It could be BMW/Merc etc are way more conservative on what the allow the system to do and when they force the driver to take over. In certain contexts Merc is actually willing to assert and stand by a higher level of autonomy than any other manufacturer: (https://www.motortrend.com/news/mercedes-benz-drive-pilot-le...). Taking that at face value it’s possible they can do it and choose not to because they don’t want the liability. Whatever systems are in regular cars are then either borked or deliberately have less hardware.

Tesla is uniquely risk tolerant for better or worse. You also don’t hear about people getting into accidents in a BMW on self driving because they don’t make the same claims and have tons of safeguards.

It’s only marginally less useful to actual biology than full on X-ray structures anyway.

I'm not sure what you're implying here. Are you saying both types of structures are useful, but not as useful as the hype suggests, or that an X-Ray Crystal (XRC) and low confidence structures are both very useful with the XRC being marginally more so?

An XRC structure is great, but it's a very (very) long way from getting me to a drug. Observe the long history of fully crystalized proteins still lacking a good drug. Or this piece on the general failure of purely structure guided efforts in drug discovery for COVID (https://www.science.org/content/blog-post/virtual-screening-...). I think this tech will certainly be helpful, but for most problems I don't see it being better than a slightly-more-than-marginal gain in our ability to find medicines.

Edit: To clarify, if the current state of the field is "given a well understood structure, I often still can't find a good medicine without doing a ton of screening experiments" then it's hard to see how much this helps us. I can also see several ways in which a less than accurate structure could be very misleading.

FWIW I can see a few ways in which it could be very useful for hypothesis generation too, but we're still talking pretty early stage basic science work with lots of caveats.

Source: PhD Biochemist and CEO of a biotech.

Such a database would be hugely helpful across chemistry. Right now it’s extremely expensive to access databases like Reaxys or Scifinder, and they’re not usually programmatically searchable at scale. Some databases do exist based on the patent literature (https://depth-first.com/articles/2019/01/28/the-nextmove-pat...) but they’re not as well curated or complete. A pubchem like database for reactions would be really awesome.

As a fellow perfectionist who has started a company, one thing that has helped me is realizing that most decisions are a lot more reversible than they appear. Even in the legal and financial domain, most things that you might obsess over are fixable if you make a mistake, and decent lawyers will tell you which ones you really have to avoid. Sometimes it'll cost you money and time, but the biggest cost is avoiding making decisions.

Always remember, no decision is a decision. Usually that's the worst choice because almost any decision, even a wrong one, at least moves the ball in some direction, allowing you to gather more information. The only guarantee in this game is that stasis will kill you, so bias towards action. When in doubt, try to evaluate "most probable bad outcome" which is different from "worst possible outcome".

Good luck!

There's a lot that can be learned with building-block based experiments. If you do a building block based experiment then train a model, then predict new compounds, the models do generalize meaningfully outside the original set of building blocks into other sets of building blocks (including variations on different ways of linking the building blocks). Granted that's not the "fully novel scaffold" test, however it suggests that there should be some positive predictive value on novel scaffolds.

We've done work in this area and will be publishing some results later in the year.

This is true. Getting datasets with the necessary quality and scale for molecular ML is hard and uncommon. Experimental design is also a huge value add, especially given the enormous search space (estimates suggest there are more possible drug-like structures than there are stars in the universe). The challenge is figuring out how to do computational work in a tight marriage with the lab work to support and rapidly explore the hypotheses generated by the computational predictions. Getting compute and lab to mesh productively is hard. Teams and projects have to be designed to do so from the start to derive maximum benefit.

Also shameless plug: I started a company to do just that, anchored to generating custom million-to-billion point datasets and using ML to interpret and design new experiments at scale.

Not a chiphead, but saw this in the article that might be a reason ARM is better for this kind of thing:

"The theory goes that arm64’s fixed instruction length and relatively simple instructions make implementing extremely wide decoding and execution far more practical for Apple, compared with what Intel and AMD have to do in order to decode x86-64’s variable length, often complex compound instructions."

Not sure it's true, not an expert. But it doesn't sound wrong!

I drive a lot, thanks. If you can prove me level-5 or even very good level 4 autonomous driving, and that a computational driver makes radically fewer fatal mistakes than a human, then I'm with you. In other words, if you can satisfy a good regulatory regime like the say, airplanes or drugs, then great.

Short of that, it's a luxury and a danger.

I'm not sure what you're arguing in terms of acceptable risk. Biotech is incredibly regulated, specifically because the risks are so high, effectively there is very little acceptable risk. In biotech, a patient dying due to your drug is a Big Problem that will at best cause you to put a disclaimer on the package (see Black Box Warning) and at worst immediately end your drug's prospects. We can argue about trade-offs (if you've got terminal cancer, maybe a rare heart event is a worthwhile risk, probably less so if you have a rash), but this is exactly the way it should be.

Self driving cars are a nice luxury, especially in city driving, not something that radically improves our world. You get to read your phone instead of paying attention, and the trade-off is someone might get killed. It's like treating a rash with a drug that could give you a heart attack. That's a far cry from, "with this technology something that took days and $$$$ now takes hours and $" as was the case with all the older examples you listed.

If self driving cars were more like airplanes, I'd have a little more faith. Tesla's marketing BS doesn't inspire me with lots of faith.

On black boxes: https://health.clevelandclinic.org/what-does-it-mean-if-my-m...

It’s pretty clearly the current best way to do deep learning on molecules in chemistry. Transformers do really well in predicting reaction outcomes but GCNNs are still better when it comes to “is this a good molecule according to this label” type questions. That said, we have added attention layers to GCNNs so I think the lines get blurry.

We already do all of this in drug discovery research before putting things into animals. We try computer modeling (very inaccurate but a good pre filter) then cell culture, and only after that do we use animals. Not only is it ethically the right thing, it’s the economical thing. Animal studies are crazy expensive (each mouse costs a couple dollars per day just to keep alive and you need a lot of them) compared to cell culture and modeling.

The problem is that in your scenario, many many more humans would be harmed. While animal models are flawed, there are no cell culture tools that even come close to recapitulating the system wide effects you get in a whole organism. Without animal tests, all the drugs that hurt a mouse (and there are very many that fail in mouse tests for this reason) would instead do it to humans.

You can ban animal research, but if you do, know that virtually none of the medicines you take would exist. It really is that stark.

Even if the amyloid hypothesis is true (and most neuroscientists I know think it's not), this is a terrible decision.

1.) We don't know when to give the drug. Maybe giving it even earlier would help, but these trials don't tell us that. Answer: Run a new trial.

2.) We don't know how much drug to give. The drug was approved on the (bad, weak) evidence that in one high dose arm of one of 2 trials, there might have been an effect. Is that the right dose? Who knows! In the other trial, the high dose may have actually been worse. You can't titrate dosage in Alzheimers like you do in cancer where you can just watch how the tumor is shrinking. Answer: Run a new trial.

Giving this drug is not without downsides. You will have side-effects, including serious ones such as potentially brain swelling. Some people may be seriously injured or killed as a result of taking this drug. You have to make sure that the benefits outweigh that downside, and the trials show us a very dubious, weak effect.

The FDA should have said, "Good work, maybe there's an effect with this dosage, go run a new trial with the revised protocol." That's the right call.

Instead, they added, "Oh and you can sell the drug in the meantime and you don't have to tell us for 9 years."

How are you going to recruit patients for the trial? "We could give you the drug, but might give you a placebo." How many patients sign up for that instead of saying, "OR I could go out and buy the drug (which you claim totally works, and the FDA agrees!) independently."?

What if the trial fails in 9 years after you have tons of anecdotal reports (remember placebo shows an effect in past trials)? Now you have to take if off the market, imagine the loss of credibility that will entail for the FDA.

How about other drugs that reduced amyloid but showed no cognitive effect? (Eli Lilly's Solazenumab among many, many others) Should they get approved now too?

This is a mess for everyone and benefits Biogen. Everyone else loses, even the well-meaning patient advocates who created the political pressure for this decision.

Additional sources:

https://www.clinicaltrials.gov/ct2/show/NCT01900665

https://blogs.sciencemag.org/pipeline/archives/2019/12/06/th...

That linked post is fundamentally wrong because it rests on an assumption that has essentially zero human evidence: amyloid causes Alzheimer’s. There have been several drugs that very efficiently reduce amyloid, strictly zero (including this one) have ever shown any benefit patient health even when running long trials (the Biogen trials started in 2015 and were halted for futility). There’s reason to believe the amyloid hypothesis is flawed, meaning that approving a drug that reduce amyloid is not going to help anyone, and will likely hurt people (through side effects).

If competing experts are the question, note that 3 actual experts have resigned from what are coveted positions in protest. Nearly every part of the pharma industry (including the press, investors, other companies) who doesn’t stand to profit (I.e. not Biogen) has been up in arms saying this is an awful decision using words such as “horrifying”. There is no expert disagreement.

People can try to ret-con this by saying it’s like HIV, but note that viral load is a pretty good marker for disease morbidity in most viral infections. Amyloid is nothing like that as a validated marker for disease burden.

Useful sources:

https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5797629/#s3titl...

https://endpts.com/what-does-a-clear-majority-of-the-biophar...

My biggest worry with this would be the low level of off target edits and the number of recombination events that yielded an unwelcome product. Looks like those were very low, but with an N of 4, hard to know long term. The reason being that when you screw around with DNA you can get cancer. This has been an issue in a variety of cases with gene therapy, though is clearly getting much better. This is really cool though, exciting times!

Crazy New Ideas 5 years ago

The first iPhone famously only had 2G cellular internet. At the time 3G phones were not ubiquitous but were readily available (had one as a mid priced Nokia flip phone), so I’ve always thought this was a fascinating aspect of the iPhone story. I bet the decision was made to go forward despite the product being slightly hobbled precisely because of timing.

A couple other perhaps more useful quotes from the same article:

“Despite Trump’s assertions, few close observers of Obama’s and Biden’s response to H1N1 consider it a “full scale disaster.” And Biden, despite his early messaging problems, played a role in mobilizing the administration and ensuring enough resources were devoted to defeating the pandemic.”

And just some of the “lessons learned” section:

“ To keep Ebola from spreading further beyond Africa, the administration, which already had dispatched 3,000 troops to West Africa to help contain the spread, had to send public health workers to the affected countries via commercial airlines. This would not be dangerous unless a person was exposed to the blood or other bodily fluids of an Ebola victim. But pilots, passengers, airport workers and others in American cities from which the workers came and went had to be put at ease about the possible spread of the contagion.”

“Fauci was dispatched to cable news shows. Employing another lesson of the H1N1 days, Klain recruited the CDC's Frieden to join him in briefings to add medical credibility to the administration’s assertions.”

2009 may have been luck. They learned things, put together a plan, that plan was trashed by the new administration.

That’s a pretty disingenuous link. The Bush administration did have a plan for Flu and was thinking about the issue. In 2015 (your article recounts stories of H1N1 in 2009) the Obama administration established NSC level officials responsible for pandemic work, Trump disbanded it in 2018. Nobody knows whether those plans would have worked, but A plan is better than NO plan every time. Even so, president “inject detergent”, “liberate America”or whatever, may have screwed it all up. But no, it is not rewriting history to say there were plans that might have helped if implemented. We’ll never know due to institutional failure in part encouraged by the last administration.

https://jamanetwork.com/channels/health-forum/fullarticle/27...

https://www.ncbi.nlm.nih.gov/pmc/articles/PMC1283304/

I can't upvote this enough. We HAD a plan, both Bush and Obama put together pandemic response plans. We had serious issues in logistics and deployment for sure. We had some screw-ups in early testing for sure. Maybe we should grant some margin for error when there's a lot of uncertainty early on, but after about March or so, the errors are on pure government failure. The CDC isn't solely responsible here, but it has been an embarrassment given the skills it should have had and claimed to have.

Those who think the CDC has done way too much because the virus was overly-exaggerated are simply wrong. I have my quibbles with the CDC, especially early reactions and subsequent messaging. I don't think people who say it exaggerated the threat have any evidence in their column. For data, I'd say look around, see India, see 500,000+ dead Americans, and the latter in the context of at least some efforts to contain the virus.

The headline is fine, those who are enraged at the CDC probably should be, it failed us in so many ways, though probably gets a little too much of the blame. Those who are enraged because it did too much? Those people don't deserve consideration at this stage given the fairly obvious science.

I don't think that's quite right. Even if you have a flu strain with a particular name (H1N1) there are a number of variants within the H1 and/or N1 proteins that the vaccine will be targeted to as well, so the number is quite a bit larger than 198/131.

It's not that hard to make a new vaccine, the process for flu is very accelerated because they're just a variation on the original theme which has been proven to be safe (though unclear on effective until AFTER the flu season hits). If you think about it, we make a new flu vaccine every year, and develop it in 6 months. Each one contains a few guesses as to which of the flus are going to be an issue, it's not just one antigen in the vaccine. Those guesses are just that, an informed prediction, which is why the vaccines tend to be fairly ineffective (40-60%) at preventing disease altogether, though perhaps better at preventing serious disease.

Moderna is claiming to be quicker, so they'd have a more accurate read on what the REAL flu strain this year is going to be, and so it should be more efficacious. The key variable will be dosing. mRNA vaccines can only deliver a certain amount of mRNA so it might be impractical to deliver very many antigens at once. Also delivery issues, the current flu vaccine is super easy to make and deliver to patients, mRNA vaccines with their complicated cold chains aren't. Obviously it's not an impossible problem, but it's less convenient.

Time will tell, it's certainly very promising!

Helpful: https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3619640/

It has very short half life, there’s not a ton of it and it’s likely to be wrapped around a nucleosome. Uniformly, sequences in blood are quite short (<200bp) which is far too small to code for anything meaningful (for reference COVID genome is 29kbp, and the spike protein version in the vaccine is about 4kbp). The bloodstream shreds free DNA pretty effectively, but yes, some short DNA can be detected. That’s a good bit different than a large amount of mRNA/DNA floating around though, especially if it’s a long sequence.

Useful reference:

https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4715266/

One thing to note is that mRNA therapeutics are a really tough area in part because:

1.) RNA and the lipids used to get it inside a cell are inherently pretty immunogenic (a huge plus for a vaccine). If you think about it, you basically never see free mRNA/DNA in the blood stream other than if something has gone wrong (usually a virus), so the immunogenicity here is pretty ancient.

2.) RNA gets shunted to the liver and chopped up. Hence most RNA based therapeutics target the liver.

Vaccines are a really great use case, not just a, "we could help out here too," side-case. They're delivered intra-muscularly so there's less "go to the liver!", and the immunogenicity is a feature not a bug.

Lots of this comes from siRNA therapeutic research that is older than mRNA work, but the principles are very similar. Some older articles that touch on some of these issues:

https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3378126/

https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3269031/

https://www.sciencedirect.com/science/article/pii/S016836591...

Agreed. Boosters will probably be required yearly or so, and those will then be the “second shot” for lots of people.

The only caveat is that partial immunization can create the conditions for viral escape, but hopefully we’ll have boosters by then.

Burying the lede:

“ While millions of people have missed their second shots, the overall rates of follow-through, with some 92 percent getting fully vaccinated, are strong by historical standards. Roughly three-quarters of adults come back for their second dose of the vaccine that protects against shingles.”

Things are going relatively well, headline is overly alarmist.