There have been efforts to standardize antibody reagent testing that are sorely underfunded/undervalued, https://ycharos.com/ (https://www.nature.com/articles/s41596-024-01095-8)
HN user
cing
@jsci http://www.proteinqure.com
Github issues will be the real social network for AI agents, no humans allowed!
Author made it clear this was an educational essay, but concluding the problem has very limited therapeutic applications comes across like a bit of a take down for Atomic AI's platform.
Heck yes. Music pirates have done a whole lot to help preserve hip-hop history by ripping/archiving countless rare/underground 12" and cassettes that are definitely not available on streaming platforms; we're talking alternative versions of songs like clean edits, remixes, acapellas, instrumentals which supported hip-hop producers and DJs in the early days.
For example, check the posts in the early 2010's of sites like this, https://hiphop-thegoldenera.blogspot.com/, of course all the links are broken now.
This has been a previous area of research for Google (https://ai.googleblog.com/2017/04/predicting-properties-of-m...). It remains routine to benchmark GNNs and other molecular machine learning models on predicting quantum mechanical properties including energies (which speed up geometry optimization)
I agree with the sentiment of this paper (AF can enable drug discovery), but in this specific instance, the authors had a real opportunity contribute a general finding to the scientific community but instead they put in the lowest amount of effort (to a point where they're almost saying nothing at all).
The target had dozens of related structures in the protein databank, including relatives with ~40% sequence identity. This target family has a very similar structure, and conserved active site residues. It's relevant that this target has approved cross-CDK family inhibitors (and thousands of data points of CDK family binders on ChEMBL). The conventional way to enable structure-based design is to build a homology model using a similar structure (see here: https://swissmodel.expasy.org/repository/uniprot/Q8IZL9?temp...), and in this case, there is very low deviation from the AF2 model and this "old fashioned" approach.
To recap, this target had a decent model that would have likely sufficed for drug discovery. The community already knows that "homology models" can be used for structure-based drug design, so any methodological hypotheses of this paper are not supported by evidence.
Not saying it's easy but ribosomal synthesis of consecutive non-canonical amino acids has been achieved by some groups (https://www.cell.com/cell-chemical-biology/fulltext/S2451-94..., https://pubs.acs.org/doi/10.1021/jacs.8b07247), and many of these peptides are extremely short. Figure 5 of this manuscript describes successful applications of this for hit identification often in the context of massive peptide library screens, https://pubs.acs.org/doi/10.1021/acs.accounts.1c00391
Our startup routinely orders the synthesis of hundreds of peptides for technology validation and drug discovery research (not suitable for human consumption). Costs for a small quantity through a contract research organizion can be $200-$2000 USD per peptide depending on desired purity, length, and chemical complexity. For some applications peptide arrays are suitable, and can drive the costs down to $10 USD per peptide or lower. In both cases, the turnaround time is 4-6 weeks, even though a peptide chemist could do the job in about half the time for a rush order.
There is innovation in this space. From green chemistry initiatives to replace hazardous solvents by CROs/industry invested in large-scale production of peptides, https://www.bachem.com/news/bachem-novo-nordisk-redesign-spp..., to routine solid-phase synthesis of peptides greater than 100 amino acids in hours (https://www.science.org/doi/10.1126/science.abb2491, being commercialized here https://www.amidetech.com/)
One of the reasons we don't have them all is that individual genes can encode for multiple protein isoforms through alternative splicing. AlphaFold was only run on one. Otherwise, there's lots of important biochemical/biophysical processes that impact structure, as cells are only about 50% protein by weight.
Just in case you're not joking, it's worth noting that the majority of distributed molecular simulation (past and present) is spent studying "folded proteins" to discover structures of proteins that are often hidden from methods like AlphaFold (currently). For example, https://www.nature.com/articles/s41557-021-00707-0
Everything between the BRCT and RING domains of BRCA1 is an intrinsically unstructured region which DeepMind correctly predicts, https://pubmed.ncbi.nlm.nih.gov/15571721/
Another famous one would be R-domain of CFTR, which was not resolved in experimental structure determination, and AlphaFold models correctly show disorder there. Nothing to be done in those cases except perform molecular simulation or other experiments to assess dynamic ensembles, https://alphafold.ebi.ac.uk/entry/P13569
Yet, there were still 136 human teams who competed in CASP14 (https://predictioncenter.org/casp14/docs.cgi?view=groupsbyna...), including DeepMind. Even if a significant fraction of these projects were done piggy-backing another grant, this work does receive research funding.
The process is described in Supplementary, but where do you see the code to train the model? The repository is the inference pipeline.
There's quite a nice plot from a review paper of D.E. Shaw Research that lists the timescale of several biological processes (and compares it to other experimental methods), https://www.annualreviews.org/doi/full/10.1146/annurev-bioph... (Figure 2). Anton has been extremely helpful for studying the basic science of protein dynamics in academia and has been applied in industry (namely at Relay Therapeutics), but drug discovery is a long process so we still haven't seen the fruits of those long simulations yet.
So what you're saying is: https://xkcd.com/1831/, except that CS/ML practitioners have a negative impact by trying to contribute without understanding the nuance. I think the next logical question is: how many years of education should you have in order to contribute? 10 years? We'll all be killed by a virus by then :)
Best of luck. Although not in the majority, there have been many academic recruiting posts on HN in the past.
I know that docking using GPU is about an order of magnitude faster than CPU (see today's Schrodinger 2019-1 release notes, https://youtu.be/K4AYdBvuOe4?t=90). Is there a way of doing GPU accelerated precomputation though?
The majority of ongoing Folding@Home tasks are not aimed at structure determination, but rather simulating the conformational dynamics of folded proteins (exploring the energy landscape rather than searching for the global minimum). Very few of the CASP algorithms are well-suited for this problem.
The path is most likely through reliable structure prediction of drug targets. That would open up rational drug design projects that may have previously been impossible. The only problem is that experimental structure determination is so good in pharma, that it's hard to compete. For example, on a structure-enabled project, it may be possible to experimentally solve multiple high-resolution 3D models per week with an order of magnitude higher accuracy than predicted models. Once you can routinely get structures, there's still the rest of the drug discovery pipeline left to go.
This reminded me of a nice article about how blindness evolved in the Mexican cavefish, http://seedmagazine.com/content/article/pz_myers_on_how_the_...
I'm a cofounder of a start-up working on near-term applications of quantum computing in biology, specifically on the protein structure side of things (https://www.proteinqure.com). There are many self-contained subproblems in this space which are not limited by data because models are accurate enough to inform experiments, but there's probably not a scientist in the world who would say we have a sufficient understanding of biology to make predictive models with generality.
ProteinQure | Machine Learning Engineer, Computational Biologist | On-site, full-time | Toronto, Canada
ProteinQure is an early stage deep techy startup building the next generation of computational drug design tools helping to reimagine how we design therapeutics. We exist to foster innovation that enables design at the atomic scale; combining biophysical models, quantum computing algorithms, and reinforcement learning. Working with us involves having the courage to reinvent the status quo and the determination to see it through. We're seeking scientists and engineers to help us build the software infrastructure to drive drug discovery. That will include inventing novel machine learning algorithms, contributing to open source software, and using hybrid quantum/classical algorithms to fold proteins. Find out more on our website: https://www.proteinqure.com
Biology experience is not required. Our software stack is Python-centric. Please email hiring@proteinqure.com and mention "[HN]" in the subject line.
I'm confused as to why you're addressing this commenter using "argument from authority" when you seem to be weakening your position, suggesting that studying protein dynamics has led to significant advances in the field. It doesn't change the fact that you made a flippant remark that a team of some of the most experienced drug discovery scientists in the industry are wasting their time using this approach (despite not having worked in drug discovery, i.e. not an expert in the field), instead of just explaining why you believe this. Didn't mean to make it personal, that's just how I interpret your comment.
Relay is not doing protein engineering or working on predicting protein structure. They are making models of protein dynamics to assist in drug discovery (often using already determined structures). We both disagree with the parent commenter that it's a waste, and to claim that Murcko and D.E. Shaw are going in "blindly" would be ignoring decades of research on protein dynamics of some of the hardest drug targets out there. The fact remains that there aren't many success stories of using simulations of protein dynamics to accelerate drug discovery. Computational chemistry protocols used routinely in pharma drug discovery typically do not include this type of detail.
The PDB represents the best we have, but I wouldn't call it a great dataset for learning. The 150,000 known structures are a drop in the ocean when it comes to the space of possible sequences/structures.
DeepMind and others are trying. "Hassabis said the company is now planning to apply an algorithm based on AlphaGo Zero to other domains with real-world applications, starting with protein folding."
[1] https://www.bloomberg.com/news/articles/2017-10-18/deepmind-...
Not sure what OP is thinking, but you can look at this example of a commercial product designed for the prediction of off-target effects (https://cyclicarx.com/ligandexpress/).
"San Francisco, IL"? It's quite a commute to an alternate dimension =)
"evidence linking amyloid beta and tau to AD is quite good"
About 100 discontinued clinical trials targeting amyloid beta disagree with you there, https://www.nature.com/articles/nrd.2017.194