Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5660_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
15.09.2026
Размер:
11 Мб
Скачать
☆
90 3 DrugBank Online: A How-to Guide
https://t.me/medicina_free
(a)
(b)
Figure 3.9 Using the DrugBank Online sequence search. This figure provides an overview of using the DrugBank Online sequence searching tool. (a) The sequence input form, which accepts one or more FASTA-formatted nucleic acid or protein sequences. Users can adjust a subset of BLAST parameters using the entry fields and radio buttons. The filters allow users to restrict the search to sequences associated with subsets of drugs based on approval status and to specific types of sequences (target, enzyme, carrier, or transporter; see Section
3.2.2.6). (b) The first result displayed after searching using the human C-X-C chemokine receptor type 5. Note the hit metrics in the top right and the BLAST output alignment present below; exact matches are denoted by the one-letter code between sequences, while similar residues are denoted with a plus symbol (“+”). The bottom table lists the drugs with which the identified sequence has known interactions in DrugBank.
2. Adjust the BLAST parameters, if desired. Note that not all parameters available in BLAST, such as the choice of substitution matrix, are available to change. The “Expectation value” controls the cuto for returning hits; increasing this value will result in more hits but many more will be only slightly similar to the target sequence. The default gap opening cost is set at one (as opposed to the normal BLAST default of 11); this may result in hits with more gaps than otherwise
3.3 Protocols 91
https://t.me/medicina_free
expected. For a full explanation of BLAST parameters, see the ocial manual (https://www.ncbi.nlm.nih.gov/books/NBK279690/).
3. Adjust the “Drug Types” lter. Only sequences associated with drugs of the selected type will be considered when searching. This is useful if, for example, you wish to consider only approved drugs, whose protein binding and MoA are more likely to be known in considerable detail. Alternatively, ltering to all but approved drugs provides insight on scaolds currently under investigation.
4. Adjust the “Protein Types” lter. Setting this can narrow the search to the most relevant type of sequence, given the starting query (see Section 3.2.2.6 for a full explanation of each protein type).
5. Run the search (press the “Search” button).
Continuing the scenario from above, you run a search with your unknown sequence using the default BLAST parameters, including approved, withdrawn, investiga­tional, and experimental drugs, and limiting the protein types to targets. BLAST returns 73 matches, the rst of which is shown in Figure 3.9b. Each hit contains the name of the hit, together with the hit E value, bit score, and alignment length. Briey, BLAST identies small local matches between query and target sequence, which it attempts to extend in either direction while obeying set cuto parameters. The longest such alignment is used to score the hit, and is also provided as part of the hit itself; in the view here, the alignment is shown on a single line and may be scrolled to the left or right if it does not fully t within the hit table. Lastly, all relevant drug–protein interactions that t the protein types lter are included.
Inspecting the hits for the unknown sequence, it is clear that the top hits all belong
to the CXC and CC chemokine receptor families, with E values ranging from e
−33
e
. Indeed, the unknown query is the human C-X-C chemokine receptor type 5
−48
(UniProt ID P32302). Although the next hit, the type 1 angiotensin II receptor, has a good E value (e
−27
), it represents a clear departure from the cluster of top hits and
will not be considered.
To get a sense of what kinds of chemical scaolds can eectively target CXC/CC chemokine receptors, we can more closely investigate the hits. There are nine small molecules in the top hits, which are listed as either antagonist or inhibitor and for which a full structure complete with chemical classication (provided in the Chem- ical Taxonomy eld of the Categories section in the relevant drug card) is present in DrugBank (Figure 3.10a). Although some similarities are apparent across scaolds, such as a generally extended conformation and the presence of phenyl groups and amines, the scaolds appear diverse.
It is possible to conduct a rudimentary chemical similarity analysis using data extracted directly from DrugBank. For all nine structures identied, the “Substituents” provided as part of the chemical taxonomy were extracted and used to produce a 9 × 94 matrix of one-hot encoded features. The most common chemical features across all molecules (present in at least three drugs) are shown in Figure 3.10b. Conrming the visual inspection of the compounds, the various nitrogen-containing functional groups make up a large proportion of the results, together with heteroaromatics, various oxygen-containing groups, and alkyl uorides.
to
92 3 DrugBank Online: A How-to Guide
https://t.me/medicina_free
Framycetin Plerixafor
AMD-070
MSX-122
(a)
8
7
6
5
4
Frequency
3
2
1
0
Amine
Hydrocarbon derivative
Organonitrogen compound
Organic nitrogen compound
Organoheterocyclic compound
(b)
Azacycle
Aralkylamine
Carboxylic acid derivative
Heteroaromatic compound
Aromatic heteropolycyclic compound
Vicriviroc
Tertiary amine
Carboxamide group
Amino acid or derivatives
Organic oxygen compound
Organic oxide
Tertiary aliphatic amine
Organooxygen compound
Organopnictogen compound
Aromatic heteromonocyclic compound
Alkyl halide
Alkyl fluoride
Azole
Dialkyl ether
Organofluoride
Carbonyl group
Organohalogen compound
Ketoprofen
Maraviroc
Drug Similarity
–3
–2
–1
0
1
Ether
Piperidine
2
3
Cenicriviroc
2
1
0
–1
–2
(c)
INCB-9471
2
1 0
–1
–2
3
AMD-070 Cenicriviroc Framycetin INCB-9471 Ketoprofen MSX-122 Maraviroc Plerixafor Vicriviroc
Figure 3.10 A simple structural analysis of sequence search results.Thisfigureshows an example of how searching for similar sequences to a putative target can help to inform compound design. (a) Molecules identified within DrugBank to interact with top hits in a resulting search (see text for more details). (b) A histogram showing the prevalence of chemical features in the molecules identified in (A). For ease of visualization, features present in two or fewer molecules are not included. (c) 3D visualization of the full constituent feature matrix following PCA projection onto three components. It is clear that there are two clusters, which may serve as starting points for further analysis.
Furthermore, by using principal component analysis to project the full chemical feature matrix onto three components, it is possible to visualize the relation­ships between these drugs (Figure 3.10c). The resulting image reveals separation between most drugs, though INCB-9471 and vicriviroc are similar, as antici­pated based on their structures. There is a single larger cluster of cenicriviroc, plerixafor, AMD-070, and MSX-122, which is not immediately apparent from a visual inspection of their structures. Furthermore, all but cenicriviroc inter­act with the most similar hit in the original BLAST search, C-X-C chemokine receptor type 4. This suggests that focusing on these structures, and similar ones to them, might be a reasonable place to start when identifying a new bioactive compound.
3.3 Protocols 93
https://t.me/medicina_free
Although small-scale and highly simplistic, this example provides insight into how sequence searching may assist in the discovery of targets and the prioritization of potential scaolds for downstream discovery work.
3.3.3 Extracting DrugBank Datasets for ML
Machine learning (ML) is increasingly applied across healthcare, from the analysis of imaging data to a myriad of applications within drug discovery pipelines [10, 27]. Although model development in these relatively new and exciting areas remains an important consideration, we argue that data quality is crucial, as health informat­ics represents a high-stakes domain [28]. Combined with the generally recognized “unreasonable eectiveness of data” [29], assuming a data-centric approach [30] to model training and continuous deployment practices may reap signicant benets to organizations and patients alike.
Although we do not aim to provide a comprehensive overview of relevant ML techniques here, we do highlight several ways in which public users may obtain large, focused datasets for use in building and evaluating their models. These may be used in model training, in validation of models trained on experimental or in-house datasets, or in some combination of training and testing. The datasets discussed below require a free account to access, which academic users may request using a simple form (available at: https://go.drugbank.com/public_users/sign_up; account requests require approval, which may take up to two business days to process).
The Advanced Search functionality discussed in Section 3.3.1.2 provides a pow­erful mechanism for querying and ltering the complete sets of drugs and targets within DrugBank. The results of a search may be exported as a CSV le, by clicking on the “Export” button at the top of the search results list (Figure 3.6b). By combin­ing search ltering with a number of display eld selections it is possible to create a focused custom dataset that can easily be loaded into an ML system or relational database system as tabular data.
Whole datasets may also be accessed under the “Downloads” tab of the main menu bar (or by navigating to https://go.drugbank.com/releases/latest). The “Complete Database” provides a wealth of information for all current DrugBank drugs in XML format with an associated schema. Relevant attributes for drug discovery include drug approval status, structural information including classications, experimen­tal and predicted properties, and detailed target and metabolism information; other information is also provided, which may be useful depending on the desired use case.
Simplied scientic data extracts are available through the headings at the top of the download page in SDF,CSV, and FASTA formats. These datasets have the advan­tage of being presorted into various categories, such as those based on the drug type and approval status. In addition to their potential use in ML applications, these other formats provide additional compatibility with existing software. Drug structures in SDF format may serve as the basis for cheminformatic studies or for in silico struc­tural work. Sequences in FASTA format are easily used in comparative methods to nd similar sequences (e.g. Section 3.3.2.2) or to study target relationships using phylogenetics.
94 3 DrugBank Online: A How-to Guide
https://t.me/medicina_free
A key advantage of the DrugBank datasets is the combination of breadth, depth, and accuracy, made possible by the combination of automated and curated data intake within DrugBank (Sections 3.2.1 and 3.2.2.7, Figure 3.1). Studies have con­sistently demonstrated the importance of data completeness and accuracy in the performance of a wide variety of ML models (e.g. see [31]). The highly structured nature of these datasets allows users to easily experiment with the addition of new features to their models, while our commitment to depth (completeness) and accu­racy of data, and the inclusion of valid scientic references (Section 3.2.2.7) ensures data quality and transparency (accuracy).
3.4 Research Using DrugBank
As the pace of data and studies being released continues to grow at a breakneck rate, researchers are struggling with time-consuming work and an increasingly compet­itive environment. In order to stay ahead of the competition, many are turning to DrugBank for reliable, high-quality data.
Recently, a large team of researchers from Wuhan, Beijing, and Shenzhen developed a virtual screening tool using DrugBank’s database to help accelerate the drug discovery process relating to COVID-19 [32]. The team used DrugBank to lter out FDA-approved drugs as well as stage 3 clinical trial drugs. Several active sites of viral proteins were then chosen to use as ligand targets for a screening process. Compounds with high binding anities to these viral proteins were identied through in silico molecular docking experiments. The results included a number of drugs that were already being studied as treatments for COVID-19, but also identied a number of new possible candidates. The list included drugs that are used to treat HIV, HCV, cancer, and asthma, as well as inuenza virus antagonists. Through in silico screening such as this, researchers can quickly identify candidate treatments for emerging diseases such as COVID-19.
Turning to Sweden, a research team there has identied lead drug compounds using DrugBank to screen against viral targets responsible for COVID-19 [33]. They compiled lists of approved, investigational, and experimental drugs to screen against four COVID-19 targets: 3C-like protease (3CLpro), papain-like protease (PLpro), RNA-dependent RNA polymerase (RdRp), and the spike (S) protein. For this study, structures of drugs identied as having high binding anities were retrieved from DrugBank and were computationally docked to these four protein targets. The compounds were validated through a double-scoring approach using molecular dynamics and a molecular mechanics-generalized Born surface area (MM-GBSA) strategy. They found drugs that were already under review in COVID-19 clinical trials, conrming their methodology. To widen the pool of candidates, they also screened for compounds that could potentially act on multiple targets. DrugBank captures many dierent categorizations of drugs (such as approved, investigational, and experimental; see Section 3.2.2.2) allowing the user to lter down to those most important to their research. This is an optimal way to repurpose drugs given a vast database and known targets; researchers can readily identify leading compounds to accelerate the drug discovery process.
3.5 Discussion and Conclusions 95
https://t.me/medicina_free
Another instance where DrugBank helped to speedup drug repurposing strate­gies involved gene networking and bioinformatic analysis. Researchers from Taiwan and Indonesia uncovered potential treatments for atopic dermatitis (AD) by integrat­ing genetic and drug information using open data sources [34]. They gathered and mapped drug target genes to DrugBank and used parameters to lter out potential candidates based on pharmacological activity, approval status, as well as the pres­ence of clinical and experimental drugs. After running the data, the results showed dupilumab as an eective treatment for AD. As dupilumab is already approved for this indication, this nding provided evidence for the accuracy of their methodology. The researchers found 10 more potential candidates that had preclinical and clinical trial evidence linking them through genetic interactions with AD.
Another important step in discovery and repurposing studies is the experimental validation of predicted drugs, however, it is not always performed due to resource constraints or a variety of other factors. In one illustrative example, researchers from Argentina trained 1000 linear classiers on random subsets of independent vari­ables (molecular descriptors) to discriminate between known active and inactive inhibitors of the Plasmodium falciparum protease falcipain-2 [35]. Ensemble learn­ing was used to improve the predictive power over individual models, which was subsequently applied to the DrugBank and SWEETLEAD [36] databases to identify putative falcipain-2 inhibitors based on their positive predictive value (PPV). Of the 157 hits, four were tested for in vitro activity against puried falcipain-2. Methacy­cline, a tetracycline antibiotic, and odanacatib, an abandoned cathepsin K inhibitor investigated for use in osteoporosis, both inhibited the ability of falcipain-2 to cleave the peptidic substrate Z-LR-AMC and inhibited P. falciparum growth in culture. Interestingly, only odanacatib was able to inhibit proteolysis of the physiological substrate hemoglobin, highlighting the nuance of drug MoA vs. therapeutic eect.
As highlighted in this section, the accessibility of DrugBank’s extensive database can oer dierent solutions to accelerate the drug discovery pipeline. Researchers are able to use it for in silico research by identifying potential drug candidates as well as for repurposing molecules. Our vast range of interconnected information creates an ideal environment for streamlined drug discovery and continues to be a strong resource for researchers to use and validate drug prediction strategies. These examples represent just a few use cases, where DrugBank provided reli­able data to help nd treatments for a particular disease, some of which can be emerging.
3.5 Discussion and Conclusions
The increasing importance of in silico methodology across healthcare, and specif­ically within the domain of drug discovery, has the potential to revolutionize the manner in which we deliver care. To fully realize these benets it is necessary to have complete, well-structured, and accurate data. In this chapter,we have discussed the DrugBank database, highlighting several key datasets, providing workows to accomplish common tasks, and discussed several research studies that demonstrate the value of this data.
96 3 DrugBank Online: A How-to Guide
https://t.me/medicina_free
The datasets discussed herein represent useful resources for drug discovery work, but do not represent an exhaustive set of those within DrugBank. Clinical trial data, as an example, can be used eectively to interrogate drug repurposing opportunities and identify underserved areas within the scope of druggable targets. Though these data are available as part of DrugBank Online (Table 3.1), expansion of this dataset and the construction of tools to assist with its interrogation are current focus areas for improvement. Similarly, although the pharmacology (and specically pharma­cokinetic) information provided in drug cards (see Section 3.2.2.3), is exceptionally detailed, it is largely in the form of unstructured text. Future eorts to create struc­tured entries from this dataset may assist in the use of these parameters as input to various algorithms and ML models.
As mentioned in the introduction section, drug target identication is often conducted through the use of genetic associations, or is strengthened by such ndings [37]. In general, the integration of genetic information in clinical diag­nosis and care, though challenging, remains a source of great interest [38, 39]. The association of genomic changes with an alteration in the safety, ecacy, or other properties of a drug with respect to the individual is usually referred to as pharmacogenomics/pharmacogenetics (PGx) [40]. Although the utility of PGx data in a clinical setting has been demonstrated, the exploration of its use in other elds, such as drug discovery [41], remains to be fully evaluated. In keeping with this exciting potential, DrugBank will be investing in expanding our PGx dataset to empower new discoveries in the area of genomic medicine (see Section 3.2.2.3).
The importance of evolutionary context in drug discovery is largely limited to the use of homology searching and phylogenetic methods to ensure orthologues of puta­tive targets are present within an animal model of choice [42]. With the recent break­through in in silico protein structural prediction [43, 44], increasing power of in silico structural analysis methods, and a rm emphasis on validated drug–target interac­tions, it is likely that this conversation will expand from a purely sequence-focused view to include important structural elements. Assisting users with analytic work­ows centered around sequence and structural homology represents a fascinating possibility for future work.
Critically, though there are numerous avenues currently under development to expand DrugBank’s oering, our existing data and infrastructure already represent an invaluable resource for the drug discovery community. Although commercial licensing is available, we provide many important datasets for drug discovery free of charge; by making these data freely available to researchers, we aim to empower health informatics research and democratize the research process such that an indi­vidual’s ability to discover novel insights is not tied to resourcing. Future eorts to expand on these data and tools will continue this motivation, and ensure the con­tinued success of health informatics research.
References
1 Eder, J. and Herrling, P.L. (2015). Trends in modern drug discovery. In: New
Approaches to Drug Discovery, vol. 232, 3–22. Cham: Springer International Publishing.
References 97
https://t.me/medicina_free
2 Blay, V., Tolani, B., Ho, S.P., and Arkin, M.R. (2020). High-throughput screening:
today’s biochemical and cell-based approaches. Drug Discovery Today 25 (10): 1807–1821.
3 Dowden, H. and Munro, J. (2019). Trends in clinical success rates and therapeu-
tic focus. Nature Reviews. Drug Discovery 18 (7): 495–496.
4 Lewis, K. (2020). The science of antibiotic discovery. Cell 181 (1): 29–45. 5 Hughes, J., Rees, S., Kalindjian, S., and Philpott, K. (2011). Principles of early
drug discovery: principles of early drug discovery. British Journal of Pharmacol­ogy 162 (6): 1239–1249.
6 Heifetz, A., Southey, M., Morao, I. et al. (2018). Computational methods used
in hit-to-lead and lead optimization stages of structure-based drug discovery. In: Computational Methods for GPCR Drug Discovery, vol. 1705 (ed. A. Heifetz), 375–394. New York, NY: Springer, New York.
7 Fourches, D. and Ash, J. (2019). 4D- quantitative structure–activity relation-
ship modeling: making a comeback. Expert Opinion on Drug Discovery 14 (12): 1227–1235.
8 Kubota, K., Funabashi, M., and Ogura, Y. (2019). Target deconvolution from
phenotype-based drug discovery by using chemical proteomics approaches. Biochimica et Biophysica Acta, Proteins and Proteomics 1867 (1): 22–27.
9 Batool, M., Ahmad, B., and Choi, S. (2019). A structure-based drug discovery
paradigm. International Journal of Molecular Sciences 20 (11): 2783.
10 Chan, H.C.S., Shan, H., Dahoun, T. et al. (2019). Advancing drug discovery via
articial intelligence. Trends in Pharmacological Sciences 40 (8): 592–604.
11 Sanchez-Lengeling, B. and Aspuru-Guzik, A. (2018). Inverse molecular design
using machine learning: generative models for matter engineering. Science 361 (6400): 360–365.
12 WHO Collaborating Centre for Drug Statistics Methodology (2022). ATC classi-
index with DDDs.
cation
13 National Center for Biotechnology Information (2022). MeSH (Medical Subject
Headings). National Library of Medicine. Available at www.ncbi.nlm.nih.gov/ mesh.
14 CDER Manual of Policies and Procedures (2018). MAPP 7400.13: Determining
the Established Pharmacologic Class for Use in the Highlights of Prescribing Information. US FDA Center for Drug Evaluation and Research.
15 Djoumbou Feunang, Y. , Eisner, R., Knox, C. et al. (2016). ClassyFire: automated
chemical classication with a comprehensive, computable taxonomy. Journal of Cheminformatics 8 (1): 61.
16 Tetko, I.V. and Tanchuk, V.Y. (2002). Application of associative neural networks
for prediction of lipophilicity in ALOGPS 2.1 program. Journal of Chemical Information and Computer Sciences 42 (5): 1136–1145.
17 Cheng, F., Li, W. , Zhou, Y. et al. (2012). admetSAR: a comprehensive source and
free tool for assessment of chemical ADMET properties. Journal of Chemical Information and Modeling 52 (11): 3099–3105.
18 The UniProt Consortium, Bateman, A., Martin, M.-J. et al. (2021). UniProt:
the universal protein knowledgebase in 2021. Nucleic Acids Research 49 (D1): D480–D489.
98 3 DrugBank Online: A How-to Guide
https://t.me/medicina_free
19 Berman, H.M. (2000). The protein data bank. Nucleic Acids Research 28 (1):
235–242.
20 Burley, S.K., Bhikadiya, C., Bi, C. et al. (2021). RCSB Protein Data Bank: pow-
erful new tools for exploring 3D structures of biological macromolecules for basic and applied research and education in fundamental biology, biomedicine, biotechnology, bioengineering and energy sciences. Nucleic Acids Research 49 (D1): D437–D451.
21 Frolkis, A., Knox, C., Lim, E. et al. (2010). SMPDB: the small molecule pathway
database. Nucleic Acids Research 38 (Database issue): D480–D487.
22 Wishart Research Group (2010). Small Molecule Pathway Database. 23 (2018). The use of stems in the selection of International Nonproprietary Names
(INN) for pharmaceutical substances, World Health Organization, Geneva.
24 Chemaxon Marvin JS User’s Guide. Chemaxon Docs. 25 Chemaxon JChem Base Query Guide: Similarity search. Chemaxon Docs. 26 Altschul, S. (1997). Gapped BLAST and PSI-BLAST: a new generation of protein
database search programs. Nucleic Acids Research 25 (17): 3389–3402.
27 Freedman, D.H. (2019). Hunting for new drugs with AI. Nature 576 (7787):
S49–S53.
28 Sambasivan, N., Kapania, S., Highll, H. et al. (2021). “Everyone wants to do the
model work, not the data work”: data cascades in high-stakes AI. In: Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems, 1–15.
29 Halevy, A., Norvig, P., and Pereira, F. (2009). The unreasonable eectiveness of
data. IEEE Intelligent Systems 24 (2): 8–12.
30 Miranda, L. (2021). “Towards data-centric machine learning: a short review”.
ljvmiranda921.github.io.
31 Budach, L., Feuerpfeil, M., Ihde, N., Nathansen, A., Noack, N., Patzla, H.,
Harmouch, H., and Naumann, F. (2022). The Eects of Data Quality on Machine Learning Performance.
32 Xu, C., Ke, Z., Liu, C. et al. (2020). Systemic In Silico screening in drug discov-
ery for coronavirus disease (COVID-19) with an online interactive web server. Journal of Chemical Information and Modeling 60 (12): 5735–5745.
33 Murugan, N.A., Kumar, S., Jeyakanthan, J., and Srivastava, V. (2020). Searching
for target-specic and multi-targeting organics for Covid-19 in the Drugbank database with a double scoring approach. Scientic Reports 10 (1): 19125.
34 Adikusuma, W. , Irham, L.M., Chou, W.-H. et al. (2021). Drug repurposing for
atopic dermatitis by integration of gene networking and genomic information. Frontiers in Immunology 12: 724277.
35 Alberca, L.N., Chuguransky, S.R., Álvarez, C.L. et al. (2019). In silico guided
drug repurposing: discovery of new competitive and non-competitive inhibitors of falcipain-2. Frontiers in Chemistry 7: 534.
36 Novick, P.A., Ortiz, O.F., Poelman, J. et al. (2013). SWEETLEAD: an in Sil-
ico database of approved drugs, regulated chemicals, and herbal isolates for computer-aided drug discovery. PLoS One 8 (11): e79568.
References 99
https://t.me/medicina_free
37 Schmidt, A.F., Finan, C., Gordillo-Marañón, M. et al. (2020). Genetic drug tar-
get validation using Mendelian randomisation. Nature Communications 11 (1):
3255.
38 Jordan, D.M. and Do, R. (2018). Using full genomic information to predict dis-
ease: breaking down the barriers between complex and mendelian diseases. Annual Review of Genomics and Human Genetics 19 (1): 289–301.
39 Burke, W. (2021). Utility and diversity: challenges for genomic medicine. Annual
Review of Genomics and Human Genetics 22 (1): 1–24.
40 Roden, D.M., McLeod, H.L., Relling, M.V. et al. (2019). Pharmacogenomics. The
Lancet 394 (10197): 521–532.
41 Roses, A.D. (2008). Pharmacogenetics in drug discovery and development: a
translational perspective. Nature Reviews. Drug Discovery 7 (10): 807–817.
42 Holbrook, J.D. and Sanseau, P. (2007). Drug discovery and computational evolu-
tionary analysis. Drug Discovery Today 12 (19–20): 826–832.
43 Jumper, J., Evans, R., Pritzel, A. et al. (2021). Highly accurate protein structure
prediction with AlphaFold. Nature 596 (7873): 583–589.
44 Baek, M., DiMaio, F., Anishchenko, I. et al. (2021). Accurate prediction of pro-
tein structures and interactions using a three-track neural network. Science 373 (6557): 871–876.