Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5580_Библиотеки_им_академика_М_И_Перельмана
.pdf
Chapter 9
https://t.me/med1917
Drug Discovery and Drug Repositioning
Using Computational Methods
Yoshihiro Yamanishi
9.1 Introduction
The success rate of drug discovery has been extremely low recently. It costs more
than one billion U.S. dollars and takes more than ten years. As an efficient strategy
to overcome this slump in new drug discovery, drug repositioning (also called drug
repurposing or drug rescue) has been attracting attention. The aim of the drug repositioning is to discover new efficacy of existing drugs (already approved drugs and
compounds whose development failed in the past due to lack of efficacy) for different
diseases. For existing drugs, the information on human safety, pharmacokinetics,
and manufacturing processes can be used, and some of the ordinary drug development processes can be skipped, allowing for rapid, low-risk, and low-cost drug
development [
Looking back at history, many new drugs were developed by discovering new
indications for existing drugs. For example, minoxidil was originally developed as a
drug for hypertension, but is now used as a hair growth drug. Sildenafil was developed
as a treatment for angina pectoris, but is now used as a treatment for male dysfunction
and pulmonary arterial hypertension. Bupropion was an antidepressant agent, but
has been successfully developed as a smoking cessation aid. However, past success
stories have largely relied on human inspiration and serendipity, and most of the
additional drug effects were discovered by chance.
In recent biomedical science, it has become possible to obtain omics information
such as the genome, transcriptome, proteome, metabolome, phenome, and interactome, enabling us to comprehensively analyze various molecules and diseases. At
the same time, advances in technologies such as combinatorial chemistry and highcontent screening have led to the accumulation of chemical and physiological activity
1].
Y. Yamanishi (B)
Department of Complex Systems Science, Graduate School of Informatics, Nagoya University,
Nagoya, Japan
e-mail: yamanishi@i.nagoya-u.ac.jp
© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2024
H. Satoh et al. (eds.), Drug Development Supported by Informatics,
https://doi.org/10.1007/978-981-97-4828-0_9
165

166 Y. Yamanishi
https://t.me/med1917
information on a vast number of compounds and drugs. Such biomedical big data
would be useful resources for drug discovery and drug repositioning. This chapter
reviews the recent trends in polypharmacology research and computational methods
for drug repositioning using various big data and machine learning (a fundamental
technology of artificial intelligence).
9.2 Polypharmacology
9.2.1 Drug–protein Interactions
In drug discovery, identification of interactions between drugs (or drug candidate compounds) and proteins is an important issue. Many drugs exert their efficacy against diseases by interacting with therapeutic target proteins and other
biomolecules and inhibiting or activating their functions. Drugs interact not only
with a single therapeutic target protein but also with multiple other proteins (offtargets), causing not only the desired effects but also various side effects. However,
certain side effects may be effective for patients with other diseases. If all the target
proteins of a drug, including off-targets, are clarified, it will lead to a system-level
understanding of the drug’s mechanism of action and make it possible to predict
potential efficacy and side effects.
The idea of considering not only a single target protein but also all proteins, that
is, the entire proteome, is called polypharmacology. Polypharmacology is becoming
a major research theme in drug discovery research. However, it is difficult to experimentally identify interactions between drugs and all proteins because it requires a
huge amount of expenditure and time. The use of big data on drugs and proteins for
predicting unknown drug–protein interactions at the genome-wide scale is expected
to narrow down the candidates for experiments. There is an incentive to develop
computational screening methods for drugs against various families of proteins such
as G protein-coupled receptors, ion channels, enzymes, transporters, and nuclear
receptors.
In the concept of polypharmacology, the interaction between drugs and target
proteins is considered to be a many-to-many relationship rather than a one-to-one
relationship. Considering the problem of predicting drug–protein interactions from a
machine learning perspective, it can be formulated as a problem of classifying drug–
protein pairs into interaction classes or other classes. Machine learning is a fundamental technology of artificial intelligence, and is a method of extracting patterns
useful for classification and regression from data, learning a predictive model, and
using the model to make predictions for new data. Various algorithms have been
proposed to predict drug–protein interactions. Figure
of a classification method for predicting drug–protein interactions using machine
learning.
9.1 shows a conceptual diagram

9 Drug Discovery and Drug Repositioning Using Computational Methods 167
https://t.me/med1917
Fig. 9.1 Machine learning to predict drug-protein interaction pairs
As methods for learning predictive models, a verity of algorithms such as
logistic regression, support vector machines, matrix decomposition, and deep neural
networks have been proposed. The information on known drug–protein interactions
for model training is available from various databases (e.g., KEGG [
3], SuperTarget [4], Matador [4], DrugBank [5], BindingDB [6], and Therapuetic
[
Target Database [
7]).
2], ChEMBL
9.2.2 Methodological Framework Based on Data
Characteristics
The performance of computational methods largely depends on the nature and
comprehensiveness of the data representing drugs and proteins. The methods of
previous research can be categorized into “chemogenomics”, which uses chemical
structure information, “phenomics”, which uses phenotypic information about drugs
on the human body, and “transcriptomics”, which uses transcriptome information
with drug administration. Figure
In chemogenomics, the basic strategy is to explore the correlation between the
chemical space of drugs and the genome space of proteins. Drugs with similar chemical structures are predicted to interact with similar proteins [
seen as an extension of the concept of structure–activity relationships based on ligand
information from a single target protein to multiple target proteins. Examples of
descriptors for drugs and proteins include the drug’s chemical structure and physicochemical properties, the protein’s amino acid sequence, structure, and functional site
(domains, motifs, ligand-binding pockets, etc.). However, the prediction accuracy
largely depends on the descriptor of the chemical structure and proteins [
The basic principle of phenomics is to analyze the phenotypes that drugs have on
the human body (various patient reactions such as drug efficacy and side effects upon
drug administration). Drugs with similar phenotypes are predicted to interact with
16–18
similar proteins [
lowered blood pressure, changes in biomarkers, and tumor shrinkage/expansion.
]. Examples of phenotypes include elevated mood, increased/
9.2 shows these three frameworks.
8–14]. It can also be
15].

168 Y. Yamanishi
https://t.me/med1917
Fig. 9.2 Methodological framework for predicting drug-protein interaction pairs
For example, the use of thousands of types of phenotypes listed in drug–package
inserts and post-marketing surveillance reports has been proposed. Because it does
not use information on the drug chemical structures, it may be possible to discover
drug–protein interactions that cannot be imagined from the chemical structures.
The basic strategy of transcriptomics is to analyze drug-induced gene expression
profiles when drugs are exposed to various human cell lines. Drugs with similar
gene expression patterns are predicted to interact with similar proteins [
19–21].
In recent years, databases of drug-induced gene expression information have been
established around the world. For example, the Connectivity Map (CMap) contains
gene expression profiles obtained when approximately 1300 drugs were exposed to
four types of human cell lines [
22]. Its successor database, the Library of Integrated
Network-based Cellular Signatures (LINCS), contains gene expression profiles for
23
approximately 20,000 drugs and 77 human cell lines [
]. The Toxicogenomics
Project-Genomics Assisted Toxicity Evaluation System (TG-GATEs) contains gene
expression profiles of approximately 150 drugs exposed to individual rats and rat/
24
human hepatocytes [
]. These databases are useful resources in drug discovery
research.

9 Drug Discovery and Drug Repositioning Using Computational Methods 169
https://t.me/med1917
9.3 Drug Repositioning Approach
9.3.1 Inverse Correlation Method Based on Gene Expression
Profiles
Drug repositioning can be viewed as a problem in predicting the potential efficacy
of drugs for various diseases. One widely used computational method is to compare
the drug-induced gene expression profile and the disease-specific gene expression
profile of patients [
Since the biological system consists of the coordinated expression of many genes
encoded in the genome, the pathology of disease can be viewed as a disorder of
the gene expression pattern. The ideal role of a therapeutic drug for a disease is to
recover the gene expression pattern from the disease state to the normal state. Thus,
it is desired that a therapeutic drug counteracts the disease-specific gene expression
pattern. From the viewpoint, the selection of drugs with gene expression profiles that
are inversely correlated with the disease-specific gene expression profile as potential
therapeutic agents for that disease has been proposed (see Fig.
In fact, the discovery of drugs (or drug candidate compounds) effective against
Alzheimer’s disease, inflammatory bowel disease, prostate cancer, and colon cancer
has been reported. However, it is necessary to keep in mind that even for the same drug
or disease, gene expression profiles vary between cell lines, measurement conditions,
and individuals in practical applications.
25–29].
9.3).
Fig. 9.3 Prediction of new applicable diseases of drugs using inverse correlation

170 Y. Yamanishi
https://t.me/med1917
Fig. 9.4 Prediction of new
applicable diseases of drugs
using machine learning
9.3.2 Prediction of Drug–disease Networks
Here, we will consider an approach using machine learning. Drug efficacy can be
viewed as a network of relationships between drugs and diseases. From a machine
learning perspective, it can be formulated as a problem of predicting the presence
or absence of relationships between drug–disease pairs. For example, drug–profiles
(e.g., chemical structure descriptors, physicochemical features, target proteins, and
drug-induced gene expression information) and disease profiles (e.g., pathogenic
genes, pathway abnormalities, environmental factors, diagnostic markers, and patient
gene expression information) can be analyzed [
have been developed to classify drug–disease pairs into “related” or “unrelated”
classes, as shown in Fig.
In fact, our group has applied this method to the analysis of 2349 drugs in Japan,
Europe, and the United States and made large-scale predictions for 858 diseases
defined by the International Classification of Diseases (e.g., cancers, immune system
diseases, neurodegenerative diseases, and psychiatric disorders) [
alendronate is a drug for osteoporosis, but it was predicted to be effective against
breast cancer. This is consistent with recent clinical research reports. The validity of
many other drug–disease pairs with high prediction scores was confirmed in recent
literature and clinical reports. The predicted results that could not be confirmed in
the literature may represent new discoveries, and are considered to be of high value
to be verified in detail experimentally and clinically.
9.4.
30, 31]. Machine learning algorithms
31]. For example,
9.3.3 Prediction of Drug–target Protein–disease Networks
Since the prediction process of machine learning methods for predicting drug–disease
networks is a black box, it is difficult to obtain knowledge of the mechanisms related
to efficacy. Here, we introduce a method based on polypharmacological information
that considers target proteins of a drug.

9 Drug Discovery and Drug Repositioning Using Computational Methods 171
https://t.me/med1917
There are many drugs whose mechanisms of action are unknown, and the target
proteins involved in drug efficacy are unknown for more than half of the approved
drugs. Furthermore, little is known about the off-target proteins of drugs. Therefore,
based on the concept of polypharmacology, we aim to estimate the potential target
proteins of drugs, including off-targets, toward the prediction of drug efficacy.
Figure 9.5 shows the procedure for predicting applicable diseases based on drug–
target protein information [
32]. Suppose there is drug X that is effective against
disease A. If the target protein of drug X is unknown, we estimate the target protein.
If the estimated target protein has the potential to be a therapeutic target for disease
B based on the similarity of the molecular mechanisms of the disease, drug X is
predicted to be effective for disease B as well. Furthermore, if the estimated offtarget protein of drug X is a therapeutic target for disease C, drug X is also predicted
to be effective against disease C.
In fact, our group applied chemogenomics, phenomics, and transcriptomics
methods to analyze 8270 drugs in Japan, Europe, and the United States, and
estimated the potential drug–target proteins on a genome-wide scale. Then, for
1401 diseases, we made large-scale predictions of new indications of drugs [
32].
For example, pioglitazone, a type 2 diabetes treatment that targets peroxisome
proliferator-activated receptor (PPARγ), was predicted to interact with monoamine
oxidase (MAOB) as an off-target. The neurotransmitter dopamine is reduced in
patients with Parkinson’s disease, and MAOB is a dopamine degrading enzyme,
so suppressing dopamine degradation through inhibition of MAOB may have a therapeutic effect on Parkinson’s disease. In fact, clinical studies demonstrating the effectiveness of pioglitazone for Parkinson’s disease have been reported in recent years,
suggesting the validity of the computational prediction results.
Fig. 9.5 Prediction of new applicable diseases of drugs based on polypharmacological information

172 Y. Yamanishi
https://t.me/med1917
In an example using transcriptomics, it was predicted that the antipsychotic drug
phenothiazine interacts with the androgen receptor (AR), so the drug might also
be effective against prostate cancer [
21]. The drug whose gene expression pattern
was most similar to the training data was enzalutamide. As a result of pathway
enrichment analysis of a group of genes whose expression is actually regulated by
drugs, we were able to suggest that both enzalutamide and phenothiazine activate
the apoptotic pathway, suggesting that they may be effective against prostate cancer
through a similar mechanism. Furthermore, when the inhibitory effect on AR was
experimentally confirmed in vitro, strong inhibitory activity was observed. In this
way, clarifying the potential target proteins of existing drugs, including off-targets,
will directly lead to expanding the indications of existing drugs.
9.4 Conclusion
In this chapter, we introduced computational methods for drug discovery and drug
repositioning using various biomedical big data on diseases, genes, proteins, drugs,
and small compounds. Although this chapter introduced only some examples for
approved drugs, the computational methods can be applied to compounds other
than approved drugs as long as the data representations are available, and can also
be used for screening new drug candidate compounds. However, all methods have
advantages and disadvantages, so it is necessary to use them appropriately depending
on the purpose and nature of the data. Due to innovative advances in experimental
and measurement techniques in recent years, the types and amounts of data continue
to increase year by year, but actual data has unique difficulties because it contains
a lot of noise, missing values, and large biases. The importance of computational
methods based on statistics and machine learning will continue to increase in order
to efficiently extract useful information from such huge amounts of data toward drug
discovery.
References
1. Chong CR, Sullivan, DJ (2007) New uses for old drugs. Nature 448:645–646. https://doi.org/
10.1038/448645a
2. Kanehisa M, Goto S, Furumichi M, Tanabe M, Hirakawa M (2009) KEGG for representation and analysis of molecular networks involving diseases and drugs. Nucleic Acids Res
38(SUPPL_1):D355–D360.
3. Gaulton A, Bellis LJ, Bento AP, Chambers J, Davies M, Hersey A, et al (2012) ChEMBL:
A large-scale bioactivity database for drug discovery. Nucleic Acids Res 40:D1100–D1107.
https://doi.org/10.1093/nar/gkr777
4. Günther S, Kuhn M, Dunkel M, Campillos M, Senger C, Petsalaki E, et al (2008) SuperTarget
and Matador: Resources for exploring drug-target relationships. Nucleic Acids Res 36(SUPPL_
1):919–922.
https://doi.org/10.1093/nar/gkp896
https://doi.org/10.1093/nar/gkm862

9 Drug Discovery and Drug Repositioning Using Computational Methods 173
https://t.me/med1917
5. Knox C, Law V, Jewison T, Liu P, Ly S, Frolkis A, et al (2011) DrugBank 3.0: A comprehensive
resource for “Omics” research on drugs. Nucleic Acids Res 39(SUPPL_1):1035–1041.
doi.org/10.1093/nar/gkq1126
6. Liu T, Lin Y, Wen X, Jorissen R N, Gilson MK (2007) BindingDB: A web-accessible database
of experimentally determined protein-ligand binding affinities. Nucleic Acids Res 35(SUPPL_
1):D198–D201.
7. Qin C, Zhang C, Zhu F, Xu F, Chen SY, Zhang P, et al (2014) Therapeutic target database
update 2014: A resource for targeted therapeutics. Nucleic Acids Res 42:1118–1123.
doi.org/10.1093/nar/gkt1129
8. Nagamine N, Sakakibara Y (2007) Statistical prediction of protein–chemical interactions based
on chemical structure and mass spectrometry data. Bioinformatics 23:2004–2012.
org/10.1093/bioinformatics/btm266
9. Yamanishi Y, Araki M, Gutteridge A, Honda W, Kanehisa M (2008) Prediction of drug–target
interaction networks from the integration of chemical and genomic spaces. Bioinformatics
24:i232–i240.
10. Faulon J-L, Misra M, Martin S, Sale K, Sapra R (2008) Genome scale enzyme–metabolite and
drug–target interaction predictions using the signature molecular descriptor. Bioinformatics
24:225–233.
11. Jacob L, Hoffmann B, Stoven V, Vert J-P (2008) Virtual screening of GPCRs: an in silico
chemogenomics approach. BMC Bioinformatics 9:1–16.
9-363
12. Keiser MJ, Setola V, Irwin JJ, Laggner C, Abbas AI, Hufeisen SJ, et al (2009) Predicting new
molecular targets for known drugs. Nature 462:175–181.
13. Tabei Y, Pauwels E, Stoven V, Takemoto K, Yamanishi Y (2012) Identification of chemogenomic features from drug–target interaction networks using interpretable classifiers. Bioinformatics 28:i487–i494.
14. Meslamani J, Rognan D (2011) Enhancing the accuracy of chemogenomic models with a threedimensional binding site kernel. J Chem Inf Model 51:1593–1603.
00166t
15. Sawada R, Kotera M, Yamanishi Y (2014) Benchmarking a wide range of chemical descriptors
for drug-target interaction prediction using a Chemogenomic approach. Mol Inform 33:719–
731.
16. Campillos M, Kuhn M, Gavin A-C, Jensen LJ, Bork P (2008) Drug target identification using
17. Yamanishi Y, Kotera M, Kanehisa M, Goto S (2010) Drug-target interaction prediction from
18. Takarabe M, Kotera M, Nishimura Y, Goto S, Yamanishi Y (2012) Drug target prediction using
19. Wang K, Sun J, Zhou S, Wan C, Qin S, Li C, et al (2013) Prediction of drug-target interac-
20. Hizukuri Y, Sawada R, Yamanishi Y (2015) Predicting target proteins for drug candidate
21. Iwata, M, Sawada R, Iwata H, Kotera M, Yamanishi Y (2017) Elucidating the modes of
22. Lamb J, Crawford ED, Peck D, Modell JW, Blat IC, Wrobel MJ, et al (2006) The Connectivity
23. Subramanian A, Narayan R, Corsello SM, Peck DD, Natoli TE, Lu X, et al (2017) A next
https://doi.org/10.1002/minf.201400066
side-effect similarity. Science 321:263–266.
chemical, genomic and pharmacological data in an integrated framework. Bioinformatics
26:i246–i254.
adverse event report systems: a pharmacogenomic approach. Bioinformatics 28:i611–i618.
https://doi.org/10.1093/bioinformatics/bts413
tions for drug repositioning only based on genomic expression similarity. PLoS Comput Biol
9:e1003315.
compounds based on drug-induced gene expression data in a chemical structure-independent
manner. BMC Med Genomics 8:1–10.
action for bioactive compounds in a cell-specific manner by large-scale chemically-induced
transcriptomics. Scientific Reports 7:40164.
Map: using gene-expression signatures to connect small molecules, genes, and disease. Science
313:1929–1935.
generation connectivity map: L1000 platform and the first 1,000,000 profiles. Cell 171:1437–
1452.
https://doi.org/10.1016/j.cell.2017.10.049
https://doi.org/10.1093/nar/gkl999
https://doi.org/10.1093/bioinformatics/btn162
https://doi.org/10.1093/bioinformatics/btm580
https://doi.org/10.1186/1471-2105-
https://doi.org/10.1038/nature08506
https://doi.org/10.1093/bioinformatics/bts412
https://doi.org/10.1021/ci2
https://doi.org/10.1126/science.1158140
https://doi.org/10.1093/bioinformatics/btq176
https://doi.org/10.1371/journal.pcbi.1003315
https://doi.org/10.1186/s12920-015-0158-1
https://doi.org/10.1038/srep40164
https://doi.org/10.1126/science.1132939
https://
https://
https://doi.

174 Y. Yamanishi
https://t.me/med1917
24. Igarashi Y, Nakatsu N, Yamashita T, Ono A, Ohno Y, Urushidani T, et al (2015) Open TGGATEs: a large-scale toxicogenomics database. Nucleic Acids Res 43:D921–D927.
doi.org/10.1093/nar/gku955
25. Dudley JT, Sirota M, Shenoy M, Pai RK, Roedder S, Chiang AP, et al. (2011) Computational
repositioning of the anticonvulsant topiramate for inflammatory bowel disease. Sci Transl Med
3:96ra76–96ra76.
26. Sirota M, Dudley JT, Kim J, Chiang AP, Morgan AA, Sweet-Cordero A, et al (2011) Discovery
and preclinical validation of drug indications using c ompendia of public gene expression data.
Sci Transl Med 3:96ra77–96ra77.
27. Kosaka T, Nagamatsu G, Saito S, Oya M, Suda T, Horimoto K (2013) Identification of drug
candidate against prostate cancer from the aspect of somatic cell reprogramming. Cancer Sci
104:1017–1026.
28. van Noort V, Schölch S, Iskar M, Zeller G, Ostertag K, Schweitzer C, et al (2014) Novel
drug candidates for the treatment of metastatic colorectal cancer through global inverse geneexpression profiling. Cancer Res 74:5690–5699.
3540
29. Iwata M, Kosai K, Ono Y, Oki S, Mimori K, Yamanishi Y (2022) Regulome-based characterization of drug activity across the human diseasome. NPJ Syst Biol Appl 8:44.
10.1038/s41540-022-00255-4
30. Gottlieb A, Stein GY, Ruppin E, Sharan R (2011), PREDICT: A method for inferring novel
drug indications with application to personalized medicine. Mol Syst Biol 7:496.
org/ https://doi.org/10.1038/msb.2011.26
31. Iwata H, Sawada R, Mizutani S, Kotera M, Yamanishi Y (2015) Systematic drug repositioning
for a wide range of diseases with integrative analyses of phenotypic and molecular data. J Chem
Inf Model 55(2): 446–459.
32. Sawada R, Iwata H, Mizutani S, Yamanishi Y (2015) Target-based drug repositioning using
large-scale chemical-protein interactome data. J Chem Inf Model 55(12):2717–2730.
doi.org/10.1021/acs.jcim.5b00330
https://doi.org/10.1126/scitranslmed.3002648
https://doi.org/10.1126/scitranslmed.3001318
https://doi.org/10.1111/cas.12183
https://doi.org/10.1158/0008-5472.CAN-13-
https://doi.org/10.1021/ci500670q
https://
https://doi.org/
https://doi.
https://
Соседние файлы в папке Библиотека им академика М.И. Перельмана
