Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5884_Библиотеки_им_академика_М_И_Перельмана.pdf
X
- •Foreword
- •Acknowledgments
- •Contents
- •1.1 Structure-Based Drug Discovery (SBDD)
- •1.2 Ligand-Based Drug Design (LBDD)
- •1.3 Echoes from the Past, Visions from the Future
- •References
- •1 Introduction
- •2.2 Second Step: Data Curation
- •2.4 Fourth Step: Updating and Maintenance
- •2 Databases and Curation
- •8 Perspectives
- •9 Conclusion
- •References
- •1 Introduction
- •2.1 Making and Matching Protein Models
- •2.2 Simulating Protein Movements
- •2.3 Analyzing Changes in Protein Shape
- •3 Pharmacogenomics in Drug Development
- •4 Case Studies of Genomics-Based Drug Design
- •References
- •1 Historical Background
- •1.1 Timeline
- •2 Methodology Overview
- •2.1 Neural Networks
- •2.1.1 Perceptron
- •2.1.2 Multilayer Neural Networks
- •2.1.3 Types of Neural Networks
- •Feedforward
- •Recurrent Neural Networks
- •LSTM
- •2.2 Deep Learning
- •3 Using Machine Learning
- •3.2 Data Collection
- •3.3 Data Preprocessing
- •3.4 Model Selection
- •3.5 Model Training
- •3.6 Validation
- •3.7 Tuning
- •3.8 Prediction
- •4 Limitations
- •4.1 Bias
- •4.3 Interpretability
- •4.4 Computational Cost
- •4.5 Data Dependency
- •4.6 Robustness
- •5 Applications in Drug Discovery
- •5.2 Lead Discovery
- •5.3 Preclinical and Clinical Development
- •6 Resources and Tools
- •7 Challenges and Perspectives
- •7.1 Future Trends
- •9 Conclusions
- •References
- •1 Historical Background
- •1.1 Applications in Drug Discovery
- •2 Validations and Controls
- •2.1 Internal Validation
- •2.2 External Validation
- •2.3 Relative Cluster Validation
- •3 Challenges and Perspectives
- •4 Conclusions
- •References
- •1 Historical Background
- •2 OECD Principles
- •2.1 A Defined Endpoint
- •2.2 An Unambiguous Algorithm
- •2.5 A Mechanistic Interpretation, if Possible
- •3 Software and Tools
- •4 Validations and Controls
- •4.1 Internal and External Validation
- •4.1.1 Regression Metrics
- •4.2 Applicability Domain
- •4.3 Randomization Tests
- •5 Interpretation
- •6 Practical Advice During QSAR Modeling
- •7 Application
- •8 Challenges and Perspectives
- •References
- •1 Molecular Docking
- •2 Advances in Scoring Functions and Search Algorithms
- •2.2 Critical Characteristics of Search Algorithms
- •2.3 Docking Programs and Scoring Functions
- •3 Calculations Performed During Docking Simulations
- •4 Essential Components for a Good Docking Program
- •5 Limitations of the Docking Technique
- •6 Validation of Docking Results
- •7 Inappropriate Use of Validation Methods in Docking
- •9 Use of Machine Learning in Molecular Docking
- •11 Challenges
- •12 Conclusions
- •References
- •3 System Preparation for MD Simulations
- •3.1 Solvation and Microensemble
- •3.2 Force Fields: General Concept and Relevant Choices
- •3.3 The Concept of Replicas and Timescale
- •4.1.2 Protein Root Mean Square Fluctuation (RMSF)
- •4.1.4 Protein Secondary Structure Analysis
- •4.1.5 Principal component Analysis (PCA)
- •4.1.6 Markov State Modelling
- •4.1.7 Distance Calculations
- •4.1.8 Angle and Plane Calculations
- •4.2.2 Distances and Ligand-Induced Geometry Rearrangements
- •4 Molecular Dynamics Analysis
- •4.1 Protein Perspective
- •4.1.1 Protein Root Mean Square Deviation (RMSD)
- •4.3 Ligand Perspective
- •4.3.1 Ligand Properties
- •4.3.2 Ligand Root Mean Square Deviation
- •4.3.3 Ligand Root Mean Square Fluctuation
- •4.3.4 Angles and Dihedrals
- •5.1 Protein Structure Prediction and Preparation
- •5.2 Molecular Docking
- •6 Concluding Remarks and Outlook
- •Glossary
- •References
- •1 Introduction
- •2.1 MDeNM
- •2.2 Collective Molecular Dynamics (coMD)
- •2.3 ClustENM and ClustENMD
- •3 Ensemble Docking
- •References
- •1 Introduction
- •1.1 Advantages, Disadvantages, Innovations, and Challenges
- •1.2 Recent Advances in Accessible FEP Software Tools
- •1.3 Applications of FEP in Industry and Consortiums
- •2 Expanding the Potential of FEP Calculations
- •2.1 Validating Binding Poses
- •2.2 Dealing with Solvent
- •2.3 FEP and Allostery
- •2.4 FEP and Covalent Ligands
- •2.5 Applications of FEP in Scaffold Hopping
- •2.6 Positional Analogue Scanning
- •2.7 Combinations and Alternative Approaches
- •3 Machine Learning for FEP
- •3.4 Implications for ML in FEP Calculations
- •4 Final Considerations
- •5 First Steps to FEP Simulations
- •References
- •1 Background
- •2 Ultra-Large Screening Libraries and Chemical Spaces
- •3.1 Implications of Dataset Size
- •4 Ligands on the Ultra-Large Scale
- •4.1 Ultra-Large 2D Similarity Searches
- •7 Challenges and Future Perspectives
- •7.1 Hit Triage: An Old Problem on a New Dimension
- •8 Conclusions
- •Appendix
- •References
- •1 Introduction
- •2 Enzymatic Activity Evaluations
- •3 Cytotoxicity Evaluation and Cell Viability
- •4 Antiviral Assays in Experimental Validation
- •6 In Vivo Evaluation of Compounds
- •7 Conclusions
- •References
- •1 Introduction
- •3.1 Data Collection
- •3.2 Data Preprocessing
- •3.4 Model Choice
- •3.5 Model Training
- •3.6 Model Assessment
- •3.7 External Validation
- •3.8 Implementation and Availability
- •3.9 Continuous Update
- •5 Conclusions and Perspectives
- •References
- •1 Experimental Approaches to Obtain Protein Structure
- •1.1 X-Ray Crystallography
- •1.2 Nuclear Magnetic Resonance
- •1.3 Cryo-EM
- •1.4 Hybrid Methods
- •2 Modeling Approaches to Obtain Protein Structure
- •2.1 Homology Modeling
- •2.2 Ab Initio Modeling
- •2.3 New Approaches
- •3 Conformational Diversity of Proteins
- •3.1 Characterization of Protein Conformational States
- •3.2 Experimental Methods to Study Protein Dynamics and Conformations
- •3.4 Molecular Dynamics Simulation
- •3.5 Sampling Strategies
- •4 Remarks and Perspectives
- •References
- •1 Introduction
- •2 Structure-Based Drug Design of HIV Protease Inhibitors
- •2.1 HIV-1 Protease as a Therapeutic Target
- •2.2.1 Saquinavir
- •2.2.2 Indinavir
- •2.3.1 Lopinavir
- •2.3.2 Darunavir
- •6 Conclusions
- •References
- •4 Experimental Methods to Analyze NR Activity
- •4.2 Coregulator-Recruitment
- •5 Concluding Remarks and Outlook
- •References

Data retrival
and curatio
132 P. O. Fernandes and V. G. Maltarollo
QSAR pipeline
Molecular
n
Fig. 6.1 Simplified process to build and validate a QSAR model
representation
of the dataset
Aplication of the
mathematical /
statistical
methods
Model
validation
nineteenth century was the correlation between melting- and boiling points described
by Edmund J. Mills in 1884 [2]. Since then, the field expanded, and in 1962, the
work from Hansch et al. [3] was considered a milestone for the modern QSAR
period. A relationship between phenoxyacetic acid derivatives and their biological
activity as plant growth regulators was established. In this work, they reported the
famous equation relating the biological activity ( y) and the descriptors (X) multiplied
by specific coefficients that could be interpreted as individual importance to modulate the modeled effect. In the following years, Corwin Hansch participated in two
other works [4, 5] that establish the basis of QSAR modeling used today, especially
the work from Fujita et al. [5] were described as a new substituent constant.
Pioneering works of these using multi-parametric regression are worth mentioning
to establish the structure–activity correlation. Later, the evolution of computers
allowed the use of new ways of representing molecules through theoretical descriptors, and today QSAR is an inseparable part of the drug design/discovery process [6].
As a mature research field, QSAR modeling has well-defined procedures for
exploring the relationship of chemical compounds to biological activity or other
properties. In this sense, the model that predicts other properties than the biological
activity is also called the quantitative structure–property relationship (QSPR) [7–9],
and other names are used for different purposes such as quantitative structure–
reactivity relationships (QSRR) [10–12] and quanti tative structure–toxicity relationship (QSTR) [ 13–15].
The construction of a model to predict a biological activity as a function of the
molecular structure is the primary goal, and it can be explained simply in four
general steps (Fig. 6.1). The first one is obtaining the data set, a group of molecules
with the endpoint (the biological activity or other property that will be modeled)
experimentally measured. These data can be obtained by in-house assays, nevertheless, it is more common to be retrieve datasets from the literature. Later, these data
should be carefully inspected to ensure their quality. Fourches et al. [16] in the paper
“Trust, but verify: on the importance of chemical structure curation in
cheminformatics and QSAR modeling research” discuss the errors in structural
representations observed in medicinal chemistry publications and how erroneous
structures and/or duplicate entries negatively influence the QSAR models, reducing
their statistical power and even leading to a complete failure. This and the follow-up

6 QSAR and Machine Learning Predictors 133
Fig. 6.2 Descriptors and fingerprints. (a) Representation of the molecular descriptors classification
by the information described by the class. (b) Schematic representation of a molecular fingerprint,
where the molecule is compared to specific fragments, and this information is combined in a binary
vector, called a fingerprint
work [17] described a pipeline to process the information to properly perform a data
curation.
Later, a mathematical representation of the molecular structure should be done as
a descriptor, a fingerprint, physicochemical properties, a graph, etc. A descriptor is a
numerical representation of a molecule using predefined rules and it can be obtained
from the molecular formula, bidimensional, or three-dimensional structure
[18, 19]. A way to classify them is related to the information obtained by using
the descriptor such as geometrical descriptors that reflect the 3D structure of the
molecule [ 19] (Fig. 6.2a). Similar to a descriptor, a fingerprint is also a representation
of the molecule based on specific rules, but in a vector, where each position (bin)
carries a different information [20] (Fig. 6.2b).
The descriptors can be obtained from the molecular formula, bidimensional, or
three-dimensional structure. Usually, the descriptors classify the QSAR model

134 P. O. Fernandes and V. G. Maltarollo
Table 6.1 Classification of the QSAR models according to their dimensionality of the molecular
descriptor
Class Examples of used descriptors
1D-QSAR Descriptors based on global molecular properties such as pk
weight, and others
2D-QSAR Descriptors based on structural patterns such as connectivity indices, 2D
pharmacophores, and molecular fragments
3D-QSAR Descriptors based on noncovalent interaction fields around 3D molecular models
4D-QSAR Descriptors that include multiple ligand conformations to the 3D QSAR model,
where the conformations can be generated by molecular dynamics
5D-QSAR Explicitly includes several induced-fit models to the 4D QSAR model
6D-QSAR Adds the solvation model to the 5D QSAR model
, LogP, molecular
a
according to their dimensionality [21, 22] (Table 6.1). For example, a 2D-QSAR
model is built using descriptors calculated from the bidimensional structure of the
molecules. A 3D-QSAR model uses the three-dimensional representation of a
molecule, while higher-dimensional QSAR (4D, 5D, 6D, and 7D) includes multiple
conformations and other structural information [23].
Once represented, the set of the endpoint vector ( y) and descriptors matrix (X)
representing the compounds (samples) are used to build a model, but first, they must
be split into training test sets. The training set compounds are used to teach the
algorithm what descriptors are important to predict a given endpoint, in other words,
to tune the relationship between X and y variables; and the test set is used to validate
the model built [24] (more about this process will be discussed in further sections).
QSAR modeling already has well-defined protocols and procedures to perform
the application of the techniques, especially with the growing collection of biological data available in databa ses such as ChEMBL [25] and PubChem [26]. The work
titled “Best practices for QSAR model development, validation and exploration”
developed by Alexander Tropsha [24] is an important starting point to adequately
build a QSAR model. Furthermore, the Organisation for Economic Co-operation and
Development (OECD) [27] principles also provide valuable checkpoints that must
be followed by the QSAR model for regulatory purposes.
From the Hansch [ 3 , 4] and Fujita [5] works, the use of linear models was well
adopted to build the relationships. One application example from that time is the
work of Di Paolo [28] who had built QSAR models to correlate the anesthetic
activity of ether derivatives. Another contemporaneous example is the one published
by Hansch and Klein [29] where they explore enzyme ligand interactions applying
QSAR techniques.
Over time, several QSAR methodologies have been developed and some of them
are worth mentioning. The first one is the one based on the Molecular Interaction
Fields (MIF) Comparative Molecular Field Analysis (CoMFA) [30, 31], Comparative Molecular Similarity Indices Analysis (CoMSIA) [32], and hologram quantitative structure-activity relationship (HQSAR) [33, 34]. These three QSAR

a
the
calculation
Virtual probe to p
o
a
cut
-of
f
val
Coulomb potential
(
)
Co
(opposite charges)
Len
nar
d-J
ones
pot
l
6 QSAR and Machine Learning Predictors 135
)
ue
entia
identical charges
ulomb potential
f the potenti
ls in the grid points
b)
Fragment
O
Generation
Biologica Activity x
36 2051081
Molecular Hologram
Fig. 6.3 Different QSAR methods: (a) schematic example of the grid-box used to calculate the
MIFs, a crucial step for performing CoMFA and CoMSIA QSAR. (b) Pipeline of the HQSAR
process. A molecular structure is fragmented, and these portions are combin ed into an array. Then, it
is correlated to the biological activity using a PLS model
O
PLS
HQSAR
Model
methodologies were successfully used in the field and they became a product
implemented in the SYBYL software, commercialized by Tripos Inc.
The CoMFA QSAR is based on the correlation of the steric and electrostatic fields
around the molecule and its biological activity. These two fields are calculated using
a virtual probe atom (usually a sp
3C+
atom) to model the stereochemical interactions
by the Len ard-Jones potential and to model the electrostatic interactions by the
Coulombic potential in specific x, y, z coordinates points into a lattice intersection
(also known as the grid box) (Fig. 6.3a). These two sterical and electrostatic MIFs in
the absence of a molecular target knowledge could be interpreted as “the negative”
of ligand-receptor representation, giving insights for molecular optimization aiming
to increase the potency of novel compounds. Due to the three-dimensional requirements and the need to analyze all compounds in the same position, molecular

136 P. O. Fernandes and V. G. Maltarollo
alignment is a key step to building a CoMFA model. A PLS regression is done using
the energy calculated from these fields in each point of the grid (X) and the endpoint
to be modeled (usually, a biological activity, y)[30]. Similarly, the CoMSIA QSAR
is based on the same principles of the CoMFA but computing similarity indices at
each point of the same grid, related to steric, electrostatic, and hydrophobic potentials obtained by a Gaussian-type function. Also, the PLS regression is used in the
CoMSIA methodology [32].
Both CoMFA and CoMSIA are dependent on the three-dimensional structure of
the compounds and their alignment and, therefore are considered 3D-QSAR
methods. The HQSAR was an alternative to work with bidimensional structures,
using structural fragments of a set of molecules (Fig. 6.3b). This is a proprietary
QSAR method implemented by Tripos in the SYBYL platform [33]. In this methodology the molecular structure is described as a hologram iteratively generated with
a fingerprint composed of molecular fragments built from each molecule structure.
Then, the PLS regression is applied to generate the QSAR model. Despite the
evolution of the QSAR field and the popularity of machine learning-based QSAR,
CoMFA, CoMSIA, and HQSAR are still used in drug design campaigns [35–40].
The QSAR field is continuously growing through the development of new
methods and applications. Since 2007 more than 1000 relevant articles have been
published annually [41]. Machine learning methods (see Chap. 4) have significantly
advanced QSAR research by enabling the understanding of complex relationships
not observed by linear models. This evolution did not just happen in the academia,
but today QSAR are valuable instrument in the industry [42]. For that reason, it is
important to follow the best practices such as the OECD principles, and to promote
high quality in QSAR modeling, those guidelines will be discussed in the next
section.
2 OECD Principles
Usually, the QSAR modeling prioritizes optimizing the model’s performance to
accurately replicate experimental outcomes. This approach leads to the development
of models with limited transparency and interpretability, commonly referred to as
“black boxes” such as Artificial Neural Networks (ANN), algorithms inspired by the
natural neural networks [43]. An extensive review of ANN-based algorithms is
available in Chap. 4. Despite their utility in predictive applications, these models
have not sufficiently supported the generation of reliable decisions within regulatory
applications [44], due to this behavior. This behavior and other misconceptions of
the QSAR modeling resulted in the publication of the “Guidance document on the
validation of (quantitative)structure-activity relationships [(Q)SAR] models” by the
OECD [45]. In this document, five principles to be addressed before the application
of a QSAR model (Fig. 6.4) were established. This publication marked a notable
advancement in the development of in silico models, emphasizing the necessity to
explicitly explain the model’s intended purpose, construction methodology, and

1.
A defined endpoint
An unambiguous algorithm
3
y
4.
Appropriate measures of goodness-of-fit, robustness and predictivity
5
6 QSAR and Machine Learning Predictors 137
OECD Principles
. A defined domain of applicabilit
. A mechanistic interpretation, if possible
Fig. 6.4 Five principles defined by the OECD to develop and validate a QSAR model for
reglementary purposes
performance metrics [44]. The following section will present in a more detailed, but
not intended to be exhaustive, way.
2.1 A Defined Endpoint
The first principle defines that a QSAR model should have a defined endpoint and
ensure the clarity of the prediction. The OECD suggests how specific an endpoint
should be to ensure reli ability based on the data used to build the model. It also
describes examples associated with the OECD Test Guidelines such as physicochemical properties (e.g., melting point, boiling point, and water solubility) environmental fate (e.g., biodegradation, hydrolysis, and bioaccumulation), ecological
effects (e.g., acute fish toxicity, and alga toxicity), and human health effects (e.g.,
acute oral toxicity, acute dermal toxicity, genotoxicity, and organ toxicity).
This transparency in the data is an essential requirement for regulatory agenci es to
validate the application of the QSAR model for a particular problem. This principle
also highlighted the importance of the quality of the data obtained and recommends
that ideally all QSAR should be developed using experimental data generated by a
single experimental protocol. This recomendation tries to avoid the introduction of
errors related to the interexperimental variability. For more details on experimental
assays, please refer to Chap. 12. Cronin and Schultz [46] also suggested that
variations in the analyst and the equipment could introduce noise to the biological
data. When it is not possible, they recommend a training set repres entative of the
different protocols [27].
For the endpoints that may result from diverse chemical mechanisms, QSAR
models should either be developed separately for each mechanism and applied to
narrowly defined classes of chemicals, or a broader QSAR relationship can be
established based on shared observations across noncongeneric chemical classes
with distinct mechanisms. Either approaches can be integrated, or a statistical

138 P. O. Fernandes and V. G. Maltarollo
approach capable of simultaneous global modeling across multiple mechanisms
must be employed to expand the domain of application on a broader scale.
2.2 An Unambiguous Algorithm
A QSAR model defines a mathematical relationship between the chemical structures
and the modeled endpoint, the way this relationship is built is the algorithm of the
model. For example, it can be a mathematical model, such as a linear regression or
PLS, or a set of knowledge-based rules, such as a decision tree. In this sense, an
unambiguous algorithm is capable of describing how the value was estimated and
can be reproduced if desired.
The OECD document presents some algorithms that can perform regression,
classification, or clustering (discussed in Chap. ). The ability to understand how
the estimation is performed contributes to the transparency of the model. Another
important aspect is the distinction between the transparency of the algorithm and the
ability to interpret it. For example, a linear regression can be transparent in terms of
the variables’ coefficients and how they are used to perform a prediction, however
the variables themselves may not have an understandable physicochemical meaning
based on the endpoint modeled. Two-dimensional autocorrelations and 3D Morse
descriptors [47] are one example of this problem. These descriptors are widely used
to build QSAR models and, despite the authors trying to give them a physical
meaning, it is very shallow and hard to understand how they affect the biological
activity.
Machine learning methods are more complex in comparison with traditional
regression methods (such as linear regression) due to the number of mathematical
operations and complex functions applied in the original descriptor values making it
hard to perform a straightforward interpretation. Due to this machine learning
researchers are working to develop strategies to easily interpret the variables of
these so-called black-box models. Recently, Rodríguez-Pérez and Bajorath [48]
described in a perspective article methodologies for better understanding the
machine learning models and their individual predictions, as well the current challenges for the integration of these techniques in medicinal chemistry in this field.
2.3 A Defined Domain of Applicabilit
This principle relies on the necessity to establish the scope and limitations of a
model, and it is based on the structure and physicochemical properties of the training
set. It is expected that a model can give reliable predictions for compounds similar to
those used to build the model and predictions outside this boundary are less likely to
be reliable because they are extrapolations of the model.

6 QSAR and Machine Learning Predictors 139
Due to the nature of the applicability domain (AD), each model has a specific one
based on the data used to train it. Also, the ou tcome of an AD can be categorical
(e.g., yes or no) or quantitative, determining the degree of similarity between the
predicted compound and the training data. There are different methods of similarity
to determine if a compound falls within the AD of a model. The OECD document
suggests some strategies to assess the AD, but the simplest one is observing the
range of the descriptors used in the model. If the predicted compound is between the
minimum and maximum values observed for each descriptor in the training set, it is
inside the domain. For the model displaying few descriptors, it is easy to observe;
nevertheless, it can be more compl icated in complex models.
2.4 Appropriate Measures of Goodness-of-Fit, Robustness,
and Predictivity
This OECD principle advocates for the necessity of statistical validation to ensure
the predictability of the model. The OECD document divides the assessment of
model performance (or statistical validation) into two stages: the internal performance (goodness-of-fit and robustness) and the external performance (predictivity).
The combination of these two methods is used to analyze the model, avoiding being
overfitted and underfitted. The target is a model not so simple that lacks information
and is not so complex to model the noise in the data [49, 50].
The statistical validation enables comparison between different models in order to
select the most predictive model. Another function of this process is to avoid
“spurious” models based on correlations by chance, which are not meaningful and
not predictive [51–53]. The inte rnal validation is a measure of the model’s performance using only the training set samples, which is the data used to build the model.
In this sense, models that could not predict the data used in the training process could
be discarded in the early stages of QSAR modeling. External validation is a way to
measure the performance using compounds unseen by the model (the section
validation and controls will further discuss how statistical metrics can be applied
to perform these validations) and, therefore, represents an application of QSAR in
actual drug discovery pipeline mimicking the prediction of compounds in commercial libraries. For that reason, it is important to assess the applicability domain to
ensure a test set representative of the training set.
2.5 A Mechanistic Interpretation, if Possible
Historically, the first QSAR models were developed using congeneric compounds
(molecules from the same chemical class with minor variations in substituents), and
the activity was related to biological activities produced by the compounds’

140 P. O. Fernandes and V. G. Maltarollo
molecular structure. The statistical methods were applied to describe the relationships and support chemical knowledge, beginning the idea of mechanistic interpretation. This OECD principle is not mandatory but is desirable, once a QSAR model
is consistent if the knowledge of the chemical/toxicological process provides more
credibility and acceptance of the predictions [27]. Furthermore, the interpret ation of
a model allows an “extra layer” of validation since the proposed mechanism could be
checked according to its chemical meaning and if it is expected or not. For example,
it is expected that hydrogen bond descriptors positively influence the aqueous
solubility of a QSPR model [54]; otherwise, there is a high possibility that the
model is not correct .
The “if possible” clause added in the principle is due to the interactive nature of
the modeling process. Usually, data exploration and modeling lead to the generation
and testing of a hypothesis, and a useful QSAR model may lack mechanistic
interpretation due to different reasons. Nevertheless, this principle motivates the
modeler to seek a mechanistic interpretation that can contribute to the understanding
of the statistic validation [27].
3 Software and Tools
Unless you are performing your own assays, the first step to performing QSAR
modeling is retrieving information from the literature. This process can be made
manually by compiling published works relevant to the project or using available
databases. The second approach is very common today since there are several
initiatives available. One popular example is the ChEMBL database [55], a largescale database comprising several bioactivity information. The compounds’ chemical structure and detailed information about the assays are provided, comprising
binding measurements, functional assays, pharmacokinetic properties, and toxicity.
Besides ChEMBL there are other large databases such as PubChem [56] and
BindingDB [57]. A detailed explanation of available chemical information sources
is described in Chap. 2.
The next importan t step in the development of a QSAR model is the representation of the molecules. To obtain these representations, the molecular descriptors
and/or fingerprints are calculated using specific software and web servers. A
nonexhaustive list of available tools to perform descriptors/fingerprint calculations
is described in Table 6.2.
Another emerging way to represent molecules is using learned embeddings such
as convolutional and graph encoding to represent the molecular structures. This
approach is focused on models capable of learning generalizable representations
from a molecular data set. Convolutional encodings perform mathematical operations compacting the space containing the variables. One example is the work from
Coley et al. [74] where atom and bond attributes were used in a convolutional neural
network to represent the molecules. The authors applied the methodology to predict
aqueous solubility, octanol solubility, melting point, and toxicity. Another

6 QSAR and Machine Learning Predictors 141
Table 6.2 Software and web servers available for descriptors and fingerprint calculation
Name Desc. FP Source
Dragon [58] 5270 1 https://www.talete.mi.it/
CDK [59] 275 9 https://cdk.github.io/
E-dragon [60] 1600 – https://vcclab.org/lab/edragon/
Mold2 [61] 779 – https://www.fda.gov/science-research/bioinformatics-tools/
Pybel [62]244https://github.com/pybel/pybel
PaDEL [63] 1875 12 https://www.yapcwsoft.com/dd/padeldescriptor/
RDKit 198 8 https://www.rdkit.org/
PyDPI [64] 615 7 https://pypi.org/project/pydpi/
Chemopy [65] 1135 7 https://github.com/ifyoungnet/Chemopy
ChemDes [66] 3679 59 http://www.scbdd.com/chemdes/
Rcpi [67] 308 10 https://bioconductor.org/packages/release/bioc/html/Rcpi.html
BioTriangle
[68]
ChemSAR [69] 783 10 http://chemsar.scbdd.com/
Mordred [70] 1825 – https://github.com/mordred-descriptor/mordred
PyBioMed [71] 775 19 https://github.com/gadsbyfly/PyBioMed
alvaDesc [72] 566 3 https://www.alvascience.com/alvadesc/
BioMedR [73] 293 13 https://github.com/wind22zhu/BioMedR
Desc. number of calculated descriptors; FP number of calculated fingerprint sets
540 7 http://biotriangle.scbdd.com/
mold2
Fig. 6.5 Simplified
example of a graph
representation of a molecule
node
Node properties:
atom type
chirality
hybridization
aromaticity
edge
Edge properties:
bond type
ring
stereochemistry
application of a convolutional encoding is the use of 3D information to predict
protein-ligand binding affinity [75]. In the graph approach, atoms are represented as
nodes and the bonds are represented by edges; both nodes and edges have associated
features such as the atomic number, formal charge, and bond type (Fig. 6.5). It
produces a 2D data structure; nevertheless, 3D information such as chirality or
stereochemistry can be included in nodes or edges [76]. One example of graph
usage is the Chemprop [77], a machine-learned package that has been used to build
models to predict biological activities [78, 79], molecular properties [80, 81], and
other chemical endpoints [82]. The work from McGibbon et al. [18] titled “From
intuition to AI: evolution of small molecule representations in drug discovery”
Соседние файлы в папке Библиотека им академика М.И. Перельмана
