Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5606_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
02.09.2026
Размер:
21 Мб
Скачать
Data retrival
and curatio
132 P. O. Fernandes and V. G. Maltarollo
QSAR pipeline
Molecular
n
Fig. 6.1 Simplied process to build and validate a QSAR model
representation
of the dataset
Aplication of the
mathematical /
statistical
methods
Model
validation
nineteenth century was the correlation between melting- and boiling points described by Edmund J. Mills in 1884 [2]. Since then, the eld expanded, and in 1962, the work from Hansch et al. [3] was considered a milestone for the modern QSAR period. A relationship between phenoxyacetic acid derivatives and their biological activity as plant growth regulators was established. In this work, they reported the famous equation relating the biological activity ( y) and the descriptors (X) multiplied by specic coefcients that could be interpreted as individual importance to modu­late the modeled effect. In the following years, Corwin Hansch participated in two other works [4, 5] that establish the basis of QSAR modeling used today, especially the work from Fujita et al. [5] were described as a new substituent constant. Pioneering works of these using multi-parametric regression are worth mentioning to establish the structure–activity correlation. Later, the evolution of computers allowed the use of new ways of representing molecules through theoretical descrip­tors, and today QSAR is an inseparable part of the drug design/discovery process [6].
As a mature research eld, QSAR modeling has well-dened procedures for exploring the relationship of chemical compounds to biological activity or other properties. In this sense, the model that predicts other properties than the biological activity is also called the quantitative structure–property relationship (QSPR) [79], and other names are used for different purposes such as quantitative structure– reactivity relationships (QSRR) [1012] and quanti tative structure–toxicity relation­ship (QSTR) [ 1315].
The construction of a model to predict a biological activity as a function of the molecular structure is the primary goal, and it can be explained simply in four general steps (Fig. 6.1). The rst one is obtaining the data set, a group of molecules with the endpoint (the biological activity or other property that will be modeled) experimentally measured. These data can be obtained by in-house assays, neverthe­less, it is more common to be retrieve datasets from the literature. Later, these data should be carefully inspected to ensure their quality. Fourches et al. [16] in the paper Trust, but verify: on the importance of chemical structure curation in cheminformatics and QSAR modeling researchdiscuss the errors in structural representations observed in medicinal chemistry publications and how erroneous structures and/or duplicate entries negatively inuence the QSAR models, reducing their statistical power and even leading to a complete failure. This and the follow-up
6 QSAR and Machine Learning Predictors 133
Fig. 6.2 Descriptors and ngerprints. (a) Representation of the molecular descriptors classication by the information described by the class. (b) Schematic representation of a molecular ngerprint, where the molecule is compared to specic fragments, and this information is combined in a binary vector, called a ngerprint
work [17] described a pipeline to process the information to properly perform a data curation.
Later, a mathematical representation of the molecular structure should be done as a descriptor, a ngerprint, physicochemical properties, a graph, etc. A descriptor is a numerical representation of a molecule using predened rules and it can be obtained from the molecular formula, bidimensional, or three-dimensional structure [18, 19]. A way to classify them is related to the information obtained by using the descriptor such as geometrical descriptors that reect the 3D structure of the molecule [ 19] (Fig. 6.2a). Similar to a descriptor, a ngerprint is also a representation of the molecule based on specic rules, but in a vector, where each position (bin) carries a different information [20] (Fig. 6.2b).
The descriptors can be obtained from the molecular formula, bidimensional, or three-dimensional structure. Usually, the descriptors classify the QSAR model
134 P. O. Fernandes and V. G. Maltarollo
Table 6.1 Classication of the QSAR models according to their dimensionality of the molecular descriptor
Class Examples of used descriptors 1D-QSAR Descriptors based on global molecular properties such as pk
weight, and others
2D-QSAR Descriptors based on structural patterns such as connectivity indices, 2D
pharmacophores, and molecular fragments 3D-QSAR Descriptors based on noncovalent interaction elds around 3D molecular models 4D-QSAR Descriptors that include multiple ligand conformations to the 3D QSAR model,
where the conformations can be generated by molecular dynamics 5D-QSAR Explicitly includes several induced-t models to the 4D QSAR model 6D-QSAR Adds the solvation model to the 5D QSAR model
, LogP, molecular
a
according to their dimensionality [21, 22] (Table 6.1). For example, a 2D-QSAR model is built using descriptors calculated from the bidimensional structure of the molecules. A 3D-QSAR model uses the three-dimensional representation of a molecule, while higher-dimensional QSAR (4D, 5D, 6D, and 7D) includes multiple conformations and other structural information [23].
Once represented, the set of the endpoint vector ( y) and descriptors matrix (X) representing the compounds (samples) are used to build a model, but rst, they must be split into training test sets. The training set compounds are used to teach the algorithm what descriptors are important to predict a given endpoint, in other words, to tune the relationship between X and y variables; and the test set is used to validate the model built [24] (more about this process will be discussed in further sections).
QSAR modeling already has well-dened protocols and procedures to perform the application of the techniques, especially with the growing collection of biolog­ical data available in databa ses such as ChEMBL [25] and PubChem [26]. The work titled Best practices for QSAR model development, validation and exploration developed by Alexander Tropsha [24] is an important starting point to adequately build a QSAR model. Furthermore, the Organisation for Economic Co-operation and Development (OECD) [27] principles also provide valuable checkpoints that must be followed by the QSAR model for regulatory purposes.
From the Hansch [ 3 , 4] and Fujita [5] works, the use of linear models was well adopted to build the relationships. One application example from that time is the work of Di Paolo [28] who had built QSAR models to correlate the anesthetic activity of ether derivatives. Another contemporaneous example is the one published by Hansch and Klein [29] where they explore enzyme ligand interactions applying QSAR techniques.
Over time, several QSAR methodologies have been developed and some of them are worth mentioning. The rst one is the one based on the Molecular Interaction Fields (MIF) Comparative Molecular Field Analysis (CoMFA) [30, 31], Compara­tive Molecular Similarity Indices Analysis (CoMSIA) [32], and hologram quantita­tive structure-activity relationship (HQSAR) [33, 34]. These three QSAR
a
the
calculation
Virtual probe to p
o
a
cut
-of
f
val
Coulomb potential
(
)
Co
(opposite charges)
Len
nar
d-J
ones
pot
l
6 QSAR and Machine Learning Predictors 135
)
ue
entia
identical charges
ulomb potential
f the potenti
ls in the grid points
b)
Fragment
O
Generation
Biologica Activity x
36 2051081
Molecular Hologram
Fig. 6.3 Different QSAR methods: (a) schematic example of the grid-box used to calculate the MIFs, a crucial step for performing CoMFA and CoMSIA QSAR. (b) Pipeline of the HQSAR process. A molecular structure is fragmented, and these portions are combin ed into an array. Then, it is correlated to the biological activity using a PLS model
O
PLS
HQSAR
Model
methodologies were successfully used in the eld and they became a product implemented in the SYBYL software, commercialized by Tripos Inc.
The CoMFA QSAR is based on the correlation of the steric and electrostatic elds around the molecule and its biological activity. These two elds are calculated using a virtual probe atom (usually a sp
3C+
atom) to model the stereochemical interactions by the Len ard-Jones potential and to model the electrostatic interactions by the Coulombic potential in specic x, y, z coordinates points into a lattice intersection (also known as the grid box) (Fig. 6.3a). These two sterical and electrostatic MIFs in the absence of a molecular target knowledge could be interpreted as the negative of ligand-receptor representation, giving insights for molecular optimization aiming to increase the potency of novel compounds. Due to the three-dimensional require­ments and the need to analyze all compounds in the same position, molecular
136 P. O. Fernandes and V. G. Maltarollo
alignment is a key step to building a CoMFA model. A PLS regression is done using the energy calculated from these elds in each point of the grid (X) and the endpoint to be modeled (usually, a biological activity, y)[30]. Similarly, the CoMSIA QSAR is based on the same principles of the CoMFA but computing similarity indices at each point of the same grid, related to steric, electrostatic, and hydrophobic poten­tials obtained by a Gaussian-type function. Also, the PLS regression is used in the CoMSIA methodology [32].
Both CoMFA and CoMSIA are dependent on the three-dimensional structure of the compounds and their alignment and, therefore are considered 3D-QSAR methods. The HQSAR was an alternative to work with bidimensional structures, using structural fragments of a set of molecules (Fig. 6.3b). This is a proprietary QSAR method implemented by Tripos in the SYBYL platform [33]. In this meth­odology the molecular structure is described as a hologram iteratively generated with a ngerprint composed of molecular fragments built from each molecule structure. Then, the PLS regression is applied to generate the QSAR model. Despite the evolution of the QSAR eld and the popularity of machine learning-based QSAR, CoMFA, CoMSIA, and HQSAR are still used in drug design campaigns [3540].
The QSAR eld is continuously growing through the development of new methods and applications. Since 2007 more than 1000 relevant articles have been published annually [41]. Machine learning methods (see Chap. 4) have signicantly advanced QSAR research by enabling the understanding of complex relationships not observed by linear models. This evolution did not just happen in the academia, but today QSAR are valuable instrument in the industry [42]. For that reason, it is important to follow the best practices such as the OECD principles, and to promote high quality in QSAR modeling, those guidelines will be discussed in the next section.

2 OECD Principles

Usually, the QSAR modeling prioritizes optimizing the models performance to accurately replicate experimental outcomes. This approach leads to the development of models with limited transparency and interpretability, commonly referred to as black boxessuch as Articial Neural Networks (ANN), algorithms inspired by the natural neural networks [43]. An extensive review of ANN-based algorithms is available in Chap. 4. Despite their utility in predictive applications, these models have not sufciently supported the generation of reliable decisions within regulatory applications [44], due to this behavior. This behavior and other misconceptions of the QSAR modeling resulted in the publication of the Guidance document on the validation of (quantitative)structure-activity relationships [(Q)SAR] modelsby the OECD [45]. In this document, ve principles to be addressed before the application of a QSAR model (Fig. 6.4) were established. This publication marked a notable advancement in the development of in silico models, emphasizing the necessity to explicitly explain the models intended purpose, construction methodology, and
1.
A defined endpoint
An unambiguous algorithm
3
y
4.
Appropriate measures of goodness-of-fit, robustness and predictivity
5
6 QSAR and Machine Learning Predictors 137
OECD Principles
. A defined domain of applicabilit
. A mechanistic interpretation, if possible
Fig. 6.4 Five principles dened by the OECD to develop and validate a QSAR model for reglementary purposes
performance metrics [44]. The following section will present in a more detailed, but not intended to be exhaustive, way.
2.1 A Dened Endpoint
The rst principle denes that a QSAR model should have a dened endpoint and ensure the clarity of the prediction. The OECD suggests how specic an endpoint should be to ensure reli ability based on the data used to build the model. It also describes examples associated with the OECD Test Guidelines such as physico­chemical properties (e.g., melting point, boiling point, and water solubility) envi­ronmental fate (e.g., biodegradation, hydrolysis, and bioaccumulation), ecological effects (e.g., acute sh toxicity, and alga toxicity), and human health effects (e.g., acute oral toxicity, acute dermal toxicity, genotoxicity, and organ toxicity).
This transparency in the data is an essential requirement for regulatory agenci es to validate the application of the QSAR model for a particular problem. This principle also highlighted the importance of the quality of the data obtained and recommends that ideally all QSAR should be developed using experimental data generated by a single experimental protocol. This recomendation tries to avoid the introduction of errors related to the interexperimental variability. For more details on experimental assays, please refer to Chap. 12. Cronin and Schultz [46] also suggested that variations in the analyst and the equipment could introduce noise to the biological data. When it is not possible, they recommend a training set repres entative of the different protocols [27].
For the endpoints that may result from diverse chemical mechanisms, QSAR models should either be developed separately for each mechanism and applied to narrowly dened classes of chemicals, or a broader QSAR relationship can be established based on shared observations across noncongeneric chemical classes with distinct mechanisms. Either approaches can be integrated, or a statistical
138 P. O. Fernandes and V. G. Maltarollo
approach capable of simultaneous global modeling across multiple mechanisms must be employed to expand the domain of application on a broader scale.
2.2 An Unambiguous Algorithm
A QSAR model denes a mathematical relationship between the chemical structures and the modeled endpoint, the way this relationship is built is the algorithm of the model. For example, it can be a mathematical model, such as a linear regression or PLS, or a set of knowledge-based rules, such as a decision tree. In this sense, an unambiguous algorithm is capable of describing how the value was estimated and can be reproduced if desired.
The OECD document presents some algorithms that can perform regression, classication, or clustering (discussed in Chap. ). The ability to understand how the estimation is performed contributes to the transparency of the model. Another important aspect is the distinction between the transparency of the algorithm and the ability to interpret it. For example, a linear regression can be transparent in terms of the variablescoefcients and how they are used to perform a prediction, however the variables themselves may not have an understandable physicochemical meaning based on the endpoint modeled. Two-dimensional autocorrelations and 3D Morse descriptors [47] are one example of this problem. These descriptors are widely used to build QSAR models and, despite the authors trying to give them a physical meaning, it is very shallow and hard to understand how they affect the biological activity.
Machine learning methods are more complex in comparison with traditional regression methods (such as linear regression) due to the number of mathematical operations and complex functions applied in the original descriptor values making it hard to perform a straightforward interpretation. Due to this machine learning researchers are working to develop strategies to easily interpret the variables of these so-called black-box models. Recently, Rodríguez-Pérez and Bajorath [48] described in a perspective article methodologies for better understanding the machine learning models and their individual predictions, as well the current chal­lenges for the integration of these techniques in medicinal chemistry in this eld.
2.3 A Dened Domain of Applicabilit
This principle relies on the necessity to establish the scope and limitations of a model, and it is based on the structure and physicochemical properties of the training set. It is expected that a model can give reliable predictions for compounds similar to those used to build the model and predictions outside this boundary are less likely to be reliable because they are extrapolations of the model.
6 QSAR and Machine Learning Predictors 139
Due to the nature of the applicability domain (AD), each model has a specic one based on the data used to train it. Also, the ou tcome of an AD can be categorical (e.g., yes or no) or quantitative, determining the degree of similarity between the predicted compound and the training data. There are different methods of similarity to determine if a compound falls within the AD of a model. The OECD document suggests some strategies to assess the AD, but the simplest one is observing the range of the descriptors used in the model. If the predicted compound is between the minimum and maximum values observed for each descriptor in the training set, it is inside the domain. For the model displaying few descriptors, it is easy to observe; nevertheless, it can be more compl icated in complex models.
2.4 Appropriate Measures of Goodness-of-Fit, Robustness,
and Predictivity
This OECD principle advocates for the necessity of statistical validation to ensure the predictability of the model. The OECD document divides the assessment of model performance (or statistical validation) into two stages: the internal perfor­mance (goodness-of-t and robustness) and the external performance (predictivity). The combination of these two methods is used to analyze the model, avoiding being overtted and undertted. The target is a model not so simple that lacks information and is not so complex to model the noise in the data [49, 50].
The statistical validation enables comparison between different models in order to select the most predictive model. Another function of this process is to avoid spuriousmodels based on correlations by chance, which are not meaningful and not predictive [5153]. The inte rnal validation is a measure of the models perfor­mance using only the training set samples, which is the data used to build the model. In this sense, models that could not predict the data used in the training process could be discarded in the early stages of QSAR modeling. External validation is a way to measure the performance using compounds unseen by the model (the section validation and controls will further discuss how statistical metrics can be applied to perform these validations) and, therefore, represents an application of QSAR in actual drug discovery pipeline mimicking the prediction of compounds in commer­cial libraries. For that reason, it is important to assess the applicability domain to ensure a test set representative of the training set.
2.5 A Mechanistic Interpretation, if Possible
Historically, the rst QSAR models were developed using congeneric compounds (molecules from the same chemical class with minor variations in substituents), and the activity was related to biological activities produced by the compounds
140 P. O. Fernandes and V. G. Maltarollo
molecular structure. The statistical methods were applied to describe the relation­ships and support chemical knowledge, beginning the idea of mechanistic interpre­tation. This OECD principle is not mandatory but is desirable, once a QSAR model is consistent if the knowledge of the chemical/toxicological process provides more credibility and acceptance of the predictions [27]. Furthermore, the interpret ation of a model allows an extra layerof validation since the proposed mechanism could be checked according to its chemical meaning and if it is expected or not. For example, it is expected that hydrogen bond descriptors positively inuence the aqueous solubility of a QSPR model [54]; otherwise, there is a high possibility that the model is not correct .
The if possibleclause added in the principle is due to the interactive nature of the modeling process. Usually, data exploration and modeling lead to the generation and testing of a hypothesis, and a useful QSAR model may lack mechanistic interpretation due to different reasons. Nevertheless, this principle motivates the modeler to seek a mechanistic interpretation that can contribute to the understanding of the statistic validation [27].

3 Software and Tools

Unless you are performing your own assays, the rst step to performing QSAR modeling is retrieving information from the literature. This process can be made manually by compiling published works relevant to the project or using available databases. The second approach is very common today since there are several initiatives available. One popular example is the ChEMBL database [55], a large­scale database comprising several bioactivity information. The compoundschem­ical structure and detailed information about the assays are provided, comprising binding measurements, functional assays, pharmacokinetic properties, and toxicity. Besides ChEMBL there are other large databases such as PubChem [56] and BindingDB [57]. A detailed explanation of available chemical information sources is described in Chap. 2.
The next importan t step in the development of a QSAR model is the representa­tion of the molecules. To obtain these representations, the molecular descriptors and/or ngerprints are calculated using specic software and web servers. A nonexhaustive list of available tools to perform descriptors/ngerprint calculations is described in Table 6.2.
Another emerging way to represent molecules is using learned embeddings such as convolutional and graph encoding to represent the molecular structures. This approach is focused on models capable of learning generalizable representations from a molecular data set. Convolutional encodings perform mathematical opera­tions compacting the space containing the variables. One example is the work from Coley et al. [74] where atom and bond attributes were used in a convolutional neural network to represent the molecules. The authors applied the methodology to predict aqueous solubility, octanol solubility, melting point, and toxicity. Another
6 QSAR and Machine Learning Predictors 141
Table 6.2 Software and web servers available for descriptors and ngerprint calculation
Name Desc. FP Source Dragon [58] 5270 1 https://www.talete.mi.it/ CDK [59] 275 9 https://cdk.github.io/ E-dragon [60] 1600 https://vcclab.org/lab/edragon/ Mold2 [61] 779 https://www.fda.gov/science-research/bioinformatics-tools/
Pybel [62]244https://github.com/pybel/pybel PaDEL [63] 1875 12 https://www.yapcwsoft.com/dd/padeldescriptor/ RDKit 198 8 https://www.rdkit.org/ PyDPI [64] 615 7 https://pypi.org/project/pydpi/ Chemopy [65] 1135 7 https://github.com/ifyoungnet/Chemopy ChemDes [66] 3679 59 http://www.scbdd.com/chemdes/ Rcpi [67] 308 10 https://bioconductor.org/packages/release/bioc/html/Rcpi.html BioTriangle
[68] ChemSAR [69] 783 10 http://chemsar.scbdd.com/ Mordred [70] 1825 https://github.com/mordred-descriptor/mordred PyBioMed [71] 775 19 https://github.com/gadsbyy/PyBioMed alvaDesc [72] 566 3 https://www.alvascience.com/alvadesc/ BioMedR [73] 293 13 https://github.com/wind22zhu/BioMedR
Desc. number of calculated descriptors; FP number of calculated ngerprint sets
540 7 http://biotriangle.scbdd.com/
mold2
Fig. 6.5 Simplied example of a graph representation of a molecule
node
Node properties:
atom type chirality hybridization aromaticity
edge
Edge properties:
bond type ring stereochemistry
application of a convolutional encoding is the use of 3D information to predict protein-ligand binding afnity [75]. In the graph approach, atoms are represented as nodes and the bonds are represented by edges; both nodes and edges have associated features such as the atomic number, formal charge, and bond type (Fig. 6.5). It produces a 2D data structure; nevertheless, 3D information such as chirality or stereochemistry can be included in nodes or edges [76]. One example of graph usage is the Chemprop [77], a machine-learned package that has been used to build models to predict biological activities [78, 79], molecular properties [80, 81], and other chemical endpoints [82]. The work from McGibbon et al. [18] titled From intuition to AI: evolution of small molecule representations in drug discovery