Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5849_Библиотеки_им_академика_М_И_Перельмана
.pdf
2.4.6 Novartis
Novartis applies computer-aided methods to support the lead discovery process [300].
The in silico group is integrated into the global lead-finding department and is used to
work in parallel with the HTS process with the aim of improving the quality of the
output. In some cases, in silico processes can successfully replace HTS experiments.
Novartis procedures focus on designing screening libraries with greater ring diversity,
low molecular weight (<700), low hydrophobicity (log P < 7.5), and good permeability (PSA < 200 A˚). The Lipinski RO5 seems to be further applied during lead
optimization. Natural products are not excluded since such compounds can act as
excellent molecular probes or inspire chemical synthesis.
2.4.7 Schering AG
At Schering AG, a dedicated hit-to-lead team exists and uses chemoinformatics to
assist the lead generation step [301]. Besides improving the lead process by evaluating
and implementing software tools for property predictions, the team provides the
followingguidelinesto design drug-likelibraries: suitable molecular properties such as
MW between 200 and 500, log P/log D between 1 and 5, H-bond donors between 0
and 5, H-bond acceptors <10, favorable pharmacodynamics and kinetics (in vivo,
in vitro, and in silico) properties like permeability in Caco2 cells >100 cm s
1
107,
rat plasma clearance <50 mL min
1kg1
, rat oral bioavailability >25%, toxicity
assessed by the commercial DEREK package, good chemical optimization potential,
and patentability.
2.4.8 Vertex Pharmaceuticals
Chemoinformatics approaches are used at Vertex and the company has published
several methods. Drugs are distinguished from nondrugs by using machine learning
methods [111] such as decision trees that help the hit-to-lead decision process. The
REOS (rapid elimination of swill) program to filter out molecules that might be
problematic was developed [24, 108, 110]. REOS combines a set of SMARTS-based
chemical functional group filters and a set of RO5-like physicochemical property
filter.The default values for the property filter are MW 200–500, log P (from 5 to 5),
HBD 0 to 5, HBA 0 to 10, formal charge between 2 and 2, rotatable bonds 0–8, and
15–20 heavy atoms.
2.5 CHALLENGING ADME/Tox PREDICTIONS
Several drugs have been withdrawn from the market in past few years such as
Astemizole or Glafenine for hERG blocking, or Mibefradil responsible of hepatoxicity via the CYP450 enzymatic system [18, 302]. Below is an illustration for the utility
of ADME/tox prediction for the investigations of a compound and some assistance
to decision-making.
88 IN SILICO ADME/Tox PREDICTIONS
https://t.me/medicina_free

2.5.1 Tolcapone
Tolcapone (tasmar, Figure 2.26) is an inhibitor of catechol-O-methyltransferase
(COMT) and can be used as anti-Parkinsonian agent. This molecule has the ability
to cross the blood–brain barrier and exerts its COMT inhibitory effects in the CNS as
well as in the periphery. During the clinical trials, the molecule showed acute
hepatotoxicity with three fatalities. Due to significant hepatotoxicity, the drug’s
therapeutic utility is limited to a drug of last resort. Tolcapone possesses one nitro
group, which has been reported to be hepatotoxic and hepatocarcinogen [303].
Nitroaromatics can be reduced to form reactive, nitroanion radical, nitroso intermediate, and N-hydroxy derivatives [304]. These reactive metabolites are usually not
desired in drug discovery projects and molecules containing nitroaromatic groups are
in general removed from a compound collection [198].
2.5.2 Factor V Inhibitors
Inhibitors of blood coagulation can be serine protease inhibitors, molecules
blocking the catalytic site of coagulation enzymes such as factor Xa or thrombin.
A prospective study was undertaken by Segers et al. [305] in order to find potential
inhibitors of the coagulation factor Va–membrane interaction. The prothrombinase
complex, the macromolecular system that generates thrombin, is essential to the
process. This mechanism requires transient interactions with the cell membrane. To
discover small molecules that could impede factor Va–membrane interaction, a
structure-based virtual ligand screening study was performed. More than 300,000
molecules were docked into factor Va C2 domain and seven hits were identified. The
molecules were shown to bind directly to the right factor Va domain using surface
plasmon resonance (SPR). In vitro experiments indicated that those compounds do
not inhibit FVa procoagulant activity in a plasma-based assay system. Further
experiments confirmed by SPR analysis showed that as albumin concentration
increased, the inhibitory activity of the compound gradually decreased. Retrospectively, by using the commercial software q-Mol, marketed by Quantum Pharmaceuticals and capable to predict human serum albumin binding for the most potent
compound was confirmed (Figure 2.27).
Figure 2.26 Tolcapone. Molecular weight ¼ 273.24; X Log P3 ¼ 3.3; tPSA ¼ 97.51; H-bond
donors ¼ 2; H-bond acceptors ¼ 5; rotatable bond ¼ 3; rigid bond ¼ 14.
CHALLENGING ADME/Tox PREDICTIONS 89
https://t.me/medicina_free

2.5.3 CRF-1 Receptor Antagonists
In2003, at Neurogen Corporation,Hodgetts et al. described the development,synthesis,
and structure–activitystudy for discoveringa novel series of 2-arylpyrimidin-4-onesas
CRF-1 receptor antagonists [306] (Figure 2.28). The group showed that chemical
optimizationof the compounds could lead to potent CRF-1 antagonists and commented
about the interest of using early ADME/Toxpredictions.ADME/Toxpropertiessuch as
log P and polar surface area were calculated to design CNS agents with appropriate
physicalproperties. Compounds were optimizedwith physical chemistry valueswithin
the range that lead to a reasonable membrane permeability and brain penetration.
2.6 STATISTICAL METHODS
2.6.1 Principal Component Analysis
2.6.1.1 Aim The goal of PCA [90] is to project the data into a subspace made of
linear combinations of the original descriptors so that this subspace is the best-
Figure 2.27 Coagulation factor V–membrane interaction inhibitor. MW ¼ 490.48; X Log
P3 ¼ 4.19; tPSA ¼ 105.17; H-bond donors ¼ 2; H-bond acceptors ¼ 6; rotatable bond ¼ 9;
rigid bond ¼ 27.
Figure 2.28 CRF-1 receptor antagonist. MW ¼ 376.54; X Log P3 ¼ 4.32; tPSA ¼ 48.05; H-
bond donors ¼ 1; H-bond acceptors ¼ 2; rotatable bond ¼ 8; rigid bond ¼ 13.
90
IN SILICO ADME/Tox PREDICTIONS
https://t.me/medicina_free

simplified image, in a small dimension, of the original data in terms of variation (i.e., a
projection that best represents the data). Data may then be explored in a small space
spanning the most informative view (according to data variance) of the original
features. Visually, according to the final chosen dimensions, successive 2D or 3D
graphics can be obtained. In this space, the new axes (called components) are
orthogonal to each other. This space can be used to globally study for possible
outliers, clusters of individuals, and coverage between descriptor spaces of several
groups of observations (Figure 2.29). One goal would be to use fewer principal
components for subsequent data analysis than in the original high-dimensional space.
In a second step and if required (in clustering studies, for example), the axes of the
new space (called components) can be interpreted in terms of the original data
according to the correlations between these axes and the original descriptors.
Examples of ADME/Tox Applications PCA can b e used in several ways in ADME/
Tox studies; PCA can be used to study redundancy between descriptors because there
is a connection between PCA components and original descriptors [307]. As
components are orthogonal, if two components are correlated to different descriptors,
those descriptors are unlikely to provide the same information. If two descriptors are
highly correlated to the same component, their correlation has to be studied.
Orthogonality of components as well as dimension decrease allows PCA to become
a good preliminary step to classification or regression methods requiring uncorrelated
features as input variables (logistic regression for example). In the study of Benigni
and Bossa, PCA [308] was applied to obtain 160 principal components (explaining
Figure 2.29 Illustration of chemical space coverage between learning compounds (in green)
and test compounds (in red). See insert for color representation of this figure.
STATISTICAL METHODS 91
https://t.me/medicina_free

94% of the total variance), which were used in stepwise linear discriminant analysis
aiming at predicting chemical carcinogenicity of different compounds. As PCA can
provide a low-dimensional image of quite complex data, the first components can be
used to study the global space of studied data, which can be linked to the applicability
domain. Kortagere et al. [309] predicted BBB using support vector machine models.
As several datasets were available (learning and test sets), applying PCA on one
sample and projecting the second one in the same space allowed to study the coverage
between the chemical space of the training set and the space spanned by the test
compounds. As the first two components accounted for 79% of variability, a
2D representation provided a good qualitative way to graphically compare those
spaces. If the test observations were outside the chemical space of the learning
material, any prediction is quite unreliable.
Usage Warning The most important point to consider before any use of PCA
results is the amount of variability provided by the successive components. This is
particularly important when reducing the space dimension. Using all the components
provided by PCA ensures that all the information is not loss. Usually 80–90% of the
components are considered as a good representativeness of the global information. In
the case of complex data with numerous descriptors and depending on the objective,
a smaller amount of global variability can be considered.
In order to obtain meaningful principal components, if descriptors have different
variances, scaling the data is required. However, careful stud y of descriptors is
necessary, for example to remove noise variables before giving them the same
importance as important variables.
Finally, PCA remains a representation in the space of descriptors. The new
components are only linear combinations of the original variables, no new information is provided.
2.6.1.2 Technical Description
Input Data This method is based on the description of n individuals by p descriptors
collected in a matrix having n rows and p columns.
Output New coordinates for each individual in the new subspace are provided. The
coordinates are collected into a matrix having k columns (where k is the dimension
of the new subspace) and called scores. PCA also produces the coefficients, called
loadings or components, applied to the original descriptors to obtain the successive
new descriptors spanning the new subspace.
Description Principal components are linear combinations of the original descriptors obtained so that the first principal component is the linear combination best
explaining the data total variance. The second principal com ponent is orthogonal and
describes the remaining variance. It aims at maximizing the amount of variance
explained by the second linear combination.
These components are obtained by diagonalization of the matrix giving the
variances of the original descriptors (in diagonal) and the covarianc es between
92 IN SILICO ADME/Tox PREDICTIONS
https://t.me/medicina_free

them. The diagonalization of this matrix provides both eigenvectors and eigenvalues. These eigenvalues represent the empirical variances of the successive components. Once sorted in decreasing order, the highest eigenvalue represents the
variance associated with the first component. The corresponding successive eigenvectors are the loadings or coefficients applied to the original variables in the
successive linear combinations.
The final number of components, k, is chosen so that a reasonable amount of
variation is explained on the k axes (80–90%). The decomposition of the total
variance is a breaking point in the histogram of eigenvalues and can be an indication
of where to stop. If graphical representation is the main objective, two or three
components can be considered if enough variation is accounted.
2.6.1.3 Possible Extensions PCA can be applied in a nonlinear way to provide
more complex combinations between original variables. This is performed by transforming the original descriptors (more or less explicitly) to nonlinear transformations
(e.g., product of descriptors) and by applying PCA on the new obtained variables.
Kernel-PCA [310] uses a kernel to implicitly project original descriptors in a higher
dimension space (just as in support vector machine) before applying classical PCA.
2.6.2 Partial Least Square
2.6.2.1 Aim Partial least square regression [311] is a technique allowing the
prediction of one or more response variables generally accounting for numerous
descriptors. It is adapted when the dataset contains more descriptors than observations
and multicollinearity between descriptors. It is one of the rare methods that allow for
the prediction of more than one variable at the same time. PLS generalizes and
combines features from both PCA and multiple regression. Linear combinations of
descriptors are used to build new factors (identical to PCA) called latent variables. The
objective is not to maximize variance but to obtain the best predictive power.PLS finds
the linear combinations of predictors having the maximum covariance with the
response variables. The obtained decomposition of the descriptors is used to predict
the response values. Similar to PCA, the number of components can be chosen to
explain enough variation of descriptors and responses while avoiding overfitting. This
choice can be performed through cross-validation or use of a test dataset.
Both a prediction of the response values and graphical interpretation tools are
provided. It is possible to study correlations between the original variables (descriptors and responses) and the latent ones. Figure 2.30 represents this type of correlation
for descriptors (X1, ..., X5) and responses (Y1, Y2). The first latent variable (PLS1)
is mainly predictive of Y1 values, which are highly related to X2 and X3 (more
specifically Y1 tends to increase with X2 and when X3 decreases). The second latent
variable (PLS2) mostly accounts for Y2 prediction, which is essentially due to X1
and X5. Y2 tends to increase when X1 and X5 decrease. Descriptor X2 provides
information on Y2. This theoretical example shows the interpretation possibilities
offered by PLS. Observations can be plotted in the new space to discover clusters and
outliers.
STATISTICAL METHODS 93
https://t.me/medicina_free

Examples of ADME/Tox Applications PLS is mostlyused to predictthe values of one
continuous response variable.The interpretation of latent variables helps to understand
which predictors are most involved in the prediction accuracy. In Chohan et al. [312],
PLSisusedamongotherregressionmethodstopredictCytochromeP450 1A2 inhibition
(pIC
50
) from in-house computed descriptors accounting for topological, geometrical,
and electronic features of molecules. Feature selection (according to variance, redundancy, and predictivity)wasperformed before PLS application.Thefinal model isbased
on two components (chosen by minimizing the cross-validation error rate). Interpretation of correlations between the latent variables and original variables (pIC
50
and the
descriptors) suggested that lipophilicity and aromaticity were the most important
features to describe CYP1A2 inhibition. A more precise link can be derived showing
that decreasing these two features should decrease inhibition.
In Luco [313], PLS is used to predict brain–blood distribution through computation
and investigation of 25 structural descriptors. The variable importance for the
projection (VIP) has been used to select the most relevant descriptors while strong
outliers were removed. Three components (accounting for 85% of the variance in
log BB) were retained (threefold cross-validation). A validation set was used with a
careful verification of the applicability domain. The BBB and aqueous solubility in
Obrezanova [314] was also modeled through PLS.
PLS is used in ADME/Tox context to perform discrimination. In Kriegl et al. [315],
a classification model was trained on 2D, 3D, and quantum-mechanical descriptors
to predict two (low versus high) or three (low, medium, high) classes of human
cytochrome P450 3A4 inhibitors. PLS was shown to be less efficient than real
discrimination methods.
Usage Warning The amount of variability provided by the successive latent
variables has to be taken into consideration. The user should choose fewer components to obtain a model that could be generalized to new data for predictions. Using a
Figure 2.30 Example of simultaneous representation of correlation between the latent
variables and original variables (descriptors are X
1–X5
and responses are Y1and Y2).
94
IN SILICO ADME/Tox PREDICTIONS
https://t.me/medicina_free

few latent variables provides graphical representations of the chosen ones. Interpretation of coefficients is not possible as descriptors are generally not independent.
PLS is often used to perform classification. It is simply obtained by providing the
group labels as response variable. As PLS is a regression method designed to predict
continuous variables, the qualitative side of group labels will not be taken into
account. This can lead to negative values. Thresholds have to be defined to obtain the
predicted groups. Group labels have to be coded into a dummy matrix containing as
many columns as groups and filled with zeros and ones.
Although PLS can handle more variables than observations, more predictive
models are obtained when irrelevant descriptors have been removed. Before applying
PLS, careful preproce ssing of data is needed.
2.6.2.2 Technical Description
Input Data This method is based on the description of n individuals by p descriptors
collected in a matrix having n rows and p columns, called X, and of the values of
q response values measured on the same n individuals and collected in Y.
Output The prediction of Y is provided. New coordinates for each individual in the
new subspace are given and are collected into a matrix having k columns (where k is
the dimension of the new subspace) and called latent variables. PLS produces the
coefficients applied to X and Y to build these latent variables.
Description PLS is a stepwise algorithm where at each step, one latent variable is
provided. At the first step, one linear combination is defined for X and another for Y so
that the covariance between them is maximum. This result is obtained via an iterative
algorithm. The obtained new feature is called the first latent variable. It is possible
to compute the proportion of variance of X and Y explained by the first latent
variable. Information by the first latent variable is subtracted from both X and Y (or
equivalently only from X) and new linear combinations are obtained in the same way,
operating on the deflated matrices. The same process is repeated until the covariance
is null.
PLS provides both graphical representation of the data and prediction of the
response values. Graphical representations allow for both observations and predictors
to be viewed on the same graph. The relative position of observations with regard to
the different predictors is readily explained. Correlations between response and latent
variables can be represented.
The choice of the number of latent variables to be used can be obtained through
cross-validation or use of a test dataset. The objective is to obtain a precise prediction
while avoiding overfitting (difficulties in generalizing to new observations). The
optimal number of latent variables is the one resulting in the best prediction error.
Possible Extensions Most extensions of PLS use nonlinear regressions to introduce
more complex combinations of predictors (see ASPLS [316] using spline functions to
perform regression). Multiblock PLS where blocks of variables (or observations) are
applied to the same operations during the latent variables construction is also possible.
STATISTICAL METHODS 95
https://t.me/medicina_free

2.6.3 Support Vector Machine
2.6.3.1 Aim The support vector machine classifier [98] was originally designed
to perform discrimination between two groups (Figure 2.31). It has been extended to
classify more than two groups and to perform regression (i.e., modeling one
continuous variable). SVM is supported by two main ideas. A transformation
performed by means of a function called kernel. The kernel projects the observations
into a new space (generally of higher dimension than the original one) where the two
groups are likely to be (or almost be) linearly separable. Second, building the best
linear rule (a hyperplane) separating the two groups in the new space. In SVM,
according to the Vapnik–Chervonenkis theory, this hyperplane is chosen so that the
distance between it and the nearest observations on either side is maximized. It is
called the maximum-margin hyperplane. Observations located on the margins are
called support vectors.
2.6.3.2 Examples of ADME/Tox Applications Vapnik introduced SVM in 1995
with emergence to ADME/Tox work in the early 2000s. SVM can be used when
several predictors (usually quantitative but there are some kernels to deal with
qualitative ones) are used to predict one other variable. The predictive variable can
be qualitative or quantitative.
For qualitative data, the objective is to predict membership of groups. A recent
application of SVM is described by Hou et al. [317]. The goal was to predict HIA by
defining two levels of absorption according to a fractional absorption threshold value.
The prediction issue was converted into a binary classification problem that particularly fits SVM use. One SVM model was built with each of the 10 chosen descriptors
as input. This allowed to investigate individual performances of the descriptors and to
rank them according to their predictive power (evaluated by a validation process
recursively splitting the sample into learning and test ones). The usefulness of
descriptor combinations was addressed. Systematic search was performed to find
the best combination. In this work, SVM parameters (kernels, error penalty, etc.) were
very carefully chosen by a validation process. A very good prediction rate was
achieved for both classes (100% for the poor-absorption class and 97.8% for the goodabsorption class).
Figure 2.31 Illustration of SVM principle. Starting from the left, observations are illustrated in
their original space, and then are projected into the new space and finally separated by the
maximum-margin hyperplane (circled points are called support vectors). See insert for color
representation of this figure.
96
IN SILICO ADME/Tox PREDICTIONS
https://t.me/medicina_free

For quantitative data, any continuous (or discrete with really numerous values)
variable can be chosen to be predicted. In Kortagere et al. [309], an SVM model was
applied with shape signatures descriptors to predict BBB.
Usage Warning The main issue with SVM is to choose the different parameters. The
kernel obviously is the first one. Despite some known recommendations, it is not an
easy task and depends on the goals. When using new data, several kernels should be
tried or an original one can be designed. The problem is exactly the same for kernel
parameters even if some advices have been posted. A cross-validation choice has to be
performed most of the time. The same work has to be done for the error penalty
parameter that controls balance between overfitting (very precise modeling of the
learning sample) and generalization error (ability to generalize the model to new
observations). By setting an adequate value, it is always possible to make a perfect
classification of the learning dataset. This is likely due to overtraining and no
satisfactory results will be obtained on a test set. Any parameter choice has to be
done through cross-validation. Eventually, obtaining a precise and robust model
requires several choices that involve advanced analyses. As the transformed space is
never explicitly known, it is not possible to have any interpretation of the original
descriptors results. As numerous descriptors are often available, knowledge of the
descriptors that account for the explanation of the predicted variable would be of
interest. In the case of a “black box” method such as SVM, prior work on feature
selection is required for any interpretation. To conclude, SVM is a powerful but not
straightforward method allowing to model quite complex relationships.
2.6.3.3 Technical Description
Input Data SVM requires two objects, the description of n individuals by
p descriptors and the vector of length n containing the data to be predicted (either
group numbers or any continuous variable).
Output The modeling of the vector to be predicted contains either the predicted
groups or the predicted values of the target variable.
Description This description will focus on the original SVM issue, the discrimination of two groups.
The first step is the kernel transformation. The goal is to project data into a higher
dimension space where linearly dividing the data into two groups is easier. The kernel
function allows combining original features leading to a new space describing the
data in a higher dimension. According to the chosen transformation, the dimension
can even be infinite. In the new space, computations may become difficult or even
impossible. The so-called kernel trick indicates that if the kernel fulfils some
properties, distances in the new space can be handled without computing the
coordinates of its observations. Many different kernels can be used and the most
common ones include the linear kernel (especially used in the case of large sparse
data, such as found in text categorization), the Gaussian radial basis kernel (generalpurpose kernel to be used in the case of no particular prior knowledge), the polynomial
STATISTICAL METHODS 97
https://t.me/medicina_free
Соседние файлы в папке Библиотека им академика М.И. Перельмана
