Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5849_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
30.08.2026
Размер:
49 Мб
Скачать
2.4.6 Novartis
Novartis applies computer-aided methods to support the lead discovery process [300]. The in silico group is integrated into the global lead-finding department and is used to work in parallel with the HTS process with the aim of improving the quality of the output. In some cases, in silico processes can successfully replace HTS experiments. Novartis procedures focus on designing screening libraries with greater ring diversity, low molecular weight (<700), low hydrophobicity (log P < 7.5), and good perme­ability (PSA < 200 A˚). The Lipinski RO5 seems to be further applied during lead optimization. Natural products are not excluded since such compounds can act as excellent molecular probes or inspire chemical synthesis.
2.4.7 Schering AG
At Schering AG, a dedicated hit-to-lead team exists and uses chemoinformatics to assist the lead generation step [301]. Besides improving the lead process by evaluating and implementing software tools for property predictions, the team provides the followingguidelinesto design drug-likelibraries: suitable molecular properties such as MW between 200 and 500, log P/log D between 1 and 5, H-bond donors between 0 and 5, H-bond acceptors <10, favorable pharmacodynamics and kinetics (in vivo, in vitro, and in silico) properties like permeability in Caco2 cells >100 cm s
1
107,
rat plasma clearance <50 mL min
1kg1
, rat oral bioavailability >25%, toxicity
assessed by the commercial DEREK package, good chemical optimization potential, and patentability.
2.4.8 Vertex Pharmaceuticals
Chemoinformatics approaches are used at Vertex and the company has published several methods. Drugs are distinguished from nondrugs by using machine learning methods [111] such as decision trees that help the hit-to-lead decision process. The REOS (rapid elimination of swill) program to filter out molecules that might be problematic was developed [24, 108, 110]. REOS combines a set of SMARTS-based chemical functional group filters and a set of RO5-like physicochemical property filter.The default values for the property filter are MW 200–500, log P (from 5 to 5), HBD 0 to 5, HBA 0 to 10, formal charge between 2 and 2, rotatable bonds 0–8, and 15–20 heavy atoms.
2.5 CHALLENGING ADME/Tox PREDICTIONS
Several drugs have been withdrawn from the market in past few years such as Astemizole or Glafenine for hERG blocking, or Mibefradil responsible of hepatoxi­city via the CYP450 enzymatic system [18, 302]. Below is an illustration for the utility of ADME/tox prediction for the investigations of a compound and some assistance to decision-making.
88 IN SILICO ADME/Tox PREDICTIONS
https://t.me/medicina_free
2.5.1 Tolcapone
Tolcapone (tasmar, Figure 2.26) is an inhibitor of catechol-O-methyltransferase (COMT) and can be used as anti-Parkinsonian agent. This molecule has the ability to cross the blood–brain barrier and exerts its COMT inhibitory effects in the CNS as well as in the periphery. During the clinical trials, the molecule showed acute hepatotoxicity with three fatalities. Due to significant hepatotoxicity, the drug’s therapeutic utility is limited to a drug of last resort. Tolcapone possesses one nitro group, which has been reported to be hepatotoxic and hepatocarcinogen [303]. Nitroaromatics can be reduced to form reactive, nitroanion radical, nitroso interme­diate, and N-hydroxy derivatives [304]. These reactive metabolites are usually not desired in drug discovery projects and molecules containing nitroaromatic groups are in general removed from a compound collection [198].
2.5.2 Factor V Inhibitors
Inhibitors of blood coagulation can be serine protease inhibitors, molecules blocking the catalytic site of coagulation enzymes such as factor Xa or thrombin. A prospective study was undertaken by Segers et al. [305] in order to find potential inhibitors of the coagulation factor Va–membrane interaction. The prothrombinase complex, the macromolecular system that generates thrombin, is essential to the process. This mechanism requires transient interactions with the cell membrane. To discover small molecules that could impede factor Va–membrane interaction, a structure-based virtual ligand screening study was performed. More than 300,000 molecules were docked into factor Va C2 domain and seven hits were identified. The molecules were shown to bind directly to the right factor Va domain using surface plasmon resonance (SPR). In vitro experiments indicated that those compounds do not inhibit FVa procoagulant activity in a plasma-based assay system. Further experiments confirmed by SPR analysis showed that as albumin concentration increased, the inhibitory activity of the compound gradually decreased. Retrospec­tively, by using the commercial software q-Mol, marketed by Quantum Pharma­ceuticals and capable to predict human serum albumin binding for the most potent compound was confirmed (Figure 2.27).
Figure 2.26 Tolcapone. Molecular weight ¼ 273.24; X Log P3 ¼ 3.3; tPSA ¼ 97.51; H-bond donors ¼ 2; H-bond acceptors ¼ 5; rotatable bond ¼ 3; rigid bond ¼ 14.
CHALLENGING ADME/Tox PREDICTIONS 89
https://t.me/medicina_free
2.5.3 CRF-1 Receptor Antagonists
In2003, at Neurogen Corporation,Hodgetts et al. described the development,synthesis, and structure–activitystudy for discoveringa novel series of 2-arylpyrimidin-4-onesas CRF-1 receptor antagonists [306] (Figure 2.28). The group showed that chemical optimizationof the compounds could lead to potent CRF-1 antagonists and commented about the interest of using early ADME/Toxpredictions.ADME/Toxpropertiessuch as log P and polar surface area were calculated to design CNS agents with appropriate physicalproperties. Compounds were optimizedwith physical chemistry valueswithin the range that lead to a reasonable membrane permeability and brain penetration.
2.6 STATISTICAL METHODS
2.6.1 Principal Component Analysis
2.6.1.1 Aim The goal of PCA [90] is to project the data into a subspace made of linear combinations of the original descriptors so that this subspace is the best-
Figure 2.27 Coagulation factor V–membrane interaction inhibitor. MW ¼ 490.48; X Log P3 ¼ 4.19; tPSA ¼ 105.17; H-bond donors ¼ 2; H-bond acceptors ¼ 6; rotatable bond ¼ 9;
rigid bond ¼ 27.
Figure 2.28 CRF-1 receptor antagonist. MW ¼ 376.54; X Log P3 ¼ 4.32; tPSA ¼ 48.05; H- bond donors ¼ 1; H-bond acceptors ¼ 2; rotatable bond ¼ 8; rigid bond ¼ 13.
90
IN SILICO ADME/Tox PREDICTIONS
https://t.me/medicina_free
simplified image, in a small dimension, of the original data in terms of variation (i.e., a projection that best represents the data). Data may then be explored in a small space spanning the most informative view (according to data variance) of the original features. Visually, according to the final chosen dimensions, successive 2D or 3D graphics can be obtained. In this space, the new axes (called components) are orthogonal to each other. This space can be used to globally study for possible outliers, clusters of individuals, and coverage between descriptor spaces of several groups of observations (Figure 2.29). One goal would be to use fewer principal components for subsequent data analysis than in the original high-dimensional space.
In a second step and if required (in clustering studies, for example), the axes of the new space (called components) can be interpreted in terms of the original data according to the correlations between these axes and the original descriptors.
Examples of ADME/Tox Applications PCA can b e used in several ways in ADME/ Tox studies; PCA can be used to study redundancy between descriptors because there is a connection between PCA components and original descriptors [307]. As components are orthogonal, if two components are correlated to different descriptors, those descriptors are unlikely to provide the same information. If two descriptors are highly correlated to the same component, their correlation has to be studied. Orthogonality of components as well as dimension decrease allows PCA to become a good preliminary step to classification or regression methods requiring uncorrelated features as input variables (logistic regression for example). In the study of Benigni and Bossa, PCA [308] was applied to obtain 160 principal components (explaining
Figure 2.29 Illustration of chemical space coverage between learning compounds (in green) and test compounds (in red). See insert for color representation of this figure.
STATISTICAL METHODS 91
https://t.me/medicina_free
94% of the total variance), which were used in stepwise linear discriminant analysis aiming at predicting chemical carcinogenicity of different compounds. As PCA can provide a low-dimensional image of quite complex data, the first components can be used to study the global space of studied data, which can be linked to the applicability domain. Kortagere et al. [309] predicted BBB using support vector machine models. As several datasets were available (learning and test sets), applying PCA on one sample and projecting the second one in the same space allowed to study the coverage between the chemical space of the training set and the space spanned by the test compounds. As the first two components accounted for 79% of variability, a 2D representation provided a good qualitative way to graphically compare those spaces. If the test observations were outside the chemical space of the learning material, any prediction is quite unreliable.
Usage Warning The most important point to consider before any use of PCA results is the amount of variability provided by the successive components. This is particularly important when reducing the space dimension. Using all the components provided by PCA ensures that all the information is not loss. Usually 80–90% of the components are considered as a good representativeness of the global information. In the case of complex data with numerous descriptors and depending on the objective, a smaller amount of global variability can be considered.
In order to obtain meaningful principal components, if descriptors have different variances, scaling the data is required. However, careful stud y of descriptors is necessary, for example to remove noise variables before giving them the same importance as important variables.
Finally, PCA remains a representation in the space of descriptors. The new components are only linear combinations of the original variables, no new informa­tion is provided.
2.6.1.2 Technical Description
Input Data This method is based on the description of n individuals by p descriptors collected in a matrix having n rows and p columns.
Output New coordinates for each individual in the new subspace are provided. The coordinates are collected into a matrix having k columns (where k is the dimension of the new subspace) and called scores. PCA also produces the coefficients, called loadings or components, applied to the original descriptors to obtain the successive new descriptors spanning the new subspace.
Description Principal components are linear combinations of the original descrip­tors obtained so that the first principal component is the linear combination best explaining the data total variance. The second principal com ponent is orthogonal and describes the remaining variance. It aims at maximizing the amount of variance explained by the second linear combination.
These components are obtained by diagonalization of the matrix giving the variances of the original descriptors (in diagonal) and the covarianc es between
92 IN SILICO ADME/Tox PREDICTIONS
https://t.me/medicina_free
them. The diagonalization of this matrix provides both eigenvectors and eigenva­lues. These eigenvalues represent the empirical variances of the successive com­ponents. Once sorted in decreasing order, the highest eigenvalue represents the variance associated with the first component. The corresponding successive eigen­vectors are the loadings or coefficients applied to the original variables in the successive linear combinations.
The final number of components, k, is chosen so that a reasonable amount of variation is explained on the k axes (80–90%). The decomposition of the total variance is a breaking point in the histogram of eigenvalues and can be an indication of where to stop. If graphical representation is the main objective, two or three components can be considered if enough variation is accounted.
2.6.1.3 Possible Extensions PCA can be applied in a nonlinear way to provide more complex combinations between original variables. This is performed by trans­forming the original descriptors (more or less explicitly) to nonlinear transformations (e.g., product of descriptors) and by applying PCA on the new obtained variables. Kernel-PCA [310] uses a kernel to implicitly project original descriptors in a higher dimension space (just as in support vector machine) before applying classical PCA.
2.6.2 Partial Least Square
2.6.2.1 Aim Partial least square regression [311] is a technique allowing the prediction of one or more response variables generally accounting for numerous descriptors. It is adapted when the dataset contains more descriptors than observations and multicollinearity between descriptors. It is one of the rare methods that allow for the prediction of more than one variable at the same time. PLS generalizes and combines features from both PCA and multiple regression. Linear combinations of descriptors are used to build new factors (identical to PCA) called latent variables. The objective is not to maximize variance but to obtain the best predictive power.PLS finds the linear combinations of predictors having the maximum covariance with the response variables. The obtained decomposition of the descriptors is used to predict the response values. Similar to PCA, the number of components can be chosen to explain enough variation of descriptors and responses while avoiding overfitting. This choice can be performed through cross-validation or use of a test dataset.
Both a prediction of the response values and graphical interpretation tools are provided. It is possible to study correlations between the original variables (descrip­tors and responses) and the latent ones. Figure 2.30 represents this type of correlation for descriptors (X1, ..., X5) and responses (Y1, Y2). The first latent variable (PLS1) is mainly predictive of Y1 values, which are highly related to X2 and X3 (more specifically Y1 tends to increase with X2 and when X3 decreases). The second latent variable (PLS2) mostly accounts for Y2 prediction, which is essentially due to X1 and X5. Y2 tends to increase when X1 and X5 decrease. Descriptor X2 provides information on Y2. This theoretical example shows the interpretation possibilities offered by PLS. Observations can be plotted in the new space to discover clusters and outliers.
STATISTICAL METHODS 93
https://t.me/medicina_free
Examples of ADME/Tox Applications PLS is mostlyused to predictthe values of one continuous response variable.The interpretation of latent variables helps to understand which predictors are most involved in the prediction accuracy. In Chohan et al. [312], PLSisusedamongotherregressionmethodstopredictCytochromeP450 1A2 inhibition (pIC
50
) from in-house computed descriptors accounting for topological, geometrical, and electronic features of molecules. Feature selection (according to variance, redun­dancy, and predictivity)wasperformed before PLS application.Thefinal model isbased on two components (chosen by minimizing the cross-validation error rate). Interpre­tation of correlations between the latent variables and original variables (pIC
50
and the descriptors) suggested that lipophilicity and aromaticity were the most important features to describe CYP1A2 inhibition. A more precise link can be derived showing that decreasing these two features should decrease inhibition.
In Luco [313], PLS is used to predict brain–blood distribution through computation and investigation of 25 structural descriptors. The variable importance for the projection (VIP) has been used to select the most relevant descriptors while strong outliers were removed. Three components (accounting for 85% of the variance in log BB) were retained (threefold cross-validation). A validation set was used with a careful verification of the applicability domain. The BBB and aqueous solubility in Obrezanova [314] was also modeled through PLS.
PLS is used in ADME/Tox context to perform discrimination. In Kriegl et al. [315], a classification model was trained on 2D, 3D, and quantum-mechanical descriptors to predict two (low versus high) or three (low, medium, high) classes of human cytochrome P450 3A4 inhibitors. PLS was shown to be less efficient than real discrimination methods.
Usage Warning The amount of variability provided by the successive latent variables has to be taken into consideration. The user should choose fewer compo­nents to obtain a model that could be generalized to new data for predictions. Using a
Figure 2.30 Example of simultaneous representation of correlation between the latent variables and original variables (descriptors are X
1–X5
and responses are Y1and Y2).
94
IN SILICO ADME/Tox PREDICTIONS
https://t.me/medicina_free
few latent variables provides graphical representations of the chosen ones. Interpre­tation of coefficients is not possible as descriptors are generally not independent.
PLS is often used to perform classification. It is simply obtained by providing the group labels as response variable. As PLS is a regression method designed to predict continuous variables, the qualitative side of group labels will not be taken into account. This can lead to negative values. Thresholds have to be defined to obtain the predicted groups. Group labels have to be coded into a dummy matrix containing as many columns as groups and filled with zeros and ones.
Although PLS can handle more variables than observations, more predictive models are obtained when irrelevant descriptors have been removed. Before applying PLS, careful preproce ssing of data is needed.
2.6.2.2 Technical Description
Input Data This method is based on the description of n individuals by p descriptors collected in a matrix having n rows and p columns, called X, and of the values of
q response values measured on the same n individuals and collected in Y.
Output The prediction of Y is provided. New coordinates for each individual in the
new subspace are given and are collected into a matrix having k columns (where k is the dimension of the new subspace) and called latent variables. PLS produces the coefficients applied to X and Y to build these latent variables.
Description PLS is a stepwise algorithm where at each step, one latent variable is provided. At the first step, one linear combination is defined for X and another for Y so that the covariance between them is maximum. This result is obtained via an iterative algorithm. The obtained new feature is called the first latent variable. It is possible to compute the proportion of variance of X and Y explained by the first latent variable. Information by the first latent variable is subtracted from both X and Y (or equivalently only from X) and new linear combinations are obtained in the same way, operating on the deflated matrices. The same process is repeated until the covariance is null.
PLS provides both graphical representation of the data and prediction of the response values. Graphical representations allow for both observations and predictors to be viewed on the same graph. The relative position of observations with regard to the different predictors is readily explained. Correlations between response and latent variables can be represented.
The choice of the number of latent variables to be used can be obtained through cross-validation or use of a test dataset. The objective is to obtain a precise prediction while avoiding overfitting (difficulties in generalizing to new observations). The optimal number of latent variables is the one resulting in the best prediction error.
Possible Extensions Most extensions of PLS use nonlinear regressions to introduce more complex combinations of predictors (see ASPLS [316] using spline functions to perform regression). Multiblock PLS where blocks of variables (or observations) are applied to the same operations during the latent variables construction is also possible.
STATISTICAL METHODS 95
https://t.me/medicina_free
2.6.3 Support Vector Machine
2.6.3.1 Aim The support vector machine classifier [98] was originally designed to perform discrimination between two groups (Figure 2.31). It has been extended to classify more than two groups and to perform regression (i.e., modeling one continuous variable). SVM is supported by two main ideas. A transformation performed by means of a function called kernel. The kernel projects the observations into a new space (generally of higher dimension than the original one) where the two groups are likely to be (or almost be) linearly separable. Second, building the best linear rule (a hyperplane) separating the two groups in the new space. In SVM, according to the Vapnik–Chervonenkis theory, this hyperplane is chosen so that the distance between it and the nearest observations on either side is maximized. It is called the maximum-margin hyperplane. Observations located on the margins are called support vectors.
2.6.3.2 Examples of ADME/Tox Applications Vapnik introduced SVM in 1995 with emergence to ADME/Tox work in the early 2000s. SVM can be used when several predictors (usually quantitative but there are some kernels to deal with qualitative ones) are used to predict one other variable. The predictive variable can be qualitative or quantitative.
For qualitative data, the objective is to predict membership of groups. A recent application of SVM is described by Hou et al. [317]. The goal was to predict HIA by defining two levels of absorption according to a fractional absorption threshold value. The prediction issue was converted into a binary classification problem that partic­ularly fits SVM use. One SVM model was built with each of the 10 chosen descriptors as input. This allowed to investigate individual performances of the descriptors and to rank them according to their predictive power (evaluated by a validation process recursively splitting the sample into learning and test ones). The usefulness of descriptor combinations was addressed. Systematic search was performed to find the best combination. In this work, SVM parameters (kernels, error penalty, etc.) were very carefully chosen by a validation process. A very good prediction rate was achieved for both classes (100% for the poor-absorption class and 97.8% for the good­absorption class).
Figure 2.31 Illustration of SVM principle. Starting from the left, observations are illustrated in their original space, and then are projected into the new space and finally separated by the maximum-margin hyperplane (circled points are called support vectors). See insert for color representation of this figure.
96
IN SILICO ADME/Tox PREDICTIONS
https://t.me/medicina_free
For quantitative data, any continuous (or discrete with really numerous values) variable can be chosen to be predicted. In Kortagere et al. [309], an SVM model was applied with shape signatures descriptors to predict BBB.
Usage Warning The main issue with SVM is to choose the different parameters. The kernel obviously is the first one. Despite some known recommendations, it is not an easy task and depends on the goals. When using new data, several kernels should be tried or an original one can be designed. The problem is exactly the same for kernel parameters even if some advices have been posted. A cross-validation choice has to be performed most of the time. The same work has to be done for the error penalty parameter that controls balance between overfitting (very precise modeling of the learning sample) and generalization error (ability to generalize the model to new observations). By setting an adequate value, it is always possible to make a perfect classification of the learning dataset. This is likely due to overtraining and no satisfactory results will be obtained on a test set. Any parameter choice has to be done through cross-validation. Eventually, obtaining a precise and robust model requires several choices that involve advanced analyses. As the transformed space is never explicitly known, it is not possible to have any interpretation of the original descriptors results. As numerous descriptors are often available, knowledge of the descriptors that account for the explanation of the predicted variable would be of interest. In the case of a “black box” method such as SVM, prior work on feature selection is required for any interpretation. To conclude, SVM is a powerful but not straightforward method allowing to model quite complex relationships.
2.6.3.3 Technical Description
Input Data SVM requires two objects, the description of n individuals by p descriptors and the vector of length n containing the data to be predicted (either
group numbers or any continuous variable).
Output The modeling of the vector to be predicted contains either the predicted groups or the predicted values of the target variable.
Description This description will focus on the original SVM issue, the discrimi­nation of two groups.
The first step is the kernel transformation. The goal is to project data into a higher dimension space where linearly dividing the data into two groups is easier. The kernel function allows combining original features leading to a new space describing the data in a higher dimension. According to the chosen transformation, the dimension can even be infinite. In the new space, computations may become difficult or even impossible. The so-called kernel trick indicates that if the kernel fulfils some properties, distances in the new space can be handled without computing the coordinates of its observations. Many different kernels can be used and the most common ones include the linear kernel (especially used in the case of large sparse data, such as found in text categorization), the Gaussian radial basis kernel (general­purpose kernel to be used in the case of no particular prior knowledge), the polynomial
STATISTICAL METHODS 97
https://t.me/medicina_free