Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5606_Библиотеки_им_академика_М_И_Перельмана.pdf
X
- •Foreword
- •Acknowledgments
- •Contents
- •1.1 Structure-Based Drug Discovery (SBDD)
- •1.2 Ligand-Based Drug Design (LBDD)
- •1.3 Echoes from the Past, Visions from the Future
- •References
- •1 Introduction
- •2.2 Second Step: Data Curation
- •2.4 Fourth Step: Updating and Maintenance
- •2 Databases and Curation
- •8 Perspectives
- •9 Conclusion
- •References
- •1 Introduction
- •2.1 Making and Matching Protein Models
- •2.2 Simulating Protein Movements
- •2.3 Analyzing Changes in Protein Shape
- •3 Pharmacogenomics in Drug Development
- •4 Case Studies of Genomics-Based Drug Design
- •References
- •1 Historical Background
- •1.1 Timeline
- •2 Methodology Overview
- •2.1 Neural Networks
- •2.1.1 Perceptron
- •2.1.2 Multilayer Neural Networks
- •2.1.3 Types of Neural Networks
- •Feedforward
- •Recurrent Neural Networks
- •LSTM
- •2.2 Deep Learning
- •3 Using Machine Learning
- •3.2 Data Collection
- •3.3 Data Preprocessing
- •3.4 Model Selection
- •3.5 Model Training
- •3.6 Validation
- •3.7 Tuning
- •3.8 Prediction
- •4 Limitations
- •4.1 Bias
- •4.3 Interpretability
- •4.4 Computational Cost
- •4.5 Data Dependency
- •4.6 Robustness
- •5 Applications in Drug Discovery
- •5.2 Lead Discovery
- •5.3 Preclinical and Clinical Development
- •6 Resources and Tools
- •7 Challenges and Perspectives
- •7.1 Future Trends
- •9 Conclusions
- •References
- •1 Historical Background
- •1.1 Applications in Drug Discovery
- •2 Validations and Controls
- •2.1 Internal Validation
- •2.2 External Validation
- •2.3 Relative Cluster Validation
- •3 Challenges and Perspectives
- •4 Conclusions
- •References
- •1 Historical Background
- •2 OECD Principles
- •2.1 A Defined Endpoint
- •2.2 An Unambiguous Algorithm
- •2.5 A Mechanistic Interpretation, if Possible
- •3 Software and Tools
- •4 Validations and Controls
- •4.1 Internal and External Validation
- •4.1.1 Regression Metrics
- •4.2 Applicability Domain
- •4.3 Randomization Tests
- •5 Interpretation
- •6 Practical Advice During QSAR Modeling
- •7 Application
- •8 Challenges and Perspectives
- •References
- •1 Molecular Docking
- •2 Advances in Scoring Functions and Search Algorithms
- •2.2 Critical Characteristics of Search Algorithms
- •2.3 Docking Programs and Scoring Functions
- •3 Calculations Performed During Docking Simulations
- •4 Essential Components for a Good Docking Program
- •5 Limitations of the Docking Technique
- •6 Validation of Docking Results
- •7 Inappropriate Use of Validation Methods in Docking
- •9 Use of Machine Learning in Molecular Docking
- •11 Challenges
- •12 Conclusions
- •References
- •3 System Preparation for MD Simulations
- •3.1 Solvation and Microensemble
- •3.2 Force Fields: General Concept and Relevant Choices
- •3.3 The Concept of Replicas and Timescale
- •4.1.2 Protein Root Mean Square Fluctuation (RMSF)
- •4.1.4 Protein Secondary Structure Analysis
- •4.1.5 Principal component Analysis (PCA)
- •4.1.6 Markov State Modelling
- •4.1.7 Distance Calculations
- •4.1.8 Angle and Plane Calculations
- •4.2.2 Distances and Ligand-Induced Geometry Rearrangements
- •4 Molecular Dynamics Analysis
- •4.1 Protein Perspective
- •4.1.1 Protein Root Mean Square Deviation (RMSD)
- •4.3 Ligand Perspective
- •4.3.1 Ligand Properties
- •4.3.2 Ligand Root Mean Square Deviation
- •4.3.3 Ligand Root Mean Square Fluctuation
- •4.3.4 Angles and Dihedrals
- •5.1 Protein Structure Prediction and Preparation
- •5.2 Molecular Docking
- •6 Concluding Remarks and Outlook
- •Glossary
- •References
- •1 Introduction
- •2.1 MDeNM
- •2.2 Collective Molecular Dynamics (coMD)
- •2.3 ClustENM and ClustENMD
- •3 Ensemble Docking
- •References
- •1 Introduction
- •1.1 Advantages, Disadvantages, Innovations, and Challenges
- •1.2 Recent Advances in Accessible FEP Software Tools
- •1.3 Applications of FEP in Industry and Consortiums
- •2 Expanding the Potential of FEP Calculations
- •2.1 Validating Binding Poses
- •2.2 Dealing with Solvent
- •2.3 FEP and Allostery
- •2.4 FEP and Covalent Ligands
- •2.5 Applications of FEP in Scaffold Hopping
- •2.6 Positional Analogue Scanning
- •2.7 Combinations and Alternative Approaches
- •3 Machine Learning for FEP
- •3.4 Implications for ML in FEP Calculations
- •4 Final Considerations
- •5 First Steps to FEP Simulations
- •References
- •1 Background
- •2 Ultra-Large Screening Libraries and Chemical Spaces
- •3.1 Implications of Dataset Size
- •4 Ligands on the Ultra-Large Scale
- •4.1 Ultra-Large 2D Similarity Searches
- •7 Challenges and Future Perspectives
- •7.1 Hit Triage: An Old Problem on a New Dimension
- •8 Conclusions
- •Appendix
- •References
- •1 Introduction
- •2 Enzymatic Activity Evaluations
- •3 Cytotoxicity Evaluation and Cell Viability
- •4 Antiviral Assays in Experimental Validation
- •6 In Vivo Evaluation of Compounds
- •7 Conclusions
- •References
- •1 Introduction
- •3.1 Data Collection
- •3.2 Data Preprocessing
- •3.4 Model Choice
- •3.5 Model Training
- •3.6 Model Assessment
- •3.7 External Validation
- •3.8 Implementation and Availability
- •3.9 Continuous Update
- •5 Conclusions and Perspectives
- •References
- •1 Experimental Approaches to Obtain Protein Structure
- •1.1 X-Ray Crystallography
- •1.2 Nuclear Magnetic Resonance
- •1.3 Cryo-EM
- •1.4 Hybrid Methods
- •2 Modeling Approaches to Obtain Protein Structure
- •2.1 Homology Modeling
- •2.2 Ab Initio Modeling
- •2.3 New Approaches
- •3 Conformational Diversity of Proteins
- •3.1 Characterization of Protein Conformational States
- •3.2 Experimental Methods to Study Protein Dynamics and Conformations
- •3.4 Molecular Dynamics Simulation
- •3.5 Sampling Strategies
- •4 Remarks and Perspectives
- •References
- •1 Introduction
- •2 Structure-Based Drug Design of HIV Protease Inhibitors
- •2.1 HIV-1 Protease as a Therapeutic Target
- •2.2.1 Saquinavir
- •2.2.2 Indinavir
- •2.3.1 Lopinavir
- •2.3.2 Darunavir
- •6 Conclusions
- •References
- •4 Experimental Methods to Analyze NR Activity
- •4.2 Coregulator-Recruitment
- •5 Concluding Remarks and Outlook
- •References

80 D. S. de Sousa et al.
Fig. 4.7 Basic steps to
build a machine learning
model. Data management
(blue), model management
(green), validations and
controls (gray), and final
stage (purple)
Problem definition
Data collection
Data preprocessing
Model selection
Model training
Validation
Model tuning
Prediction
Let us take an example. Initially, understanding the nature of the target disea se, its
molecular characteristics, and underlying biological mechanisms is elementary.
From this understanding, the decision about what to do follows target discovery,
ligand search, prediction of biological activity, or evaluation of ADMETox (Absorption, Distribution, Metabolism, Excretion, and Toxicology) properties. Careful analysis will determine whether it is appropriate to employ ML for any of these
objectives.
Furthermore, it is necessary to define whether the appropriate approach would be
regression, classification, or another method. For example, in ligand search, classification techniques may be suitable, while in property prediction, regression may be
more indicated.
The decision about the type of problem and the specific task will guide the choice
of the most appropriate ML algorithm. This selection is fundamental to solidifying
the foundations of the successful develo pment of the scientific project at hand.

4 Machine Learning and Neural Network Methods Applied to Drug Discovery 81
3.2 Data Collection
Data collection is a fundamental element in the construction and performance of any
model. The quality and quantity of the provided data are primary in the effectiveness
of the resulting model. It is essential to consider various factors during this process to
ensure the robustness and relevance of the information used [60–64].
Firstly, the quantity of data is a determining factor. The more data made available
to the model, the greater its learning and generalization capabilities. However, the
pursuit of quantity should not compromise quality. It is important to ensure that the
collected data are relevant to the specific context of the problem at hand [60, 61].
Data representativeness is equally essential, especially in sampling cases. The
data should faithfully reflect the diversity and complexity of the environment or
phenomenon that the model aims to address. Otherwise, the model may develop
biases and may not be able to handle situations outside the scope of the provided data
[60, 61].
The choice of the database is a strategic step in data collection. Different types of
data can be stored in various formats, and the selection of the appropriate database
will directly influence the efficiency of the model. The structure and flexibility of the
database should be considered to ensure effective data handling throughout the
process. In the field of drug discovery, there are various databases designed for a
variety of applications, which are highlighted in Table 4.6 under the “Resources and
Tools” section.
3.3 Data Preprocessing
Data preprocessing is an integral step in ML, as the quality of data and the useful
information derived from it directly affect the model’s learning ability. In general, all
raw data undergo several processing stages to enhance its quality. These stages can
be simple steps that can be automated or even performed manually in some cases.
These steps can be listed as follows:
Data Cleansing This step involves removing missing, inconsistent, or irrelevant
data from data sets. Tasks within this phase may include eliminating records with
missing values, addressing outliers, and rectifying data errors. For example, in a
database of active compounds with various properties, some molecules may lack
Log P values. It is essential to remove them because inferring that these values are
null can significantly impair the model’s performance. Moreo ver, outliers in molecular data can distort the overall analysis. For instance, an outlier could represent a
compound with exceptionally high or low bindi ng affinity to a target protein [65].
Normalization or Standardization It is often necessary to normalize or standardize data to ensure uniform scaling across all characteristics. This measure is important in preventing features with vastly different magnitudes from exerting undue

82 D. S. de Sousa et al.
Table 4.2 Example of
one-hot coding for types of
molecules
Molecule type Small molecule Peptide Antibody
Small molecule 1 0 0
Peptide 0 1 0
Antibody 0 0 1
influence during the training process. Consider a data set that includes molecular
descriptors such as molecular weight, lipophilicity, and binding affinity to a target
protein. These descriptors may naturally exhibit different scales, wi th molecular
weight measured in Daltons, lipophilicity represented by log P values, and binding
affinity often quantified in terms of IC
or Kivalues. During the normalization
50
phase, each molecular descriptor is transformed to a common scale, scaling the
values of each descriptor to a range between 0 and 1, making them comparable
regardless of their original units [66]. This is very important because features with
larger scales might otherwise dominate the learning process, potentially
overshadowing the significance of other descriptors.
Encoding Categorical Variables When the data set includes categorical variables,
such as different molecule types or protein categories, it becomes necessary to
encode them into numerical values, often achieved through techniques like
one-hot coding. Consider a data set containing information about various drug
molecules, where a categorical variable represents the type of molecule, such as
small molecules, peptides, or antibodies. To incorporate this categorical information
into an ML model, one-hot coding is applied. In this technique, each category is
assigned a binary value, and a new binary column is created for each category. The
presence of a specific category is indicated by a “1” in the corresponding column,
while the absence is denoted by a “0” [67]. An example of this process is shown in
Table 4.2.
In this example, the categorical variable “Molecule Type” has been encoded into
three binary columns using one-hot coding. Each row now represents a drug
molecule and indicates its type through the binary columns. This encoding ensures
that categorical variables do not introduce ordinal relationships that may mislead the
ML model. By representing categorical information in a numerical format, the model
can effectively incorporate these features into its learning process, contributing to
more accurate predictions.
Dimensionality Reduction In high-dimensional data sets, such as those generated
in genomics studies, it can be useful to apply dimensionality reduction techniques
such as PCA (Principal Component Analysis) or t-SNE (t-distributed Stochastic
Neighbor Embedding). For instance, consider a genomics data set where each
sample is characterized by the expression levels of thousands of genes. The high
dimensionality of these data can make it computationally intensive and may lead to
the curse of dimensionality, where the model’s performance decreases as the number
of features increases. In the case of PCA, the algorithm identifies the principal
components, which are linear combinations of the original featu res that capture the

4 Machine Learning and Neural Network Methods Applied to Drug Discovery 83
Table 4.3 An example of the original data set
Sample Gene 1 Gene 2 ... Gene 1000
1 0.5 1.0 ... 0.8
2 0.8 0.7 ... 0.2
(continued)
Table 4.4 The same data set after applying PCA (reduced to 2 PCs)
Sample PC1 PC2
1 0.3 0.5
2 0.1 0.8
... ... ...
N 0.6 0.2
Table 4.3 (continued)
Sample Gene 1 Gene 2 ... Gene 1000
... ... ... ... ...
N 0.9 0.6 ... 1.0
maximum variance in the data. By retaining a subset of these principal components,
one can effectively reduce the dimensionality of the data set while preserving the
most critical information. This not only makes the data set more manageable but also
helps in identifying patterns and relationships between samples. Similarly, t-SNE is
a nonlinear dimensionality reduction technique that focuses on preserving the
pairwise similarities between data points. It is particularly useful for visualizing
high-dimensional data in lower-dimensional spaces, revealing the underlying structure and clusters within the data set [68, 69]. Tables 4.3 and 4.4 show a simplified
example with an original data set and after applying PCA, respectively.
After applying PCA, the data set might be represented using a reduced set of
principal components.
In this transformed representation, the dimensionality has been reduced to two
principal components (PC1 and PC2).
Data Sampling This technique is used to enhance the performance of ML models
when there is a significant imbalance between classes of interest. For instance, when
predicting the biological activity of chemical compounds, it is not uncommon to
encounter data sets where inactive compounds signific antly outnumber their active
counterparts. To mitigate this issue, two prevalent data sampling techniques are
commonly used: undersampling and oversampling. Undersampling involves reducing the number of instances belonging to the majority class (in this case, inactive
compounds) to match the number of instances of the minority class (active compounds). This approach helps balance the class proportions in the data set, ensuring
that the model is not overwhelmed by the abundance of instances from the majority

84 D. S. de Sousa et al.
class. While undersampling can lead to a reduction in the amount of data available
for training, it helps prevent the model from being biased toward the dominant class.
Conversely, oversampling entails generating additional copies of instances from the
minority class (active compounds) to balance the class proportions. This measure is
implemented to ensure that the model has sufficient exposure to instances of the
minority class, preventing it from exhibiting bias toward the majority class. Techniques such as duplicating existing instances, generating synthetic samples (using
methods like SMOTE—Synthetic Minority Over-sampling Technique), or other
sophisticated oversampling methods are employed to augment the representation
of the minority class in the data set [70, 71]. Consider a data set for predicting the
biological activity of chemical compounds, where only 10% of the compounds are
labeled as active, while the remaining 90% are inactive. To address this class
imbalance, an undersampling approach would involve randomly selecting 10% of
the instances from the inactive class, creating a balanced data set. On the other hand,
an oversampling approach might generate additional synthetic instances for the
active class to match the size of the inactive class, ensuring that the model is exposed
to a more balanced representation of both classes during training.
Feature Selection It is the process of discerning the most pertinent features for the
specific task at hand . The primary goals of feature selection are to reduce data
dimensionality and enhance model performance by focusing on the most informative
attributes. In various data sets, especially those with a large number of features, not
all features contribute equally to the predictive power of a model. Some features may
be redundant, irrelevant, or even introduce noise, leading to overfitting. Feature
selection helps address these issues by retaining only the most significant features,
thereby streamlining the data representation and improving the efficiency and
effectiveness of the ML model [72, 73]. For example, consider a data set for
predicting disea se outcomes based on patient profiles. The data set may include
numerous features such as age, gender, blood pressure, cholesterol levels, and
genetic markers. Feature selection would involve identifying which subset of these
features is most informative for accurately predicting the disease outcome. By
focusing on the most relevant features, the model becomes more interpretable,
computationally efficient, and less prone to overfitting. The process of feature
selection can be approached in various ways, including filter methods, wrapper
methods, and embedded methods. Filter methods assess the relevance of features
independently of the chosen ML algorithm, wrapper methods use the model’s
performance as a criterion for feature selection, and embedded methods incorporate
feature selection as an integral part of the model training process.
Separation of Training, Validation, and Test Sets This practice is important to
ensure robust evaluation and optimization of model performance. In this process, the
data set is divided into three distinct subsets. The Training Set is used to train the
model, allowing it to learn patterns and relationships in the data. During this stage,
the model’s parameters are iteratively adjusted to improve the accuracy of predictions. The Validation Set plays a key role in hyperparameter tuning and model
evaluation during training, which will be discussed later. The Test Set is reserved for

4 Machine Learning and Neural Network Methods Applied to Drug Discovery 85
the final evaluation of the model’s performance. It provides an unbiased assessment
of how well the model generalizes to new, unseen data. After training and
hyperparameter tuning, the model is evaluated on the Test Set to measure its ability
to make accurate predictions in real-world scenarios. In drug discovery, allocating
approximately 70– 80% of the data to the training set, 10–15% to the validation set,
and another 10–15% to the test set is a common practice [74 ]. The specific percentages may vary based on data characteristics, but this general split provides a balance
between training the model effectively and assessing its generalization capabilities in
drug discovery contexts.
Treatment of Temporal Data When dealing with temporal data, such as time
series records of biological activity over time, it is essential to appropriately handle
the data, taking into account temporal dependencies. Time series data often exhibit
patterns and dependencies over sequential observations. To address this, techniques
like lag features, rolling statistics, and time-based cross-validation can be employed.
Lag features capture historical values, rolling statistics summarize trends, and timebased cross-validation ensure evaluation reflects the temporal nature of the data.
Additionally, considering seasonality, trend decomposition, and incorporating timeaware models, like RNNs or LSTMs, enhances the modeling of temporal dynamics
in biological activity data sets [75].
3.4 Model Selection
In the application of ML in the context of drug discovery, the model selection stage
is important for identifying promising drug candidates. Let us consider a data set
describing the molecular properties of different compounds and their effectiveness in
treating a specific disease. In this scenario, choosing the appropriate model can be
crucial for accurately predicting the biological activity of new compounds.
Simpler models, such as linear regression, may be used for problems where the
relationships between molecular features and drug efficacy are predominantly linear.
However, in more complex cases where molecular interactions are nonlinear and
involve a variety of factors, more sophisticated models, such as NNs or DL methods,
may be more appropriate.
Model selection should consider the ability to handle molecular nuances such as
specific protein interactio ns, the three-dimensional structure of the molecule, and
other complexities inherent in biochemistry. Experimentation with different algorithms and parameter adjustments becomes even more relevant, as drug disco very
often requires highly speci alized models tailored to the specific characteristics of the
problem at hand [2–5, 38, 73]. The topic of applications addresses which environments in drug discovery are suitable for the most appropriate ML algorithms.

86 D. S. de Sousa et al.
3.5 Model Training
The training stage in ML models is fundamental for empowering the algorithm to
perform specific tasks based on the provided data. During this phase, the model is
exposed to labeled data sets, where inputs (features) are associated with known
outputs. The goal is to adjust the model’s parameters to minimize the difference
between the predicted outputs and the actual outputs present in the training data.
The training process can be divided into several iterations known as epochs. In
each epoch, the model goes through the enti re training set, makes predictions for
each input, and compares these predictions with the actual outputs. Based on this
discrepancy, an optimization algorithm adjusts the model’s weights and biases,
aiming to reduce error and improve predi ction accuracy [38, 46, 74].
The loss function is notable in this context, quantifying the discrepancy between
the model’s predictions and the actual labels. During training, the objective is to
minimize this loss function, resulting in a more accurate and generalizable
model [76].
The backpro pagation technique is widely used in this process. It involves calculating the gradient of the loss function with respect to the model’s parameters,
allowing weighted adjustments during the optimi zation phase [77]. This is essential
for updating the weights of connections between unit s in a neural network.
It is important to mention that the batch size is also a critical aspect. Training can
be performed in batches of data, where the model is adjusted based on a subset of the
training set [78]. This not only reduces computational requirements but also introduces a form of regularization that can benefit the model’s generalization to
new data.
Suppose we have a data set of molecular information about compounds and their
biological activities concerning a specific target such as a protein associated with a
disease. During training, the model would be exposed to this set, adjusting its
parameters to learn complex patterns that relate molecular features to biological
activities. The loss function would be applied to assess the discrepancy between the
activities predicted by the model and the actual activities observed in the training
data. A common example would be the use of Mean Squared Error (MSE) as the loss
function.
3.6 Validation
During the model evaluation phase, various methods are applied to ensure a comprehensive and accurate analysis of performance on unseen data. Initially, test data
sets, consisting of examples not used during training, are employed. The model
makes predictions for these data, and the discrepancy between predictions and actual
labels provides a direct measure of its generalization capability.

4 Machine Learning and Neural Network Methods Applied to Drug Discovery 87
In the context of classification, metrics such as precision, recall, and F1-score are
fundamental. Precision measures the proportion of correct predictions to the total
predictions, while recall (or sensitivity) assesses the proportion of true positives to
the total actual positives. The F1-score provides a harmonic average between
precision and recall, proving useful in situations with class imbalance. The confusion
matrix is a visual tool that exposes correct predictions and confusion between
classes [79].
For regression problems, metrics such as Mean Absolute Error (MAE) and Mean
Squared Error (MSE), along with Y-randomization, are employed. MAE represents
the average of absolute differences between predictions and actual labels, while
MSE measures the average of squared differences between predictions and actual
labels. Y-randomization is a technique used to asses s the robustness of a regression
model, evaluating whether the model captures the relationship between independent
and dependent variables or merely adjusts to the training data. The technique
involves randomizing the values of the dependent variable while keepin g the values
of independent variables unchanged. The model is then trained on the randomized
data, and performance is compared with the model trained on the original data. If the
model trained on randomized data performs similarly to the model trained on the
original data, it suggests that the model is not capturing the relationship between
independent and dependent variables [73, 79–82].
Cross-validation, such as the K-Fold Cross-Validation method, is another indispensable approach. This method divides the data set into K parts, trains the model on
K-1 parts, and evaluates one part. This process is repeated K times, and the average
of performance metrics provi des a more robust insight into how well the model
generalizes [60, 61].
Learning curves are valuable for understanding how the model’s performance
varies with the size of the training set. This helps determine whether the model
would benefit from more data or has reached a saturation point.
ROC (Receiver Operating Characteristic) curves and AUC (Area under the
Curve) are often used in classification problems. The ROC curve represents the
true positive rate versus the false-positive rate for different decision thresholds. The
AUC is a metric quantifying the discriminative ability of the model, with a higher
AUC being desirable [83].
Residual analysis is applied to understand the differences between the model’s
predictions and actual values. Observing patterns in residuals can indicate trends or
structures not captured by the model. Additionally, Bootstrap is a resampli ng
technique that generates multiple samples with replacemen ts from the original data
set. This approach can be useful for evaluating the stability of performance estimates
[80, 81].
Statistical tests are employed to assess whether performance differences between
models are statistically significant. This is especially relevant when comparing
alternative models to identify the most effective one.
In drug discovery, these methods are essential to assess the model’s effectiveness.
The choice of specific methods depends on the characteristics of the problem at
hand, but the joint application of these techniques is necessary for a robust and
reliable model evaluation.

88 D. S. de Sousa et al.
3.7 Tuning
The model tuning stage, also known as hyperparameter tuning, takes place after the
validation phase in the development process of the ML model. While validation is
crucial for assessing the model’s performance on a separate data set not used during
training, the tuning stage aims to further optimize the model’s performance by
adjusting its hyperparameters [84].
Hyperparameters are external elements to ML models, essential for influencing
their performance but not adjusted during training. Proper selection of these parameters is cruci al to ensure a well-performing model. Below are some common
hyperparameters and considerations on how to adjust them. The learning rate
controls the size of steps taken during optimization. It can be manually adjusted,
starting with a small value. To find a learning rate that results in stable and efficient
training, grid search or random search methods can be used. Grid search tests
combinations of hyperparameters in a predefined grid, while random search tests
combinations randomly. The number of epochs represents how many times the
algorithm goes through the entire training set during training. If the model is
overfitting, it may be necessary to reduce the number of epochs. Batch size refers
to the numbe r of training samples used in one iteration. Larger sizes can speed up
training but demand more memory. Smaller sizes may result in smoother convergence [84 – 86].
In the neural network architecture, the number of layers and neurons in each layer
is a significant hyperparameter. It is often advisable to initiate with a simple
architecture and gradually escalate complexity based on the specific demands of
the task. Employing cross-validation becomes essential in this iterative process to
identify the architecture that exhibits optimal generalization performance across
diverse data sets [87].
To mitigate the risk of overfitting in NNs, regularization-related hyperparameters,
such as L1 and L2 regularization terms, are highly recommended. These terms
introduce penalties for large weight magnitudes, promoting a more generalized
model. Experimentation with different values for regularization terms is
recommended, with larger values intensifying the penalty and aiding in the prevention of overfitting [87– 89].
Another pertinent hyperparameter is the dropout rate, denoting the percentage of
neurons randomly “turned off” during training to enhance robustness and prevent
overfitting. Commonly ranging from 0.2 to 0.5, the dropout rate can be finetuned
empirically based on the characteristics of the data. Careful adjustment of this
parameter contributes to the network’s ability to learn robust features while avoiding
reliance on specific neurons, thereby promoting better generalization to unseen
data [ 87 ].
In ensemble models like RF or gradient boosting, the number of trees is a
significant hyperparameter. A larger number generally improves performance, but
there is a point of diminishing returns. The maximum tree depth is relevant in
ensembles. Deeper depths in a model can potentially result in over fitting. Therefore,

4 Machine Learning and Neural Network Methods Applied to Drug Discovery 89
it is advisable to experiment with various depth values and employ cross-validation
as a means of assessing performance [90].
In the context of KNN (K-Nearest Neighbors) algorithms, the choice of the
number of neighbors is an elementary hyperparameter. Opting for small values can
make the model excessively sensitive to noise, potentially leading to overfitting. On
the other hand, selecting large values may result in an overly smoothed model,
potentially missing important patterns in the data [91].
In the case of SVM, the kernel type (e.g., linear, polynomial, and radial) and its
associated param eters are determined in model performance. The kernel defines the
transformation applied to the input data, and variations in kernel types and parameters can significantly affect the model’s ability to capture complex relationships
within the data. Therefore, a thoughtful exploration of different kernel types and
their respective parameter settings, possibly through techniques like grid search, is
essential to ensure the SVM model is appropriately configured for the specific
characteristics of the data set at hand [92].
It is important to emphasize that the evaluation and parameter tuning process take
place through iterative cycles. These cycles encompass ongoing analysis of the
model’s performance on validation data sets, facilitating the detection of potential
enhancements. Following each cycle, the validation process is concluded by evaluating the final model using an independent test set. The inclusion of this test set is
crucial to guarantee that the model not only adapts well to the training data but also
demonstrates effective generalization to new data. This ensures the robustness and
reliability of the model across diverse scenarios [73 , 80, 85, 86].
3.8 Prediction
The prediction phase in the ML process is the step where the trained model is used to
make predictions or inferences on unseen data. After completing the training,
validation, and hyperparameter tuning steps, the model is ready to be applied to
new data for making predictions or classifications.
During the prediction phase, input data are fed into the model, and the model
utilizes the patterns learned during training to generate predictions or inferences. The
specific nature of the prediction task can vary widely, depending on the type of
problem being addressed. Additionally, the interpretability of the model can be
explored to understand how the model makes decisions. In critical tasks, such as
in healthcare, the interpretability of the model is often as important as predictive
performance. The prediction phase is the culmination of the ML process,
transforming the knowledge acquired during training into practical insights and
actions for real-world applications.
Соседние файлы в папке Библиотека им академика М.И. Перельмана
