Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5606_Библиотеки_им_академика_М_И_Перельмана.pdf
X
- •Foreword
- •Acknowledgments
- •Contents
- •1.1 Structure-Based Drug Discovery (SBDD)
- •1.2 Ligand-Based Drug Design (LBDD)
- •1.3 Echoes from the Past, Visions from the Future
- •References
- •1 Introduction
- •2.2 Second Step: Data Curation
- •2.4 Fourth Step: Updating and Maintenance
- •2 Databases and Curation
- •8 Perspectives
- •9 Conclusion
- •References
- •1 Introduction
- •2.1 Making and Matching Protein Models
- •2.2 Simulating Protein Movements
- •2.3 Analyzing Changes in Protein Shape
- •3 Pharmacogenomics in Drug Development
- •4 Case Studies of Genomics-Based Drug Design
- •References
- •1 Historical Background
- •1.1 Timeline
- •2 Methodology Overview
- •2.1 Neural Networks
- •2.1.1 Perceptron
- •2.1.2 Multilayer Neural Networks
- •2.1.3 Types of Neural Networks
- •Feedforward
- •Recurrent Neural Networks
- •LSTM
- •2.2 Deep Learning
- •3 Using Machine Learning
- •3.2 Data Collection
- •3.3 Data Preprocessing
- •3.4 Model Selection
- •3.5 Model Training
- •3.6 Validation
- •3.7 Tuning
- •3.8 Prediction
- •4 Limitations
- •4.1 Bias
- •4.3 Interpretability
- •4.4 Computational Cost
- •4.5 Data Dependency
- •4.6 Robustness
- •5 Applications in Drug Discovery
- •5.2 Lead Discovery
- •5.3 Preclinical and Clinical Development
- •6 Resources and Tools
- •7 Challenges and Perspectives
- •7.1 Future Trends
- •9 Conclusions
- •References
- •1 Historical Background
- •1.1 Applications in Drug Discovery
- •2 Validations and Controls
- •2.1 Internal Validation
- •2.2 External Validation
- •2.3 Relative Cluster Validation
- •3 Challenges and Perspectives
- •4 Conclusions
- •References
- •1 Historical Background
- •2 OECD Principles
- •2.1 A Defined Endpoint
- •2.2 An Unambiguous Algorithm
- •2.5 A Mechanistic Interpretation, if Possible
- •3 Software and Tools
- •4 Validations and Controls
- •4.1 Internal and External Validation
- •4.1.1 Regression Metrics
- •4.2 Applicability Domain
- •4.3 Randomization Tests
- •5 Interpretation
- •6 Practical Advice During QSAR Modeling
- •7 Application
- •8 Challenges and Perspectives
- •References
- •1 Molecular Docking
- •2 Advances in Scoring Functions and Search Algorithms
- •2.2 Critical Characteristics of Search Algorithms
- •2.3 Docking Programs and Scoring Functions
- •3 Calculations Performed During Docking Simulations
- •4 Essential Components for a Good Docking Program
- •5 Limitations of the Docking Technique
- •6 Validation of Docking Results
- •7 Inappropriate Use of Validation Methods in Docking
- •9 Use of Machine Learning in Molecular Docking
- •11 Challenges
- •12 Conclusions
- •References
- •3 System Preparation for MD Simulations
- •3.1 Solvation and Microensemble
- •3.2 Force Fields: General Concept and Relevant Choices
- •3.3 The Concept of Replicas and Timescale
- •4.1.2 Protein Root Mean Square Fluctuation (RMSF)
- •4.1.4 Protein Secondary Structure Analysis
- •4.1.5 Principal component Analysis (PCA)
- •4.1.6 Markov State Modelling
- •4.1.7 Distance Calculations
- •4.1.8 Angle and Plane Calculations
- •4.2.2 Distances and Ligand-Induced Geometry Rearrangements
- •4 Molecular Dynamics Analysis
- •4.1 Protein Perspective
- •4.1.1 Protein Root Mean Square Deviation (RMSD)
- •4.3 Ligand Perspective
- •4.3.1 Ligand Properties
- •4.3.2 Ligand Root Mean Square Deviation
- •4.3.3 Ligand Root Mean Square Fluctuation
- •4.3.4 Angles and Dihedrals
- •5.1 Protein Structure Prediction and Preparation
- •5.2 Molecular Docking
- •6 Concluding Remarks and Outlook
- •Glossary
- •References
- •1 Introduction
- •2.1 MDeNM
- •2.2 Collective Molecular Dynamics (coMD)
- •2.3 ClustENM and ClustENMD
- •3 Ensemble Docking
- •References
- •1 Introduction
- •1.1 Advantages, Disadvantages, Innovations, and Challenges
- •1.2 Recent Advances in Accessible FEP Software Tools
- •1.3 Applications of FEP in Industry and Consortiums
- •2 Expanding the Potential of FEP Calculations
- •2.1 Validating Binding Poses
- •2.2 Dealing with Solvent
- •2.3 FEP and Allostery
- •2.4 FEP and Covalent Ligands
- •2.5 Applications of FEP in Scaffold Hopping
- •2.6 Positional Analogue Scanning
- •2.7 Combinations and Alternative Approaches
- •3 Machine Learning for FEP
- •3.4 Implications for ML in FEP Calculations
- •4 Final Considerations
- •5 First Steps to FEP Simulations
- •References
- •1 Background
- •2 Ultra-Large Screening Libraries and Chemical Spaces
- •3.1 Implications of Dataset Size
- •4 Ligands on the Ultra-Large Scale
- •4.1 Ultra-Large 2D Similarity Searches
- •7 Challenges and Future Perspectives
- •7.1 Hit Triage: An Old Problem on a New Dimension
- •8 Conclusions
- •Appendix
- •References
- •1 Introduction
- •2 Enzymatic Activity Evaluations
- •3 Cytotoxicity Evaluation and Cell Viability
- •4 Antiviral Assays in Experimental Validation
- •6 In Vivo Evaluation of Compounds
- •7 Conclusions
- •References
- •1 Introduction
- •3.1 Data Collection
- •3.2 Data Preprocessing
- •3.4 Model Choice
- •3.5 Model Training
- •3.6 Model Assessment
- •3.7 External Validation
- •3.8 Implementation and Availability
- •3.9 Continuous Update
- •5 Conclusions and Perspectives
- •References
- •1 Experimental Approaches to Obtain Protein Structure
- •1.1 X-Ray Crystallography
- •1.2 Nuclear Magnetic Resonance
- •1.3 Cryo-EM
- •1.4 Hybrid Methods
- •2 Modeling Approaches to Obtain Protein Structure
- •2.1 Homology Modeling
- •2.2 Ab Initio Modeling
- •2.3 New Approaches
- •3 Conformational Diversity of Proteins
- •3.1 Characterization of Protein Conformational States
- •3.2 Experimental Methods to Study Protein Dynamics and Conformations
- •3.4 Molecular Dynamics Simulation
- •3.5 Sampling Strategies
- •4 Remarks and Perspectives
- •References
- •1 Introduction
- •2 Structure-Based Drug Design of HIV Protease Inhibitors
- •2.1 HIV-1 Protease as a Therapeutic Target
- •2.2.1 Saquinavir
- •2.2.2 Indinavir
- •2.3.1 Lopinavir
- •2.3.2 Darunavir
- •6 Conclusions
- •References
- •4 Experimental Methods to Analyze NR Activity
- •4.2 Coregulator-Recruitment
- •5 Concluding Remarks and Outlook
- •References

Chapter 12
Experimental Assays: Chemical Properties,
Biochemical and Cellular Assays,and In
Vivo Evaluations
Mateus Sá Magalhães Serafim, Erik Vinicius de Sousa Reis,
Jordana Grazziela Alves Coelho-dos-Reis, Jônatas Santos Abrahão,
and Anthony John O’Donoghue
Abstract The design and discovery of new bioactive compounds have been essen-
tial for the development of potential new inhibitors and drug candidates. In this
regard, the use of computational simulations has proven to play an important role in
achieving new drugs. Throughout history and most recently, new drugs, e.g.,
protease inhibitors, have benefited from the so-called computer-aided drug discovery
(CADD) approaches, providing available therapeutic options to emerging or
re-emerging diseases, such as the coronavirus disease 2019 (COVID-19). These in
silico models and methods can be employed for different purpos es, such as prediction of various biological activities, toxicity, pharmacokinetics, target specificity,
and even the synthesis of new analogs. Ultimately, such predictions can select or
disregard a given compound for an in vitro or in vivo evaluation. However, translating a simulation to an experimental validation may be chall enging. For instance,
one should consider chemical properties and solubility, different biochemical and
cellular assays, and the availability of data or methods to assess bioactive compounds and potentially reach a successful candidate. This chapter aims to provide a
detailed overview of the many computational possibilities to achieve or improve
experimental feasibility. Furthermore, we address the challenges and pitfalls regarding such approaches, which may contribute to a successful drug design and discovery campaign in the field.
Keywords Biological activity · Computational simulation · Experimental
validation · In silico · In vitro · In vivo
M. Sá Magalhães Serafim(✉) · E. V. de Sousa Reis · J. G. Alves Coelho-dos-Reis ·
J. Santos Abrahão
Department of Microbiology, Institute of Biological Sciences, Federal University of Minas
Gerais, Belo Horizonte, Minas Gerais, Brazil
A. J. O’Donoghue
Center for Discovery and Innovation in Parasitic Diseases, Skaggs School of Pharmacy and
Pharmaceutical Sciences, University of California San Diego, La Jolla, CA, USA
© The Author(s), under exclusive license to Springer Nature Switzerland AG 2024
V. G. Maltarollo (ed.), Computer-Aided and Machine Learning-Driven Drug
Design, Computer-Aided Drug Discovery and Design 3,
https://doi.org/10.1007/978-3-031-76718-0_12
347

348 M. Sá Magalhães Serafim et al.
1 Introduction
The field of drug design and discovery is marked by how expensive it is—about
$300 million to $2.1 billion per product [1], up to estimates of over $2.8 billion [2]—
and how long it takes, rangi ng from 5.8 to over 15.2 years, to design, develop, and
license a small-molecule drug for the market [3]. In addition, less than 30 new nextin-class drugs are expected to be introduced to the market over the next decade
(2030–2039) [4]. Arguably, cost- and time-saving opportunities are intrinsic to the
early steps of design and discovery, up to the preclinical stages, which account for
90% of clinical drug development failing [5]. Considering target selection, hit
identification and optimization, and the selection of a potential clinical candidate,
the current optimization of lead compounds usually emphasizes either potency,
specificity, or both [5, 6]. These steps may include different computational
approaches, such as structure-activity relationship (SAR) analysis [7], that may
integrate molecular information and biological outcomes [8].
The current state of the art is supported by an increasing number of different
computational simulations that can be implemented as a single model or in combination with other methods to properly predict general or specific outcomes, such as
biological activities [9]. One can cite the integration of machine learning (ML)
techniques as predictive and classifi cation models for potency [10], solubility [11],
and absorption, distribution, metabolism, excretion, and toxicity (ADMET) [12]
predictions, as well as the implementation of artificial intelligence (AI) in drug
discovery approaches [13]. Moreover, with the development of AlphaFold [14],
computational approaches have proven their ability to support and directly affect
drug discovery [15]. Therefore, the design and search for fast, reliable, and accessible ways to enhance drug discovery [16] could improve the discovery of diverse
hits and leads with optimal drug-like parameters, such as pharmacokinetics (PK) and
ADMET [17, 18], drug stability in solution [19, 20], and even alternative formulations [21]. In addition, these parameters could also improve the expected outcomes
in various stages of design and development in the preclinical phase, thereby
lowering associated costs and obtaining more effective and safer drugs [6, 16, 22].
In this sense, companies, institutes, and universities expand the use of computational approaches, as seen in the economic impact of AI usage [23]. For example, AI
can be used in synthesis planning [24], virtually screening ultra-large libraries with
billions of compounds at once [25], or filtering over a billion commercially available
compounds from an open-source library [26]. Furthermore, studies in the field of
drug discovery have already proven how fast, feasible, and reliable computational
approaches can be, such as in the coronavirus disease 2019 (COVID-19) pandemic
scenario, which urged computational approaches in a common effort for new
therapeutic options [27, 28]. For instance, in 2022, a successful approach resulted
in the design and discovery of the antiviral drug ensitrelvir, a severe acute respiratory
syndrome-related coronavirus 2 (SARS-CoV-2) main protease (M
Herein, hundreds of thousands of compounds were virtually screened for a lead
candidate, which was ultimately optimized with structure-based approaches [29].
pro
) inhibitor.

12 Experimental Assays: Chemical Properties, Biochemical and... 349
The various computer-aided drug discovery (CADD) approaches employed in
this pandemic scenario led to the rapid availability of potential drug candidates in
different stages of validation, but it also led to the surge of compounds supported
only by simulations [30]. In this sense, translating simulations to experimental
validations is a crucial step for drug discovery, which may avoid unrealistic results
from predictive models by providing reliable data from predictions complemented
with experimental evidence [31]. Experimental validation is referred to as the
procedure (i.e., in vitro assays) that is able to reproduce a scientific result obtained
using computational models or methods (i.e. , in silico predictions) [32], such as
those regarding biological activities [33]. However, validations may not be simple
and can face pharmaceutical bottlenecks [34], as in the case of nucleic acid drugs
[35], that is, drugs that control functions of cells based on nucleotide sequence
information (e.g., genome expression and gene regulation) [36]. Although difficult,
it is important to employ feasible experimental data to support, corroborate, or
demonstrate that a proposed simulation model or method is accurate and reproducible, thus validating a given study’s hypothesis [31, 37].
Furthermore, it is also important to consider the expected outcomes (e.g., negative or positive data) from an initial hypothesis, which may not have yet produced
the desired results [38], such as weak inhibitors being considered in optimization
studies [39], while also considering the disclosure of negative results to the scientific
community [40, 41]. Herein, avoiding false negatives and false positives is one of the
major challenges of screening approaches [42], and some elements may be potential
limitations that can affect the performance of any computational simulation, such as
the differences in datasets (e.g., size and diversity) [43]. Therefore, reiterating rigor
in good practices is of critical importance to CADD methods and their predicted
outcomes [30], allowing for better development, exploitation, and validation of
models. An example is the use of quantitative SAR (QSAR) analysis [44], in
which rigorous models’ design improve its ability to quantify the influence of each
chemical structure fragment toward a biological activity [7]. This is especially
important when considering various studies assessing the same inhibitor, such as
different IC
CoV-2 M
values obtained against the same target (e.g., GC376 against SARS-
50
pro
)[45]. Moreover, with the increasing abundance of available data (e.g.,
AI and predictive ML algorithms [46]) and methods for ultra-large [26] and accelerated [47] screening, the curation of data [48] is also necessary to assure the quality
of predictions for experi mental determinations.
The ability of in silico tools to accurately predict a bioactive compound may be
reaching the point of CADD turning to a computer-driven drug discovery [49]. Subsequent experimental determination is essential, such as in hybrid silico-vitro
approaches, which require a combination of multiple computational simulations
and in silico screening with specific and more sensitive experimental validation
in vitro [6]. For example, the experimental determination of various small binding
fragments complexed to enzymes during the design of potential inhibitors [50]. Hereupon, one could also cite the applicability of molecular dynamics (MD) simulations
in predicting binding sites as druggable pockets [51], which may vary in accuracy to
support ligands and the discovery of potential inhibitors for different targets (e.g.,

350 M. Sá Magalhães Serafim et al.
G-protein coupled receptors; GPCR) [52]. Predicted hits can be confirmed by
experimental determination in methods such as cryo-electron microscopy, X-ray
crystallography, and nuclear magnetic resonance (NMR) [51], thus supporting a
successful inhibitor [50], as well as determining binding affinity [ 53], or even
correlating potency [54]. Additionally, they can be performed using nanodifferential scanning fluorimetry (nanoDSF) and microscale thermophoresis
(MST), which can evaluate ligand binding by assessing protein denaturation and
stability at varied temperatures, and by changes in the fluorescence of tagged proteins induced by temperat ure, respectively [55 ].
However, achieving true inhibitors from initial simulations or fragment-based
models may be a long and costly effort that usually involves reassessing the ligand’s
design, synthesis, and experimental validation [56]. On the other hand, employing a
combination of ligand- and structure-based simulations with virtual screening (VS)
could optimize ligands as target-hits that better reproduce experimentally determined
conformations of know n inhibitors. Thus, a more cost-effective approach could be
directed towards experimental validation in vitro [57, 58], including potential
applications to other fields of research [59]. This combination of methods and
simulations may increase the overall accuracy of a given approach, as in the
combination of ligand-based drug design (LBDD) methods and QSAR [38, 60,
61], for example, combining molecular docking and QSAR models to discover
SARS-CoV-2 protease inhibitors [38]. In addition, the combination of ligandbased models in a consensus VS approach may benefit from the calculation of
decoys (i.e. , putative inactive compounds), which aim to increase the predictive
ability of identifying true negatives [62] in a VS [63]. Further, structure-based drug
design (SBDD) approaches may benefit from the design of a pharmacophore model
[64], which can improve the comparison of true inhibitors and designed ligands [65],
and increase the success rate of selected compounds against proposed targe ts [66].
It is also important to mention that computational approaches may never reach a
point where all predictions are correct [6]. Overall, VS campaigns may not even
reach a substantial number of hits accurately confirmed in experimental validation
(e.g., inactive compounds or false positive hits) [67]. For instance, this was observed
from SARS-CoV-2 M
pro
consensus VS approaches, which only one to three true
positive hits resulted from dozens to thousands of screened compounds (hit rate
ranging from 0.066% to 7.14%) [38]. Such a rate is still higher than usual expected
from high-throughput screening (HTS) approaches (between 0.01% and 0.15%)
[68], which are automated screenings of thousands of compounds in vitro. Nonetheless, high hit rates are also achievable in different VS approaches, typically in the
range between 1% and over 25% (median value of 13% in 421 studies) [68], as they
may be designed for specific targets and ligands and can be improved from available
computational and experimental data [69]. For example, a VS can identify compounds whose structures are complementary or appropriate for binding to a target
enzyme, depending on its binding site, and hits can account for up to 20% of a given
chemical library [70].
Notwithstanding, such simulations must be followed by experimental validation
in vitro and/or in vivo that can either verify the predictions or at least improve the

12 Experimental Assays: Chemical Properties, Biochemical and... 351
Computational approaches
ADMET
(Q)SAR
Target
Protein
Affinity
DL
Hit
Cell/Tissue
Organism
Fig. 12.1 An artificial network for translating computational approaches in drug design. A protein
or enzyme, a specific cell lineage or tissue, and a specific organism (e.g., virus) may be selected as
an initial target. Various parameters can be simulated with computational approaches before
experimental validation, such as absorption, distribution, metabolism, excretion, and toxicity
(ADMET) and pharmacokinetics (PK) properties of a given compound, in addition to its stability
in solution, affinity, or inhibitory activity against a given target. These data can be predicted with
various models and methods, such as quantitative structure-activity relationship (QSAR) analysis,
molecular docking, molecular dynamics (MD) simulations, as well as deep learning
(DL) approaches, including modeling from different software. Ultimately, a hit compound can be
obtained for subsequent experimental validation in vitro and/or in vivo
Inhibition
PK
Stability
Docking
MD
Modeling
quality of the model focusing on a proposed target. This may lead to better estimations of compounds’ properties such as ligand affinity and ADMET that translate the
computational approaches (Fig. 12.1) to more accurate in vitro and in vivo hit
results, thereby reducing test requirements [71]. However, aiming for the construction of these hybrid silico-vitro workflows requires extensive research teams and
laboratory structure, a large data management and curation system, as well as the
integration of academic researchers and pharmaceutical companies [72, 73].
As an example, hybrid silico-vitro drug discovery approaches against SARS-
CoV-2 M
pro
identified low-affinity binding fragments (IC50values between 180 μM
and 1 mM) from a combined crystallographic screening, that is, a screening of
potential binding fragments experimentally determining each of them by crystallography. This experimental screening was followed by a VS searching for similar
fragments as those that did bind to the target structure, identifying fragments that
displayed IC
values as low as 0.4 μM[74]. This study was followed by a similar
50
approach, which designed ligands by virtually linking small binding fragments into a

352 M. Sá Magalhães Serafim et al.
single scaffold and searching for similar compounds. Additionally, 450 million
compounds were screened using docking approaches searching for similar lead
binding fragments, which ultimately identified a 1.7 μM hit [75]. Despite being
successful, these studies took at least 2 years, required multiple different approaches,
and still did not reach the potency of the approved drug nirmatrelvir (IC
equals to
50
2.5 nM) [76]. In another example, from the COVID Moonshot initiative [77],
numerous research groups and crowdsourcing using computational approaches
focused on the design and discovery of novel inhibitors based on a noncovalent
and non-peptide inhibitor scaffold. After 2 years of this complex collaborative effort,
2400 new compounds yielded were assessed in more than 10,000 assays [77],
obtaining potent lead candidates such as MAT-POS-e194df51-1 (IC
equals to
50
37 nM).
However, the successful approach so far was the discovery of the antiviral drug
ensitrelvir, which was obtained in a collaboration of researchers after the desig n of a
lead candidate virtually screened from hundreds of thousands of compounds, which
was further optimized with structure-based approaches to the current drug [29]. Furthermore, the extensive COVID-19 drug discovery effort also led to failed clinical
trials [78], such as ebselen [79], reinforcing the discussion of the gaps, challenges,
and bottlenecks between computational approaches, proposed targets, and subsequent evaluation in vitro and in vivo. In this regard, the following topics discuss
potential pitfalls between simulations and experimental validation, aiming to understand and address potential issues, as well as the need to develop or improve
consensus approaches, including existing in silico models and methods in drug
design and discovery.
2 Enzymatic Activity Evaluations
Translating computational simulations to obtain bioactive compounds in a drug
discovery approach may start with the validation of a proposed target [80], such as
an enzyme. Approaches may consider choosing “one target to one drug” in different
models or methods, in the sense of predicting a potential desired (e.g., enzymatic
inhibition) or undesired effect (e.g., cytotoxicity or promiscuity) from a designed
inhibitor against the proposed target [81, 82]. In addition, one may also consider
alternatives that rely on polypharmacology, that is, the potential biological activity
of a small molecule interacti ng with multiple targets, which can be effective toward
multifactorial diseases and inflammation, such as producing synergistic effects
[83]. However, multi-target approaches can raise concerns due to limitations in the
ability of simulations to predict a target specificity or an unwanted promiscuity, as
well as potential adverse effects [80], such as the potential toxicity effects of some
agonists and antagonists over other enzymes [84] (not covered in this chapter). In
this sense, with the increasing rate of clinical trial failures [5], an ideal scenario
would comprise assessing large-scale approaches, considering drug-target networks,
prospective and retrospective drug-target relationships, and experimentally

12 Experimental Assays: Chemical Properties, Biochemical and... 353
quantifying the relation between a compound and a target to validate the performance of a model [81, 85]. For instance, assessing datasets of drug-target interactions and employing statistical validation, combined with the use of ML models for
predicting new compounds, and ultimately experimentally validating hits to reassess
the model can be used to improve the discovery of inhibitors [80].
Experimental determination of a given target in vitro is pivotal to predictive
models that aim to evaluate the potential interaction of a predicted ligand with a
given target [ 53 ]. However, these validation methods can be expensive and timeconsuming [85], or sometimes not possible to be performed due to the lack of an
available expressed and purified enzyme [86]. In addition, they may face difficulties
in establishing or standardizing assay concentrations, protein stability, substrate
specificity, and incubation periods, which may impair measurement reliability [87]
(e.g., in HTS applications [88]) and ultimately lead to incorrect reported affinities
[53]. Protein stability, for example, can be enhanced by consensus structural design
approaches, which focus on predicting and building well-folded or stabilized proteins that retain their original biological activities [89]. Additionally, some practical
considerations are to be sought for subsequent binding measurement assays, which
can help from predictions toward experimental validation, such as the enzyme
turnover numbe r (k
), which is essential to understand a given protein efficiency
cat
and its role in the metabolism [90]. Simulations may be tricky when predicting the
sparse data from experimentally measured k
values from different studies [91],
cat
especially when taking into consideration that only ~10% of all enzyme-catalyzed
reactions are known for Es cherichia coli [92].
As k
prediction models that assess comprehensive k
estimates are mostly unavailable for enzymatic reactions, computational
cat
data are desirable to simulate the
cat
assessment of metabolic models [93]. These can be expensive (e.g., production,
maintenance, and testing) [94] and require an intrinsic pre-requisite for modeling
(e.g., enzyme itself) [95]. High-throughput experimental validation assays, for
example, are costly and not time-effective, therefore models that can simulate
enzyme reactions are desirable, such as TurNuP, a web server that can generalize
and predict reactions from enzymes [93]. In addition, deep learning (DL) models
such as DLKcat use thousands of different substrate structures (>3,000) and enzyme
sequences (>300,000) to predict k
enzyme catalytic rates can be experimentally validated (e.g., k
values [91]. Thus, the characterization of
cat
measurements),
cat
assessing their potential correspondence to the in silico predictions [95]. However,
they may be susceptible to interference and variations when considering that changes
in protein conformation or structure could result in changes in their enzymatic
activity [96].
Elucidating this influence on altered activities or selec tivity issues from computational simulations is a challenge that can be explored by obtaining large and
diverse datasets, as well as defining suitable parameters for a given enzyme family
or class [97]. To this regard, one could access large databases, such as carbohydrateenzymes [98], kinome panels [99], network relations of various kinases and their
inhibitors [100], non-kinase enzyme assay panels (e.g., the CEREP diversity profile)
[101], or compiling general properties and information about a given enzyme class,

354 M. Sá Magalhães Serafim et al.
such as serine proteases [102]. Additionally, one could explore enzymes’ interactions and scaffolds by molecular docking analysis [103], or predict their ligand
accessibility and affinity by MD simulations [104], as in target-specific drug discovery approaches [43 ].
Notwithstanding, predicting whether an uncharacterized protein is an enzyme or
not [97], or predicting functions of uncharacterized enzymes, may benefit from
simulations and experimentation aiming for comprehensive data regarding translated
protein conformers and their function [105]. Further, models that incorporate a
single feature may also limit the applicability of a computational approach into
predicting enzymatic activity, suggesting the importance of consensus approac hes
or multiple methods to enhance the accuracy of an enzyme characterization, such as
substrate specificity [96, 106]. In this sense, different DL algorithm may extend an
ability of predicting single functions to multifunctional enzyme or multiple functions
predictions, as shown by the use and improvement of DEEPre [107] to the
mlDEEPre [108]. However, simulating potential substrates that an enzyme may
promote its catalytic activity is challenging when translating predicted data to
an experimental validation [109]. Herein, like k
predictions [96, 97], the combi-
cat
nation of existing data from experimental activity and structural information of
enzymes can be useful for proposing substrates using docking analysis, as well as
for calculating their physicochemical properties, thus favoring more accurate
predictions [109].
Moreover, one should account for existing different or variable enzyme domains,
which must be considered in simulations aiming to generate selectivity toward
substrates, as each domain may influence binding or site accessibility
[110, 111]. Hereupon, predicting a binding site preference or affinity to certain
molecular probes (i.e., functional groups or small molecules) in a multi-domain
structure can be helpful to support an enzyme’s specific substrate or inhibitor. Such
simulations, as in the prediction using molecular probes by FTMap [112] can be
considered a close computational analog of the experimental determination methods
[112], such as crystallography and NMR [113]. These may even be useful for further
predicting substrate fragment interactions in MD simulations (please, see Chap. 8)
based on previous crystal data [104]. Herein, the experimental validation of these
ligands can be assessed using designed or selected existing substrates, such as
accessing a physicochemically diverse library of peptides in a global identification
of substrate specificity [114]. This can be performed by employing an established
peptide cleavage assay yielded by quantitative multiplex substrate profiling by mass
spectrometry (qMSP-MS), which is able to quantitatively measure cleavage fluorescence units of each proposed substrate [115]. This approach can also be a valuable
tool to translate the experi mental validation to designed ligands from different
substrate preferences (e.g., supported by docking analysis) [116], be beneficial to
other enzymes and fields of research [117], and also be included in the design of
potential protease inhibitors in drug discovery [118].
In view of the binding affinity and specificity of a given substrate or inhibitor, it is
important to consider the binding equilibrium of an enzymatic reaction, such as
the dissociation constants (K
values), that is, the dissociation of an inhibitor or
D

12 Experimental Assays: Chemical Properties, Biochemical and... 355
substrate from the complex, accounting for the reaction velocity and bonding
formation between the target and the ligand [119, 120]. Experimental determination
of binding equilibrium would require demonstrating an absence of change for
complex formation over time and systematically varying the concentration of an
assay component, which may provide a solid method to display this behavior in vitro
[53]. Validating such inhibition behavior can be performed with dilution assays,
characterizing the nature of a given covalent bond, either it being reversible or
irreversible [121]. Nevertheless, as reversible behaviors usually tend to build complexes that block a given substrate proteolysis [122], they may also favor
noncovalent bonds [78]. In this sense, the formation of covalent or noncovalent
bonds can be simulated in MD simulations, as in the successful predictions of
covalent and noncovalent inhibitors for SARS-CoV-2 M
pro
[65].
Investigating the time-dependence of an inhibitor interaction with an enzyme
target can initially be investigated by assessing a compound with and without
pre-incubation with the enzyme target [119]. This might indicate or help in differentiating covalent from noncovalent bonds, as covalent bonds are usually timedependent, as opposed to noncovalent interactions [119]. In addition, a covalent
behavior can be reversible or irreversible, which can also be simulated in MD
simulations and experimentally validated with dilution assays aiming to investigate
reversibility. For instance, the covalent reversible behavior of nirmatrelvir was
initially predicted in MD simulations and confirmed in reversibility assays against
SARS-CoV-2 M
pro
[123]. Moreover, MD simulations can also predict the reversibility of protein recognition, interactions, and potential binding [124], as well as
suggesting other potential covalent or noncovalent inhibitors [125]. As covalent
bonds would confer an overall additional affinity in comparison to noncovalent
interactions in potential drug candidates binding [126], simulating and validating
such candidates could facilitate success in the drug discovery field [127].
Furthermore, complementary analysis may be of benefit for new drug candidates,
such as integrating proteomics platforms, focusing on the selectivity profiling of
covalent-bond ligand-target prediction [128], as well as predicting the reactivity
profiling of cysteines, which may be assessed for their potential catalysis influence
[129]. However, this could be less suitable when addressing the relative affinities of
particular enzymatic systems, such as DNA as a target in CRISPR-Cas systems,
which is a defense mechanism in some prokaryotes against viruses that have been
repurposed as an RNA-guided DNA targeting for genome editing [130], and thus
would require specifically designed simulations and experimental validation
[131]. One should also bear in min d the mechanism-based for inhibition in
predicting inhibitors, which are frequently quantified in terms of their half-maximal
inhibitory concentrations (IC
the rate of an enzyme inactivation (k
Predictions from k
inact
) values or may rely on an inhibition constant (KI) and
50
), depending on the types of inhibition [132].
inact
, KI, or their relation as k
, can be critical to SAR
inact/KI
analysis and PK param eters, and to accurately define a proper selectivity from
biochemical assays [133]. These can be illustrated with kinetic simulations, such
as covalent inhibitions predicted from SAR and their comparison to the potency of
known inhibitors, followed by detailed experimental validation in kinetics [134]. In

356 M. Sá Magalhães Serafim et al.
addition, IC50values are not simply predicted by chemical structure comparison or
available data regarding the mechanism and inhibition of similar ligands against
targets (e.g., structural analogs), as shown for a series of acrylamides and
propionamides found to not provide predictable responses in vitro [135]. Similarly,
predictions that do not comprise rigor in good practices may fail in experimental
determinations, even with the existence of an enormous amount of available data, as
seen for SARS-CoV-2 M
pro
inhibitors planned by QSAR analysis (see Chap. 6 for
methodological details on QSAR methods and proper validation protocols) [38].
It is also important to address whether a given compound with inhibitory activity
is in fact a true inhibitor or a false positive, aside from compounds’ solubility issues
[136, 137], or natural interferents [138] such as those presenting self-fluorescence
[139], which are covered in the following topics. For instance, an artifactual inhibition behavior can be observed in enzymatic assays due to colloidal aggregation
[140, 141], that is, compounds causing promiscuous inhibition as aggregators. Some
conditions can be assessed to detect aggregators, such as the addition of detergents
(e.g., tween or triton) to the assay buffer, which generally disrupts aggregates
[142]. In addition, compounds can be pre-incubated with bovine serum albumin
(BSA) to assess whether it can saturate the target enzyme binding capacity of the
potential aggregates [143]. Additionally, increasing the concentration of an enzyme
while maintaining the inhibitor concentration tends to reduce the percentage of
inhibition by aggregators [143].
Furthermore, dynamic light scattering (DLS) can also be employed to detect
aggregates in a solution. DLS is usually suitable to assess macromolecules degradation or disassembly, as well as enzyme-catalyzed polymerization [144], as these
measurements describe the ease with which a molecule displaces another one by
diffusion, expressly measuring the intensity of light scattering in an enzyme complex
formation [145]. Lastly, a counter-screening test using an unrelated enzyme, assayed
in the same buffer and inhibitor conditions, could also be employed [103]. Predicting
aggregation can be performed with reasonable accuracy using computational
methods [146, 147], but one should consider that even approved drugs can aggregate
at high concentrations [148]. Thus, it is important to make an educated decision
regarding the most suitable computational and validation methods to predict and
determine inhibitory activities and their mechanisms, such as potential aggregate
behavior [134], especially considering buffer, compound, enzyme, and substrate
concentrations [140, 148] (Fig. 12.2).
3 Cytotoxicity Evaluation and Cell Viability
Experimental validation of compounds in cell viability assays for cultured cells has
been based on different assays, including colorimetric, fluorometric, and
luminometric, as well as dye exclusion assays [149]. For instance, the reduction of
tetrazolium salts to colored formazans (e.g.,MTT[150 ] and XTT [151]), uptake and
incorporation of thymidine analogs for measuring DNA synthesis [152, 153],
Соседние файлы в папке Библиотека им академика М.И. Перельмана
