Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5884_Библиотеки_им_академика_М_И_Перельмана.pdf
X
- •Foreword
- •Acknowledgments
- •Contents
- •1.1 Structure-Based Drug Discovery (SBDD)
- •1.2 Ligand-Based Drug Design (LBDD)
- •1.3 Echoes from the Past, Visions from the Future
- •References
- •1 Introduction
- •2.2 Second Step: Data Curation
- •2.4 Fourth Step: Updating and Maintenance
- •2 Databases and Curation
- •8 Perspectives
- •9 Conclusion
- •References
- •1 Introduction
- •2.1 Making and Matching Protein Models
- •2.2 Simulating Protein Movements
- •2.3 Analyzing Changes in Protein Shape
- •3 Pharmacogenomics in Drug Development
- •4 Case Studies of Genomics-Based Drug Design
- •References
- •1 Historical Background
- •1.1 Timeline
- •2 Methodology Overview
- •2.1 Neural Networks
- •2.1.1 Perceptron
- •2.1.2 Multilayer Neural Networks
- •2.1.3 Types of Neural Networks
- •Feedforward
- •Recurrent Neural Networks
- •LSTM
- •2.2 Deep Learning
- •3 Using Machine Learning
- •3.2 Data Collection
- •3.3 Data Preprocessing
- •3.4 Model Selection
- •3.5 Model Training
- •3.6 Validation
- •3.7 Tuning
- •3.8 Prediction
- •4 Limitations
- •4.1 Bias
- •4.3 Interpretability
- •4.4 Computational Cost
- •4.5 Data Dependency
- •4.6 Robustness
- •5 Applications in Drug Discovery
- •5.2 Lead Discovery
- •5.3 Preclinical and Clinical Development
- •6 Resources and Tools
- •7 Challenges and Perspectives
- •7.1 Future Trends
- •9 Conclusions
- •References
- •1 Historical Background
- •1.1 Applications in Drug Discovery
- •2 Validations and Controls
- •2.1 Internal Validation
- •2.2 External Validation
- •2.3 Relative Cluster Validation
- •3 Challenges and Perspectives
- •4 Conclusions
- •References
- •1 Historical Background
- •2 OECD Principles
- •2.1 A Defined Endpoint
- •2.2 An Unambiguous Algorithm
- •2.5 A Mechanistic Interpretation, if Possible
- •3 Software and Tools
- •4 Validations and Controls
- •4.1 Internal and External Validation
- •4.1.1 Regression Metrics
- •4.2 Applicability Domain
- •4.3 Randomization Tests
- •5 Interpretation
- •6 Practical Advice During QSAR Modeling
- •7 Application
- •8 Challenges and Perspectives
- •References
- •1 Molecular Docking
- •2 Advances in Scoring Functions and Search Algorithms
- •2.2 Critical Characteristics of Search Algorithms
- •2.3 Docking Programs and Scoring Functions
- •3 Calculations Performed During Docking Simulations
- •4 Essential Components for a Good Docking Program
- •5 Limitations of the Docking Technique
- •6 Validation of Docking Results
- •7 Inappropriate Use of Validation Methods in Docking
- •9 Use of Machine Learning in Molecular Docking
- •11 Challenges
- •12 Conclusions
- •References
- •3 System Preparation for MD Simulations
- •3.1 Solvation and Microensemble
- •3.2 Force Fields: General Concept and Relevant Choices
- •3.3 The Concept of Replicas and Timescale
- •4.1.2 Protein Root Mean Square Fluctuation (RMSF)
- •4.1.4 Protein Secondary Structure Analysis
- •4.1.5 Principal component Analysis (PCA)
- •4.1.6 Markov State Modelling
- •4.1.7 Distance Calculations
- •4.1.8 Angle and Plane Calculations
- •4.2.2 Distances and Ligand-Induced Geometry Rearrangements
- •4 Molecular Dynamics Analysis
- •4.1 Protein Perspective
- •4.1.1 Protein Root Mean Square Deviation (RMSD)
- •4.3 Ligand Perspective
- •4.3.1 Ligand Properties
- •4.3.2 Ligand Root Mean Square Deviation
- •4.3.3 Ligand Root Mean Square Fluctuation
- •4.3.4 Angles and Dihedrals
- •5.1 Protein Structure Prediction and Preparation
- •5.2 Molecular Docking
- •6 Concluding Remarks and Outlook
- •Glossary
- •References
- •1 Introduction
- •2.1 MDeNM
- •2.2 Collective Molecular Dynamics (coMD)
- •2.3 ClustENM and ClustENMD
- •3 Ensemble Docking
- •References
- •1 Introduction
- •1.1 Advantages, Disadvantages, Innovations, and Challenges
- •1.2 Recent Advances in Accessible FEP Software Tools
- •1.3 Applications of FEP in Industry and Consortiums
- •2 Expanding the Potential of FEP Calculations
- •2.1 Validating Binding Poses
- •2.2 Dealing with Solvent
- •2.3 FEP and Allostery
- •2.4 FEP and Covalent Ligands
- •2.5 Applications of FEP in Scaffold Hopping
- •2.6 Positional Analogue Scanning
- •2.7 Combinations and Alternative Approaches
- •3 Machine Learning for FEP
- •3.4 Implications for ML in FEP Calculations
- •4 Final Considerations
- •5 First Steps to FEP Simulations
- •References
- •1 Background
- •2 Ultra-Large Screening Libraries and Chemical Spaces
- •3.1 Implications of Dataset Size
- •4 Ligands on the Ultra-Large Scale
- •4.1 Ultra-Large 2D Similarity Searches
- •7 Challenges and Future Perspectives
- •7.1 Hit Triage: An Old Problem on a New Dimension
- •8 Conclusions
- •Appendix
- •References
- •1 Introduction
- •2 Enzymatic Activity Evaluations
- •3 Cytotoxicity Evaluation and Cell Viability
- •4 Antiviral Assays in Experimental Validation
- •6 In Vivo Evaluation of Compounds
- •7 Conclusions
- •References
- •1 Introduction
- •3.1 Data Collection
- •3.2 Data Preprocessing
- •3.4 Model Choice
- •3.5 Model Training
- •3.6 Model Assessment
- •3.7 External Validation
- •3.8 Implementation and Availability
- •3.9 Continuous Update
- •5 Conclusions and Perspectives
- •References
- •1 Experimental Approaches to Obtain Protein Structure
- •1.1 X-Ray Crystallography
- •1.2 Nuclear Magnetic Resonance
- •1.3 Cryo-EM
- •1.4 Hybrid Methods
- •2 Modeling Approaches to Obtain Protein Structure
- •2.1 Homology Modeling
- •2.2 Ab Initio Modeling
- •2.3 New Approaches
- •3 Conformational Diversity of Proteins
- •3.1 Characterization of Protein Conformational States
- •3.2 Experimental Methods to Study Protein Dynamics and Conformations
- •3.4 Molecular Dynamics Simulation
- •3.5 Sampling Strategies
- •4 Remarks and Perspectives
- •References
- •1 Introduction
- •2 Structure-Based Drug Design of HIV Protease Inhibitors
- •2.1 HIV-1 Protease as a Therapeutic Target
- •2.2.1 Saquinavir
- •2.2.2 Indinavir
- •2.3.1 Lopinavir
- •2.3.2 Darunavir
- •6 Conclusions
- •References
- •4 Experimental Methods to Analyze NR Activity
- •4.2 Coregulator-Recruitment
- •5 Concluding Remarks and Outlook
- •References

398 J. E. Gonçalves
31. Pelletier, D. J., Gehlhaar, D., Tilloy-Ellul, A., Johnson, T. O., & Greene, N. (2007). Evaluation
of a published in silico model and construction of a novel Bayesian model for predicting
phospholipidosis inducing potential. Journal of Chemical Information and Modeling, 47(3),
1196–1205.
32. Zhu, X. W., Sedykh, A., Zhu, H., et al. (2013). The use of pseudo-equilibrium constant affords
improved QSAR models of human plasma protein binding. Pharmaceutical Research, 30,
1790–1798.
33. Berellini, G., Springer, C., Waters, N. J., & Lombardo, F. (2009). In silico prediction of volume
of distribution in human using linear and nonlinear models on a 669 compound data set. Journal
of Medicinal Chemistry, 52(14), 4488–4495.
34. Jones, H. M., Dickins, M., Youdim, K., Gosset, J. R., Attkins, N. J., Hay, T. L., Gurrell, I. K.,
Logan, Y. R., Bungay, P. J., Jones, B. C., & Gardner, I. B. (2012). Application of PBPK
modelling in drug discovery and development at Pfizer. Xenobiotica, 42(1), 94–106.
35. Chiann, C., Rama, E. M., & Crsitofoletti, R. (2011). Técnicas Computacionais em
Farmacocinética. In S. Storpirtis, M. N. Gai, D. R. Campos, & J. E. Gonçalves (Eds.),
Farmacocinética Básica e Aplicada. Guanabara Koogan. cap. 21.
36. Chicco, D. (2017). Ten quick tips for machine learning in computational biology. Biodata
Mining, 10, 35.
37. Shang, J., Sun, H., Liu, H., Chen, F., Tian, S., Pan, P., et al. (2017). Comparative analyses of
structural features and scaffold diversity for purchasable compound libraries. Journal of
Cheminformatics, 9, 25.
38. Schmidt, U., Struck, S., Gruening, B., Hossbach, J., Jaeger, I. S., Parol, R., et al. (2009).
SuperToxic: A comprehensive database of toxic compounds. Nucleic Acids Research, 37,
D295–D299.
39. Cao, D., Wang, J., Zhou, R., Li, Y., Yu, H., & Hou, T. (2012). ADMET evaluation in drug
discovery. 11. Pharmacokinetics knowledge base (PKKB): A comprehensive database of
pharmacokinetic and toxic properties for drugs. Journal of Chemical Information and Model-
ing, 52, 1132–1137.
40. Williams, A. J., Grulke, C. M., Edwards, J., McEachran, A. D., Mansouri, K., Baker, N. C.,
et al. (2017). The comptox chemistry dashboard: A community data resource for environmental
chemistry. Journal of Cheminformatics, 9, 61.
41. Lombardo, F., Desai, P. V., Arimoto, R., Desino, K. E., Fischer, H., Keefer, C. E., Petersson, C.,
Winiwarter, S., & Broccatelli, F. (2017). In silico absorption, distribution, metabolism, excretion, and pharmacokinetics (ADME-PK): Utility and best practices. An industry perspective
from the international consortium for innovation through quality in pharmaceutical development. Journal of Medicinal Chemistry, 60(22), 9097–9113.
42. Chi, C., Lee, M., Weng, C., & Leong, M. K. (2019). In silico prediction of PAMPA effective
permeability using a two-QSAR approach. International Journal of Molecular Sciences,
20, 3170.
43. Cai, X., Patel, S., Huang, C., Paiva, A., Sun, Y., Barker, G., Weller, H., & Shou, W. (2022).
Comprehensive characterization and optimization of Caco-2 cells enabled the development of a
miniaturized 96-well permeability assay. Xenobiotica, 52, 742–750.
44. Stéen, E. J. L., Vugts, D. J., & Windhorst, A. D. (2022). The application of in silico methods for
prediction of blood-brain barrier permeability of small molecule PET tracers. Frontiers in
Nuclear Medicine, 2, 853475.
45. Sasahara, K., Shibata, M., Sasabe, H., Suzuki, T., Takeuchi, K., Umehara, K., & Kashiyama,
E. (2021). Predicting drug metabolism and pharmacokinetics features of in-house compounds
by a hybrid machine-learning model. Drug Metabolism and Pharmacokinetics, 39, 100395.
46. Votano, J. R., Parham, M., Hall, L. M., Hall, L. H., Kier, L. B., Oloff, S., & Tropsha, A. (2006).
QSAR modeling of human serum protein binding with several modeling techniques utilizing
structure-information representation. Journal of Medicinal Chemistry, 49, 7169–7181.

13 Challenges Faced in the Development of Computational Methods... 399
47. Danishuddin, Kumar, V., Faheem, M., & Lee, K. W. (2022). A decade of machine learningbased predictive models for human pharmacokinetics: Advances and challenges. Drug Discov-
ery Today, 27(2), 529–537. ISSN 1359-6446.
48. Lee, C. H., & Yoon, H.-J. (2017). Medical big data: Promise and challenges. Kidney Research
and Clinical Practice, 36,3.
49. Schneider, P., Walters, W. P., Plowright, A. T., Sieroka, N., Listgarten, J., Goodnow, R. A.,
Fisher, J., Jansen, J. M., Duca, J. S., & Rush, T. S. (2020). Rethinking drug design in the
artificial intelligence era. Nature Reviews. Drug Discovery, 19, 353–364.
50. ISO 20691:2022 - Biotechnology — Requirements for data formatting and description in the
life sciences. https://www.iso.org/standard/68848.html
51. Wang, W., & Ouyang, D. (2022). Opportunities and challenges of physiologically based
pharmacokinetic modeling in drug delivery. Drug Discovery Today, 27(8), 2100–2120. ISSN
1359-6446.
52. Scannell, J. W., Bosley, J., Hickman, J. A., et al. (2022). Predictive validity in drug discovery:
What it is, why it matters and how to improve it. Nature Reviews. Drug Discovery, 21, 915–931.
53. Hooijmans, C. R., de Vries, R., Leenaars, M., Curfs, J., & Ritskes-Hoitinga, M. (2011).
Improving planning, design, reporting and scientific quality of animal experiments by using
the Gold Standard Publication Checklist, in addition to the ARRIVE guidelines. British Journal
of Pharmacology, 162(6), 1259–1260.
54. Fagerholm, U., Hellberg, S., Alvarsson, J., & Spjuth, O. (2023). In silico prediction of human
clinical pharmacokinetics with ANDROMEDA by Prosilico: Predictions for an established
benchmarking data set, a modern small drug data set, and a comparison with laboratory
methods. Alternatives to Laboratory Animals, 51(1), 39–54.
55. Sadybekov, A. V., & Katritch, V. (2023). Computational approaches streamlining drug discovery. Nature, 616, 673–685.

Chapter 14
Exploring the Significance of Experimental
and Computational Methods in Protein
Structure Determination
Adolfo Henrique Moraes, Diego Magno Martins,
and Marcelo Andrade Chaga s
Abstract This chapter describes X-ray crystallography, nuclear magnetic resonance
(NMR) spectroscopy, and cryo-electron microscopy (cryo-EM), the most used
experimental techniques to determine protein structure. X-ray crystallography is
highlighted as a powerful technique for resolving high-resolution structures of
crystallized proteins, while NMR spectroscopy offers insights into protein dynamics
and structures in solution. Cryo-EM, on the other hand, is emphasized for its ability
to characterize large protein complexes and membrane proteins without the need for
crystallization, providing near-atomic resolution. Meanwhile, computational
methods such as homology modeling are discussed as a method that leverages
known structures of related proteins to predict the structure of a target protein. The
chapter also highlights the transformative impact of AI, particularly AlphaFold and
RoseTTAFold, which has achieved unprecedented accuracy in predicting protein
structures from amino acid sequences. Additionally, molecular dynamics simulations are described as powerful tools for studying protein flexibility and conformational changes over time, providing dynamic insights that complement static
structural predictions. These experimental and computational approaches have
been advancing our understanding of protein structure and function.
Keywords X-ray crystallography · Nuclear magnetic resonance · Cryo-EM ·
Homology modeling · AlphaFold · Molecular dynamics · Protein structure
A. H. Moraes (✉) · D. M. Martins
Departamento de Química, Universidade Federal de Minas Gerais, Belo Horizonte, Minas
Gerais, Brazil
e-mail: adolfohmoraes@ufmg.br
M. A. Chagas
Departamento de Química, Universidade Federal de Minas Gerais, Belo Horizonte, Minas
Gerais, Brazil
Departamento de Ciências Exatas, Universidade do Estado de Minas Gerais (UEMG), João
Monlevade, Minas Gerais, Brazil
© The Author(s), under exclusive license to Springer Nature Switzerland AG 2024
V. G. Maltarollo (ed.), Computer-Aided and Machine Learning-Driven Drug
Design, Computer-Aided Drug Discovery and Design 3,
https://doi.org/10.1007/978-3-031-76718-0_14
401

402 A. H. Moraes et al.
1 Experimental Approaches to Obtain Protein Structure
Determining protein structure was one of the challenges of twentieth-century science
[1]. The development of experimental methodologies enabling the determination of
protein structures at atomic or near-atomic resolution has been a major contributor
for understanding different protein properties and functions, such as protein molecular recognition, enzyme catalysis, and others. The first technique developed was
X-ray crystallography, and the first protein crystal structure was published in 1958,
when John Kendrew and Max Perutz determined the structure of myoglobin [2]. In
1969, Max Perutz presented the hemog lobin crystal structure with 2.8 Å of resolution [3]. The development of X-ray crystallography and the many methodologies
related to protein crystallization and structure resolution provided a groundbreaking
contribution to our understanding of protein function and structure determination.
During the 1970s and 1980s, another technique emerged to study protein structure:
nuclear magnetic resonance, NMR. Initially developed after World War II, NMR
spectroscopy played a significant role in protein structure characterization after
advancements in Fourier transformation, pulsed NMR, and the 2D NMR [4]. In
1985, the first solution structure of a protein was determined by NMR spectroscopy
[5]. Since then, NMR spectroscopy has characterized several protein structures in
solution. Besides the structure determination of proteins, NMR spectroscopy has
been widely used to characterize protein dynamics at atomic resolution [6]. In the
2000s, another technique was developed and applied to determine supramolecular
protein structures, cryoelectronic microscopy, Cryo-EM [7]. Cryo-EM allowed the
structure characterization of large proteins and supramolecular structures, such as
virus particles, amyloid fibers, and protein-protein complexes [8–10]. Another valuable protein structure characterization and popularization tool was the Protein Data
Bank (PDB) organization in 1971 [11]. As of August 2024, the PDB hosts more than
200,000 protein structures. The PDB organization has played a pivotal role in
standardizing protein structure representation and developing quality control methodologies. Over the past few decades, these methodologies have evolved, ensuring
the reliability and accuracy of protein structures obtained using different experimental and compu tational strategies.
1.1 X-Ray Crystallography
X-ray crystallography was the first technique to obtain protein structure at the atomic
scale. Its application expanded over the last decades, mainly because of the development of protein crystallization methodologies and new computational strategies to
process experimental data and model the protein structures. These approaches led to
automation in protein structure determination by X-ray crystallography [12]. X-ray
crystallography is the technique responsible for most of the protein structure entries
in the PDB. The simplified steps of protein structure determination are usually the
following [12, 13]:

14 Exploring the Significance of Experimental and Computational Methods .. . 403
Sample Preparation: Protein purification is crucial for successful crystallization. The
protein must be highly pure and stable in solution. Various purification techniques, such as chromatography, are employed. Once purified, the protein is
concentrated to a suitable level for crystallization. Crystallization conditions are
refined by screening different precipitants, buffers, and pH values. Crystals
suitable for diffraction are grown by controlling temperature, pH, and concentration. Crystallization involves coaxing the protein molecules to form a highly
ordered array, or crystal lattice, whi ch can be subjected to X-ray diffraction.
X-ray Diffraction: X-rays are directed at the crystal and are diffracted by the electron
clouds of the atoms within the crystal lattice. The resulting diffraction pattern is
captured on a detector. The protein crystal is slightly rotated 360° along one axis
during the acquisition to ensure the recording of a diffraction pattern for each
position. X-rays generate protein crystal diffraction patterns because their wavelength (100 pm < λ < 10, 000 pm) is comparable to the size and atomic
constitution of proteins. The X-ray source most used today is synchrotron
radiation.
Data Analysis: The diffraction pattern can be used to infer the positions of the atoms
within the crystal. Complex mathematical algorithms based on Fourier transformation are used to analyze the diffraction pattern and reconstruct the protein’s
electron density map. The process of converting the reciprocal space representation of the crystal into an interpretable electron density map is known as phasing.
In simple terms, each spot of the diffraction image is indexed, integrated, merged,
and scaled. The position of each spot reflects the atomic constitution and threedimensional structure of a protein.
Model Building: Using the electron density map as a guide, researchers built a model
of the protein’s atomic structure, placing atoms in positions that best fit the
experimental data.
Refinement: The initial model is refined iteratively against the experimental data to
improve its accuracy and reliability. This procedure is simplified if a sufficient
resolution (less than 1.5 Å) is achieved. In this case, it is possible to automatically
generate a model based on the electron density map, with correct bond angles and
lengths. When the crystallography data are not of such high quality, molecular
visualization software is used to fit the protein structure model to the electron
density data. Validation: After the initial model is obtained, it undergoes rigorous
validation. This includes assessing the quality of the electron density maps,
checking for steric clashes, and evaluating geometric parameters such as bond
lengths and angles. Additionally, the model is confirmed against the protein’s
known biochemical and biophysical properties. Validation tools such as
MolProbity [ 14 ] and PROCHECK [15] are commonly used to assess the quality
of protein structures obtained by X-ray crystallography.

404 A. H. Moraes et al.
1.2 Nuclear Magnetic Resonance
NMR spectroscopy was the second experimental technique used to determine
protein structures. NMR spectroscopy’s great advantage is its ability to provide
protein structures in solution. This is especially important for proteins that are not
easily crystallized, such as highly dynamic proteins or whose crystals do not have the
quality needed to obtain the protein structures with atomic resolution. Unlike X-ray
crystallography, protein structure determination by NMR spectroscopy is achieved
by the simulation of the protein structure using minimization methodologies under
experimental constraints [16]. These experimental constraints are obtained from
different NMR experiments called NMR spectra. NMR spectroscopy is based on
measuring the precession frequency of the magnetic moment of nuclei in the
presence of a strong magnetic field. This frequency is called the Larmor frequency
and is modulated by the chemic al environment of each nucleus; therefore, it is a
probe of protein structure, as it reflects the electronic environment surrounding the
nucleus. Nuclei such as
them easily detected and analyzed by NMR spectroscopy.
Although it cannot be applied to large molecular systems (with molecular
weight > 100 kDa) [17], such as those studied by cryo-EM, NMR has emerged as
the foremost experimental technique for investigating protein dynamics with atomic
resolution [18]. NMR o ffers detailed structural insights through various parameters,
including chemical shifts, residual dipolar couplings, chemical shift anisotropy, and
paramagnetic relaxation enhancement. An overview of protein structure determination by NMR can be found in [19]. Additionally, rate constants can be determined by
assessing magnetization exchange, saturation transfer, and relaxation
dispersion [20].
The steps for protein structure calculation by NMR spectroscopy are usually the
following:
1H,13
C, and15N have a spin number of 1/2, which makes
Sample Preparation: For NMR studies, the protein must be dissolved in a suitable
solvent buffer at a concentration typically ranging from ~0.1 to 4 mM. The
solvent choice depends on factors such as protein stability and solubility. Isotopic
labeling with
assignment. The protein
15
Nor13C can enhance spectral resolution and help resonance
13
C and15N-labeling is achieved by producing the
protein using heterologous expression and by cultivating it in a unique medium,
known as minimal medium, where the only source of the elements nitrogen and
carbon are
15
N-ammonium chloride and13C-glucose, respectively [21–23]. Care
must be taken to ensure the protein remains in its native state without aggregation
or degradation during the NMR experiment. The quality and stability of the
protein sample can be monitored by
1
H 1D and1H–15N 2D NMR correlation
spectra.
NMR Experiment: The samp le is placed in a strong magnetic field, causing the nuclei
of certain atoms (e.g., hydrogen, carbon, and nitrogen) to align with the magnetic
field. Radiofrequency pulses are applied to perturb this alignment, and the

14 Exploring the Significance of Experimental and Computational Methods .. . 405
resulting responses, or NMR signals, are detected. Each nucleus will exhibit a
specific frequency value, which indicates a different chemical environment.
Data Collection: By using various pulse sequences and recording the frequencies
and intensities of the NMR signals, multidimensional spectra are obtained, which
contain infor mation about the spatial relationships between atoms in the protein.
Spectral Analysis: NMR spectra are analyzed to extract parameters such as chemical
shifts, coupling constants, and relaxation rates, which provide insights into the
protein’s structure, dynamics, and interactions [24, 25].
Structure Calculation: Computational methods, such as distance geometry, molec-
ular dynamics (MD) simulations (see Chap. 8), or simulated annealing, are used
to interpret the NMR data and generate a three-dimensional (3D) model of the
protein structure that satisfies all experimental constraints. The most common
experimental constraints are pair distance constraints obtained from nuclear
Overhauser effect (NOE) spectroscopy (NOESY) [5] experiments and the dihedral angles (ϕ and ψ ), which can be obtained from the
1H,15
N, and13C backbone
chemical shifts and the scalar coupling constants [19, 26].
Validation: NMR-derived structures are confirmed using a variety of techniques.
This includes assessing the precision and accuracy of the structure based on
experimental data such as NOE (Nuclear Overhauser Effect) intensities, chemical
shifts, and coupling constants. Structure validation tools like PROCHECK-NMR
[27], MolProbity [14], PSVS [28], and CING [29] are used to analyze the quality
of the NMR-derived structures, checking for stereochemical correctness and
overall structural quality.
1.3 Cryo-EM
Cryo-electron microscopy (Cryo-EM) has enhanced our ability to visualize biological macromolecules at near-atomic resolution [30]. Traditional methods such as
X-ray crystallography and NMR spectroscopy have long been the cornerstones of
protein structure determination. However, Cryo-EM has overcome many of the
limitations associated with these techniques for understanding the intricate architecture and function of biomolecules [31]. Historically, protein crystallization has been
a significant bottleneck in X-ray crystallography, often requiring laborious optimization and sometimes proving unattainable for specific proteins [13]. In contrast,
Cryo-EM enables the study of proteins in their native states without crystallization.
This breakthrough has democratized structural biology, allowing researchers to
tackle previously intractable targets and dynamic biological assemblies. Moreover,
Cryo-EM has redefined the concept of resolution in structural biology. While early
Cryo-EM structures were limited to low-resolution reconstructions (4–6 Å), recent
advancements in detector technology [8, 32], computational algorithms [33], and
sample preparation techniques [32, 34] have pushed the achievable resolution to
higher levels Cryo-EM now rivals X-ray crystallography in resolution (<1.5 Å)
[35, 36], offering insights into molecular structures with exquisite detail

406 A. H. Moraes et al.
[31, 37]. The impact of Cryo-EM extends beyond static structures; it provides a
dynamic view of biological processes. By capturing snapshots of biomolecules in
different conformations and functional states, Cryo-EM elucidates the mechanisms
underlying fundamental cellular processes, such as membrane transport, protein
synthesis, and signal transduction [38].
The cryo-EM technique involves capturing electron microscopy images of bio-
molecules embedded in a thin layer of vitreous ice, which are then used to generate
precise 3D reconstructions. These reconstructions offer intricate structural models,
shedding light on the functionalities of macromolecules and their involvement in
biological processes [32]. Noteworthy applications of Cryo-EM include the elucidation of tau filaments [39] and amyloid fibrils [40], providing crucial insights into
the mechanisms underlying Alzheimer’s disease. Moreover, Cryo-EM has been used
to determine the structure of the SARS-CoV-2 spike protein at a resolution of 3.5 Å
[41]. Cryo-EM‘s impact on the study of membrane proteins is particularly striking,
overcoming experimental limitations inherent in other techniques [42, 43]. Advances
in cryo-electron microscopy enable the solving of high-resolution structures of large,
>1 megadalton (MDa), and small, <100 kDa drug targets in near-native conditions,
routinely reachi ng resolutions around or below 3 Å [44]. One example of these
advances is the elucidation of the structures of native type A γ-aminobutyric acid
receptors (GABA
Rs) assemblies and their interactions with FDA-approved drugs,
A
such as those used to treat insomnia (zolpidem (ZOL) and flurazepam) and postpartum depression (the neurosteroid allopregnanolone (APG)) [45].
The steps of protein structure determination by Cryo-EM are usually the follow-
ing [46 ]:
Sample Preparation: Sample preparation for cryo-EM involves applying a small
volume (typically 3–5 μL) of the protein solution onto a holey carbon grid, which
is then blotted to remove excess liquid. The grid is rapidly plunged into liquid
ethane or propane, rapidly freezing the sample and forming a thin layer of
vitreous ice. Grid preparation should be carefully performed to ensure that the
protein particles are evenly distributed and not aggregated or absorbed into the
grid surface.
Data Acquisition: The frozen sample is imaged using an electron microscope under
cryogenic conditions. Electron micrographs are recorded at various tilt angles to
capture multiple 2D views of the protein particles embedded in the ice.
Image Processing: Advanced image processing techniques, such as single-particle
analysis or electron tomography, are used to align and combine the 2D images,
correcting for imperfections in the microscope and variations in particle orientation. The result is a 3D reconstruction of the protein density.
Model Building: Atomic models of the protein are fitted into the density map using
computational tools. This process involves manually or computationally positioning the protein coordinates within the density map and refining their positions
to maximize agreement with the experimental data. AI can be applied in this step
[33, 47].

14 Exploring the Significance of Experimental and Computational Methods .. . 407
Validation: Cryo-EM structures are confirmed through a combination of methods.
This includes assessing the resolution of the reconstructed density map using
criteria such as Fourier Shell Correlation (FSC) [48], which measures how two
similar signals are by comparing them in the frequency domain, focusing on
corresponding shells. It is widely used in microscopy, especially in structural
biology, to validate results, determine resolution, and enhance signals [48]. The
model is confirmed against the density map, ensuring that the atomic model fits
well within the density and does not clash with neighboring molecules. Other
validation measures include cross-validation against independent datasets,
assessment of local map quality, and comparison with existing structural data
or biochemical experiments related to the protein’s function [9]. A good discussion on the steps for checking the quality of protein structures determined by
Cryo-EM can be found in [49].
1.4 Hybrid Methods
Integrating multiple techniques allows researchers to cross-validate structural information from different experimental approaches, enhancing con fidence in the
resulting models. Moreover, hybrid methodologies enable the investigation of protein dynamics, interactions, and conformational changes that may be challenging to
capture using a single technique alone. Other techniques can provide complementary
information to X-ray crystallography, NMR spectroscopy, and Cryo-EM. Smallangle X-ray Scattering (SAXS) is a solution-based technique that gives information
about macromolecules’ overall shape, size, and conformation in solution [50]. SAXS
is particularly useful for studying the overall shape and conformational changes in
proteins and protein complexes in solution, including their flexibility and oligomeric
states. SAXS data can be integrated with high-resolution structures obtained from
X-ray crystallography, NMR spectroscopy, or Cryo-EM to generate pseudo-atomic
models that combine the detailed local information from these techniques with the
overall shape information from SAXS [51]; Hydrogen-Deuterium Exchange Mass
Spectrometry (HDX-MS) provides information about the solvent accessibility and
dynamics of protein structures by measuring the exchange of backbone amide
hydrogen atoms with deuterium atoms in solution. HDX-MS is used to study protein
folding, conformational dynamics, ligand binding, and protein–protein interactions
[52]. HDX-MS data can be integrated with high-resolution structures obtained from
X-ray crystallography, NMR spectroscopy, or Cryo-EM to provide insights into the
dynamics and flexibility of proteins and protein complexes. This information can
help refine structural models and elucidate dynamic aspects of protein function.
Chemical Cross-Linking Mass Spectrometry (XL-MS) finds cross-linked residues
within protein or protein complex structures induced by the covalent linkage of
reactive chemical probes . XL-MS is used to study protein–protein interactions,
protein structure, and conformational changes within protein complexes
[53]. XL-MS data can be integrated with structural information from other

408 A. H. Moraes et al.
techniques to identify interacting regions within protein complexes and confirm
protein–protein interfaces. XL-MS can also provide distance constraints for modeling protein structures, especially in regions where other techniques may be less
informative.
Additionally, computational methods can play an important role in hybrid
approaches for protein structure determination by integrating and interpret ing experimental data from multiple sources. These methods encompass various approaches,
including molecular modeling and bioinformatics algorithms, that can be used to
refine experimental structures obtained from X-ray crystallography, NMR spectroscopy, or Cryo-EM techniques by fitting structural models into experimental density
maps, refining protein–ligand interactions, and predicting protein dynamics and
conformational changes. Furthermore, computational tools help with the integration
of complementary experimental data, such as small-angle X-ray scattering (SAXS),
hydrogen-deuterium exchange mass spectrometry (HDX-MS), or chemical crosslinking mass spectrometry (XL-MS), into structural models, providing a more
comprehensive understanding of protein structure and function.
2 Modeling Approaches to Obtain Protein Structure
Many challenges persist despite innovations in experimental methods for determining protein atomic structure. Transitioning from the primary sequence to a protein’s
3D structure is complex. The rapid advancements in genomics have widened the gap
between the number of identified primary protein sequences and the experimentally
validated structures [54]. Predicting the 3D structure of a protein from its amino acid
sequence has been a major research challenge for over 50 years [55]. High-precision
computational approaches have made significant strides in bridging this gap and
advancing large-scale structural bioinformatics, even when existing experimental
methods fail to achieve atomic precision. Moreover, proteins and nucleic acids are
flexible structures, and their movements can play a fundamental role in their function
[56]. The methods and techniques used in computational modeling contribute to
understanding the molecular mechanisms of proteins at the atomic level.
The “Critical Assessment of Techniques for Protein Structure Prediction”
(CASP) [57] event resume is published every 2 years. This event aims to evaluate
the theoretical methods for predicting protein structure . During the event, protein
crystallography and NMR experts provide participants with prote in sequences
whose tertiary structures have been recently determined and not yet disclosed.
Several groups specializing in protein modeling by homology, ab initio methods,
and protein folding try to elucidate the structures obtained experimentally. The
results are discussed, and the findings are published in PROTEINS [57]. The last
edition, CASP15, occurred between December 10th and 13th, 2022. In this event,
for the first time, studies through homology modeling of RNA structures and
protein-ligand complexes were included, where classical methods showed results
in line with experimental data rather than techniques employing machine learning
[57, 58].
Соседние файлы в папке Библиотека им академика М.И. Перельмана
