Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5606_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
02.09.2026
Размер:
21 Мб
Скачать
14 Exploring the Signicance of Experimental and Computational Methods .. . 409
Given the challenges and advances, experimental and computational methods are
essential for understanding protein structures. Techniques like homology and ab initio modeling drive innovation in structural bioinformatics, offering insights with therapeutic potential. These methods not only deepen our understanding of protein structures but also inspire hope for the development of novel therapeutic strategies. The following sections will explore these promising avenues of modeling approaches.
2.1 Homology Modeling
Protein modeling by homology is one of the most common approaches for modeling protein structure and is based on molecular evolution. Protein evolution has involved complex mechanisms related to gene duplication and random DNA mutations [59]. These molecular changes have generated slightly different amino acid sequences, forming families of structurally related proteins [60]. Proteins with a common ancestry are called homologous [61]. The concept of homology encom­passes two related denitions: orthology and paralogy. Orthologous genes and proteins share a common ancestry and diverge due to speciation events, while paralogous genes and proteins differentiate due to gene duplication [62]. The sim­ilarity between amino acid sequences in homologous proteins, as expressed by the degree of identity, is less conserved than the similarity of their 3D structures [63]. Additionally, the similarity of the protein backbone increases with relative sequence identity. An identity above 50% is considered high, resulting in a small difference in RMSD (Å) between the 3D structures, while an identity below 30% is considered low, which signicantly impacts homology modeling [64]. Some histor­ical aspects of homology modeling are shown in Fig. 14.1 and discussed in several reviews, highlighting the information acquired in recent years [65].
Homology modeling uses previous information from structural similarities with
proteins from the same family [66]. The following steps (Fig. 14.2) are generally employed:
Selection of the protein primary sequence; Identication of structural homologies in
the PDB or another database system, such as the universal protein database (UniProt), currently the worlds largest database with information on proteins [67];
Alignment of the target protein with the homologous protein sequences: Sequence
alignment software such as Clustal [68], Omega [69], or MUSCLE [70] can be used;
Modelling structured regions: The structure is obtained by homology using the
alignment as a starting point to build the target protein model. This procedure can be performed using molecular modeling software such as MODELLER [63] or SWISS-MODEL [ 71 ];
410 A. H. Moraes et al.
Fig. 14.1 Historical timeline of major developments in homology modeling, CASP Critical Assessment of Protein Structure Prediction, GPU Graphics Processing Unit, TPU Tensor Processing Unit. (Figure adapted from reference [66])
Fig. 14.2 The seven classical steps of homology modeling. The shown structures exemplify events that inuence and are critical for modeling: (1) Selection of the protein primary sequence; (2) Alignment correction; (3) Modeling structured regions; (4) Loop modeling; (5) Side chains modeling; (6) Optimization, and (7) Validation. (Figure adapted from reference [66])
14 Exploring the Signicance of Experimental and Computational Methods .. . 411
Loop modeling: Loop region modeling can be done through database search
approaches or conformational analysis methods (ab initio). This is a critical step because loop coordinates ca n be crucial for protein functions;
Side chains modeling: The conformations of the side chains are also modeled based
on steric and energetic data from similar structures obtained experimentally. This process uses querying rotamer libraries, a scoring function, and a scanning method. Several software programs have been developed to sample rotational angles of the side chains, such as OPUS-Rota2 [72], SCWRL [73], and FASPR [74];
Optimization and renement: energy minimization process using molecular
mechanics force elds is applied to eliminate atomic overlaps and collisions. MD and Monte Carlo simulations can be performed as additional optimization measures [1];
Evaluation and Validation: Cross-validation methods and available tools can be
used, such as Distance-matrix alignment (DALI, http://ekhidna2.biocenter.
helsinki./dali/) or Verify3D (https://servicesn.mbi.ucla.edu/Verify3D/).
It is worth noting that most online servers and programs for homology modeling
automate all steps from the primary sequence to the 3D structure without user control over parameters. Other tools allow manual inspection, requiring more user knowl­edge about the system and program theories [66].
Approximately 75% of proteins are formed by multiple domains (about 2.1
domains in eukaryotic systems and 1.5 domains in prokaryotic systems) [75]. Under­standing the 3D structure of proteins with multiple domains can provide insights into several mechanisms, including the impact of mutations on the functions of multiple domains and their relationship with different types of diseases. The re are tools for modeling the structure of proteins with multiple domains. For example, an auto­mated method for homology modeling was developed using probabilistic data from multiple domains of primary sequence [76].
The following inherent aspects of homology modeling can be highlighted: (1) it
offers several signicant advantages and is efcient in generating 3D structures of proteins more quickly than through experimental techniques such as X-ray crystal­lography or NMR; (2) it allows an expansion of the range of structural coverage, as it allows the prediction of structures that have not yet been determined experimentally; (3) homology-generated models can be adapted for molecular docking simulations, enabling theor etical investigation of protein-substrate interactions in a variety of biological contexts. On the other hand, we can also highlight the following problems when using protein homology modeling: (a) model s may have limited accuracy since they are estimates based on structures of homologous proteins, mainly in regions of low primary sequence sim ilarity; (b) the quality of the structural model prediction depends on the quality of the alignment of the primary amino acid sequences used as a reference; (c) problems with modeling highly variable regions such as loops and side chain residues are very dependent on rotamer angle informa­tion, resulting in inaccu rate structures in these regions [66, 77, 78].
412 A. H. Moraes et al.
2.2 Ab Initio Modeling
It is possible to predict a proteins 3D structure from its primary sequence using ab initio protein modeling. This method relies on conformational analyses guided by a potential energy function based on the structural environment [54]. In ab initio modeling, a series of possible conformations (decoy structures) are generated and subsequently ranked and selected using conformational energy evaluation method­ologies. The success of ab initio modeling is related to three factors: (1) a function that accurately describes the potential energy of the proteins native structure and that leads to the most stable thermodynamic state through comparison with all possibilities of decoy structures; (2) an efcient search algorithm that can quickly characterize the lowest energy states through conformational analysis; and (3) strat­egies that select near-native structures from a set of decoy structures. It is important to highlight that ab initio modeling requires broader conformational sampling due to the complexity of structural space and the diversity of energy landscapes associated with different protein folding mechanisms [79 ].
2.3 New Approaches
Deep learning-based approaches such as AlphaFold [58] and trRosetta [80] were introduced in the CASP editions CASP14 and CASP15. These articial intelligence approaches utilize protein structure data available in the PDB. The computational experiments conducted during these CASP editions demonstrated that these approaches outperform traditional methods, such as genetic algorithms, graph the­ory, ML, and neural networks in protein folding [81]. AlphaFold2, developed by Googles DeepMind, showcased exceptional performance at CASP14 [82]. Concom­itantly, RoseTTAFold was developed, whose results were presented at CASP15 and incorporated into the Continuous Automated Model Evaluation (CAMEO) experi­ment [83]. CAMEO blindly evaluates structure prediction servers like the programs tested in the CASP editions using protein structures available in the PDB.
AlphaFold was the rst computational method capable of predicting protein
structures with atomic-level accuracy, even for cases without known homologous structures exist [58, 84]. The method was validated during the CASP14 and is currently in its second version. AlphaFold3: The basis of this new version of AlphaFold is the machine learning approach incorporating physical and biological information of protein structures, using alignments of multiple primary sequences. A two-way network architecture is employed, in which information from the sequence and the 2D distance map levels is transformed and transmitted interactively. The method implemented end-to-end learning, from the nal 3D coordinates generated along all network layers to the input sequence [58].
On the other hand, RoseTTAFold is based on a three-way network architecture,
which processes sequence alignment information and 2D distance matrix
14 Exploring the Signicance of Experimental and Computational Methods .. . 413
information in parallel, with better predictions than those provided by trRosetta, the second overall-best ranked method after AlphaFold2 in CASP14 [85]. The proposed three-way information architecture network might improve performance using an architecture operating in the 3D coordinate space. This innovation has provided a closer connection between sequence, distances between residues, and atomic orien­tations and coordinates. In this architecture model, structure prediction arises from the relationships between the amino acid sequence information, the 2D distance map, and the 3D coordinates. This allows the network to collectively reason about relationships within and between sequences, distances, and coordinates. This inter­active mechanism, therefore, contrasts with the processing of 3D atomic coordinates in the two-network architecture of AlphaFold2, which occurs after the processing of sequence and distance maps is completed. One of the limitations, due to computer hardware memory problems, was the impossibility, at the time, of training the model for large protein systems [80].

3 Conformational Diversity of Proteins

Proteins do not exist in a single, xed conformation. Instead, they are best described as conformational ensembles, wherein multiple conformations coexist in equilibrium [86, 87]. Among these conformations, some states are more frequently populated than others. The extent of conformational diversity varies among proteins. While some proteins exhibit only minor conformational changes, others, known as intrin­sically disordered proteins (IDPs), occupy many conformations, making them highly exible and dynamic [88, 89].
In one-domain proteins, the movements responsible for conformational diversity
include the rotation of amino acid residue side chains, loops, and the collective movement of connected parts [87]. The rotation of amino acid side chains plays a crucial role in the protein folding process and the transmission of information through allostery [9092]. Loops, which can be well-dened or disordered random­coil-like structures, generally serve as exible connectors between more regular secondary structures (α-helices and β-strands). However, it is important to highlight that loops are more than connecting elements as they interact with the solvent, ligands, and other biomolecules. They also participate in the conformational transi­tions involved in protein regulation [93].
The different modules comprising multidomain proteins are commonly
interconnected by short or lengthy sequences of amino acids, typically ranging from approximately 5 to 25, known as linkers. Variations in linker sequences inuence their conformational dynamics, potentially affecting cooperative adjust­ments in interdomain and protein–protein interactions. These changes can facilitate long-range communication between different functional modules in multidomain protein [93]. Rigid domains can also be connected by exible joints known as hinges. During hinge bending motions, structural units shift relative to each other,
414 A. H. Moraes et al.
disrupting the interfacing packing while preserving the compact arrangement within the protein subunit [87].
3.1 Characterization of Protein Conformational States
Proteins can adopt various conformational states in solution, which may differ from those observed in the crystal lattice. This discrepancy arises because crystall ization conditions can impose constraints and stabilize specic conformations, potentially altering the native structural dynamics of the protein [94]. Various ensemble tech­niques have been developed to investigate protein conformations and dynamics using experimental and computational methodologies. Prominent experimental approaches include NMR spectroscopy, SAXS, single molecule spectroscopy, and Cryo-EM. These experimental techniques are frequently complemented by compu­tational methods that use empirical data as structural constraints. The computational methods include REMD simulations, metadynamics, steered MD, accelerated MD, and Markov state models (Fig. 14.3)[87].
3.2 Experimental Methods to Study Protein Dynamics and Conformations
Experimental determination methods can be employed to determine the conforma­tions of proteins. In X-ray crystallography, conformational heterogeneity hinders crystal formation. Even when crystals are successfully grown, this variability can cause disorder in regions of the structure that are more mobile [95]. Conversely, in cryo-EM structural determination, computational techniques often aid in identifying such heterogeneity. This may involve focusing analyses solely on structurally homogeneous segments of the complex or deriv ing an ensemble of structures when discrete subpopulations are present. Importantly, delineating conformational diversity often provides invaluable insights into the underlying functional mechanisms [96].
NMR offers various experimental strategi es to examine protein kinetics and interactions across various time scales and binding afnities [97]. For instance, the dissociation constants (K femtomolar for proteases, RNases, and DNases to millimolar for protein­carbohydrate molecular recognition [20, 98]. NMR is a potent tool for interface mapping, facilitating the investigation of weak and transient complexes that may pose challenges for other experimental methodologies [99].
Unlike NMR and Cryo-EM, small-angle X-ray scattering (SAXS) does not offer atomic resolution. It is typically considered a low-resolution structural techni que for investigating conformational changes, ligand binding, or protein interactions over time. Nevertheless, SAXS can provide valuable insights into protein behavior,
) associated with protein interactions range from
d
14 Exploring the Signicance of Experimental and Computational Methods .. . 415
Fig. 14.3 Timescale of dynamic processes in proteins and the experimental methods that can detect uctuations on each time scale. Local motions, including bond vibrations, methyl rotations, loop motions, and sidechain rotamers, occur on timescales ranging from femtoseconds (fs) to nanosec­onds (ns). Global motions, such as larger domain movements, range from nanoseconds (ns) to seconds (s). The timeline depicted spans from fs to s. The methods used to investigate these motions include MD simulations (fs–ms), NMR (ps–s), smFRET (ns–ms), AFM/OT (μs–ms), hydrogen­deuterium exchange (ms–s), SAXS (μs–s), Cryo-EM (ms–s), IR/Raman (fs–s), and uorescence (fs–s). The gure highlights the temporal ranges of these motions, and the corresponding techniques used for their study
including subtle domain movements [100]. Notably, SAXS is a rapid structural approach that can be directly applied to solutions of macromolecules under various conditions, including quasi-native environments [101]. SAXS provides a convenient means to evaluate alterations in the extension or exibility of a protein. Among its most impactful applications is the assessment of conformational transitions triggered by ligand or protein partner binding, allowing for the characterization of the overall shape and oligomerization state of the biological macromolecule [101, 102]. Time­resolved probing of conformational changes can be achieved through various mixing methods (e.g., stopped-ow, laminar ow, chaotic ow) and pump-probe techniques [100].
SAXS is a well-establ ished technique for structural characterization of samples, offering resolutions ranging from 5 nm to 1000 nm. Notably, SAXS exhibits sensitivity to both ordered and disordered features within the sample and circum­vents the need for crystallization, xation, or vitrication procedures. In a typical SAXS experiment, a collimated, monochromatic X-ray beam is directed onto the sample, with the radiation scattered at low angles (usually a few degrees) being
416 A. H. Moraes et al.
captured by a detector [103]. Due to the generally random orientations of particles within the solution, a rotational averaging occurs, resulting in a one-dimensional scattering prole expressed as I(q) versus q, where I represents the net intensity and q denotes the momentum transfer variable. This variable is dened as q = (4πsinϴ)/
radiation. Interpreting the solution-scattering prole in terms of average global parameters, such as radius of gyration (Rg), molecular volume, and the interatomic distance distribution P(r) versus r (which is related to I(q) through a Fourier transform), is relatively straightforward. However, contemporary approaches increasingly involve interpreting SAXS data using 3D models in shapes or atomistic representations. This needs careful consideration of how to accurately constrain the structural landscape to nd the correct solution while acknowledging the inherent ambiguity if more than one 3D model can t the same one-dimensional scattering prole [104].
Single-molecule spectroscopy methods include force-based techniques such as atomic force microscopy (AFM) [105] and optical tweezers [106], as well as uorescence-based techniques, such as single-molecule uorescence resonance energy transfer (smFRET) [107 ]. These methodologies have gained widespread use in quantifying the conformational heterogeneity and structural dynamics of biomolecules, both in vitro and in vivo. They enable the observation of transient intermediates and the characterization of static and dynamic heterogeneity [87, 108].
Several instances highlight the efcacy of these single-molecule techniques in elucidating the dynamics of interactions among biological molecules. Notably, riboswitches RNA structures that trigger conformational and functional changes upon ligand binding have been investigated using these methods. Optical twee­zers have facilitated the study of numerous riboswitches, demonstrating how ligand adenine binding orchestrates the coordination of domains within the three-loop structure [109, 110]. This observation is supported by steered MD simulations [111]. Furthermore, single-molecule Förster resonance energy transfer (smFRET) has enhanced our understanding of membrane transporter dynamics [112], including ATP-binding cassette (ABC) transporters [113
, 114 ].
3.3 Computational Methods for Exploring Protein
Conformations
Currently, the computational study of macromolecules at the molecular level involves steps and methodologies [115, 116]. Among the most used methods, we can mention molecular docking, classical MD, hybrid methods quantum-mechanics/ molecular-mechanics (QM/MM) [117], and quantum-mechanics/molecular­mechanics/molecular-dynamics (QM/MM/M D) [118], which can be used to inves­tigate the interactions between ligands and proteins, between proteins, proteins and membranes, proteins and DNA, and RNA.
14 Exploring the Signicance of Experimental and Computational Methods .. . 417
3.4 Molecular Dynamics Simulation
MD simulations use the equations of motion from classical mechanics to describe molecular events by analyzing conformational changes over time. Variations in this methodology have been successfully employed to study: (1) protein stability, (2) conformational changes, (3) interactions with natural biological cofactors, (4) aiding in the rational design of new drugs, (5) ab initio structure prediction of peptides and small proteins, (6) renement of protein structure models, (7) and structure modeling assisted by experimental or bioinformatics data [119].
MD simulation, developed in the late 1970s, has advanced from initially being restricted to molecules with hundreds of atoms to systems with millions, such as proteins [120 ]. Nowadays, it is possible to study whole proteins in solution using explicit solvation, membrane proteins, and large macromolecular complexes such as nucleosomes and ribosomes. Computational simulations of systems with ~50,000–100,000 atoms are now routine, and simulations of approxi mately 500,000 atoms are common when proper data processing computational facilities are available [121]. Thus, signicant computational development and major advances in MD algorithms have been made in recent decades, providing greater efciency in molecular simulations. This notable improvement is primarily a con­sequence of using high-performance parallel computing architectures (HPC) and the simplicity of the MD algorithm. Recently, progress has been made in developing software allowing routine MD simulations on graphics processing units (GPUs). MD simulations of hundreds of nanoseconds (ns) per day can now be easily achieved for small protein systems in explicit or implicit solvents.
Some of the most used software in MD are AMBER [122, 123], CHARMM [124], OPLS-AA [125], and GROMACS [118]. The MD is based on the principles of classical mechanics, where each atom is considered a point mass body whose motion is inuenced by the forces acting on it from all other atoms. These forces acting on each body are calculated at each MD step using force elds. Force elds are compl ete interaction potentials between atoms that combine to comprise the systems total energy. These mathematical models include potentials for bonded terms (bond lengths, bond angles, dihedral angles) and non-bonded terms (van der Waals and Coulomb interactions) [54]
UR
V
n
1 þ cos nϕ - δðÞðÞþ
dihedrals
2
The systems total energy UR
=
bonds
Krr - r
ðÞ2þ
0
non - bonded terms
is obtained by summing the energies of the
K0ðÞ
angles
bonded and non-bonded terms. In this equation, K
r
2
þ
A
ij
-
12
r
ij
, K
, and Vnare the force
qiq
B
ij
j
þ
6
ϵR
r
ij
ij
constants for bond stretching, angle bending, and torsional potential, respectively.
418 A. H. Moraes et al.
r, θ, and ϕ are bond length, bond angle, and dihedral angle. The subscripted zero terms stand for the equilibrium values for the individual variable. For example, the AMBER program calculates the force eld, chemical bond stretching, and angular strain terms using equations with harmonic potentials described by quadratic terms. The terms related to torsional potentials are determined by truncated Fourier series; Van der Waals interactions are approximated by the 12–6 Lennard–Jones (LJ) potential. Coulomb’s equation is used to calculate the electrostatic interaction potential. The types of charges vary with the force eld version, but for most AMBER force elds, partial charges are obtained using the RESP (Restrained Electrostatic Potential) method. Some force elds include other terms to obtain better agreement with vibrational spectra, for example, terms that better specify hydrogen bonds or the coupling of oscillations between bond angles and bond lengths [126].
The choice of the force eld gene rally depends on the properties and systems to be studied, making it an important step in MD simulation. In the AMBER, there are a variety of force elds available. For instance, lipid14 is indicated to study lipid bilayers [127], GLYCAM_06j for carbohydrates [128], ff14SB for proteins [123], OL15 for DNA [129], and OL3 for RNA [130]. The solvent effect of water plays a fundamental role in the protein folding process and in determining the structure and movement of protein molecules . Water force eld models such as the TIP3P [131] and TIP4P [132] have accurately described various physicochemical characteristics of the solvent medium and continue to be widely used in MD simulations with explicit solvents. These different force elds can be used together to describe different systemscomponents. For example, in a simulation of a protein in water as the solvent and with an organic ligand, it is usual to combine the following force elds: ff14SB (for the protein), GAFF [133] (for the organic compound), and TIP3P (for water). In these explicit solvent simulations, much of the computational resources are consumed in calculating the forces acting on water molecules. Methods were developed to implicitly treat the water solvent to reduce the computational cost with the generalized Born (GB) solvation model [134]. An additional energy term related to the proteins solvent-accessible surface areas (SA) can approximate the nonpolar contributions to solvation.
High-frequency motions can limit the size of the integration step in molecular dynamics (MD) simulations, but these motions g enerally do not signicantly impact the systems overall behavior. In molecular systems, particle motions occur on different time scales, inuenced by intermolecular and intramolecular forces. The SHAKE algorithm is the most widely used methodology for maintaining or constraining geometry in MD simulations. This method modies the equations of motion for atoms of interest (such as those involved in chemical bonds with hydrogen atoms) by introducing force constraints that preserve the lengths of these bonds. It is typically applied to constrain hydrogen bonds with other atoms in the system [135].