Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5606_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
02.09.2026
Размер:
21 Мб
Скачать
8 Drug Design in Motion: Concepts and Applications of Classical... 213

4 Molecular Dynamics Analysis

The vast spectrum of MD analysis methods can be divided into three perspectives, which encompass the joint analysis of protein and ligand, then the ligand and the protein analysed separately (Fig. 8.4). Further, each subcategory is described with examples of frequently applied techniques.
4.1 Protein Perspective
4.1.1 Protein Root Mean Square Devi ation (RMSD)
RMSD is a measurement of the deviation of group coordinates from a reference position. The RMSD variation between given groups of atoms, such as Cα atoms or protein backbone, is commonly used in analyses of MD [84]. RMSD provides insights into the deviations in the protein conformation along the MD trajectory, which correlates with the indication of protein stability. Namely, the values are calculated for atomic coordinates at each frame of the MD trajectory, which allows seeing the value variation along the timescale (Fig. 8.5a, blue). One crucial factor in RMSD is the choice of the reference point to which the subsequent RMSD values are compared. This is because the quality and reliability of the reference point determine how accurate the RMSD calculations will be. For instance, the simulations starting frame is commonly designated as the reference point in MD. Further, if the starting frame contains poor contacts, such as clashes or ambiguities in the ligand pose, one will obtain high RMSD values. High values in this regard point back to the poor
Fig. 8.4 Three approaches of Molecular Dynamics trajectory analysis and their sample methodol­ogies. These different techniques are discussed in detail in the next sessions
214 E. Shevchenko et al.
Fig. 8.5 Illustrative showcase examples of RMSD (a), RMSF (b), and secondary structure analysis (c) on a molecular dynamic trajectory.R1–5 indicates simulation replicas 1–5 accordingly. In (c) the replicas are shown with the shades of overlayed grey. In (b) green vertical lines indicate protein residues that are in contact with the ligand
choice of the reference point rather than necessarily indicating ligand instability or poor model quality. Moreover, RMSD indicates the approximate timeline of system equilibration, which can be observed from higher-than-average RMSD values at the beginning of MD. RMSD can be calculated in the majority of drug discovery software, such as GROMACS [46], Maestro Simulation interaction analysis tool, or for complex selections can be scripted with MDAnalysis.analysis.rms [85] mod­ule for Python.
4.1.2 Protein Root Mean Square Fluctuation (RMSF)
RMSF is a measurement of deviation from a reference position of a selected group of atoms that occurs along the MD trajectory and is averaged over the total amount of atoms (Fig. 8.5b)[86]. Thus, the RMSF values indicate which structural elements of a protein uctuate most compared to their reference positions. An illustrative example of such a structural element is the activation loop in protein kinases. Its exible behaviour is often evident in MD trajectories and can be identied through RMSF calculations (Fig. 8.5b, activation loop region highlighted with yellow). Conceptually, uctuations of RMSF values in a particular part of the protein align with B-factor values from the experimental X-ray structures, allowing direct com­parison between the computational and empirical data. Generally, RMSF and RMSD calculations are bound together as an essential basis for trajectory analysis; therefore, they can be calculated within the same software.
Despite using high-quality and resolution crystals, one has to bear in mind that diffraction data is not homogeneous across the system and, specically in the ligand binding pocket, it is important to verify whether certain atoms/bonds are correctly assigned and supported by proper electron densities. When working with
8 Drug Design in Motion: Concepts and Applications of Classical... 215
low-available datas projects, where structural information is sparse, this may serve as a warning that lowly supported and dened regions will lead to articially highly exible in the simulations.
4.1.3 Radius of Gyration (R
)
g
Since its original postulation on crystal structures, the radius of gyration determines the protein structure compactness. The lowest radius of gyration and, accordingly, the tightest packing would correlate with a higher number of intra-molecular protein contacts and could be used as a folding measurement [87]. The R
can be used as a
g
numerical statistical indicator of large conformational change along the protein trajectory (specially in bre-like and semi-exible polymer proteins), but also to monitor the protein stability as a quality control, for globular proteins.
4.1.4 Protein Secondary Structure Analysis
The initial classication of the secondary structure was based on distinct hydrogen bonding modes. It encompassed three elements: helix, strand, and coil [88], further, this classication was extended to eight states: α-helix, 3
helices, π-helix, β-strand,
10
β-bridge, β-turn, bend, and loop or others [89]. As the physical process of protein folding is tightened up with the various biological events and processes [90], deriving seconda ry structure information from MD trajectories can provide insights into the structural behaviour of the protein (Fig. 8.5c). Although to date, the possibility of computational prediction is limited to helix, strand, or loop (coil), the results can highlight valuable changes in the investigated structure correlated with the impact of the ligand, mutation or normal protein function. For instance, a notable shift in the secondary structure composition of a specic helix could be observed when the protein interacts with an agonist (Fig. 8.5c, left chart; the helix maintains a helical structure for around 50% of the simulation time across 5 replicas) compared to antagonist (Fig. 8.5c, right chart; the helix keeps a helical structure for about 95% of the simulation time across 5 replicas). This observation can provide a hint for further in-depth system analysis. Computational results of the secondary structure prediction can be further supported with circular dichroism (CD), which provides an experimental evaluation of proteinssecondary structure. Up to this point, secondary structure calculations along MD trajectories are somewhat rare and frequently require custom coding for a system.
4.1.5 Principal component Analysis (PCA)
PCA is a prominent approach for evaluating large datasets with several dimensions or variables per observation. It enhances the datas readabi lity, while maintaining the most information, besides facilitating visualisation. In other words, PCA retains the
216 E. Shevchenko et al.
greatest amount of statistically signicant data within a given dataset that results in dimensionality while reducing the noisy data that causes valuable statistical trends to be overseen. This is conducted by the linear translation of the dataset to a new coordinate system, which produces a two- or three-dimensional plot with data points that allow deriving the clusters of nearby data points visually. The number of principle components chosen for the visualisation is the main difference between 2D and 3D PCA. The construction of principle components in PCA aims to capture the most variety in the dataset: PC1 depicts the greatest variance, PC2 describes the next-greatest variety, and so forth. Most of the variation is usually captured by the rst two to three PCAs, and the remaining ones can be ignored without losing valuable data. The linear translation (linear reduction) is a process of nding new uncorrelated variables that represent linear functions of the initial set and sequen­tially optimising the variance [91]. In addition, PCAs are widely used in a variety of disciplines where large datasets must be statistically analysed. Among these disci­plines are genomics, astrophysics, machine learning, and data science.
Moreover, PCA is widely used to analyse MD trajectories, where it applies the same principle: a linear transformation that diagonalises the covariance matrix and eliminates the spontaneous linear correlations between the coordinates [92, 93]. PCA results can be utilised for the free energy landscape construction, which describes the metastable conformational states and the transition states [94]. These states contrib­ute to understanding functionally relevant motions within the investigated system, which can be hard to derive by visual examination due to numerous other motions simultaneously occurring within the MD trajectory. In other words, PCA’s applica- tion to MD trajectory is instrumental in identifying dominant protein motions, such as hinge-like movements or domain rotations. By projecting MD frames onto eigenvectors corresponding to principal components, one can visualise how protein conformations evolve over time and uncover correlations among atoms. The PCA analysis can be conducted, for instance, with GROMACS [46] covariance tool (gmx covar and gmx anaeig scripts).
While PCA provides quantitative analysis, direct visual observation remains valuable for comprehending protein dynamics. The input selection for PCA impacts the scope and specicity of the results and can range from all atomic coordinates to Cα atoms or particular residues. While using all atomic coordinates provides a comprehensive view of global motions and conformational changes, this will incur higher computational costs, due to the curse of dimensionality. While it captures global motions, interpreting such high-dimensional data might require advanced visualisation and analysis techniques. Therefore, to gain insights into the overall protein dynamics, one may consider selecting alpha carbon atoms as an input for PCA analysis. This strategy simplies the analysis while capturing essential back­bone motions, for instance, effectively highlighting hinge-like movements and other signicant conformational changes, while disregarding a ne representation of side­chain dynamics.
Meanwhile, focusing on specic residues for PCA narrows down the analysis to regions of interest. This approach can well capture local dynamics, bearing in mind that global motions can still be inuenced, to some extent, by these regions, even if
8 Drug Design in Motion: Concepts and Applications of Classical... 217
they are not directly analysed. Thus, this choice offers insi ghts into the behaviour of selected regions, the choice of which is usually connected with a portion of specic protein machinery. Examples of such regions can be α-C helix in protein kinases, switch-I and II in RAS proteins, or transmem brane helixes 5– 6 in class A GPCRs.
PCA can be used to investigate variations in the dynamic behaviour across different systems. For instance, our work illustrates the use of PCA on different simulated kinase states (multimerisation and phosphorylation) to unravel the impact of different phosphorylation states on kinase behaviour [95]. As a result, in PCA score plots each system showcased a distinct trajectory, highlighting the impact of phosphorylation on the kinase conformational landscape. The phosphorylation states exhibited a direct correlation with the proles observed in the principal components . Further, complementing the quantitative PCA results with a visual inspection of the simulation trajectories revealed system-specic shifts in the activation segment, distinctly deviating from the crystal conformation. Therefore, generating the basis for future protein analysis.
4.1.6 Markov State Modelling
Markov State Modelling (MSM) plays a signicant role in the modelling and interpretation of MD simulation data, while it allows pinpointing of statistically signicant events in the MD trajectories [96]. MSM is a robust framework for analysing long-scale MDs, where it becomes more challenging to visually observe changes in protein dynam ics and substructural geometrical variations.
MSM stands for a master equation framework, indicating that the systems complete dynamics can be highlighted using just the MSM. The master equation formalism has also been applied in a wide range of scientic disciplines [97]. By denition, MSM is a nxnsquare matrix (often referred to as transition probability matrix), in which the n states represent the total number of congurations that the system may exist in [98]. The dynamical change of the system can be observed by dening the states at time point s separated by the lag of time (τ). This separation of the states ensures the lag of time to be a Markov process [98]. The term Markov processrefers to a random process in which, given the present, the future is independent of the past [96]. In terms of MD, the Markov process means that the probability of the system transition from one state to another does not depend on where the system was before this initial state.
The obstacles in building an MSM may be divided into two main categories: the denition of the states in a kinetically relevant order and the effective employment of the state decomposition for building an efcient transition matrix [99]. Additionally, a bias-variance problem related to the choice of the number of states is a dilemma with MSMs. When utilising a limited number of states, one knows analytically that the predicted value of the relaxation timescales is less than the real value and that this bias gets less as the number of states increases. On the other hand, with a given
218 E. Shevchenko et al.
dataset, as the number of states rises, the statistical error in the MSM grows [99]. Although algorithms have been developed to deal with this problem, to date, there are no algorithms in the literature that automatically and effectively balance these conicting sources of errors [100].
As in application of PCA on MD trajectory, MSM requires careful consideration of the input data. There are three common input choices for MSM calculation: atomic coordinates, alpha carbon atoms, and functional groups. Using atomic coordinates as input data offers a potentially detailed review of protein dynamics. However, this level of granularity and complexion can lead to computational challenges due to the sheer volume of data. Alternatively, selecting alpha carbon atoms simplies the input data while retaining backbone motions. This approach is particularly useful for capturing large-scale conformational changes while reducing computational complexity. Another strategy involves focusing on specic functional groups or key residues. This allows one to focus the calculation on the role of these groups in driving conformational transitions. Reg ardless of the input choice, constructing an MSM involves estimating transition probabilities between different conformational states. These probabilities help to understand the kinetics of state-to­state transitions and ha ve a glance into the proteins conformational landscape.
MSMs have been applied effectively to various molecular dynamics studies and beyond. For instance, the study by Meng et al. [101] generated Markov state models for the inhibited Abl Tyrosine Kinase [101]. Through extensive simulations totalling around 400 μs, achieved through multiple runs of 400–750 ns each, they revealed novel conformations previously not disclosed by X-ray crystallography.
In a similar vein, the study by Maltarollo et al. [102] demonstrated the application of MSM to explore dynamic interactions and collaborative events withi n protein binding sites for the FabI protein [102]. By utilising replicas of apo structures with varying timescales and starting conformations, they generated a total concatenated trajectory of 60 μs. Analysing this trajectory using MSMs revealed notable aspects such as the most probable states, relative free energies, and key structural features. This analysis underscored the dynamic interplay and cooperation between binding sites, elaborating their functional dynamics and sequence of geometrical reorganisations. One has to bear in mind that even this long timescale might not be sufcient to capture all the large conformational changes, and their work mostly focused on describing C-terminal and interface conformational changes, where larger full-domain rearrangements would require even longer timescales or enhanced sampling. Interestingly, the process of MSM generation can highlight under sampled zones of the phase space, suggesting where to focus further simulation, adapting the sampling.
In a separate study, Rouxs group demonstrated the utility of MSMs in the study of tyrosine kinases [103, 104]. They contrasted models from apostructure simula­tions with those from a diverse ensemble of initial structures, namely DFG-in, DFG-out, and apostructure states. Notably, even apo-kinase simulations unveiled novel structures in uncharacterised kinases and accurately restored known 3D inhibited-like conformations. This approach, grounded in comprehensive sampling and equilibrium probabilities, expands the utility of MSMs beyond traditional
8 Drug Design in Motion: Concepts and Applications of Classical... 219
methods. As an additional conclu sion, MSM can complement virtual screening strategies. By identifying potentially druggable pockets through Markov states, MSMs enable the detection of sites that might be overlooked by conventional methods, elevating the potency and selectivity of screened target compo unds.
4.1.7 Distance Calculations
Distance analysis is a powerful tool that enables the exploration of changes in relevant distances within a system across the course of MD simulation. This tech­nique involves measuring the distances between specic pairs of atoms, residues, or structural elements within a biomolecular system. For instance, these selections could involve comparing the distances between two atoms (such as Cα atoms),
Fig. 8.6 Example of distance calculation applied on MD trajectory to track various events within protein structure (a) and ways of its statistical representation (b). AV average, SD standard deviation. Statistical boxplot representation of angle values, derived from MD trajectory. Boxes display the quartiles of the dataset (25–75%) and whiskers the rest of the data within 1.5 times the interquartile range (IQR). (Adapted from Shevchenko et al. [95])
220 E. Shevchenko et al.
two residues, or even larger structural features like helices, chains, or lobes within a protein (F ig. 8.6a)[95, 102, 105]. Such analysis provides insights into the dynamic behaviour of various components of the system.
Distance analysis can be employed to understand the dynamic behaviour of multimeric assemblies and the interdependence of their constituent monomeric units across different systems (Fig. 8.5a)[106]. Selecting several intervals for calculating distance can provide comprehensive data to track the co-dependent movements of monomeric substructures (Fig. 8.6a,d1–d4). Another useful applica- tion of distance analysis is in understanding the impact of mutations. For instance, d5 on Fig. 8.6a describes the change of N-lobe composition upon mutation.
It is important to note that distance analysis generates lots of data, which can rapidly accumulate with longer simulation times. To extract meaningful trends from this data and avoid delving into the noise, statistical met hods are often employed to explore and visualise the information (Fig. 8.6b). Common approaches include calculating averages and standard deviations to provide a basic numerical assess­ment of dynamic processes (Fig. 8.6b, upper panel). However, for a more detailed examination, other strategies like box plots can be utilised to demonstrate the distribution of motions across various systems (Fig. 8.6b, box plots highlighted in yellow, and blue showcase the difference between wild type and mutant protein).
When tracking distance changes over time, a straightforward plotting of MD output can some times result in a multitude of peaks, potentially obscuring the primary statistical trends. In such cases, employing a moving average (MA) technique can prove benecial. MA, frequently used in analysing time-series data such as nancial trends, calculates averages over different segments of the complete dataset. This technique effectively captures short-term uctuations while focusing on the broader trendsaligning perfectly with the object ives of MD trajectory analysis.
There are various types of moving averages, depending on various factors such as data structurefor our purpose we employ simple moving average (SMA), which uses a sliding window to take the average over a set number of time periods. Once applied and plotted, one can observe the main trend in the distance change. As demonstrated in Fig. 8.6b, the SMA applied along the timescale shows the gradual opening of specic protein substructures (indicated by the yellow line), suggesting studying the origin of this diverging movement.
In practical terms, with GROMACS [107] gmx distance script offers a broad spectrum of options for conducting distance calculations. This script permits the denition of reference positions based on the centre of mass or centre of geometry for various selections, including atoms, residues, or custom-dened entities. This variety allows extracting relevant results from dynamically changing reference selection, which is highly important for the MD trajector y analysis.
8 Drug Design in Motion: Concepts and Applications of Classical... 221
Fig. 8.7 Example of angle calculation applied on MD trajectory to track movement within a protein structure. (a) selection of residue intervals for angle construction. Selected residue intervals represent protein substructures, which movements should be tracked (B stable apex point (A centre of geometry (cog). Cog is visualised on the plane projection with semi-transparent ovals of the same colours as displayed residue intervals. (c) Statistical boxplot representation of angle values, derived from MD trajectory. Boxes display the quartiles of the dataset (25–75%) and whiskers the rest of the data within 1.5 times of the interquartile range (IQR). The reference angle value is highlighted in yellow in (b) and (c) to underline the dynamic change one can observe, when angle calculation is applied on MD trajectory. (Adapted from Shevchenko et al. [95])
). (b) Plane projection of selected residue intervals. Apex points dened by
1–A2
and C1–C2) with a
1–B2
4.1.8 Angle and Plane Calculations
Angle or plane calculations are benecial to support visually perceived motions within the MD trajectory with the numerical values. There are various examples of possible practical applications of this method. For instance, angle calculations can be utilised to dene the protruding outward movement of αC-helix in kinases quanti­tatively, justify the opening or closure of Switch-II in RAS proteins or assist the hit­to-lead identication by dening the degree of pocket opening in MD simulations [108, 109].
The reference positions of the vertex and apex points can be dened using several methods, as seen with the distance calculationsnamely, the centre of mass or geometry for an atom, residue, or substructure. For instance, Fig. 8.7a shows the selection of residue intervals (A
1–A2,B1–B2,
and C1–C2) for further angle calcula­tion. Next, these intervals are projected on a plane, where the apex points are calculated as centres of geometry for each interval (Fig. 8.7b). The validation of the residues chosen for the computation of the angle value is an optional step that allows to verify if the calculated value is biased or whether a particular selection has no impact on the average angle value along the trajectories. This is accomplished by selecting the neighbouring to the reference selection residues and repeating the
222 E. Shevchenko et al.
calculation. The average values from the reference and validation calculations should appear to be in the same range. Another way to produce reliable angle calculation between moving helices or domains is to choose the stable residue interval for the apex. On Fig. 8.7a the hinge region (A
)defines the apex of
1–A2
the angle, as one of the most rigid regions of a protein kinase [95]. Moreover, the MD-derived RMSD values can be used to support the interval stability. Python is the primary tool for carrying out these computations for an MD trajectory since they are highly dependent on the investigated system. Additionally, the classes required to access data in the MD trajectories are provided by the MDAnalysis package for Python [110, 111]. Further, to represent the most valuable data and reduce the noise, the derived angle values, similarly to distance analysis, are best represented with statistical visualisation methods, such as box plots (Fig. 8.7c) or simple moving average (Fig. 8.6b).
4.2 Protein–Ligand Perspective
4.2.1 Protein–Ligands Interaction Pattern Determination
The analysis of the interaction network that the ligand is forming with the protein is the fundamental part of the simulation analysis as the non-covalent interactions, namely hydrogen bonds and hydrophobic interactions, are the core of the ligand­protein binding [112]. Unlike docking, MD-derived interaction analysis gives not a single-point calculation but the statistical range of the simulation time when the ligand is in contact with a particular protein residue. Moreover, the interaction can be plotted as a function of time. The derived know ledge can be further applied for SAR analysis and hit-to-lead optimisation. Finally, distinct software may have different denitions of distance and angle thresholds that characterise specic interactions, particularly in the case of hydrophobicity.
4.2.2 Distances and Ligand-Induced Geometry Rearrangements
As in protein perspective, distance analysis comes in hand as a tool for understand­ing ligand-induced geometry rearrangements within proteins [113]. By quantifying distances between key residues for protein functioning and ligands, it is possible to unlock additional insights into the dynamic interactions that govern binding events [102]. These distance uctuations not only reveal the stability and exibility of protein–ligand complexes but also highlight allosteric effects and potential druggable sites that can be further utilised in the drug design process [114]. For instance, in the context of G protein-coupled receptors, distance analyses help to spot communication between ligands and the receptors transmembrane helices, which