Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5319_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
29.08.2026
Размер:
92 Мб
Скачать
Chapter  • Three-Dimensional Structure of Biomolecules
. Fig. 14.19 The DNA base pairs
of cytosine(C) with guanine(G) and thymine(T) with adenine(A) on the individual steps are formed by complementary hydrogen bonds. Each base carries asugar–phos­phate group that is coupled with the polymer chain. It affords adou­ble-helical construction with aminor (green) and major (yellow) groove (cf. . Fig.14.18). If viewed from parallel to the steps, four groups can be seen in the major groove that possess either hydrogen-bond donor (blue), hydrogen-bond acceptor (red), or hydrophobic properties (gray). Three such groups are aligned in the minor groove. If an attempt is made to read the interaction pattern from this side, an GC or CG pair and an AT or TA pair will be recognized as identical. Here the orientation of the interaction pattern cannot be distin­guished. In the major groove, on the other hand, the pattern of exposed interaction is unambiguous. There­fore, proteins read information about the DNA from the major groove
14
magnesium ions. Because of its important role in trans­mitting genetic information, several important drugs act on DNA. Two examples are briey mentioned here. Cisplatin 14.4 is areactive metal complex that can co­valently bind to the nitrogen atoms of two nucleobases on two adjacent steps of the DNA by replacing both chlorine substituents (. Fig.14.20). This cross-linking distorts the DNA in such away that the sequence infor­mation is no longer readable. Cisplatin and analogues such as carboplatin are used as potent chemotherapeutic agents in cancer therapy. Daunorubicin 14.5 is arepre­sentative with aslightly different mechanism of action, but it also prevents the reading of DNA base pairs. The planar molecular moiety of 14.5 slips largely between two adjacent base pairs, causing astructural distortion of the DNA (intercalation). This intravenously admin­istered cytostatic drug is used as acombination ther­apy for the treatment of acute leukemias. Many natural products also use this intercalation mechanism for their antibacterial activity spectrum. Other pharmaceutical research approaches attempt to use segments of DNA itself for therapy. Such modied oligonucleotide thera­peutics are discussed in Sect.32.4.
Every third bond in the polymer chain of aprotein is
-
an amide bond. It is the fundamental building block in the protein backbone and the mutual spatial ar-
rangement of the sequential planar amide bonds de­termines the overall architecture of aprotein.
Typical arrangements involving the amide NH and
-
CO groups in hydrogen bonds lead to α-helical and
β
-pleated sheet structures. Reversal of the polymer chain in space is achieved in turns that can adopt avariety of distinct geometries.
Helices, sheets, and turns, the secondary structural
-
elements, assemble into motifs and domains to form the tertiary and quaternary structure of proteins.
The function of aprotein is not necessarily coupled
-
to aparticular folding pattern; however, the catalytic and ligand-functional sites within afolding class are found at the same position.
Nature separates fold-stabilizing residues from func-
-
tion-carrying amino acids to keep the individual steps of the dual optimization problem separated.
Proteases recognize peptide sequences specically via
-
the binding in well-tailored pockets on both sides of the cleavage site.
Peptide libraries with an attached photometric or
-
uorescent label that can be cleaved by the protease reaction help to elucidate the substrate prole of dif­ferent proteases.
Structural arrangements of molecular portions found
-
in multiple crystal structures can be arranged sequen­tially in akinematic order to provide an idea of ady­namic process.
The spatial arrangement of amino acid residues ex-
-
erting aparticular chemical transformation is highly

Bibliography and further reading


. Fig. 14.20 Crystal structure of an oligomeric DNA segment after
areaction with cisplatin 14.4 (left) or intercalation with daunorubi­cin 14.5 (right). In both cases, the DNA molecule is severely distorted and the genetic information on the DNA cannot be read for cell di­vision. Cisplatin reacts with the nitrogen atoms of two nucleobases (here guanine) of the DNA on neighboring steps with substitution of both chlorine atoms in achemical reaction. With its planar tet­racyclic ring system, daunorubicin intercalates between two neigh-
conserved and can reside on protein architectures with similar geometry that are constructed from de­viating folds.
The DNA molecule encodes our genetic information
-
and forms adouble helix of two opposing strands wrapped like ahandrail by sugar–phosphate polymer chains. They carry the complementary base pairs on the stacked steps of the double helix. The individual base pairs form atypical H-bonding pattern in the center of the helix.
The typical H-bonding pattern of bases allows each
-
single strand of DNA to be complementary to the second strand. When DNA is amplied, the dou­ble-stranded DNA can be completed from the sin­gle-stranded DNA. Asmall and alarge groove are formed between the sugar–phosphate backbone. In the large groove, the coding of the base pairs on each step can be read from the side. By complexing bases on adjacent steps with metal ions or by intercalating planar agents between these steps, the DNA is geo­metrically distorted which will prevent reading of the
boring base pairs by spreading the DNA along the helix axis. The compound’s amino sugar accommodates in the DNA minor groove. (7 https://sn.pub/PMGNV7)
genetic information, which consequently inhibits cell growth.
General Literature
C. Branden and J. Tooze, Introduction to Protein Structure, 2nd Ed.,
Garland Publ. Inc., New York (1999)
G. A. Jeffrey and W. Saenger, Hydrogen Bonding in Biological Struc-
tures, Springer Verlag, Berlin (1991)
H. B. Bürgi and J. D. Dunitz (Eds.), Structure Correlation, Vol. 1 and
2, VCH, Weinheim (1994)
G. E. Schulz and R. H. Schirmer, Principles of Protein Structure,
Springer Edition, New York (1978)
Special Literature
I. Schechter, A. P. Berger etal., On the size of the active site in prote-
ases, Biochem. Biophys. Res. Commun., 27,157–162 (1967)
P. I. Lario and A. Vrielink, Atomic Resolution Density Maps Reveal
Secondary Structure Dependent Differences in Electronic Distri­bution, J. Am. Chem. Soc., 125, 12787–12794 (2003)
Chapter  • Three-Dimensional Structure of Biomolecules
F. A. Allen, O. Kennard and R. Taylor, Systematic Analysis of Struc-
tural Data as a Research Technique in Organic Chemistry, Acc. Chem. Res., 16, 146–153 (1983)
C. A. Orengo, D. T. Jones and J. M. Thornton, Protein Superfamilies
and Domain Superfolds, Nature, 372, 631–634 (1994)
K. Vyas, H. Monahar and K. Venkatesan, Thermally Induced O to N
Acyl Migration in Salicylamides. Thermal Motion Analysis of the Reactants, J. Phys. Chem., 94, 6069–6073 (1990)
G. Klebe, The Use of Composite Crystal-Field Environments in Mo-
lecular Recognition and the De Novo Design of Protein Ligands, J. Mol. Biol., 237, 212–235 (1994)
W. J. L. Wood, A. W. Patterson, H. Tsuruoka, R. K. Jain, and J. A.
Ellman, Substrate Activity Screening: A Fragment-Based Method for the Rapid Identication of Nonpeptidic Protease Inhibitors, J. Am. Chem. Soc., 127, 15521–15527 (2005)
O. Koch and G. Klebe, Turns Revisited: A Uniform and Comprehen-
sive Classication of Normal, Open, and Reverse Turn Families Minimizing Unassigned Random Chain Portions, Proteins: Struct., Funct. Bioinform., 74, 353–367 (2008)
PDB Database: http://www.rcsb.org/pdb/home/home.do (Last accessed
Nov. 17, 2024)
CSD Database: http://www.ccdc.cam.ac.uk/products/csd/ (Last ac-
cessed Nov. 17, 2024)
IsoStar Database: https://www.ccdc.cam.ac.uk/solutions/software/iso-
star/ (Last accessed Nov. 17, 2024)
Mogul, assessment of molecular conformations: https://www.ccdc.cam.
ac.uk/solutions/software/mogul/ (Last accessed Nov. 17, 2024)
14

Molecular Modeling

Contents
15.1 3D Structural Models as Well-Established Tools in Chemistry – 234
15.2 Strategies in Molecular Modeling – 234
15.3 Knowledge-Based Approaches – 235
15.4 Force Field Methods – 236
15.5 Quantum Chemical Methods – 237
15.6 Computing and Analyzing Molecular Properties – 239
15.7 Molecular Dynamics: Simulation of Molecular Motion – 239


15.8 Dynamics of aFlexible Protein in Water – 242
15.9 Model and Simulation: Where Are the Dierences? – 243
15.10 Synopsis – 244
Bibliography and further reading – 245
© The Author(s), under exclusive license to Springer-Verlag GmbH, DE, part of Springer Nature 2024 G. Klebe, Drug Design, https://doi.org/10.1007/978-3-662-68998-1_15
Chapter  • Molecular Modeling
15
Molecules are most commonly communicated in chem­istry as two-dimensional molecular representations that show their topology. This formalism is well established and has proven enormously fruitful. The ability of achemist to quickly grasp and intellectually process such structures should not be underestimated. However, the notation has its limitations. In particular, the three-di­mensional shape of amolecule is not immediately ap­parent from the chemical formula. However, geometry is of great importance for the physical, chemical, and biological properties of drugs and thus for drug design. Therefore, 3D structure determination (Chap.13) is of particular importance. Whenever possible, the experi­mentally determined 3D structure of the compound and the target protein is used to explain the mode of action and the structure–activity relationship. Notwithstanding this, there is often the problem that these structures are not always available. In these cases, the explanation of experimental results is limited to the structural consider­ation of generated models. In some cases where it is even not possible to determine an experimental structure, an attempt is made to calculate amodel. Often the proper­ties of molecules can be estimated more easily by calcu­lations than by time-consuming measurements. In many cases, molecules can adopt several stable geometries. Through extensive simulations, it is possible to calculate these geometries and derive ideas about the dynamic transformations between different states. Computational models are, therefore, acrucial link to correlate experi­mental data with design ideas. Computations also allow scanning the structural property space around agiven ge­ometry and assessing its relevance for agiven molecular property. The computational methods available for drug design, the assumptions and approximations on which they are based, and what can be deduced from them for the development of anovel drug will be the subject of this chapter.
15.1 3D Structural Models as Well-
Established Tools in Chemistry
Three-dimensional structure models have been used since JacobusH. van’t Hoff and Joseph Le Bel. Emil Fischer reported in his book Aus meinem Leben about avacation in Italy:
“In the previous winter 1890/91 Iwas busy with the
»
task of clarifying the conguration of sugar, without entirely achieving my goal. Then the thought came to me in Bordighera that the decision about the congu­ration of pentose has to do with its relation to trioxy­glutaric acid. Unfortunately for lack of amodel Icould not tell to what extent such acids are possible according to theory and Itherefore posed the question to Baeyer. He picks up such things with great enthusiasm, and di-
rectly constructed carbon atoms from balls of bread and toothpicks. But after many attempts he gave the cause up, ostensibly because it was too hard. Later in Würzburg after considering good models at length, Imanaged to nd the conclusive solution.”
Linus Pauling was the rst to propose the α-helix as asecondary structure in proteins.
“The key to Linus’s success was his reliance on the simple
»
laws of structural chemistry. The α-helix had not been found by only staring at X-ray pictures. The essential trick, instead, was to ask which atoms like to sit next to each other. In place of pencil and paper, the main work­ing tools for this work were aset of molecular models supercially resembling the toys of preschool children.”
These were the words Nobel Laureate James Watson used to describe Pauling’s approach in his book The Dou- ble Helix. Pauling’s success was also based on asound knowledge of theoretical chemistry. For example, Paul­ing knew that an amide bond is rigid and at, whereas his competitors, William Bragg, Max Perutz, and John Kendrew, mistakenly believed that it was exible. James Watson and Francis Crick followed the same course as Pauling in their search for the structure of DNA:
“We could thus see no reason why we should not solve the
»
DNA problem in the same way [as Pauling]. All we had to do was build aset of molecular models and begin to play—with luck the structure would be aspiral.”
Based on this background, the achievement of Wat­son and Crick seems even more impressive. They were awarded the Nobel Prize in 1962 for the elucidation of the double-helix structure of DNA. This example should underscore the importance of models in science. To end with aword from Francis Crick: “Agood model is worth its weight in gold.”

15.2 Strategies in Molecular Modeling

In contrast to the 1950s and 1960s, today’s computers have impressive graphical and computational capabili­ties. Correspondingly, programs for working with mo­lecular models are available. The new eld of molecular modeling has emerged. This term encompasses the rep­resentation and manipulation of realistic three-dimen­sional molecular structures along with the calculation of their physicochemical properties. The main methods used in molecular modeling are summarized in . Table15.1.
In principle, molecular modeling can be approached from two directions. One approach is to extrapolate the geometry and physicochemical properties from known experimental data to the structure under investigation.
. • Knowledge-Based Approaches


The other approach is to try to obtain as accurate acom­putational prediction as possible by starting from rst principles. Quantum chemical methods and force eld calculations are part of this strategy. In practice, both approaches are used in parallel and are increasingly cou­pled. If relevant experimentally determined structures are available, it would be stupid not to use them for model building. On the other hand, quantum chemical and mo­lecular mechanical approaches are widely applicable and provide reliable results. Most importantly, models can be changed gradually, one parameter at atime. Experiments usually do not allow this. Instead, experimental investi­gations typically change several parameters at the same time, often unavoidably.
The construction of astructural model is achieved
in three steps:
Generation of astarting model,
-
Optimization and analysis, and
-
Work with the model.
-
It is advisable to stay as close as possible to experimen­tal structures when creating the initial starting model.
. Table 15.1 Overview of the most important molecular
modeling approaches in pharmaceutical research
Technique Objective
Interactive computer graphics
Modeling small molecules
Comparing molecules
Protein model­ing
Modeling of protein–ligand interactions
Ligand design Searches in 3D databases
Display of 3D structures
3D Structure generation (CONCORD, CORINA)
Molecular mechanics—force elds
Molecular dynamics
Quantum mechanical techniques
Conformational analysis
Calculation of physicochemical properties
Superimposition of molecules according to their similarity
Volume comparisons
3D-QSAR (e.g., CoMFA methods)
Sequence comparisons
Protein homology modeling
Protein-folding simulations
Binding constant calculations
Ligand docking
Structure-based ligand design
de novo design
Virtual screening
This can be done by consulting the crystal structure of acompound. The Cambridge Crystallographic Data­base, which stores experimentally determined crystal structures of small molecules, is searched and the geom­etry of the resulting hits that most closely resembles the query molecule is used. The next step is to optimize the molecule using aforce eld calculation.
There are also off-the-shelf programs available for generating starting models, which translate a2D struc­tural formula into a3D spatial structure according to the principle of amolecular construction kit. These “elec- tronic molecular construction kits” store lists of bond lengths and angles, as well as preferred fragment geom­etries, and build molecules according to asophisticated set of rules. In fractions of asecond, they determine the 3D spatial structure for the 2D structural formula. The rst programs in this eld were CONCORD by Rob­ert Pearlman in Austin, Texas, USA, and CORINA by Johann Gasteiger and Jens Sadowski at the University of Erlangen, Germany. Both programs can be used to generate 3D structures of small molecules. In addition, CCDC’s program Mogul provides a comprehensive source of information for comparing and validating acomputed molecular model with experimental data. However, these programs cannot build the 3D structure of aprotein. More sophisticated techniques are required to predict the spatial structure of aprotein from its se­quence (Sect.20.6).

15.3 Knowledge-Based Approaches

Perhaps the most widely used techniques in molecular modeling are knowledge-based approaches. They at­tempt to use the vast amount of knowledge accumulated from experimentally determined molecular structures, crystal packing, protein structures, protein sequences, and structure–activity relationships of protein–ligand complexes, etc. to efciently solve a given problem. Basically, the computer program mimics the approach that a conscientious scientist would also take. First, as much experimental data as possible are collected and analyzed. Important sources of information are the Cambridge Crystallographic Database, which contains more than 1.25 million crystal structures of small mole­cules, and the Protein Databank (PDB), which contains more than 227,000 protein and DNA structures. Phys­icochemical properties are also available in databases. For example, the Beilstein database of nearly 10million chemical structures contains pKa values for more than 20,000 compounds. For amino acid residues in proteins, the PKAD database has been developed (7 http://com-
pbio.clemson.edu/PKAD-2/). The challenge lies in the
extraction of the necessary data for the question at hand from the enormous plethora of electronically available information. Furthermore, it must be considered that the
15
12
ij
6
ij
6 ij
Chapter  • Molecular Modeling
. Fig. 15.1 E is the total energy of amolecule or complex of mole-
cules. It is made up of several contributions. The rst term describes the energy change when achemical bond is stretched or compressed. In this example, it describes the so-called harmonic potential with the force constant Kb and the equilibrium bond length b0 as aparameter. The second term describes the energy as afunction of the bond an­gleΘ. Again, the harmonic potential is used with the force constants
KΘ and the equilibrium constant Θ0. The third contribution describes
the change in energy when the dihedral angle is changed, and the last term represents noncovalent interactions. The sum of three terms is used for this last contribution. The rst term tive and increases rapidly with decreasing distance. It describes the re­pulsion between atoms that come too close together. The contribution
Cij=r
from distance rij, but not as fast as the repulsive term. It describes attractive interactions, which are also called dispersion interactions. Other at­tractive interactions exist between polar molecules which are also pro­portional to
. Fig.18.5). The last term q
tions based on Coulomb’s law, which works with apoint charge model. The dielectric constant isε. The noncovalent contribution to the total energy, without the electrostatic term, is called van der Waals energy
is always negative and tends to zero with increasing
r
(for adescription of the potentials see Sect.18.12,
/εrij describes the electrostatic interac-
iqj
=r
ij
is always posi-
data come from different sources and could be partially erroneous or were measured and collected under barely comparable conditions.
The largest growth in electronically available data recently has occurred in the area of DNA sequences. Hundreds of genomes have been sequenced, and new
ones are added weekly. The nearly endless number of se­quences can only be conquered with intelligent searching protocols. Modeling of protein structures is now being conducted on alarge scale with novel knowledge-based approaches using machine learning and articial intelli­gence (Sect.20.6). The generated models are collected in adatabase of computed structure models and contains
more than amillion models (7 https://www.rcsb.org/
news/6304ee57707ecd4f63b3d3db).

15.4 Force Field Methods

Force eld methods, also known as molecular mechanics, are empirical techniques for calculating molecular geom­etries. The goal of aforce eld calculation is to determine an energetically favorable three-dimensional structure of amolecule or complex of molecules. The forces acting between the atoms are described in an analytical form, and appropriate parameters have to be assigned to the different terms. In principle, alarge number of such analytical functions can be imagined. We will consider asimple form that has been used very often. It considers covalent and noncovalent forces. The central idea of mo­lecular mechanics is the assumption that the bond lengths and angles take on values that are close to the standard values in molecules. Steric interactions, this means the repulsion of two atoms that are not directly bonded to each other, can cause some bond lengths and angles not to take their ideal values. These repulsive interactions are also known as van der Waals interactions. In the simplest case, the deviations from the ideal values are described by aparabolic potential (the so-called “harmonic” potential, which applies to the motion of amass on afreely sus­pended spring). However, this form ignores the fact that very strong distortions lead to bond breaking. To better describe such strong distortions in the potential prole, adistance-dependent function with an exponential char­acteristic (e.g., Morse potential) is often used.
In 1946, three terms, van der Waals interaction, bond stretching, and angular deformation, were rst proposed to be sufcient to calculate the structure and energy of molecules. At that time, however, performing such cal­culations was extremely difcult. It was not until the availability of computers increased that molecular me­chanics calculations gained importance. In addition to the three terms originally proposed, atypical force eld in use today contains at least one additional contribution that takes into account rotations about the dihedral angles (. Fig.15.1). In addition, many force elds use terms for electrostatic interactions. To do this, each atom must be assigned apartial charge. The sum of these charges gives the formal charge of the whole molecule. This is usually set to zero for uncharged particles.
Coulomb’s law is used to describe the forces that oc­cur between charges. This law states that the product of interacting charges is inversely related to the square of the distance between them, or, considering the poten tial, inversely related to the distance. The assignment of charges and the correct choice of the dielectric constant are critical to the correct treatment of electrostatic en­ergy contributions. These values are in the denominator of Coulomb’s law and can take values between ε = 80
-
. • Quantum Chemical Methods


for water and ε = 1 for vacuum. This dampens the elec­trostatic interactions in water very rapidly, while in avacuum they tend to reach much further. Choosing the correct dielectric constant for force eld calculations in proteins is very difcult. Many values between ε = 4 and ε = 20 have been tried. The constant is sometimes assumed to be environment dependent, so that larger val­ues are chosen near the surface than for the interior of the protein. The van der Waals interactions are modeled by the Lennard–Jones potential. This interaction has an attractive term falling at arate of 1/r6 and arepulsive term falling at arate of 1/r12 (. Fig.15.1). The result of the combination of these terms is agradient that is very large near the atoms and approaches zero as the distance increases. In between, it passes through apotential en­ergy minimum (. Fig.18.5). In addition to A/r6–C/r12 as the functional form, alternative distance dependencies with other potentials or exponential gradients have been used in force elds.
Aforce eld is derived by calibration to experimental data or to the results of high-level quantum mechanical calculations. The 3D structures of small molecules and force constants derived from infrared and Raman spec­troscopy are used. It is clear that different parameters must be used for asingle bond between two carbon atoms than for adouble bond. Therefore, several different atom types per element are dened in aforce eld. The crystal packing of small organic molecules can be consulted to parameterize nonbonded interactions. Amino acids and many functional groups of active compounds can exist in either aprotonated or deprotonated state depending on the applied pH conditions (so-called titratable groups). The strength of the resulting interactions is strongly dependent on the charge state of the functional groups involved. The acidity or basicity of agiven functional group is determined by its pKa value. This indicates how easily agroup accepts or releases aproton. This property, in turn, depends heavily on the partial charge that the group carries and what other charges are in the immedi­ate vicinity of the group. Thus, the pKa will change when afunctional group is placed in adifferent environment. For example, carboxylic acids become more acidic when placed near apositive charge. On the other hand, their acidic nature will change if apartially negatively charged group is in the vicinity. This effect must be taken into ac­count in areliable force eld calculation. An attempt can be made to predict the protonation state in protein–li­gand complexes with such calculations. The contribution to the energy content of the complex is determined by evaluating all possible combinations of protonated states of titratable groups. In this way, the shift in pKa values of functional groups can be estimated. To better describe the shift of charges in molecules, so-called polarizable force elds can be applied. They adjust the charge distri­bution to the local distances and take into account that charges are displaced or “polarized” in the molecule by
attractive or repulsive effects. Certainly, these force elds give amore realistic picture. However, the disadvantage is that they are computationally much more demanding and require many more parameters to be adjusted.
The importance of water as abinding partner in the formation of protein–ligand complexes was emphasized in Sect.4.6. The formation of aprotein–ligand complex causes achange in the solvation conditions for the mol­ecules involved. This has to be taken into account in the force eld calculations. In order to do this, aforce eld is combined with estimates for the contribution of solva­tion. Approaches such as the MM-PBSA or MM-GBSA methods try to sum up these contributions over the lo­cal environment in asurface-dependent manner. Newer methods such as the Grid Inhomogeneous Solvation Theory (GIST) method can calculate thermodynamic contributions such as solvation enthalpy and entropy from the interaction contributions collected during amo­lecular dynamics simulation (Sect.15.7). These values are then mapped onto agrid for subsequent graphical analysis.
The choice of arelevant initial starting geometry is important for any force eld calculation. A force eld calculation involves energy minimization. If one starts with an energetically unfavorable geometry, the force eld will travel “downhill” to the next local minimum on the multidimensional energy surface (Sect.16.2). If one starts with two different geometries, the structures obtained at the end of the minimization can be different, depending on which local minimum is reached. Many molecules, especially protein–ligand complexes, can adopt many en­ergetically favorable conformations. It is, therefore, rec­ommended to perform multiple force eld calculations starting from different geometries. Molecular dynamics simulations, discussed below, also provide asolution to this problem of getting trapped in local minima.

15.5 Quantum Chemical Methods

In quantum mechanical approaches, the electronic struc­ture of molecules is calculated using the Schrödinger equation. However, its mathematically closed solution is only possible for simple cases such as the hydrogen atom or the molecular ion of hydrogen, +. For molecules with more than one electron, approximate methods must be used to solve the quantum mechanical “many-body problem.” The most commonly used approximation is the Hartree–Fock method. Here the many-body prob- lem is reduced to several single-body problems. The sum of electron–electron interactions within amolecule is replaced by an effective eld, which can be iteratively rened and optimized. This is where the popular name SCF (self-consistent eld) comes from. In this model, each electron “sees,” in addition to the potential of the nuclei, an averaged potential of the remaining electrons.
Chapter  • Molecular Modeling
15
The state of each electron in amolecule is, thus, described by asingle-particle function called the atomic orbital (AO) or, in amolecule, the molecular orbital (MO). The wave function of the entire molecule is the antisymmetric product of the many orbitals considered. The Hartree– Fock equation is then obtained under the condition that the optimally chosen orbitals lead to aminimum overall energy. The main shortcoming of the Hartree–Fock ap­proach, namely the neglect of the electron correlation, can be corrected by more sophisticated methods, which, however, increase the computational time considerably.
Quantum mechanical ab initio calculations allow the calculation of the molecular structure and electron density distribution as well as molecular properties without the assumptions necessary for force eld calculations. In many cases, it is difcult to make apriori predictions based on the hybridization state of the atoms. For example, in the case of amines and sulfonamides, it is often impossible to predict whether the atoms bonded to the nitrogen in these compounds are all in the same plane or whether the nitrogen deviates with apyramidal environment. In aforce eld calculation, this must be specied at the outset of the calculation by assigning which atom type to which atom in amolecule (i.e., in the above case, whether the nitrogen should be in aplanar or pyramidal local geometry). Of course, if the wrong atom type is chosen, the resulting structure will be meaningless. As an advantage, quantum mechanical calculations do not require such assumptions.
Most currently applied force elds use apoint-charge
model to describe the electrostatic interactions. One op-
tion to derive the atomic charges is to calculate the electro­static potential of asmall molecule containing the group of interest using quantum mechanical methods. Aset of partial charges is then assigned to the different nuclei in order to reproduce the quantum mechanical potential as accurately as possible. These charges can then be trans­ferred to force eld calculations for use in alarge system.
Another important application of quantum me­chanical calculations in drug design is the calculation of conformational energies of small molecules to calibrate force elds. The force elds developed for proteins and peptides are based on conformational energies calculated quantum-mechanically for small peptides.
In contrast to force eld methods, quantum mechan­ical techniques are able to take into account the polar­ization of the electron density caused by the inuence of neighboring groups. For example, the amide bond dipoles in an α-helix are all oriented in the same direction, so they add up to asignicant total dipole moment (Sect.14.2). As aconsequence, such large and enhanced dipoles can polarize other groups located at the end of the helix. In fact, induced dipoles are incompletely described by stan­dard force eld methods (see polarizable force elds). This is not aproblem for quantum mechanical methods. An­other important area of application is chemical reactions, for which force elds are hardly parameterized, except for
afew special cases. Here, quantum mechanical methods are the only option for areliable theoretical description.
Quantum mechanical methods are considerably more elaborate than force eld methods. The most accurate meth­ods, which also consume the most computational time, are the so-called ab initio methods. However, these techniques quickly reach their limits for very large systems. Therefore, other less computationally intensive methods have been developed. In these so-called semiempirical methods, cer­tain integrals, the determination of which represents the rate-determining step in ab initio methods, are replaced by adequate approximations that can be computed quickly. The resulting drastic reduction in computational time, at the expense of accuracy, allows the routine application of semiempirical calculations to active molecules and pro­teins. Density functional theory is another faster ab initio technique. In this method, the position-dependent electron density distribution is calculated in the ground state for amany-body system, avoiding the complete solution of the Schrödinger equation for a many-body system. All interesting properties are then derived from the electron density. For large protein–ligand systems, techniques have been developed that treat the interesting regions, such as the binding site or the catalytic reaction center, quantum mechanically. The surrounding regions are approximated with afaster force eld method (QM/MM methods).
Recently, there has been alot of discussion about new methods using articial intelligence and machine learn- ing. This is the attempt to transfer the way humans learn and think to computers. The idea is to give computer algorithms some intelligence. The goal is to use learning algorithms that allow the computer to nd answers and solve problems on its own, without having to develop anew special program for each case. Neural networks are an important tool in this area. In 2024, the Nobel Prize in Physics was jointly awarded to John J. Hopeld (Princeton Univ., New Jersey, USA) and Geoffrey E. Hinton (Univ. of Toronto, Canada) for their fundamen­tal discoveries and inventions that enable machine learn­ing with articial neural networks. Such methods are es­pecially powerful for independent data analysis. Today’s computers have the storage capacity and speed to crunch through vast amounts of data in avery short time and “remember” it all in away that no human brain could capture, correlate, and evaluate in the same amount of time. In the eld of drug development, such methods are not new. Data analysis methods and procedures have always been used to identify correlations in often high-di­mensional data spaces. As the data have become larger and more complex, the algorithms have become better, and the computers have become faster. We will return to these quantitative assessments of structure–property correlations several times in Chaps.18,19, 20. Despite the excitement of the ever-increasing ood of data to be evaluated, it is important to remember that every evalua­tion, and ultimately every insight, is already contained in
. • Molecular Dynamics: Simulation of Molecular Motion


the data. If the data are not reliable, meaning that if they are too “noisy” due to bad and erroneous information, even the best machine learning or articial intelligence algorithm will hardly bring any new insights to light.
But what can articial intelligence contribute to the use of quantum chemistry in drug design? As mentioned above, quantum chemistry is characterized by high ac­curacy. But its computational cost is still prohibitive for many problems. Today, we are increasingly taking the ap­proach of breaking molecules into small building blocks and performing sophisticated quantum chemical calcula­tions on the resulting building blocks to determine their energy and geometry. When alarge dataset of molecules is processed in this way and care is taken to ensure that the building blocks of interest are repeatedly embedded in different molecular environments, it is possible to generate adataset with inherently redundant information. From such adataset, one can then use articial intelligence to infer parameters for agiven force eld. Typically, neural networks are used for this purpose. Force elds obtained in this way represent areal advance. They approach the accuracy of quantum chemical calculations, but can be computed as quickly as typical empirical force elds.
15.6 Computing and Analyzing Molecular
Properties
The result of amolecular mechanics or quantum chem­ical calculation is initially aset of atomic coordinates that dene the three-dimensional shape of the molecule. What can be done with this? An important application of the calculations is the determination of conformational energies: this is the relative energy of one molecular con­formation compared to another (Sect.16.1).
Two other molecular properties can be calculated: the shape and size of amolecule, along with its electronic characteristics. All current graphics programs have sev­eral ways of displaying such properties with the spatial structure of molecules. The most important ones are summarized in . Fig.15.2.
The most commonly used representation is aline or stick representation (Dreiding models), sometimes atoms are displayed as small spheres. Usually acolor coding is used to represent the atoms: nitrogen is blue, oxygen is red, sulfur is yellow, uorine is turquoise, chlorine is green, bromine is brown, and iodine is purple. Hydrogen atoms are shown in white, but are usually omitted for clarity. Carbon atoms are usually shown in black or gray. In most gures in this book, carbon atoms that belong to the protein are shown in orange, and carbon atoms that belong to the ligand are shown in gray, but sometimes in adifferent color is applied to distinguish them in differ­ent molecules. Another display option is the space-lling model, which shows van der Waals surfaces. In this repre­sentation, each atomic nucleus is represented by asphere
whose size corresponds to the van der Waals radius. Val­ues for these radii are derived from crystal packing or from very accurate ab initio calculations. Such represen­tations are also known as CPK models (named after the scientists Corey, Pauling, and Koltun).
There are other ways to represent surfaces (. Fig.15.3). The solvent-accessible surface has proven particularly valuable for proteins. The most commonly used way to display a protein in this book is atranspar­ent opaque white surface. The van der Waals surfaces in
. Fig.15.3 (left) give the impression that there is agap at
the location marked by the arrow. However, this gap is so narrow that no other atom can t into it. Therefore, the solvent-accessible surface (. Fig.15.3, center) is less mis- leading. It is created by rolling asphere with aradius of
1.4 Å, the size of awater molecule, over the surface of the studied molecule. This surface appears much smoother. The depressions that are still present mean that small mol­ecules—at least one water molecule—can actually t in there. The Lee–Richards surface is less commonly used, but very helpful. It is chosen so that ligand atoms that come into contact with atoms of the protein under inves­tigation will lie directly on this surface (. Fig.15.3, right).
The surface can also be colored. For example, each atom type can be assigned acolor, and then the color of the next closest atom is used in that part of the surface. Similar representations where the surface of the molecule is colored according to other properties, such as electro­static or hydrophobic potential, are very instructive and often used.
15.7 Molecular Dynamics:
Simulation of Molecular Motion
None of the processes we are interested in take place at 0 K, but rather at body temperature, which is about 310 K. It is, therefore, obvious that not only the potential energy but also the kinetic energy must be considered. Molecules also move at room temperature, which is close to body temperature (about 295 K). They diffuse and change shape by adopting different conformations. The exibility and adaptability of both binding partners play an important role in protein–ligand interactions. Apre­requisite for protein binding is that the ligand can adopt aconformation with ashape that ts into aprotein–bind­ing pocket. On the other hand, the protein is exible to some extent. For example, side chains on the surface can adopt different conformations or entire domains can move relative to one another. The mutual adaptation of protein and ligand shapes plays an important role in the formation of protein–ligand complexes.
Molecular dynamics simulations (MD) are atheo­retical method for describing these effects. Molecular dynamics simulations follow the motion of atoms and molecules under the inuence of aselected force eld. It