Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5319_Библиотеки_им_академика_М_И_Перельмана.pdf
X
- •Preface and Acknowledgement
- •Chemical Structures of Amino Acids,Molecular Graphics and Introduction
- •Introduction
- •Literature
- •Chapter Abstract Videos
- •Contents
- •About the author
- •1.10 Synopsis
- •1.3 The Battle Against Infectious Disease
- •1.4 Biological Concepts in Drug Research
- •Bibliography and Further Reading
- •2.8 A Long List of Accidents
- •2.10 Synopsis
- •Bibliography and Further Reading
- •3. Classical Drug Research
- •3.2 Malaria: Success and Failure
- •3.6 Synopsis
- •Bibliography and Further Reading
- •4.1 The Lock-and-Key Principle
- •4.2 The Essential Role of the Membrane
- •4.6 Blame It All on Water!
- •4.11 Lessons for Drug Design
- •4.12 Synopsis
- •Bibliography and Further Reading
- •5.1 Louis Pasteur Sorts Crystals
- •5.2 Structural Basis of Optical Activity
- •5.4 Lipases Separate Racemates
- •5.8 Synopsis
- •Bibliography and Further Reading
- •6.2 Lead Structures from Plants
- •6.9 Synopsis
- •Bibliography and Further Reading
- •7.2 Color Change Demonstrates Activity
- •7.7 Biophysics Supports Screening
- •7.11 Synopsis
- •Bibliography and Further Reading
- •8.1 Strategies for Drug Optimization
- •8.5 From Agonists to Antagonists
- •8.9 Synopsis
- •Bibliography and Further Reading
- •9. Designing Prodrugs
- •9.1 Foundations of Drug Metabolism
- •9.2 Esters Are Ideal Prodrugs
- •9.6 Synopsis
- •Bibliography and Further Reading
- •10. Peptidomimetics
- •10.1 Therapeutic Relevance of Peptides
- •10.2 Designing Peptidomimetics
- •Bibliography and Further Reading
- •11.4 What Is Contained in Chemical Space?
- •Bibliography and Further Reading
- •12.7 Silencing Genes by RNA Interference
- •12.9 Proteomics and Metabolomics
- •Bibliography and Further Reading
- •13.3 Crystal Lattices Diffract X-Rays
- •Bibliography and Further Reading
- •Bibliography and further reading
- •15. Molecular Modeling
- •15.2 Strategies in Molecular Modeling
- •15.3 Knowledge-Based Approaches
- •15.4 Force Field Methods
- •15.5 Quantum Chemical Methods
- •Bibliography and further reading
- •16. Conformational Analysis
- •16.8 Synopsis
- •Bibliography and Further Reading
- •Bibliography and Further Reading
- •18.4 Lipophilicity and Biological Activity
- •Bibliography and Further Reading
- •19.3 The Role of Hydrogen Bonds
- •19.5 Absorption Profiles of Acids and Bases
- •19.8 From In Vitro to In Vivo Activity
- •Bibliography and Further Reading
- •Bibliography and Further Reading
- •21.5 LUDI Discovers the First Leads
- •Bibliography and Original Papers
- •22.1 The Druggable Genome
- •22.4 Enzymes and Their Inhibitors
- •22.9 Resistance and Its Origin
- •Bibliography and Further Reading
- •23.1 Serine-Dependent Hydrolases
- •23.10 Synopsis
- •Bibliography and Further Reading
- •24. Aspartic Protease Inhibitors
- •24.2 Design of Renin Inhibitors
- •24.8 Synopsis
- •Bibliography and Further Reading
- •25.1 Structure of Zinc Metalloproteases
- •25.9 What Zinc Can Do, Iron Can Too
- •25.11 Synopsis
- •Bibliography and Further Reading
- •26. Transferase Inhibitors
- •26.1 The Kinase “Gold Rush”
- •Bibliography and Further Reading
- •27. Oxidoreductase Inhibitors

Chapter • Protein Modeling and Structure-Based Drug Design
20
discovery. These methods can lead to completely novel,
nonpeptidic structures.
An essential requirement for the success of structure-based drug design is an iterative approach, as illustrated in Sect.7.6 and. Fig.7.3. Further examples of
this approach are given later in this chapter. In all cases,
however, the existence of a3D structure of the protein is
the prerequisite and starting point for the structure-based
design of aligand, which is then synthesized and tested.
How do you get such a3D structure of the protein, and
does it have to be determined experimentally in every
single case (Chap.13)?
20.3 Search Tools for Databases
of Experimentally Determined
Protein Complexes
The number of experimentally determined protein structures has grown exponentially in the last years. In 1988,
200 3D structures were found in the Protein Data Bank
(PDB), in the meantime there are more than 227,000 entries, mainly from proteins and protein–ligand complexes.
This rapid growth of known spatial protein structures
is stimulating the development of methods to use this
structural information for the design of new active compounds. Most of the available examples are still predominantly globular, water-soluble enzymes. However, the
number of novel membrane-bound proteins is steadily
increasing. To really exploit this wealth of structures,
database tools are needed that can retrieve, correlate,
and analyze structures and structural motifs. There are
many programs that can compare the sequence and
folded structure of proteins. The Relibase database was
one of the rst tools developed specically for the analysis of protein–ligand complexes. Such adatabase can
be used to search for sequence patterns in proteins and
also to compare the connectivity of bound small-molecule ligands. The database automatically superimposes
proteins using an iterative process to nd optimal superposition, especially of binding pocket regions. Structures aligned in this way can be systematically evaluated.
Which amino acids are involved in ligand interactions?
What functional groups do the ligands use to interact
with the amino acids of the protein? Which residues in
the binding pocket occur repeatedly with identical geometry or are highly exible? The water structure at the
protein–ligand interface can be studied in detail. Surprisingly, astatistical analysis revealed that in about two
thirds of all protein–ligand complexes, at least one crystallographically determined water molecule is involved in
the binding of aligand. This underscores the importance
of including water molecules in modeling efforts. However, it is at this point that concepts for the treatment of
water need to be signicantly improved and extended.
20.4 Comparison of Protein-Binding
Pockets
Another important question concerns the shape and
composition of the binding pocket. Are there other proteins with similar amino acid compositions in which the
pocket has an analogous shape? The actual amino acids
are less important here. Rather, it is the analogous physicochemical properties of the exposed groups, such as
hydrogen bond donors or acceptors, that are oriented
towards the binding pocket. Programs that enable these
comparisons describe the shape and surface of protein
pockets along with the exposed properties. The function
of proteins is often coupled to the recognition and binding of small-molecule ligands or segments of peptide sequences (e.g., proteases). Once bound, these molecules
are chemically modied in the case of enzymes. In the
case of receptors, the ligands are able to induce an effect within the receptor, for example, stabilizing an active
or inactive conformation of the protein or changing its
dynamic behavior. In this way, asignal is transmitted.
The discovery of similarities in binding pockets can lead
to the discovery of functional similarities between proteins. This is independent of whether there is sequence
or folding homology between the proteins. There is also
achance of nding unexpected cross-reactivity through
similarities in the shape and properties of binding pockets. Such unexpected binding is often the cause of adverse effects. By evaluating similarities and differences in
such pockets, it is also possible to identify how ligands
should be modied to achieve the desired selectivity for
the given target protein. Valuable ideas for the design
of new or modied protein ligands can be generated by
studying and comparing bound ligands or ligand building blocks in similar pockets. This provides valuable ideas
for isosteric replacements in the structure-based optimization of initial lead structures. The Cavbase search
engine, implemented in the Relibase database, enables
such pocket comparisons. Ruben Abagyan’s group at
UCSD in San Diego, USA, has developed Pocketome,
acomprehensive encyclopedia that can be used to search
for related binding pockets in protein families. The recent tool SiteMine for binding site comparisons has been
developed by Matthias Rarey’s group at the Hamburg
university in Germany.
20.5 High Sequence Identity Facilitates
Model Generation
An indispensable prerequisite for the use of the method
arsenal of structure-based drug design is the existence
of aspatial structure. Acrystal structure cannot always
be obtained. Under what conditions can amodel of an
unknown protein be constructed from agiven sequence?

. • High Sequence Identity Facilitates Model Generation
. Fig. 20.4 The primary sequences of three cytochromeC proteins
arranged using the typical one-letter code are shown from (a)the denitrifying bacterium Paracoccus denitricans (134 amino acids), (b)the
proteobacterium Rhodospirillum rubrum (112 amino acids), and (c)the
mitochondria of atuna sh (103 amino acids). The proteins vary in
their length and composition. The sequence comparison shows the
Proteins with similar functions from different species differ in their amino acid sequences. As the distance up the phylogenetic tree increases, these differences become more pronounced. Consider the example
of cytochromeC (. Fig.20.4). This widely distributed
protein in mitochondria plays acentral role in the respiratory chain. It consists of apolypeptide chain of about
100 ± 20amino acids. Three cytochromes are shown in
. Fig.20.5 which, despite their different peptide chain
lengths and compositions, have avery similar folding pattern. The proteins from the phylogenetically related spe-
cies human and chimpanzee have 100% sequence identity. In contrast, the yeast enzyme has only 45% identity
with that of these mammals. If the homology is very high
and only afew mutations are present, model construction
will be relatively easy. When sequence identity is greater
than 90%, models can be constructed with uncertainties
approaching the error margins of experimental structure
determinations (Sect.13.5). As sequence identity further
decreases, model building becomes less accurate. At 50%,
the average coordinate error can be afew angstroms. Below an identity of 25–30%, recognition of structural relationships becomes very problematic.
The vast majority of sequence differences between
homologous proteins are located on the surface of the
protein in loop regions that are not critical for the folding
of the protein backbone (Sect.14.4). Exchanges in the interior of the protein can have amuch greater effect on its
architecture. They are usually limited to amino acids of
similar volume and very similar physicochemical proper-
alignment with the best agreement. Invariable or conserved positions
in the sequence are marked in bold. Dashes stand for areas in which
other proteins carry additional amino acids (insertions). The red bars
underscoring the sequences show preferred helical areas. AAla, CCys,
DAsp, EGlu, FPhe, GGly, HHis, IIle, KLys, LLeu, MMet, NAsn,
PPro, QGln, RArg, SSer, TThr, VVal, WTrp, YTyr
ties, such as the exchange of aleucine for an isoleucine.
Often the exchange of one amino acid is coupled with the
complementary exchange of one or more other amino
acids in the immediate vicinity. This is especially true
when polar amino acids are exchanged inside the protein, which are internally saturated, e.g., by salt bridges.
In the newly mutated protein variants, these amino acids form astable orientation. Since the spatial proximity
of amino acid residues in the fold does not necessarily
correspond to their sequential proximity in the protein
chain, the recognition of such structural relationships is
considerably complicated. Mutations in the protein core
can lead to expansion, spatial shifts, or twisting of the
structural building blocks of the protein.
If the identity is very high, only afew amino acid
side chains need to be exchanged. The conformations of
the involved side chains can be deduced from acomparison with the structurally resolved proteins showing these
amino acids in asimilar environment. With decreasing
identity, insertions and deletions in loop regions, i.e., an ex-
pansion or contraction of the polypeptide chain, must be
considered. To predict the conformations of these loops
during model building, libraries of known protein structures have been compiled. Based on length and sequence,
these loops are classied into conformational families.
They can be retrieved by the computer and support the
construction of the spatial arrangement of amodied
loop. The validation of the relevance of these protein
models follows empirical rules. It is checked whether
the constructed geometry agrees with experimental evi-

Chapter • Protein Modeling and Structure-Based Drug Design
20
. Fig. 20.5 Left Superposition of the folded structures of the three
cytochromeC proteins from . Fig.20.4 based on a ribbon model:
Paracoccus denitricans in blue, Rhodospirillum rubrum in red, and
tuna sh in yellow. The cytochromes bind via ahistidine and amethionine to an iron–heme center. The structures were determined by
X-ray crystallography. Structural deviations occur particularly in the
loop regions. Right The same superposition is shown, only here the
individual amino acids are color-coded. The same colors in all three
ribbon models show identical amino acids at different positions (color
coding: Ala: light gray, Val: chartreuse, Gly: white, Ile: bright green,
dence. For example, it must be ensured that hydrophobic
groups are oriented inwards and hydrophilic groups are
oriented mainly outwards. The contact between amino
acid groups is checked, and the chosen torsion angles are
compared with those typically observed.
20.6 Secondary Structure Prediction
and Amino Acid Replacement
Propensities Support Model Building
at Low Sequence Identity
When the sequence identity between the known and
modeled protein falls below 30%, determining structural
homology becomes increasingly difcult. All additional
information must be employed as aresource. An attempt
is made to estimate where in the polymer chain of the
modeled protein certain secondary structure elements
are expected to occur (Sect.14.2). When evaluating the
frequency with which individual amino acids occur in
helices, pleated sheets, or loops, signicant differences
are found. For example, proline is considered a“helix
Leu: olive green, Pro: pink, Phe: violet, Tyr: dark purple, Trp: light vio-
let, Asp: dark red, Glu: wine red, Asn: turquoise, Gln: cyan, Lys: blue,
His: light blue, Arg: medium blue, Ser: light orange, Thr: dark orange,
Cys: light yellow, Met: dark yellow). (7 https://sn.pub/UmXxlS)
breaker.” It occurs at most in the rst turn of ahelix;
at other positions, it disrupts the geometry and induces
akink. To determine whether aparticular sequence segment folds as ahelix, pleated sheet, or loop, the information about positional preferences is evaluated for several
neighboring amino acids in an overlapping fashion.
Once analyzed, the primary sequence is compared
to areference protein of known geometry. Since its 3D
structure is known, the assignment of the sequence to the
secondary structural elements is straightforward. If not
only one, but several 3D structures of members of ahomologous protein family are known, multiple sequence
alignments can be used to construct arepresentative prole of the expected secondary structure. This prole then
serves as areference for the mutual comparison of the
sequences of structurally known and unknown proteins.
For this purpose, in addition to the physicochemical
properties of the amino acids, their local conformational
properties, their main and side chain orientations, their
accessibility for solvent molecules, or their involvement
in hydrogen bonds are analyzed. At the same time, the
probabilities with which such an amino acid substitution can occur at the DNA sequence level are taken into

. • Secondary Structure Prediction and Amino Acid Replacement Propensities Support Model Building
account (Sect. 32.7). These parameters can be easily
determined for proteins with aknown spatial structure.
Comparing the structures within aset of homologous
proteins yields probabilities for mutual amino acid substitutions. For example, unlike all other amino acids,
glycine has no side chain (see pageIX). Therefore, it can
adopt conformations in the polymer chain that are sterically impossible for other amino acids. The polymer
chain assumes such conformations in regions close to the
protein surface, where it reverses its orientation. This is
where the conformationally exible glycines play an important role. Such glycines that are exposed to the solvent
and have unusual torsion angles are largely conserved
between the folding of homologous proteins. The conserved glycines can be searched for during the sequence
alignment of aprotein to be modeled. They thus provide anchor points for sequence alignment. Many similar
rules have been established. They serve as criteria for the
recognition of sequence segments of structural importance. They are then applied to the protein sequence to
be modeled. Even with relatively low sequence identity,
structural homologies between aprimary sequence and
aprotein of known 3D structure can be detected in this
way. They are used as criteria for homology modeling.
Before the breakthrough in successful experimental
structure determination of G-protein-coupled recep-
tors, the most important group of membrane receptors
(Chap.29), it was necessary to rely on their model building. Additional criteria were used for this purpose. The
modeling had to ensure that the hydrophobic amino acids in the helical areas were embedded in the membrane
and oriented towards the membrane environment. Meanwhile, the homology modeling programs have achieved
ahigh degree of automation. At the Biozentrum in Basel,
the SWISS-MODEL server which transforms submitted
sequences fully automatically into 3D structures was set
up. The program Modeller from the group of Andrej Šali
in San Francisco is able to assemble protein models from
the sequences of entire genomes in silico. Although the
structures are certainly rough, and many may be incorrect,
such an approach allows asearch for similarities in the
recognition determinants of proteins. In this way, possible
interactions between proteins can be discovered, or commonalities in metabolic pathways become transparent.
The modeling of proteins gives good results especially
when they have ahigh homology. This is given in regions
that determine the folding scaffold. The binding pockets are mostly located in the loop regions (Sect.14.4).
This is where even homologous proteins differ greatly.
Therefore, the model constructions do not achieve the
desired accuracy in these regions. An improvement can
be achieved if aligand is already placed in the assumed
binding region during model construction. Model and
placement must be optimized in an iterative process using
appropriate energy functions.
As mentioned above, the problem of predicting the
structure of aprotein from its amino acid sequence using theoretical concepts has not yet been successfully
solved. Therefore, the development of methods for prediction using data from experimentally determined structures has been the subject of intensive research for many
years. Frustratingly, progress has been very incremental.
The breakthrough to signicantly improve the quality
of predictions came in 2020, when the researchers at the
Google subsidiary DeepMind, who had already attracted
attention with their articial intelligence (AI)-based programs AlphaGo and AlphaZero for simulating board
games, achieved surprising success. Under the leadership
of John Jumper, they won the CASP14 competition (Critical Assessment of Protein Structure Prediction) by awide
margin with their machine learning-based program Alpha-
Fold2.0. The competition, which has been held since 1994,
is performed to compare approaches to protein structure
prediction. Research groups are provided with amino acid
sequences of proteins. Their structures have been experimentally determined but not yet published. For these
sequences, the participants have to predict structures using their computational models. AlphaFold2.0 is based
on aneural network (deep learning). It is trained to learn
structure predictions based on experimentally determined
protein structures from the PDB database. It evaluates input data by incorporating multiple sequence comparisons,
evolutionarily related sequences, and physical and geometric constraints such as the preferred spatial pairing of
amino acid residues. It applies rules similar to those used
in the homology modeling described above. The scientists
at DeepMind use the “attention mechanism” known from
transformer architectures in their neural network. This
involves highlighting parts of the input data in the context of the overall environmental data, while attenuating
others. For language, this would be words in the context
and meaning of asentence or text. In this way, the neural
network automatically focuses on and weights the critical
information. Similar approaches have revolutionized natural language processing in recent years and are the basis
for signicant performance improvements in current AI
systems for language and translation programs. As AlphaFold iterates through the entire structure of the protein to
be folded, it simultaneously understands how to locally
improve each part of the structure through repeated application. In doing so, it takes into account information
already learned from previous steps. But how AlphaFold
ultimately achieves its impressive results is not easy to
deduce. With AI approaches, it is by no means trivial to
deduce the system’s rules and decision trees after the fact.
This is especially true for analyses in which the decisions
of the neural network are self-organizing and in which
humans no longer intervene during the learning process.
DeepMind left the other participants far behind in
the CASP14 competition. But academic research quickly
caught up. David Baker’s group in Seattle has since in-

Chapter • Protein Modeling and Structure-Based Drug Design
20
tegrated DeepMind’s AI approaches into the RoseTTA
Fold software, which now delivers equivalent results.
DeepMind has released AlphaFold2.0 for free via acloud
computing tool. It can also be installed locally on acomputer. The free use appeared only logical, as DeepMind
has trained its AI on the publicly available structures in
the PDB database. With the release, huge databases are
now lling up with structures predicted by AlphaFold2.0.
There are now more than amillion models available in
the PDB. Meanwhile, anew version, AlphaFold3, with
enhanced functionality has been released and published
in Nature. It is described to be not limited to single-chain
proteins, it can also predict the structures of protein complexes with DNA, RNA, post-translational modications,
and selected ligands and ions. Unfortunately, this time its
publication was without freely accessible executables or
source code and for noncommercial trial applications it
can be accessed through a server. In the interest of free
scientic exchange and despite the predictable high level
of commercial interest, it would be greatly appreciated
if DeepMind could consider a more open approach to
the release of the tool. It is believed that only through
frequent use and feedback from the community, who provided the required database to train the neural network,
will the computer tool achieve further improvements.
It seems reasonable to conclude that AI programs
have already made a notable impact on protein structure
research, and it seems probable that they will also inuence predictions in future drug design. It is important
to remember that structures for targeted ligand design
require high resolution and the modeled geometries must
be accurate to less than an Angstrom. Accurately predicting folding patterns to within afew angstroms may
be sufcient for some applications. However, drug design
researchers have learned the hard way from arenin model
derived from the related endothiapepsin that small deviations can mislead targeted ligand design (Sect.24.2).
The tools have been met with considerable enthusiasm and were recognized in 2024 with the Nobel Prize
in Chemistry, which was bestowed upon the two leading
DeepMind scientists, John Jumper and Demis Hassabis. The long-standing achievements of David Baker's
group were also acknowledged, particularly given that
David Baker has succeeded in predicting entirely novel
folded proteins with tailored properties, which were subsequently validated through experiment. This opens up
new perspectives for countless applications in biochemistry, structural biology, and drug design.
20.7 Ligand Design: Seeding, Expanding,
and Linking
The next step after the analysis of the binding pocket of the
experimentally determined or modeled protein is the actual
ligand design. There are several computer-aided design ap-
proaches available to suggest new protein ligands. Adocking program can be used to place successively preselected ligands from adatabase into the binding pocket (. Fig.20.6).
Typically, the database is populated with molecular candidates that closely resemble common drug-like molecules.
This method, which has become known as “virtual screen-
ing” (Sect.7.6), can now screen many millions of candidate
molecules in afew hours using parallel computers.
Another approach starts with a“seed” in the binding pocket. Starting from this point, the ligand gradually
grows into the binding pocket. This is the principle followed by most de novo design programs. The placement of
the rst seed is critical. Such approaches are parti cularly
successful when there is aspecic hot spot in the binding pocket from which further optimization starts. Salt
bridges to charged amino acids or coordination of metal
ion centers are particularly well suited for this approach.
This concept has been successfully applied, for example,
to the serine proteases trypsin and thrombin (Sect.23.4)
and the zinc-containing carbonic anhydrases (Sect.25.7).
Another approach starts with small fragments that are
placed into the protein binding pocket. Increasingly, the
results of experimental fragment screening (Sect.7.9) are
used to provide reliable starting points for further growth
of these candidate molecules in the binding pocket. The
approach is an ideal symbiosis of experimental work
and subsequent computational design. It combines
the strengths of both techniques. Another approach is
to chemically link several of the tting fragments together. This strategy has been successfully applied, for
example, in the SAR-by-NMR method (Sect.7.8). The
cover of this book shows an example of how ananomolar protein kinaseA inhibitor was developed through
several design steps, starting with the crystal structure
of a millimolar phenol fragment (see . Fig.02 page
XI,7 https://sn.pub/4NppHn).
20.8 Docking Ligands into Binding
Pockets
Docking attempts to t potential protein ligands into
abinding pocket using acomputer. Adocking program
takes one candidate at atime from aprecompiled library
of molecules. For each entry, a3D structure is generated.
If aexible molecule is encountered, either multiple conformations will be generated and docked individually,
or they will be generated on the y during docking. The
next step is to t each molecule into the binding pocket.
First, the structures that cannot bind to the protein are
discarded. In addition, other structures that cause obvious problems, such as electrostatic repulsion with the
protein in the assumed docking mode, are eliminated.
Typically, adocking program generates several solutions.
These are evaluated on the basis of the generated binding
geometries and their afnity is estimated.

. • Docking Ligands into Binding Pockets
. Fig. 20.6 Potential strategies for ligand design. The complete 3D
structures of putative ligands are docked into the binding pocket (left
part of the image). The construction of new molecules is sketched in
the middle and right part of the picture. Basically there are two pos-
Irwin Kuntz is apioneer in the eld of docking programs; DOCK was developed in his group at UCSF in
San Francisco, USA. The original version in 1982 evaluated only the steric complementarity of ligands and
proteins. The shape of the binding pocket was approximated by aset of different spheres so that the pocket was
completely lled. Amathematical method was then used
to place the test ligands on this distribution of spheres.
Complementarity, ameasure of direct protein–ligand
contacts, served as ascoring function. Since the rst version, DOCK has evolved considerably. The program now
uses aforce eld for scoring and calculates the contributions for desolvation. Even the placement of the ligands is
sibilities. Afragment can be placed as aseed and other groups can be
attached step by step (middle). Alternatively, several small-molecule
fragments can be placed in the binding pocket independently of one
another and later linked or merged together (right)
exible by considering rotatable bonds. Another docking
prototype was developed at the GMD in Bonn, Germany,
by Matthias Rarey. The program FlexX was the rst program that could quickly handle ligand exibility during
docking. It decomposes the test ligands into individual
fragments and then uses an algorithm very similar to the
positioning algorithm implemented in the de novo design
program LUDI (Sect.20.10). After placement of the rst
building block, the ligand is successively reconstructed
in the binding pocket. Different conformers along the
rotatable bonds are considered. The program maintains
stored tables of preferred torsion angles, similar to those
described in Sect.16.6. The energetic evaluation of the

Chapter • Protein Modeling and Structure-Based Drug Design
20
placement is performed in this step. The program AutoDock from the group of Art Olson at Scripps in La Jolla,
San Diego, USA, uses alattice-based algorithm for placement. Aforce eld function similar to that of the GRID
program (Sect.17.10) is used to place potential values
on agrid embedded in the binding pocket. Starting from
arandom orientation, the ligand is moved over the grid
until an optimum is found. In doing so, it “feels” the interaction potential with the protein. Since the potential has
already been precalculated on the lattice, this evaluation is
particularly fast. At the same time, twisting around rotatable bonds is performed. The program GOLD, developed
by Gerrith Jones in the group of Peter Willett in Shefeld,
England, also uses alattice for placement. However, the
interaction potentials are parameterized using crystal
data. GOLD uses agenetic algorithm to optimize the geometry. Over time, aplethora of docking programs has
been developed. All follow aslightly different strategy, but
are based on the concepts described for the prototypes
mentioned above. Some follow the idea that it is better to
generate awell-distributed number of rigid ligand conformers and then dock them quickly as rigid bodies.
Today there are three main problems that impose limits
on docking. One is the energetic evaluation of the generated geometries. This will be specially addressed in the
next section. Another is that water plays adecisive role in
ligand binding (Sect.20.3). Even today no really convincing solution to the handling of water molecules during
docking has been found. The third problem is the exible
adaptation of the protein (Sect.15.8). Usually there are
small adaptations on the side of the protein that slightly
change the shape of the binding pocket. In fact, they are
big enough to make the docking programs look for the
proverbial red herrings (means: they lead to wrong tracks).
20.9 Scoring Functions: Ranking
of Constructed Binding Geometries
Arelevant scoring of the generated binding geometries is
essential for all docking and de novo design approaches
in structure-based drug design. Among the numerous
geometrically plausible placements, only those that reasonably approximate the experimentally found situations
must be ltered out. The enthalpic and entropic contributions, which determine the afnity of aligand to its target
protein, were described in Chap.4. The goal of ascoring
function is to quickly estimate the expected binding afnity
from agiven interaction geometry. Theoretically, asingle
geometry is not sufcient to solve this problem. Molar energies are determined by anite set of conformations. They
are distributed over an “ensemble” of multiple states. These
states are differently populated according to their energy
content. One group of methods tries to account for this
fact in the calculations. This is the most theoretically sound
approach. The energy contributions of the ensemble (usu-
ally taken from the trajectories of amolecular dynamics
simulation, Sect.15.7) are summed. The Gibbs free energy
∆G (Sect.4.3) can be estimated from the resulting partition function. However, the computations required for such
evaluations are very time consuming. As aresult, they are
usually not feasible for screening large amounts of data in
astructure-based drug design session.
Instead, regression-based scoring functions are used as
an alternative approach. Assuming that aparticular state
is predominantly populated, it may be justied to consider
only one state (or conformation) in the scoring function.
The enthalpy and entropy contributions that most likely determine the binding afnity are considered. The approach
is reminiscent of setting up aQSAR equation (Sect.18.2).
The terms are combined into an energy function. Molecular
descriptors that correctly reect the contributions to these
terms are sought. In doing so, the erroneous assumption
is made that the individual contributions to the description of the free enthalpy are additive (Sect.4.10). The indi-
vidual terms of the equation are each given an adjustable
weighting factor. As with QSAR equations, amathematical
technique is used to optimally t these weighting factors to
atraining dataset. This set consists of crystallographically
determined protein–ligand complexes for which experimental binding afnities are available.
Athird approach follows aso-called knowledge-based
concept. As already discussed in Sect.17.10, the frequency
of individual contact geometries in the crystal structures
of protein–ligand complexes is evaluated. Akind of “normal distribution” is dened as areference state. Then,
all contacts that occur more frequently than average are
classied as energetically favorable. All contacts occurring
less frequently than the average are classied as unfavorable. These scores are then applied to all contacts present
in acomputationally generated protein–ligand complex
and combined for aranking. So far, afunction derived
in this way can be used for the relative energetic ranking
of aset of ligands binding to the same reference protein.
However, afnity prediction will also be successful here if
training is performed against areference dataset of known
geometry and binding afnity in amanner analogous to
regression-based scoring functions. The evaluation with
the regression-based or knowledge-based function is very
fast. In the meantime, alarge number of scoring functions
have been developed. So far, none of them has proven to
be ideal and generally applicable. Therefore, in each case
it has to be checked which function provides the best performance for the protein under investigation.
20.10 De Novo Design: From LUDI
to the Automated Assembly
of Novel Ligands
The rst program for stepwise de novo design was GROW
by Jeffrey Howe and Joseph Moon at Upjohn. It focuses

. • The Feasibility of Designing Ligands In Silico
. Fig. 20.7 Concept of the program LUDI for de novo design of
protein ligands. In the rst step, the interaction sites are determined
(left). Donor sites are represented by blue lines, acceptor sites by red
lines. The green dots symbolize lipophilic sites. Subsequently, small
molecules from adatabase are tted into the binding pocket in that
on peptides as lead structures. An amide group is positioned in afavorable orientation in the binding pocket.
The next amino acids are added to the starting amide
group in astepwise fashion. At each step, alarge number
of different conformations of all 20proteinogenic amino
acids are attached to the seed on the y. The “best” solutions for each are followed. In this way, GROW constructs
apeptide ligand of increasing length in the binding pocket.
In the early 1990s, Hans-Joachim Böhm developed
the program LUDI at BASF, Ludwigshafen, Germany.
The idea was to read small molecules or molecular fragments from adatabase with precalculated spatial geometries and position them in the binding pocket so that
hydrogen bonds are formed with the protein and hydrophobic pockets are lled with nonpolar groups. As input, the program requires the coordinates of the protein
and alibrary of 3D structures of fragments or drug-like
molecules.
The precalculation of interaction sites is crucial.
These are placed in the binding pocket around the
amino acid residues in the form of tting points or directional vectors (. Fig.20.7). The program uses rules
derived from the nonbonded interactions found in the
crystal packing of small organic molecules (Sects.14.7
and17.10). LUDI then extracts small molecules or molecular fragments from the 3D library. For each entry, an
attempt is made to position it in the binding pocket of
the protein so that as many of these interaction sites as
possible are satised (. Figs.20.7 and17.12). All successfully placed fragments are then ranked. The scoring
function used takes into account the number and quality
of H-bonds and ionic interactions formed, hydrophobic
contact surfaces shared by the protein and ligand, and
unfavorable contributions due to the number of rotatable
bonds in the ligands. An example of the successful application of this program is described in Sect.21.5.
As the rst prototype, LUDI was the gold standard for
many de novo design programs that were developed later.
These approaches implemented enhanced scoring functions and improved fragment libraries. The programs were
also taught synthesis rules so that the chemical accessi-
they are matched with the interaction sites (middle). Finally LUDI can
chemically link groups or parts of molecules to larger structures to
match the remaining unsatised interaction sites and to ll the entire
binding pocket (right)
bility of the generated molecules was taken into account.
The search space of the programs has also been expanded
to include multiple conformations and congurations.
Ade novo design program is an idea generator. Its value
is, of course, determined by the concepts that went into its
development. On the other hand, its value also depends
strongly on the user and how the suggestions of such
aprogram are interpreted and used for further design.
20.11 The Feasibility of Designing
Ligands In Silico
Certainly, many examples have been presented that
demonstrate the power of de novo design, virtual screening, and docking. The example described in Chap.21 was
successful because of the use of such methods. However,
it is critical that the computational methods are tightly
integrated into an iterative process of synthesis, biological testing, and experimental structure determination. It
is important to keep in mind that not all hits from computational screening are based on the correct assumptions, and not all hits are picked up for the right reasons.
The predictive power of available methods is still limited. The synthetic accessibility of aproposed molecule is
often not sufciently taken into account, the exibility of
the protein is frequently neglected, and the methods for
estimating binding afnity are still too inaccurate. This is
because the process and factors responsible for molecular
recognition and ligand binding are still poorly understood. The correct description of solvation effects, the
involvement of water molecules in the binding process,
and the change of protonation states are major problems.
The contribution of hydrogen bonding to binding afnity is still only an estimate, despite efforts to the contrary.
In the case of lipophilic interactions, it can at least be assumed that the lling of an unoccupied lipophilic pocket
with additional nonpolar substituents is in most cases
accompanied by an increase in binding afnity.
How to account for the changes in binding entropy
that contribute to the free energy of ligand binding is still

Chapter • Protein Modeling and Structure-Based Drug Design
20
insufciently considered. Certainly, there is evidence that
the oversimplied assumption that entropic contributions will be constant within aset of congeneric ligands
is denitely not valid.
However, there are other fundamental limitations of
this approach. The most important is that the technique is
limited to optimizing direct interactions with the protein.
Successful binding to atarget protein is essential for any
drug. However, to be suitable as adrug, additional requirements must be met. These include good selectivity, metabolic stability, adequate duration of action, low addictive
potential, and negligible toxicity. Today, at least the selectivity of acompound towards members of astructurally
related protein family can be estimated with some certainty.
Fully automated molecular design on acomputer is
not yet possible, nor is it likely to be in the long term. The
methods of structure-based design are valuable as idea
generators. The resulting suggestions must be checked
and modied if necessary. Time will tell whether these
methods will gradually approach the “holy grail” of drug
design: the design of drug molecules from scratch.
20.12 Synopsis
In structure-based drug design, attempts are made
-
to design small-molecule ligands by docking them
directly into the binding pocket of atarget protein.
This requires a3D structure of the reference protein.
The goal is to optimally ll the binding pocket with
aligand by satisfying nonbonded interactions with
the functional groups of the binding-site residues.
Structure-based design starts with adetailed analysis
-
of the binding pocket to elucidate hot spots for puta-
tive interactions with the protein. Either experimental
methods or computational tools can be used to per-
form an active site mapping with molecular probes or
small solvent-like molecules.
In an iterative process of structure determination,
-
modeling of modied ligands, docking and screen-
ing, synthesis, and biological testing, the properties
of small-molecule ligands are improved to optimize
binding to the target protein.
Databases have been developed to retrieve and com-
-
pare structural information about the exponentially
growing body of structural data on protein–ligand
complexes. They allow comparison of binding poses,
active-site interaction geometries, protein–ligand
binding motifs, and the original solvation structures
in the protein’s binding pocket.
Proteins can be compared in terms of their exposed
-
binding pockets. The shape and the exposure of
groups which have particular physicochemical prop-
erties in binding pockets are compared and help to
design small-molecule ligands with the desired selec-
tivity. Ideas for isosteric replacements on the ligand
scaffold can also be generated in this way.
If an experimentally determined structure of the tar-
-
get protein is unavailable, ahomology model can be
constructed by using arelated protein of known architecture as atemplate. The accuracy and the success
of such homology modeling depend strongly on the
sequence homology with the template structure.
Tools for secondary structure prediction and amino
-
acid replacement propensities have been developed to
improve the reliability of the sequence assignment of
proteins to the 3D structure of the reference template.
Recently, AI programs have been developed based on
neural networks to generate 3D structures from sequence data. They were trained to learn structure predictions based on experimentally determined protein
structures.
Ligand design approaches screen large databases of
-
candidate molecules by docking them into the binding pocket.
Alternatively, de novo design starts with asmall mol-
-
ecule seed or fragment and grow them into putative
ligands in the pocket. Two nonoverlapping fragments
can also be linked together to form alarger ligand
with improved binding properties.
The geometry of aconstructed protein–ligand com-
-
plex must be evaluated in terms of the expected
binding afnity in all structure-based design strategies. Alarge variety of scoring functions are used to
predict the binding afnity based on the geometry of
the formed complex.
Bibliography and Further Reading
General Literature
C. Branden, J. Tooze, Introduction to Protein Structure, Garland Pub-
lishing, Inc. New York, 1991, 2nd Edition, 1999
T. J. P. Hubbard, A. M. Lesk, Modelling Protein Structures, in Com-
puter Modelling in Molecular Biology, J. M. Goodfellow, Eds.,
VCH, Weinheim, 1995
C. Hutchins, J. Greer, Comparative Modeling of Proteins in the Design
of Novel Renin Inhibitors, Crit. Reviews Biochem. Molec. Biol.,
26, 77–127 (1991)
P. Goodford, Drug Design by the Method of Receptor Fit, J. Med.
Chem., 27, 557–564 (1984)
J. Greer, J. W. Erickson, J. J. Baldwin, M. D. Varney, Application of
the Three-Dimensional Structures of Protein Target Molecules
in Structure-Based Drug Design, J. Med. Chem., 37, 1035–1054
(1994)
C. R. Beddell, Ed., The Design of Drugs to Macromolecular Targets,
Wiley, Chichester, 1992
I. D. Kuntz, Structure-Based Strategies for Drug Design and Discovery,
Science 257, 1078–1082 (1992)
S. Borman, New 3D Search and De Novo Design Techniques Aid Drug
Development, Chem. & Eng. News, 10, 18–26 (1992)
Y. C. Martin, 3D Database Searching in Drug Design, J. Med. Chem.,
35, 2145–2154 (1992)
I. D. Kuntz, E. C. Meng, B. K. Shoichet, Structure-Based Molecular
Design, Acc. Chem. Res., 27, 117–123 (1994)

Bibliography and Further Reading
H. J. Böhm, Ligand Design, in: 3D QSAR in Drug Design, H. Kubinyi,
Eds., Escom, Leiden, 1993, pp. 386–405
K. Müller, Ed., De Novo Design, Persp. Drug Discov. Design, Vol. 3,
Escom, Leiden, 1995
H. J. Böhm, G. Schneider, Molecular Recognition in Protein–Ligand
Interactions (Vol. 19, in Methods and Principles in Medicinal
Chemistry, R. Mannhold, H. Kubinyi und G. Folkers, Eds.), Wi-
ley-VCH, Weinheim, 2006
G. Schneider, K. H. Baringhaus, Molecular Design, Wiley-VCH, Wein-
heim, 2008
M. Eguida, D. Rognan, Estimating the Similarity between Protein
Pockets, Int. J. Mol. Sci., 23, 12462 (2022)
K. A. Carpenter, R. B. Altman, Databases of ligand-binding pockets
and protein-ligand interactions, 23, 1320–1338 (2024)
Special Literature
C. R. Bedell, P. J. Goodford, F. E. Norrington, S. Wilkinson, R. Woot-
ton, Compounds Designed to Fit a Site of Known Structure in
Human Haemoglobin, Brit. J. Pharmacol., 57, 201–209 (1976)
M. Hendlich, A. Bergner, J. Günther, G. Klebe, Design and Devel-
opment of Relibase—a Database for Comprehensive Analysis of
Protein–Ligand Interactions, J. Mol. Biol., 326, 607–620 (2003)
S. Schmitt, D. Kuhn, G. Klebe, A new method to detect related function
among proteins independent of sequence and fold homology, J.
Mol. Biol., 323, 387–406 (2002)
A. Weber etal., Unexpected Nanomolar Inhibition of Carbonic An-
hydrase by COX-2-Selective Celecoxib: New Pharmacological
Opportunities due to Related Binding Site Recognition, J. Med.
Chem., 47 550–557 (2004)
C. Gerlach etal., KNOBLE: a knowledge-based approach for the de-
sign and synthesis of readily accessible small molecule chemical
probes to test protein binding, Angew. Chem. Int. Ed. Engl., 46,
9105–9109 (2007)
I. Kufareva, A. V. Ilatovskiy, R. Abagyan, Pocketome: An Encyclope-
dia of Small-molecule Binding Sites in 4D. Nucl. Acid. Res., 40,
D535–D540 (2012)
K. N. Allen, C. R. Bellamacina, X. Ding, C. J. Jeffery, C. Mattos, G.
A. Petsko, and D. Ringe An Experimental Approach to Mapping
the Binding Surfaces of Crystalline Proteins, J. Phys. Chem. 100,
2605–2611 (1996)
J. Overington, M. S. Johnson, A. Sali and T. L. Blundell, Tertiary Struc-
tural Constraints on Protein Evolutionary Diversity: Templates,
Key Residues and Structure Prediction, Proc. Royal Soc. Lond. B
241, 132–145 (1990)
M. S. Johnson, J. P. Overington, T. L. Blundell, Alignment and Search-
ing for Common Protein Folds Using a Data Bank of Structural
Templates, J. Mol. Biol., 231, 735–752 (1993)
A. Waterhouse etal., SWISS-MODEL: Homology Modelling of Pro-
tein Structures and Complexes, Nucl. Acid. Res., 46, W296–W303
(2018)
A. Sali and T. L. Blundell, Denition of General Topological Equiv-
alence in Protein Structures, J. Mol. Biol., 212, 403–428 (1990)
A. Sali and T. L. Blundell, Comparative Protein Modelling by Satis-
faction of Spatial Restraints. J. Mol. Biol., 234, 779–815 (1993)
J. Jumper etal., Highly accurate protein structure prediction with Al-
phaFold, Nature 596, 583–589 (2021)
J. Abramson, J. Adler, J. Dunger etal. Accurate structure prediction of
biomolecular interactions with AlphaFold3. Nature 630, 493–500
(2024)
M. Baek etal. Accurate prediction of protein structures and interac-
tions using a three-track neural network, Science, 373, 871–876
(2021)
E. Callaway, It will change everything: DeepMind’s AI makes gigantic
leap in solving protein structures, Nature 588, 203–204 (2020)
M. Hibert, S. Trumpp-Kallmeyer, J. Hoack, A. Bruinvels, This is Not
a G-Protein-Coupled Receptor, Trends Pharm. Sci., 14, 7–12 (1993)
J. Hoack, S. Trumpp-Kallmeyer and M. Hibert, Re-evaluation of
Bacteriorhodopsin as a Model for G-Protein-Coupled Receptors,
Trends Pharm. Sci., 15, 7–9 (1994)
R. Henderson, J. M. Baldwin, T. A. Ceska, F. Zemlin, E. Beckmann, K.
H. Downing, Model of the Structure of Bacteriorhodopsin Based
on High-Resolution Electron Cryo-Microscopy, J. Mol. Biol., 213,
899–929 (1990)
G. F. X. Schertler, C. Villa, R. Henderson, Projection Structure of Rho-
dopsin, Nature, 362, 770–772 (1993)
J. Travis, Proteins and Organic Solvents Make an Eye-Opening Mix,
Science, 262, 1374 (1993)
C. S. Ring etal., Structure-based Inhibitor Design by Using Protein
Models for the Development of Antiparasitic Agents, Proc. Natl.
Acad. Sci., 90, 3583–3587 (1993)
I. D. Kuntz, J. M. Blaney, S. J. Oatley, R. Langridge, T. E. Ferrin, A
Geometric Approach to Macromolecule-Ligand Interactions, J.
Mol. Biol., 161, 269–288 (1982)
M. Rarey, B. Kramer, T. Lengauer, G. Klebe, A Fast Flexible Docking
Method using an Incremental Construction Algorithm, J. Mol.
Biol., 261 470–489 (1996)
D. S. Goodsell, A. J. Olson, Automated Docking of Substrates to Pro-
teins by Simulated Annealing, Proteins, 8, 195–202 (1990)
G. Jones, P. Willett, R. C. Glen, A. R. Leach, R. Taylor, Development
and Validation of a Genetic Algorithm for Flexible Docking, J.
Mol. Biol., 267, 727–748 (1997)
Pocketome: https://ngdc.cncb.ac.cn/databasecommons/database/id/501
(Last accessed Nov. 19, 2024)
SiteMine: https://uhh.de/naomi (Last accessed Nov. 19, 2024)
Explore Computed Structure Models Alongside PDB Data: https://
www.rcsb.org/news/6304ee57707ecd4f63b3d3db (Last accessed
Nov. 19, 2024)
SWISS-Model web server: https://swissmodel.expasy.org/ (Last ac-
cessed Nov. 19, 2024)
Modeller: https://salilab.org/modeller/ (Last accessed Nov. 19, 2024)
AlphaFold: https://neurosnap.ai/academic?gclid=CjwKCAiAqNS-
sBhAvEiwAn_tmxbN3V2G2yrU1LJABBGDpMRMS-
gu_2u6c-8v76nQEU-tDj1OMPXPFoQxoCX7gQAvD_BwE (Last
accessed Nov. 19, 2024)
Соседние файлы в папке Библиотека им академика М.И. Перельмана
