Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5319_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
29.08.2026
Размер:
92 Мб
Скачать
Chapter  • Protein Modeling and Structure-Based Drug Design
20
discovery. These methods can lead to completely novel, nonpeptidic structures.
An essential requirement for the success of struc­ture-based drug design is an iterative approach, as illus­trated in Sect.7.6 and. Fig.7.3. Further examples of this approach are given later in this chapter. In all cases, however, the existence of a3D structure of the protein is the prerequisite and starting point for the structure-based design of aligand, which is then synthesized and tested. How do you get such a3D structure of the protein, and does it have to be determined experimentally in every single case (Chap.13)?
20.3 Search Tools for Databases
of Experimentally Determined Protein Complexes
The number of experimentally determined protein struc­tures has grown exponentially in the last years. In 1988, 200 3D structures were found in the Protein Data Bank (PDB), in the meantime there are more than 227,000 en­tries, mainly from proteins and protein–ligand complexes. This rapid growth of known spatial protein structures is stimulating the development of methods to use this structural information for the design of new active com­pounds. Most of the available examples are still predom­inantly globular, water-soluble enzymes. However, the number of novel membrane-bound proteins is steadily increasing. To really exploit this wealth of structures, database tools are needed that can retrieve, correlate, and analyze structures and structural motifs. There are many programs that can compare the sequence and folded structure of proteins. The Relibase database was one of the rst tools developed specically for the anal­ysis of protein–ligand complexes. Such adatabase can be used to search for sequence patterns in proteins and also to compare the connectivity of bound small-mole­cule ligands. The database automatically superimposes proteins using an iterative process to nd optimal su­perposition, especially of binding pocket regions. Struc­tures aligned in this way can be systematically evaluated. Which amino acids are involved in ligand interactions? What functional groups do the ligands use to interact with the amino acids of the protein? Which residues in the binding pocket occur repeatedly with identical ge­ometry or are highly exible? The water structure at the protein–ligand interface can be studied in detail. Sur­prisingly, astatistical analysis revealed that in about two thirds of all protein–ligand complexes, at least one crys­tallographically determined water molecule is involved in the binding of aligand. This underscores the importance of including water molecules in modeling efforts. How­ever, it is at this point that concepts for the treatment of water need to be signicantly improved and extended.
20.4 Comparison of Protein-Binding
Pockets
Another important question concerns the shape and composition of the binding pocket. Are there other pro­teins with similar amino acid compositions in which the pocket has an analogous shape? The actual amino acids are less important here. Rather, it is the analogous phys­icochemical properties of the exposed groups, such as hydrogen bond donors or acceptors, that are oriented towards the binding pocket. Programs that enable these comparisons describe the shape and surface of protein pockets along with the exposed properties. The function of proteins is often coupled to the recognition and bind­ing of small-molecule ligands or segments of peptide se­quences (e.g., proteases). Once bound, these molecules are chemically modied in the case of enzymes. In the case of receptors, the ligands are able to induce an ef­fect within the receptor, for example, stabilizing an active or inactive conformation of the protein or changing its dynamic behavior. In this way, asignal is transmitted. The discovery of similarities in binding pockets can lead to the discovery of functional similarities between pro­teins. This is independent of whether there is sequence or folding homology between the proteins. There is also achance of nding unexpected cross-reactivity through similarities in the shape and properties of binding pock­ets. Such unexpected binding is often the cause of ad­verse effects. By evaluating similarities and differences in such pockets, it is also possible to identify how ligands should be modied to achieve the desired selectivity for the given target protein. Valuable ideas for the design of new or modied protein ligands can be generated by studying and comparing bound ligands or ligand build­ing blocks in similar pockets. This provides valuable ideas for isosteric replacements in the structure-based opti­mization of initial lead structures. The Cavbase search engine, implemented in the Relibase database, enables such pocket comparisons. Ruben Abagyan’s group at UCSD in San Diego, USA, has developed Pocketome, acomprehensive encyclopedia that can be used to search for related binding pockets in protein families. The re­cent tool SiteMine for binding site comparisons has been developed by Matthias Rarey’s group at the Hamburg university in Germany.
20.5 High Sequence Identity Facilitates
Model Generation
An indispensable prerequisite for the use of the method arsenal of structure-based drug design is the existence of aspatial structure. Acrystal structure cannot always be obtained. Under what conditions can amodel of an unknown protein be constructed from agiven sequence?
. • High Sequence Identity Facilitates Model Generation


. Fig. 20.4 The primary sequences of three cytochromeC proteins
arranged using the typical one-letter code are shown from (a)the de­nitrifying bacterium Paracoccus denitricans (134 amino acids), (b)the proteobacterium Rhodospirillum rubrum (112 amino acids), and (c)the mitochondria of atuna sh (103 amino acids). The proteins vary in their length and composition. The sequence comparison shows the
Proteins with similar functions from different spe­cies differ in their amino acid sequences. As the dis­tance up the phylogenetic tree increases, these differ­ences become more pronounced. Consider the example of cytochromeC (. Fig.20.4). This widely distributed protein in mitochondria plays acentral role in the respi­ratory chain. It consists of apolypeptide chain of about 100 ± 20amino acids. Three cytochromes are shown in
. Fig.20.5 which, despite their different peptide chain
lengths and compositions, have avery similar folding pat­tern. The proteins from the phylogenetically related spe-
cies human and chimpanzee have 100% sequence iden­tity. In contrast, the yeast enzyme has only 45% identity with that of these mammals. If the homology is very high and only afew mutations are present, model construction will be relatively easy. When sequence identity is greater than 90%, models can be constructed with uncertainties approaching the error margins of experimental structure determinations (Sect.13.5). As sequence identity further decreases, model building becomes less accurate. At 50%, the average coordinate error can be afew angstroms. Be­low an identity of 25–30%, recognition of structural re­lationships becomes very problematic.
The vast majority of sequence differences between homologous proteins are located on the surface of the protein in loop regions that are not critical for the folding of the protein backbone (Sect.14.4). Exchanges in the in­terior of the protein can have amuch greater effect on its architecture. They are usually limited to amino acids of similar volume and very similar physicochemical proper-
alignment with the best agreement. Invariable or conserved positions in the sequence are marked in bold. Dashes stand for areas in which other proteins carry additional amino acids (insertions). The red bars underscoring the sequences show preferred helical areas. AAla, CCys,
DAsp, EGlu, FPhe, GGly, HHis, IIle, KLys, LLeu, MMet, NAsn, PPro, QGln, RArg, SSer, TThr, VVal, WTrp, YTyr
ties, such as the exchange of aleucine for an isoleucine. Often the exchange of one amino acid is coupled with the complementary exchange of one or more other amino acids in the immediate vicinity. This is especially true when polar amino acids are exchanged inside the pro­tein, which are internally saturated, e.g., by salt bridges. In the newly mutated protein variants, these amino ac­ids form astable orientation. Since the spatial proximity of amino acid residues in the fold does not necessarily correspond to their sequential proximity in the protein chain, the recognition of such structural relationships is considerably complicated. Mutations in the protein core can lead to expansion, spatial shifts, or twisting of the structural building blocks of the protein.
If the identity is very high, only afew amino acid side chains need to be exchanged. The conformations of the involved side chains can be deduced from acompari­son with the structurally resolved proteins showing these amino acids in asimilar environment. With decreasing identity, insertions and deletions in loop regions, i.e., an ex- pansion or contraction of the polypeptide chain, must be considered. To predict the conformations of these loops during model building, libraries of known protein struc­tures have been compiled. Based on length and sequence, these loops are classied into conformational families. They can be retrieved by the computer and support the construction of the spatial arrangement of amodied loop. The validation of the relevance of these protein models follows empirical rules. It is checked whether the constructed geometry agrees with experimental evi-
Chapter  • Protein Modeling and Structure-Based Drug Design
20
. Fig. 20.5 Left Superposition of the folded structures of the three
cytochromeC proteins from . Fig.20.4 based on a ribbon model: Paracoccus denitricans in blue, Rhodospirillum rubrum in red, and tuna sh in yellow. The cytochromes bind via ahistidine and ame­thionine to an iron–heme center. The structures were determined by X-ray crystallography. Structural deviations occur particularly in the loop regions. Right The same superposition is shown, only here the individual amino acids are color-coded. The same colors in all three ribbon models show identical amino acids at different positions (color coding: Ala: light gray, Val: chartreuse, Gly: white, Ile: bright green,
dence. For example, it must be ensured that hydrophobic groups are oriented inwards and hydrophilic groups are oriented mainly outwards. The contact between amino acid groups is checked, and the chosen torsion angles are compared with those typically observed.
20.6 Secondary Structure Prediction
and Amino Acid Replacement Propensities Support Model Building at Low Sequence Identity
When the sequence identity between the known and modeled protein falls below 30%, determining structural homology becomes increasingly difcult. All additional information must be employed as aresource. An attempt is made to estimate where in the polymer chain of the modeled protein certain secondary structure elements are expected to occur (Sect.14.2). When evaluating the frequency with which individual amino acids occur in helices, pleated sheets, or loops, signicant differences are found. For example, proline is considered a“helix
Leu: olive green, Pro: pink, Phe: violet, Tyr: dark purple, Trp: light vio- let, Asp: dark red, Glu: wine red, Asn: turquoise, Gln: cyan, Lys: blue, His: light blue, Arg: medium blue, Ser: light orange, Thr: dark orange, Cys: light yellow, Met: dark yellow). (7 https://sn.pub/UmXxlS)
breaker.” It occurs at most in the rst turn of ahelix; at other positions, it disrupts the geometry and induces akink. To determine whether aparticular sequence seg­ment folds as ahelix, pleated sheet, or loop, the informa­tion about positional preferences is evaluated for several neighboring amino acids in an overlapping fashion.
Once analyzed, the primary sequence is compared to areference protein of known geometry. Since its 3D structure is known, the assignment of the sequence to the secondary structural elements is straightforward. If not only one, but several 3D structures of members of aho­mologous protein family are known, multiple sequence alignments can be used to construct arepresentative pro­le of the expected secondary structure. This prole then serves as areference for the mutual comparison of the sequences of structurally known and unknown proteins.
For this purpose, in addition to the physicochemical properties of the amino acids, their local conformational properties, their main and side chain orientations, their accessibility for solvent molecules, or their involvement in hydrogen bonds are analyzed. At the same time, the probabilities with which such an amino acid substitu­tion can occur at the DNA sequence level are taken into
. • Secondary Structure Prediction and Amino Acid Replacement Propensities Support Model Building


account (Sect. 32.7). These parameters can be easily determined for proteins with aknown spatial structure. Comparing the structures within aset of homologous proteins yields probabilities for mutual amino acid sub­stitutions. For example, unlike all other amino acids, glycine has no side chain (see pageIX). Therefore, it can adopt conformations in the polymer chain that are ste­rically impossible for other amino acids. The polymer chain assumes such conformations in regions close to the protein surface, where it reverses its orientation. This is where the conformationally exible glycines play an im­portant role. Such glycines that are exposed to the solvent and have unusual torsion angles are largely conserved between the folding of homologous proteins. The con­served glycines can be searched for during the sequence alignment of aprotein to be modeled. They thus pro­vide anchor points for sequence alignment. Many similar rules have been established. They serve as criteria for the recognition of sequence segments of structural impor­tance. They are then applied to the protein sequence to be modeled. Even with relatively low sequence identity, structural homologies between aprimary sequence and aprotein of known 3D structure can be detected in this way. They are used as criteria for homology modeling.
Before the breakthrough in successful experimental
structure determination of G-protein-coupled recep-
tors, the most important group of membrane receptors (Chap.29), it was necessary to rely on their model build­ing. Additional criteria were used for this purpose. The modeling had to ensure that the hydrophobic amino ac­ids in the helical areas were embedded in the membrane and oriented towards the membrane environment. Mean­while, the homology modeling programs have achieved ahigh degree of automation. At the Biozentrum in Basel, the SWISS-MODEL server which transforms submitted sequences fully automatically into 3D structures was set up. The program Modeller from the group of Andrej Šali in San Francisco is able to assemble protein models from the sequences of entire genomes in silico. Although the structures are certainly rough, and many may be incorrect, such an approach allows asearch for similarities in the recognition determinants of proteins. In this way, possible interactions between proteins can be discovered, or com­monalities in metabolic pathways become transparent.
The modeling of proteins gives good results especially when they have ahigh homology. This is given in regions that determine the folding scaffold. The binding pock­ets are mostly located in the loop regions (Sect.14.4). This is where even homologous proteins differ greatly. Therefore, the model constructions do not achieve the desired accuracy in these regions. An improvement can be achieved if aligand is already placed in the assumed binding region during model construction. Model and placement must be optimized in an iterative process using appropriate energy functions.
As mentioned above, the problem of predicting the structure of aprotein from its amino acid sequence us­ing theoretical concepts has not yet been successfully solved. Therefore, the development of methods for pre­diction using data from experimentally determined struc­tures has been the subject of intensive research for many years. Frustratingly, progress has been very incremental. The breakthrough to signicantly improve the quality of predictions came in 2020, when the researchers at the Google subsidiary DeepMind, who had already attracted attention with their articial intelligence (AI)-based pro­grams AlphaGo and AlphaZero for simulating board games, achieved surprising success. Under the leadership of John Jumper, they won the CASP14 competition (Crit­ical Assessment of Protein Structure Prediction) by awide margin with their machine learning-based program Alpha- Fold2.0. The competition, which has been held since 1994, is performed to compare approaches to protein structure prediction. Research groups are provided with amino acid sequences of proteins. Their structures have been exper­imentally determined but not yet published. For these sequences, the participants have to predict structures us­ing their computational models. AlphaFold2.0 is based on aneural network (deep learning). It is trained to learn structure predictions based on experimentally determined protein structures from the PDB database. It evaluates in­put data by incorporating multiple sequence comparisons, evolutionarily related sequences, and physical and geo­metric constraints such as the preferred spatial pairing of amino acid residues. It applies rules similar to those used in the homology modeling described above. The scientists at DeepMind use the “attention mechanism” known from transformer architectures in their neural network. This involves highlighting parts of the input data in the con­text of the overall environmental data, while attenuating others. For language, this would be words in the context and meaning of asentence or text. In this way, the neural network automatically focuses on and weights the critical information. Similar approaches have revolutionized nat­ural language processing in recent years and are the basis for signicant performance improvements in current AI systems for language and translation programs. As Alpha­Fold iterates through the entire structure of the protein to be folded, it simultaneously understands how to locally improve each part of the structure through repeated ap­plication. In doing so, it takes into account information already learned from previous steps. But how AlphaFold ultimately achieves its impressive results is not easy to deduce. With AI approaches, it is by no means trivial to deduce the system’s rules and decision trees after the fact. This is especially true for analyses in which the decisions of the neural network are self-organizing and in which humans no longer intervene during the learning process.
DeepMind left the other participants far behind in the CASP14 competition. But academic research quickly caught up. David Baker’s group in Seattle has since in-
Chapter  • Protein Modeling and Structure-Based Drug Design
20
tegrated DeepMind’s AI approaches into the RoseTTA Fold software, which now delivers equivalent results.
DeepMind has released AlphaFold2.0 for free via acloud computing tool. It can also be installed locally on acom­puter. The free use appeared only logical, as DeepMind has trained its AI on the publicly available structures in the PDB database. With the release, huge databases are now lling up with structures predicted by AlphaFold2.0. There are now more than amillion models available in the PDB. Meanwhile, anew version, AlphaFold3, with enhanced functionality has been released and published in Nature. It is described to be not limited to single-chain proteins, it can also predict the structures of protein com­plexes with DNA, RNA, post-translational modications, and selected ligands and ions. Unfortunately, this time its publication was without freely accessible executables or source code and for noncommercial trial applications it can be accessed through a server. In the interest of free scientic exchange and despite the predictable high level of commercial interest, it would be greatly appreciated if DeepMind could consider a more open approach to the release of the tool. It is believed that only through frequent use and feedback from the community, who pro­vided the required database to train the neural network, will the computer tool achieve further improvements.
It seems reasonable to conclude that AI programs have already made a notable impact on protein structure research, and it seems probable that they will also inu­ence predictions in future drug design. It is important to remember that structures for targeted ligand design require high resolution and the modeled geometries must be accurate to less than an Angstrom. Accurately pre­dicting folding patterns to within afew angstroms may be sufcient for some applications. However, drug design researchers have learned the hard way from arenin model derived from the related endothiapepsin that small de­viations can mislead targeted ligand design (Sect.24.2).
The tools have been met with considerable enthusi­asm and were recognized in 2024 with the Nobel Prize in Chemistry, which was bestowed upon the two leading DeepMind scientists, John Jumper and Demis Hassa­bis. The long-standing achievements of David Baker's group were also acknowledged, particularly given that David Baker has succeeded in predicting entirely novel folded proteins with tailored properties, which were sub­sequently validated through experiment. This opens up new perspectives for countless applications in biochem­istry, structural biology, and drug design.
20.7 Ligand Design: Seeding, Expanding,
and Linking
The next step after the analysis of the binding pocket of the experimentally determined or modeled protein is the actual ligand design. There are several computer-aided design ap-
proaches available to suggest new protein ligands. Adock­ing program can be used to place successively preselected li­gands from adatabase into the binding pocket (. Fig.20.6). Typically, the database is populated with molecular candi­dates that closely resemble common drug-like molecules. This method, which has become known as “virtual screen- ing” (Sect.7.6), can now screen many millions of candidate molecules in afew hours using parallel computers.
Another approach starts with a“seed” in the bind­ing pocket. Starting from this point, the ligand gradually grows into the binding pocket. This is the principle fol­lowed by most de novo design programs. The placement of the rst seed is critical. Such approaches are parti cularly successful when there is aspecic hot spot in the bind­ing pocket from which further optimization starts. Salt bridges to charged amino acids or coordination of metal ion centers are particularly well suited for this approach. This concept has been successfully applied, for example, to the serine proteases trypsin and thrombin (Sect.23.4) and the zinc-containing carbonic anhydrases (Sect.25.7).
Another approach starts with small fragments that are placed into the protein binding pocket. Increasingly, the results of experimental fragment screening (Sect.7.9) are used to provide reliable starting points for further growth of these candidate molecules in the binding pocket. The approach is an ideal symbiosis of experimental work and subsequent computational design. It combines the strengths of both techniques. Another approach is to chemically link several of the tting fragments to­gether. This strategy has been successfully applied, for example, in the SAR-by-NMR method (Sect.7.8). The cover of this book shows an example of how ananomo­lar protein kinaseA inhibitor was developed through several design steps, starting with the crystal structure of a millimolar phenol fragment (see . Fig.02 page XI,7 https://sn.pub/4NppHn).
20.8 Docking Ligands into Binding
Pockets
Docking attempts to t potential protein ligands into abinding pocket using acomputer. Adocking program takes one candidate at atime from aprecompiled library of molecules. For each entry, a3D structure is generated. If aexible molecule is encountered, either multiple con­formations will be generated and docked individually, or they will be generated on the y during docking. The next step is to t each molecule into the binding pocket. First, the structures that cannot bind to the protein are discarded. In addition, other structures that cause ob­vious problems, such as electrostatic repulsion with the protein in the assumed docking mode, are eliminated. Typically, adocking program generates several solutions. These are evaluated on the basis of the generated binding geometries and their afnity is estimated.
. • Docking Ligands into Binding Pockets


. Fig. 20.6 Potential strategies for ligand design. The complete 3D
structures of putative ligands are docked into the binding pocket (left part of the image). The construction of new molecules is sketched in the middle and right part of the picture. Basically there are two pos-
Irwin Kuntz is apioneer in the eld of docking pro­grams; DOCK was developed in his group at UCSF in San Francisco, USA. The original version in 1982 eval­uated only the steric complementarity of ligands and proteins. The shape of the binding pocket was approxi­mated by aset of different spheres so that the pocket was completely lled. Amathematical method was then used to place the test ligands on this distribution of spheres. Complementarity, ameasure of direct protein–ligand contacts, served as ascoring function. Since the rst ver­sion, DOCK has evolved considerably. The program now uses aforce eld for scoring and calculates the contribu­tions for desolvation. Even the placement of the ligands is
sibilities. Afragment can be placed as aseed and other groups can be attached step by step (middle). Alternatively, several small-molecule fragments can be placed in the binding pocket independently of one another and later linked or merged together (right)
exible by considering rotatable bonds. Another docking prototype was developed at the GMD in Bonn, Germany, by Matthias Rarey. The program FlexX was the rst pro­gram that could quickly handle ligand exibility during docking. It decomposes the test ligands into individual fragments and then uses an algorithm very similar to the positioning algorithm implemented in the de novo design program LUDI (Sect.20.10). After placement of the rst building block, the ligand is successively reconstructed in the binding pocket. Different conformers along the rotatable bonds are considered. The program maintains stored tables of preferred torsion angles, similar to those described in Sect.16.6. The energetic evaluation of the
Chapter  • Protein Modeling and Structure-Based Drug Design
20
placement is performed in this step. The program Auto­Dock from the group of Art Olson at Scripps in La Jolla,
San Diego, USA, uses alattice-based algorithm for place­ment. Aforce eld function similar to that of the GRID program (Sect.17.10) is used to place potential values on agrid embedded in the binding pocket. Starting from arandom orientation, the ligand is moved over the grid until an optimum is found. In doing so, it “feels” the inter­action potential with the protein. Since the potential has already been precalculated on the lattice, this evaluation is particularly fast. At the same time, twisting around rotat­able bonds is performed. The program GOLD, developed by Gerrith Jones in the group of Peter Willett in Shefeld, England, also uses alattice for placement. However, the interaction potentials are parameterized using crystal data. GOLD uses agenetic algorithm to optimize the ge­ometry. Over time, aplethora of docking programs has been developed. All follow aslightly different strategy, but are based on the concepts described for the prototypes mentioned above. Some follow the idea that it is better to generate awell-distributed number of rigid ligand con­formers and then dock them quickly as rigid bodies.
Today there are three main problems that impose limits on docking. One is the energetic evaluation of the gener­ated geometries. This will be specially addressed in the next section. Another is that water plays adecisive role in ligand binding (Sect.20.3). Even today no really convinc­ing solution to the handling of water molecules during docking has been found. The third problem is the exible adaptation of the protein (Sect.15.8). Usually there are small adaptations on the side of the protein that slightly change the shape of the binding pocket. In fact, they are big enough to make the docking programs look for the proverbial red herrings (means: they lead to wrong tracks).
20.9 Scoring Functions: Ranking
of Constructed Binding Geometries
Arelevant scoring of the generated binding geometries is essential for all docking and de novo design approaches in structure-based drug design. Among the numerous geometrically plausible placements, only those that rea­sonably approximate the experimentally found situations must be ltered out. The enthalpic and entropic contribu­tions, which determine the afnity of aligand to its target protein, were described in Chap.4. The goal of ascoring function is to quickly estimate the expected binding afnity from agiven interaction geometry. Theoretically, asingle geometry is not sufcient to solve this problem. Molar en­ergies are determined by anite set of conformations. They are distributed over an “ensemble” of multiple states. These states are differently populated according to their energy content. One group of methods tries to account for this fact in the calculations. This is the most theoretically sound approach. The energy contributions of the ensemble (usu-
ally taken from the trajectories of amolecular dynamics simulation, Sect.15.7) are summed. The Gibbs free energy G (Sect.4.3) can be estimated from the resulting parti­tion function. However, the computations required for such evaluations are very time consuming. As aresult, they are usually not feasible for screening large amounts of data in astructure-based drug design session.
Instead, regression-based scoring functions are used as an alternative approach. Assuming that aparticular state is predominantly populated, it may be justied to consider only one state (or conformation) in the scoring function. The enthalpy and entropy contributions that most likely de­termine the binding afnity are considered. The approach is reminiscent of setting up aQSAR equation (Sect.18.2). The terms are combined into an energy function. Molecular descriptors that correctly reect the contributions to these terms are sought. In doing so, the erroneous assumption is made that the individual contributions to the descrip­tion of the free enthalpy are additive (Sect.4.10). The indi- vidual terms of the equation are each given an adjustable weighting factor. As with QSAR equations, amathematical technique is used to optimally t these weighting factors to atraining dataset. This set consists of crystallographically determined protein–ligand complexes for which experimen­tal binding afnities are available.
Athird approach follows aso-called knowledge-based concept. As already discussed in Sect.17.10, the frequency of individual contact geometries in the crystal structures of protein–ligand complexes is evaluated. Akind of “nor­mal distribution” is dened as areference state. Then, all contacts that occur more frequently than average are classied as energetically favorable. All contacts occurring less frequently than the average are classied as unfavor­able. These scores are then applied to all contacts present in acomputationally generated protein–ligand complex and combined for aranking. So far, afunction derived in this way can be used for the relative energetic ranking of aset of ligands binding to the same reference protein. However, afnity prediction will also be successful here if training is performed against areference dataset of known geometry and binding afnity in amanner analogous to regression-based scoring functions. The evaluation with the regression-based or knowledge-based function is very fast. In the meantime, alarge number of scoring functions have been developed. So far, none of them has proven to be ideal and generally applicable. Therefore, in each case it has to be checked which function provides the best per­formance for the protein under investigation.
20.10 De Novo Design: From LUDI
to the Automated Assembly of Novel Ligands
The rst program for stepwise de novo design was GROW by Jeffrey Howe and Joseph Moon at Upjohn. It focuses
. • The Feasibility of Designing Ligands In Silico


. Fig. 20.7 Concept of the program LUDI for de novo design of
protein ligands. In the rst step, the interaction sites are determined (left). Donor sites are represented by blue lines, acceptor sites by red lines. The green dots symbolize lipophilic sites. Subsequently, small molecules from adatabase are tted into the binding pocket in that
on peptides as lead structures. An amide group is posi­tioned in afavorable orientation in the binding pocket. The next amino acids are added to the starting amide group in astepwise fashion. At each step, alarge number of different conformations of all 20proteinogenic amino acids are attached to the seed on the y. The “best” solu­tions for each are followed. In this way, GROW constructs apeptide ligand of increasing length in the binding pocket.
In the early 1990s, Hans-Joachim Böhm developed the program LUDI at BASF, Ludwigshafen, Germany. The idea was to read small molecules or molecular frag­ments from adatabase with precalculated spatial geom­etries and position them in the binding pocket so that hydrogen bonds are formed with the protein and hydro­phobic pockets are lled with nonpolar groups. As in­put, the program requires the coordinates of the protein and alibrary of 3D structures of fragments or drug-like molecules.
The precalculation of interaction sites is crucial.
These are placed in the binding pocket around the
amino acid residues in the form of tting points or di­rectional vectors (. Fig.20.7). The program uses rules derived from the nonbonded interactions found in the crystal packing of small organic molecules (Sects.14.7 and17.10). LUDI then extracts small molecules or mo­lecular fragments from the 3D library. For each entry, an attempt is made to position it in the binding pocket of the protein so that as many of these interaction sites as possible are satised (. Figs.20.7 and17.12). All suc­cessfully placed fragments are then ranked. The scoring function used takes into account the number and quality of H-bonds and ionic interactions formed, hydrophobic contact surfaces shared by the protein and ligand, and unfavorable contributions due to the number of rotatable bonds in the ligands. An example of the successful appli­cation of this program is described in Sect.21.5.
As the rst prototype, LUDI was the gold standard for many de novo design programs that were developed later. These approaches implemented enhanced scoring func­tions and improved fragment libraries. The programs were also taught synthesis rules so that the chemical accessi-
they are matched with the interaction sites (middle). Finally LUDI can chemically link groups or parts of molecules to larger structures to match the remaining unsatised interaction sites and to ll the entire binding pocket (right)
bility of the generated molecules was taken into account. The search space of the programs has also been expanded to include multiple conformations and congurations.
Ade novo design program is an idea generator. Its value is, of course, determined by the concepts that went into its development. On the other hand, its value also depends strongly on the user and how the suggestions of such aprogram are interpreted and used for further design.
20.11 The Feasibility of Designing
Ligands In Silico
Certainly, many examples have been presented that demonstrate the power of de novo design, virtual screen­ing, and docking. The example described in Chap.21 was successful because of the use of such methods. However, it is critical that the computational methods are tightly integrated into an iterative process of synthesis, biologi­cal testing, and experimental structure determination. It is important to keep in mind that not all hits from com­putational screening are based on the correct assump­tions, and not all hits are picked up for the right reasons.
The predictive power of available methods is still lim­ited. The synthetic accessibility of aproposed molecule is often not sufciently taken into account, the exibility of the protein is frequently neglected, and the methods for estimating binding afnity are still too inaccurate. This is because the process and factors responsible for molecular recognition and ligand binding are still poorly under­stood. The correct description of solvation effects, the involvement of water molecules in the binding process, and the change of protonation states are major problems. The contribution of hydrogen bonding to binding afn­ity is still only an estimate, despite efforts to the contrary. In the case of lipophilic interactions, it can at least be as­sumed that the lling of an unoccupied lipophilic pocket with additional nonpolar substituents is in most cases accompanied by an increase in binding afnity.
How to account for the changes in binding entropy that contribute to the free energy of ligand binding is still
Chapter  • Protein Modeling and Structure-Based Drug Design
20
insufciently considered. Certainly, there is evidence that the oversimplied assumption that entropic contribu­tions will be constant within aset of congeneric ligands is denitely not valid.
However, there are other fundamental limitations of this approach. The most important is that the technique is limited to optimizing direct interactions with the protein. Successful binding to atarget protein is essential for any drug. However, to be suitable as adrug, additional require­ments must be met. These include good selectivity, meta­bolic stability, adequate duration of action, low addictive potential, and negligible toxicity. Today, at least the selec­tivity of acompound towards members of astructurally related protein family can be estimated with some certainty.
Fully automated molecular design on acomputer is not yet possible, nor is it likely to be in the long term. The methods of structure-based design are valuable as idea generators. The resulting suggestions must be checked and modied if necessary. Time will tell whether these methods will gradually approach the “holy grail” of drug design: the design of drug molecules from scratch.
20.12 Synopsis
In structure-based drug design, attempts are made
-
to design small-molecule ligands by docking them
directly into the binding pocket of atarget protein.
This requires a3D structure of the reference protein.
The goal is to optimally ll the binding pocket with
aligand by satisfying nonbonded interactions with
the functional groups of the binding-site residues.
Structure-based design starts with adetailed analysis
-
of the binding pocket to elucidate hot spots for puta-
tive interactions with the protein. Either experimental
methods or computational tools can be used to per-
form an active site mapping with molecular probes or
small solvent-like molecules.
In an iterative process of structure determination,
-
modeling of modied ligands, docking and screen-
ing, synthesis, and biological testing, the properties
of small-molecule ligands are improved to optimize
binding to the target protein.
Databases have been developed to retrieve and com-
-
pare structural information about the exponentially
growing body of structural data on protein–ligand
complexes. They allow comparison of binding poses,
active-site interaction geometries, protein–ligand
binding motifs, and the original solvation structures
in the protein’s binding pocket.
Proteins can be compared in terms of their exposed
-
binding pockets. The shape and the exposure of
groups which have particular physicochemical prop-
erties in binding pockets are compared and help to
design small-molecule ligands with the desired selec-
tivity. Ideas for isosteric replacements on the ligand scaffold can also be generated in this way.
If an experimentally determined structure of the tar-
-
get protein is unavailable, ahomology model can be constructed by using arelated protein of known ar­chitecture as atemplate. The accuracy and the success of such homology modeling depend strongly on the sequence homology with the template structure.
Tools for secondary structure prediction and amino
-
acid replacement propensities have been developed to improve the reliability of the sequence assignment of proteins to the 3D structure of the reference template. Recently, AI programs have been developed based on neural networks to generate 3D structures from se­quence data. They were trained to learn structure pre­dictions based on experimentally determined protein structures.
Ligand design approaches screen large databases of
-
candidate molecules by docking them into the bind­ing pocket.
Alternatively, de novo design starts with asmall mol-
-
ecule seed or fragment and grow them into putative ligands in the pocket. Two nonoverlapping fragments can also be linked together to form alarger ligand with improved binding properties.
The geometry of aconstructed protein–ligand com-
-
plex must be evaluated in terms of the expected binding afnity in all structure-based design strate­gies. Alarge variety of scoring functions are used to predict the binding afnity based on the geometry of the formed complex.

Bibliography and Further Reading

General Literature
C. Branden, J. Tooze, Introduction to Protein Structure, Garland Pub-
lishing, Inc. New York, 1991, 2nd Edition, 1999
T. J. P. Hubbard, A. M. Lesk, Modelling Protein Structures, in Com-
puter Modelling in Molecular Biology, J. M. Goodfellow, Eds., VCH, Weinheim, 1995
C. Hutchins, J. Greer, Comparative Modeling of Proteins in the Design
of Novel Renin Inhibitors, Crit. Reviews Biochem. Molec. Biol., 26, 77–127 (1991)
P. Goodford, Drug Design by the Method of Receptor Fit, J. Med.
Chem., 27, 557–564 (1984)
J. Greer, J. W. Erickson, J. J. Baldwin, M. D. Varney, Application of
the Three-Dimensional Structures of Protein Target Molecules in Structure-Based Drug Design, J. Med. Chem., 37, 1035–1054 (1994)
C. R. Beddell, Ed., The Design of Drugs to Macromolecular Targets,
Wiley, Chichester, 1992
I. D. Kuntz, Structure-Based Strategies for Drug Design and Discovery,
Science 257, 1078–1082 (1992)
S. Borman, New 3D Search and De Novo Design Techniques Aid Drug
Development, Chem. & Eng. News, 10, 18–26 (1992)
Y. C. Martin, 3D Database Searching in Drug Design, J. Med. Chem.,
35, 2145–2154 (1992)
I. D. Kuntz, E. C. Meng, B. K. Shoichet, Structure-Based Molecular
Design, Acc. Chem. Res., 27, 117–123 (1994)
Bibliography and Further Reading


H. J. Böhm, Ligand Design, in: 3D QSAR in Drug Design, H. Kubinyi,
Eds., Escom, Leiden, 1993, pp. 386–405
K. Müller, Ed., De Novo Design, Persp. Drug Discov. Design, Vol. 3,
Escom, Leiden, 1995
H. J. Böhm, G. Schneider, Molecular Recognition in Protein–Ligand
Interactions (Vol. 19, in Methods and Principles in Medicinal
Chemistry, R. Mannhold, H. Kubinyi und G. Folkers, Eds.), Wi-
ley-VCH, Weinheim, 2006 G. Schneider, K. H. Baringhaus, Molecular Design, Wiley-VCH, Wein-
heim, 2008 M. Eguida, D. Rognan, Estimating the Similarity between Protein
Pockets, Int. J. Mol. Sci., 23, 12462 (2022)
K. A. Carpenter, R. B. Altman, Databases of ligand-binding pockets
and protein-ligand interactions, 23, 1320–1338 (2024)
Special Literature
C. R. Bedell, P. J. Goodford, F. E. Norrington, S. Wilkinson, R. Woot-
ton, Compounds Designed to Fit a Site of Known Structure in
Human Haemoglobin, Brit. J. Pharmacol., 57, 201–209 (1976) M. Hendlich, A. Bergner, J. Günther, G. Klebe, Design and Devel-
opment of Relibase—a Database for Comprehensive Analysis of
Protein–Ligand Interactions, J. Mol. Biol., 326, 607–620 (2003) S. Schmitt, D. Kuhn, G. Klebe, A new method to detect related function
among proteins independent of sequence and fold homology, J.
Mol. Biol., 323, 387–406 (2002)
A. Weber etal., Unexpected Nanomolar Inhibition of Carbonic An-
hydrase by COX-2-Selective Celecoxib: New Pharmacological
Opportunities due to Related Binding Site Recognition, J. Med.
Chem., 47 550–557 (2004)
C. Gerlach etal., KNOBLE: a knowledge-based approach for the de-
sign and synthesis of readily accessible small molecule chemical
probes to test protein binding, Angew. Chem. Int. Ed. Engl., 46,
9105–9109 (2007)
I. Kufareva, A. V. Ilatovskiy, R. Abagyan, Pocketome: An Encyclope-
dia of Small-molecule Binding Sites in 4D. Nucl. Acid. Res., 40,
D535–D540 (2012)
K. N. Allen, C. R. Bellamacina, X. Ding, C. J. Jeffery, C. Mattos, G.
A. Petsko, and D. Ringe An Experimental Approach to Mapping
the Binding Surfaces of Crystalline Proteins, J. Phys. Chem. 100,
2605–2611 (1996) J. Overington, M. S. Johnson, A. Sali and T. L. Blundell, Tertiary Struc-
tural Constraints on Protein Evolutionary Diversity: Templates,
Key Residues and Structure Prediction, Proc. Royal Soc. Lond. B
241, 132–145 (1990) M. S. Johnson, J. P. Overington, T. L. Blundell, Alignment and Search-
ing for Common Protein Folds Using a Data Bank of Structural
Templates, J. Mol. Biol., 231, 735–752 (1993)
A. Waterhouse etal., SWISS-MODEL: Homology Modelling of Pro-
tein Structures and Complexes, Nucl. Acid. Res., 46, W296–W303
(2018)
A. Sali and T. L. Blundell, Denition of General Topological Equiv-
alence in Protein Structures, J. Mol. Biol., 212, 403–428 (1990)
A. Sali and T. L. Blundell, Comparative Protein Modelling by Satis-
faction of Spatial Restraints. J. Mol. Biol., 234, 779–815 (1993)
J. Jumper etal., Highly accurate protein structure prediction with Al-
phaFold, Nature 596, 583–589 (2021) J. Abramson, J. Adler, J. Dunger etal. Accurate structure prediction of
biomolecular interactions with AlphaFold3. Nature 630, 493–500
(2024)
M. Baek etal. Accurate prediction of protein structures and interac-
tions using a three-track neural network, Science, 373, 871–876
(2021)
E. Callaway, It will change everything: DeepMind’s AI makes gigantic
leap in solving protein structures, Nature 588, 203–204 (2020)
M. Hibert, S. Trumpp-Kallmeyer, J. Hoack, A. Bruinvels, This is Not
a G-Protein-Coupled Receptor, Trends Pharm. Sci., 14, 7–12 (1993)
J. Hoack, S. Trumpp-Kallmeyer and M. Hibert, Re-evaluation of
Bacteriorhodopsin as a Model for G-Protein-Coupled Receptors, Trends Pharm. Sci., 15, 7–9 (1994)
R. Henderson, J. M. Baldwin, T. A. Ceska, F. Zemlin, E. Beckmann, K.
H. Downing, Model of the Structure of Bacteriorhodopsin Based on High-Resolution Electron Cryo-Microscopy, J. Mol. Biol., 213, 899–929 (1990)
G. F. X. Schertler, C. Villa, R. Henderson, Projection Structure of Rho-
dopsin, Nature, 362, 770–772 (1993)
J. Travis, Proteins and Organic Solvents Make an Eye-Opening Mix,
Science, 262, 1374 (1993)
C. S. Ring etal., Structure-based Inhibitor Design by Using Protein
Models for the Development of Antiparasitic Agents, Proc. Natl. Acad. Sci., 90, 3583–3587 (1993)
I. D. Kuntz, J. M. Blaney, S. J. Oatley, R. Langridge, T. E. Ferrin, A
Geometric Approach to Macromolecule-Ligand Interactions, J. Mol. Biol., 161, 269–288 (1982)
M. Rarey, B. Kramer, T. Lengauer, G. Klebe, A Fast Flexible Docking
Method using an Incremental Construction Algorithm, J. Mol. Biol., 261 470–489 (1996)
D. S. Goodsell, A. J. Olson, Automated Docking of Substrates to Pro-
teins by Simulated Annealing, Proteins, 8, 195–202 (1990)
G. Jones, P. Willett, R. C. Glen, A. R. Leach, R. Taylor, Development
and Validation of a Genetic Algorithm for Flexible Docking, J. Mol. Biol., 267, 727–748 (1997)
Pocketome: https://ngdc.cncb.ac.cn/databasecommons/database/id/501
(Last accessed Nov. 19, 2024) SiteMine: https://uhh.de/naomi (Last accessed Nov. 19, 2024) Explore Computed Structure Models Alongside PDB Data: https://
www.rcsb.org/news/6304ee57707ecd4f63b3d3db (Last accessed
Nov. 19, 2024) SWISS-Model web server: https://swissmodel.expasy.org/ (Last ac-
cessed Nov. 19, 2024) Modeller: https://salilab.org/modeller/ (Last accessed Nov. 19, 2024) AlphaFold: https://neurosnap.ai/academic?gclid=CjwKCAiAqNS-
sBhAvEiwAn_tmxbN3V2G2yrU1LJABBGDpMRMS-
gu_2u6c-8v76nQEU-tDj1OMPXPFoQxoCX7gQAvD_BwE (Last
accessed Nov. 19, 2024)