Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5319_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
29.08.2026
Размер:
92 Мб
Скачать
. • Synopsis


developed the Feature-Trees method, which can search large databases using topological criteria. However, it does not compare the connectivity of chemical formu­las. Rather, the database entries are rst classied by the topological sequences of certain features, such as the presence of an H-bond donor group or ahydrophobic cyclic molecular building block. In this way, molecules can be compared and candidates with pharmacophore properties in acomparable topological connectivity can be found extremely quickly.
Databases that contain 3D molecular geometries al­low the search for the spatial pattern of the pharmaco­phore. For example, the Cambridge Structural Data­base of crystal structures of small organic molecules (Sect.13.9) can be used for such asearch. Molecules are found with experimentally determined geometries that satisfy the pharmacophore. In the search for ligands for HIV protease (Sect.24.3), apharmacophore pattern was derived from the known crystal structure of the enzyme, and the Cambridge Database was searched for molecules that match this pattern. The results of this search are presented in Sect.24.4 (. Fig.24.16) in detail. It inspired the researchers at Dupont–Merck with rst ideas that led to the development of an entirely new class of nonpep­tidic HIV protease inhibitors.
Today, databases containing 3D structures of mol­ecules generated from 2D structural formulas are com­monly used alongside experimental structural databases. In other approaches, the spatial structure of candidate molecules is generated on the y during the search (Sect.15.2). Here, as with most entries in the Cambridge Database, each molecule exists in only one conforma­tion. However, molecules can adopt many different con­formations (Chap.16). It is usually the exception rather than the rule that aexible molecule will be stored in the “correct” conformation required for the search. There­fore, conformational exibility must be taken into ac­count during the search. An exhaustive search, such as the “active analog approach,” would require too much computational time. Therefore, fast algorithms have been developed to identify whether certain pharmacophoric groups on the molecules could fall within predened dis­tance ranges. It is sufcient to estimate the minimum or maximum achievable distances. This concept has been realized, for example, in the program UNITY from the company Tripos. It is possible to start from adatabase that contains several precalculated conformers. In this case, it is very important that the distribution of the con­formers in the conformational space is as representative as possible (Sect.16.6). The individual conformers will then be checked to see whether they t the dened phar­macophore pattern. This concept has been implemented in the database search engine Catalyst from Accelrys. Asimilar concept is followed by the program Ligand- Scout of Thierry Langer in Vienna, Austria and Gerhard Wolber in Berlin, Germany.
Such database searches are not expected to imme­diately yield candidates for clinical testing. However, as agenerator of ideas, they can lead drug discovery scientist to new lead structures and take his or her syn­thesis plans down completely different paths. Database searches are now widely used as part of virtual screening (Sect.7.6). This involves screening proprietary collec­tions of compounds or searching compilations of com­mercially available compounds. John Irwin and Brian Shoichet at UCSF in San Francisco, USA, have taken the initiative to continuously store commercially available compounds in the ZINC database and make them avail­able for database searching. Preset lters help to extract the desired subset from the collection of several million compounds for the user’s own search. Hits found in this way can be obtained commercially and tested experimen­tally in an assay. Many candidates for new lead structures have already been discovered through this lead discovery by shopping (for an example, see Sect.21.7).
17.12 Synopsis
The structure of the binding pocket determines which
-
functional groups are necessary on the ligand side for
successful protein binding. Either the ligand or the
protein structure can be used as the starting point
from which apharmacophore is derived.
The superposition of active and inactive small-mole-
-
cule ligands from aseries of related compounds can
be used to dene the allowed and forbidden areas in
ahypothetical binding pocket. Logical operations of
volume differences are indicative for the design of op-
timized ligands.
Flexible molecules that can adopt different conforma-
-
tions present aspecial challenge in the mutual super-
positions. The molecules must be energy-minimized
as part of the superposition procedure or, alterna-
tively, multiple conformations must be evaluated.
Alternatively, aset of molecules can be superimposed
-
by assigning pharmacophoric groups, and through
systematic rotations about all open-chain single
bonds acommon alignment is found in the “active
analog approach.”
Care must be taken to not be deceived by molecules
-
that look similar with respect to their chemical for-
mulas. Instead, the interacting functional groups are
important for the molecular recognition at the bind-
ing pocket and not the scaffold itself. The role of wa-
ter in the binding must not be underestimated.
Molecular recognition properties can also be consid-
-
ered to mutually superimpose molecules.
The synthesis of a structurally rigid analogue (or
-
analogues) can help to dene and validate the phar-
macophore assignment and the determination of the
biologically active conformation.
Chapter  • Pharmacophore Hypotheses and Molecular Comparisons
1
17
Binding “hot spots” can be found by examining the
-
protein by mapping the binding pocket with small molecules or molecular probes with different proper­ties. These give some ideas as to what sort of molecule might successfully bind to the target protein.
The Cambridge Database of crystal structures pro-
-
vides valuable insights into preferred interaction geometries and motifs. Such information is of high relevance for protein–ligand complexes because the forces that are responsible for crystal packing are the same as for nonbonding interactions between active substances and proteins.
A variety of databases are available that can be
-
screened by using a3D pharmacophore as asearch query. Usually, commercially available compounds are screened rst. If they show activity on acertain protein of interest, they can be purchased and tested, and will hopefully provide astarting point for further lead discovery.

Bibliography and Further Reading

General Literature
T. Langer and R. D. Hoffmann, Pharmacophores and Pharmacophore
Searches (Vol. 32 in Methods and Principles in Medicinal Chemis­try, R. Mannhold, H. Kubinyi and G. Folkers, Eds.), Wiley-VCH, Weinheim (2006)
G. R. Marshall, Computer-Aided Drug Design, in: Computer-Aided
Molecular Design, W. G. Richards, Ed., IBC Technical Services Ltd, London, pp. 91–104 (1989)
G. Klebe, Structural Alignment of Molecules, in: 3D-QSAR in Drug
Design. Theory, Methods and Application, H. Kubinyi, Ed., ES­COM, Leiden, pp. 173–199 (1993)
Y. C. Martin, 3D Database Searching in Drug Design, J. Med. Chem.,
35, 2145–2154 (1992)
Special Literature
C. G. Wermuth, C. R. Ganellin, P. Lindberg, L. A. Mitscher, Glossary
of terms used in medicinal chemistry (IUPAC Recommendations
1998), Pure Appl. Chem., 70, 1129–1143 (1998)
W. E. Klunk, B. L. Kalman, J. A. Ferrendelli and D. F. Covey, Com-
puter-Assisted Modeling of the Picrotoxinin and γ-Butyrolactone Receptor Site, Mol. Pharmacol., 23, 511–518 (1983)
M. F. Mackay and M. Sadek, The Crystal and Molecular Structure of
Picrotoxinin, Austr. J. Chem., 36, 2111–2117 (1983)
G. R. Marshall, C. D. Barry, H. E. Bossard, R. A. Dammkoehler and
D. A. Dunn, The Conformational Parameter in Drug Design: The Active Analog Approach, in: Computer-Assisted Drug Design, ACS Symp. Series 112, E. C. Olson and R. E. Christoffersen, Eds., Amer. Chem. Soc., Washington DC., pp. 205–226 (1979)
D. Mayer, C. B. Naylor, I. Motoc and G. R. Marshall, A Unique Ge-
ometry of the Active Site of Angiotensin-Converting Enzyme Consistent with Structure-Activity Studies, J. Comput.-Aided Mol. Design, 1, 3–16 (1987)
D. J. Kuster and G. R. Marshall, Validated Ligand Mapping of ACE
Active Site, J. Comput.-Aided Mol. Design, 19, 609–615 (2005)
J. T. Bolin, D. J. Filman, D. A. Matthews, R. C. Hamlin and J. Kraut,
Crystal Structure of Eschericha coli and Lactobacillus casei Dihy­drofolate Reductase Rened at 1.7 Å Resolution, J. Biol. Chem., 257, 13650–13662 (1982)
S. K. Kearsley and G. M. Smith, An Alternative Method for the Align-
ment of Molecular Structures: Maximizing Electrostatic and Steric
Overlap, Tetrahedron Comput. Methodol., 3, 615–633 (1990) P. C. D. Hawkins, A. G. Skillman and A. Nicholls. Comparison of
shape-matching and docking as virtual screening tools. J. Med.
Chem., 50, 74–82 (2007) https://www.eyesopen.com/rocs (Last ac-
cessed Nov. 18, 2024) C. Lemmen, T. Lengauer and G. Klebe, FlexS: A Method for Fast Flex-
ible Ligand Superposition, J. Med. Chem., 41, 4502–4520 (1998)
https://www.biosolveit.de/wp-content/uploads/2021/01/FlexS.pdf
(Last accessed Nov. 18, 2024) G. Klebe, T. Mietzner and F. Weber, Different Approaches Toward an
Automatic Structural Alignment of Drug Molecules: Applications
to Sterol Mimics, Thrombin and Thermolysin Inhibitors, J. Com-
put.-Aided Mol. Design, 8, 751–778 (1995) W. Seidel, H. Meyer, L. Born, S. Kazda and W. Dompert, Rigid Cal-
cium Antagonists of the Nifedipine-Type: Geometric Require-
ments for the Dihydropyridine Receptor, in: QSAR as Strategies
in the Design of Bioactive Compounds, J. K. Seydel, Ed., VCH,
Weinheim, pp. 366–369 (1984) P. Goodford, Drug design by the method of receptor t, J. Med. Chem.,
27, 557–564 (1984) https://www.moldiscovery.com/software/grid/
(Last accessed Nov. 18, 2024) I. J. Bruno, J. C. Cole, J. P. Lommerse, R. S. Rowland, R. Taylor and
M. L. Verdonk, IsoStar: a library of information about nonbonded
interactions, J. Comput.-Aided Mol. Design, 11, 525–537 (1997) IsoStar: https://www.ccdc.cam.ac.uk/solutions/software/isostar/ (Last
accessed Nov. 18, 2024) M. L. Verdonk, J. C. Cole, R. Taylor, SuperStar: A Knowledge-based
Approach for Identifying Interaction Sites in Proteins, J. Mol.
Biol., 289, 1093–1108 (1999) SuperStar: https://www.ccdc.cam.ac.uk/solutions/software/superstar/
(Last accessed Nov. 18, 2024) G. Klebe, The Use of Composite Crystal-Field Environments in Molec-
ular Recognition and the “De-Novo” Design of Protein Ligands,
J. Mol. Biol., 237 212–235 (1994) R. Taylor and P. A. Wood, A Million Crystal Structures: The Whole
Is Greater than the Sum of Its Parts, Chem. Rev., 119, 9427–9477
(2019) H. Gohlke, M. Hendlich and G. Klebe, Knowledge-based Scoring
Function to Predict Protein-Ligand Interactions, J. Mol. Biol., 295,
337–356 (2000) https://www.fz-juelich.de/en/ibg/ibg-4/expertise/da-
tabases-softwares-and-webservers-in-the-gohlke-group/drugscore
(Last accessed Nov. 18, 2024) A. Caisch, A. Miranker, and M. Karplus, Multiple copy simultaneous
search and construction of ligands in binding sites: application
to inhibitors of HIV-1 aspartic proteinase, J. Med. Chem., 36,
2142–2167 (1993) D. Joseph-McCarthy, J. M. Hogle and M. Karplus, Use of the multiple
copy simultaneous search (MCSS) method to design a new class of
picornavirus capsid binding drugs, Proteins, Struct, Funct, Bioin-
form., 29, 32–58 (1997) M. Rarey and J. S. Dixon, Feature trees: A new molecular similarity
measure based on tree matching, J Comput.-Aided Mol. Des.,
12, 471–490 (1998) https://www.biosolveit.de/wp-content/up-
loads/2022/03/FTrees.pdf (Last accessed Nov. 18, 2024)
G. Wolber and T. Langer, LigandScout: 3-D pharmacophores derived
from protein-bound ligands and their use as virtual screening l-
ters, J. Chem. Inf. Model., 45, 160–169 (2005) https://ligandscout.
software.informer.com/ (Last accessed Nov. 18, 2024)
J.J. Irwin and B.K. Shoichet, ZINC—A Free Database of Commer-
cially Available Compounds for Virtual Screening, J. Chem. Inf.
Model., 45, 177–182 (2005) https://zinc15.docking.org/ (Last ac-
cessed Nov. 18, 2024)
Quantitative Structure–
Activity Relationships
Contents
18.1 How It All Began: Structure–Activity Relationships of Alkaloids – 274
18.2 From Richet, Meyer, and Overton to Hammett and Hansch – 274
18.3 The Determination and Calculation of Lipophilicity – 275
18.4 Lipophilicity and Biological Activity – 275
18.5 The Hansch Analysis and the Free–Wilson Model – 276


18.6 Structure–Activity Relationships of Molecules in Space – 278
18.7 Structural Alignment as aPrerequisite for the Relative Comparison of Molecules – 278
18.8 Binding Anities as Compound Properties – 278
18.9 How Is aCoMFA Analysis Performed? – 279
18.10 Molecular Fields as Criteria of aComparative Analysis – 280
18.11 3D-QSAR: Correlation of Molecular Fields with Biological Properties – 280
18.12 Results of aComparative Molecular Field Analysis and Their Graphical Interpretation – 282
18.13 Scope, Limitations, and Possible Expansions of the CoMFA Analysis – 283
18.14 A Glimpse Behind the Scenes: Comparative Molecular Field Analysis of Carbonic Anhydrase Inhibitors – 284
18.15 Synopsis – 287
Bibliography and Further Reading – 288
© The Author(s), under exclusive license to Springer-Verlag GmbH, DE, part of Springer Nature 2024 G. Klebe, Drug Design, https://doi.org/10.1007/978-3-662-68998-1_18
18
Chapter  • Quantitative Structure–Activity Relationships
Quantitative structure–activity relationships, QSAR (usually pronounced [kyü:sar]), attempt to describe and quantify the correlation between chemical structure and biological activity. The investigated substances should come from achemically uniform series and must interact with the same biological target. They should also dis­play the same mode of action. For example, structurally analo gous inhibitors of aparticular protein can be com­pared among themselves, but not different blood pressure lowering drugs that have diverse modes of action on dif­ferent target proteins. The correlation between biological activity and physicochemical properties is always related to relative potency in atest model, but not to different modes of action.
The basis for quantitative correlations between chem­ical structure and biological effect is the entirely reason­able assumption that the differences in physicochemical properties are responsible for the relative potency of the interactions of the drug with biological macromolecules. In arst approximation, these are assumed to contrib­ute additively to the afnity of adrug for its receptor. The concept of describing the biological activity of sub­stances with mathematical models is derived from this approach.
For the system under investigation, it can be assumed that the simpler it is, the more likely it is that aquanti­tative structure–activity relationship can be derived. To acertain extent, this is true for in vitro systems, such as enzyme inhibition or receptor binding, where the assay records only the binding of acompound to aprotein. The more complex the system, e.g., effects on the central nervous system of an animal after oral administration, the more different processes have to be considered. In this case, absorption, distribution, blood–brain barrier pen­etration, transport to the target tissue, metabolism, and excretion overlap with one another and with the actual effect at the receptor. In principle, an individual struc­ture–activity relationship is required for each of these events. In order to establish valid and relevant models for each of these steps, appropriate test systems are needed to study the different steps separately. In favorable cases, it may be possible to characterize acomplex multistep process by asingle equation. This will only be feasible if one step, e.g., the penetration of the blood–brain barrier, dominates the entire structure–activity relationship.
18.1 How It All Began: Structure–Activity
Relationships of Alkaloids
The South American arrow poison tubocurarine (Sect.6.2) was the rst therapeutic principle for which the exact mode of action was elucidated. In 1852, Claude Bernard realized that this quaternary alkaloid causes muscle paralysis, but that both the nerve and the mus­cle remain independently excitable. Curare must, there-
. Fig. 18.1 The protonation of atertiary amine depends on the pH
value of the medium (left). On the other hand, the quaternization of anitrogen atom leads to apermanently positively charged compound (right)
fore, act on the coupling between nerve and muscle. The Scottish pharmacologists Alexander Crum-Brown and Thomas Fraser studied in more detail whether the qua­ternization of the nitrogen atom of various alkaloids (. Fig.18.1) inuences their biological effects. In 1868, on the basis of very different effects observed before and after the transformation of alkaloids, they formulated ageneral equation to describe structure–activity relation­ships (Eq.18.1).
(18.1)
This equation is ingeniously simple, but it says only thatΦ, the biological activity, is a function ofC, the chemical structure. At that time, the tetrahedral struc­ture of the carbon atom had not been elucidated, and the composition of many organic compounds, especially complex natural products, was completely unknown.
18.2 From Richet, Meyer, and Overton
to Hammett and Hansch
In 1893, Charles Richet published astudy on the toxicity of organic compounds. Comparing the water solubility of ethanol, diethyl ether, urethane, paraldehyde, amyl alcohol, and absinthe extract(!) to the lethal dose in the dog, he concluded plus ils sont subles, moins ils sont tox- iques, that is, the better the solubility, the less the toxicity. This was the rst evidence of alinear inverse relationship between water solubility and biological activity.
At the turn of the last century, the pharmacologist Hans Horst Meyer and the botanist Charles Ernest Over ton independently established the lipid theory of anesthe- sia, which combines three important statements:
All chemically unreactive substances that are lipo-
-
philic and can be distributed in biological systems
have anesthetic effects.
The biological effect occurs in nerve cells because fats
-
play an important role in their function.
The relative potency of anesthetics depends on their
-
partition coefcient (Sect.19.2) in amixture of fats
and water.
-
RX
−log K
RH
−log P
H
. • Lipophilicity and Biological Activity


The work of Crum-Brown, Fraser and Richet or the contribution of Meyer and Overton can be considered as the origin of quantitative structure–activity relation­ships. In fact, after the formulation of the anesthetic theory, numerous other linear and later nonlinear depen­dencies on the lipophilicity, the “fat afnity” of drugs, were found. However, these were all relatively unspecic “membrane” effects.
In the mid-1930s, Louis P. Hammett formulated arelationship between the electronic properties of sub­stituents and the reactivity of aromatic compounds. Ac­cording to this relationship, the relative contributions of electron-withdrawing and electron-donating substituents to the electron density of the aromatic ring are always constant. They are determined by the electronic param­eter of the substituent, the Hammett constant,σ. Elec­tron-accepting substituents with positive σvalues are, among others, the nitro group, the cyano group, and the halogens. Electron-donating substituents with negative σvalues are hydroxyl and amino groups, the methoxy group, and alkyl substituents. Acceptor substituents enhance the acidity of benzoic acids and phenols, they reduce the basicity of anilines, and they accelerate the basic hydrolysis of benzoic ethers. Electron-donating substituents exert an opposite inuence.
However, an individual reaction constantρ must be applied for each reaction type of aromatic compounds. By using Eq.18.2, later generally called the Hammett equation, the equilibrium constantK for an arbitrary reaction can be calculated fromρ andσ. R–X and R–H represent the relevant aromatic compounds substituted with the groupX, or unsubstituted, respectively.
(18.2)
Acceptor and donor substituents inuence the electron density on the heteroatoms and reduce or increase the abil­ity to form hydrogen bonds. This explains, among other things, the electronic inuence of aromatic substituents on the biological activity of drug molecules. The Ham­mett equation has, therefore, been seen as achallenge for medicinal chemists and biologists to derive quantitative structure–activity relationships from this concept. Many groups have attempted to nd relationships between bio­logical activity and the Hammet constantsσ, or betweenσ and/or ρ-analogous substituents, and to derive test pa­rameters for biological systems. Despite some interesting results, no general concept could be established.
It was Corwin Hansch and Toshio Fujita who pub­lished apaper in 1964 that laid the foundation for quan- titative structure–activity relationships. In it they describe:
The denition of alipophilicity parameterπ, analo-
-
gous to the electronic termσ in the Hammett equation.
The combination of several parameters in one model.
-
The formulation of a parabolic model to describe
-
nonlinear lipophilicity–activity relationships.
18.3 The Determination and Calculation
of Lipophilicity
Corwin Hansch had previously studied the structure– activity relationship of phenoxyacetic acids, which have growth-stimulating effects in plants. In addition to their biological activity, he was particularly interested in their lipophilicity, which can be measured by the partition co­efcient in an octanol/water system (Sect.19.1). While analyzing the data, he realized that lipophilicity is an additive molecular parameter. The logarithm of the oc­tanol/water partition coefcientP is given by the sum of the group contributions of each part of the molecule. Hansch dened alipophilicity parameterπ (Eq. 18.3) analogous to the Hammett equation. R–X and R–H have the same meaning as in Eq.18.2. The absence of areac­tion-specic ρterm in Eq.18.3 results from relating the πvalues to asingle distribution system, the two-phase mixture of n-octanol and water.
R−X
n-Octanol was chosen for theoretical and practical rea­sons. It has along aliphatic chain and ahydroxyl group that is an H-bond donor as well as an acceptor. Its structure, therefore, resembles the membrane lipids to some extent. It dissolves alarge number of organic com­pounds, it has alow vapor pressure, but can nonetheless be easily removed. Its UV transparence over an extremely wide range is particularly advantageous.
With the help of the lipophilicity parameterπ, the logP values of new compounds, and therefore their li­pophilicity, can be calculated. For this, the lipophilicity of the basic scaffold and the πvalues of the substituents must be known. In this way, the biological activity can be correlated without the tedious experimental measure­ments of each individual partition coefcient. In addi­tion to the πvalues of all important substituents, avery large number of experimentally determined octanol/ water partition coefcients are available in the literature.

18.4 Lipophilicity and Biological Activity

Lipophilicity plays an overwhelming role in describing the dependence of biological effects on chemical struc­ture and, therefore, explains many quantitative structure– activity relationships. This is easily understood because biological systems consist of aqueous phases separated by lipid membranes. The transport and distribution of small molecules in such systems must, therefore, depend on their lipophilicity. For polar substances, the lipid membrane is an insurmountable barrier. Only substances with moderate lipophilicity have agood chance of “mi­grating” into both the aqueous and lipid phases to reach the target tissue in adequate concentrations (Chap.19).
(18.3)
R
1
.log P/2+ k2log P + k3 + :::k
Chapter  • Quantitative Structure–Activity Relationships
18
Although soluble proteins carry predominantly polar amino acid residues on their surface, the more or less buried binding sites for ligands are composed of polar and nonpolar regions. The hydrophobic parts of the li­gand bind to the hydrophobic parts of these pockets. The size of these hydrophobic surfaces is always limited. The size and shape of the lipophilic part of the ligand must t the hydrophobic surfaces in the binding pocket. Since the natural ligands normally bound in these pockets are themselves sufciently water soluble, the lipophilic re­gions in the binding pockets are of limited size. This fact is another explanation for the complex, generally nonlin­ear, lipophilicity–activity relationships.
Many linear and nonlinear lipophilicity–activity
relationships describe relatively unspecic biological
. Table 18.1 The biological activity of meta- and para-sub-
stituents of phenethylamines 18.1 (i.v. application in the rat; Cin mol/kg rat)
meta para log 1/C
H H 7.46
H F 8.16
H Cl 8.68
H Br 8.89
H I 9.25
H Me 9.30
F H 7.52
Cl H 8.16
Br H 8.30
I H 8.40
Me H 8.46
Cl F 8.19
Br Cl 8.57
Me F 8.82
Cl Cl 8.89
Br Cl 8.92
Me Cl 8.96
Cl Br 9.00
Br Br 9.35
Me Br 9.22
Me Me 9.30
Br Me 9.52
effects, such as anesthetic, bactericidal, fungicidal, and hemolytic effects. They will not be discussed further here. Other relationships describe the transport and distribu­tion in abiological system. Such structure–activity rela­tionships are discussed in Chap.19.
18.5 The Hansch Analysis and
the Free–Wilson Model
In 1964 Corwin Hansch and Toshio Fujita derived amathematical model more intuitively than theoreti­cally that can quantitatively describe structure–activity relationships, the Hansch analysis (Eq.18.4).
(18.4)
In Eq.18.4, C is amolar concentration that produces aparticular biological effect. When related to aseries of substances, it is the equieffective molar dose. LogP is the logarithm of the octanol/water partition coefcientP, and σis the Hammett constant. The square of the logP term allows the quantitative description of nonlinear lipo­philicity–activity relationships. This term is omitted when the dependence is linear. Other terms such as polarizability and steric parameters can additionally occur.
The coefcients k1, k2,… andk are determined using the method of regression analysis. The Hansch analysis, therefore, establishes ahypothetical model for quantita­tive relationships between biological activity and physi­cochemical parameters. Biological data are awed, and the same is true for physicochemical properties. Despite this, the reliability of the latter parameters is usually greater than those of the biological data. The result of acalculation is judged by the squared differences between the measured biological data and the values that were calculated from the model. The sum must be as small as possible over all the compounds investigated. It is an important criterion for judging the quality of amodel or for comparing different models of different quality.
The quantitative structure–activity relationship of the antiadrenergic effect of N,N-dimethyl-β-bro- mophenethylamines 18.1 (. Table18.1) is considered as an example. According to their structure, these com­pounds more or less reverse the agonistic effect of an adrenaline dose. The valueC is the dose of an antagonist that blocks the adrenaline effect by 50%. The data can be described using the Hansch model shown in . Fig.18.2.
The entire dataset can be described by amathemati­cal model using the derived equations. When bromine is cleaved, acarbocation is formed and the substances bind irreversibly to the adrenergic receptor. Accordingly, the
+
σ
term is found in the Hansch equation (. Fig.18.2), which describes this type of reaction particularly well. Lipophilic substituents increase the biological activity (positive π term) and electron withdrawing substitu-
. • The Hansch Analysis and the Free–Wilson Model
. Fig. 18.2 A QSAR equation delivers individual parameters
for aquantitative model for the prediction of biological activity, in this case from substituted N,N-dimethyl-β-bromophenethyl­amines (. Table18.1)
ents decrease it (negative σ+ term). Therefore, lipophilic electron-donating substituents, such as large alkyl sub­stituents, should be optimal for activity. Second, within certain limits, the effect of other compounds can be pre­dicted. Interpolations (i.e., conclusions based on very similar substituents) are generally much more reliable than attempts at extrapolations (i.e., predictions made outside the parameter space, e.g., for considerably more lipophilic, more polar, or larger substituents). As arst approximation for the statistical parametersr, s, andF (. Fig.18.2), it can be said that the correlation coef­cientr should have values close to 1.00, the standard deviations should be as small as possible, and the Fvalue should be as large as possible. The better these criteria are met, the better the quantitative model will be, in other words, the better the agreement between the experimen­tal and calculated values.
Also in 1964, and independent of Hansch and Fu jita, S.R. Free and J.W. Wilson developed acompletely different model for structure–activity analysis. Since the original approach is confusingly formulated and difcult to use, only avariant, which was later proposed by Fu­jita and T.Ban, will be discussed here, the Free–Wilson analysis. The Free–Wilson analysis assumes that within aset of chemically related substances, areference com­pound, usually the unsubstituted parent compound, per se makes aspecic contributionμ to the biological effect. Each substituent on this scaffold makes an “ad­ditive and constitutive” contribution ai to the biological activity (. Fig.18.3)—additive, because there is no con­sideration of structural variation at other positions in the molecule, and constitutive, because it matters where in the molecule the specic structural change is made. Despite these relatively simple assumptions, Free–Wil­son analysis provides good quantitative models for many structure–activity relationships.
In contrast to the Hansch analysis, which compares properties, the Free–Wilson analysis is areal “structure–

. Fig. 18.3 The Free–Wilson analysis uses the additive nature of the
group contributions to describe the biological activity. Accordingly, the biological activity in the displayed equation is made up of the ac­tivityμ of the basic scaffold and the constant group contributions ai of the substituents X
activity analysis,” because the parameter that codes for the structural information (1for present, 0for absent) correlates with biological effects. It is easily carried out, but the structures and the biological data must be known.
-
Unfortunately, the Free–Wilson analysis also has disad­vantages:
The structural variation must be present on at least
-
two different substitution sites, because otherwise there will not be enough degrees of freedom to use statistical methods.
The usually large number of variables diminishes the
-
predictive value and reliability of the analyses.
Predictions are only possible for combinations of
-
substituents that have already been considered in the analysis, and not for new substituents.
When the Free–Wilson analysis is applied to the above example of the antiadrenergic phenethylamine, the val­ues for the scaffold and substituent contributions shown in . Table18.2 are obtained. At rst glance, an increase in the values fromF to Cl and from Br toI, i.e., the in­uence of lipophilicity, is obvious. Despite having almost the same lipophilicity, the methyl and chloro substitu­ents are different. This is due to their different electronic properties. Differences in the meta and para positions on
i

Chapter  • Quantitative Structure–Activity Relationships
18
. Table 18.2 Free–Wilson group contributions for
phenethylamines
Atom type Position
meta
para
µ = 7.82 (n = 22; r = 0.97; s = 0.19)
a
For an explanation of these values see . Fig.18.2
H F Cl Br I Me
0.00 −0.30 0.21 0.43 0.58 0.45
0.00 0.34 0.77 1.02 1.43 1.26
a
the electronic inuence can also be followed. Therefore, the Free–Wilson analysis indeed has advantages for the analysis of substituent effects.
18.6 Structure–Activity Relationships of
Molecules in Space
As shown in the previous section, an attempt is made to correlate structure–activity relationships with com­pound-specic parameters. These parameters, such as volume, polarizability, or lipophilicity, are properties that are calculated or measured for the entire molecule or for specic groups of substituents. The 3D structure of the molecules is only conditionally taken into account by these descriptors. Therefore, in the context of increas­ing knowledge of the spatial structure of protein–ligand complexes, QSAR methods focus on parameters that can be derived from the 3Dstructure. In general, the goal of these approaches is to calculate binding afnity. The techniques can also be used to describe other biological properties such as bioavailability, toxicity, or metabolic reactivity (Chap.19). To distinguish them from the clas­sical QSAR techniques described above, they are referred to as 3D-QSAR methods.
Ideally, parameters that can be read directly from the 3D structure of acompound and used to infer its binding afnity would be desirable. However, the interplay be­tween these parameters and activity is very complex and still far from being fully understood. In addition, there are many other biological systems to which one would like to apply 3D-QSAR methods, but the structures of the relevant target proteins are unknown. Many phar­macologically relevant receptors are membrane-bound and their structure determination has proven to be ex­tremely difcult. However, the knowledge of their struc­tures is aprerequisite for areasonable estimation of the binding afnity of a ligand from the geometry of the formed complex (Chap.4). Therefore, instead of trying to calculate the absolute values of the binding afnities from these incomplete data, we will focus on the relative afnity differences between compounds in adataset. The
gradual changes in the compound-specic parameters are then correlated with the biological data.
18.7 Structural Alignment as
aPrerequisite for the Relative Comparison of Molecules
Assumptions about the spatial structure of molecules are already taken into account in classical QSAR techniques. Different positions of substituents, e.g., in the meta or para position of an aromatic ring, are often described by individual parameters. In this form, they are considered in the Hansch equation as well as in the Free–Wilson analysis (Sect.18.5). Furthermore, in classical QSAR models, indicator variables are dened for different con­gurations of substituents, e.g., the conguration of stereoisomers. The use of these parameters assumes an analogous orientation of the molecules in ahypothetical binding pocket. For example, in aseries of ortho-substi- tuted derivatives, it is assumed that all ortho substituents are oriented towards the “same side.” Structure–activity relationships that correlate biological activity with prop­erties of the 3D structure require aspatial superposition of the compounds. This superposition should approxi­mate the relative orientation in the binding pocket as closely as possible. Methods for calculating these spatial superpositions were discussed in Chap.17.
18.8 Binding Affinities as Compound
Properties
What compound-specic properties can be used to cor­relate the properties of the 3D structure with the bind­ing afnity? As discussed in Chap.4, binding afnity is composed of enthalpic and entropic components. The former includes everything that depends on direct en­ergetic interactions. These are mainly of steric (van der Waals potentials, Sect.15.4) or electrostatic (Coulomb potentials) nature. The second contribution focuses on the degree of order and the distribution of energy over the different degrees of freedom of the system under in­vestigation. The ligands as well as the binding pockets of aprotein are solvated by water molecules in the un­complexed state. Upon complex formation, the enthal­pic interactions to these water molecules are lost. They are replaced by direct interactions between the ligand and the protein. Since only relative differences between the molecules of adataset are of interest, effects that are the same for all ligands are not considered. This in­cludes all inuences that affect the protein. This omis­sion is certainly an oversimplication, since the protein changes its degree of solvation upon ligand binding and is polarized differently by ligands. Water molecules are displaced from the binding site. Ligand-induced adapta-
. • How Is aCoMFA Analysis Performed?


tions of side chains in the binding pocket or changes in the rotational degrees of freedom of methyl groups and side chains (Sect.4.10) are conceivable. These effects are either not considered or are assumed to be the same for all molecules in the dataset. This assumption is proba­bly valid in many cases. However, many recent investiga­tions clearly show that changes affecting the protein or the dynamics of the ligand are often not constant within aseries of compounds. This is where the methods fail.
Initially, only the steric and electrostatic interactions of acompound in the binding pocket should be consid­ered. How can these properties be compared for aset of ligands? Arst approach has been the hypothetical interaction models developed by Hans-Dieter Höltje and LemontB. Kier. Akey assumption of these models was the selection and spatial positioning of amino acid side chains around the ligands. When the molecules are embedded in alattice and systematically scanned with an interaction probe, these assumptions are no longer necessary. Richard Cramer and M.Milne proposed such amodel in 1978 (DYLOMMS). It took another 10years before the generally applicable CoMFA (comparative mo- lecular eld analysis) method was established. Despite many theoretical and practical shortcomings in its ap­plication, the method was quickly accepted. Today, it is applied in many different variations.
Before performing such an analysis in practice, some basic considerations should be made. Do steric and elec­trostatic interactions account for all contributions to ligand binding that ultimately lead to acorrect relative ranking of binding afnities? As mentioned above, bind­ing afnity is composed of enthalpic and entropic contri­butions. Sampling properties via probes to map interac­tions certainly provides ameasure of how well amolecule can undergo energetically favorable interactions. But how well are the entropic contributions accounted for? Asig­nicant part of this is due to solvation and desolvation processes (Sect.4.6). In the dissolved state, in the im- mediate vicinity of the hydrophobic surface portion of aligand, the water structure must assume amore ordered state compared to the bulk water phase. The transfer of such aligand from water to the protein-binding pocket, thus, requires that acertain number of water molecules in the water phase change to asignicantly less ordered state. This increases the entropy of the system and fa­vors the spontaneous occurrence of the binding event. The number of water molecules involved in this process depends on the size of the hydrophobic surface of the ligand. Furthermore, the displacement of bound water molecules out of the binding pocket by the ligand to be accommodated increases the disorder of the system un­der consideration and, thus, also the entropy of the sys­tem. In the approximation discussed above, it is assumed that these effects are the same for all molecules in the dataset and are not important in arelative comparison. In addition, rotational, translational, and internal con-
formational degrees of freedom are frozen. As aresult, the entropy of the system decreases. For the afnities to be correctly considered, all these effects would have to be taken into account.
In 2019, Tobias Hüfner in Marburg performed MD simulations on enzyme complexes using the GIST method (Sect.15.4) to predict binding afnities considering en­thalpic and entropic solvation contributions. He made avery interesting observation. Aset of crystal structures of different ligands with atarget protein was used. All the structures were superimposed in their experimentally observed geometry. Thus, the problem of ligand align­ment, which is essential for performing aCoMFA analy­sis, could be solved using experimental data. Tobias Hüf­ner then used different mathematical models to calculate the GIST contributions to the dataset. He also analyzed adataset in which only the ligands were considered, but in their protein-bound conformations. In this way, the ligands are practically sampled for their potential inter­actions with water molecules using an MD simulation. The calculated solvation contributions have been depos­ited on alattice surrounding all ligands. Accordingly, the result is very similar to aCoMFA analysis. Surprisingly, this simulation, performed with only the ligands, gave an excellent afnity prediction for the ligands. How can this be understood? It is possible that asignicant part of the relative differences in binding afnities is already ac­counted for by the desolvation properties of the ligands. Contributions due to the properties of the protein and its desolvation seem to cancel each other out to arst ap­proximation in the relative comparison. Certainly, there are water molecules in the binding pockets whose enthal­pic and entropic properties are very different from those in asurrounding water phase. Apparently, alarge part of these differences in the individual desolvation contribu­tions required to displace the water molecules from the binding site are in turn compensated for by the individual functional groups of the ligands binding to these sites and achieving comparably graded binding contributions with the residues of the protein. Therefore, simply scan­ning the ligands with their different functional groups in the correct bound geometry with awater probe al­ready provides arelevant picture to reasonably predict the relative differences in the afnity data. Perhaps this is aclue as to why comparative eld analysis works so surprisingly well.
18.9 How Is aCoMFA Analysis Performed?
The most important and widely used 3D structure–ac­tivity method is the CoMFA method. The rst step in performing aCoMFA study is to select adataset of suitable compounds. This dataset should contain about 50–100 compounds with related overall geometry. It should also be ensured that all compounds bind to the
Chapter  • Quantitative Structure–Activity Relationships
18
same protein at the same site and that a binding af­nity is known for all of them. The ligands must have acertain diversity of their structural variation. Their binding afnities should be spread over at least three orders of magnitude. Conformations are generated for all molecules (Chap.16) and superimposed using one of the techniques discussed in Chap.17. In general, one refers to the spatial structure of the target protein, if available, and ts the ligands of interest into the binding pocket. Of course, it will be optimals if acrystal struc­ture with the protein is available for many, if not all, bound ligands. Finally, the superimposed molecules are embedded in alattice (. Fig.18.4) that surrounds them by asufciently large margin. The intersections of the lattice should have aspacing of 1 or 2 Å. Aprobe, this is an atom with the properties of hydrogen, carbon, or oxygen, or aparticle with aformal charge, is placed at each of the lattice intersections. The interaction energies between this probe and each molecule in the dataset are calculated. The collective interaction contributions on the lattice are referred to as the interaction eld of the molecule. This is where the name of the method comes from. Finally, the elds of the molecules in the dataset are compared. With abox size of 10–20 Å and agrid spacing of 1–2 Å, there are many thousands of eld values per molecule in the dataset to be processed. This huge amount of data means that eld evaluation can be computationally rather intensive.
18.10 Molecular Fields as Criteria of
aComparative Analysis
Steric and electrostatic interactions are described by aLennard-Jones or Coulomb potential (. Fig.18.5) in force elds (Sect.15.4). As the distance between aprobe and an atom of the molecule approaches zero, the Len- nard-Jones and Coulomb potentials increase towards in- nity. For like-charged particles, the Coulomb potential approaches innity; for oppositely charged particles, it approaches negative innity. These values reach extremely high eld contributions at grid points near the surface or inside amolecule. They must be avoided in aCoMFA analysis. Therefore, the eld contributions above and below acertain threshold are set to apredened cut-off value. Following these procedures, aLennard-Jones or Coulomb potential can be calculated. For example, ali­phatic carbon atoms can be used as probes. These probes are given apositive or negative charge to study the elec­trostatic properties of the molecules. The program GRID by Peter Goodford was introduced in Sect.17.10. With this program, molecular elds can be calculated for nu­merous probes describing different functional groups. For each predened probe, there are regions in space where favorable or unfavorable interactions between the probe and the studied molecule are expected.
In addition, other elds can be dened besides those that probe the steric and electrostatic properties of mole­cules. It was discussed in Sect.18.8 that the hydrophobic surface of amolecule is ameasure of the entropic con­tribution, especially in the transition from the bulk wa­ter phase. In the group of Donald Abraham, at Virginia Commonwealth University, Richmond, USA, molecular elds have been developed which allow the hydrophobic properties of molecules (program HINT) to be studied. These are calculated using avery similar distance-depen­dent function. The resulting molecular eld describes the lipophilicity distribution on the surface of amolecule.
18.11 3D-QSAR: Correlation of Molecular
Fields with Biological Properties
Let us assume that multiple molecular elds have been calculated for each molecule in adataset, and a cor­relation of their differences with binding afnity is at­tempted. How are these differences expressed? For this, we will consider three hypothetical examples of substi­tuted phenyl derivatives.
First, all substituents on the phenyl ring in aseries
-
of compounds should be varied so that increasingly
larger eld contributions occur in the vicinity of the
substituent when scanned with apositively charged
probe. If the binding afnities increase as the eld
contributions increase, this will be reected in the
quantitative analysis. They indicate that ligands with
increasingly negatively charged groups in this region
of the molecule lead to more potent compounds.
This is explained by the fact that the more negatively
charged groups interact better with the positively
charged probe.
The second example is abit different. Now the sub-
-
stituents on the phenyl ring are given positive or neg-
ative partial charges. Their variation has no inuence
on the potency of the compounds. The quantitative
analysis shows that the changes in the electrostatic
eld contributions have no correlation with the bi-
ological activity. A possible explanation could be
that this effect and another property, e.g., the size of
the substituents, cancel each other out. It could also
be that the biological activity is inuenced by other
properties of the substituents, such as their hydropho-
bic character.
In the third case, the electrostatic properties of the
-
substituents that are important for binding to the re-
ceptor should not vary much at the examined position.
There may be different substituents present, but they
all have comparable partial charges. The model that
analyzes the eld contributions in the vicinity of these
groups does not recognize differences and, therefore,
does not nd acorrelation with binding afnity. It
may be that aclass of substituents at aparticular posi-