Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5440_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
10.10.2026
Размер:
9 Мб
Скачать
☆
References 163
34 Günther, S., Kuhn, M., Dunkel, M. et al. (2008). SuperTarget and Matador: resources for exploring
drug-target relationships. Nucleic Acids Research 36 (Suppl. 1): 919–922.
35 Hettne, K.M., Williams, A.J., Van Mulligen, E.M. et al. (2010). Automatic vs. manual curation of a
multi-source chemical dictionary: the impact on text mining. Journal of Cheminformatics
2 (1): 10–11.
36 Von Eichborn, J., Murgueitio, M.S., Dunkel, M. et al. (2011). PROMISCUOUS: a database for
network-based drug-repositioning. Nucleic Acids Research 39 (Suppl. 1): 1060–1066.
37 Lavecchia, A. (2015). Machine-learning approaches in drug discovery: methods and applications.
Drug Discovery Today [Internet] 20 (3): 318–331. https://doi.org/10.1016/j.drudis.2014.10.012.
38 Danishuddin, K.A.U. (2016). Descriptors and their selection methods in QSAR analysis: paradigm
for drug design. Drug Discovery Today [Internet] 21 (8): 1291–1302. https://doi.org/10.1016/
j.drudis.2016.06.013.
39 Nettles, J.H., Jenkins, J.L., Bender, A. et al. (2006). Bridging chemical and biological space: “Target
fishing” using 2D and 3D molecular descriptors. Journal of Medicinal Chemistry 49 (23):
6802–6810.
40 Hert, J., Willett, P., Wilton, D.J. et al. (2004). Comparison of fingerprint-based methods for virtual
screening using multiple bioactive reference structures. Journal of Chemical Information and
Computer Sciences 44 (3): 1177–1185.
41 Raymond, J.W. and Willett, P. (2002). Effectiveness of graph-based and fingerprint-based similarity
measures for virtual screening of 2D chemical structure databases. Journal of Computer-Aided
Molecular Design 16 (1): 59–71.
42 Gao, K., Nguyen, D.D., Sresht, V. et al. (2020). Are 2D fingerprints still valuable for drug discovery?
Physical Chemistry Chemical Physics 22 (16): 8373–8390.
43 Ibrahim, K.A., Helmy, O.M., Kashef, M.T. et al. (2020). Identification of potential drug targets in
Helicobacter pylori using in silico subtractive proteomics approaches and their possible inhibition
through drug repurposing. Pathogens 9 (9): 1–21.
44 Gfeller, D., Grosdidier, A., Wirth, M. et al. (2014). SwissTargetPrediction: a web server for target
prediction of bioactive small molecules. Nucleic Acids Research 42 (W1): 32–38.
45 Dunkel, M., Günther, S., Ahmed, J. et al. (2008). SuperPred: drug classification and target
prediction. Nucleic Acids Research 36 (Web Server Issue): 55–59.
46 Awale, M. and Reymond, J.L. (2017). The polypharmacology browser: a web-based multi
fingerprint target prediction tool using ChEMBL bioactivity data. Journal of Cheminformatics
9 (1): 1–10.
47 Liu, X., Vogt, I., Haque, T., and Campillos, M. (2013). HitPick: a web server for hit identification
and target prediction of chemical screenings. Bioinformatics 29 (15): 1910–1912.
48 Peón, A., Li, H., Ghislat, G. et al. (2019). MolTarPred: a web tool for comprehensive target
prediction with reliability estimation. Chemical Biology & Drug Design 94 (1): 1390–1401.
49 Alberga, D., Trisciuzzi, D., Montaruli, M. et al. (2019). A new approach for drug target and
bioactivity prediction: the multifingerprint similarity search algorithm (MuSSeL). Journal of
Chemical Information and Modeling 59 (1): 586–596.
50 Koszła, O., Stępnicki, P., Zięba, A. et al. (2021). Current approaches and tools used in drug
development against Parkinson’s disease. Biomolecules 11 (6): 897.
51 Piñero, J., Ramírez-Anguita, J.M., Saüch-Pitarch, J. et al. (2020). The DisGeNET knowledge
platform for disease genomics: 2019 update. Nucleic Acids Research 48 (D1): D845–D855.
52 Wang, J., Wolf, R.M., Caldwell, J.W. et al. (2004). 20035_Ftp. Journal of Computational Chemistry
56531 (9): 1157–1174.
7 In Silico Modeling and Drug Design164
53 Van Westen, G.J.P., Wegner, J.K., Ijzerman, A.P. et al. (2011). Proteochemometric modeling as a
tool to design selective compounds and for extrapolating to novel targets. MedChemComm
2 (1): 16–30.
54 Singh, U.C., Brown, F.K., Bash, P.A., and Kollman, P.A. (1987). An approach to the application of
free energy perturbation methods using molecular dynamics: applications to the transformations
of methanol .fwdarw. ethane, oxonium .fwdarw. ammonium, glycine .fwdarw. alanine, and alanine
.fwdarw. phenylalanine in aqueous solution and to H
3
O+(H
2
O)
3
.fwdarw. NH
4
+(H
2
O)
3
in the gas
phase. Journal of the American Chemical Society [Internet] 109 (6): 1607–1614. https://doi.
org/10.1021/ja00240a001.
55 Miyamoto, S. and Kollman, P.A. (1993). Absolute and relative binding free energy calculations of
the interaction of biotin and its analogs with streptavidin using molecular dynamics/free energy
perturbation approaches. Proteins: Structure, Function, and Bioinformatics 16 (3): 226–245.
56 Cortés-Ciriano, I., Ain, Q.U., Subramanian, V. et al. (2015). Polypharmacology modelling using
proteochemometrics (PCM): recent methodological developments, applications to target families,
and future prospects. MedChemComm 6 (1): 24–50.
57 Zhang, S. (2011). Computer-aided drug discovery and development. Methods in Molecular Biology
716: 23–38.
58 Lemkul, J., Genheden, S., Ryde, U. et al. (2015). Assessing the performance of the MM_PBSA and
MM_GBSA methods. 1. The accuracy.pdf. Journal of Chemical Information and Modeling [Internet]
10 (7): 449–461. Available from: http://www.gromacs.org/@api/deki/files/198/=gmx-tutorial.pdf.
59 Liu, W., Schmidt, B., Voss, G., and Müller-Wittig, W. (2008). Accelerating molecular dynamics
simulations using graphics processing units with CUDA. Computer Physics Communications
179 (9): 634–641.
60 Yuriev, E., Agostino, M., and Ramsland, P.A. (2011). Challenges and advances in computational
docking: 2009 in review. Journal of Molecular Recognition 24 (2): 149–164.
61 Matsoukas, M.T., Cordomí, A., Ríos, S. et al. (2013). Ligand binding determinants for angiotensin
II type 1 receptor from computer simulations. Journal of Chemical Information and Modeling
53 (11): 2874–2883.
62 Wu, B., Chien, E.Y.T., Mol, C.D. et al. (2010). Structures of the CXCR4 chemokine GPCR with
small-molecule and cyclic peptide antagonists. Science (80–) 330 (6007): 1066–1071.
63 Schwede, T., Kopp, J., Guex, N., and Peitsch, M.C. (2003). SWISS-MODEL: an automated protein
homology-modeling server. Nucleic Acids Research 31 (13): 3381–3385.
64 Myers, S. and Baker, A. (2001). Drug discovery– an operating model for a new era. Despite the
advent of new science and technologies, drug developers will need to make radical changes in their
operations if they are to remain competitive and innovative. Nature Biotechnology 19 (8): 727–730.
65 Hecker, N., Ahmed, J., Von Eichborn, J. et al. (2012). SuperTarget goes quantitative: update on
drug–target interactions. Nucleic Acids Research 40 (D1): 1113–1117.
66 Davis, A.P., Grondin, C.J., Johnson, R.J. et al. (2021). Comparative toxicogenomics database (CTD):
update 2021. Nucleic Acids Research 49 (D1): D1138–D1143.
67 Chen, B., Dong, X., Jiao, D. et al. (2010). Chem2Bio2RDF: a semantic framework for linking and
data mining chemogenomic and systems chemical biology data. BMC Bioinformatics 11.
68 Wishart, D., Arndt, D., Pon, A. et al. (2015). T3DB: the toxic exposome database. Nucleic Acids
Research 43 (D1): D928–D934.
69 Lindsay, M.A. (2003). Target discovery. Nature Reviews. Drug Discovery 2 (10): 831–838.
70 Noori, H.R. and Spanagel, R. (2013). In silico pharmacology: drug design and discovery’s gate to the
future. Silico Pharmacology 1 (1): 1–2.
References 165
71 Na, D., Rouf, M., O’Kane, C.J. et al. (2013). NeuroGeM, a knowledgebase of genetic modifiers in
neurodegenerative diseases. BMC Medical Genomics 6 (1): 1–14.
72 Drews, J. (2000). Drug discovery: a historical perspective. Science (80–) 287 (5460): 1960–1964. 73
Venkatesh, S. and Lipper, R.A. (2000). Role of the development scientist in compound lead
selection and optimization. Journal of Pharmaceutical Sciences 89 (2): 145–154.
74 Kennedy, T. (1997). Managing the drug discovery/development interface. Drug Discovery Today
[Internet] 2 (10): 436–444. https://doi.org/10.1016/S1359-6446(97)01099-4.
75 Zhu, T., Cao, S., Su, P.C. et al. (2013). Hit identification and optimization in virtual screening:
practical recommendations based on a critical literature analysis. Journal of Medicinal Chemistry
56 (17): 6560–6572.
76 Tuccinardi, T., Poli, G., Romboli, V. et al. (2014). Extensive consensus docking evaluation for ligand
pose prediction and virtual screening studies. Journal of Chemical Information and Modeling
54 (10): 2980–2986.
77 Moustakas, D.T., Lang, P.T., Pegg, S. et al. (2006). Development and validation of a modular,
extensible docking program: DOCK 5. Journal of Computer-Aided Molecular Design 20 (10–11):
601–619.
78 Jain, A.N. (2003). Surflex: fully automatic flexible molecular docking using a molecular similarity
based search engine. Journal of Medicinal Chemistry 46 (4): 499–511.
79 Daina, A., Michielin, O., and Zoete, V. (2017). SwissADME: a free web tool to evaluate
pharmacokinetics, drug-likeness and medicinal chemistry friendliness of small molecules.
Scientific Reports 7 (March): 1–13.
80 Friesner, R.A., Murphy, R.B., Repasky, M.P. et al. (2006). Extra precision glide: docking and scoring
incorporating a model of hydrophobic enclosure for protein–ligand complexes. Journal of
Medicinal Chemistry 49 (21): 6177–6196.
81 Houston, D.R. and Walkinshaw, M.D. (2013). Consensus docking: improving the reliability of
docking in a virtual screening context. Journal of Chemical Information and Modeling 53 (2): 384–
390.
82 Wolber, G. and Langer, T. (2005). LigandScout: 3-D pharmacophores derived from protein-bound
ligands and their use as virtual screening filters. Journal of Chemical Information and Modeling
45 (1): 160–169.
83 Allouche, A. (2012). Software news and updates gabedit – a graphical user interface for
computational chemistry softwares. Journal of Computational Chemistry 32: 174–182.
84 Mueller, R., Rodriguez, A.L., Dawson, E.S. et al. (2010). Identification of metabotropic glutamate
receptor subtype 5 potentiators using virtual high-throughput screening. ACS Chemical
Neuroscience 1 (4): 288–305.
167

8.1 Introduction

The landscape of drug discovery and development unfolds as a multifaceted and resource-intensive
journey spanning a duration exceeding a decade [1]. The intricacies of drug design and exploration
demand a holistic embrace of interdisciplinary strategies. Within this context, computer-aided
drug design (CADD) methodologies come to the fore, predominantly orchestrating the early to
intermediate phases of the drug discovery continuum. The expansion of computational prowess,
data reservoirs, software innovations, and algorithmic sophistication [2–4] has synergistically pro-
pelled the significant amplification of CADD’s role within the realms of drug discovery.
The CADD’s imprint resonates across diverse domains encompassing the pursuit of target iden-
tification, the validation of prospective targets, the pursuit of promising hits, the judicious curation
of lead compounds, and the meticulous refinement of their therapeutic potential [5]. This dis-
course converges its focus on the meticulous scrutiny of pharmacophore modeling situated within
the pantheon of CADD methodologies, thereby accentuating its prominence and implications.

8.1.1 The Role of Pharmacophore Modeling in Drug Design

A pharmacophore delineates the inherent molecular attributes of a small molecule that govern its
effective engagement within the cellular environment for biological or pharmacological functions.
These attributes encompass facets such as geometry, spatial orientation, physiological relevance,
and various other characteristics that collectively contribute to its functional role [6]. Serving as a
prime exemplar, a pharmacophore offers a tangible link between the three-dimensional (3D) struc-
ture and the activity of a molecule. Investigating the pharmacokinetic, pharmacodynamic, and
ADMET (absorption, distribution, metabolism, excretion, and toxicity) properties of a compound
constitute an indispensable stride toward establishing its candidacy as a potential drug-like entity.
The assessment of ADMET profiles holds pivotal sway, wielding significance for both synthetic
and endogenous molecules alike. Essential criteria encompassing favorable, adverse, and inhibi-
tory responses of ligands – small molecules – constitute pivotal landmarks for achieving targeted
ligand–receptor interactions [7].
8

Pharmacophore Modeling in Drug Design

Rahul Ghosh, Sharanya Roy, Gourav Rakshit, Nisha Kumari Singh,
and Nigam Jyoti Maiti
Department of Pharmaceutical Sciences & Technology, Birla Institute of Technology, Ranchi, Jharkhand, India
     168
In drug development, there is a pivotal phase dedicated to the thorough investigation of interac-
tions between ligands and proteins [8]. The application of pharmacophore modeling enhances the
precision of this approach. Pharmacophores were initially proposed in 1970 as a method to gain a
better understanding of the interactions between small molecules (ligands) and receptors, particu-
larly in cases where detailed structural information was unavailable. Pharmacophore modeling
plays a prominent role in CADD, seamlessly integrating with similarity analysis and quantitative
structure–activity relationship (QSAR) studies [9].
Fundamentally, the objective of pharmacophore characterization is to shed light on the spatial
arrangement of functional groups or fragments within an active compound that primarily contrib-
utes to its biological efficacy. The identification of pharmacophore properties derived from ligand–
receptor interactions at binding sites enables the screening of extensive chemical libraries,
encompassing numerous potential novel entities. The construction of a robust pharmacophore
model hinges on the accurate depiction of compounds with biological activity [10–12].
A distinguishing hallmark of the pharmacophore concept lies in the amalgamation of atom-
centric models with their pharmacophoric properties. This amalgamation intentionally incorpo-
rates critical atoms or groups essential for biological interactions. It encompasses features such as
positively or negatively charged moieties, aromatic motifs, hydrophobic clusters, hydrogen bond
acceptors (HBAs), and hydrogen bond donors (HBDs). To facilitate comprehension, pharmacoph-
ore aspects are often represented as points, such as the centroid of five- or six-membered rings.
Inter-feature distances, crucial for the structural definition of pharmacophore models, delineate
these feature sites. The amalgamation of different features and the distances between them eluci-
date the chemical properties and spatial configurations of a pharmacophore.
Pharmacophore modeling stands as a pivotal stride within the intricate tapestry of drug design,
serving as a potent tool for sifting through potential inhibitors by leveraging the intricate tapestry
of their pharmacophore attributes [13]. The journey of drug design traverses a multifaceted land-
scape that commands substantial financial investments, often reaching into the millions, in pur-
suit of the elusive “magic bullet” capable of targeting and remedying specific diseases. Pioneers
within the biopharmaceutical sphere are driven to harness the formidable potential embedded
within the pharmacophore modeling paradigm, aiming to discern bioactive entities through a sub-
strate of foundational biological insights. This endeavor not only augments the precision of the
process but also compresses the temporal horizons of the drug design narrative [8, 14–18].
At the core of drug design, the pharmacophore model is delineated into two primary facets:
structure-based pharmacophore modeling and ligand-based pharmacophore modeling, which is
depicted in Figure 8.1. The former hinges upon preexisting cognizance of ligand properties or
documented inhibitor–receptor complexes. In contrast, the latter thrives on the diversity of phar-
macophore attributes, encompassing hydrogen bonding, hydrophobic interactions, aromatic con-
tacts, metal-mediated associations, and charged interactions, all marshaled in the pursuit of
fashioning novel therapeutic agents [19].
A wealth of online bioinformatics tools stands ready to empower researchers in their pharmaco-
phore modeling ventures, making these sophisticated methods widely accessible. Present-day, a
generalized pharmacophore-based computational approach unfolds through several sequential
steps. This journey commences with the quest for the 3D structure of a biological target implicated
in a specific ailment. Pharmacophore modeling enters the scene, complemented by virtual screen-
ings (VSs) of compound databases. This dual strategy endeavors to unearth molecules displaying
propitious attributes for future therapeutic endeavors. Hits are then subjected to thorough physico-
chemical evaluations, as these attributes emerge as pivotal determinants in establishing their
potential as potent inhibitors [20].
8.1 Introduction 169
Expanding the scope, the biological activities of the chosen molecules undergo scrutiny using
predictive servers that illuminate enzyme-catalyzed metabolic pathways, facilitating the identifi-
cation of high-scoring hits. The following assessments use the Lipinski Rule of Five to examine
bioavailability and ADME-Tox characteristics. As a result, compounds with promising bioavaila-
bility, bioactivity, and ADMET features should be given priority for in vitro testing. The pharmaco-
phore concept thus catalyzes hastening the drug design trajectory [2].
Aligning medications with individual genetic profiles is one area where the pharmacophore
method has expanded into modern personalized medicine. The pharmacophore idea has several
practical uses, including but not limited to target discernment, off-target prediction, VS, and
ADME-Tox modeling. The use of the pharmacophore paradigm in molecular docking simulations
improves the accuracy of binding posture prediction even more. A vital subset within the sphere
of CADD, pharmacophore modeling encapsulates innovation and potential at the crossroads of
modern scientific advancement [21].
Data collection
Ligand-based
pharmacophore
modelling
Structure based
pharmacophore
modelling
Model generation
Virtual screening
Post processing
Best performing
model
Renement and
validation of the model
In vitro and in vivo
validation
No
Yes
Figure 8.1 Flowchart illustrating the procedure of computational drug design.
     170

8.1.2 Historical Perspective and Evolution of Pharmacophore Concepts

The origin of the pharmacophore concept is attributed to Paul Ehrlich, who introduced a novel
approach to developing dyes by examining chromophores, the molecular components responsible
for coloration. In 1890, Ehrlich provided the initial definition of a pharmacophore as “a molecular
framework that bears (phoros) the essential attributes contributing to the biological activity (phar-
macon) of a drug.” In contemporary scientific terminology, a pharmacophore is defined, as estab-
lished by Peter Günd, as “a collection of structural features within a molecule that is recognized at
a receptor site and is accountable for the molecule’s biological activity” [22]. Peter Günd has fur-
ther explored the evolutionary history of the pharmacophore concept in his review [22].
The pharmacophore concept reached its full potential when 3D database search software became
available in the 1990s. The pioneering computer program MOLPAT [6], designed to identify phar-
macophore patterns, was created by Günd, Wipke, and Langridge at Princeton University in 1974.
The demand for 3D structure search software grew alongside the development of rapid 3D struc-
ture generation programs like CONCORD [23], CORINA [24, 25], AIMB [26], and WIZARD [27].
Pharmaceutical companies contributed to the development of 3D search software, with examples
such as ALADDIN [28] (originally by Abbott Laboratories, later commercialized by Daylight
Chemical Information Systems, Inc.) and 3D-Search [29] (developed by Lederle Laboratories),
while academic and government institutions introduced CAST-3D [30] (Chemical Abstract
Services), DOCK [31] (University of California at San Francisco), and CAVEAT (University of
California at Berkeley). Subsequently, the Marshall group devised a pharmacophore method
rooted in ligand structures, known as the “active analog” approach [22]. They applied this approach
to a group of ACE inhibitors [32] and validated the pharmacophore model against available experi-
mental data, demonstrating a strong correlation [33].
The advent of commercial 3D searching systems marked a significant milestone in the field with
the release of MACCS-3D by Güner et al. [22]. Over the subsequent four years, critical advance-
ments were made, paving the way for the technology available today. This period saw the develop-
ment of key 3D searching technologies, including ChemDBS-3D [34] (Chemical Design Inc.,
USA), UNITY [35] (Tripos Inc., USA), and Catalyst [36] (Accelrys Inc., USA). The demand for
pharmacophore development software surged as these 3D searching technologies became widely
accessible. While many of these 3D searching software tools included built-in query generation
capabilities, specialized pharmacophore generation software also emerged. Notable examples
included DISCO [37] by Martin et al. (Tripos Inc., USA), HipHop [36] by Barnum et al. (Accelrys
Inc., USA), and GASP [38] by Jones et al. (Tripos Inc., USA). Simultaneously, predictive models
rooted in QSAR, such as CoMFA (Tripos Inc., USA) [39] by Cramer et al., Apex-3D (Accelrys Inc.,
USA) by Golander and Vorpagel, and HypoGen [40] by Teig et al. (Accelrys Inc., USA), also came
into existence. A comprehensive exploration of the usage and validation of pharmacophore devel-
opment software can be found in the pharmacophore book [8, 28, 41–43]. In Table 8.1, we have
given commonly employed server names with descriptions.

8.2 Essential Concepts in Pharmacophore Hypothesis Generation

Ideally, IC50 values serve as a reliable metric derived from in vitro experimentation, encompassing
target-based and cell-free methodologies, with certain factors mitigated, such as cell efflux, cellular
uptake, and metabolic influences. Cell-based assays, on the other hand, rely on discerning target
interactions with ligands, which may induce modifications in protein expression, protein
       171
Table 8.1 Programs and servers used in pharmacophore modeling.
Server name Description
CATALYST-HipHop [36] CATALYST has been incorporated into the BIOVIA Discovery Studio and
comprises essential algorithms for pharmacophore generation, specifically
HipHop and HypoGen. HipHop is employed for aligning active ligands
concerning a particular target, enabling the identification of 3D
conformations featuring common pharmacophoric elements through the
superimposition of diverse molecular structures
CATALYST-HypoGen [40] By assimilating data derived from biological analyses, pharmacophore
modeling establishes hypotheses that facilitate the quantitative estimation of
molecular activity. This integration enables a more streamlined approach to
pharmacophore modeling by establishing meaningful connections between
structural features and activity data
GASP [44] GASP is a component of the SYBYL package that employs a genetic
algorithm to identify pharmacophores. Unlike conventional pharmacophore
determination methods, GASP conducts conformational searches in real
time and seamlessly integrates this process into its program workflow. Before
superimposing them onto each input chemical, the analysis involves the
examination of conformational changes utilizing a singular low-energy
structure and random spinning
LigandScout [45] While LigandScout offers the capability to conduct structure-based and
ligand-based pharmacophore modeling, it stands out as one of the
pioneering software applications primarily designed for specialized
structure-based pharmacophore modeling. It is particularly favored for
applications where the 3D structure of the target protein in complex with
its ligands is available
GALAHAD [46] This software utilizes an adapted genetic algorithm that addresses specific
limitations found in the GASP program, resulting in enhanced performance.
It accelerates computational processing by employing preconstructed
molecular structures as an initial reference point
MOE [47] MOE possesses the capability to conduct both ligand-based and structure-
based pharmacophore modeling. The process of constructing these models
involves pairwise alignment of active ligands. For optimal results, it is
advisable to reduce the size of the training dataset by clustering molecules
with similar characteristics
PHASE [48] This tool is included in the Schrodinger package. It serves as a practical
method employed in the field of drug discovery, whether the receptor
structure is available. It generates a hypothesis based on one or more ligands,
protein–ligand complexes, and apoproteins. It features a specialized
algorithm specifically tailored for optimizing lead compounds and
conducting virtual screening
PharmaGist [49] This is an openly accessible web server utilized for the generation of
ligand-based pharmacophores. This online tool identifies pharmacophores
through numerous flexible alignments of the input molecules
Pharmer [50] In contrast to conventional molecular library screening methods, this
pharmacophore technique conducts searches based on the breadth and
intricacy of the query. The source code is accessible under an open-source
license, and the technique is well-regarded for its exceptional speed
PharmMapper [51] This publicly accessible web service is employed to identify potential targets
for input ligands. Utilizing semirigid pharmacophore mapping, the
methodology involves the calculation of pharmacophores
     172
degradation, protein trafficking, and perturbations in cell membrane stability. When dealing with
inactive compounds, it is imperative to elucidate the underlying reasons for their lack of activity,
whether attributed to the absence of target interactions, metabolic processes, or efflux pump
activity. It is noteworthy that in vitro data, often acquired through fluorometry-based or
spectrophotometry-based techniques, may be prone to inaccuracies, partially due to the presence
of chromophoric groups. Therefore, meticulous scrutiny is essential when interpreting such data.
In the contemporary scientific landscape, an immense volume of data, spanning both in silico and
in vitro realms, has been exponentially accumulating and populating various databases. For
instance, activity data can be readily accessed from publicly available sources, including but not
limited to open PHACT [52] and PubChem [53]. However, when extracting activity data from
these sources, caution must be exercised, as a single compound may exhibit activity against
multiple targets and be the result of various bioassays. It is essential to acknowledge that minor
discrepancies exist in virtually every database; for example, a typical release of chEMBL may con-
tain approximately 5% inaccuracies in chemical structures, 3% erroneous target information, and
1% errors in data deposition. Hence, users should exercise due diligence when selecting and utiliz-
ing data from these repositories.

8.2.1.1 Partitioning Initial Data into Distinctive Datasets

In accordance with contemporary scientific practices, the raw initial dataset is meticulously
partitioned into three distinct subsets, namely, the training set, the test set, and the decoy set. This
systematic division serves as a pivotal step in the process of constructing a robust pharmacophore
model, as elucidated by some research (Figure 8.2).
Collection of data
Compounds
with no known
IC50 (decoys)
Compounds with
known IC50
(actives)
Training of the
model
Test set
Training set
Validation of model
Virtual screening for
unknown databases to
find the potential hits
Figure 8.2 An overarching procedure for partitioning the data into many datasets.