Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5345_Библиотеки_им_академика_М_И_Перельмана.pdf

Ligand- and Structure-Based Drug Design of NSAIs in Breast Cancer
Zhao, Y., Agarwal, V. R., Mendelson, C. R., & Simpson, E. R. (1997). Transcriptional regulation of CYP19
gene (aromatase) expression in adipose stromal cells in primary culture. The Journal of Steroid Biochem-
istry and Molecular Biology, 61(3-6), 203–210. doi:10.1016/S0960-0760(97)80013-1 PMID:9365191
Zhou, C., Zhou, D., Esteban, J., Murai, J., Siiteri, P. K., Wilczynski, S., & Chen, S. (1996). Aromatase
gene expression and its exon I usage in human breast tumors. Detection of aromatase messenger RNA
by reverse transcription polymerase chain reaction (RT-PCR). The Journal of Steroid Biochemistry and
Molecular Biology, 59(2), 163–171. doi:10.1016/S0960-0760(96)00100-8 PMID:9010331
ADDITIONAL READING
Balunas, M. J., Su, B., Brueggemeier, R. W., & Kinghorn, A. D. (2008). Xanthones from the botanical
dietary supplement mangosteen (Garcinia mangostana) with aromatase inhibitory activity. Journal of
Natural Products, 71(7), 1161–1166. doi:10.1021/np8000255 PMID:18558747
Balunas, M. J., Su, B., Landini, S., Brueggemeier, R. W., & Kinghorn, A. D. (2006). Interference by
naturally occurring fatty acids in a noncellular enzyme-based aromatase bioassay. Journal of Natural
Products, 69(4), 700–703. doi:10.1021/np050513p PMID:16643058
Bhatnagar, A. S. (2007). The discovery and mechanism of action of letrozole. Breast Cancer Research
and Treatment, 105(1), 7–17. doi:10.1007/s10549-007-9696-3 PMID:17912633
Boccardo, F., Rubagotti, A., Puntoni, M., Guglielmini, P., Amoroso, D., & Fini, A. etal. (2005). Switching to anastrozole versus continued tamoxifen treatment of early breast cancer: Preliminary results of the
Italian Tamoxifen Anastrozole Trial. Journal of Clinical Oncology, 23(22), 5138–5148. doi:10.1200/
JCO.2005.04.120 PMID:16009955
Brodie, A. H., & Mouridsen, H. T. (2003). Applicability of the intratumor aromatase preclinical model
to predict clinical trial results with endocrine therapy. American Journal of Clinical Oncology, 26(4),
S17–S26. doi:10.1097/00000421-200308001-00004 PMID:12902873
Caporuscio, F., Rastelli, G., Imbriano, C., & Del Rio, A. (2011). Structure-Based Design of Potent Aromatase Inhibitors by High-Throughput Docking. Journal of Medicinal Chemistry, 54(12), 4006–4017.
doi:10.1021/jm2000689 PMID:21604760
Coupez, B., & Lewis, R. A. (2006). Docking and scoring--theoretically easy, practically impossible?
Current Medicinal Chemistry, 13(25), 2995–3003. doi:10.2174/092986706778521797 PMID:17073642
Doiron, J., Soultan, A. H., Richard, R., Touré, M. M., Picot, N., & Richard, R. etal. (2011). Synthesis
and structureeactivity relationship of 1- and 2-substituted-1,2,3-triazole letrozole-based analogues as
aromatase inhibitors. European Journal of Medicinal Chemistry, 46(9), 4010–4024. doi:10.1016/j.
ejmech.2011.05.074 PMID:21703734
Friesner, R. A., Banks, J. L., Murphy, R. B., Halgren, T. A., Klicic, J. J., & Mainz, D. T. etal. (2004).
GLIDE: A new approach for rapid, accurate, docking and scoring method and assessment of docking
accuracy. Journal of Medicinal Chemistry, 47(7), 1739–1749. doi:10.1021/jm0306430 PMID:15027865
466
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use

Ligand- and Structure-Based Drug Design of NSAIs in Breast Cancer
Gobbi, S., Hu, Q., Negri, M., Zimmer, C., Belluti, F., & Rampa, A. etal. (2013). Modulation of cytochromes P450 with xanthone-based molecules: From aromatase to aldosterone synthase and steroid
11β-hydroxylase inhibition. Journal of Medicinal Chemistry, 56(4), 1723–1729. doi:10.1021/jm301844q
PMID:23363058
Goss, P., & von Eichel, L. (2007). Summary of aromatase inhibitor trials: The past and future. The Jour-
nal of Steroid Biochemistry and Molecular Biology, 106(1-5), 40–48. doi:10.1016/j.jsbmb.2007.05.023
PMID:17627816
Hu, Q., Jagusch, C., Hille, U. E., Haupenthal, J., & Hartmann, R. W. (2010). Replacement of Imidazolyl
by Pyridyl in Biphenylmethylenes Results in Selective CYP17 and Dual CYP17/CYP11B1 Inhibitors
for the Treatment of Prostate Cancer. Journal of Medicinal Chemistry, 53(15), 5749–5758. doi:10.1021/
jm100317b PMID:20684610
Hu, Q., Yin, L., & Hartmann, R. W. (2013). Selective Dual Inhibitors of CYP19 and CYP11B2: Targeting Cardiovascular Diseases Hiding in the Shadow of Breast Cancer. Journal of Medicinal Chemistry,
55(16), 7080–7089. doi:10.1021/jm3004637 PMID:22861193
Hu, Q., Yin, L., Jagusch, C., Hille, U. E., & Hartmann, R. W. (2010). Isopropylidene Substitution Increases Activity and Selectivity of Biphenylmethylene 4-Pyridine Type CYP17 Inhibitors. Journal of
Medicinal Chemistry, 53(13), 5049–5053. doi:10.1021/jm100400a PMID:20550118
Huey, R., Morris, G. M., Olson, A. J., & Goodsell, D. S. (2007). A semiempirical free energy force field
with charge-based desolvation. Journal of Computational Chemistry, 28(6), 1145–1152. doi:10.1002/
jcc.20634 PMID:17274016
Jain, A. (2003). Surflex: Fully automatic flexible molecular docking using a molecular similarity-based
search engine. Journal of Medicinal Chemistry, 46(4), 499–511. doi:10.1021/jm020406h PMID:12570372
Jakesz, R., Greil, R., Gnant, M., Schmid, M., Kwasny, W., & Kubista, E. etal. (2007). Extended Adjuvant
Therapy With Anastrozole Among Postmenopausal Breast Cancer Patients: Results From the Randomized
Austrian Breast and Colorectal Cancer Study Group Trial 6a. Journal of the National Cancer Institute,
99(24), 1845–1853. doi:10.1093/jnci/djm246 PMID:18073378
Kondratyuk, T. P., Park, E. J., Marler, L. E., Ahn, S., Yuan, Y., & Choi, Y. etal. (2011). Resveratrol
derivatives as promising chemopreventive agents with improved potency and selectivity. Molecular
Nutrition & Food Research, 55(8), 1249–1265. doi:10.1002/mnfr.201100122 PMID:21714126
Krammer, A., Kirchhoff, P. D., Jiang, X., Venkatachalam, C. M., & Waldman, M. (2005). LigScore: A
novel scoring function for predicting binding affinities. Journal of Molecular Graphics & Modelling,
23(5), 395–407. doi:10.1016/j.jmgm.2004.11.007 PMID:15781182
Lu, W. J., Xu, C., Pei, Z., Mayhoub, A. S., Cushman, M., & Flockhart, D. A. (2012). The tamoxifen
metabolite norendoxifen is a potent and selective inhibitor of aromatase (CYP19) and a potential lead
compound for novel therapeutic agents. Breast Cancer Research and Treatment, 133(1), 99–109.
doi:10.1007/s10549-011-1699-4 PMID:21814747
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
467

Ligand- and Structure-Based Drug Design of NSAIs in Breast Cancer
Maiti, A., Reddy, P. V., Sturdy, M., Marler, L., Pegan, S. D., & Mesecar, A. D. etal. (2009). Synthesis
of casimiroin and optimization of its quinone reductase 2 and aromatase inhibitory activities. Journal
of Medicinal Chemistry, 52(7), 1873–1884. doi:10.1021/jm801335z PMID:19265439
Mayhoub, A. S., Marler, L., Kondratyuk, T. P., Park, E. J., Pezzuto, J. M., & Cushman, M. (2012). Optimization of the aromatase inhibitory activities of pyridylthiazole analogues of resveratrol. Bioorganic
& Medicinal Chemistry, 20(7), 2427–2434. doi:10.1016/j.bmc.2012.01.047 PMID:22386564
Mayhoub, A. S., Marler, L., Kondratyuk, T. P., Park, E. J., Pezzuto, J. M., & Cushman, M. (2012).
Optimizing thiadiazole analogues of resveratrol versus three chemopreventive targets. Bioorganic &
Medicinal Chemistry, 20(1), 510–520. doi:10.1016/j.bmc.2011.09.031 PMID:22115839
Mooij, W. T., & Verdonk, M. L. (2005). General and targeted statistical potentials for protein-ligand
interactions. Proteins, 61(2), 272–287. doi:10.1002/prot.20588 PMID:16106379
Oda, A., Tsuchida, K., Takakura, T., Yamaotsu, N., & Hirono, S. (2006). Comparison of consensus scoring strategies for evaluating computational models of protein-ligand complexes. Journal of Chemical
Information and Modeling, 46(1), 380–391. doi:10.1021/ci050283k PMID:16426072
Rarey, M., Kramer, B., Lengauer, T., & Klebe, G. (1996). A fast flexible docking method using an
incremental construction algorithm. Journal of Molecular Biology, 261(3), 470–489. doi:10.1006/
jmbi.1996.0477 PMID:8780787
Raub, S., Steffen, A., Kamper, A., & Marian, C. M. (2008). AIScore chemically diverse empirical scoring function employing quantum chemical binding energies of hydrogen-bonded complexes. Journal
of Chemical Information and Modeling, 48(7), 1492–1510. doi:10.1021/ci7004669 PMID:18597446
Stauffer, F., Furet, P., Floersheimer, A., & Lang, M. (2012). New aromatase inhibitors from the 3-pyridyl
arylether and 1-aryl pyrrolo[2,3-c]pyridine series. Bioorganic & Medicinal Chemistry Letters, 22(5),
1860–1863. doi:10.1016/j.bmcl.2012.01.076 PMID:22335894
Velec, H. F. G., Gohlke, H., & Klebe, G. (2005). DrugScore knowledge-based scoring function derived
from small molecule crystal data with superior recognition rate of near-native ligand poses and better affinity prediction. Journal of Medicinal Chemistry, 48(20), 6296–9303. doi:10.1021/jm050436v
PMID:16190756
Voets, M., Antes, I., Scherer, C., Muller-Vieira, U., Biemel, K., & Barassin, C. etal. (2005). Heteroarylsubstituted naphthalenes and structurally modified derivatives: Selective inhibitors of CYP11B2 for the
treatment of congestive heart failure and myocardial fibrosis. Journal of Medicinal Chemistry, 48(21),
6632–6642. doi:10.1021/jm0503704 PMID:16220979
Voets, M., Antes, I., Scherer, C., Muller-Vieira, U., Biemel, K., Marchais-Oberwinkler, S., & Hartmann,
R. W. (2006). Synthesis and evaluation of heteroaryl-substituted dihydronaphthalenes and indenes: Potent
and selective inhibitors of aldosterone synthase (CYP11B2) for the treatment of congestive heart failure
and myocardial fibrosis. Journal of Medicinal Chemistry, 49(7), 2222–2231. doi:10.1021/jm060055x
PMID:16570918
468
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use

Ligand- and Structure-Based Drug Design of NSAIs in Breast Cancer
Wang, R., Lai, L., & Wang, S. (2002). Further development and validation of empirical scoring functions
for structure-based binding affinity prediction. Journal of Computer-Aided Molecular Design, 16(1),
11–26. doi:10.1023/A:1016357811882 PMID:12197663
Wang, R., Shi, H. F., Zhao, J. F., He, Y. P., Zhang, H. B., & Liu, J. P. (2013). Design, synthesis and aromatase inhibitory activities of novel indole-imidazole derivatives. Bioorganic & Medicinal Chemistry
Letters, 23(6), 1760–1762. doi:10.1016/j.bmcl.2013.01.045 PMID:23403081
Woo, L. W. L., Wood, P. M., Bubert, C., Thomas, M. P., Purohit, A., & Potter, B. V. (2013). Synthesis and structure-activity relationship studies of derivatives of the dual aromatase-sulfatase inhibitor
4-{[(4-cyanophenyl)(4H-1,2,4-triazol-4-yl)amino]methyl}phenyl sulfamate. ChemMedChem, 8(5),
779–799. doi:10.1002/cmdc.201300015 PMID:23495205
Yahiaoui, S., Fagnere, C., Pouget, C., Buxeraud, J., & Chulia, A. (2008). New 7,8-benzoflavanones as
potent aromatase inhibitors: Synthesis and biological evaluation. Bioorganic & Medicinal Chemistry,
16(3), 1474–1480. doi:10.1016/j.bmc.2007.10.057 PMID:18042388
Yang, C., Wang, R., & Wang, S. (2006). M-score: A knowledge-based potential scoring function accounting for protein atom mobility. Journal of Medicinal Chemistry, 49(20), 5903–5911. doi:10.1021/
jm050043w PMID:17004706
Yin, L., Hu, Q., & Hartmann, R. W. (2013). Tetrahydropyrroloquinolinone type dual inhibitors of aromatase/aldosterone synthase as a novel strategy for breast cancer patients with elevated cardiovascular
risks. Journal of Medicinal Chemistry, 56(2), 460–470. doi:10.1021/jm301408t PMID:23281812
Yin, S., Biedermannova, L., Vondrasek, J., & Dokholyan, N. V. (2008). MedusaScore: An accurate force
field-based scoring function for virtual drug screening. Journal of Chemical Information and Modeling,
48(8), 1656–1662. doi:10.1021/ci8001167 PMID:18672869
Zhang, C., Liu, S., Zhu, Q., & Zhou, Y. (2005). A knowledge-based energy function for protein-ligand,
protein-protein and protein-DNA complexes. Journal of Medicinal Chemistry, 48(7), 2325–2335.
doi:10.1021/jm049314d PMID:15801826
Zhao, X., Liu, X., Wang, Y., Chen, Z., Kang, L., & Zhang, H. etal. (2008). An improved PMF scoring
function for universally predicting the interactions of a ligand with protein, DNA and RNA. Journal
of Chemical Information and Modeling, 48(7), 1438–1447. doi:10.1021/ci7004719 PMID:18553962
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
469

Ligand- and Structure-Based Drug Design of NSAIs in Breast Cancer
KEY TERMS AND DEFINITIONS
Aromatase: Aromatase is a multienzymatic complex which is mostly overexpressed in breast cancer
tissue and may be responsible for estrogen production and proliferation of breast tumors.
CoMFA: Comparative molecular field analysis is a ligand based drug design approach to find im-
portant steric and electrostatic fields of the ligands crucial for the biological activity.
Competitive Inhibitors: Competitive inhibitors compete with the substrate through non-covalent
binding to the enzyme active site and block the enzymatic action.
CoMSIA: Comparative molecular similarity analysis is similar like CoMFA. Along with the impor-
tant steric and electrostatic fields of the ligands, hydrophobic parameter is determined that can influence
the biological activity.
Docking: A structure-based drug design approach where receptor/enzyme/protein and the ligand
interactions are determined.
Mechanism-Based Inhibitors: Mechanism-based inhibitors mimic the substrate and converted to a
reactive intermediate by the enzyme that inactivate the enzymatic action.
Molecular Properties: Those are the functional as well as the physicochemical characteristics of a
molecule that is supposed to be responsible for the biological activity of molecules.
Pharmacophore Mapping: Important structural and physicochemical parameters responsible for
potential biological activity may be determined.
QSAR: Quantitative structure-activity relationship is a ligand-based drug design approach where
the important structural and physicochemical properties of the ligands are correlated with the biological
activity to determine the important features.
470
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use

Chapter 12
Computational Techniques
Application in Environmental
Exposure Assessment
Karolina Jagiello
University of Gdansk, Poland
Tomasz Puzyn
University of Gdansk, Poland
471
ABSTRACT
In this chapter, the application of computational techniques in environmental exposure assessment was
described. The most important groups of these techniques are Multimedia Mass-balance (MM) modelling and Quantitative Structure-Activity/Structure-Property Relationships (QSAR/QSPR) modelling.
Multimedia Mass-balance models have been widely utilized for studying Long-Range Transport Potential
(LRTP) and overall persistence (P
tional and international acts, including the Stockholm Convention on POPs. Recently, a novel modelling
methodology that links QSPR and MM has been implemented. According to this approach, the physical/
chemical properties required as the input variables for multimedia modelling can be calculated directly
from appropriate QSPR models. QSPR models must be previously developed based on the relationships
between the chemical structure and the modelled properties (QSPR).
) of Persistent Organic Pollutants (POPs), regulated by many na-
OV
INTRODUCTION
The commercialisation of new products that contain growing volumes of various chemical compounds,
is the result of progress in civilisation. Certain chemicals, however, may have a negative impact on the
human body and natural environment. Furthermore, noxious chemical compounds may be generated
as the side products of the manufacturing processes. Consequently, development of reliable methods of
risk assessment, which would allow eliminating potentially dangerous chemical substances at the stage
of synthesis or final product planning, seems to be increasingly essential. Such methods should ensure
good pace and cost-efficiency so as to avoid additional expenses.
DOI: 10.4018/978-1-4666-8136-1.ch012
Copyright © 2015, IGI Global. Copying or distributing in print or electronic forms without written permission of IGI Global is prohibited.
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use

Computational Techniques Application in Environmental Exposure Assessment
y f X=
The provisions contained in the REACH Regulation (EC, 2006) that requires the entities to file
every new chemical substance to be commercialised on the European market in large quantities (either
manufactured or imported) with the European Chemicals Agency in Helsinki, Finland, have applied to
the European Union countries since recently. The registration process must be preceded by the phase of
evaluation of the chemical risk. Article 13 of the REACH Regulation determines that information on
the substance should be generated, if only possible, using so called “alternative methods” with regard
to ethically doubtful and expensive tests on animals. Computer assisted methods, specifically those
using the quantitative modelling of relationships between the chemical structure and activity (Quantitative Structure-Activity Relationships, QSAR) and physical properties (Quantitative Structure-Property
Relationships, QSPR) may be listed from among “alternative methods”.
Special attention is given to the chemical substances which are defined by the acronym PBT (Persistent,
Bioaccumulative and Toxic) and classified as: persistent (P), bioaccumulative (B), and toxic (T) either
in the European REACH Regulation or in the legislation concerning chemical safety and applicable to
Canada (CEPA, 1999), United States of America (USEPA, 1999), or Japan (METI, 1973). It should be
mentioned here that the criteria of eligibility of a substance for PBT, Table 1, originate from previous
research into the group of Persistent Organic Pollutants (POPs) (Gobas, 2009; van Wijk, 2009). The
awareness of the risk connected with the presence of POPs in the natural environment at the turn of the
1960’s and 1970’s allowed launching many initiatives of multinational nature, with the aim to decrease
the production and usage of such substances in order to reduce the emissions of chemical pollutants. In
this context, POPs may be considered as “model pollutants” the tests of which have allowed developing
many theories of transport and deposits in natural environment as well as toxicity and ecotoxicity of
pollutants, such theories being largely applied these days. Recently, more and more attention is given to
brominated and bromochlorinated POPs (Br-POPs and Br/Cl-POPs) given their increasing emissions to
the environment and high levels of Br-POPs and Br/Cl-POPs observed in environmental matrices (Batterman, 2007; Du, 2010; Hutson, 2009), as well as their confirmed toxicity (Birnbaum, 2003).
The present chapter describes the existing scientific knowledge in two areas of concern, that is:
1. Fundamental elements of QSAR/QSPR methodologies, and
2. Development of the methodology of predicting transport of chemical substances in the environ-
ment on the basis of combination of Multimedia Mass Balance Models (MM) with QSPR models
(QSPR-MM hybrid models).
BASIC IDEA AND INCREASING IMPORTANCE OF QSPR/QSAR METHODS
QSAR/QSPR models rely on a shared assumption that the variance of modelled quantity y (either physi-
cal/chemical property or biological activity, respectively) in a group of chemical compounds featuring
similar structures is determined by the variance connected with differences in chemical structure of such
compounds, whereas the mutability of chemical structure is expressed numerically by means of structural
descriptors X. Hence, if we have experimentally measured values of a physical/chemical quantity or of
activity y for a suitably number of compounds as well as the values of structural descriptors X for all the
compounds, we will be able to build a suitable model in the format shown in Equation (1).
( ). (1)
472
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use

Computational Techniques Application in Environmental Exposure Assessment
Table 1. Criteria for the identification of PBT/vPvB under annex XIII of the REACH regulation (EC, 2006)
Persistence (P) Bioaccumulation (B) Toxicity (T)
PBT Substances t1/2 in marine water
> 60 days
or
t1/2 in fresh / estuarine water
> 40 days
or
t1/2 in marine sediment
> 180 days
or
t1/2 in fresh / estuarine
sediment
> 120 days
or
t1/2 in soil
> 120 days
vPvB Substances t1/2 in marine, fresh/estuarine
water
> 60 days
or
t1/2 in marine, fresh/estuarine
sediment
> 180 days
or
t1/2 in soil
> 180 days
BCF > 2000 NOEC (long-term) < 0.01 mg/L
• Substance is classified as
carcinogenic (category 1 or 2),
mutagenic (category 1 or 2), or
toxic for reproduction (category 1,
2 or 3), or
• There is other evidence of
chronic toxicity, as identified by the
classifications: T, R48, or Xn, R48
according to Directive 67/548/EEC
BCF > 5000 N/A
The formula allows predicting the modelled quantity for the remaining compounds from the group,
on the basis of their structural descriptors (Cronin, 2010). In QSAR/QSPR, the modelling procedure is
composed of three stages:
1. Model calibration on the basis of a part of compounds assigned to a “training set”;
2. Model validation which accounts for test of model functioning in the case of compounds which are
not covered by the process of calibration and which have a defined quantity y (compounds from
the validation set); and
3. Model application for predicting the quantity y in the case of new compounds (Figure 1) (Cronin,
2010; Puzyn, 2009a).
In compliance with the recommendations of validation (OECD, 2007) adopted in OECD countries,
QSAR/QSPR model should feature the following characteristics:
1. Well defined modelled quantity y (endpoint);
2. An unambiguous algorithm which allows repeating the tests;
3. Defined applicability domain of the model, which means a very well-defined group of chemical
compounds for which quantity y is predicted by the model as a result of interpolation and hence,
the value of the endpoint may be deemed correct;
4. Appropriately used and interpreted statistical quantities which describe the goodness-of-fit, flex-
ibility (robustness) and predictive ability of the model; and
5. Physical interpretation of applied combination of structural descriptors, if possible.
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
473

Computational Techniques Application in Environmental Exposure Assessment
Figure 1. Scheme of QSAR/QSPR models development
If the model meets all the five criteria, it may be used as a reliable tool in the procedure of new
chemical compound risk assessment.
QSPR/QSAR MODELLING ALGORITHM
QSPR/QSAR modelling procedure is composed of three stages:
1. Model calibration;
2. Model validation, and
3. Model application for the purpose of predicting the quantity y for new compounds (Cronin, 2010;
Puzyn, 2009b).
Model Calibration
The first stage is initiated by the construction of molecular models of all the molecules of a group, and
is followed by the computation of the matrix of structural descriptors X on the basis of such molecular
models. The matrix contains (in columns) the descriptors for each of the compounds (matrix rows)
(Cronin, 2010; Puzyn, 2009b).
The values of modelled y are simultaneously collected (or measures) for as large part of compounds
from the group as possible. Similarly as in the case of each mathematical model, QSPR/QSAR modelling
will never lead to predictions on the quality higher (smaller error) than the quality of data used for such
474
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use

Computational Techniques Application in Environmental Exposure Assessment
model calibration. Occurrence of additional variability sources that are the result of different conditions
(e.g. temperature and pressure) of the experimental study, laboratory-specific differences in the individual studies, another methodology applied for the purpose of determination of the same quantity, etc.,
should be as minimal as possible. If not, the modelling process will apply firstly to the impact of such
other features rather than to the relationship between the quantity y and the chemical structure (Puzyn,
2009b; Cronin & Schultz, 2003; Dearden, 2009).
If matrix X and vector y are available, the construction of QSPR/QSAR model will commence with
three chemometric methods being the most commonly used. These are: MLR method (Multiple Linear
Regression) as well as two methods employing the latent vectors which are linear combinations of
original descriptors, that is: PCR method (Principal Component Regression) and PLS method (Partial
Least Squares) (Puzyn, 2009b; Walczak, 2009). The complexity and the type of actual method should
depend on the specificity of the problem, which means that the simplest method should be always
applied, whenever possible. The two last methods will be explicitly useful if individual descriptors in
matrix X demonstrate high inter-correlation (Puzyn, 2009a).
Furthermore, selection of the optimum combination of descriptors for the purpose of modelling
seems to be an important issue. Two alternative approaches to the problems are used, conditional on
the availability of information connected with the mechanism which decides on the specificity of
either structure-property or structure-activity relationship. If the mechanism and structural factors
having an impact on quantity y are known, the descriptors might be selected arbitrarily. If not, special
algorithms need to be employed in order to test thousands of descriptor combinations and select such
combinations that lead to the model featuring the lowest-error-possible predictivity (Puzyn, 2009b;
Gramatica, 2010). Final selection, however, must be made by the researcher responsible for construction of the model, who takes into account the possibility of assignment of the physical sense to the
combination of descriptors as proposed by the application. Moreover, selection of descriptors for a
model should take into account the fact that excessive number of descriptors (variables) compared
with the number of compounds used for the purpose of calibration almost certainly leads to excessive model matching (overfitting) and hence, to the loss of information generalisation capacity and
of correct predictions of variable y for new compounds. Empirical studies allow making an assumption that at least 5 calibration compounds should be assigned to a single descriptor (Topliss ratio) as
a standard in QSPR/QSAR modelling (Gramatica, 2007, 2010; Holland, 1992; Livingstone & Salt
2005; Topliss & Edwards, 1979).
Model Validation
The second stage of QSPR/QSAR modelling consists in a reliable validation. The validation contains the
evaluation of goodness-of-fit, evaluation of model robustness, evaluation of predictive capacity in case
of new compounds (predictivity), and verification of the model applicability domain (Puzyn, 2009b;
Gramatica, 2010).
Evaluation of model goodness-of-fit is conducted on the basis of values of residuals for the compounds
used for calibration. Distribution of residuals should be similar to the normal (random) distribution.
The most frequently used measurements of ability of QSAR models to reproduce the training data are
2
determination coefficient R
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
(2) and RMSEC (root mean square error of calibration) (3):
475
Соседние файлы в папке Библиотека им академика М.И. Перельмана
