Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5587_Библиотеки_им_академика_М_И_Перельмана.pdf

An Introduction to the Basic Concepts in QSAR-Aided Drug Design
12
Frog (FRee Online druG) is a tool for generating small molecular conformations by using 1D or
2D descriptions. http://bioserv.rpbs.jussieu.fr/cgi-bin/Frog
13
Smi23d or (3D Cordinate Generation) program is able to convert one or more SMILIES into 3D
structure.
http://www.chembiogrid.org/cheminfo/smi23d/
14
ADRIANA.Code is used to compute molecular structure descriptors in drug design and predict
the ADME/toxic properties. https://www.molecular-networks.com/products/adrianacode
15
Molconn-Z is a software which creates Molecular Connectivity, Shape, and Information Indices
for (QSAR) Analyses. The new E-State parameter is also included. http://www.edusoft-lc.com/
molconn/
16
A software that calculates 1875 molecular descriptors (i.e., 1444 1D, 2D descriptors and 431 3D
descriptors and fingerprints). http://padel.nus.edu.sg/software/padeldescriptor/
17
KNIME is an open source software for intelligent algorithms including the data mining, machine
learning, and others. https://www.knime.org/
18
RapidMiner is mostly used for data integration, data transformation, data modeling, and data
visualization.
http://rapidminer.com/
19
Weka (Waikato Environment for Knowledge Analysis) is a general machine learning software
including some common statistical, clustering, and intelligent system algorithms written in Java.
http://www.cs.waikato.ac.nz/ml/weka/
20
An Open source program based on visual programming or Python used for data visualization and
analysis.
http://orange.biolab.si/
21
TANAGRA is a free software which includes some data mining methods and is capable of ana-
lyzing the through statistical and, machine learning algorithms. http://eric.univ-lyon2.fr/~ricco/
tanagra/en/tanagra.html
22
MATLAB (matrix laboratory) is a numerical computing environment which inclides various types
of toolboxes for artificial neural networks, image processing, bioinformatics, parallel computing,
optimization procedures and etc. http://www.mathworks.com/products/matlab/
23
R is a free software environment alike MATLAB for statistical and machine learning computing
and graphical illustrations. Many packages over the world are being written for the purpose of
general and expert researches. http://www.r-project.org/
24
SYBYL® is considered as the heart of Tripos. This software has various applications including
construction, editing, and visualization tools, modeling and simulation of the molecules and lead
identification, and lead optimization. SYBYL-X is the latest version.
25
Discovery Studio is developed by Accelrys used on small/macro molecule and simulation.
http://accelrys.com/products/discovery-studio/
26
MOE is a software used generally for life and material sciences by comprehensively performing
the visualization, simulations on molecular structures.
http://www.chemcomp.com/MOE-Molecular_Operating_Environment.htm
46
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use

An Introduction to the Basic Concepts in QSAR-Aided Drug Design
27
CODESSA is a tool for obtaining descriptors to develop QSAR/QSPR model through quantum
mechanical results from sources such as AMPAC, Gaussian, AIMALL. http://www.semichem.
com/codessa/
28
It is a 3D-QSAR program calculating of total interaction energies between different chemical
probes and molecules. It determines energetically favorable binding sites on known molecular
structures as well as surface properties. GRID 22c is the latest version. http://www.moldiscovery.
com/software/grid
29
Molecular Discovery is one of the companies in the field of structure-based drug design. All the
developed programs, based on GRID force field, are designed for 3D-QSAR studies, virtual screening, scaffold-hopping, ADME and pharmacokinetic modelling, optimization of metabolic stability
and metabolite prediction. http://www.moldiscovery.com/
30
Pentacle is a alignment-independent 3D-QSAR software. It computes interaction energies between
the molecule and chemical probes (GRID based Molecular Interaction Fields, or MIFs). The latest
version of Pentacle is 1.0.7.http://www.moldiscovery.com/software/pentacle
31
GOLPE (Generating Optimal Linear PLS Estimations) is an advanced chemometric toolbox for
building, validating and interpreting the 3D QSAR models.
http://www.miasrl.com/golpe.htm
32
The toolbox is especially suitable for chemical researchers to assess the toxicity of the chemical
compounds. The latest version is 3.2. http://www.oecd.org/chemicalsafety/risk-assessment/theoecdqsartoolbox.htm
33
REACH stands for Registration, Evaluation, Authorisation and Restriction of Chemicals which
has been approved by European Union in 2006 for assessing the chemical effects on health and
environment.
34
TOXNET (TOXicology Data NETwork) is a group of toxicology dataset in various sections in-
cluding the chemicals, drugs and environmental effects. http://toxnet.nlm.nih.gov/
35
SIDER stands for Side Effect Resource which includes the information about side effect frequency
of drugs. And the latest version of SIDER is 2.0 and can be found at the link ttp://sideeffects.embl.
de/
36
ACToR is a tool for the interested researches in chemical toxicity to extract data about the possible
risk of current chemicals on human health. http://www.epa.gov/actor/
37
DailyMed provides the most comprehensive information about the marketed drugs which are ap-
proved by Food and Drug Administration (FDA). http://dailymed.nlm.nih.gov/dailymed/about.cfm
38
Distributed Structure-Searchable Toxicity (DSSTox) Database Network is a project of EPA for
publicly presenting the data about structure-activity and predictive toxicology capabilities
http://www.epa.gov/ncct/dsstox/.
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
47

48
Chapter 2
The “ETA” Indices in QSAR/
QSPR/QSTR Research
Kunal Roy
Jadavpur University, India
Rudra Narayan Das
Jadavpur University, India
ABSTRACT
Descriptors are one of the most essential components of predictive Quantitative Structure-Activity/
Property/Toxicity Relationship (QSAR/QSPR/QSTR) modeling analysis, as they encode chemical information of molecules in the form of quantitative numbers, which are used to develop mathematical
correlation models. The quality of a predictive model not only depends on good modeling statistics, but
also on the extraction of chemical features. A significant amount of research since the beginning of QSAR
analysis paradigm has led to the introduction of a large number of predictor variables or descriptors.
The Extended Topochemical Atom (ETA) indices, developed by the authors’ group, successfully address
the aspects of molecular topology, electronic information, and different types of bonded interactions,
and have been extensively employed for the modeling of different types of activity/property and toxicity
endpoints. This chapter provides explicit information regarding the basis, algorithm, and applicability
of the ETA indices for a predictive modeling paradigm.
INTRODUCTION
Chemicals are the inseparable components of human life, and thus they can affect a wide variety of industrial, agricultural as well as household processes if not monitored properly. Elucidation of the chemistry
of compounds is an essential task of the designer so as to enable easy tuning of their properties. The
diverse manifestation of chemicals can be systematically explored when the study of their chemistry is
suitably attended with the related biology (including toxicology), mathematics and accompanying statistics. Since the past, mathematics has been considered as an effective media of natural science (Wigner,
1960) and has been employed in different branches of science to derive logical relationship employing
abstract ideas. Chemistry has a very good coherence with mathematics that leads to the quantification
DOI: 10.4018/978-1-4666-8136-1.ch002
Copyright © 2015, IGI Global. Copying or distributing in print or electronic forms without written permission of IGI Global is prohibited.
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use

The “ETA” Indices in QSAR/QSPR/QSTR Research
of chemical information by providing suitable platform for the generation of various algorithms. How
different chemicals act differently and even the same chemical elicits different responses give the hint
to nurture the ‘chemistry’ of compounds and explore the possible quantitative phenomena involved. The
attempt to quantify the chemistry of compounds has lead to the development of several independent
parameters termed as predictor variables or descriptors. These descriptors are further utilized to bring
under a quantitative relationship with the responses of the chemicals commonly termed as the quantitative structure-activity/property/toxicity relationship (QSAR/QSPR/QSTR) studies. Mathematics plays
a great role in deriving such predictor variables as well as in the development of correlations.
Through the journey of QSAR analysis, several predictor variables have evolved by exploring the
concepts of chemistry, mathematics and physics. However, the ‘graph theory’ has contributed much to
the development of initial ideas about descriptors. The graph theory, one of the branches in mathematics,
can be traced back to Euler’s Königsberg bridges problem (Euler, 1736) which led to the idea of solving
problems in different fields of science namely the study of electrical circuit by Kirchhoff (Kirchhoff,
1847), chemical isomer enumeration by Cayley (Cayley, 1875) etc. In a true mathematical sense, the
term ‘graph’ also denotes cartesian plot of data, although it was Sylvester (Sylvester, 1878) who coined
the term from a contemporary chemical perspective. Application of the graph theory in chemistry leads
to the concept of ‘chemical graph’ which was supposed to have its origin in the late eighteen century
holding the hand of the Scottish chemist Cullen (Crosland, 1959). William Cullen introduced the application of ‘affinity diagrams’, although it remains controversial with similar type of contributions by Black
(Crosland, 1959). Later Higgins (Higgins, 1789) employed the use of graphical diagrams to denote the
forces between atoms in a molecule. It was Dalton (Dalton, 1808) in the nineteenth century who started
the concept of ball-and-stick representation of chemical models. The idea of diagrammatic presentation
of chemical compounds was very primitive during that age and through various notable discoveries the
chemistry got explored gradually. The famous scientist Kekulé brought the idea of three-dimensional
orientation (Fischer, 1974) and later idea of chemical valence came likewise. Considering the scope of
this chapter, we would like to limit this discussion at this point.
It was in the twentieth century when the chemical graphs were explored properly, and utilizing the concepts
of topological distance measures the scientific fraternity started developing theoretical variables to correlate
activity of chemicals with their chemistry. Although the history of QSAR starts back in the past when Cross
(Cross, 1863) reported his idea of possible correlation between chemical composition and toxicity of alcohols,
the concept of theoretical descriptor did not come in action properly until the year 1947 when the Wiener index
(Wiener, 1947) and the Platt number (Platt, 1947) were introduced as the first two graph theoretical descriptors
for modeling boiling point of hydrocarbons. Following these milestone achievements in the realm of quantification of chemistry paradigm, a new wave was hit towards the exploration of graph theoretic parameters, and
in the 1960s and 1970s the pathway was systematically imputed by the following pioneering contributions
namely Simon’s minimal topological difference (MTD) concept (Simon, 1974), Balaban J (Balaban, 1982),
Gordon and Scantlebury index (Gordon & Scantlebury, 1964), Schultz molecular topological index (MTI)
(Schultz, 1989), Hosoya Z (Hosoya, 1971), Randić’s branching parameter (Randić, 1975), Kier and Hall’s
molecular connectivity, Kappa shape, sub-graph count, flexibility and electrotopological state (E-state) atom
indices (Kier & Hall, 1986), Zagreb index by Trinajstić and co-workers (Gutman & Trinajstić, 1972), Harary
index (Plavšić et al., 1993), Verloop’s STERIMOL parameter (Verloop, 1987), etc. Graph theoretic descriptors
are purely based on a two dimensional basis, whereas concepts of three dimension was also under parallel
research and various predictor variables were also derived employing three dimensional basis. A torrential
amount of research on the descriptors has led to the development of a wide variety of predictor variables till
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
49

The “ETA” Indices in QSAR/QSPR/QSTR Research
today. The present chapter gives an account of such a group of two-dimensional novel descriptors named
extended topochemical atom (ETA) indices which have been developed in the twenty first century from the
modification of another set of topologically arrived unique (TAU) scheme descriptors of the late eighties.
The ETA indices are now available under two generations and have been used in various modeling studies to
predict property, activity as well as toxicity endpoints.
Descriptors in QSAR Studies: The Impact and Required Quality
Scientific deductions are based on logical components that aids in deriving conclusion from an observation. The chemoinformatics study also bears such inevitable components that ultimately help in developing predictive mathematical models and also leave a plausible explanation for the reasons involved.
The molecular descriptors in QSAR studies can be viewed as the quantitative presentation of chemical
structures obtained by processing the molecule using a specified mathematical algorithm. Todeschini
and Consonni (Todeschini & Consonni, 2000) have ascribed descriptors as the useful numbers or the
standardized experimental results obtained as the final result of a logic and mathematical procedure that
leads to the transformation of chemical information which is encoded within symbolic representation
of molecules. So, the mathematics is here utilized to unwrap the chemical information of compounds
in the form of numerical entities. Descriptors are the algorithmic mathematical entities that strongly
adhere to the chemistry of compounds and thereby form a core to the QSAR analysis. Since, the attributes of chemical compounds are depicted using different levels and dimensions of theory of chemistry
and physics, descriptors are also defined on the basis of such theory coupled with a definite algorithm
of mathematics. A good number of descriptors are available now-a-days corresponding to different
algorithms which can be conveniently categorized from the dimensional perspective. Figure 1 shows a
representative dimensional classification of various descriptors for QSAR analysis.
Figure 1. Dimensional perspective of descriptors used in predictive modeling (QSAR) analysis
50
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use

The “ETA” Indices in QSAR/QSPR/QSTR Research
Descriptors can be defined for a whole molecule or they may also refer to specific fragments of
the molecule. Again from the basis of computation, descriptors can be theoretical or experimental.
Experimental descriptors correspond to various derived physicochemical properties namely partition coefficient, molar refractivity, optical rotation, pK
values, reaction rate, parachor etc. and
a
they usually depict characteristics of the entire molecule while theoretical descriptors are usually
computationally derived mathematical entities and can define specific molecular fragment as well as
the whole molecule. An interesting fact about descriptors is that they evolved from lower to higher
dimensions as the practice of QSAR made its progress and so the theory. Looking back at the history of QSAR, it can be observed that Cross (Cross, 1863) was the first who observed the toxicity
of primary aliphatic alcohols to be possibly correlated with their water solubility. Crum-Brown and
Fraser (Crum-Brown & Fraser, 1868) are inexorably believed to be first to report the physiological action as the mathematical function of chemical constituents during an investigation of muscle
paralysis activity of strychnine derivatives. Richardson (Richardson, 1869) observed a proportional
relationship of narcotic effect of primary alcohols with their molecular weight and almost after two
decades Richet (Richet, 1893) found the cytotoxicity of ethers, alcohols and ketones to be inversely
related with their aqueous solubility measure. Then Meyer (Meyer, 1899) and Overton (Overton,
1901) came with their decisive work by correlating narcotic potency with lipid/water partition coefficient. After that, at the beginning of the twentieth century Trabe (Trabe, 1904) and Seidell (Seidell,
1912) came with their correlation analysis employing surface tension and solubility and partition
coefficient respectively. Then came the pioneering discovery by Hammett (Hammett, 1935) who
initially showed relationship between reaction rates with acid dissociation constant values (pK
a
and the parameter σ (sigma) which afterwards led to a “sigma-rho” culture in the field of predictive
modeling analysis. Some more notable research activities were performed during this time namely
correlation of narcotic activity with partition coefficient and exploring thermodynamics of chemicals
by Ferguson (Ferguson, 1939), study on the involvement of ionization and shape factors towards
the bacteriostatic action of aminoacridines by Albert (Albert et al., 1941) and correlation studies
on antibacterial action (E. coli) of sulfonamides with their ionization and shape parameters by Bell
and Roblin (Bell & Roblin, 1942). In this journey, it may be observed that QSAR started with the
concept of correlating toxicity of chemicals with aqueous solubility by Cross whereas the term
‘chemical constituents’ in Crum-Brown and Fraser’s hypothesis can be connected to the elemental
composition of the compounds. Richardson and Richet found toxicity to be expressible in terms of
molecular weight and aqueous solubility respectively. Hence, we can see that at the dawn of QSAR,
scientists used different physicochemical properties, i.e., experimental descriptors as variables and
at the same time the concept of zero dimensional parameter like elemental composition, molecular
weight were also used. The practice of using physicochemical properties like partition coefficient,
solubility etc. as descriptors was also done by other scientists namely Trabe and Seidell. The concept
of extracting chemical information was getting boosted during this time and Hammett led to the
discovery of a theoretical parameter σ (sigma) in connection with modeling reaction kinetics. The
contributions by Albert et al. and Bell & Roblin are also notable since they were the first to focus
on the ionization factor and shape factors as independent parameters to derive correlation. It was
in 1947, Wiener (Wiener, 1947) and Platt (Platt, 1947) came up with some more light to the path
with their first topologically derived graph theoretic descriptors namely Wiener index and Platt
index respectively by developing QSPR models for the boiling points of paraffin hydrocarbons. The
inception of graph theoretic parameters as chemical descriptors opened a big doorway in the field
)
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
51

The “ETA” Indices in QSAR/QSPR/QSTR Research
of theoretical descriptors, and since then a significant amount of research has taken place towards
the development of more informative parameters. The discovery of Hammett (Hammett, 1935) was
forwarded by Taft who introduced a steric parameter E
(Taft, 1952) that distinguishes polar, steric
s
and resonating effect exerted by chemicals and reported the first multiple regression method based
linear free energy relationship (LFER) equation. In the early 1960s, several newer developments were
made in the field of 2D-descriptors. In 1962 Hansch (Hansch et al., 1962) introduced a substituent
hydrophobicity parameter π and later in 1964 Hansch and Fujita (Hansch & Fujita, 1964) extended
the analysis by combining parameters of Hammett and Taft and developed LFER models comprising
of σ, E
and π. The analysis by Fujita and Hansch has undergone further parabolic extension and
s
later it was also modified by Kubinyi (Kubinyi, 1976). Another concept of presenting molecular
descriptors as the contribution of the parent moiety and substituents was depicted by Free-Wilson
(Free & Wilson, 1964) which gave rise to ‘indicator variables’. The Free-Wilson analysis was later
modified by Fujita-Ban (Fujita & Ban, 1971) by incorporating logarithmic terminology.
It is to be noted that from the bases of the two-dimensional chemistry, concepts of three- dimensional (3D) molecular structure representation, electronic information and the descriptors thereof
were also in the developing stage starting in 1930s. Some of the notable contributions can be cited
namely the chemical bonding theory by Pauling (Pauling, 1939) and Coulson (Coulson, 1939),
Sanderson’s (Sanderson, 1952) electronegativity parameters, electronic distribution theory by Fukui
et al. (Fukui et al., 1954) and Mulliken (Mulliken, 1955), charged partial surface area measures
of Jurs & Stanton (Stanton & Jurs, 1990), shadow indices (Rohrbaugh & Jurs, 1987), WHIM (Todeschini & Lasagni, 1994) & GETAWAY descriptors by Todeschini and co-workers (Consonni et
al., 2002), 3D-MoRSE descriptors (Schuur et al., 1996), EEVA descriptors, CoMFA and CoMSIA
analysis by Cramer et al. (Cramer et al., 1988) and the ligand-receptor based analyses including
multidimensional (4D, 5D, 6D and so on) features in the modern century (Lill, 2007). However,
considering the scope of this article we would like to focus our topic on two-dimensional topological structure representation only.
Hence it is observed that the molecular descriptors of chemical compounds are pivotal as they
play a very crucial role in predictive modeling analysis. Through the journey of QSAR paradigm,
the attempt to generate quantitative numbers using the chemistry of compounds has yielded various
mathematical descriptors. Now, the obvious question arises regarding the required quality of such
essential component of QSAR analysis. It is apparently true that quantitative predictor variables were
introduced to establish mathematical relationships so that predictions can be made when changes
(e.g., new substituents) are incorporated into structures. Since a good amount of mathematics is
involved in this aspect, care should be taken such that the descriptors do not become just random
numbers for developing correlation equations, rather they should encode proper chemical information.
The experimental properties which are used as descriptors usually do not suffer from such problem,
since they are the true reflection of the physicochemical behavior of the chemicals. However, there
is possibility of incomplete information in theoretically derived parameters if they are not judged
properly. Proper care is required while dealing with the mathematics since it only acts as a tool to
define the final index employing several algebraic and arithmetic operations while the diagnosis
of chemical information entirely belongs to the part of chemistry. Hence, two required criteria of
molecular descriptors can be identified namely predictive ability and interpretability. From practical
examples, in many cases it has been observed that descriptors giving high correlation against an
endpoint of interest suffer from the lack of chemical diagnosis potential and vice-versa. A proper
52
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use

The “ETA” Indices in QSAR/QSPR/QSTR Research
Figure 2. The required features of theoretically derived descriptors
balance between this two makes a descriptor important for the QSAR analysis. This feature has also
been indicated in the definition of descriptors prescribed by Todeschini and Consonni (Todeschini
& Consonni, 2000) mentioned earlier as the ‘useful numbers’ which states the descriptors having
the insight of the molecular attributes as well as ability to take part in predictive modeling analysis.
Four of the basic features of chemical descriptors are shown in Figure 2. Here, the term ‘easy
computation’ refers to simplicity and explicitness of the descriptor algorithm since many of the
higher dimensional variants require complex pretreatment operations. The descriptor should also be
easily reproducible irrespective of the computational environment employed. The predictive ability
refers to the skill of the descriptor in giving good correlation with an endpoint under investigation
while the interpretability focuses on the chemical attributes of the analyzed molecular structures.
Once a proper basis of chemical or biological activity of the molecule is established, the tuning
of its features becomes easier and technically more specific. Two- dimensional descriptors are
usually simple, derived using hydrogen suppressed molecular graphs with a high degree of precision and consume less computational time than the higher order indices. In the beginning of the
development of graph theoretic descriptors, much more emphasis was given on the establishment
of quantitative correlation than the search for mechanistic basis. As a matter of fact, many of the
formerly defined descriptors were found to be flawed when chemometric modeling studies started
in a much larger scale.
This article presents a group of two-dimensional descriptors, extended topochemical atom (ETA)
indices, which were developed addressing several flaws of the previously described parameters. The
ETA indices comprise a few graph theoretic parameters as well as some electronic indices to describe
the bonding pattern, unsaturation, electronegativity and hydrogen-bonding related information. Till
date, the ETA indices have been used to model several endpoints employing different categories of
chemical compounds ranging from simple alkanes to complex drug molecules and agrochemicals.
An account of the definition of various ETA indices with an overview of their modeling status is
presented in the next sections.
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
53

The “ETA” Indices in QSAR/QSPR/QSTR Research
THE ETA INDICES
Brief Background
The ETA indices use information from hydrogen suppressed graphical presentation of molecules where
atoms are denoted by vertices while edges represent covalent bonds. The ETA formalism was derived as
an extension of another group of graph theoretic descriptors called topochemically arrived unique (TAU)
scheme indices. Before getting into details of the ETA indices, we would like to provide the readers with
some basic information and the motivation of developing TAU scheme parameters.
The branching parameter by Milan Randić (Randić, 1975) gave the measure of ramification of
branching in a graph theoretical topological problem. This idea was extended by Kier and Hall (Kier &
Hall, 1986), who gave a more comprehensive description on the type of branching by incorporating the
concepts of path, cluster and chain. The molecular connectivity index of Kier and Hall’s was observed
to produce good correlational results with some physico-chemical and biological endpoints. However, in
the late eighties Pal et al. (1988, 1989, 1990) observed several limitations in Kier and Hall’s parameters
and argued that predictive feature of this group of descriptors is biased, especially when higher order
indices are used. The following drawbacks of Kier and Hall molecular connectivity or chi (χ) indices
were identified by Pal and co-workers.
1. Kier and Hall’s molecular connectivity indices chiefly portray additive information with no con-
sideration of constitutive aspects.
2. It bears a good correlation with some quantum mechanical indices and mass spectroscopic frag-
mentation pattern and thereby possesses a similar pattern like logPo/w.
3. The molecular connectivity index formalism does not include information for steric interaction,
stereoisomerism, resonance and aromaticity.
v
4. The vertex valency term δ
often shows redundancy and becomes inconsistent when atoms of dif-
ferent quantum levels are considered.
5. Unlike vertex valency, the relative weight of the edges is not properly distinguished which leads
1
to the fact that only
mχv
i.e.,
that
(m>1), were found to possess good correlation when 1χv showed similar result and reflect
mχv
is strongly correlated with 1χv.
χ and 1χv have graph theoretical meaning. Furthermore, higher order indices,
6. The sigma and pi bonds in a system are equally weighted in this scheme.
7. Suitable adjustment in the molecular structure of several compounds of varied structures which
mχv
do not possess optimum biological response in the dataset can yield the ideal value of
obtained
from the regression equation.
With the aim of circumventing such undesirable shortcomings of molecular connectivity indices, Pal
et al. (1988, 1989, 1990) proposed a new set of topological descriptors namely TAU indices emphasizing the nuclear and electronic features of the atoms. In the TAU formalism, the vertex in a connected
molecular graph has been considered to be composed of a core and a valence electronic environment.
The valence electronic environment has been further partitioned into a localized and a mobile electron
measure. Table 1 shows the formal definitions of various TAU parameters. It should be noted that other
54
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use

( )
′
θ
+ + + =π ν
( )
θ ν π= +
V
V
=
The “ETA” Indices in QSAR/QSPR/QSTR Research
Table 1. Definition of the descriptors under the topochemically arrived unique (TAU) scheme formalism
Sl.
No.
1 Core count
2 Valence
3 Vertex valency
4 Edge weight
5 Molecular
6 Molecular
7 The composite
Type Mathematical Definition Note (If Any)
electron count
index (1st
order)
index (Zero
order)
index
Parameter for the Vertex Core
v
′
,
′
ij
′
i
i j
Z Z
Z
1
2
( )
i j
v
(VEM)
θ θ
= −
= − + +
=
i
E E V V
ij ji i j
when and connected
=∑0 5.
0
0 5.
′
8
8 2 1 5 2h l.
( )
λ
i
θ
i
= = ×
=∑.
( )
≠
i j i j
0
otherwise
=
0 5T V
,
ij
i
ν
1
2
Core count = =−λ
where Z=atomic number, Zv=number of valence electron, i.e.,
number of positive charges not neutralized by core electrons.
Parameters for the Localized and Mobile Valence Electron Environment
Valence Electron Local (VEL) Valence Electron Mobile
′
= + +θ ν2 1 5 2h l.
h is the number of bonded
H-atoms, ν is the number of
sigma bonds other than with
hydrogen and l is the lone pairs
of electrons.
λ
i
′
=
i
′
θ
i
′=′
E E V V
when and connected
T E
0
T V
T E V V
=′×
ij ji i j
′
′
= = ×
( )
i j i j
≠
0
otherwise
0 5. T E
=
∑
0 5
=
.
∑
∑ ∑
ij
< <
i j
λ shows a regular change with
Pauling’s electronegativity (ε
the period and along the group in an
opposite manner. 1/λ gives a rough
measure of the strength of the positive
field of the atomic core.
gives a count of the relatively
localized valence electron that the
positive field of the atomic core enjoys
while θ represents the count of
relatively mobile electrons enjoyed by
the field of the core.
The number ‘8’ in the definition of
VEM denotes the number of electrons
in the valence octet of a bonded atom.
It was assumed that an atom sigma
bonded to another non-hydrogen atom
shares 50% of the other electron
through the sigma bond apart from its
own electrons.
For a covalently linked organic atom,
h l
number of pi-bonds.
′
So,
θ ν π8 0 5 2.
= − +
0
) across
4 , π being the
and
0 5 2. .
Isosteric group or isovertices are
characterized by similar θ or θ′ values.
Parameter for Branching and Functional Group
8 Functionality
parameter
9 Branching
parameter
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
F T T
= −
R
TR being the first order VEM molecular index of the corresponding
reference alkane.
B T T
= −
N R
TN being the first order VEM molecular index of the normal alkane.
A reference alkane of a molecule
corresponds to a structure where all
heteroatoms are replaced with carbon
atom and multiple bond (covalent) with
single bond. The branching parameter
and the functionality parameter can be
correlated as:
T T F T B F
= − = − −
R N
55
Соседние файлы в папке Библиотека им академика М.И. Перельмана
