Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5865_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
31.08.2026
Размер:
34 Мб
Скачать
An Introduction to the Basic Concepts in QSAR-Aided Drug Design
12
Frog (FRee Online druG) is a tool for generating small molecular conformations by using 1D or
2D descriptions. http://bioserv.rpbs.jussieu.fr/cgi-bin/Frog
13
Smi23d or (3D Cordinate Generation) program is able to convert one or more SMILIES into 3D
structure.
http://www.chembiogrid.org/cheminfo/smi23d/
14
ADRIANA.Code is used to compute molecular structure descriptors in drug design and predict
the ADME/toxic properties. https://www.molecular-networks.com/products/adrianacode
15
Molconn-Z is a software which creates Molecular Connectivity, Shape, and Information Indices
for (QSAR) Analyses. The new E-State parameter is also included. http://www.edusoft-lc.com/ molconn/
16
A software that calculates 1875 molecular descriptors (i.e., 1444 1D, 2D descriptors and 431 3D
descriptors and fingerprints). http://padel.nus.edu.sg/software/padeldescriptor/
17
KNIME is an open source software for intelligent algorithms including the data mining, machine
learning, and others. https://www.knime.org/
18
RapidMiner is mostly used for data integration, data transformation, data modeling, and data
visualization.
http://rapidminer.com/
19
Weka (Waikato Environment for Knowledge Analysis) is a general machine learning software
including some common statistical, clustering, and intelligent system algorithms written in Java.
http://www.cs.waikato.ac.nz/ml/weka/
20
An Open source program based on visual programming or Python used for data visualization and
analysis.
http://orange.biolab.si/
21
TANAGRA is a free software which includes some data mining methods and is capable of ana-
lyzing the through statistical and, machine learning algorithms. http://eric.univ-lyon2.fr/~ricco/ tanagra/en/tanagra.html
22
MATLAB (matrix laboratory) is a numerical computing environment which inclides various types
of toolboxes for artificial neural networks, image processing, bioinformatics, parallel computing, optimization procedures and etc. http://www.mathworks.com/products/matlab/
23
R is a free software environment alike MATLAB for statistical and machine learning computing
and graphical illustrations. Many packages over the world are being written for the purpose of general and expert researches. http://www.r-project.org/
24
SYBYL® is considered as the heart of Tripos. This software has various applications including
construction, editing, and visualization tools, modeling and simulation of the molecules and lead identification, and lead optimization. SYBYL-X is the latest version.
25
Discovery Studio is developed by Accelrys used on small/macro molecule and simulation.
http://accelrys.com/products/discovery-studio/
26
MOE is a software used generally for life and material sciences by comprehensively performing
the visualization, simulations on molecular structures.
http://www.chemcomp.com/MOE-Molecular_Operating_Environment.htm
46
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
An Introduction to the Basic Concepts in QSAR-Aided Drug Design
27
CODESSA is a tool for obtaining descriptors to develop QSAR/QSPR model through quantum
mechanical results from sources such as AMPAC, Gaussian, AIMALL. http://www.semichem. com/codessa/
28
It is a 3D-QSAR program calculating of total interaction energies between different chemical
probes and molecules. It determines energetically favorable binding sites on known molecular structures as well as surface properties. GRID 22c is the latest version. http://www.moldiscovery. com/software/grid
29
Molecular Discovery is one of the companies in the field of structure-based drug design. All the
developed programs, based on GRID force field, are designed for 3D-QSAR studies, virtual screen­ing, scaffold-hopping, ADME and pharmacokinetic modelling, optimization of metabolic stability and metabolite prediction. http://www.moldiscovery.com/
30
Pentacle is a alignment-independent 3D-QSAR software. It computes interaction energies between
the molecule and chemical probes (GRID based Molecular Interaction Fields, or MIFs). The latest version of Pentacle is 1.0.7.http://www.moldiscovery.com/software/pentacle
31
GOLPE (Generating Optimal Linear PLS Estimations) is an advanced chemometric toolbox for
building, validating and interpreting the 3D QSAR models.
http://www.miasrl.com/golpe.htm
32
The toolbox is especially suitable for chemical researchers to assess the toxicity of the chemical
compounds. The latest version is 3.2. http://www.oecd.org/chemicalsafety/risk-assessment/theo­ecdqsartoolbox.htm
33
REACH stands for Registration, Evaluation, Authorisation and Restriction of Chemicals which
has been approved by European Union in 2006 for assessing the chemical effects on health and environment.
34
TOXNET (TOXicology Data NETwork) is a group of toxicology dataset in various sections in-
cluding the chemicals, drugs and environmental effects. http://toxnet.nlm.nih.gov/
35
SIDER stands for Side Effect Resource which includes the information about side effect frequency
of drugs. And the latest version of SIDER is 2.0 and can be found at the link ttp://sideeffects.embl. de/
36
ACToR is a tool for the interested researches in chemical toxicity to extract data about the possible
risk of current chemicals on human health. http://www.epa.gov/actor/
37
DailyMed provides the most comprehensive information about the marketed drugs which are ap-
proved by Food and Drug Administration (FDA). http://dailymed.nlm.nih.gov/dailymed/about.cfm
38
Distributed Structure-Searchable Toxicity (DSSTox) Database Network is a project of EPA for
publicly presenting the data about structure-activity and predictive toxicology capabilities
http://www.epa.gov/ncct/dsstox/.
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
47
48
Chapter 2
The “ETA” Indices in QSAR/
QSPR/QSTR Research
Kunal Roy
Jadavpur University, India
Rudra Narayan Das
Jadavpur University, India
ABSTRACT
Descriptors are one of the most essential components of predictive Quantitative Structure-Activity/ Property/Toxicity Relationship (QSAR/QSPR/QSTR) modeling analysis, as they encode chemical in­formation of molecules in the form of quantitative numbers, which are used to develop mathematical correlation models. The quality of a predictive model not only depends on good modeling statistics, but also on the extraction of chemical features. A significant amount of research since the beginning of QSAR analysis paradigm has led to the introduction of a large number of predictor variables or descriptors. The Extended Topochemical Atom (ETA) indices, developed by the authors’ group, successfully address the aspects of molecular topology, electronic information, and different types of bonded interactions, and have been extensively employed for the modeling of different types of activity/property and toxicity endpoints. This chapter provides explicit information regarding the basis, algorithm, and applicability of the ETA indices for a predictive modeling paradigm.
INTRODUCTION
Chemicals are the inseparable components of human life, and thus they can affect a wide variety of indus­trial, agricultural as well as household processes if not monitored properly. Elucidation of the chemistry of compounds is an essential task of the designer so as to enable easy tuning of their properties. The diverse manifestation of chemicals can be systematically explored when the study of their chemistry is suitably attended with the related biology (including toxicology), mathematics and accompanying statis­tics. Since the past, mathematics has been considered as an effective media of natural science (Wigner,
1960) and has been employed in different branches of science to derive logical relationship employing abstract ideas. Chemistry has a very good coherence with mathematics that leads to the quantification
DOI: 10.4018/978-1-4666-8136-1.ch002
Copyright © 2015, IGI Global. Copying or distributing in print or electronic forms without written permission of IGI Global is prohibited.
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
The “ETA” Indices in QSAR/QSPR/QSTR Research
of chemical information by providing suitable platform for the generation of various algorithms. How different chemicals act differently and even the same chemical elicits different responses give the hint to nurture the ‘chemistry’ of compounds and explore the possible quantitative phenomena involved. The attempt to quantify the chemistry of compounds has lead to the development of several independent parameters termed as predictor variables or descriptors. These descriptors are further utilized to bring under a quantitative relationship with the responses of the chemicals commonly termed as the quantita­tive structure-activity/property/toxicity relationship (QSAR/QSPR/QSTR) studies. Mathematics plays a great role in deriving such predictor variables as well as in the development of correlations.
Through the journey of QSAR analysis, several predictor variables have evolved by exploring the concepts of chemistry, mathematics and physics. However, the ‘graph theory’ has contributed much to the development of initial ideas about descriptors. The graph theory, one of the branches in mathematics, can be traced back to Euler’s Königsberg bridges problem (Euler, 1736) which led to the idea of solving problems in different fields of science namely the study of electrical circuit by Kirchhoff (Kirchhoff,
1847), chemical isomer enumeration by Cayley (Cayley, 1875) etc. In a true mathematical sense, the term ‘graph’ also denotes cartesian plot of data, although it was Sylvester (Sylvester, 1878) who coined the term from a contemporary chemical perspective. Application of the graph theory in chemistry leads to the concept of ‘chemical graph’ which was supposed to have its origin in the late eighteen century holding the hand of the Scottish chemist Cullen (Crosland, 1959). William Cullen introduced the applica­tion of ‘affinity diagrams’, although it remains controversial with similar type of contributions by Black (Crosland, 1959). Later Higgins (Higgins, 1789) employed the use of graphical diagrams to denote the forces between atoms in a molecule. It was Dalton (Dalton, 1808) in the nineteenth century who started the concept of ball-and-stick representation of chemical models. The idea of diagrammatic presentation of chemical compounds was very primitive during that age and through various notable discoveries the chemistry got explored gradually. The famous scientist Kekulé brought the idea of three-dimensional orientation (Fischer, 1974) and later idea of chemical valence came likewise. Considering the scope of this chapter, we would like to limit this discussion at this point.
It was in the twentieth century when the chemical graphs were explored properly, and utilizing the concepts of topological distance measures the scientific fraternity started developing theoretical variables to correlate activity of chemicals with their chemistry. Although the history of QSAR starts back in the past when Cross (Cross, 1863) reported his idea of possible correlation between chemical composition and toxicity of alcohols, the concept of theoretical descriptor did not come in action properly until the year 1947 when the Wiener index (Wiener, 1947) and the Platt number (Platt, 1947) were introduced as the first two graph theoretical descriptors for modeling boiling point of hydrocarbons. Following these milestone achievements in the realm of quantifi­cation of chemistry paradigm, a new wave was hit towards the exploration of graph theoretic parameters, and in the 1960s and 1970s the pathway was systematically imputed by the following pioneering contributions namely Simon’s minimal topological difference (MTD) concept (Simon, 1974), Balaban J (Balaban, 1982), Gordon and Scantlebury index (Gordon & Scantlebury, 1964), Schultz molecular topological index (MTI) (Schultz, 1989), Hosoya Z (Hosoya, 1971), Randić’s branching parameter (Randić, 1975), Kier and Hall’s molecular connectivity, Kappa shape, sub-graph count, flexibility and electrotopological state (E-state) atom indices (Kier & Hall, 1986), Zagreb index by Trinajstić and co-workers (Gutman & Trinajstić, 1972), Harary index (Plavšić et al., 1993), Verloop’s STERIMOL parameter (Verloop, 1987), etc. Graph theoretic descriptors are purely based on a two dimensional basis, whereas concepts of three dimension was also under parallel research and various predictor variables were also derived employing three dimensional basis. A torrential amount of research on the descriptors has led to the development of a wide variety of predictor variables till
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
49
The “ETA” Indices in QSAR/QSPR/QSTR Research
today. The present chapter gives an account of such a group of two-dimensional novel descriptors named extended topochemical atom (ETA) indices which have been developed in the twenty first century from the modification of another set of topologically arrived unique (TAU) scheme descriptors of the late eighties. The ETA indices are now available under two generations and have been used in various modeling studies to predict property, activity as well as toxicity endpoints.
Descriptors in QSAR Studies: The Impact and Required Quality
Scientific deductions are based on logical components that aids in deriving conclusion from an observa­tion. The chemoinformatics study also bears such inevitable components that ultimately help in devel­oping predictive mathematical models and also leave a plausible explanation for the reasons involved. The molecular descriptors in QSAR studies can be viewed as the quantitative presentation of chemical structures obtained by processing the molecule using a specified mathematical algorithm. Todeschini and Consonni (Todeschini & Consonni, 2000) have ascribed descriptors as the useful numbers or the standardized experimental results obtained as the final result of a logic and mathematical procedure that leads to the transformation of chemical information which is encoded within symbolic representation of molecules. So, the mathematics is here utilized to unwrap the chemical information of compounds in the form of numerical entities. Descriptors are the algorithmic mathematical entities that strongly adhere to the chemistry of compounds and thereby form a core to the QSAR analysis. Since, the attri­butes of chemical compounds are depicted using different levels and dimensions of theory of chemistry and physics, descriptors are also defined on the basis of such theory coupled with a definite algorithm of mathematics. A good number of descriptors are available now-a-days corresponding to different algorithms which can be conveniently categorized from the dimensional perspective. Figure 1 shows a representative dimensional classification of various descriptors for QSAR analysis.
Figure 1. Dimensional perspective of descriptors used in predictive modeling (QSAR) analysis
50
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
The “ETA” Indices in QSAR/QSPR/QSTR Research
Descriptors can be defined for a whole molecule or they may also refer to specific fragments of the molecule. Again from the basis of computation, descriptors can be theoretical or experimental. Experimental descriptors correspond to various derived physicochemical properties namely parti­tion coefficient, molar refractivity, optical rotation, pK
values, reaction rate, parachor etc. and
a
they usually depict characteristics of the entire molecule while theoretical descriptors are usually computationally derived mathematical entities and can define specific molecular fragment as well as the whole molecule. An interesting fact about descriptors is that they evolved from lower to higher dimensions as the practice of QSAR made its progress and so the theory. Looking back at the his­tory of QSAR, it can be observed that Cross (Cross, 1863) was the first who observed the toxicity of primary aliphatic alcohols to be possibly correlated with their water solubility. Crum-Brown and Fraser (Crum-Brown & Fraser, 1868) are inexorably believed to be first to report the physiologi­cal action as the mathematical function of chemical constituents during an investigation of muscle paralysis activity of strychnine derivatives. Richardson (Richardson, 1869) observed a proportional relationship of narcotic effect of primary alcohols with their molecular weight and almost after two decades Richet (Richet, 1893) found the cytotoxicity of ethers, alcohols and ketones to be inversely related with their aqueous solubility measure. Then Meyer (Meyer, 1899) and Overton (Overton,
1901) came with their decisive work by correlating narcotic potency with lipid/water partition coef­ficient. After that, at the beginning of the twentieth century Trabe (Trabe, 1904) and Seidell (Seidell,
1912) came with their correlation analysis employing surface tension and solubility and partition coefficient respectively. Then came the pioneering discovery by Hammett (Hammett, 1935) who initially showed relationship between reaction rates with acid dissociation constant values (pK
a
and the parameter σ (sigma) which afterwards led to a “sigma-rho” culture in the field of predictive modeling analysis. Some more notable research activities were performed during this time namely correlation of narcotic activity with partition coefficient and exploring thermodynamics of chemicals by Ferguson (Ferguson, 1939), study on the involvement of ionization and shape factors towards the bacteriostatic action of aminoacridines by Albert (Albert et al., 1941) and correlation studies on antibacterial action (E. coli) of sulfonamides with their ionization and shape parameters by Bell and Roblin (Bell & Roblin, 1942). In this journey, it may be observed that QSAR started with the concept of correlating toxicity of chemicals with aqueous solubility by Cross whereas the term ‘chemical constituents’ in Crum-Brown and Fraser’s hypothesis can be connected to the elemental composition of the compounds. Richardson and Richet found toxicity to be expressible in terms of molecular weight and aqueous solubility respectively. Hence, we can see that at the dawn of QSAR, scientists used different physicochemical properties, i.e., experimental descriptors as variables and at the same time the concept of zero dimensional parameter like elemental composition, molecular weight were also used. The practice of using physicochemical properties like partition coefficient, solubility etc. as descriptors was also done by other scientists namely Trabe and Seidell. The concept of extracting chemical information was getting boosted during this time and Hammett led to the discovery of a theoretical parameter σ (sigma) in connection with modeling reaction kinetics. The contributions by Albert et al. and Bell & Roblin are also notable since they were the first to focus on the ionization factor and shape factors as independent parameters to derive correlation. It was in 1947, Wiener (Wiener, 1947) and Platt (Platt, 1947) came up with some more light to the path with their first topologically derived graph theoretic descriptors namely Wiener index and Platt index respectively by developing QSPR models for the boiling points of paraffin hydrocarbons. The inception of graph theoretic parameters as chemical descriptors opened a big doorway in the field
)
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
51
The “ETA” Indices in QSAR/QSPR/QSTR Research
of theoretical descriptors, and since then a significant amount of research has taken place towards the development of more informative parameters. The discovery of Hammett (Hammett, 1935) was forwarded by Taft who introduced a steric parameter E
(Taft, 1952) that distinguishes polar, steric
s
and resonating effect exerted by chemicals and reported the first multiple regression method based linear free energy relationship (LFER) equation. In the early 1960s, several newer developments were made in the field of 2D-descriptors. In 1962 Hansch (Hansch et al., 1962) introduced a substituent hydrophobicity parameter π and later in 1964 Hansch and Fujita (Hansch & Fujita, 1964) extended the analysis by combining parameters of Hammett and Taft and developed LFER models comprising of σ, E
and π. The analysis by Fujita and Hansch has undergone further parabolic extension and
s
later it was also modified by Kubinyi (Kubinyi, 1976). Another concept of presenting molecular descriptors as the contribution of the parent moiety and substituents was depicted by Free-Wilson (Free & Wilson, 1964) which gave rise to ‘indicator variables’. The Free-Wilson analysis was later modified by Fujita-Ban (Fujita & Ban, 1971) by incorporating logarithmic terminology.
It is to be noted that from the bases of the two-dimensional chemistry, concepts of three- dimen­sional (3D) molecular structure representation, electronic information and the descriptors thereof were also in the developing stage starting in 1930s. Some of the notable contributions can be cited namely the chemical bonding theory by Pauling (Pauling, 1939) and Coulson (Coulson, 1939), Sanderson’s (Sanderson, 1952) electronegativity parameters, electronic distribution theory by Fukui et al. (Fukui et al., 1954) and Mulliken (Mulliken, 1955), charged partial surface area measures of Jurs & Stanton (Stanton & Jurs, 1990), shadow indices (Rohrbaugh & Jurs, 1987), WHIM (To­deschini & Lasagni, 1994) & GETAWAY descriptors by Todeschini and co-workers (Consonni et al., 2002), 3D-MoRSE descriptors (Schuur et al., 1996), EEVA descriptors, CoMFA and CoMSIA analysis by Cramer et al. (Cramer et al., 1988) and the ligand-receptor based analyses including multidimensional (4D, 5D, 6D and so on) features in the modern century (Lill, 2007). However, considering the scope of this article we would like to focus our topic on two-dimensional topologi­cal structure representation only.
Hence it is observed that the molecular descriptors of chemical compounds are pivotal as they play a very crucial role in predictive modeling analysis. Through the journey of QSAR paradigm, the attempt to generate quantitative numbers using the chemistry of compounds has yielded various mathematical descriptors. Now, the obvious question arises regarding the required quality of such essential component of QSAR analysis. It is apparently true that quantitative predictor variables were introduced to establish mathematical relationships so that predictions can be made when changes (e.g., new substituents) are incorporated into structures. Since a good amount of mathematics is involved in this aspect, care should be taken such that the descriptors do not become just random numbers for developing correlation equations, rather they should encode proper chemical information. The experimental properties which are used as descriptors usually do not suffer from such problem, since they are the true reflection of the physicochemical behavior of the chemicals. However, there is possibility of incomplete information in theoretically derived parameters if they are not judged properly. Proper care is required while dealing with the mathematics since it only acts as a tool to define the final index employing several algebraic and arithmetic operations while the diagnosis of chemical information entirely belongs to the part of chemistry. Hence, two required criteria of molecular descriptors can be identified namely predictive ability and interpretability. From practical examples, in many cases it has been observed that descriptors giving high correlation against an endpoint of interest suffer from the lack of chemical diagnosis potential and vice-versa. A proper
52
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
The “ETA” Indices in QSAR/QSPR/QSTR Research
Figure 2. The required features of theoretically derived descriptors
balance between this two makes a descriptor important for the QSAR analysis. This feature has also been indicated in the definition of descriptors prescribed by Todeschini and Consonni (Todeschini & Consonni, 2000) mentioned earlier as the ‘useful numbers’ which states the descriptors having the insight of the molecular attributes as well as ability to take part in predictive modeling analysis.
Four of the basic features of chemical descriptors are shown in Figure 2. Here, the term ‘easy computation’ refers to simplicity and explicitness of the descriptor algorithm since many of the higher dimensional variants require complex pretreatment operations. The descriptor should also be easily reproducible irrespective of the computational environment employed. The predictive ability refers to the skill of the descriptor in giving good correlation with an endpoint under investigation while the interpretability focuses on the chemical attributes of the analyzed molecular structures. Once a proper basis of chemical or biological activity of the molecule is established, the tuning of its features becomes easier and technically more specific. Two- dimensional descriptors are usually simple, derived using hydrogen suppressed molecular graphs with a high degree of preci­sion and consume less computational time than the higher order indices. In the beginning of the development of graph theoretic descriptors, much more emphasis was given on the establishment of quantitative correlation than the search for mechanistic basis. As a matter of fact, many of the formerly defined descriptors were found to be flawed when chemometric modeling studies started in a much larger scale.
This article presents a group of two-dimensional descriptors, extended topochemical atom (ETA) indices, which were developed addressing several flaws of the previously described parameters. The ETA indices comprise a few graph theoretic parameters as well as some electronic indices to describe the bonding pattern, unsaturation, electronegativity and hydrogen-bonding related information. Till date, the ETA indices have been used to model several endpoints employing different categories of chemical compounds ranging from simple alkanes to complex drug molecules and agrochemicals. An account of the definition of various ETA indices with an overview of their modeling status is presented in the next sections.
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
53
The “ETA” Indices in QSAR/QSPR/QSTR Research
THE ETA INDICES
Brief Background
The ETA indices use information from hydrogen suppressed graphical presentation of molecules where atoms are denoted by vertices while edges represent covalent bonds. The ETA formalism was derived as an extension of another group of graph theoretic descriptors called topochemically arrived unique (TAU) scheme indices. Before getting into details of the ETA indices, we would like to provide the readers with some basic information and the motivation of developing TAU scheme parameters.
The branching parameter by Milan Randić (Randić, 1975) gave the measure of ramification of branching in a graph theoretical topological problem. This idea was extended by Kier and Hall (Kier & Hall, 1986), who gave a more comprehensive description on the type of branching by incorporating the concepts of path, cluster and chain. The molecular connectivity index of Kier and Hall’s was observed to produce good correlational results with some physico-chemical and biological endpoints. However, in the late eighties Pal et al. (1988, 1989, 1990) observed several limitations in Kier and Hall’s parameters and argued that predictive feature of this group of descriptors is biased, especially when higher order indices are used. The following drawbacks of Kier and Hall molecular connectivity or chi (χ) indices were identified by Pal and co-workers.
1. Kier and Hall’s molecular connectivity indices chiefly portray additive information with no con-
sideration of constitutive aspects.
2. It bears a good correlation with some quantum mechanical indices and mass spectroscopic frag-
mentation pattern and thereby possesses a similar pattern like logPo/w.
3. The molecular connectivity index formalism does not include information for steric interaction,
stereoisomerism, resonance and aromaticity.
v
4. The vertex valency term δ
often shows redundancy and becomes inconsistent when atoms of dif-
ferent quantum levels are considered.
5. Unlike vertex valency, the relative weight of the edges is not properly distinguished which leads
1
to the fact that only
mχv
i.e., that
(m>1), were found to possess good correlation when 1χv showed similar result and reflect
mχv
is strongly correlated with 1χv.
χ and 1χv have graph theoretical meaning. Furthermore, higher order indices,
6. The sigma and pi bonds in a system are equally weighted in this scheme.
7. Suitable adjustment in the molecular structure of several compounds of varied structures which
mχv
do not possess optimum biological response in the dataset can yield the ideal value of
obtained
from the regression equation.
With the aim of circumventing such undesirable shortcomings of molecular connectivity indices, Pal et al. (1988, 1989, 1990) proposed a new set of topological descriptors namely TAU indices emphasiz­ing the nuclear and electronic features of the atoms. In the TAU formalism, the vertex in a connected molecular graph has been considered to be composed of a core and a valence electronic environment. The valence electronic environment has been further partitioned into a localized and a mobile electron measure. Table 1 shows the formal definitions of various TAU parameters. It should be noted that other
54
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
( )
θ
+ + + =π ν
( )
θ ν π= +
V
V
=
The “ETA” Indices in QSAR/QSPR/QSTR Research
Table 1. Definition of the descriptors under the topochemically arrived unique (TAU) scheme formalism
Sl.
No.
1 Core count
2 Valence
3 Vertex valency
4 Edge weight
5 Molecular
6 Molecular
7 The composite
Type Mathematical Definition Note (If Any)
electron count
index (1st order)
index (Zero order)
index
Parameter for the Vertex Core
v
,
ij
i
i j
Z Z
Z
1 2
( )
i j
v
(VEM)
θ θ
=
= + +
=
i
E E V V
ij ji i j
when and connected
=∑0 5.
0
0 5.
8
8 2 1 5 2h l.
( )
λ
i
θ
i
= = ×
=∑.
( )
i j i j
0
otherwise
=
0 5T V
,
ij
i
ν
1 2
Core count = =−λ
where Z=atomic number, Zv=number of valence electron, i.e., number of positive charges not neutralized by core electrons.
Parameters for the Localized and Mobile Valence Electron Environment
Valence Electron Local (VEL) Valence Electron Mobile
= + +θ ν2 1 5 2h l.
h is the number of bonded H-atoms, ν is the number of sigma bonds other than with hydrogen and l is the lone pairs of electrons.
λ
i
=
i
θ
i
′=′
E E V V
when and connected
T E
0
T V
T E V V
=′×
ij ji i j
= = ×
( )
i j i j
0
otherwise
0 5. T E
=
0 5
=
.
ij
< <
i j
λ shows a regular change with Pauling’s electronegativity (ε the period and along the group in an opposite manner. 1/λ gives a rough measure of the strength of the positive field of the atomic core.
gives a count of the relatively localized valence electron that the positive field of the atomic core enjoys while θ represents the count of relatively mobile electrons enjoyed by the field of the core. The number ‘8’ in the definition of VEM denotes the number of electrons in the valence octet of a bonded atom. It was assumed that an atom sigma bonded to another non-hydrogen atom shares 50% of the other electron through the sigma bond apart from its own electrons. For a covalently linked organic atom,
h l
number of pi-bonds.
So,
θ ν π8 0 5 2.
= +
0
) across
4 , π being the
and
0 5 2. .
Isosteric group or isovertices are characterized by similar θ or θ′ values.
Parameter for Branching and Functional Group
8 Functionality
parameter
9 Branching
parameter
EBSCOhost - printed on 2/14/2023 7:16 AM via . All use subject to https://www.ebsco.com/terms-of-use
F T T
=
R
TR being the first order VEM molecular index of the corresponding
reference alkane.
B T T
=
N R
TN being the first order VEM molecular index of the normal alkane.
A reference alkane of a molecule corresponds to a structure where all heteroatoms are replaced with carbon atom and multiple bond (covalent) with single bond. The branching parameter and the functionality parameter can be correlated as:
T T F T B F
= =
R N
55