Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5443_Библиотеки_им_академика_М_И_Перельмана
.pdf
14.3 Description of the principal component
analysis method
PCA is by far the most common and wide-spread multivariate and multidimensional
statistical analysis method. According to Karl Pearson, one of the main founders of PCA,
one of the most important pillars of science is the desire to collapse multidimensional
data scattered over different (and sometimes heterogeneous) dimensions or descriptors
into a lower number of relevant dimensions [8]. This conformed to Bellman’s views in
1961, who introduced his strange idea that sometimes “Less is more” and talked about
the “curse of dimensionality” and that sometimes increasing the number of dimensions
of the data leads to loss of important information and observations [9], which can only
become obvious by reducing dimensionality [10].
The concept of reducing dimensionality in PCA by collapsing the dimensions into
best fitting linear projections is different from the tra ditional linear regression [11].
Usually, the least squares method deals with the variables on the right side of equa-
tions as independent and those on the left side as dependent. This implicates that the
minimization of the sum of squared error distances by choosing the best fitting line
deals with the dependent (y) variable only. The (x) factor “variable” is dealt with as
independent and it is the consequence of the researcher selection such as time and
dose that are suggested to be highly controlled and hence not dealt with in the least
squares computation.
The novelty of PCA lies in taking the error of all dimensions in consideration as
both “x” and “y” or rather “x1andx2 and so on,” subject to deviations or errors equally.
Therefore, the error lines in PCA are perpendicular to the best fitting linear projection
passing between the scattered points while in the classical linear regression, the error
distances are calculated normally to imaginary lines, parallel to the x-axis.
Figure 14.2: Molecular descriptors and machine learning in drug loading prediction. Reprinted from
Abd-algaleel et al. [7] under the terms and conditions of the Creative Commons Attribution (CC BY) license
(https://creativecommons.org/licenses/by/4.0/).
14 Role of principal component analysis in drug formulation and delivery 333
https://t.me/med1917

Figure 14.3 demonstrates the difference between the two concepts.
Usually in PCA, a covariance matrix is generated between the different independent
variables of factors in order to obtain an equal number of principal components.
For example, if we have three different variables, then three corresponding principal
components PC1, PC2, and PC3 will be generated. The variation in the collected data and
differences in the scales of the used variables should be prior normalized to unit variance
(1/SD) before analysis [13]. Correlation matrix is to be utilized in order to compute the
data principal components. These are actually the generated eigenvectors of the matrix.
For each eigenvector, a corresponding eigenvector is generated and accordingly,
the calculated eigenvectors are ranked with respect to their eigenvalues.
Consequently, the eigenvector with the largest eigenvalue will be considered the
main principal component, followed by the second principal component possessing
the second-highest eigenvalue, and so on [11].
The covariance matrix is usually presented as follows:
A =
cov x, xðÞcov x, yðÞcov x, zðÞ
cov y, xðÞcov y, yðÞcov y, zðÞ
cov z, xðÞcov z, yðÞcov z, zðÞ
0
B
@
1
C
A
(14:1)
It is a square and symmetrical matrix, where for an n-dimensional data set,
n!
n − 2ðÞ!
✶
2
different covariance values are generated.
In the above case, where we have three variables (three dimensions), x, y, and z,
we have three different covariances; cov(x, y), cov(x, z), and cov(y, z), bearing in mind
that cov(x, y) = cov(y, x), and so on.
Figure 14.3: A graphical comparison of the PCA optimization (left panel) versus the linear regression
optimization (right panel), as abstracted from Guiliani [12] through RightsLink copyright permission
license number:5633801180071.
334 Rania M. Hathout
https://t.me/med1917

The covariance is an important statistical tool that measures the correlation be-
tween two dimensions and is normally calculated as follows:
cov x, yðÞ=
P
n
i=1
ðX
i
−
XÞðY
i
− YÞ
n − 1ðÞ
(14:2)
The eigenvector of the covariance matrix (A) can simply be calculated after obtaining
the solution of the following characteristic equation:
det A − λIðÞ= 0 (14:3)
where I represents the identity matrix of the covariance matrix A containing three
variables (three dimensions), as follows:
100
010
001
0
B
@
1
C
A
For a 3 × 3 covariance matrix, a cubic polynomial equation is usually obtained where
its solution (its roots) leads to three values for λ, which are the matrix’ eigenvalues
(λ
1
, λ
2
, and λ
3
).
The solution can be obtained from Wolfram Alpha engine or software such as Mathe-
matica or utilizing Python
®
(Centrum Wiskunde and Informatica, Amsterdam, Netherlands).
Later, the corresponding eigenvectors can be obtained as follows:
cov x, xðÞcov x, yðÞcov x, zðÞ
cov y, xðÞcov y, yðÞcov y, zðÞ
cov z, xðÞcov z, yðÞcov z, zðÞ
0
B
@
1
C
A
a
b
c
0
B
@
1
C
A
= λ
1
a
b
c
0
B
@
1
C
A
(14:4)
These will yield three linear equations for a, b, and c, from which a relation between
them can be deduced. Any vector in the form
a
b
c
0
@
1
A
satisfying the relation between a,
b, and c is an eigenvector corresponding to the eigenvalue λ
1
.
Similarly,
cov x, xðÞcov x, yðÞcov x, zðÞ
cov y, xðÞcov y, yðÞcov y, zðÞ
cov z, xðÞcov z, yðÞcov z, zðÞ
0
B
@
1
C
A
a
b
c
0
B
@
1
C
A
= λ
2
a
b
c
0
B
@
1
C
A
(14:5)
Equation (14.5) will yield three linear equations, giving another relationship between
a, b, and c, and consequently, another eigenvector. The same calculations go for λ
3
.
A feature vector that represents the main principal component (the vector of the
highest eigenvalue) is accordingly selected.
14 Role of principal component analysis in drug formulation and delivery 335
https://t.me/med1917

The new derived data “final data” according to the obtained feature vector is de-
rived as follows:
Final data = Row feature vector × rowzero mean data (14:6)
In the same context,
Final data = Feature vector
T
× mean-adjusteddata
T
(14:7)
This is because the row feature vector is actually the feature vector transposed, and
the row zero mean data represents the mean-adjusted data but transposed.
The mean-adjusted data is obtained by calculating the mean of each dataset and
then subtracting it from all the data values.
Simplifying the whole data into only one or two dimensions (according to the fea-
ture vector representing only the first or sometimes the first and the second principal
components) is one of the most important tasks of PCA.
Reducing dimensionality may help in better visualization of the data and hence
deducing more conclusions and inferences.
Presenting the data in lower dimensions helps for clustering purposes as well, as
we can see in the next section of this chapter.
Figure 14.4 demonstrates the different steps of the method.
Figure 14.4: Summary of the principal component analysis method.
336 Rania M. Hathout
https://t.me/med1917

Figure 14.5: Characteristic plots of principal component analysis: (a) scree plot; left panel, eigenvalues
versus principal components and right panel, the proportion/cumulative variability versus principal
components; (b) score plot; observations projected onto PC1 and PC2 (obtained from Hathout et al. [14]
through RightsLink copyright permission license number: 5631520326995); and (c loadings plot; variables
projected onto PC1 and PC2 (obtained from Hathout [5] through RightsLink copyright permission license
number: 5631521229889).
14 Role of principal component analysis in drug formulation and delivery 337
https://t.me/med1917

14.4 Generated plots of the principal component
analysis method
There are different of characteristic plots that are generated from the principal com-
ponent technique that are usually generated through the different available software
and data analysis programs such as Unscrambler
®
, SAS/JMP
®
, Matlab, and XLSTAT or
through the R-programming language:
– The scree plot presents the eigenvalues versus the various generated principal
components or rather the proportion or the cumulative variance for each of
these principal components.
– The score plot presents the observations projected into the obtained main princi-
pal components (usually up to three principal components).
– The loadings plot presents the original variables projected in the obtained main
principal components (e.g., PC1 and PC2).
Figure 14.5 illustrates the different types of principal component plots.
14.5 Applications of PCA in drug formulation
and delivery
14.5.1 Exploiting PCA in determining the stability of soft
nanoscale drug carriers
The first attempts to utilize the clustering power of unsupervised machine learning meth-
odssuchasthePCAandHCAindetectingthestable microemulsion formulations (nano-
scale soft carriers possessing droplet size of 10–150 nm [15]) goes back to 2010, where
lovastatin- and glibenclamide-loaded self-microemulsifying (SMEDDS) and nanoemulsiy-
ing drug delivery systems were investigated[16,17].PCAandHCAwereusedtofurther
characterize the formed microemulsion formulations in the presence of water, with re-
spect to the similarity of particle size distribution obtained at 2 and 24 h (Figure 14.6).
Similarly, in 2015, two finasteride SMEDDS systems (one comprised of Capryol
90
®
, Tween 80
®
and Transcutol
®
while the other was composed of Labrafac cc
®
,
Tween 80
®
and Labrasol
®
) were prepared. Different formulations of the two systems
were prepared and characterized for particle size and polydispersity index for 30 min
and 24 h post preparation [18]. PCA and HCA were utilized to detect the stable formu-
lations, as indicated by clustering the same formulation for 30 min and 24 h in the
same cluster (see Figure 14.7).
It was concluded accordingly that the microemulsion formulations 2, 4, and 5 of
the first system Capryol 90
®
/Tween 80
®
/Transcutol were stable. Likewise, the micro-
338 Rania M. Hathout
https://t.me/med1917

emulsion formulations of the Labrafac cc
®
/Tween 80
®
/Labrasol
®
systems 1, 2, 4, and 6
were considered stable [19].
More research focus on utilizing machine learning methods to aid in selecting
promising formulations possessing good stability would be a great asset from the indus-
trial point of view. The concept can be projected to any drug delivery system as well.
14.5.2 Capturing the highly loaded drugs on polymeric
nanocarriers utilizing PCA
Gelatin matrices (microparticles and nanoparticles) are gaining high interests due to the
high biocompatibility, safety, availability and abundancy of the reactive functional
groups [20] that can be exploited to conjugate wide varieties of drugs and ligands [21, 22]
ofgelatinasaproteincarrier[23–26]. Accordingly, the machine learning methods – the
PLS, PCA, and HCA – were successfully utilized to cluster and /or capture the highly
loaded drugs on the gelatin matrix (Figure 14.8)fromasetof10drugs,namely,Acyclovir,
Amphotericin B, Cryptolepine, Curcumin, Doxorubicin, Indomethacin, Isoniazid, Resvera-
trol, Paclitaxel, and 5-fluorouracil. Focusing on the PCA method, the drugs were clustered
with respect to the main first two principal components generated from four main consti-
tutional, electronic and physicochemical descriptors; number of H-bond donors, number
of H-bond acceptors, molecular weight, and xLogP [27]. Interestingly, Isoniazid and 5-FU
were closely clustered together, demonstrating the least distance between any two drugs
Figure 14.6: PCA of droplet size of different lovastatin-loaded SMEDDS formulations after re-dispersion at
2 h and 24 h, as abstracted from Singh et al. [16] and under RightsLink copyright permission license
number: 5634331381387.
14 Role of principal component analysis in drug formulation and delivery 339
https://t.me/med1917

Figure 14.7: Clustering of (a) the formulations of the system Capryol 90
®
/Tween 80
®
/Transcutol (b) and
the formulations of the system Labrafac cc
®
/Tween 80
®
/Labrasol
®
. Left panel represents HCA while the
right panel represents PCA, according to the two principal components. Modified with permission from
Future Medicine Ltd.
®
through Copyright Clearance Center
®
license ID: 1074183-1.
340 Rania M. Hathout
https://t.me/med1917

in the set. Meanwhile, these specific two drugs scored the highest drug-loading on the
gelatin matrix, as proven in [28, 29]. The other artificial intelligence and machine learning
methods, whether supervised or not, such as the HCA, PLS, and molecular dynamics inte-
grated with the molecular docking experiments, confirmed these results [3].
14.5.3 PCA in drug–excipient compatibility and interactions
PCA or the factor analysis method has recently been used in detecting the incompati-
bilities between drugs and excipients in pharmaceutical formulations with high accu-
racy. In a study conducted on the interaction between theophylline and several
excipients, such as gum arabic, glucose, sorbitol, and sucrose, incompatibilities were
proven using the factor analysis method (PCA) by assessing differential scanning calo-
rimetry (DSC data) results [22]. In this study, binary mixtures of theophylline with ex-
cipients were prepared at ratios of 9:1, 7:3,1:1, 3:7, and 1:9 molar or mass ratios (where
the first value in the ratio presents the amount of theophylline in the mixtures). In
PCA analysis, the variables were the DSC parameters such as enthalpies, onset tem-
peratures, peak temperatures, peak heights, and peak widths. The samples were con-
Figure 14.8: PCA score plot of 10 several drugs, according to four constitutional, electronic and physico-
chemical descriptors. The inset is a scree plot demonstrating the cumulative variance percentage of the
four principal components (modified from [3] under Creative Commons public use license CC-BY).
14 Role of principal component analysis in drug formulation and delivery 341
https://t.me/med1917

sidered as theophylline, excipients and their mixtures, at ratios of 9:1, 7:3, 1:1, 3:7, and
1:9. Figure 14.9 shows the DSC curves of theophylline with microcrystalline cellulose
(MCC) at different ratios [30].
Sometimes, it is hard to decide whether there is compatibility or not between the
mixture materials in the DSC data, for example, theophylline mixtures with MCC. In
instances such as these, the factor analysis methods such as PCA are considered very
convenient tools in order to determine incompatibilities. The outcome of PCA can be
figured out on the ordinary two-dimensional score plot. The plot is informative about
the presence of incompatibility or not according to the locus of both the drug and the
excipient and their blends on the PCA score plot.
In interpreting the results, if the drug with mixtures at its highest amounts and
the 1:1 blend form a separate cluster and the other counterpart cluster is comprised of
Figure 14.9: (a) DSC curves of a: theophylline, g: microcrystalline cellulose and their mixtures at API/
excipient ratios, b: 9:1, c: 7:3, d: 1:1, e: 3:7, f: 1:9; (b) DSC curves of a theophylline, g: sorbitol and their
mixtures at API/excipient ratios, b: 9:1, c: 7:3, d: 1:1, e 3:7, f: 1:9 (modified from Khajavi [30] under the
terms of the Creative Commons Attribution License).
342 Rania M. Hathout
https://t.me/med1917
Соседние файлы в папке Библиотека им академика М.И. Перельмана
