Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_2617_Библиотеки_им_академика_М_И_Перельмана
.pdf
28
s
−
−
xx
ii
=−
Ü
M
=× −
()
bxx
x
Ü
https://t.me/medicina_free
M. Vidali
association of two linearly correlated variables. If they are
linked by a strong non-linear relationship (e.g., quadratic),
the correlation coefcient may be falsely low. This underscores, as noted earlier, the importance of always viewing
data before statistical analysis.
The goodness of t of the regression model to our data
can be assessed through the coefcient of determination R2,
which is the square of the correlation coefcient: R2 = r2.
Since r varies from −1 to +1, R2 varies from 0 to 1 and indicates the proportion of variability of Y explained by the
regression line.
Evaluation ofAnalytical Performance
Accuracy
The precision of a method can be dened as the degree of
agreement between different independent measurements
under dened experimental conditions. Precision can be
expressed in terms of imprecision as standard deviation, variance and relative standard deviation (RSD), or coefcient of
variation (CV). Several types of precision are dened:
• Repeatability (or within-run precision): precision under
experimental conditions that minimize other sources of
variability (same laboratory, same method, same instru-
ment, same operator, same operating conditions, with
measurements made within a short time interval).
• Intermediate precision (formerly total precision, recently
within-laboratory precision): precision under well-
dened experimental conditions (same laboratory, same
instrument, same method, with measurements made over
an extended time interval, with different calibrators, oper-
ators, and/or lots of reagents).
• Reproducibility: accuracy obtained under different condi-
tions (laboratory, operator, instrument). Reproducibility
can only be assessed by inter-laboratory studies.
Verication of repeatability with the protocol 5 days×5
replicates.
Before using a method validated by a third party (e.g., a
manufacturing company), the laboratory must perform a
rapid verication protocol. To this aim, various experimental designs are reported in the literature, including the
3×5 scheme (3 replicates for 5days); however, the most
recent literature suggests increasing the number of replicates and/or days in order to obtain a more realistic estimate of repeatability. It is therefore advisable to use a
5×5 scheme (5 replicates for 5days) or one in which the
product of replicates × days is at least ≥20. For methods
that are not validated by third parties but developed on
their own, a more extensive protocol should be adopted
(the literature suggests a 2× 20 scheme, i.e., 2 replicates
for 20days).
The repeatability verication protocol includes the fol-
lowing steps:
• Sample selection.
• Sample analysis (5 replicates × 5days) using the method
to be veried.
• Processing of any aberrant data and possible integration
with new measurements until the desired size is obtained.
• Calculation of repeatability.
• Verication of what has been declared by the producer.
The verication should be performed using at least two
controls, or pools of patients, with concentrations close to
both any decision values (and/or cut-off), and to those used
by the manufaturer in the validation phase and reported in
the data sheet.
If one or more observations from the collected data devi-
ate to a certain extent from all the others (extreme or aberrant
values), it is necessary to perform a statistical test for aberrant data, after checking that these values are not the result of
analytical or gross errors. A number of statistical tests useful
in identifying extreme values are described in the literature,
including the Dixon-Reed test, Grubbs’ test, Huber’s test,
and Horn’s method. In the Dixon-Reed test, the ratio D/R is
calculated, where D is the absolute difference between the
aberrant and the least extreme value immediately preceding
it and R is the maximum-minimum range of all the data; if
the ratio is greater than 1/3, the observation is considered
aberrant. In the Grubbs’ test, instead, we calculate the ratio
maxmean
=
or s=
mean min
, where min and max
are, respectively, the presumed minimum or maximum aberrant value, mean and s are, respectively, the mean and standard deviation of the sample. If the calculated G is greater
than a certain critical value, which can be found in the literature and depends on the sample size, the data is considered
an aberrant. Huber’s test is based instead on the median
absolute deviation (MAD), that is, the median of the absolute
differences (i.e., without sign) between the observations and
the median. The procedure is as follows: the median is calcu-
lated; then the absolute differences D
between the
median and the individual observations are calculatedand
then the MAD, that is, the median of all these absolute differences, multiplied by an appropriate constant b=1.4826(when
data follow a gaussian distribution), is obrained. Thus,
AD median
i
, where
indicates the median of
the observations. Finally, we calculate the limit value as
D
=MAD×3 (very conservative; or 2.5 moderately con-
crit

D
R
−
−
=>
3
./
−
9
x
λ
λ
− 1
xx
;;
()
=+
V
between between between within within with
= =
iin
x
j
SV
rW
=
%/
()
SV
BB
=
%/
()
100
SVV
WB
=+
%/
()
100
4 The Role ofStatistics inLaboratory Medicine
https://t.me/medicina_free
29
servative). If an observation has an absolute difference from
the median greater than this critical value (Di> D
), it is
crit
considered aberrant. Finally, in the Horn’s method, we rst
apply the Box-Cox transformation, to obtain a more Gaussian
distribution of the data, and then Tukey’s robust method, calculating the interval between Q1–1.5 × IQR and
Q3+1.5×IQR, where Q1, Q3, and IQR are the rst quartile,
the third quartile, and the interquartile range (Q3-Q1),
respectively: all data outside this interval are considered
aberrant.
Example: for the 25 replicates (5×5 design), the following data were obtained: mean=101.2; s=14.0; min=76.4;
max=150.0. The graphical analysis seems to suggest that
the value 150.0 is an aberrant. We test this hypothesis with
the four tests illustrated.
(a) Dixon-Reed test:
Given less extreme than the max: 113.2 then
150 0 113 2
..
=
150 76 4
.
0501
(b) Grubbs’ test: for N=25 G
..
150 0 101 2
=
14 0
=
.
34
aberrant.
=3.14
critical
since G > G
.
critical
we con-
clude that 150.0 is aberrant.
(c) Horn’s method: after having found through statistical
software λ =−0.8511372 and having transformed the
necessary, to transform the data. If a non-normal distribution is assumed for the data, the multiplicative constant
for Huber’s test changes from b= 1.4826 to b = 1/Q3,
where Q3 is the third quartile. Graphical analysis of the
data also allows to check for multiple aberrant data and
thus choose a suitable test. In this regard, it should be
noted that the Dixon-Reed test may not identify the aberrant data in the presence of more than one aberrant. In
addition, while Tukey’s test can identify multiple aberrants, Grubbs’ test identies one aberrant at a time, and it
is therefore necessary to apply the test to the data set several times after removing the suspected aberrant (a minimum data size of >6 is recommended for Grubbs’ test).
Finally, it is advisable not to use Tukey’s method for small
data sets.
After identifying any aberrant data, we proceed to calculate total or within-laboratory repeatability through analysis
of variance (ANOVA), as shown in the introductory section
on basic statistical methods:
jdi
11
r
ij j
2
x
d
DEVDEV
between within
DE
VVDEV DEV
totbetween within
DF DF DF DF DF
between within totbetween within
nx
jj
j
1
=− =−
ddr11;; ;
AR DEVDFVAR DEVDF
2
/; /
,
data according to Box-Cox, that is, with
the limits
of the interval Q1–1.5 × IQR and Q3 + 1.5 × IQR
(1.147–1.156) are calculated. The suspected aberrant
150.0, after transformation, is equal to 1.158, which is
outside the interval, conrming what was found with the
previous tests.
(d) Huber’s method: the median of the observations is equal
to 100.8; the absolute differences between this median
and all the observations are calculated and then the
MAD, which results to be 9.17 (already corrected for the
constant b). The critical value is D
=9.17×3=27.5.
crit
The observation 150.0, since it presents an absolute difference from the median equal to 49.2 (since Di>D
crit
49.2>27.5), is considered aberrant.
The most critical issues of these tests are the assumption of normality of the data distribution and/or the failure
to identify the aberrant data in the presence of more than
one extreme data. In fact, the reliability of the identication of an observation as aberrant depends on the distribution of the data (in particular for the Grubbs’ test and the
Dixon-Reed test). Before applying these tests, it is therefore necessary to check the Gaussian distribution and, if
with d being the number of days, r the number of replicates,
xij the i-th replicate of the j-th day,
the average of the j-th
day, and x the average of all replicates.
From the above data, we calculate the following:
The variance of repeatability (within-run or within-run):
VW=VAR
within.
The repeatability (within-run or within-run):
CV
=
Sx100.
rr
The “pure” between-run repeatability variance (i.e., cor-
rected for the contribution of VW): VB=(VAR
R
)/n0, where n0 represents the average number of rep-
within
;
licates of the runs (in the 5×5 design n0=5). If VAR
≤ VA R
, then VB=0.
within
between
The “pure” (between-run) repeatability:
CV
=
Sx
BB
Total repeatability or within the laboratory:
and
CV S
=
WL WL
.
WL
x
.
Example: we want to evaluate the repeatability of a
method for the measurement of cholesterol (5×5 design).
From the analysis of the data, we obtain:
and
− VA
between
e

30
SV
rW
= = =
SV
BB
= = =
SVV
WL WB
=+=+=
./
x =
./,
%.
.%
=
()
×=
%.
.%
=
()
×=
%.
.%
=
()
×=
Sx
pp
=
()
()
()
()
22
22
rb
S
b
2
SVVr
2
=+
SV
rw
0= =
SVVr
2
3=+ =+ =
xx tS
a
−< ×
r
62
08<× ×=
B xx=−
0
%
xx
x
−
100
x
x
0
100
%
CC
C
a
−
100
Befo
./
=+++
()
=
A
./
= +++
()
https://t.me/medicina_free
M. Vidali
DEV
DEV
Vw = 7.50; since VAR
: 132.56; DF
between
: 150; DF
within
between
: 20; VAR
within
between
: 4; VAR
within
between
: 7.50
: 33.14
(33.14) > VAR
(7.50) then
within
VB=(33.14–7,50)/5=5,13
750274../
513226../
Since the average of all 25 replicates is
mg dl
mg dl
750513 355..
mg dl
203 76
mg dl
then:
CV
CV
CV
r
B
WL
/.
274 203 76 100 134
/.
226 203 76 100 111
/.
355 203 76 100 174
The calculated total repeatability can be compared with
that stated by the manufacturer. The manufacturer can provide the total repeatability either as a standard deviation (Sp)
or as %CVp. In the latter case, to obtain the Sp, it is sufcient
to apply the formula
CV
%/100 .
In the cholesterol example above, the manufacturer provided an Sp = 3.20. If he had provided, for example, a
%CVp=1.6%, then Sp=1.6%/100× 203.76 = 3.26 would
have been obtained.
It is then calculated: T (effective degrees of freedom):
rSrS
=
−
1
r
d
1
−
⋅+⋅
2
2
⋅
+
S
()
r
2
⋅
rS
b
1
−
d
, where
2
is the “non-pure”
the documented repeatability at the introduction of the
method.
Example: Sr=2.74mg/dl (see previous example); degrees
of freedom (df) 20. Sr had been determined in a 5day × 5
replicate design, with df = days × (replicates−1)=5×(5–1)=20; if repeatability had been determined in a single run of N replicates, the df would have been
(N−1)); t
(a/2;df= 20)
x2=6, we get:
=2.086. If x1=202 and x2=196, x1 −
2 086 2748
... and thus it can be
assumed that the laboratory is operating under the established repeatability conditions.
Trueness andBias
Trueness can be dened as the closeness of agreement
between the mean value, obtained from a large number of
determinations of a sample, and the true value. Note that the
true value is dened as an accepted or assigned reference
value. Trueness is expressed as bias, or systematic error,
which can be calculated as the difference between the mean
value of the replicates and the true value or as the ratio of this
difference to the true value. Thus,
B
0
=
, where x represents the average of the val-
0
ues obtained by repeatedly analyzing the sample and x0 the
assigned value. Bias is often reported as recovery and
expressed as %R
=
, where the relation
or
between-run variance, that is, not corrected for the contribution of the within-run variance:
In our example: r = 5; d = 5;
513755 66
bBW
/. ./ . ; T = 12.475. Using
the calculated parameter T and the statistical table of the
chi- square distribution, we nd the critical value
C=23.337. Finally, we calculate a limit value (or verication value)
/ , where Sp represents the total
P
repeatability declared by the manufacturer. In our example,
Sp=3.20; V =4.38. Since the total repeatability obtained
experimentally (3.55) is lower than the limit value (4.38),
the total repeatability declared by the manufacturer is
veried.
If substantial changes are made to the method in use, it is
necessary to revalidate the method; however, even in the
absence of such changes, it is good laboratory practice to
periodically verify the repeatability demonstrated at the time
the method was introduced. This verication can be easily
conducted through the so-called duplicate test experiment:
two independent replicates of the same sample are analyzed
under repeatability conditions (x1 and x2). For the verication
to be positive, it must be
bBW
12 2
2
/ .
2
, where Sr is
/;df
75
. ;
%R=100+%B. Recovery is frequently evaluated by making repeated determinations of a sample before and after the
addition of a known amount of analyte; we have that
21
R
=
, where C1 and C2 are, respectively, the
mean value of the replicates of the sample before and after
addition of the analyte, while Ca is the expected increase in the
concentration of the analyte in the sample after addition. Ca can
be obtained from Ca=Cst(Vst/Vf), where Cst is the concentration
of the standard used for addition, Vst is the added volume of the
standard solution, and Vf is the nal volume of the sample (initial volume+added volume of the standard solution).
Example: 0.2 ml (Vst) of a standard calcium solution
21.0mg/ml (Cst) is added to a serum sample of 2.0ml volume. Four determinations are carried out before and after the
addition.
re additionC mg dl:..../
fter addition C
93 96 95 92 494
1
: ...
11 2110 10 9113 4
2
mg dl;
./
11 1
=

dl=×
()
=
%
.%
CC
C
a
−
−
sN
x tsN
aN±−
B tsN
aN±−
B =−=−
..
dl
%./.
B =−
()
×=−
.
%%
R
RB
=×=
=+=−=
...
−
9
236
1
..±× =−
05
..±× =−
x
N
−
df
()
N
2
CRM
±×
N
−
5
with
df
()
113
236
8
4 The Role ofStatistics inLaboratory Medicine
https://t.me/medicina_free
31
.
mg
=
100 89 0
21 02 22 191./ .. /
R
a
21
=
100
..
11 194
=
191
The recovery experiment allows to estimate the proportional systematic error, that is, a type of error that increases
as the concentration of the analyte increases, but not the constant systematic error; in fact, even in the presence of a
recovery close to 100%, it is not possible to exclude the presence of bias due, for example, to interferents or matrix
components.
Bias can also be assessed: (1) by measuring certied reference materials; (2) through a method comparison study.
Bias Assessment with a Certied Reference Material
A certied reference material (CRM), with an assigned concentration value (and associated measurement uncertainty)
and with a composition as close as possible to that of the test
samples, is analyzed N times. It is obtained that
x
known value
t
calc
mean and the standard deviation of the replicates. If t
greater than a tabulated t
−
=
/
, where x and s are, respectively, the
is
calc
, which depends on N and the
critical
chosen signicance level, then the bias is considered statistically signicant. The condence intervals of the mean value
of the N replicates and the bias are, respectively,
/;/21
and
/;/21
.
Example: The assigned value for a CRM for cholesterol is
250mg/dL.Ten measurements are made obtaining a mean of
236.8mg/dl and a standard deviation of 10.2mg/dl. You can
then calculate:
236 8 250 13 2
236 8 250 250 100 53
236 8 250 100 94 7
./ .
%% %
100 100 53 94 7
236 8 250
.
calc
=
10 210
./
=
40
.
/mg
%
or
For (10–1)=9 degrees of freedom and a 95% signicance level, we have that t
= 2.262; since t
a/2;9
calc
> t
crit
(4.09>2.262) we have that the bias is statistically signicant, that is, the mean measured value is different from the
known or assigned value (with a 95% condence level).
To conrm this, the 95% condence interval (95%CI) of
the mean value of the N determinations is
82262 10 210 229 5 244
.. ./
, which does not
include the assigned value of 250mg/dl. Similarly, the 95%
condence interval of the bias does not include zero (in fact:
32 2 262 10 210592
.. ./
).
A bias, although statistically signicant, could still be
acceptable if it is less than an error established on the basis
of different criteria (expert consensus, biological
variability).
When assessing bias by repeated measurements of a certied reference material, the investigator should consider both
the uncertainty of the method and that of the CRM itself. In
such a case, it follows that:
calc
degrees of freedom given by
where
rial. Note that, with
represents the uncertainty of the reference mate-
CRM
2
very small, the degrees of freedom
=
knownvalue
2
s
2
u
+
CRM
2
s
u
+
N
=
2
s
2
CRM
2
with the
2
−
N
1 ,
approximate to (N−1). The condence interval, including
the contribution of the CRM uncertainty, becomes
2
s
2
+t
u
critical CRM
.
In the previous example, the CRM has an uncertainty
equal to u
236 8 250
=
calc
10 2
.
10
.
10 2
10
=
10 2
Alsoin this situation, with t
= 1.4 mg/dl; it is then that
CRM
.
2
10
2
+
14
+
2
.
14
.
2
.
2
37
=
.
2
2
10
×−
=
calc
and t
>t
=2.160.
crit
(3.75>2.160), we
crit
have that the bias is statistically signicant. The 95% condence intervals for the mean value and the bias will be,
2
10 2
respectively,
and
32 2 160
..
82160
..
2
10 2
.
14 56 20
10
.
10
2
...±× +=− . As can be seen,
2
...±× += −
14 229 2 244 4
the condence intervals, calculated including the contribution of the CRM uncertainty, are only slightly wider than the
previous ones but the conclusion remains the same.
The Experiment of Comparison of Two Analytical
Methods
Systematic error can also be assessed through a method comparison experiment, where patient samples are measured
with two different methods or instruments. This approach
allows both calculation of the error at decision levels and

32
±+
35
Average of the 2 methods
Method 1 - Method 2
35
35
https://t.me/medicina_free
assessment of proportional (slope of the regression line) and
constant (intercept of the regression line) systematic error.
The comparison experiment between analytical methods
consists of these steps:
1. Familiarization with the new method and verication of
the analytical performance of the method under test.
2. Denition of acceptability levels for maximum allowable
error:
30
25
20
15
Method 2
10
5
M. Vidali
• Based on the combined imprecision of the methods: if
the two methods do not differ signicantly, their differences will be within the range:
196
.CVCV
2
method method
2
1
, where CV represents the
2
total repeatability of the method.
• Based on quality specications reported in the litera-
ture (e.g., TE
I
are the maximum imprecision and inaccuracy,
max
max
=B
+1.65 × I
max
, where B
max
respectively).
3. Sample selection and analysis: 40–60 patient samples
distributed over the entire reportable range (range in
which analyte concentrations can be measured with
acceptable accuracy and precision); a greater proportion
of samples should be selected around any clinical decision values and/or cut-offs.
4. Plotting and statistical analysis of the data: the most used
tools are the regression model and the Bland-Altman difference graph (Fig.4.5).
To conduct the regression analysis, the data obtained
with the reference or in-use method (x-axis) are plotted
against the data obtained with the method under test
(y-axis) together with the identity line (y = x) and the
regression line. Different regression models can be used
to nd the regression line. The widely used simple (or
ordinary) linear regression model has the merit of computational simplicity. However, it is not always correct to
use it. In fact, simple linear regression assumes that the
method in use (the one on the x-axis) is error-free and that
the error of the method under test (y-axis) is normally
distributed and is constant throughout the range of concentrations studied (homoscedasticity). When the
assumptions underlying the linear regression model are
not met, it is necessary to use another regression model
(Deming, weighted Deming, or the non-parametric
Passing-Bablok model). The non-parametric PassingBablok model requires the least number of assumptions.
In fact, in Passing-Bablok regression, extreme values
(outliers) can be included, imprecision is allowed in both
methods, and it is not necessary either that the error be
normally distributed or that it be constant over the concentration range. Regression analysis allows the assessment of the systematic constant (intercept) and
max
0
0
5 10 15 20 25 30
Method 1
1.5
1.0
and
Fig. 4.5 Passing-Bablok regression analysis and Bland-Altman graph
for the evaluation of systematic error in the comparison experiment
between methods. In the Bland-Altman graph the line corresponding to
the bias and the two lines delimiting the interval which includes 95% of
the differences between the methodsare shown
0.5
0.0
–0.5
–1.0
–1.5
–2.0
0 5 10 15 20 25 30
proportional (slope) error. It is always necessary to report
not only these two parameters, point estimates of the true
intercept and the true slope, but also their 95% condence
intervals (interval estimation), that is, the intervals in
which we have some condence (usually 95%) that the
true intercept and slope can be found. In particular, to
demonstrate the absence of constant, proportional systematic error, the relevant condence intervals must
include 0 (zero) for the intercept and 1 (one) for the slope.
Small proportional or constant systematic errors identied by regression analysis, while statistically signicant,
do not preclude the use of a particular method if the error
is less than the established maximum acceptable error. It
is also important to note that the correlation coefcient
should not be used to evaluate the agreement between two
methods; in fact, it allows us to evaluate only the association between two variables but not the agreement
(Fig.4.5).
The Bland-Altman plot, on the other hand, allows calculation of the bias between the two methods, expressed
as the mean of the differences. In this graph, the analyte
+1.96 DS
1.18
bias
–0.01
–1.96 DS
–1.21

%b
−×+
()
CCba
C
15
090
DIfference between
Average of method 1 and method 2
4 The Role ofStatistics inLaboratory Medicine
https://t.me/medicina_free
concentrations on the x-axis (expressed as the mean of the
two methods) are graphed toward the differences between
the two methods on the y-axis. Alternatively, the mean of
the two methods (x-axis) can be graphed toward the differences of the two methods expressed as a percentage of
the mean, that is, (method1 − method2)/mean × 100
(y-axis). This method is especially useful when the concentration range is wide.
The bias is then calculated as the mean of all differences and expressed along with the associated 95% condence interval. Signicant bias occurs when the value 0
(zero) is not within the relevant 95% condence interval.
In the absence of bias (bias=0), the points corresponding
to the differences between the two methods should accumulate randomly around the zero line, that is, positive
and negative differences close to zero should be observed
for each concentration level. It is also possible to observe
the presence of particular trends in the differences and/or
the widening of the agreement range with increasing concentrations, a sign of the dependence of the differences
on the concentration level. In addition to the mean difference (and the relative 95% condence interval), it is also
necessary to calculate the interval including 95% of such
differences dened by mean±1.96 × SD, where mean
and SD are, respectively, the mean (bias) and the standard
deviation of the differences found.
On the other hand, this interval, which is often overlooked and not reported, is of considerable importance in
the correct interpretation of the comparison experiment,
because it allows us to identify the magnitude of most
(95%) of the differences found. Alternatively, the interval
comprising 95% of the differences can also be calculated
using non-parametric methods, for example, as the interval between the percentiles 2.5 and 97.5 (Fig.4.5).
5. Evaluation of the acceptability of the new method: the
bias identied is compared with the maximum acceptable
error dened at the beginning of the experiment.
For acceptability based on the combined imprecision of
the two methods, we proceed as follows:
1. For each concentration level, the range dened by
0±1.96×CV%×C shall be calculated, where CV% is
the combined inaccuracy of the two methods and C is the
concentration. Since repeatability frequently varies over
the concentration range, different CV% must be used for
the two methods over the concentration range and different combined inaccuracies must be calculated.
2. Count how many method1-method2 differences fall outside the range. If the two methods are identical within the
combined imprecision, their differences should be symmetrically distributed around zero and 95% of them
should be within the interval.
33
10
5
0
–5
method 1 and method 2
–10
–15
0102030405060708
Fig. 4.6 Acceptability assessment based on combined imprecision
For a better visualization it is possible, in the same difference graph, to add the two lines representing the range
0±1.96×CV%×C (Fig.4.6).
To assess acceptability based on predened quality specications, the previously dened maximum total allowable
error is identied and a diagram (MEDx chart) is constructed,
plotting the imprecision (x-axis) against the systematic deviation (bias) (y-axis) and representing the performance of the
test under evaluation as a point. The coordinates of the point
are obtained using for the abscissa the imprecision (calculated as repeatability within the laboratory) and for the ordinate the bias (using the equation of the regression line applied
at a certain chosen concentration level, generally at clinical
decision levels or cut-offs). The chart also shows four zones
corresponding to different quality specications (not acceptable, marginal, good, excellent). For the construction of the
MEDx chart, we proceed as follows:
1. The point corresponding to the maximum allowableerror
is xed on the y-axis.
2. On the abscissa axis, we x three points corresponding to
the various quality criteria: (max/2 error), (max/3 error),
(max/4 error).
3. Three lines are drawn, starting from the point on the ordi-
nate axis to the three points on the abscissa axis, which
divide the graph space into four zones corresponding to
the different quality specications;
4. A point is represented using the performance of the
method under test as coordinates (abscissa = withinlaboratory repeatability; ordinate = bias as assessed by
regression analysis). If y=a+bx is the regression model,
at a given concentration the bias will be given by
ias =
, where C is the expected
×
100
concentration, a and b, respectively, the intercept and

16
8
Bias %
Imprecision %
()
S
nS
ii
()
∑−
()
S
+++
9999
8
22
https://t.me/medicina_free
34
14
12
10
8
Excellent
6
4
2
0
0 1 2 3 4 5 6 7
Fig. 4.7 MEDx chart for the assessment of acceptability based on
quality specications
Not acceptable
Good
Marginal
slope obtained with the regression model. Finally, we
evaluate in which region of the graph the point falls
(Fig.4.7).
Limit ofBlank (LOB), Limit ofDetection (LOD),
Limit ofQuantication (LOQ)
Determining the lower limit at which the analytical method
under study is able to accurately identify and quantify the
analyte is extremely important in many laboratory settings,
including toxicology, drug monitoring, assessment of biochemical recurrence in cancer patients after surgery, and the
assay of certain analytes (e.g., TSH, PSA, troponin).
The most recent international literature suggests adopting
a common terminology. It is denoted by the following:
• Limit of Blank (LOB): the largest result that is likely
(with a given probability) to be observed in repeated measurements of a blank sample, that is, not containing the
analyte of interest.
• Limit of Detection or Sensitivity (LOD): the smallest
amount of analyte in a sample that can be detected, but
not quantied, with a given probability.
• Limit of Quantitation (LOQ): the smallest amount of analyte that can be quantied with a given precision.
Calculation of LOB and LOD
There are several protocols to determine these parameters,
based on the signal-to-noise ratio (used for example in
chromatography), on multiples of the standard deviation of
replicates of white samples (LOD = meanB + 3 × SDB;
LOQ=meanB+10×SDB), on the regression parameters of
the calibration line (LOD=3Sa /b; LOQ=10Sa /b; where
Sa is the standard deviation of the y residuals and b the
slope).
M. Vidali
The traditional approach for the determination of LOB
and LOD, suggested by the CLSI protocol, involves the following steps:
• Analysis of blank samples, that is, samples without the
analyte (samples that do not contain the drug or substance
of abuse, samples pre-treated with resins or precipitating
agents, samples from pathological subjects): K blank
samples are analyzed r times with N=K×r.
• Assessment of the presence of aberrants.
• Calculation of LOB=meanB+Cp×SB, where meanB and
SB are the mean and standard deviation of the replicates,
respectively. Cp is a multiplicative factor related to the
type I error at 5%, that is, the error that is made if a blank
sample result exceeds the LOB and the investigator mis-
takenly concludes that the sample contains the analyte. Cp
.
−
1
1 645
4
×−
1
NK
, where N is the
is calculated as:
=
p
total number of replicates and K is the number of blank
samples. Since the results of the blank samples are often
not distributed according to a Gaussian pattern, it is pref-
erable to estimate the LOB with a non-parametric method:
(1) the N results are ordered from the smallest to the larg-
est; (2) the rank=0.5 +N×0.95 is determined and the
LOB is equal to the value of the observation correspond-
ing to the calculated rank (if the rank is not integer, inter-
polation is performed).
Example: N=45; rank=0.5+N×0.95=43.25. Since it
isnon-integer, we proceed by interpolation with neighboring
observations 43 and 44. If N43 = 2.6 and N44 = 2.9 then
LOB=N43+0.25×(N44 − N43)=2.7.
• Analysis of samples containing low concentrations of
analyte (analyte concentration from1 to 5 times the esti-
mated LOB).
• Calculation of the standard deviation for the replicates of
the J samples: Si (standard deviation of the i-th sample)
and calculation of the pooled standard deviation:
pooled
J
∑−
i
==1
=
iJi
1
2
×
1
.
1
n
Example: Four blank samples are analyzed (10 replicates
× sample). The standard deviations of the 10 replicates of the
4 samples are as follows: S1 = 3.20; S2 = 2.97; S3 = 3.54;
S4 = 3.75. The pooled standard deviation will be
22
....
×+×+×+×
9320 9297 9354 9375
=
pooled
• Calculation of the LOD=LOB+Cp×S
, where Cp is
pooled
.
=
33
.
a multiplicative factor related tothe type II error at 5%,

()
x
i
x
i
%CV= CX
C
0
LO
C
x
i
x
i
T
WL
=+
()
S
T
()
45
70
CV%
Concentration (mg/dL)
4 The Role ofStatistics inLaboratory Medicine
https://t.me/medicina_free
35
that is, the error that is made if a result from a sample
containing the analyte is less than the LOB and the experimenter mistakenly concludes that the sample does not
contain the analyte. Cp is calculated as
p
4
×−
1
LJ
−
1
, where L is the total number of rep-
.
1 645
=
licates and J is the number of samples containing the analyte at low concentrations.
LOQ Calculation
Unlike LOB and LOD, wherethe calculation is based on statistical reasoning (in terms of type I and type II error), the
determination of LOQ depends on operator-dened quality
specications (e.g., maximum allowable error, maximum
imprecision). Two, among the many approaches suggested in
the literature, are LOQ as functional sensitivity (accuracy
prole) (Fig.4.8) or as total error. In the rst approach, the
following steps can be identied:
1. The quality specication is set as total or withinlaboratory repeatability (e.g., 10%CVWL).
2. The protocol for determining total repeatability (design
5days × 5 replicates or 20days × 2 replicates) for 6–9
samples of known concentration distributed along the
lower part of the concentration range) is applied.
3. The presence of any aberrants is checked.
4. For each sample (of the 6–9 tested), the mean concentration of all replicates is calculated
and, following the
protocol described above, the relative repeatability within
the laboratory (SWLand %CVWL).
5. The precision prole is drawn, plotting the average concentration
on the x-axis and the repeatability within
the laboratory (%CVWL) on the y-axis.
6. The best power model explaining the data is found:
1
, where C0 and C1 are the two parameters of
the model (Fig.4.8).
7. Having identied C0 and C1, wend the concentration of
the analyte (X) at which the repeatability within the
Laboratory %CV = 10% (the chosen quality target):
1
/
1
Y
X
Q ==
, where Y=10%.
C
0
The experiment for determining LOQ based on total
desirable error (TE) consists of these steps:
1. The quality specication, that is, the total desirable error,
is set, based on biological variability, expert consensus,
or other criteria.
2. Four to ve CRMs are analyzed at assigned low concentrations: a 3× 5 × (2 reagents) or 5 × 5 × (2 reagents)
design is used for each CRM.
40
35
30
25
20
15
10
5
0
0 10 20 30 40 50 60
Fig. 4.8 Precision prole
y = 371.67x
–0.975
3. The presence of any aberrants is checked.
4. For each sample/reagent pair, the mean of all replicates
we calculate (
as the difference between the mean (
), the total repeatability (SWL), and the bias
) and the assigned
value for that sample.
5. For each sample/reagent pair, the total error (TE) is calculated from the above data according to one of the following models:
Ebias Westgardmodel
EbiasRMSmodel
=+
2
22
S
WL
6. The TE of each pair/reagent is expressed as a percentage
of the CRM assigned value: %TE=(TEx100)/(assigned
value).
7. For each reagent, the sample with the lowest concentration that veries the quality objective (i.e., with
TE%<desirable total error) is selected; this concentration corresponds to the LOQ for that reagent,
8. The LOQ of the analytical method is set equal to the
greater of the LOQs of the reagents.
If all of the samples used for one or more reagents do not
meet the quality objective, the protocol for determining the
LOQ must be rerun, using CRMs at higher concentrations.
Linearity
The linearity of an analytical method is its ability to give
results that are directly proportional to the concentration of
the analyte studied. Among the various methods that can be
used for the determination or verication of linearity, the
most recent literature suggests the robust method of polynomial regression in which regression models of different
orders (rst, second, and third order) are compared. The protocol consists of the following steps:
1. The maximum desirable error for non-linearity and
repeatability is set.

36
F
S
S
S
2
S
2
https://t.me/medicina_free
M. Vidali
2. Sample selection: pool of patient samples (samples to be
analyzed are produced by mixing in various ratios two
patient samples, one with known concentration of analyte
at the upper limit of the range and one with known concentration near the LOQ) or pool of patients fortied with
the analyte of interest or purchased controls/calibrators;
the number of samples depends on the purpose of the linearity experiment (5–7 levels × 2 replicates in case of previously validated method, 9–11 levels × 2–4 replicates in
case of new method to be validated).
3. Analysis of samples.
4. The results are graphed (results on the y-axis, theoretical
concentrations on the x-axis) verifying the presence of
evident non-linearity (in this case the experimental conditions should be re-evaluated).
5. The presence of possible aberrants is evaluated (in presence of more aberrant values the experimental conditionsshould be re-evaluated).
6. Using statistical software, rst-, second-, and third-order
regression models are obtained:
First order: y=b0+b1x (linear model)
Second order: y=b0+b1x+b2x2 (non-linear model)
Third order: y = b0 + b1x + b2x2 + b3x3 (non-linear
model)
7. The possible signicance for the coefcients b2 b3 is evaluated as the ratio between the non-linear coefcient (b2 or
b3) and the relative standard error (provided by the soft-
ware): t=b/SE
. t is compared with a critical value that
slope
depends on the chosen probability level (e.g., alpha=0.05)
and the degrees of freedom given by df= C × R −B,
where C is the number of samples, R the number of replicates for each sample, and B the number of coefcients of
the regression model (rst order: 2, second order: 3, third
order: 4).
8. If neither b2 or b3 are signicant, one concludes for linearity; otherwise, one selects among the signicant nonlinear models the one with the smallest standard error of
regression (also provided by the statistical software). The
difference between the value predicted by the non-linear
model and the value predicted by the linear model is calculated for each level and expressed as a percentage of
the concentration level (%diff= (100 × diff)/concentration). If all %differences are smaller than the xed desirable error, the linear model is chosen anyway (the error,
although statistically signicant, is still within the desirable one). If, on the other hand, one or more of these differences% are greater than the xed desirable error, one
concludes for non-linearity and tries to identify any
experimental conditions underlying this non-linearity;
alternatively, if the major differences% are at the limit of
the range, it is possible to run the statistical analysis again
after removing the relevant concentration level (with
obvious reduction of the linear range).
9. After verifying linearity, it is necessary to check that
repeatability is also within the selected maximum error;
to this end, compare Sr (obtained as in the repeatability
experiment from the ANOVA table) or %CVr (if repeatability is not constant along the interval) with the maximum error. If Sr or %CVr is greater than the stated
maximum error, the results of the linearity experiment are
not reliable and the experimental conditions and the
method itself must be veried.
When the repeatability varies a lot along the concentration interval (heteroscedasticity), in particular for concentration intervals that are very large at several orders of
magnitude, it is preferable to use a weighted linear regression model rather than a simple linear regression model. A
quick check for this purpose is obtained in the following
way: (1) 10 replicates at the two extremes of the interval are
analyzed (in total 10×2=20 replicates); (2) the variances at
the two concentration levels are calculated; (3) an F-test for
2
the variance is performed:
calculated
=
a
, where
2
b
and
a
b
are, respectively, the greatest and the least variance; (4) the
F
is compared with the F
calculated
level=0.05; dfa=10–1=9; dfb=10–1=9; F
(5) if F
calculated
>F
, we conclude that the repeatability is
tabulated
(signicance
tabulated
tabulated
=3.18);
not constant along the concentration interval and it is preferable to use the weighted linear regression model.
Finally, it should be noted that the correlation coefcient
is not a measure of linearity but a measure of the degree of
association between two variables.
Biological Variability
Several factors, acting on the concentrations of the different
constituents of the body’s biological uids, can inuence the
analytical result. Among these sources of biological variability, age, sex, race, pregnancy, circadian, monthly or seasonal
rhythms, and the presence of underlying medical conditions
are of particular importance. Generally, these factors are
considered controllable and included in the pre-analytical
variability. The term “controllable” means that their effect
can be controlled by appropriate measures, for example, by
taking the sample under standardized conditions and comparing the subject with individuals who are comparable in as
many characteristics as possible. Then there is an intrinsic
variability of the organism, which cannot be controlled or
further reduced, and which is the expression of random uctuations around a homeostatic point. This random variability,
different for the various components of the organism, is
called intraindividual biological variability. The difference
between the homeostatic points of different individuals is

CV
TA
=+
22
22
TT
Criti
eC
V.
AI
()
×+
22
critical
eCVC
AI
()
22
()
×+
AI
22
4 The Role ofStatistics inLaboratory Medicine
https://t.me/medicina_free
37
instead called interindividual biological variability and, with
few exceptions, is generally greater than intraindividual variability. The knowledge of these sources of variability, measurable through the application of appropriate experimental
protocols, is essential both for the denition of analytical
goals and for the correct interpretation of a test result.
Different criteria can be used to interpret laboratory test
results, such as comparing the current result with previous
results from the same subject, or with a dened reference
interval for the healthy population, with decision limits or
following a probabilistic diagnostic approach.
Comparison ofCurrent Results withPrevious
Results fortheSame Subject: TheCritical
Dierence
If, using standardized procedures, pre-analytical variability
is minimized, the total variability associated with a laboratory result is
respectively, analytical variability and intraindividual biological variability, expressed as coefcients of variation
(standard deviation/mean×100%).
For two subsequent results obtained on the same individual and in the same laboratory, the total variability will be
given by the sum of the individual variabilities and therefore
equal to
×=×CV CV
ence, that is, the maximum difference attributable to the
effect of the total variability at a chosen probability level, can
be derived:
cal differenc
For a 95% probability level (Z=1.96) the critical difference, expressed as a percentage, becomes
differenc
Differences greater than the critical difference may
depend on variations in the subject’s medical condition.
Alternatively, from the above formula, it is possible to derive
the value of Z and thus the probability that a certain percentage difference between two results from the same individual
is signicant: Z=
The critical difference is simple to understand and easy to
calculate from the values of CVI available in the literature
and of CVA obtained in the laboratory with quality controls;
however, its use presents some criticalities: (1) the intraindividual variabilities generally used are determined on healthy
volunteers and are smaller than those presented by hospitalized or critical patients; (2) successive results obtained in a
short time in the same individual tend to self-correlate and
not to follow a Gaussian distribution; (3) different individuals present different intraindividual variabilities. In spite of
CV CV
2
%
=×
%.
=× +277
difference%
2
CV CV
, where CVA and CVI are,
I
from which the critical differ-
Z
VC
2
V
.
.
these limitations, the critical difference is an important interpretative aid for the clinician, particularly in some situations
(e.g., in serial analyses of tumor markers or in monitoring the
effects or toxicity of a drug).
Denition andUse ofReference Intervals
Frequently, the result of a laboratory test is interpreted by
comparison with a range of values determined in a reference
population. The term “reference population” is preferred to
the more ambiguous terms “normal population” or “healthy
population” and indicates instead a set of reference individuals selected on the basis of specic criteria. The traditional
approach to the denition of reference intervals foresees that
from this population of reference individuals, a reference
sample is extracted from which reference values are obtained,
through measurements, which will present a certain distribution (reference distribution) and whose analysis, with appropriate statistical methods, will allow to estimate the reference
limits and the relative reference interval. The most critical
aspects of this process are the selection of the individuals,
the standardization of the pre-analytical conditions, the analytical aspects, and the statistical processing. When the biological characteristics of the analyte are known, it is
preferable to use the “a priori” selection approach, excluding
and stratifying individuals before sampling according to specic criteria dened and well documented in the procedures
(age, sex, genetic factors, physiological factors, risk factors,
presence of pathologies). When, more rarely, the analyte is
new and there is little information regarding biological variability, it is necessary to follow the “a posteriori” approach
and therefore sampling a very large number of individuals
and then excluding and stratifying on the basis of the
measurements obtained. As an alternative to the previous
direct methods, indirect methods use analytical results contained in existing databases and are based on the assumption
that most of the data produced by a laboratory, even in hospitalized patients, are normal. From the mass of extracted
data, the use of appropriate statistical methods allows to
identify and exclude the abnormal values of the minority of
“unhealthy” subjects and thus dene the reference interval.
Indirect methods are relatively simple and inexpensive and,
in particular when it is difcult to collect samples of healthy
subjects, they are the only methods that can be applied; however, they have important limitations: they are difcult to
generalize, since the relevant pre-analytical, analytical, and
clinical data are not available, and they strongly depend on
the mathematical model used to dene the interval, since the
distribution of the values of the reference population is not
known. These methods can therefore be considered to estimate or at most verify reference intervals determined by
direct methods.
Соседние файлы в папке Библиотека им академика М.И. Перельмана
