Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_2617_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
30.08.2026
Размер:
59 Мб
Скачать
28
s
xx
ii
=−
Ü
M
()
bxx
x
Ü
https://t.me/medicina_free
M. Vidali
association of two linearly correlated variables. If they are linked by a strong non-linear relationship (e.g., quadratic), the correlation coefcient may be falsely low. This under­scores, as noted earlier, the importance of always viewing data before statistical analysis.
The goodness of t of the regression model to our data can be assessed through the coefcient of determination R2, which is the square of the correlation coefcient: R2 = r2. Since r varies from 1 to +1, R2 varies from 0 to 1 and indi­cates the proportion of variability of Y explained by the regression line.
Evaluation ofAnalytical Performance
Accuracy
The precision of a method can be dened as the degree of agreement between different independent measurements under dened experimental conditions. Precision can be expressed in terms of imprecision as standard deviation, vari­ance and relative standard deviation (RSD), or coefcient of variation (CV). Several types of precision are dened:
• Repeatability (or within-run precision): precision under
experimental conditions that minimize other sources of
variability (same laboratory, same method, same instru-
ment, same operator, same operating conditions, with
measurements made within a short time interval).
• Intermediate precision (formerly total precision, recently
within-laboratory precision): precision under well-
dened experimental conditions (same laboratory, same
instrument, same method, with measurements made over
an extended time interval, with different calibrators, oper-
ators, and/or lots of reagents).
• Reproducibility: accuracy obtained under different condi-
tions (laboratory, operator, instrument). Reproducibility
can only be assessed by inter-laboratory studies.
Verication of repeatability with the protocol 5 days×5 replicates.
Before using a method validated by a third party (e.g., a manufacturing company), the laboratory must perform a rapid verication protocol. To this aim, various experi­mental designs are reported in the literature, including the 3×5 scheme (3 replicates for 5days); however, the most recent literature suggests increasing the number of repli­cates and/or days in order to obtain a more realistic esti­mate of repeatability. It is therefore advisable to use a 5×5 scheme (5 replicates for 5days) or one in which the product of replicates × days is at least 20. For methods that are not validated by third parties but developed on their own, a more extensive protocol should be adopted
(the literature suggests a 2× 20 scheme, i.e., 2 replicates for 20days).
The repeatability verication protocol includes the fol-
lowing steps:
• Sample selection.
• Sample analysis (5 replicates × 5days) using the method to be veried.
• Processing of any aberrant data and possible integration with new measurements until the desired size is obtained.
• Calculation of repeatability.
• Verication of what has been declared by the producer.
The verication should be performed using at least two
controls, or pools of patients, with concentrations close to both any decision values (and/or cut-off), and to those used by the manufaturer in the validation phase and reported in the data sheet.
If one or more observations from the collected data devi-
ate to a certain extent from all the others (extreme or aberrant values), it is necessary to perform a statistical test for aber­rant data, after checking that these values are not the result of analytical or gross errors. A number of statistical tests useful in identifying extreme values are described in the literature, including the Dixon-Reed test, Grubbs’ test, Huber’s test, and Horn’s method. In the Dixon-Reed test, the ratio D/R is calculated, where D is the absolute difference between the aberrant and the least extreme value immediately preceding it and R is the maximum-minimum range of all the data; if the ratio is greater than 1/3, the observation is considered aberrant. In the Grubbs’ test, instead, we calculate the ratio
maxmean
=
or s=
mean min
, where min and max
are, respectively, the presumed minimum or maximum aber­rant value, mean and s are, respectively, the mean and stan­dard deviation of the sample. If the calculated G is greater than a certain critical value, which can be found in the litera­ture and depends on the sample size, the data is considered an aberrant. Huber’s test is based instead on the median absolute deviation (MAD), that is, the median of the absolute differences (i.e., without sign) between the observations and the median. The procedure is as follows: the median is calcu-
lated; then the absolute differences D
between the
median and the individual observations are calculatedand then the MAD, that is, the median of all these absolute differ­ences, multiplied by an appropriate constant b=1.4826(when data follow a gaussian distribution), is obrained. Thus,
AD median
i
, where
indicates the median of
the observations. Finally, we calculate the limit value as
D
=MAD×3 (very conservative; or 2.5 moderately con-
crit
D
R
=>
3
./
9
x
λ
λ
1
xx
;;
()
=+
V
between between between within within with
= =
iin
x
j
SV
rW
=
%/
()
SV
BB
=
%/
()
100
SVV
WB
=+
%/
()
100
4 The Role ofStatistics inLaboratory Medicine
https://t.me/medicina_free
29
servative). If an observation has an absolute difference from the median greater than this critical value (Di> D
), it is
crit
considered aberrant. Finally, in the Horn’s method, we rst apply the Box-Cox transformation, to obtain a more Gaussian distribution of the data, and then Tukey’s robust method, cal­culating the interval between Q1–1.5 × IQR and Q3+1.5×IQR, where Q1, Q3, and IQR are the rst quartile, the third quartile, and the interquartile range (Q3-Q1), respectively: all data outside this interval are considered aberrant.
Example: for the 25 replicates (5×5 design), the follow­ing data were obtained: mean=101.2; s=14.0; min=76.4; max=150.0. The graphical analysis seems to suggest that the value 150.0 is an aberrant. We test this hypothesis with the four tests illustrated.
(a) Dixon-Reed test:
Given less extreme than the max: 113.2 then
150 0 113 2
..
=
150 76 4
.
0501
(b) Grubbs’ test: for N=25 G
..
150 0 101 2
=
14 0
=
.
34
aberrant.
=3.14
critical
since G > G
.
critical
we con-
clude that 150.0 is aberrant.
(c) Horn’s method: after having found through statistical
software λ =0.8511372 and having transformed the
necessary, to transform the data. If a non-normal distribu­tion is assumed for the data, the multiplicative constant for Huber’s test changes from b= 1.4826 to b = 1/Q3, where Q3 is the third quartile. Graphical analysis of the data also allows to check for multiple aberrant data and thus choose a suitable test. In this regard, it should be noted that the Dixon-Reed test may not identify the aber­rant data in the presence of more than one aberrant. In addition, while Tukey’s test can identify multiple aber­rants, Grubbs’ test identies one aberrant at a time, and it is therefore necessary to apply the test to the data set sev­eral times after removing the suspected aberrant (a mini­mum data size of >6 is recommended for Grubbs’ test). Finally, it is advisable not to use Tukey’s method for small data sets.
After identifying any aberrant data, we proceed to calcu­late total or within-laboratory repeatability through analysis of variance (ANOVA), as shown in the introductory section on basic statistical methods:
jdi
11
r
ij j
2
x
d
DEVDEV
between within
DE
VVDEV DEV
totbetween within
DF DF DF DF DF
between within totbetween within
nx
jj
j
1
=− =−
ddr11;; ;
AR DEVDFVAR DEVDF
2
/; /
,
data according to Box-Cox, that is, with
the limits
of the interval Q1–1.5 × IQR and Q3 + 1.5 × IQR (1.147–1.156) are calculated. The suspected aberrant
150.0, after transformation, is equal to 1.158, which is outside the interval, conrming what was found with the previous tests.
(d) Huber’s method: the median of the observations is equal
to 100.8; the absolute differences between this median and all the observations are calculated and then the MAD, which results to be 9.17 (already corrected for the constant b). The critical value is D
=9.17×3=27.5.
crit
The observation 150.0, since it presents an absolute dif­ference from the median equal to 49.2 (since Di>D
crit
49.2>27.5), is considered aberrant.
The most critical issues of these tests are the assump­tion of normality of the data distribution and/or the failure to identify the aberrant data in the presence of more than one extreme data. In fact, the reliability of the identica­tion of an observation as aberrant depends on the distribu­tion of the data (in particular for the Grubbs’ test and the Dixon-Reed test). Before applying these tests, it is there­fore necessary to check the Gaussian distribution and, if
with d being the number of days, r the number of replicates,
xij the i-th replicate of the j-th day,
the average of the j-th
day, and x the average of all replicates.
From the above data, we calculate the following:
The variance of repeatability (within-run or within-run):
VW=VAR
within.
The repeatability (within-run or within-run):
CV
=
Sx100.
rr
The “pure” between-run repeatability variance (i.e., cor-
rected for the contribution of VW): VB=(VAR R
)/n0, where n0 represents the average number of rep-
within
;
licates of the runs (in the 5×5 design n0=5). If VAR VA R
, then VB=0.
within
between
The “pure” (between-run) repeatability:
CV
=
Sx
BB
Total repeatability or within the laboratory:
and
CV S
=
WL WL
.
WL
x
.
Example: we want to evaluate the repeatability of a method for the measurement of cholesterol (5×5 design). From the analysis of the data, we obtain:
and
 VA
between
e
30
SV
rW
= = =
SV
BB
= = =
SVV
WL WB
=+=+=
./
x =
./,
%.
.%
=
()
×=
%.
.%
=
()
×=
%.
.%
=
()
×=
Sx
pp
=
()
()
()
()
22
22
rb
S
b
2
SVVr
2
=+
SV
rw
0= =
SVVr
2
3=+ =+ =
xx tS
a
−< ×
r
62
08<× ×=
B xx=−
0
%
xx
x
100
x
x
0
100
%
CC
C
a
100
Befo
./
=+++
()
=
A
./
= +++
()
https://t.me/medicina_free
M. Vidali
DEV DEV
Vw = 7.50; since VAR
: 132.56; DF
between
: 150; DF
within
between
: 20; VAR
within
between
: 4; VAR
within
between
: 7.50
: 33.14
(33.14) > VAR
(7.50) then
within
VB=(33.14–7,50)/5=5,13
750274../
513226../
Since the average of all 25 replicates is
mg dl
mg dl
750513 355..
mg dl
203 76
mg dl
then:
CV
CV
CV
r
B
WL
/.
274 203 76 100 134
/.
226 203 76 100 111
/.
355 203 76 100 174
The calculated total repeatability can be compared with that stated by the manufacturer. The manufacturer can pro­vide the total repeatability either as a standard deviation (Sp) or as %CVp. In the latter case, to obtain the Sp, it is sufcient to apply the formula
CV
%/100 .
In the cholesterol example above, the manufacturer pro­vided an Sp = 3.20. If he had provided, for example, a %CVp=1.6%, then Sp=1.6%/100× 203.76 = 3.26 would have been obtained.
It is then calculated: T (effective degrees of freedom):
rSrS
=
1
r
d
1
⋅+⋅
2
2
+
S
()
r
  
2
 
rS
b
1
d
, where
2
  
is the “non-pure”
the documented repeatability at the introduction of the method.
Example: Sr=2.74mg/dl (see previous example); degrees of freedom (df) 20. Sr had been determined in a 5day × 5 replicate design, with df = days × (repli­cates1)=5×(5–1)=20; if repeatability had been deter­mined in a single run of N replicates, the df would have been (N1)); t
(a/2;df= 20)
x2=6, we get:
=2.086. If x1=202 and x2=196, x1
2 086 2748
... and thus it can be
assumed that the laboratory is operating under the estab­lished repeatability conditions.
Trueness andBias
Trueness can be dened as the closeness of agreement between the mean value, obtained from a large number of determinations of a sample, and the true value. Note that the true value is dened as an accepted or assigned reference value. Trueness is expressed as bias, or systematic error, which can be calculated as the difference between the mean value of the replicates and the true value or as the ratio of this difference to the true value. Thus,
B
0
=
, where x represents the average of the val-
0
ues obtained by repeatedly analyzing the sample and x0 the assigned value. Bias is often reported as recovery and
expressed as %R
=
, where the relation
or
between-run variance, that is, not corrected for the contribu­tion of the within-run variance:
In our example: r = 5; d = 5;
513755 66
bBW
/. ./ . ; T = 12.475. Using
the calculated parameter T and the statistical table of the chi- square distribution, we nd the critical value C=23.337. Finally, we calculate a limit value (or verica­tion value)
/ , where Sp represents the total
P
repeatability declared by the manufacturer. In our example,
Sp=3.20; V =4.38. Since the total repeatability obtained
experimentally (3.55) is lower than the limit value (4.38), the total repeatability declared by the manufacturer is veried.
If substantial changes are made to the method in use, it is necessary to revalidate the method; however, even in the absence of such changes, it is good laboratory practice to periodically verify the repeatability demonstrated at the time the method was introduced. This verication can be easily conducted through the so-called duplicate test experiment: two independent replicates of the same sample are analyzed under repeatability conditions (x1 and x2). For the verication to be positive, it must be
bBW
12 2
2
/ .
2
, where Sr is
/;df
75
. ;
%R=100+%B. Recovery is frequently evaluated by mak­ing repeated determinations of a sample before and after the addition of a known amount of analyte; we have that
21
R
=
, where C1 and C2 are, respectively, the
mean value of the replicates of the sample before and after addition of the analyte, while Ca is the expected increase in the concentration of the analyte in the sample after addition. Ca can be obtained from Ca=Cst(Vst/Vf), where Cst is the concentration of the standard used for addition, Vst is the added volume of the standard solution, and Vf is the nal volume of the sample (ini­tial volume+added volume of the standard solution).
Example: 0.2 ml (Vst) of a standard calcium solution
21.0mg/ml (Cst) is added to a serum sample of 2.0ml vol­ume. Four determinations are carried out before and after the addition.
re additionC mg dl:..../
fter addition C
93 96 95 92 494
1
: ...
11 2110 10 9113 4
2
mg dl;
./
11 1
=
dl
()
=
%
.%
CC
C
a
sN
x tsN
aN±−
B tsN
aN±−
B =−=−
..
dl
%./.
B =−
()
×=
.
%%
R
RB
=
=+=−=
...
9
236
1
..±× =−
05
..±× =−
x
N
df
()
N
2
CRM
±×
N
5
with
df
()
113
236
8
4 The Role ofStatistics inLaboratory Medicine
https://t.me/medicina_free
31
.
mg
=
100 89 0
21 02 22 191./ .. /
R
a
21
=
100
..
11 194
=
191
The recovery experiment allows to estimate the propor­tional systematic error, that is, a type of error that increases as the concentration of the analyte increases, but not the con­stant systematic error; in fact, even in the presence of a recovery close to 100%, it is not possible to exclude the pres­ence of bias due, for example, to interferents or matrix components.
Bias can also be assessed: (1) by measuring certied ref­erence materials; (2) through a method comparison study.
Bias Assessment with a Certied Reference Material
A certied reference material (CRM), with an assigned con­centration value (and associated measurement uncertainty) and with a composition as close as possible to that of the test samples, is analyzed N times. It is obtained that
x
known value
t
calc
mean and the standard deviation of the replicates. If t greater than a tabulated t
=
/
, where x and s are, respectively, the
is
calc
, which depends on N and the
critical
chosen signicance level, then the bias is considered statisti­cally signicant. The condence intervals of the mean value of the N replicates and the bias are, respectively,
/;/21
and
/;/21
.
Example: The assigned value for a CRM for cholesterol is 250mg/dL.Ten measurements are made obtaining a mean of
236.8mg/dl and a standard deviation of 10.2mg/dl. You can then calculate:
236 8 250 13 2
236 8 250 250 100 53
236 8 250 100 94 7
./ .
%% %
100 100 53 94 7
236 8 250
.
calc
=
10 210
./
=
40
.
/mg
%
or
For (10–1)=9 degrees of freedom and a 95% signi­cance level, we have that t
 = 2.262; since t
a/2;9
calc
 > t
crit
(4.09>2.262) we have that the bias is statistically signi­cant, that is, the mean measured value is different from the known or assigned value (with a 95% condence level).
To conrm this, the 95% condence interval (95%CI) of the mean value of the N determinations is
82262 10 210 229 5 244
.. ./
, which does not include the assigned value of 250mg/dl. Similarly, the 95% condence interval of the bias does not include zero (in fact:
32 2 262 10 210592
.. ./
).
A bias, although statistically signicant, could still be acceptable if it is less than an error established on the basis of different criteria (expert consensus, biological variability).
When assessing bias by repeated measurements of a certi­ed reference material, the investigator should consider both the uncertainty of the method and that of the CRM itself. In
such a case, it follows that:
calc
degrees of freedom given by
where rial. Note that, with
represents the uncertainty of the reference mate-
CRM
2
very small, the degrees of freedom
=
knownvalue
2
s
2
u
+
CRM
2
s
u
+
N
=
2
s
2 CRM
 
2
with the
2
 
N
1 ,
approximate to (N1). The condence interval, including the contribution of the CRM uncertainty, becomes
2
s
2
+t
u
critical CRM
.
In the previous example, the CRM has an uncertainty equal to u
236 8 250
=
calc
10 2
.
10
.
10 2
10
=
10 2
Alsoin this situation, with t
 = 1.4 mg/dl; it is then that
CRM
.
2
10
2
+
14
+
2
.
14
.
2
 
.
2
37
=
.
2
2
 
10
×−
=
calc
and t
>t
=2.160.
crit
(3.75>2.160), we
crit
have that the bias is statistically signicant. The 95% con­dence intervals for the mean value and the bias will be,
2
10 2
respectively,
and
32 2 160
..
82160
..
2
10 2
.
14 56 20
10
.
10
2
...±× +=− . As can be seen,
2
...±× += −
14 229 2 244 4
the condence intervals, calculated including the contribu­tion of the CRM uncertainty, are only slightly wider than the previous ones but the conclusion remains the same.
The Experiment of Comparison of Two Analytical Methods
Systematic error can also be assessed through a method com­parison experiment, where patient samples are measured with two different methods or instruments. This approach allows both calculation of the error at decision levels and
32
±+
35
Average of the 2 methods
Method 1 - Method 2
35
35
https://t.me/medicina_free
assessment of proportional (slope of the regression line) and constant (intercept of the regression line) systematic error. The comparison experiment between analytical methods consists of these steps:
1. Familiarization with the new method and verication of the analytical performance of the method under test.
2. Denition of acceptability levels for maximum allowable error:
30
25
20
15
Method 2
10
5
M. Vidali
• Based on the combined imprecision of the methods: if
the two methods do not differ signicantly, their dif­ferences will be within the range:
196
.CVCV
2
method method
2
1
, where CV represents the
2
total repeatability of the method.
• Based on quality specications reported in the litera-
ture (e.g., TE
I
are the maximum imprecision and inaccuracy,
max
max
=B
+1.65 × I
max
, where B
max
respectively).
3. Sample selection and analysis: 40–60 patient samples distributed over the entire reportable range (range in which analyte concentrations can be measured with acceptable accuracy and precision); a greater proportion of samples should be selected around any clinical deci­sion values and/or cut-offs.
4. Plotting and statistical analysis of the data: the most used tools are the regression model and the Bland-Altman dif­ference graph (Fig.4.5).
To conduct the regression analysis, the data obtained with the reference or in-use method (x-axis) are plotted against the data obtained with the method under test (y-axis) together with the identity line (y = x) and the regression line. Different regression models can be used to nd the regression line. The widely used simple (or ordinary) linear regression model has the merit of compu­tational simplicity. However, it is not always correct to use it. In fact, simple linear regression assumes that the method in use (the one on the x-axis) is error-free and that the error of the method under test (y-axis) is normally distributed and is constant throughout the range of con­centrations studied (homoscedasticity). When the assumptions underlying the linear regression model are not met, it is necessary to use another regression model (Deming, weighted Deming, or the non-parametric Passing-Bablok model). The non-parametric Passing­Bablok model requires the least number of assumptions. In fact, in Passing-Bablok regression, extreme values (outliers) can be included, imprecision is allowed in both methods, and it is not necessary either that the error be normally distributed or that it be constant over the con­centration range. Regression analysis allows the assess­ment of the systematic constant (intercept) and
max
0
0
5 10 15 20 25 30
Method 1
1.5
1.0
and
Fig. 4.5 Passing-Bablok regression analysis and Bland-Altman graph for the evaluation of systematic error in the comparison experiment between methods. In the Bland-Altman graph the line corresponding to the bias and the two lines delimiting the interval which includes 95% of the differences between the methodsare shown
0.5
0.0
–0.5
–1.0
–1.5
–2.0
0 5 10 15 20 25 30
proportional (slope) error. It is always necessary to report not only these two parameters, point estimates of the true intercept and the true slope, but also their 95% condence intervals (interval estimation), that is, the intervals in which we have some condence (usually 95%) that the true intercept and slope can be found. In particular, to demonstrate the absence of constant, proportional sys­tematic error, the relevant condence intervals must include 0 (zero) for the intercept and 1 (one) for the slope. Small proportional or constant systematic errors identi­ed by regression analysis, while statistically signicant, do not preclude the use of a particular method if the error is less than the established maximum acceptable error. It is also important to note that the correlation coefcient should not be used to evaluate the agreement between two methods; in fact, it allows us to evaluate only the associa­tion between two variables but not the agreement (Fig.4.5).
The Bland-Altman plot, on the other hand, allows cal­culation of the bias between the two methods, expressed as the mean of the differences. In this graph, the analyte
+1.96 DS
1.18
bias
–0.01
–1.96 DS
–1.21
%b
−×+
()
CCba
C
15
090
DIfference between
Average of method 1 and method 2
4 The Role ofStatistics inLaboratory Medicine
https://t.me/medicina_free
concentrations on the x-axis (expressed as the mean of the two methods) are graphed toward the differences between the two methods on the y-axis. Alternatively, the mean of the two methods (x-axis) can be graphed toward the dif­ferences of the two methods expressed as a percentage of the mean, that is, (method1 method2)/mean × 100 (y-axis). This method is especially useful when the con­centration range is wide.
The bias is then calculated as the mean of all differ­ences and expressed along with the associated 95% con­dence interval. Signicant bias occurs when the value 0 (zero) is not within the relevant 95% condence interval. In the absence of bias (bias=0), the points corresponding to the differences between the two methods should accu­mulate randomly around the zero line, that is, positive and negative differences close to zero should be observed for each concentration level. It is also possible to observe the presence of particular trends in the differences and/or the widening of the agreement range with increasing con­centrations, a sign of the dependence of the differences on the concentration level. In addition to the mean differ­ence (and the relative 95% condence interval), it is also necessary to calculate the interval including 95% of such differences dened by mean±1.96 × SD, where mean and SD are, respectively, the mean (bias) and the standard deviation of the differences found.
On the other hand, this interval, which is often over­looked and not reported, is of considerable importance in the correct interpretation of the comparison experiment, because it allows us to identify the magnitude of most (95%) of the differences found. Alternatively, the interval comprising 95% of the differences can also be calculated using non-parametric methods, for example, as the inter­val between the percentiles 2.5 and 97.5 (Fig.4.5).
5. Evaluation of the acceptability of the new method: the bias identied is compared with the maximum acceptable error dened at the beginning of the experiment.
For acceptability based on the combined imprecision of
the two methods, we proceed as follows:
1. For each concentration level, the range dened by 0±1.96×CV%×C shall be calculated, where CV% is the combined inaccuracy of the two methods and C is the concentration. Since repeatability frequently varies over the concentration range, different CV% must be used for the two methods over the concentration range and differ­ent combined inaccuracies must be calculated.
2. Count how many method1-method2 differences fall out­side the range. If the two methods are identical within the combined imprecision, their differences should be sym­metrically distributed around zero and 95% of them should be within the interval.
33
10
5
0
–5
method 1 and method 2
–10
–15
0102030405060708
Fig. 4.6 Acceptability assessment based on combined imprecision
For a better visualization it is possible, in the same differ­ence graph, to add the two lines representing the range 0±1.96×CV%×C (Fig.4.6).
To assess acceptability based on predened quality speci­cations, the previously dened maximum total allowable error is identied and a diagram (MEDx chart) is constructed, plotting the imprecision (x-axis) against the systematic devi­ation (bias) (y-axis) and representing the performance of the test under evaluation as a point. The coordinates of the point are obtained using for the abscissa the imprecision (calcu­lated as repeatability within the laboratory) and for the ordi­nate the bias (using the equation of the regression line applied at a certain chosen concentration level, generally at clinical decision levels or cut-offs). The chart also shows four zones corresponding to different quality specications (not accept­able, marginal, good, excellent). For the construction of the MEDx chart, we proceed as follows:
1. The point corresponding to the maximum allowableerror
is xed on the y-axis.
2. On the abscissa axis, we x three points corresponding to
the various quality criteria: (max/2 error), (max/3 error), (max/4 error).
3. Three lines are drawn, starting from the point on the ordi-
nate axis to the three points on the abscissa axis, which divide the graph space into four zones corresponding to the different quality specications;
4. A point is represented using the performance of the
method under test as coordinates (abscissa = within­laboratory repeatability; ordinate = bias as assessed by regression analysis). If y=a+bx is the regression model, at a given concentration the bias will be given by
ias =
, where C is the expected
×
100
concentration, a and b, respectively, the intercept and
16
8
Bias %
Imprecision %
()
S
nS
ii
()
∑−
()
S
+++
9999
8
22
https://t.me/medicina_free
34
14
12
10
8
Excellent
6
4
2
0
0 1 2 3 4 5 6 7
Fig. 4.7 MEDx chart for the assessment of acceptability based on quality specications
Not acceptable
Good
Marginal
slope obtained with the regression model. Finally, we evaluate in which region of the graph the point falls (Fig.4.7).
Limit ofBlank (LOB), Limit ofDetection (LOD), Limit ofQuantication (LOQ)
Determining the lower limit at which the analytical method under study is able to accurately identify and quantify the analyte is extremely important in many laboratory settings, including toxicology, drug monitoring, assessment of bio­chemical recurrence in cancer patients after surgery, and the assay of certain analytes (e.g., TSH, PSA, troponin).
The most recent international literature suggests adopting
a common terminology. It is denoted by the following:
• Limit of Blank (LOB): the largest result that is likely (with a given probability) to be observed in repeated mea­surements of a blank sample, that is, not containing the analyte of interest.
• Limit of Detection or Sensitivity (LOD): the smallest amount of analyte in a sample that can be detected, but not quantied, with a given probability.
• Limit of Quantitation (LOQ): the smallest amount of ana­lyte that can be quantied with a given precision.
Calculation of LOB and LOD
There are several protocols to determine these parameters, based on the signal-to-noise ratio (used for example in chromatography), on multiples of the standard deviation of replicates of white samples (LOD = meanB + 3 × SDB; LOQ=meanB+10×SDB), on the regression parameters of the calibration line (LOD=3Sa /b; LOQ=10Sa /b; where
Sa is the standard deviation of the y residuals and b the
slope).
M. Vidali
The traditional approach for the determination of LOB and LOD, suggested by the CLSI protocol, involves the fol­lowing steps:
• Analysis of blank samples, that is, samples without the
analyte (samples that do not contain the drug or substance
of abuse, samples pre-treated with resins or precipitating
agents, samples from pathological subjects): K blank
samples are analyzed r times with N=K×r.
• Assessment of the presence of aberrants.
• Calculation of LOB=meanB+Cp×SB, where meanB and
SB are the mean and standard deviation of the replicates,
respectively. Cp is a multiplicative factor related to the
type I error at 5%, that is, the error that is made if a blank
sample result exceeds the LOB and the investigator mis-
takenly concludes that the sample contains the analyte. Cp
.
1
1 645
 
4
×−
1
NK
, where N is the
 
is calculated as:
=
p
total number of replicates and K is the number of blank
samples. Since the results of the blank samples are often
not distributed according to a Gaussian pattern, it is pref-
erable to estimate the LOB with a non-parametric method:
(1) the N results are ordered from the smallest to the larg-
est; (2) the rank=0.5 +N×0.95 is determined and the
LOB is equal to the value of the observation correspond-
ing to the calculated rank (if the rank is not integer, inter-
polation is performed).
Example: N=45; rank=0.5+N×0.95=43.25. Since it isnon-integer, we proceed by interpolation with neighboring observations 43 and 44. If N43 = 2.6 and N44 = 2.9 then LOB=N43+0.25×(N44 N43)=2.7.
• Analysis of samples containing low concentrations of
analyte (analyte concentration from1 to 5 times the esti-
mated LOB).
• Calculation of the standard deviation for the replicates of
the J samples: Si (standard deviation of the i-th sample)
and calculation of the pooled standard deviation:
pooled
J
∑−
i
==1
=
iJi
1
2
×
1
.
1
n
Example: Four blank samples are analyzed (10 replicates × sample). The standard deviations of the 10 replicates of the 4 samples are as follows: S1 = 3.20; S2 = 2.97; S3 = 3.54;
S4 = 3.75. The pooled standard deviation will be
22
....
×+×+×+×
9320 9297 9354 9375
=
pooled
• Calculation of the LOD=LOB+Cp×S
, where Cp is
pooled
.
=
33
.
a multiplicative factor related tothe type II error at 5%,
()
x
i
x
i
%CV= CX
C
0
LO
C
x
i
x
i
T
WL
=+
()
S
T
()
45
70
CV%
Concentration (mg/dL)
4 The Role ofStatistics inLaboratory Medicine
https://t.me/medicina_free
35
that is, the error that is made if a result from a sample containing the analyte is less than the LOB and the exper­imenter mistakenly concludes that the sample does not contain the analyte. Cp is calculated as
p
 
4
×−
1
LJ
1
, where L is the total number of rep-
 
.
1 645
=
licates and J is the number of samples containing the ana­lyte at low concentrations.
LOQ Calculation
Unlike LOB and LOD, wherethe calculation is based on sta­tistical reasoning (in terms of type I and type II error), the determination of LOQ depends on operator-dened quality specications (e.g., maximum allowable error, maximum imprecision). Two, among the many approaches suggested in the literature, are LOQ as functional sensitivity (accuracy prole) (Fig.4.8) or as total error. In the rst approach, the following steps can be identied:
1. The quality specication is set as total or within­laboratory repeatability (e.g., 10%CVWL).
2. The protocol for determining total repeatability (design 5days × 5 replicates or 20days × 2 replicates) for 6–9 samples of known concentration distributed along the lower part of the concentration range) is applied.
3. The presence of any aberrants is checked.
4. For each sample (of the 6–9 tested), the mean concentra­tion of all replicates is calculated
and, following the protocol described above, the relative repeatability within the laboratory (SWLand %CVWL).
5. The precision prole is drawn, plotting the average con­centration
on the x-axis and the repeatability within
the laboratory (%CVWL) on the y-axis.
6. The best power model explaining the data is found:
1
, where C0 and C1 are the two parameters of
the model (Fig.4.8).
7. Having identied C0 and C1, wend the concentration of the analyte (X) at which the repeatability within the Laboratory %CV = 10% (the chosen quality target):
1
/
1
Y
X
Q ==
, where Y=10%.
C
0
The experiment for determining LOQ based on total
desirable error (TE) consists of these steps:
1. The quality specication, that is, the total desirable error, is set, based on biological variability, expert consensus, or other criteria.
2. Four to ve CRMs are analyzed at assigned low concen­trations: a 3× 5 × (2 reagents) or 5 × 5 × (2 reagents) design is used for each CRM.
40 35 30 25 20 15 10
5 0
0 10 20 30 40 50 60
Fig. 4.8 Precision prole
y = 371.67x
–0.975
3. The presence of any aberrants is checked.
4. For each sample/reagent pair, the mean of all replicates we calculate ( as the difference between the mean (
), the total repeatability (SWL), and the bias
) and the assigned
value for that sample.
5. For each sample/reagent pair, the total error (TE) is calcu­lated from the above data according to one of the follow­ing models:
Ebias Westgardmodel
EbiasRMSmodel
=+
2
22
S
WL
6. The TE of each pair/reagent is expressed as a percentage of the CRM assigned value: %TE=(TEx100)/(assigned value).
7. For each reagent, the sample with the lowest concentra­tion that veries the quality objective (i.e., with TE%<desirable total error) is selected; this concentra­tion corresponds to the LOQ for that reagent,
8. The LOQ of the analytical method is set equal to the greater of the LOQs of the reagents.
If all of the samples used for one or more reagents do not meet the quality objective, the protocol for determining the LOQ must be rerun, using CRMs at higher concentrations.
Linearity
The linearity of an analytical method is its ability to give results that are directly proportional to the concentration of the analyte studied. Among the various methods that can be used for the determination or verication of linearity, the most recent literature suggests the robust method of polyno­mial regression in which regression models of different orders (rst, second, and third order) are compared. The pro­tocol consists of the following steps:
1. The maximum desirable error for non-linearity and
repeatability is set.
36
F
S
S
S
2
S
2
https://t.me/medicina_free
M. Vidali
2. Sample selection: pool of patient samples (samples to be analyzed are produced by mixing in various ratios two patient samples, one with known concentration of analyte at the upper limit of the range and one with known con­centration near the LOQ) or pool of patients fortied with the analyte of interest or purchased controls/calibrators; the number of samples depends on the purpose of the lin­earity experiment (5–7 levels × 2 replicates in case of pre­viously validated method, 9–11 levels × 2–4 replicates in case of new method to be validated).
3. Analysis of samples.
4. The results are graphed (results on the y-axis, theoretical concentrations on the x-axis) verifying the presence of evident non-linearity (in this case the experimental condi­tions should be re-evaluated).
5. The presence of possible aberrants is evaluated (in pres­ence of more aberrant values the experimental condi­tionsshould be re-evaluated).
6. Using statistical software, rst-, second-, and third-order regression models are obtained:
First order: y=b0+b1x (linear model) Second order: y=b0+b1x+b2x2 (non-linear model) Third order: y = b0 + b1x + b2x2 + b3x3 (non-linear
model)
7. The possible signicance for the coefcients b2 b3 is eval­uated as the ratio between the non-linear coefcient (b2 or
b3) and the relative standard error (provided by the soft-
ware): t=b/SE
. t is compared with a critical value that
slope
depends on the chosen probability level (e.g., alpha=0.05) and the degrees of freedom given by df= C × R B, where C is the number of samples, R the number of repli­cates for each sample, and B the number of coefcients of the regression model (rst order: 2, second order: 3, third order: 4).
8. If neither b2 or b3 are signicant, one concludes for linear­ity; otherwise, one selects among the signicant non­linear models the one with the smallest standard error of regression (also provided by the statistical software). The difference between the value predicted by the non-linear model and the value predicted by the linear model is cal­culated for each level and expressed as a percentage of the concentration level (%diff= (100 × diff)/concentra­tion). If all %differences are smaller than the xed desir­able error, the linear model is chosen anyway (the error, although statistically signicant, is still within the desir­able one). If, on the other hand, one or more of these dif­ferences% are greater than the xed desirable error, one concludes for non-linearity and tries to identify any experimental conditions underlying this non-linearity; alternatively, if the major differences% are at the limit of the range, it is possible to run the statistical analysis again after removing the relevant concentration level (with obvious reduction of the linear range).
9. After verifying linearity, it is necessary to check that repeatability is also within the selected maximum error; to this end, compare Sr (obtained as in the repeatability experiment from the ANOVA table) or %CVr (if repeat­ability is not constant along the interval) with the maxi­mum error. If Sr or %CVr is greater than the stated maximum error, the results of the linearity experiment are not reliable and the experimental conditions and the method itself must be veried.
When the repeatability varies a lot along the concentra­tion interval (heteroscedasticity), in particular for concentra­tion intervals that are very large at several orders of magnitude, it is preferable to use a weighted linear regres­sion model rather than a simple linear regression model. A quick check for this purpose is obtained in the following way: (1) 10 replicates at the two extremes of the interval are analyzed (in total 10×2=20 replicates); (2) the variances at the two concentration levels are calculated; (3) an F-test for
2
the variance is performed:
calculated
=
a
, where
2
b
and
a
b
are, respectively, the greatest and the least variance; (4) the F
is compared with the F
calculated
level=0.05; dfa=10–1=9; dfb=10–1=9; F (5) if F
calculated
>F
, we conclude that the repeatability is
tabulated
(signicance
tabulated
tabulated
=3.18);
not constant along the concentration interval and it is prefer­able to use the weighted linear regression model.
Finally, it should be noted that the correlation coefcient is not a measure of linearity but a measure of the degree of association between two variables.
Biological Variability
Several factors, acting on the concentrations of the different constituents of the body’s biological uids, can inuence the analytical result. Among these sources of biological variabil­ity, age, sex, race, pregnancy, circadian, monthly or seasonal rhythms, and the presence of underlying medical conditions are of particular importance. Generally, these factors are considered controllable and included in the pre-analytical variability. The term “controllable” means that their effect can be controlled by appropriate measures, for example, by taking the sample under standardized conditions and com­paring the subject with individuals who are comparable in as many characteristics as possible. Then there is an intrinsic variability of the organism, which cannot be controlled or further reduced, and which is the expression of random uc­tuations around a homeostatic point. This random variability, different for the various components of the organism, is called intraindividual biological variability. The difference between the homeostatic points of different individuals is
CV
TA
=+
22
22
TT
Criti
eC
V.
AI
()
×+
22
critical
eCVC
AI
()
22
()
×+
AI
22
4 The Role ofStatistics inLaboratory Medicine
https://t.me/medicina_free
37
instead called interindividual biological variability and, with few exceptions, is generally greater than intraindividual vari­ability. The knowledge of these sources of variability, mea­surable through the application of appropriate experimental protocols, is essential both for the denition of analytical goals and for the correct interpretation of a test result.
Different criteria can be used to interpret laboratory test results, such as comparing the current result with previous results from the same subject, or with a dened reference interval for the healthy population, with decision limits or following a probabilistic diagnostic approach.
Comparison ofCurrent Results withPrevious Results fortheSame Subject: TheCritical Dierence
If, using standardized procedures, pre-analytical variability is minimized, the total variability associated with a labora­tory result is respectively, analytical variability and intraindividual bio­logical variability, expressed as coefcients of variation (standard deviation/mean×100%).
For two subsequent results obtained on the same individ­ual and in the same laboratory, the total variability will be given by the sum of the individual variabilities and therefore equal to
×=×CV CV
ence, that is, the maximum difference attributable to the effect of the total variability at a chosen probability level, can be derived:
cal differenc
For a 95% probability level (Z=1.96) the critical differ­ence, expressed as a percentage, becomes
differenc
Differences greater than the critical difference may depend on variations in the subject’s medical condition. Alternatively, from the above formula, it is possible to derive the value of Z and thus the probability that a certain percent­age difference between two results from the same individual
is signicant: Z=
The critical difference is simple to understand and easy to calculate from the values of CVI available in the literature and of CVA obtained in the laboratory with quality controls; however, its use presents some criticalities: (1) the intraindi­vidual variabilities generally used are determined on healthy volunteers and are smaller than those presented by hospital­ized or critical patients; (2) successive results obtained in a short time in the same individual tend to self-correlate and not to follow a Gaussian distribution; (3) different individu­als present different intraindividual variabilities. In spite of
CV CV
2
%
%.
=× +277
difference%
2
CV CV
, where CVA and CVI are,
I
from which the critical differ-
Z
VC
2
V
.
.
these limitations, the critical difference is an important inter­pretative aid for the clinician, particularly in some situations (e.g., in serial analyses of tumor markers or in monitoring the effects or toxicity of a drug).
Denition andUse ofReference Intervals
Frequently, the result of a laboratory test is interpreted by comparison with a range of values determined in a reference population. The term “reference population” is preferred to the more ambiguous terms “normal population” or “healthy population” and indicates instead a set of reference individu­als selected on the basis of specic criteria. The traditional approach to the denition of reference intervals foresees that from this population of reference individuals, a reference sample is extracted from which reference values are obtained, through measurements, which will present a certain distribu­tion (reference distribution) and whose analysis, with appro­priate statistical methods, will allow to estimate the reference limits and the relative reference interval. The most critical aspects of this process are the selection of the individuals, the standardization of the pre-analytical conditions, the ana­lytical aspects, and the statistical processing. When the bio­logical characteristics of the analyte are known, it is preferable to use the “a priori” selection approach, excluding and stratifying individuals before sampling according to spe­cic criteria dened and well documented in the procedures (age, sex, genetic factors, physiological factors, risk factors, presence of pathologies). When, more rarely, the analyte is new and there is little information regarding biological vari­ability, it is necessary to follow the “a posteriori” approach and therefore sampling a very large number of individuals and then excluding and stratifying on the basis of the measurements obtained. As an alternative to the previous direct methods, indirect methods use analytical results con­tained in existing databases and are based on the assumption that most of the data produced by a laboratory, even in hos­pitalized patients, are normal. From the mass of extracted data, the use of appropriate statistical methods allows to identify and exclude the abnormal values of the minority of “unhealthy” subjects and thus dene the reference interval. Indirect methods are relatively simple and inexpensive and, in particular when it is difcult to collect samples of healthy subjects, they are the only methods that can be applied; how­ever, they have important limitations: they are difcult to generalize, since the relevant pre-analytical, analytical, and clinical data are not available, and they strongly depend on the mathematical model used to dene the interval, since the distribution of the values of the reference population is not known. These methods can therefore be considered to esti­mate or at most verify reference intervals determined by direct methods.