Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_2617_Библиотеки_им_академика_М_И_Перельмана
.pdf
18
https://t.me/medicina_free
G. Lippi et al.
of nancial processes and costs, and “intelligent systems”
equipped with logical activities that include calculations and
comparisons (e.g., with previous data of the same patient),
up to “expert systems” for the validation of analytical results
and the management of appropriateness.
Organizational Models oftheLaboratory
Network
The organization of laboratories in a homogeneous geographical area (corporate, district, provincial, regional, and
even national) has undergone considerable changes over the
years. Simplifying, we have passed from a model (typical in
the 70s) of extreme parceling out of laboratory diagnostics
(facilities operating below a critical performance mass) to a
model of progressive integration and consolidation, which
has led to the unication or closure of numerous territorial
realities. In Italy, this change was strongly supported by law
number 296 of December 27, 2006, which states that “the
Regions shall approve a reorganization plan of the network
of public and private accredited facilities providing specialist
services and laboratory diagnostics, in order to adjust the
organizational and personnel standards consistent with the
processes of increased efciency made possible by the use of
automated methods.” The general principles that guide the
reorganization process must be:
• Close interrelation between the type of hospital and type
of hospital laboratory
• Hospital territory continuity
• Proximity to the needs of the patient and the clinician
• The role of continuing education and research
• The role of information systems and (health) technology
assessment
• The role of external quality assurance
• Centrality of the promotion and control of
appropriateness
• The role of a targeted reporting system on laboratory
activities
ight to London from Washington, Miami, Baltimore,
Indianapolis, and Atlanta, the daily number of ights from
the Atlanta hub was increased (thus increasing the offer of
departure times), eliminating those from spoke airports and
transporting passengers by connecting ights from peripheral airports to that of Atlanta.
In general, the hub-and-spoke model has also progressively spread to the organization of health-care structures
and has created an impact in terms of services provided. In
this case, borrowing the model of aviation, there has been,
within homogeneous geographical areas, a progressive reorganization of laboratory facilities, also divided into spoke
(peripheral laboratories) and hub (central laboratories). The
former has, therefore, been assigned a much lower degree of
complexity (in relation to the types and number of tests performed) than has the latter. In this perspective, instead of
transporting passengers from a peripheral airport to the central one, it was planned to transport patients’ test tubes from
a peripheral laboratory to the central one. In the economy of
scale, this allowed considerable savings related to the abolition of duplicate analyses between two laboratories in proximity, thus saving on the number of instruments, reagents,
and personnel. This model functionally operates according
to a network in which each laboratory has a well-dened task
and function (Fig.3.3) and is qualitatively and economically
sustainable after an accurate feasibility study. Briey, it is
necessary to create a balance between two opposing models,
characterized by an unjustied multiplication of laboratories
in the same area and by an irrational concentration of exams
in laboratory “mega-structures” that, because of distance and
volume, risk losing proximity to the clinician and the patient.
In concrete terms, the facilities operating within a network
should be characterized by a proper balance between general
laboratories, also divided into specialized sectors, and specialized laboratories with organizational autonomy. This can
be achieved in relation to the discipline and the type of activity involved, even for over-company catchment areas, and in
explicit compliance with all the criteria provided in the reorganization process. Therefore, effective departmental inte-
The paradigm that has inspired most reorganization processes is the so-called hub-and-spoke model. This model was
developed in the United States following the deregulation of
commercial civil aviation and was originally introduced by
the airline Delta in 1955, with the identication of Atlanta
airport as the nodal point on which to concentrate most longdistance ights. In particular, the channeling of long-distance
connections to a reference pole, called a hub, and the creation of a series of connecting ights from peripheral airports, called spoke, allowed the company to make much
more efcient use of resources without penalizing passengers too much. In essence, instead of starting a daily direct
Spoke
Spoke
HUB
Spoke
Spoke
Fig. 3.3 Organization of a network of laboratories according to the
hub-and-spoke model. (Copyright EDISES 2021. Reproduced with
permission)
Spoke
Spoke

3 Elements ofBiomedical Laboratory Organization
https://t.me/medicina_free
19
gration of the structures in the corporate, supra-provincial,
and regional network must be developed, through a strong
level of coordination, sharing of management processes,
quality policies, and continuous staff training. Another critical aspect concerns the transport of samples, especially when
the distance between the central and peripheral laboratories
is considerable and transport is further complicated by complex orography (mountains, busy roads). The use of systems
that allow adequate storage of samples and monitoring of
transport conditions is therefore essential to ensure the integrity of samples and the quality of the results. In this highly
complex scenario, there is also the problem related to the
so-called decentralized diagnostics, carried out using pointof- care instrumentation, which is not dealt with in this
chapter.
Roles oftheLaboratory Sta
In the description of the different professional gures working in a laboratory, it is important to underline how the activity is always, in any case, carried out as a team, in a
harmonious equilibrium prodromal to the production of the
results in the form of the nal report. For the sake of simplicity, we will describe in this part of the chapter the roles and
functions of the director, the health manager, and the laboratory technicians, since only for them, in Italy, there is a professional prole well-dened by law regarding the functions
to be performed in a laboratory.
The gure of a laboratory director requires a length of
service of no less than 7years in the discipline, a specialization diploma in the discipline, or a length of service of
10 years in the discipline, participation in a management
training course, and a suitable professional case history. It
should be noted that law number 229 of June 19, 1999, identies only one level of management for access to the apical
position of director of complex operational units (UOCs) of
laboratory analysis and/or clinical pathology and/or clinical
biochemistry (or other similar denominations). They are,
therefore, provided for single competition procedures for
graduated doctors, biologists, and chemists, provided they
have a specialization, thus emphasizing the subordination of
the degree to the course of study and professional. In the
event of an absence or an impediment, the director delegates
his/her functions to a collaborating manager. In summary,
the director in charge of abiomedical laboratory performs a
series of complex tasks that include:
• Negotiation of the budget of the laboratory for which he/
she is responsible and transmission (“cascading”) of the
outcome of the negotiations
• Choice and approval of analytical methods, being person-
ally responsible for the reliability of test results
• Formulation of proposals for the acquisition of diagnostic
systems, taking into account the state of the art of technology, consistent with predened nancial sustainability
• Organization of services and quality control, being personally responsible for the suitability of equipment and
facilities
• Signature of the results of analyses and/or diagnostic
judgments
• Recording and archiving of test results
• Application of the internal regulations (hygienic state of
the premises, good functionality of the installations, and
materials used)
• Response to warnings and complaints
• Application of the rules protecting operators against the
risks arising from the specic activity
• Formulation and management of proposals for the updating and continuous training of personnel health in light of
the company’s mission and the resources actually
available
The gure of a health manager (doctor, biologist, chem-
ist, or other qualifying degrees), who works in a laboratory, is purely managerial, requiring specic skills in the
organization of activities for maintaining and improving
the quality of the results, and is also expressed through an
activity of advice to clinicians on the appropriateness of
the request, interpretation of the results, and possible continuation of the diagnostic process. A manager, however,
enjoys a relative professional autonomy, as dened by his/
her organizational or functional assignment, subordinate
to the organization of the activities dened by the director.
Within his/her sphere of responsibility, a manager organizes the activities of his/her sector in collaboration with
the director, the technical coordinator, and the technical
and auxiliary staff and participates in the solution of clinical questions using the diagnostic potential assigned.
Access to health management in the laboratory requires a
university degree, an exam qualifying the profession and
its registration, and a postgraduate specialization in a qualifying discipline (clinical biochemistry, clinical pathology,
microbiology, or equivalent). It follows, therefore, that the
prole of a health manager is that of a professional who
has in his/her curriculum vitae no less than 9–10years of
targeted studies (degree and specialization) that allow him/
her to perform and/or coordinate analytical activities, both
clinical and managerial. It is important to remember that,
following the ItalianInter-ministerial Decree number 68
of February 4, 2015, on the reorganization of specialization schools in the health area, the specializations of clinical pathology and clinical biochemistry have been merged
under the new name of specialization in clinical pathology
and clinical biochemistry. In summary, a specialist in this
discipline must:

20
https://t.me/medicina_free
G. Lippi et al.
• Develop theoretical, scientic, and professional knowledge (including the relative assistance activities) in the
eld of diagnostic clinical pathology and laboratory
methodology in cytology, cytopathology, immunohematology, and genetic pathology and in the diagnostic
application of cellular and molecular methodologies in
human pathology
• Acquire skills in the diagnostic clinical aspects of reproductive medicine and in the laboratory of medicine of the
sea and sports activities
• Acquire skills in the study of cellular pathology in the
elds of oncology, immunology, and immunopathology,
and genetic, ultrastructural, and molecular pathology
• Acquire theoretical, scientic, and professional knowledge for laboratory diagnostics on human samples related
to the problems of hygiene and preventive medicine, control and prevention of human health in relation to the environment, occupational medicine, community medicine,
forensic medicine, thermal medicine, and space medicine
• Develop theoretical, scientic, and professional knowledge in the study of biological and biochemical parameters in biological samples as well as in vivo, also in
relation to the pathophysiological states and clinical biochemistry of nutrition and motor activities, at different
levels of structural organization, from single molecules to
cells, tissues, and organs, up to the whole organism, both
in humans and in animals
• Acquire the necessary skills to study the indicators of
alterations underlying hereditary and acquired genetic
diseases
• Acquire the necessary skills for the development, use, and
quality control of (1) clinical molecular biology, molecular diagnostics, and recombinant biotechnology methodologies, including for the diagnosis and assessment of
disease susceptibility; (2) instrumental technologies,
including automated ones, which allow a quantitative and
qualitative analysis of the above parameters at high levels
of sensitivity and specicity; and (3) biochemical molecular technologies linked to human and/or veterinary clinical diagnostics and to environmental diagnostics relating
to xenobiotics, residues, and additives, including in food.
The professional prole of a laboratory technician is well-
dened by law in Italiandecree number 745 of September
26, 1994 (and subsequent law number 42 of February 26,
1999, and law number 251 of August 10, 2000), which has as
its objective the “regulation concerning the identication of
the gure and the relative professional prole of the biomedical laboratory technician.” Specically, a laboratory technician is a professional gure in the health sector in possession
of a qualifying university diploma, with duties aimed at analysis and research related to biomedical and biotechnological
analysis, in particular biochemistry, microbiology and virology, drug toxicology, immunology, clinical pathology, hematology, cytology, and histopathology. Genetics and molecular
biology have obviously been recently added to these disciplines. The activities are carried out (in public and private
facilities, authorized according to the current regulations) in
technical and professional autonomy, in direct collaboration
with the laboratory manager in charge of the various operational responsibilities. A laboratory technician is also responsible for the correct fulllment of the analytical procedures
and his/her own work within the scope of the functions in the
application of the work protocols dened by the responsible
managers. He/she also:
• Veries the correspondence of the services provided to
the indicators and the standards predened by the head of
the structure
• Checks and veries the correct functioning of the equipment used
• Provides for routine maintenance and the possible elimination of minor inconveniences
• Participates in the planning and organization of work
• Contributes to the training of support staff and directly
contributes to updating their professional prole and
research
In his/her activity, a laboratory technician must be able to
autonomously assume responsibility for processes and decisions to implement interdisciplinary and interprofessional
work in the complex care contexts in which the user expresses
his/her health needs. In addition to specic professional
skills, the gure of a laboratory technician is equipped with
transversal skills, not specic to particular roles but related
to knowing how to act in different situations. Law number 43
of February 1, 2006, regarding provisions on health professions, states that as a result of the new university curricula,
the graduate staff belonging to health professions are classied into four categories, which contemplate the functions of
health-care coordination, specialist coordination, and management, reserved to holders of a specialist degree.
The technical staff is therefore subordinate to the gure of
the technical coordinator, the person in most direct contact
with the director of the unit in the organization of laboratory
activities. The coordinator coordinates and organizes the
professional and economic resources assigned, thus full time
carrying out many management tasks within a laboratory in
which he/she operates. His/her role is crucial because,
through an analysis of the needs of the laboratory, workloads, and availability of staff, he/she has the task of planning and managing the work, even in the event of
unpredictable and unforeseen events (absences, malfunctions of technical and/or instrumental aids).

The Role ofStatistics inLaboratory
https://t.me/medicina_free
Medicine
MatteoVidali
4
Introduction
Statistics plays a crucial role in many areas of laboratory
medicine, from validation of new methods, to verication of
analytical and diagnostic performances, denition of reference intervals, and design and conduction of experiments.
The knowledge and the correct application of statistical
methods allow to deal with the variability of laboratory data,
to organize and synthesize information in order to make
objective clinical decisions.
Moreover, methodological errors in quality monitoring,
in terms of both data analysis and experimental design of
procedures, can determine a serious clinical risk for the
patient and an excessive consumption of economic resources.
Finally, the introduction of new and complex statistical
methods in the clinical laboratory, their increasing diffusion
in scientic publications in the eld and their implementation in easily available software packages, have led to consider as basic statistical tools those that until a few years ago
were advanced statistical techniques.
In this chapter, together with introductory statistical concepts, the main statistical methodologies used in the most
common scenarios of laboratory medicine are explained.
However, these data analysis techniques are not presented as
a mere list of tools, but organized into possible experiments
in order to make the reader to understand their use and impact
and to provide correct and practical solutions to the main
problems encountered in the activity of a laboratory.
Basic Statistical Tools
Samples andPopulations
In statistics, population means the nite or innite set of all
possible elementary units, homogeneous for a given characteristic, to which the statistical investigation refers. Rarely,
the experimenter has access to all the units of the population
(too costly in terms of money or time, or because the population is as innite as the innite number of measurements that
can be made on a sample under specic conditions). Usually,
the experimenter extracts a representative subset (sample) of
the population (sampling process), calculates the desired
sample statistics, and uses this information to estimate the
true parameters of the population (inference process). Our
interest is therefore not limited to the sample but to what the
latter can tell us about the population. However, the sample
estimate is affected by uncertainty (a different sample would
have provided a different estimate) and must be interpreted
in probabilistic terms.
We can therefore distinguish between descriptive statistics, which uses graphical and numerical methods to describe,
summarize, and present data, and inferential statistics, which
uses the information obtained from the sample to make more
general statements that are valid and referable to a broader
context than the data from that single experiment.
Descriptive Statistics
M. Vidali (*)
Clinical Pathology Unit, Fondazione IRCCS Ca’ Granda Ospedale
Maggiore Policlinico, Milan, Italy
© The Author(s), under exclusive license to Springer Nature Switzerland AG 2023
M. Ciaccio (ed.), Clinical and Laboratory Medicine Textbook, https://doi.org/10.1007/978-3-031-24958-7_4
The statistical analysis of the results of an experiment should
always be preceded by their visualization; for this purpose, a
valid tool is the histogram which is the graphical representation of the frequency distribution of a quantitative variable.
The frequency distribution is obtained by dividing the measurement scale into intervals or classes of equal amplitude
and counting the number of observations that fall within each
class (absolute frequency). The histogram is made up of many
21

22
Frequency Density
Cholesterol (mg/dL)
o
µ
∑ x
N
x
x
N
∑
()
N
s
xx
()
()
()
N
s
xx
()
()
s
x
https://t.me/medicina_free
0.020
0.010
0.000
140 160 180 200 220
Fig. 4.1 Distribution (relative frequency density) of 300 cholesterol
observations, displayed via histogram, frequency polygon, and boxplot
adjacent rectangles whose base is the class and whose height
is the frequency. In order to avoid loss of information (when
the number of classes is too small), or efciency in the synthesis (when the number of classes is high), it is advisable to
choose a number of classes between 5 and 20 and of equal
amplitude, or the use of formulas such as Sturges’ one
k=1+(10/3)log10N, where k and N represent the number of
classes and observations, respectively. As an alternative to
absolute frequency, it is possible to use relative frequency
(ratio of absolute frequency to the total number of observations) or relative frequency density (ratio of relative frequency
to class amplitude) (Fig.4.1). Note that in the case of relative
frequency density, the area of each bar of the histogram is
equal to the relative frequency for that class, so the area of the
entire histogram is equal to 1 (sum of all relative frequencies).
The histogram makes easy to check how the data are distributed, that is, where the observations are concentrated, the possible presence of extreme observations and/or asymmetry.
From the histogram, it is possible to construct the frequency
polygon by joining the midpoints of the upper bases of each
rectangle and closing the polygon by joining the ends of the
line thus obtained with the x-axis (Fig.4.1). The frequency
polygon, unlike the histogram, allows us to compare different
distributions in the same graph.
Although the frequency distribution is a useful tool to
illustrate a quantitative variable, it is sometimes necessary to
summarize the data with a few values: position (or central
tendency) and dispersion measures are used for this purpose.
One must remember the conceptual difference between population parameters (denoted by letters of the Greek alphabet),
which are generally unobservable but estimable quantities,
and estimates of these parameters or statistics (denoted by
letters of the Latin alphabet) calculated from sample data.
The two most widely used measures of position are the arithmetic mean and the median. The former, calculated as the
ratio of the sum of all observations to their number
(
for the population,
i
=
i
for the sample), has
=
M. Vidali
the advantage of synthesis and ease of calculation, but the
disadvantage of being inuenced by extreme values. The
second one, less dependent on extreme values, represents the
value that divides an ordered set of data into two equal parts,
so that half of the observations have a value lower than the
median and half a value higher than the median; in other
words, in an ordered set of observations, the median is the
value corresponding to the [(N+1)/2]-th observation (if the
number of observations is odd) or the value corresponding to
the arithmetic mean of the values of the two central observations (N/2) and [(N/2)+1] (if the number of observations is
even).
Measures of dispersion are useful in assessing the vari-
ability of data around their central tendency. The most widely
2
x
µ
i
x
for the population,
2
µ
i
for the population,
used are the variance (
∑−
2
=
2
i
N
for the sample) and its square root, that is,
1
−
the standard deviation (
∑−
=
2
i
N
for the sample). The advantage of the
−
1
2
Ã
σ=
=
∑−
∑−
standard deviation, compared to the variance, is that it is
expressed in the same units as the observations. Note that in
the denominator of the variance and ofthe standard deviation
does not appear N but (N− 1), that is, the degrees of freedom. In general, the degrees of freedom (df) are equal to the
difference between the number of data and the number of
estimated parameters (in the standard deviation, we have
N−1 degrees of freedom because the mean was estimated).
When we want to compare the variability of two or more sets
of data, in particular when the observations are expressed
with different units or orders of magnitude, it is preferable to
use the coefcient of variation, which expresses the standard
deviation as a percentage of the mean %CV =×
100 .
Other measures often used to summarize a distribution
are quantiles. A quantile q is a value such that, in an ordered
set, a proportion q of data is less than the value corresponding toq and a proportion 1− q of data is greater than the
value corresponding to q. In particular, a quantile q in an
ordered set is the value corresponding to the observation
i=q(N + 1) (if i is not an integer, the quantile is found by
interpolation). Particularly useful quantiles are the quartiles
(Q) and the percentiles (P) that correspond to those values
that divide an ordered set into 4 and 100 parts of equal size,
respectively. The median corresponds to the second quartile
(Q2) and the 50th percentile (P50). From the denition of
quartiles, it is clear that 25% of a data set will have a value
lower than Q1 (P25), 25% higher than Q3 (P75), and the

xe
()
σπ
z
x
−
µ
σ
Z
X
−
µ
σ
z
2
−
z
1
−
z
333
−
z
667
−
σ
4 The Role ofStatistics inLaboratory Medicine
https://t.me/medicina_free
remaining 50% will have a value within the interquartile
range (Q3-Q1 or IQR). Similarly, the interval between the
2.5 percentile (P2.5) and the 97.5 percentile (P97.5) includes
95% of the observations: in fact, 2.5% of the observations
will be lower than P2.5, 95% between P2.5 and P97.5 and
the remaining 2.5 will be greater than P97.5. The combination of ve numbers, Q1, Q2, Q3, min, and max, is a convenient and often used way in the literature to summarize a
distribution. This information can be visualized in the box
and whisker plot or boxplot, in which the box, bisected by
the median Q2, has Q1 and Q3 as its ends, with whiskers
extending from Q1 and Q3, respectively, to the minimum
and maximum values equal to or less than 1.5 times the IQR
interval. The extreme observations, not reached by the whiskers, are represented as points (Fig.4.1).
Probability Distributions andtheNormal
Distribution
As the sample size (N) increases, the class size (A) of the
histogram may be decreased; in general, as N tends to innity, A tends to zero and the histogram (or frequency polygon)
is approximated by a smooth curve. When N tends to population size (i.e., with N very large), this smooth curve becomes
the relative frequency density of the population. Just as the
total area of the histogram is equal to 1, the area under the
curve is also equal to 1. The proportion of observations that
falls between two xed limits is equal to the area under the
curve between those two limits. Thus, since the probability
of a randomly chosen observation falling between two limits
is equal to the proportion of observations falling between
these two limits, the population frequency distribution is
called the probability distribution or probability density
function.
The most common probability distribution in statistics is
the normal or Gaussian distribution (or bell curve) (Fig.4.2).
2
z
−
2
Its probability density has equation: f
()
1
=
2
with
=
.
The Gaussian distribution is completely dened by mean
and standard deviation. It is a symmetrical distribution in
which the mean, median, and mode (most frequent value)
coincide. This distribution is particularly important for at
least two reasons: many biological variables of interest to the
laboratory, including measurement errors, are approximately
normally distributed and, in addition, many statistical methods are based on the assumption of normal distribution of the
data. Knowing that a variable is normally distributed and
knowing the mean and standard deviation, we can know the
probability that a randomly selected individual will exhibit a
value greater than, less than, or within 2 xed limits. For
/
µ–3σ µ–2σ µ–1σ µµ+1σ µ+2σ µ+3σ
µ±1σ = 68.2%
µ±2σ = 95.4%
µ±3
= 99.7%
Fig. 4.2 Gaussian probability distribution
example, knowing that the variable glycemia has μ=90mg/
dl and σ=15mg/dl, we might want to know the probability
that a randomly selected individual has a glucose level
greater than 120, or less than 75, or between 85 and 100mg/
dl. To answer this question, we must determine the proportion of the area under the curve to the right of x1=120, or to
the left of x2=75, or between x3=85 and x4=100 mg/dl,
respectively. These areas have been calculated and tabulated
and can be found in statistical texts, literature, or statistical
software. However, it would not be possible to tabulate values for all the innite normal curves given by the combination of the different means and standard deviations. In fact,
only one particular curve was tabulated, called the standardized normal distribution, with μ=0 and σ= 1. To use the
tabulated area table in order to estimate the probability associated with a normal variable X, it is necessary to transform
the normal variable X into a standardized normal variable Z
with the operation
look for the relative probability value in the table. In our
example,
85 90
=
3
15
=− . ;
0
=
1
=
120 90
15
100 90
=
4
(standardization) and then
= ;
= . . The area under
15
0
2
75 90
=
15
the curve to the right of z=2 is equal to 0.023; to the left of
z=−1 is 0.159; included between z=−0.333 and z=0.667
is 0.378; this means that the probability that a subject randomly selected from this population presents a glycemia
value higher than 120 P(X>120), or lower than 75 P(X<75),
or included between 85 and 100mg/dl P(85 <X<100) is,
respectively, equal to 2.3%, 15.9%, and 37.8% (Fig.4.3). It
is easy to verify that in a Gaussian distribution 68.4% of the
23
=− ;

24
–3 –2 –1 0 +0.67 +1 +2 +3–0.33
µ = 90
Theoretical quantiles
Observed quantiles
()
λ
λλ
10
/ n
https://t.me/medicina_free
M. Vidali
σ = 15
45 60 75 90 100 105 120 13585
– µ
X
i
Z =
σ
µ = 0
σ = 1
Fig. 4.3 Standardization procedure: from normal curve to standardized normal curve with mean=0 and standard deviation=1
area under the curve is between mean ±1 standard deviation
(in the standardized normal between z=−1 and z=1), 95.4%
between mean ±2 standard deviations (between z=−2 and
z = 2), and 99.7% between mean ±3 standard deviations
(between z=−3 and z=3) (Figs.4.2 and 4.3).
Since many statistical methods can only be used if the
observations are normally distributed, it is necessary to verify this assumption. Graphical and formal statistical methods
can be used for this purpose. The visualization of the data
through histogram and boxplot and the comparison of the
position indices (mean and median) allow to easily verify the
presence of deviations from normality: since, in fact, the
mean is affected by extreme values, in distributions with
positive asymmetry (right long tail), the mean will be higher
than the median, and, in distributions with negative asymmetry (left long tail), the mean will be lower than the median.
A reliable tool, also for small samples, is represented by the
normal probability plot, a graphical method in which the
observed quantiles (y-axis) are graphed toward the theoretical quantiles (x-axis). The presence of pronounced deviations from the theoretical quantile-quantile line suggests the
200
180
160
140
–2 –1 012
Fig. 4.4 Normal probability plot to verify the normal distribution of
the data. The points are well aligned along the theoretical line that
passes through the rst and third quartiles. The graph shows only a
slight curvature to the left (bottom of the curve) with a point identied
as aberrant by Horn’s method
assumption of non-normality in the data (Fig. 4.4). The
assumption of normality can also be veried with formal statistical tests such as the Kolmogorov-Smirnov test, the
Anderson-Darling test, and the more popular Shapiro-Wilk
test (not usable, however, for large samples).
In the presence of signicant deviations, it is possible to
attempt mathematical transformations of the data in order to
proceed subsequently to a new normality check. Several
transformations can be applied to the data (including logarithm, reciprocal, square root, power-elevation), and it is
often necessary to proceed by trial and error. Instead of using
arbitrarily chosen transformations, it is possible to use the
transformation proposed by Box and Cox, a technique that
allows to determine which is the best transformation to apply
to the variable y. The transformation is dened by
y
=
log,ifif
−
/,
y
()
≠
Where W is the transformed
=
λ
0
variable, y the original variable, and λ the parameter dening
the transformation estimated using the maximum likelihood
criterion.
Inferential Statistics: Condence Intervals
andHypothesis Tests
Imagine extracting all possible samples of size n from the
population and calculating the mean of each of them. The
distribution of all possible averages is called the sampling
distribution of the mean. More generally, we speak of sampling distributions to denote the distribution of the values of
different statistics computed over all possible samples.
The sampling distribution of the mean will have its own
mean and standard deviation. It can be shown that: (1) the
mean of the sampling distribution (i.e., the mean of the averages) is equal to the population mean μ; (2) the standard
deviation of the sampling distribution of the mean, or standard error (es), is equal to
; (3) if n is sufciently large,

µµ
−<<+
61
./ÃÃnx n
x
i
µσ
±
./n
x n±
./
σ
x n±
./
σ
x n−
./
σ
x n+
./
σ
/ n
zn
/
xs
//
xs
//
x
sn
−
−
µ
x tsn
n±−
α
.;
rse=±×k
4 The Role ofStatistics inLaboratory Medicine
https://t.me/medicina_free
25
the sampling distribution is approximately normal (the more
the population deviates from normality, the larger n must be).
Condence Intervals
Just as 95% of the observations in a normal distribution are
within ±2 standard deviations (more precisely 1.96 sd) from
the mean (Fig.4.2), so, if n is sufciently large, 95% of the
averages in the sampling distribution of the mean are within
±1.96 standard errors from μ (i.e., x
19
there is a 95% probability that the average
sample is within the range
is equivalent to saying that μ is within the range
96./
), which means that
of a particular
196
); this, conversely,
196
or that there is a 95% probability that the limits of the interval delimited by
196
contain the true population
mean μ. This interval is called the 95% condence interval,
and the extremes of the interval
196
are called the 95% condence limits. In real
196
and
life, we do not have all possible samples of size n extracted
from a population, and σ is unknown; however, if n is sufciently large (n>60), not only does the sampling distribution
approximate well to a normal distribution but s (the sample
standard deviation) is a good estimate of σ (the unknown
standard deviation of the population) and
good estimate of
. This result allows us to associate a
/ is a
statistics calculated on a sample (point estimate) with an
interval in which we have some level of condence that it
contains the true population parameter (interval estimate).
The condence interval ±
σ
increases as the selected
condence level increases (z=1.64 for 90%, z = 1.96 for
95%, z = 2.58 for 99%) and decreases as n increases (the
estimate becomes less uncertain). With small samples
(n<60), s may not be a reliable estimate of σ, because s can
vary greatly from sample to sample, and the ratio
2
µ
n−
does not follow a standardized normal distribution. However, for small samples, it can be shown, that
if the observations come from a normal distribution, the ratio
2
µ
n−
follows a Student’s t distribution with
(n− 1) degrees of freedom. The t distribution is atter and
has thicker tails than the normal distribution. It is actually a
family of distributions whose shape depends on the number
of degrees of freedom. In general, for the same proportion of
observations under the curve or the same probability, t has a
larger value than z, resulting in wider condence intervals
and hence greater uncertainty in the estimate. As n, and
hence degrees of freedom, increase, t tends to z.
Hypothesis Testing
A second type of inference is represented by statistical
hypothesis testing. Hypothesis testing checks the validity of
a hypothesis and involves rst formulating a hypothesis relative to a characteristic of the population and then evaluating
the probability of obtaining the observed sample (and therefore the statistics calculated on the sample) if the hypothesis
is true. The hypothesis to be tested is called null hypothesis
and is indicated with H0, while the alternative hypothesis,
complementary to the null hypothesis, is indicated with H1.
Since hypothesis testing is a procedure based on probability,
two errors may occur: an error of the rst type, or alpha error
(α), which is the one committed if we reject H0 when it is
true, and an error of the second type, or beta error (β), which
is committed if we accept H0 when it is false.
In general, we proceed as follows: (1) we x H0 and H1; (2)
we calculate the desired sample statistic and the value of the
statistical test; (3) using the tables of the appropriate distribution, we calculate the probability p of obtaining a statistic like
the observed one, or more extreme, under the hypothesis that
H0 is true; (4) we compare p with α (xed in general equal to
0.05): if p is less than α, we consider the test statistically signicant and reject the null hypothesis (the null hypothesis is not
consistent with the observed results) otherwise we accept it.
Example: To test the hypothesis that smoking affects
blood pressure, we design a study in which the mean blood
pressure of 20 heavy smokers is measured. The working
hypothesis or alternative H1 is “heavy smokers have a different mean blood pressure than the non-smoking population,”
while the null hypothesis H0 is “heavy smokers have the
same mean blood pressure as the non-smoking population.”
The value of the mean arterial pressure of the population
(obtained from previous studies) is μ= 148 mmHg; we set
α=0.05; H0:μ=148mmHg; H1:μ≠148mmHg; the sample
of subjects studied (n=20) has mean x=152mmHg and
standard deviation s=8mmHg; we calculate the value of the
statistical test that in this case is
=
n−
1
152 148
=
//
820
=
.
224
; using the distribution t (non-large sample) with (n−1=19)
degrees of freedom, we look for the probability value corresponding to t
= 2.24, which is p = 0.037. This means
α/2;n−1
that, if the null hypothesis is true (μ=148mmHg), the probability of obtaining by chance a result like the observed
(x=152mmHg), or even more extreme, is 3.7%. Therefore,
since p<α, we have sufcient evidence to reject H0 and conclude that the mean arterial pressure of smokers is different
from the pressure of non-smokers. With the data from the
previous example, we can also calculate the 95% condence
interval of the mean:
152 820 152 2 093 179 148 3 155 7
±× =± ×= −t
0 025 19
.;
/...
/;/21
, that is,
conrming what previously concluded, the condence interval does not contain the value 148mmHg (we are 95% condent that the observed sample does not come from the
population with mean μ=148mmHg).
In general, from a sample statistic (a mean, a difference
between means, or some other estimator), it is possible to construct a condence interval and to calculate the test statistic:
5%CI Estimatedparamete

26
T
rse=
with
−
=≠
d
HH
δ
µδ
::
0121
00
s
2
s
2
2
F ss=
222
s s
2
>
±+
µµ
()
()
ns
22
11
s
1
s
2
2
pp
22
11
()
z
pp
nn
−
()
−−
12
12
11
()
ππ
np np
nn
+
22
12
()
=
=
i
xx
2
∑−
()
=mj
()
=mj
2
within
within
F F>
within between
https://t.me/medicina_free
M. Vidali
estEstimated paramete
/
with the standard error (se) calculated in different ways
depending on the estimator and k depending on the distribution considered. Let us see some examples (z can substitute t
for large n), where the statistical test,calculations for CI and
the specic statistics are reported, together with H0 and H1:
1. Test for two paired samples (e.g., subjects measured
before and after a treatment): t-test for paired samples
with (n−1) degrees of freedom, where n is the number of
pairs of observations
CI
=± =
dt sn t
α
=−
δµ
−
n
/;
21
d
/;
sn
d
and
/
with d e sd mean and standard deviation of the differences, respectively, and δ the true difference between
population means
2. Test for two independent samples (n1 and n2): Student’s
t-test with (n1+n2 − 2) degrees of freedom
(a) With variances
and
1
equal (hypothesis of homogeneity or homoscedasticity of variances veriable by
tests for variance
/ with
1
122
and F following a Fisher distribution with (n1–1) and (n2–1)
degrees of freedom (other tests that can be used for
the same purpose are Bartlett’s test and Levene’s test):
2
p
with
.
11
nn
12
H
01 2
:
=
2
2
CI
=−
xx ts
()
xx
()
t
=
−−
−
12 12
11
//
sn
p
1
and
H
+−
nn12 22
α
/;
12
µµ
()
nn
+
2
:
≠
µµ
11 2
where μ1 andμ2 are the population means and s
is the common or pooled variance: s
2
−
nn
12
+−
+−
2
2
e
non-homogeneous (heterosce-
ns
11
=
(b) With variances
dasticity): other tests must be used (e.g., Welch test)
3. Test for two proportions (p1 and p2 calculated in two samples of size n1 and n2): test z
(( )(
CI =−
pp z
()
±
12 2
α
/
−
pp
11
n
1
−
+
;
n
2
4. Tests for three or more independent samples: ANOVA
test or analysis of variance
The null hypothesis H0 in this case is that the samples are
from the same population (same means). One approach
would be to perform as many two-way comparisons with as
many Student’s t-tests. However, with m groups, the possible
two-way comparisons are m(m−1)/2 and, as the number of
groups, and hence comparisons, increases, the probability
that at least one comparison is statistically signicant
increases even if H0 is true (i.e.greater probability of committing a Type I error: if for one comparison the probability
of not rejecting H0, when true, is 0.95, for k comparisons it is
(0.95)k and so the probability of rejecting H0 in at least one
comparison, if H0 is true, is 1–0.95k; for example, if m=5
groups, k =10 comparisons and the probability of rejecting
H0, when true, in at least one comparison is 1–0.9510=0.40
much larger than α xed).
The ANOVA test instead of considering averages considers variances. It is in fact based on the reasoning that when
dealing with several populations (e.g., different treatments),
the total variability depends on the variability of individual
values with respect to the mean of their population, or withingroup variance (due to measurement error, individual characteristics, or uncontrollable factors), and on the variability
of their population averages with respect to the overall mean,
or between-group variance. If the variability within populations (within-group variance) is small compared to the variability between population averages (between-group
variance), we conclude that the averages are different. With
m groups, we proceed in this way by calculating:
1
• Deviance within groups:
n11 degrees of freedom, which measures the vari-
j
S
within
1
=∑∑−
mjn
ij j
j
, with
ability of the data around the group mean
p
• Deviance between groups:
p
(m−1) degrees of freedom, which measures the variabil-
S
between
1
=∑ −
nx x
jj
ity of group averages around the overall mean
• Total failure: SS
df
total
=df
within
+df
• Variance within groups: MS
• Variance between groups: MS
total
between
= SS
within
within
between
+ SS
SS
=
df
=
within
SS
df
between
between
between
, with
, with
12
=
pp
(
−
1
()
H0:π1=π2 and H1:π1≠π2, whereπ1 andπ2 are the popu-
lation proportions, assuming that, n1p, n1(1-p), n2p and
n2(1-p)are all ≥5.
with p
+
MS
11
=
,
+
• Fisher’s test: F=
if
α
;;df df
between
MS
, then p<α and we reject H0: at least one
of the m groups is signicantly differenttinct from the others.
A signicant F-test indicates that not all averages are equal
but does not allow us to know which and how many are dif-

αα
==
()
12
n
i
()
()
=
∑
∑
1
xx
ini
∑−
()
=
s
xx
i
()
=∑1
n
()
y
i
b
b
()
yy
ii
()
yi
yy
ii
()
r
xxyy
∑−
()
−
()
∑−
()
()
==
4 The Role ofStatistics inLaboratory Medicine
https://t.me/medicina_free
27
ferent from each other. To answer this question, we need to
run tests for multiple comparisons, that is, comparisons for
each pair of averages. To avoid, as shown before, that increasing the number of comparisons leads to an increase in the
probability of committing a type I error, various approaches
or various corrections can be used. For example, Bonferroni’s
correction suggests xing the type I error by calculating a
modied α as
∗
numberofcomparisons mm
,
−
/
where m is the number of groups. The main drawback of the
Bonferroni correction is the decrease in statistical power as
the number of comparisons increases (i.e., the probability of
rejecting H0,when it is false, decreases). Several other tests
for multiple comparisons are described in the literature
including the Student-Newman-Keuls test, Tukey’s method,
and the Holm-Bonferroni method.
Like the t-test, two assumptions must be met for the
ANOVA test: (1) the variable is normally distributed; (2) the
variances of the groups are equal (homoscedasticity). Both
assumptions can be veried through graphical methods (histogram, normal probability plot, scatter plot) and formal statistical methods (Shapiro-Wilk test for normality and Levene’s
test for homoscedasticity). While moderate deviations from
normality can be ignored (especially for groups of equal
numerosity), non-homogeneity of variances can greatly inuence the results. In the presence of violations of these assumptions, it is advisable to try a mathematical transformation
(which can both normalize the distribution and make the variance similar across groups) or to use non- parametric methods, that is, methods that do not depend on the shape of the
distribution and do not involve the estimation of statistical
parameters: the Mann-Whitney test for two independent samples, the Wilcoxon test for two paired samples, and the
Kruskal-Wallis test for three or more samples. Although nonparametric tests do not require the assumptions underlying
parametric tests, it should be remembered that, if these
assumptions are met, non-parametric tests have less statistical
power.
Correlation andLinear Regression
Sometimes we are interested in evaluating the relationship
between two quantitative variables. In this situation, it is necessary to initially display the data in a scatter plot, that is, a
dot plot whose coordinates are represented by the two variables being studied. Linear regression allows us to estimate
the mathematical relationship between an independent variable (x-axis) and a dependent variable (y-axis). Correlation,
on the other hand, allows us to estimate the strength of the
association between the two variables.
The regression provides the equation of the line that best
describes the dependent variable as a function of the independent variable: Y=a+bX+E, where a and b are, respec-
tively, the intercept and the slope coefcient (or regression
coefcient) of the line and E is the error, that is, the part of
the variability of Y not explained by the line. The parameters
of the line are estimated through the method of least squares,
that is, looking for the values of the parameters that minimize the sum of the squares of the vertical distances of the
points from the line. The following formulas are used:
−
=
n
()
xxyy
ii
i
=
1
xx
−
i
−
aybx
=−
;
2
The parameters can then also be used to make predictions
by the equation of the line.
As with other estimators, we can calculate condence
intervals and conduct hypothesis tests for a and b.
The condence intervals are as follows:
2
for a: CI=a± t
for b: CI=b±t
∑−
inii
with
s
=1
=
α/2; n−2
α/2; n−2
yy
−
SE(a) with
SE(b) with
2
, where
2
E as
()
E( )b
1
=+
n
=
are the corresponding
x
1
n
i
2
2
−
values of yi predicted by the regression line. The degrees of
freedom are (n−2) because we estimated two parameters a
and b.
It is also possible to test the null hypothesis that the population regression coefcient is 0, that is, that there is no linear
relationship between X and Y.
So for H0:β=0 we have
=
SE
.
The use of the least squares method and the t distribution
for hypothesis testing and condence intervals are based on
−
the assumption that the residuals
are normally distributed and have equal variance. These assumptions can be
veried by a normal probability plot of the residuals
andbygraphing the predicted values
−
. Analysis of the residuals can also reveal extreme
toward the residuals
points that are highly inuential on the regression line.
Instead, the strength of the association between two quantitative variables is expressed by the linear correlation coefcient r estimated by:
inii
=
=
()
1
2
xx yy
∑−
()
inii
1
n
1
2
i
The coefcient r ranges from −1 (perfect association, as
one variable increases, the other decreases) to +1 (perfect
association, as one variable increases, the other also
increases). If r=0, there is no relationship between the variables. The correlation coefcient measures the strength of
Соседние файлы в папке Библиотека им академика М.И. Перельмана
