Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_6034_Библиотеки_им_академика_М_И_Перельмана.pdf
X
- •The Lumbar Spine
- •Contents
- •Contributing Authors
- •Preface
- •Acknowledgments
- •Epidemiology and the Economics of Low Back Pain
- •Pathophysiology of Nerve Root Pain in Disc Herniation and Spinal Stenosis
- •Biomechanical Considerations of Disc Degeneration
- •Clinical Spinal Instability Resulting from Injury and Degeneration
- •Morphologic Changes of End Plates in Degenerative Disc Disease
- •Spinal Instrumentation
- •Fracture and Repair of Lumbar Vertebrae
- •Genetic Transmission of Common Spinal Disorders
- •Genetic Applications to Lumbar Disc Disease
- •Clinical Neurophysiologic and Electrodiagnostic Testing in Disorders of the Lumbar Spine
- •Sensorimotor Control of the Lumbar Spine
- •Outcomes Assessment: Overview and Specific Tools
- •The Role of Outcomes and How to Integrate Them into Your Practice
- •Manual Therapy in Patients with Low Back Pain
- •Acupuncture and Reflexology
- •Returning Workers to Gainful Employment
- •Occupational Ergonomics
- •Preparation for Surgery
- •Surgical Approaches to the Thoracolumbar Spine
- •Surgical Approaches to the Lumbar Spine: Anterior and Posterior
- •Posterior and Anterior Surgical Approaches to the Lumbosacral Junction
- •Endoscopic Anterior Lumbar Procedures
- •Biology of Bone Grafting: Autograft and Allograft
- •Bone Graft Substitutes in Spinal Surgery
- •Spinal Instrumentation Overview in Lumbar Degenerative Disorders: Cages
- •Translaminar Screw Fixation
- •Lumbar Disc Disorders
- •Facet Joint Denervation: A Minimally Invasive Treatment for Low Back Pain in Selected Patients
- •Intradiscal Electrothermal Therapy
- •Operative Management of the Degenerative Disc: Posterior and Posterolateral Procedures
- •Posterior Lumbar Interbody Fusion
- •Operative Treatment of Anterior Procedures
- •Operative Treatment of Anterior and Posterior Fusion
- •Degenerative Disc Disease: Fusion Cages and Dowels
- •Minimally Invasive Procedures for Anterior Column Fusion and Reconstruction
- •Degenerative Disc Disease: Complications of Surgery
- •Dynamic Stabilization in the Treatment of Low Back Pain Due to Degenerative Disorders
- •Lumbar Artificial Disc Replacement: Rationale and Biomechanics
- •Lumbar Disc Replacement: Current Model, Results, and the Future
- •Disc Herniation: Definition and Types
- •Disc Herniation: Imaging
- •Disc Herniation: Nonoperative Treatment
- •Operative Treatment of Disc Herniation: Natural History and Indications for Surgery
- •Operative Treatment of Disc Herniation: Laminotomy
- •Chymopapain and Chemonucleolysis
- •Microscopic Lumbar Discectomy
- •Classification, Natural History, and Clinical Evaluation
- •Imaging of Spinal Stenosis and Degenerative Lumbar Spondylolisthesis with Stenosis

CHAPTER 12/OUTCOMES ASSESSMENT / 133
(usually two) occasions under the assumptions that the
first assessment does not affect, or is independent of, the
second or subsequent assessments. In many cases, including clinical settings, this can be problematic because: (a)
a second assessment shortly after the first assessment is
not likely to be independent because the patient is likely
to remember what he or she said and try for consistency
rather than an accurate estimate of the current state, or (b)
a second assessment a few days after the first assessment
may be influenced by treatments that resulted from the
initial visit. Thus, the most accurate method for estab lishing test-retest reliability for an instrument is to have the
patient respond to the instrument on separate occasions
with nothing but a short natural history intervening
between the events. In this situation, which is rarely evaluated in the clinical setting, low reliabilities, meaning
large discrepancies within a patient across time, could
reflect: (a) an instrument that does not have adequate psychometric properties, or (b) a construct under assessment
(e.g., pain, function, or attitudes) that are more state-like
than trait-like, meaning that the construct may not be stable. The consequence of this result is that the instrument
is not useful in evaluating treatment progress because the
score might vary independently of treatment effectiveness on any given day. Thus, the single most important
aspect of the reliability of an instrument used in the clinical setting is that it demonstrates stability, meaning that
the score measures a condition or construct that is not
amenable to quick and unpredictable changes across
short time increments. In practice, internal consistency
reliability estimates often are used as proxies for testretest reliability . Because instrument construction often is
guided by internal consistency estimates, the general
sense is that internal consistency reliabilities are positively biased estimates of test-retest reliability; therefore,
the amount of error expected in a score that is most relevant to the clinical setting is often underestimated.
This notion of stability , of course, be gs the issue of validity because it suggests that the specific constructs chosen to
be of interest need to be reasonably stable in the short r un.
Strictly speaking, instruments do not have validity, but the
scores they generate have validity to the extent that those
scores provide information that aids in making appropriate
decisions. If a test produces a score that is likely to be interpreted as high today but lo w tomorro w, then the lack of stability in that score indicates that it does not provide information useful in making appropriate decisions.
Measures of reliability include Cronbach’s alpha for
internal consistency, kappa for categorical outcomes, and
intraclass correlations or generalizability coefficients
(30) when outcomes are reasonably continuous. It has
been argued that generalizability coefficients are preferable to kappa coefficients under circumstances when
kappa struggles, such as when the number of categories
becomes large (e.g., more than four levels); when the
number of raters scores being compared are greater than
two (although generalized kappa statistics can be computed); and when the prevalence of one of the categories
is low or the sample size is small.
The value of generalizability coefficients is that they can
be developed in ways that isolate potential culprits or candidates for lack of reliability , and that they allo w estimation
of what some consider the most relev ant form of reliability
in the clinical setting: the estimate of the reliability for a
single rater on a single occasion; in other words, the reliability in the classical clinical situation where one clinician
is evaluating a score based on a single reading. A problem
with reliability coefficients, including generalizability
coefficients, always has been that the meaning of the magnitude of the coefficient is not set. A reliability of .7 in
some fields is considered good to excellent, whereas .8 is
considered dismal in other fields.
To better understand the relationship between the reliability coefficient and the consequent differences in
assessments between raters, 100,000 pairs of random normal scores were generated, transformed to have exact
reliabilities of .30, .40, .50, .60, .70, .80, .90, .95, and .99,
again transformed to have the mean and standard deviation associated with selected SF-36 scores and then broken down into six mutually exclusive and exhaustive categories based on the tenth, 25th, 50th, 75th, and 90th
percentiles, as summarized in Table 12-1. These six cate-
TABLE 12-1. Descriptive statistics amd cut points for selected SF-36 scales
Statistics PCS MCS SF PF Example: PCS
Mean 30.6 45.9 39.1 40.8
SD 7.2 14.5 25.0 22.4
Min 13.0 16.0 0 0
P10 22.0 21.5 12.5 15.0 Category 1:(13–21.9) Min–P10;10% of scores
P25 26.0 35.0 25.0 25.0 Category 2:(22–25.9) P10–P25;15% of scores
P50 30.0 52.5 37.5 40.0 Category 3:(26–29.9) P25–P50;25% of scores
P75 35.0 58.0 50.0 55.0 Category 4:(30–34.9) P50–P75;25% of scores
P90 39.0 60.0 75.0 75.0 Category 5:(35–38.9) P75–P90;15% of scores
Max 58.0 63.0 100 100 Category 6:(39–58) P90–Max;10% of scores
MCS, mental component summary; PCS, physical component summary;PF , ph ysical functioning;SD,
standard deviation; SF, social functioning.
Description statistics are based on the initial visit for 376 University of Iowa patients presenting with
low back troubles.

134 /SECTION I/BASIC SCIENCE
gories for each pair of ratings with the specified reliability were then crossed and the level of agreement and disagreement determined based on differences in classif ication group. The percentage of same and different
categorization is summarized in Table 12-2 for each reliability level. This pattern is consistent for all outcome
measures under the assumption of normality; therefore,
separate tables for each outcome were unnecessary.
By way of interpreting the information provided in
Tables 12-1 and 12-2, suppose a test-retest reliability of
.8. In this situation, 42.63% of scores from that instrument are expected to remain in the score category;
45.56% are expected to be in adjacent categories; and
1.87% to differ by two categories, .915% to differ by
three categories, and .026% to differ by four categories.
With a test-retest reliability of .8, a miss by two or more
categories is expected 11.81% of the time. Across 400
patients, just considering two category differences, this
means that approximately 44 patients might have a score:
In category 1 (13 to 21.9), but a true score in category 3
(26 to 29.9); worst case difference 16.9, best case 4.1
In category 2 (22 to 25.9), but a true score in category 4
(30 to 34.9); worst case difference 13.9, best case 4.1
In category 3 (26 to 29.9), but a true score in category 5
(35 to 38.9); worst case difference 12.9, best case 5.1
In category 4 (30 to 34.9), but a true score in category 6
(39 to 58); worst case difference 28, best case 4.1
result in a differential diagnosis suggesting a less aggressive or immediate treatment path (e.g., w atchful waiting),
the lack of reliability of the diagnostic tool clearly affects
quality of care.
It should be noted that a test-retest reliability estimated
based on the clinically relevant generalizability coefficient of one clinician making one rating is likely to be
lower than the .8 estimate used in this example. Further
note that if dropping to a reliability of .7, only 36.18% of
scores are expected to be in the same category if a second
independent evaluation is done at the same time. Thus,
the stability of many outcome measures employed at initial visits, which are often used as ancillary diagnostic
tools in clinical research, may have marginal value when
applied to clinical practice.
In sum, the reliability of many outcome tools used to
evaluate patients presenting with low back pain have limited test-retest reliability evidence. The proxy internal
consistency reliability estimates used are likely to be positivel y biased, suggesting that the differences betw een the
observed and true scores are likely to be larger than
expected; therefore, clinical decisions based on these
scores are likely to be based on unreliable information.
This may be one explanation for the often repeated, generally accepted, but not necessarily well-documented
cliché that if you do not like the opinion of your clinician,
get a second opinion because it will probably be dif ferent.
Of course, the reverse pattern is equally likely: An
observed score in category 3 (26 to 29.9) might reflect a
true score in category 1 (13 to 21.9). Thus, overtreatment
or undertreatment might result if the observed score overestimates or underestimates severity. To the extent that
overestimates of symptom severity result in a differential
diagnosis suggesting a more aggressive treatment path
(e.g., surgery), or underestimates of symptom severity
TABLE 12-2. Magnitudes of disagreement for selected reliabilities
Number of categories of disagreement
Reliability Match 1 2 3 4 5
0.30 23.40 37.99 24.34 10.63 3.06 0.586
0.40 25.65 39.55 23.41 8.966 2.079 0.348
0.50 28.24 41.35 22.00 6.957 1.293 0.158
0.60 31.62 42.99 19.79 4.890 0.663 0.052
0.70 36.18 44.73 16.12 2.765 0.201 0.010
0.80 42.63 45.56 10.87 0.915 0.026 —
0.90 54.15 42.14 3.665 0.046 — —
0.95 65.79 33.58 0.630 — — —
0.99 84.45 15.55 — — — —
For categories of disagreement:
Match indicates no disagreement (i.e., 1-1, 2-2, 3-3, 4-4, 5-5, 6-6)
1 indicates disagreement by 1 category (i.e., 1-2, 2-3, 3-4, 4-5, 5-6)
2 indicated disagreement by 2 categories (i.e., 1-3, 2-4, 3-5, 4-6)
3 indicated disagreement by 3 categories (i.e., 1-4, 2-5, 3-6)
4 indicated disagreement by 4 categories (i.e., 1-5, 2-6)
5 indicates disagreement by 5 categories (i.e, 1-6)
GENERIC VERSUS CONDITION-SPECIFIC
INSTRUMENTS
Outcomes associated with evaluating patients with low
back pain employ two broad types of measures: (a)
generic outcomes typically assessing general health that
were developed with the general population in mind, and
(b) condition-specific outcomes typically constructed by

CHAPTER 12/OUTCOMES ASSESSMENT / 135
practitioners in a particular field to more carefully assess
the outcomes thought to be relevant to the specific condition under consideration. The SF-36 is a well-known
example of a generic health outcome tool. The Roland
and Morris Disability Questionnaire and the Oswestry
Disability Index are well-known examples of conditionspecific instruments. In practice, the title of conditionspecific instr ument is a misnomer because these instruments are not meant to be linked to a particular condition
or diagnosis, but rather to a particular region of the body.
The Roland and Morris Disability Questionnaire, for
example, is a list of 24 statements associated with actions
or activities, such as, “Because of m y back or le g I stay at
home,” and, “Because of my back or leg I sit down for
most of the day.”
In practice, the differences between generic and condition-specific instruments are more in name than anything
else. The intent of using both types of instruments ma y be
to differentiate between non–back/leg and back/leg-specific symptoms. In practice, judging by the cor relations
in the .7 to .9 range between generic and condition-specific outcome instr uments, respondents either can not or
do not differentiate between the sources or causes of their
pain and symptoms. These high correlations can be worse
news for the researcher when considering the notion of
correlational attenuation. In short, this concept of attenuation allows the researcher to estimate the true correlation between two scores adjusting for the unreliability in
each measure. The logic is that the unreliability in a measure reflects random error, and random error is uncorrelated. Thus, when one cor relates two scores that are not
perfectly reliable, the resultant correlation is an underestimate of the true correlation because of the error in each
of the two measures. The formula for correcting for attenuation is given by:
r
12
=
R
12
r11r22
where R12is the disattenuated correlation between 1 and
is the attenuated correlation between 1 and 2; r11is
2; r
12
the reliability associated with instrument 1; and r
22
is the
reliability associated with instrument 2.
From this formula, as shown in Table 12-3, a consequence of a high correlation between two instruments
with moderate reliabilities is a disattenuated correlation
that approaches or even exceeds 1.0. This makes logical
sense when one considers, for example, that two instruments with reliabilities of .7 that correlate with each other
at .8, in essence, are correlated higher with each other
than they are with themselves, because reliability can be
thought of as the extent that a score correlates with itself.
Thus, from Table 12-3, if instruments 1 and 2 both have
reliabilities of .7, and the observed correlation between
instruments 1 and 2 (r
relation between instruments 1 and 2 (R
) = .8, then the disattenuated cor-
12
) is 1.14, indi-
12
cating the unlikely event that two distinct instruments
correlate higher with each other than with themselves.
This situation suggests that the two instruments are not
distinct. In a more common example, consider the situation were two instruments both have test-retest reliabilities of .8 and the intercorrelation between the two instruments is .7. In this case the disattenuated correlation is
.88, which is higher than the reliabilities for either of the
two instruments and suggests that they should not be considered distinct.
Perhaps the worst consequence of labeling instruments
as condition specific when, in fact, they are region specific
(i.e., the back or leg areas) is that clinicians have not tried
to establish stronger links between outcome measures and
specific conditions. Spratt (31), in discussing the clinical
model for health care, argues that the classic DiagnosisTreatment model should be expanded to an AssessmentDiagnosis-Treatment-Outcome (ADTO) model, which
should be viewed as an iterative cycle where: (a) Assessment leads to diagnosis; (b) diagnosis leads to treatment;
(c) treatment goals suggest relevant outcomes; and (d) out-
TABLE 12-3. The relationship between instrument reliability and disattenuated correlations among instruments
r
r
11
0.30 0.30 0.30 1.00 0.50 1.67 0.70 2.33 0.80 2.67 0.85 2.83 0.90 3.00 0.95 3.17
0.50 0.50 0.30 0.60 0.50 1.00 0.70 1.40 0.80 1.60 0.85 1.70 0.90 1.80 0.95 1.90
0.70 0.70 0.30 0.43 0.50 0.71 0.70 1.00 0.80 1.14 0.85 1.21 0.90 1.29 0.95 1.36
0.75 0.75 0.30 0.40 0.50 0.67 0.70 0.93 0.80 1.07 0.85 1.13 0.90 1.20 0.95 1.27
0.80 0.80 0.30 0.38 0.50 0.63 0.70 0.88 0.80 1.00 0.85 1.06 0.90 1.13 0.95 1.19
0.85 0.85 0.30 0.35 0.50 0.59 0.70 0.82 0.80 0.94 0.85 1.00 0.90 1.06 0.95 1.12
0.90 0.90 0.30 0.33 0.50 0.56 0.70 0.78 0.80 0.89 0.85 0.94 0.90 1.00 0.95 1.06
0.95 0.95 0.30 0.32 0.50 0.53 0.70 0.74 0.80 0.84 0.85 0.89 0.90 0.95 0.95 1.00
0.30 0.50 0.30 0.77 0.50 1.29 0.70 1.81 0.80 2.07 0.85 2.19 0.90 2.32 0.95 2.45
0.50 0.70 0.30 0.51 0.50 .085 0.70 1.18 0.80 1.35 0.85 1.44 0.90 1.52 0.95 1.61
0.70 0.80 0.30 0.40 0.50 0.67 0.70 0.94 0.80 1.07 0.85 1.14 0.90 1.20 0.95 1.27
0.75 0.90 0.30 0.37 0.50 0.61 0.70 0.85 0.80 0.97 0.85 1.03 0.90 1.10 0.95 1.16
0.80 0.95 0.30 0.34 0.50 0.57 0.70 0.80 0.80 0.92 0.85 0.98 0.90 1.03 0.95 1.09
0.85 0.50 0.30 0.46 0.50 0.77 0.70 1.07 0.80 1.23 0.85 1.30 0.90 1.38 0.95 1.46
0.90 0.50 0.30 0.45 0.50 0.75 0.70 1.04 0.80 1.19 0.85 1.27 0.90 1.34 0.95 1.42
0.95 0.50 0.30 0.44 0.50 0.73 0.70 1.02 0.80 1.16 0.85 1.23 0.90 1.31 0.95 1.38
r
22
R
12
r
12
R
12
r
12
R
12
r
12
R
12
r
12
R
12
r
12
R
12
r
12
R
12
12

136 /SECTION I/BASIC SCIENCE
comes lead to reassessment of the patient’s condition,
which suggests a potential shift in diagnosis. In this framework, an outcome might reasonably be considered an
extension of the assessment conducted to determine diagnosis. In this way, outcomes of interest are in fact condition
specific because these aspects of the patient’s health status
are evaluated, presumably, because the patient’s states and
traits (e.g., pain or symptom location, magnitude, stability,
progression, radiographic evidence of degeneration or
lesions, etc.) are fundamental to determining what is
wrong with the patient. Logically, effective treatment
results in changes in these conditions; therefore, the very
assessments and diagnostic tests done to establish a specific diagnosis seem to provide the basis for determining
the condition- or diagnosis-specific outcomes of interest
for evaluating a patient’s health status. Within this framework, a clinically relevant change in outcome could be
defined as a change in diagnosis based on changes in the
assessments of the patient’s health status originally used to
inform the diagnosis.
CLINICALLY RELEVANT DIFFERENCES
Currently, a concern among clinicians wishing to use
outcome instruments to help understand a patient’s
progress is determining the minimum clinically significant difference; that is, the smallest change in a patient’s
score from time 1 to time 2 than can be considered a clinically relevant change. Unfortunately, a real change need
not be synonymous with one that is greater than can be
expected by chance, is clinically relevant, and reflects a
meaningful amount of movement in terms of resulting
clinical decisions. A real change at the group level is a
simple statistical procedure. One compares the two
groups’ distribution of scores, determines the appropriate
statistical test based on those distributions, performs the
test, and obtains the probability that such a difference is
likely to ha ve occurred by chance. If the likelihood is low
(<1/20 or .05), the decision is typically that the difference
is too large to expect under the assumption that there was
no change and the decision is to reject the notion that
there was no change. If the change is for the better, the
conclusion is that the patient improved over time, or if
there were two treatment groups, that one g roup did better than the other.
However, at the clinical level, where the sample size is
1 (i.e., the individual patient), inferential statistical procedures are not of much practical value. In this case, ho wever, the notion of a reliable dif ference is still quantifiable
within measurement theory by estimating the standard
error of measurement defined as:
SEM = S
1 − rxx
x
where SEM is the standard er ror of measurement; SXis
the standard deviation of instrument X; and r
is the reli-
xx
ability associated with instrument X.
From this equation, it should be clear that the standard
error of measurement (SEM) for an instrument X is
smaller as the standard deviation of the scores decreases
and the reliability (range, 0 to 1) increases.
Under small sample probability theory (n = 1 surely
qualifies as small), multiplying the obtained standard
error of measurement by two and adding it to a patient’s
score would approximate a 95% confidence interval
around that score. From Table 12-1, consider the SF-36
for the PCS outcome: Mean = 30.6, SD = 7.2, and assume
test-retest reliability of .85. In this situation the SEM
computes to be 2.79. Thus, a follow-up score of 30.6 + 2
× 2.79 ≥36.18 or a score ≤25.03 (30.6 − 2 × 2.79) indicates a reliable change in the score; that is, a score different from the initial one by more than could be expected
because of unreliability. In other words, the approximate
95% confidence inter val around the initial score of 30.6
is (25.03 − 36.18). If the reliability of this score was
lower, say .75, this 95% confidence interval would
become (23.4 − 37.8), indicating a change of around 7.2
points for the difference to be considered greater than
might be expected because of lack of reliability. This
analysis assumes that the underlying scores are continuous, unimodal, and reasonably symmetrical, all fairly reasonable for the SF-36 PCS and MCS scales.
However, consider the SF-36 Social Functioning (SF)
scale: a two-item scale with each item effectively scored
as 0, 25, 50, 75, and 100 so that possible averaged scores
are 0, 12.5, 25, 37.5, 50, 62.5, 75.0, 87.5, and 100. In this
case, the assumption of continuous data clearly does not
hold. If one ignores the volition of assumptions and
applies the SEM formula to these data, assuming a generous test-retest reliability of .7, one obtains a SEM of
13.7 and a 95% confidence inter val around the mean of
39.1 of (11.7 − 66.5). However, these calculations mean
virtually nothings because the changes that can occur on
the SF scale are such that a single change of one level on
one item will move the score 12.5 points, which is close
to 1 SEM. Thus, a “reliable” difference can be obtained
by consistent changes on one or more units on each of the
two items making up the scale. This example should
highlight the need to have a large enough number of
items to allow a reasonable assumption of continuous
data and range of responses to an item to allow adequate
discrimination among responses.
In sum, the notion of SEM provides a coherent theoretical framework for establishing how much change in
an outcome is required to be comfortable that the
observed change is more than could be expected
because of an inherent error in the assessment. This represents a necessary condition for establishing whether
or not an observed change in outcome is clinically relevant. A second necessary condition, based on the logic
provided in the previous section regarding the true concept of condition-specific outcomes guided by the
ATDO clinical model is to demonstrate that the changes

CHAPTER 12/OUTCOMES ASSESSMENT / 137
observed in the outcome(s) of interest are sufficient to
change the diagnostic status of the patient. It seems that
these two necessary conditions for establishing the clinical relevance of changes in outcome represent a suff icient condition for establishing the clinical relevance of
an observed difference.
THE FUTURE OF OUTCOMES
In the early 1980s, when clinicians interested in studying and treating low back pain began to embrace patient
self-report as an important tool in evaluating patients’
health status, and by proxy treatment eff icacy, I felt that
clinicians had found the path to righteousness (i.e., finding the truth about treatment efficacy). More than 20
years later the path remains a long a winding road and the
potential is unfulfilled.
Einstein defined insanity as doing the same thing over
and over again and expecting different outcomes. Over
the course of the last 20 years three common themes
recur that seem to support Einstein’s notion of insanity as
it applies to the use of outcomes by clinicians treating
patients with back troubles.
Shorter is better. If a set of 10 items have been identified as providing a reasonably reliable and valid score,
then shortening this instrument to five items will make it
twice as good. Hopefully, the potential consequences of
shorting a scale on the SEM, as illustrated with the social
functioning subscale on the SF-36, provide compelling
reasons why shorter is not always better.
The best way to implement an outcomes program is to
start with a core set of items and expand this core as
needed. The thinking is that, in theory , it is relati vely easy
to establish a reasonably small information system, get
the technology up and running, and train the clinician
users in the system. Once this system is in operation, clinicians will demand improvements in reporting, which
will expand the core. In practice, the changes in technology (e.g., additional programming and reconciling
reports) required to expand questionnaires usually are
perceived by those who administer the system to outweigh the potential gains to the clinician users that would
result from adding to the core set. Conversely, but consistent with the aforementioned short is better bias, it is
generally perceived to be easy to remove items from a
core, even though this act also requires changes in programming and reconciling reports.
The clinical community is unab le to appreciate and the
measurement community to clarify the distinction between outcomes as applied to groups of patients as
opposed to a given patient. In general, the current state of
clinical outcomes measures is adequate in many ways
when used to evaluate treatment efficacy within a randomized controlled trial. However, these tools typically
are not sufficiently reliable and valid for use in clinical
practice, nor does reporting at the patient level typically
provide clinicians with adequate warnings about the
potential magnitude of error in the score they are using to
inform their decision making. Efforts to improve the
accuracy and precision of outcome instruments and the
score reporting to the point that these scores would be
appropriate in clinical situations would make these tools
better in the aggregate sense as well, thus indicating a
win-win situation. The catch? Clinically relevant outcomes require more rather than fewer items to achie v e the
accuracy and precision demanded at the individual patient level.
REFERENCES
1. Million R, Hall W, Haavik Nilsen K, et al. Assessment of the progress
of the back-pain patient. Spine 1982;7:204–212.
2. Roland M, Morris R. A study of the natural history of back pain. Part
I: Development of a reliable and sensiti ve measure of disability in lowback pain. Spine 1983;8:141–144.
3. Gatchel RJ. Compendium of outcome instruments for assessment and
research of spinal disorders. La Grange, IL: North American Spine
Society, 2001.
4. Beck A, Ward C, Mendelson M, et al. An inventory for measuring
depression. Arch Gen Psychiatry 1961;4:561–571.
5. Zung WWK. A self-rating depression scale. Arch Gen Psychiatry
1965;12:63–70.
6. Keller L, Butcher J. Assessment of chronic pain patients with the
MMPI-2. Minneapolis: University of Minnesota Press, 1991.
7. Derogatis L. Symptom checklist-90-R: administration, scoring and
procedures manual. Minneapolis: National Computer Systems, 1994.
8. Ware JE Jr, Sherbourne CD. The MOS 36-item short-form health survey (SF-36): I. conceptual framework and item selection. Med Care
1992;30:473–483.
9. Ware JEJ, Kosinski M, Keller SD. SF-36 Physical and mental health
summary scales: a user’s manual. Boston: The Health Institute, New
England Medical Center, 1994.
10. Ware JEJ, Snow KK, Kosinski M, et al. SF-36 Health survey manual
and interpretation guide. Boston: The Health Institute, New England
Medical Center, 1993.
11. Jensen M, Turner J, Romano J, et al. The Chronic Pain Inventory:
development and preliminary validation. Pain 1995;60:203–216.
12. Rosenstiel A, Keefe F. The use of coping strategies in low back pain
patients: relationship to patient characteristics and current adjustment.
Pain 1983;17:33–40.
13. Melzack R. The McGill Pain Questionnaire: major properties and scoring methods. Pain 1975;1:277–299.
14. Fairbank J. Revised Oswestry Disability Questionnaire (comment).
Spine 2000;25(19):2552.
15. Fairbank J. Use of Oswestry Disability Index (comment). Spine
1995;20(13):1535–1537.
16. Fairbank JC, Cooper J, Davies JB, et al. The Oswestry low back pain
disability questionnaire. Physiotherapy 1980;66(8):271–273.
17. Kerns RD , Turk DC, Rudy TE. The West Haven-Yale Multidimensional
Pain Inventory. Pain 1985;23:345–356.
18. Kopec J, Esdaile J, Abrahamowicz M, et al. The Quebec Back Pain Disability Scale: conceptualization and development. J Clin Epidemiol
1996;49:151–161.
19. Kopec J, Esdaile J, Abrahamowicz M, et al. The Quebec Back Pain Disability Scale: measurement properties. Spine 1995;20:341–352.
20. Bergner M, Bobbitt RA, Carter WB , et al. The Sickness Impact Profile:
development and final revision of a health status measure. Med Care
1981;19(8):787–805.
21. Bergner M, Bobbitt RA, P ollard WE, et al. The sickness impact profile:
validation of a health status measure. Med Care 1976;14(1):57–67.
22. Katz S. Index of Independence in Activities of Daily Living. In: Ward
MJ, Lindeman CA, eds. Instruments for measuring nursing practice
and other health care variables. Washington, DC: US Government
Printing Office, 1979:275–228; 275–280.
23. Katz S, Akpom CA. Index of ADL. Med Care 1976;14:116–118.
24. Katz S, Ford AB, Moskowitz RW, et al. Studies of illness in the aged.

138 /SECTION I/BASIC SCIENCE
The Index of ADL: a standardized measure of biological and psychosocial function. JAMA 1963;185:914–919.
25. First M, Spitzer R, Gibbon M, et al. Structured clinical interview for
DSM-IV axis I disorders: nonpatient version 2. New York: New York
State Psychiatric Institute, 1995.
26. Waddell G, McCulloch JA, Kummell E, et al. Nonorganic physical
signs in low-back pain. Spine 1980;5:117–125.
27. Spratt KF, Weinstein JN. Measuring clinical outcomes. In: Weisel S, ed.
The lumbar spine. Philadelphia: WB Saunders, 1996:1313–1338.
28. Feldt LS, Brennan RL. Reliability. In: Linn RL, ed. Educational measurement. New York: Macmillan, 1989:105–146.
29. Cronbach LJ. Test validation. In: Thorndike RL, ed. Educational measurement. Washington, DC: American Council on Education, 1971:
443–507.
30. Brennan RL. Generalizability theory. New York: Springer, 2001.
31. Spratt KF. Statistical relevance. In: Fardon DF, Garfin SR, Abitbol J-J,
et al, eds. Orthopaedic knowledge update: spine 2. Rosemont, IL: The
American Academy of Orthopaedic Surgeons, 2002:497–505.

CHAPTER 13
The Role of Outcomes and How to Integrate Them into Your Practice
Richard A. Deyo
Outcomes research became a buzzword in the 1990s,
although it seems to mean different things to different people. In general, it refers to a strategy of assessing clinical
practices according to patient outcomes, rather than to
some prespecified, often arbitrary, set of criteria for the
process of care. For e xample, we might judge the quality of
care for a patient with metastatic cancer to the lumbar
spine by his neurologic function, activities of daily living,
and survival, rather than by whether the patient received a
particular surgical implant or a particular diagnostic test.
Several important trends hav e led to the increasing interest in outcome assessment. First, medical care costs are rising much more rapidly than inflation, and the employers
and government agencies who pay the bills are asking if
they are getting their money’s worth. This seems to be an
important question, because per capita costs for medical
care in the United States are well above any other country
in the world, and yet measures of population health, such
as morbidity and mortality, are substantially worse in the
U.S. than in many other developed countries (1).
A second trend has been the observation that medical
practices vary widely from place to place, even among
very small geographic areas (2). At an international level,
the United States appears to perform roughly twice as
much back surgery as most developed countries, and five
times more back surgery than the United Kingdom (3). No
one knows which rate is optimal, but it seems unlik ely that
differences in surgical rates reflect any signif icant differences in the prevalence of back pain or disc disease. Thus,
explanations often in voke dif ferences in training, surgeons’
beliefs, public attitudes, financial incentives, imaging
strategies, and professional uncertainty. Unfortunately, the
differences do not seem to be based on evidence about
which style of practice produces the best patient outcomes.
One implication of these findings is that some medical
services may be unnecessary. Without information on
patient outcomes, however, it is impossible to know
whether, and when, this is the case. Recent data from the
Maine Lumbar Spine Study (MLSS) suggest that outcomes do vary from one geographic area to the next. In
fact, within the state of Maine, the best surgical outcomes—in terms of pain relief, functional status, disability compensation, and patient satisfaction—all occur in
the areas with the lowest surgical rates. In contrast, the
region of the state with the highest surgical rates reports
the worst surgical outcomes. The area of the state with
intermediate surgical rates has intermediate outcomes by
every measure (4). Thus, it seems clear that more is not
necessarily better. Such observations have led to greater
calls for accountability by the medical profession.
TYPES OF OUTCOME MEASURES
The Problem of Surrogate Outcomes
Traditionally, many research studies have focused on
physiologic outcomes or anatomic outcomes as indicators
of success. Examples would be whether a solid fusion is
achieved in a patient who undergoes lumbar spine fusion.
Other examples would be spinal range of motion as a
measure of improvement, spinal fluid endorphins as a
measure of possible pain relief, or surface electromyography as an indicator of “muscle spasm.” Unfortunately,
as suggested in Table 13-1, these are all intermediate or
“surrogate” outcomes, that may or ma y not reflect the end
results in which we—and our patients—are most interested. That is, these outcomes do not necessaril y correlate
well with pain relief, return to work, or improvement in
daily function. The implication is that if we are interested
in the outcomes of pain relief, return to work, and daily
functioning, we must measure them directly rather than
try to infer them from these physiologic or anatomic surrogates (5).
139

140 /SECTION I/BASIC SCIENCE
TABLE 13-1. Contrasting results for “surrogate” outcomes versus end results
Treatment (reference) Surrogate outcome End result
Lumbar fusion for Solid fusion achieved Many patients with solid fusion continue to have pain; many
degenerative discs patients without solid fusion have good pain relief
Biofeedback (28) Reduced paraspinal EMG No change in pain
Antidepressant drugs (29) No change in spinal fluid Better pain relief than placebo
Surgical discectomy (20) Recovery of motor deficits equal, Better pain relief with surgery
EMG, electomyogram.
activity
endorphins
with or without surgery
Dissociations among Outcomes
Another problem has been that, in the past, much of the
research on back problems focused only on measuring
pain, to the exclusion of other dimensions of outcome.
However, in a modern understanding of chronic pain
management, it has become apparent that both clinical
and research work may need to focus more on patient
functioning than on pain reports. Some clinical trials have
shown that it is possible to improve pain reports without
improving functional status scores, suggesting that even
though pain reports diminish, behavior may not change in
any significant way. This highlights the fact that even
among the results most important to us, there are often
dissociations among outcomes.
As one example, in the MLSS, patients treated surgically for herniated discs were compared to others treated
nonsurgically. Even after statistically adjusting for many
baseline characteristics to produce more nearly equivalent
groups, surgical outcomes were substantially better than
nonsurgical outcomes regarding pain and daily functioning. However, return to work was equivalent between the
two arms (6). If one focused only on return to work, one
might erroneously conclude that surgery was not helpful
for herniated discs. If one focused on pain and function,
however, a large advantage of surgery would be apparent.
In another example, we studied long-term outcomes in
a longitudinal cohort of primary care patients seen in a
managed care organization. After 2 years of follow-up,
even among those with the worst pain ratings (6 to 10 on
a 10-point scale), the vast majority of patients were still
working. Onl y 11.3% w ere unemplo y ed despite their high
levels of pain (Table 13-2). Conversely, among those who
reported no pain at all, 6.5% remained unemployed. In
other words, those with the least pain had a higher
employment rate, but even there, a substantial fraction
remained unemployed (7). Thus, even among the outcomes that may be most relevant to doctors and patients,
there are often dissociations, and these different dimensions of outcome must be measured independently.
This is one reason why the traditional outcome scale of
“excellent/good/fair/poor” is often inadequate. Howe and
Frymoyer noted that different definitions of these terms
result in dramatically different conclusions about the
efficacy of surgical procedures, even with the same data
in hand (8). Furthermore, as the examples given earlier
suggest, any attempt to combine pain, function, and
employment status into a single scale may be misleading,
because the different outcomes can move in different
directions, or one may improve while others do not.
In studying patients with degenerative spinal disorders,
death and cure are generally not relev ant measures of outcome. Very few patients die from back pain or disc disease. Thus, unlike the study of heart disease or cancer,
death rate is a poor outcome measure for spinal degenerative conditions. Furthermore, patients are rarely cured of
these degenerative conditions, because the degenerative
process continues even after successful surgical intervention. In most studies of both surgical and nonsurgical
treatments, a substantial proportion of patients continue
to have pain symptoms, although the symptoms may be
improved by the treatments under study. Unlike infectious diseases or the surgical treatment of appendicitis, it
is generally inappropriate to talk about “cure.”
Modern Outcome Questionnaires
All of these factors help to explain the growing interest
in questionnaire-based measures of a patient’s pain, backspecific functioning, general health status, and work disability. Indeed, these are the dimensions of outcome recommended for routine measurement by an international
working group (9) and in an update, by the participants in
TABLE 13-2. Dissociations among outcomes at 2-year
Worst pain rating (6–10): 11.3% unemployed
Best pain rating (0): 6.5% unemployed
Worst modified Roland score (37.6–100%): 11.4%
unemployed
Best modified Roland score (0): 4.3% unemployed
a
Excludes subjects keeping house, retired, or otherwise
outside the work force.
From Dionne CE, Von Korf M, Koepsell TD, et al. A comparison of pain, functional limitations, and work status indices as outcome measures in back pain research. Spine
1999;24:2339–2345, with permission.
follow-up of patients in primary care
a

CHAPTER 13/ THE ROLE OF OUTCOMES AND HOW TO INTEGRATE THEM INTO YOUR PRACTICE / 141
TABLE 13-3. Reproducibility of patient self-repor ts and physician observations
Test-retest reliability of patient Interobserver agreement
self-reports over several weeks Kappa
Health history questionnaire 0.79 Ankle reflexes nor mal 0.50
Daily function: sickness impact profile 0.87 Soft tissue tender ness 0.24
Physical function: SF-36 0.89 Lumbar spine x-ray, normal or abnormal 0.51
Pain: visual analog scale 0.94 Presence of osteophytes on x-ray 0.64
a
Data are from Deyo (5), Deyo, et al. (12), Deyo, et al. (13), Patrick (14), Pecoraro (15), McCombe (16),
with permission.
b
Kappa quantifies agreement on two measures after adjusting for chance agreements.
a “focus” issue of Spine (10). Reliable and valid measures
of each of these dimensions are av ailable because of fusing
clinical expertise with social science methodology.
A common concern about questionnaire measures is
that they are “soft data.” Physiologic measures are attractive in part because they seem “harder.” However, the
boundary between hard and soft data is indistinct at best.
Feinstein pointed out that we might judge the hardness of
data by their objectivity (physician finding versus patient
report); preservability (e.g., radiologic or histologic specimen); or by the ability to quantify (e.g., a hematocrit versus the observation that a patient is pale). However, he
concluded that the essence of “hardness” was the reproducibility of data when measured repeatedly under the
same circumstances (11). By this measure, many modern
questionnaires are at least as hard as the clinical observations with which we are more familiar. For example,
Table 13-3 shows the test-retest reliability of several selfreport questionnaires, and contrasts these with interobserver agreement on several clinical measures (12–16). In
many cases, the reproducibility of the questionnaire measures substantially exceeds that of the clinical measures.
In addition to reproducibility, these measures have
demonstrable validity, as judged by comparison with
other more objective measures of health. For example, in
a national survey, middle-aged men responded to a single
question about whether their health was excellent, very
a
b
by expert clinicians Kappa
good, good, fair, or poor. The responses to this single
question predicted 10-year mortality. In fact, the survival
rates fell in perfect order according to the initial self-evaluation of health, and ranged from about 60% survival
among those who indicated poor health to 95% survival
among those who indicated excellent health (17).
Although we know little about how subjects made their
self-evaluations at baseline, this simple self-rating obviously had important prognostic ability.
Questionnaires for studying back-related dysfunction
have been validated against a variety of clinical measures,
with reassuring results. Table 13-4 provides an example that
compares the Roland Disability Questionnaire, the Short
Form 36 (SF-36), and some “disability day” measures from
U.S. population surveys (14). While some association is
expected betw een a valid questionnaire and other measures
of health status (e.g., opioid use, or physical examination
findings), such an association would not necessarily be
expected to be a strong association. In the absence of a
“gold standard” for daily functioning, this sort of cumulative “construct validation” is generally the best we can do.
Performance Measures
Tests of patient performance have sometimes been
used to evaluate physical capacity. Such tests may
include, for example, computerized dynamometry, or
b
TABLE 13-4. Construct validation of several patient self-report measures by comparison
Outcome questionnaire Yes No Yes No Yes No
Modified Roland scale
SF-36 physical function
SF-36 pain
Days of reduced activity
a
Mean scores at baseline for sciatica patients seen in surgical practices. All differences are
significant at p ≤ .005. Higher scores on the Roland scale represent worse function; range
0–2.4.
b
All differences are significant at p ≤ .005. Higher scores on the SF-36 represent better function, range 0–100.
c
All differences are significant at p ≤ .005 except for those with and without abnormal SLR.
From Patrick DL, Deyo RA, Atlas SJ, et al. Assessing health related quality of life in patients
with sciatica. Spine 1995;20:1899–1909, with permission.
b
a
b
with other clinical phenomena
Opioids in past mo Workers’ comp leg raising
17.5 14.2 17.6 14.9 16.8 13.9
7.6 16.8 3.1 16.8 9.4 18.4
c
19.2 33.0 21.4 28.7 23.4 32.3
21.3 18.1 24.3 17.4 20.2 18.7
a
Abnormal straight

142 /SECTION I/BASIC SCIENCE
timed or measured performance on a standardized set of
tasks. While this approach has the attraction of seeming
objective, patient mood, motivation, and other factors
affect performance. Further, they require in-person evaluation (rather than by mail or telephone) and may require
special equipment, making them less practical and
affordable than questionnaire measures. How these performance measures may compare with self-report measures in terms of validity and responsiveness remains
unclear and is an area for further investigation.
OUTCOME ASSESSMENT FOR QUALITY
IMPROVEMENT
One application of outcome measurement is for
improving clinical practices. In this circumstance, patients would use outcome questionnaires in the course of
routine clinical care. The results might be used to evaluate changes in clinical practice, surgical technique,
staffing patterns, or other aspects of care.
As one example, Zucherman et al. reported their experience of measuring patient outcomes with different surgical
implants for performing spinal fusions. Unfortunately , they
found that with successive w a v es of ne w sur gical implants,
their outcomes became worse rather than better (18). Only
with the most recent implants at the time did their results
finally improve, though they remained worse than their
results using bone grafts without surgical implants. It
seems unlikely that these surgeons were less technically
skilled than other surgeons. Instead, they seemed to identify an important trend in their own practice that might otherwise have gone unnoticed. Even though their report did
not make use of some of the newer outcome instruments,
their data were sufficient for monitoring and improving
their own clinical practice.
Such quality improvement efforts may have been part
of the motivation for developing the Musculoskeletal
Outcomes Data Evaluation and Management System
(MODEMS) program by the American Academy of
Orthopaedic Surgeons (AAOS). That program was designed to have surgeons implement outcome measures
routinely in their own practices, with data submitted to a
central database. Unfortunately, the effort was not highly
successful, and this experience may point to some of the
problems with outcome assessment in routine practice.
Measuring outcomes in routine care requires real
resources. In addition to identifying and duplicating questionnaires, one must have some means of entering the data
into a computerized database, calculating scores, and reporting the results. Some practices have succeeded in doing
this at baseline for most new patients, but hav e found it difficult to obtain uniform follow-up. Some patients do not
return to clinic, some return at unexpected intervals, and
some simply do not respond to mail or telephone surveys.
Obtaining a high rate of follow-up at consistent time intervals is likely to require dedicated personnel who are able to
conduct multiple mailings or phone calls, and this adds to
the expense. These activities are not a part of routine care as
it is currently conceived, and therefore are not reimbursed
by patients or insurance companies. Thus, many practices
find it difficult to collect uniform outcome data.
In part for this reason, we have proposed a simple set
of outcome questions that might be used in routine care
and would require minimal resources (9). This is a set of
just six questions, all derived from well-validated outcome questionnaires, and covering several important
dimensions of outcome: pain, back-related functioning,
general health status, work disability, and satisfaction
with care (Table 13-5). Even this short set of measures
appears to be a substantial improvement compared with
measuring pain severity alone, or b y “excellent/good/fair/
poor” standards. The questions w ere intended to be e xamined individually, without generating an overall score, in
part to avoid obscuring the situation in which one dimension improves while others do not. A 1-week period for
measuring symptoms was suggested because it allo ws the
patient to integrate recent experience for a long enough
interval to be meaningful, but short enough to avoid
important problems with recall and to identify relatively
short-term improvements. Many of these items are
included in the lumbar cluster of the AAOS outcome
instruments used for the MODEMS program.
While these outcome measures may be useful for monitoring quality improvement over time, they should be used
with caution to compare individual physicians, clinics, or
hospitals. Such cross-system comparisons may be misleading because of important demographic or clinical differences in the patient populations served. Thus, for example, a hospital serving low-income patients with low levels
of literacy, language barriers, high levels of comorbidity,
poor health insurance, and menial if any work, is likely to
have worse outcomes than a health care system serving
well-insured, affluent, and well-educated patients. Similarly, a physician with a reputation for excellence ma y ha v e
the most difficult cases referred to him or her , while a lessskilled physician may see patients with less severe problems. The patients of the more skilled ph ysician might ha ve
worse outcomes despite higher quality care simpl y because
of a worse initial prognosis. Comparing outcomes under
these circumstances could lead to an erroneous conclusion
about the quality of care. Having measures of baseline
demographic and clinical characteristics for patients would
help to avoid such mistak en conclusions, but ma y not completely adjust for all the differences in patient populations.
Furthermore, if financial incentives are tied to outcome measures, there is a substantial risk of “gaming”
the results. This could occur if a health care system made
only nominal efforts to collect data from patients with
low literacy, limited English fluency, the most severe illness, or simply removed “outliers” from calculations or
adjusted inclusion and exclusion criteria to optimize their
apparent outcomes. The gaming that has been well
Соседние файлы в папке Библиотека им академика М.И. Перельмана
