Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_6034_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
02.09.2026
Размер:
20 Мб
Скачать
CHAPTER 12/OUTCOMES ASSESSMENT / 133
(usually two) occasions under the assumptions that the first assessment does not affect, or is independent of, the second or subsequent assessments. In many cases, includ­ing clinical settings, this can be problematic because: (a) a second assessment shortly after the first assessment is not likely to be independent because the patient is likely to remember what he or she said and try for consistency rather than an accurate estimate of the current state, or (b) a second assessment a few days after the first assessment may be influenced by treatments that resulted from the initial visit. Thus, the most accurate method for estab lish­ing test-retest reliability for an instrument is to have the patient respond to the instrument on separate occasions with nothing but a short natural history intervening between the events. In this situation, which is rarely eval­uated in the clinical setting, low reliabilities, meaning large discrepancies within a patient across time, could reflect: (a) an instrument that does not have adequate psy­chometric properties, or (b) a construct under assessment (e.g., pain, function, or attitudes) that are more state-like than trait-like, meaning that the construct may not be sta­ble. The consequence of this result is that the instrument is not useful in evaluating treatment progress because the score might vary independently of treatment effective­ness on any given day. Thus, the single most important aspect of the reliability of an instrument used in the clin­ical setting is that it demonstrates stability, meaning that the score measures a condition or construct that is not amenable to quick and unpredictable changes across short time increments. In practice, internal consistency reliability estimates often are used as proxies for test­retest reliability . Because instrument construction often is guided by internal consistency estimates, the general sense is that internal consistency reliabilities are posi­tively biased estimates of test-retest reliability; therefore, the amount of error expected in a score that is most rele­vant to the clinical setting is often underestimated.
This notion of stability , of course, be gs the issue of valid­ity because it suggests that the specific constructs chosen to be of interest need to be reasonably stable in the short r un. Strictly speaking, instruments do not have validity, but the
scores they generate have validity to the extent that those scores provide information that aids in making appropriate decisions. If a test produces a score that is likely to be inter­preted as high today but lo w tomorro w, then the lack of sta­bility in that score indicates that it does not provide infor­mation useful in making appropriate decisions.
Measures of reliability include Cronbach’s alpha for internal consistency, kappa for categorical outcomes, and intraclass correlations or generalizability coefficients (30) when outcomes are reasonably continuous. It has been argued that generalizability coefficients are prefer­able to kappa coefficients under circumstances when kappa struggles, such as when the number of categories becomes large (e.g., more than four levels); when the number of raters scores being compared are greater than two (although generalized kappa statistics can be com­puted); and when the prevalence of one of the categories is low or the sample size is small.
The value of generalizability coefficients is that they can be developed in ways that isolate potential culprits or can­didates for lack of reliability , and that they allo w estimation of what some consider the most relev ant form of reliability in the clinical setting: the estimate of the reliability for a single rater on a single occasion; in other words, the relia­bility in the classical clinical situation where one clinician is evaluating a score based on a single reading. A problem with reliability coefficients, including generalizability coefficients, always has been that the meaning of the mag­nitude of the coefficient is not set. A reliability of .7 in some fields is considered good to excellent, whereas .8 is considered dismal in other fields.
To better understand the relationship between the reli­ability coefficient and the consequent differences in assessments between raters, 100,000 pairs of random nor­mal scores were generated, transformed to have exact reliabilities of .30, .40, .50, .60, .70, .80, .90, .95, and .99, again transformed to have the mean and standard devia­tion associated with selected SF-36 scores and then bro­ken down into six mutually exclusive and exhaustive cat­egories based on the tenth, 25th, 50th, 75th, and 90th percentiles, as summarized in Table 12-1. These six cate-
TABLE 12-1. Descriptive statistics amd cut points for selected SF-36 scales
Statistics PCS MCS SF PF Example: PCS
Mean 30.6 45.9 39.1 40.8 SD 7.2 14.5 25.0 22.4 Min 13.0 16.0 0 0 P10 22.0 21.5 12.5 15.0 Category 1:(13–21.9) Min–P10;10% of scores P25 26.0 35.0 25.0 25.0 Category 2:(22–25.9) P10–P25;15% of scores P50 30.0 52.5 37.5 40.0 Category 3:(26–29.9) P25–P50;25% of scores P75 35.0 58.0 50.0 55.0 Category 4:(30–34.9) P50–P75;25% of scores P90 39.0 60.0 75.0 75.0 Category 5:(35–38.9) P75–P90;15% of scores Max 58.0 63.0 100 100 Category 6:(39–58) P90–Max;10% of scores
MCS, mental component summary; PCS, physical component summary;PF , ph ysical functioning;SD,
standard deviation; SF, social functioning.
Description statistics are based on the initial visit for 376 University of Iowa patients presenting with
low back troubles.
134 /SECTION I/BASIC SCIENCE
gories for each pair of ratings with the specified reliabil­ity were then crossed and the level of agreement and dis­agreement determined based on differences in classif ica­tion group. The percentage of same and different categorization is summarized in Table 12-2 for each reli­ability level. This pattern is consistent for all outcome measures under the assumption of normality; therefore, separate tables for each outcome were unnecessary.
By way of interpreting the information provided in Tables 12-1 and 12-2, suppose a test-retest reliability of .8. In this situation, 42.63% of scores from that instru­ment are expected to remain in the score category;
45.56% are expected to be in adjacent categories; and
1.87% to differ by two categories, .915% to differ by three categories, and .026% to differ by four categories. With a test-retest reliability of .8, a miss by two or more categories is expected 11.81% of the time. Across 400 patients, just considering two category differences, this means that approximately 44 patients might have a score:
In category 1 (13 to 21.9), but a true score in category 3
(26 to 29.9); worst case difference 16.9, best case 4.1 In category 2 (22 to 25.9), but a true score in category 4
(30 to 34.9); worst case difference 13.9, best case 4.1 In category 3 (26 to 29.9), but a true score in category 5
(35 to 38.9); worst case difference 12.9, best case 5.1 In category 4 (30 to 34.9), but a true score in category 6
(39 to 58); worst case difference 28, best case 4.1
result in a differential diagnosis suggesting a less aggres­sive or immediate treatment path (e.g., w atchful waiting), the lack of reliability of the diagnostic tool clearly affects quality of care.
It should be noted that a test-retest reliability estimated based on the clinically relevant generalizability coeffi­cient of one clinician making one rating is likely to be lower than the .8 estimate used in this example. Further note that if dropping to a reliability of .7, only 36.18% of scores are expected to be in the same category if a second independent evaluation is done at the same time. Thus, the stability of many outcome measures employed at ini­tial visits, which are often used as ancillary diagnostic tools in clinical research, may have marginal value when applied to clinical practice.
In sum, the reliability of many outcome tools used to evaluate patients presenting with low back pain have lim­ited test-retest reliability evidence. The proxy internal consistency reliability estimates used are likely to be pos­itivel y biased, suggesting that the differences betw een the observed and true scores are likely to be larger than expected; therefore, clinical decisions based on these scores are likely to be based on unreliable information. This may be one explanation for the often repeated, gen­erally accepted, but not necessarily well-documented cliché that if you do not like the opinion of your clinician, get a second opinion because it will probably be dif ferent.
Of course, the reverse pattern is equally likely: An observed score in category 3 (26 to 29.9) might reflect a true score in category 1 (13 to 21.9). Thus, overtreatment or undertreatment might result if the observed score over­estimates or underestimates severity. To the extent that overestimates of symptom severity result in a differential diagnosis suggesting a more aggressive treatment path (e.g., surgery), or underestimates of symptom severity
TABLE 12-2. Magnitudes of disagreement for selected reliabilities
Number of categories of disagreement
Reliability Match 1 2 3 4 5
0.30 23.40 37.99 24.34 10.63 3.06 0.586
0.40 25.65 39.55 23.41 8.966 2.079 0.348
0.50 28.24 41.35 22.00 6.957 1.293 0.158
0.60 31.62 42.99 19.79 4.890 0.663 0.052
0.70 36.18 44.73 16.12 2.765 0.201 0.010
0.80 42.63 45.56 10.87 0.915 0.026
0.90 54.15 42.14 3.665 0.046
0.95 65.79 33.58 0.630
0.99 84.45 15.55
For categories of disagreement: Match indicates no disagreement (i.e., 1-1, 2-2, 3-3, 4-4, 5-5, 6-6) 1 indicates disagreement by 1 category (i.e., 1-2, 2-3, 3-4, 4-5, 5-6) 2 indicated disagreement by 2 categories (i.e., 1-3, 2-4, 3-5, 4-6) 3 indicated disagreement by 3 categories (i.e., 1-4, 2-5, 3-6) 4 indicated disagreement by 4 categories (i.e., 1-5, 2-6) 5 indicates disagreement by 5 categories (i.e, 1-6)
GENERIC VERSUS CONDITION-SPECIFIC INSTRUMENTS
Outcomes associated with evaluating patients with low back pain employ two broad types of measures: (a) generic outcomes typically assessing general health that were developed with the general population in mind, and (b) condition-specific outcomes typically constructed by
CHAPTER 12/OUTCOMES ASSESSMENT / 135
practitioners in a particular field to more carefully assess the outcomes thought to be relevant to the specific con­dition under consideration. The SF-36 is a well-known example of a generic health outcome tool. The Roland and Morris Disability Questionnaire and the Oswestry Disability Index are well-known examples of condition­specific instruments. In practice, the title of condition­specific instr ument is a misnomer because these instru­ments are not meant to be linked to a particular condition or diagnosis, but rather to a particular region of the body. The Roland and Morris Disability Questionnaire, for example, is a list of 24 statements associated with actions or activities, such as, “Because of m y back or le g I stay at home,” and, “Because of my back or leg I sit down for most of the day.”
In practice, the differences between generic and condi­tion-specific instruments are more in name than anything else. The intent of using both types of instruments ma y be to differentiate between non–back/leg and back/leg-spe­cific symptoms. In practice, judging by the cor relations in the .7 to .9 range between generic and condition-spe­cific outcome instr uments, respondents either can not or do not differentiate between the sources or causes of their pain and symptoms. These high correlations can be worse news for the researcher when considering the notion of correlational attenuation. In short, this concept of attenu­ation allows the researcher to estimate the true correla­tion between two scores adjusting for the unreliability in each measure. The logic is that the unreliability in a mea­sure reflects random error, and random error is uncorre­lated. Thus, when one cor relates two scores that are not perfectly reliable, the resultant correlation is an underes­timate of the true correlation because of the error in each of the two measures. The formula for correcting for atten­uation is given by:
r
12
=
R
12
r11r22
where R12is the disattenuated correlation between 1 and
is the attenuated correlation between 1 and 2; r11is
2; r
12
the reliability associated with instrument 1; and r
22
is the
reliability associated with instrument 2.
From this formula, as shown in Table 12-3, a conse­quence of a high correlation between two instruments with moderate reliabilities is a disattenuated correlation that approaches or even exceeds 1.0. This makes logical sense when one considers, for example, that two instru­ments with reliabilities of .7 that correlate with each other at .8, in essence, are correlated higher with each other than they are with themselves, because reliability can be thought of as the extent that a score correlates with itself. Thus, from Table 12-3, if instruments 1 and 2 both have reliabilities of .7, and the observed correlation between instruments 1 and 2 (r relation between instruments 1 and 2 (R
) = .8, then the disattenuated cor-
12
) is 1.14, indi-
12
cating the unlikely event that two distinct instruments correlate higher with each other than with themselves. This situation suggests that the two instruments are not distinct. In a more common example, consider the situa­tion were two instruments both have test-retest reliabili­ties of .8 and the intercorrelation between the two instru­ments is .7. In this case the disattenuated correlation is .88, which is higher than the reliabilities for either of the two instruments and suggests that they should not be con­sidered distinct.
Perhaps the worst consequence of labeling instruments as condition specific when, in fact, they are region specific (i.e., the back or leg areas) is that clinicians have not tried to establish stronger links between outcome measures and specific conditions. Spratt (31), in discussing the clinical model for health care, argues that the classic Diagnosis­Treatment model should be expanded to an Assessment­Diagnosis-Treatment-Outcome (ADTO) model, which should be viewed as an iterative cycle where: (a) Assess­ment leads to diagnosis; (b) diagnosis leads to treatment; (c) treatment goals suggest relevant outcomes; and (d) out-
TABLE 12-3. The relationship between instrument reliability and disattenuated correlations among instruments
r
r
11
0.30 0.30 0.30 1.00 0.50 1.67 0.70 2.33 0.80 2.67 0.85 2.83 0.90 3.00 0.95 3.17
0.50 0.50 0.30 0.60 0.50 1.00 0.70 1.40 0.80 1.60 0.85 1.70 0.90 1.80 0.95 1.90
0.70 0.70 0.30 0.43 0.50 0.71 0.70 1.00 0.80 1.14 0.85 1.21 0.90 1.29 0.95 1.36
0.75 0.75 0.30 0.40 0.50 0.67 0.70 0.93 0.80 1.07 0.85 1.13 0.90 1.20 0.95 1.27
0.80 0.80 0.30 0.38 0.50 0.63 0.70 0.88 0.80 1.00 0.85 1.06 0.90 1.13 0.95 1.19
0.85 0.85 0.30 0.35 0.50 0.59 0.70 0.82 0.80 0.94 0.85 1.00 0.90 1.06 0.95 1.12
0.90 0.90 0.30 0.33 0.50 0.56 0.70 0.78 0.80 0.89 0.85 0.94 0.90 1.00 0.95 1.06
0.95 0.95 0.30 0.32 0.50 0.53 0.70 0.74 0.80 0.84 0.85 0.89 0.90 0.95 0.95 1.00
0.30 0.50 0.30 0.77 0.50 1.29 0.70 1.81 0.80 2.07 0.85 2.19 0.90 2.32 0.95 2.45
0.50 0.70 0.30 0.51 0.50 .085 0.70 1.18 0.80 1.35 0.85 1.44 0.90 1.52 0.95 1.61
0.70 0.80 0.30 0.40 0.50 0.67 0.70 0.94 0.80 1.07 0.85 1.14 0.90 1.20 0.95 1.27
0.75 0.90 0.30 0.37 0.50 0.61 0.70 0.85 0.80 0.97 0.85 1.03 0.90 1.10 0.95 1.16
0.80 0.95 0.30 0.34 0.50 0.57 0.70 0.80 0.80 0.92 0.85 0.98 0.90 1.03 0.95 1.09
0.85 0.50 0.30 0.46 0.50 0.77 0.70 1.07 0.80 1.23 0.85 1.30 0.90 1.38 0.95 1.46
0.90 0.50 0.30 0.45 0.50 0.75 0.70 1.04 0.80 1.19 0.85 1.27 0.90 1.34 0.95 1.42
0.95 0.50 0.30 0.44 0.50 0.73 0.70 1.02 0.80 1.16 0.85 1.23 0.90 1.31 0.95 1.38
r
22
R
12
r
12
R
12
r
12
R
12
r
12
R
12
r
12
R
12
r
12
R
12
r
12
R
12
12
136 /SECTION I/BASIC SCIENCE
comes lead to reassessment of the patient’s condition, which suggests a potential shift in diagnosis. In this frame­work, an outcome might reasonably be considered an extension of the assessment conducted to determine diag­nosis. In this way, outcomes of interest are in fact condition specific because these aspects of the patient’s health status are evaluated, presumably, because the patient’s states and traits (e.g., pain or symptom location, magnitude, stability, progression, radiographic evidence of degeneration or lesions, etc.) are fundamental to determining what is wrong with the patient. Logically, effective treatment results in changes in these conditions; therefore, the very assessments and diagnostic tests done to establish a spe­cific diagnosis seem to provide the basis for determining the condition- or diagnosis-specific outcomes of interest for evaluating a patient’s health status. Within this frame­work, a clinically relevant change in outcome could be defined as a change in diagnosis based on changes in the assessments of the patient’s health status originally used to inform the diagnosis.
CLINICALLY RELEVANT DIFFERENCES
Currently, a concern among clinicians wishing to use outcome instruments to help understand a patient’s progress is determining the minimum clinically signifi­cant difference; that is, the smallest change in a patient’s score from time 1 to time 2 than can be considered a clin­ically relevant change. Unfortunately, a real change need not be synonymous with one that is greater than can be expected by chance, is clinically relevant, and reflects a meaningful amount of movement in terms of resulting clinical decisions. A real change at the group level is a simple statistical procedure. One compares the two groups’ distribution of scores, determines the appropriate statistical test based on those distributions, performs the test, and obtains the probability that such a difference is likely to ha ve occurred by chance. If the likelihood is low (<1/20 or .05), the decision is typically that the difference is too large to expect under the assumption that there was no change and the decision is to reject the notion that there was no change. If the change is for the better, the conclusion is that the patient improved over time, or if there were two treatment groups, that one g roup did bet­ter than the other.
However, at the clinical level, where the sample size is 1 (i.e., the individual patient), inferential statistical pro­cedures are not of much practical value. In this case, ho w­ever, the notion of a reliable dif ference is still quantifiable within measurement theory by estimating the standard error of measurement defined as:
SEM = S
1 − rxx
x
where SEM is the standard er ror of measurement; SXis the standard deviation of instrument X; and r
is the reli-
xx
ability associated with instrument X.
From this equation, it should be clear that the standard error of measurement (SEM) for an instrument X is smaller as the standard deviation of the scores decreases and the reliability (range, 0 to 1) increases.
Under small sample probability theory (n = 1 surely qualifies as small), multiplying the obtained standard error of measurement by two and adding it to a patient’s score would approximate a 95% confidence interval around that score. From Table 12-1, consider the SF-36 for the PCS outcome: Mean = 30.6, SD = 7.2, and assume test-retest reliability of .85. In this situation the SEM computes to be 2.79. Thus, a follow-up score of 30.6 + 2 × 2.79 36.18 or a score 25.03 (30.6 2 × 2.79) indi­cates a reliable change in the score; that is, a score differ­ent from the initial one by more than could be expected because of unreliability. In other words, the approximate 95% confidence inter val around the initial score of 30.6 is (25.03 36.18). If the reliability of this score was lower, say .75, this 95% confidence interval would become (23.4 37.8), indicating a change of around 7.2 points for the difference to be considered greater than might be expected because of lack of reliability. This analysis assumes that the underlying scores are continu­ous, unimodal, and reasonably symmetrical, all fairly rea­sonable for the SF-36 PCS and MCS scales.
However, consider the SF-36 Social Functioning (SF) scale: a two-item scale with each item effectively scored as 0, 25, 50, 75, and 100 so that possible averaged scores are 0, 12.5, 25, 37.5, 50, 62.5, 75.0, 87.5, and 100. In this case, the assumption of continuous data clearly does not hold. If one ignores the volition of assumptions and applies the SEM formula to these data, assuming a gen­erous test-retest reliability of .7, one obtains a SEM of
13.7 and a 95% confidence inter val around the mean of
39.1 of (11.7 66.5). However, these calculations mean virtually nothings because the changes that can occur on the SF scale are such that a single change of one level on one item will move the score 12.5 points, which is close to 1 SEM. Thus, a “reliable” difference can be obtained by consistent changes on one or more units on each of the two items making up the scale. This example should highlight the need to have a large enough number of items to allow a reasonable assumption of continuous data and range of responses to an item to allow adequate discrimination among responses.
In sum, the notion of SEM provides a coherent theo­retical framework for establishing how much change in an outcome is required to be comfortable that the observed change is more than could be expected because of an inherent error in the assessment. This rep­resents a necessary condition for establishing whether or not an observed change in outcome is clinically rele­vant. A second necessary condition, based on the logic provided in the previous section regarding the true con­cept of condition-specific outcomes guided by the ATDO clinical model is to demonstrate that the changes
CHAPTER 12/OUTCOMES ASSESSMENT / 137
observed in the outcome(s) of interest are sufficient to change the diagnostic status of the patient. It seems that these two necessary conditions for establishing the clin­ical relevance of changes in outcome represent a suff i­cient condition for establishing the clinical relevance of an observed difference.
THE FUTURE OF OUTCOMES
In the early 1980s, when clinicians interested in study­ing and treating low back pain began to embrace patient self-report as an important tool in evaluating patients’ health status, and by proxy treatment eff icacy, I felt that clinicians had found the path to righteousness (i.e., find­ing the truth about treatment efficacy). More than 20 years later the path remains a long a winding road and the potential is unfulfilled.
Einstein defined insanity as doing the same thing over and over again and expecting different outcomes. Over the course of the last 20 years three common themes recur that seem to support Einstein’s notion of insanity as it applies to the use of outcomes by clinicians treating patients with back troubles.
Shorter is better. If a set of 10 items have been identi­fied as providing a reasonably reliable and valid score, then shortening this instrument to five items will make it twice as good. Hopefully, the potential consequences of shorting a scale on the SEM, as illustrated with the social functioning subscale on the SF-36, provide compelling reasons why shorter is not always better.
The best way to implement an outcomes program is to start with a core set of items and expand this core as needed. The thinking is that, in theory , it is relati vely easy
to establish a reasonably small information system, get the technology up and running, and train the clinician users in the system. Once this system is in operation, clin­icians will demand improvements in reporting, which will expand the core. In practice, the changes in tech­nology (e.g., additional programming and reconciling reports) required to expand questionnaires usually are perceived by those who administer the system to out­weigh the potential gains to the clinician users that would result from adding to the core set. Conversely, but con­sistent with the aforementioned short is better bias, it is generally perceived to be easy to remove items from a core, even though this act also requires changes in pro­gramming and reconciling reports.
The clinical community is unab le to appreciate and the measurement community to clarify the distinction be­tween outcomes as applied to groups of patients as opposed to a given patient. In general, the current state of
clinical outcomes measures is adequate in many ways when used to evaluate treatment efficacy within a ran­domized controlled trial. However, these tools typically are not sufficiently reliable and valid for use in clinical practice, nor does reporting at the patient level typically
provide clinicians with adequate warnings about the potential magnitude of error in the score they are using to inform their decision making. Efforts to improve the accuracy and precision of outcome instruments and the score reporting to the point that these scores would be appropriate in clinical situations would make these tools better in the aggregate sense as well, thus indicating a win-win situation. The catch? Clinically relevant out­comes require more rather than fewer items to achie v e the accuracy and precision demanded at the individual pa­tient level.
REFERENCES
1. Million R, Hall W, Haavik Nilsen K, et al. Assessment of the progress of the back-pain patient. Spine 1982;7:204–212.
2. Roland M, Morris R. A study of the natural history of back pain. Part I: Development of a reliable and sensiti ve measure of disability in low­back pain. Spine 1983;8:141–144.
3. Gatchel RJ. Compendium of outcome instruments for assessment and research of spinal disorders. La Grange, IL: North American Spine Society, 2001.
4. Beck A, Ward C, Mendelson M, et al. An inventory for measuring depression. Arch Gen Psychiatry 1961;4:561–571.
5. Zung WWK. A self-rating depression scale. Arch Gen Psychiatry 1965;12:63–70.
6. Keller L, Butcher J. Assessment of chronic pain patients with the MMPI-2. Minneapolis: University of Minnesota Press, 1991.
7. Derogatis L. Symptom checklist-90-R: administration, scoring and procedures manual. Minneapolis: National Computer Systems, 1994.
8. Ware JE Jr, Sherbourne CD. The MOS 36-item short-form health sur­vey (SF-36): I. conceptual framework and item selection. Med Care 1992;30:473–483.
9. Ware JEJ, Kosinski M, Keller SD. SF-36 Physical and mental health summary scales: a user’s manual. Boston: The Health Institute, New England Medical Center, 1994.
10. Ware JEJ, Snow KK, Kosinski M, et al. SF-36 Health survey manual and interpretation guide. Boston: The Health Institute, New England Medical Center, 1993.
11. Jensen M, Turner J, Romano J, et al. The Chronic Pain Inventory: development and preliminary validation. Pain 1995;60:203–216.
12. Rosenstiel A, Keefe F. The use of coping strategies in low back pain patients: relationship to patient characteristics and current adjustment. Pain 1983;17:33–40.
13. Melzack R. The McGill Pain Questionnaire: major properties and scor­ing methods. Pain 1975;1:277–299.
14. Fairbank J. Revised Oswestry Disability Questionnaire (comment). Spine 2000;25(19):2552.
15. Fairbank J. Use of Oswestry Disability Index (comment). Spine 1995;20(13):1535–1537.
16. Fairbank JC, Cooper J, Davies JB, et al. The Oswestry low back pain disability questionnaire. Physiotherapy 1980;66(8):271–273.
17. Kerns RD , Turk DC, Rudy TE. The West Haven-Yale Multidimensional Pain Inventory. Pain 1985;23:345–356.
18. Kopec J, Esdaile J, Abrahamowicz M, et al. The Quebec Back Pain Dis­ability Scale: conceptualization and development. J Clin Epidemiol 1996;49:151–161.
19. Kopec J, Esdaile J, Abrahamowicz M, et al. The Quebec Back Pain Dis­ability Scale: measurement properties. Spine 1995;20:341–352.
20. Bergner M, Bobbitt RA, Carter WB , et al. The Sickness Impact Profile: development and final revision of a health status measure. Med Care 1981;19(8):787–805.
21. Bergner M, Bobbitt RA, P ollard WE, et al. The sickness impact profile: validation of a health status measure. Med Care 1976;14(1):57–67.
22. Katz S. Index of Independence in Activities of Daily Living. In: Ward MJ, Lindeman CA, eds. Instruments for measuring nursing practice and other health care variables. Washington, DC: US Government Printing Office, 1979:275–228; 275–280.
23. Katz S, Akpom CA. Index of ADL. Med Care 1976;14:116–118.
24. Katz S, Ford AB, Moskowitz RW, et al. Studies of illness in the aged.
138 /SECTION I/BASIC SCIENCE
The Index of ADL: a standardized measure of biological and psy­chosocial function. JAMA 1963;185:914–919.
25. First M, Spitzer R, Gibbon M, et al. Structured clinical interview for DSM-IV axis I disorders: nonpatient version 2. New York: New York State Psychiatric Institute, 1995.
26. Waddell G, McCulloch JA, Kummell E, et al. Nonorganic physical signs in low-back pain. Spine 1980;5:117–125.
27. Spratt KF, Weinstein JN. Measuring clinical outcomes. In: Weisel S, ed. The lumbar spine. Philadelphia: WB Saunders, 1996:1313–1338.
28. Feldt LS, Brennan RL. Reliability. In: Linn RL, ed. Educational mea­surement. New York: Macmillan, 1989:105–146.
29. Cronbach LJ. Test validation. In: Thorndike RL, ed. Educational mea­surement. Washington, DC: American Council on Education, 1971: 443–507.
30. Brennan RL. Generalizability theory. New York: Springer, 2001.
31. Spratt KF. Statistical relevance. In: Fardon DF, Garfin SR, Abitbol J-J, et al, eds. Orthopaedic knowledge update: spine 2. Rosemont, IL: The American Academy of Orthopaedic Surgeons, 2002:497–505.
CHAPTER 13

The Role of Outcomes and How to Integrate Them into Your Practice

Richard A. Deyo
Outcomes research became a buzzword in the 1990s, although it seems to mean different things to different peo­ple. In general, it refers to a strategy of assessing clinical practices according to patient outcomes, rather than to some prespecified, often arbitrary, set of criteria for the process of care. For e xample, we might judge the quality of care for a patient with metastatic cancer to the lumbar spine by his neurologic function, activities of daily living, and survival, rather than by whether the patient received a particular surgical implant or a particular diagnostic test.
Several important trends hav e led to the increasing inter­est in outcome assessment. First, medical care costs are ris­ing much more rapidly than inflation, and the employers and government agencies who pay the bills are asking if they are getting their money’s worth. This seems to be an important question, because per capita costs for medical care in the United States are well above any other country in the world, and yet measures of population health, such as morbidity and mortality, are substantially worse in the U.S. than in many other developed countries (1).
A second trend has been the observation that medical practices vary widely from place to place, even among very small geographic areas (2). At an international level, the United States appears to perform roughly twice as much back surgery as most developed countries, and five times more back surgery than the United Kingdom (3). No one knows which rate is optimal, but it seems unlik ely that differences in surgical rates reflect any signif icant differ­ences in the prevalence of back pain or disc disease. Thus, explanations often in voke dif ferences in training, surgeons’ beliefs, public attitudes, financial incentives, imaging strategies, and professional uncertainty. Unfortunately, the differences do not seem to be based on evidence about which style of practice produces the best patient outcomes.
One implication of these findings is that some medical services may be unnecessary. Without information on
patient outcomes, however, it is impossible to know whether, and when, this is the case. Recent data from the Maine Lumbar Spine Study (MLSS) suggest that out­comes do vary from one geographic area to the next. In fact, within the state of Maine, the best surgical out­comes—in terms of pain relief, functional status, disabil­ity compensation, and patient satisfaction—all occur in the areas with the lowest surgical rates. In contrast, the region of the state with the highest surgical rates reports the worst surgical outcomes. The area of the state with intermediate surgical rates has intermediate outcomes by every measure (4). Thus, it seems clear that more is not necessarily better. Such observations have led to greater calls for accountability by the medical profession.
TYPES OF OUTCOME MEASURES The Problem of Surrogate Outcomes
Traditionally, many research studies have focused on physiologic outcomes or anatomic outcomes as indicators of success. Examples would be whether a solid fusion is achieved in a patient who undergoes lumbar spine fusion. Other examples would be spinal range of motion as a measure of improvement, spinal fluid endorphins as a measure of possible pain relief, or surface electromyog­raphy as an indicator of “muscle spasm.” Unfortunately, as suggested in Table 13-1, these are all intermediate or “surrogate” outcomes, that may or ma y not reflect the end results in which we—and our patients—are most inter­ested. That is, these outcomes do not necessaril y correlate well with pain relief, return to work, or improvement in daily function. The implication is that if we are interested in the outcomes of pain relief, return to work, and daily functioning, we must measure them directly rather than try to infer them from these physiologic or anatomic sur­rogates (5).
139
140 /SECTION I/BASIC SCIENCE
TABLE 13-1. Contrasting results for “surrogate” outcomes versus end results
Treatment (reference) Surrogate outcome End result
Lumbar fusion for Solid fusion achieved Many patients with solid fusion continue to have pain; many
degenerative discs patients without solid fusion have good pain relief
Biofeedback (28) Reduced paraspinal EMG No change in pain
Antidepressant drugs (29) No change in spinal fluid Better pain relief than placebo
Surgical discectomy (20) Recovery of motor deficits equal, Better pain relief with surgery
EMG, electomyogram.
activity
endorphins
with or without surgery
Dissociations among Outcomes
Another problem has been that, in the past, much of the research on back problems focused only on measuring pain, to the exclusion of other dimensions of outcome. However, in a modern understanding of chronic pain management, it has become apparent that both clinical and research work may need to focus more on patient functioning than on pain reports. Some clinical trials have shown that it is possible to improve pain reports without improving functional status scores, suggesting that even though pain reports diminish, behavior may not change in any significant way. This highlights the fact that even among the results most important to us, there are often dissociations among outcomes.
As one example, in the MLSS, patients treated surgi­cally for herniated discs were compared to others treated nonsurgically. Even after statistically adjusting for many baseline characteristics to produce more nearly equivalent groups, surgical outcomes were substantially better than nonsurgical outcomes regarding pain and daily function­ing. However, return to work was equivalent between the two arms (6). If one focused only on return to work, one might erroneously conclude that surgery was not helpful for herniated discs. If one focused on pain and function, however, a large advantage of surgery would be apparent.
In another example, we studied long-term outcomes in a longitudinal cohort of primary care patients seen in a managed care organization. After 2 years of follow-up, even among those with the worst pain ratings (6 to 10 on a 10-point scale), the vast majority of patients were still working. Onl y 11.3% w ere unemplo y ed despite their high levels of pain (Table 13-2). Conversely, among those who reported no pain at all, 6.5% remained unemployed. In other words, those with the least pain had a higher employment rate, but even there, a substantial fraction remained unemployed (7). Thus, even among the out­comes that may be most relevant to doctors and patients, there are often dissociations, and these different dimen­sions of outcome must be measured independently.
This is one reason why the traditional outcome scale of “excellent/good/fair/poor” is often inadequate. Howe and Frymoyer noted that different definitions of these terms result in dramatically different conclusions about the
efficacy of surgical procedures, even with the same data in hand (8). Furthermore, as the examples given earlier suggest, any attempt to combine pain, function, and employment status into a single scale may be misleading, because the different outcomes can move in different directions, or one may improve while others do not.
In studying patients with degenerative spinal disorders, death and cure are generally not relev ant measures of out­come. Very few patients die from back pain or disc dis­ease. Thus, unlike the study of heart disease or cancer, death rate is a poor outcome measure for spinal degener­ative conditions. Furthermore, patients are rarely cured of these degenerative conditions, because the degenerative process continues even after successful surgical interven­tion. In most studies of both surgical and nonsurgical treatments, a substantial proportion of patients continue to have pain symptoms, although the symptoms may be improved by the treatments under study. Unlike infec­tious diseases or the surgical treatment of appendicitis, it is generally inappropriate to talk about “cure.”
Modern Outcome Questionnaires
All of these factors help to explain the growing interest in questionnaire-based measures of a patient’s pain, back­specific functioning, general health status, and work dis­ability. Indeed, these are the dimensions of outcome rec­ommended for routine measurement by an international working group (9) and in an update, by the participants in
TABLE 13-2. Dissociations among outcomes at 2-year
Worst pain rating (6–10): 11.3% unemployed Best pain rating (0): 6.5% unemployed Worst modified Roland score (37.6–100%): 11.4%
unemployed
Best modified Roland score (0): 4.3% unemployed
a
Excludes subjects keeping house, retired, or otherwise
outside the work force.
From Dionne CE, Von Korf M, Koepsell TD, et al. A com­parison of pain, functional limitations, and work status in­dices as outcome measures in back pain research. Spine 1999;24:2339–2345, with permission.
follow-up of patients in primary care
a
CHAPTER 13/ THE ROLE OF OUTCOMES AND HOW TO INTEGRATE THEM INTO YOUR PRACTICE / 141
TABLE 13-3. Reproducibility of patient self-repor ts and physician observations
Test-retest reliability of patient Interobserver agreement
self-reports over several weeks Kappa
Health history questionnaire 0.79 Ankle reflexes nor mal 0.50 Daily function: sickness impact profile 0.87 Soft tissue tender ness 0.24 Physical function: SF-36 0.89 Lumbar spine x-ray, normal or abnormal 0.51 Pain: visual analog scale 0.94 Presence of osteophytes on x-ray 0.64
a
Data are from Deyo (5), Deyo, et al. (12), Deyo, et al. (13), Patrick (14), Pecoraro (15), McCombe (16),
with permission.
b
Kappa quantifies agreement on two measures after adjusting for chance agreements.
a “focus” issue of Spine (10). Reliable and valid measures of each of these dimensions are av ailable because of fusing clinical expertise with social science methodology.
A common concern about questionnaire measures is that they are “soft data.” Physiologic measures are attrac­tive in part because they seem “harder.” However, the boundary between hard and soft data is indistinct at best. Feinstein pointed out that we might judge the hardness of data by their objectivity (physician finding versus patient report); preservability (e.g., radiologic or histologic spec­imen); or by the ability to quantify (e.g., a hematocrit ver­sus the observation that a patient is pale). However, he concluded that the essence of “hardness” was the repro­ducibility of data when measured repeatedly under the same circumstances (11). By this measure, many modern questionnaires are at least as hard as the clinical observa­tions with which we are more familiar. For example, Table 13-3 shows the test-retest reliability of several self­report questionnaires, and contrasts these with interob­server agreement on several clinical measures (12–16). In many cases, the reproducibility of the questionnaire mea­sures substantially exceeds that of the clinical measures.
In addition to reproducibility, these measures have demonstrable validity, as judged by comparison with other more objective measures of health. For example, in a national survey, middle-aged men responded to a single question about whether their health was excellent, very
a
b
by expert clinicians Kappa
good, good, fair, or poor. The responses to this single question predicted 10-year mortality. In fact, the survival rates fell in perfect order according to the initial self-eval­uation of health, and ranged from about 60% survival among those who indicated poor health to 95% survival among those who indicated excellent health (17). Although we know little about how subjects made their self-evaluations at baseline, this simple self-rating obvi­ously had important prognostic ability.
Questionnaires for studying back-related dysfunction have been validated against a variety of clinical measures, with reassuring results. Table 13-4 provides an example that compares the Roland Disability Questionnaire, the Short Form 36 (SF-36), and some “disability day” measures from U.S. population surveys (14). While some association is expected betw een a valid questionnaire and other measures of health status (e.g., opioid use, or physical examination findings), such an association would not necessarily be expected to be a strong association. In the absence of a “gold standard” for daily functioning, this sort of cumula­tive “construct validation” is generally the best we can do.
Performance Measures
Tests of patient performance have sometimes been used to evaluate physical capacity. Such tests may include, for example, computerized dynamometry, or
b
TABLE 13-4. Construct validation of several patient self-report measures by comparison
Outcome questionnaire Yes No Yes No Yes No Modified Roland scale
SF-36 physical function SF-36 pain Days of reduced activity
a
Mean scores at baseline for sciatica patients seen in surgical practices. All differences are significant at p .005. Higher scores on the Roland scale represent worse function; range 0–2.4.
b
All differences are significant at p .005. Higher scores on the SF-36 represent better func­tion, range 0–100.
c
All differences are significant at p .005 except for those with and without abnormal SLR.
From Patrick DL, Deyo RA, Atlas SJ, et al. Assessing health related quality of life in patients
with sciatica. Spine 1995;20:1899–1909, with permission.
b
a
b
with other clinical phenomena
Opioids in past mo Workers’ comp leg raising
17.5 14.2 17.6 14.9 16.8 13.9
7.6 16.8 3.1 16.8 9.4 18.4
c
19.2 33.0 21.4 28.7 23.4 32.3
21.3 18.1 24.3 17.4 20.2 18.7
a
Abnormal straight
142 /SECTION I/BASIC SCIENCE
timed or measured performance on a standardized set of tasks. While this approach has the attraction of seeming objective, patient mood, motivation, and other factors affect performance. Further, they require in-person eval­uation (rather than by mail or telephone) and may require special equipment, making them less practical and affordable than questionnaire measures. How these per­formance measures may compare with self-report mea­sures in terms of validity and responsiveness remains unclear and is an area for further investigation.
OUTCOME ASSESSMENT FOR QUALITY IMPROVEMENT
One application of outcome measurement is for improving clinical practices. In this circumstance, pa­tients would use outcome questionnaires in the course of routine clinical care. The results might be used to evalu­ate changes in clinical practice, surgical technique, staffing patterns, or other aspects of care.
As one example, Zucherman et al. reported their experi­ence of measuring patient outcomes with different surgical implants for performing spinal fusions. Unfortunately , they found that with successive w a v es of ne w sur gical implants, their outcomes became worse rather than better (18). Only with the most recent implants at the time did their results finally improve, though they remained worse than their results using bone grafts without surgical implants. It seems unlikely that these surgeons were less technically skilled than other surgeons. Instead, they seemed to iden­tify an important trend in their own practice that might oth­erwise have gone unnoticed. Even though their report did not make use of some of the newer outcome instruments, their data were sufficient for monitoring and improving their own clinical practice.
Such quality improvement efforts may have been part of the motivation for developing the Musculoskeletal Outcomes Data Evaluation and Management System (MODEMS) program by the American Academy of Orthopaedic Surgeons (AAOS). That program was de­signed to have surgeons implement outcome measures routinely in their own practices, with data submitted to a central database. Unfortunately, the effort was not highly successful, and this experience may point to some of the problems with outcome assessment in routine practice.
Measuring outcomes in routine care requires real resources. In addition to identifying and duplicating ques­tionnaires, one must have some means of entering the data into a computerized database, calculating scores, and re­porting the results. Some practices have succeeded in doing this at baseline for most new patients, but hav e found it dif­ficult to obtain uniform follow-up. Some patients do not return to clinic, some return at unexpected intervals, and some simply do not respond to mail or telephone surveys. Obtaining a high rate of follow-up at consistent time inter­vals is likely to require dedicated personnel who are able to
conduct multiple mailings or phone calls, and this adds to the expense. These activities are not a part of routine care as it is currently conceived, and therefore are not reimbursed by patients or insurance companies. Thus, many practices find it difficult to collect uniform outcome data.
In part for this reason, we have proposed a simple set of outcome questions that might be used in routine care and would require minimal resources (9). This is a set of just six questions, all derived from well-validated out­come questionnaires, and covering several important dimensions of outcome: pain, back-related functioning, general health status, work disability, and satisfaction with care (Table 13-5). Even this short set of measures appears to be a substantial improvement compared with measuring pain severity alone, or b y “excellent/good/fair/ poor” standards. The questions w ere intended to be e xam­ined individually, without generating an overall score, in part to avoid obscuring the situation in which one dimen­sion improves while others do not. A 1-week period for measuring symptoms was suggested because it allo ws the patient to integrate recent experience for a long enough interval to be meaningful, but short enough to avoid important problems with recall and to identify relatively short-term improvements. Many of these items are included in the lumbar cluster of the AAOS outcome instruments used for the MODEMS program.
While these outcome measures may be useful for moni­toring quality improvement over time, they should be used with caution to compare individual physicians, clinics, or hospitals. Such cross-system comparisons may be mis­leading because of important demographic or clinical dif­ferences in the patient populations served. Thus, for exam­ple, a hospital serving low-income patients with low levels of literacy, language barriers, high levels of comorbidity, poor health insurance, and menial if any work, is likely to have worse outcomes than a health care system serving well-insured, affluent, and well-educated patients. Simi­larly, a physician with a reputation for excellence ma y ha v e the most difficult cases referred to him or her , while a less­skilled physician may see patients with less severe prob­lems. The patients of the more skilled ph ysician might ha ve worse outcomes despite higher quality care simpl y because of a worse initial prognosis. Comparing outcomes under these circumstances could lead to an erroneous conclusion about the quality of care. Having measures of baseline demographic and clinical characteristics for patients would help to avoid such mistak en conclusions, but ma y not com­pletely adjust for all the differences in patient populations.
Furthermore, if financial incentives are tied to out­come measures, there is a substantial risk of “gaming” the results. This could occur if a health care system made only nominal efforts to collect data from patients with low literacy, limited English fluency, the most severe ill­ness, or simply removed “outliers” from calculations or adjusted inclusion and exclusion criteria to optimize their apparent outcomes. The gaming that has been well