Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_6022_Библиотеки_им_академика_М_И_Перельмана
.pdf
FIGURE8.2Therelationshipofreliabilityandvalidity.
FA,functionalassessment.
Ifthesamepatients areratedbythesamecliniciantwice,wesimilarlycan
calculate“intra-raterreliability,”thedegreetowhichheorsheagreeswithhisor
her own earlier ratings. There are two scenarios for doing this—either the
patients’performanceisvideotaped,ortheraterobservesthepatientsdoingtheir
morninggroomingroutinetwice.Inthelattercase,itisofcourseimportantthat
wearesurethesepatientshavenotchangedinthemeantime.Inbothinstances,
theclinicianshould“forget”abouthisorherratingthefirsttimearound—which
isnotthatdifficultiflargenumbersofratingsaretobemade.
Functional status is a fairly broadand abstract construct and itis unlikely
thatasingleitemsuchasGroomingcanrepresentitsentirescope.Typically,we
selectmultipleindicators(items)andcombinethemtoadequatelyoperationalize
thedefinitionof“functionalstatus”wemayhave.Useofmultipleindicatorshas
anotheradvantage—randommeasurement error in quantifying any one item is
likely offset by the random error in another item. (For that reason, the more
items there are in an instrument, the more reliable it will be, ceteris paribus,
because the chance of random error being eliminated is increased. Any
systematicerrorwillremain,however,andpracticalityissuescomeintoplayif
instruments are too long). Because each item in an FAinstrument is a repeat
measurement of the construct, like the two raters for Grooming are repeat
“measures,” we can calculate the agreement between items, as yet another
estimate of reliability. Several formulas to estimate “internal consistency
reliability” exist, of which the most frequently used is Cronbach’s coefficient
alpha. “Split-half” and “parallel forms” reliability arerelated formulas. All of
themtakevaluesbetween0.00and1.00.
Theminimalreliabilityameasureneedstohavedependsonthepurposeto
whichthedataaretobeput.Aminimumof0.90forsituationswheredecisions

onanindividual patientneed tobe made(dischargeMrs. Jones orextendher
stayanother week?) is often quoted, while 0.70 or 0.80 isa typicalminimum
required for group applications, such as in program evaluation and research.
Instruments with only “moderate” test–retest or inter-rater reliability are no
longerseen as an acceptable option (27). Longer instruments tendto be more
reliable,butthetrendistowardtheuseofshortforms,suchastheSF-12andthe
CHART-SF,ratherthanlongones,liketheirparentsSF-36(8)andCHART(6).
With better construction, new short instruments can offer reliability
approximatingthatofolderlongones.ArelevantdevelopmentisCAT,inwhich
onlythosequestionsthataretargetedtotheabilitylevelofthepersonarebeing
asked(seethesectionComputerAdaptiveTestingforfurtherdiscussion)(21).
Validity
Validitycannotbeestimated in suchasimple way asreliabilitycan, exceptin
one unusual situation: there is an existing instrument that we are certain is
perfectlyvalid and thus perfectly reliable. In that case,wecanadministerthis
existing instrument and a newly proposed one to a sample, calculate the
correlationbetweenthetwoscores,andusethatcorrelationastheestimateofthe
validityofthenewmeasure.Theissueofcourseis,ifthereisaperfectlygood
measure (a “gold standard”), why is there a need for a new one? Having a
shorterorcheaperinstrumentmaybetheonlyacceptablereason.Lesspowerful
methodstoestimatevalidityareusedinthemorecommonsituation:thereisno
existingmeasure oftheconstruct weare interested in,or the existingones are
problematicinthemselves.
In the absence of a gold standard instrument, correlations with existing
measure(s) are used to validate a new one; it is hoped that the old and new
measure will correlate strongly, providing evidence of “convergent validity.”
Sometimescorrelationswithcharacteristicsthatareseenasunrelatedtotheone
the new measure is operationalizing are also computed; the expected low
correlationisseenasevidencefor“divergentvalidity.”
“Face validity” is (in the eyes of most authorities) not a form of validity
assessment,butananswertothequestion:doestheinstrument“onthefaceofit”
measure what those completing it expect to see—does a measure of trait X
actuallyhavequestionsaboutXthatpatients/subjectsrecognizeassuch?Some
instrumentshavenoorlittlefacevalidity,butareperfectlyvalid—forinstance,
theMinnesotaMultiphasicPersonalityInventory(MMPI).However,instruments

that lack face validity may not be completed, or not be completed correctly,
becausethepatientfailstoseetheirrelevance.InthearenaofFA,facevalidityis
hardlyanissue,becausetheactivitiesthat are used as indicators of functional
ability have a fairly low level of abstraction and are recognized by
patients/clientsasrelevanttotheirlife.
The closely related term “content validity” refers to a measure actually
covering the entire width of the construct the developer is targeting. It is
generallydeterminedbyhavingexpertsdrawuplistsofnecessarycontentsfora
measure of X, or their checking the content of a draft measure against their
unwrittenexpectations.Ofcourse,thispresumesthereisacleardescriptionby
thetestdevelopersoftheconcepttheywouldliketooperationalize—itmakesno
sensecriticizingtheveracityofapaintedportraitifyoudonotknowtheperson
depicted.Therearenostandardformulasforcalculatingthisvalidityaspect.
“Predictive validity” concerns the ability of a measure to predict a future
stateor event that is inherently linked to the characteristic being measured. A
collegeentranceexaminationissaidtohavepredictivevalidityifitcanbeused
toaccuratelypredictwhoin4(5,6)yearswillgraduate.AparallelinFAwould
be the ability of a measure to predict which rehabilitation patients will be
successfully discharged home versus to a nursing home. One problem with
predictivevalidityassessmentisthefactthattherearenohardandfastrulesas
to what should bethe minimum level ofsuccess in predicting. We know that
manyfactorsaffectsuccessfulindependentliving—theaccessibilityofthehome,
familysupportavailable,theperson’sdeterminationandtoleranceforrisk, and
soforth.DoesanFAinstrumenthaveadequatepredictivevalidityifpredictions
basedonitarecorrectatleast50%ofthetime?Atleast80%?
“Knowngroupvalidity”isbasedondifferencesinscalescoresbetweentwo
groups that are known to differ in the characteristic the instrument aims to
measure. The average score of persons with SCI on a measure of physical
functioning should be lower than the average of persons with traumatic brain
injury(TBI);incaseofameasureofcognitivefunctioning,thesituationshould
bereversed. Ifthedata donot paralleltheseexpectations, itisquite likelythe
instrumentisnotmeasuringwhatwethinkitismeasuring.Alternatively,alotof
systematicerror(bias)isreflectedinthedata.Asimilarproblemasmentioned
earlieroccurswithdeterminingknowngroupvalidity:howmuchofthevariation
inthe functional statusofthe overall groupshouldbe explained bydiagnostic
category,SCIversusTBI?IfeverypersonwithSCIisknowntohaveahigher
cognitivefunctioninglevelthaneverypersonwithaTBI,thingswouldbeeasy:

thevariationexplainedshouldbe100%,andeverything lowerthan thatwould
mean less than perfect validity of a proposed FA instrument. However, the
distributions of functioning ability (both motor and cognitive) of theTBI and
SCIgroups overlap somewhat.Statingthat a good FAmeasureshouldexplain
between1%and 100%of thedifferencebetween groupsisnotveryhelpfulin
selectingordevelopinganinstrument.
“Construct validity” concerns the relationships between the measurement
dataofa(highlyabstract)constructanddataforotherconstructs.Sometimeswe
havea basisin theoryto predictthatconstructKshould bestrongly relatedto
(yet not identical with) construct L and be independent of construct M. (For
instance, “ADL ability is related to community integration, but unrelated to
political party preference.”) If the data bring this out, the measurement of K
likelyisvalid(andsimilarlytheoperationalizationsofLandM).Ifthepredicted
associationbetweenKandLisminimalorabsent,however,wedonotknowif
theproblemiswiththetheory,orwiththeoperationalizationofK,orwiththe
measurementofL.Anditisanunusualtheorythatspecifiestheexactstrengthof
the relationship between K and L, predicated on perfect measurement of the
constructsinvolved.“Strong”or“verystrong”isthebestwegetfromtheorists,
andthosearenotverygoodstartingpointsforevaluatingthelevelofvalidityof
theinstrumentsinvolved.
“Ecological validity” does not concern an instrument’s validity perse, but
the relevance of assessment data to real-life situations outside the testing
situation. Testing ambulation skills in situations that resemble the real world
(morethantheparallelbarsinthephysicaltherapygymdo) providesdatathat
are more “ecologically valid,” but the standardization of testing might suffer.
Standardizationoftestinghasalwaysbeenakeystoneofpsychometricmethods
of assessing reliability and validity of neuropsychological instruments and of
instrumentsquantifyingtheabilitiesofindividualpatientsorclients,allofwhich
are capacity measures. However, standardized environments tend to be
dissimilar from the settings where people perform self-care, communicate, do
work,andallotherthingscapturedundertheumbrellaoffunctioning.Testingin
a standardized environment almost always means in an optimal environment
(15), and the results tend therefore to be more indicative of capacity than of
performance,whichmaymeanthatthetestingdatawillnotbeverypredictiveof
real-liferoutine.
The above discussion should make clear that estimating the validity of
instruments is always less straightforward than the quantification of their

reliability.Findinghighvaluesparallelto,forexample,a0.91leveloftest-retest
reliability just does not happen; validity coefficients are almost always much
lowerbecauseallmethodsofvalidityestimationareroundabout.Themostdirect
assessment method is convergent validity, for which a minimum correlation
betweentwomeasuresofthesameconstructof0.60issuggestedasminimally
adequate (27). In practice, it is almost always necessary to use all available
methodsofestimatingvalidity,andbasedonmultiplefindings“patchtogether”
evidence supporting validity—which never will be iron-clad. Finding
encouraginglevelsofthevarioustypesofvaliditydistinguishedhere,inmultiple
studies, with patterns of correlations that make sense based on expert
knowledge, is what typically occurs. Fortunately, in the case of FA, the
specialists involved have extensive knowledge of the determinants, correlates,
intergroup differences, and so forth, of various aspects of functional status,
makingthematterofappraisingthequalityofspecificmeasureslessproblematic
thantheprecedinglistofissuesmightsuggest.
Sensitivity
It is easy to see that if an FA “measure” has just two categories, “able” and
“unable,” it lacks sensitivity: it cannot reflect fine distinctions in
capacity/performance, and it cannot be used to record minor but clinically
significant changes in the performance of an individual or group. Sensitivity
refers to the ability of an instrument to capture, across the full range of
functional ability of the subjects/patients to be measured, distinctions that are
clinicallyrelevantorsmall enoughto stillbeof importancein research.When
sensitivityisdiscussedinrelationtochangeover time(e.g.,fromadmissionto
discharge),thetermresponsivenessisfrequentlyused.
Flooreffectsand ceilingeffectsare oneissue insensitivity.Thefirstterms
refertothelowestmeasurablelevelofperformanceonanFAinstrumentbeing
higherthanthestatusoftheleastable person to be measured. All individuals
who have ability equal to the lowest measurable level or lower are lumped
togetherand given the corresponding score. Vice versa, a ceiling effectmeans
thatthehighestmeasurablelevelislowerthantheperformancelevelofatleast
someofthemoreablepatients.Itshouldbenotedthatveryoftenmeasuresare
developedforonepopulationinwhichtheyhavenofloororceilingeffects,but
thenareappliedtoanothergroup inwhichtheydo.Forinstance,theFIM was
designedtoquantifyfunctionalstatusofrehabilitationinpatients,andanypatient

whoachievesthe maximum score on dischargeprobablywasan inappropriate
admission.However,afewyearsafteronsetofincompleteparaplegia(e.g.,C4
orbelow,ASIAImpairmentScaleD),manypersonswillscoreatthemaximum
oftheFIMMotorsubscale.TheFIMwasneverdesignedtodistinguishbetween
peoplewithminimaldeficitsthatdonotaffectfunctioning,othermeremortals,
andSuperman.Thus,“lackofresponsiveness”sometimesisaproblemforwhich
theinstrumentuserisresponsible,nottheinstrumentdeveloper.
Quantification of responsiveness is not done using formulas resulting in
simple coefficients ranging from 0.0 (not responsive at all) to 1.0 (maximum
responsiveness possible). All quantification methods are mostly useful for
comparingtheresponsivenessofonemeasurewiththatofanother,allowingone
toselectthemostresponsiveone.Avarietyofindicesareused,includingeffect
sizes (the mean change between time 1 and time 2 divided by the standard
deviationattime1),thestandardizedresponsemean(themeanchangebetween
time1andtime2,dividedbythestandarddeviationofchangescores),receiver
operating characteristic (ROC) analysis, andmany others. Discussion ofthese
indicesisbeyondthescopeofthischapter;thereaderisreferredtotheextensive
literature(28,29).
A related issue is: what is the smallest change (over time) or difference
(between two cases) that a measure allows one to detect? This is not just a
questionofthemetricused(ageindaysofcoursecanreflectsmallerdifferences
thanageinyears),butalsoinvolvesmeasurementerror.Everymeasuredvalueis
an approximation, and it may be appropriate toindicatethe likely error range
involved:theweightofpatientXonadmissionwas53.7±0.2kg.Information
like that indicates that the scale that was used cannot reliably detect changes
smaller than 0.2 kg. In clinical epidemiology and related areas, the term
Minimum Detectable Difference (MDD) may be used; in psychology, the
corresponding term is reliable change, and a Reliable Change Index might be
offered.OthertermsincludeSmallestRealDifferenceandMinimumDetectable
Change.
The MDD should not be confused with a second concept: the Minimal
ClinicallyImportantDifference,orMCID(sometimesdesignatedtheMinimum
[or Minimal] Important Difference, Clinically Important Change or Minimal
Important Change). MCID refers to the smallest score difference that patients
consider to be of value. The MCID by definition should be larger than the
correspondingMDD—adifferencethatisdetectableisnotnecessarilysufficient
to reflect meaningful change in everyday functioning for the person being

assessed. Witha very sensitive scale, someone who wants to lose weight can
determineheorshehaslost0.013(±0.002)kgover4weeks,butdoesthatmean
heishappywiththeresultofhisdieting?Similarly,inaclinicaltrialofanew
methodforrehabilitatingtheupperextremities(UE)afterSCIthemeanscorefor
thetreatmentgrouponaUEmeasuresuchastheCUE-T(CUEtest[ratherthan
self-report] version) (30) may be 102, versus 99 for a usual care comparison
group(withp=.03indicatingthatthedifferenceisstatisticallysignificant),but
does that difference matter clinically—especially in light of potentially much
higherresourceexpenditureassociatedwiththenewintervention?Thus,MCID
reflectsclinicalsignificance,mostlyasseenfromtheperspectiveofpatients,and
assuch can belinked to effectsizes, numberneededto treat (NNT)and other
waysofquantifyinghowmuchdifferenceismeaningful.Thereareavarietyof
methodsfor determiningthe MCIDfora particularFAmeasure,(31) which—
becausetheyare basedon differentassumptions—tendtogivedifferentMCID
values.Wuetal.(32)giveacogentdiscussionofissuesinvolvedindetermining
MDDsandMCIDsforuseinSCIresearch,emphasizingthequestionableuseof
theseparameterswithordinalFAmeasures.
OtherMetricCharacteristics
Beyondvalidity,reliability,andsensitivity,thereareafewothercharacteristics
ofanFAmeasurethatarerelevanttoitsuseinclinical,programevaluation,and
researchapplications,mostofwhichhavetodowithpracticality:
• Language:Carefulwordingisespeciallyrelevantforself-administered
instrumentssuchastheCUE,(7)butmayalsobeanissuewithobservational
andothermeasures(33).Boththetext’sreadinglevelandatranslationintoa
languagetheuserisfamiliarwithareofconcern.InSCI,thewordsusedin
instruments(suchas“walk”intheSF-36)maybeproblematicinthatthey
arenotapplicabletothosewhouseawheelchairformobility,andevenmay
beinterpretedasreflectinginsensitivityonthepartoftheresearcher(34).
• Trainingrequired:ManyobservationalFAinstrumentsandtest-typemeasures
requiretheusertobetrained,andsometimescertified,toproducereliable
data.
• Availabilityandcosts:Somemeasuresarecopyrighted,andmaynotbe
availableatall,oronlyforaone-timeorper-usepayment.
• Timeandequipmentrequired:Measuresthattakeinordinatetimeonthepart

ofthesubjectsortheadministrator,orthatusespecialequipment,maynotbe
suitableoutsideresearchapplications.
• Alternativeversions:Availabilityofaversionforcompletionbyaproxymay
beusefulforself-reportmeasuresusedwithchildren,adultswithhigh
tetraplegia,orindividualswithcognitive-communicativedeficits.Similarly,
equivalentversionsareofuseinsituations(e.g.,psychologicaltesting)where
subjectsmay“learnthetest,”andwouldappeartogaininskillsonrepeat
administrationofasingleversion.
• Patientsafety:Iftestingtheabilitytodriveacarisdisturbedbythetest
administrator’sinterventioneverytimethereseemstobedanger,thetest
resultisnotveryrealistic.Ontheotherhand,notinterferingatallisnotto
thebenefitofthepersontested–orthetester.Simulationssuchasvirtual
reality(seelater)havebeendevelopedtomake“realistic”testingin
situationslikethispossible.However,theymaybringoutecologicalvalidity
concerns.
• Clinicalutility:Ifadministrationandscoringofameasuredoesnotaddtothe
clinician’sknowledgebase,oriftheinformationdoesnothelphimorherto
makedecisionswithrespecttoaparticularpatientoraclassofpatients,the
instrumentlacksclinicalutility.Interpretabilityofthescoresmaycontribute
toclinicalutility;availabilityofnorms(forallpersonswithSCI,or
preferablyforsubgroupswhoarecomparableintermsofage,leveland
completenessofinjury,etc.)alsoisofbenefitinsomeinstances.
USESOFFUNCTIONALASSESSMENT
Theorigin of FA(definednarrowlyas measurement ofActivityLimitations)is
foundinattemptsbyclinicianstoexpressquantitativelythedeficitspatientshad
on admission to rehabilitation, and to monitor their progress, or at least
determinedischargestatus,sothattherewassome“proof”oftheeffectiveness
of treatment beyond the patient’s simple report that he or she now could do
things that were impossible or difficult before admission. The range of
applicationsofFAinstrumentshasexpandedtremendously,sothatwenowcan
describeusesinthecareofindividualpatients,inprogramsadministration,for
reimbursement,andforresearch.
CareofIndividualPatients

Decisionsonadmissiontoinpatientandoutpatientprogramsareoftenbasedona
formal FA to see if the person has the types and degrees of deficits that the
programisqualifiedandauthorizedtotreat,eitheringeneralorforthespecific
personinquestion.The“baseline”assessmentthereforeisoftencommunicated
to the third-party payor,who mayuse it to approve program admission anda
certain duration or intensity of treatment. The pre-admission or admission
assessmentisfrequentlythebasisforaprognosis,whichiscommunicatedtothe
patient and the payor, and ideally underlies goal setting. Many rehabilitation
programsuseanFAinstrumentsuchastheFIMtosetexpectedoutcomes,either
forclassesofpatientsorforindividualpatients.Softwareapplicationshavebeen
developedtoassistcasemanagerstomakesuchpredictions;theyarefounded,in
large part, on an FA database that contains information on the admission and
discharge status of many previous patients with the same rehabilitation
diagnosis,age,gender,andcomorbidities.
Treatmentmonitoring using an FAmeasure is done in many rehabilitation
programs. Team rounds often consist of the reporting by “most responsible/
knowledgeable therapies” (nursing for bladder; speech for expression, etc.) of
the current status of the patient on the numeric items offered by the FA
instrument used in the facility. Although treatment termination decisions are
increasingly triggered by “external” criteria (e.g., a maximal length of stay
approved by a third-party payor), ideally they are founded on either the
accomplishment of goals or the plateauing of the patient in terms of overall
functional ability. In both instances, measurement of patient status should be
performedusinganinstrumentthathashighsensitivitysothat(lackof)change
canbereliablydetermined.
Unfortunately,casemanagersandmedicalinsurancecompaniesmaydemand
thatimprovementismeasuredusinginstrumentsthatdonotadequatelycapture
theextentofimprovementthatmaybeoccurring.Anexampleistheuseofthe
Motor subscale of the FIMin an SCI patient with high-leveltetraplegia. This
measure is unlikely to adequately document improvement due to the FIM’s
insensitivityto changeinthis group.A handfunctiontest, ortheQuadriplegia
IndexofFunction(QIF) maybe abetter choice(35). In theoutpatient setting,
theFIMalsomaynotcaptureimprovementsinfunctionduetoaceilingeffect.A
broad understanding of the pitfalls of available scales allows the healthcare
providertobestapplythesemeasurementsandeducateinsurers.
FA information may also be used to communicate about progress and
outcomesoftreatmentwithpersonswhoarenotpartoftherehabilitationteam.

Patientsthemselves, theirfamilymembers, referral sources,and payors havea
stronginterestinthefunctionalaspectsofthepatient’sstatus,especiallywhereit
concerns Activities and Participation. One additional use of FA is long-term
monitoringof apatient’s status. Especiallyin the caseofprogressive diseases,
suchasmultiplesclerosis,thisinformationisimportanttomakedecisionsona
needfornewtreatments,changesinpatientenvironments,andsoforth.Infact,
thisuseofFAhasledtoadesignationoffunctionalstatusinformationasasixth
(afterthestandardfourandpain)vitalsign(36).
ProgramAdministrationandEvaluation
Whetherpartofoutcomesmanagement,continuousqualityimprovement(CQI),
ortotalqualitymanagement(TQM),programevaluationaimstoassesstowhat
degree a program indeed accomplishes what it sets out to do—improve the
functional status of people with disabilities. Basic questions of program
evaluationare:dopatientschangeforthebetter(programeffectiveness),andif
so, are resources used optimally in accomplishing this (program efficiency)?
Program (self-)evaluation is required by the Commission on Accreditation of
Rehabilitation Facilities (CARF) (37), a widely recognized not-for-profit that
accreditsorganizationsandprograms.Routinelycollectedoutcomedatacanand
shouldbecommunicatedtostakeholders, includingcurrentandfuturepatients,
third-partypayors,andthelocalcommunity.TheU.S.CentersforMedicareand
MedicaidServices (CMS)has started topost comparativefunctionaloutcomes
fornursinghomesandhomehealthagenciesonitswebsite,andsimilar“report
cards” including FA information will be published in the future for other
facilitiesthatofferrehabilitationservices(38).
Althoughchangefromprogramadmissiontodischargeiscommon,itisnot
easy to offer proof that the program deserves credit. Aperson with SCI may
scorehigheronapost-testthanonapre-testforreasonsthathavenothingtodo
with the selection, timing, quality, and quantity of services received. Positive
changemaybe dueto naturalrecovery,improvedtest-takingability,andmany
otherfactors(39,40).Unfortunately,routineprogramevaluationdatatendtobe
insufficient to indicate what factors are contributing to success or failure;
additionalstudiesmaybeneededtoobtainthatinformation.
Allrehabilitationprogramsfacethesameproblemofprovingeffectiveness,
andone(partial)solutionthathasbeenfoundistocompareoutcomesbetween
programs through a “minimum data base” that includes demographics, time
Соседние файлы в папке Библиотека им академика М.И. Перельмана
