Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_6022_Библиотеки_им_академика_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
31.08.2026
Размер:
36 Мб
Скачать
FIGURE8.2Therelationshipofreliabilityandvalidity.
FA,functionalassessment.
Ifthesamepatients areratedbythesamecliniciantwice,wesimilarlycan calculate“intra-raterreliability,”thedegreetowhichheorsheagreeswithhisor her own earlier ratings. There are two scenarios for doing this—either the patients’performanceisvideotaped,ortheraterobservesthepatientsdoingtheir morninggroomingroutinetwice.Inthelattercase,itisofcourseimportantthat wearesurethesepatientshavenotchangedinthemeantime.Inbothinstances, theclinicianshould“forget”abouthisorherratingthefirsttimearound—which isnotthatdifficultiflargenumbersofratingsaretobemade.
Functional status is a fairly broadand abstract construct and itis unlikely thatasingleitemsuchasGroomingcanrepresentitsentirescope.Typically,we selectmultipleindicators(items)andcombinethemtoadequatelyoperationalize thedefinitionof“functionalstatus”wemayhave.Useofmultipleindicatorshas anotheradvantage—randommeasurement error in quantifying any one item is likely offset by the random error in another item. (For that reason, the more items there are in an instrument, the more reliable it will be, ceteris paribus, because the chance of random error being eliminated is increased. Any systematicerrorwillremain,however,andpracticalityissuescomeintoplayif instruments are too long). Because each item in an FAinstrument is a repeat measurement of the construct, like the two raters for Grooming are repeat “measures,” we can calculate the agreement between items, as yet another estimate of reliability. Several formulas to estimate “internal consistency reliability” exist, of which the most frequently used is Cronbach’s coefficient alpha. “Split-half” and “parallel forms” reliability arerelated formulas. All of themtakevaluesbetween0.00and1.00.
Theminimalreliabilityameasureneedstohavedependsonthepurposeto whichthedataaretobeput.Aminimumof0.90forsituationswheredecisions
onanindividual patientneed tobe made(dischargeMrs. Jones orextendher stayanother week?) is often quoted, while 0.70 or 0.80 isa typicalminimum required for group applications, such as in program evaluation and research. Instruments with only “moderate” test–retest or inter-rater reliability are no longerseen as an acceptable option (27). Longer instruments tendto be more reliable,butthetrendistowardtheuseofshortforms,suchastheSF-12andthe CHART-SF,ratherthanlongones,liketheirparentsSF-36(8)andCHART(6). With better construction, new short instruments can offer reliability approximatingthatofolderlongones.ArelevantdevelopmentisCAT,inwhich onlythosequestionsthataretargetedtotheabilitylevelofthepersonarebeing asked(seethesectionComputerAdaptiveTestingforfurtherdiscussion)(21).
Validity
Validitycannotbeestimated in suchasimple way asreliabilitycan, exceptin one unusual situation: there is an existing instrument that we are certain is perfectlyvalid and thus perfectly reliable. In that case,wecanadministerthis existing instrument and a newly proposed one to a sample, calculate the correlationbetweenthetwoscores,andusethatcorrelationastheestimateofthe validityofthenewmeasure.Theissueofcourseis,ifthereisaperfectlygood measure (a “gold standard”), why is there a need for a new one? Having a shorterorcheaperinstrumentmaybetheonlyacceptablereason.Lesspowerful methodstoestimatevalidityareusedinthemorecommonsituation:thereisno existingmeasure oftheconstruct weare interested in,or the existingones are problematicinthemselves.
In the absence of a gold standard instrument, correlations with existing measure(s) are used to validate a new one; it is hoped that the old and new measure will correlate strongly, providing evidence of “convergent validity.” Sometimescorrelationswithcharacteristicsthatareseenasunrelatedtotheone the new measure is operationalizing are also computed; the expected low correlationisseenasevidencefor“divergentvalidity.”
“Face validity” is (in the eyes of most authorities) not a form of validity assessment,butananswertothequestion:doestheinstrument“onthefaceofit” measure what those completing it expect to see—does a measure of trait X actuallyhavequestionsaboutXthatpatients/subjectsrecognizeassuch?Some instrumentshavenoorlittlefacevalidity,butareperfectlyvalid—forinstance, theMinnesotaMultiphasicPersonalityInventory(MMPI).However,instruments
that lack face validity may not be completed, or not be completed correctly, becausethepatientfailstoseetheirrelevance.InthearenaofFA,facevalidityis hardlyanissue,becausetheactivitiesthat are used as indicators of functional ability have a fairly low level of abstraction and are recognized by patients/clientsasrelevanttotheirlife.
The closely related term “content validity” refers to a measure actually covering the entire width of the construct the developer is targeting. It is generallydeterminedbyhavingexpertsdrawuplistsofnecessarycontentsfora measure of X, or their checking the content of a draft measure against their unwrittenexpectations.Ofcourse,thispresumesthereisacleardescriptionby thetestdevelopersoftheconcepttheywouldliketooperationalize—itmakesno sensecriticizingtheveracityofapaintedportraitifyoudonotknowtheperson depicted.Therearenostandardformulasforcalculatingthisvalidityaspect.
“Predictive validity” concerns the ability of a measure to predict a future stateor event that is inherently linked to the characteristic being measured. A collegeentranceexaminationissaidtohavepredictivevalidityifitcanbeused toaccuratelypredictwhoin4(5,6)yearswillgraduate.AparallelinFAwould be the ability of a measure to predict which rehabilitation patients will be successfully discharged home versus to a nursing home. One problem with predictivevalidityassessmentisthefactthattherearenohardandfastrulesas to what should bethe minimum level ofsuccess in predicting. We know that manyfactorsaffectsuccessfulindependentliving—theaccessibilityofthehome, familysupportavailable,theperson’sdeterminationandtoleranceforrisk, and soforth.DoesanFAinstrumenthaveadequatepredictivevalidityifpredictions basedonitarecorrectatleast50%ofthetime?Atleast80%?
“Knowngroupvalidity”isbasedondifferencesinscalescoresbetweentwo groups that are known to differ in the characteristic the instrument aims to measure. The average score of persons with SCI on a measure of physical functioning should be lower than the average of persons with traumatic brain injury(TBI);incaseofameasureofcognitivefunctioning,thesituationshould bereversed. Ifthedata donot paralleltheseexpectations, itisquite likelythe instrumentisnotmeasuringwhatwethinkitismeasuring.Alternatively,alotof systematicerror(bias)isreflectedinthedata.Asimilarproblemasmentioned earlieroccurswithdeterminingknowngroupvalidity:howmuchofthevariation inthe functional statusofthe overall groupshouldbe explained bydiagnostic category,SCIversusTBI?IfeverypersonwithSCIisknowntohaveahigher cognitivefunctioninglevelthaneverypersonwithaTBI,thingswouldbeeasy:
thevariationexplainedshouldbe100%,andeverything lowerthan thatwould mean less than perfect validity of a proposed FA instrument. However, the distributions of functioning ability (both motor and cognitive) of theTBI and SCIgroups overlap somewhat.Statingthat a good FAmeasureshouldexplain between1%and 100%of thedifferencebetween groupsisnotveryhelpfulin selectingordevelopinganinstrument.
“Construct validity” concerns the relationships between the measurement dataofa(highlyabstract)constructanddataforotherconstructs.Sometimeswe havea basisin theoryto predictthatconstructKshould bestrongly relatedto (yet not identical with) construct L and be independent of construct M. (For instance, “ADL ability is related to community integration, but unrelated to political party preference.”) If the data bring this out, the measurement of K likelyisvalid(andsimilarlytheoperationalizationsofLandM).Ifthepredicted associationbetweenKandLisminimalorabsent,however,wedonotknowif theproblemiswiththetheory,orwiththeoperationalizationofK,orwiththe measurementofL.Anditisanunusualtheorythatspecifiestheexactstrengthof the relationship between K and L, predicated on perfect measurement of the constructsinvolved.“Strong”or“verystrong”isthebestwegetfromtheorists, andthosearenotverygoodstartingpointsforevaluatingthelevelofvalidityof theinstrumentsinvolved.
“Ecological validity” does not concern an instrument’s validity perse, but the relevance of assessment data to real-life situations outside the testing situation. Testing ambulation skills in situations that resemble the real world (morethantheparallelbarsinthephysicaltherapygymdo) providesdatathat are more “ecologically valid,” but the standardization of testing might suffer. Standardizationoftestinghasalwaysbeenakeystoneofpsychometricmethods of assessing reliability and validity of neuropsychological instruments and of instrumentsquantifyingtheabilitiesofindividualpatientsorclients,allofwhich are capacity measures. However, standardized environments tend to be dissimilar from the settings where people perform self-care, communicate, do work,andallotherthingscapturedundertheumbrellaoffunctioning.Testingin a standardized environment almost always means in an optimal environment (15), and the results tend therefore to be more indicative of capacity than of performance,whichmaymeanthatthetestingdatawillnotbeverypredictiveof real-liferoutine.
The above discussion should make clear that estimating the validity of instruments is always less straightforward than the quantification of their
reliability.Findinghighvaluesparallelto,forexample,a0.91leveloftest-retest reliability just does not happen; validity coefficients are almost always much lowerbecauseallmethodsofvalidityestimationareroundabout.Themostdirect assessment method is convergent validity, for which a minimum correlation betweentwomeasuresofthesameconstructof0.60issuggestedasminimally adequate (27). In practice, it is almost always necessary to use all available methodsofestimatingvalidity,andbasedonmultiplefindings“patchtogether” evidence supporting validity—which never will be iron-clad. Finding encouraginglevelsofthevarioustypesofvaliditydistinguishedhere,inmultiple studies, with patterns of correlations that make sense based on expert knowledge, is what typically occurs. Fortunately, in the case of FA, the specialists involved have extensive knowledge of the determinants, correlates, intergroup differences, and so forth, of various aspects of functional status, makingthematterofappraisingthequalityofspecificmeasureslessproblematic thantheprecedinglistofissuesmightsuggest.
Sensitivity
It is easy to see that if an FA “measure” has just two categories, “able” and “unable,” it lacks sensitivity: it cannot reflect fine distinctions in capacity/performance, and it cannot be used to record minor but clinically significant changes in the performance of an individual or group. Sensitivity refers to the ability of an instrument to capture, across the full range of functional ability of the subjects/patients to be measured, distinctions that are clinicallyrelevantorsmall enoughto stillbeof importancein research.When sensitivityisdiscussedinrelationtochangeover time(e.g.,fromadmissionto discharge),thetermresponsivenessisfrequentlyused.
Flooreffectsand ceilingeffectsare oneissue insensitivity.Thefirstterms refertothelowestmeasurablelevelofperformanceonanFAinstrumentbeing higherthanthestatusoftheleastable person to be measured. All individuals who have ability equal to the lowest measurable level or lower are lumped togetherand given the corresponding score. Vice versa, a ceiling effectmeans thatthehighestmeasurablelevelislowerthantheperformancelevelofatleast someofthemoreablepatients.Itshouldbenotedthatveryoftenmeasuresare developedforonepopulationinwhichtheyhavenofloororceilingeffects,but thenareappliedtoanothergroup inwhichtheydo.Forinstance,theFIM was designedtoquantifyfunctionalstatusofrehabilitationinpatients,andanypatient
whoachievesthe maximum score on dischargeprobablywasan inappropriate admission.However,afewyearsafteronsetofincompleteparaplegia(e.g.,C4 orbelow,ASIAImpairmentScaleD),manypersonswillscoreatthemaximum oftheFIMMotorsubscale.TheFIMwasneverdesignedtodistinguishbetween peoplewithminimaldeficitsthatdonotaffectfunctioning,othermeremortals, andSuperman.Thus,“lackofresponsiveness”sometimesisaproblemforwhich theinstrumentuserisresponsible,nottheinstrumentdeveloper.
Quantification of responsiveness is not done using formulas resulting in simple coefficients ranging from 0.0 (not responsive at all) to 1.0 (maximum responsiveness possible). All quantification methods are mostly useful for comparingtheresponsivenessofonemeasurewiththatofanother,allowingone toselectthemostresponsiveone.Avarietyofindicesareused,includingeffect sizes (the mean change between time 1 and time 2 divided by the standard deviationattime1),thestandardizedresponsemean(themeanchangebetween time1andtime2,dividedbythestandarddeviationofchangescores),receiver operating characteristic (ROC) analysis, andmany others. Discussion ofthese indicesisbeyondthescopeofthischapter;thereaderisreferredtotheextensive literature(28,29).
A related issue is: what is the smallest change (over time) or difference (between two cases) that a measure allows one to detect? This is not just a questionofthemetricused(ageindaysofcoursecanreflectsmallerdifferences thanageinyears),butalsoinvolvesmeasurementerror.Everymeasuredvalueis an approximation, and it may be appropriate toindicatethe likely error range involved:theweightofpatientXonadmissionwas53.7±0.2kg.Information like that indicates that the scale that was used cannot reliably detect changes smaller than 0.2 kg. In clinical epidemiology and related areas, the term Minimum Detectable Difference (MDD) may be used; in psychology, the corresponding term is reliable change, and a Reliable Change Index might be offered.OthertermsincludeSmallestRealDifferenceandMinimumDetectable Change.
The MDD should not be confused with a second concept: the Minimal ClinicallyImportantDifference,orMCID(sometimesdesignatedtheMinimum [or Minimal] Important Difference, Clinically Important Change or Minimal Important Change). MCID refers to the smallest score difference that patients consider to be of value. The MCID by definition should be larger than the correspondingMDD—adifferencethatisdetectableisnotnecessarilysufficient to reflect meaningful change in everyday functioning for the person being
assessed. Witha very sensitive scale, someone who wants to lose weight can determineheorshehaslost0.013(±0.002)kgover4weeks,butdoesthatmean heishappywiththeresultofhisdieting?Similarly,inaclinicaltrialofanew methodforrehabilitatingtheupperextremities(UE)afterSCIthemeanscorefor thetreatmentgrouponaUEmeasuresuchastheCUE-T(CUEtest[ratherthan self-report] version) (30) may be 102, versus 99 for a usual care comparison group(withp=.03indicatingthatthedifferenceisstatisticallysignificant),but does that difference matter clinically—especially in light of potentially much higherresourceexpenditureassociatedwiththenewintervention?Thus,MCID reflectsclinicalsignificance,mostlyasseenfromtheperspectiveofpatients,and assuch can belinked to effectsizes, numberneededto treat (NNT)and other waysofquantifyinghowmuchdifferenceismeaningful.Thereareavarietyof methodsfor determiningthe MCIDfora particularFAmeasure,(31) which— becausetheyare basedon differentassumptions—tendtogivedifferentMCID values.Wuetal.(32)giveacogentdiscussionofissuesinvolvedindetermining MDDsandMCIDsforuseinSCIresearch,emphasizingthequestionableuseof theseparameterswithordinalFAmeasures.
OtherMetricCharacteristics
Beyondvalidity,reliability,andsensitivity,thereareafewothercharacteristics ofanFAmeasurethatarerelevanttoitsuseinclinical,programevaluation,and researchapplications,mostofwhichhavetodowithpracticality:
Language:Carefulwordingisespeciallyrelevantforself-administered
instrumentssuchastheCUE,(7)butmayalsobeanissuewithobservational andothermeasures(33).Boththetext’sreadinglevelandatranslationintoa languagetheuserisfamiliarwithareofconcern.InSCI,thewordsusedin instruments(suchas“walk”intheSF-36)maybeproblematicinthatthey arenotapplicabletothosewhouseawheelchairformobility,andevenmay beinterpretedasreflectinginsensitivityonthepartoftheresearcher(34).
Trainingrequired:ManyobservationalFAinstrumentsandtest-typemeasures
requiretheusertobetrained,andsometimescertified,toproducereliable data.
Availabilityandcosts:Somemeasuresarecopyrighted,andmaynotbe
availableatall,oronlyforaone-timeorper-usepayment.
Timeandequipmentrequired:Measuresthattakeinordinatetimeonthepart
ofthesubjectsortheadministrator,orthatusespecialequipment,maynotbe suitableoutsideresearchapplications.
Alternativeversions:Availabilityofaversionforcompletionbyaproxymay
beusefulforself-reportmeasuresusedwithchildren,adultswithhigh tetraplegia,orindividualswithcognitive-communicativedeficits.Similarly, equivalentversionsareofuseinsituations(e.g.,psychologicaltesting)where subjectsmay“learnthetest,”andwouldappeartogaininskillsonrepeat administrationofasingleversion.
Patientsafety:Iftestingtheabilitytodriveacarisdisturbedbythetest
administrator’sinterventioneverytimethereseemstobedanger,thetest resultisnotveryrealistic.Ontheotherhand,notinterferingatallisnotto thebenefitofthepersontested–orthetester.Simulationssuchasvirtual reality(seelater)havebeendevelopedtomake“realistic”testingin situationslikethispossible.However,theymaybringoutecologicalvalidity concerns.
Clinicalutility:Ifadministrationandscoringofameasuredoesnotaddtothe
clinician’sknowledgebase,oriftheinformationdoesnothelphimorherto makedecisionswithrespecttoaparticularpatientoraclassofpatients,the instrumentlacksclinicalutility.Interpretabilityofthescoresmaycontribute toclinicalutility;availabilityofnorms(forallpersonswithSCI,or preferablyforsubgroupswhoarecomparableintermsofage,leveland completenessofinjury,etc.)alsoisofbenefitinsomeinstances.
USESOFFUNCTIONALASSESSMENT
Theorigin of FA(definednarrowlyas measurement ofActivityLimitations)is foundinattemptsbyclinicianstoexpressquantitativelythedeficitspatientshad on admission to rehabilitation, and to monitor their progress, or at least determinedischargestatus,sothattherewassome“proof”oftheeffectiveness of treatment beyond the patient’s simple report that he or she now could do things that were impossible or difficult before admission. The range of applicationsofFAinstrumentshasexpandedtremendously,sothatwenowcan describeusesinthecareofindividualpatients,inprogramsadministration,for reimbursement,andforresearch.
CareofIndividualPatients
Decisionsonadmissiontoinpatientandoutpatientprogramsareoftenbasedona formal FA to see if the person has the types and degrees of deficits that the programisqualifiedandauthorizedtotreat,eitheringeneralorforthespecific personinquestion.The“baseline”assessmentthereforeisoftencommunicated to the third-party payor,who mayuse it to approve program admission anda certain duration or intensity of treatment. The pre-admission or admission assessmentisfrequentlythebasisforaprognosis,whichiscommunicatedtothe patient and the payor, and ideally underlies goal setting. Many rehabilitation programsuseanFAinstrumentsuchastheFIMtosetexpectedoutcomes,either forclassesofpatientsorforindividualpatients.Softwareapplicationshavebeen developedtoassistcasemanagerstomakesuchpredictions;theyarefounded,in large part, on an FA database that contains information on the admission and discharge status of many previous patients with the same rehabilitation diagnosis,age,gender,andcomorbidities.
Treatmentmonitoring using an FAmeasure is done in many rehabilitation programs. Team rounds often consist of the reporting by “most responsible/ knowledgeable therapies” (nursing for bladder; speech for expression, etc.) of the current status of the patient on the numeric items offered by the FA instrument used in the facility. Although treatment termination decisions are increasingly triggered by “external” criteria (e.g., a maximal length of stay approved by a third-party payor), ideally they are founded on either the accomplishment of goals or the plateauing of the patient in terms of overall functional ability. In both instances, measurement of patient status should be performedusinganinstrumentthathashighsensitivitysothat(lackof)change canbereliablydetermined.
Unfortunately,casemanagersandmedicalinsurancecompaniesmaydemand thatimprovementismeasuredusinginstrumentsthatdonotadequatelycapture theextentofimprovementthatmaybeoccurring.Anexampleistheuseofthe Motor subscale of the FIMin an SCI patient with high-leveltetraplegia. This measure is unlikely to adequately document improvement due to the FIM’s insensitivityto changeinthis group.A handfunctiontest, ortheQuadriplegia IndexofFunction(QIF) maybe abetter choice(35). In theoutpatient setting, theFIMalsomaynotcaptureimprovementsinfunctionduetoaceilingeffect.A broad understanding of the pitfalls of available scales allows the healthcare providertobestapplythesemeasurementsandeducateinsurers.
FA information may also be used to communicate about progress and outcomesoftreatmentwithpersonswhoarenotpartoftherehabilitationteam.
Patientsthemselves, theirfamilymembers, referral sources,and payors havea stronginterestinthefunctionalaspectsofthepatient’sstatus,especiallywhereit concerns Activities and Participation. One additional use of FA is long-term monitoringof apatient’s status. Especiallyin the caseofprogressive diseases, suchasmultiplesclerosis,thisinformationisimportanttomakedecisionsona needfornewtreatments,changesinpatientenvironments,andsoforth.Infact, thisuseofFAhasledtoadesignationoffunctionalstatusinformationasasixth (afterthestandardfourandpain)vitalsign(36).
ProgramAdministrationandEvaluation
Whetherpartofoutcomesmanagement,continuousqualityimprovement(CQI), ortotalqualitymanagement(TQM),programevaluationaimstoassesstowhat degree a program indeed accomplishes what it sets out to do—improve the functional status of people with disabilities. Basic questions of program evaluationare:dopatientschangeforthebetter(programeffectiveness),andif so, are resources used optimally in accomplishing this (program efficiency)? Program (self-)evaluation is required by the Commission on Accreditation of Rehabilitation Facilities (CARF) (37), a widely recognized not-for-profit that accreditsorganizationsandprograms.Routinelycollectedoutcomedatacanand shouldbecommunicatedtostakeholders, includingcurrentandfuturepatients, third-partypayors,andthelocalcommunity.TheU.S.CentersforMedicareand MedicaidServices (CMS)has started topost comparativefunctionaloutcomes fornursinghomesandhomehealthagenciesonitswebsite,andsimilar“report cards” including FA information will be published in the future for other facilitiesthatofferrehabilitationservices(38).
Althoughchangefromprogramadmissiontodischargeiscommon,itisnot easy to offer proof that the program deserves credit. Aperson with SCI may scorehigheronapost-testthanonapre-testforreasonsthathavenothingtodo with the selection, timing, quality, and quantity of services received. Positive changemaybe dueto naturalrecovery,improvedtest-takingability,andmany otherfactors(39,40).Unfortunately,routineprogramevaluationdatatendtobe insufficient to indicate what factors are contributing to success or failure; additionalstudiesmaybeneededtoobtainthatinformation.
Allrehabilitationprogramsfacethesameproblemofprovingeffectiveness, andone(partial)solutionthathasbeenfoundistocompareoutcomesbetween programs through a “minimum data base” that includes demographics, time