Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:

Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_205_библиотеки_им_акад_М_И_Перельмана

.pdf
Скачиваний:
0
Добавлен:
15.09.2026
Размер:
13 Мб
Скачать
☆
72 The APA Publishing Textbook of Mood Disorders, Second Edition
https://t.me/med1917
TABLE 4–1. Validated functional impairment scales
No. of
Validated scales items Domains
Social and Occupational Functioning
Assessment Scale (SOFAS) and the related Personal and Social Performance (PSP) scale
Global Assessment of Functioning
(GAF) scale
Work and Social Adjustment Scale
(WSAS)
Longitudinal Interval Follow-Up
Evaluation Range of Impaired Functioning Tool (LIFE-RIFT)
Assessment of Impairment
of Functioning/Disability from the Mini International Neuropsychiatric Interview for DSM-5
Life Functioning Questionnaire
(LFQ)
Endicott Work Productivity Scale
(EWPS)
World Health Organization
Disability Assessment Schedule (WHODAS 2.0)
1 Social, occupational
1 Severity of symptoms and functional
alteration
activities, family life/home responsibilities
5 Work, home management, social and
personal leisure activities, relationships
9 Work, interpersonal relations, general
satisfaction with functioning, recreation
12 Work, social life, leisure ac tivities, family
life/home responsibilities, relationships, getting along with others, ability to communicate, self-care ability, interpersonal aggression/disruption, financial impairment, physical impairment, spiritual/religious impairment, impairment of family by symptoms
14 Tasks at work/school, tasks in the home,
leisure time with family and friends
25 Work efficiency and productivity
36 Cognition, mobility, self-care, interpersonal
actions, activities, participation in society
international clinical trials. The SDS is very sensitive to change effects and is very sen­sitive as an efficacy signal detector, often detecting such signals when the other out­come measures fail. It has been used as the key measure for regulatory agency approvals for improving functional impairment with several medications for mood disorders. The SDS, which has been translated into more than 80 languages/language variants, assesses the three most important domains of functioning: work and school, social life/leisure activities, and family life/home responsibilities. It captures days lost and days underproductive at work. It accommodates 13 different time frames. It is very short, simple, and clear, appearing on one page. The SDS can be rated by pa­tient or clinician, although it is much more frequently used as a patient-rated scale. It uses a unique discretized-analog (Discan) metric that was specifically designed to be
role, parent role, family role
73 Rating Scales and Structured Diagnostic Interviews for Mood Disorders
https://t.me/med1917
sensitive in detecting efficacy signals in treatment studies. The Discan metric is an­chored numerically (0–10), verbally-descriptively, and visually-spatially to enhance sensitivity across a population of subjects. Several variants of the scale with addi tional domains, including a child and adolescent version, are available for specific set­tings and needs (Sheehan and Sheehan 2008; Sheehan et al. 1996). The SDS has been expanded to cover 12 domains in an add-on module to the MINI 7.0.2 for DSM-5 (Sheehan 2020).
The five-item Work and Social Adjustment Scale (WSAS) (Mundt et al. 2002) is very similar to the SDS, and its design was largely based on an earlier SDS. It assesses five domains and uses a 0–8 scale of response options. It has acceptable internal con sistency, test-retest reliability, and sensitivity to change.
The nine-item Longitudinal Interval Follow-up Evaluation Range of Impaired Functioning Tool (LIFE-RIFT; P. Fisher, A.C. Leon, and M.E. Coles, unpublished man uscript, New York, New York State Psychiatric Institute, 2002) assesses functioning in the domains of work, interpersonal relations, life satisfaction, and recreation. This scale has been found to have good reliability and validity (Leon et al. 1999). It has been used in a longitudinal study of bipolar I disorder with good results (Leon et al.
2000) and has been adapted for children and adolescents (P. Fisher, A.C. Leon, and M.E. Coles, “The Longitudinal Interval Follow-up Evaluation Range of Impaired Functioning Tool (LIFE-RIFT) Adapted for Children and Adolescents,” unpublished manuscript, New York, New York State Psychiatric Institute, 2002).
The 14-item Life Functioning Questionnaire (LFQ) (Altshuler et al. 2002) was de­signed to assess functioning in four domains: workplace, duties at home, leisure time with family, and leisure time with friends. It was found to be a reliable, consistent, and valid assessment of function at work and home in a study of patients with a mood disorder.
The 25-item Endicott Work Productivity Scale (EWPS) (Endicott and Nee 1997) fo­cuses on work productivity and efficiency.
The 48-item self-rated Social Adjustment Scale (SAS) (Weissman and Bothwell
1976) is the oldest and most detailed of the functioning scales. It was used in many epidemiology studies and found to be very informative across a wide range of do­mains. However, because of its length, it is less suitable for routine clinical use or as an outcome measure in treatment studies. However, it remains the mother of most functional impairment scales.
-
-
-
Assessment Instruments for Suicidal Ideation and Behavior
More attention has been paid to the development of assessment instruments for sui­cidal ideation and behavior in the past decade than to scale development in any other area of mood disorders. Meyer et al. (2010) reviewed the status of the 29 most import­ant assessment scales following a large international consensus conference on suicid­ality and suicide risk in 2009. Nine of the 29 scales were considered to have achieved a higher level of development than the others. Several of the more important scales in that review have undergone significant updates and revisions since then. The stan­dards against which they were judged were the categories in the 2010 Draft Guidance
74 The APA Publishing Textbook of Mood Disorders, Second Edition
https://t.me/med1917
of the U.S. Food and Drug Administration (2010). This guidance underwent signifi­cant revisions in 2012 (U.S. Food and Drug Administration, Center for Drug Evalua­tion and Research 2012) and is undergoing further revision at the time of this writing in 2021.
Some of these scales were developed for adults, others for adolescents, and still others for children. Some are self-report measures, others are clinician administered, and some accommodate both routes of administration. Some focus on the assessment of suicidal ideation, others evaluate suicidal behaviors or self-injury, and others as­sess a wider range of suicidality phenomena across the spectrum of suicidality.
To be acceptable for use in a clinical trial or as an outcome measure of safety or ef­ficacy, a suicidality assessment scale must map closely or precisely to the categories outlined in the 2012 FDA draft guidance document (U.S. Food and Drug Administra tion, Center for Drug Evaluation and Research 2012). In 2021, the following five scales were deemed to meet this standard: the Columbia–Suicide Severity Rating Scale (C SSRS) (Posner et al. 2011), the Suicide Ideation and Behavior Assessment Tool (SIBAT) (Alphs et al. 2017), the Sheehan–Suicidality Tracking Scale (S-STS) (Sheehan et al. 2014b), the Sheehan–Suicidality Tracking Scale Clinically Meaningful Change Mea­sure (S-STS CMCM) version (Sheehan et al. 2014a), and the InterSePT Scale for Sui­cidal Thinking (ISST Plus) (Lindenmayer et al. 2003). All five scales can be clinician rated. The S-STS and the S-STS CMCM can also be patient rated. The SIBAT has some patient-rated components. The S-STS is the only suicidality assessment scale that has been linguistically validated for children and adolescents (Amado et al. 2014). It is available in three versions for three different age groups.
The C-SSRS was originally developed as a safety assessment questionnaire with bi­nary (yes/no) response options. It is therefore a questionnaire rather than a dimen­sional rating scale and consequently does not lend itself to use as a sensitive treatment outcome measure capable of detecting an efficacy signal. Although it was a welcome improvement at the time over preexisting suicidality assessment methods in assess­ing treatment-emergent suicidality in clinical trials, it has many well-documented limitations (Giddens and Sheehan 2014b; Giddens et al. 2014).
The Beck Scale for Suicidal Ideation (BSI) (Beck et al. 1979) has also been used as an efficacy outcome measure. The BSI assesses a narrower range of the spectrum of suicidality than many of the other suicidality scales. One result is that it does not de­tect some important treatment-emergent suicidality phenomena. Overall, the BSI has been the least sensitive assessment in detecting treatment effects and in discriminat ing between active treatment and placebo for suicidality. Both item 10 on the MADRS and the S-STS have been more sensitive in detecting these efficacy signals in double­blind, placebo-controlled trials (Canuso et al. 2018; Khan et al. 2011). Suicidality as­sessment measures continue to evolve rapidly (Chappell et al. 2014).
There is much more to suicidality than suicidal ideation and behavior. Suicidality also includes suicidal impulses, suicidal hallucinations, delusional suicidality, and suicidal dreams. A good suicidality assessment scale should ask about the full range of suicidality phenomena. It should be sensitive in simultaneously detecting treat­ment efficacy and treatment-emergent safety effects (e.g., increased suicidality) and in detecting differences between drug and placebo. Such sensitivity is ethically nec­essary to identify efficacy and safety signals in as small a sample size as possible, while exposing the fewest possible number of patients to risk in clinical trials.
-
-
-
75 Rating Scales and Structured Diagnostic Interviews for Mood Disorders
https://t.me/med1917
Using one or two suicidality items from a depression scale (e.g., item 3 of the HAM­D, item 10 of the MADRS, or item 9 on the PHQ-9) to screen for suicidality is a mistake. If you want to screen for suicidality, screen for suicidality directly rather than screen ing for one of its proxies, and assess the full range of suicidality phenomena in the pro­cess (Preti et al. 2013). This matters because the progression of any one suicidality phenomenon over time is not related to any other in a precise linear fashion (Sheehan and Giddens 2015). Their relationship to each other is almost always polynomial to the fourth order or higher. You cannot accurately predict the presence or magnitude of one by the presence or absence of another (Sheehan and Giddens 2015). As a result, you cannot rely on any one phenomenon to tell you whether any other suicidality phenom enon is present or absent or to what extent within a time frame. Time spent per day in suicidal ideation/behavior and the total S-STS score may be more accurate and sensi tive measures of global suicidality than a global severity of suicidality measure, espe­cially at the more serious end of the spectrum (Giddens and Sheehan 2014a, 2014c).
When using a suicidality assessment scale, you should adopt the mindset that you are assessing the presence or absence of suicidal phenomena within a recent past time frame while at the same time avoiding the temptation to make predictions about future suicidality. You cannot accurately predict suicidality at the individual level within a narrow time frame. The reason is that the progression of suicidality over time within a single individual is nonlinear when mathematically modeled (Sheehan and Giddens
2015). You do not use a schizophrenia scale to predict if and when someone is going to get a future hallucination or delusion, and you do not use a panic disorder scale to predict if or when someone is going to have another panic attack. Instead, you mea sure what has already occurred and you make your management decisions accord­ingly. Assessment of suicidality calls for the same approach as assessment of other areas of mental health. My recommendation is that you get copies of the various scales, try them out with patients, get patient feedback on them, and see whether they meet your needs and address your clinical and research questions.
-
-
-
-
Global Severity and Improvement Scales
The Clinical Global Impression—Severity (CGI-S) and Clinical Global Impression— Improvement (CGI-I) scales have been widely used to assess illness severity and global improvement or change. Both are seven-point scales. The CGI-S rates severity from 1 (normal) through 7 (among the most severely ill patients) (Guy 1976). On the CGI-I, there are three points of deterioration and three points of improvement, with a point denoting no change in between (Guy 1976). The CGI-I met the needs for an im provement scale when it was first developed and used. However, with increasing numbers of treatments and an increasing need to detect smaller differences in efficacy between multiple treatments in the same study, it became necessary to develop more sensitive global improvement scales. To meet this need, the Sheehan Clinical Global Improvement (S-CGI) scale and the Sheehan Patient Global Improvement (S-PGI) scale were developed. These measures of improvement are 21-point scales using a Dis­can metric for the response options. There are 10 points of improvement (1–10) and 10 points of deterioration (–1 to –10), with a zero (0) point in the middle, reflecting the baseline starting point score in the study or no change at any visit (Sheehan et al. 1988).
-
76 The APA Publishing Textbook of Mood Disorders, Second Edition
https://t.me/med1917
Comparison of Patient-Rated and Clinician-Rated Scales of Mood Disorders
Self-rated scales take less clinician time, but are they comparable to clinician-rated scales? Early comparisons in the adult depression scale literature (Carroll et al. 1973; Prusoff et al. 1972) found only moderate correlation between self-rated scales and cli nician-rated scales (r=0.11–0.63 and r=0.42, respectively); this may be because differ­ent scales were used in the comparisons. The Carroll (Depression) Rating Scale (CRS) is a scale specifically designed to correspond with the 17-item Hamilton Depression Rating Scale (HAM-D). Even though they are very closely related instruments, a com parison between the two yielded a correlation that was still far from ideal, although it was higher (r=0.67) than that usually found between self-rated and clinician-rated scales. The conclusion was that self-rated depression scales are less reliable than cli nician-rated depression scales (Carroll et al. 1981). That conclusion is open to debate. Biases (sometimes self-serving) are inherent in clinician-rated scales. Clinicians using a clinician scale sometimes have their own rating prejudices and preconceived ideas that are at variance with other raters. Similarly, patients using a patient-rated scale may not be able to accurately rate the severity of their symptoms in relation to how other patients rate the same symptoms. In the absence of information about how other people experience depression, one patient may overvalue or underestimate the severity of one or another symptom in the depressive cluster, compared with another patient’s ratings. However, both clinician and patient perspectives have merits and limitations.
Experience has taught me that as patients became more familiar over time with completing these scales and reviewing and discussing them with their clinicians, their scores grow closer to the clinicians’ ratings. I have more confidence in patient ratings than many other clinicians do. Clinicians should use them more frequently and should spend more time up front training patients in the proper and accurate completion of rating scales.
-
-
-
Interpretation of Scale Scores
Use and Misuse of Factor Scores
Factor scores have been extracted from existing scales based on factor analytic studies. For example, somatic anxiety/somatization, psychic anxiety, pure depression, and an orexia factors have been identified for the HAM-D. However, these factors are not sta­ble across different samples in different settings. Identifying items belonging to these named factors and then studying the effect of treatments on them in subsequent stud­ies is a misleading process. For example, in a very large dataset used to investigate anxiety in clinical trials with depressed patients, we (Goldberger et al. 2011) found that the “use of Cleary and Guy’s anxiety/somatisation factor from the HAMD unstable across studies and is not an appropriate routine measure of anxiety in de­pressed patients” (p. 48). We suggested that “in clinical trials, other scales specifically
-
is
17
77 Rating Scales and Structured Diagnostic Interviews for Mood Disorders
https://t.me/med1917
designed and validated for anxiety assessment should be used. An alternative strat­egy is to assess the factorial structure of the HAMD and use for outcome assessment the factors identified by this analysis [within each study sample]. Such ‘tailored’ factor analysis would retrospectively allow the identi fication of the set of items appropriate for each specific study sample” (p. 48).
scale for each studied sample
17
Response and Remission Cutoffs
Response for each individual on any scale is usually defined as a 50% or greater reduc­tion from the baseline score. Remission for each individual on any scale is usually de­fined as a 70% or greater reduction from the baseline score. Remission is typically anchored to a reduction of 2 standard deviations from the baseline mean. Patients are said to be in remission when their scores reflect their own expectations of a good out come, even if the score is less than perfect. Remission is also said to reflect an experi­enced clinician’s expectations of a good treatment result, even if the outcome is not yet perfect. In mood disorders there is remarkable confluence and consistency across all four of these anchors of remission. These four anchors of remission line up with remission scores of 7 or lower on the HAM-D17 and 10 or lower on the MADRS.
In recent years, some pharmaceutical companies and researchers have used a score of 12 or lower on the MADRS as a remission cutoff score. Clinicians should beware of cutoffs like this. Such reports can be self-serving, instigated without justification by marketing divisions so that a larger percentage of patients can be declared in remission in the interest of making an investigational product appear more effective than it really is. Instead, such reports have the opposite effect by diminishing the companies’ image in the eyes of the reader. Inflating results in this way has negative consequences. These include reducing interest in investing in new treatments by suggesting 1) that existing treatments are better than they are, 2) that we do not have unmet treatment needs, and
3) that we do not have a long way to go in improving the efficacy of current treatments. No clinicians should participate in these charades.
-
-
Challenges in Developing Scales
Based on my experience in developing several scales, at least $2 million is spent per scale before it is widely adopted, translated into many languages, and used and ac­cepted internationally. If a scale is good, it takes about 7 years of work before it gains traction in the field and 15 years before its use grows exponentially. The hurdles in volve psychometric testing, validation and reliability studies, extensive patient and expert opinion input, cognitive debriefing, overcoming hurdles and expectations of regulatory agencies, and linguistic validation and cognitive debriefing for all transla­tions into other languages. That is why there are so few very good and widely used scales across the entire field of mental health. Many developers give up long before they reach the goal, because they get discouraged at the time and/or expense in­volved. Scale development requires a thick skin and endless persistence. The results are rewarding, however. What is important is to get started and to stick with it. The biggest progress is made in the first 2–3 years, if the scale developer stays focused on designing a scale that is “fit for purpose.”
-
78 The APA Publishing Textbook of Mood Disorders, Second Edition
https://t.me/med1917
Advice on Use of Scales in Clinical Practice
You as a clinician may hope to administer scales and structured diagnostic interviews during clinic visits, but you will not have enough time. If you expect success, then have the patient complete a self-report scale before the start of the visit. Schedule pa tients early so they have time to complete the battery of scales before they see you. Arrange for front office staff to hand these scales to patients on a clipboard to com plete in the waiting room before their appointments. Alternatively, email the patients the battery of scales individualized for them to complete on the day before their ap pointments. Ask patients to bring their completed scales to the visit. Let them know they must complete the scales before seeing the clinician. This practice will signifi cantly improve the success and adherence rate. At the start of the visit, quickly review the completed scales, thank patients for completing them, and tell them how helpful this information is to you in your decision making. Use the scale responses as a basis for more in-depth probing or follow-up questions.
Ecological Momentary Assessment
Ecological momentary assessment (EMA) involves repeated sampling of an individ­ual’s symptoms, behaviors, functioning, and cognitions in real time. This can be done using mobile devices or microprocessors to sample subjects in their real-world and natural environment for a few moments during their day, often using random timing. It minimizes recall bias and maximizes ecological validity—that is in the individuals’ ability to relate to themselves, to other people, and to their environment.
EMA most frequently utilizes mobile phones or wearable devices like heart moni­tors, electronic wrist bands, or rings with built-in technology. The information is often linked to a mobile phone app, which in turn synchronizes data to remote sites, where the data are gathered across large study samples. The use of EMA is rising rapidly.
-
-
-
-
Advice on Use of Scales in Research Settings
Number of Scales
A common problem in a research study is using too many scales. As you do more re­search across time, you will learn to minimize the number of scales you will use. In your early-career studies, try to resist participating in “fishing expedition studies” with a wide range of scales. Stay focused on addressing your narrow research questions.
Statistical Analysis
When you plan a study involving rating scales and structured diagnostic interviews, you need to clearly state your plans for how to collect and analyze your scale data. Your study design and analysis should be precisely aligned to address your research question. Your hypotheses should be clearly and explicitly stated a priori. If you have several hypotheses, they should be prioritized. All this is best done in consultation with an experienced statistician. Presenting your dataset for the first time to statisti­cians after the study is completed and then expecting them to find something of in-
79 Rating Scales and Structured Diagnostic Interviews for Mood Disorders
https://t.me/med1917
terest by fishing around in your dataset elicits sighs of despair and is an all-too­frequent occurrence. You cannot rescue by analysis what you have already damaged in your research design. There are many fine books on this topic. My favorite is Krath wohl’s superb classic on preparing a research proposal (Krathwohl and Smith 2005).
Reporting Results
It is important for clinical researchers to follow the Consolidated Standards of Report­ing Trials (CONSORT) statement proposed by the Journal of the American Medical Asso- ciation when reporting results of studies that use rating scales as endpoints (American Educational Research Association 2014). The figures and tables in a research report should clearly identify the specific statistical test and the dataset—whether it is a com pleter analysis, an observed case, a last observation carried forward, or a modeled dataset (as in MMRM [Mallinckrodt et al. 2001] or ET Rank analysis [Entsuah 1996]). There are different levels of scientific confidence associated with each (Begg et al.
1996). In addition to the sources cited above, readers seeking further guidance on these different datasets are referred to Elobeid et al. (2009).
Validation
Validation is a fundamental and necessary step in scale development. It refers to the evidence in support of the interpretation of the scale scores for the proposed use and the extent to which it is “fit for purpose.” However, there can be an overemphasis on validating any new scale against some existing “gold standard” scale. The problem is that there may not be gold in any so-called gold standard scale. When two scales are compared or “validated against” each other in this manner, it is like studying the cor­relation or agreement between two clouds floating across the sky. Neither cloud is an­chored to anything on the ground. If the so-called gold standard is itself flawed, then showing a high agreement or correlation with it is not necessarily a positive attribute. When I am contacted about one of my scales or structured interviews and the first question is “Has your scale been validated?” I know I am dealing with someone who is not sophisticated in scale selection or evaluation. Many scales in mental health are flawed in one way or another and need to be improved. In the validation testing of a new scale, being able to demonstrate discrepancies or nonagreement with a preexist ing standard may be a positive attribute rather than a negative one. Indeed, it may indicate that the new scale has improved on an older reference scale in an area of known weakness in the earlier scale. When there is no standard scale in a new area of investigation, it is difficult to validate a new scale in this way. Standards for Educational and Psychological Testing (American Educational Research Association 2014) is manda tory reading for everyone involved in scale development and evaluation.
-
-
-
-
Reliability
In contrast to the tendency to overemphasize validation in scale development, there is an underemphasis, in my opinion, on the importance of reliability testing. Inter­rater, intrarater, and test-retest reliability are all critically important properties in the use of scales in mental health research. Good reliability lends confidence to adminis­tering a scale in a large multisite study involving many raters over time. For a detailed
80 The APA Publishing Textbook of Mood Disorders, Second Edition
https://t.me/med1917
discussion of reliability, readers are referred to the Standards for Educational and Psy­chological Testing (American Educational Research Association 2014).
Linguistic Validation
Linguistic validation is also too often neglected, misunderstood, and undervalued. It needs to be much more highly valued and much more frequently included in the planning of scale development and international study planning. Linguistic valida tion is very expensive and time consuming, but it is a necessary step to ensure good interrater reliability across collaborating study sites and international investigators. This process is well delineated in publications by the International Society for Phar macoeconomics and Outcomes Research (ISPOR) Task Force for Translation and Cul­tural Adaptation (Wild et al. 2005, 2009). The process involves a forward translation of the source scale by one research team. A second team, blind to the original lan guage’s source scale, then back-translates the forward translation into the original source language. The back-translation is then compared with the original source lan guage scale by a third team, which includes expert bilingual consultants who have ex­pertise in the mental health specialties in collaboration with input and consultation from a consultant at a translation and linguistic validation service, such as Mapi Re search Trust in Lyon, France. Final reconciliation between all involved, including the scale author or developer, completes the process. Every step by all involved in every stage of the process is documented in a spreadsheet. For example, a typical spread­sheet for the translation and linguistic validation of the MINI into one language is six columns wide and 10,000–15,000 rows deep. When the process is complete, the lin guistic validation service provides certification of the process. This documentation is then available for audit and review by regulatory agencies, research organizations, and sponsors with oversight responsibilities for the study.
-
-
-
-
-
-
Importance of Sensitivity in Detecting Change Over Time and in Identifying Differences Between Effective Medications and Placebo
Efficacy signal detection, or the ability to detect a statistically significant difference between an active medication and placebo in a clinical trial, is a highly valued and essential property of any scale used in monitoring treatment outcomes. Scales that are consistently successful in this regard are more widely used, are more highly regarded, and survive longer than scales that do not have this property. When scales are found to lack this property, their use thereafter goes into decline.
Virtual Administration of Scales and Structured Diagnostic Interviews During Pandemics Such as COVID-19
During the COVID-19 pandemic, I regularly encountered problems while consulting on the use of scales and structured interviews during virtual administration. These
81 Rating Scales and Structured Diagnostic Interviews for Mood Disorders
https://t.me/med1917
issues merit further attention, investigation, and study to restore confidence in the integrity of the process. The following are some concerns regarding and recommen dations for improving the issues that became apparent in virtual administration of scales and structured interviews:
1. The strengths, weaknesses, limitations, sensitivity, and specificity of administration of scales and structured interviews conducted during virtual administration versus live face-to-face interviews need to be carefully studied. Researchers need to inves tigate telephone interviewing; virtual interviewing on Zoom, Doximity, Skype, and WhatsApp; and the different technologies on which apps are used, such as mobile devices (phones) or desktops, laptops, or tablet computers. The use of the various platforms may result in alterations in sensitivity, specificity, and efficacy and safety signal detection. Screen sharing on several of these platforms may come very close to in-person interviews in sensitivity and specificity. There may even be some ad­vantages in signal detection during virtual administration. However, there may be a loss of sensitivity and efficacy and safety signal detection while using information technology (IT) versions or during virtual administration of rating scales.
2. Because these virtual modes of administration are likely to become more common in the future, we need more data to justify the pooling of data captured in these diverse ways across sites and over time in a single study in order to have confi­dence in the integrity of the process.
3. During the pandemic, restrictions on security concerns in the collection of health care information were relaxed in the interest of safety and to ensure proper social distancing. Policies to ensure proper compliance with the Health Insurance Porta bility and Accountability Act of 1996 (HIPAA) and the General Data Protection Regulation (GDPR) (Politou et al. 2018; see also Pew Research Center 2014) need to be reviewed and revised for future research studies.
4. The IT version of a scale or interview on a software platform may not be exactly the same as a paper or PDF version. The layout of a scale on a page can affect both response and error rates.
5. In reviewing software implementations of scales and structured interviews, I reg­ularly encountered software coding errors by IT developers and have learned that the scale authors are frequently not involved or consulted in the certification of the IT versions. Closer collaboration is necessary.
6. The use of scales and structured interviews in data entry/capture software has too frequently fallen short of ideal implementation. IT developers and their sponsors have frequently refused to update or to correct updated versions of the scales and structured interviews and have failed to update these to the correct and current linguistically validated translations. This is a particular problem in ensuring con­sistency with updates to DSM.
7. Training of raters in the use of remote data entry software is incomplete when ver­sions of scales suitable for virtual administration are used. During a pandemic, it is not possible to have safe in-person investigator meetings to ensure good reliabil ity and consistent implementation across sites.
8. Missing data, double entries, and input errors can be avoided by careful software design on the platform used for data capture and entry. Many data capture soft­ware programs are not designed to minimize these problems.
-
-
-
-