Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5431_Библиотеки_им_академика_М_И_Перельмана.pdf
X
- •Contents
- •Preface
- •1. Introduction
- •1.1. Central Questions
- •1.2. Potential Outcomes
- •1.3. Estimand
- •1.3.1. The PROTECT checklist
- •1.3.4. Internal validity and external validity
- •1.4. Probability and Statistics
- •1.4.1. Probability
- •1.4.2. Directed acyclic graphs
- •1.4.3. Statistics
- •1.3.2. Estimand for a given population
- •1.3.3. Estimand for a given super-population
- •1.5. Exercises
- •2.1. Randomization and Blinding
- •2.2. Estimand
- •2.2.1. Causal estimand
- •2.2.2. Statistical estimand
- •2.3. Estimator
- •2.3.1. Expectation of the estimator
- •2.3.2. Variance of the estimator
- •2.3.3. Statistical inference
- •2.4. Common Types of Randomization
- •2.4.1. Simple randomization
- •2.4.2. Block randomization
- •2.5. Exercises
- •3. Missing Data Handling
- •3.1. Missing Data
- •3.2.1. Scenario one
- •3.2.2. Scenario two
- •3.4. Sources of Missing Data
- •3.4.1. Intercurrent events
- •3.4.2. Missing data that are consequences of ICEs
- •3.4.3. Missing data that are not consequences of ICEs
- •3.5. Appendix
- •3.6. Exercises
- •4. Intercurrent Events Handling
- •4.1. Five Strategies
- •4.1.1. The treatment policy strategy
- •4.1.2. The hypothetical strategy
- •4.1.3. The composite variable strategy
- •4.1.4. The while on treatment strategy
- •4.1.5. The principal stratum strategy
- •4.2. Combinations of Strategies
- •4.3. Time-to-event Outcome
- •4.3.1. Censoring
- •4.3.2. The treatment policy strategy
- •4.3.3. The hypothetical strategy
- •4.3.4. The composite variable strategy
- •4.3.5. The while on treatment strategy
- •4.3.6. The principal stratum strategy
- •4.3.7. The competing risk strategy
- •4.4. Sample Size Calculation
- •4.4.1. The treatment policy strategy
- •4.4.2. The hypothetical strategy
- •4.4.3. The composite variable strategy
- •4.4.4. The while on treatment strategy
- •4.4.5. The principal stratum strategy
- •4.5. Exercises
- •5. Longitudinal Studies
- •5.1. Continuous or Binary Outcome
- •5.2. Time-to-event Outcome
- •5.3. Treatment Regimes
- •5.3.1. Dynamic treatment regimes
- •5.3.2. SMART design
- •5.4. Exercises
- •6. Real-World Evidence Studies
- •6.1. RWE Studies
- •6.1.1. Pragmatic RCTs
- •6.1.2. Observational studies
- •6.1.3. Externally controlled trials
- •6.2. Confounding Bias
- •6.2.1. No unmeasured confounder
- •6.2.2. Unmeasured confounders
- •6.2.3. Proxy variables
- •6.3. Longitudinal Cohort Studies
- •6.3.1. Causal estimand
- •6.4. Externally Controlled Trials
- •6.4.1. Causal estimand
- •6.5. Appendix
- •6.6. Exercises
- •7.1. Introduction
- •7.2. M-estimation
- •7.2.1. M-estimator
- •7.2.2. Asymptotic linearity
- •7.2.3. Regularity
- •7.3. G-computation Estimator
- •7.3.1. Plug-in estimator
- •7.3.2. MLE
- •7.3.3. Asymptotic variance
- •7.4. Inverse Probability Weighted Estimator
- •7.4.1. IPW estimator
- •7.4.2. Asymptotic variance
- •7.5. Augmented Inverse Probability Weighted Estimator
- •7.5.1. A class of estimators
- •7.5.2. Asymptotic variances
- •7.6. Exercises
- •8.1. Semiparametric Statistics
- •8.1.1. Semiparametric estimators
- •8.1.2. Super learner
- •8.1.3. Semiparametric estimators based on super learner
- •8.2. Asymptotic Variances of Semiparametric Estimators
- •8.2.1. Parametric submodels
- •7.5.3. AIPW estimator
- •7.5.4. Double robustness
- •8.2.2. The fundamental theorem of regularity
- •8.2.5. Double robustness of AIPW-SL estimator
- •8.3. The Targeted Learning Framework
- •8.3.1. Mini-roadmap
- •8.3.2. TMLE
- •8.3.3. Double robustness
- •8.4.3. Missing data due to analysis dropout
- •8.5. Discussion
- •8.5.1. How to select covariates?
- •8.5.2. How to handle missing covariates?
- •8.5.3. How to use TMLE for RCTs?
- •8.5.4. How to implement TMLE?
- •8.6. Exercises
- •9.1. Longitudinal Cohort Studies
- •9.1.1. Causal estimand
- •9.1.4. LTMLE
- •9.1.5. ATE estimand
- •9.2. Missing Data
- •9.2.1. Monotone missing
- •9.2.2. Non-monotone missing
- •9.3. Implementation
- •9.4. Exercises
- •10. Sensitivity Analysis
- •10.1. Introduction
- •10.2.1. The consistency assumption
- •10.2.2. The exchangeability assumption
- •10.2.3. The positivity assumption
- •10.3. Sensitivity Analysis for the MAR Assumption
- •10.3.1. A class of reference-based imputation models
- •10.3.2. Sequential modeling
- •10.4. Appendix
- •10.5. Exercises
- •11.1. Introduction
- •11.2. Roadmap
- •11.2.1. Study protocol
- •11.2.2. Data collection
- •11.2.3. Statistical analysis plan
- •11.2.4. Clinical study report
- •11.3. A Plasmode Case Study
- •11.3.1. Research question
- •11.3.2. Study design
- •11.3.3. Causal estimand
- •11.3.4. Data
- •11.3.5. Statistical estimand
- •11.3.6. Estimator
- •11.3.7. Estimate
- •11.3.8. Sensitivity analysis
- •11.3.9. Evidence
- •11.4. Exercises
- •12. Applications of the Roadmap
- •12.1. Introduction
- •12.2. Applications to RCTs
- •12.2.1. RCTs with a single follow-up
- •12.2.2. Longitudinal RCTs
- •12.2.3. RCTs with time-to-event outcome
- •12.3. Applications to Cohort Studies
- •12.3.1. Cohort studies with a single follow-up
- •12.3.2. Externally controlled trials
- •12.3.3. Longitudinal cohort studies
- •12.4. Exercises
- •Bibliography
- •Index

96 Longitudinal Studies
R
Treatment 1
Treatment 0
Response?
Response?
Yes
No
Yes
No
R
R
R
R
Treatment 1
NULL
Treatment 1
Treatment 1+
Treatment 0
NULL
Treatment 0
Treatment 0+
FIGURE 5.5
An example of SMART design where ‘R’ indicates randomization
treatment 1 or treatment 0. Whether or not a patient responds, which may be
defined according to whether an outcome is improved by a certain percentage
from baseline, is ascertained at the end of stage 0. At stage 1, the responders to
treatment 1 are randomized to continuing treatment 1 or discontinuing treat-
ment 1 (denoted as NULL), while non-responders to treatment 1 are random-
ized to continuing treatment 1 or starting treatment 1 with add-on (denoted
as treatment 1+). Meanwhile, at stage 1, the responders to treatment 0 are
randomized to continuing treatment 0 or discontinuing treatment 0 (denoted
as NULL), while non-responders to treatment 0 are randomized to continuing
treatment 0 or starting treatment 0 with add-on (denoted as treatment 0+).
In this SMART design, two stages of randomization are conducted sequen-
tially, with the second stage of randomization depending on the outcome of
the first stage of randomization.
Then, we are able to compare the following eight DTRs: (1) start
with treatment 1 at t = 0 and continue treatment 1 at t =1;(2)
start with treatment 1 at t = 0, then continue treatment 1 if response or
enhance with treatment 1+ if non-response; (3) start with treatment 1 at
t = 0, then discontinue treatment 1 if response or continue treatment 1 if
non-response; (4) start with treatment 1 at t = 0, then discontinue treat-
ment 1 if response or enhance with treatment 1+ if non-response; (5) start
with treatment 0 at t = 0 and continue treatment 0 at t = 1; (6) start with
treatment 0 at t = 0, then continue treatment 0 if response or enhance with

Exercises 97
treatment 0+ if non-response; (7) start with treatment 0 at t =0,thendis-
continue treatment 0 if response or continue treatment 0 if non-response; (8)
start with treatment 0 at t = 0, then discontinue treatment 0 if response or
enhance with treatment 0+ if non-response.
d
(1)
= {d
(1)
0
=1,d
(1)
1
=1},
d
(2)
= {d
(2)
0
=1,d
(2)
1
(response) = 1; d
(2)
1
(non-response) = 1+},
d
(3)
= {d
(3)
0
=1,d
(3)
1
(response) = NULL; d
(3)
1
(non-response) = 1},
d
(4)
= {d
(4)
0
=1,d
(4)
1
(response) = NULL; d
(4)
1
(non-response) = 1+},
d
(5)
= {d
(5)
0
=0,d
(5)
1
=0},
d
(6)
= {d
(6)
0
=0,d
(6)
1
(response) = 0; d
(6)
1
(non-response) = 0+},
d
(7)
= {d
(7)
0
=0,d
(7)
1
(response) = NULL; d
(7)
1
(non-response) = 0},
d
(8)
= {d
(8)
0
=0,d
(8)
1
(response) = NULL; d
(8)
1
(non-response) = 0+}.
After the above DRTs are defined, we are able to compare any two given
DTRs, and furthermore, we are able to select the optimal DTR. Refer to
Tsiatis et al. (2020) for a unified discussion.
5.4 Exercises
Ex 5.1
In Subsection 5.1.1, we use the standardization strategy to translate causal
estimand θ
∗
ITT
into a statistical estimand θ
ITT
. Alternatively, we can use the

98 Longitudinal Studies
weighting strategy to do the job. We can show that, for T =2,
E
I(Z = j, Δ(1) = 0, Δ(2) = 0)Y
P{Z = j|X (0)}P{Δ(1) = 0|X (0),Z = j}P{Δ(2) = 0|X(0),Z = j, X(1)}
(a)
= E
I(Z = j,
Δ=0)Y
P{Z = j|X (0)}π
j
{X(0)}π
j
{X(0),X(1)}
(b)
=E
I(Z = j,
Δ=0)Y
z=j,δ=0
P{Z = j|X (0)}π
j
{X(0)}π
j
{X(0),X
z=j,δ(1)=0
(1)}
(c)
=E
E
I(Z = j,
Δ=0)Y
z=j,δ=0
P{Z = j|X (0)}π
j
{X(0)}π
j
{X(0),X
z=j,δ(1)=0
(1)}
X(0),W
z=j
(d)
= E
P{Z = j,
Δ=0|X(0),X
z=j,δ(1)=0
(1)}Y
z=j,δ=0
P{Z = j|X (0)}π
j
{X(0)}π
j
{X(0),X
z=j,δ(1)=0
(1)}
(e)
=E
)
Y
z=j,δ=0
*
.
Write down the reasons why equalities (a)–(e)hold.
Ex 5.2
In Subsection 5.1.1, based on the population shown in Table 5.1, we use the
standardization strategy to calculate the value of statistical estimand θ
ITT
.
Based on the same population, use the weighting strategy proposed in Ex 5.1
to calculate the value of the following statistical estimand θ
ITT
,
E
I(Z =1,
Δ=0)Y
P{Z =1|X(0)}π
j
{X(0)}π
j
{X}
− E
I(Z =0,
Δ=0)Y
P{Z =0|X(0)}π
j
{X(0)}π
j
{X}
.
Ex 5.3
In Section 5.2, we propose a method to discretize a time-to-event outcome into
a longitudinal binary outcome. Let Y be the time-to-death in months, C be
the censoring time in months, Y
∗
=min(Y, C) be the observed variable, and
Δ be the censoring indicator. Use this method to discretize five time-to-event
observations, with (Y
∗
i
, Δ
i
), i =1,...,5, being (5, 1), (7, 0), (3, 0), (4, 0), (4, 1),
into longitudinal binary observations, (Y
i
(1),Y
i
(2),...,Y
i
(7)), i =1,...,5,
respectively. Use NA to denote missing data.

6
Real-World Evidence Studies
6.1 RWE Studies
Randomized controlled clinical trials (RCTs) are the gold standard for eval-
uating the safety and efficacy of pharmaceutical drugs, as discussed in the
first five chapters. However, in many cases their costs, duration, limited gen-
eralizability, and ethical or technical feasibility have caused some to look for
real-world evidence (RWE) studies as alternatives.
In 2018, the Food and Drug Administration (FDA) released Framework for
FDA’s Real-World Evidence Program and provided definitions for real-world
data (RWD) and RWE:
“Real-World Data (RWD) are data relating to patient health status
and/or the delivery of health care routinely collected from a variety of
sources. Real-World Evidence (RWE) is the clinical evidence about the
usage and potential benefits or risks of a medical product derived from
analysis of RWD.”—Framework for FDA’s RWE Program
The framework for FDA’s RWE program covered many types of studies
that generate RWE, which are referred to as RWE studies in this book. In the
following three subsections, we describe three main types of RWE studies.
6.1.1 Pragmatic RCTs
As in the framework for FDA’s RWE program, we refer to the randomized
controlled clinical trials (RCTs) discussed in the previous five chapters as
traditional RCTs and refer to the other more flexible RCTs conducted in the
real-world setting as pragmatic RCTs. As pointed out by Thorpe et al. (2009),
traditional RCTs seek to answer the question, “Can this intervention work un-
der ideal conditions?”, whereas pragmatic RCTs seek to answer the question,
“Does this intervention work under usual conditions?”.
Thorpe et al. (2009) summarized the differences between pragmatic RCTs
and traditional RCTs in 10 domains. For simplicity, here we compare them
following the PROTECT checklist.
DOI: 10.1201/9781003433378-6 99

100 Real-World Evidence Studies
Population: In traditional RCTs, the population is defined via a set of
inclusion/exclusion criteria, whereas in pragmatic RCTs, the population is
often defined loosely, e.g., defined as a population consisting of all the pa-
tients who have certain of conditions. Response/Outcome: In traditional RCTs,
the response/outcome variables may need specialized training or testing,
whereas in pragmatic RCTs, the response/outcome variables can be assessed
under the usual conditions. Treatment/Exposure: The treatment/exposure
variable—a binary variable indicating either the investigative intervention or
the comparator—is defined with more restrictions in traditional RCTs than in
pragmatic RCTs. Often the comparator is placebo in traditional RCTs to en-
sure the trial is double-blinded, whereas the pragmatic RCTs are open-label,
in which the comparator is often “standard of care” or “usual practice.” Thus,
in pragmatic RCTs, there are bigger proportions of intercurrent events (ICEs)
such as treatment dropout and treatment switching than in traditional RCTs.
Counterfactual thinking: Since there are more missing data and ICEs in prag-
matic RCTs than in traditional RCTs, counterfactual thinking and causal
thinking play a more important role. Time: In traditional RCTs, patients are
followed with more frequent visits and more extensive data are collected than
in pragmatic RCTs conducted in routine practice.
Despite these differences between them, the causal inference methods that
have been developed for traditional RCTs are applicable to pragmatic RCTs.
Moreover, the causal inference methods are highly demanded for pragmatic
RCTs, given that pragmatic RCTs are conducted in routine practice and there
would be more missing data and ICEs.
6.1.2 Observational studies
In RCTs (traditional RCTs or pragmatic RCTs), the treatments are assigned
by the investigators and the treatment assignment function is known to the
investigators. Let Z be the treatment assignment variable, with Z = j stand-
ing for assignment to treatment j, j =1, 0. For example, in a 1:1 RCT,
P(Z =1)=0.5.
In observational studies, the treatments are determined by the patients
and their physicians. Let A be the treatment received by the patient and X
be the pre-treatment variables, based on which the treatment decision is made.
The treatment assignment function, g(1|X )=P(A =1|X), is referred to as
the propensity score function (Rosenbaum and Rubin 1983). The major differ-
ence between RCTs and observational studies is that in RCTs the propensity
score function is known whereas in observational studies it is unknown. In
observational studies, the investigators only know that A may depend on X,
without knowing which variables are in X and how A depend on X.
Observational studies include but are not limited to cohort studies, cross-
sectional studies, and case-control studies. In this chapter, we focus on cohort
studies. As discussed by Fang et al. (2020), there are two scenarios of cohort
studies: (1) prospective cohort studies that are designed to generate RWD;

Confounding Bias 101
and (2) retrospective cohort studies that are designed based on preexisting
RWD sources (e.g., electronic health records, claims data, or registries).
6.1.3 Externally controlled trials
Externally controlled trials (ECTs) are single-arm clinical trials with an ex-
ternal control group. To understand external controls, we need to understand
concurrent controls first. ICH E10 “Choice of control group in clinical trials”
provided the following definition.
“A concurrent control group is one chosen from the same population
as the test group and treated in a defined way as part of the same trial
that studies the test treatment, and over the same period of time. The
test and control groups should be similar with regard to all baseline and
on-treatment variables that could influence outcome, except for the study
treatment.”—ICH E10
Consequently, ICH E10 provided the definition of ECT.
“An externally controlled trial is one in which the control group con-
sists of patients who are not part of the same randomized study as the
group receiving the investigational agent; i.e., there is no concurrently
randomized control group. The control group is thus not derived from
exactly the same population as the treated population.”—ICH E10
There are two major sources of external controls. We may choose external
controls from the placebo arms of some completed RCTs and refer to them
as historical controls. We may choose external controls from real-world data
and referred as external RWD controls.
In the remaining of this chapter, we focus on observational studies and
ECTs. In the first five chapters, we used directed acyclic graphs (DAGs)
occasionally to display the relationship among variables. For observational
studies, we will rely more on DAGs. Revisit Subsection 1.4.2 for a brief intro-
duction to DAGs or refer to Pearl (2009) and Hern´an and Robins (2020) for
a comprehensive discussion on DAGs.
6.2 Confounding Bias
We often hear people say something like: “There is bias due to non-
randomization.” Is this true? What is bias?

102 Real-World Evidence Studies
ܣ
ܺ
ܻ
FIGURE 6.1
A well-conducted cohort study
We turn to ICH E9 “Statistical principles for clinical trials”, which pro-
vided the following definition of bias:
“As used in this guidance, the term ‘bias’ describes the systematic
tendency of any factors associated with the design, conduct, analysis and
interpretation of the results of clinical trials to make the estimate of a
treatment effect deviate from its true value.”—ICH E9
In the following three subsections, in three steps, we discuss the source of
“confounding bias”, using the example of a cohort study.
6.2.1 No unmeasured confounder
Assume there is one well-designed and well-conducted cohort study as dis-
played in Figure 6.1, assuming all the variables are measured without mea-
surement bias, there are no missing data, and the sample size is large enough.
In Figure 6.1, X is a vector of pre-treatment covariates, A is the treatment
variable (taking on 1 or 0), and Y is the outcome variable (either continuous
or binary).
Now let’s see whether there is any confounding bias if we do causal in-
ference in a certain way. When we speak of bias, we should first define the
estimand of interest. Without referring to an estimand, any statement about
bias is meaningless.
Causal estimand
Let Y
a=j
be the potential outcome that one patient would have had if the
patient had been treated by treatment j = 1 or 0.
Recall that in an RCT, we start with an intent-to-treat (ITT) population
of size N, which consists of data (Y
a=1
i
,Y
a=0
i
) in the potential world, i =
1,...,N, which are assumed to be independent and identically distributed

Confounding Bias 103
(i.i.d.) drawn from a super-population. That is, we say (Y
a=1
i
,Y
a=0
i
), i =
1,...,N, are i.i.d. with (Y
a=1
,Y
a=0
).
Also, recall the population-sample duality discussed in Chapter 1. There-
fore, in an observational study, we start with a sample of size n, which con-
sists of data (Y
a=1
i
,Y
a=0
i
) in the potential world, i =1,...,n, which are
assumed to be i.i.d. drawn from a population. We could construct such a
population conceptually following the same way by which we construct the
super-population in an RCT. Thus, we say (Y
a=1
i
,Y
a=0
i
), i =1,...,n,are
i.i.d. with (Y
a=1
,Y
a=0
).
After we understand the above relationship between sample and popula-
tion in an observational study, we are able to define an estimand of interest.
For example, we are interested in the following causal estimand,
θ
∗
ATE
= E(Y
a=1
) − E(Y
a=0
), (6.1)
where the expectation is over the population. As suggested by the subscript,
this causal estimand is referred to as the average treatment effect (ATE),
which measures the difference between the mean of the potential outcomes if
all the patients in the population had been treated by treatment 1 and that
if all the patients in the population had been treated by treatment 0.
Identification
In the potential world, potential outcomes (Y
a=1
i
,Y
a=0
i
), i =1,...,n,are
i.i.d. with (Y
a=1
,Y
a=0
). In the real world, observations (X
i
,A
i
,Y
i
), i =
1,...,n, are i.i.d. with (X, A, Y ).
We make the following three identifiability assumptions.
1. The consistency assumption:
Y = Y
a=j
if A = j, for j =1, 0; (6.2)
2. The exchangeability assumption:
(Y
a=0
,Y
a=1
)
|=
A|X; (6.3)
3. The positivity assumption:
P(A = j|X = x) > 0, for j =1, 0; x ∈ supp(X). (6.4)
We are familiar with the consistency assumption and the positivity as-
sumption, since we have seen them several times in some of the previous chap-
ters. The exchangeability assumption is also known as the no-unmeasured-
confounding (NUC) assumption (Hern´an and Robins 2020). With the help
of the structural causal model (SCM) in Appendix 6.5, we show that the
exchangeability assumption is implied by the DAG in Figure 6.1.

104 Real-World Evidence Studies
Under these assumptions, we can translate the causal estiamnd into a sta-
tistical estimand. There are two strategies for this task: the standardization
strategy and the weighting strategy (Fang 2020; Hern´an and Robins 2020).
The standardization strategy
Under the identifiability assumptions, we can show that
θ
∗
ATE
= E(Y
a=1
) − E(Y
a=0
)
(a)
= E[E(Y
a=1
|X)] − E[E(Y
a=0
|X)]
(b)
= E[E(Y
a=1
|A =1,X)] −E[E(Y
a=0
|A =0,X)]
(c)
= E[E(Y |A =1,X)] − E[E(Y |A =0,X)],
where (a) holds using the law of iterated expectation, (b) holds under the
exchangeability assumption and the positivity assumption, and (c) holds under
the consistency assumption. Thus, we translate the causal estimand to the
following statistical estimand under the identifiability assumptions,
θ
ATE
= E[E(Y |A =1,X)] − E[E(Y |A =0,X)]. (6.5)
As in Hern´an and Robins (2020), we call the above identification method
the standardization strategy. The other method is the weighting strategy.
The weighting strategy
Under the identifiability assumptions, we can show that
E
I(A = a)
P (A = a|X)
Y
(a)
= E
I(A = a)
P(A = a|X)
Y
a
(b)
= E
E
I(A = a)
P(A = a|X)
Y
a
X
(c)
= E
E
I(A = a)
P(A = a|X)
X
E{Y
a
X}
(d)
= E [E{Y
a
|X}]
(e)
= E(Y
a
),
where the left-hand-side of (a) is well-defined under the positivity assumption,
the right-hand-side of (a) is using the consistency assumption, (b) holds using
the law of iterated expectations, (c) holds under the exchangeability assump-
tion, (d) holds because E{I(A = a|X)} = P(A = a|X), and (e) holds using
the law of iterated expectations Thus, under the identifiability assumptions,
θ
∗
ATE
= E
I(A =1)
P (A =1|X)
Y
− E
I(A =0)
P (A =0|X)
Y
.

Confounding Bias 105
Thus, by the weighting strategy, we translate the causal estimand to the
following statistical estimand,
θ
ATE
= E
I(A =1)
P (A =1|X)
Y
− E
I(A =0)
P (A =0|X)
Y
, (6.6)
which is equivalent to the statistical etsimand defined in (6.5) under the same
three identifiability assumptions.
Confounding bias
After we define a causal estimand and accomplish the task of identification,
we are ready to discuss the issue of confounding bias.
The causal estimand, θ
∗
ATE
, is our ultimate target, while the statistical
estimand, θ
ATE
, is our realistic target.
In the next three chapters, we will discuss methods for constructing con-
sistent estimators of θ
ATE
. We denote such estimator as
θ
ATE ,n
.Wesaythat
θ
ATE ,n
is a consistent estimator of θ
ATE
if
θ
ATE ,n
p
−→ θ
ATE
, as n →∞.
Sincewehaveshownthatθ
∗
ATE
= θ
ATE
under the identifiability assump-
tions, we conclude that if
θ
ATE ,n
is a consistent estimator of θ
ATE
then it is
also a consistent estimator of θ
∗
ATE
. Thus, if we use consistent estimator
θ
ATE ,n
to estimate θ
∗
ATE
, we say that there is no confounding bias under the iden-
tifiability assumptions.
However, what happens if we use the following naive estimator,
θ
n
=
n
i=1
I(A
i
=1)Y
i
n
i=1
I(A
i
=1)
−
n
i=1
I(A
i
=0)Y
i
n
i=1
I(A
i
=0)
,
to estimate θ
∗
ATE
? In this case, we say that there is statistical bias if we use
θ
n
to estimate statistical estimand θ
ATE
, because this naive estimator is a
consistent estimator of the following estimand,
θ
= E(Y |A =1)−E(Y |A =0),
which is not equal to θ
ATE
. Furthermore, we say that there is confounding
bias if we use
θ
n
to estimate causal estimand θ
∗
ATE
, because θ
is not equal to
θ
∗
ATE
even under the identifiability assumptions.
6.2.2 Unmeasured confounders
As displayed in Figure 6.2, assume that there is a vector of unmeasured con-
founders, denoted as U, along with measured variables X, A,andY .
Let’s see whether there is any confounding bias if we do causal inference in
a certain way. Again, when we speak of bias, we should first define the estimand
Соседние файлы в папке Библиотека им академика М.И. Перельмана
