Добавил:
Sekretar
kiopkiopkiop18@yandex.ru
t.me/Prokururor I Вовсе не секретарь, но почту проверяю
Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз:
Предмет:
Файл:Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5431_Библиотеки_им_академика_М_И_Перельмана.pdf
X
- •Contents
- •Preface
- •1. Introduction
- •1.1. Central Questions
- •1.2. Potential Outcomes
- •1.3. Estimand
- •1.3.1. The PROTECT checklist
- •1.3.4. Internal validity and external validity
- •1.4. Probability and Statistics
- •1.4.1. Probability
- •1.4.2. Directed acyclic graphs
- •1.4.3. Statistics
- •1.3.2. Estimand for a given population
- •1.3.3. Estimand for a given super-population
- •1.5. Exercises
- •2.1. Randomization and Blinding
- •2.2. Estimand
- •2.2.1. Causal estimand
- •2.2.2. Statistical estimand
- •2.3. Estimator
- •2.3.1. Expectation of the estimator
- •2.3.2. Variance of the estimator
- •2.3.3. Statistical inference
- •2.4. Common Types of Randomization
- •2.4.1. Simple randomization
- •2.4.2. Block randomization
- •2.5. Exercises
- •3. Missing Data Handling
- •3.1. Missing Data
- •3.2.1. Scenario one
- •3.2.2. Scenario two
- •3.4. Sources of Missing Data
- •3.4.1. Intercurrent events
- •3.4.2. Missing data that are consequences of ICEs
- •3.4.3. Missing data that are not consequences of ICEs
- •3.5. Appendix
- •3.6. Exercises
- •4. Intercurrent Events Handling
- •4.1. Five Strategies
- •4.1.1. The treatment policy strategy
- •4.1.2. The hypothetical strategy
- •4.1.3. The composite variable strategy
- •4.1.4. The while on treatment strategy
- •4.1.5. The principal stratum strategy
- •4.2. Combinations of Strategies
- •4.3. Time-to-event Outcome
- •4.3.1. Censoring
- •4.3.2. The treatment policy strategy
- •4.3.3. The hypothetical strategy
- •4.3.4. The composite variable strategy
- •4.3.5. The while on treatment strategy
- •4.3.6. The principal stratum strategy
- •4.3.7. The competing risk strategy
- •4.4. Sample Size Calculation
- •4.4.1. The treatment policy strategy
- •4.4.2. The hypothetical strategy
- •4.4.3. The composite variable strategy
- •4.4.4. The while on treatment strategy
- •4.4.5. The principal stratum strategy
- •4.5. Exercises
- •5. Longitudinal Studies
- •5.1. Continuous or Binary Outcome
- •5.2. Time-to-event Outcome
- •5.3. Treatment Regimes
- •5.3.1. Dynamic treatment regimes
- •5.3.2. SMART design
- •5.4. Exercises
- •6. Real-World Evidence Studies
- •6.1. RWE Studies
- •6.1.1. Pragmatic RCTs
- •6.1.2. Observational studies
- •6.1.3. Externally controlled trials
- •6.2. Confounding Bias
- •6.2.1. No unmeasured confounder
- •6.2.2. Unmeasured confounders
- •6.2.3. Proxy variables
- •6.3. Longitudinal Cohort Studies
- •6.3.1. Causal estimand
- •6.4. Externally Controlled Trials
- •6.4.1. Causal estimand
- •6.5. Appendix
- •6.6. Exercises
- •7.1. Introduction
- •7.2. M-estimation
- •7.2.1. M-estimator
- •7.2.2. Asymptotic linearity
- •7.2.3. Regularity
- •7.3. G-computation Estimator
- •7.3.1. Plug-in estimator
- •7.3.2. MLE
- •7.3.3. Asymptotic variance
- •7.4. Inverse Probability Weighted Estimator
- •7.4.1. IPW estimator
- •7.4.2. Asymptotic variance
- •7.5. Augmented Inverse Probability Weighted Estimator
- •7.5.1. A class of estimators
- •7.5.2. Asymptotic variances
- •7.6. Exercises
- •8.1. Semiparametric Statistics
- •8.1.1. Semiparametric estimators
- •8.1.2. Super learner
- •8.1.3. Semiparametric estimators based on super learner
- •8.2. Asymptotic Variances of Semiparametric Estimators
- •8.2.1. Parametric submodels
- •7.5.3. AIPW estimator
- •7.5.4. Double robustness
- •8.2.2. The fundamental theorem of regularity
- •8.2.5. Double robustness of AIPW-SL estimator
- •8.3. The Targeted Learning Framework
- •8.3.1. Mini-roadmap
- •8.3.2. TMLE
- •8.3.3. Double robustness
- •8.4.3. Missing data due to analysis dropout
- •8.5. Discussion
- •8.5.1. How to select covariates?
- •8.5.2. How to handle missing covariates?
- •8.5.3. How to use TMLE for RCTs?
- •8.5.4. How to implement TMLE?
- •8.6. Exercises
- •9.1. Longitudinal Cohort Studies
- •9.1.1. Causal estimand
- •9.1.4. LTMLE
- •9.1.5. ATE estimand
- •9.2. Missing Data
- •9.2.1. Monotone missing
- •9.2.2. Non-monotone missing
- •9.3. Implementation
- •9.4. Exercises
- •10. Sensitivity Analysis
- •10.1. Introduction
- •10.2.1. The consistency assumption
- •10.2.2. The exchangeability assumption
- •10.2.3. The positivity assumption
- •10.3. Sensitivity Analysis for the MAR Assumption
- •10.3.1. A class of reference-based imputation models
- •10.3.2. Sequential modeling
- •10.4. Appendix
- •10.5. Exercises
- •11.1. Introduction
- •11.2. Roadmap
- •11.2.1. Study protocol
- •11.2.2. Data collection
- •11.2.3. Statistical analysis plan
- •11.2.4. Clinical study report
- •11.3. A Plasmode Case Study
- •11.3.1. Research question
- •11.3.2. Study design
- •11.3.3. Causal estimand
- •11.3.4. Data
- •11.3.5. Statistical estimand
- •11.3.6. Estimator
- •11.3.7. Estimate
- •11.3.8. Sensitivity analysis
- •11.3.9. Evidence
- •11.4. Exercises
- •12. Applications of the Roadmap
- •12.1. Introduction
- •12.2. Applications to RCTs
- •12.2.1. RCTs with a single follow-up
- •12.2.2. Longitudinal RCTs
- •12.2.3. RCTs with time-to-event outcome
- •12.3. Applications to Cohort Studies
- •12.3.1. Cohort studies with a single follow-up
- •12.3.2. Externally controlled trials
- •12.3.3. Longitudinal cohort studies
- •12.4. Exercises
- •Bibliography
- •Index

76 Intercurrent Events Handling
Example 1: We continue the discussion of Example 0. Furthermore, we
assume that r
1
= 15% and r
0
= 10%. By the treatment policy strategy, the
original comparison of p
1
=0.5vs.p
0
=0.3 is diluted as the comparison of
p
1
=0.5(1 −0.15) +0.3(0.15) = 0.47 vs. p
0
=0.3. For this diluted comparison,
therequiredsamplesizeisN
∗
=2×128 = 256.
4.4.2 The hypothetical strategy
Assume that there is only one ICE to be handled by the hypothetical strategy.
Assume that the proportion of subjects who are expected to have ICE in arm
Z = j is r
j
, j =1, 0. By the hypothetical strategy, the estimand of interest is
θ
∗
PP
and the proportion of ICE in each arm is the proportion of missing data
in each arm. Thus, the conventional approach is appropriate.
Example 2: We continue the discussion of Example 0. Furthermore, we
assume that r
1
= 15% and r
0
= 10%. Without adjusting for missing data or
ICEs, the required sample size is N = 188. Using the conventional approach,
the required sample size is adjusted as N
∗
= N/[1 − (r
1
+ r
0
)/2] = 188/(1 −
0.125) = 216. Note that the sample size is rounded up to an even number such
that it can be divided into two arms.
4.4.3 The composite variable strategy
Assume that there is only one ICE to be handled by the composite variable
strategy. Assume that the proportion of subjects who are expected to have
ICE in arm Z = j is r
j
, j =1, 0. If Y is binary (say, 1 stands for failure and 0
stands for success), define a new outcome variable
&
Y ,suchas
&
Y =1ifY =1
or an ICE occurs. Then, we have
P(
&
Y
z=j
=1)=(1− r
j
)P(Y
z=j
=1)+r
j
,j =1, 0.
Example 3: We continue the discussion of Example 0. In this example, Y is
binary with 1 standing for failure, so the comparison becomes q
1
=1−p
1
=0.5
vs. q
0
=1−p
0
=0.7. Furthermore, we assume that r
1
= 15% and r
0
= 10%.
By the composite variable strategy, the original comparison of q
1
=0.5vs.
q
0
=0.7 becomes the new comparison of q
1
=[0.5(1 −0.15)+ 0.15] = 0.575 vs.
q
0
=0.7(1 − 0.1) + 0.1=0.73. For the new comparison, the required sample
size is N
∗
=2×147 = 294.
4.4.4 The while on treatment strategy
Assume that there is only one ICE to be handled by the while on treatment
strategy. Assume that the proportion of subjects who are expected to have
ICE in arm Z = j is r
j
, j =1, 0. By the while on treatment strategy, we
define the new outcome variable
&
Y as the outcome variable measured at the
time immediately prior to the ICE occurrence if an ICE occurs. Furthermore,

Exercises 77
assume that, in the treatment arm, the treatment effect at a given time is
proportional to the treatment duration up to that time, while in the control
arm, the treatment effect is the same regardless of the ICE occurrence. More-
over, assume the time of ICE occurrence is following a uniform distribution
between 0 and T . Under these assumptions, we have
P(
&
Y
z=1
=1)=(1− r
1
)P(Y
z=1
=1)+r
1
P(Y
z=1
=1)+P(Y
z=0
=1)
(
2.
Example 4: We continue the discussion of Example 0. Furthermore, we
assume that r
1
= 15% and r
0
= 10%. By the while on treatment strategy, the
original comparison of p
1
=0.5vs.p
0
=0.3 is diluted as the new comparison
of p
1
=(1−0.15)(0.5) +0.15(0.5+0.3)/2=0.485 vs. p
0
=0.3. For this diluted
comparison, the required sample size is N
∗
=2×109 = 218.
4.4.5 The principal stratum strategy
Assume that there is only one ICE to be handled by the principal stratum
strategy. Assume that the proportion of subjects who are expected to have
ICE in arm Z = j is r
j
, j =1, 0. Assume we are interested in the principal
stratum PS
00
= {E
z=1
=0,E
z=0
=0}.Wehave
P(E
z=1
=0,E
z=0
=0)=1−P(E
z=1
=1orE
z=0
=1)
≥ 1 − [P(E
z=1
=1)+P(E
z=0
=1)]
=1−(r
1
+ r
0
).
Thus, the required sample size is N
∗
= N/(1 − r
1
− r
0
).
Example 5: We continue the discussion of Example 0. Furthermore, we
assume that r
1
= 15% and r
0
= 10%. Thus, if the principal stratum PS
00
is of
interest, the needed sample size is N
∗
= N/(1−0.15 −0.1) = 188/0.75 = 252.
4.5 Exercises
In R, generate a data frame that has 5 columns (subject ID SID,covariateX,
treatment assignment Z being0or1,ICEindicatorE being0or1,andobserved
incomplete outcome Yobs) and 100 rows, using the following R codes:
1 set. seed (6)
2 SID <- 1:100
3 X <- rno rm( n=100 , mean =0 , sd=2)
4 Z <- rbinom (n=100 , size=1, prob=0.5)
5
6 E1 <- rbinom ( n=100 , size =1, prob=exp ( -2.5+ X) /(1+ exp ( -2.5+X )))
7 E0 <- rbinom ( n=100 , size =1, prob=exp ( -3.0+ X) /(1+ exp ( -3.0+X )))
8
9 E<-Z*E1+(1-Z)*E0

78 Intercurrent Events Handling
10 mean( E) # ICE rate
11
12 Y00<- rnorm (n=100 , mean =0, sd =1)
13 Y01<- rnorm (n=100 , mean = -0.3, sd=1)
14 Y11<- rnorm (n=100 , mean =1, sd =1)
15 Y10< - rnorm ( n=100 , mean =0.5 , sd=1)
16
17 Yobs <- (1-Z )*(1- E)* Y00 +(1-Z )*E*Y01+Z*(1-E)*Y10+Z*E*Y11
18 summary( Yobs ) # summary of obs erved outcome
19
20 dataset 6 <- data. frame (SID =SID , X =X, Z=Z, E=E , Y= Yobs )
Note that in the above population, E1 is E
z=1
, E0 is E
z=0
, Y00 is Y
z=0,e=0
,
Y01 is Y
z=0,e=1
, Y10 is Y
z=1,e=0
,andY11 is Y
z=1,e=1
. The following three
exercises are based on this population.
Ex 4.1
Use the treatment policy strategy to handle ICE E and find the value of the
following estimand:
θ
∗
TP
=
1
N
N
i=1
Y
z=1
i
− Y
z=0
i
.
Ex 4.2
Use the hypothetical strategy to handle ICE E and find the value of the fol-
lowing estimand:
θ
∗
H
=
1
N
N
i=1
)
Y
z=1,e=0
i
− Y
z=0,e=0
i
*
.
Ex 4.3
Use the principal stratum strategy to handle ICE E. Find the value of the
following estimand:
θ
∗
PS
00
=
1
#(PS
00
)
i∈PS
00
Y
z=1
i
− Y
z=0
i
,
where #(PS
00
) is the size of PS
00
.

5
Longitudinal Studies
5.1 Continuous or Binary Outcome
In the first four chapters, we focused on studies where there is only one follow-
up time and there is no time-dependent covariates, as illustrated in Figure 1.2.
In this chapter, we will discuss longitudinal studies. In longitudinal studies
with time-dependent treatments or intercurrent events (ICEs), we need to
incorporate time explicitly in the definition of treatment (Hern´an and Robins
2020).
Assume that there is one longitudinal study starting from baseline t =0,
along with follow-up visits, t =1,...,T, as illustrated in Figure 5.1. Assume
that the primary outcome variable, Y , which is either continuous or binary, is
the outcome variable measured at the final visit T .
At baseline t =0,letZ be the treatment assignment, with Z = 1 and
Z = 0 standing for assignment to treatment 1 and treatment 0, respectively,
either by complete randomization or stratified randomization conditional on
factor S.LetX(0) be the vector containing baseline characteristics, including
stratification factor S,ifany.
To describe the time-dependent treatment explicitly, we introduce the no-
tation A(t), t =0,...,T −1. Between baseline t = 0 and follow-up visit t =1,
let A(0) be the treatment that is actually taken by the patient. Some possible
values that A(t),t=0,...,T − 1, may take on are:
• 1: Treatment 1,
• 0: Treatment 0,
• NULL: No treatment,
• 1+: Treatment 1 with some add-on,
• 0+: Treatment 0 with some add-on,
• 2: Alternative treatment different from 1 and 0.
At follow-up visit t =1,letX(1) be the vector containing time-dependent
covariates and/or the intermediate outcome variable measured at t =1.
Similarly, let A(t − 1) be the treatment that is actually taken between
visit t − 1 and visit t,andletX(t) be the vector containing time-dependent
covariates and/or the intermediate outcome variable measured at visit t,f
or
t =1,...,T − 1.
DOI: 10.1201/9781003433378-5 79

80 Longitudinal Studies
=0
Baseline Endpoint
=
…
= −1
=1 =2
FIGURE 5.1
Baseline and follow-up visits
Finally, let A(T −1) be the treatment that is actually taken between visit
T − 1 and visit T ,andletY be the primary outcome measured at visit T .
Let
A(t)=(A(0),...,A(t)) be the actually taken treatment sequence up
to t.Let
X(t)=(X(0),...,X(t)) be the vector consisting of all the ob-
served history—including baseline covariates, time-dependent covariates, and
intermediate outcomes—up to time t, t =0,...,T − 1. In particular, define
A =(A(0),...,A(T − 1)) and X =(X(0),...,X(T − 1)).
5.1.1 The intent-to-treat effect
Causal estimand
Let Δ(t) be the indicator of analysis dropout at time t, t =1,...,T.For
the purpose of demonstration, we assume that the analysis dropout pattern
is monotone; that is, if Δ(t)=1,thenΔ(t
)=1fort
= t +1,...,T.The
methods to be discussed in this chapter can be applied to both monotone
patterns and non-monotone patterns. Let
Δ = (Δ(1),...,Δ(T )) and Δ(t)=
(Δ(1),...,Δ(t)), t =1,...,T.
Let Y
z=j,δ=0
be the potential outcome had the patient been assigned to
treatment j and the data been collected throughout the study period, j =1, 0.
Similar to primary outcome Y , X(t), t =1,...,T − 1, are post-treatment
initiation variables, and therefore we can define their potential outcomes. Let
X
z=j,δ(t)=0
(t) be the potential outcome had the patient been assigned to treat-
ment j and the data been collected up to time t, j =1, 0. Here, by convention,
0 is a vector of all zeros and its dimension is the same as the one of δ(t). Thus,
define the set of all potential outcomes as
W
z=j
= {X
z=j,δ(t)=0
(t),t=1,...,T − 1; Y
z=j,δ=0
},j =1, 0. (5.1)
Following the discussion in Chapter 4, as illustrated in Figure 5.2, if we use
the treatment policy strategy to handle treatment dropouts if any, along with

Continuous or Binary Outcome 81
0
ܶ =5
ICE occurrence
Measured outcome Missing outcome
Treatment policy strategy
1
2 3 4
0
ܶ =5
1
2 3 4
݅ =1
݅ =2
݅ =3
݅ =4
݅ =5
FIGURE 5.2
The treatment policy strategy to handle treatment dropouts
the hypothetical strategy to handle analysis dropouts (the source of missing
data), we are interested in the intent-to-treat (ITT) estimand,
θ
∗
ITT
= E
Y
z=1,δ=0
− Y
z=0,δ=0
. (5.2)
Statistical estimand
In order to translate the above causal estimand into a statistical estimand, we
make the following three assumptions:
1. The consistency assumption:
X(t)=X
z=j,δ(t)=0
(t), if Z = j, Δ(t)=0,
Y = Y
z=j,δ=0
, if Z = j, Δ=0;
2. The missing at random (MAR) assumption:
(W
z=1
,W
z=0
)
|=
Δ(1)|(Z, X(0)) ,
(W
z=1
,W
z=0
)
|=
Δ(t)|
Z, X(t − 1), Δ(t − 1) = 0
,t=1,...,T − 1;
3. The positivity assumption:
P
Δ(t)=0|Z = j,
X(t − 1) = x(t − 1), Δ(t − 1) = 0
> 0,
for t =1,...,T − 1; j =1, 0;
x(t − 1) ∈ supp(X (t − 1)).
Under these three assumptions, we will show that
E
Y
z=j,δ=0
= E
X∼F
j
E(Y |
X,Z = j, Δ=0)
, (5.3)

82 Longitudinal Studies
where the outer expectation on the right-hand-side (RHS) is over
X ∼F
j
and
F
j
is a distribution function of (X(0),...,X(T − 1)) defined as
P
F
j
(X(0) = x(0),...,X(T − 1) = x(T − 1))
=P{X(0) = x(0)}
T −1
+
t=1
P
X(t)=x(t)|X(t − 1) = x(t − 1),Z = j, Δ(t)=0
.
For a study where T = 1, (5.3) becomes
E
Y
z=j,δ(1)=0
= E
X(0)
[E(Y |X(0),Z = j, Δ(1) = 0)] ,
which has been shown in Chapter 3.
If (5.3) is proved, then we can complete the task of translating the causal
estimand into a statistical estimand:
θ
∗
ITT
= E
Y
z=j,δ=0
− E
Y
z=j,δ=0
= E
X∼F
1
E(Y |
X,Z =1, Δ=0)
− E
X∼F
0
E(Y |
X,Z =0, Δ=0)
θ
ITT
. (5.4)
Before we provide formal proof for (5.3), let’s understand the meaning
of RHS of (5.3) using a numerical example where T = 2. Table 5.1 shows
the observed data of a population of 20 subjects—imagine that one subject
represents one million subjects.
Based on the data in Table 5.1, the complete-case subset consists of sub-
jects 4–10, 12–13, 15, 17, and 10–20. Using the data from the complete-case
subset, we obtain
E(Y |X(0) = 1,Z =1,X(1) = 1,
Δ=0) = 1/2,
E(Y |X(0) = 1,Z =1,X(1) = 0,
Δ=0) = 1,
E(Y |X(0) = 0,Z =1,X(1) = 1,
Δ=0) = 0,
E(Y |X(0) = 0,Z =1,X(1) = 0,
Δ=0) = 1,
E(Y |X(0) = 1,Z =0,X(1) = 1,
Δ=0) = 1/2,
E(Y |X(0) = 1,Z =0,X(1) = 0,
Δ=0) = 0,
E(Y |X(0) = 0,Z =0,X(1) = 1,
Δ=0) = 1/2,
E(Y |X(0) = 0,Z =0,X(1) = 0,
Δ=0) = 0.
Based on all the observed data, we obtain distributions F
1
and F
0
:
P
F
1
(X(0) = 1,X(1) = 1)
= P(X(0) = 1)P(X(1) = 1|X(0) = 1,Z =1, Δ(1) = 0)
=(12/20)(3/4) = 9/20,

Continuous or Binary Outcome 83
TABL E 5 .1
Data from a population
SID X(0) ZX(1) Y
1111NA
211NA NA
311NA NA
41101
51110
61111
71000
81000
91011
10 1 0 1 0
11 1 0 NA NA
12 1 0 0 0
13 0 1 1 0
14 0 1 1 NA
15 0 1 0 1
16 0 1 0 NA
17 0 0 1 0
18 0 0 NA NA
19 0 0 1 1
20 0 0 0 0
P
F
1
(X(0) = 1,X(1) = 0)
= P(X(0) = 1)P(X(1) = 0|X(0) = 1,Z =1, Δ(1) = 0)
=(12/20)(1/4) = 3/20,
P
F
1
(X(0) = 0,X(1) = 1)
= P(X(0) = 0)P(X(1) = 1|X(0) = 0,Z =1, Δ(1) = 0)
=(8/20)(2/4) = 1/5,
P
F
1
(X(0) = 0,X(1) = 0)
= P(X(0) = 0)P(X(1) = 0|X(0) = 0,Z =1, Δ(1) = 0)
=(8/20)(2/4) = 1/5;
and
P
F
0
(X(0) = 1,X(1) = 1)
= P(X(0) = 1)P(X(1) = 1|X(0) = 1,Z =0, Δ(1) = 0)
=(12/20)(2/5) = 6/25,

84 Longitudinal Studies
P
F
0
(X(0) = 1,X(1) = 0)
= P(X (0) = 1)P(X(1) = 0|X (0) = 1,Z =0, Δ(1) = 0)
=(12/20)(3/4) = 9/25,
P
F
0
(X(0) = 0,X(1) = 1)
= P(X (0) = 1)P(X(1) = 1|X (0) = 0,Z =0, Δ(1) = 0)
=(8/20)(2/3) = 4/15,
P
F
0
(X(0) = 0,X(1) = 0)
= P(X (0) = 1)P(X(1) = 0|X (0) = 0,Z =0, Δ(1) = 0)
=(8/20)(1/3) = 2/15.
Thus, we have
E
X∼F
1
E(Y |
X,Z =1, Δ=0)
=9/20(1/2) + 3/20(1) + 1/5(0) + 1/5(1)
=23/40,
E
X∼F
0
E(Y |
X,Z =0, Δ=0)
=6/25(1/2) + 9/25(0) + 4/15(1/2) + 2/15(0)
=19/75.
Formula like (5.3) is usually referred to as the g-computation algorithm,
also known as the g-formula (Hern´an and Robins 2020; Robins 1986).
“The ‘g’ stands for ‘generalized’.”—Hern´an and Robins (2020)
Now we are ready to prove g-formula (5.3). Although it is tedious, the
proof is important for us to understand why we are able to translate causal
estimand to statistical estimand in a longitudinal study. For simplicity, we
consider T = 2 and consider the setting where all the variables are discrete.
The result can be easily extended to any T ≥ 2 and the setting where there is
a mixture of continuous variables with density functions and discrete variables
with probability mass functions.
Proof of g-formula (5.3)
:ForT =2,wehave
P{Y = y|X(0) = x(0),Z = j, Δ(1) = 0,X(1) = x(1), Δ(2) = 0}
(a)
= P{Y
z=j,δ(2)=0
= y|X(1) = x(1),Z = j, Δ(1) = 0, Δ(2) = 0}
(b)
=P{Y
z=j,δ(2)=0
= y|X(0) = x(0),X(1) = x(1),Z = j, Δ(1) = 0}
(c)
=P{Y
z=j,δ(2)=0
= y|X(0) = x(0),X
z=j,δ(1)=0
(1) = x(1),Z = j, Δ(1) = 0}
(d)
=
P{Y
z=j,δ(2)=0
= y, X
z=j,δ(1)=0
(1) = x(1)|X (0) = x(0),Z = j, Δ(1) = 0}
P{X
z=j,δ(1)=0
(1) = x(1)|X (0) = x(0),Z = j, Δ(1) = 0}
(e)
=
P{Y
z=j,δ(2)=0
= y, X
z=j,δ(1)=0
(1) = x(1)|X (0) = x(0),Z = j}
P{X
z=j,δ(1)=0
(1) = x(1)|X (0) = x(0),Z = j}
(f )
=
P{Y
z=j,δ(2)=0
= y, X
z=j,δ(1)=0
(1) = x(1)|X (0) = x(0)}
P{X
z=j,δ(1)=0
(1) = x(1)|X (0) = x(0)}
(g)
= P{Y
z=j,
δ(2)=0
= y|X(0) = x(0),X
z=j,δ(1)=0
(1) = x(1)},

Continuous or Binary Outcome 85
where (a) holds under the consistency assumption, (b) holds under the MAR
assumption, (c) holds under the consistency assumption again, (d) holds us-
ing the relationship among joint, marginal, and conditional distributions, (e)
holds applying the MAR assumption in both the numerator and denominator,
(f) holds because of the randomization, and (g) holds using the relationship
among joint, marginal, and conditional distributions again.
We also have
P{X(1) = x(1)|X(0) = x(0),Z = j, Δ(1) = 0}
(a)
= P{X
z=j,δ(1)=0
(1) = x(1)|X(0) = x(0),Z = j, Δ(1) = 0}
(b)
=P{X
z=j,δ(1)=0
(1) = x(1)|X(0) = x(0),Z = j}
(c)
=P{X
z=j,δ(1)=0
(1) = x(1)|X(0) = x(0)},
where (a) holds under the consistency assumption, (b) holds under the MAR
assumption, and (c) holds because of the randomization.
Combining the above two results, we have
P{X(0) = x(0),X
z=j,δ(1)=0
(1) = x(1),Y
z=j,δ(2)=0
= y}
=P{X(0) = x(0)}×P{X
z=j,δ(1)=0
(1) = x(1)|X(0) = x(0)}
× P{Y
z=j,δ(2)=0
= y|X(0) = x(0),X
z=j,δ(1)=0
(1) = x(1)}
=P{X(0) = x(0)}×P{X(1) = x(1)|X(0) = x(0),Z = j, Δ(1) = 0}
× P{Y = y|X(0) = x(0),Z = j, Δ(1) = 0,X(1) = x(1), Δ(2) = 0}.
Writing the above finding in terms of expectation, we have
E
Y
z=j,δ(2)=0
= E
X(1)∼F
j
E(Y |
X(1),Z = j, Δ(2) = 0)
,
where the outer expectation on the right-hand-side is over
X(1) ∼F
j
and F
j
is a distribution function of (X(0),X(1)) defined as
P{X(0) = x(0)}×P{X(1) = x(1)|X(0) = x(0),Z = j, Δ(1) = 0}.
This completes the proof for T =2.
Estimator
Define the following regression function,
Q
Δ=0
(Z, X)=E(Y |X, Z =1, Δ=0). (5.5)
Then the statistical estimand defined in (5.4) can be written as
θ
ITT
= E
X∼F
1
Q
Δ=0
(1, X)
− E
X∼F
0
Q
Δ=0
(0, X)
.
Соседние файлы в папке Библиотека им академика М.И. Перельмана
