Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5431_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
10.10.2026
Размер:
10 Мб
Скачать
☆
166 The Art of Estimation (II): TMLE
Ex 8.2
Use R function “tmle” to obtain an AIPW estimate for statistical estimand
θ
ATE
, fitting a linear model for outcome regression function Q and a logistic
model for propensity score function g.
Ex 8.3
Use R function “tmle” to obtain a TMLE estimate for statistical estimand
θ
ATE
, using the default SL library for Q and g, respectively.
Ex 8.4
Use R function “tmle” to obtain a TMLE estimate for statistical estimand
θ
ATE
, using the default SL library for Q and g, respectively.
9
The Art of Estimation (III):
LTMLE

9.1 Longitudinal Cohort Studies

In Chapter 8, we described the targeted learning framework (van der Laan
and Rose 2011) for cohort studies; in particular, we reviewed the targeted
maximum likelihood estimator or targeted minimum loss estimator (TMLE).
In this chapter, we will describe the longitudinal TMLE (LTMLE) for longi-
tudinal cohort studies discussed in Chapter 6.

9.1.1 Causal estimand

As in Chapter 6, let A(t)=(A(0),...,A(t)) be the observed treatment se-
quence up to t,andlet
X(t)=(X(0),...,X(t)) be the vector consisting of all
the observed history up to time t including baseline covariates, time-dependent
covariates, and intermediate outcomes, t =0,...,T − 1. In particular, define
A =(A(0),...,A(T −1)) and X =(X(0),...,X(T −1)). Let Y be the primary
outcome variable measured at time T .
Let Y
a
be the potential outcome had the patient been treated by treatment
sequence
a =(a(0),a(1),...,a(T − 1)). Let a(t)=(a(0),a(1),...,a(t)). Two
particular treatment sequences are
a = 1=(1,...,1) and a = 0=(0,...,0).
After defining the above potential outcomes, we can define the causal estimand
of interest. In this chapter, we focus on the following estimand,
v
∗
(a)=E(Y
a
), (9.1)
whichisreferredtoasthevalue of
a (Tsiatis et al. 2020).
DOI: 10.1201/9781003433378-9 167
168 The Art of Estimation (III): LTMLE
The reason why we focus on the above estimand is that many other esti-
mands can be defined consequently; e.g., the average treatment effect (ATE)
of treatment sequence
1 compared with treatment sequence 0,
θ
∗
ATE
= E(Y
a=1
) − E(Y
a=0
)=v
∗
(1) − v
∗
(0). (9.2)
9.1.2 Identification
In Chapter 6, we showed that, under the identifiability assumptions (consis-
tency, sequential exchangeability, and positivity), we can translate the ATE
causal estimand (9.2) into a statistical estimand, by either the standardiza-
tion strategy or the weighting strategy. In this chapter, we revisit the task
of identification, but for the value causal estimand (9.1), via the weighting
strategy and a series of standardization strategies.
The weighting strategy
In Chapter 6, we showed the following result for the setting where T =2,
v
∗
(a)=E(Y
a
)
= E
I(
A = a)Y
P{A(0) = a(0)|X(0)}P{A(1) = a(1)|X(0),A(0),X(1)}
= v(a),
under the identifiability assumptions. We can generate it to any T>1,
v(
a)=E
I(
A = a)Y
P{A(0) = a(0)|X(0)}
-
T −1
t=1
P{A(t)=a(t)|X(t), A(t − 1)}
=E
I(
A = a)Y
-
T −1
t=0
P{A(t)=a(t)|X(t), A(t − 1)}
, (9.3)
where, by convention,
X(0) = X(0), A(0) = A(0), and X(−1) = A(−1) = ∅.
Define the following propensity score function at time t,
g
t
a(t)
X(t), A(t − 1)
= P
A(t)=a(t)
X(t), A(t − 1)
, (9.4)
where t =0,...,T − 1. Then, (9.3) can be expressed as
v(
a)=E
I(
A = a)Y
-
T −1
t=0
g
t
a(t)|
X(t), A(t − 1)
. (9.5)
Longitudinal Cohort Studies 169
A series of standardization strategies
We consider a backward procedure. Starting from (9.5), we have
v(
a)
(a)
= E
E
I(
A = a)Y
-
T −1
t=0
g
t
a(t)|
X(t), A(t − 1)
X(T − 1), A(T − 2)

(b)
=E
E
I(
A = a)Y
a
-
T −1
t=0
g
t
a(t)|
X(t), A(t − 1)
X(T − 1), A(T − 2)

(c)
=E
I
A(T − 2) = a(T − 2)
-
T −2
t=0
g
t
a(t)|
X(t), A(t − 1)
× E
I
A(T − 1) = a(T − 1)
Y
a
g
T −1
a(T − 1)|
X(T − 1), A(T − 2)
X(T − 1), A(T − 2)

(d)
= E
I
A(T − 2) = a(T − 2)
-
T −2
t=0
g
t
a(t)|
X(t), A(t − 1)
E
Y
a
X(T − 1), A(T − 2)
(e)
=E
I
A(T − 2) = a(T − 2)
-
T −2
t=0
g
t
a(t)|
X(t), A(t − 1)
× E
Y
a
X(T − 1), A(T − 2),A(T − 1) = a(T −1)
(f)
= E
E
I
A(T − 2) = a(T − 2)
Y
a
-
T −2
t=0
g
t
a(t)|
X(t), A(t − 1)
X(T − 1), A(T − 2),a(T − 1)

(g)
= E
E
I
A(T − 2) = a(T − 2)
Y
-
T −2
t=0
g
t
a(t)|
X(t), A(t − 1)
X(T − 1), A(T − 2),a(T − 1)

(h)
= E
I
A(T − 2) = a(T − 2)
-
T −2
t=0
g
t
a(t)|
X(t), A(t − 1)
E
Y
X(T − 1), A(T − 2),a(T − 1)
(i)
=E
I
A(T − 2) = a(T − 2)
-
T −2
t=0
g
t
a(t)|
X(t), A(t − 1)
Q
T −1
X(T − 1), A(T − 2),a(T − 1)
,
where (a) is using the law of iterated expectations, (b) holds under the consis-
tency assumption, (c) holds because of the conditional expectation, (d)holds
because g
T −1
is canceled out, (e) holds under the sequential exchangeability
assumption, (f ) holds because of the conditional expectation, (g) holds under
the consistency assumption, (h) holds because of the conditional expectation,
and in (i)
Q
T −1
X(T − 1), A(T − 1)
= E
Y
X(T − 1), A(T − 1)
(9.6)
is the outcome regression function of Y ∼
X(T − 1) + A(T − 1).
170 The Art of Estimation (III): LTMLE
Next, we can proceed one step backward. In the previous step the outcome
is Y . In this step, consider the following tentative “outcome.”
&
Y
T −1
= Q
T −1
X(T − 1), A(T − 2),a(T − 1)
,
which equals the regression function evaluated at
X(T − 1), A(T − 2), and
A(T −1) = a(T −1). Then following the similar arguments, we can show that
v(
a)
=E
I
A(T − 2) = a(T − 2)
-
T −2
t=0
g
t0
a(t)|
X(t), A(t − 1)
Q
T −1
X(T − 1), A(T − 2),a(T − 1)
=E
I
A(T − 2) = a(T − 2)
-
T −2
t=0
g
t0
a(t)|
X(t), A(t − 1)
&
Y
T −1
=E
I
A(T − 3) = a(T − 3)
-
T −3
t=0
g
t0
a(t)|
X(t), A(t − 1)
Q
X(T − 2), A(T − 3),a(T − 2)
,
where
Q
T −2
X(T − 2), A(T − 2)
= E
$
&
Y
T −1
X(T − 2), A(T − 2)
%
(9.7)
is the outcome regression function of
&
Y
T −1
∼ X(T − 2) + A(T − 2).
We can repeat the above by T −2 steps backward until we reach the last
step. In the last step, consider the following tentative “outcome,”
&
Y
1
= Q
1
X(1),A(0),a(1)
,
which equals the regression function Q
1
(X(1), A(1)) evaluated at X(T −
1),
A(T − 2), and A(T − 1) = a(T − 1). Similarly, we can show that
v(
a)
=E
I (A(0) = a(0))
g
0
(a(0)|X(0))
Q
1
X(1),A(0),a(1)
=E
I (A(0) = a(0))
g
0
(a(0)|X(0))
&
Y
1
=E {Q
0
(X(0),a(0))},
where
Q
0
(X(0),A(0)) = E
$
&
Y
1
X(0),A(0)
%
(9.8)
is the outcome regression function of
&
Y
1
∼ X(0)+A(0). The last step is similar
to applying the standardization strategy to the study.
Longitudinal Cohort Studies 171
9.1.3 Efficient influence function
We will use the shortcut proposed in Chapter 8 to find the efficient influence
function for estimating v(
a). For this aim, we center all the above T versions
of the statistical estimand. This time, let’s take these steps forward.
At Step t = 0, for the following version of statistical estimand,
v(
a)=E {Q
0
(X(0),a(0))}, (9.9)
we define
φ
0
(O)=Q
0
(X(0),a(0)) − v(a).
At Step t =2,...,T − 1, for the following version of statistical estimand,
v(
a)=E
I
A(t − 1) = a(t − 1)
-
t−1
s=0
g
s
a(s)|
X(s), A(s − 1)
Q
t
X(t), A(t − 1),a(t)
, (9.10)
we define
φ
t
(O)=
I
A(t − 1) = a(t − 1)
-
t−1
s=0
g
s
a(s)|
X(s), A(s − 1)
×
Q
t
X(t), A(t − 1),a(t)
− Q
t−1
X(t − 1), A(t − 1)

.
At Step t = T , for the following version of statistical estimand,
v(
a)=E
I(
A = a)Y
-
T −1
t=0
g
t
a(t)|
X(t), A(t − 1)
, (9.11)
we define
φ
T
(O)=
I(
A = a)
-
T −1
t=0
g
t
a(t)|
X(t), A(t − 1)
Y −Q
T −1
X(T − 1), A(T − 1)

.
By the shortcut of finding the efficient influence function, we can show
that the efficient influence function of v(
a)is
φ
EIF-a
(O)=
T
t=0
φ
t
(O). (9.12)
That is,
φ
EIF-a
(O)=[Q
0
(X(0),a(0)) − v(a)] +
T −1
t=1
I
A(t − 1) = a(t − 1)
-
t−1
s=0
g
s
a(s)|
X(s), A(s − 1)
×
Q
t
X(t), A(t − 1),a(t)
− Q
t−1
X(t − 1), A(t − 1)

+
I(
A = a)
-
T −1
t=0
g
t
a(t)|
X(t), A(t − 1)
Y −Q
T −1
X(T − 1), A(T − 1)

.
172 The Art of Estimation (III): LTMLE
Proof of the validity of the shortcut
:
Consider any regular parametric submodel indicated by =(
X
,
g
t
,
Q
t
,t=
0,...,T−1)
T
,where
X
is for f
X
(x),
g
t
is for g
t
,and
Q
t
is for Q
t
, t =0,...,T−
1. Assume that when = 0 =(0,...,0)
T
, these parametric submodels are the
same as the true underlying models. Let v(
a; ) be the statistical estimand
which is defined via any of the above T versions, which is a function of had
the data been generated by the parametric submodel.
We can verify that function φ
EFF-a
(O) defined via the shortcut satisfies the
condition of the fundamental theorem regularity provided in Chapter 8,
∂v(
a; )/∂|
=0
= E[φ
EFF-a
(O)S
],
where S
=(S
X
,S
g
0
,S
Q
0
,...,S
g
T −1
,S
Q
T −1
)
T
is the corresponding score.
For
X
,weuseversion(9.9) to define v(a; ). Thus, we have
∂v(
a; )/∂
X
|
=0
= E[φ
0
(O)S
X
],
E[φ
s
(O)S
X
]=0,s =1,...,T.
This implies that ∂v(
a; )/∂
X
|
=0
= E[φ
EFF-a
(O)S
X
].
For
g
0
, we also use version (9.9) to define v(a; ). Thus,
∂v(
a; )/∂
g
0
|
=0
= E[φ
0
(O)S
g
0
],
E[φ
s
(O)S
g
0
]=0,s =1,...,T.
This implies that ∂v(
a; )/∂
g
0
|
=0
= E[φ
EFF-a
(O)S
g
0
].
For
g
t
,t≥ 1, we use version (9.10) to define v(a; ). Thus,
∂v(
a; )/∂
g
t
|
=0
= E[φ
t
(O)S
g
t
],
E[φ
s
(O)S
g
t
]=0,s =0,...,T and s = t.
This implies that ∂v(
a; )/∂
g
t
|
=0
= E[φ
EFF-a
(O)S
g
t
], t =1,...,T − 1.
For
Q
t
,t≥ 0, we use version (9.11) to define v(a; ). Thus,
∂v(
a; )/∂
Q
t
|
=0
= E[φ
T
(O)S
Q
t
],
E[φ
s
(O)S
Q
t
]=0,s =0,...,T − 1.
This implies that ∂v(
a; )/∂
Q
t
|
=0
= E[φ
EFF-a
(O)S
Q
t
], t =0,...,T − 1.
Therefore, we show that ∂v(
a; )/∂|
=0
= E[φ
EFF-a
(O)S
].

9.1.4 LTMLE

After we obtain the efficient influence function, φ
EFF-a
(O), we are ready to
describe the LTMLE procedure. Simply put, LTMLE is a combination of a
series of TMLE procedures that are applied backward.
Longitudinal Cohort Studies 173
Using super learner, we can obtain initial estimators of g
t
and Q
t
, t =
0,...,T−1. The first step is to obtain the targeted estimator of Q
T −1
, denoted
as
Q
∗
T −1
, from the regression modeling of Y ∼ (X, A), using TMLE with the
following clever covariate,
H
T −1
X,A
=
I(
A = a)
-
T −1
t=0
g
t
A(t)|
X(t), A(t − 1)
.
In the next step, denote the tentative “outcome” as
&
Y
T −1
=
Q
∗
T −1
X(T − 1), A(T − 2),a(T − 1)
.
Then obtain the targeted estimator of Q
T −2
, denoted as
Q
∗
T −2
,fromthe
regression modeling of
&
Y
T −1
∼ (X(T −2), A(T − 2)), using TMLE with the
following clever covariate,
H
T −2
X(T − 2), A(T − 2)
=
I
A(T − 2) = a(T − 2)
-
T −2
t=0
g
t
A(t)|
X(t), A(t − 1)
.
Repeat the above TMLE procedures until the last step. At the last step,
denote the tentative “outcome” as
&
Y
1
=
Q
∗
1
X(1),A(0),a(1)
.
Then obtain the targeted estimator of Q
0
, denoted as
Q
∗
0
, from the regres-
sion modeling of
&
Y
1
∼ (X (0),A(0)), using TMLE with the following clever
covariate,
H
0
(X(0),A(0)) =
I
A(0) = a(0)
g
0
(A(0)|X(0))
.
Finally, we obtain the LTMLE estimator,
v(
a)
LTM L E
=
1
n
n
i=1
Q
∗
0
(X
i
(0),a(0)) . (9.13)

9.1.5 ATE estimand

After we discuss the efficient influence function for the value estimand, we
can obtain the efficient influence function for the ATE estimand easily. Under
the identifiability assumptions, we can translate causal estimand θ
∗
ATE
in the
following statistical estimand,
θ
ATE
= v(1) −v(0).
We can easily verify that the efficient influence function for estimating θ
ATE
is
φ
EIF-ATE
(O)=φ
EFF-1
(O) − φ
EFF-0
(O). (9.14)
174 The Art of Estimation (III): LTMLE
Thus, the LTMLE estimator of θ
ATE
is
θ
LTM L E- AT E
= v(1)
LTM L E
− v(0)
LTM L E
. (9.15)
And the asymptotic variance of
√
n(
θ
LTM L E- AT E
−θ
ATE
)isV(φ
EIF-ATE
(O)), which
can be estimated by its sample analog.

9.2 Missing Data

9.2.1 Monotone missing

In this subsection, we only consider monotone missing data due to analysis
dropouts. Later we will discuss how to handle missing data and intercurrent
events (ICEs) in general.
Let Δ(t) be the indicator of analysis dropout at time t, t =1,...,T.Let
Δ = (Δ(1),...,Δ(T )) and Δ(t) = (Δ(1),...,Δ(t)), t =1,...,T.Inthis
subsection, we assume that the analysis dropout pattern is monotone; that is,
if Δ(t)=1,thenΔ(t
)=1fort
= t +1,...,T. Therefore, we can understand
such monotone missing as censoring.
Estimand
Let Y
a,δ=0
be the potential outcome if the patient had followed treatment
sequence
a and the data had been collected throughout the study.
After defining the potential outcome, we can define the causal estimand of
interest. We start with the following estimand,
v
∗
H
(a)=E(Y
a,δ=0
). (9.16)
Consequently, many other estimands can be defined; for example,
θ
∗
H
= E(Y
a=1,δ=0
) − E(Y
a=0,δ=0
)=v
∗
H
(1) − v
∗
H
(0). (9.17)
Identification
Besides those three identifiability assumptions (consistency, sequential ex-
changeability, and positivity), we assume the missing at random assumption.
If we think of “
A = a, Δ = 0” as an action parallel to action “A = a, ”we
can revise the statements in the preceding section to accomplish the task of
identification. For this aim, define the following non-missing (or non-censoring)
probability function,
h
t
X(t), A(t)
= P
Δ(t +1)=0
X(t), A(t), Δ(t)=0
,
where t =0,...,T − 1 and by convention,
Δ(0) = ∅.
Missing Data 175
To simply the notation, let
g
t
=
t
+
s=0
g
s
A(s)|
X(s), A(s − 1)
,
h
t
=
t
+
s=0
h
s
X(s), A(s)
,
I
t
= I(A(t)=a(t), Δ(t +1)=0),
where t =0,...,T − 1.
Similarly, the statistical estimand to be developed has T equivalent forms.
First, we obtain the following form via the weighting strategy,
v
H
(a)=E
I
T −1
Y
g
T −1
h
T −1
.
Second, consider Y as the outcome variable and define
&
Q
T −1
X(T − 1), A(T − 1)
= E
Y |X(T − 1), A(T − 1), Δ=0
as the outcome regression function in the modeling of Y ∼ (
X(T −1), A(T −1))
conditional on
Δ=0. Thus, we obtain the second form,
v
H
(a)=E
I
T −2
g
T −2
h
T −2
&
Q
T −1
X(T − 1), A(T − 2),a(T − 1)
.
Third, consider the following variable as the tentative “outcome,”
&
Y
∗
T −1
=
&
Q
T −1
X(T − 1), A(T − 2),a(T − 1)
,
and define
&
Q
T −2
X(T − 2), A(T − 2)
= E
)
&
Y
∗
T −1
|X(T − 2), A(T − 2), Δ(T − 1) = 0
*
as the outcome regression function in the modeling of
&
Y
∗
T −1
∼ (X(T −2), A(T−
2)) conditional on
Δ(T − 1) = 0. Thus, we obtain the third form,
v
H
(a)=E
I
T −3
g
T −3
h
T −3
&
Q
T −2
X(T − 2), A(T − 3),a(T − 2)
.
Repeat the above steps and obtain the following forms,
v
H
(a)=E
I
t−1
g
t−1
h
t−1
&
Q
t
X(t), A(t − 1),a(t)
,
where t = T − 1,T − 2,...,0. This includes the last form of the estimand,
v
H
(
a)=E
$
&
Q
0
(X(0),a(0))
%
.