Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5431_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
10.10.2026
Размер:
10 Мб
Скачать
☆
146 The Art of Estimation (II): TMLE
g(a|x;0) = g
0
(a|x), and f
Y
(y|a, x;0) = f
Y 0
(y|a, x). Note that the outcome
regression function is then Q(x, a;
3
)=
.
yf
Y
(y|x, a;
3
)dy and Q(x, a;0) =
Q
0
(x, a). To be more accurate, we only imagine some regular parametric model,
the one that is smooth with a non-singular information matrix. But we omit
such technical detail here, which can be found in Newey (1990).
Under a given parametric submodel, we can calculate the corresponding
score S
=(S
1
,S
2
,S
3
)
T
,where=(
1
,
2
,
3
)
T
and
S
1
= ∂ log f
X
(x;
1
)/∂
1
|
1
=0
,
S
2
= ∂ log g(a|x;
2
)/∂
2
|
2
=0
,
S
3
= ∂ log f
Y
(y|x, a;
3
)/∂
3
|
3
=0
.
Because the expectation of any score is zero, S
1
is a function of X with
E(S
1
)=0,S
2
is a function of (X, A)withE(S
2
|X) = 0, and S
3
is a function
of (X, A, Y )withE(S
3
|X, A) = 0. Because we could imagine any number of
parametric submodels, we put all the possible S
j
’s into a set named as T
j
,
j =1, 2, 3. Thus, we show that
T
1
= {α
1
(X):E[α
1
(X)] = 0},
T
2
= {α
2
(X, A):E[α
2
(X, A)|X]=0},
T
3
= {α
3
(X, A, Y ):E[α
3
(X, A, Y )|X, A]=0}.
Here is a remark on T
2
. For any α
2
(X, A) such that E[α
2
(X, A)|X]=0,
because A is binary, we have α
2
(X, 1)g
0
(1|X)+α
2
(X, 0)g
0
(0|X) = 0, im-
plying that α
2
(X, 1) = [−α
2
(X, 0)/g
0
(1|X)](1 − g
0
(1|X)andα
2
(X, 0) =
[−α
2
(X, 0)/g
0
(0|X)](−g
0
(1|X). Then, α
2
(X, A) can be written as a
2
(X)[A −
g
0
(1|X)], where a
2
(X)=−α
2
(X, 0)/g
0
(1|X). Thus, we have
T
2
= {[A − g
0
(1|X)]a
2
(X) : any a
2
(X)}.
We are interested in estimating the following statistical estimand:
θ
0
= E [Q
0
(X, 1) −Q
0
(X, 0)]
=
yf
Y 0
(y|1,x)dy −
yf
Y 0
(y|0,x)dy
f
X0
(x)dx.
or equivalently,
θ
0
= E
(2A − 1)Y
g
0
(A|X)
=
yf
Y 0
(y|1,x)dy −
yf
Y 0
(y|0,x)dy
f
X0
(x)dx.
Asymptotic Variances of Semiparametric Estimators 147
For any given regular parametric submodel indicated by =(
1
,
2
,
3
)
T
,the
estimand associated with the parametric submodel is
θ()=E
[Q(Y |X, 1;
3
) − Q(Y |X, 0;
3
)] = E
(2A − 1)Y
g(A|X;
2
)
=
yf
Y
(y|x, 1;
3
)dy −
yf
Y
(y|x, 0;
3
)dy
f
X
(x;
1
)dx.
To summarize, we consider a semiparametric model where θ is the pa-
rameter of interest and there are no constraints on f
X0
, g
0
,andf
Y 0
.For
any imagined parametric submodel, the estimand becomes θ = θ(), where
=(
1
,
2
,
3
)
T
.When= 0 =(0, 0, 0)
T
, θ
0
= θ(0).
We use the idea of parametric submodels to define the Cramer-Rao bound
for a semiparametric model. For any parametric submodel, we can obtain
the Cramer-Rao bound following the discussion in Chapter 7. We aim for
constructing semiparametric estimators whose asymptotic variances are com-
parable to the Cramer-Rao bound of the semiparametric model, which is to
be defined. For any given parametric submodel, the asymptotic variances of
these semiparametric estimators are no smaller than the Cramer-Rao bound of
the given parametric submodel. Therefore, the asymptotic variances of these
semiparametric estimators are no smaller than the supremum of the Cramer-
Rao bounds of all parametric submodels—this supremum is defined as the
Cramer-Rao bound of the semiparametric model.

8.2.2 The fundamental theorem of regularity

We also use the idea of parametric submodels to define the regularity. An
estimator is said to be regular if it is regular in every regular parametric
submodel and the limiting distribution does not depend on the parametric
submodel (Newey 1990).
Newey (1990) presented the following fundamental theorem of regularity:
Theorem Suppose that
θ is an asymptotically linear estimator for θ with
influence function φ(O). Suppose that, for any regular parametric submodel,
θ() is differentiable and E
[φ
2
(O)] exists and is continuous on a neighborhood
of 0.Then
θ is regular if and only if, for any regular parametric submodel,
∂θ()/∂
=0
= E[φ(O)S
], (8.12)
where S
is the score of the corresponding parametric submodel.
The first direct application of the fundamental theorem of regularity
Let U
1
,...,U
n
be observations which are i.i.d. with U ∼ f
0
(u). The estimand
of interest is β
0
= E(U ). Let the distribution of U be unrestricted, but with
finite variance, V(U ) < ∞.
148 The Art of Estimation (II): TMLE
Thesamplemean,
β =
n
i=1
U
i
/n, is an asymptotically linear estimator
for β with influence function φ(U )=U − E(U ), knowing that
√
n(
β − β
0
)=n
−1/2
n
i=1
[U
i
− E(U )].
Consider any parametric submodel f(u; ) such that f(u;
0
) is the true
distribution f
0
(u). The estimand of interest corresponding to the parametric
submodel becomes β()=
.
uf(u; )du.Thus,
∂β(
0
)
∂
=
u
∂f(u;
0
)
∂
du =
u
∂ log f(u;
0
)
∂
f(u;
0
)du
= E[US
]=E[(U − EU)S
]=E[φ(U)S
],
where S
= ∂ log f (u;
0
)/∂ is the score. Therefore, by the fundamental the-
orem of regularity,
β is regular and therefore it is an RAL estimator.
This simple example has a lot of applications. For example, if the propen-
sity score function, g
0
(A|X), is known, then
θ
rand
=
1
n
n
i=1
(2A
i
− 1)Y
i
g
0
(A
i
|X
i
)
is an RAL estimator for the following estimand,
θ
0
= E
(2A − 1)Y
g
0
(A|X)
.
Here subscript “rand” stands for randomization, because the propensity score
function is known in randomized studies. Then it implies that the influence
function is
φ
rand
(O)=
(2A − 1)Y
g
0
(A|X)
− θ
0
.
8.2.3 Influence function of MLE-SL estimator
Assume that
Q
SL
is a consistent estimator of Q
0
(x, a). Then under any regular
parametric submodel {f
X
(x;
1
},g(a|x;
2
),f
Y
(y|x, a;
3
)}, the limit of
Q
SL
is
Q(x, a;
3
). Then, under the parametric submodel, the limit of
θ
MLE-SL
is
θ()=E
[Q(X, 1;
3
) − Q(X, 0;
3
)] . (8.13)
By differentiation under the integral and the chain rule,
∂θ()/∂
=0
=∂E
[Q
0
(X, 1) −Q
0
(X, 0)]/∂
=0
(8.14)
+ ∂E [Q(X,1;
3
) − Q(X, 0;
3
)] /∂
=0
. (8.15)
Asymptotic Variances of Semiparametric Estimators 149
Using the first direct application of the fundamental theorem of regularity,
the first term, (8.14), becomes
∂E
[Q
0
(X, 1) −Q
0
(X, 0)]/∂
=0
= E {[Q
0
(X, 1) −Q
0
(X, 0)]S
}
where S
=(S
1
,S
2
,S
3
)
T
. Note that, by the law of iterated expectations,
E {[Q
0
(X, 1) −Q
0
(X, 0)]S
2
} =0,
E {[Q
0
(X, 1) −Q
0
(X, 0)]S
3
} =0.
For the second term, (8.15), its derivative over
1
or
2
is zero and
∂E [Q(X, 1;
3
) − Q(X, 0;
3
)] /∂
3
3
=0
=
y
∂f
Y
(y|x, 1; 0)
∂
3
dy −
y
∂f
Y
(y|x, 0; 0)
∂
3
dy
f
X0
(x)dx.
We can show that

y
∂f
Y
(y|x, 1; 0)
∂
3
dyf
X0
(x)dx
(a)
=

[y − Q
0
(x, 1)]
∂f
Y
(y|x, 1; 0)
∂
3
dyf
X0
(x)dx
(b)
=

[y − Q
0
(x, 1)]
g
0
(1|x)
∂f
Y
(y|x, 1; 0)
∂
3
g
0
(1|x)+
0
g
0
(0|x)
g
0
(0|x)
dyf
X0
(x)dx
(c)
=

E
A[y − Q
0
(x, A)]
g
0
(A|x)
∂f
Y
(y|x, A;0)
∂
3
X = x
dyf
X0
(x)dx
(d)
=

E
A[y − Q
0
(x, A)]
g
0
(A|x)
∂ log f
Y
(y|x, A;0)
∂
3
X = x
f
Y 0
(y|x, A)dyf
X0
(x)dx
(e)
=
E
A[y − Q
0
(x, A)]
g
0
(A|x)
∂ log f
Y
(y|x, A;0)
∂
3
f
Y 0
(y|x, A)dy
X = x
f
X0
(x)dx
(f )
=
E
E
A[Y − Q
0
(x, A)]
g
0
(A|x)
S
3
A, X = x
X = x
f
X0
(x)dx
(g)
=
E
A[Y − Q
0
(x, A)]
g
0
(A|x)
S
3
X = x
f
X0
(x)dx
(h)
= E
A[Y − Q
0
(X, A)]
g
0
(A|X)
S
3
,
where (a)–(b) hold because the newly added terms equal zero, (c)holdsby
writing summation as expectation, (d) holds by forming a score function,
(e) holds after switching expectation and integration, (f) holds by writing
integration as expectation, (g) holds by the law of iterated expectations, and
(h) holds by writing integration as expectation. Similarly, we can show that

y
∂f
Y
(y|x, 0; 0)
∂
3
dyf
X0
(x)dx = E
(1 − A)[Y − Q
0
(X, A)]
g
0
(A|X)
S
3
.
150 The Art of Estimation (II): TMLE
Thus,
∂E [Q(X, 1;
3
) − Q(X, 0;
3
)] /∂
3
3
=0
= E
(2A − 1)[Y − Q
0
(X, A)]
g
0
(A|X)
S
3
.
In addition, because E{[Q(X, 1) − Q(X, 0)]S
3
} = 0 using the law of iter-
ated expectations, we have
∂E [Q(X, 1;
3
) − Q(X, 0;
3
)] /∂
3
3
=0
=E

Q(X, 1) −Q(X, 0) +
(2A − 1)[Y − Q
0
(X, A)]
g
0
(A|X)
− θ
0
S
3
.
And because E{(2A − 1)[Y −Q
0
(X, A)]/g
0
(A|X)S
1
}=0 also using the law of
iterated expectations, we have
E {[Q
0
(X, 1) −Q
0
(X, 0)]S
1
}
=E

Q(X, 1) −Q(X, 0) +
(2A − 1)[Y − Q
0
(X, A)]
g
0
(A|X)
− θ
0
S
1
.
Moreover, by the law of iterated expectations, we can verify that
E

Q(X, 1) −Q(X, 0) +
(2A − 1)[Y − Q
0
(X, A)]
g
0
(A|X)
− θ
0
S
2
=0.
Therefore, combining the above three results, we have
∂θ(0)
∂
= E

Q(X, 0) −Q(X, 0) +
(2A − 1)[Y − Q
0
(X, A)]
g
0
(A|X)
− θ
0
S
=0,
for any regular parametric submodels. Thus, by the fundamental theorem of
regularity, we show that the influence function of
θ
MLE-SL
is
φ
MLE-SL
(O)=Q(X, 0) − Q(X, 0) +
(2A − 1)[Y − Q
0
(X, A)]
g
0
(A|X)
− θ
0
. (8.16)
8.2.4 Influence function of IPW-SL estimator
Assume that g
SL
is a consistent estimator of g
0
(a|x). Then under any regular
parametric submodel {f
X
(x;
1
},g(a|x;
2
),f
Y
(y|x, a;
3
)}, the limit of g
SL
is
g(a|x;
2
). Then, under the parametric submodel, the limit of
θ
IPW-SL
is
θ()=E
(2A − 1)Y
g(A|X;
2
)
. (8.17)
By differentiation under the integral and the chain rule,
∂θ()
∂
=0
= ∂E
(2A − 1)Y
g
0
(A|X)
/∂
=0
+ ∂E
(2A − 1)Y
g(A|X;
2
)
/∂
=0
. (8.18)
Asymptotic Variances of Semiparametric Estimators 151
Using the direct application of the fundamental theorem of regularity, the
first term on the right-hand-side (RHS) of (8.18) becomes
∂E
(2A − 1)Y
g
0
(A|X)
/∂
=0
= E
(2A − 1)Y
g
0
(A|X)
S
.
For the second term on RHS of (8.18), the derivatives over
1
are
3
are
equal to zero, respectively, and
∂E
(2A − 1)Y
g(A|X;
2
)
/∂
2
2
=0
= −E
(2A − 1)Y
g
2
0
(A|X)
∂g(A|X;0)
∂
2
= −E
(2A − 1)Y
g
0
(A|X)
∂ log g(A|X ;0)
∂
2
= −E
(2A − 1)Y
g
0
(A|X)
S
2
.
Combining these two results, we have
∂θ(0)
∂
=
E
(2A − 1)Y
g
0
(A|X)
S
1
, 0, E
(2A − 1)Y
g
0
(A|X)
S
3

T
.
Therefore, by the first direct application of the fundamental theorem of regu-
larity again, we see that the influence function of
θ
IPW-SL
is
φ
IPW-SL
(O)=
(2A − 1)Y
g
0
(A|X)
+ h(X, A) −θ
0
, (8.19)
where h(X, A) is a function to be determined, which satisfies
E [h(X,A)S
1
]=0, (8.20)
E

(2A − 1)Y
g
0
(A|X)
+ h(X, A)
S
2
=0, (8.21)
E [h(X,A)S
3
]=0, (8.22)
for any regular parametric submodels.
First, (8.22) holds for any regular parametric submodels, using the law of
iterated expectations and the fact that E(S
3
|X, A)=0.
Second, in order to satisfy (8.20) for any regular parametric submodels, it
requires that E [h(X, A)|X] = 0. Thus, it requires h(X, 1)g
0
(1|X)+h(X, 0)[1−
g
0
(1|X)] = 0. That is, it requires h(X, 1) = −h(X,0)/g
0
(1|X)[1 − g(1|X )],
along with an apparent fact that h(X,0) = −h(X,0)/g
0
(1|X)[0 − g(1|X)].
Therefore, we see that h(X, A)={A − g
0
(1|X)}h
∗
(X), where h
∗
(X)=
h(X, 0)/g
0
(1|X) is to be determined.
Third, in order to satisfy (8.21),
E

(2A − 1)Y
g
0
(A|X)
+ {A − g
0
(1|X)}h
∗
(X)
S
2
=0,
152 The Art of Estimation (II): TMLE
for any regular parametric submodels, it requires that
E

(2A − 1)Y
g
0
(A|X)
+ {A − g
0
(1|X)}h
∗
(X)
[{A − g
0
(1|X)}α
2
(X)]
=0,
for any α
2
(X), using the membership of T
2
. Thus, by the law of iterated
expectations, it requires that
E

(2A − 1){A − g
0
(1|X)}Q
0
(X, A)
g
0
(A|X)
+ {A − g
0
(1|X)}
2
h
∗
(X)
α
2
(X)
=0,
for any α
2
(X). By the law of iterated expectations again, it requires that
E {[g
0
(0|X)Q
0
(X, 1) + g
0
(1|X)Q
0
(X, 0) + g
0
(1|X)g
0
(0|X)h
∗
(X)] α
2
(X)} =0,
for any α
2
(X). Therefore, it requires that
g
0
(0|X)Q
0
(X, 1) + g
0
(1|X)Q
0
(X, 0) + g
0
(1|X)g
0
(0|X)h
∗
(X)=0,
leading to
h
∗
(X)=−
Q
0
(X, 1)
g
0
(1|X)
+
Q
0
(X, 0)
g
0
(0|X)
.
Therefore, plugging the resulting h
∗
(X) into (8.19), we show that
φ
IPW-SL
(O)=
(2A − 1)Y
g
0
(A|X)
−{A − g
0
(1|X)}
1
j=0
Q
0
(X, j)
g
0
(j|X )
− θ
0
. (8.23)
Aha! This influence function is the same as the one we developed in the
construction of the AIPW estimator in Chapter 7 using only the tools of
elementary calculus, without using the fundamental theorem of regularity.

8.2.5 Double robustness of AIPW-SL estimator

The performance of the MLE-SL estimator depends on whether
Q
SL
is a con-
sistent estimator of Q
0
. Thus, the MLE-SL estimator is singly robust. If
Q
SL
is a consistent estimator of Q
0
, we show that the asymptotic variance can be
estimated based on the influence function (8.16). But the influence function
depends on g
0
, which should be estimated, say by g
SL
.
The performance of the IPW-SL estimator depends on whether g
SL
is a
consistent estimator of g
0
. Thus, the IPW-SL estimator is singly robust. If g
SL
is a consistent estimator of g
0
, we show that the asymptotic variance can be
estimated based on the influence function (8.23). But the influence function
depends on Q
0
, which should be estimated, say by
Q
SL
.
Since, in order to estimate the asymptotic variance of either MLE-SL or
IPW-SL, we need to estimate both g
0
and Q
0
no matter what, then why not
Asymptotic Variances of Semiparametric Estimators 153
consider the AIPW-SL estimator to begin with? The AIPW-SL estimator,
(8.10), is doubly robust, in the sense that it is consistent if either g
SL
is a
consistent estimator of g
0
or
Q
SL
is a consistent estimator of Q
0
.
Simply put, the construction of AIPW-SL depends on g
SL
and
Q
SL
,and
the asymptotic variance of AIPW-SL can be estimated by
1
n(n − 1)
n
i=1
⎡
⎣
(2A
i
− 1)Y
i
g
SL
(A
i
|X
i
)
+ {A
i
− g
SL
(1|X
i
)}
1
j=0
Q
SL
(X
i
,j)
g
SL
(j|X
i
)
−
θ
AIPW-SL
⎤
⎦
2
.
8.2.6 Efficient influence function
There is a geometrical viewpoint of influence functions (Bickel et al. 1993).
Let H be the Hilbert space of random functions, h(O), with mean zero and
finite variance, equipped with the inner product h
1
,h
2
= E(h
1
h
2
). Therefore,
influence functions are members of H.
Our goal is to identify the “optimal” influence function, referred to as
the efficient influence function, φ
EIF
(O) ∈H, such that the variance of an
RAL estimator with φ
EIF
(O) achieves the Cramer-Rao bound. The efficient
influence function can be expressed as
φ
EIF
=arg min
φ∈S
RAL
V(φ(O)), (8.24)
where S
RAL
is the set of all the influence functions associated with RAL esti-
mators for θ.
The equivalence between the above two influence functions
We have shown that the influence functions of the MLE-SL estimator and the
IPW-SL estimator are, respectively,
Q
0
(X, 0) −Q
0
(X, 0) +
(2A − 1)
g
0
(A|X)
[Y −Q
0
(X, A)] − θ
0
. (8.25)
and
(2A − 1)Y
g
0
(A|X)
−
Q
0
(X, 1)
g
0
(1|X)
+
Q
0
(X, 0)
g
0
(0|X)
[A − g
0
(1|X)] − θ
0
. (8.26)
These two influence functions are equal—they are just two forms of the
same thing. To show the above two forms, (8.25) and (8.26), are equivalent,
we check the case where A = 1 and the case where A = 0, respectively. In
fact, when A = 1, both (8.25) and (8.26) become
Y
g
0
(1|X)
− Q
0
(X, 1)
g
0
(0|X)
g
0
(1|X)
− Q
0
(X, 0) −θ
0
,
154 The Art of Estimation (II): TMLE
while when A = 0, both (8.25) and (8.26) become
−
Y
g
0
(0|X)
+ Q
0
(X, 1) + Q
0
(X, 0)
g
0
(1|X)
g
0
(0|X)
− θ
0
.
The fact that the above two influence functions are the same is not sur-
prising because we will soon show that, when there is no constraint on f
X0
,
g
0
,andf
Y 0
, all the RAL estimators have the same influence function; that is,
there is only one influence function that satisfies the fundamental theorem of
regularity, if there is no constraint on the underlying distributions. Therefore,
we will call this influence function the efficient influence function:
φ
EIF
(O)=Q
0
(X, 1) −Q
0
(X, 0) +
(2A − 1)
g
0
(A|X)
[Y −Q
0
(X, A)] − θ
0
=
(2A − 1)Y
g
0
(A|X)
−
Q
0
(X, 1)
g
0
(1|X)
+
Q
0
(X, 0)
g
0
(0|X)
[A − g
0
(1|X)] − θ
0
. (8.27)
The efficient influence function
For any function h(O) ∈H, we have the following decomposition,
h(O)=E[h(O)|X ]+{E[h(O)|X, A] −E[h(O)|X]} + {h(O) −E[h(O)|X, A]}
= α
1
(X)+α
2
(X, A)+α
3
(X, A, Y ),
where E[α
1
(X)]=0,E[α
2
(X, A)|X]=0,andE[α
3
(X, A, Y )|X, A] = 0. Hence,
H = T
1
⊕T
2
⊕T
3
. (8.28)
Previously we have shown that φ
EIF
(O) satisfies the fundamental theorem
of regularity; that is,
∂θ()/∂
=0
= E[φ
EIF
(O)S
].
Assume there is another influence function
&
φ(O) that also satisfies the funda-
mental theorem of regularity,
∂θ()/∂
=0
= E[
&
φ(O)S
].
Thus, E{[
&
φ(O) − φ
EIF
(O)]S
]} = 0 for all regular parametric submodels. By
(8.28), E{[
&
φ(O)−φ
EIF
(O)]h(O)} = 0, for any h ∈H, implying
&
φ(O)=φ
EIF
(O).
Therefore, there is only one influence function that satisfies the fundamental
theorem of regularity and this influence function achieves the minimum of
(8.24) automatically.

8.3 The Targeted Learning Framework

In this section, we describe the targeted learning framework briefly. Refer to
van der Laan and Rose (2011) for more detail.
The Targeted Learning Framework 155

8.3.1 Mini-roadmap

A full roadmap for conducting causal inference will be discussed in Chapter 11.
In this subsection, we provide a mini-roadmap of using the targeted learning
approach to estimate the ATE estimand.
Step 1: The ATE estimand
Assume we target at the following ATE estimand,
θ = θ(Q, F
X
)=
[Q(x, 1) − Q(x, 0)]dF
X
(x). (8.29)
Step 2: Initial estimators of Q and g
Obtain an initial estimator,
Q
SL
, for the outcome regression function Q.Also
obtain an initial estimator, g
SL
, for the propensity score function g.
Step 3: The efficient influence function of the ATE estimand
Derive the following efficient influence function,
φ
EIF
(O)=Q
0
(X, 1) −Q
0
(X, 0) +
(2A − 1)
g
0
(A|X)
[Y −Q
0
(X, A)] − θ
0
.
Step 4: TMLE
Update the initial estimator to obtain the targeted maximum likelihood esti-
mator/targeted minimum loss estimator (TMLE) using the procedure to be
described in the next section.
8.3.2 TMLE
We start with the following initial estimator for θ,
θ
MLE-SL
= θ(
Q
SL
,
F
X
)=
1
n
n
i=1
$
Q
SL
(X
i
, 1) −
Q
SL
(X
i
, 0)
%
.
Then, using the efficient influence function, we are able to update the initial
estimator
Q
SL
to
Q
∗
, producing an updated estimator for the target estimand
(van der Laan and Rose 2011), which is referred to as TMLE, denoted as
θ
SL-MLE
;thatis,
θ

TMLE

= θ(
Q
∗
,
F
X
)=
1
n
n
i=1
$
Q
∗
(X
i
, 1) −
Q
∗
(X
i
, 0)
%
. (8.30)