Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5431_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
10.10.2026
Размер:
10 Мб
Скачать
☆
136 The Art of Estimation (I): M-estimation
where D
0
is the same as the one defined in the preceding subsection and
V
0
(h)=E
AY
g(1|X; γ
0
)
−
(1 − A)Y
g(0|X; γ
0
)
−{A − g(1|X; γ
0
)}h(X) − θ
0
2
=E
AY
g(1|X; γ
0
)
−
(1 − A)Y
g(0|X; γ
0
)
−{A − g(1|X; γ
0
)}h(X)
2
− θ
2
0
=E
AY
g(1|X; γ
0
)
−
(1 − A)Y
g(0|X; γ
0
)
2
+ E [{A − g(1|X; γ
0
)}h(X)]
2
−2E

AY
g(1|X; γ
0
)
−
(1 − A)Y
g(0|X; γ
0
)
{A − g(1|X; γ
0
)}h(X)
− θ
2
0
. (7.62)
The bread term in the sandwich formula is
E
∂M(O, μ
0
)
∂μ
T
=
−1 E
0
(h)
0 −F
0
,
where F
0
is the same as the one defined in the preceding subsection and
E
0
(h)=E

−
AY g(0|X
0
; γ
0
)
g(1|X; γ
0
)
−
(1 − A)Yg(1|X; γ
0
)
g(0|X; γ
0
)
+ g(1|X ; γ
0
)g(0|X; γ
0
)h(X)
&
X
T
=E

− E(Y |X, A =1)g(0|X
0
; γ
0
) − E(Y |X, A =0)g(1|X ; γ
0
)
+ g(1|X ; γ
0
)g(0|X; γ
0
)h(X)
&
X
T
.
Thus, if g(a|x; γ) is a correct model for g(a|x), we derive the asymptotic
normality of
θ(h) for any given h,
√
n(
θ(h) − θ
0
) →N(0,σ
2
0
(h)), (7.63)
where
σ
2
0
(h)=V
0
(h)+E
0
(h)F
−1
0
D
0
F
−1
0
E
T
0
(h). (7.64)

7.5.3 AIPW estimator

Our goal is to find an optimal h
∗
that minimizes σ
2
0
(h). Note that the second
term on the right-hand-side (RHS) of (7.64) is non-negative. Therefore, if we
could find an h
∗
that minimizes V
0
(h) and satisfies E
0
(h
∗
) = 0, then this h
∗
minimizes σ
2
0
(h).
Augmented Inverse Probability Weighted Estimator 137
In the expansion of V
0
(h) in (7.62), there are four terms in the foremost
RHSof(7.62),amongwhichonlytwotermsdependonh(X). Thus, applying
the law of iterated expectations to those two terms respectively, we have
E {[A − g(1|X ; γ
0
)]h(X)}
2
=E
E {[A − g(1|X ; γ
0
)]h(X)}
2
|X
=E
h
2
(X)E{[A − g(1|X; γ
0
)]
2
|X}
=E
g(1|X; γ
0
)g(0|X; γ
0
)h
2
(X)
(7.65)
and
E

AY
g(1|X; γ
0
)
−
(1 − A)Y
g(0|X; γ
0
)
{A − g(1|X; γ
0
)}h(X)
=E
E

AY
g(1|X; γ
0
)
−
(1 − A)Y
g(0|X; γ
0
)
{A − g(1|X; γ
0
)}h(X)
X

=E
E

Y
g(1|X; γ
0
)
g(0|X; γ
0
)
X, A =1
P(A =1|X)h(X)
+ E
E

Y
g(0|X; γ
0
)
g(1|X; γ
0
)
X, A =0
P(A =0|X)h(X)
=E {[E(Y |X, A =1)g(0|X; λ
0
)+E(Y |X, A =0)g(1|X; λ
0
)]h(X)}. (7.66)
Combining the above two results, (7.65) and (7.66), we see that the optimal
h
∗
(X) that minimizes V
0
(h)isequalto
h
∗
(X)=argmin
h(·)
V
0
(h)=argmin
h(·)
E
)
g(1|X; γ
0
)g(0|X; γ
0
)h
2
(X)
−2
E(Y |X, A =1)g(0|X; λ
0
)+E(Y |X, A =0)g(1|X; λ
0
)
h(X)
*
. (7.67)
Note the solution of the minimization of quadratic function f(x)=ax
2
+bx+c
is x
∗
= −b/(2a), as illustrated in Figure 7.1. Thus, we obtain
h
∗
(X)=
E(Y |X, A =1)
g(1|X; γ
0
)
+
E(Y |X, A =0)
g(0|X; γ
0
)
. (7.68)
In addition, we can easily verify that h
∗
(X) satisfies E
0
(h
∗
) = 0. Therefore,
we complete the proof that h
∗
(X) defined in (7.68) minimizes σ
2
0
(h). Thus,
the optimal “estimator” of the form (7.59) with the smallest variance is
θ(h
∗
)=
1
n
n
i=1
A
i
Y
i
g(1|X
i
; γ)
−
(1 − A
i
)Y
i
g(0|X
i
; γ)
− [A
i
− g(1|X
i
; γ)]
×
Q(X
i
, 1)
g(1|X
i
,γ
0
)
+
Q(X
i
, 0)
g(0|X
i
; γ
0
)
. (7.69)
138 The Art of Estimation (I): M-estimation
= 
+  +
∗
= −
2
FIGURE 7.1
The axis of symmetry
This is not an estimator because it depends on the unknown outcome regres-
sion function Q(x, a). If we specify a parametric model, say Q(X, A; β), and
estimate β using
β, we construct the augmented inverse probability weighted
estimator (AIPW),
θ
AIPW
=
1
n
n
i=1
A
i
Y
i
g(1|X
i
; γ)
−
(1 − A
i
)Y
i
g(0|X
i
; γ)
− [A
i
− g(1|X
i
; γ)]
×
Q(X
i
, 1;
β)
g(1|X
i
, γ)
+
Q(X
i
, 0;
β)
g(0|X
i
, γ)

. (7.70)

7.5.4 Double robustness

Let’s put the construction of
θ
AIPW
within the framework of M-estimation.
Assume the logistic model (7.45) for propensity score function g(A|X); that
is, logit[g(A|X; β)] = γ(0) + γ
T
(1)X = γ
T
&
X,whereγ =(γ(0),γ
T
(1))
T
.
In addition, assume a model for the outcome regression function, say linear
model Q(X, A;β)=β(0)+β(1)A+β
T
(2)X + β
T
(3)AX for continuous outcome
and logistic model logit[Q(X, A; β)] = β(0) + β(1)A + β
T
(2)X + β
T
(3)AX for
binary outcome, respectively, where β =(β(0),β
T
(2),β
T
(3))
T
. Hereafter, we
consider the setting of continuous outcome only.
Let μ =(θ, β
T
,γ
T
)
T
and M (O; μ)=(M
1
(O; μ),M
T
2
(O; μ),M
T
3
(O; μ))
T
with
M
1
(O; μ)=
AY
g(1|X; γ)
−
(1 − A)Y
g(0|X; γ)
−{A − g(1|X; γ)}
×
Q(X, 1;β)
g(1|X; γ)
+
Q(X, 0;β)
g(0|X; γ)
− θ, (7.71)
M
2
(O; μ)=
∂Q(X,A; β)
∂β
T
[Y −Q(X,A; β)], (7.72)
M
3
(O; μ)=
&
X[A − g(1|X; γ)]. (7.73)
Augmented Inverse Probability Weighted Estimator 139
If the propensity score model g is correctly specified, by the M-estimation
theory, we have γ → γ
0
. Then, we can verify that
E [M
1
(O; θ
0
,β,γ
0
)] = 0, (7.74)
for any β.Thus,
θ
AIPW
is consistent if g is correctly specified.
If the outcome regression model Q is correctly specified, by the M-
estimation theory,
β → β
0
. Then, we can verify that
E [M
1
(O; θ
0
,β
0
,γ)] = 0, (7.75)
for any γ.Thus,
θ
AIPW
is consistent if Q is correctly specified.
The above property is called double robustness because
θ
AIPW
is consistent
if either Q or g is correctly specified.
When both Q and g are correctly specified, by the M-estimation theory,
we can show that (the details are omitted; see Ex 7.4)
√
n(
θ
AIPW
− θ
0
) →N(0,σ
2
0,
AIPW
), (7.76)
where
σ
2
0,
AIPW
=E
AY
g(1|X; γ
0
)
−
(1 − A)Y
g(0|X; γ
0
)
− [A − g(1|X; γ
0
)]
×
Q(X, 1;β
0
)
g(1|X; γ
0
)
+
Q(X, 0;β
0
)
g(0|X; γ
0
)
− θ
0
2
, (7.77)
which can be estimated by
σ
2
AIPW
=
1
n
n
i=1
A
i
Y
i
g(1|X
i
; γ)
−
(1 − A
i
)Y
i
g(0|X
i
; γ)
− [A
i
− g(1|X
i
; γ)]
×
Q(X
i
, 1;
β)
g(1|X
i
; γ)
+
Q(X
i
, 0;
β)
g(0|X
i
; γ)
−
θ
AIPW
2
.
7.5.5 Influence function
To summarize, we show that
θ
AIPW
is an RAL estimator,
√
n(
θ
AIPW
− θ
0
)=
1
√
n
n
i=1
φ
AIPW
(O
i
)+o
p
(1), (7.78)
where φ
AIPW
(O) is the corresponding influence function and
φ
AIPW
(O)=
AY
g(1|X; γ
0
)
−
(1 − A)Y
g(0|X; γ
0
)
− [A − g(1|X; γ
0
)]
×
Q(X, 1;β
0
)
g(1|X; γ
0
)
+
Q(X, 0;β
0
)
g(0|X; γ
0
)
− θ
0
. (7.79)
140 The Art of Estimation (I): M-estimation
TABL E 7 .2
A small dataset
ID XAY ID XAY
100011 1 0 1
2001
12 1 0 0
3001
13 1 1 0
4000
14 1 1 1
5011
15 1 1 1
6011
16 1 1 0
7010
17 1 1 0
8011
18 1 1 0
9100
19 1 1 1
10 1 0 0
20 1 1 0
Since
θ
AIPW
is an efficient estimator if both g and Q are correctly specified,
we call the above influence function, φ
AIPW
(O), the efficient influence function
(EIF). In the next chapter, we will see that this EIF plays a crucial role in
the targeted learning framework.
In this chapter, we use only elementary calculus to derive the EIF (7.79)
without stepping into the territory of the semiparametric theory. Please refer
to Bickel et al. (1993) or Tsiatis (2006) for a much more elegant theory of
deriving the EIF (7.79).

7.6 Exercises

Ex 7.1
Consider the small dataset in Table 7.2. Since all the three variables are binary,
we can fit a saturate model for Q. Obtain a point estimate for the ATE
estimand using the g-computation estimator
θ
g-comp
.
Ex 7.2
Consider the small dataset in Table 7.2. Since all the three variables are bi-
nary, we can fit a saturate model for g. Obtain a point estimate for the ATE
estimand using the IPW estimator
θ
IPW
.
Exercises 141
Ex 7.3
Consider the small dataset in Table 7.2. Since all the three variables are binary,
we can fit saturate models for Q and g, respectively. Obtain a point estimate
for the ATE estimand using the AIPW estimator
θ
AIPW
.
Ex 7.4
Following the M-estimation theory and (7.71)–(7.73), verify the result (7.76).
8
The Art of Estimation (II):
TMLE

8.1 Semiparametric Statistics

8.1.1 Semiparametric estimators

Continue the discussion in Chapter 7. Consider a cohort study, in which we
are interested in estimating the average treatment effect (ATE),
θ
∗
= E(Y
a=1
− Y
a=0
), (8.1)
and we have data {O
i
=(X
i
,A
i
,Y
i
),i=1,...,n}.
Recall that, under the identifiability assumptions—consistency, exchange-
ability, and positivity—we can translate the causal estimand θ
∗
to a statistical
estimand θ and there are two strategies to do so: the standardization strategy
and the weighting strategy. By the standardization strategy, we derive the
following statistical estimand,
θ = E[Q(X, 1) − Q(X, 0)], (8.2)
where Q(x, a)=E(Y |X = x, A = a) is the outcome regression function. By
the weighting strategy, we derive the following statistical etsimand,
θ = E
(2A − 1)Y
g(A|X)
, (8.3)
where g(a|x)=P(A = a|X = x)] is the propensity score function.
In Chapter 7, within the framework of M-estimation, we reviewed three
methods: MLE, IPW, and AIPW. In the construction of the MLE, we assume
a parametric model Q(x, a; β)forQ(x, a); if Q(x, a; β) is a correct model for
Q(x, a), then
θ
MLE
is a consistent estimator for θ. In the construction of the
IPW estimator, we assume a parametric model g(a|x; γ)forg(a|x); if g(a|x; γ)
is a correct model for g(a|x; γ), then
θ
IPW
is a consistent estimator for θ.Inthe
construction of the AIPW estimator, we assume a parametric model Q(x, a; β)
for Q(x, a) and a parametric model g(a|x; γ)forg(a|x); if either Q(x, a; β)or
g(a|x; γ) is correct, then
θ
AIPW
is a consistent estimator for θ.
DOI: 10.1201/9781003433378-8 142
Semiparametric Statistics 143
In practice, since Q(x, a)andg(a|x) are unknown functions, it is unreal-
istic to specify a parametric model for either Q(x, a)org(a|x). Actually, we
don’t even know which variables should be included in X in the first place.
This is the problem of misspecification. Therefore, to solve the problem of mis-
specification, it is more realistic to consider nonparametric or semiparametric
models for Q(x, a)andg(a|x).
Let
Q(x, a)andg(a|x) be some estimators of Q(x, a)andg(a|x), respec-
tively, via some nonparametric or semiparametric predictive modeling pro-
cedures (Hastie, Tibshirani, and Friedman 2009; van der Laan, Polley, and
Hubbard 2007). Based on
Q(x, a) and/or g(a|x) , we can construct the corre-
sponding MLE, IPW estimator, and AIPW estimator:
θ
MLE
=
1
n
n
i=1
$
Q(X
i
, 1) −
Q(X
i
, 0)
%
,
θ
IPW
=
1
n
n
i=1
(2A
i
− 1)Y
i
g(A
i
|X
i
)
,
θ
AIPW
=
1
n
n
i=1
⎡
⎣
(2A
i
− 1)Y
i
g(1|X
i
)
−{A
i
− g(1|X
i
)}
1
j=0
Q(X
i
,j)
g(j|X
i
)
⎤
⎦
.
In semiparametric statistics, a semiparametric estimator
θ for θ is one that
is consistent and asymptotically normal:
θ
p
−→ θ,
√
n(
θ − θ
0
)
d
−→ N (0,σ
2
).
We will demonstrate that these estimators,
θ
MLE
,
θ
IPW
,and
θ
AIPW
,are
semiparametric estimators if
Q(x, a) and/or g(a|x) are consistent. We will
also demonstrate there is still room to improve their performance, which will
be improved by the targeted learning approach (van der Laan and Rose 2011).

8.1.2 Super learner

There are a variety of methods for obtaining nonparametric or semiparametric
estimators for Q(x, a)andg(a|x) (Hastie, Tibshirani, and Friedman 2009).
Among them, super learner is a promising method (van der Laan, Polley, and
Hubbard 2007).
Unlike parametric modeling that relies on a pre-specified parametric
model, super learner utilizes a library of predictive models—statistical learning
algorithms or machine learning algorithms (Hastie, Tibshirani, and Friedman
2009). Super learner attempts to choose the “best” algorithm that achieves the
“best” performance in terms of a loss function. For example, for continuous
outcome, we consider the L
2
loss function,
L(O, Q)=[Y − Q(X, A)]
2
, (8.4)
144 The Art of Estimation (II): TMLE
and for binary outcome, we consider the negative-log-likelihood loss function,
L(O, Q)=−log
Q(X, A)
Y
[1 − Q(X, A)]
1−Y
. (8.5)
For treatment variable A, which is a binary variable, we also consider the
negative-log-likelihood loss function,
L(O, g)=−log
g(1|X)
A
[1 − g(0|X)]
1−A
. (8.6)
Super learner uses cross-validation to choose the “best” algorithm. With-
out cross-validation, we would choose an algorithm that overfits. For ex-
ample, for continuous outcome, we may choose an algorithm that outputs
Q(X
i
,A
i
)=Y
i
for any i, i =1,...,n, but such an algorithm is useless. With
cross-validation, we will choose an algorithm that achieves the “best” com-
promise between goodness-of-fit and model complexity.
The K-fold cross-validation procedure divides the data into K subgroups,
with each subgroup having about n/K observations. Leaving out one of the
K subgroups as the validation set, denoted as O
(k)
, we fit each algorithm
indicated by j in the library consisting of J algorithms to the remaining
K − 1 subgroups, providing a fitted model denoted as
Q
(j,k)
, k =1,...,K
and j =1,...,J. For algorithm j, the cross-validated risk is estimated by
CV
(j)
=
1
K
K
k=1
1
#(O
(k)
)
O∈O
(k)
L(O,
Q
(j,k)
). (8.7)
Discrete super learner (van der Laan, Polley, and Hubbard 2007) chooses
the algorithm with the smallest cross-validated risk,
j
∗
=arg min
j=1,...,J
CV
(j)
.
van der Laan, Polley, and Hubbard (2007) further proposed super learner
to improve discrete super learner. Super learner searches for an optimally
weighted combination of J algorithms in the library, by minimizing the fol-
lowing weighted cross-validated risk,
CV (w
1
,...,w
J
)=
1
K
K
k=1
1
#(O
(k)
)
O∈O
(k)
L(O,
J
j=1
w
j
Q
(j,k)
), (8.8)
where w
j
≥ 0and
J
j=1
w
j
=1.
Both discrete super learner and super learner can be implemented using
R package “SuperLearner.”

8.1.3 Semiparametric estimators based on super learner

Let
Q
SL
(x, a)andg
SL
(a|x) be estimators of Q(x, a)andg(a|x), respectively,
via super learner. Based on
Q
SL
(x, a) and/or g
SL
(a|x), we can construct the

Asymptotic Variances of Semiparametric Estimators 145

corresponding MLE-SL, IPW-SL, and AIPW-SL estimators:
θ
MLE-SL
=
1
n
n
i=1
$
Q
SL
(X
i
, 1) −
Q
SL
(X
i
, 0)
%
, (8.9)
θ
IPW-SL
=
1
n
n
i=1
(2A
i
− 1)Y
i
g
SL
(A
i
|X
i
)
, (8.10)
θ
AIPW-SL
=
1
n
n
i=1
⎡
⎣
(2A
i
− 1)Y
i
g
SL
(A
i
|X
i
)
−{A
i
− g
SL
(1|X
i
)}
1
j=0
Q
SL
(X
i
,j)
g
SL
(j|X
i
)
⎤
⎦
. (8.11)
Despite that we can use the bootstrap method—which is time-
consuming—to obtain the variances of these estimators, it is desirable to derive
explicit formulae for their asymptotic variances, which will be derived in the
next section.
8.2 Asymptotic Variances of Semiparametric Estimators
“Knowledge of the asymptotic variance of an estimator is important
for large sample inference, efficiency, and as a guide to the specification
of regularity conditions.”—Newey (1994)

8.2.1 Parametric submodels

In Chapter 7, we studied the asymptotic variances of parametric M-estimators.
To study the asymptotic variances of semiparametric M-estimators, we con-
sider the idea of parametric submodels (Stein 1956).
“One could imagine that the data are generated by a parametric model
that satisfies the semiparametric assumptions and contains the truth. Such
a model is referred to as a parametric submodel, where the ‘sub’ prefix
refers to the fact that it is a subset of the model consisting of all distri-
butions satisfying the assumptions.”—Newey (1990)
Assume the true model for X is f
X0
(x), the true model for A|(X = x)
is g
0
(a|x), and the true model for Y |(X = x, A = a)isf
Y 0
(y|a, x). These
distributions are unknown and we don’t put any constraints on them. Note
that the true outcome regression function is Q
0
(x, a)=
.
yf
Y 0
(y|x, a)dy.
Imagine any parametric submodel: X ∼ f
X
(x;
1
), A|(X = x) ∼ g(a|x;
2
),
and Y |(X = x, A = a) ∼ f
Y
(y|x, a;
3
), such that f
X
(x;0) = f
X0
(x),