Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_5431_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
10.10.2026
Размер:
10 Мб
Скачать
☆
126 The Art of Estimation (I): M-estimation
Consider the extreme example that μ = μ
0
.Ifμ = and
n
=
0
+ δ/
√
n,
then
√
n[μ − μ(
n
)] =
√
n[
0
− (
0
+ δ/
√
n)] = −δ, which is a degenerate
distribution that depends on the LDGP with {
n
}. Therefore, this extreme
estimator is not regular.
Newey (1990) presented the fundamental theorem of regularity:
Theorem Suppose that μ is an asymptotically linear estimator with influ-
ence function φ(O). Suppose that μ() is differentiable and E
[φ
2
(O)] exists
and is continuous on a neighborhood of
0
.Thenμ isregularifandonlyif
∂μ(
0
)/∂
T
= E[φ(O)S
T
]. (7.19)
By the fundamental theorem of regularity, from (7.18), we see that μ
MLE
is regular. Given that we have shown that μ
MLE
is asymptotically linear, we
conclude that μ
MLE
is a regular and asymptotically linear (RAL) estimator.
Regularity of M-estimator
Now we are ready to show that the M-estimator, which has the property of
linearity (7.12), is regular, leading to the conclusion that the M-estimator is
an RAL estimator. For this aim, after taking the derivative over on both
sides of equation (7.7), which is equivalent to the following equation,
M(o; μ())f(o; )do =0,
we obtain
∂M(o; μ)
∂μ
T
∂μ()
∂
T
f(o; )do +
M(o; μ())
∂f(o; )
∂
T
do =0.
Thus, we have
∂M(o; μ)
∂μ
T
f(o; )do
∂μ()
∂
T
+
M(o; μ())
∂ log f(o; )
∂
T
f(o; )do =0,
which implies
∂M(o; μ
0
)
∂μ
T
f(o;
0
)do
∂μ(
0
)
∂
T
+
M(o; μ
0
)
∂ log f(o;
0
)
∂
T
f(o;
0
)do =0,
which can be written as
E
∂M(O; μ
0
)
∂μ
T
∂μ(
0
)
∂
T
+ E
M(O; μ
0
)S
T
=0.
Thus, we have
∂μ(
0
)
∂
T
= E
−
E
∂M(O; μ
0
)
∂μ
T

−1
M(O; μ
0
)S
T
,

G-computation Estimator 127

which is equivalent to
∂μ(
0
)/∂
T
= E
φ(O)S
T
,
where φ(O) is the influence function defined in (7.13). Therefore, by the fun-
damental theorem of regularity, we prove that the M-estimator is regular.
7.3 G-computation Estimator

7.3.1 Plug-in estimator

Consider the statistical estimand in terms of (7.2), which depends on the
outcome regression function Q(x, a). Thus, we can rewrite (7.2) as
θ = E[Q(X, 1)] − E[Q(X, 0)]. (7.20)
Hence, if we can obtain an estimator of Q(x,a), denoted as
Q(x, a), and
use the sample mean to estimate the population mean, then we can obtain a
g-computation estimator of the statistical estimand in (7.20),
θ
g-comp
=
1
n
n
i=1
Q(X
i
, 1) −
1
n
n
i=1
Q(X
i
, 0), (7.21)
where subscript “g-comp” stands for “g-computation.”
In order to conduct statistical inference based on
θ
g-comp
, we need to have an
estimator of its variance. Although we can always use the bootstrap procedure
(Efron and Tibshirani 1994) to estimate its variance without attempting to
derive an explicit expression for it, it is desirable to derive an explicit formula
for the asymptotic variance of
θ
g-comp
, which will be derived soon.
To see that estimator
θ
g-comp
is a plug-in estimator, we rewrite the statis-
tical estimand in (7.20) as
θ =
[Q(x, 1) − Q(x, 0)]dF
X
(x) θ(Q, F
X
), (7.22)
where F
X
is the probability distribution function of X .
It is well-known that the empirical distribution of {X
1
,...,X
n
} is the non-
parametric maximum likelihood estimator (NPMLE) of F
X
. The empirical
distribution is defined as
F
X
(x)=
1
n
n
i=1
I(X
i
≤ x),
128 The Art of Estimation (I): M-estimation
where X
i
=(X
i
(1),...,X
i
(p)) ≤ x =(x(1),...,x(p)) means X
i
(j) ≤ x(j),
j =1,...,p. One property of the empirical distribution function is that, for
any function h(x), we have
h(x)d
F
X
(x)=
1
n
n
i=1
h(X
i
).
Utilizing this property, we see that
θ
g-comp
=
1
n
n
i=1
$
Q(X
i
, 1) −
Q(X
i
, 0)
%
=
$
Q(x, 1) −
Q(x, 0)
%
d
F
X
(x)=θ(
Q,
F
X
).
Therefore,
θ
g-comp
= θ(
Q,
F
X
) is a plug-in estimator of θ = θ(Q, F
X
).

7.3.2 MLE

We have seen that the g-computation estimator is a plug-in estimator. Based
on this idea, we can construct any number of g-computation estimators: (1)
obtain an estimator
Q
via a generalized linear model (GLM) or any other pre-
dictive model (Hastie, Tibshirani, and Friedman 2009); (2) consider empirical
distribution
F
X
; (3) obtain a g-computation estimator
θ
g-comp
= θ(
Q
,
F
X
).
In this subsection, we focus on g-computation estimator when Q is esti-
mated via GLM. We start with specifying a GLM, Q(x, a; β), for Q(x, a).
Continuous outcome
For continuous outcome Y , we consider the following linear model,
Q(X, A; β)=β(0) + β(1)A + β
T
(2)X + β
T
(3)AX, (7.23)
where β =(β(0),β
T
(2),β
T
(3))
T
. We can obtain the minimum loss estimator
(MLE) of β,
β =argmin
n
i=1
[Y
i
− Q(X
i
,A
i
; β)]
2
, (7.24)
where the loss function is the L
2
loss function. Estimator (7.24) is equivalent
to the solution of the following estimating equation,
n
i=1
∂Q(X
i
,A
i
; β)
∂β
[Y
i
− Q(X
i
,A
i
; β)] = 0, (7.25)
G-computation Estimator 129
where
∂Q(X,A; β)
∂β
=
⎛
⎜
⎜
⎝
1
A
X
AX
⎞
⎟
⎟
⎠
. (7.26)
Binary outcome
For binary outc o m e Y , we consider the following logistic model,
logit[Q(X, A; β)] = β(0) + β(1)A + β
T
(2)X + β
T
(3)AX, (7.27)
where logit[Q]=log[Q/(1 −Q)] and β =(β(0),β
T
(2),β
T
(3))
T
. We can obtain
the maximum likelihood estimator (MLE) of β,
β =argmax
n
+
i=1
Q(X
i
, 1; β)
Y
i
Q(X
i
, 0; β)
1−Y
i
, (7.28)
where the maximization objective function is the likelihood function. It is
equivalent to the minimum loss estimator (MLE),
β =argmin
n
i=1
−[Y
i
log{Q(X
i
, 1; β)} +(1− Y
i
)log{Q(X
i
, 0; β)}] , (7.29)
where the loss function is the negative-log-likelihood loss function. Estimator
(7.29) is equivalent to the solution of the following estimating equation,
n
i=1
∂Q(X
i
,A
i
; β)/∂β
Q(X
i
,A
i
; β)[1 − Q(X
i
,A
i
; β)]
[Y
i
− Q(X
i
,A
i
; β)]=0, (7.30)
where
∂Q(X,A; β)/∂β
Q(X, A; β)[1 − Q(X, A;β)]
=
⎛
⎜
⎜
⎝
1
A
X
AX
⎞
⎟
⎟
⎠
. (7.31)
After we obtain an estimator of β, via linear model for continuous out-
come or logistic model for binary outcome, we can obtain an estimator of Q,
Q
MLE
(X, A)=Q(X, A;
β). Consequently, we obtain a plug-in estimator for θ,
θ
MLE
= θ(
Q
MLE
,
F
X
)=
1
n
n
i=1
$
Q(X
i
, 1;
β) − Q(X
i
, 0;
β)
%
. (7.32)
Here is a remark on why
θ
MLE
is MLE of θ. This is due to the invariance
property of MLE: if h(μ) is the parameter of interest and μ
MLE
is MLE of μ,
then h(μ
MLE
) is MLE of h(μ). Using the invariance property, since
β is MLE
of β,
Q
MLE
(X, A)=Q(X, A;
β) is MLE of Q(X, A). Using the invariance
property one more time, since
Q
MLE
is MLE of Q and
F
X
is NPMLE of F
X
,
θ
MLE
= θ(
Q
MLE
,
F
X
) is MLE of θ = θ(Q, F
X
).
130 The Art of Estimation (I): M-estimation

7.3.3 Asymptotic variance

To derive an explicit expression of the asymptotic variance of
θ
MLE
, we consider
M-estimation. Here we consider the setting of continuous outcome, and the
discussion can be extended to binary outcome easily.
Let μ =(θ, β
T
)
T
and M (O; μ)=(M
1
(O, μ),M
T
2
(O, μ))
T
,where
M
1
(O, μ)=Q(X, 1; β) − Q(X, 0; β) − θ, (7.33)
M
2
(O, μ)=[Y − Q(X, A; β)]∂Q(X,A; β)/∂β. (7.34)
Let θ
0
bethetruevalueofθ and β
0
bethetruevalueofβ if the parametric
model Q(X,A; β) is a correctly specified model for the outcome regression
function. Then μ
0
=(θ
0
,β
T
0
)
T
is the true value of μ.
The meat term in the sandwich formula is
E
M
⊗2
(O; μ
0
)
=
E[Q(X, 1;β
0
) − Q(X, 0; β
0
) − θ
0
]
2
0
T
0 A
0
, (7.35)
where
A
0
= E
[Y −Q(X,A; β
0
)]
2
[∂Q(X,A; β
0
)/∂β]
⊗2
.
The bread term in the sandwich formula is
E
∂M(O, μ
0
)
∂μ
T
=
−1 B
0
0 −C
0
, (7.36)
whose inverse is
E
∂M(O, μ
0
)
∂μ
T

−1
=
−1 −B
0
C
−1
0
0 −C
−1
0
, (7.37)
where
B
0
= E
∂Q(X,1; β
0
)/∂β
T
− ∂Q(X,0; β
0
)/∂β
T
= E
(1, 1,X
T
,X
T
) − (1, 0,X
T
, 0)
=(0, 1, 0, E{X
T
})
and
C
0
= E{[∂Q(X,A; β
0
)/∂β]
⊗2
}.
Thus, we derive the consistency and asymptotic normality of
θ
MLE
,
θ
MLE
→ θ
0
, (7.38)
√
N(
θ
MLE
− θ
0
) →N(0,σ
2
0,
MLE
), (7.39)
where
σ
2
0,
MLE
= E[Q(X, 1;β
0
) − Q(X, 0; β
0
) − θ
0
]
2
+ B
0
C
−1
0
A
0
C
−1
0
B
T
0
, (7.40)

Inverse Probability Weighted Estimator 131

which can be estimated by
σ
2
MLE
=
1
n
n
i=1
$
Q(X
i
, 1;
β) − Q(X
i
, 0;
β) −
θ
MLE
%
2
+
B
C
−1
A
C
−1
B
T
,
where
A =
1
n
n
i=1
)
[Y
i
− Q(X
i
,A
i
;
β)]
2
[∂Q(X
i
,A
i
;
β)/∂β]
⊗2
*
,
B =
1
n
n
i=1
)
∂Q(X
i
, 1;
β)/∂β
T
− ∂Q(X
i
, 0;
β)/∂β
T
*
,
C =
1
n
n
i=1
[∂Q(X
i
,A
i
;
β)/∂β]
⊗2
.
7.3.4 Influence function
To summarize, we show that
θ
MLE
is an RAL estimator,
√
n(
θ
MLE
− θ
0
)=
1
√
n
n
i=1
φ
MLE
(O
i
)+o
p
(1), (7.41)
where φ
MLE
(O) is the corresponding influence function and
φ
MLE
(O)=Q(X, 1; β
0
) − Q(X, 0; β
0
)
+ B
0
C
−1
0
∂Q(X,A; β
0
)
∂β
[Y −Q(X,A; β
0
)] − θ
0
. (7.42)
7.4 Inverse Probability Weighted Estimator

7.4.1 IPW estimator

Consider the statistical estimand in (7.4), which depends on the propensity
score function g(a|x). Thus, we can rewrite (7.4) as
θ = E
I(A =1)
g(1|X)
Y
− E
I(A =0)
g(0|X)
Y
. (7.43)
Hence, if we can obtain an estimator of g(a|x), denoted as g(a|x), and use
the sample mean to estimate the population mean, then we can obtain an
IPW estimator of (7.43),
θ
IPW
=
1
n
n
i=1
I(A
i
=1)
g(1|X
i
)
Y
i
−
1
n
n
i=1
I(A
i
=0)
g(0|X
i
)
Y
i
. (7.44)
132 The Art of Estimation (I): M-estimation

7.4.2 Asymptotic variance

To derive an explicit expression of the asymptotic variance of
θ
IPW
, we consider
M-estimation. We start with specifying a parametric model for the propensity
score function. For example, we may specify a logistic regression model,
g(1|X; γ)=1− g(0|X; γ)=
exp(γ
T
&
X)
1+exp(γ
T
&
X)
, where
&
X =
1
X
. (7.45)
Based on the observed data {O
i
=(X
i
,A
i
,Y
i
),i=1,...,n}, we can obtain
MLE of γ,
γ =argmax
n
+
i=1
g(1|X
i
; γ)
A
i
g(0|X
i
; γ)
1−A
i
,
which is equivalent to
γ =argmin
n
i=1
−[A
i
log{g(1|X
i
; γ)} +(1− A
i
)log{g(0|X
i
; γ)}] ,
which is further equivalent to the solution to the following estimating equation,
n
i=1
∂g(1|X
i
; γ)/∂γ
g(1|X
i
; γ)g(0|X
i
; γ)
[A
i
− g(1|X
i
; γ)] = 0, (7.46)
which can be simplified as
n
i=1
&
X
i
[A
i
− g(1|X
i
; γ)]=0, (7.47)
noting that
∂g(1|X; γ)/∂γ
g(1|X; γ)g(0|X; γ)
=
&
X. (7.48)
Since g(a|X)=g(a|X, γ),
θ
IPW
in (7.44) becomes
θ
IPW
=
1
n
n
i=1
I(A
i
=1)
g(1|X
i
; γ)
Y
i
−
1
n
n
i=1
I(A
i
=0)
g(0|X
i
; γ)
Y
i
. (7.49)
Now we are ready to derive the asymptotic normality of
θ
IPW
in (7.49)
within the framework of M-estimation. Let μ =(θ, γ
T
)
T
,andM(O; μ)=
(M
1
(O, μ),M
T
2
(O, μ))
T
,where
M
1
(O, μ)=
AY
g(1|X; γ)
−
(1 − A)Y
g(0|X; γ)
− θ, (7.50)
M
2
(O, μ)=[A − g(1|X; γ)]
&
X. (7.51)
Inverse Probability Weighted Estimator 133
Let θ
0
be the true value of θ and γ
0
be the true value of γ if the parametric
model g(a|X; γ) is a correctly specified model for the propensity score function.
Then μ
0
=(θ
0
,γ
T
0
)
T
is the true value of μ.
The meat term in the sandwich formula is
E
M
⊗2
(O; μ
0
)
=
E
$
AY
g(1|X;γ
0
)
−
(1−A)Y
g(0|X;γ
0
)
− θ
%
2
0
T
0 D
0
,
where the off-diagonal elements are equal to zero by the law of iterated ex-
pectations and
D
0
= E
)
[A − g(1|X; γ
0
)]
2
&
X
⊗2
*
. (7.52)
The bread term in the sandwich formula is
E
∂M(O, μ
0
)
∂μ
T
=
−1 E
0
0 −F
0
,
whose inverse is
E
∂M(O, μ
0
)
∂μ
T

−1
=
−1 −E
0
F
−1
0
0 −F
−1
0
, (7.53)
where
E
0
= −E
[A − g(1|X; γ
0
)]
2
Y
g(1|X; γ
0
)g(0|X; γ
0
)
&
X
T
,
F
0
= E
)
g(1|X; γ
0
)g(0|X; γ
0
)
&
X
⊗2
*
.
Thus, if g(a|x; γ) is a correct model for g(a|x), we derive the consistency
and asymptotic normality of
θ
IPW
,
θ
IPW
→ θ
0
, (7.54)
√
n(
θ
IPW
− θ
0
) →N(0,σ
2
0,
IPW
), (7.55)
where
σ
2
0,
IPW
=
AY
g(1|X; γ
0
)
−
(1 − A)Y
g(0|X; γ
0
)
− θ
0
2
+ E
0
F
−1
0
D
0
F
−1
0
E
T
0
, (7.56)
which can be estimated by
σ
2
IPW
=
1
n
n
i=1
A
i
Y
i
g(1|X
i
; γ)
−
(1 − A
i
)Y
i
g(0|X
i
; γ)
−
θ
IPW
2
+
E
F
−1
D
F
−1
E
T
,
134 The Art of Estimation (I): M-estimation
where
D =
1
n
n
i=1
[A
i
− g(1|X
i
; γ)]
2
&
X
⊗2
i
,
E = −
1
n
n
i=1
[A
i
− g(1|X
i
; γ)]
2
Y
i
g(1|X
i
; γ)g(0|X
i
; γ)
&
X
T
i
,
F =
1
n
n
i=1
g(1|X
i
; γ)g(0|X
i
; γ)
&
X
⊗2
i
.
7.4.3 Influence function
To summarize, we show that
θ
IPW
is an RAL estimator,
√
n(
θ
IPW
− θ
0
)=
1
√
n
n
i=1
φ
IPW
(O
i
)+o
p
(1), (7.57)
where φ
IPW
(O) is the corresponding influence function and
φ
IPW
(O)=
AY
g(1|X; γ
0
)
−
(1 − A)Y
g(0|X; γ
0
)
+ E
0
F
−1
0
&
X[A − g(1|X; γ
0
)] − θ
0
. (7.58)

7.5 Augmented Inverse Probability Weighted Estimator

7.5.1 A class of estimators

We h ave shown that
θ
IPW
is consistent and asymptotically normal if the
propensity score function is correctly specified. Furthermore, it is desirable
to construct some estimator that is not only consistent and asymptotically
normal but also has the smallest asymptotic variance. We refer to such an
estimator as an efficient estimator (Bickel et al. 1993; Tsiatis 2006).
Let’s start with the following class of estimators,
θ
(H)=
1
n
n
i=1
A
i
Y
i
g(1|X
i
; γ
0
)
−
(1 − A
i
)Y
i
g(0|X
i
; γ
0
)
− H(A
i
,X
i
)
,
where H(A, X) is any function satisfying E[H(A, X)] = 0. Because of (7.43),
we can easily verify that
θ(H) is an unbiased estimator of θ
0
for any H(A, X)
satisfying that E[H(A, X)] = 0.
In addition, we can know the form of H(A, X) to some extent. In fact,
because A is binary and E[H (A, X)] = 0, we have
H(1,X)g(1|X; γ
0
)+H(0,X)g(0|X; γ
0
)=0,
Augmented Inverse Probability Weighted Estimator 135
implying that H(1,X)=−H(0,X)g(0|X ; γ
0
)/g(1|X; γ
0
). Thus, we have
H(1,X)=−H(0,X)[1 − g(1|X; γ
0
)]/g(1|X; γ
0
),
H(0,X)=−H(0,X)[0 − g(1|X; γ
0
)]/g(1|X; γ
0
).
Letting h(X )=−H (0,X)/g(1|X; γ
0
), we show that
H(A, X)=[A − g(1|X; γ
0
)]h(X).
Therefore, class {
θ(H)} is equivalent to the following class of estimators,
θ
(h)=
1
n
n
i=1
A
i
Y
i
g(1|X
i
; γ
0
)
−
(1 − A
i
)Y
i
g(0|X
i
; γ
0
)
− [A − g(1|X; γ
0
)]h(X
i
)
,
where h(X ) is any function of X .
If the propensity score model is correctly specified and γ
0
is estimated by
γ, the above class of estimators becomes the following class of estimators,
θ(h)=
1
n
n
i=1
A
i
Y
i
g(1|X
i
; γ)
−
(1 − A
i
)Y
i
g(0|X
i
; γ)
− [A
i
− g(1|X
i
; γ)]h(X
i
)
, (7.59)
where h(X ) is any arbitrary function of X . Note that
θ
IPW
is the special case
of (7.59) with h(X) ≡ 0.
We can easily show that for any h,
θ(h) is consistent if the propensity score
function is correctly specified. In fact, if γ
0
is the true value of γ,
E {[A − g(1|X ; γ
0
)]h(X)}
(a)
= E {E[A − g(1|X ; γ
0
)]h(X)|X}
(b)
=E {h(X)E[A − g(1|X; γ
0
)]|X} =0,
where (a) is using the law of iterated expectations and (b) holds because
E(A|X)=g(1|X; γ
0
).

7.5.2 Asymptotic variances

We can apply the M-estimation theory to derive the asymptotic variance of
θ(h). Assume logistic model (7.45) for the propensity score function. Let μ =
(θ, γ
T
)
T
and M
h
(O; μ)=(M
h1
(O; μ),M
T
h2
(O; μ))
T
,where
M
h1
(O; μ)=
AY
g(1|X; γ)
−
(1 − A)Y
g(0|X; γ)
− [A − g(1|X; γ)]h(X) − θ, (7.60)
M
h2
(O; μ)=[A − g(1|X ; γ)]
&
X. (7.61)
The meat term in the sandwich formula is
E
M
⊗2
h
(O; μ
0
)
=
V
0
(h)0
T
0 D
0
,