Doubly Robust Estimator
정의
Doubly Robust (DR) Estimator는 결과 회귀(outcome regression)와 성향점수(propensity score) 모델을 결합한 추정량(estimator)으로, 두 모델 중 하나만 올바르게 설정(specified)되어도 일치성(consistency)을 갖는다.
ATE에 대한 DR Estimator:
여기서 pseudo-outcome은 효율적 영향함수(EIF)다:
또는 동치 형태로도 쓸 수 있다:
직관적 이해
세 가지 추정 전략의 결합:
- 결과 회귀 (OR):
- 역성향 가중 (IPW):
- 이중 강건 (DR): OR에 IPW 보정항(correction term)을 더한 형태
OR만 사용: μ̂ 틀리면 biased
IPW만 사용: π̂ 틀리면 biased
DR 사용: μ̂ OR π̂ 중 하나만 맞아도 consistent!
왜 “Doubly Robust”인가?
- (결과 모델이 맞을 때): augmentation 항의 기댓값이 0이 된다.
- (성향 모델이 맞을 때): 가중이 정확해 편향(bias)이 상쇄된다.
- 둘 다 틀려도 편향이 두 오차의 곱에 비례한다:
핵심 성질
이중 강건성 (Double Robustness)
정리: 다음 두 조건 중 하나만 성립해도 는 일치성을 갖는다.
- 결과 모델이 올바르게 설정된 경우:
- 성향 모델이 올바르게 설정된 경우:
준모수 효율성 (Semiparametric Efficiency)
DR 추정량은 **준모수적으로 효율적(semiparametrically efficient)**이다. 효율적 영향함수를 기반으로 구성되어 준모수 효율성 한계(semiparametric efficiency bound)를 달성하며, 가장 낮은 점근 분산(asymptotic variance)을 갖는다.
속도 이중 강건성 (Rate Double Robustness)
곱 속도 조건(product rate condition)에서 -일치성을 갖는다:
예를 들어 각 모델이 속도(rate)면 충분하다.
수학적 유도
효율적 영향함수
ATE 의 효율적 영향함수는 다음과 같다:
성질:
- 준모수 분산 한계
- Neyman 직교성(Neyman orthogonal):
편향 분석
이처럼 편향이 두 오차의 곱(product of errors) 형태로 나타난다.
비교: OR vs IPW vs DR
| Aspect | Outcome Regression | IPW | Doubly Robust |
|---|---|---|---|
| Model needed | Both | ||
| Consistency | If correct | If correct | If either correct |
| Efficiency | Not efficient | Not efficient | Semiparametrically efficient |
| Variance | Low if good | High with extreme | Best of both |
| With ML | Regularization bias | Variance issues | Robust to both |
확장
CATE 추정
DR-Learner는 DR pseudo-outcome을 에 대해 회귀한다.
ATT 추정
종단 설정 (Longitudinal Settings)
시변 처치(time-varying treatment)에 적용하며, g-computation과 결합한다.
관련 개념
- Pseudo-outcome - DR 추정량의 핵심 구성요소
- DR-Learner - CATE를 위한 DR 확장
- Influence Function - DR의 이론적 기반
- Neyman-Orthogonal Score - 직교성(orthogonality) 성질
- Propensity Score - 처치 배정 확률
- Double-Debiased ML - 관련 프레임워크
역사적 배경
- Robins, Rotnitzky, Zhao (1994): 최초의 이중 강건 추정량.
- Bang & Robins (2005): “Doubly Robust Estimation”이라는 이름이 여기서 붙었다.
- Scharfstein, Rotnitzky, Robins (1999): 준모수 이론과의 연결을 보였다.
- Chernozhukov et al. (2018): ML과 결합한 DML로 이어진다.
구현
Python (econml):
from econml.dr import LinearDRLearner
dr = LinearDRLearner()
dr.fit(Y, T, X=X, W=W)
ate = dr.ate(X)
R (AIPW package):
library(AIPW)
AIPW_SL <- AIPW$new(Y = Y, A = A, W = W,
Q.SL.library = c("SL.glm", "SL.ranger"),
g.SL.library = c("SL.glm", "SL.ranger"))
AIPW_SL$fit()
AIPW_SL$summary()
참고 문헌
- Robins, Rotnitzky, Zhao (1994) - Original DR estimator
- kennedyOptimalDoublyRobust2023 - Optimal DR for CATE
- chernozhukovDoubleDebiasedMachine2018 - DML framework
- Bang & Robins (2005) - “Doubly Robust Estimation in Missing Data and Causal Inference Models”