Decision-Making under Uncertainty
bandits · RL · OPE · DTR/OTR · policy learning
추정된 효과를 결정으로 — optimal policy learning, bandits·reinforcement learning, off-policy evaluation, dynamic/optimal treatment regimes.
노트 21개
-
OPE 배포 게이트 플레이북 — 참값 없이 "믿는다/못 믿는다/A/B"를 판정하기
참값 V(π_e) 를 아무도 모르는 실무 조건에서, 추천 로그만으로 배포를 판정하는 OPE 게이트의 운영 규칙 — 진단 3종의 3-way 게이트, validity battery(필요조건 검사), Λ-감도 증명서, 비교형 우선 원칙, 그리고 관측 동등성이라는 원리적 한계까지. ope-to-decision 벤치마크의 플레이북 큐레이션.
-
From Estimation to Action — How HTE Drives Personalized Policy Across Domains
One methodological spine — estimate heterogeneous treatment effects and turn them into individual-level policies — powers both clinical sequential treatment decisions and industrial targeting, pricing, and recommendation.
-
Marketing Attribution at Scale — From Simulation to Causal Inference
A case study comparing 10+ multi-touch attribution methods against a known-ground-truth simulator, then scaling them on the public Criteo dataset, closing the loop with budget off-policy evaluation for channel allocation.
-
RTB Bidding Strategy via Causal ML — From Prediction to Optimization
A five-stage case study on the public iPinYou RTB dataset that moves from pCTR/pCVR prediction through causal effect estimation (CATE, SCM) to budget-constrained optimal bidding and off-policy policy evaluation.
-
Sequential and Adaptive Decision-Making — From Bandits to Dynamic Treatment Regimes
A synthesis essay tracing one methodological spine through sequential decision-making under uncertainty — exploration–exploitation in bandits, off-policy evaluation, and optimal/dynamic treatment regimes — that powers clinical adaptive trials and real-time bidding alike.
-
Anytime-Valid Inference Overview
Anytime-valid 추론은 고정 표본 가설검정의 "peeking" 문제를 푸는 game-theoretic statistics다. 식별-타당성 drift를 실시간으로 모니터링하는 안전 추론의 수학적 기초를 제공한다.
-
Anytime-Valid OPE
Anytime-valid OPE는 임의의 정지 시점에서도 유효한(time-uniform) 오프폴리시 가치 신뢰 수열(confidence sequence)을 e-process 기반으로 제공하는 오프폴리시 평가(OPE) 방법이다.
-
Confidence Sequence
신뢰 수열(confidence sequence, CS) $(Ct){t\ge1}$은 time-uniform 커버리지(coverage)를 갖는 신뢰구간의 열이다. 즉 모든 시점에서 동시에 다음을 만족한다.
-
Decision-Making Overview
불확실성하 (순차) 의사결정은 bandit 후회(regret)부터 RL, 오프폴리시 평가(OPE), 동적·최적 처치 요법(DTR/OTR)까지를 아우르는 방법론이다. 임상(DTR/OTR)과 산업(타겟팅·bidding) 개인화(personalization)를 모두 받친다.
-
Dynamic Treatment Regimes (DTR / OTR)
동적 처치 요법(DTR)은 누적 이력 $Ht$(공변량·이전 처치·중간결과)를 처치로 사상하는 결정규칙 열 $\{dt(Ht)\}{t=1}^T$이다. 최적 처치 요법(optimal treatment regime, OTR)은 기대 장기결과 $E[Y^{d}]$를 최대화하는 요법을 가리킨다. 다음 방법으로 추정한다.
-
e-process (e-value)
e-value $E$는 귀무가설 $H0$가 참일 때 $EP[E]\le 1$ ($\forall P\in H0$)을 만족하는 비음 확률변수다. e-process $(Et)$는 임의의 정지시각 $\tau$에 대해 $E\tau$가 e-value가 되는 비음 과정($E[E\tau]\le1$)으로, 보통 귀무가설 아래에서…
-
Multi-Armed Bandits
multi-armed bandit은 $K$개의 arm 중 매 라운드 $t$에 하나의 arm $At$를 당겨 보상(reward)을 관측하는 순차적 의사결정 문제다. 목표는 누적 후회(cumulative regret)를 최소화하는 것이다.
-
Off-Policy Evaluation (OPE)
오프폴리시 평가(OPE)는 다른 행동 정책(behavior policy) $\pib$로 수집한 로그만으로 목표 정책(target policy) $\pie$의 가치 $V(\pie)=E{\pie}[\sum r]$를 추정하는 문제다.
-
A/B Testing
A/B 테스트는 무작위 대조 실험(RCT)을 온라인에 적용한 것으로, 둘 이상의 변형(variant)을 사용자에게 무작위로 노출해 인과 효과를 추정하는 방법이다.
-
Contextual Bandits
Contextual Bandits는 맥락(context)에 따라 최적의 행동(arm)이 달라지는 다중 슬롯머신 문제다. 즉 매 라운드 관측되는 맥락에 맞춰 행동을 고르고, 그 행동의 보상(reward)만 피드백으로 받으며 정책을 학습한다.
-
CUPED
CUPED (Controlled-experiment Using Pre-Experiment Data)는 사전 실험 데이터로 A/B 테스트 추정량의 분산(variance)을 줄여 검정 민감도를 높이는 분산 감소(variance reduction) 기법이다.
-
Design Effect
설계 효과(Design Effect, DEFF)는 복잡한 표집 설계가 단순 무작위 표집에 비해 분산을 얼마나 키우는지를 나타내는 배율이다.
-
MDP (Markov Decision Process)
마르코프 결정 과정(Markov Decision Process, MDP)은 순차적 의사결정 문제를 다루는 수학적 프레임워크다. 어떤 상태(state)에서 행동(action)을 선택하면 다음 상태와 보상(reward)이 확률적으로 결정되며, 이런 상호작용이 시간에 걸쳐 반복되는 구조를 기술한다.
-
Policy Trees
정책 트리(policy tree)는 Athey & Wager (2021)가 제안한 해석 가능한 정책 학습(policy learning) 방법이다. 개인 수준의 처치효과(treatment effect) 추정값(estimate)을 입력으로 받아, 결정 트리 형태의 명시적 정책 규칙을 학습한다.
-
Statistical Power
통계적 검정력(statistical power)은 효과가 실제로 존재할 때 그것을 탐지할 확률이다.
-
Thompson Sampling
Thompson sampling은 보상에 대한 베이지안 사후 분포에서 표본을 뽑아 탐색(exploration)과 활용(exploitation)의 균형을 맞추는 bandit 알고리즘이다.