#contextual-bandits
노트 2개
- Contextual Bandits Contextual Bandits는 맥락(context)에 따라 최적의 행동(arm)이 달라지는 다중 슬롯머신 문제다. 즉 매 라운드 관측되는 맥락에 맞춰 행동을 고르고, 그 행동의 보상(reward)만 피드백으로 받으며 정책을 학습한다.
- RTB Bidding Strategy via Causal ML — From Prediction to Optimization A five-stage case study on the public iPinYou RTB dataset that moves from pCTR/pCVR prediction through causal effect estimation (CATE, SCM) to budget-constrained optimal bidding and off-policy policy evaluation.