Tae Hyun Kim (Lowell)
Data Scientist · Data Lab, Hanwha General Insurance · Assistant Manager · Seoul, Korea
Personalized decision-making under uncertainty, made causal.
I am a data scientist whose research centers on causal inference and data-driven decision-making for high-stakes, individualized recommendations. To that end I draw on statistical and probabilistic modeling, optimization, ML/AI, and LLM-based agentic systems — approaching applied problems, across both industry and medical research, through counterfactual reasoning and policy learning rather than pure prediction.
Research Interests
Research Pillars →Causal Inference
CATE · counterfactual · causal discovery · SCM · semiparametric · partial ID
Identifying what causes what — heterogeneous treatment effects and counterfactuals — with structural causal models, semiparametric estimation, and sensitivity / partial-identification under unobserved confounding.
51 notes → 02Decision-Making under Uncertainty
bandits · RL · OPE · DTR/OTR · policy learning
Turning estimated effects into decisions — optimal policy learning, bandits and reinforcement learning, off-policy evaluation, and dynamic / optimal treatment regimes.
20 notes → 03Personalization
HTE · targeting · recommendation · pricing
The through-line — individualized clinical treatment decisions and industry targeting, recommendation, and pricing as two sides of one methodological core.
20 notes →Selected Projects
All 3 projects →- 01
Customer Segmentation & Causal Targeting
An end-to-end analysis on the public Dunnhumby retail dataset — NMF + K-Means segmentation feeding meta-learner / Causal-Forest HTE and an OPE-validated targeting policy.
- 02
Causal Multi-Touch Attribution
A simulation study unifying Incremental·Shapley channel credit and a path-level decomposition on a single Inhomogeneous Poisson Process — answering channel budget, journey design, and population causal effect under one efficiency identity (18-method benchmark, ground-truth MAE 0.016,
- 03
The Chatbot You're Talking To
The grounded RAG assistant on this site — a safe LLM on a static Cloudflare edge that answers only from the published notes. The demo is the button at the bottom-right of this page.
Play & Learn
All demos →Touch the counter-intuitive results in one-minute interactive demos.
Recent Notes
All 82 notes →- 01
Dunnhumby — Track 1: Latent-Factor Customer Segmentation
NMF latent factors (92.44% explained variance) + K-Means yield 7 stable behavioral segments (Bootstrap ARI 0.77) with per-segment marketing actions. Illustrative case study on the public Dunnhumby retail dataset.
- 02
Dunnhumby — Track 2: Causal Targeting via Heterogeneous Treatment Effects
Meta-learner / Causal Forest CATE under severe positivity violation (PS AUC 0.989); an OPE-validated policy targets ~31% of customers and surfaces counter-intuitive negative-CATE segments. Hypothesis-generating on public data.
- 03
Applied Causal Inference for Pricing — CATE & SCM Across Public Datasets
An applied case study using only public datasets (LendingClub, iPinYou) that combines CATE estimation for price-sensitivity heterogeneity with SCM-based moderator analysis to design individual-level, risk-based pricing and RTB bidding policies — all findings illustrative and projected, not proprietary.
- 04
Causal Inference Under Partial Identification — Sensitivity and Evidence Hierarchies
When real-world data fail strong ignorability, point identification gives way to bounds, proxies, and sensitivity analysis — an honest hierarchy of evidence that connects credible causal claims to semiparametric efficiency.
- 05
Customer Segmentation
Customer Segmentation is the unsupervised task of partitioning customers into a finite set of segments by similarity in behavior, value, and preference. A common recipe is latent-factor decomposition followed by clustering: behavioral features → NMF (non-negative, parts-based decomposition) → factor scores → K-Means → segments.
- 06
Customer Segmentation & Causal Targeting — An Applied Case Study
An end-to-end applied case study on the public Dunnhumby dataset — NMF latent factors and K-Means segmentation feeding meta-learner / Causal Forest HTE and an OPE-validated optimal targeting policy, with a candid look at positivity violation and counter-intuitive "sleeping dog" segments.
Contact
About · CV →taehyun9573@gmail.com ·Seoul, Republic of Korea