Tae Hyun Kim (Lowell)
← All projects
Decision-Making under UncertaintyCausal Inference

OPE → Decision — A Deployment Gate That Never Sees the Truth

An OPE deployment gate that rules "trust / distrust / send it to A/B" from recommendation logs alone, under the real-world condition that nobody knows the true value $V(\pi_e)$ — plus a benchmark that scores those signals' forecasting power and in-principle blind spots on a truth-holding backstage.

2026 · Solo · end-to-end (estimator implementations → 21 experiment axes → deployment-gate playbook)
PythonNumPyscikit-learnobp (cross-validation)Open Bandit Dataset

⏱️ TL;DR (30 seconds)


validity battery detection matrix — the failures it catches and the blanks it cannot see What the battery catches (misrecording · support deficiency — firing rate 1.0) and the blank it cannot see in principle (calibrated confounding — 0 detection-arm firings against an actual large-error rate of 0.55). The blank is exhibited as-is, not hidden.

ZOZO real-log 1-page decision card — a verdict with no reveal One round for real: the decision card on the ZOZO OBD production log. Estimates+CIs, diagnostics, battery, the Λ\Lambda fan, and the verdict — all from the log and the candidate distribution alone. On this axis, no answer sheet (no reveal) exists.

🎯 Key Results at a Glance

Key metricValueNotes
Experiment axes21 (12 synthetic + 3 real-data + 1 sensitivity + 5 GT-unknown frontstage)axis 13 honestly dropped on a probe NO-GO
Estimator cross-validation7 estimators × 2 tracks, rel_diff ≤ 1e-8adversarial checks against obp (py3.9) · sb-obp (py3.12)
Gate forecasting power (backstage)large-error rate 4.6% on trust vs 44.4% on A/B fallbackpre-registered thresholds · evaluated with no tuning
Decision value (contaminated log)naive false-go 0.9 → protocol 0.0candidate ladder includes a genuinely bad policy · regret 0.21
Observational-equivalence boundarybattery firings 0/240 · bias growing to −0.073a constructive exhibit of what it cannot see
Real-data replicationnoised · support detection replicates + 1 expectation refutedinjections into the 4 c2b datasets — the refutation logged as-is
ZOZO real-log verdictAB_FALLBACK · fragile (Λ\Lambda^* 1.34)harmonic firing (a recorded-propensity inconsistency signal)
Tests100 greenincluding blindness, oracle-ban, and checksum barriers

Number status — Every number on this page is an abbreviated transcription of values registered in the repo’s canonical number ledger (docs/LEDGER.md). The repo is canonical for raw values, precision, and generating experiments (figure ↔ CSV 1:1), and the gate/battery rules are this repo’s proposal — a systematization of scattered folklore practice, not an established standard.

🧩 The Two-Stage Framework — frontstage / backstage

frontstage / backstage structure diagram

A practitioner only ever holds the left side (the frontstage): the logged layer → gate + battery + Λ\Lambda^* → verdict and decision — no ground truth needed. The grounds for trusting the left side are manufactured on the right (the backstage): on truth-holding stages (synthetic DGP · fully-labeled c2b · approximate-GT ZOZO), the frontstage signals were scored blind-then-reveal. And the line dividing the two — the observational-equivalence boundary — rides along as a disclaimer on every frontstage “pass” verdict.

💡 Key Findings — Three Counter-Intuitive Results

1. Failures the diagnostics cannot see are real — and in practice, the failure itself is invisible.

confounding blind spot — diagnostics flat, bias alone growing As confounding strength grows, ESS computed from recorded propensities stays flat (0.823→0.822) while the bias alone grows (0→−0.057). In practice, all you ever see is the flat top panel — the bottom panel (the bias) is visible only on a backstage that knows the truth. This figure is the motivation for the entire GT-unknown protocol.

2. The battery narrows the blind spot but cannot eliminate it — and that is a provable boundary. A confounded world whose recorded pscores are the true marginals is exactly identical, in observed distribution, to a sound world (observational equivalence). Measured: in this world the battery fires 0/240 across all arms while the bias grows monotonically. The only log-computable signal that moves is the contraction of Λ\Lambda^* (1.31→1.05) — not a confounding detection, but a report that the conclusion is turning fragile without knowing why.

3. A pre-registered expectation was refuted on real data — and that is exactly why the replication axis exists. In the synthetic world, in-sample estimated propensities were quasi-null (the auto-correction trap — no firing); but the real-data (c2b) double-softmax policy geometry cannot be recovered by an in-sample LR, so E[w] fired on every dataset. The refutation of a pre-registered expectation is logged as-is — a refuted expectation is a finding too.

📊 Results Summary

The frontstage (GT-unknown) track — computed and ruled from logs alone:

AxisFinding
Real-log gateZOZO production log ruled DISTRUST (ESS/n 0.034 · max weight 278) — the approximate GT (±32% CI) backs the verdict
Sensitivity certificateThe breakdown Λ\Lambda^* at which ranking assertions collapse has median ≈1.07/≈1.04 — the computation needs only the log; Λ\Lambda is an assumption not identifiable from data
Detection matrixMisrecording · support deficiency caught (per-action harmonic recovers even where every global signal dies) · small-sample false-alarm rate 0.275 recorded honestly
Boundary exhibit0/240 firings · bias −0.073 · Λ\Lambda^* contraction — “neither recording mode fires, only the bias grows”
Decision valuenaive false-go 0.9 vs protocol 0.0 · decisive on healthy logs at go 0.95/no-go 0.9 (zero errors)
Real-log cardAB_FALLBACK · fragile — the very absence of a reveal file is part of the exhibit
Replicationinjections into the 4 c2b datasets — detection replicates + 1 expectation refuted + over-caution (fires even on harmless injections — a deferral cost)

The backstage — the grounds for trusting the signals: on the 12 truth-holding synthetic axes, the estimator regime map (the DR family trading dominance · diagnostic power going missing in small samples — a cell ruled trust while its IPS MSE is 70.6× the winner’s), gate forecasting power 4.6%/44.4%, DR misspecification robustness replicating 4/4 on real data, and the demonstration that bootstrap CIs cannot catch structural bias (9/28 fail to cover) — every number flows through a single ledger (LEDGER).

⚠️ Limitations & Lessons Learned

LimitationDetail
The battery is a necessary-condition screenPassing is not proof of soundness — blind in principle to calibrated confounding (observational equivalence). This disclaimer is attached to every battery claim
Thresholds are a proposal, untunedGate and battery thresholds were pre-registered and then only evaluated — porting them to another domain without recalibration is unfounded
Small-sample trapsGate power goes missing in small samples, and the harmonic arm false-alarms in small samples (0.275) — a trap in both directions
Over-cautionThe battery fires even on harmless injections, producing A/B deferrals — the price of safety is paid in deferral
Out of scopeThe confounding-correction mainline (proximal methods and kin), OPL, and slate/RL OPE are not covered — this repo deliberately stops at the boundary

Lessons

“In a world without ground truth there are only three honest answers — a necessary-condition screen for the failures you can catch, a sensitivity interval for the failures you cannot, and the courage to output I don’t know (fall back to A/B).”


Artifacts