OPE → Decision — A Deployment Gate That Never Sees the Truth
An OPE deployment gate that rules "trust / distrust / send it to A/B" from recommendation logs alone, under the real-world condition that nobody knows the true value $V(\pi_e)$ — plus a benchmark that scores those signals' forecasting power and in-principle blind spots on a truth-holding backstage.
⏱️ TL;DR (30 seconds)
- Problem — We want to evaluate a new recommendation policy from logs alone, before any A/B test (off-policy evaluation). But practice adds one more condition — the real one: nobody knows the true value . In a place where no one will ever grade your estimate, what can you compute, and when must you stop?
- Approach — Assemble the signals computable from logs alone into a pre-registered protocol: three diagnostics → a 3-way gate, a validity battery (E[w] · per-action harmonic · placebo · disagreement — a necessary-condition screen), and the sensitivity certificate. The signals’ forecasting power and blind spots are scored blind-then-reveal on a truth-holding backstage (synthetic DGP · UCI · ZOZO production logs). All seven estimators are direct numpy implementations, cross-validated against obp (rel_diff ≤ 1e-8).
- Three headline results
- On a real log, the protocol stops itself — judging the ZOZO production log from the log alone: the battery fires → AB_FALLBACK, fragile. No answer sheet (no ground truth) exists — and that is exactly what practice looks like.
- The decision value is quantified — on a contaminated log, naive IPS point estimation green-lights a genuinely bad candidate with false-go 0.9 (mean regret 0.21); the protocol’s rate is 0.0 (fall back to A/B). On healthy logs it stays decisive (go 0.95 · no-go 0.9, zero errors).
- The boundary of what it cannot see is mapped too — confounding whose recorded propensities stay marginally calibrated is indistinguishable by any log statistic: battery firings sit at 0/240 while the bias alone grows to −0.073. The exit is not a point estimate but the width of a -sensitivity interval, plus abstention.
- Impact in one line — A verdict procedure that someone holding only logs can actually walk through, plus a map of where that procedure may — and may not — be trusted.
What the battery catches (misrecording · support deficiency — firing rate 1.0) and the blank it
cannot see in principle (calibrated confounding — 0 detection-arm firings against an actual
large-error rate of 0.55). The blank is exhibited as-is, not hidden.
One round for real: the decision card on the ZOZO OBD production log. Estimates+CIs,
diagnostics, battery, the fan, and the verdict — all from the log and the candidate
distribution alone. On this axis, no answer sheet (no reveal) exists.
🎯 Key Results at a Glance
| Key metric | Value | Notes |
|---|---|---|
| Experiment axes | 21 (12 synthetic + 3 real-data + 1 sensitivity + 5 GT-unknown frontstage) | axis 13 honestly dropped on a probe NO-GO |
| Estimator cross-validation | 7 estimators × 2 tracks, rel_diff ≤ 1e-8 | adversarial checks against obp (py3.9) · sb-obp (py3.12) |
| Gate forecasting power (backstage) | large-error rate 4.6% on trust vs 44.4% on A/B fallback | pre-registered thresholds · evaluated with no tuning |
| Decision value (contaminated log) | naive false-go 0.9 → protocol 0.0 | candidate ladder includes a genuinely bad policy · regret 0.21 |
| Observational-equivalence boundary | battery firings 0/240 · bias growing to −0.073 | a constructive exhibit of what it cannot see |
| Real-data replication | noised · support detection replicates + 1 expectation refuted | injections into the 4 c2b datasets — the refutation logged as-is |
| ZOZO real-log verdict | AB_FALLBACK · fragile ( 1.34) | harmonic firing (a recorded-propensity inconsistency signal) |
| Tests | 100 green | including blindness, oracle-ban, and checksum barriers |
Number status — Every number on this page is an abbreviated transcription of values registered in the repo’s canonical number ledger (
docs/LEDGER.md). The repo is canonical for raw values, precision, and generating experiments (figure ↔ CSV 1:1), and the gate/battery rules are this repo’s proposal — a systematization of scattered folklore practice, not an established standard.
🧩 The Two-Stage Framework — frontstage / backstage
A practitioner only ever holds the left side (the frontstage): the logged layer → gate + battery + → verdict and decision — no ground truth needed. The grounds for trusting the left side are manufactured on the right (the backstage): on truth-holding stages (synthetic DGP · fully-labeled c2b · approximate-GT ZOZO), the frontstage signals were scored blind-then-reveal. And the line dividing the two — the observational-equivalence boundary — rides along as a disclaimer on every frontstage “pass” verdict.
💡 Key Findings — Three Counter-Intuitive Results
1. Failures the diagnostics cannot see are real — and in practice, the failure itself is invisible.
As confounding strength grows, ESS computed from recorded propensities stays flat
(0.823→0.822) while the bias alone grows (0→−0.057). In practice, all you ever see is the flat
top panel — the bottom panel (the bias) is visible only on a backstage that knows the truth.
This figure is the motivation for the entire GT-unknown protocol.
2. The battery narrows the blind spot but cannot eliminate it — and that is a provable boundary. A confounded world whose recorded pscores are the true marginals is exactly identical, in observed distribution, to a sound world (observational equivalence). Measured: in this world the battery fires 0/240 across all arms while the bias grows monotonically. The only log-computable signal that moves is the contraction of (1.31→1.05) — not a confounding detection, but a report that the conclusion is turning fragile without knowing why.
3. A pre-registered expectation was refuted on real data — and that is exactly why the replication axis exists. In the synthetic world, in-sample estimated propensities were quasi-null (the auto-correction trap — no firing); but the real-data (c2b) double-softmax policy geometry cannot be recovered by an in-sample LR, so E[w] fired on every dataset. The refutation of a pre-registered expectation is logged as-is — a refuted expectation is a finding too.
📊 Results Summary
The frontstage (GT-unknown) track — computed and ruled from logs alone:
| Axis | Finding |
|---|---|
| Real-log gate | ZOZO production log ruled DISTRUST (ESS/n 0.034 · max weight 278) — the approximate GT (±32% CI) backs the verdict |
| Sensitivity certificate | The breakdown at which ranking assertions collapse has median ≈1.07/≈1.04 — the computation needs only the log; is an assumption not identifiable from data |
| Detection matrix | Misrecording · support deficiency caught (per-action harmonic recovers even where every global signal dies) · small-sample false-alarm rate 0.275 recorded honestly |
| Boundary exhibit | 0/240 firings · bias −0.073 · contraction — “neither recording mode fires, only the bias grows” |
| Decision value | naive false-go 0.9 vs protocol 0.0 · decisive on healthy logs at go 0.95/no-go 0.9 (zero errors) |
| Real-log card | AB_FALLBACK · fragile — the very absence of a reveal file is part of the exhibit |
| Replication | injections into the 4 c2b datasets — detection replicates + 1 expectation refuted + over-caution (fires even on harmless injections — a deferral cost) |
The backstage — the grounds for trusting the signals: on the 12 truth-holding synthetic axes, the estimator regime map (the DR family trading dominance · diagnostic power going missing in small samples — a cell ruled trust while its IPS MSE is 70.6× the winner’s), gate forecasting power 4.6%/44.4%, DR misspecification robustness replicating 4/4 on real data, and the demonstration that bootstrap CIs cannot catch structural bias (9/28 fail to cover) — every number flows through a single ledger (LEDGER).
⚠️ Limitations & Lessons Learned
| Limitation | Detail |
|---|---|
| The battery is a necessary-condition screen | Passing is not proof of soundness — blind in principle to calibrated confounding (observational equivalence). This disclaimer is attached to every battery claim |
| Thresholds are a proposal, untuned | Gate and battery thresholds were pre-registered and then only evaluated — porting them to another domain without recalibration is unfounded |
| Small-sample traps | Gate power goes missing in small samples, and the harmonic arm false-alarms in small samples (0.275) — a trap in both directions |
| Over-caution | The battery fires even on harmless injections, producing A/B deferrals — the price of safety is paid in deferral |
| Out of scope | The confounding-correction mainline (proximal methods and kin), OPL, and slate/RL OPE are not covered — this repo deliberately stops at the boundary |
Lessons
“In a world without ground truth there are only three honest answers — a necessary-condition screen for the failures you can catch, a sensitivity interval for the failures you cannot, and the courage to output I don’t know (fall back to A/B).”