M2W— Model-to-World LabBlanchet Research Group · Stanford MS&E
Menu

Decision Lab 03 · Robust planning under preference shift

Pay a little now. Protect the decision later.

Freeze a synthetic nominal choice model, declare how far future preferences may move, and choose the assortment with the strongest protected floor—before the stress test is revealed.

Known synthetic MNLGlobal KL prior ballExact small-instance search

M2W · Model the world → simulate uncertainty → decide before reveal → test toward transfer

  1. 01Freezenominal model input
  2. 02Declareglobal KL shift budget
  3. 03Protectmaximize the worst-case floor
  4. 04Stressreveal one shared future
  5. 05Judgeprice, floor, and shifted benefit

Choose the protection. Then reveal the future.

The robust planner uses the frozen baseline and your KL budget; it does not use the stress endpoint or slider. Both policies are then judged in the same shifted MNL world.

1 · Declare the global preference-shift budget

Larger budgets protect against a wider set of globally coherent MNL preference priors. The selected value is an assumption, not estimated here; 0.25 stays below this model's 0.270 feasibility limit.

2 · Reveal one shared preference shift

The policy is already frozen. This control changes evaluation only.

Global KL distance0.078 / 0.10
Inside declared setKL(qα || p₀)

Choose blind. Evaluate after reveal.

Nominal and robust decisions are frozen without using α or the stress endpoint. The shifted oracle is displayed only to measure the opportunity left after the reveal.

PolicyAssortmentNominalProtected floorRevealed stressRole
Nominal leaderPro + Everyday$44.24$27.27$35.31Optimizes the frozen nominal baseline
Robust leaderPlus + Everyday$42.60$27.72$44.87Optimizes the declared floor
Shifted oraclePlus + Everyday$44.87Evaluator only · sees the stress

Nominal price of robustness$1.643.7% of nominal value

Protected-floor gain+$0.45against the nominal policy

Shifted benefit at α = 1.00+$9.56synthetic comparison

Follow both frozen policies.

α changes demand, not the decision
Nominal policyRobust policy

The stress endpoint softens Pro demand while strengthening Plus and Entry. It is curated to interrogate the policy—not fitted from customer data.

How the adversary moves choice.

OutcomeNominalWorst caseMovement
No purchase29.2%52.5% 23.3%
Plus12.3%5.4% 6.9%
Everyday58.5%42.1% 16.4%

Effective conditional radius for this shelf: 0.125. The adverse mix exhausts the allowed KL budget unless zero revenue is already reachable.

Watch the decision switch.

Each card re-optimizes before any stress is revealed
ρ₀ = 0.00Pro + Everyday

Protected floor$44.24

Nominal value$44.24

ρ₀ = 0.04Pro + Everyday

Protected floor$33.41

Nominal value$44.24

ρ₀ = 0.05Plus + Everyday

Protected floor$32.16

Nominal value$42.60

ρ₀ = 0.10Plus + Everyday

Protected floor$27.72

Nominal value$42.60

ρ₀ = 0.20Plus + Everyday

Protected floor$21.39

Nominal value$42.60

ρ₀ = 0.25Plus + Everyday

Protected floor$18.84

Nominal value$42.60

Nothing hidden in the stress path.

Outside attraction stays at 1. Capacity stays at 2. At intermediate α, each preference weight moves linearly between its two disclosed values.

Pro$120.00

Nominal v0.300

Stress-end v0.045

falls
Plus$85.00

Nominal v0.420

Stress-end v0.714

rises
Everyday$55.00

Nominal v2.000

Stress-end v1.600

falls
Entry$20.00

Nominal v0.500

Stress-end v0.650

rises

Every feasible nonempty shelf.

Ranked by protected value at ρ₀ = 0.10
RankAssortmentNominalProtectedStress
01Plus + Everyday$42.60$27.72$44.87
02Pro + Everyday$44.24$27.27$35.31
03Everyday$36.67$22.23$33.85
04Everyday + Entry$34.29$21.87$31.08
05Pro + Plus$41.69$8.88$37.57
06Plus + Entry$23.80$5.11$31.17
07Pro + Entry$25.56$3.32$10.86
08Entry$6.67$0.46$7.88
09Plus$25.14$0.26$35.41
10Pro$27.69$0.00$5.17

Enumeration is exact for 4 products and capacity 2; it is not a claim of scalable optimization.

One uncertain prior. Coherent choices across every shelf.

The experiment implements the known-model planning objective in Example 2.2 of the robust assortment work by Miao Lu, Yuxuan Han, Han Zhong, Zhengyuan Zhou, and José Blanchet. It perturbs one global preference prior inside a Kullback–Leibler ball, then conditions that prior on each offered shelf.

The product names and prices provide continuity with Decision Lab 02. This lab specifies a new synthetic preference prior so the robustness mechanism can be examined in isolation. Connecting D2's learned model directly to the robust planner is the next end-to-end integration step. The global construction preserves multinomial-logit coherence across assortments; exhaustive search is exact here only because there are four products.

Read the robustness paper

The nouns change. The decision architecture survives.

A later autonomy prototype can apply the same freeze–protect–stress pattern to warehouse task and charging decisions under operating-regime drift.

RetailAssortmentAutonomy: task or route set
RetailPreference priorAutonomy: operating regime
RetailRevenueAutonomy: mission utility
RetailDemand shiftAutonomy: model or environment drift

To establish this transfer, the next prototype must specify the warehouse model, validate the adapter, and evaluate operating-regime shifts.

What the robustness prototype establishes—and what comes next.

Established here

A global preference-shift budget can change the selected assortment; its nominal cost and protected floor are computed exactly for this tiny model; a separate in-set stress path tests the frozen policies.

Next validation

Learn the nominal model from logged choices, calibrate the shift budget, extend the optimizer beyond this small MNL instance, and validate the resulting policy with held-out and experimental evidence.

Test whether an observed improvement survives sequential experimentation.