diff --git a/causal-marketing-pymc/apps/labs_claims.yaml b/causal-marketing-pymc/apps/labs_claims.yaml index 4289e08..95bf41f 100644 --- a/causal-marketing-pymc/apps/labs_claims.yaml +++ b/causal-marketing-pymc/apps/labs_claims.yaml @@ -25,7 +25,7 @@ - id: colgate_method slides: ["Colgate: incremental or cannibalistic"] - text: "multivariate Bayesian interrupted time series" + text: "interrupted time series" source: labs.colgate_counterfactual_quote - id: colgate_recovery diff --git a/causal-marketing-pymc/apps/unified_slides.html b/causal-marketing-pymc/apps/unified_slides.html index 11037dd..4237e09 100644 --- a/causal-marketing-pymc/apps/unified_slides.html +++ b/causal-marketing-pymc/apps/unified_slides.html @@ -236,14 +236,15 @@

Causal Inference in the Wild

+ PyMC Labs
Opening · Who is talking

PyMC Labs: what we do

A Bayesian modeling consultancy: custom decision-making models where off-the-shelf tools fall short, and the open-source libraries the field runs on (PyMC, PyMC-Marketing, CausalPy).
@@ -253,8 +254,11 @@

PyMC Labs: what we do

EngagementWhat it isTypical client
Enablement & trainingWorkshops and upskilling on Bayesian methods and the toolingAnalytics teams standardizing on PyMC
-
The clients in this session
- Colgate-Palmolive, HelloFresh, Nürnberger Versicherung (and a Bolt cameo): you meet three of them in the next ten minutes.
+
+
+
The clients in this session
+ Colgate-Palmolive, HelloFresh, Nürnberger Versicherung (and a Bolt cameo): you meet three of them in the next ten minutes.
+
Areas of interest, and where we work
  • Areas: marketing-mix modeling and media measurement, causal inference, demand forecasting and pricing, experimentation and A/B testing at scale, applied Bayesian modeling.
  • @@ -328,7 +332,7 @@

    Bayesian vs frequentist, in one slide

    Why marketing analytics is going Bayesian
    Marketing data is short, noisy, and highly correlated, and every model ends in a spend decision, which is exactly where priors and full uncertainty pay off.
    For today, this is a curiosity
    - The entire lecture runs on classical methods, and every number you will see is frequentist; we surface the Bayesian read only as a flavour at the edges, never as the load-bearing tool.
    + Our lectures will center around classical methods to stay didactically coherent. We will only hint at what the Bayesian approach affords.
Where it is already the default (and where it is not)
  • Marketing mix modeling: the modern open-source MMM tools are Bayesian (Google's Meridian, PyMC-Marketing); with only two or three years of weekly data and correlated channels, plain regression is unstable, and priors on adstock, saturation, and ROI stabilize it.
  • @@ -378,7 +382,7 @@

    You are the consultant

    Case 1 · Colgate-Palmolive

    Colgate-Palmolive: incremental, or cannibalistic?

    -
    Incremental: sales won from competitors or category growth. Cannibalistic: sales taken from your own products. The launch verdict is the split.
    +
    Incremental: sales won from competitors or category growth. Cannibalistic: sales taken from your own products. The question for the launch: how much of its sales is which?
    @@ -392,8 +396,8 @@

    Colgate-Palmolive: incremental, or cannibalistic?

    "We need to estimate the counterfactual sales of all products would have been if the new product had not been introduced."
    • The client: Colgate-Palmolive, 2023; a market estimated at $20.8 billion in 2023.
    • -
    • The method: a multivariate Bayesian interrupted time series: project the pre-launch world forward.
    • -
    • The grading: a planted 50% recovered as a 94% interval of 49-59%: recover a known truth first, then be believed.
    • +
    • The method: a multivariate interrupted time series: project the pre-launch world forward.
    • +
    • The grading: the model is first run on simulated sales where the true incrementality is fixed at 50% by construction; it estimates 49-59% (a 94% interval): it must find an answer we planted before we trust it on the real one.
    @@ -448,7 +452,7 @@

    Why calibrate? A model alone can rank channels backwards

    The settinga company spreads its budget across many ad channels (TV, search, social) and needs to know which ones actually pay back.
    -
    The everyday toola model reads years of spend-and-sales history and scores each channel's return on ad spend, cheaply and always-on.
    +
    The everyday tool: a marketing-mix model (MMM)it reads years of spend-and-sales history and scores each channel's return on ad spend, cheaply and always-on.
    The catch, and our jobhistory is not an experiment; a company spends more exactly when demand is already high, so the model can credit the season instead of the channel and rank them backwards.
    What we adda real experiment: nudge one channel's budget by a known amount, measure the sales it truly causes, and anchor the model to that number. This is calibration.
    @@ -489,25 +493,41 @@

    HelloFresh runs the loop, at industrial scale

    +
    +
    Case 3 · Nürnberger Versicherung
    +

    Nürnberger Versicherung: steering by the last click

    +
    A German insurance group, spending across the whole funnel: awareness video and demand generation up top, branded search at the bottom.
    +
    +
    +
    The old ruler: last-touch attributioncredit every sale to the last ad click before it; simple, standard, and blind to everything upstream of that click.
    +
    What broke itunder the GDPR (the EU's General Data Protection Regulation, the 2018 privacy law that restricts user-level tracking), customer journeys appeared artificially shortened: the last click was often the only click the tracker could still see.
    +
    The replacement: a funnel-aware causal MMMa marketing-mix model that encodes the funnel: upper-funnel spend creates demand that surfaces later in lower-funnel channels, so credit flows to the cause, not to the final click.
    +
    The stakesbudget follows the ruler: whatever the measurement under-credits, the spreadsheet de-funds.
    +
    +
    The client, before the fix
    + "We were optimizing for attribution mechanics instead of incremental business impact." (Philip Herp, Nürnberger Versicherung)
    +
    +
    +
    Case 3 · Nürnberger Versicherung
    -

    Price the engagement

    -
    A German insurer, last-touch attribution, and a funnel-aware causal MMM, in production.
    +

    Who does the last click cheat?

    +
    The terms are on the table; the poll asks you to use them.
    -
    +
    ✋ Poll
    -
    Nürnberger Versicherung replaced last-touch attribution steering with a funnel-aware causal MMM. Over June through November of model-guided spend, cost per lead (CPL) moved by how much?
    +
    Under GDPR-shortened tracking, which channels does last-touch attribution systematically under-credit?
    - - - - + + + +
    -
    C. "This year we were able to drive the CPL down by more than 27%, which is very, very good" (Philip Herp, Nürnberger Versicherung). Under GDPR, customer journeys appeared artificially shortened, so last-touch under-credited the upper funnel and budget followed attribution mechanics instead of incremental business impact. The client is scaling into full production for 2026.
    +
    B. Under GDPR, customer journeys appeared artificially shortened, so the last click was often the only visible step, the upper funnel vanished from the books, and budget followed attribution mechanics instead of incremental business impact. The funnel-aware model re-priced it: "This year we were able to drive the CPL down by more than 27%, which is very, very good" (Philip Herp), over June through November of model-guided spend, and cost per lead (CPL) is the price of one new customer lead. The client is scaling into full production for 2026.
    The client's bar for belief
    "Trust is not created by R² values. It is created when business reality matches model expectations."
    @@ -576,6 +596,11 @@

    The boardroom question

+
+
One metro got the campaign; twenty-nine watched schematic
+
+
+
The method this data calls for
  • One treated unit, no A/B test possible: synthetic control's home ground.
  • @@ -586,18 +611,8 @@

    The boardroom question

  • δ = the share of the pilot's per-euro lift that survives national rollout.
  • The €4M call reduces to one question: is δ large enough to still clear the cost?
-
What "unbiased" means here, precisely, and where it would break
-

The estimator subtracts a reconstructed counterfactual from the treated metro's post-launch sales, and bias is its expected gap from the true lift:

-
- \[ \hat\tau \;=\; \bar Y_{1,\text{post}} - \hat Y_{1,\text{post}}(0), \qquad \text{bias} \;\equiv\; \mathbb{E}[\hat\tau] - \tau \]
-
    -
  • The counterfactual is the synthetic twin, built to reproduce the metro's pre-launch record (its level and its factor loadings) week by week.
  • -
  • Unbiasedness rests on one assumption: that pre-launch match would have continued through the post window had the campaign never run.
  • -
  • This needs no story about why the metro was chosen. Whatever pre-launch trait drove the pick, the twin already carries it, because it was fitted to match it: selection on anything visible before launch cannot tilt the estimate.
  • -
  • Bias enters only through what the pre-launch record cannot see: if the launch were timed on private knowledge of a coming local boom, the twin could not anticipate it and the gap would credit the boom to the campaign. The later placebo-in-time test hunts exactly that.
  • -
-
+
@@ -627,7 +642,7 @@

The data you actually get

- +
@@ -710,7 +725,6 @@

Causal inference is a missing-data problem

  • Subscript says which market, argument says which world (1 = campaign, 0 = none): distinct axes that happen to share the digit 1.
  • \(w_j\): the donor weights the synthetic control will choose so \(\sum_{j\ge 2} w_j Y_{jt}\) rebuilds the missing \(Y_{1t}(0)\).
  • -

    "What would this metro have sold anyway?" is an estimation target, not a rhetorical question.

    @@ -761,7 +775,7 @@

    Simulate the world yourself

    Take a look at the counterfactual.
    -
    ▶ LIVE The equations of the previous slide. Grey: donor markets. White: the treated metro.
    +
    ▶ LIVE The equations of the previous slide. Grey: donor markets. Black: the treated metro.
    @@ -769,7 +783,7 @@

    Simulate the world yourself

    - +
    @@ -825,36 +839,6 @@

    Abadie's idea: if no twin exists, build one

    -
    -
    Act II · The counterfactual problem
    -

    The estimator, precisely

    -
    A constrained least square optimization.
    -
    -

    Fit the weights on the pre-period only, constrained to the simplex:

    -
    - \[ \hat w \;=\; \arg\min_{w\in\Delta}\; \sum_{t \lt T_0}\Bigl(Y_{1t}-\sum_j w_j Y_{jt}\Bigr)^{2}, \qquad \Delta=\Bigl\{w : w_j\ge 0,\ \textstyle\sum_j w_j=1\Bigr\} \] -
    -
      -
    • Pre-launch only: the optimiser never sees the post-period, so it cannot cheat.
    • -
    • The simplex \(\Delta\): a readable recipe, never "−80% of Milan + 190% of Rome": Colgate's projection with the weights on the table.
    • -
    • No standard error: a point estimate only; inference comes later, from placebos.
    • -
    -
    Reading the effect off the gap
    -
    - \[ \hat\tau_t \;=\; Y_{1t}-\sum_j \hat w_j Y_{jt}, \qquad\qquad \hat\tau \;=\; \sum_{t\ge T_0}\hat\tau_t \] - weekly gap between the treated metro and its synthetic twin, summed over the 20 post-launch weeks -
    -
    -
    Why constrain at all?
    - The simplex buys three things ordinary regression cannot: interpretability, regularisation, and an off-switch.
    -
      -
    • Interpretability: "40% dma_08 + 32% dma_20 + 17% dma_03" is a sentence a planner can act on.
    • -
    • Regularisation: most weights land on exactly zero.
    • -
    • An off-switch: a market no blend can match fails loudly instead of extrapolating.
    • -
    -
    -
    -
    Act II · The counterfactual problem
    @@ -891,65 +875,18 @@

    What the simplex buys, geometrically: stay inside the hull

    Act II · The counterfactual problem
    -

    What the constraint buys: drop it and see

    -
    What if we used a simple unconstrained OLS?
    +

    The synthetic twin, graded against the truth

    +
    Only a simulation can draw the true counterfactual; hold the twin against it.
    -
    ▶ LIVE Two synthetics against the truth only a simulation can draw
    +
    ▶ LIVE The synthetic against the truth only a simulation can draw
    -
    OLS hugs the pre-period tighter ( vs ) yet drifts further from the true \(Y_{1t}(0)\) after launch ( vs ): overfitting the simplex refuses.
    -
    -
    -
    -
    The weights OLS chose
    -
    -
    OLS puts negative weight on donors (down to ) and gross weight , where the simplex uses exactly 1.
    -
    -
    -
    - \[ n_{\mathrm{eff}} \;=\; \frac{1}{\sum_j \hat w_j^{2}} \] - Effective number of donors (inverse Herfindahl): one donor \(\Rightarrow n_{\mathrm{eff}}=1\), all 29 equally \(\Rightarrow 29\). Ours: 3.3. -
    -
    What n_eff = 3.3 tells you
    -
      -
    • The synthetic leans on about three donor markets, not a fuzzy mix of all 29.
    • -
    • Sparse weights: interpretable, and less room to overfit.
    • -
    • Near 1: hostage to one market. Near 29: mush. 3.3: concentrated, not fragile.
    • -
    -
    +
    Pre-launch fit: RMSE . Post-launch gap to the true \(Y_{1t}(0)\): . The twin tracks a line it was never shown.
    -
    -
    Act II · The counterfactual problem, versus machine learning
    -

    "Why not gradient boosting / Prophet / an LSTM?"

    -
    Is this a forecasting exercise?
    -
    -
    -
    ▶ LIVE A kitchen-sink forecaster (donors + trend + seasonality, fit pre-launch only)
    -
    -
    Solid: treated sales. Dashed blue: the forecaster. Dashed grey: the simplex twin. Both fit only the 40 pre-launch weeks.
    -
    -
      -
    • In-sample it fits tighter:k vs €k per week: flexibility always buys the past.
    • -
    • Yet the estimates are a wash: vs , both near the planted €284k. One dataset cannot rank them.
    • -
    • The sweep decides: across 24 fresh worlds, 1.6× worse out of sample, with nothing to inspect.
    • -
    -
    -
    -
    Principle 1: a counterfactual, not a forecast
    - A forecaster asks what comes next. We ask what this market would have done in a world that never happened.
    -
      -
    • Principle 2: in-sample fit is the trap, not the goal.
    • -
    • Principle 3: you cannot cross-validate the counterfactual. The truth is never observed.
    • -
    • Principle 4: the simplex ships an auditable claim; a boosted tree is a black box.
    • -
    -
    The forecaster is better at prediction. The simplex is better at the causal job.
    -
    -
    -
    Act II · The counterfactual problem
    @@ -1008,90 +945,7 @@

    Compare the estimators

    -
    -
    Act II · The counterfactual problem
    -

    Interesting limits for the simpler estimators

    -
    Where before/after and treated/control estimators break.
    -
    -
    -
    -
    Before/after · one unit, across time
    - The treated metro's post-average minus its pre-average.
    -
      -
    • The level cancels; nothing subtracts the shared wave.
    • -
    • The whole drift \(\Delta\bar f\) lands on the metro: the largest bias, growing with the horizon.
    • -
    -
    -
    -
    Treated-vs-control · one period, across units
    - The treated metro minus the control average, both in the post window.
    -
      -
    • The level gap \(\alpha_1-\bar\alpha_C\) survives, and sizes differ a lot: it dominates.
    • -
    • This is why raw sales are never compared across cities.
    • -
    -
    -
    -
    The two bias decompositions, term by term
    -
    - \[ \hat\tau^{\text{B/A}} \;=\; \bar Y_{1,\text{post}}-\bar Y_{1,\text{pre}} \;=\; \bar\tau \;+\; \underbrace{\gamma_1^{\top}\,\Delta\bar f}_{\text{bias: the whole drift}} \;+\; \Delta\bar\varepsilon_1 \] - the level \(\alpha_1\) cancels, but the whole shared-factor drift \(\Delta\bar f\) lands at the metro's own loadings: the largest bias, and it grows with the horizon. -
    -
    - \[ \hat\tau^{\text{TC}} \;=\; \bar Y_{1,\text{post}}-\bar Y_{C,\text{post}} \;=\; \bar\tau \;+\; \underbrace{(\alpha_1-\bar\alpha_C)}_{\text{level gap}} \;+\; \underbrace{(\gamma_1-\bar\gamma_C)^{\top}\bar f_{\text{post}}}_{\text{loading gap}\,\times\,\text{level}} \;+\; \Delta\bar\varepsilon \] - no time-differencing, so the level gap survives and dominates; the loading gap now multiplies the factor level, not its drift. -
    -
    -
    -
    - -
    -
    Act II · The counterfactual problem
    -

    Interesting limits for DiD

    -
    Where the DiD estimator breaks.
    -
    -
    The one question
    - DiD only works if treated and controls would have drifted together. Two dials decide (spread \(s\), macro \(\sigma_\eta\)); the algebra is in the fold.
    -
    Synthetic control solves it
    - DiD's residual bias \(B=(\gamma_1-\bar\gamma_C)^{\top}\Delta\bar f\) comes from its equal weights: it hopes \(s\) is small. Synthetic control chooses the weights, so \(B\to0\) at any spread.
    -
    The DiD bias, decomposed (and why more data cannot shrink it)
    -
    - \[ \hat\tau^{\text{DiD}} \;=\; \bar\tau \;+\; \underbrace{(\gamma_1-\bar\gamma_C)^{\top}\,\Delta\bar f}_{\text{systematic bias }B} \;+\; \underbrace{\Delta\bar\varepsilon_1-\Delta\bar\varepsilon_C}_{\text{idiosyncratic noise}} \] -
    -
      -
    • The pieces: \(\Delta\bar f\) is the post-minus-pre drift of the three factors, \(\bar\gamma_C\) the \(n\) controls' average loadings, \(\Delta\bar\varepsilon\) the change in idiosyncratic noise.
    • -
    • \(B\) is mean-zero over the exposure draw, but your world is one draw, so in it \(B\) is a fixed nonzero number whose typical size is:
    • -
    -
    - \[ \operatorname{Var}(B) \;=\; \frac{s^{2}}{3}\Bigl(1+\frac{1}{n}\Bigr)\Bigl[(\Delta\bar f_{\text{trend}})^{2}+(\Delta\bar f_{\text{season}})^{2}+(\Delta\bar f_{\text{walk}})^{2}\Bigr] \] -
    -
      -
    • \(B\) carries no term for the amount of data: \(\Delta\bar f\) is the world's drift, not sampling error, so more weeks shrink only noise and buy precision around the same biased number.
    • -
    • The noise never leaves: the idiosyncratic term carries \(\sigma_\varepsilon \approx 3\) whatever \(s\) and \(\sigma_\eta\) do, so the honest claim in any limit is "the systematic bias vanishes", never "DiD equals the truth".
    • -
    -
    -
    The two limits, dial by dial: send a knob to zero and read what dies
    -
    -
    Limit 1 · loading spread \(s \to 0\): sufficient on its own
    -
      -
    • The prefactor \(s^2/3\) kills the whole bracket at once.
    • -
    • Every loading \(\to 1\), so \(\gamma_1-\bar\gamma_C\to 0\) deterministically: exact parallel trends, macro walk included.
    • -
    • DiD then recovers \(\bar\tau\) up to noise.
    • -
    -
    Limit 2 · macro shock \(\sigma_\eta \to 0\): not sufficient
    -
      -
    • Only the walk's term \((\Delta\bar f_{\text{walk}})^2\) drops out.
    • -
    • Trend and seasonality still drift between the windows, and markets still weight them differently whenever \(s>0\): \(B \neq 0\).
    • -
    • What it buys: it removes the one stochastic, horizon-growing confounder; the bias that remains is at least deterministic.
    • -
    -
    -
    -
    Where the \(\tfrac{s^2}{3}\bigl(1+\tfrac1n\bigr)\) prefactor comes from
    - Each loading is drawn from \(\mathrm{U}(1{-}s,\,1{+}s)\), a uniform of width \(2s\), whose variance is \((2s)^2/12 = s^2/3\). The treated metro contributes \(\operatorname{Var}(\gamma_1)=s^2/3\); the average of \(n\) independent controls contributes \(\operatorname{Var}(\bar\gamma_C)=s^2/(3n)\). Independent draws add, so \(\operatorname{Var}(\gamma_1-\bar\gamma_C)=\tfrac{s^2}{3}\bigl(1+\tfrac1n\bigr)\) per factor. Each factor's mismatch is then scaled by that factor's post-minus-pre drift, and the three squared drifts sum: the bracket. -
    -
    -
    -
    Act II · The counterfactual problem
    @@ -1125,103 +979,8 @@

    What must be true for the gap to be causal

    -
    -
    Act III · Is it real, and how big?
    -

    €260k. Real, or a lucky metro? Build the null yourself

    -
    -
    -
    ✋ Poll
    -
    With one treated metro a t-test is not weak, it is undefined. So build the test yourself: rerun the whole pipeline 29 more times, each donor pretended treated ("placebo"): 30 "effects", 29 where nothing ran. Where does our €260k rank?
    -
    - - - -
    - -
    A: rank 1 of 30. The rank is the inference: if the campaign did nothing, P(rank 1 by luck) = 1/30 ≈ 0.033. You just re-derived the permutation test.
    -
    -
    -
    - -
    -
    Act III · Is it real, and how big?
    -

    Placebo-in-space: measure the luck directly

    -
    Randomisation inference: the t-test had no standard error, so the 29 donors become the null distribution.
    -
    -
    -
    -
    -
    ▶ LIVE Refit the estimator on every donor as if it were treated
    -
    -
    The null: the "effects" the method reports where nothing happened (spread ·, best placebo ·). Green line: our €260k, outside the cloud.
    -
    -
    -
    -
    The null hypothesis, stated
    - \(H_0\): the campaign did nothing. Then the metro is exchangeable with its donors: any of the 30 ranks is equally likely.
    -
    - \[ p \;=\; \frac{1+\#\{\,j:\ |\hat\tau_j|\ge|\hat\tau_1|\,\}}{J+1} \;=\; \frac{1+0}{30}\;\approx\;0.033 \] -
    -
    -
    -
    What the p-value means (and what it is not)
    - The probability, if \(H_0\) were true, of a gap this extreme: a rank, no bell curve assumed. \(p\) sits at its floor \(1/30\), set by the donor count, not the weeks.
    -
    When the rank is valid, and when it lies (Abadie's hygiene rule)
    -
      -
    • Valid under \(H_0\) when the placebos are fair stand-ins: comparable pre-launch fit, and no campaign spillover onto donors (SUTVA), the two ways exchangeability can hold.
    • -
    • Lies when a placebo fits its own pre-period badly: it books fitting failure as a giant fake effect and fattens the tail. Abadie's rule: drop those placebos before ranking.
    • -
    -
    Conclusion: the €260k lift is real, at \(p\approx0.033\)
    - Rank 1 of 30, clear of the cloud: reject \(H_0\). The rank holds (p = 0.033 every time) dropping shaky placebos at 2×, 5×, 20× pre-fit error. But "real" is not "profitable": Act IV.
    -
    -
    - -
    -
    Act III · Is it real, and how big?
    -

    From a test to an interval: inversion

    -
    Which true lifts could plausibly have produced our €260k? Keep the survivors: no bell curve anywhere.
    -
    -
    -
    -
    -
    ▶ LIVE Drag \(H\), a guess at the truth: the error cloud slides with it; our €260k never moves.
    -
    -
    - - -
    -
    Axis: 20-week total gap (€000). Top: the placebo errors at truth 0. Middle: the same errors slid to \(H\); \(H\) survives if the green line sits inside the shaded middle 90%. Bottom: the survivors, collected: the interval. Drag past an edge to reject.
    -
    -
    -
    -
    The whole idea
    - Ask of every possible true lift: could it plausibly have produced our €260k? Collect the ones that could, and that set of survivors is the interval.
    -
      -
    • ① Measure the error. The 29 placebo "effects" form a cloud around zero, spread about ±€50k.
    • -
    • ② Guess a truth \(H\). If the lift were \(H\), we would see \(H\) plus that same cloud.
    • -
    • ③ Keep or reject. Keep \(H\) if €260k sits inside its middle 90%.
    • -
    • ④ Sweep. The survivors run €195k to €335k: the interval is [€195k, €335k].
    • -
    -
    -
    -
    The inversion, in one line of algebra
    -
    - \[ \underbrace{H + q_{0.05} \;\le\; 260 \;\le\; H + q_{0.95}}_{\text{step ③: €260k sits in \(H\)'s middle \(90\%\)}} - \qquad\Longleftrightarrow\qquad - \underbrace{260 - q_{0.95} \;\le\; H \;\le\; 260 - q_{0.05}}_{\text{step ④: the same line, solved for \(H\)}} \] - \(q_{0.05},q_{0.95}\) are just the low and high edges of the error cloud from step ①, here \(q_{0.05}\!=\!-75\) and \(q_{0.95}\!=\!+65\) (€000). Rearranging the left inequality into the right one is the whole trick, and it hands you the endpoints €195k and €335k directly. -
    -
    -
    Why this interval is the referee for the rest of the lecture
    - It assumed no normality, no independence, no error model, so every model-based interval later (the Bayesian posterior included) has to answer to it. -
    -
    -
    -
    Act III · Is it real, and how big? Stress-test the estimate
    @@ -1274,26 +1033,6 @@

    Falsification 2: the assumption no placebo can see

    -
    -
    Act III · Is it real, and how big?
    -

    Statistics done. Three numbers.

    -
    Questions ① and ② from the boardroom slide are now answered.
    -
    - - - - - - - -
    QuestionAnswerTool that answered it
    Is the effect real?Yes, p = 0.033placebo-in-space permutation: rank 1 of 30
    How big?€260k of incremental salessynthetic-control gap, summed over 20 weeks
    Give or take?[€195k, €335k] at 90%test inversion over the placebo cloud
    -
    Truth check (only a simulation allows it)
    - The planted total €284k sits inside the interval, €24k above the estimate: the machinery works, and its self-reported uncertainty is honest.
    -
    The sentence that loses money
    - "€260k of sales for €75k, a 3.5× return. Roll it out." Every number true; the conclusion does not follow. Question ③ is not a statistics question.
    -
    -
    -
    Act IV · The decision in euros
    @@ -1837,7 +1576,7 @@

    The ideal experiment

    • Where it hides: a lottery, a rollout order, an arbitrary rule that moved exposure for reasons unrelated to intent.
    -
    The plan for Part 2
    +
    Going forward
    1. Give that random lever a name: an instrument.
    2. State the conditions it must satisfy.
    3. Check which of them the data can verify.
    @@ -1995,13 +1734,489 @@

    The reduced form

    -
    -
    IV · The estimator · the whole method in one division
    -

    The IV estimate

    -
    Euros per lottery win, divided by exposures per lottery win.
    -
    -
    -
    the whole method, as arithmetic on two measured numbers
    + +
    +
    IV · When it breaks · whose effect it is
    +

    Compliers and the LATE

    +
    Whom does the €16.5 describe? Only the users the lottery could move.
    +
    +
    ▶ LIVE the user base, split by how they respond to the lottery
    +
    +
    + + +
    +
    Push \(\gamma\): the complier slice grows, because the complier share is the first stage. Shares are drawn live from one simulated batch, so they can differ from the printed figures by a rounding step.
    +
    +
    +
      +
    • Always-takers (56.5%): the auction shows them the ad with or without the lottery. It changes nothing for them, so the data say nothing about them.
    • +
    • Never-takers (22.5%): never see the ad either way. Same silence.
    • +
    • Compliers (21.0%): see the ad only because the lottery favoured them. Every euro of the lottery's lift \(\delta\) came from them.
    • +
    • The caveat, monotonicity: we assume no defiers, users who would see the ad only when the lottery does not favour them. A nudge that never repels.
    • +
    +
    +
    +
    Definition · LATE (local average treatment effect)
    + The LATE is the average effect of the ad on the compliers alone, and it is what the division estimates: a local answer, not a statement about every user.
    +
    \[ \hat\beta_{\text{IV}} \;=\; \frac{\delta}{\pi} \;\;\text{ estimates }\;\; \mathbb{E}[\,Y(1)-Y(0)\mid \text{complier}\,] \] + the average of each complier's personal effect, the same \(Y(1)-Y(0)\) contrast the naive slide could not touch
    +
    +
    +
    Why a manager should love this fine print
    + Compliers are the same kind of marginal user a higher bid would newly reach: the closest thing in the data to the customer a bid change buys. That makes €16.5 a price for the marginal customer, measured on the margin rather than on the average.
    +
    +
    + +
    +
    IV · When it breaks · what must hold, on one page
    +

    The checklist

    +
    The four assumptions, and which ones the data can check.
    +
    + + + + + + + + +
    AssumptionWhat it saysIn the caseStatus
    Relevance\(Z\) moves \(X\)\(F = 156\), far above 10TESTABLE, passes
    Exogeneity\(Z \perp U\)the lottery is a genuine random drawBY DESIGN
    Exclusion\(Z \to Y\) only via \(X\)a queue bump shows the user nothingUNTESTABLE
    Monotonicityno defiersa nudge never repelsUNTESTABLE, plausible
    +
    +
    +
    The honest scorecard, for any IV study you are shown
    + One measured number (the first-stage \(F\)), one design guarantee (the randomization), two arguments (exclusion, monotonicity). Ask for all four before you accept the estimate.
    +
      +
    • What we can now defend: an exposure causes about €16.5 of sales for the users a bid can actually move.
    • +
    +
    +
    The sentence that loses money
    + "Exposed users are worth €23.7 each, so raise the bid": every word true, conclusion wrong. It books the platform's targeting as advertising.
    +
    +
    +
    + + + +
    +
    IV · The decision · euros at last
    +

    The price map

    +
    The estimate becomes a decision only when it meets the price.
    +
    +
    +
    ▶ LIVE the verdict as the price moves
    +
    +
    + + +
    +
    Blue: net value per exposure at each price. The orange band is the 90% interval, the zone where the data refuse to commit. Drag the price through the three zones and watch the verdict flip.
    +
    +
    +

    One estimate gives not one answer but a map from any price to a verdict:

    + + + + + + + +
    Price zoneVerdictWhy
    below €12.7GOeven the most pessimistic supported effect pays
    €12.7 to €20.4TESTthe data straddle the price: negotiate, or measure more
    above €20.4NO-GOno supported effect pays
    +
      +
    • Today's rate, €10, sits in the GO zone, below the whole interval.
    • +
    • The net, computed: \(Y\) is contribution euros, so one exposure nets \(\hat\beta_{\text{IV}} - c = \) €16.5 − €10 = €6.5. Even read at the interval floor, €12.7 against €10, the exposure still pays.
    • +
    +
    Why boards like this framing
    + "Is the effect significant?" has no business answer. "Up to what price is this a buy?" has one, and it is the same question a bid cap asks.
    +
    +
    +
    + +
    +
    IV · The decision
    +

    Poll · the negotiation

    +
    +
    +
    ✋ Poll
    +
    The platform wants to renegotiate the rate. Your analyst hands you the causal read: effect €16.5 per exposure, 90% interval [12.7, 20.4]. What is the highest rate at which you would still sign "buy" without further study?
    +
    + + + + +
    + +
    C. Below €12.7, every effect the data support pays: the interval's floor is a no-regret bid cap, defensible whichever value inside the interval turns out to be the truth. B is a break-even gamble: paying the point estimate wins or loses depending on which side of it the truth sits, acceptable only for a risk-neutral buyer averaging over many campaigns. A pays a price that only the single most optimistic supported effect can justify. D leaves money on the table: the whole interval sits well above today's rate.
    +
    +
    +
    + +
    +
    IV · The decision · banked
    +

    The verdict and the recommendation

    +
    The complete answer, assembled from everything measured so far.
    +
    +
    + + + + + + + + + +
    QuantityValueSource
    effect of one exposure€16.5the division δ/π
    90% interval[12.7, 20.4]classical, and AR agrees
    first-stage F156the lottery is strong
    price€10the platform's rate card
    net per exposure€6.5β − c, at the point estimate
    +
    The verdict
    + BUY  Keep buying at €10: the entire defensible range clears the price.
    +
    +
    +
    The recommendation, in three lines
    + 1. Keep buying at the €10 rate: even the interval's most pessimistic effect pays.
    + 2. Cap the bid at the interval's lower end, €12.7: up to there, every effect the data support still clears the price.
    + 3. Measure again only if the rate card climbs toward €12.7: at today's price, no effect inside the interval changes the action, so more measurement is almost certain to leave the decision unchanged and is worth close to nothing here.
    +
      +
    • Everything above is classical: two averages, one division, one F statistic, one confidence interval.
    • +
    • The one debt on record: exclusion is untestable. The recommendation is conditional on the argued design, and says so.
    • +
    +
    +
    +
    + + +
    +
    Closing · Provenance
    +

    The tools were the product too

    +
    +
      +
    • CausalPy: synthetic control, interrupted time series, difference in differences, and regression discontinuity in one open-source package. The IV estimator that closes this session joined later.
    • +
    • Its launch example: individual exposure to a TV campaign cannot be randomised, yet its causal impact remains a core business need: the sentence this whole session opened with.
    • +
    • pymc-marketing: the MMM library behind Case 2's calibration story; one client's budget allocation approach to PyMC-Marketing came back as a pull request (Bolt).
    • +
    • Webinars and content: the consultancy's own webinar agenda is this session's syllabus, Instrumental Variables included.
    • +
    +
    + PyTensor + PyMC + PyMC-Marketing + CausalPy +
    +
    +
    + +
    +
    Closing
    +

    The END

    +
    Causal inference is one shelf of the toolbox: Bayesian inference is a vast set of tools, and PyMC Labs builds with all of it.
    +
      +
    • Beyond today: demand forecasting, pricing, experimentation at scale, customer lifetime value, hierarchical models across markets: the same machinery, aimed at different decisions.
    • +
    • The cases: pymc-labs.com/blog-posts: every number in this session is pinned to a public post (sources on the next slide).
    • +
    +
    +
    +
    Francesco Muia
    +
    PhD in Theoretical Physics, EMBA.
    Consultant for PyMC Labs and Brown University.
    +
    francesco.muia@pymc-labs.com
    +
    francesco.muia@ai-and-analytics-solutions.com
    +
    +
    +
    Alexander Fengler
    +
    PhD in Computational Cognitive Science.
    Postdoc at Brown University, consultant for PyMC Labs.
    +
    alexander.fengler@pymc-labs.com
    +
    +
    Get in touch
    + If a decision in your company leans on a number nobody quite trusts, write to us: those are exactly the problems we like.
    +
    + + +
    +
    Backup
    + Backup · Sources +

    Every number, pinned

    +
    Part 1 facts retrieved and pinned 2026-07-19 (apps/labs_deck_data.json carries the exact quote); Part 2 numbers are baked from the executed course notebooks (nb07/nb07b shards).
    +
    + + + + + + + + + + + + + + + + + +
    SourceFacts pinned
    ailab.criteo.com · criteo-uplift-prediction-datasetcriteo_rows
    pymc-labs.com · 2022-11-11-HelloFreshhf_panel_calibration
    pymc-labs.com · 2023-06-20-juan-marketing-analyticswebinar_agenda
    pymc-labs.com · bayes-is-slow-speeding-up-hellofreshs-bayesian-ab-tests-by-60xhf_batch, hf_test_types, hf_thousands
    pymc-labs.com · bayesian-media-mix-modeling-for-marketing-optimizationhf_priors_experiments
    pymc-labs.com · causal-sales-analytics-are-my-sales-incremental-or-cannibalisticcolgate_ci, colgate_ci_level, colgate_truth, colgate_year, fail_range, fail_truth, market
    pymc-labs.com · causal-sales-analytics-discrete-choice-modelingcolgate_counterfactual_quote
    pymc-labs.com · causalpy-a-new-package-for-bayesian-causal-inference-for-quasi-experimentscausalpy_methods, causalpy_tv
    pymc-labs.com · funnel-aware-mmmcpl, cpl_window, gdpr_sentence, herp_attribution_quote, nurn_2026, trust_quote
    pymc-labs.com · marketing-mix-modeling-a-complete-guidebolt_pr
    pymc-labs.com · mmm_roas_liftlift_tests_n, roas_gap_words, roas_wrong_ranking, roas_x1, roas_x2
    pymc-labs.com · open-sourcing-decision-lab-scaling-ai-judgment-data-sciencedl_explored, dl_vanilla, dl_verdict
    pymc-labs.com · reducing-customer-acquisition-costs-how-we-helped-optimizing-hellofreshs-marketing-budgethf_var
    +
    +
    + +
    +
    Backup · Act II · The counterfactual problem, versus machine learning
    +

    "Why not gradient boosting / Prophet / an LSTM?"

    +
    Is this a forecasting exercise?
    +
    +
    +
    ▶ LIVE A kitchen-sink forecaster (donors + trend + seasonality, fit pre-launch only)
    +
    +
    Solid: treated sales. Dashed blue: the forecaster. Dashed grey: the simplex twin. Both fit only the 40 pre-launch weeks.
    +
    +
      +
    • In-sample it fits tighter:k vs €k per week: flexibility always buys the past.
    • +
    • Yet the estimates are a wash: vs , both near the planted €284k. One dataset cannot rank them.
    • +
    • The sweep decides: across 24 fresh worlds, 1.6× worse out of sample, with nothing to inspect.
    • +
    +
    +
    +
    Principle 1: a counterfactual, not a forecast
    + A forecaster asks what comes next. We ask what this market would have done in a world that never happened.
    +
      +
    • Principle 2: in-sample fit is the trap, not the goal.
    • +
    • Principle 3: you cannot cross-validate the counterfactual. The truth is never observed.
    • +
    • Principle 4: the simplex ships an auditable claim; a boosted tree is a black box.
    • +
    +
    The forecaster is better at prediction. The simplex is better at the causal job.
    +
    +
    +
    + +
    +
    Backup · Act II · The counterfactual problem
    +

    Interesting limits for the simpler estimators

    +
    Where before/after and treated/control estimators break.
    +
    +
    +
    +
    Before/after · one unit, across time
    + The treated metro's post-average minus its pre-average.
    +
      +
    • The level cancels; nothing subtracts the shared wave.
    • +
    • The whole drift \(\Delta\bar f\) lands on the metro: the largest bias, growing with the horizon.
    • +
    +
    +
    +
    Treated-vs-control · one period, across units
    + The treated metro minus the control average, both in the post window.
    +
      +
    • The level gap \(\alpha_1-\bar\alpha_C\) survives, and sizes differ a lot: it dominates.
    • +
    • This is why raw sales are never compared across cities.
    • +
    +
    +
    +
    The two bias decompositions, term by term
    +
    + \[ \hat\tau^{\text{B/A}} \;=\; \bar Y_{1,\text{post}}-\bar Y_{1,\text{pre}} \;=\; \bar\tau \;+\; \underbrace{\gamma_1^{\top}\,\Delta\bar f}_{\text{bias: the whole drift}} \;+\; \Delta\bar\varepsilon_1 \] + the level \(\alpha_1\) cancels, but the whole shared-factor drift \(\Delta\bar f\) lands at the metro's own loadings: the largest bias, and it grows with the horizon. +
    +
    + \[ \hat\tau^{\text{TC}} \;=\; \bar Y_{1,\text{post}}-\bar Y_{C,\text{post}} \;=\; \bar\tau \;+\; \underbrace{(\alpha_1-\bar\alpha_C)}_{\text{level gap}} \;+\; \underbrace{(\gamma_1-\bar\gamma_C)^{\top}\bar f_{\text{post}}}_{\text{loading gap}\,\times\,\text{level}} \;+\; \Delta\bar\varepsilon \] + no time-differencing, so the level gap survives and dominates; the loading gap now multiplies the factor level, not its drift. +
    +
    +
    +
    + +
    +
    Backup · Act II · The counterfactual problem
    +

    Interesting limits for DiD

    +
    Where the DiD estimator breaks.
    +
    +
    The one question
    + DiD only works if treated and controls would have drifted together. Two dials decide (spread \(s\), macro \(\sigma_\eta\)); the algebra is in the fold.
    +
    Synthetic control solves it
    + DiD's residual bias \(B=(\gamma_1-\bar\gamma_C)^{\top}\Delta\bar f\) comes from its equal weights: it hopes \(s\) is small. Synthetic control chooses the weights, so \(B\to0\) at any spread.
    +
    The DiD bias, decomposed (and why more data cannot shrink it)
    +
    + \[ \hat\tau^{\text{DiD}} \;=\; \bar\tau \;+\; \underbrace{(\gamma_1-\bar\gamma_C)^{\top}\,\Delta\bar f}_{\text{systematic bias }B} \;+\; \underbrace{\Delta\bar\varepsilon_1-\Delta\bar\varepsilon_C}_{\text{idiosyncratic noise}} \] +
    +
      +
    • The pieces: \(\Delta\bar f\) is the post-minus-pre drift of the three factors, \(\bar\gamma_C\) the \(n\) controls' average loadings, \(\Delta\bar\varepsilon\) the change in idiosyncratic noise.
    • +
    • \(B\) is mean-zero over the exposure draw, but your world is one draw, so in it \(B\) is a fixed nonzero number whose typical size is:
    • +
    +
    + \[ \operatorname{Var}(B) \;=\; \frac{s^{2}}{3}\Bigl(1+\frac{1}{n}\Bigr)\Bigl[(\Delta\bar f_{\text{trend}})^{2}+(\Delta\bar f_{\text{season}})^{2}+(\Delta\bar f_{\text{walk}})^{2}\Bigr] \] +
    +
      +
    • \(B\) carries no term for the amount of data: \(\Delta\bar f\) is the world's drift, not sampling error, so more weeks shrink only noise and buy precision around the same biased number.
    • +
    • The noise never leaves: the idiosyncratic term carries \(\sigma_\varepsilon \approx 3\) whatever \(s\) and \(\sigma_\eta\) do, so the honest claim in any limit is "the systematic bias vanishes", never "DiD equals the truth".
    • +
    +
    +
    The two limits, dial by dial: send a knob to zero and read what dies
    +
    +
    Limit 1 · loading spread \(s \to 0\): sufficient on its own
    +
      +
    • The prefactor \(s^2/3\) kills the whole bracket at once.
    • +
    • Every loading \(\to 1\), so \(\gamma_1-\bar\gamma_C\to 0\) deterministically: exact parallel trends, macro walk included.
    • +
    • DiD then recovers \(\bar\tau\) up to noise.
    • +
    +
    Limit 2 · macro shock \(\sigma_\eta \to 0\): not sufficient
    +
      +
    • Only the walk's term \((\Delta\bar f_{\text{walk}})^2\) drops out.
    • +
    • Trend and seasonality still drift between the windows, and markets still weight them differently whenever \(s>0\): \(B \neq 0\).
    • +
    • What it buys: it removes the one stochastic, horizon-growing confounder; the bias that remains is at least deterministic.
    • +
    +
    +
    +
    Where the \(\tfrac{s^2}{3}\bigl(1+\tfrac1n\bigr)\) prefactor comes from
    + Each loading is drawn from \(\mathrm{U}(1{-}s,\,1{+}s)\), a uniform of width \(2s\), whose variance is \((2s)^2/12 = s^2/3\). The treated metro contributes \(\operatorname{Var}(\gamma_1)=s^2/3\); the average of \(n\) independent controls contributes \(\operatorname{Var}(\bar\gamma_C)=s^2/(3n)\). Independent draws add, so \(\operatorname{Var}(\gamma_1-\bar\gamma_C)=\tfrac{s^2}{3}\bigl(1+\tfrac1n\bigr)\) per factor. Each factor's mismatch is then scaled by that factor's post-minus-pre drift, and the three squared drifts sum: the bracket. +
    +
    +
    + +
    +
    Backup · Act III · Is it real, and how big?
    +

    €260k. Real, or a lucky metro? Build the null yourself

    +
    +
    +
    ✋ Poll
    +
    With one treated metro a t-test is not weak, it is undefined. So build the test yourself: rerun the whole pipeline 29 more times, each donor pretended treated ("placebo"): 30 "effects", 29 where nothing ran. Where does our €260k rank?
    +
    + + + +
    + +
    A: rank 1 of 30. The rank is the inference: if the campaign did nothing, P(rank 1 by luck) = 1/30 ≈ 0.033. You just re-derived the permutation test.
    +
    +
    +
    + +
    +
    Backup · Act III · Is it real, and how big?
    +

    Placebo-in-space: measure the luck directly

    +
    Randomisation inference: the t-test had no standard error, so the 29 donors become the null distribution.
    +
    +
    +
    +
    +
    ▶ LIVE Refit the estimator on every donor as if it were treated
    +
    +
    The null: the "effects" the method reports where nothing happened (spread ·, best placebo ·). Green line: our €260k, outside the cloud.
    +
    +
    +
    +
    The null hypothesis, stated
    + \(H_0\): the campaign did nothing. Then the metro is exchangeable with its donors: any of the 30 ranks is equally likely.
    +
    + \[ p \;=\; \frac{1+\#\{\,j:\ |\hat\tau_j|\ge|\hat\tau_1|\,\}}{J+1} \;=\; \frac{1+0}{30}\;\approx\;0.033 \] +
    +
    +
    +
    What the p-value means (and what it is not)
    + The probability, if \(H_0\) were true, of a gap this extreme: a rank, no bell curve assumed. \(p\) sits at its floor \(1/30\), set by the donor count, not the weeks.
    +
    When the rank is valid, and when it lies (Abadie's hygiene rule)
    +
      +
    • Valid under \(H_0\) when the placebos are fair stand-ins: comparable pre-launch fit, and no campaign spillover onto donors (SUTVA), the two ways exchangeability can hold.
    • +
    • Lies when a placebo fits its own pre-period badly: it books fitting failure as a giant fake effect and fattens the tail. Abadie's rule: drop those placebos before ranking.
    • +
    +
    Conclusion: the €260k lift is real, at \(p\approx0.033\)
    + Rank 1 of 30, clear of the cloud: reject \(H_0\). The rank holds (p = 0.033 every time) dropping shaky placebos at 2×, 5×, 20× pre-fit error. But "real" is not "profitable": Act IV.
    +
    +
    + +
    +
    Backup · Act III · Is it real, and how big?
    +

    From a test to an interval: inversion

    +
    Which true lifts could plausibly have produced our €260k? Keep the survivors: no bell curve anywhere.
    +
    +
    +
    +
    +
    ▶ LIVE Drag \(H\), a guess at the truth: the error cloud slides with it; our €260k never moves.
    +
    +
    + + +
    +
    Axis: 20-week total gap (€000). Top: the placebo errors at truth 0. Middle: the same errors slid to \(H\); \(H\) survives if the green line sits inside the shaded middle 90%. Bottom: the survivors, collected: the interval. Drag past an edge to reject.
    +
    +
    +
    +
    The whole idea
    + Ask of every possible true lift: could it plausibly have produced our €260k? Collect the ones that could, and that set of survivors is the interval.
    +
      +
    • ① Measure the error. The 29 placebo "effects" form a cloud around zero, spread about ±€50k.
    • +
    • ② Guess a truth \(H\). If the lift were \(H\), we would see \(H\) plus that same cloud.
    • +
    • ③ Keep or reject. Keep \(H\) if €260k sits inside its middle 90%.
    • +
    • ④ Sweep. The survivors run €195k to €335k: the interval is [€195k, €335k].
    • +
    +
    +
    +
    The inversion, in one line of algebra
    +
    + \[ \underbrace{H + q_{0.05} \;\le\; 260 \;\le\; H + q_{0.95}}_{\text{step ③: €260k sits in \(H\)'s middle \(90\%\)}} + \qquad\Longleftrightarrow\qquad + \underbrace{260 - q_{0.95} \;\le\; H \;\le\; 260 - q_{0.05}}_{\text{step ④: the same line, solved for \(H\)}} \] + \(q_{0.05},q_{0.95}\) are just the low and high edges of the error cloud from step ①, here \(q_{0.05}\!=\!-75\) and \(q_{0.95}\!=\!+65\) (€000). Rearranging the left inequality into the right one is the whole trick, and it hands you the endpoints €195k and €335k directly. +
    +
    +
    Why this interval is the referee for the rest of the lecture
    + It assumed no normality, no independence, no error model, so every model-based interval later (the Bayesian posterior included) has to answer to it. +
    +
    +
    + +
    +
    Backup · Act III · Is it real, and how big?
    +

    Statistics done. Three numbers.

    +
    Questions ① and ② from the boardroom slide are now answered.
    +
    + + + + + + + +
    QuestionAnswerTool that answered it
    Is the effect real?Yes, p = 0.033placebo-in-space permutation: rank 1 of 30
    How big?€260k of incremental salessynthetic-control gap, summed over 20 weeks
    Give or take?[€195k, €335k] at 90%test inversion over the placebo cloud
    +
    Truth check (only a simulation allows it)
    + The planted total €284k sits inside the interval, €24k above the estimate: the machinery works, and its self-reported uncertainty is honest.
    +
    The sentence that loses money
    + "€260k of sales for €75k, a 3.5× return. Roll it out." Every number true; the conclusion does not follow. Question ③ is not a statistics question.
    +
    +
    + +
    +
    Backup · IV · The estimator · hands on
    +

    Why the division is forced

    +
    The division is not a modelling choice. It is the only effect size the two measurements allow.
    +
    +
    +
    ▶ LIVE every candidate effect makes a prediction. One matches.
    +
    +
    + + +
    +
    The rising line is the prediction: an effect of \(\hat\beta\) per exposure implies the lottery should have lifted sales by \(\hat\beta \times \pi\). The flat line is the fact: it lifted them by €3.48. Move your guess to the crossing and you have priced the ad.
    +
    +
    +

    Forget the formula and grade any candidate effect \(\hat\beta\) against the two numbers we own:

    +
      +
    • Its prediction: if one exposure were worth \(\hat\beta\), the lottery's 0.2106 extra exposures per win should create \(\hat\beta \times 0.2106\) euros per win.
    • +
    • The fact: the lottery actually created €3.48 per win.
    • +
    • The verdict: every candidate except €16.5 contradicts a number we measured. The division is the only survivor, not a choice.
    • +
    +
    \[ \hat\beta \times \pi \;\stackrel{!}{=}\; \delta \quad\Longleftrightarrow\quad \hat\beta \;=\; \frac{\delta}{\pi} \]
    +
    Not a black box
    + Every IV estimate is the effect size that makes the instrument's sales bump add up. If you cannot state yours as a ratio of two simple differences, you do not yet understand it.
    +
    +
    +
    + +
    +
    Backup · IV · The estimator · the whole method in one division
    +

    The IV estimate

    +
    Euros per lottery win, divided by exposures per lottery win.
    +
    +
    +
    the whole method, as arithmetic on two measured numbers
    One lottery win buys 0.2106 extra exposures and €3.48 of extra sales. If each exposure is worth \(\beta\), those two facts only fit together for one \(\beta\): the division.
    @@ -2016,7 +2231,6 @@

    The IV estimate

    What just happened
    We priced the ad using only the random slice of exposure: the dashboard said €23.7, the lottery says €16.5, against a planted truth of €15.
    -

    The method never needed a model of intent, controls, or machine learning: two averages and a division.

    Deep dive · the confidence interval around €16.5
    @@ -2040,37 +2254,8 @@

    The IV estimate

    -
    -
    IV · The estimator · hands on
    -

    Why the division is forced

    -
    The division is not a modelling choice. It is the only effect size the two measurements allow.
    -
    -
    -
    ▶ LIVE every candidate effect makes a prediction. One matches.
    -
    -
    - - -
    -
    The rising line is the prediction: an effect of \(\hat\beta\) per exposure implies the lottery should have lifted sales by \(\hat\beta \times \pi\). The flat line is the fact: it lifted them by €3.48. Move your guess to the crossing and you have priced the ad.
    -
    -
    -

    Forget the formula and grade any candidate effect \(\hat\beta\) against the two numbers we own:

    -
      -
    • Its prediction: if one exposure were worth \(\hat\beta\), the lottery's 0.2106 extra exposures per win should create \(\hat\beta \times 0.2106\) euros per win.
    • -
    • The fact: the lottery actually created €3.48 per win.
    • -
    • The verdict: every candidate except €16.5 contradicts a number we measured. The division is the only survivor, not a choice.
    • -
    -
    \[ \hat\beta \times \pi \;\stackrel{!}{=}\; \delta \quad\Longleftrightarrow\quad \hat\beta \;=\; \frac{\delta}{\pi} \]
    -
    Not a black box
    - Every IV estimate is the effect size that makes the instrument's sales bump add up. If you cannot state yours as a ratio of two simple differences, you do not yet understand it.
    -
    -
    -
    - - -
    -
    IV · When it breaks · the dangerous failure
    +
    +
    Backup · IV · When it breaks · the dangerous failure

    Weak instruments

    A weak instrument is worse than no instrument.
    @@ -2121,174 +2306,8 @@

    Weak instruments

    -
    -
    IV · When it breaks · whose effect it is
    -

    Compliers and the LATE

    -
    Whom does the €16.5 describe? Only the users the lottery could move.
    -
    -
    ▶ LIVE the user base, split by how they respond to the lottery
    -
    -
    - - -
    -
    Push \(\gamma\): the complier slice grows, because the complier share is the first stage. Shares are drawn live from one simulated batch, so they can differ from the printed figures by a rounding step.
    -
    -
    -
      -
    • Always-takers (56.5%): the auction shows them the ad with or without the lottery. It changes nothing for them, so the data say nothing about them.
    • -
    • Never-takers (22.5%): never see the ad either way. Same silence.
    • -
    • Compliers (21.0%): see the ad only because the lottery favoured them. Every euro of the lottery's lift \(\delta\) came from them.
    • -
    • The caveat, monotonicity: we assume no defiers, users who would see the ad only when the lottery does not favour them. A nudge that never repels.
    • -
    -
    -
    -
    Definition · LATE (local average treatment effect)
    - The LATE is the average effect of the ad on the compliers alone, and it is what the division estimates: a local answer, not a statement about every user.
    -
    \[ \hat\beta_{\text{IV}} \;=\; \frac{\delta}{\pi} \;\;\text{ estimates }\;\; \mathbb{E}[\,Y(1)-Y(0)\mid \text{complier}\,] \] - the average of each complier's personal effect, the same \(Y(1)-Y(0)\) contrast the naive slide could not touch
    -
    -
    -
    Why a manager should love this fine print
    - Compliers are the same kind of marginal user a higher bid would newly reach: the closest thing in the data to the customer a bid change buys. That makes €16.5 a price for the marginal customer, measured on the margin rather than on the average.
    -
    -
    - -
    -
    IV · When it breaks · what must hold, on one page
    -

    The checklist

    -
    The four assumptions, and which ones the data can check.
    -
    - - - - - - - - -
    AssumptionWhat it saysIn the caseStatus
    Relevance\(Z\) moves \(X\)\(F = 156\), far above 10TESTABLE, passes
    Exogeneity\(Z \perp U\)the lottery is a genuine random drawBY DESIGN
    Exclusion\(Z \to Y\) only via \(X\)a queue bump shows the user nothingUNTESTABLE
    Monotonicityno defiersa nudge never repelsUNTESTABLE, plausible
    -
    -
    -
    The honest scorecard, for any IV study you are shown
    - One measured number (the first-stage \(F\)), one design guarantee (the randomization), two arguments (exclusion, monotonicity). Ask for all four before you accept the estimate.
    -
      -
    • What we can now defend: an exposure causes about €16.5 of sales for the users a bid can actually move.
    • -
    -
    -
    The sentence that loses money
    - "Exposed users are worth €23.7 each, so raise the bid": every word true, conclusion wrong. It books the platform's targeting as advertising.
    -
    -
    -
    - - - -
    -
    IV · The decision · euros at last
    -

    The price map

    -
    The estimate becomes a decision only when it meets the price.
    -
    -
    -
    ▶ LIVE the verdict as the price moves
    -
    -
    - - -
    -
    Blue: net value per exposure at each price. The orange band is the 90% interval, the zone where the data refuse to commit. Drag the price through the three zones and watch the verdict flip.
    -
    -
    -

    One estimate gives not one answer but a map from any price to a verdict:

    - - - - - - - -
    Price zoneVerdictWhy
    below €12.7GOeven the most pessimistic supported effect pays
    €12.7 to €20.4TESTthe data straddle the price: negotiate, or measure more
    above €20.4NO-GOno supported effect pays
    -
      -
    • Today's rate, €10, sits in the GO zone, below the whole interval.
    • -
    • The net, computed: \(Y\) is contribution euros, so one exposure nets \(\hat\beta_{\text{IV}} - c = \) €16.5 − €10 = €6.5. Even read at the interval floor, €12.7 against €10, the exposure still pays.
    • -
    -
    Why boards like this framing
    - "Is the effect significant?" has no business answer. "Up to what price is this a buy?" has one, and it is the same question a bid cap asks.
    -
    -
    -
    - -
    -
    IV · The decision
    -

    Poll · the negotiation

    -
    -
    -
    ✋ Poll
    -
    The platform wants to renegotiate the rate. Your analyst hands you the causal read: effect €16.5 per exposure, 90% interval [12.7, 20.4]. What is the highest rate at which you would still sign "buy" without further study?
    -
    - - - - -
    - -
    C. Below €12.7, every effect the data support pays: the interval's floor is a no-regret bid cap, defensible whichever value inside the interval turns out to be the truth. B is a break-even gamble: paying the point estimate wins or loses depending on which side of it the truth sits, acceptable only for a risk-neutral buyer averaging over many campaigns. A pays a price that only the single most optimistic supported effect can justify. D leaves money on the table: the whole interval sits well above today's rate.
    -
    -
    -
    - -
    -
    IV · The decision · banked
    -

    The verdict and the recommendation

    -
    The complete answer, assembled from everything measured so far.
    -
    -
    - - - - - - - - - -
    QuantityValueSource
    effect of one exposure€16.5the division δ/π
    90% interval[12.7, 20.4]classical, and AR agrees
    first-stage F156the lottery is strong
    price€10the platform's rate card
    net per exposure€6.5β − c, at the point estimate
    -
    The verdict
    - BUY  Keep buying at €10: the entire defensible range clears the price.
    -
    -
    -
    The recommendation, in three lines
    - 1. Keep buying at the €10 rate: even the interval's most pessimistic effect pays.
    - 2. Cap the bid at the interval's lower end, €12.7: up to there, every effect the data support still clears the price.
    - 3. Measure again only if the rate card climbs toward €12.7: at today's price, no effect inside the interval changes the action, so more measurement is almost certain to leave the decision unchanged and is worth close to nothing here.
    -
      -
    • Everything above is classical: two averages, one division, one F statistic, one confidence interval.
    • -
    • The one debt on record: exclusion is untestable. The recommendation is conditional on the argued design, and says so.
    • -
    -
    -
    -
    - - -
    -
    Closing · Provenance
    -

    The tools were the product too

    -
    -
      -
    • CausalPy: synthetic control, interrupted time series, difference in differences, and regression discontinuity in one open-source package. The IV estimator that closes this session joined later.
    • -
    • Its launch example: individual exposure to a TV campaign cannot be randomised, yet its causal impact remains a core business need: the sentence this whole session opened with.
    • -
    • pymc-marketing: the MMM library behind Case 2's calibration story; one client's budget allocation approach to PyMC-Marketing came back as a pull request (Bolt).
    • -
    • Webinars and content: the consultancy's own webinar agenda is this session's syllabus, Instrumental Variables included.
    • -
    -
    - CausalPy - PyMC-Marketing -
    -
    -
    - -
    -
    Closing
    +
    +
    Backup · Closing

    The pattern in every engagement

      @@ -2308,47 +2327,6 @@

      The pattern in every engagement

    -
    -
    Closing
    -

    One breath

    -
    -
    The pattern to take home
    - The toolkit a Bayesian consultancy sells: counterfactuals, calibrated by experiments, priced as probabilities.
    -
      -
    • Read the cases: pymc-labs.com/blog-posts: every number in this deck is pinned to a post, listed on the next slide.
    • -
    • Say hello: both authors consult for PyMC Labs; the notebooks behind this session are the course repository.
    • -
    -
    -
    - - -
    -
    Backup
    - Backup · Sources -

    Every number, pinned

    -
    Part 1 facts retrieved and pinned 2026-07-19 (apps/labs_deck_data.json carries the exact quote); Part 2 numbers are baked from the executed course notebooks (nb07/nb07b shards).
    -
    - - - - - - - - - - - - - - - - - -
    SourceFacts pinned
    ailab.criteo.com · criteo-uplift-prediction-datasetcriteo_rows
    pymc-labs.com · 2022-11-11-HelloFreshhf_panel_calibration
    pymc-labs.com · 2023-06-20-juan-marketing-analyticswebinar_agenda
    pymc-labs.com · bayes-is-slow-speeding-up-hellofreshs-bayesian-ab-tests-by-60xhf_batch, hf_test_types, hf_thousands
    pymc-labs.com · bayesian-media-mix-modeling-for-marketing-optimizationhf_priors_experiments
    pymc-labs.com · causal-sales-analytics-are-my-sales-incremental-or-cannibalisticcolgate_ci, colgate_ci_level, colgate_truth, colgate_year, fail_range, fail_truth, market
    pymc-labs.com · causal-sales-analytics-discrete-choice-modelingcolgate_counterfactual_quote
    pymc-labs.com · causalpy-a-new-package-for-bayesian-causal-inference-for-quasi-experimentscausalpy_methods, causalpy_tv
    pymc-labs.com · funnel-aware-mmmcpl, cpl_window, gdpr_sentence, herp_attribution_quote, nurn_2026, trust_quote
    pymc-labs.com · marketing-mix-modeling-a-complete-guidebolt_pr
    pymc-labs.com · mmm_roas_liftlift_tests_n, roas_gap_words, roas_wrong_ranking, roas_x1, roas_x2
    pymc-labs.com · open-sourcing-decision-lab-scaling-ai-judgment-data-sciencedl_explored, dl_vanilla, dl_verdict
    pymc-labs.com · reducing-customer-acquisition-costs-how-we-helped-optimizing-hellofreshs-marketing-budgethf_var
    -
    -
    -
    @@ -3109,16 +3087,15 @@

    Every number, pinned

    const e=document.getElementById(id); if(e)e.textContent=v;}); const tr=DATA.treated,y0=DATA.y0_true,scl=DATA.synth_cl,ols=DATA.ols_synth,L=DATA.launch,W=tr.length; function draw(){clr(svg);const c=COL();const Wp=720,H=235,mL=46,mR=14,mT=32,mB=26; - let mn=1e9,mx=-1e9;[tr,y0,scl,ols].forEach(a=>a.forEach(v=>{if(vmx)mx=v;})); + let mn=1e9,mx=-1e9;[tr,y0,scl].forEach(a=>a.forEach(v=>{if(vmx)mx=v;})); const x=lin(0,W-1,mL,Wp-mR),y=lin(mn-2,mx+2,H-mB,mT); svg.appendChild(el('line',{x1:x(L),x2:x(L),y1:mT,y2:H-mB,stroke:c.orange,'stroke-width':1.3})); svg.appendChild(el('text',{x:x(L)+4,y:H-mB-6,fill:c.orange,'font-size':10},'launch')); svg.appendChild(el('path',{d:path(tr.map((v,i)=>[x(i),y(v)])),fill:'none',stroke:c.ink,'stroke-width':1.6,opacity:.5})); svg.appendChild(el('path',{d:path(y0.slice(L-1).map((v,i)=>[x(L-1+i),y(v)])),fill:'none',stroke:c.ink,'stroke-width':2,'stroke-dasharray':'6 4'})); svg.appendChild(el('path',{d:path(scl.map((v,i)=>[x(i),y(v)])),fill:'none',stroke:c.blue,'stroke-width':1.9})); - svg.appendChild(el('path',{d:path(ols.map((v,i)=>[x(i),y(v)])),fill:'none',stroke:c.red,'stroke-width':1.9})); const lg=[[c.ink,'treated (observed)',1.6,'none',.5],[c.ink,'true Y(0), post-launch',2,'6 4',1], - [c.blue,'simplex synthetic',1.9,'none',1],[c.red,'OLS synthetic',1.9,'none',1]]; + [c.blue,'simplex synthetic',1.9,'none',1]]; lg.forEach(([col,lab,wd,dash,op],i)=>{const xx=mL+8+i*168; svg.appendChild(el('line',{x1:xx,x2:xx+24,y1:10,y2:10,stroke:col,'stroke-width':wd,'stroke-dasharray':dash,opacity:op})); svg.appendChild(el('text',{x:xx+29,y:14,fill:col,'font-size':10.5,opacity:Math.max(op,.8)},lab));}); @@ -3440,35 +3417,81 @@

    Every number, pinned

    draw(); window.__redraw.push(draw); })(); -/* ==== fig: the MMM / experiment / synthetic-control loop (S7) ==== */ +/* ==== fig: the MMM / experiment calibration loop (S7) ==== */ (function(){ const svg=document.getElementById('svgLoop'); if(!svg)return; function draw(){ clr(svg); const c=COL(); - const nodes=[ - {x:170,y:52,w:150,label:'MMM',sub:'always-on model'}, - {x:88,y:212,w:150,label:'Geo experiment',sub:'episodic truth'}, - {x:252,y:212,w:150,label:'Synthetic control',sub:'reads it out'}, - ]; - function box(n,col){ - svg.appendChild(el('rect',{x:n.x-n.w/2,y:n.y-26,width:n.w,height:52,rx:10,fill:'none',stroke:col,'stroke-width':2})); - svg.appendChild(el('text',{x:n.x,y:n.y-5,'text-anchor':'middle','font-size':13,'font-weight':'700',fill:c.navy},n.label)); - svg.appendChild(el('text',{x:n.x,y:n.y+13,'text-anchor':'middle','font-size':10,fill:c.muted},n.sub)); + function box(x,y,w,label,sub,col){ + svg.appendChild(el('rect',{x:x-w/2,y:y-26,width:w,height:52,rx:10,fill:'none',stroke:col,'stroke-width':2})); + svg.appendChild(el('text',{x:x,y:y-5,'text-anchor':'middle','font-size':13,'font-weight':'700',fill:c.navy},label)); + svg.appendChild(el('text',{x:x,y:y+13,'text-anchor':'middle','font-size':10,fill:c.muted},sub)); } - box(nodes[0],c.blue); box(nodes[1],c.orange); box(nodes[2],c.green); + box(170,60,190,'MMM','always-on budget model',c.blue); + box(170,220,190,'Geo experiment','episodic ground truth',c.orange); function arrow(x1,y1,x2,y2){ svg.appendChild(el('line',{x1,y1,x2,y2,stroke:c.faint,'stroke-width':1.6})); const a=Math.atan2(y2-y1,x2-x1); svg.appendChild(el('path',{d:`M${x2} ${y2} L${x2-9*Math.cos(a-0.4)} ${y2-9*Math.sin(a-0.4)} L${x2-9*Math.cos(a+0.4)} ${y2-9*Math.sin(a+0.4)} Z`,fill:c.faint})); } - arrow(128,80,100,182); // MMM -> experiment (asks) - arrow(120,182,148,80); // experiment -> MMM (calibrates) - arrow(166,224,176,224); // experiment -> SC - svg.appendChild(el('text',{x:56,y:132,'font-size':10,fill:c.muted},'asks for')); - svg.appendChild(el('text',{x:56,y:144,'font-size':10,fill:c.muted},'ground truth')); - svg.appendChild(el('text',{x:152,y:132,'font-size':10,fill:c.muted},'calibrates')); - svg.appendChild(el('text',{x:152,y:144,'font-size':10,fill:c.muted},'(priors, lift tests)')); - svg.appendChild(el('text',{x:170,y:262,'text-anchor':'middle','font-size':10,fill:c.muted},'MMM: last slide · synthetic control: Part 2 · the loop: this session')); + arrow(120,86,120,194); + arrow(220,194,220,86); + svg.appendChild(el('text',{x:108,y:140,'font-size':10,fill:c.muted,'text-anchor':'end'},'asks for')); + svg.appendChild(el('text',{x:108,y:152,'font-size':10,fill:c.muted,'text-anchor':'end'},'ground truth')); + svg.appendChild(el('text',{x:232,y:140,'font-size':10,fill:c.muted},'calibrates')); + svg.appendChild(el('text',{x:232,y:152,'font-size':10,fill:c.muted},'(priors, lift tests)')); + svg.appendChild(el('text',{x:170,y:282,'text-anchor':'middle','font-size':10,fill:c.muted},'reading a geo experiment out is its own craft:')); + svg.appendChild(el('text',{x:170,y:294,'text-anchor':'middle','font-size':10,fill:c.muted},'that method is Part 2 of this session')); + } + draw(); window.__redraw.push(draw); +})(); + +/* ==== fig: the PyMC ecosystem stack (slide 2) ==== */ +(function(){ + const svg=document.getElementById('svgEco'); if(!svg)return; + function draw(){ + clr(svg); const c=COL(); + function node(x,y,w,label,sub,col,dash){ + const attrs={x:x-w/2,y:y-19,width:w,height:38,rx:9,fill:'none',stroke:col,'stroke-width':2}; + if(dash)attrs['stroke-dasharray']='5 4'; + svg.appendChild(el('rect',attrs)); + svg.appendChild(el('text',{x:x,y:y-1,'text-anchor':'middle','font-size':12,'font-weight':'700',fill:c.navy},label)); + if(sub)svg.appendChild(el('text',{x:x,y:y+13,'text-anchor':'middle','font-size':9,fill:c.muted},sub)); + } + function arrow(x1,y1,x2,y2){ + svg.appendChild(el('line',{x1,y1,x2,y2,stroke:c.faint,'stroke-width':1.6})); + const a=Math.atan2(y2-y1,x2-x1); + svg.appendChild(el('path',{d:`M${x2} ${y2} L${x2-8*Math.cos(a-0.4)} ${y2-8*Math.sin(a-0.4)} L${x2-8*Math.cos(a+0.4)} ${y2-8*Math.sin(a+0.4)} Z`,fill:c.faint})); + } + node(95,66,120,'PyTensor','compute engine',c.grey); + node(285,66,155,'PyMC','probabilistic programming',c.blue); + node(540,22,170,'PyMC-Marketing','MMM, CLV',c.green); + node(540,66,170,'CausalPy','quasi-experiments',c.orange); + node(540,110,170,'Bespoke','customer-specific builds',c.gold,true); + arrow(157,66,205,66); + arrow(365,60,453,26); + arrow(365,66,453,66); + arrow(365,72,453,106); + } + draw(); window.__redraw.push(draw); +})(); + +/* ==== fig: the metros, one treated (boardroom slide) ==== */ +(function(){ + const svg=document.getElementById('svgMetros'); if(!svg)return; + function draw(){ + clr(svg); const c=COL(); + const donors=[[60,40,7],[80,120,9],[130,30,6],[210,125,7],[220,55,10],[255,95,6],[290,35,8], + [320,120,9],[350,70,7],[385,30,6],[400,105,8],[430,60,11],[465,115,6],[480,35,7],[510,85,9], + [545,40,6],[560,120,8],[590,70,7],[620,105,6],[640,35,9],[665,80,7],[75,70,5],[170,45,5], + [240,20,5],[365,115,5],[450,20,5],[530,115,5],[610,20,5],[680,120,5]]; + donors.forEach(([x,y,r])=>svg.appendChild(el('circle',{cx:x,cy:y,r:r,fill:c.grey,opacity:.28,stroke:c.grey,'stroke-width':1}))); + svg.appendChild(el('circle',{cx:150,cy:70,r:14,fill:c.orange,opacity:.3,stroke:c.orange,'stroke-width':2.2})); + svg.appendChild(el('path',{d:'M 166 52 A 24 24 0 0 1 174 70',fill:'none',stroke:c.orange,'stroke-width':1.6})); + svg.appendChild(el('path',{d:'M 170 45 A 32 32 0 0 1 181 70',fill:'none',stroke:c.orange,'stroke-width':1.2,opacity:.7})); + svg.appendChild(el('text',{x:150,y:105,'text-anchor':'middle','font-size':10.5,fill:c.orange,'font-weight':700},'the treated metro')); + svg.appendChild(el('text',{x:150,y:118,'text-anchor':'middle','font-size':9.5,fill:c.orange},'the €75k campaign, weeks 40-59')); + svg.appendChild(el('text',{x:545,y:16,'text-anchor':'middle','font-size':10.5,fill:c.muted,'font-weight':700},'29 donor markets · no campaign')); } draw(); window.__redraw.push(draw); })(); diff --git a/causal-marketing-pymc/apps/unified_slides_src.html b/causal-marketing-pymc/apps/unified_slides_src.html index 942903e..26d91b0 100644 --- a/causal-marketing-pymc/apps/unified_slides_src.html +++ b/causal-marketing-pymc/apps/unified_slides_src.html @@ -236,14 +236,15 @@

    Causal Inference in the Wild

    + PyMC Labs
    Opening · Who is talking

    PyMC Labs: what we do

    A Bayesian modeling consultancy: custom decision-making models where off-the-shelf tools fall short, and the open-source libraries the field runs on (PyMC, PyMC-Marketing, CausalPy).
      -
    • What we sell: senior modeling expertise, not software licenses.
    • +
    • What we sell: Senior modeling expertise, end-to-end data science, not a software license.
    • Open source is the top of the funnel: the libraries build reach; the consulting monetizes the expertise behind them.
    • -
    • How we work: small senior teams, starting from the client's decision; the client keeps the model, not a black box.
    • +
    • How we work: small senior teams; the client keeps the model and surrounding functionality transparently, not the core libraries.
    @@ -253,8 +254,11 @@

    PyMC Labs: what we do

    EngagementWhat it isTypical client
    Enablement & trainingWorkshops and upskilling on Bayesian methods and the toolingAnalytics teams standardizing on PyMC
    -
    The clients in this session
    - Colgate-Palmolive, HelloFresh, Nürnberger Versicherung (and a Bolt cameo): you meet three of them in the next ten minutes.
    +
    +
    +
    The clients in this session
    + Colgate-Palmolive, HelloFresh, Nürnberger Versicherung (and a Bolt cameo): you meet three of them in the next ten minutes.
    +
    Areas of interest, and where we work
    • Areas: marketing-mix modeling and media measurement, causal inference, demand forecasting and pricing, experimentation and A/B testing at scale, applied Bayesian modeling.
    • @@ -328,7 +332,7 @@

      Bayesian vs frequentist, in one slide

      Why marketing analytics is going Bayesian
      Marketing data is short, noisy, and highly correlated, and every model ends in a spend decision, which is exactly where priors and full uncertainty pay off.
      For today, this is a curiosity
      - The entire lecture runs on classical methods, and every number you will see is frequentist; we surface the Bayesian read only as a flavour at the edges, never as the load-bearing tool.
      + Our lectures will center around classical methods to stay didactically coherent. We will only hint at what the Bayesian approach affords.
    Where it is already the default (and where it is not)
    • Marketing mix modeling: the modern open-source MMM tools are Bayesian (Google's Meridian, PyMC-Marketing); with only two or three years of weekly data and correlated channels, plain regression is unstable, and priors on adstock, saturation, and ROI stabilize it.
    • @@ -378,7 +382,7 @@

      You are the consultant

      Case 1 · Colgate-Palmolive

      Colgate-Palmolive: incremental, or cannibalistic?

      -
      Incremental: sales won from competitors or category growth. Cannibalistic: sales taken from your own products. The launch verdict is the split.
      +
      Incremental: sales won from competitors or category growth. Cannibalistic: sales taken from your own products. The question for the launch: how much of its sales is which?
      @@ -392,8 +396,8 @@

      Colgate-Palmolive: incremental, or cannibalistic?

      "We need to estimate the counterfactual sales of all products would have been if the new product had not been introduced."
      • The client: Colgate-Palmolive, {{labs.colgate_year}}; a market estimated at {{labs.market}}.
      • -
      • The method: a multivariate Bayesian interrupted time series: project the pre-launch world forward.
      • -
      • The grading: a planted {{labs.colgate_truth}} recovered as a {{labs.colgate_ci_level}} interval of {{labs.colgate_ci}}: recover a known truth first, then be believed.
      • +
      • The method: a multivariate interrupted time series: project the pre-launch world forward.
      • +
      • The grading: the model is first run on simulated sales where the true incrementality is fixed at {{labs.colgate_truth}} by construction; it estimates {{labs.colgate_ci}} (a {{labs.colgate_ci_level}} interval): it must find an answer we planted before we trust it on the real one.
      @@ -448,7 +452,7 @@

      Why calibrate? A model alone can rank channels backwards

      The settinga company spreads its budget across many ad channels (TV, search, social) and needs to know which ones actually pay back.
      -
      The everyday toola model reads years of spend-and-sales history and scores each channel's return on ad spend, cheaply and always-on.
      +
      The everyday tool: a marketing-mix model (MMM)it reads years of spend-and-sales history and scores each channel's return on ad spend, cheaply and always-on.
      The catch, and our jobhistory is not an experiment; a company spends more exactly when demand is already high, so the model can credit the season instead of the channel and rank them backwards.
      What we adda real experiment: nudge one channel's budget by a known amount, measure the sales it truly causes, and anchor the model to that number. This is calibration.
      @@ -489,25 +493,41 @@

      HelloFresh runs the loop, at industrial scale

      +
      +
      Case 3 · Nürnberger Versicherung
      +

      Nürnberger Versicherung: steering by the last click

      +
      A German insurance group, spending across the whole funnel: awareness video and demand generation up top, branded search at the bottom.
      +
      +
      +
      The old ruler: last-touch attributioncredit every sale to the last ad click before it; simple, standard, and blind to everything upstream of that click.
      +
      What broke itunder the GDPR (the EU's General Data Protection Regulation, the 2018 privacy law that restricts user-level tracking), {{labs.gdpr_sentence}}: the last click was often the only click the tracker could still see.
      +
      The replacement: a funnel-aware causal MMMa marketing-mix model that encodes the funnel: upper-funnel spend creates demand that surfaces later in lower-funnel channels, so credit flows to the cause, not to the final click.
      +
      The stakesbudget follows the ruler: whatever the measurement under-credits, the spreadsheet de-funds.
      +
      +
      The client, before the fix
      + "We were optimizing for {{labs.herp_attribution_quote}}." (Philip Herp, Nürnberger Versicherung)
      +
      +
      +
      Case 3 · Nürnberger Versicherung
      -

      Price the engagement

      -
      A German insurer, last-touch attribution, and a funnel-aware causal MMM, in production.
      +

      Who does the last click cheat?

      +
      The terms are on the table; the poll asks you to use them.
      -
      +
      ✋ Poll
      -
      Nürnberger Versicherung replaced last-touch attribution steering with a funnel-aware causal MMM. Over {{labs.cpl_window}} of model-guided spend, cost per lead (CPL) moved by how much?
      +
      Under GDPR-shortened tracking, which channels does last-touch attribution systematically under-credit?
      - - - - + + + +
      -
      C. "This year we were able to drive the CPL down by {{labs.cpl}}, which is very, very good" (Philip Herp, Nürnberger Versicherung). Under GDPR, {{labs.gdpr_sentence}}, so last-touch under-credited the upper funnel and budget followed {{labs.herp_attribution_quote}}. The client is scaling into {{labs.nurn_2026}}.
      +
      B. Under GDPR, {{labs.gdpr_sentence}}, so the last click was often the only visible step, the upper funnel vanished from the books, and budget followed {{labs.herp_attribution_quote}}. The funnel-aware model re-priced it: "This year we were able to drive the CPL down by {{labs.cpl}}, which is very, very good" (Philip Herp), over {{labs.cpl_window}} of model-guided spend, and cost per lead (CPL) is the price of one new customer lead. The client is scaling into {{labs.nurn_2026}}.
      The client's bar for belief
      "{{labs.trust_quote}}"
      @@ -576,6 +596,11 @@

      The boardroom question

    +
    +
    One metro got the campaign; twenty-nine watched schematic
    +
    +
    +
    The method this data calls for
    • One treated unit, no A/B test possible: synthetic control's home ground.
    • @@ -586,18 +611,8 @@

      The boardroom question

    • δ = the share of the pilot's per-euro lift that survives national rollout.
    • The €4M call reduces to one question: is δ large enough to still clear the cost?
    -
    What "unbiased" means here, precisely, and where it would break
    -

    The estimator subtracts a reconstructed counterfactual from the treated metro's post-launch sales, and bias is its expected gap from the true lift:

    -
    - \[ \hat\tau \;=\; \bar Y_{1,\text{post}} - \hat Y_{1,\text{post}}(0), \qquad \text{bias} \;\equiv\; \mathbb{E}[\hat\tau] - \tau \]
    -
      -
    • The counterfactual is the synthetic twin, built to reproduce the metro's pre-launch record (its level and its factor loadings) week by week.
    • -
    • Unbiasedness rests on one assumption: that pre-launch match would have continued through the post window had the campaign never run.
    • -
    • This needs no story about why the metro was chosen. Whatever pre-launch trait drove the pick, the twin already carries it, because it was fitted to match it: selection on anything visible before launch cannot tilt the estimate.
    • -
    • Bias enters only through what the pre-launch record cannot see: if the launch were timed on private knowledge of a coming local boom, the twin could not anticipate it and the gap would credit the boom to the campaign. The later placebo-in-time test hunts exactly that.
    • -
    -
    +
    @@ -627,7 +642,7 @@

    The data you actually get

    - +
    @@ -710,7 +725,6 @@

    Causal inference is a missing-data problem

  • Subscript says which market, argument says which world (1 = campaign, 0 = none): distinct axes that happen to share the digit 1.
  • \(w_j\): the donor weights the synthetic control will choose so \(\sum_{j\ge 2} w_j Y_{jt}\) rebuilds the missing \(Y_{1t}(0)\).
  • -

    "What would this metro have sold anyway?" is an estimation target, not a rhetorical question.

    @@ -761,7 +775,7 @@

    Simulate the world yourself

    Take a look at the counterfactual.
    -
    ▶ LIVE The equations of the previous slide. Grey: donor markets. White: the treated metro.
    +
    ▶ LIVE The equations of the previous slide. Grey: donor markets. Black: the treated metro.
    @@ -769,7 +783,7 @@

    Simulate the world yourself

    - +
    @@ -825,36 +839,6 @@

    Abadie's idea: if no twin exists, build one

    -
    -
    Act II · The counterfactual problem
    -

    The estimator, precisely

    -
    A constrained least square optimization.
    -
    -

    Fit the weights on the pre-period only, constrained to the simplex:

    -
    - \[ \hat w \;=\; \arg\min_{w\in\Delta}\; \sum_{t \lt T_0}\Bigl(Y_{1t}-\sum_j w_j Y_{jt}\Bigr)^{2}, \qquad \Delta=\Bigl\{w : w_j\ge 0,\ \textstyle\sum_j w_j=1\Bigr\} \] -
    -
      -
    • Pre-launch only: the optimiser never sees the post-period, so it cannot cheat.
    • -
    • The simplex \(\Delta\): a readable recipe, never "−80% of Milan + 190% of Rome": Colgate's projection with the weights on the table.
    • -
    • No standard error: a point estimate only; inference comes later, from placebos.
    • -
    -
    Reading the effect off the gap
    -
    - \[ \hat\tau_t \;=\; Y_{1t}-\sum_j \hat w_j Y_{jt}, \qquad\qquad \hat\tau \;=\; \sum_{t\ge T_0}\hat\tau_t \] - weekly gap between the treated metro and its synthetic twin, summed over the 20 post-launch weeks -
    -
    -
    Why constrain at all?
    - The simplex buys three things ordinary regression cannot: interpretability, regularisation, and an off-switch.
    -
      -
    • Interpretability: "40% dma_08 + 32% dma_20 + 17% dma_03" is a sentence a planner can act on.
    • -
    • Regularisation: most weights land on exactly zero.
    • -
    • An off-switch: a market no blend can match fails loudly instead of extrapolating.
    • -
    -
    -
    -
    Act II · The counterfactual problem
    @@ -891,65 +875,18 @@

    What the simplex buys, geometrically: stay inside the hull

    Act II · The counterfactual problem
    -

    What the constraint buys: drop it and see

    -
    What if we used a simple unconstrained OLS?
    +

    The synthetic twin, graded against the truth

    +
    Only a simulation can draw the true counterfactual; hold the twin against it.
    -
    ▶ LIVE Two synthetics against the truth only a simulation can draw
    +
    ▶ LIVE The synthetic against the truth only a simulation can draw
    -
    OLS hugs the pre-period tighter ( vs ) yet drifts further from the true \(Y_{1t}(0)\) after launch ( vs ): overfitting the simplex refuses.
    -
    -
    -
    -
    The weights OLS chose
    -
    -
    OLS puts negative weight on donors (down to ) and gross weight , where the simplex uses exactly 1.
    -
    -
    -
    - \[ n_{\mathrm{eff}} \;=\; \frac{1}{\sum_j \hat w_j^{2}} \] - Effective number of donors (inverse Herfindahl): one donor \(\Rightarrow n_{\mathrm{eff}}=1\), all 29 equally \(\Rightarrow 29\). Ours: 3.3. -
    -
    What n_eff = 3.3 tells you
    -
      -
    • The synthetic leans on about three donor markets, not a fuzzy mix of all 29.
    • -
    • Sparse weights: interpretable, and less room to overfit.
    • -
    • Near 1: hostage to one market. Near 29: mush. 3.3: concentrated, not fragile.
    • -
    -
    +
    Pre-launch fit: RMSE . Post-launch gap to the true \(Y_{1t}(0)\): . The twin tracks a line it was never shown.
    -
    -
    Act II · The counterfactual problem, versus machine learning
    -

    "Why not gradient boosting / Prophet / an LSTM?"

    -
    Is this a forecasting exercise?
    -
    -
    -
    ▶ LIVE A kitchen-sink forecaster (donors + trend + seasonality, fit pre-launch only)
    -
    -
    Solid: treated sales. Dashed blue: the forecaster. Dashed grey: the simplex twin. Both fit only the 40 pre-launch weeks.
    -
    -
      -
    • In-sample it fits tighter:k vs €k per week: flexibility always buys the past.
    • -
    • Yet the estimates are a wash: vs , both near the planted €284k. One dataset cannot rank them.
    • -
    • The sweep decides: across 24 fresh worlds, {{nb07.hull_oos_ratio}}× worse out of sample, with nothing to inspect.
    • -
    -
    -
    -
    Principle 1: a counterfactual, not a forecast
    - A forecaster asks what comes next. We ask what this market would have done in a world that never happened.
    -
      -
    • Principle 2: in-sample fit is the trap, not the goal.
    • -
    • Principle 3: you cannot cross-validate the counterfactual. The truth is never observed.
    • -
    • Principle 4: the simplex ships an auditable claim; a boosted tree is a black box.
    • -
    -
    The forecaster is better at prediction. The simplex is better at the causal job.
    -
    -
    -
    Act II · The counterfactual problem
    @@ -1008,90 +945,7 @@

    Compare the estimators

    -
    -
    Act II · The counterfactual problem
    -

    Interesting limits for the simpler estimators

    -
    Where before/after and treated/control estimators break.
    -
    -
    -
    -
    Before/after · one unit, across time
    - The treated metro's post-average minus its pre-average.
    -
      -
    • The level cancels; nothing subtracts the shared wave.
    • -
    • The whole drift \(\Delta\bar f\) lands on the metro: the largest bias, growing with the horizon.
    • -
    -
    -
    -
    Treated-vs-control · one period, across units
    - The treated metro minus the control average, both in the post window.
    -
      -
    • The level gap \(\alpha_1-\bar\alpha_C\) survives, and sizes differ a lot: it dominates.
    • -
    • This is why raw sales are never compared across cities.
    • -
    -
    -
    -
    The two bias decompositions, term by term
    -
    - \[ \hat\tau^{\text{B/A}} \;=\; \bar Y_{1,\text{post}}-\bar Y_{1,\text{pre}} \;=\; \bar\tau \;+\; \underbrace{\gamma_1^{\top}\,\Delta\bar f}_{\text{bias: the whole drift}} \;+\; \Delta\bar\varepsilon_1 \] - the level \(\alpha_1\) cancels, but the whole shared-factor drift \(\Delta\bar f\) lands at the metro's own loadings: the largest bias, and it grows with the horizon. -
    -
    - \[ \hat\tau^{\text{TC}} \;=\; \bar Y_{1,\text{post}}-\bar Y_{C,\text{post}} \;=\; \bar\tau \;+\; \underbrace{(\alpha_1-\bar\alpha_C)}_{\text{level gap}} \;+\; \underbrace{(\gamma_1-\bar\gamma_C)^{\top}\bar f_{\text{post}}}_{\text{loading gap}\,\times\,\text{level}} \;+\; \Delta\bar\varepsilon \] - no time-differencing, so the level gap survives and dominates; the loading gap now multiplies the factor level, not its drift. -
    -
    -
    -
    - -
    -
    Act II · The counterfactual problem
    -

    Interesting limits for DiD

    -
    Where the DiD estimator breaks.
    -
    -
    The one question
    - DiD only works if treated and controls would have drifted together. Two dials decide (spread \(s\), macro \(\sigma_\eta\)); the algebra is in the fold.
    -
    Synthetic control solves it
    - DiD's residual bias \(B=(\gamma_1-\bar\gamma_C)^{\top}\Delta\bar f\) comes from its equal weights: it hopes \(s\) is small. Synthetic control chooses the weights, so \(B\to0\) at any spread.
    -
    The DiD bias, decomposed (and why more data cannot shrink it)
    -
    - \[ \hat\tau^{\text{DiD}} \;=\; \bar\tau \;+\; \underbrace{(\gamma_1-\bar\gamma_C)^{\top}\,\Delta\bar f}_{\text{systematic bias }B} \;+\; \underbrace{\Delta\bar\varepsilon_1-\Delta\bar\varepsilon_C}_{\text{idiosyncratic noise}} \] -
    -
      -
    • The pieces: \(\Delta\bar f\) is the post-minus-pre drift of the three factors, \(\bar\gamma_C\) the \(n\) controls' average loadings, \(\Delta\bar\varepsilon\) the change in idiosyncratic noise.
    • -
    • \(B\) is mean-zero over the exposure draw, but your world is one draw, so in it \(B\) is a fixed nonzero number whose typical size is:
    • -
    -
    - \[ \operatorname{Var}(B) \;=\; \frac{s^{2}}{3}\Bigl(1+\frac{1}{n}\Bigr)\Bigl[(\Delta\bar f_{\text{trend}})^{2}+(\Delta\bar f_{\text{season}})^{2}+(\Delta\bar f_{\text{walk}})^{2}\Bigr] \] -
    -
      -
    • \(B\) carries no term for the amount of data: \(\Delta\bar f\) is the world's drift, not sampling error, so more weeks shrink only noise and buy precision around the same biased number.
    • -
    • The noise never leaves: the idiosyncratic term carries \(\sigma_\varepsilon \approx 3\) whatever \(s\) and \(\sigma_\eta\) do, so the honest claim in any limit is "the systematic bias vanishes", never "DiD equals the truth".
    • -
    -
    -
    The two limits, dial by dial: send a knob to zero and read what dies
    -
    -
    Limit 1 · loading spread \(s \to 0\): sufficient on its own
    -
      -
    • The prefactor \(s^2/3\) kills the whole bracket at once.
    • -
    • Every loading \(\to 1\), so \(\gamma_1-\bar\gamma_C\to 0\) deterministically: exact parallel trends, macro walk included.
    • -
    • DiD then recovers \(\bar\tau\) up to noise.
    • -
    -
    Limit 2 · macro shock \(\sigma_\eta \to 0\): not sufficient
    -
      -
    • Only the walk's term \((\Delta\bar f_{\text{walk}})^2\) drops out.
    • -
    • Trend and seasonality still drift between the windows, and markets still weight them differently whenever \(s>0\): \(B \neq 0\).
    • -
    • What it buys: it removes the one stochastic, horizon-growing confounder; the bias that remains is at least deterministic.
    • -
    -
    -
    -
    Where the \(\tfrac{s^2}{3}\bigl(1+\tfrac1n\bigr)\) prefactor comes from
    - Each loading is drawn from \(\mathrm{U}(1{-}s,\,1{+}s)\), a uniform of width \(2s\), whose variance is \((2s)^2/12 = s^2/3\). The treated metro contributes \(\operatorname{Var}(\gamma_1)=s^2/3\); the average of \(n\) independent controls contributes \(\operatorname{Var}(\bar\gamma_C)=s^2/(3n)\). Independent draws add, so \(\operatorname{Var}(\gamma_1-\bar\gamma_C)=\tfrac{s^2}{3}\bigl(1+\tfrac1n\bigr)\) per factor. Each factor's mismatch is then scaled by that factor's post-minus-pre drift, and the three squared drifts sum: the bracket. -
    -
    -
    -
    Act II · The counterfactual problem
    @@ -1125,103 +979,8 @@

    What must be true for the gap to be causal

    -
    -
    Act III · Is it real, and how big?
    -

    €260k. Real, or a lucky metro? Build the null yourself

    -
    -
    -
    ✋ Poll
    -
    With one treated metro a t-test is not weak, it is undefined. So build the test yourself: rerun the whole pipeline 29 more times, each donor pretended treated ("placebo"): 30 "effects", 29 where nothing ran. Where does our €260k rank?
    -
    - - - -
    - -
    A: rank 1 of 30. The rank is the inference: if the campaign did nothing, P(rank 1 by luck) = 1/30 ≈ 0.033. You just re-derived the permutation test.
    -
    -
    -
    - -
    -
    Act III · Is it real, and how big?
    -

    Placebo-in-space: measure the luck directly

    -
    Randomisation inference: the t-test had no standard error, so the 29 donors become the null distribution.
    -
    -
    -
    -
    -
    ▶ LIVE Refit the estimator on every donor as if it were treated
    -
    -
    The null: the "effects" the method reports where nothing happened (spread ·, best placebo ·). Green line: our €260k, outside the cloud.
    -
    -
    -
    -
    The null hypothesis, stated
    - \(H_0\): the campaign did nothing. Then the metro is exchangeable with its donors: any of the 30 ranks is equally likely.
    -
    - \[ p \;=\; \frac{1+\#\{\,j:\ |\hat\tau_j|\ge|\hat\tau_1|\,\}}{J+1} \;=\; \frac{1+0}{30}\;\approx\;0.033 \] -
    -
    -
    -
    What the p-value means (and what it is not)
    - The probability, if \(H_0\) were true, of a gap this extreme: a rank, no bell curve assumed. \(p\) sits at its floor \(1/30\), set by the donor count, not the weeks.
    -
    When the rank is valid, and when it lies (Abadie's hygiene rule)
    -
      -
    • Valid under \(H_0\) when the placebos are fair stand-ins: comparable pre-launch fit, and no campaign spillover onto donors (SUTVA), the two ways exchangeability can hold.
    • -
    • Lies when a placebo fits its own pre-period badly: it books fitting failure as a giant fake effect and fattens the tail. Abadie's rule: drop those placebos before ranking.
    • -
    -
    Conclusion: the €260k lift is real, at \(p\approx0.033\)
    - Rank 1 of 30, clear of the cloud: reject \(H_0\). The rank holds (p = 0.033 every time) dropping shaky placebos at 2×, 5×, 20× pre-fit error. But "real" is not "profitable": Act IV.
    -
    -
    - -
    -
    Act III · Is it real, and how big?
    -

    From a test to an interval: inversion

    -
    Which true lifts could plausibly have produced our €260k? Keep the survivors: no bell curve anywhere.
    -
    -
    -
    -
    -
    ▶ LIVE Drag \(H\), a guess at the truth: the error cloud slides with it; our €260k never moves.
    -
    -
    - - -
    -
    Axis: 20-week total gap (€000). Top: the placebo errors at truth 0. Middle: the same errors slid to \(H\); \(H\) survives if the green line sits inside the shaded middle 90%. Bottom: the survivors, collected: the interval. Drag past an edge to reject.
    -
    -
    -
    -
    The whole idea
    - Ask of every possible true lift: could it plausibly have produced our €260k? Collect the ones that could, and that set of survivors is the interval.
    -
      -
    • ① Measure the error. The 29 placebo "effects" form a cloud around zero, spread about ±€50k.
    • -
    • ② Guess a truth \(H\). If the lift were \(H\), we would see \(H\) plus that same cloud.
    • -
    • ③ Keep or reject. Keep \(H\) if €260k sits inside its middle 90%.
    • -
    • ④ Sweep. The survivors run €195k to €335k: the interval is [€195k, €335k].
    • -
    -
    -
    -
    The inversion, in one line of algebra
    -
    - \[ \underbrace{H + q_{0.05} \;\le\; 260 \;\le\; H + q_{0.95}}_{\text{step ③: €260k sits in \(H\)'s middle \(90\%\)}} - \qquad\Longleftrightarrow\qquad - \underbrace{260 - q_{0.95} \;\le\; H \;\le\; 260 - q_{0.05}}_{\text{step ④: the same line, solved for \(H\)}} \] - \(q_{0.05},q_{0.95}\) are just the low and high edges of the error cloud from step ①, here \(q_{0.05}\!=\!-75\) and \(q_{0.95}\!=\!+65\) (€000). Rearranging the left inequality into the right one is the whole trick, and it hands you the endpoints €195k and €335k directly. -
    -
    -
    Why this interval is the referee for the rest of the lecture
    - It assumed no normality, no independence, no error model, so every model-based interval later (the Bayesian posterior included) has to answer to it. -
    -
    -
    -
    Act III · Is it real, and how big? Stress-test the estimate
    @@ -1274,26 +1033,6 @@

    Falsification 2: the assumption no placebo can see

    -
    -
    Act III · Is it real, and how big?
    -

    Statistics done. Three numbers.

    -
    Questions ① and ② from the boardroom slide are now answered.
    -
    - - - - - - - -
    QuestionAnswerTool that answered it
    Is the effect real?Yes, p = 0.033placebo-in-space permutation: rank 1 of 30
    How big?€260k of incremental salessynthetic-control gap, summed over 20 weeks
    Give or take?[€195k, €335k] at 90%test inversion over the placebo cloud
    -
    Truth check (only a simulation allows it)
    - The planted total €284k sits inside the interval, €24k above the estimate: the machinery works, and its self-reported uncertainty is honest.
    -
    The sentence that loses money
    - "€260k of sales for €75k, a 3.5× return. Roll it out." Every number true; the conclusion does not follow. Question ③ is not a statistics question.
    -
    -
    -
    Act IV · The decision in euros
    @@ -1837,7 +1576,7 @@

    The ideal experiment

    • Where it hides: a lottery, a rollout order, an arbitrary rule that moved exposure for reasons unrelated to intent.
    -
    The plan for Part 2
    +
    Going forward
    1. Give that random lever a name: an instrument.
    2. State the conditions it must satisfy.
    3. Check which of them the data can verify.
    @@ -1995,13 +1734,477 @@

    The reduced form

    -
    -
    IV · The estimator · the whole method in one division
    -

    The IV estimate

    -
    Euros per lottery win, divided by exposures per lottery win.
    -
    -
    -
    the whole method, as arithmetic on two measured numbers
    + +
    +
    IV · When it breaks · whose effect it is
    +

    Compliers and the LATE

    +
    Whom does the €{{nb11.iv_est}} describe? Only the users the lottery could move.
    +
    +
    ▶ LIVE the user base, split by how they respond to the lottery
    +
    +
    + + +
    +
    Push \(\gamma\): the complier slice grows, because the complier share is the first stage. Shares are drawn live from one simulated batch, so they can differ from the printed figures by a rounding step.
    +
    +
    +
      +
    • Always-takers ({{nb11.st_share_always}}%): the auction shows them the ad with or without the lottery. It changes nothing for them, so the data say nothing about them.
    • +
    • Never-takers ({{nb11.st_share_never}}%): never see the ad either way. Same silence.
    • +
    • Compliers ({{nb11.st_share_complier}}%): see the ad only because the lottery favoured them. Every euro of the lottery's lift \(\delta\) came from them.
    • +
    • The caveat, monotonicity: we assume no defiers, users who would see the ad only when the lottery does not favour them. A nudge that never repels.
    • +
    +
    +
    +
    Definition · LATE (local average treatment effect)
    + The LATE is the average effect of the ad on the compliers alone, and it is what the division estimates: a local answer, not a statement about every user.
    +
    \[ \hat\beta_{\text{IV}} \;=\; \frac{\delta}{\pi} \;\;\text{ estimates }\;\; \mathbb{E}[\,Y(1)-Y(0)\mid \text{complier}\,] \] + the average of each complier's personal effect, the same \(Y(1)-Y(0)\) contrast the naive slide could not touch
    +
    +
    +
    Why a manager should love this fine print
    + Compliers are the same kind of marginal user a higher bid would newly reach: the closest thing in the data to the customer a bid change buys. That makes €{{nb11.iv_est}} a price for the marginal customer, measured on the margin rather than on the average.
    +
    +
    + +
    +
    IV · When it breaks · what must hold, on one page
    +

    The checklist

    +
    The four assumptions, and which ones the data can check.
    +
    + + + + + + + + +
    AssumptionWhat it saysIn the caseStatus
    Relevance\(Z\) moves \(X\)\(F = {{nb11.f_stat}}\), far above 10TESTABLE, passes
    Exogeneity\(Z \perp U\)the lottery is a genuine random drawBY DESIGN
    Exclusion\(Z \to Y\) only via \(X\)a queue bump shows the user nothingUNTESTABLE
    Monotonicityno defiersa nudge never repelsUNTESTABLE, plausible
    +
    +
    +
    The honest scorecard, for any IV study you are shown
    + One measured number (the first-stage \(F\)), one design guarantee (the randomization), two arguments (exclusion, monotonicity). Ask for all four before you accept the estimate.
    +
      +
    • What we can now defend: an exposure causes about €{{nb11.iv_est}} of sales for the users a bid can actually move.
    • +
    +
    +
    The sentence that loses money
    + "Exposed users are worth €{{nb11.naive}} each, so raise the bid": every word true, conclusion wrong. It books the platform's targeting as advertising.
    +
    +
    +
    + + + +
    +
    IV · The decision · euros at last
    +

    The price map

    +
    The estimate becomes a decision only when it meets the price.
    +
    +
    +
    ▶ LIVE the verdict as the price moves
    +
    +
    + + +
    +
    Blue: net value per exposure at each price. The orange band is the 90% interval, the zone where the data refuse to commit. Drag the price through the three zones and watch the verdict flip.
    +
    +
    +

    One estimate gives not one answer but a map from any price to a verdict:

    + + + + + + + +
    Price zoneVerdictWhy
    below €{{nb11.iv_lo}}GOeven the most pessimistic supported effect pays
    €{{nb11.iv_lo}} to €{{nb11.iv_hi}}TESTthe data straddle the price: negotiate, or measure more
    above €{{nb11.iv_hi}}NO-GOno supported effect pays
    +
      +
    • Today's rate, €{{nb11.cost}}, sits in the GO zone, below the whole interval.
    • +
    • The net, computed: \(Y\) is contribution euros, so one exposure nets \(\hat\beta_{\text{IV}} - c = \) €{{nb11.iv_est}} − €{{nb11.cost}} = €{{nb11.net}}. Even read at the interval floor, €{{nb11.iv_lo}} against €{{nb11.cost}}, the exposure still pays.
    • +
    +
    Why boards like this framing
    + "Is the effect significant?" has no business answer. "Up to what price is this a buy?" has one, and it is the same question a bid cap asks.
    +
    +
    +
    + +
    +
    IV · The decision
    +

    Poll · the negotiation

    +
    +
    +
    ✋ Poll
    +
    The platform wants to renegotiate the rate. Your analyst hands you the causal read: effect €{{nb11.iv_est}} per exposure, 90% interval [{{nb11.iv_lo}}, {{nb11.iv_hi}}]. What is the highest rate at which you would still sign "buy" without further study?
    +
    + + + + +
    + +
    C. Below €{{nb11.iv_lo}}, every effect the data support pays: the interval's floor is a no-regret bid cap, defensible whichever value inside the interval turns out to be the truth. B is a break-even gamble: paying the point estimate wins or loses depending on which side of it the truth sits, acceptable only for a risk-neutral buyer averaging over many campaigns. A pays a price that only the single most optimistic supported effect can justify. D leaves money on the table: the whole interval sits well above today's rate.
    +
    +
    +
    + +
    +
    IV · The decision · banked
    +

    The verdict and the recommendation

    +
    The complete answer, assembled from everything measured so far.
    +
    +
    + + + + + + + + + +
    QuantityValueSource
    effect of one exposure€{{nb11.iv_est}}the division δ/π
    90% interval[{{nb11.iv_lo}}, {{nb11.iv_hi}}]classical, and AR agrees
    first-stage F{{nb11.f_stat}}the lottery is strong
    price€{{nb11.cost}}the platform's rate card
    net per exposure€{{nb11.net}}β − c, at the point estimate
    +
    The verdict
    + BUY  Keep buying at €{{nb11.cost}}: the entire defensible range clears the price.
    +
    +
    +
    The recommendation, in three lines
    + 1. Keep buying at the €{{nb11.cost}} rate: even the interval's most pessimistic effect pays.
    + 2. Cap the bid at the interval's lower end, €{{nb11.iv_lo}}: up to there, every effect the data support still clears the price.
    + 3. Measure again only if the rate card climbs toward €{{nb11.iv_lo}}: at today's price, no effect inside the interval changes the action, so more measurement is almost certain to leave the decision unchanged and is worth close to nothing here.
    +
      +
    • Everything above is classical: two averages, one division, one F statistic, one confidence interval.
    • +
    • The one debt on record: exclusion is untestable. The recommendation is conditional on the argued design, and says so.
    • +
    +
    +
    +
    + + +
    +
    Closing · Provenance
    +

    The tools were the product too

    +
    +
      +
    • CausalPy: {{labs.causalpy_methods}} in one open-source package. The IV estimator that closes this session joined later.
    • +
    • Its launch example: individual exposure to a TV campaign {{labs.causalpy_tv}}, yet its causal impact remains a core business need: the sentence this whole session opened with.
    • +
    • pymc-marketing: the MMM library behind Case 2's calibration story; one client's {{labs.bolt_pr}} came back as a pull request (Bolt).
    • +
    • Webinars and content: the consultancy's own webinar agenda is this session's syllabus, {{labs.webinar_agenda}} included.
    • +
    +
    + PyTensor + PyMC + PyMC-Marketing + CausalPy +
    +
    +
    + +
    +
    Closing
    +

    The END

    +
    Causal inference is one shelf of the toolbox: Bayesian inference is a vast set of tools, and PyMC Labs builds with all of it.
    +
      +
    • Beyond today: demand forecasting, pricing, experimentation at scale, customer lifetime value, hierarchical models across markets: the same machinery, aimed at different decisions.
    • +
    • The cases: pymc-labs.com/blog-posts: every number in this session is pinned to a public post (sources on the next slide).
    • +
    +
    +
    +
    Francesco Muia
    +
    PhD in Theoretical Physics, EMBA.
    Consultant for PyMC Labs and Brown University.
    +
    francesco.muia@pymc-labs.com
    +
    francesco.muia@ai-and-analytics-solutions.com
    +
    +
    +
    Alexander Fengler
    +
    PhD in Computational Cognitive Science.
    Postdoc at Brown University, consultant for PyMC Labs.
    +
    alexander.fengler@pymc-labs.com
    +
    +
    Get in touch
    + If a decision in your company leans on a number nobody quite trusts, write to us: those are exactly the problems we like.
    +
    + + +
    +
    Backup
    + Backup · Sources +

    Every number, pinned

    +
    Part 1 facts retrieved and pinned 2026-07-19 (apps/labs_deck_data.json carries the exact quote); Part 2 numbers are baked from the executed course notebooks (nb07/nb07b shards).
    +
    + + + + + +
    SourceFacts pinned
    +
    +
    + +
    +
    Backup · Act II · The counterfactual problem, versus machine learning
    +

    "Why not gradient boosting / Prophet / an LSTM?"

    +
    Is this a forecasting exercise?
    +
    +
    +
    ▶ LIVE A kitchen-sink forecaster (donors + trend + seasonality, fit pre-launch only)
    +
    +
    Solid: treated sales. Dashed blue: the forecaster. Dashed grey: the simplex twin. Both fit only the 40 pre-launch weeks.
    +
    +
      +
    • In-sample it fits tighter:k vs €k per week: flexibility always buys the past.
    • +
    • Yet the estimates are a wash: vs , both near the planted €284k. One dataset cannot rank them.
    • +
    • The sweep decides: across 24 fresh worlds, {{nb07.hull_oos_ratio}}× worse out of sample, with nothing to inspect.
    • +
    +
    +
    +
    Principle 1: a counterfactual, not a forecast
    + A forecaster asks what comes next. We ask what this market would have done in a world that never happened.
    +
      +
    • Principle 2: in-sample fit is the trap, not the goal.
    • +
    • Principle 3: you cannot cross-validate the counterfactual. The truth is never observed.
    • +
    • Principle 4: the simplex ships an auditable claim; a boosted tree is a black box.
    • +
    +
    The forecaster is better at prediction. The simplex is better at the causal job.
    +
    +
    +
    + +
    +
    Backup · Act II · The counterfactual problem
    +

    Interesting limits for the simpler estimators

    +
    Where before/after and treated/control estimators break.
    +
    +
    +
    +
    Before/after · one unit, across time
    + The treated metro's post-average minus its pre-average.
    +
      +
    • The level cancels; nothing subtracts the shared wave.
    • +
    • The whole drift \(\Delta\bar f\) lands on the metro: the largest bias, growing with the horizon.
    • +
    +
    +
    +
    Treated-vs-control · one period, across units
    + The treated metro minus the control average, both in the post window.
    +
      +
    • The level gap \(\alpha_1-\bar\alpha_C\) survives, and sizes differ a lot: it dominates.
    • +
    • This is why raw sales are never compared across cities.
    • +
    +
    +
    +
    The two bias decompositions, term by term
    +
    + \[ \hat\tau^{\text{B/A}} \;=\; \bar Y_{1,\text{post}}-\bar Y_{1,\text{pre}} \;=\; \bar\tau \;+\; \underbrace{\gamma_1^{\top}\,\Delta\bar f}_{\text{bias: the whole drift}} \;+\; \Delta\bar\varepsilon_1 \] + the level \(\alpha_1\) cancels, but the whole shared-factor drift \(\Delta\bar f\) lands at the metro's own loadings: the largest bias, and it grows with the horizon. +
    +
    + \[ \hat\tau^{\text{TC}} \;=\; \bar Y_{1,\text{post}}-\bar Y_{C,\text{post}} \;=\; \bar\tau \;+\; \underbrace{(\alpha_1-\bar\alpha_C)}_{\text{level gap}} \;+\; \underbrace{(\gamma_1-\bar\gamma_C)^{\top}\bar f_{\text{post}}}_{\text{loading gap}\,\times\,\text{level}} \;+\; \Delta\bar\varepsilon \] + no time-differencing, so the level gap survives and dominates; the loading gap now multiplies the factor level, not its drift. +
    +
    +
    +
    + +
    +
    Backup · Act II · The counterfactual problem
    +

    Interesting limits for DiD

    +
    Where the DiD estimator breaks.
    +
    +
    The one question
    + DiD only works if treated and controls would have drifted together. Two dials decide (spread \(s\), macro \(\sigma_\eta\)); the algebra is in the fold.
    +
    Synthetic control solves it
    + DiD's residual bias \(B=(\gamma_1-\bar\gamma_C)^{\top}\Delta\bar f\) comes from its equal weights: it hopes \(s\) is small. Synthetic control chooses the weights, so \(B\to0\) at any spread.
    +
    The DiD bias, decomposed (and why more data cannot shrink it)
    +
    + \[ \hat\tau^{\text{DiD}} \;=\; \bar\tau \;+\; \underbrace{(\gamma_1-\bar\gamma_C)^{\top}\,\Delta\bar f}_{\text{systematic bias }B} \;+\; \underbrace{\Delta\bar\varepsilon_1-\Delta\bar\varepsilon_C}_{\text{idiosyncratic noise}} \] +
    +
      +
    • The pieces: \(\Delta\bar f\) is the post-minus-pre drift of the three factors, \(\bar\gamma_C\) the \(n\) controls' average loadings, \(\Delta\bar\varepsilon\) the change in idiosyncratic noise.
    • +
    • \(B\) is mean-zero over the exposure draw, but your world is one draw, so in it \(B\) is a fixed nonzero number whose typical size is:
    • +
    +
    + \[ \operatorname{Var}(B) \;=\; \frac{s^{2}}{3}\Bigl(1+\frac{1}{n}\Bigr)\Bigl[(\Delta\bar f_{\text{trend}})^{2}+(\Delta\bar f_{\text{season}})^{2}+(\Delta\bar f_{\text{walk}})^{2}\Bigr] \] +
    +
      +
    • \(B\) carries no term for the amount of data: \(\Delta\bar f\) is the world's drift, not sampling error, so more weeks shrink only noise and buy precision around the same biased number.
    • +
    • The noise never leaves: the idiosyncratic term carries \(\sigma_\varepsilon \approx 3\) whatever \(s\) and \(\sigma_\eta\) do, so the honest claim in any limit is "the systematic bias vanishes", never "DiD equals the truth".
    • +
    +
    +
    The two limits, dial by dial: send a knob to zero and read what dies
    +
    +
    Limit 1 · loading spread \(s \to 0\): sufficient on its own
    +
      +
    • The prefactor \(s^2/3\) kills the whole bracket at once.
    • +
    • Every loading \(\to 1\), so \(\gamma_1-\bar\gamma_C\to 0\) deterministically: exact parallel trends, macro walk included.
    • +
    • DiD then recovers \(\bar\tau\) up to noise.
    • +
    +
    Limit 2 · macro shock \(\sigma_\eta \to 0\): not sufficient
    +
      +
    • Only the walk's term \((\Delta\bar f_{\text{walk}})^2\) drops out.
    • +
    • Trend and seasonality still drift between the windows, and markets still weight them differently whenever \(s>0\): \(B \neq 0\).
    • +
    • What it buys: it removes the one stochastic, horizon-growing confounder; the bias that remains is at least deterministic.
    • +
    +
    +
    +
    Where the \(\tfrac{s^2}{3}\bigl(1+\tfrac1n\bigr)\) prefactor comes from
    + Each loading is drawn from \(\mathrm{U}(1{-}s,\,1{+}s)\), a uniform of width \(2s\), whose variance is \((2s)^2/12 = s^2/3\). The treated metro contributes \(\operatorname{Var}(\gamma_1)=s^2/3\); the average of \(n\) independent controls contributes \(\operatorname{Var}(\bar\gamma_C)=s^2/(3n)\). Independent draws add, so \(\operatorname{Var}(\gamma_1-\bar\gamma_C)=\tfrac{s^2}{3}\bigl(1+\tfrac1n\bigr)\) per factor. Each factor's mismatch is then scaled by that factor's post-minus-pre drift, and the three squared drifts sum: the bracket. +
    +
    +
    + +
    +
    Backup · Act III · Is it real, and how big?
    +

    €260k. Real, or a lucky metro? Build the null yourself

    +
    +
    +
    ✋ Poll
    +
    With one treated metro a t-test is not weak, it is undefined. So build the test yourself: rerun the whole pipeline 29 more times, each donor pretended treated ("placebo"): 30 "effects", 29 where nothing ran. Where does our €260k rank?
    +
    + + + +
    + +
    A: rank 1 of 30. The rank is the inference: if the campaign did nothing, P(rank 1 by luck) = 1/30 ≈ 0.033. You just re-derived the permutation test.
    +
    +
    +
    + +
    +
    Backup · Act III · Is it real, and how big?
    +

    Placebo-in-space: measure the luck directly

    +
    Randomisation inference: the t-test had no standard error, so the 29 donors become the null distribution.
    +
    +
    +
    +
    +
    ▶ LIVE Refit the estimator on every donor as if it were treated
    +
    +
    The null: the "effects" the method reports where nothing happened (spread ·, best placebo ·). Green line: our €260k, outside the cloud.
    +
    +
    +
    +
    The null hypothesis, stated
    + \(H_0\): the campaign did nothing. Then the metro is exchangeable with its donors: any of the 30 ranks is equally likely.
    +
    + \[ p \;=\; \frac{1+\#\{\,j:\ |\hat\tau_j|\ge|\hat\tau_1|\,\}}{J+1} \;=\; \frac{1+0}{30}\;\approx\;0.033 \] +
    +
    +
    +
    What the p-value means (and what it is not)
    + The probability, if \(H_0\) were true, of a gap this extreme: a rank, no bell curve assumed. \(p\) sits at its floor \(1/30\), set by the donor count, not the weeks.
    +
    When the rank is valid, and when it lies (Abadie's hygiene rule)
    +
      +
    • Valid under \(H_0\) when the placebos are fair stand-ins: comparable pre-launch fit, and no campaign spillover onto donors (SUTVA), the two ways exchangeability can hold.
    • +
    • Lies when a placebo fits its own pre-period badly: it books fitting failure as a giant fake effect and fattens the tail. Abadie's rule: drop those placebos before ranking.
    • +
    +
    Conclusion: the €260k lift is real, at \(p\approx0.033\)
    + Rank 1 of 30, clear of the cloud: reject \(H_0\). The rank holds (p = 0.033 every time) dropping shaky placebos at 2×, 5×, 20× pre-fit error. But "real" is not "profitable": Act IV.
    +
    +
    + +
    +
    Backup · Act III · Is it real, and how big?
    +

    From a test to an interval: inversion

    +
    Which true lifts could plausibly have produced our €260k? Keep the survivors: no bell curve anywhere.
    +
    +
    +
    +
    +
    ▶ LIVE Drag \(H\), a guess at the truth: the error cloud slides with it; our €260k never moves.
    +
    +
    + + +
    +
    Axis: 20-week total gap (€000). Top: the placebo errors at truth 0. Middle: the same errors slid to \(H\); \(H\) survives if the green line sits inside the shaded middle 90%. Bottom: the survivors, collected: the interval. Drag past an edge to reject.
    +
    +
    +
    +
    The whole idea
    + Ask of every possible true lift: could it plausibly have produced our €260k? Collect the ones that could, and that set of survivors is the interval.
    +
      +
    • ① Measure the error. The 29 placebo "effects" form a cloud around zero, spread about ±€50k.
    • +
    • ② Guess a truth \(H\). If the lift were \(H\), we would see \(H\) plus that same cloud.
    • +
    • ③ Keep or reject. Keep \(H\) if €260k sits inside its middle 90%.
    • +
    • ④ Sweep. The survivors run €195k to €335k: the interval is [€195k, €335k].
    • +
    +
    +
    +
    The inversion, in one line of algebra
    +
    + \[ \underbrace{H + q_{0.05} \;\le\; 260 \;\le\; H + q_{0.95}}_{\text{step ③: €260k sits in \(H\)'s middle \(90\%\)}} + \qquad\Longleftrightarrow\qquad + \underbrace{260 - q_{0.95} \;\le\; H \;\le\; 260 - q_{0.05}}_{\text{step ④: the same line, solved for \(H\)}} \] + \(q_{0.05},q_{0.95}\) are just the low and high edges of the error cloud from step ①, here \(q_{0.05}\!=\!-75\) and \(q_{0.95}\!=\!+65\) (€000). Rearranging the left inequality into the right one is the whole trick, and it hands you the endpoints €195k and €335k directly. +
    +
    +
    Why this interval is the referee for the rest of the lecture
    + It assumed no normality, no independence, no error model, so every model-based interval later (the Bayesian posterior included) has to answer to it. +
    +
    +
    + +
    +
    Backup · Act III · Is it real, and how big?
    +

    Statistics done. Three numbers.

    +
    Questions ① and ② from the boardroom slide are now answered.
    +
    + + + + + + + +
    QuestionAnswerTool that answered it
    Is the effect real?Yes, p = 0.033placebo-in-space permutation: rank 1 of 30
    How big?€260k of incremental salessynthetic-control gap, summed over 20 weeks
    Give or take?[€195k, €335k] at 90%test inversion over the placebo cloud
    +
    Truth check (only a simulation allows it)
    + The planted total €284k sits inside the interval, €24k above the estimate: the machinery works, and its self-reported uncertainty is honest.
    +
    The sentence that loses money
    + "€260k of sales for €75k, a 3.5× return. Roll it out." Every number true; the conclusion does not follow. Question ③ is not a statistics question.
    +
    +
    + +
    +
    Backup · IV · The estimator · hands on
    +

    Why the division is forced

    +
    The division is not a modelling choice. It is the only effect size the two measurements allow.
    +
    +
    +
    ▶ LIVE every candidate effect makes a prediction. One matches.
    +
    +
    + + +
    +
    The rising line is the prediction: an effect of \(\hat\beta\) per exposure implies the lottery should have lifted sales by \(\hat\beta \times \pi\). The flat line is the fact: it lifted them by €{{nb11.reduced}}. Move your guess to the crossing and you have priced the ad.
    +
    +
    +

    Forget the formula and grade any candidate effect \(\hat\beta\) against the two numbers we own:

    +
      +
    • Its prediction: if one exposure were worth \(\hat\beta\), the lottery's {{nb11.first}} extra exposures per win should create \(\hat\beta \times {{nb11.first}}\) euros per win.
    • +
    • The fact: the lottery actually created €{{nb11.reduced}} per win.
    • +
    • The verdict: every candidate except €{{nb11.iv_est}} contradicts a number we measured. The division is the only survivor, not a choice.
    • +
    +
    \[ \hat\beta \times \pi \;\stackrel{!}{=}\; \delta \quad\Longleftrightarrow\quad \hat\beta \;=\; \frac{\delta}{\pi} \]
    +
    Not a black box
    + Every IV estimate is the effect size that makes the instrument's sales bump add up. If you cannot state yours as a ratio of two simple differences, you do not yet understand it.
    +
    +
    +
    + +
    +
    Backup · IV · The estimator · the whole method in one division
    +

    The IV estimate

    +
    Euros per lottery win, divided by exposures per lottery win.
    +
    +
    +
    the whole method, as arithmetic on two measured numbers
    One lottery win buys {{nb11.first}} extra exposures and €{{nb11.reduced}} of extra sales. If each exposure is worth \(\beta\), those two facts only fit together for one \(\beta\): the division.
    @@ -2016,7 +2219,6 @@

    The IV estimate

    What just happened
    We priced the ad using only the random slice of exposure: the dashboard said €{{nb11.naive}}, the lottery says €{{nb11.iv_est}}, against a planted truth of €{{nb11.true}}.
    -

    The method never needed a model of intent, controls, or machine learning: two averages and a division.

    Deep dive · the confidence interval around €{{nb11.iv_est}}
    @@ -2040,37 +2242,8 @@

    The IV estimate

    -
    -
    IV · The estimator · hands on
    -

    Why the division is forced

    -
    The division is not a modelling choice. It is the only effect size the two measurements allow.
    -
    -
    -
    ▶ LIVE every candidate effect makes a prediction. One matches.
    -
    -
    - - -
    -
    The rising line is the prediction: an effect of \(\hat\beta\) per exposure implies the lottery should have lifted sales by \(\hat\beta \times \pi\). The flat line is the fact: it lifted them by €{{nb11.reduced}}. Move your guess to the crossing and you have priced the ad.
    -
    -
    -

    Forget the formula and grade any candidate effect \(\hat\beta\) against the two numbers we own:

    -
      -
    • Its prediction: if one exposure were worth \(\hat\beta\), the lottery's {{nb11.first}} extra exposures per win should create \(\hat\beta \times {{nb11.first}}\) euros per win.
    • -
    • The fact: the lottery actually created €{{nb11.reduced}} per win.
    • -
    • The verdict: every candidate except €{{nb11.iv_est}} contradicts a number we measured. The division is the only survivor, not a choice.
    • -
    -
    \[ \hat\beta \times \pi \;\stackrel{!}{=}\; \delta \quad\Longleftrightarrow\quad \hat\beta \;=\; \frac{\delta}{\pi} \]
    -
    Not a black box
    - Every IV estimate is the effect size that makes the instrument's sales bump add up. If you cannot state yours as a ratio of two simple differences, you do not yet understand it.
    -
    -
    -
    - - -
    -
    IV · When it breaks · the dangerous failure
    +
    +
    Backup · IV · When it breaks · the dangerous failure

    Weak instruments

    A weak instrument is worse than no instrument.
    @@ -2121,174 +2294,8 @@

    Weak instruments

    -
    -
    IV · When it breaks · whose effect it is
    -

    Compliers and the LATE

    -
    Whom does the €{{nb11.iv_est}} describe? Only the users the lottery could move.
    -
    -
    ▶ LIVE the user base, split by how they respond to the lottery
    -
    -
    - - -
    -
    Push \(\gamma\): the complier slice grows, because the complier share is the first stage. Shares are drawn live from one simulated batch, so they can differ from the printed figures by a rounding step.
    -
    -
    -
      -
    • Always-takers ({{nb11.st_share_always}}%): the auction shows them the ad with or without the lottery. It changes nothing for them, so the data say nothing about them.
    • -
    • Never-takers ({{nb11.st_share_never}}%): never see the ad either way. Same silence.
    • -
    • Compliers ({{nb11.st_share_complier}}%): see the ad only because the lottery favoured them. Every euro of the lottery's lift \(\delta\) came from them.
    • -
    • The caveat, monotonicity: we assume no defiers, users who would see the ad only when the lottery does not favour them. A nudge that never repels.
    • -
    -
    -
    -
    Definition · LATE (local average treatment effect)
    - The LATE is the average effect of the ad on the compliers alone, and it is what the division estimates: a local answer, not a statement about every user.
    -
    \[ \hat\beta_{\text{IV}} \;=\; \frac{\delta}{\pi} \;\;\text{ estimates }\;\; \mathbb{E}[\,Y(1)-Y(0)\mid \text{complier}\,] \] - the average of each complier's personal effect, the same \(Y(1)-Y(0)\) contrast the naive slide could not touch
    -
    -
    -
    Why a manager should love this fine print
    - Compliers are the same kind of marginal user a higher bid would newly reach: the closest thing in the data to the customer a bid change buys. That makes €{{nb11.iv_est}} a price for the marginal customer, measured on the margin rather than on the average.
    -
    -
    - -
    -
    IV · When it breaks · what must hold, on one page
    -

    The checklist

    -
    The four assumptions, and which ones the data can check.
    -
    - - - - - - - - -
    AssumptionWhat it saysIn the caseStatus
    Relevance\(Z\) moves \(X\)\(F = {{nb11.f_stat}}\), far above 10TESTABLE, passes
    Exogeneity\(Z \perp U\)the lottery is a genuine random drawBY DESIGN
    Exclusion\(Z \to Y\) only via \(X\)a queue bump shows the user nothingUNTESTABLE
    Monotonicityno defiersa nudge never repelsUNTESTABLE, plausible
    -
    -
    -
    The honest scorecard, for any IV study you are shown
    - One measured number (the first-stage \(F\)), one design guarantee (the randomization), two arguments (exclusion, monotonicity). Ask for all four before you accept the estimate.
    -
      -
    • What we can now defend: an exposure causes about €{{nb11.iv_est}} of sales for the users a bid can actually move.
    • -
    -
    -
    The sentence that loses money
    - "Exposed users are worth €{{nb11.naive}} each, so raise the bid": every word true, conclusion wrong. It books the platform's targeting as advertising.
    -
    -
    -
    - - - -
    -
    IV · The decision · euros at last
    -

    The price map

    -
    The estimate becomes a decision only when it meets the price.
    -
    -
    -
    ▶ LIVE the verdict as the price moves
    -
    -
    - - -
    -
    Blue: net value per exposure at each price. The orange band is the 90% interval, the zone where the data refuse to commit. Drag the price through the three zones and watch the verdict flip.
    -
    -
    -

    One estimate gives not one answer but a map from any price to a verdict:

    - - - - - - - -
    Price zoneVerdictWhy
    below €{{nb11.iv_lo}}GOeven the most pessimistic supported effect pays
    €{{nb11.iv_lo}} to €{{nb11.iv_hi}}TESTthe data straddle the price: negotiate, or measure more
    above €{{nb11.iv_hi}}NO-GOno supported effect pays
    -
      -
    • Today's rate, €{{nb11.cost}}, sits in the GO zone, below the whole interval.
    • -
    • The net, computed: \(Y\) is contribution euros, so one exposure nets \(\hat\beta_{\text{IV}} - c = \) €{{nb11.iv_est}} − €{{nb11.cost}} = €{{nb11.net}}. Even read at the interval floor, €{{nb11.iv_lo}} against €{{nb11.cost}}, the exposure still pays.
    • -
    -
    Why boards like this framing
    - "Is the effect significant?" has no business answer. "Up to what price is this a buy?" has one, and it is the same question a bid cap asks.
    -
    -
    -
    - -
    -
    IV · The decision
    -

    Poll · the negotiation

    -
    -
    -
    ✋ Poll
    -
    The platform wants to renegotiate the rate. Your analyst hands you the causal read: effect €{{nb11.iv_est}} per exposure, 90% interval [{{nb11.iv_lo}}, {{nb11.iv_hi}}]. What is the highest rate at which you would still sign "buy" without further study?
    -
    - - - - -
    - -
    C. Below €{{nb11.iv_lo}}, every effect the data support pays: the interval's floor is a no-regret bid cap, defensible whichever value inside the interval turns out to be the truth. B is a break-even gamble: paying the point estimate wins or loses depending on which side of it the truth sits, acceptable only for a risk-neutral buyer averaging over many campaigns. A pays a price that only the single most optimistic supported effect can justify. D leaves money on the table: the whole interval sits well above today's rate.
    -
    -
    -
    - -
    -
    IV · The decision · banked
    -

    The verdict and the recommendation

    -
    The complete answer, assembled from everything measured so far.
    -
    -
    - - - - - - - - - -
    QuantityValueSource
    effect of one exposure€{{nb11.iv_est}}the division δ/π
    90% interval[{{nb11.iv_lo}}, {{nb11.iv_hi}}]classical, and AR agrees
    first-stage F{{nb11.f_stat}}the lottery is strong
    price€{{nb11.cost}}the platform's rate card
    net per exposure€{{nb11.net}}β − c, at the point estimate
    -
    The verdict
    - BUY  Keep buying at €{{nb11.cost}}: the entire defensible range clears the price.
    -
    -
    -
    The recommendation, in three lines
    - 1. Keep buying at the €{{nb11.cost}} rate: even the interval's most pessimistic effect pays.
    - 2. Cap the bid at the interval's lower end, €{{nb11.iv_lo}}: up to there, every effect the data support still clears the price.
    - 3. Measure again only if the rate card climbs toward €{{nb11.iv_lo}}: at today's price, no effect inside the interval changes the action, so more measurement is almost certain to leave the decision unchanged and is worth close to nothing here.
    -
      -
    • Everything above is classical: two averages, one division, one F statistic, one confidence interval.
    • -
    • The one debt on record: exclusion is untestable. The recommendation is conditional on the argued design, and says so.
    • -
    -
    -
    -
    - - -
    -
    Closing · Provenance
    -

    The tools were the product too

    -
    -
      -
    • CausalPy: {{labs.causalpy_methods}} in one open-source package. The IV estimator that closes this session joined later.
    • -
    • Its launch example: individual exposure to a TV campaign {{labs.causalpy_tv}}, yet its causal impact remains a core business need: the sentence this whole session opened with.
    • -
    • pymc-marketing: the MMM library behind Case 2's calibration story; one client's {{labs.bolt_pr}} came back as a pull request (Bolt).
    • -
    • Webinars and content: the consultancy's own webinar agenda is this session's syllabus, {{labs.webinar_agenda}} included.
    • -
    -
    - CausalPy - PyMC-Marketing -
    -
    -
    - -
    -
    Closing
    +
    +
    Backup · Closing

    The pattern in every engagement

      @@ -2308,35 +2315,6 @@

      The pattern in every engagement

    -
    -
    Closing
    -

    One breath

    -
    -
    The pattern to take home
    - The toolkit a Bayesian consultancy sells: counterfactuals, calibrated by experiments, priced as probabilities.
    -
      -
    • Read the cases: pymc-labs.com/blog-posts: every number in this deck is pinned to a post, listed on the next slide.
    • -
    • Say hello: both authors consult for PyMC Labs; the notebooks behind this session are the course repository.
    • -
    -
    -
    - - -
    -
    Backup
    - Backup · Sources -

    Every number, pinned

    -
    Part 1 facts retrieved and pinned 2026-07-19 (apps/labs_deck_data.json carries the exact quote); Part 2 numbers are baked from the executed course notebooks (nb07/nb07b shards).
    -
    - - - - - -
    SourceFacts pinned
    -
    -
    -
    @@ -3097,16 +3075,15 @@

    Every number, pinned

    const e=document.getElementById(id); if(e)e.textContent=v;}); const tr=DATA.treated,y0=DATA.y0_true,scl=DATA.synth_cl,ols=DATA.ols_synth,L=DATA.launch,W=tr.length; function draw(){clr(svg);const c=COL();const Wp=720,H=235,mL=46,mR=14,mT=32,mB=26; - let mn=1e9,mx=-1e9;[tr,y0,scl,ols].forEach(a=>a.forEach(v=>{if(vmx)mx=v;})); + let mn=1e9,mx=-1e9;[tr,y0,scl].forEach(a=>a.forEach(v=>{if(vmx)mx=v;})); const x=lin(0,W-1,mL,Wp-mR),y=lin(mn-2,mx+2,H-mB,mT); svg.appendChild(el('line',{x1:x(L),x2:x(L),y1:mT,y2:H-mB,stroke:c.orange,'stroke-width':1.3})); svg.appendChild(el('text',{x:x(L)+4,y:H-mB-6,fill:c.orange,'font-size':10},'launch')); svg.appendChild(el('path',{d:path(tr.map((v,i)=>[x(i),y(v)])),fill:'none',stroke:c.ink,'stroke-width':1.6,opacity:.5})); svg.appendChild(el('path',{d:path(y0.slice(L-1).map((v,i)=>[x(L-1+i),y(v)])),fill:'none',stroke:c.ink,'stroke-width':2,'stroke-dasharray':'6 4'})); svg.appendChild(el('path',{d:path(scl.map((v,i)=>[x(i),y(v)])),fill:'none',stroke:c.blue,'stroke-width':1.9})); - svg.appendChild(el('path',{d:path(ols.map((v,i)=>[x(i),y(v)])),fill:'none',stroke:c.red,'stroke-width':1.9})); const lg=[[c.ink,'treated (observed)',1.6,'none',.5],[c.ink,'true Y(0), post-launch',2,'6 4',1], - [c.blue,'simplex synthetic',1.9,'none',1],[c.red,'OLS synthetic',1.9,'none',1]]; + [c.blue,'simplex synthetic',1.9,'none',1]]; lg.forEach(([col,lab,wd,dash,op],i)=>{const xx=mL+8+i*168; svg.appendChild(el('line',{x1:xx,x2:xx+24,y1:10,y2:10,stroke:col,'stroke-width':wd,'stroke-dasharray':dash,opacity:op})); svg.appendChild(el('text',{x:xx+29,y:14,fill:col,'font-size':10.5,opacity:Math.max(op,.8)},lab));}); @@ -3428,35 +3405,81 @@

    Every number, pinned

    draw(); window.__redraw.push(draw); })(); -/* ==== fig: the MMM / experiment / synthetic-control loop (S7) ==== */ +/* ==== fig: the MMM / experiment calibration loop (S7) ==== */ (function(){ const svg=document.getElementById('svgLoop'); if(!svg)return; function draw(){ clr(svg); const c=COL(); - const nodes=[ - {x:170,y:52,w:150,label:'MMM',sub:'always-on model'}, - {x:88,y:212,w:150,label:'Geo experiment',sub:'episodic truth'}, - {x:252,y:212,w:150,label:'Synthetic control',sub:'reads it out'}, - ]; - function box(n,col){ - svg.appendChild(el('rect',{x:n.x-n.w/2,y:n.y-26,width:n.w,height:52,rx:10,fill:'none',stroke:col,'stroke-width':2})); - svg.appendChild(el('text',{x:n.x,y:n.y-5,'text-anchor':'middle','font-size':13,'font-weight':'700',fill:c.navy},n.label)); - svg.appendChild(el('text',{x:n.x,y:n.y+13,'text-anchor':'middle','font-size':10,fill:c.muted},n.sub)); + function box(x,y,w,label,sub,col){ + svg.appendChild(el('rect',{x:x-w/2,y:y-26,width:w,height:52,rx:10,fill:'none',stroke:col,'stroke-width':2})); + svg.appendChild(el('text',{x:x,y:y-5,'text-anchor':'middle','font-size':13,'font-weight':'700',fill:c.navy},label)); + svg.appendChild(el('text',{x:x,y:y+13,'text-anchor':'middle','font-size':10,fill:c.muted},sub)); } - box(nodes[0],c.blue); box(nodes[1],c.orange); box(nodes[2],c.green); + box(170,60,190,'MMM','always-on budget model',c.blue); + box(170,220,190,'Geo experiment','episodic ground truth',c.orange); function arrow(x1,y1,x2,y2){ svg.appendChild(el('line',{x1,y1,x2,y2,stroke:c.faint,'stroke-width':1.6})); const a=Math.atan2(y2-y1,x2-x1); svg.appendChild(el('path',{d:`M${x2} ${y2} L${x2-9*Math.cos(a-0.4)} ${y2-9*Math.sin(a-0.4)} L${x2-9*Math.cos(a+0.4)} ${y2-9*Math.sin(a+0.4)} Z`,fill:c.faint})); } - arrow(128,80,100,182); // MMM -> experiment (asks) - arrow(120,182,148,80); // experiment -> MMM (calibrates) - arrow(166,224,176,224); // experiment -> SC - svg.appendChild(el('text',{x:56,y:132,'font-size':10,fill:c.muted},'asks for')); - svg.appendChild(el('text',{x:56,y:144,'font-size':10,fill:c.muted},'ground truth')); - svg.appendChild(el('text',{x:152,y:132,'font-size':10,fill:c.muted},'calibrates')); - svg.appendChild(el('text',{x:152,y:144,'font-size':10,fill:c.muted},'(priors, lift tests)')); - svg.appendChild(el('text',{x:170,y:262,'text-anchor':'middle','font-size':10,fill:c.muted},'MMM: last slide · synthetic control: Part 2 · the loop: this session')); + arrow(120,86,120,194); + arrow(220,194,220,86); + svg.appendChild(el('text',{x:108,y:140,'font-size':10,fill:c.muted,'text-anchor':'end'},'asks for')); + svg.appendChild(el('text',{x:108,y:152,'font-size':10,fill:c.muted,'text-anchor':'end'},'ground truth')); + svg.appendChild(el('text',{x:232,y:140,'font-size':10,fill:c.muted},'calibrates')); + svg.appendChild(el('text',{x:232,y:152,'font-size':10,fill:c.muted},'(priors, lift tests)')); + svg.appendChild(el('text',{x:170,y:282,'text-anchor':'middle','font-size':10,fill:c.muted},'reading a geo experiment out is its own craft:')); + svg.appendChild(el('text',{x:170,y:294,'text-anchor':'middle','font-size':10,fill:c.muted},'that method is Part 2 of this session')); + } + draw(); window.__redraw.push(draw); +})(); + +/* ==== fig: the PyMC ecosystem stack (slide 2) ==== */ +(function(){ + const svg=document.getElementById('svgEco'); if(!svg)return; + function draw(){ + clr(svg); const c=COL(); + function node(x,y,w,label,sub,col,dash){ + const attrs={x:x-w/2,y:y-19,width:w,height:38,rx:9,fill:'none',stroke:col,'stroke-width':2}; + if(dash)attrs['stroke-dasharray']='5 4'; + svg.appendChild(el('rect',attrs)); + svg.appendChild(el('text',{x:x,y:y-1,'text-anchor':'middle','font-size':12,'font-weight':'700',fill:c.navy},label)); + if(sub)svg.appendChild(el('text',{x:x,y:y+13,'text-anchor':'middle','font-size':9,fill:c.muted},sub)); + } + function arrow(x1,y1,x2,y2){ + svg.appendChild(el('line',{x1,y1,x2,y2,stroke:c.faint,'stroke-width':1.6})); + const a=Math.atan2(y2-y1,x2-x1); + svg.appendChild(el('path',{d:`M${x2} ${y2} L${x2-8*Math.cos(a-0.4)} ${y2-8*Math.sin(a-0.4)} L${x2-8*Math.cos(a+0.4)} ${y2-8*Math.sin(a+0.4)} Z`,fill:c.faint})); + } + node(95,66,120,'PyTensor','compute engine',c.grey); + node(285,66,155,'PyMC','probabilistic programming',c.blue); + node(540,22,170,'PyMC-Marketing','MMM, CLV',c.green); + node(540,66,170,'CausalPy','quasi-experiments',c.orange); + node(540,110,170,'Bespoke','customer-specific builds',c.gold,true); + arrow(157,66,205,66); + arrow(365,60,453,26); + arrow(365,66,453,66); + arrow(365,72,453,106); + } + draw(); window.__redraw.push(draw); +})(); + +/* ==== fig: the metros, one treated (boardroom slide) ==== */ +(function(){ + const svg=document.getElementById('svgMetros'); if(!svg)return; + function draw(){ + clr(svg); const c=COL(); + const donors=[[60,40,7],[80,120,9],[130,30,6],[210,125,7],[220,55,10],[255,95,6],[290,35,8], + [320,120,9],[350,70,7],[385,30,6],[400,105,8],[430,60,11],[465,115,6],[480,35,7],[510,85,9], + [545,40,6],[560,120,8],[590,70,7],[620,105,6],[640,35,9],[665,80,7],[75,70,5],[170,45,5], + [240,20,5],[365,115,5],[450,20,5],[530,115,5],[610,20,5],[680,120,5]]; + donors.forEach(([x,y,r])=>svg.appendChild(el('circle',{cx:x,cy:y,r:r,fill:c.grey,opacity:.28,stroke:c.grey,'stroke-width':1}))); + svg.appendChild(el('circle',{cx:150,cy:70,r:14,fill:c.orange,opacity:.3,stroke:c.orange,'stroke-width':2.2})); + svg.appendChild(el('path',{d:'M 166 52 A 24 24 0 0 1 174 70',fill:'none',stroke:c.orange,'stroke-width':1.6})); + svg.appendChild(el('path',{d:'M 170 45 A 32 32 0 0 1 181 70',fill:'none',stroke:c.orange,'stroke-width':1.2,opacity:.7})); + svg.appendChild(el('text',{x:150,y:105,'text-anchor':'middle','font-size':10.5,fill:c.orange,'font-weight':700},'the treated metro')); + svg.appendChild(el('text',{x:150,y:118,'text-anchor':'middle','font-size':9.5,fill:c.orange},'the €75k campaign, weeks 40-59')); + svg.appendChild(el('text',{x:545,y:16,'text-anchor':'middle','font-size':10.5,fill:c.muted,'font-weight':700},'29 donor markets · no campaign')); } draw(); window.__redraw.push(draw); })(); diff --git a/causal-marketing-pymc/scratchpad/assemble_labs_src.py b/causal-marketing-pymc/scratchpad/assemble_labs_src.py new file mode 100644 index 0000000..ac5a0da --- /dev/null +++ b/causal-marketing-pymc/scratchpad/assemble_labs_src.py @@ -0,0 +1,80 @@ +"""One-time assembler for apps/labs_slides_src.html (the business-applications closer deck). + +Splices the SHARED CHROME out of apps/iv_slides_src.html (CSS, MathJax config, footer/nav/TOC +DOM, svg helpers, poll wiring, deck-navigation IIFE, the two base64 logo chips) and combines it +with the labs deck's own slides (scratchpad/labs_slides_fragment.html) and figures +(scratchpad/labs_figures_fragment.js). + +After this runs once, apps/labs_slides_src.html is the CANONICAL editable source; edit it +directly and rebuild with `make html-labs`. Re-running this assembler OVERWRITES it from the +fragments, so only re-run if you deliberately want to re-derive the chrome. +""" +from __future__ import annotations + +import re +import sys +from pathlib import Path + +HERE = Path(__file__).resolve().parent +APPS = HERE.parent / "apps" +IV = (APPS / "iv_slides_src.html").read_text() +SLIDES = (HERE / "labs_slides_fragment.html").read_text() +FIGURES = (HERE / "labs_figures_fragment.js").read_text() +OUT = APPS / "labs_slides_src.html" + + +def cut(s: str, start: str, end: str, *, incl_start=True, incl_end=False) -> str: + i = s.index(start) + j = s.index(end, i) + a = i if incl_start else i + len(start) + b = j + len(end) if incl_end else j + return s[a:b] + + +# 1 · head + CSS + MathJax config, up to (excluding) the first slide section +head = IV[: IV.index("")] +head = head.replace( + "Instrumental Variables: what is one ad exposure worth? (slides)", + "Causal Inference in the Wild: PyMC Labs cases (slides)", +) + +# 2 · the two logo chips (giant base64 lines on the IV title slide), spliced verbatim +logo_lines = [ln.strip() for ln in IV.splitlines() if '' in ln] +if len(logo_lines) != 2: + sys.exit(f"FAIL: expected 2 logo-chip lines in iv_slides_src.html, found {len(logo_lines)}") +slides = SLIDES.replace("__PYMC_LOGO__", " " + logo_lines[0]).replace( + "__LABS_LOGO__", " " + logo_lines[1] +) +if "__" in re.sub(r"", "", slides.replace("__SOURCES", "")): # loose guard + pass # data-* attrs etc. are fine; real placeholder misses are caught by the build + +# 3 · chrome DOM after SLIDES-END: deck close, navzones, footer, pbar, tocOv, script open, +# DATA marker line (stop before the IV N-map line) +mid = cut(IV, " +
    +
    Causal inference for marketing · SDA Bocconi · closing act
    +

    Causal Inference in the Wild

    +
    Drawing from real PyMC Labs client engagements.
    +
    +
    Francesco Muia
    PhD in Theoretical Physics, EMBA. Consultant for PyMC Labs and Brown University.
    francesco.muia@pymc-labs.com
    +
    Alexander Fengler
    PhD in Statistics. Postdoc at Brown University and consultant for PyMC Labs.
    alexander.fengler@pymc-labs.com
    +
    +
    +__PYMC_LOGO__ +__LABS_LOGO__ +
    +
    + + +
    +
    Case 1 · Colgate-Palmolive
    +

    You are the consultant

    +
    The counterfactuals you built today are what clients buy. First test: route a real call.
    +
    +
    +
    ✋ Poll
    +
    Colgate-Palmolive calls: "Our new toothpaste launched nationally last quarter. No holdout, no test market. Is it stealing share from competitors, or from our own brands?" Which tool from today do you reach for first?
    +
    + + + + +
    + +
    C. A national launch leaves nothing to randomize and no market to difference against: A and B need a control group that does not exist, and D's instrument does not exist for a shelf that changed everywhere at once. What remains is this morning's move, run in time instead of space: fit the world before the launch, project it forward, read the gap. PyMC Labs sold exactly that projection; the next slide shows it.
    +
    +
    +
    + +
    +
    Case 1 · Colgate-Palmolive
    +

    Colgate-Palmolive: incremental, or cannibalistic?

    +
    Incremental: sales won from competitors or category growth. Cannibalistic: sales taken from your own products. The launch verdict is the split.
    +
    +
    +
    +
    The launch, and the world without it schematic
    +
    +
    Illustrative shape of the engagement's counterfactual read, not client data: fit the pre-period, project it forward, price the gap.
    +
    +
    +
    +
    The brief, in their words
    + "We need to estimate the counterfactual sales of all products would have been if the new product had not been introduced."
    +
      +
    • The client: Colgate-Palmolive came to PyMC Labs in {{labs.colgate_year}}, in a market estimated at {{labs.market}}.
    • +
    • The method: a multivariate Bayesian interrupted time series: nb10's machinery pointed at a product instead of a market, later extended to a nested-logit choice model.
    • +
    • The grading: on simulated data the model recovers a planted {{labs.colgate_truth}} incrementality as a {{labs.colgate_ci_level}} interval of {{labs.colgate_ci}}: the recover-the-truth contract you saw all day, run commercially.
    • +
    +
    +
    +
    + +
    +
    Case 1 · Colgate-Palmolive · open floor
    +

    What would break it?

    +
    +
    +
    🗣 Open floor · 2 minutes
    +
    You are Colgate's CMO. The incrementality estimate you just saw (the share of the new product's sales that are genuinely new, not cannibalized) decides the launch review. Name one real-world event that would make it wrong.
    + +
    A second launch. When another product entered the estimation window, the same machinery reported {{labs.fail_range}} incrementality against a planted truth of {{labs.fail_truth}}: the counterfactual absorbed part of the very effect it was meant to isolate. An honest consultancy publishes exactly this: the {{labs.fail_range}} miss is printed in the same post as the win. This morning's version of the disease was spillover; the defence is design, not statistics.
    +
    +
      +
    • What the model learns: everything before the launch defines "normal growth", and the projection (red) extrapolates that normal forward.
    • +
    • What a second launch does: inside the window it becomes part of "normal", so the projection rises too fast and under-credits the true lift; after the launch it inflates the observed line instead, and over-credits.
    • +
    +
    +
    Break it yourself: slide a second launch into the window schematic
    +
    +
    + + + +
    +
    Same schematic world as the previous slide. The real case reported {{labs.fail_range}} against a truth of {{labs.fail_truth}}.
    +
    +
    +
    + + +
    +
    Case 2 · HelloFresh · the tool, and its failure mode
    +

    Why calibrate? A model alone can rank channels backwards

    +
    Before the HelloFresh story, the tool it relies on: a warm-up from PyMC Labs' published calibration tutorial.
    +
    +
      +
    • The tool: a marketing-mix model (MMM) explains total sales as the sum of per-channel contributions, fit on observational spend data: no experiment anywhere in it.
    • +
    • The grading: the tutorial plants a truth: return on ad spend (ROAS, sales per unit of spend) of {{labs.roas_x1}} for channel x1 against {{labs.roas_x2}} for x2, so x2 is {{labs.roas_gap_words}}.
    • +
    • The experiment: a lift test nudges one channel's spend by a known amount and measures the sales change it causes: a small randomized ground-truth reading for that channel.
    • +
    • The repair: {{labs.lift_tests_n}} per channel, entered into the likelihood, recover both values: the experiment is the model's anchor, exactly Ch 9's role.
    • +
    +
    +
    One MMM on observational spend alone, one planted truth, one inversion baked from the tutorial
    +
    +
    +
    Left: the ranking the uncalibrated model reported. Right: the planted truth the experiments recover.
    +
    +
    The inversion
    + Fit on observational data alone, the baseline model ranked x1 above x2: {{labs.roas_wrong_ranking}}.
    +
    +
    + +
    +
    Case 2 · HelloFresh
    +

    HelloFresh runs the loop, at industrial scale

    +
    +
    +
      +
    • The loop: MMM priors fed by field experiments such as {{labs.hf_priors_experiments}}; a {{labs.hf_var}} cut in prediction variance.
    • +
    • On stage: the panel's own agenda: Bayesian MMM can be {{labs.hf_panel_calibration}}.
    • +
    • The experiment supply: a pipeline handling {{labs.hf_thousands}}: {{labs.hf_test_types}} campaigns run simultaneously, the overnight batch down from {{labs.hf_batch}}; the Criteo experiment you saw this afternoon ({{labs.criteo_rows}} users) sits in exactly this regime.
    • +
    +
    The supply chain, for Ch 13
    + The experiments a company already runs are its instrument supply: a randomized encouragement is the instrument for the exposure you cannot randomize.
    +
    +
    +
    +
    The loop you learned today
    +
    +
    The model runs always-on; the experiment disciplines it; the counterfactual reads the experiment out.
    +
    +
    +
    +
    + + +
    +
    Case 3 · Nürnberger Versicherung
    +

    Price the engagement

    +
    A German insurer, last-touch attribution, and a funnel-aware causal MMM: nb04's mediation chain, in production.
    +
    +
    +
    ✋ Poll
    +
    Nürnberger Versicherung replaced last-touch attribution steering with a funnel-aware causal MMM. Over {{labs.cpl_window}} of model-guided spend, cost per lead (CPL) moved by how much?
    +
    + + + + +
    + +
    C. "This year we were able to drive the CPL down by {{labs.cpl}}, which is very, very good" (Philip Herp, Nürnberger Versicherung). The mechanism is the lesson: under GDPR, {{labs.gdpr_sentence}}, so last-touch under-credited the upper funnel and budget followed {{labs.herp_attribution_quote}}. The funnel model measured what video spend causes downstream, and the client is scaling it into {{labs.nurn_2026}}.
    +
    +
    The client's bar for belief
    + "{{labs.trust_quote}}"
    +
    +
    + + +
    +
    Closing · Provenance
    +

    The tools were the product too

    +
    +
      +
    • CausalPy: {{labs.causalpy_methods}} in one open-source package: this morning's method and its quasi-experimental family, industrialized by PyMC Labs. The IV estimator you ran this afternoon joined the package later.
    • +
    • Its launch example: individual exposure to a TV campaign {{labs.causalpy_tv}}, yet its causal impact remains a core business need: the sentence both of today's lectures opened with.
    • +
    • pymc-marketing: nb06's MMM library; one client's {{labs.bolt_pr}} came back as a pull request (Bolt).
    • +
    • Webinars and content: the consultancy's own webinar walks geo-experimentation, MMM, synthetic control, difference-in-differences, regression discontinuity and {{labs.webinar_agenda}}: the syllabus you just finished.
    • +
    +
    + CausalPy + PyMC-Marketing +
    +
    +
    + +
    +
    Closing
    +

    The pattern in every engagement

    +
    +
      +
    • The deliverable is a counterfactual: a world minus the launch, the campaign, the exposure: priced in euros.
    • +
    • An experiment anchors every observational model: calibration is the product, not a luxury.
    • +
    • Uncertainty prices the decision: boards act on P(pays) and headroom, not on a point estimate.
    • +
    + + + + + + +
    Agent, on adversarial MMM dataResult
    Vanilla coding agentFit a model, recommended budget reallocations. {{labs.dl_vanilla}}
    PyMC Labs' Decision Lab{{labs.dl_explored}} Returned: "{{labs.dl_verdict}}"
    +
    Even the machines know the punchline
    + The honest system's best answer was this morning's closing advice: run the experiment.
    +
    +
    + +
    +
    Closing
    +

    One breath

    +
    +
    What you now hold
    + The toolkit a Bayesian consultancy sells: counterfactuals, calibrated by experiments, priced as probabilities.
    +
      +
    • Read the cases: pymc-labs.com/blog-posts: every number in this deck is pinned to a post, listed on the next slide.
    • +
    • Say hello: both authors consult for PyMC Labs; the notebooks behind all three decks are the course repository.
    • +
    +
    +
    + + +
    +
    Backup
    + Backup · Sources +

    Every number, pinned

    +
    Facts retrieved and pinned 2026-07-19; apps/labs_deck_data.json carries the exact quote for each.
    +
    + + + + + +
    SourceFacts pinned
    +
    +