A media-mix model is supposed to tell you what each of your marketing channels is worth. It does so by fitting a statistical model to historical spend, GRPs, impressions, and revenue. The math is straightforward. The interpretation is not. Most commercial MMM — including the kind sold by name-brand vendors — commits at least one of three causal errors that will systematically mis-attribute lift, in directions that are convenient for the channel owners and inconvenient for the CFO.
This dispatch walks through the three traps, what they look like in your data, and what to do about them.
Trap one: correlated spend
Most large advertisers spend more on every channel at the same time. Q4 brings Black Friday, holiday gifting, and end-of-fiscal budget flushes. Spend on TV, paid search, social, and display all move together. When you fit a regression of revenue against spend across all those channels, the coefficients on each individual channel are not identified. The model literally cannot tell which dollar caused which outcome, because the dollars all moved at once.
What the model produces in this situation is not zero. It produces numbers. The numbers are essentially noise — or, worse, they are systematically biased toward whichever channel happens to have a slightly different temporal pattern. Vendors will report those numbers as “your TV ROAS is 2.4x and your paid search ROAS is 1.8x.” The standard errors on those point estimates, if anyone bothered to compute them, frequently span zero.
How to detect it
Check the variance inflation factors on your spend regressors. If any VIF is above 10, your channel coefficients are not identified in a meaningful sense. (For context: a VIF of 10 corresponds to an R-squared of 0.9 between that channel and the other channels. Most CPG MMMs have VIFs in the 15–40 range.)
What to do
Ridge regularization helps stabilize the estimates but does not fix the identification problem; it just smooths the noise. The real fix is to introduce exogenous variation — randomized holdouts, geographic experiments, or the natural experiments created by media-platform glitches and policy changes. Geo-experiments are the workhorse here. If you can hold out paid search in three matched DMAs for four weeks, you get an estimate of paid search lift that does not depend on disentangling correlated spend.
The most common MMM failure mode is not bad math. It is good math fit to data that cannot answer the question. First law of media-mix modeling
Trap two: omitted confounders
Revenue is caused by marketing spend. It is also caused by everything else — pricing, distribution, weather, macro, competitor moves, organic search trends, product launches, and a hundred other things that move week to week. If any of those omitted confounders is correlated with your marketing spend (and many of them are), the marketing coefficients in your MMM will absorb their effects.
A concrete example: a beverage client’s MMM reported a TV ROAS of 3.1x. Six months later we re-fit the model with retailer-level distribution as a covariate. TV ROAS dropped to 1.7x. The difference was real: the brand had been expanding distribution into new convenience accounts during the same windows when TV pressure peaked. The original MMM was attributing the distribution-driven volume to TV.
This is not a vendor failing in a moral sense. It is a methodological failing. Most MMM vendors do not have visibility into the retailer-level distribution data, the syndicated pricing data, or the competitive context that would let them properly control for the omitted variables. They fit the model with what they have. The fact that the model returns a number does not mean the number is correct.
How to detect it
The honest signal is the R-squared. If your MMM reports R-squared above 0.95 on weekly data with 5–7 spend channels, the model is almost certainly overfit and the channel coefficients are absorbing variance from things that are not in the model. A well-specified MMM on commercial data has an R-squared in the 0.65–0.85 range. Anything higher should make you suspicious.
Trap three: the saturation curve identification problem
Marketing has diminishing returns. Spending a second million dollars on TV does not produce twice the impressions or twice the lift of the first million. MMM accounts for this by fitting nonlinear “saturation curves” — Hill functions, S-curves, log-transforms — to the spend variables.
The problem: the curvature parameter of those saturation curves is identified almost entirely by the variance in spend across the observation window. If your TV spend ranges from $4M to $6M per quarter across three years of data, the model can only estimate the shape of the response curve in that range. Anything the curve says about what happens at $2M or $10M is functionally an extrapolation, and the extrapolations are highly sensitive to the parametric form you chose.
This matters because the curve is exactly what gets used to make budget allocation recommendations. The recommendation “reallocate $3M from TV to paid search” depends on what the model thinks happens to TV at lower spend levels — and the model knows essentially nothing about lower spend levels if you never spent that little.
What to do instead
Three structural shifts produce more defensible MMM.
- Treat the model as an estimator, not an oracle. Every channel coefficient gets a credible interval. If the interval crosses zero, you do not have a point estimate; you have an uncertain answer. Budget recommendations should be made against the full posterior, not the point estimate.
- Use Bayesian priors informed by experiments. If you ran a geo-experiment on paid search and estimated a CAC of $42 with a 90% interval of $38–$48, that interval becomes the prior on the paid search coefficient in the MMM. The model is then constrained to be consistent with the experimental evidence. This is the single biggest improvement most MMM workflows can make.
- Validate against holdouts. Set aside the last 4–8 weeks of data when you fit. Predict that window. If the prediction error is larger than the credible interval predicted, the model is overfit, and you should either simplify or get more data.
The combination of those three — uncertainty quantification, experimental priors, and holdout validation — produces MMM that is harder to oversell, slower to build, and substantially more useful when the pricing committee starts asking pointed questions about why the model says what it says.
The vendor incentive problem
The structural reason most MMM is broken is not statistical. It is commercial. The vendor selling your MMM has a strong incentive to report numbers that are confident, interpretable, and actionable. Reporting wide credible intervals, refusing to recommend reallocation in regions of the spend space that are unobserved, and validating against holdouts all reduce the apparent confidence and actionability of the output. The vendor who sells you a tighter, more confident-looking model wins the next renewal.
None of that is the vendor’s fault — they are responding to what buyers reward. But it does mean the buyer needs to be the one asking for credible intervals, asking for the holdout validation, and asking which coefficients are identified by experimental variance vs. observational correlation. If the vendor cannot answer those questions, the MMM should not be making budget recommendations.
The honest version
Done right, MMM is one of the more useful things a commercial team can have. It quantifies what is otherwise an argument; it forces channel owners to defend their reach claims against hard data; it converts brand spend into a number the CFO can plan against. It does all of that when the math is honest. When it is not, it produces confidence-shaped output that the organization mistakes for evidence, and the budget moves in directions that the data did not actually support.
The next time a vendor reports your paid search ROAS as 2.1x, ask three questions: what is the 90% credible interval, what omitted variables are not in the model, and which channels have experimental evidence supporting their coefficients. If the answers are “we don’t report intervals,” “we don’t know,” and “none of them,” the MMM is probably lying to you.