The client’s pricing committee had read the slides three times. The slides said the right things. They argued for moderate, targeted SRP reductions on a defined set of priority SKUs. The slides did not say what those reductions should be, or which SKUs, or what would happen if you actually shipped them. So nothing shipped. For eighteen months.

The brief that landed on our desk in March 2026 was a single sentence: tell us what to actually do with our prices, and prove it. Eight weeks later we delivered a defensible 14-SKU price change. Q3 revenue came in $1.84M above the counterfactual. The 90% credible interval was $1.43M to $2.27M. The pricing committee shipped without amendment.

This is what happened in between.

The brief

The client is a top-five beverage operator in the US. The category is mature, the margins are thin, and the SRP gap to private label had been widening for six quarters. The previous engagement — a large strategy consultancy whose name you can guess — had delivered what they called a “pricing intelligence transformation roadmap.” The deliverable was a 184-slide deck recommending a center-of-excellence, a vendor RFP, and a three-phase rollout culminating eighteen months in the future.

None of those slides answered the actual question the operator was trying to answer, which was: if we drop SRP on Brand X by some amount, what happens to volume, revenue, and margin, with what confidence, on what timescale?

That question is a quantitative question. It has a numerical answer. The previous consultancy treated it as an organizational question. That was the first thing we noticed and the first thing we said out loud.

“If we drop SRP on Brand X by some amount, what happens to volume, revenue, and margin, with what confidence, on what timescale?” The actual question. Not in the 184-slide deck.

The math

The model is not exotic. It is a hierarchical log-log demand system with cross-price terms, fit on 36 months of weekly scanner data across 14 priority SKUs and their primary competitive set. The hierarchy lets us pool information across SKUs in the same category — which matters because three of the 14 SKUs had under 18 months of clean post-relaunch history.

ln Qjt = αj + βj ln Pjt + γj ln Pjtc + εjt The own-price elasticity (β) is the term that does the work. Cross-price (γ) controls for competitor moves. Hierarchy pools α and β across SKUs in the same category.

The own-price elasticities came back tighter than we expected. Eleven of the 14 SKUs had 90% credible intervals fully below −1 (elastic). Three were inelastic; we recommended not dropping price on those. The last consultancy’s deck had assumed uniform elastic response across the portfolio — an assumption that, had it been actioned, would have destroyed margin on those three SKUs without producing volume.

Why the previous estimate was wrong

The previous team had estimated elasticities using a single-equation OLS regression of log-volume on log-price, no hierarchy, no cross-price terms, no time fixed effects, fit on a 26-week window. Three problems with that, in increasing order of severity:

  1. Omitted variable bias. Without competitor price terms, β picks up shared price movement across the category. The estimate is biased toward zero for some SKUs and away from zero for others, in directions that depend on retailer assortment.
  2. No partial pooling. For SKUs with thin histories, the un-pooled estimates were dominated by promotional noise. Two SKUs had estimated elasticities of −0.3 and −4.1 respectively. The pooled hierarchical estimates were −1.6 and −2.2, with credible intervals that overlapped substantially.
  3. No uncertainty quantification. The deck reported point estimates with no intervals. A pricing committee asked to defend a 4.2% price change cannot defend a point estimate. They need to know how much the answer could move and still be consistent with the data.

The third problem was the political one. Once the committee had Bayesian credible intervals in hand, the conversation moved from “are these numbers right” to “given these numbers, what should we do.” That is a different conversation. That is the conversation we were hired to enable.

What we found

Across the 14 priority SKUs, the recommended SRP reductions ranged from 2.1% to 5.8%, with a weighted average of 4.2%. The expected Q3 revenue lift (in counterfactual terms — what we expected to see vs. holding price flat) was $1.84M, with a 90% credible interval of $1.43M to $2.27M.

The actual Q3 result, when it came in, was $1.91M above the synthetic-control counterfactual. That sits inside the 90% interval and within a quarter of a standard deviation of the posterior mean. The model performed as advertised. The pricing committee re-ran the model against Q4 actuals and re-committed to the methodology for FY27 planning.

Quarterly revenue: counterfactual vs. actual (Q1–Q4 2026) $0 $1M $2M Q1 Q2 Q3 Q4 PRICE CHANGE Actual: +$1.84M Counterfactual
Figure 1: Quarterly revenue with and without the recommended SRP changes. Counterfactual estimated via synthetic control against matched non-priority SKUs.

Why the last consultancy got it wrong

This is the part of the dispatch that’s closer to a rant than a methodology piece, so we’ll keep it tight. The previous engagement got it wrong for three reasons, none of which were technical, all of which were structural.

1. They were selling a transformation, not an answer.

The economics of a Tier-1 consultancy reward selling 18-month, multi-million-dollar transformations. They do not reward selling 8-week, deliverable-bounded studies that produce a defensible number. The work we do is structurally smaller and structurally faster, which means we cannot survive on the consultancy economic model and we don’t try to.

2. They staffed it with strategy consultants, not data scientists.

The team on the previous engagement was a partner, two senior managers, and four analysts. Of the seven, exactly one had a quantitative degree (an econ undergrad). The actual model fitting was outsourced to a marketing analytics vendor whose work the engagement team could not meaningfully review. If your team cannot defend the math, you have not done the work.

3. They optimized for committee comfort over committee defensibility.

The slides were beautiful. They had been through six rounds of polish. The pricing committee enjoyed reading them. The committee could not, however, defend a specific number to the CFO using those slides — because the slides did not contain a specific number. There is a structural difference between work that is comfortable to receive and work that is defensible to act on. A pricing committee that has to defend a 4.2% SRP change needs the latter.

The handoff

We delivered three things at the end of the eight weeks: the model and its code (fully reproducible), the recommendation document (12 pages, no decks), and the operational handoff — what we call the “defending document.” That last piece is the one most engagements never produce. It is the answer to the question what does the pricing committee say to the CFO when the CFO asks how confident we are in this number. It contains:

  • The full posterior distributions for every parameter that drives the recommendation
  • The sensitivity analysis against the three biggest data quality concerns (one of which we couldn’t fully resolve)
  • The list of assumptions that, if violated, would change the recommendation by more than 50 basis points
  • The monitoring protocol — the metrics the operator should watch in the eight weeks after launch to detect whether the model is breaking

The defending document is not a sales artifact. It exists for the same reason a scientific paper has a methods section: so that the work can be checked, defended, and reproduced. It is the artifact that distinguishes a study from a slide deck.

If the work cannot survive that level of scrutiny, it shouldn’t have shipped.

What this generalizes to

The pricing case is specific. The pattern is general. Across the engagements we’ve run in the last twelve months — pricing, trade promo, demand forecasting, customer LTV — the same three failures show up in the work we’re asked to replace:

  1. The previous team estimated the right quantity using the wrong method.
  2. The previous team did not quantify uncertainty in a way that survives committee challenge.
  3. The previous team did not produce a defending document.

None of those failures are caused by the previous team being bad at their jobs. They are caused by the structure of the engagement they were sold. We are trying to sell a different structure.

That structure is: 6–10 weeks, senior operators end to end, the math is in the deliverable, and the work is defensible when the room gets uncomfortable. If that’s what you need, we’d like to talk.