Calculating probabilities of an outcome is not the same as making a prediction.

What is the probability that a potential customer will buy your product rather than an alternative?

What is the probability that they will choose one SKU rather than another from your range?

Will your new advertising tagline remain in their memory long enough to influence a purchase?

Will a price increase reduce volume by more than the additional margin generated?

These are probability questions, but that does not make them the same statistical problem. They are not predictions.

Marketers have always tried to put numbers around questions like these, because the accountants and engineers who generally run the place distrust anything they cannot squeeze into a spreadsheet. A number, particularly one that looks like a detailed calculation,  attracts far less scrutiny than an honest admission of uncertainty.

Probability models cannot remove uncertainty. They can, however, show us what normal buying behaviour looks like, and provide a benchmark against which we can test our assumptions.

Two drivers in any repeat purchase market

In established FMCG and other repeat-purchase markets, two forces largely shape buying behaviour.

The first is how often someone buys from the category.

Take the yoghurt market as an example. In every case, the behaviour and volume of buying will vary from heavy users to occasional buyers, from brand and sub category agnostic price buyers to brand advocates, and every point between.

Statisticians describe the pattern created by these different buying rates with the eye watering name of Negative Binomial Distribution, usually shortened to NBD.

NBD does not tell us that an individual will buy yoghurt next Thursday. It describes how purchase frequency spreads across the whole population: a few heavy buyers, many light buyers and a group who buy nothing during the measurement period.

The second force is brand choice.

Most buyers choose from a brand and variety repertoire. One brand may dominate their purchases, but price, availability, flavour, pantry stock and the occasionally volatile demands of the household influence each decision.

This is the Dirichlet model which describes how buyers divide their purchases across that repertoire. It is a weighted average of buyer behaviour across the market for each individual possible choice and combination of choices.

Think of each buyer rolling a set of weighted dice. Every brand appears on the dice, but some brands occupy more faces than others. The result remains uncertain, but it does not remain completely random.

Combine category purchase frequency with brand-choice probabilities and you can build a picture of the market. Statisticians would call it an NBD–Dirichlet market model.

From untidy households to stable market models

Continuing the yoghurt example.

The market contains multiple brands and many SKUs covering plain, fruit, Greek, low-fat, full-fat, lactose-free and some emerging and specialty products that represent a purchase choice in the wider yoghurt market.

Each household behaves differently. One buys frequently and moves between several brands. Another buys occasionally and nearly always chooses the same product. A third selects whichever brand carries the discount sticker.

Individual purchases look erratic. Add thousands of them together over time, and recognisable patterns emerge.

Large brands usually win because more people buy them, not because their customers display dramatically greater loyalty.

Smaller brands suffer a form of ‘double jeopardy’. They attract fewer buyers, and those buyers tend to purchase them slightly less often.

Buyers also share their purchases across competing brands. Your customers do not belong to you. They belong to the category and sometimes, when it suits them, buy your product.

Having operated in many FMCG categories over the years, those observations have held true in every case.

A model is not a crystal ball

The NBD–Dirichlet model works best in relatively stable markets where customers make repeat purchases and treat the competing brands as reasonable substitutes.

Defining the boundaries of the market is therefore crucial.

A parent may not see a child’s yoghurt pouch and a tub of plain Greek yoghurt as alternatives. Combining them in one model will produce statistical anomalies. The analysis should separate meaningful subcategories where buyer behaviour shows clear partitions.

The model cannot tell you whether a new tagline will lodge in buyers’ memories. That requires creative testing and evidence of memory and behavioural effects.

It cannot calculate the price elasticity, the contribution on margins of price changes, and likely competitor responses.

It cannot reliably forecast a genuinely new category for which no pattern yet exists. That task remains in the hands of the creative marketer, an increasingly valuable person in this age of ‘AI everything’

Importantly, the model cannot explain why every change occurred. A promotion, stockout, new distribution agreement, competitor withdrawal or advertising campaign may shift the observed probabilities. The model provides the baseline that helps us recognise when something unusual has happened.

Marketers should use probability models to challenge assumptions, establish realistic benchmarks and identify deviations and outliers worth investigating.

They should not be used as a crutch for decision making, as they cannot tell you why change has occurred.

Note: the combination of NBD and the Dirichlet models to reflect behaviour in a market comes from the Ehrenberg-Bass institute for marketing science.