AITraining2U

Programs

Resources

Case Studies

Quick Links

Enquire Now
AI Analytics

Correlation vs Causation: Why the Campaign That "Worked" May Have Done Nothing

Sales rose 18% after the campaign. That sentence has convinced more boards to approve more budget than almost any other — and on its own it proves nothing at all.

By AITraining2U Editorial Team 2026-08-26 10 min read
Analysing cause and effect in business data

A Malaysian retailer runs a Hari Raya campaign. Sales rise 18%. The campaign is declared a success and the budget is increased next year.

But sales rise every Hari Raya. Competitors also ran campaigns. There was a salary adjustment that month. The question that was never asked: what would sales have been without the campaign?

That counterfactual is the whole of causal reasoning, and it is the question business analysis most consistently skips.

Three explanations for any correlation

When two things move together, exactly three explanations exist, and only one of them justifies spending money.

1. A genuinely causes B. The useful case.

2. Something else causes both. The classic example: ice cream sales and drowning incidents correlate strongly. Hot weather drives both. Banning ice cream would save nobody.

The business version is subtler and far more common. Customers who use your mobile app spend more. Therefore the app drives spending? Or your most engaged customers were always going to spend more, and they are also the ones who bother installing apps? These have opposite investment implications.

3. Coincidence. Test enough variable pairs and some will correlate by chance. Comparing fifty metrics against revenue will surface several impressive-looking relationships that mean nothing whatsoever.

The third explanation: something else causes both

NOT causalIce creamsalesDrowningincidentsHot weathercausescausesBoth rise together. Neither causes the other.The business version: “app users spend more” — or do engaged customers install apps?
When two things move together, only one of the three possible explanations justifies spending money.

The traps with names

Selection effects

Loyalty-programme members spend more than non-members. Does the programme cause spending? Partly — but customers who join loyalty programmes were already your frequent buyers. You are measuring who selected in, not what the programme did.

Simpson's paradox

A pattern that holds across the whole dataset can reverse within every subgroup. Branch A appears to outperform Branch B overall, yet B outperforms A in every product category — because A sells more of the high-margin categories. The aggregate is real and the conclusion drawn from it is backwards. Any time an aggregate surprises you, split it.

Regression to the mean

Your worst-performing branch improves after an intervention. Extreme results tend to be followed by less extreme ones regardless of what you do, simply because part of the original extreme was luck. Without a comparison group you cannot separate your intervention from ordinary reversion — and this pattern quietly validates a great deal of management action that did nothing.

Survivorship bias

Analysing only current customers to understand retention misses everyone who already left — the exact group with the answer.

Simpson's paradox: the aggregate can point the wrong way

In aggregate: A beats BIn every category: B beats ABranch A62%Branch B54%Fresh44%58%Frozen51%66%Dry goods38%49%Both are true. A sells more of the high-margin categories, which lifts its aggregate.Any time an aggregate surprises you — split it before you act on it.
Branch A wins overall. Branch B wins in every category. Both facts are correct.

Methods that actually establish cause

Randomised experiment (A/B test) — the gold standard

Randomly split into treatment and control. Because assignment is random, the two groups are comparable on everything, measured or not. The difference is the causal effect. Where feasible this settles the question; see the A/B testing article.

Difference-in-differences — when you cannot randomise

Roll a change out to some branches and not others. Compare the change in treated branches against the change in untreated ones. Comparing changes rather than levels removes the effect of anything affecting both groups equally — the festive season, the economy, a competitor's national campaign.

This is the most practical method for Malaysian businesses that cannot randomise customers but can stagger a rollout across outlets or regions. Its one real requirement: before the change, both groups should have been trending similarly. Check that.

Holdout groups

Deliberately exclude a random 5–10% from every campaign. Sacrificing a little short-term revenue buys you the ability to know what your marketing actually does. Most organisations that adopt this discover at least one long-running programme with no measurable effect.

Before-and-after with a comparison series

The weakest of the credible options, but far better than before-and-after alone: track a similar metric you did not intervene on, as a rough control.

Where AI helps — and where it will mislead you

AI tools are excellent at finding correlations, testing many relationships quickly, and flagging confounders you had not considered. Asking a capable model "what else could explain this result?" is a genuinely useful five-minute exercise that surfaces alternative explanations a busy team will not generate on its own.

The caution is real, though. AI writes fluent causal language by default — "the campaign drove an 18% increase" — because that is how business writing sounds. Fluency is not evidence. The model has no privileged knowledge of what caused what; it is describing a correlation in causal grammar.

Ask explicitly: what would have happened without this? What else could explain it? What comparison group would settle it?

The practical standard

You will rarely have experimental proof for every decision, and demanding it would paralyse the business. A workable bar:

  • Small, reversible decisions — correlation plus sensible reasoning is fine.
  • Significant recurring spend — require a holdout or a difference-in-differences comparison. If a campaign runs annually, you can afford to measure it properly once.
  • Major strategic commitments — run a pilot with a genuine control group before scaling.

And when someone presents a chart claiming X caused Y, the most valuable question in the room remains: compared to what?

Our AI Analytics programme covers causal reasoning, experiment design and attribution for Malaysian business teams — HRD Corp SBL-KHAS claimable.

Frequently Asked Questions

Ask what would have happened without the intervention. That counterfactual is the entire question. If you cannot answer it — because you have no comparison group, no holdout, and no untreated branches — then you have a correlation, however large the observed change. 'Sales rose 18% after the campaign' is not evidence the campaign worked, because sales may well have risen anyway.

You roll a change out to some branches or regions and not others, then compare the change in treated units against the change in untreated ones. Comparing changes rather than levels cancels out anything affecting both groups equally — the festive season, the economy, a competitor's national campaign. It is the most practical causal method for businesses that cannot randomise customers but can stagger a rollout. Its requirement is that both groups were trending similarly beforehand.

A pattern that holds in aggregate can reverse within every subgroup. Branch A may appear to outperform Branch B overall while B outperforms A in every single product category — because A happens to sell more of the high-margin categories. Both facts are true; the aggregate conclusion is backwards. The practical rule: any time an aggregate number surprises you, split it by the obvious dimensions before acting on it.

Because it is the only way to know what your marketing actually does. Excluding a random 5–10% from every campaign sacrifices a small amount of short-term revenue in exchange for a genuine control group. Most organisations that adopt this discover at least one long-running programme with no measurable effect — a finding that pays for the holdout many times over.

No, and it will sound as though it has. AI tools are genuinely useful for finding correlations, testing many relationships quickly and — if you ask directly — suggesting confounders you had not considered. But they write fluent causal language by default because that is how business writing sounds. The model has no privileged knowledge of what caused what. Ask it explicitly what else could explain the result and what comparison group would settle the question.

Build AI systems that hold up in production

Evaluation, observability and guardrails are what separate a demo from a system your business can depend on. AITraining2U runs hands-on, HRD Corp SBL-KHAS claimable AI training for Malaysian organisations — tool-agnostic and mapped to your actual stack.