A Malaysian retailer runs a Hari Raya campaign. Sales rise 18%. The campaign is declared a success and the budget is increased next year.
But sales rise every Hari Raya. Competitors also ran campaigns. There was a salary adjustment that month. The question that was never asked: what would sales have been without the campaign?
That counterfactual is the whole of causal reasoning, and it is the question business analysis most consistently skips.
Three explanations for any correlation
When two things move together, exactly three explanations exist, and only one of them justifies spending money.
1. A genuinely causes B. The useful case.
2. Something else causes both. The classic example: ice cream sales and drowning incidents correlate strongly. Hot weather drives both. Banning ice cream would save nobody.
The business version is subtler and far more common. Customers who use your mobile app spend more. Therefore the app drives spending? Or your most engaged customers were always going to spend more, and they are also the ones who bother installing apps? These have opposite investment implications.
3. Coincidence. Test enough variable pairs and some will correlate by chance. Comparing fifty metrics against revenue will surface several impressive-looking relationships that mean nothing whatsoever.
The third explanation: something else causes both
The traps with names
Selection effects
Loyalty-programme members spend more than non-members. Does the programme cause spending? Partly — but customers who join loyalty programmes were already your frequent buyers. You are measuring who selected in, not what the programme did.
Simpson's paradox
A pattern that holds across the whole dataset can reverse within every subgroup. Branch A appears to outperform Branch B overall, yet B outperforms A in every product category — because A sells more of the high-margin categories. The aggregate is real and the conclusion drawn from it is backwards. Any time an aggregate surprises you, split it.
Regression to the mean
Your worst-performing branch improves after an intervention. Extreme results tend to be followed by less extreme ones regardless of what you do, simply because part of the original extreme was luck. Without a comparison group you cannot separate your intervention from ordinary reversion — and this pattern quietly validates a great deal of management action that did nothing.
Survivorship bias
Analysing only current customers to understand retention misses everyone who already left — the exact group with the answer.
Simpson's paradox: the aggregate can point the wrong way
Methods that actually establish cause
Randomised experiment (A/B test) — the gold standard
Randomly split into treatment and control. Because assignment is random, the two groups are comparable on everything, measured or not. The difference is the causal effect. Where feasible this settles the question; see the A/B testing article.
Difference-in-differences — when you cannot randomise
Roll a change out to some branches and not others. Compare the change in treated branches against the change in untreated ones. Comparing changes rather than levels removes the effect of anything affecting both groups equally — the festive season, the economy, a competitor's national campaign.
This is the most practical method for Malaysian businesses that cannot randomise customers but can stagger a rollout across outlets or regions. Its one real requirement: before the change, both groups should have been trending similarly. Check that.
Holdout groups
Deliberately exclude a random 5–10% from every campaign. Sacrificing a little short-term revenue buys you the ability to know what your marketing actually does. Most organisations that adopt this discover at least one long-running programme with no measurable effect.
Before-and-after with a comparison series
The weakest of the credible options, but far better than before-and-after alone: track a similar metric you did not intervene on, as a rough control.
Where AI helps — and where it will mislead you
AI tools are excellent at finding correlations, testing many relationships quickly, and flagging confounders you had not considered. Asking a capable model "what else could explain this result?" is a genuinely useful five-minute exercise that surfaces alternative explanations a busy team will not generate on its own.
The caution is real, though. AI writes fluent causal language by default — "the campaign drove an 18% increase" — because that is how business writing sounds. Fluency is not evidence. The model has no privileged knowledge of what caused what; it is describing a correlation in causal grammar.
Ask explicitly: what would have happened without this? What else could explain it? What comparison group would settle it?
The practical standard
You will rarely have experimental proof for every decision, and demanding it would paralyse the business. A workable bar:
- Small, reversible decisions — correlation plus sensible reasoning is fine.
- Significant recurring spend — require a holdout or a difference-in-differences comparison. If a campaign runs annually, you can afford to measure it properly once.
- Major strategic commitments — run a pilot with a genuine control group before scaling.
And when someone presents a chart claiming X caused Y, the most valuable question in the room remains: compared to what?
Our AI Analytics programme covers causal reasoning, experiment design and attribution for Malaysian business teams — HRD Corp SBL-KHAS claimable.