CRISP-DM — the Cross Industry Standard Process for Data Mining — was published in 1999. It has outlived nearly every methodology that tried to replace it, for a simple reason: it describes what actually happens rather than what a vendor wishes happened.
It remains the right structure for an analytics project in 2026. But the cost profile across its six phases has changed dramatically, and running it the way you would have in 2015 now misallocates most of your effort.
CRISP-DM: the lifecycle, and where AI actually changes it
Phase 1 — Business Understanding
What are we actually trying to decide?
Unchanged, still the phase that determines whether anything downstream was worth doing, and still where projects most often go wrong — not by executing badly but by executing something nobody needed.
The test: can you finish the sentence "when this is done, someone will decide X differently"? If not, you are producing a report, not an analysis. A churn model that accurately predicts churn nobody can prevent is a technically successful failure.
AI contributes nothing here, and that is not a limitation of current tools. It requires knowing what the business can actually change — information that does not exist in the dataset.
Phase 2 — Data Understanding
What have we got, and can we trust it?
Profiling, distributions, missingness, obvious anomalies. AI accelerates this substantially — asking "what is unusual about this dataset" now returns something useful in seconds.
The trap is stopping at what the tool surfaces. AI profiles the data in front of it; it does not know that a branch stopped reporting in March, or that a definition changed mid-year. Those gaps are found by asking people, not tools.
Phase 3 — Data Preparation
Getting it into a usable state.
Historically 60–80% of project effort. AI cuts it meaningfully but does not remove it — and adds a specific new risk: a tool will confidently make a cleaning decision you did not intend and will not mention it.
The practice that matters: ask for a change log, not a cleaned file. Which columns changed, under what rule, how many rows affected, with an example. Without that, you have inherited decisions you cannot audit.
Phase 4 — Modelling
Building something that predicts or classifies.
This is where the change is dramatic. The phase that defined the profession is now often the cheapest — frequently an afternoon rather than weeks.
The practical implication is one most teams have not absorbed: stop budgeting your project around modelling. If your plan allocates most of the timeline here, it was written for a world that no longer exists.
Phase 5 — Evaluation
Does this actually solve the business problem?
Distinct from technical validation, and the distinction matters. A model with 94% accuracy that performs worse than the existing rule-of-thumb on the cases that matter has passed technical validation and failed evaluation.
This phase has grown in importance precisely because modelling got cheap. When producing a model took weeks, its scarcity filtered quality. Now that anyone can generate several, something has to decide which is worth trusting — and that something is this phase.
Phase 6 — Deployment
Getting it into the business, and keeping it working.
Now the bottleneck, and the phase where most projects quietly stop. Deployment means more than shipping code: who acts on the output, how they access it, what happens when it degrades, who notices.
Monitoring belongs here and is routinely omitted. A model deployed and never checked is a liability accruing quietly — data drifts, definitions change, and the output stays confident throughout.
The iteration point people miss
CRISP-DM is drawn as a cycle deliberately. The arrows back are as important as the arrows forward, and the most common productive loop is Evaluation back to Business Understanding — the analysis reveals the original question was slightly wrong.
Teams running this as a waterfall treat that loop as failure and push a mediocre answer to deployment instead. It is not failure; it is the method working.
How to run it in 2026
- Spend disproportionate time on Phase 1. It is free and it determines everything.
- Use AI aggressively in Phases 2 and 3 — but demand change logs.
- Timebox Phase 4. If modelling takes more than a few days, the problem is usually the data or the question, not the algorithm.
- Over-invest in Phase 5. Cheap models mean the evaluation gate is doing more work.
- Plan Phase 6 from the start. Deciding how something deploys after it is built is how projects join the 95% that never ship.
Our AI Analytics programme runs this lifecycle end-to-end on real data with AI tooling — HRD Corp SBL-KHAS claimable for eligible Malaysian employers.