AITraining2U

Programs

Resources

Case Studies

Quick Links

Enquire Now
AI Analytics

CRISP-DM in the Age of AI: The Lifecycle Still Holds

A methodology from 1999 remains the clearest way to run an analytics project. What has changed is not the phases — it is which of them costs you anything.

By AITraining2U Editorial Team 2026-09-16 10 min read
Planning an analytics project lifecycle

CRISP-DM — the Cross Industry Standard Process for Data Mining — was published in 1999. It has outlived nearly every methodology that tried to replace it, for a simple reason: it describes what actually happens rather than what a vendor wishes happened.

It remains the right structure for an analytics project in 2026. But the cost profile across its six phases has changed dramatically, and running it the way you would have in 2015 now misallocates most of your effort.

CRISP-DM: the lifecycle, and where AI actually changes it

BusinessUnderstandingDataUnderstandingDataPreparationModelingEvaluationDeploymentDATA(iterate)Business UnderstandingFrame the decision, not the dataset. AI cannot do this for you.Data UnderstandingProfile, spot gaps. AI accelerates this heavily.Data PreparationStill 60–80% of effort. AI cuts it, does not remove it.ModelingNow the cheapest phase. This is the change.EvaluationAgainst the business criterion, not just accuracy.DeploymentWhere most projects quietly stop.
The phases have not changed since 1999. What has changed is which of them is expensive.

Phase 1 — Business Understanding

What are we actually trying to decide?

Unchanged, still the phase that determines whether anything downstream was worth doing, and still where projects most often go wrong — not by executing badly but by executing something nobody needed.

The test: can you finish the sentence "when this is done, someone will decide X differently"? If not, you are producing a report, not an analysis. A churn model that accurately predicts churn nobody can prevent is a technically successful failure.

AI contributes nothing here, and that is not a limitation of current tools. It requires knowing what the business can actually change — information that does not exist in the dataset.

Phase 2 — Data Understanding

What have we got, and can we trust it?

Profiling, distributions, missingness, obvious anomalies. AI accelerates this substantially — asking "what is unusual about this dataset" now returns something useful in seconds.

The trap is stopping at what the tool surfaces. AI profiles the data in front of it; it does not know that a branch stopped reporting in March, or that a definition changed mid-year. Those gaps are found by asking people, not tools.

Phase 3 — Data Preparation

Getting it into a usable state.

Historically 60–80% of project effort. AI cuts it meaningfully but does not remove it — and adds a specific new risk: a tool will confidently make a cleaning decision you did not intend and will not mention it.

The practice that matters: ask for a change log, not a cleaned file. Which columns changed, under what rule, how many rows affected, with an example. Without that, you have inherited decisions you cannot audit.

Phase 4 — Modelling

Building something that predicts or classifies.

This is where the change is dramatic. The phase that defined the profession is now often the cheapest — frequently an afternoon rather than weeks.

The practical implication is one most teams have not absorbed: stop budgeting your project around modelling. If your plan allocates most of the timeline here, it was written for a world that no longer exists.

Phase 5 — Evaluation

Does this actually solve the business problem?

Distinct from technical validation, and the distinction matters. A model with 94% accuracy that performs worse than the existing rule-of-thumb on the cases that matter has passed technical validation and failed evaluation.

This phase has grown in importance precisely because modelling got cheap. When producing a model took weeks, its scarcity filtered quality. Now that anyone can generate several, something has to decide which is worth trusting — and that something is this phase.

Phase 6 — Deployment

Getting it into the business, and keeping it working.

Now the bottleneck, and the phase where most projects quietly stop. Deployment means more than shipping code: who acts on the output, how they access it, what happens when it degrades, who notices.

Monitoring belongs here and is routinely omitted. A model deployed and never checked is a liability accruing quietly — data drifts, definitions change, and the output stays confident throughout.

The iteration point people miss

CRISP-DM is drawn as a cycle deliberately. The arrows back are as important as the arrows forward, and the most common productive loop is Evaluation back to Business Understanding — the analysis reveals the original question was slightly wrong.

Teams running this as a waterfall treat that loop as failure and push a mediocre answer to deployment instead. It is not failure; it is the method working.

How to run it in 2026

  1. Spend disproportionate time on Phase 1. It is free and it determines everything.
  2. Use AI aggressively in Phases 2 and 3 — but demand change logs.
  3. Timebox Phase 4. If modelling takes more than a few days, the problem is usually the data or the question, not the algorithm.
  4. Over-invest in Phase 5. Cheap models mean the evaluation gate is doing more work.
  5. Plan Phase 6 from the start. Deciding how something deploys after it is built is how projects join the 95% that never ship.

Our AI Analytics programme runs this lifecycle end-to-end on real data with AI tooling — HRD Corp SBL-KHAS claimable for eligible Malaysian employers.

Frequently Asked Questions

CRISP-DM (Cross Industry Standard Process for Data Mining) is a six-phase methodology for analytics projects, published in 1999: Business Understanding, Data Understanding, Data Preparation, Modelling, Evaluation and Deployment. It is drawn as a cycle because iteration between phases is expected. It has outlasted most alternatives because it describes what actually happens on real projects rather than an idealised sequence.

The phases are, entirely. What changed is the cost profile across them. Modelling collapsed from the most expensive phase to often the cheapest, data preparation got faster without disappearing, and deployment became the bottleneck. Running the lifecycle with a 2015 effort allocation — most of the timeline on modelling — now misallocates the majority of your project.

Business Understanding and Deployment. The first is free, requires no tooling, and determines whether anything downstream was worth doing. The second is now where most projects stall. Modelling should be timeboxed — if it takes more than a few days, the real problem is usually the data or the question rather than the algorithm.

Data Understanding and Data Preparation most substantially — profiling, spotting anomalies, normalising formats, deduplicating, joining messy sources. It also collapsed Modelling. It contributes essentially nothing to Business Understanding, because that requires knowing what the business can actually change, which is not in the dataset. In Preparation, always ask for a change log rather than a cleaned file so the decisions stay auditable.

Because the most common productive discovery in an analytics project is that the original question was slightly wrong. The analysis reveals that churn was not the problem, onboarding was. Teams running CRISP-DM as a waterfall treat that loop as failure and push a mediocre answer to deployment instead. It is not failure — it is the method working as designed.

Build AI systems that hold up in production

Evaluation, observability and guardrails are what separate a demo from a system your business can depend on. AITraining2U runs hands-on, HRD Corp SBL-KHAS claimable AI training for Malaysian organisations — tool-agnostic and mapped to your actual stack.