Build notes · Goh Kun Ming
Customer Segmentation
AI · Notebook experiment
Jan 2025 – Feb 2025
Project overview
Exploring customer groups through clustering.
A customer-segmentation workflow comparing clustering approaches for age, income and spending data. It includes input validation, reusable model artifacts and interpretable segment summaries for exploratory marketing analysis.
Inside this build
- Validated finite, unique customer identifiers and documented categorical values, with a 1.5×IQR income-outlier rule.
- Standardised age, income and spending features while reserving gender for post-clustering interpretation.
- Compared K-Means, agglomerative clustering and DBSCAN, with and without PCA.
- Built a configurable six-segment K-Means workflow that saves its scaler, model, assignments and metrics.
- Evaluated silhouette, Davies–Bouldin and Calinski–Harabasz scores alongside inertia, and added 15 tests with 95% coverage.
- Documented the absence of external validation, stability studies and measured campaign lift; the segments remain exploratory.
Project files
Original project files are in English.
-
Original coursework presentation
Document preview
Use the page controls to browse. Enlarge a page to read the details.
Loading document…
When enlarged, scroll within the page to see more.
Page text
Text is extracted from the original English document. Charts, images and reading order may need the PDF for context.












