Build notes · Goh Kun Ming
Driver Safety Analytics Platform
Data · Public project
Oct 2025 – Dec 2025
Project overview
A driver-safety data pipeline combining validation, feature engineering, dimensional modelling and analytical reporting.
A driver-safety data-engineering pipeline moving driver, sensor and safety-label records from MySQL staging through Python and PySpark processing into a star-schema warehouse. It prepares data and Power BI artefacts for exploratory analysis, with safe command-line orchestration and offline checks rather than claims of production scale or safety impact.
Inside this build
- Modelled driver and trip-sensor dimensions linked to booking-level safety labels.
- Applied schema standardisation, deduplication, missing-value handling and rule-based validation.
- Derived acceleration magnitude, jerk, trip aggregates and behaviour indicators.
- Implemented IQR and PCA-reconstruction diagnostics for unusual sensor records.
- Added environment validation, dry runs and opt-in live database access.
- Preserved notebooks, SQL, Power BI artefacts and architecture, privacy and operations documentation.












