Machine Learning2 min read

Student Dropout Prediction

Student Dropout Prediction

Problem Statement

Higher education dropout is a costly problem for institutions and students alike. Using a UCI Machine Learning dataset of 4,424 students with 36 enrollment-time features, this project builds a classifier to identify at-risk students early enough for intervention.

The dataset captures academic path, demographics, and socio-economic factors known at enrollment - meaning the model could flag students before they ever attend a class.

Methodology

Preprocessing & Class Imbalance

The original three-class target (Dropout, Enrolled, Graduated) was collapsed into a binary problem: Dropout vs. Retained. The resulting class imbalance was addressed with SMOTE oversampling on the training set.

Class distribution before and after SMOTE balancing
Class distribution before and after SMOTE resampling, ensuring balanced training data.

Feature Engineering

Continuous variables were scaled using Min-Max normalization, fit on the training set and applied consistently to the test set to prevent data leakage.

Results

Model Performance

A CatBoost Classifier achieved strong performance on the held-out test set:

Accuracy Precision Recall
89% 87% 91%
Confusion matrix for the CatBoost model
Confusion matrix shows strong classification performance with few false negatives.
ROC and Precision-Recall curves for the model
ROC and Precision-Recall curves confirm robust discrimination between dropout and retained students.

Feature Importance

Top features driving dropout prediction
Top predictive features reveal that tuition status and curricular performance dominate the model.

Financial & Demographic Risk Factors

Students with tuition debt, without scholarships, and from lower-income backgrounds showed meaningfully higher dropout rates - suggesting targeted financial aid could be a powerful retention tool.

Dropout rate by tuition payment status
Students with outstanding tuition have significantly higher dropout rates.
Dropout rate by scholarship status
Scholarship holders are substantially more likely to complete their degree.
Dropout rate by gender
Gender-based dropout rate comparison reveals differences in retention patterns.

Next Steps

  • Deploy the model as a scoring API for real-time enrollment risk assessment
  • Apply SHAP values for per-student explainability to guide targeted interventions
  • Explore ensemble methods (stacking, blending) for incremental accuracy gains

Tools and technologies

  • Python
  • CatBoost
  • SMOTE
  • scikit-learn

Data source: KaggleStatus: Completed

  • Classification
  • CatBoost
  • Education