Predicting Credit Default (Lending Club)

Problem Statement
Peer-to-peer lending platforms approve thousands of loans daily with limited underwriting infrastructure. Using 38,971 historical loans from Lending Club (2007–2011), this project benchmarks multiple classification algorithms to predict which borrowers will default - and quantifies just how difficult consumer credit modeling remains even with 38 attributes.
Methodology
Data & Preprocessing
Our analysis uses data from 2007-2011, comprising 38,971 observations and 38 attributes. Four categories of models were developed to predict loan status:
- Logistic Regression (Backwards Selection)
- LASSO model
- Elastic Net model
- Random Forest model
Model Training Pipeline
The model training process involved four steps:
- Choosing performance metrics: RMSE, MAE, and R-squared
- Benchmarking with 7 different algorithms (Linear Regression, KNeighbors, AdaBoost, etc.)
- Fine-tuning the 2 best models using k-fold cross validation
- Exporting the best model for application on the test set
Results
The final validation AUC of 69% underscores the challenge of consumer credit default prediction - even with 38 attributes, borrower behavior remains difficult to model with high accuracy.
Next Steps
- Engineer additional features from existing attributes to capture non-linear relationships
- Explore gradient boosting methods (XGBoost, LightGBM) for potentially better performance
- Incorporate temporal features to capture economic cycle effects on default rates