Predicting Next Month's Unemployment Rate

Problem Statement
The Federal Reserve communicates its economic outlook through carefully worded FOMC meeting minutes. This project tests whether the language in those minutes can predict the U.S. unemployment rate one month ahead - and deploys the model as a live Streamlit web tool.
The premise: the Fed responds to real-time market data and telegraphs its policy response through deliberate word choices. If we can quantify that language, we can extract a leading signal on labor market conditions.
Methodology
NLP Feature Engineering (77 Features)
229 historical FOMC meeting minutes (1993–2024) were scraped from the Federal Reserve website, cleaned, and transformed into 77 quantitative features across seven categories: document statistics (word count, vocabulary richness, sentence complexity), sentiment analysis (TextBlob polarity and subjectivity), economic keyword frequencies across 10 curated categories (employment, inflation, recession, monetary policy, etc.), TF-IDF with SVD dimensionality reduction (20 latent components), NMF topic modeling (8 topics), lagged unemployment values (lags, rolling statistics, differences), and cyclical month encoding.
Model Selection
Eight candidate models were evaluated using time-series cross-validation (5-fold TimeSeriesSplit, no future data leakage): Linear Regression, Ridge, Lasso, Elastic Net, Random Forest, Gradient Boosting, XGBoost, and an ensemble of the top three performers. The best model was selected based on cross-validation MAE, then retrained on all available data. Bootstrap resampling (50 iterations) provides 90% prediction intervals at inference time.
Results
The winning model was Lasso regression (alpha=0.01), achieving a test MAE of 0.608 percentage points and R-squared of 0.496. Regularized linear models outperformed tree ensembles on this small dataset, where the high feature-to-sample ratio benefits from strong regularization. The model struggles most during regime shifts (e.g., the COVID-19 spike in 2020), where historical patterns break down.
Sentiment Analysis
Feature Importance
Model Performance
Streamlit Web Application
The multi-page Streamlit app includes a prediction tool, model performance dashboard, data explorer with sentiment and keyword trend charts, and a methodology page. The prediction tool requires two user inputs:
- Meeting date (used to fetch lagged unemployment values from FRED)
- Copy and paste of FOMC meeting minutes text
Predictions include a 90% bootstrap confidence interval and a historical context chart showing where the forecast falls relative to recent unemployment trends.
Future Directions
- Add VADER sentiment analysis for richer sentiment features alongside TextBlob
- Expand the model to predict other macroeconomic indicators from FOMC language
- Expand keyword dictionaries to capture emerging economic themes (e.g., AI, supply chain, tariffs)