Economics & Finance3 min read

Predicting Next Month's Unemployment Rate

Predicting Next Month's Unemployment Rate

Problem Statement

The Federal Reserve communicates its economic outlook through carefully worded FOMC meeting minutes. This project tests whether the language in those minutes can predict the U.S. unemployment rate one month ahead - and deploys the model as a live Streamlit web tool.

The premise: the Fed responds to real-time market data and telegraphs its policy response through deliberate word choices. If we can quantify that language, we can extract a leading signal on labor market conditions.

Methodology

NLP Feature Engineering (77 Features)

229 historical FOMC meeting minutes (1993–2024) were scraped from the Federal Reserve website, cleaned, and transformed into 77 quantitative features across seven categories: document statistics (word count, vocabulary richness, sentence complexity), sentiment analysis (TextBlob polarity and subjectivity), economic keyword frequencies across 10 curated categories (employment, inflation, recession, monetary policy, etc.), TF-IDF with SVD dimensionality reduction (20 latent components), NMF topic modeling (8 topics), lagged unemployment values (lags, rolling statistics, differences), and cyclical month encoding.

Model Selection

Eight candidate models were evaluated using time-series cross-validation (5-fold TimeSeriesSplit, no future data leakage): Linear Regression, Ridge, Lasso, Elastic Net, Random Forest, Gradient Boosting, XGBoost, and an ensemble of the top three performers. The best model was selected based on cross-validation MAE, then retrained on all available data. Bootstrap resampling (50 iterations) provides 90% prediction intervals at inference time.

Results

U.S. unemployment rate history showing sensitivity to credit cycles
Historical U.S. unemployment rate showing high sensitivity to changes in the credit cycle.

The winning model was Lasso regression (alpha=0.01), achieving a test MAE of 0.608 percentage points and R-squared of 0.496. Regularized linear models outperformed tree ensembles on this small dataset, where the high feature-to-sample ratio benefits from strong regularization. The model struggles most during regime shifts (e.g., the COVID-19 spike in 2020), where historical patterns break down.

Sentiment Analysis

FOMC statement sentiment score vs unemployment rate over time
FOMC statement sentiment inversely tracks unemployment - hawkish language precedes labor market weakness.

Feature Importance

Feature importance from the NLP prediction model
Top predictive features: lagged unemployment values dominate, with text-derived features (TF-IDF components, keywords, sentiment) providing incremental improvement.

Model Performance

Model predictions vs actual unemployment rates
Model predictions closely track actual unemployment rates across the test period.

Streamlit Web Application

The multi-page Streamlit app includes a prediction tool, model performance dashboard, data explorer with sentiment and keyword trend charts, and a methodology page. The prediction tool requires two user inputs:

  • Meeting date (used to fetch lagged unemployment values from FRED)
  • Copy and paste of FOMC meeting minutes text

Predictions include a 90% bootstrap confidence interval and a historical context chart showing where the forecast falls relative to recent unemployment trends.

Streamlit web application interface for unemployment prediction
The Streamlit web tool interface - users paste FOMC minutes text and receive a prediction with confidence intervals.

Future Directions

  • Add VADER sentiment analysis for richer sentiment features alongside TextBlob
  • Expand the model to predict other macroeconomic indicators from FOMC language
  • Expand keyword dictionaries to capture emerging economic themes (e.g., AI, supply chain, tariffs)

Tools and technologies

  • Python
  • Streamlit
  • scikit-learn
  • NLP

Data source: Federal Reserve (FOMC Statements)Status: Active

  • NLP
  • Time Series
  • Streamlit
  • Federal Reserve