Lesson 15
Multicollinearity and VIF
Big question
Why can highly related regressors make estimates imprecise?
Lesson progress
Complete checkpoints as you learn
Learning objectives
- Explain multicollinearity and vif in plain language.
- Use ceteris paribus correctly in an interpretation.
- Connect the lesson idea to a formula, graph, Python result, or real example.
Simple explanation
Multicollinearity means explanatory variables share information. VIF summarizes how much the variance of a coefficient is inflated by that shared information.
Key terms
- ceteris paribus
- Other included factors held fixed for interpretation.
- multicollinearity
- Strong but not perfect relationships among explanatory variables.
- residual
- The sample prediction error y minus fitted y.
Core formula
Use plain-language interpretation before algebra.
Example
A high VIF is a warning about precision, not automatic proof that a model is invalid or that a control should be removed.
Interactive visual
VIF calculator and multicollinearity visualizer
Original Module 3 visual for Multicollinearity and VIF.
y variable
wage
The dependent variable. It is the outcome students want to explain.
x variable
education
The explanatory variable. It is used to describe changes in wage.
Live Python
Multicollinearity and VIF Python example
Multicollinearity and VIF Python example
Stdout
Run Python to see results here.
Status / stderr
Ready to run Python in your browser.
Line-by-line guide
- Line 1Load a Python library needed for data work or regression.
- Line 2Load a Python library needed for data work or regression.
- Line 3Load a Python library needed for data work or regression.
- Line 5Load the dataset into a pandas DataFrame.
- Line 6Keep rows that have the variables required for this model.
- Line 7Add an intercept column to the regression design matrix.
- Line 8Create or update a Python object used in the analysis.
- Line 9Run this Python instruction as part of the lesson workflow.
- Line 10Run this Python instruction as part of the lesson workflow.
- Line 11Run this Python instruction as part of the lesson workflow.
- Line 12Display a result so students can inspect the output.
Python walkthrough
- 1Load the dataset and keep only variables needed for the current model.
- 2Construct the dependent variable and explanatory-variable matrix.
- 3Add a constant so the regression includes an intercept.
- 4Fit OLS and interpret coefficients with the controls held fixed.
Live notebook
Run this lesson as a notebook
Open an editable notebook cell-by-cell, run Python in the browser, and download the `.ipynb` file for later.
Related dataset
WAGE1
Estimated time
25 to 40 min
Packages
pandas, numpy, statsmodels, patsy
Expected output
Printed Python results that can be compared with the lesson explanation.
Learning goals
- Load and inspect WAGE1.
- Run the Python cells connected to Multicollinearity and VIF.
- Interpret the output using VIF and multicollinearity.
Common errors
- File not found: check that WAGE1.DTA is installed or use the course data folder.
- Package import error: use the browser notebook first, then download for local Jupyter if your local packages differ.
- Column name error: compare your variable names with the dataset variables listed for this notebook.
Dataset path helper
import pandas as pd
df = pd.read_stata("/data/WAGE1.DTA")
df.head()Interactive activity
MulticollinearityVisualizer
Calculate VIF
Translate shared variation among regressors into a precision warning.
Controlled prediction
Interpretation sentence
Holding experience and tenure fixed, one more year of education changes predicted log wage by about 0.075 in this teaching setup.
Residual sketch
Variable role check
In a wage model with education as the focal x, classify experience.
Precision check
Larger samples tighten estimates. Under the Gauss-Markov assumptions, OLS has the smallest variance among linear unbiased estimators.
VIF calculator
VIF: 2.22
SE index: 0.136
Try it yourself
Write one plain-English sentence explaining the main idea from this lesson.
Common mistakes
Check these before you move on.
A regression coefficient describes a pattern unless the assumptions or research design support a causal interpretation.
Quick quiz
What is the safest interpretation focus in Multicollinearity and VIF?
Quick quiz
Which interpretation habit is most important in Multicollinearity and VIF?
Quick quiz
Why is HTV a reasonable practice dataset here?
Key takeaway
Multicollinearity and VIF helps turn multiple regression output into a careful ceteris paribus interpretation.