Lesson 18
Prediction under Heteroskedasticity
Big question
Why can prediction intervals widen at high-variance x values?
Lesson progress
Complete checkpoints as you learn
Learning objectives
- Explain prediction under heteroskedasticity in plain language.
- Use homoskedasticity correctly in an interpretation.
- Connect the lesson idea to a formula, graph, Python result, or real example.
Simple explanation
Heteroskedasticity is about changing uncertainty. In prediction under heteroskedasticity, students focus on heteroskedastic prediction bands and learn how to describe that pattern without confusing it with coefficient bias.
Key terms
- Homoskedasticity
- A constant conditional error variance assumption used by conventional OLS standard errors.
- Breusch-Pagan test
- A diagnostic test that regresses squared residuals on variables thought to explain error variance.
- Feasible GLS
- A weighted estimator that first estimates the variance function and then uses predicted weights.
Core formula
Use plain-language interpretation before algebra.
Example
MODULE8_HOUSING_VARIANCE_SYNTHETIC supports original Ceteris Lab practice for prediction under heteroskedasticity. Synthetic files are clearly labeled as synthetic, and installed course datasets are used only when the file is present.
Interactive visual
HeteroskedasticPredictionIntervalVisualizer: adjust the controls, classify the diagnostic evidence, and write one sentence explaining the implication for inference.
Original Module 8 visual for Prediction under Heteroskedasticity.
y variable
wage
The dependent variable. It is the outcome students want to explain.
x variable
education
The explanatory variable. It is used to describe changes in wage.
Live Python
Prediction under Heteroskedasticity Python example
Prediction under Heteroskedasticity Python example
Stdout
Run Python to see results here.
Status / stderr
Ready to run Python in your browser.
Line-by-line guide
- Line 1Load a Python library needed for data work or regression.
- Line 2Load a Python library needed for data work or regression.
- Line 3Load a Python library needed for data work or regression.
- Line 4Load a Python library needed for data work or regression.
- Line 6Load the dataset into a pandas DataFrame.
- Line 7Add an intercept column to the regression design matrix.
- Line 8Estimate an ordinary least squares regression.
- Line 9Create or update a Python object used in the analysis.
- Line 10Add an intercept column to the regression design matrix.
- Line 11Create or update a Python object used in the analysis.
- Line 12Create or update a Python object used in the analysis.
- Line 13Run this Python instruction as part of the lesson workflow.
- Line 14Run this Python instruction as part of the lesson workflow.
- Line 15Run this Python instruction as part of the lesson workflow.
- Line 16Run this Python instruction as part of the lesson workflow.
- Line 17Run this Python instruction as part of the lesson workflow.
- Line 18Display a result so students can inspect the output.
Python walkthrough
- 1Load the Python packages needed for data, regression, diagnostics, or plotting.
- 2Read an installed Ceteris Lab dataset from a browser-safe public path.
- 3Estimate the baseline model before changing the covariance method or weights.
- 4Print diagnostic evidence or a coefficient comparison so students can inspect the result.
- 5Interpret the output as practice evidence and avoid making real empirical claims from synthetic data.
Live notebook
Run this lesson as a notebook
Open an editable notebook cell-by-cell, run Python in the browser, and download the `.ipynb` file for later.
Related dataset
MODULE8_HOUSING_VARIANCE_SYNTHETIC
Estimated time
25 to 40 min
Packages
pandas, numpy, statsmodels, patsy
Expected output
Printed Python results that can be compared with the lesson explanation.
Learning goals
- Load and inspect MODULE8_HOUSING_VARIANCE_SYNTHETIC.
- Run the Python cells connected to Prediction under Heteroskedasticity.
- Interpret the output using heteroskedasticity and robust standard errors.
Common errors
- File not found: check that module8_housing_variance_synthetic.csv is installed or use the course data folder.
- Package import error: use the browser notebook first, then download for local Jupyter if your local packages differ.
- Column name error: compare your variable names with the dataset variables listed for this notebook.
Dataset path helper
import pandas as pd
df = pd.read_csv("/data/module-8/module8_housing_variance_synthetic.csv")
df.head()Interactive activity
Heteroskedastic Prediction Interval Visualizer
Prediction under Heteroskedasticity
change the inputs, inspect the feedback, and decide whether robust inference, diagnostics, WLS, or reporting caution is needed.
Inputs
Visual preview
Changing variance across fitted values
A wider fan means the uncertainty changes across observations. Robust standard errors adjust inference; WLS needs a defensible variance model.
Try it yourself
Write one plain-English sentence explaining the main idea from this lesson.
Common mistakes
Check these before you move on.
A regression coefficient describes a pattern unless the assumptions or research design support a causal interpretation.
Quick quiz
What should a careful Module 8 report include for Prediction under Heteroskedasticity?
Quick quiz
What should a careful Module 8 report include for Prediction under Heteroskedasticity?
Quick quiz
Why is MEAP00 a reasonable practice dataset here?
Key takeaway
Prediction under Heteroskedasticity helps students diagnose changing variance and choose inference or weighting methods without overclaiming what those methods can fix.