Lesson 1

What Is Heteroskedasticity?

Big question

How can regression uncertainty change across values of x?

Lesson progress

Complete checkpoints as you learn

0% complete0 checkpoint streak
Progress tracking available after sign-in. Sign in to save your work.
Big question
Concept
Activity
Quiz

Learning objectives

  • Explain what is heteroskedasticity? in plain language.
  • Use heteroskedasticity correctly in an interpretation.
  • Connect the lesson idea to a formula, graph, Python result, or real example.

Simple explanation

Heteroskedasticity is about changing uncertainty. In what is heteroskedasticity?, students focus on changing error spread and learn how to describe that pattern without confusing it with coefficient bias.

Key terms

Heteroskedasticity
A pattern where the conditional variance of the error changes with regressors or groups.
Robust standard error
A standard error designed to remain asymptotically valid under heteroskedasticity.
Weighted least squares
A least-squares method that gives lower weight to observations with higher error variance.

Core formula

Var(uixi1,...,xik)=sigma2(xi)Var(u_i | x_{i1},...,x_{ik}) = sigma^2(x_i)

Use plain-language interpretation before algebra.

Example

MODULE8_INCOME_SAVINGS_SYNTHETIC supports original Ceteris Lab practice for what is heteroskedasticity?. Synthetic files are clearly labeled as synthetic, and installed course datasets are used only when the file is present.

Interactive visual

VariancePatternExplorer: adjust the controls, classify the diagnostic evidence, and write one sentence explaining the implication for inference.

Original Module 8 visual for What Is Heteroskedasticity?.

wage_sample.csv

y variable

wage

The dependent variable. It is the outcome students want to explain.

x variable

education

The explanatory variable. It is used to describe changes in wage.

Live Python

What Is Heteroskedasticity? Python example

What Is Heteroskedasticity? Python example

Stdout

Run Python to see results here.

Status / stderr

Ready to run Python in your browser.

Line-by-line guide

  1. Line 1Load a Python library needed for data work or regression.
  2. Line 2Load a Python library needed for data work or regression.
  3. Line 4Load the dataset into a pandas DataFrame.
  4. Line 5Display a result so students can inspect the output.
  5. Line 6Create or update a Python object used in the analysis.
  6. Line 7Run this Python instruction as part of the lesson workflow.
  7. Line 8Run this Python instruction as part of the lesson workflow.
  8. Line 9Run this Python instruction as part of the lesson workflow.
  9. Line 10Run this Python instruction as part of the lesson workflow.

Python walkthrough

  1. 1Load the Python packages needed for data, regression, diagnostics, or plotting.
  2. 2Read an installed Ceteris Lab dataset from a browser-safe public path.
  3. 3Estimate the baseline model before changing the covariance method or weights.
  4. 4Print diagnostic evidence or a coefficient comparison so students can inspect the result.
  5. 5Interpret the output as practice evidence and avoid making real empirical claims from synthetic data.

Live notebook

Run this lesson as a notebook

Open an editable notebook cell-by-cell, run Python in the browser, and download the `.ipynb` file for later.

Related dataset

MODULE8_INCOME_SAVINGS_SYNTHETIC

Estimated time

25 to 40 min

Packages

pandas, numpy, statsmodels, patsy

Expected output

Printed Python results that can be compared with the lesson explanation.

Learning goals

  • Load and inspect MODULE8_INCOME_SAVINGS_SYNTHETIC.
  • Run the Python cells connected to What Is Heteroskedasticity?.
  • Interpret the output using heteroskedasticity and robust standard errors.

Common errors

  • File not found: check that module8_income_savings_heteroskedastic.csv is installed or use the course data folder.
  • Package import error: use the browser notebook first, then download for local Jupyter if your local packages differ.
  • Column name error: compare your variable names with the dataset variables listed for this notebook.

Dataset path helper

import pandas as pd

df = pd.read_csv("/data/module-8/module8_income_savings_heteroskedastic.csv")
df.head()

Interactive activity

Variance Pattern Explorer

What Is Heteroskedasticity?

change the inputs, inspect the feedback, and decide whether robust inference, diagnostics, WLS, or reporting caution is needed.

MODULE8_INCOME_SAVINGS_SYNTHETIC

Inputs

Visual preview

Changing variance across fitted values

fitted valueresidual spread

A wider fan means the uncertainty changes across observations. Robust standard errors adjust inference; WLS needs a defensible variance model.

Try it yourself

Write one plain-English sentence explaining the main idea from this lesson.

Common mistakes

Check these before you move on.

A regression coefficient describes a pattern unless the assumptions or research design support a causal interpretation.

Quick quiz

Which statement is most accurate for What Is Heteroskedasticity??

Quick quiz

What should a careful Module 8 report include for What Is Heteroskedasticity??

Quick quiz

Why is MODULE8_INCOME_SAVINGS_SYNTHETIC a reasonable practice dataset here?

Key takeaway

What Is Heteroskedasticity? helps students diagnose changing variance and choose inference or weighting methods without overclaiming what those methods can fix.