Lesson 9
Comparing Non-Base Categories
Big question
How can we compare two included categories when neither is the base group?
Lesson progress
Complete checkpoints as you learn
Learning objectives
- Explain comparing non-base categories in plain language.
- Use robust standard error correctly in an interpretation.
- Connect the lesson idea to a formula, graph, Python result, or real example.
Simple explanation
Subtract the two category coefficients or change the base group so the comparison appears directly.
Key terms
- Robust standard error
- A standard error designed to remain valid under heteroskedasticity.
- Binary variable
- A variable that equals one when a condition is present and zero otherwise.
- Intercept shift
- A group difference that changes the expected outcome level while leaving slopes unchanged.
Core formula
Use plain-language interpretation before algebra.
Example
MODULE7_CATEGORY_EFFECTS_SYNTHETIC is an original Ceteris Lab synthetic teaching dataset for comparing non-base categories. It lets students practice non-base test without presenting fabricated real-world empirical findings.
Interactive visual
CategoryDifferenceTester: use the controls to code a group, choose a base group, and write one correct interpretation.
Original Module 7 visual for Comparing Non-Base Categories.
y variable
wage
The dependent variable. It is the outcome students want to explain.
x variable
education
The explanatory variable. It is used to describe changes in wage.
Live Python
Comparing Non-Base Categories Python example
Comparing Non-Base Categories Python example
Stdout
Run Python to see results here.
Status / stderr
Ready to run Python in your browser.
Line-by-line guide
- Line 1Load a Python library needed for data work or regression.
- Line 2Load a Python library needed for data work or regression.
- Line 4Load the dataset into a pandas DataFrame.
- Line 5Create or update a Python object used in the analysis.
- Line 6Create or update a Python object used in the analysis.
- Line 7Add an intercept column to the regression design matrix.
- Line 8Display a result so students can inspect the output.
- Line 9Display a result so students can inspect the output.
Python walkthrough
- 1Load pandas and statsmodels so the workflow is reproducible.
- 2Read the installed Ceteris Lab synthetic CSV from the public data folder.
- 3Create or inspect dummy variables before estimating the model.
- 4Estimate OLS with an intercept and the selected regressors.
- 5Print coefficients or summaries, then interpret them as associations unless the design supports causality.
Live notebook
Run this lesson as a notebook
Open an editable notebook cell-by-cell, run Python in the browser, and download the `.ipynb` file for later.
Related dataset
MODULE7_CATEGORY_EFFECTS_SYNTHETIC
Estimated time
25 to 40 min
Packages
pandas, numpy, statsmodels, patsy
Expected output
Printed Python results that can be compared with the lesson explanation.
Learning goals
- Load and inspect MODULE7_CATEGORY_EFFECTS_SYNTHETIC.
- Run the Python cells connected to Multiple Categories.
- Interpret the output using dummy variables and qualitative information.
Common errors
- File not found: check that module7_category_effects_synthetic.csv is installed or use the course data folder.
- Package import error: use the browser notebook first, then download for local Jupyter if your local packages differ.
- Column name error: compare your variable names with the dataset variables listed for this notebook.
Dataset path helper
import pandas as pd
df = pd.read_csv("/data/module-7/module7_category_effects_synthetic.csv")
df.head()Interactive activity
CategoryDifferenceTester
Comparing Non-Base Categories
CategoryDifferenceTester: choose the coding rule, base group, or probability interpretation before reading the coefficient.
Inputs
Try it yourself
Write one plain-English sentence explaining the main idea from this lesson.
Common mistakes
Check these before you move on.
A regression coefficient describes a pattern unless the assumptions or research design support a causal interpretation.
Quick quiz
Which interpretation is most careful for Comparing Non-Base Categories?
Quick quiz
What should students check before trusting the result in Comparing Non-Base Categories?
Quick quiz
Why is GPA1 a reasonable practice dataset here?
Key takeaway
Comparing Non-Base Categories helps students convert qualitative information into transparent, testable regression comparisons.