Lesson 6
Dummy Variables in Log-Dependent Models
Big question
How do we interpret a dummy coefficient when the dependent variable is logged?
Lesson progress
Complete checkpoints as you learn
Learning objectives
- Explain dummy variables in log-dependent models in plain language.
- Use interaction term correctly in an interpretation.
- Connect the lesson idea to a formula, graph, Python result, or real example.
Simple explanation
The quick percent interpretation is 100 times the coefficient, but the exact conversion uses the exponential function.
Key terms
- Interaction term
- A regressor formed by multiplying variables so one effect can depend on another variable.
- Self-selection
- A situation where people choose treatment or participation based on factors related to outcomes.
- Dummy variable
- A binary regressor used to represent qualitative information in a regression.
Core formula
Use plain-language interpretation before algebra.
Example
MODULE7_WAGE_GROUPS_SYNTHETIC is an original Ceteris Lab synthetic teaching dataset for dummy variables in log-dependent models. It lets students practice log dummy percent without presenting fabricated real-world empirical findings.
Interactive visual
DummyPercentEffectCalculator: use the controls to code a group, choose a base group, and write one correct interpretation.
Original Module 7 visual for Dummy Variables in Log-Dependent Models.
y variable
wage
The dependent variable. It is the outcome students want to explain.
x variable
education
The explanatory variable. It is used to describe changes in wage.
Live Python
Dummy Variables in Log-Dependent Models Python example
Dummy Variables in Log-Dependent Models Python example
Stdout
Run Python to see results here.
Status / stderr
Ready to run Python in your browser.
Line-by-line guide
- Line 1Load a Python library needed for data work or regression.
- Line 2Load a Python library needed for data work or regression.
- Line 3Load a Python library needed for data work or regression.
- Line 5Load the dataset into a pandas DataFrame.
- Line 6Add an intercept column to the regression design matrix.
- Line 7Create or update a Python object used in the analysis.
- Line 8Display a result so students can inspect the output.
- Line 9Display a result so students can inspect the output.
Python walkthrough
- 1Load pandas and statsmodels so the workflow is reproducible.
- 2Read the installed Ceteris Lab synthetic CSV from the public data folder.
- 3Create or inspect dummy variables before estimating the model.
- 4Estimate OLS with an intercept and the selected regressors.
- 5Print coefficients or summaries, then interpret them as associations unless the design supports causality.
Live notebook
Run this lesson as a notebook
Open an editable notebook cell-by-cell, run Python in the browser, and download the `.ipynb` file for later.
Related dataset
MODULE7_WAGE_GROUPS_SYNTHETIC
Estimated time
25 to 40 min
Packages
pandas, numpy, statsmodels, patsy
Expected output
Printed Python results that can be compared with the lesson explanation.
Learning goals
- Load and inspect MODULE7_WAGE_GROUPS_SYNTHETIC.
- Run the Python cells connected to Dummy Variables in Log-Dependent Models.
- Interpret the output using dummy variables and qualitative information.
Common errors
- File not found: check that module7_wage_groups_synthetic.csv is installed or use the course data folder.
- Package import error: use the browser notebook first, then download for local Jupyter if your local packages differ.
- Column name error: compare your variable names with the dataset variables listed for this notebook.
Dataset path helper
import pandas as pd
df = pd.read_csv("/data/module-7/module7_wage_groups_synthetic.csv")
df.head()Interactive activity
DummyPercentEffectCalculator
Dummy Variables in Log-Dependent Models
DummyPercentEffectCalculator: choose the coding rule, base group, or probability interpretation before reading the coefficient.
Inputs
Try it yourself
Write one plain-English sentence explaining the main idea from this lesson.
Common mistakes
Check these before you move on.
A regression coefficient describes a pattern unless the assumptions or research design support a causal interpretation.
Quick quiz
What should students check before trusting the result in Dummy Variables in Log-Dependent Models?
Quick quiz
What should students check before trusting the result in Dummy Variables in Log-Dependent Models?
Quick quiz
Why is MODULE7_DISCRETE_OUTCOME_SYNTHETIC a reasonable practice dataset here?
Key takeaway
Dummy Variables in Log-Dependent Models helps students convert qualitative information into transparent, testable regression comparisons.