Lesson 20
Module 7 Capstone
Big question
How do we build a complete qualitative-variable analysis?
Lesson progress
Complete checkpoints as you learn
Learning objectives
- Explain module 7 capstone in plain language.
- Use linear probability model correctly in an interpretation.
- Connect the lesson idea to a formula, graph, Python result, or real example.
Simple explanation
A complete analysis codes categories, chooses a base group, estimates models, tests differences, checks LPM limits, and reports uncertainty.
Key terms
- Linear probability model
- An OLS model with a binary dependent variable interpreted as a probability.
- Categorical variable
- A variable whose values name groups rather than measure amounts.
- Dummy-variable trap
- Perfect collinearity caused by including an intercept and every category indicator.
Core formula
Use plain-language interpretation before algebra.
Example
MODULE7_STUDENT_COMPLETION_SYNTHETIC is an original Ceteris Lab synthetic teaching dataset for module 7 capstone. It lets students practice capstone workflow without presenting fabricated real-world empirical findings.
Interactive visual
Module7CapstoneWorkspace: use the controls to code a group, choose a base group, and write one correct interpretation.
Original Module 7 visual for Module 7 Capstone.
y variable
wage
The dependent variable. It is the outcome students want to explain.
x variable
education
The explanatory variable. It is used to describe changes in wage.
Live Python
Module 7 Capstone Python example
Module 7 Capstone Python example
Stdout
Run Python to see results here.
Status / stderr
Ready to run Python in your browser.
Line-by-line guide
- Line 1Load a Python library needed for data work or regression.
- Line 3Create or update a Python object used in the analysis.
- Line 4Run this Python instruction as part of the lesson workflow.
- Line 5Run this Python instruction as part of the lesson workflow.
- Line 6Run this Python instruction as part of the lesson workflow.
- Line 7Display a result so students can inspect the output.
Python walkthrough
- 1Load pandas and statsmodels so the workflow is reproducible.
- 2Read the installed Ceteris Lab synthetic CSV from the public data folder.
- 3Create or inspect dummy variables before estimating the model.
- 4Estimate OLS with an intercept and the selected regressors.
- 5Print coefficients or summaries, then interpret them as associations unless the design supports causality.
Live notebook
Run this lesson as a notebook
Open an editable notebook cell-by-cell, run Python in the browser, and download the `.ipynb` file for later.
Related dataset
MODULE7_STUDENT_COMPLETION_SYNTHETIC
Estimated time
25 to 40 min
Packages
pandas, numpy, statsmodels, patsy
Expected output
Printed Python results that can be compared with the lesson explanation.
Learning goals
- Load and inspect MODULE7_STUDENT_COMPLETION_SYNTHETIC.
- Run the Python cells connected to Module 7 Capstone.
- Interpret the output using dummy variables and qualitative information.
Common errors
- File not found: check that module7_student_completion_synthetic.csv is installed or use the course data folder.
- Package import error: use the browser notebook first, then download for local Jupyter if your local packages differ.
- Column name error: compare your variable names with the dataset variables listed for this notebook.
Dataset path helper
import pandas as pd
df = pd.read_csv("/data/module-7/module7_student_completion_synthetic.csv")
df.head()Interactive activity
Module7CapstoneWorkspace
Module 7 Capstone
Module7CapstoneWorkspace: choose the coding rule, base group, or probability interpretation before reading the coefficient.
Inputs
Try it yourself
Write one plain-English sentence explaining the main idea from this lesson.
Common mistakes
Check these before you move on.
A regression coefficient describes a pattern unless the assumptions or research design support a causal interpretation.
Quick quiz
What should students check before trusting the result in Module 7 Capstone?
Quick quiz
What should students check before trusting the result in Module 7 Capstone?
Quick quiz
Why is MODULE7_WAGE_GROUPS_SYNTHETIC a reasonable practice dataset here?
Key takeaway
Module 7 Capstone helps students convert qualitative information into transparent, testable regression comparisons.