Lesson 11

Histograms, Normality, and Transformations

Big question

Can a log transformation make inference easier?

Lesson progress

Complete checkpoints as you learn

0% complete0 checkpoint streak
Progress tracking available after sign-in. Sign in to save your work.
Big question
Concept
Activity
Quiz

Learning objectives

  • Explain histograms, normality, and transformations in plain language.
  • Use zero correlation correctly in an interpretation.
  • Connect the lesson idea to a formula, graph, Python result, or real example.

Simple explanation

Histograms of outcomes and residuals help diagnose skewness and nonnormality. Logs can help, but not always.

Key terms

zero correlation
A variable has zero covariance with the error.
asymptotic confidence interval
An interval justified by a large-sample approximation.
n-R-squared statistic
The LM statistic form n times R-squared.

Core formula

Comparehistogramsofy,log(y),residuals,andlogmodelresidualsCompare histograms of y, log(y), residuals, and log-model residuals

Use plain-language interpretation before algebra.

Example

WAGE1 gives students a real-data setting for histograms, normality, and transformations. The lesson reports code and diagnostics only after the student runs the live Python lab.

Interactive visual

HistogramNormalityExplorer

Original Module 5 visual for Histograms, Normality, and Transformations.

wage_sample.csv

y variable

wage

The dependent variable. It is the outcome students want to explain.

x variable

education

The explanatory variable. It is used to describe changes in wage.

Live Python

Histograms, Normality, and Transformations Python example

Histograms, Normality, and Transformations Python example

Stdout

Run Python to see results here.

Status / stderr

Ready to run Python in your browser.

Line-by-line guide

  1. Line 1Load a Python library needed for data work or regression.
  2. Line 2Load a Python library needed for data work or regression.
  3. Line 3Load a Python library needed for data work or regression.
  4. Line 4Load a Python library needed for data work or regression.
  5. Line 6Load the dataset into a pandas DataFrame.
  6. Line 7Keep rows that have the variables required for this model.
  7. Line 8Add an intercept column to the regression design matrix.
  8. Line 9Estimate an ordinary least squares regression.
  9. Line 10Create a log version of the variable so coefficients can be read approximately as percentages.
  10. Line 11Estimate an ordinary least squares regression.
  11. Line 12Create or update a Python object used in the analysis.
  12. Line 13Create or update a Python object used in the analysis.
  13. Line 14Run this Python instruction as part of the lesson workflow.
  14. Line 15Run this Python instruction as part of the lesson workflow.
  15. Line 16Create or update a Python object used in the analysis.
  16. Line 17Run this Python instruction as part of the lesson workflow.
  17. Line 18Run this Python instruction as part of the lesson workflow.
  18. Line 19Run this Python instruction as part of the lesson workflow.
  19. Line 20Display a result so students can inspect the output.
  20. Line 21Display a result so students can inspect the output.

Python walkthrough

  1. 1Load libraries and data or set a simulation seed.
  2. 2Build the model or simulation that matches the lesson question.
  3. 3Compute the statistic, graph, or summary table.
  4. 4Interpret the result as large-sample evidence, not automatic causality.

Live notebook

Run this lesson as a notebook

Open an editable notebook cell-by-cell, run Python in the browser, and download the `.ipynb` file for later.

Related dataset

WAGE1

Estimated time

25 to 40 min

Packages

pandas, numpy, matplotlib

Expected output

A printed summary plus a chart in the output panel.

Learning goals

  • Load and inspect WAGE1.
  • Run the Python cells connected to Histograms, Normality, and Transformations.
  • Interpret the output using histograms and skewness.

Common errors

  • File not found: check that WAGE1.csv is installed or use the course data folder.
  • Package import error: use the browser notebook first, then download for local Jupyter if your local packages differ.
  • Column name error: compare your variable names with the dataset variables listed for this notebook.

Dataset path helper

import pandas as pd

df = pd.read_csv("/data/module-5/WAGE1.csv")
df.head()

Interactive activity

HistogramNormalityExplorer

Choose the histogram to inspect

Compare raw, logged, and residual distributions before making an inference claim.

WAGE1

Inputs

Pick the next report ingredient

Try it yourself

Write one plain-English sentence explaining the main idea from this lesson.

Common mistakes

Check these before you move on.

A regression coefficient describes a pattern unless the assumptions or research design support a causal interpretation.

Quick quiz

What should students remember when reading residual histograms?

Quick quiz

Which reporting habit is most important in Histograms, Normality, and Transformations?

Quick quiz

Why is HPRICE2 a reasonable practice dataset here?

Key takeaway

Histograms, Normality, and Transformations helps students separate large-sample approximation from valid research design.