Lesson 19

Predictions when the Dependent Variable Is Logged

Big question

Why is exp(predicted log y) usually not the best prediction of y?

Lesson progress

Complete checkpoints as you learn

0% complete0 checkpoint streak
Progress tracking available after sign-in. Sign in to save your work.
Big question
Concept
Activity
Quiz

Learning objectives

  • Explain predictions when the dependent variable is logged in plain language.
  • Use semi-elasticity correctly in an interpretation.
  • Connect the lesson idea to a formula, graph, Python result, or real example.

Simple explanation

Exponentiating the predicted log outcome gives a median-style prediction under common assumptions. To predict the mean of y, students need a retransformation adjustment such as a smearing factor.

Key terms

Semi-elasticity
A coefficient interpretation involving a level change in one variable and a percent change in another.
Interaction
A product of variables that lets one slope depend on another variable.
Precision control
A safe control that can reduce residual variation and improve precision.

Core formula

y_hat_mean = exp(log_y_hat) * mean(exp(residual))

Use plain-language interpretation before algebra.

Example

M6_HOUSING_LOGS is a synthetic teaching dataset for predictions when the dependent variable is logged. It is designed to practice log retransformation without presenting fabricated real-world empirical findings.

Interactive visual

Apply a smearing factor and compare naive and adjusted predictions.

Original Module 6 visual for Predictions when the Dependent Variable Is Logged.

wage_sample.csv

y variable

wage

The dependent variable. It is the outcome students want to explain.

x variable

education

The explanatory variable. It is used to describe changes in wage.

Live Python

Predictions when the Dependent Variable Is Logged Python example

Predictions when the Dependent Variable Is Logged Python example

Stdout

Run Python to see results here.

Status / stderr

Ready to run Python in your browser.

Line-by-line guide

  1. Line 1Load a Python library needed for data work or regression.
  2. Line 2Load a Python library needed for data work or regression.
  3. Line 3Load a Python library needed for data work or regression.
  4. Line 5Load the dataset into a pandas DataFrame.
  5. Line 6Add an intercept column to the regression design matrix.
  6. Line 7Create or update a Python object used in the analysis.
  7. Line 8Create or update a Python object used in the analysis.
  8. Line 9Display a result so students can inspect the output.
  9. Line 10Display a result so students can inspect the output.

Python walkthrough

  1. 1Load the synthetic teaching dataset from the Module 6 public data folder.
  2. 2Create transformed variables only after checking their meaning and valid support.
  3. 3Fit a regression that matches the lesson's interpretation target.
  4. 4Print coefficient or prediction summaries that students can connect to the formula.
  5. 5Use comments and output labels so no empirical result is presented without context.

Live notebook

Run this lesson as a notebook

Open an editable notebook cell-by-cell, run Python in the browser, and download the `.ipynb` file for later.

Related dataset

M6_HOUSING_LOGS

Estimated time

25 to 40 min

Packages

pandas, numpy

Expected output

Printed Python results that can be compared with the lesson explanation.

Learning goals

  • Load and inspect M6_HOUSING_LOGS.
  • Run the Python cells connected to Predictions when the Dependent Variable Is Logged.
  • Interpret the output using log prediction and smearing.

Common errors

  • File not found: check that M6_HOUSING_LOGS.csv is installed or use the course data folder.
  • Package import error: use the browser notebook first, then download for local Jupyter if your local packages differ.
  • Column name error: compare your variable names with the dataset variables listed for this notebook.

Dataset path helper

import pandas as pd

df = pd.read_csv("/data/module-6/M6_HOUSING_LOGS.csv")
df.head()

Interactive activity

LogPredictionSmearingLab

Apply smearing correction

Compare naive and smearing-adjusted retransformed predictions.

M6_HOUSING_LOGS

Inputs

Try it yourself

Write one plain-English sentence explaining the main idea from this lesson.

Common mistakes

Check these before you move on.

A regression coefficient describes a pattern unless the assumptions or research design support a causal interpretation.

Quick quiz

Which interpretation is most careful for Predictions when the Dependent Variable Is Logged?

Quick quiz

What is the main mistake to avoid in Predictions when the Dependent Variable Is Logged?

Quick quiz

Why is M6_WAGE_SCALING a reasonable practice dataset here?

Key takeaway

Predictions when the Dependent Variable Is Logged helps students make multiple regression more flexible while keeping interpretation precise and honest.