Tutorial 6: Difference-in-Differences (Finance and Entrepreneurship Tracks)
Tutorial 6
Course evaluation
We would really like you to take the time to fill out the course evaluation form. Your feedback is important to us and helps us improve the course for future students.
You can fill it in on https://entry.caracal.uu.nl/44007 under “Course Evaluations” or scan the QR code:

Recapitulation of the lecture
Causal inference needs a counterfactual
A core challenge in economics is to separate causal effects from correlations. The potential outcomes framework makes the problem precise. For an individual \(i\) and a treatment, say a job training program, \(Y_i(1)\) is the outcome if \(i\) receives the treatment and \(Y_i(0)\) is the outcome if \(i\) does not.
Only one of the two is ever realised. We observe \(Y_i(1)\) for the treated and \(Y_i(0)\) for the untreated, never both for the same individual at the same time. The outcome we do not observe is the counterfactual, which is why the individual causal effect \(\tau_i=Y_i(1)-Y_i(0)\) can never be computed directly and why every empirical strategy is, in the end, a way of filling in a counterfactual.
What a difference in means actually contains
A simple comparison of average outcomes between the treated and the untreated, \(E[Y_i \mid T_i=1] - E[Y_i \mid T_i=0]\), is misleading because it combines the treatment effect with pre-existing differences between the two groups. Adding and subtracting the counterfactual mean \(E[Y_i(0)\mid T_i=1]\) splits it into two interpretable pieces:
\[ \begin{align*} E[Y_i\mid T_i=1] - E[Y_i\mid T_i=0] = &\underbrace{E[Y_i(1) - Y_i(0) \mid T_i=1]}_{\text{ATT}} +\\ &\underbrace{E[Y_i(0)\mid T_i=1] - E[Y_i(0)\mid T_i=0]}_{\text{Selection bias}} \end{align*} \]
Only the first term is the object of interest. The second is zero only if the groups would have looked the same without the treatment.
Difference-in-differences
Difference-in-differences (DiD) removes selection bias in levels by using two groups, treatment and control, observed in two periods, pre and post. Take the change in the outcome for the treatment group, \(\Delta_T\), and subtract the change for the control group over the same period, \(\Delta_C\), which stands in for the trend that the treated group would have followed without the treatment:
\[ \tau_{DiD} = \Delta_T - \Delta_C = (\hat{Y}_{T,Post} - \hat{Y}_{T,Pre}) - (\hat{Y}_{C,Post} - \hat{Y}_{C,Pre}). \]
The permanent gap between the groups drops out of the first difference, and the common trend drops out of the second. What remains is the treatment effect, provided the two groups’ untreated outcomes really would have moved together.
DiD in practice
The same estimate is the coefficient \(\beta_3\) in the regression
\[ Y_{it} = \beta_0 + \beta_1 D_i + \beta_2 Post_t + \beta_3(D_i \times Post_t) + u_{it}, \]
where \(D_i\) is a dummy for the treatment group and \(Post_t\) a dummy for the post-treatment period. Writing the estimator as a regression is what gives us standard errors for inference and room for control variables \(X_{it}\), which make parallel trends a more plausible assumption.
With more than two periods the model extends to an event study that traces the effect over time,
\[ y_{it} = \alpha_i + \lambda_t + \sum_{k=-K}^{L} \delta_k D_{it}^k + u_{it}, \]
where the coefficients \(\delta_k\) for \(k<0\) report the pre-treatment differences that a credible design should not find.
Retrieval practice
Before opening your notes, answer these questions in one sentence each.
- What are the two potential outcomes for individual \(i\), and which of them is the counterfactual?
- Into which two terms does a simple difference in means between treated and untreated groups decompose?
- Write the DiD estimator as a difference of two changes, and say what the control group’s change is meant to represent.
- Which coefficient in the regression \(Y_{it}=\beta_0+\beta_1 D_i+\beta_2 Post_t+\beta_3(D_i\times Post_t)+u_{it}\) is the DiD estimate?
- What does parallel trends assume, and why can it never be verified directly?
If the quiz does not load, open it directly in Wooclap.
Questions
1. The effect of union membership on wages
This question uses the classic nlswork.dta dataset, which contains panel data on young working women. We want to estimate the causal effect of joining a union on wages.1
Define the treatment group as women who were non-union in
year77 but became union members byyear78, and the control group as women who were non-union in both years. Create the dummy variablesTreatandPostfor this 2x2 setup (Pre = 77, Post = 78).Calculate the four means of the 2x2 DiD table by hand (\(\hat{Y}_{T,Pre}\), \(\hat{Y}_{T,Post}\), \(\hat{Y}_{C,Pre}\), \(\hat{Y}_{C,Post}\)) and compute the DiD estimate.
Estimate the DiD effect by running the regression \(Y_{it} = \beta_0 + \beta_1 Treat_i + \beta_2 Post_t + \beta_3(Treat_i \times Post_t) + \epsilon_{it}\). Confirm that \(\hat{\beta}_3\) matches your manual calculation, then report and interpret the result.
2. Event study and parallel trends
Duflo (2001) analyses a major program that built over 61,000 new primary schools across Indonesia. The number of schools built in a region depended on the number of school-aged children living there in 1972, so some regions received many new schools and others very few. The study asks whether children in regions that received more schools ended up with more education and higher wages.2 This supports a difference-in-differences analysis comparing high-construction with low-construction regions, for cohorts that went to school before and after the program started in 1989.
Import the data into your statistical software of choice.
Plot average educational attainment (
yeduc) for the high-construction and low-construction regions in the pre-treatment cohorts (-4 until -1). Does this visual evidence support the parallel trends assumption for a simple DiD analysis around 1989? Discuss what you see in the cohorts leading up to the treatment.Explain what parallel trends means in this context and outline how you would formally test for pre-trends in a regression framework. Which coefficients would you look at, and what would you hope to find?
Conduct that test, normalising \(\beta_{-1}=0\). Report and interpret your findings.
3. Decomposing the naive estimator
The lecture shows that the simple difference-in-means estimator, \(E[Y\mid T=1] - E[Y\mid T=0]\), is a biased estimator of the average treatment effect on the treated (ATT).
Starting from the identity \(E[Y\mid T=1] - E[Y\mid T=0] = E[Y(1)\mid T=1] - E[Y(0)\mid T=0]\), derive the decomposition into the ATT and the selection bias term. Explain each step.
Give an intuitive explanation of the selection bias term, \(E[Y(0)\mid T=1] - E[Y(0)\mid T=0]\), using the lecture’s job training example. What would a positive selection bias imply there?
4. The 2x2 calculation
A researcher evaluates the impact of a new subway line on local house prices. She collects data from a neighbourhood that received the line (treatment) and a similar neighbourhood that did not (control), both before and after the line opened. Average house prices, in thousands of dollars, are
| Before period (Pre) | After period (Post) | |
|---|---|---|
| Treatment group | 450 | 520 |
| Control group | 420 | 460 |
Calculate the difference-in-differences estimate \(\tau_{DiD}\) for the effect of the subway line on house prices, showing your calculations step by step.
What is the secular trend in house prices according to these data?
Based on your result, what do you conclude about the causal effect of the new subway line?
5. Applying potential outcomes
Consider a study of the effect of a new irrigation system (the treatment) on the crop yield of farms. Let \(Y_i\) be the crop yield of farm \(i\).
Define the potential outcomes \(Y_i(1)\) and \(Y_i(0)\) in the context of this example.
What is the individual causal effect \(\tau_i\) for a single farm \(i\)?
Explain the fundamental problem of causal inference using this irrigation example. Why can we not calculate \(\tau_i\) directly for any farm?
6. The parallel trends assumption
The validity of the DiD estimator hinges on the parallel trends assumption.
Write down the parallel trends assumption mathematically, using the potential outcomes notation from the lecture.
Explain in plain English what the assumption means. Why is it essential for DiD to identify the ATT, and which unobservable counterfactual does it allow us to fill in?
If you had data from several years before the treatment was implemented, how could you build visual evidence for the credibility of this assumption?
7. Interpreting DiD regression output
A researcher estimates the effect of a province-level environmental policy on air quality using the model below, where AirQuality is an index, Treat is a dummy for provinces that adopted the policy, and Post is a dummy for the period after adoption:
\[ Y_{it} = \beta_0 + \beta_1 Treat_i + \beta_2 Post_t + \beta_3(Treat_i \times Post_t) + \epsilon_{it}. \]
The estimated model is
\[ \widehat{\text{AirQuality}}_{it} = 75.2 + 5.5 \cdot Treat_i - 8.1 \cdot Post_t - 4.3 \cdot (Treat_i \times Post_t). \]
What is the average air quality for the control group in the pre-treatment period?
What does the coefficient \(\beta_1 = 5.5\) represent? Does it indicate that the policy was assigned randomly?
What does the coefficient \(\beta_2 = -8.1\) represent?
What is the DiD estimate of the policy’s effect? Interpret \(\beta_3 = -4.3\) in the context of the study.
Self-study
8. Threats to identification
Imagine you are designing a study to evaluate the impact of a scholarship program, awarded to students in City A, on university enrollment rates. City B, which has no such program, serves as your control.
Policy anticipation. How might anticipation effects violate the parallel trends assumption? Give a concrete example related to this scholarship program.
Spillovers. How might spillover effects contaminate your control group? Give a concrete example.
Ashenfelter’s dip. Describe what Ashenfelter’s dip would look like in this context and why it would lead to a biased estimate of the program’s effect.
9. Direction of bias
A researcher evaluates the effect of a new fertilizer on crop yields, using a DiD design that compares farms which adopted the fertilizer (treatment) to farms which did not (control).
After the results are presented, a skeptical audience member points to the pre-treatment trend graph. It shows that in the years leading up to adoption the yields of the treatment farms were already declining, while the yields of the control farms were stable.
Does this observation support or violate the parallel trends assumption?
Assuming the researcher found a positive DiD estimate, so that the fertilizer appears to increase yields, is this estimate likely an overestimate or an underestimate of the true causal effect? Write the DiD estimand in potential outcomes notation, isolate the bias term, and sign it using what the graph tells you about the counterfactual trend for the treatment group.
Takeaways
What should you be able to do now?
- State a causal effect in potential-outcomes notation and show why a simple difference in means mixes the ATT with selection bias.
- Build a 2x2 difference-in-differences estimate by hand from four group means, and reproduce exactly that number as an interaction coefficient in a regression.
- Estimate a DiD and an event-study specification in R or Python, and interpret every coefficient in the model rather than only the interaction.
- Say precisely what parallel trends assumes, test for pre-trends, and explain why passing that test is supporting evidence rather than proof.
- Name concrete threats to identification – anticipation, spillovers, Ashenfelter’s dip – and reason about the direction in which each would bias the estimate.