Session 2
September 18, 2026
\(\newcommand{\E}{\mathbb{E}} \newcommand{\var}{\mathrm{var}} \newcommand{\cov}{\mathrm{cov}} \newcommand{\Var}{\mathrm{var}} \newcommand{\Cov}{\mathrm{cov}} \newcommand{\Corr}{\mathrm{corr}} \newcommand{\corr}{\mathrm{corr}} \newcommand{\L}{\mathrm{L}} \renewcommand{\P}{\mathrm{P}} \newcommand{\independent}{{\perp\!\!\!\perp}} \newcommand{\indicator}[1]{ \mathbf{1}\{#1\} } \newcommand{\ATT}{\text{ATT}} \newcommand{\ACR}{\text{ACR}}\)Session 1: interpreting comparisons across doses as causal effects required strong parallel trends — a substantially stronger assumption than most researchers realize they are invoking
This session: can we use pre-treatment data to say anything about whether (local) strong parallel trends is plausible in a given application?
Setup and discussion of strong parallel trends
Close comparison groups
Empirical application
Same staggered panel data continuous treatment design as Session 1:
The most common way to “more primitively” motivate the parallel trends assumption is as a combination of the following two assumptions (Ashenfelter and Card 1985; Angrist and Krueger 1999; Ghanem et al. 2024; Marx et al. 2025):
Latent Unconfoundedness
\[\big( Y_{it=1}(0), Y_{it=2}(0) \big) \;\independent\; D_i \;\mid\; \eta_i\]
where \(\eta_i\) is time-invariant unobserved heterogeneity (i.e., a unit fixed effect)
The most common way to “more primitively” motivate the parallel trends assumption is as a combination of the following two assumptions (Ashenfelter and Card 1985; Angrist and Krueger 1999; Ghanem et al. 2024; Marx et al. 2025):
Latent Unconfoundedness
\[\big( Y_{it=1}(0), Y_{it=2}(0) \big) \;\independent\; D_i \;\mid\; \eta_i\]
where \(\eta_i\) is time-invariant unobserved heterogeneity (i.e., a unit fixed effect)
The most common way to “more primitively” motivate the parallel trends assumption is as a combination of the following two assumptions (Ashenfelter and Card 1985; Angrist and Krueger 1999; Ghanem et al. 2024; Marx et al. 2025):
Linear Model for Untreated Potential Outcomes
\[Y_{it}(0) = \theta_t + \eta_i + e_{it}\]
We can extend the notion of latent unconfoundedness as follows:
Latent Unconfoundedness for Dose
\[\big\{Y_{i,t=1}(0),\, Y_{i,t=2}(d) : d \in \mathcal{D}\big\} \;\independent\; D_i \;\mid\; \eta_i\]
We can extend the notion of latent unconfoundedness as follows:
Latent Unconfoundedness for Dose
\[\big\{Y_{i,t=1}(0),\, Y_{i,t=2}(d) : d \in \mathcal{D}\big\} \;\independent\; D_i \;\mid\; \eta_i\]
We can also consider linear models for all potential outcomes:
Linear Model for Potential Outcomes
\[Y_{i,t}(d) = \theta_{t,d} + \beta_{t,d} \,\eta_i + e_{i,t}(d)\]
Back to Session 1’s running example where \(dose\) = length of Covid school closure, \(outcome\) = test scores
\(\eta_i\): school’s unobserved student ability
\(\theta_{t=2,d}\): the ability-neutral part of the dose response, e.g., the component of lost instructional time/disruption that hits any school closed for length \(d\), regardless of its students’ baseline ability
\(\beta_{t=2,d}\): how strongly baseline ability \(\eta_i\) translates into test scores, given closure length \(d\) — i.e., varies with \(d\) if high ability schools respond differently to different lengths of closures than low ability schools
Applying the same argument as in Session 1: \[ \begin{aligned} \E[\Delta Y \mid D=d_h] - \E[\Delta Y \mid D=d_l] &= \ATT(d_h|d_h) - \ATT(d_l|d_l) \\ &= \underbrace{(\theta_{t=2,d_h} - \theta_{t=2,d_l}) + (\beta_{t=2,d_h} - \beta_{t=2,d_l}) \E[\eta \mid D=d_h]}_{\textrm{Causal Response}} \\ &+ \underbrace{\beta_{t=2,d_l} \Big( \E[\eta \mid D=d_h] - \E[\eta \mid D=d_l] \Big)}_{\textrm{Selection Bias}} \end{aligned} \]
Learning about \(\beta_{t=2,d_l}\) directly does not seem possible—e.g., \(\E[\Delta Y \mid D=d_l] - \E[\Delta Y \mid D=0]\) includes \(\beta_{t=2,d_l}\) but also differences in \(\E[\eta \mid D=d_l]\) and \(\E[\eta \mid D=0]\), which are unobserved
Making assumptions about \(\beta_{t=2,d_l}\) would also be difficult to justify and infeasible to check in practice
Applying the same argument as in Session 1: \[ \begin{aligned} \E[\Delta Y \mid D=d_h] - \E[\Delta Y \mid D=d_l] &= \ATT(d_h|d_h) - \ATT(d_l|d_l) \\ &= \underbrace{(\theta_{t=2,d_h} - \theta_{t=2,d_l}) + (\beta_{t=2,d_h} - \beta_{t=2,d_l}) \E[\eta \mid D=d_h]}_{\textrm{Causal Response}} \\ &+ \underbrace{\beta_{t=2,d_l} \Big( \E[\eta \mid D=d_h] - \E[\eta \mid D=d_l] \Big)}_{\textrm{Selection Bias}} \end{aligned} \]
In Callaway et al. (2025), we introduce a notion of “close comparison groups”
Setting of that paper: binary treatment, multiple pre-treatment periods, multiple comparison groups, looking for a way to relax the parallel trends assumption, and we look for pockets of simple identification
Groups that are “alike” in terms of pre-treatment levels of outcomes are “close comparison groups”
If you can find such groups, you can get very robust identification of causal effects (i.e., robustness to the identification strategy)
As discussed above, difference-in-differences attempts to allow for selection on time-invariant unobservables by differencing them out.
But we can look at the level of pre-treatment outcomes to inform us about the distribution of unobservables across dose groups.
For some dose group \(l\), its mean of untreated potential outcomes is given by:
\[ \begin{aligned} \E[Y_{t=1} \mid D=l] &= \E[Y_{t=1}(0) \mid D=l] \hspace{150pt} \end{aligned} \]
For some dose group \(l\), its mean of untreated potential outcomes is given by:
\[ \begin{aligned} \E[Y_{t=1} \mid D=l] &= \E[Y_{t=1}(0) \mid D=l] \hspace{150pt}\\ &= \theta_1 + \E[\eta \mid D=l] \end{aligned} \]
Since \(\theta_1\) is the same for every \(l\):
\[\E[Y_{t=1} \mid D=l] = \E[Y_{t=1} \mid D=d] \quad\iff\quad \E[\eta \mid D=l] = \E[\eta \mid D=d]\]
In other words, dose groups with the same pre-treatment mean of outcomes have the same mean of unobservables \(\eta_i\)…
If they have the same mean of \(\eta_i\), then the selection-bias terms above vanish, and strong parallel trends holds (at least locally)
In a similar fashion to using pre-treatment periods to validate the parallel trends assumption, we can use pre-treatment periods to validate the “no selection”
It rationalizes interpreting those regions as being regions where strong parallel trends is plausible, and suggests that we can recover locally comparable causal effects in those regions, e.g., \(\ACR(d)\) in these regions.
Same empirical setting as Session 1:
Two of three fail — 1980-81 is close to 0, but 1981-82 (\(\widehat{\ATT}=-0.22\)) and 1982-83 (\(\widehat{\ATT}=+0.20\)) are both significant, in opposite directions
Next, we check for regions of close comparison groups
The flat region is remarkably stable across three separate years: roughly dose \(0.35\) to \(0.70\) in every one of 1980, 1981, and 1982
Same ACR(d) curves as before, now with the canonical region (dose \(0.35\)-\(0.70\)) highlighted, plus its density-weighted average (i.e., a local \(\ACR\) in the close comparison region)
It appears that:
Using pre-treatment periods to assess the plausibility of strong parallel trends is challenging because only untreated potential outcomes are observed in those periods.
However, the levels of outcomes in pre-treatment periods across different dose groups provides evidence about difference in the unobserved heterogeneity driving their outcomes.
If differences in observed pre-treatment levels of outcomes are small, it suggests that differences in unobservables are small too.
If differences in unobservables are small, then strong parallel trends is more plausible, and cross-dose comparisons of trends in outcomes can be interpreted as approximately causal in those regions.