Assessing Strong Parallel Trends via Close Comparison Groups

Session 2

Brantly Callaway

University of Georgia

September 18, 2026

Introduction

\(\newcommand{\E}{\mathbb{E}} \newcommand{\var}{\mathrm{var}} \newcommand{\cov}{\mathrm{cov}} \newcommand{\Var}{\mathrm{var}} \newcommand{\Cov}{\mathrm{cov}} \newcommand{\Corr}{\mathrm{corr}} \newcommand{\corr}{\mathrm{corr}} \newcommand{\L}{\mathrm{L}} \renewcommand{\P}{\mathrm{P}} \newcommand{\independent}{{\perp\!\!\!\perp}} \newcommand{\indicator}[1]{ \mathbf{1}\{#1\} } \newcommand{\ATT}{\text{ATT}} \newcommand{\ACR}{\text{ACR}}\)Session 1: interpreting comparisons across doses as causal effects required strong parallel trends — a substantially stronger assumption than most researchers realize they are invoking

  • We discussed some (sub-)cases where it might be plausible and/or weakened, but it seemed fundamentally difficult to assess as only untreated potential outcomes are observed in pre-treatment periods

This session: can we use pre-treatment data to say anything about whether (local) strong parallel trends is plausible in a given application?

  • Main idea: check for close comparison groups — neighboring dose groups that have the same level of outcomes in pre-treatment periods, argue that strong parallel trends is more plausible for these groups

Plan for This Session




  1. Setup and discussion of strong parallel trends

  2. Close comparison groups

  3. Empirical application

Setup

Same staggered panel data continuous treatment design as Session 1:

  • Two periods \(t=1,2\);
  • No one treated at \(t=1\);
  • Continuous dose \(D_i\) realized at \(t=2\)

School Closures Example

Back to Session 1’s running example where \(dose\) = length of Covid school closure, \(outcome\) = test scores

  • \(\eta_i\): school’s unobserved student ability

  • \(\theta_{t=2,d}\): the ability-neutral part of the dose response, e.g., the component of lost instructional time/disruption that hits any school closed for length \(d\), regardless of its students’ baseline ability

  • \(\beta_{t=2,d}\): how strongly baseline ability \(\eta_i\) translates into test scores, given closure length \(d\) — i.e., varies with \(d\) if high ability schools respond differently to different lengths of closures than low ability schools

    • Would be increasing in \(d\) if high ability schools are more resilient to long closures than low ability schools.

2. Close Comparison Groups

Background

In Callaway et al. (2025), we introduce a notion of “close comparison groups”

  • Setting of that paper: binary treatment, multiple pre-treatment periods, multiple comparison groups, looking for a way to relax the parallel trends assumption, and we look for pockets of simple identification

  • Groups that are “alike” in terms of pre-treatment levels of outcomes are “close comparison groups”

  • If you can find such groups, you can get very robust identification of causal effects (i.e., robustness to the identification strategy)

CCGs Adapted to a Continuous Treatment Setting

As discussed above, difference-in-differences attempts to allow for selection on time-invariant unobservables by differencing them out.

But we can look at the level of pre-treatment outcomes to inform us about the distribution of unobservables across dose groups.

Same Pre-Treatment Mean Implies Same Mean of

For some dose group \(l\), its mean of untreated potential outcomes is given by:

\[ \begin{aligned} \E[Y_{t=1} \mid D=l] &= \E[Y_{t=1}(0) \mid D=l] \hspace{150pt} \end{aligned} \]

Same Pre-Treatment Mean Implies Same Mean of

For some dose group \(l\), its mean of untreated potential outcomes is given by:

\[ \begin{aligned} \E[Y_{t=1} \mid D=l] &= \E[Y_{t=1}(0) \mid D=l] \hspace{150pt}\\ &= \theta_1 + \E[\eta \mid D=l] \end{aligned} \]

Since \(\theta_1\) is the same for every \(l\):

\[\E[Y_{t=1} \mid D=l] = \E[Y_{t=1} \mid D=d] \quad\iff\quad \E[\eta \mid D=l] = \E[\eta \mid D=d]\]

In other words, dose groups with the same pre-treatment mean of outcomes have the same mean of unobservables \(\eta_i\)

If they have the same mean of \(\eta_i\), then the selection-bias terms above vanish, and strong parallel trends holds (at least locally)

Implications

In a similar fashion to using pre-treatment periods to validate the parallel trends assumption, we can use pre-treatment periods to validate the “no selection”

  • We can do this globally,
  • Find regions where the pre-treatment mean is flat, or
  • Find regions where the pre-treatment mean is flat conditional on covariates

It rationalizes interpreting those regions as being regions where strong parallel trends is plausible, and suggests that we can recover locally comparable causal effects in those regions, e.g., \(\ACR(d)\) in these regions.

3. Empirical Illustration

Application

Same empirical setting as Session 1:

  • Acemoglu and Finkelstein (2008)’s study of the 1983 Medicare PPS reform.
  • Dose = hospital’s pre-reform Medicare share
  • outcome = capital-labor ratio (depreciation share)
  • In session 1, we mainly looked at average outcomes before and after the treatment, but here we will use adjacent years instead, using 1982 as the base period.
    • Treatment effects: 1982 vs. 1983, 1984, 1985, 1986
    • Plus two more placebos marching up to 1982: 1980 vs. 1981, 1981 vs. 1982

Results: ATT(d|d)

Results: ATT(d|d)

Results: ATT(d|d)

Results: ACR(d)

Results: ACR(d)

Results: ACR(d)

Pre-Test: ATT(d|d)

Pre-Test: ATT(d|d)

Pre-Test: ATT(d|d)

Pre-Test: ATT(d|d)

Two of three fail — 1980-81 is close to 0, but 1981-82 (\(\widehat{\ATT}=-0.22\)) and 1982-83 (\(\widehat{\ATT}=+0.20\)) are both significant, in opposite directions

  • 1982-83 could reflect anticipation: PPS was legislated in April 1983, months before its October 1983 effective date
  • 1981-82 has no such story — both years are firmly pre-treatment. This looks like genuine year-to-year instability, not just anticipation

Pre-Test: Same Pre-Treatment Levels

Next, we check for regions of close comparison groups

Pre-Test: Same Pre-Treatment Levels

Pre-Test: Same Pre-Treatment Levels

Pre-Test: Same Pre-Treatment Levels

Pre-Test: Same Pre-Treatment Levels

The flat region is remarkably stable across three separate years: roughly dose \(0.35\) to \(0.70\) in every one of 1980, 1981, and 1982

  • This is what we’d expect if the flatness reflects a persistent hospital characteristic (\(\eta_i\)), not a fluke of any one year
  • Use 1982’s (largest) flat region as a single canonical “close comparison group” definition, applied everywhere next

ACR(d): Close Comparison Region

Same ACR(d) curves as before, now with the canonical region (dose \(0.35\)-\(0.70\)) highlighted, plus its density-weighted average (i.e., a local \(\ACR\) in the close comparison region)

ACR(d): Close Comparison Region

ACR(d): Close Comparison Region

ACR(d): Close Comparison Region

ACR(d): Close Comparison Region

ACR(d): Close Comparison Region

ACR(d): Close Comparison Region

Summary of Application

It appears that:

  1. Parallel trends seems to be violated in pre-treatment periods
  2. Checking the levels of pre-treatment outcomes as a function of the dose suggested that there were pockets of no selection
  3. Cross-dose comparisons are more likely to deliver causal effects in those pockets (i.e., a local version of strong parallel trends)
  4. The estimates were noisy, but often negative. They don’t provide any evidence that the Medicare policy increased capital labor ratios.

Conclusion

Using pre-treatment periods to assess the plausibility of strong parallel trends is challenging because only untreated potential outcomes are observed in those periods.

However, the levels of outcomes in pre-treatment periods across different dose groups provides evidence about difference in the unobserved heterogeneity driving their outcomes.

If differences in observed pre-treatment levels of outcomes are small, it suggests that differences in unobservables are small too.

If differences in unobservables are small, then strong parallel trends is more plausible, and cross-dose comparisons of trends in outcomes can be interpreted as approximately causal in those regions.

Appendix

References

Acemoglu, Daron, and Amy Finkelstein. 2008. “Input and Technology Choices in Regulated Industries: Evidence from the Health Care Sector.” Journal of Political Economy 116 (5): 837–80.
Angrist, Joshua D, and Alan B Krueger. 1999. “Empirical Strategies in Labor Economics.” In Handbook of Labor Economics, vol. 3. Elsevier.
Ashenfelter, Orley, and David Card. 1985. “Using the Longitudinal Structure of Earnings to Estimate the Effect of Training Programs.” The Review of Economics and Statistics 67 (4): 648–60.
Callaway, Brantly, Derek Dyal, Pedro HC Sant’Anna, and Emmanuel S Tsyawo. 2025. “Beyond Parallel Trends: An Identification-Strategy-Robust Approach to Causal Inference with Panel Data.” Unpublished manuscript.
Ghanem, Dalia, Pedro H. C. Sant’Anna, and Kaspar Wüthrich. 2024. “Selection and Parallel Trends.” Unpublished manuscript.
Hirano, Keisuke, and Guido W Imbens. 2004. “The Propensity Score with Continuous Treatments.” Applied Bayesian Modeling and Causal Inference from Incomplete-Data Perspectives 226164: 73–84.
Marx, Philip, Elie Tamer, and Xun Tang. 2025. “Heterogeneous Treatment Effects via Linear Dynamic Panel Data Models.” Unpublished manuscript.