Session 1
September 18, 2026
\(\newcommand{\E}{\mathbb{E}} \newcommand{\E}{\mathbb{E}} \newcommand{\var}{\mathrm{var}} \newcommand{\cov}{\mathrm{cov}} \newcommand{\Var}{\mathrm{var}} \newcommand{\Cov}{\mathrm{cov}} \newcommand{\Corr}{\mathrm{corr}} \newcommand{\corr}{\mathrm{corr}} \newcommand{\L}{\mathrm{L}} \renewcommand{\P}{\mathrm{P}} \newcommand{\ACR}{\mathrm{ACR}} \newcommand{\ATT}{\mathrm{ATT}} \newcommand{\ACRT}{\mathrm{ACRT}} \newcommand{\independent}{{\perp\!\!\!\perp}} \newcommand{\indicator}[1]{ \mathbf{1}\{#1\} }\)
Difference-in-differences with a continuous treatment (or multi-valued discrete) treatment
Main reference: Callaway et al. (2026)
Most recent work on DiD has focused on settings with a binary, staggered treatment
This session: Move from a binary treatment to a continuous treatment (“dose”)
Some of the arguments involve extending ideas from the binary, staggered treatment case to a setting with continuous treatment
Running Example:
For today, we mostly emphasize a continuous treatment, but our results also apply to other settings (with trivial modifications):
Multi-valued treatments (e.g., effect of state-level minimum wage policies on employment)
Binary policy variable with differential exposure (application in the paper: a binary Medicare policy where different hospitals had more exposure to the policy)
Fuzzy DiD - binary treatment but the researcher observes aggregate data, often using policy changes as an instrument
We will focus on a setting with a panel data pre-post research design
We will return to extensions where
Identification: What’s the same as for a binary treatment?
Identification: What’s different from a binary treatment?
Reverse Engineering TWFE Regressions
Empirical Application
Extensions
Two time periods: \(t=1\) and \(t=2\)
Treatment: \(D_i\)
Potential outcomes: \(Y_{it}(d)\)
Observed outcomes: \(Y_{it=2}\) and \(Y_{it=1}\)
\[Y_{it=2}=Y_{it=2}(D_i) \quad \textrm{and} \quad Y_{it=1}=Y_{it=1}(0)\]
Level Effects (Average Treatment Effect on the Treated)
\[\ATT(d \mid d) := \E[Y_{t=2}(d) - Y_{t=2}(0) \mid D=d]\]
Interpretation: The average effect of dose \(d\) relative to not being treated local to the group that actually experienced dose \(d\)
This is the natural analogue of \(\ATT\) in the binary treatment case
Slope Effects (Average Causal Response on the Treated)
\[\ACRT(d \mid d) := \frac{\partial \ATT(l \mid d)}{\partial l} \Big|_{l=d}\]
Notice that \(\ATT(d \mid d)\) and \(\ACRT(d \mid d)\) are functional parameters
We can view \(\ATT(d \mid d)\) and \(\ACRT(d \mid d)\) as the “building blocks” for a more aggregated parameter.
Aggregated versions of these (into a single number) are \[\begin{align*} \ATT^{\texttt{loc}} := \E\Big[\ATT(D \mid D)\Bigm|D>0\Big] \qquad \qquad \ACRT^{\texttt{loc}} := \E\Big[\ACRT(D \mid D)\Bigm|D>0\Big] \end{align*}\]
\(\ATT^{\texttt{loc}}\) averages \(\ATT(d \mid d)\) over the population distribution of the dose
\(\ACRT^{\texttt{loc}}\) averages \(\ACRT(d \mid d)\) over the population distribution of the dose
Parallel Trends Assumption
For all \(d \in \mathcal{D}_+\),
\[\E[\Delta Y(0) \mid D=d] = \E[\Delta Y(0) \mid D=0]\]
Parallel Trends Assumption
For all \(d \in \mathcal{D}_+\),
\[\E[\Delta Y(0) \mid D=d] = \E[\Delta Y(0) \mid D=0]\]
Parallel Trends Assumption
For all \(d \in \mathcal{D}_+\),
\[\E[\Delta Y(0) \mid D=d] = \E[\Delta Y(0) \mid D=0]\]
Under parallel trends,
\[ \begin{aligned} \ATT(d \mid d) &= \E[Y_{t=2}(d) - Y_{t=2}(0) \mid D=d] \hspace{150pt} \end{aligned} \]
Parallel Trends Assumption
For all \(d \in \mathcal{D}_+\),
\[\E[\Delta Y(0) \mid D=d] = \E[\Delta Y(0) \mid D=0]\]
Under parallel trends,
\[ \begin{aligned} \ATT(d \mid d) &= \E[Y_{t=2}(d) - Y_{t=2}(0) \mid D=d] \hspace{150pt}\\ &= \E[Y_{t=2}(d) - Y_{t=1}(0) \mid D=d] - \E[Y_{t=2}(0) - Y_{t=1}(0) \mid D=d] \end{aligned} \]
Parallel Trends Assumption
For all \(d \in \mathcal{D}_+\),
\[\E[\Delta Y(0) \mid D=d] = \E[\Delta Y(0) \mid D=0]\]
Under parallel trends,
\[ \begin{aligned} \ATT(d \mid d) &= \E[Y_{t=2}(d) - Y_{t=2}(0) \mid D=d] \hspace{150pt}\\ &= \E[Y_{t=2}(d) - Y_{t=1}(0) \mid D=d] - \E[Y_{t=2}(0) - Y_{t=1}(0) \mid D=d]\\ &= \E[\Delta Y \mid D=d] - \E[\Delta Y \mid D=0] \end{aligned} \]
This is exactly what you would expect
Parallel Trends Assumption
For all \(d \in \mathcal{D}_+\),
\[\E[\Delta Y(0) \mid D=d] = \E[\Delta Y(0) \mid D=0]\]
Using similar arguments, you can show \[\ATT^{\texttt{loc}} = \E[\Delta Y \mid D>0] - \E[\Delta Y \mid D=0]\]
Unfortunately, no
Most empirical work with a continuous treatment wants to think about how causal responses vary across dose
Plot treatment effects as a function of dose and ask: does more dose cause outcomes to increase/decrease/no effect?
Important because distinguishing between different theories often involves making comparisons across doses
Unfortunately, no
Most empirical work with a continuous treatment wants to think about how causal responses vary across dose
Unfortunately, no
Most empirical work with a continuous treatment wants to think about how causal responses vary across dose
There are new issues related to comparing \(\ATT(d \mid d)\) at different doses and interpreting these differences as causal effects
Consider comparing \(\ATT(d \mid d)\) for two different doses
\[ \begin{aligned} & \ATT(d_h \mid d_h) - \ATT(d_l \mid d_l) \hspace{350pt} \end{aligned} \]
Consider comparing \(\ATT(d \mid d)\) for two different doses
\[ \begin{aligned} & \ATT(d_h \mid d_h) - \ATT(d_l \mid d_l) \hspace{350pt}\\ & \hspace{25pt} = \E[Y_{t=2}(d_h)-Y_{t=2}(d_l) \mid D=d_h] + \E[Y_{t=2}(d_l) - Y_{t=2}(0) \mid D=d_h] - \E[Y_{t=2}(d_l) - Y_{t=2}(0) \mid D=d_l] \end{aligned} \]
Consider comparing \(\ATT(d \mid d)\) for two different doses
\[ \begin{aligned} & \ATT(d_h \mid d_h) - \ATT(d_l \mid d_l) \hspace{350pt}\\ & \hspace{25pt} = \E[Y_{t=2}(d_h)-Y_{t=2}(d_l) \mid D=d_h] + \E[Y_{t=2}(d_l) - Y_{t=2}(0) \mid D=d_h] - \E[Y_{t=2}(d_l) - Y_{t=2}(0) \mid D=d_l]\\ & \hspace{25pt} = \underbrace{\E[Y_{t=2}(d_h) - Y_{t=2}(d_l) \mid D=d_h]}_{\textrm{Causal Response}} + \underbrace{\ATT(d_l \mid d_h) - \ATT(d_l \mid d_l)}_{\textrm{Selection Bias}} \end{aligned} \]
“Standard” Parallel Trends is not strong enough to rule out the selection bias terms here
Implication: If you want to interpret differences in treatment effects across different doses, then you will need stronger assumptions than standard parallel trends
This problem spills over into identifying \(\ACRT(d|d)\)
Intuition
Difference-in-differences identification strategies result in \(\ATT(d|d)\) parameters. These are local parameters and difficult to compare to each other
This explanation is similar to thinking about LATEs with two different instruments
Thus, comparing \(\ATT(d \mid d)\) across different values is tricky and not for free
What can you do?
One idea, just recover \(\ATT(d \mid d)\) and interpret it cautiously (interpret it by itself not relative to different values of \(d\))
If you want to compare them to each other, it will come with the cost of additional (structural) assumptions
To be able to make comparisons across different doses, we need target parameters that are “less local”:
\[ \ATT(d) := \E[Y_{t=2}(d) - Y_{t=2}(0) \mid D>0] \qquad \qquad \ACRT(d) := \frac{\partial \ATT(d)}{\partial d} \]
These involve the level effect or slope effect at \(d\) (like the previous parameters), but they are causal effect parameters for the entire treated population (not just the local group that actually experienced dose \(d\))
We can also consider aggregated versions of these parameters:
\[\ATT^{\texttt{glob}} = \E[ATT(D) \mid D>0] \quad \quad \ACRT^{\texttt{glob}} = \E[\ACRT(D) \mid D>0]\]
Strong Parallel Trends Assumption
For all doses \(d \in \mathcal{D}\),
\[\E[Y_{t=2}(d) - Y_{t=1}(0) \mid D>0] = \E[Y_{t=2}(d) - Y_{t=1}(0) \mid D=d]\]
If we maintain standard parallel trends, strong parallel trends is equivalent to a certain restriction on treatment effect heterogeneity. Notice:
\[ \begin{aligned} \ATT(d|d) &= \E[Y_{t=2}(d) - Y_{t=1}(0) \mid D=d] - \E[Y_{t=2}(0) - Y_{t=1}(0) \mid D=d] \hspace{200pt} \ \end{aligned} \]
If we maintain standard parallel trends, strong parallel trends is equivalent to a certain restriction on treatment effect heterogeneity. Notice:
\[ \begin{aligned} \ATT(d|d) &= \E[Y_{t=2}(d) - Y_{t=1}(0) \mid D=d] - \E[Y_{t=2}(0) - Y_{t=1}(0) \mid D=d] \hspace{200pt} \\\ &= \E[Y_{t=2}(d) - Y_{t=1}(0) \mid D>0] - \E[Y_{t=2}(0) - Y_{t=1}(0) \mid D > 0] \ \end{aligned} \]
If we maintain standard parallel trends, strong parallel trends is equivalent to a certain restriction on treatment effect heterogeneity. Notice:
\[ \begin{aligned} \ATT(d|d) &= \E[Y_{t=2}(d) - Y_{t=1}(0) \mid D=d] - \E[Y_{t=2}(0) - Y_{t=1}(0) \mid D=d] \hspace{200pt} \\\ &= \E[Y_{t=2}(d) - Y_{t=1}(0) \mid D>0] - \E[Y_{t=2}(0) - Y_{t=1}(0) \mid D > 0] \\\ &= \E[Y_{t=2}(d) - Y_{t=2}(0) \mid D>0] = \ATT(d) \end{aligned} \]
Thus, under strong parallel trends, we have that
\[\ATT(d) = \E[\Delta Y|D=d] - \E[\Delta Y|D=0]\]
RHS is exactly the same expression as for \(\ATT(d|d)\) under standard PT, but here
ATT(d) parameters do not suffer from the same issues as ATT(d|d) parameters when making comparisons across dose
\[ \begin{aligned} \ATT(d_h) - \ATT(d_l) &= \E[Y_{t=2}(d_h) - Y_{t=2}(0) | D>0] - \E[Y_{t=2}(d_l) - Y_{t=2}(0)|D>0] \end{aligned} \]
ATT(d) parameters do not suffer from the same issues as ATT(d|d) parameters when making comparisons across dose
\[ \begin{aligned} \ATT(d_h) - \ATT(d_l) &= \E[Y_{t=2}(d_h) - Y_{t=2}(0) | D>0] - \E[Y_{t=2}(d_l) - Y_{t=2}(0)|D>0]\\ &= \underbrace{\E[Y_{t=2}(d_h) - Y_{t=2}(d_l)|D>0]}_{\textrm{Causal Response}} \end{aligned} \]
Thus, recovering \(\ATT(d)\) side-steps the issues about comparing treatment effects across doses, but it comes at the cost of needing a (potentially very strong) extra assumption
Given that we can compare \(\ATT(d)\)’s across dose, we can recover slope effects in this setting
\[ \begin{aligned} \ACRT(d) := \frac{\partial \ATT(d)}{\partial d} \qquad &\textrm{or} \qquad \ACRT^{\texttt{glob}} := \E\Big[\ACRT(D) \Big| D>0\Big] \end{aligned} \]
Can You Relax Strong Parallel Trends?
No Untreated Units
Binarizing the Treatment
Pre-Testing
Some ideas:
It’s possible to do some versions of DiD with a continuous treatment without having access to a fully untreated group.
In this case, it is not possible to recover level effects like \(\ATT(d|d)\).
However, notice that \[\begin{aligned}& \E[\Delta Y | D=d_h] - \E[\Delta Y | D=d_l] \\ &\hspace{20pt}= \Big(\E[\Delta Y | D=d_h] - \E[\Delta Y(0) | D=d_h]\Big) - \Big(\E[\Delta Y | D=d_l]-\E[\Delta Y(0) | D=d_l]\Big) \\ &\hspace{20pt}= \ATT(d_h|d_h) - \ATT(d_l|d_l)\end{aligned}\]
In words: comparing path of outcomes for those that experienced dose \(d_h\) to path of outcomes among those that experienced dose \(d_l\) (and not relying on having an untreated group) delivers the difference between their \(\ATT\)’s.
Still face issues related to selection bias / strong parallel trends though
Strategies like binarizing the treatment can still work (though be careful!)
If you classify units as being treated or untreated, you can recover \(\ATT^{\texttt{loc}}\) by comparing mean outcome for treated relative to untreated.
On the other hand, if you classify units as being “high” treated, “low” treated, or untreated — our arguments imply that selection bias terms can come up when comparing effects for “high” to “low”
That the expressions for \(\ATT(d)\) and \(\ATT(d|d)\) are exactly the same also means that we cannot use pre-treatment periods to try to distinguish between “standard” and “strong” parallel trends.
Or, alternatively, we only observe untreated potential outcomes in pre-treatment periods, so pre-treatment periods do not seem useful for validating strong parallel trends.
It is straightforward/familiar to identify \(\ATT(d \mid d)\) parameters with a multi-valued or continuous dose
However, comparisons of \(\ATT(d\mid d)\) parameters across different doses are hard to interpret
This suggests targeting \(\ATT(d)\) parameters
Consider the same TWFE regression as before: \[\begin{align*} Y_{it} = \theta_t + \eta_i + \beta^{twfe} D_{it} + e_{it} \end{align*}\] We show that \[\begin{align*} \beta^{twfe} = \int_{\mathcal{D}_+} w(l) m'_\Delta(l) \, dl \end{align*}\] where \(m_\Delta(l) := \E[\Delta Y \mid D=l] - \E[\Delta Y \mid D=0]\) and \(w(l)\) are weights
About the weights: they are all positive, but have some strange properties (e.g., always maximized at \(l = \E[D]\) (even if this is not a common value for the dose))
Other issues can arise in more complicated cases
For example, suppose you have a staggered continuous treatment, then you will additionally get issues that are analogous to the ones that arise with a binary staggered treatment
In general, things get worse for TWFE regressions with more complications
Level Effects - no issues related to selection bias
Slope Effects - must deal with selection bias
Changing the estimation strategy helps with the weights, but it does not fix the issues related to standard vs. strong parallel trends
Application: Simplified version of Acemoglu and Finkelstein (2008)
Setting: 1983 Medicare reform that eliminated labor subsidies for hospitals
Dose: Hospital’s exposure to reform by fraction of Medicare patients pre-treatment
Outcome: Hospital’s capital-labor ratio
Focus: Comparison of estimates of \(\ACRT^{\texttt{glob}}\) to \(\beta^{twfe}\)
This is a simplified version of Acemoglu and Finkelstein (2008)
1983 Medicare reform that eliminated labor subsidies for hospitals
Medicare moved to the Prospective Payment System (PPS) which replaced “full cost reimbursement” with “partial cost reimbursement” which eliminated reimbursements for labor (while maintaining reimbursements for capital expenses)
Rough idea: This changes relative factor prices which suggests hospitals may adjust by changing their input mix. Could also have implications for technology adoption, etc.
In the paper, we provide some theoretical arguments concerning properties of production functions that suggests conditions for strong parallel trends to hold.
Annual hospital-reported data from the American Hospital Association, 1980-1986
Outcome is capital/labor ratio
Dose is “exposure” to the policy
\(\widehat{\ATT}^{\texttt{loc}} = 0.80~~(\textrm{s.e}.=0.05)\)
| Approach | Estimate | Std. Error | Data |
|---|---|---|---|
| TWFE | full | ||
| TWFE | no untreated | ||
| \(\ACRT^{\texttt{glob}}\) | either |
| Approach | Estimate | Std. Error | Data |
|---|---|---|---|
| TWFE | 1.14 | 0.10 | full |
| TWFE | 0.25 | 0.18 | no untreated |
| \(\ACRT^{\texttt{glob}}\) | -0.08 | 0.19 | either |
Staggered treatment adoption has received a lot of attention in the literature on DiD with a binary treatment (Goodman-Bacon 2021; de Chaisemartin and D’Haultfoeuille 2020; Callaway and Sant’Anna 2021)
Extending the previous results to settings with a continuous, staggered treatment (i.e., once it turns on, the treatment keeps the same amount) is relatively straightforward:
Including covariates in the parallel trends assumption often makes it more plausible (Heckman et al. 1997; Abadie 2005; Caetano and Callaway 2025)
Conceptually, it’s straightforward to extend the previous results to include covariates, but the main challenge is operationalizing it.
Haddad et al. (2024) use kernel-based estimators combined with double/debiased machine learning, allowing flexible, data-adaptive adjustment for covariates.
The discussion above used level effects (\(\ATT\)’s) and slope effects (\(\ACRT\)’s) as the “building block” parameters.
An alternative building block parameter is the “scaled level” parameter:
\[ \frac{\ATT(d \mid d)}{d} = \frac{\E[Y_{t=2}(d) - Y_{t=2}(0) \mid D=d]}{d} \]
This is the average treatment effect per unit of dose. This parameter shows up in (de Chaisemartin et al. 2024, 2025).
Above, we considered the case where \(D_{t=1} = 0\) for all units. It’s common to have continuous treatments where all units experience some dose across all periods.
The key issue in this case is whether or not there are any stayers—units that do not change their dose over time.
Previously, even in the staggered adoption case, we assumed that the dose was constant across time. In many applications, the dose may change across time too.
The previous ideas basically carry over to this case too, but there are some complications:
There are a number of challenges to implementing/interpreting DiD with a continuous treatment
In my view, the main new issue here is that justifying interpreting comparisons across different doses as causal effects requires stronger assumptions than most researchers probably think that they are making
Code: R contdid package available on CRAN