Difference-in-Differences with a Continuous Treatment

Session 1

Brantly Callaway

University of Georgia

September 18, 2026

This Session

\(\newcommand{\E}{\mathbb{E}} \newcommand{\E}{\mathbb{E}} \newcommand{\var}{\mathrm{var}} \newcommand{\cov}{\mathrm{cov}} \newcommand{\Var}{\mathrm{var}} \newcommand{\Cov}{\mathrm{cov}} \newcommand{\Corr}{\mathrm{corr}} \newcommand{\corr}{\mathrm{corr}} \newcommand{\L}{\mathrm{L}} \renewcommand{\P}{\mathrm{P}} \newcommand{\ACR}{\mathrm{ACR}} \newcommand{\ATT}{\mathrm{ATT}} \newcommand{\ACRT}{\mathrm{ACRT}} \newcommand{\independent}{{\perp\!\!\!\perp}} \newcommand{\indicator}[1]{ \mathbf{1}\{#1\} }\)

Difference-in-differences with a continuous treatment (or multi-valued discrete) treatment

Main reference: Callaway et al. (2026)

Introduction

Most recent work on DiD has focused on settings with a binary, staggered treatment

This session: Move from a binary treatment to a continuous treatment (“dose”)

Some of the arguments involve extending ideas from the binary, staggered treatment case to a setting with continuous treatment

  • But we will also face new conceptual issues in this case that do not show up in a setting with a binary treatment

Running Example:

  • Effect of \(\underbrace{\textrm{length of school closures}}_{\textrm{continuous treatment}}\) (during Covid) on \(\underbrace{\textrm{students' test scores}}_{\textrm{outcome}}\)
  • e.g., Ager et al. (2024), Gillitzer and Prasad (2023), among others

What Exactly Is a Continuous Treatment?

For today, we mostly emphasize a continuous treatment, but our results also apply to other settings (with trivial modifications):

  • Multi-valued treatments (e.g., effect of state-level minimum wage policies on employment)

  • Binary policy variable with differential exposure (application in the paper: a binary Medicare policy where different hospitals had more exposure to the policy)

What’s Not a Continuous Treatment?

Fuzzy DiD - binary treatment but the researcher observes aggregate data, often using policy changes as an instrument

  • Example: Study union wage-premium (at individual level) using state-level data and exploiting variation in the amount of unionization across different locations
  • See de Chaisemartin and D’Haultfoeuille (2018) and Miyaji (2025)

Narrowing the Scope of the Talk

We will focus on a setting with a panel data pre-post research design

  • No units are treated in the first period
  • Some units become treated, with possibly different doses, in the second period

We will return to extensions where

  • More periods and variation in treatment timing
  • All units become treated in the second period
  • Units can be already treated in the first period

Plan for This Session



  1. Identification: What’s the same as for a binary treatment?

  2. Identification: What’s different from a binary treatment?

  3. Reverse Engineering TWFE Regressions

  4. Empirical Application

  5. Extensions

1. Identification: What’s the Same as for a Binary Treatment?

Continuous Treatment Notation

  • Two time periods: \(t=1\) and \(t=2\)

    • No one treated until period \(t=2\)
    • Some units remain untreated in period \(t=2\)
  • Treatment: \(D_i\)

  • Potential outcomes: \(Y_{it}(d)\)

  • Observed outcomes: \(Y_{it=2}\) and \(Y_{it=1}\)

    \[Y_{it=2}=Y_{it=2}(D_i) \quad \textrm{and} \quad Y_{it=1}=Y_{it=1}(0)\]

Parameters of Interest (ATT-type)

Level Effects (Average Treatment Effect on the Treated)

\[\ATT(d \mid d) := \E[Y_{t=2}(d) - Y_{t=2}(0) \mid D=d]\]

  • Interpretation: The average effect of dose \(d\) relative to not being treated local to the group that actually experienced dose \(d\)

  • This is the natural analogue of \(\ATT\) in the binary treatment case

Parameters of Interest (ACRT-type)

Slope Effects (Average Causal Response on the Treated)

\[\ACRT(d \mid d) := \frac{\partial \ATT(l \mid d)}{\partial l} \Big|_{l=d}\]

  • Interpretation: \(\ACRT(d \mid d)\) is the causal effect of a marginal increase in dose local to units that actually experienced dose \(d\)

Aggregated Parameters

Notice that \(\ATT(d \mid d)\) and \(\ACRT(d \mid d)\) are functional parameters

  • This is different from targeting a single number, like \(\beta^{twfe}\) in \(Y_{it} = \theta_t + \eta_i + \beta^{twfe} D_{it} + e_{it}\)

We can view \(\ATT(d \mid d)\) and \(\ACRT(d \mid d)\) as the “building blocks” for a more aggregated parameter.

Aggregated versions of these (into a single number) are \[\begin{align*} \ATT^{\texttt{loc}} := \E\Big[\ATT(D \mid D)\Bigm|D>0\Big] \qquad \qquad \ACRT^{\texttt{loc}} := \E\Big[\ACRT(D \mid D)\Bigm|D>0\Big] \end{align*}\]

  • \(\ATT^{\texttt{loc}}\) averages \(\ATT(d \mid d)\) over the population distribution of the dose

  • \(\ACRT^{\texttt{loc}}\) averages \(\ACRT(d \mid d)\) over the population distribution of the dose

    • \(\ACRT^{\texttt{loc}}\) is the natural target parameter for the TWFE regression in this case

Identification

Parallel Trends Assumption

For all \(d \in \mathcal{D}_+\),

\[\E[\Delta Y(0) \mid D=d] = \E[\Delta Y(0) \mid D=0]\]

Identification

Parallel Trends Assumption

For all \(d \in \mathcal{D}_+\),

\[\E[\Delta Y(0) \mid D=d] = \E[\Delta Y(0) \mid D=0]\]

  • This is the natural way to extend the parallel trends assumption from the binary treatment setting to the continuous treatment setting
  • In Acemoglu and Finkelstein (2008):

Identification

Parallel Trends Assumption

For all \(d \in \mathcal{D}_+\),

\[\E[\Delta Y(0) \mid D=d] = \E[\Delta Y(0) \mid D=0]\]

Under parallel trends,

\[ \begin{aligned} \ATT(d \mid d) &= \E[Y_{t=2}(d) - Y_{t=2}(0) \mid D=d] \hspace{150pt} \end{aligned} \]

Identification

Parallel Trends Assumption

For all \(d \in \mathcal{D}_+\),

\[\E[\Delta Y(0) \mid D=d] = \E[\Delta Y(0) \mid D=0]\]

Under parallel trends,

\[ \begin{aligned} \ATT(d \mid d) &= \E[Y_{t=2}(d) - Y_{t=2}(0) \mid D=d] \hspace{150pt}\\ &= \E[Y_{t=2}(d) - Y_{t=1}(0) \mid D=d] - \E[Y_{t=2}(0) - Y_{t=1}(0) \mid D=d] \end{aligned} \]

Identification

Parallel Trends Assumption

For all \(d \in \mathcal{D}_+\),

\[\E[\Delta Y(0) \mid D=d] = \E[\Delta Y(0) \mid D=0]\]

Under parallel trends,

\[ \begin{aligned} \ATT(d \mid d) &= \E[Y_{t=2}(d) - Y_{t=2}(0) \mid D=d] \hspace{150pt}\\ &= \E[Y_{t=2}(d) - Y_{t=1}(0) \mid D=d] - \E[Y_{t=2}(0) - Y_{t=1}(0) \mid D=d]\\ &= \E[\Delta Y \mid D=d] - \E[\Delta Y \mid D=0] \end{aligned} \]

This is exactly what you would expect

Identification

Parallel Trends Assumption

For all \(d \in \mathcal{D}_+\),

\[\E[\Delta Y(0) \mid D=d] = \E[\Delta Y(0) \mid D=0]\]

Using similar arguments, you can show \[\ATT^{\texttt{loc}} = \E[\Delta Y \mid D>0] - \E[\Delta Y \mid D=0]\]

2. Identification: What’s Different from a Binary Treatment?

Are We Done?

Unfortunately, no

Most empirical work with a continuous treatment wants to think about how causal responses vary across dose

  • Plot treatment effects as a function of dose and ask: does more dose cause outcomes to increase/decrease/no effect?

  • Important because distinguishing between different theories often involves making comparisons across doses

    • Smoking Example

Are We Done?

Unfortunately, no

Most empirical work with a continuous treatment wants to think about how causal responses vary across dose

  • Empirical work very often interprets TWFE regressions in terms of causal responses (slope parameters):
    • Goodman-Bacon (2018): “After Medicaid, nonwhite child mortality fell by 1.4 percent (s.e. = 0.34; table 3) for each percentage point difference in initial AFDC rates.”
  • Average causal response parameters inherently involve comparisons across slightly different doses

Are We Done?

Unfortunately, no

Most empirical work with a continuous treatment wants to think about how causal responses vary across dose

There are new issues related to comparing \(\ATT(d \mid d)\) at different doses and interpreting these differences as causal effects

  • Unlike the staggered, binary treatment case: No easy fixes here!

Interpretation Issues

Consider comparing \(\ATT(d \mid d)\) for two different doses

\[ \begin{aligned} & \ATT(d_h \mid d_h) - \ATT(d_l \mid d_l) \hspace{350pt} \end{aligned} \]

Interpretation Issues

Consider comparing \(\ATT(d \mid d)\) for two different doses

\[ \begin{aligned} & \ATT(d_h \mid d_h) - \ATT(d_l \mid d_l) \hspace{350pt}\\ & \hspace{25pt} = \E[Y_{t=2}(d_h)-Y_{t=2}(d_l) \mid D=d_h] + \E[Y_{t=2}(d_l) - Y_{t=2}(0) \mid D=d_h] - \E[Y_{t=2}(d_l) - Y_{t=2}(0) \mid D=d_l] \end{aligned} \]

Interpretation Issues

Consider comparing \(\ATT(d \mid d)\) for two different doses

\[ \begin{aligned} & \ATT(d_h \mid d_h) - \ATT(d_l \mid d_l) \hspace{350pt}\\ & \hspace{25pt} = \E[Y_{t=2}(d_h)-Y_{t=2}(d_l) \mid D=d_h] + \E[Y_{t=2}(d_l) - Y_{t=2}(0) \mid D=d_h] - \E[Y_{t=2}(d_l) - Y_{t=2}(0) \mid D=d_l]\\ & \hspace{25pt} = \underbrace{\E[Y_{t=2}(d_h) - Y_{t=2}(d_l) \mid D=d_h]}_{\textrm{Causal Response}} + \underbrace{\ATT(d_l \mid d_h) - \ATT(d_l \mid d_l)}_{\textrm{Selection Bias}} \end{aligned} \]

“Standard” Parallel Trends is not strong enough to rule out the selection bias terms here

  • Implication: If you want to interpret differences in treatment effects across different doses, then you will need stronger assumptions than standard parallel trends

  • This problem spills over into identifying \(\ACRT(d|d)\)

Interpretation Issues

Intuition

  • Difference-in-differences identification strategies result in \(\ATT(d|d)\) parameters. These are local parameters and difficult to compare to each other

  • This explanation is similar to thinking about LATEs with two different instruments

  • Thus, comparing \(\ATT(d \mid d)\) across different values is tricky and not for free

Examples

Interpretation Issues

What can you do?

  • One idea, just recover \(\ATT(d \mid d)\) and interpret it cautiously (interpret it by itself not relative to different values of \(d\))

  • If you want to compare them to each other, it will come with the cost of additional (structural) assumptions

Less Local Target Parameters

To be able to make comparisons across different doses, we need target parameters that are “less local”:

\[ \ATT(d) := \E[Y_{t=2}(d) - Y_{t=2}(0) \mid D>0] \qquad \qquad \ACRT(d) := \frac{\partial \ATT(d)}{\partial d} \]

These involve the level effect or slope effect at \(d\) (like the previous parameters), but they are causal effect parameters for the entire treated population (not just the local group that actually experienced dose \(d\))

We can also consider aggregated versions of these parameters:

\[\ATT^{\texttt{glob}} = \E[ATT(D) \mid D>0] \quad \quad \ACRT^{\texttt{glob}} = \E[\ACRT(D) \mid D>0]\]

Introduce Stronger Assumptions

Strong Parallel Trends Assumption

For all doses \(d \in \mathcal{D}\),

\[\E[Y_{t=2}(d) - Y_{t=1}(0) \mid D>0] = \E[Y_{t=2}(d) - Y_{t=1}(0) \mid D=d]\]

  • This is notably different from “Standard” Parallel Trends
  • It involves potential outcomes for all values of the dose (not just untreated potential outcomes)
  • Rules out selection into the amount of treatment
  • All treated dose groups would have experienced the same path of outcomes had they been assigned the same dose
  • Non-nested with “Standard” Parallel Trends, but likely stronger

Introduce Stronger Assumptions

If we maintain standard parallel trends, strong parallel trends is equivalent to a certain restriction on treatment effect heterogeneity. Notice:

\[ \begin{aligned} \ATT(d|d) &= \E[Y_{t=2}(d) - Y_{t=1}(0) \mid D=d] - \E[Y_{t=2}(0) - Y_{t=1}(0) \mid D=d] \hspace{200pt} \ \end{aligned} \]

Introduce Stronger Assumptions

If we maintain standard parallel trends, strong parallel trends is equivalent to a certain restriction on treatment effect heterogeneity. Notice:

\[ \begin{aligned} \ATT(d|d) &= \E[Y_{t=2}(d) - Y_{t=1}(0) \mid D=d] - \E[Y_{t=2}(0) - Y_{t=1}(0) \mid D=d] \hspace{200pt} \\\ &= \E[Y_{t=2}(d) - Y_{t=1}(0) \mid D>0] - \E[Y_{t=2}(0) - Y_{t=1}(0) \mid D > 0] \ \end{aligned} \]

Introduce Stronger Assumptions

If we maintain standard parallel trends, strong parallel trends is equivalent to a certain restriction on treatment effect heterogeneity. Notice:

\[ \begin{aligned} \ATT(d|d) &= \E[Y_{t=2}(d) - Y_{t=1}(0) \mid D=d] - \E[Y_{t=2}(0) - Y_{t=1}(0) \mid D=d] \hspace{200pt} \\\ &= \E[Y_{t=2}(d) - Y_{t=1}(0) \mid D>0] - \E[Y_{t=2}(0) - Y_{t=1}(0) \mid D > 0] \\\ &= \E[Y_{t=2}(d) - Y_{t=2}(0) \mid D>0] = \ATT(d) \end{aligned} \]

Thus, under strong parallel trends, we have that

\[\ATT(d) = \E[\Delta Y|D=d] - \E[\Delta Y|D=0]\]

RHS is exactly the same expression as for \(\ATT(d|d)\) under standard PT, but here

  • assumptions are different
  • parameter interpretation is different

Comparisons across Dose

ATT(d) parameters do not suffer from the same issues as ATT(d|d) parameters when making comparisons across dose

\[ \begin{aligned} \ATT(d_h) - \ATT(d_l) &= \E[Y_{t=2}(d_h) - Y_{t=2}(0) | D>0] - \E[Y_{t=2}(d_l) - Y_{t=2}(0)|D>0] \end{aligned} \]

Comparisons across Dose

ATT(d) parameters do not suffer from the same issues as ATT(d|d) parameters when making comparisons across dose

\[ \begin{aligned} \ATT(d_h) - \ATT(d_l) &= \E[Y_{t=2}(d_h) - Y_{t=2}(0) | D>0] - \E[Y_{t=2}(d_l) - Y_{t=2}(0)|D>0]\\ &= \underbrace{\E[Y_{t=2}(d_h) - Y_{t=2}(d_l)|D>0]}_{\textrm{Causal Response}} \end{aligned} \]

Thus, recovering \(\ATT(d)\) side-steps the issues about comparing treatment effects across doses, but it comes at the cost of needing a (potentially very strong) extra assumption

Given that we can compare \(\ATT(d)\)’s across dose, we can recover slope effects in this setting

\[ \begin{aligned} \ACRT(d) := \frac{\partial \ATT(d)}{\partial d} \qquad &\textrm{or} \qquad \ACRT^{\texttt{glob}} := \E\Big[\ACRT(D) \Big| D>0\Big] \end{aligned} \]

Additional Discussion

  1. Can You Relax Strong Parallel Trends?

  2. No Untreated Units

  3. Binarizing the Treatment

  4. Pre-Testing

No Untreated Units

It’s possible to do some versions of DiD with a continuous treatment without having access to a fully untreated group.

  • In this case, it is not possible to recover level effects like \(\ATT(d|d)\).

  • However, notice that \[\begin{aligned}& \E[\Delta Y | D=d_h] - \E[\Delta Y | D=d_l] \\ &\hspace{20pt}= \Big(\E[\Delta Y | D=d_h] - \E[\Delta Y(0) | D=d_h]\Big) - \Big(\E[\Delta Y | D=d_l]-\E[\Delta Y(0) | D=d_l]\Big) \\ &\hspace{20pt}= \ATT(d_h|d_h) - \ATT(d_l|d_l)\end{aligned}\]

  • In words: comparing path of outcomes for those that experienced dose \(d_h\) to path of outcomes among those that experienced dose \(d_l\) (and not relying on having an untreated group) delivers the difference between their \(\ATT\)’s.

  • Still face issues related to selection bias / strong parallel trends though

Binarizing the Treatment

Strategies like binarizing the treatment can still work (though be careful!)

  • If you classify units as being treated or untreated, you can recover \(\ATT^{\texttt{loc}}\) by comparing mean outcome for treated relative to untreated.

  • On the other hand, if you classify units as being “high” treated, “low” treated, or untreated — our arguments imply that selection bias terms can come up when comparing effects for “high” to “low”

Pre-Testing

That the expressions for \(\ATT(d)\) and \(\ATT(d|d)\) are exactly the same also means that we cannot use pre-treatment periods to try to distinguish between “standard” and “strong” parallel trends.

Or, alternatively, we only observe untreated potential outcomes in pre-treatment periods, so pre-treatment periods do not seem useful for validating strong parallel trends.

Summarizing

It is straightforward/familiar to identify \(\ATT(d \mid d)\) parameters with a multi-valued or continuous dose

However, comparisons of \(\ATT(d\mid d)\) parameters across different doses are hard to interpret

  • They include selection bias terms
  • This issue extends to identifying \(\ACRT\) parameters
  • These issues extend to TWFE regressions

This suggests targeting \(\ATT(d)\) parameters

  • Comparisons across doses do not contain selection bias terms
  • But identifying \(\ATT(d)\) parameters requires stronger assumptions

3. Reverse Engineering TWFE Regressions

TWFE Regressions in This Context

Consider the same TWFE regression as before: \[\begin{align*} Y_{it} = \theta_t + \eta_i + \beta^{twfe} D_{it} + e_{it} \end{align*}\] We show that \[\begin{align*} \beta^{twfe} = \int_{\mathcal{D}_+} w(l) m'_\Delta(l) \, dl \end{align*}\] where \(m_\Delta(l) := \E[\Delta Y \mid D=l] - \E[\Delta Y \mid D=0]\) and \(w(l)\) are weights

  • Under standard parallel trends, \(m'_{\Delta}(l) = \ACRT(l\mid l) + \textrm{local selection bias}\)
  • Under strong parallel trends, \(m'_{\Delta}(l) = \ACRT(l)\).

About the weights: they are all positive, but have some strange properties (e.g., always maximized at \(l = \E[D]\) (even if this is not a common value for the dose))

  • \(\implies\) even under strong parallel trends, \(\beta^{twfe} \neq \ACRT^{\texttt{glob}}\).

TWFE Regressions in This Context

Other issues can arise in more complicated cases

  • For example, suppose you have a staggered continuous treatment, then you will additionally get issues that are analogous to the ones that arise with a binary staggered treatment

  • In general, things get worse for TWFE regressions with more complications

Estimation — What Should You Do?

Level Effects - no issues related to selection bias

  • For \(\ATT^{\texttt{loc}}\): Binarize treatment, \(\ATT^{\texttt{loc}} = \E[\Delta Y \mid D > 0] - \E[\Delta Y \mid D=0]\).
  • For \(\ATT(d \mid d)\): Nonparametrically estimate \(m_\Delta(d) = \E[\Delta Y \mid D=d]-\E[\Delta Y \mid D=0]\)
    • This is not actually too hard to estimate. No curse-of-dimensionality, etc.

Slope Effects - must deal with selection bias

  • Nonparametrically estimate derivative of \(m_\Delta(d)\)
  • For \(\ACRT(d)\): Under strong parallel trends, derivative is equal to \(\ACRT(d)\)
  • For \(\ACRT^{\texttt{glob}}\): Average \(\ACRT(D)\) over \(D>0\); i.e., \(\ACRT^{\texttt{glob}} = \E[\ACRT(D)|D>0]\)

Additional Comments about Estimation

Changing the estimation strategy helps with the weights, but it does not fix the issues related to standard vs. strong parallel trends

4. Empirical Application

Empirical Application

Application: Simplified version of Acemoglu and Finkelstein (2008)

Setting: 1983 Medicare reform that eliminated labor subsidies for hospitals

Dose: Hospital’s exposure to reform by fraction of Medicare patients pre-treatment

Outcome: Hospital’s capital-labor ratio

Focus: Comparison of estimates of \(\ACRT^{\texttt{glob}}\) to \(\beta^{twfe}\)

  • They are both weighted averages of \(\ACRT(d)\), only the weights differ

Empirical Application

This is a simplified version of Acemoglu and Finkelstein (2008)

1983 Medicare reform that eliminated labor subsidies for hospitals

  • Medicare moved to the Prospective Payment System (PPS) which replaced “full cost reimbursement” with “partial cost reimbursement” which eliminated reimbursements for labor (while maintaining reimbursements for capital expenses)

  • Rough idea: This changes relative factor prices which suggests hospitals may adjust by changing their input mix. Could also have implications for technology adoption, etc.

  • In the paper, we provide some theoretical arguments concerning properties of production functions that suggests conditions for strong parallel trends to hold.

Data

Annual hospital-reported data from the American Hospital Association, 1980-1986

Outcome is capital/labor ratio

  • proxy using the depreciation share of total operating expenses (avg. 4.5%)
  • our setup: collapse to two periods by taking average in pre-treatment periods and average in post-treatment periods

Dose is “exposure” to the policy

  • the fraction of Medicare patients in the period before the policy was implemented
  • roughly 15% of hospitals are untreated (have essentially no Medicare patients)
    • AF report results both with and without these hospitals: having untreated hospitals is useful, but they are fairly different (federal, long-term, psychiatric, children’s, and rehabilitation hospitals)

Bin Scatter

ATT(d|d) Plot

\(\widehat{\ATT}^{\texttt{loc}} = 0.80~~(\textrm{s.e}.=0.05)\)

ACRT(d) Plot

Results


Approach Estimate Std. Error Data
TWFE full
TWFE no untreated
\(\ACRT^{\texttt{glob}}\) either

Results


Approach Estimate Std. Error Data
TWFE 1.14 0.10 full
TWFE 0.25 0.18 no untreated
\(\ACRT^{\texttt{glob}}\) -0.08 0.19 either

Density Weights vs. TWFE Weights

TWFE Weights with and without Untreated Group

5. Extensions

1. Multiple Periods and Variation in Treatment Timing

Staggered treatment adoption has received a lot of attention in the literature on DiD with a binary treatment (Goodman-Bacon 2021; de Chaisemartin and D’Haultfoeuille 2020; Callaway and Sant’Anna 2021)

Extending the previous results to settings with a continuous, staggered treatment (i.e., once it turns on, the treatment keeps the same amount) is relatively straightforward:

  • Following Callaway and Sant’Anna (2021), can focus on identifying “group-time-dose” parameters, e.g., \(\ATT(d, g, t \mid d, g)\) (under parallel trends) or \(\ACRT(d, g, t \mid g)\) (under strong parallel trends)
  • The parameters can be averaged together into an event study-type parameter or an aggregated treatment effect parameter as a function of the dose.
  • TWFE regressions end up suffering from a combination of the issues that arise in the staggered binary treatment case and the issues that arise in the continuous treatment case.

\(\ATT^{\texttt{loc}}(e)\) Event Study

\(\ACRT^{\texttt{glob}}(e)\) Event Study

2. Including Covariates

Including covariates in the parallel trends assumption often makes it more plausible (Heckman et al. 1997; Abadie 2005; Caetano and Callaway 2025)

Conceptually, it’s straightforward to extend the previous results to include covariates, but the main challenge is operationalizing it.

Haddad et al. (2024) use kernel-based estimators combined with double/debiased machine learning, allowing flexible, data-adaptive adjustment for covariates.

3. Alternative Building Block Parameters

The discussion above used level effects (\(\ATT\)’s) and slope effects (\(\ACRT\)’s) as the “building block” parameters.

An alternative building block parameter is the “scaled level” parameter:

\[ \frac{\ATT(d \mid d)}{d} = \frac{\E[Y_{t=2}(d) - Y_{t=2}(0) \mid D=d]}{d} \]

This is the average treatment effect per unit of dose. This parameter shows up in (de Chaisemartin et al. 2024, 2025).

  • You can average these to get an overall “per unit of dose” effect. Since it doesn’t involve comparisons across dose, it does not contain selection bias.
  • Can show that the TWFE regression coefficient is a weighted average of these scaled level parameters but the weights can be negative.

4. Units Treated in the First Period

Above, we considered the case where \(D_{t=1} = 0\) for all units. It’s common to have continuous treatments where all units experience some dose across all periods.

The key issue in this case is whether or not there are any stayers—units that do not change their dose over time.

  • If yes, then the previous results effectively carry over to this case. You can proceed by, for each initial dose, applying the same sort of arguments as above but with initial dose as the “untreated” group.
    • The main complication is how to pool information across different initial doses.
    • See de Chaisemartin et al. (2025) for one approach.
  • If no, then at least arguably, this is no longer DiD as it does not fit into the pre-post/staggered research design.

5. Doses Change across Time

Previously, even in the staggered adoption case, we assumed that the dose was constant across time. In many applications, the dose may change across time too.

The previous ideas basically carry over to this case too, but there are some complications:

  • What’s the target parameter? If the dose changes over time, are we going for the treatment effect relative to being untreated or the treatment effect relative to the previous dose? This affects what comparison group to use.
  • Applying the same types of arguments will result in very high dimensional target parameters, e.g., an \(\ATT\) for an entire treatment sequence. How should these be aggregated?

Conclusion

  • There are a number of challenges to implementing/interpreting DiD with a continuous treatment

  • In my view, the main new issue here is that justifying interpreting comparisons across different doses as causal effects requires stronger assumptions than most researchers probably think that they are making

  • Link to paper: [published] [arXiv]

  • Code: R contdid package available on CRAN

Thank You

brantly.callaway@uga.edu

Appendix

References

Abadie, Alberto. 2005. “Semiparametric Difference-in-Differences Estimators.” The Review of Economic Studies 72 (1): 1–19.
Acemoglu, Daron, and Amy Finkelstein. 2008. “Input and Technology Choices in Regulated Industries: Evidence from the Health Care Sector.” Journal of Political Economy 116 (5): 837–80.
Ager, Philipp, Katherine Eriksson, Ezra Karger, Peter Nencka, and Melissa A Thomasson. 2024. “School Closures During the 1918 Flu Pandemic.” Review of Economics and Statistics 106 (1): 266–76.
Caetano, Carolina, and Brantly Callaway. 2025. “Difference-in-Differences When Parallel Trends Holds Conditional on Covariates.” Unpublished manuscript.
Callaway, Brantly, Andrew Goodman-Bacon, and Pedro HC Sant’Anna. 2026. “Difference-in-Differences with a Continuous Treatment.” American Economic Review.
Callaway, Brantly, and Pedro HC Sant’Anna. 2021. “Difference-in-Differences with Multiple Time Periods.” Journal of Econometrics 225 (2): 200–230.
de Chaisemartin, Clément, Diego Ciccia, Xavier D’Haultfoeuille, and Felix Knau. 2024. “Two-Way Fixed Effects and Differences-in-Differences Estimators in Heterogeneous Adoption Designs.” Unpublished manuscript.
de Chaisemartin, Clement, and Xavier D’Haultfoeuille. 2018. “Fuzzy Differences-in-Differences.” The Review of Economic Studies 85 (2): 999–1028.
de Chaisemartin, Clement, and Xavier D’Haultfoeuille. 2020. “Two-Way Fixed Effects Estimators with Heterogeneous Treatment Effects.” American Economic Review 110 (9): 2964–96.
de Chaisemartin, Clément, Xavier D’Haultfoeuille, Félix Pasquier, Doulo Sow, and Gonzalo Vazquez-Bare. 2025. “Difference-in-Differences for Continuous Treatments and Instruments with Stayers.” Unpublished manuscript.
Gillitzer, Christian, and Nalini Prasad. 2023. “The Effect of School Closures on Standardized Test Scores: Evidence from a Zero-COVID Environment.” Unpublished manuscript.
Goodman-Bacon, Andrew. 2018. “Public Insurance and Mortality: Evidence from Medicaid Implementation.” Journal of Political Economy 126 (1): 216–62.
Goodman-Bacon, Andrew. 2021. “Difference-in-Differences with Variation in Treatment Timing.” Journal of Econometrics 225 (2): 254–77.
Haddad, Michel F. C., Martin Huber, and Lucas Z. Zhang. 2024. “Difference-in-Differences with Time-Varying Continuous Treatments Using Double/Debiased Machine Learning.” https://arxiv.org/abs/2410.21105.
Heckman, James, Hidehiko Ichimura, and Petra Todd. 1997. “Matching as an Econometric Evaluation Estimator: Evidence from Evaluating a Job Training Programme.” The Review of Economic Studies 64 (4): 605–54.
Miyaji, Sho. 2025. “Instrumented Difference-in-Differences with Heterogeneous Treatment Effects.” Unpublished manuscript.