Session 3
September 18, 2026
\(\newcommand{\E}{\mathbb{E}} \newcommand{\var}{\mathrm{var}} \newcommand{\cov}{\mathrm{cov}} \newcommand{\Var}{\mathrm{var}} \newcommand{\Cov}{\mathrm{cov}} \newcommand{\Corr}{\mathrm{corr}} \newcommand{\corr}{\mathrm{corr}} \newcommand{\L}{\mathrm{L}} \renewcommand{\P}{\mathrm{P}} \newcommand{\independent}{{\perp\!\!\!\perp}} \newcommand{\indicator}[1]{ \mathbf{1}\{#1\} } \newcommand{\ATT}{\text{ATT}} \newcommand{\ACR}{\text{ACR}}\) Shift-share (aka Bartik) variables are created as combinations of
Combining these creates a single variable capturing how “exposed” a unit is to some shocks
This session: An overview of shift-share designs for causal inference
Introduction to shift-share designs and an application
Constructing shift-share variables
Causal effects
Extensions
Question: Did the expansion of labor unions contribute to the American Baby Boom?
Setting: The 1935 National Labor Relations Act (NLRA) — the largest expansion of unionism in U.S. history
Treatment \(D_{it}\): county-level union membership rate
Outcome \(Y_{it}\): county-level birth rates
Target Parameter: Not always clear in the literature (i.e., often stated as a regression coefficient), but something like \(\ACR(d) = \frac{\partial \E[Y_{it}(d)]}{\partial d}\)—how much a marginal increase in unionization would raise fertility
Downes (2024) is representative of many shift-share papers, but let me mention a few more well known ones:
| Paper | Unit | Treatment \(D\) | Outcome \(Y\) |
|---|---|---|---|
| Bartik (1991), Blanchard and Katz (1992) | Region | \(\Delta\) local employment | \(\Delta\) local wages |
| Card (2001) | City \(\times\) skill group | Immigrant inflows | Native wages and employment |
| Autor et al. (2013) | Commuting zone | \(\Delta\) exposure to Chinese imports | \(\Delta\) manufacturing employment |
| Nunn and Qian (2014) | Country | US food aid (wheat) received | Civil conflict |
| Greenstone et al. (2020) | County | \(\Delta\) bank credit supply | \(\Delta\) employment |
| Xu (2022) | Region | Exposure to bank failures | \(\Delta\) exports |
| Franklin et al. (2024) | Local labor market | Exposure to public works program | Wages |
Start with a simplified setting similar to the one we have been considering:
Endogeneity: locations with tight labor markets may experience higher unionization but a tight labor market may also affect the trend in county fertility even without unionization.
This is directly related to our earlier discussion of continuous treatments. Here the worry is that parallel trends and/or strong parallel trends is violated:
\(\text{Parallel Trends: } \E[\Delta Y(0) \mid D=d] \neq \E[\Delta Y(0) \mid D=0]\)
or
\(\text{Strong Parallel Trends: } \E[\Delta Y(d) \mid D=d'] \neq \E[\Delta Y(d) \mid D=d]\)
To start with, notice that
\[ D_{it=2} = \sum_{j=1}^J s_{ijt=2} \times r_{ijt=2}\]
where
This is a decomposition and holds by definition.
Next, we can decompose the industry share variable \(s_{ijt=2}\) and the industry unionization rate \(r_{ijt=2}\) as follows:
\[s_{ijt=2} = \color{red}{s_{ijt=1}} + \Delta s_{ij}\]
where \(\color{red}{s_{ijt=1}}\) is the share of county \(i\)’s employment in industry \(j\) at time \(t=1\) (pre-NLRA) and \(\Delta s_{ij}\) is the change in that share from \(t=1\) to \(t=2\).
\[r_{ijt=2} = \color{blue}{r_{jt=2}} + \tilde{r}_{ijt=2}\]
where \(\color{blue}{r_{jt=2}}\) is the national unionization rate in industry \(j\) at time \(t=2\) and \(\tilde{r}_{ijt=2}\) is the unionization rate remainder in county \(i\) for industry \(j\) at time \(t=2\).
Roughly, the idea of a shift-share variable is that \(s_{ijt=1}\) and/or \(r_{jt=2}\) are more likely to be exogenous than the treatment \(D_{it=2}\) itself.
We can now construct a shift-share variable as follows: \[ B_{i} = \sum_{j=1}^J s_{ijt=1} \times r_{jt=2} \]
This is following the notation in Goldsmith-Pinkham et al. (2020), you can think of \(B\) standing for “Bartik”
Downes (2024) writes the same equation as: \[SSIV_{it} = \sum_{j=1}^J \underbrace{\textrm{IndShare}^{1910}_{ij}}_{\textrm{share}} \times \underbrace{\textrm{NatlUnionRate}_{jt}}_{\textrm{shift}}\]
\(B_i\) represents the “exposure” of different counties to unionization:
Downes (2024) gives a concrete two-county example:
Cameron County, PA and Forest County, PA — similar 1930 population (\(\approx 5{,}200\)), neither unionized before the NLRA
But Cameron had more transportation, more coal mining, and more of the manufacturing subindustries that later unionized heavily
This implies that these two counties were differently exposed to the NLRA shock.
The way that Goldsmith-Pinkham et al. (2020) discuss exogeneity of the shift-share variable is through exogeneity of the shares.
For simplicity, think about the case with two industries \(j=1,2\):
Notation:
The way that Goldsmith-Pinkham et al. (2020) discuss exogeneity of the shift-share variable is through exogeneity of the shares.
For simplicity, think about the case with two industries \(j=1,2\):
IV-type assumptions:
These are exactly the same type of assumption as you would see in the IV literature, applied to the shares.
Under exogeneity of the shares, it immediately holds that, e.g.,
\[ \E[\Delta Y \mid S_1 = s_1, S_2 = s_2] - \E[\Delta Y \mid S_1 = s_1', S_2 = s_2'] = \underbrace{\E[Y_{t=2}(s_1, s_2) - Y_{t=2}(s_1', s_2')]}_{\text{causal effect of shares}} \]
where \(Y_{t=2}(s_1, s_2) := Y_{t=2}(D(s_1, s_2))\) is the potential outcome if the shares were set to \(s_1\) and \(s_2\).
i.e., we can compare the trend in outcomes across counties with different shares and interpret that as a causal effect of different shares.
However, the expression on the previous slide
Instead, a natural estimand to consider is:
\[ \E[\Delta Y \mid B = b+1] - \E[\Delta Y \mid B = b] \]
which is the difference in trends in outcomes across counties with different values of the shift-share variable.
It is straighforward to show that, under the previous assumptions, the previous estimand is equal to \[\E\Big[ \E[Y_{t=2}(S_1,S_2)] \Bigm| B = b+1\Big] - \E\Big[ \E[Y_{t=2}(S_1,S_2)] \Bigm| B = b\Big]\]
This is not exactly what we want though:
We can additionally assume that
\[D(s_1,s_2) = D(b)\]
for all \(s_1, s_2\) such that \(s_1 r_1 + s_2 r_2 = b\). In other words, the amount of unionization in a county is invariant to the particular shares, as long as the shift-share variable is the same.
This allows us to causally interpret comparisons of counties with \(B=b+1\) to counties with \(B=b\) even if they have different substantially different underlying shares
Then, it immediately follows that
\[\E[\Delta Y \mid B = b+1] - \E[\Delta Y \mid B = b] =\E[Y_{t=2}(b+1) - Y_{t=1}(b)]\]
which is the causal effect of a marginal increase in the shift-share variable.
There’s a close connection between the previous result and strong parallel trends.
Notice that we are making comparisons across dose, so we need more than just parallel trends in “unexposed” potential outcomes.
\[\text{Parallel Trends: } \E[\Delta Y(0) \mid B=b] = \E[\Delta Y(0) \mid B=0]\]
But the previous assumptions imply a version of strong parallel trends in terms of the shift-share variable:
\[\text{Strong Parallel Trends: } \E[\Delta Y(b) \mid B=b'] = \E[\Delta Y(b) \mid B=b]\]
This is exactly what rationalizes comparing trends in outcomes across counties with different values of the shift-share variable as being causal effects of exposure.
To recover the causal effect of the treatment, a natural estimand is the derivatve version of the Wald estimand:
\[ \tau(b) := \frac{\partial \E[\Delta Y \mid B=b]}{\partial b} \bigg/ \frac{\partial \E[D \mid B=b]}{\partial b} \]
At this point, we need to introduce the other main IV assumptions
Relevance typically holds in shift-share IV applications (the shift-share variable is typically a strong instrument too).
Returning to the two counties we considered before. Cameron County was more exposed to the NLRA than Forest County. By 1960, union membership was 26.3% in Cameron vs. 7.7% in Forest.
Unfortunately, with a continuous treatment and continuous instrument, we know that (Angrist et al. 2000), \[\tau(b) = \frac{\E[Y'(D(b)) D'(b)]}{\E[D'(b)]} \neq ACR(d)\]
nor is \(\E[\tau(B)] = \E[ACR(D)]\).
i.e., \(\tau(b)\) is a weighted average of the causal response to the treatment (the \(Y'\) term in the expression), but the weights are \(D'(b)/\E[D'(b)]\), which are not constant across units.
It is more common in empirical work to run the use \(B_i\) as an instrument for \(D_i\) and estimate a linear model by two-stage least squares (2SLS). In this case,
\[\beta_{2sls} = \int w(b) \; \tau(b) \; db\]
where \(w(b)\) is a weight function that depends on the distribution of \(B\) and the conditional variance of \(D\) given \(B\), i.e., \(\beta_{2sls}\) is a weighted average of \(\tau(b)\).
Under the (likely mild and directly testable) condition that \(\frac{\partial \E[D \mid B=b]}{\partial b} \geq 0\), \(\beta_{2sls}\) is also a weighted average of causal responses, with weights that are all positive.
In settings where units are already treated in the first period, my understanding of the literature is that it has basically replaced \(d\) with \(\Delta d\) in the potential outcomes and used arguments similar to those I discussed above.
I emphasized the “exogenous shares” approach to identification from Goldsmith-Pinkham et al. (2020), but this is really only half of the literature.
The other main approach is to assume the shifts are exogenous and the shares are allowed to be endogenous, as in Borusyak et al. (2022)
Borusyak et al. (2025) provide a checklist for deciding if you should argue for exogenous shares or shifts.
Count the shifts—few suggests shares path, while many (dozens+) suggests shifts path
Are shares tailored or generic?—tailored to this channel suggests shares path, shares that could proxy many shocks suggestsshifts path
Which story can you say with a straight face: “each share alone would satisfy parallel trends” or “this shock is as-good-as-random and only reaches units through the treatment”?
Example: Autor et al. (2013) “China shock”
Example: Card (1990), Mariel Boatlift / immigrant-enclave designs