Recent Advances in Difference in Differences Methods for Policy Research — Transcript
Full transcript
- 0:04one for being here this afternoon. Uh so
- 0:08today I will be talking about uh recent
- 0:10advances in difference indifference
- 0:12methods and I will touch um some uh
- 0:16limitations that has been highly debated
- 0:18in uh recent literature and also the
- 0:22potential solutions. Here is uh an
- 0:24overview of today's talk. Um I will
- 0:27start with a brief review of the classic
- 0:30did setup and then move to the
- 0:32limitations of the standard two-way fix
- 0:35effect estimators especially when
- 0:37treatment effects are heterogeneous. I
- 0:39will then present alternative estimators
- 0:42that address these issues and finally I
- 0:44will touch on a few additional concerns
- 0:47related to the methods more broadly.
- 0:51So let's start with a quick overview of
- 0:53the basic of DID.
- 0:56So when can we use DID? Um so DID is
- 1:00used when we want to evaluate the effect
- 1:03of a policy or program and we have two
- 1:06groups treated and control and also two
- 1:09time periods before and after. So the
- 1:12key identification assumption is the
- 1:15parallel trend assumption which is the
- 1:17two group treated and control group
- 1:19would have follow parallel trends in the
- 1:22absence of treatment.
- 1:24So in that case we can infer the
- 1:27counterfactual treatment um outcome uh
- 1:31for the treatment group in the post
- 1:33treatment period and then uh the
- 1:36difference between actual and the
- 1:38counterfactual outcome can reveal the
- 1:41treatment effect. A historical example
- 1:43comes from Jon Snow's uh study of the uh
- 1:471854 colera outbreak in London uh which
- 1:51works a lot like a did uh study. Uh I
- 1:54guess you may be more familiar with me
- 1:56about this case. Um so uh for more
- 2:00context in 1854
- 2:02there was a cholera outbreak in London.
- 2:05So at that time people didn't know what
- 2:07caused it. Many believe it was due to uh
- 2:10bad air. But Jones Snow suspected that
- 2:13contaminated water caused the problem.
- 2:16So he noted that there are two water
- 2:18companies served the area. So one
- 2:20company is Danbat uh that it had started
- 2:24getting clean water from uh upstreams
- 2:27while the other company uh Southwalk and
- 2:30uh Ball still use dirty water. So the
- 2:33death rate uh in these two areas uh uh
- 2:36are very different. So in Dambeath area
- 2:39is only 10 deaths per 10,000 people
- 2:42compared to 150 in the other uh area. So
- 2:47um however this difference might not be
- 2:49caused by the water. Maybe the areas
- 2:52just different in other ways. So um snow
- 2:55looked back to an earlier outbreak in
- 2:581849 before the water source changed in
- 3:02Lambbeath. So back then both area had
- 3:05high death rate 150 versus 125. So that
- 3:09comparison helped isolate the effect of
- 3:12cleaner water. The fact that only lambus
- 3:15death rate dropped after switching water
- 3:18suggests the water change was the key
- 3:20factor. And that's the basic idea behind
- 3:24difference in difference which is
- 3:26comparing change over time between a
- 3:28group that gets a treatment and a group
- 3:31that doesn't. Here is another classic
- 3:33example of the which comes from a famous
- 3:36study by Cart and Krueger in 1994. This
- 3:40study is was considered as the
- 3:42foundation of modern did. So their
- 3:45research question is does higher minimum
- 3:48wage reduce jobs. So um for context in
- 3:521992 New Jersey in United States raised
- 3:56its minimum wage from 4.25 to 5.05 per
- 4:00hour. So across the border in
- 4:02Pennsylvania, the minimum wage stayed
- 4:04the same. So in this case, New Jersey is
- 4:07a treaty state and uh Pennsylvania is a
- 4:10control state. So these two researchers
- 4:13collected data from fast food uh
- 4:15restaurants uh in both state before and
- 4:18after the wage change. Um so and they
- 4:21found that they compare the trends uh in
- 4:24employment in both state over time and
- 4:27they found that employment did not fall
- 4:30in New Jersey in fact is slightly
- 4:32increased. So this suggests that raising
- 4:36the minimum wage didn't hurt low wage
- 4:38jobs in that context. So again uh this
- 4:41is the basic DID idea comparing before
- 4:45and after differences between a treated
- 4:47and the control group to estimate the
- 4:50effect of a policy.
- 4:52So now let's take a look at what the
- 4:55regression setup looks like for the
- 4:57study. On the top is the basic 2x2 did
- 5:01where all treated units uh receive the
- 5:04policy or intervention at the same time.
- 5:07So for this figure, the x-axis shows
- 5:09time and yaxis shows individual units.
- 5:13Each square represents the treatment
- 5:15status of a unit at a specific time. In
- 5:18this example, all treated units receive
- 5:21treatment starting at time 16. Sorry,
- 5:24the text is too small. Um but so in the
- 5:28basic 2x two did setup the regression
- 5:32includes a uh indicator for treatment
- 5:35group uh post treatment indicator uh a
- 5:39post treatment period indicator and also
- 5:41their interaction which captures the
- 5:44treatment effect. So this is the classic
- 5:46before and after comparison between
- 5:48treated and control group. But in many
- 5:52real world settings uh we know that
- 5:54treatment doesn't happen all at once.
- 5:56Instead different groups adopt the
- 5:58policy at different times. This is
- 6:01called uh staggered treatment or
- 6:03staggered adoption. Um that is what the
- 6:06bottom graph shows. Uh again time is on
- 6:09the x-axis and the uh units are on the
- 6:12y-axis. But now you can see that
- 6:15treatment starts earlier for some units
- 6:18and later for others. So in this case we
- 6:21often use generalized DID model uh
- 6:24estimated using two-way fixed effects.
- 6:27So this model include unit and the time
- 6:31fix effect and a treatment indicator
- 6:34that turns on when a unit is treated. So
- 6:37actually this uh DGT is equivalent to
- 6:40the interaction turn uh here in the
- 6:43simple 2x2 DID.
- 6:45So now let's look at uh how the ID is
- 6:48actually used in real world policy
- 6:50research more recently. Um a well-known
- 6:54example comes from the affordable the
- 6:56affordable care act ACA Medicaid
- 6:58expansion. Um for context Medicaid is a
- 7:02public health insurance program in the
- 7:04US that provides coverage for low-income
- 7:07individuals and families. It's jointly
- 7:10run by federal and state governments.
- 7:13Under the ACA, which was passed in 2010,
- 7:17states were given the option to expand
- 7:19Medicaid to cover more low-income
- 7:21adults, but not all state expanded that
- 7:24at the same time. Some adopted in 2014
- 7:28as shown in figure A and some states
- 7:31adopted later and a few didn't uh expand
- 7:34it uh at all until 2020. So this creates
- 7:39a natural experiment. So we can compare
- 7:42expansion state to non-expansion state
- 7:45before and after treatment and depending
- 7:48on the timing researchers can use both
- 7:50type both type of did setup for example
- 7:53in early studies say in 2015 uh when
- 7:57most extension state just adopted the
- 7:59policy at the same time uh in 2014. So a
- 8:03single 2x2 did uh is okay. So in later
- 8:07years as more states gradually adopted
- 8:10expansion researchers became using uh
- 8:12stagger did to account for variation in
- 8:15treatment timing.
- 8:18So as mentioned earlier uh parot train
- 8:21assumption is the key for the study. Um
- 8:24there are multiple ways to assess that.
- 8:28For example we can uh plot the uh raw
- 8:31data uh for uh by treatment status. uh
- 8:35for example uh meter and the wary 2017
- 8:39estimates the effect of the first two
- 8:41years of ACA Medicaid expansion on
- 8:44health and access to care. So um they um
- 8:49plots the raw data uh for expansion
- 8:53state and the non-expansion state uh for
- 8:56all the outcomes. And in this figure we
- 8:59can see that before 2014 uh when the
- 9:03policy was effective we see that lines
- 9:06were moving in parallel between uh
- 9:09treatment and control states. So that
- 9:11which support the parot trend
- 9:13assumption. So that means it's more cred
- 9:16credible to attribute the post204
- 9:19divergence to policy itself.
- 9:24Another useful tool for checking
- 9:26transumption is the event study design
- 9:29which is a quantitative way to examine
- 9:32parot transumption and the virtualized
- 9:34uh dynamic effects. So we have seen this
- 9:37uh did regressions uh earlier. So um
- 9:42instead of having a post indicator uh an
- 9:46event study replace that with a set of
- 9:49event time indicators uh which are
- 9:52variables that measure time relative to
- 9:55the treatment both before and after it
- 9:57happens. So for example, if a state
- 10:00expanded Medicaid in 2014, uh then 2013
- 10:04will correspond to event time uh equals
- 10:07to minus one indicating one year before
- 10:09treatment and 2015 will correspond to
- 10:12event time uh one indicating one year
- 10:15after treatment. So the event study
- 10:18allows us to do two important things.
- 10:21One is to diagnose pre-trans. we can
- 10:23look at the coefficients for the uh
- 10:26periods before treatment and the second
- 10:29thing is it can it allows us to see how
- 10:31the treatment effect evolves over time
- 10:34by looking at post treatment periods. So
- 10:38uh on the right we have two event study
- 10:40plots from uh meter at all 19 uh 2019 uh
- 10:45which estimates the Medicaid expansion's
- 10:48effect on uninsured rate and annual
- 10:51mortality. The uninsured rate is on the
- 10:54top panel and annual mortal mortality is
- 10:56on the bottom panel. Uh so for these two
- 10:59figure x6 shows event time uh which is
- 11:02time relative to when each state
- 11:05implemented uh medic expansion yaxxis
- 11:08shows the estimated effect. So in both
- 11:12cases uh we can see that the
- 11:14pre-treatment coefficients are close to
- 11:16zero which supports the parot
- 11:18transumption and after the treatment we
- 11:21see clear policy effect uh a decline in
- 11:25uninsur rate and also gradually a
- 11:28reduction in mortality. So uh besides
- 11:32parotrren assumption there are also many
- 11:34other important things to check to make
- 11:37the result more convincing. uh for
- 11:40example uh we want to look out for
- 11:43compositional change. So um repeated
- 11:46cross-sectional data are mostly used in
- 11:49DID studies. Um in cross-sections uh the
- 11:52individuals um sampled in each period
- 11:55may not be the same. So if the
- 11:57composition of the sample change say the
- 12:00average wage or um income level shift
- 12:04between pre-post periods that will
- 12:06confirm the result. Um to check for this
- 12:09we can examine whether the distribution
- 12:11of observable characteristic like age,
- 12:14income or race, education uh remains
- 12:16stable over time within each group. If
- 12:19not, we may need to adjust for those
- 12:22coariates. Next we have uh robustness
- 12:24checks. So it's always good to do as
- 12:27many robustness check as possible. Um so
- 12:30one of them is falsification
- 12:32uh test. Um
- 12:35so this is about using an alternative
- 12:38control group that shouldn't be affected
- 12:40by the policy or intervention. For
- 12:43example, uh in Medicaid expansion study,
- 12:46uh we can look at individuals age 65 and
- 12:49over who are already covered by Medicare
- 12:53what is which is uh a universal health
- 12:55insurance for elder people in the United
- 12:58States. or we can uh look at uh
- 13:00highinccome individuals who are not
- 13:03eligible for Medicaid. So um after
- 13:06comparing you know state differences uh
- 13:09within these groups if we see no
- 13:11treatment effects in this group that
- 13:13supports our identification assumption
- 13:17strategy sorry. Um so another approach
- 13:20is to use a placebo outcome um something
- 13:24that logically shouldn't respond to the
- 13:27policy. So for example um um a policy
- 13:31that expands health insurance probably
- 13:33shouldn't affect outcomes like type 1
- 13:36diabetes which is genetic. Um if we
- 13:38found no effect on the placebo outcomes
- 13:41that again supports the validity of our
- 13:44design. So this kind of check help us uh
- 13:47rule out spirious associations and
- 13:50increase our confidence that we are
- 13:53identify a true causal effect. Now let's
- 13:56move to next section. um the limitations
- 13:58of two-way fixed DID uh with
- 14:01heterogeneous treatsment effect which is
- 14:03a major topic of recent debate in the
- 14:06literature.
- 14:08So recent economics uh literature on DID
- 14:11has shown that using a traditional
- 14:14two-way face DID in settings with
- 14:17staggered treatment timing can lead to
- 14:19bias result especially when treatment
- 14:22effects vary across unit or over time.
- 14:25what we refer to treatment effect
- 14:27heterogeneity.
- 14:30Um so before we discuss the limitations
- 14:33uh what are those limitations it's
- 14:35important to understand how the two-way
- 14:38fixed backd is actually implement
- 14:41implemented uh and then we can see the
- 14:43problem. So actually uh two-way fixed
- 14:47feedback DID model uh contains many many
- 14:50simple 2x two did comparisons and the
- 14:54final estimate is obtained by
- 14:57aggregating those uh treatment effect
- 15:00calculated from all these pair wise
- 15:02comparisons. So I will use an example to
- 15:06illustrate that. So running a two-way
- 15:09basic back uh model is equivalent to
- 15:13performing the following procedure.
- 15:16First uh we will identify switchers. Uh
- 15:19these are groups that switch from
- 15:21control to treatment. Uh for example,
- 15:24consider we have four groups and four
- 15:27time periods. Uh group NT is never
- 15:30treated. Group one, two, three are
- 15:33treated at time one, time two and time
- 15:36three. So we identified group one two
- 15:39three are switchers at some point. So
- 15:42for each switcher uh we will compute the
- 15:46prepost change using the prepost uh
- 15:48treatment periods and then uh we will
- 15:51also find every possible control case
- 15:54where treatment status doesn't change
- 15:57over those two period and compute the
- 15:59pre-post change and then uh we will
- 16:03compute the treatment effect for all
- 16:05possible 2x2s and then aggregate them
- 16:09together. So uh for example uh for group
- 16:13one uh which was treated at time one. If
- 16:17we want to calculate a treatment effect
- 16:19at time one uh we can use never treated
- 16:22group as a control and compare the
- 16:25differences uh between them before and
- 16:28after time uh uh uh treat uh treat uh
- 16:33before and after time zero. Uh similarly
- 16:37we can calculate the treatment effect
- 16:39for group one at time two uh at time
- 16:42three using never control uh never
- 16:44treated uh group and this is just a few
- 16:47example there are many many other
- 16:49examples as long as the control group
- 16:53doesn't change their you know treatment
- 16:55status um uh the two that will include
- 16:59them in the comparison and later we will
- 17:02see why that is problematic.
- 17:06Um so uh as uh Goodman Bacon 2021 points
- 17:12out um the
- 17:15uh problem like the bias uh estimate
- 17:17actually sterns from the forbidden
- 17:20comparisons. So we mentioned there are
- 17:22many many comparisons um but some are
- 17:25clean and some are forbidden. So um I
- 17:29will uh so now we can see and for
- 17:32forbidden comparisons um time varying
- 17:34treatment effect can create big uh
- 17:37problems. So uh we can take a closer
- 17:40look at the clean comparisons and the
- 17:42back comparisons and see how the problem
- 17:45happens.
- 17:47So clean comparisons refers to uh cases
- 17:51where a treated group is compared to a
- 17:55not yet treated or never treated group.
- 17:57So for example, a good comparison
- 17:59include using never treated group as a
- 18:02control when calculate calculating the
- 18:05treatment effect for group one across
- 18:08all periods which we just talked about.
- 18:10Similarly, uh never treated group can
- 18:12also be used when calculating the
- 18:15treatment effect for group two and group
- 18:17three uh as shown in the second row. Um
- 18:21a group comparison can also include not
- 18:24yet treated groups. For example, for
- 18:26group two, uh group two and group three
- 18:29can serve as control when calculating
- 18:32the treatment effect for group one at
- 18:34time one because group two and three
- 18:37haven't been treated yet. And also uh
- 18:40group three can serve as a control when
- 18:42calculating the treatment effect for
- 18:45group two since group three since group
- 18:48three hasn't been treated yet at that
- 18:50time. So these comparisons are valid
- 18:53because the control groups are truly
- 18:55untreated during the relevant time
- 18:58window. Uh which helps ensure that the
- 19:00parallot trend assumption holds.
- 19:03Uh now let's take a look at the
- 19:05forbidden comparisons. Um forbidden
- 19:08comparisons refer to cases where a late
- 19:11treated group is compared to an early
- 19:14treated group which is has already been
- 19:17exposed to the treatment. For example, a
- 19:20bad comparison involves using group one
- 19:23which has already been treated earlier
- 19:25as control group when computing the
- 19:27treatment effect for group two and group
- 19:30three uh as shown here. So in in the
- 19:34next slide you will easily see how this
- 19:36type of comparison um can distort the
- 19:39result when treatment effect change over
- 19:41time due to violations of paral trend
- 19:44assumption.
- 19:46So from this figure we can see that B is
- 19:50early treated C is uh later late
- 19:53treated. So if we use B the early
- 19:56treated group as control group to
- 19:58calculate treatment effect for uh C it's
- 20:02clear that the perot transumption uh is
- 20:04violated. So even though we know that
- 20:07policy actually has a positive effect on
- 20:09C the forbidden comparison will give
- 20:12wrong answer um if if we compare you
- 20:14know pre-post difference uh among these
- 20:17two two units um the the wrong answer
- 20:21will even show now or even uh negative
- 20:24effects in this case. So um actually
- 20:29like even treatment time uh it even
- 20:31treatment effects remain constant over
- 20:34time but if it differ across group it
- 20:37can still introduce some bias but that
- 20:40is less concerning than the effects uh
- 20:42vary over time and we will also see this
- 20:45uh later.
- 20:48So um so far we have talked through the
- 20:51problems in an intuitive way. Um there's
- 20:54a lot of technical and mathematical
- 20:56discussion in the economics literature.
- 20:59Um I won't go into that. Uh that is very
- 21:02very complicated but I think it's
- 21:04helpful to uh know the general takeaways
- 21:07from this literature. So um basically
- 21:12the two-way fixed effect did DID
- 21:13estimate est estimator tends to
- 21:16downweight the effects for groups that
- 21:18are treated for longer periods and for
- 21:22time periods when more groups are uh
- 21:24treated. So this means if the treatment
- 21:28effects are larger among early treated
- 21:31groups or larger in later periods then
- 21:34two-way fixed bet may underestimate the
- 21:37average treatment effect. So and also in
- 21:41some extreme cases two-way fix effect
- 21:43can even produce negative estimate even
- 21:46if the treatment effect is positive uh
- 21:49for each group in each time period.
- 21:53So here is a an example from my own work
- 21:56which found the two-way fixing that the
- 21:59ID underestimate the treatment effect.
- 22:02Um so the policy contact is school
- 22:04desegregation in the United States. Um
- 22:07in 1954 the US Supreme Court ruled that
- 22:11school racial segregation was
- 22:13unconstitutional.
- 22:15Um, this led to a decades
- 22:19of court ordered desegregation aimed at
- 22:22improving school quality and assets for
- 22:25black children. However, since 1991, the
- 22:29Supreme Court issued several rulings
- 22:31that made it easier for school district
- 22:33to be released from this desegregation
- 22:36orders. After that, um, racial
- 22:38segregation in schools began to rise
- 22:40again. So in this study, we examined the
- 22:44impact of this policy reversal on black
- 22:46children's health. Um since school
- 22:49districts were gradually released over
- 22:51time, the setting naturally fits within
- 22:54um staggered adoption framework.
- 22:57So for the first stage we tracked
- 22:59whether release from court oversight
- 23:02actually led to increased school
- 23:04segregation and we compared traditional
- 23:06two-way fixed effect uh estimator with
- 23:09uh newer heterogeneity robust uh methods
- 23:13which we will talk about in the next
- 23:14section. So we found that the two-way
- 23:17fixed effect estimator produce smaller
- 23:20and less significant effect size.
- 23:25Um so so far we have seen the issue of
- 23:28two-way basic bed uh when treatment
- 23:30effect is heterogeneous. Um in practice
- 23:33this kind of uh heterogeneity is very
- 23:36common. Uh for example um the effect of
- 23:39a policy might be different across uh
- 23:42groups due to different group
- 23:44characteristic or timing of being uh
- 23:47treated. Also the effect might change
- 23:50over time. uh for example there may be
- 23:52lacked policy effects or the effect
- 23:55might depend on timing for example a
- 23:58policy or program might work better in a
- 24:00good economy than during a recession. So
- 24:03this tell us that to produce more robust
- 24:06result we need method that handle uh
- 24:09heterogeneity.
- 24:12So given the limitations the field has
- 24:15developed a range of new estimators that
- 24:18are robust to treatment effect
- 24:20hathogenity. Um in this section I will
- 24:23walk through some widely used uh
- 24:25methods.
- 24:29Um so uh popular estimators include uh
- 24:33for example colorway and s Anna 2021
- 24:36buzzac uh BJS 2024
- 24:40s and abrehan 2021 exacture. So there
- 24:44are many many uh methods proposed but
- 24:47they all have the key uh their key
- 24:49strategy is very similar. So in the
- 24:52first step um they use uh clean control
- 24:55group that we just mentioned earlier to
- 24:58infer the counterfactual outcomes for
- 25:00treated units. So avoid using forbidden
- 25:03comparisons. Uh and the second step um
- 25:06they will compute the treatment effect
- 25:08for each unit and then aggregate them to
- 25:11the final target uh parameter.
- 25:15Sorry.
- 25:18Um so uh as mentioned there are so many
- 25:21uh uh methods um uh the those method
- 25:26generally can like they vary uh a lot in
- 25:29implementation but generally they can be
- 25:31classified into three uh general
- 25:34approaches. Uh the first one is group
- 25:37time uh estimator approach proposed by
- 25:40uh colorway and sana 2021. So this is
- 25:44very similar to what we just uh dis uh
- 25:46talked about before like finding those
- 25:48uh clean comparisons calculating every
- 25:51possible 2x2s
- 25:54uh and then aggregate them together. Um
- 25:56but uh we I need to mention that for
- 26:00this method CS method um the the the
- 26:04reference they just use the last
- 26:06pre-treatment period uh as the reference
- 26:10periods against the all like
- 26:12pre-treatment periods. So uh in this
- 26:15case um this method relies on weaker
- 26:19parotren assumption um it only require
- 26:22paral trends between the last
- 26:24pre-treatment periods and the post
- 26:27treatment uh periods.
- 26:29Uh so the second approach is imputation
- 26:33approach proposed by BJS 2024.
- 26:37So um uh the first step they in their
- 26:41this method is to feed um two-way facing
- 26:46only observations for units and the
- 26:49periods not yet treated to impute a
- 26:52counterfactual outcome for each treated
- 26:54unit in the absence of uh treatment. So
- 26:58this is for example this is the
- 27:00regression. uh the sample only use not
- 27:03yet treated uh units uh to run this
- 27:07regression. And after obtaining
- 27:09parameters, we can use these parameters
- 27:11to predict the counterfactual outcomes
- 27:14for treated units. And then we can you
- 27:17know compare the uh true you know actual
- 27:20outcome with this counter counterfactual
- 27:23outcomes for treated units to uh
- 27:26calculate the treatment effect. And then
- 27:28finally like we will uh also aggregate
- 27:31them to and over OAT. So different from
- 27:35uh CS BJS uh method relies on stronger
- 27:39PTA uh which is they require parallel
- 27:42trans hold for all groups and all
- 27:45periods uh because this their approach
- 27:47use average outcomes across all
- 27:50pre-treatment periods as a baseline. So
- 27:53there are some tradeoff. So uh if we use
- 27:57like all the pre-treatment time periods
- 28:00um if paral trans like the strong PTA
- 28:03host then the BJS method is more
- 28:05efficient and more precise but if PTA is
- 28:09violated then these methods might be
- 28:11like more more biased. Um also like um
- 28:15there are uh the third approach is about
- 28:18regression based approach uh which
- 28:20basically means using regression models
- 28:23uh specifically designed to avoid the
- 28:26issue arise from forbidden comparison.
- 28:29Um I'm less familiar with this approach
- 28:32and thus won't go into the details but I
- 28:35want to mention that um all statistical
- 28:39package are available online and pretty
- 28:41easy to implement and empirical evidence
- 28:45have uh shown that those method
- 28:48generally produce very similar result.
- 28:52So um in our method method
- 28:55methodological paper on recent advances
- 28:58in DID aimed at researchers in public
- 29:01health and epidemiology
- 29:03we conducted a simulation study to
- 29:06compare the performance of different
- 29:08estimators. Um the results are shown
- 29:11here. Uh we compared the traditional
- 29:15two-way fix effect estimator with four
- 29:17heterogeneity robust alternatives. Um so
- 29:21the first row um and we we simulate like
- 29:24based on different scenarios. Um so the
- 29:27first row represent scenarios where
- 29:30treatment effect are constant over time.
- 29:34Um the second rows represent scenarios
- 29:37uh with dynamic treatment effect which
- 29:39is treatment effect change over time.
- 29:43The first column shows case where
- 29:45treatment effects are the same across
- 29:48groups. So heterogeneous treatment
- 29:50effect. Second group Colin introduced
- 29:53random variation in effects across group
- 29:56which means like effect may be large for
- 29:59some group small for other group but the
- 30:01generally it's different it's random and
- 30:04for the last column it reflect case
- 30:07where early treated group experience uh
- 30:10larger effect.
- 30:12So this setup allows this setup allows
- 30:16us to assess how each uh estimator
- 30:18performs under different types of
- 30:21heterogeneity.
- 30:23So um in this figure um the x row
- 30:27represent the distribution of bias which
- 30:30is um like the difference uh um between
- 30:33the true value and the estimated
- 30:35treatment effect using those estimators.
- 30:39So we found that two-way fixed effect
- 30:42estimator is the most efficient option
- 30:44when treatment effects are constant
- 30:46across both group and time as shown in
- 30:50uh scenario one a. So uh it's very
- 30:54precise and less um the variation of
- 30:56bias is very low. Um so um however when
- 31:01treatment effect change over time as
- 31:04shown in the second row two fix uh
- 31:09estimator produce uh is very biased. Uh
- 31:12while heterogeneity robust estimators uh
- 31:16mators yield more accurate and a
- 31:18consistent result. Um we can also see
- 31:22that robust estimators tend to produce
- 31:25very similar estimates.
- 31:28Then uh we can also look at scenario 1B
- 31:31and 1 C where treatment effects are
- 31:34constant over time but vary across
- 31:37groups. We see that when difference
- 31:40across group are random two-way B effect
- 31:43still performs reasonably well. But when
- 31:47effects are systematically larger among
- 31:50earlier treated group two-way physic
- 31:52become more biased.
- 31:56So up to this point we have focused on
- 32:00the limitations of the traditional
- 32:02two-way fix effect estimator um
- 32:05particularly how it can go wrong when
- 32:07treatment effect are heterogeneous and
- 32:09we also have discussed about servo uh
- 32:13heterogeneity robust estimators. Um in
- 32:16the final section I want to uh shift
- 32:19gears uh slightly to discuss some other
- 32:22concerns that apply to DID designs more
- 32:25broadly.
- 32:30So an ongoing discussion is the
- 32:34conditional paral assumption. Uh so in
- 32:37practice it's more uh plausible to
- 32:40assume that parot train assumption holds
- 32:43conditional uncertain observed
- 32:45characteristic. Um a common approach is
- 32:48to include those these covariates in the
- 32:51regression uh such as in the in a di uh
- 32:54two-way fix did model.
- 32:57However this approach has some
- 32:59limitations as pointed out by recent
- 33:02literature. Um so the first one is the
- 33:05back control problem. Um so researchers
- 33:09often include time varying covariants in
- 33:12two-way basic back regressions. But in
- 33:15that case if the treatment effects um if
- 33:19the treatment affects those covariants
- 33:22uh then including those time varying
- 33:25coariants in the model would introduce
- 33:27bias. So this happens because um if the
- 33:32treatment has an effect on covariates
- 33:36then um the covariants may act as both a
- 33:40confounder and a mediator uh meaning
- 33:42that the uh estimated effect captures
- 33:45not only the direct effect of the
- 33:47treatment but also the indirect effect
- 33:50through the co-variate. Um so um there
- 33:54are pos a few possible solutions uh for
- 33:58example uh do not adjust for that
- 34:00covariate but this may make the parot
- 34:03trend assumption less likely to hold um
- 34:07also we can use uh pre-treatment values
- 34:10of the time varying covariate instead
- 34:13and also a recent working paper by
- 34:16kayatano uh at all 2022 um they propose
- 34:21an approach approach where paral
- 34:23assumption is conditioned on the
- 34:26untreated potential value of the
- 34:28co-variates. Uh but this method is still
- 34:31under development.
- 34:35Um second uh two-way feedback
- 34:37regressions do not condition uh the
- 34:40parot train assumption on time invariant
- 34:43coariates uh which we know is absorbed
- 34:46by group and time uh fix usually uh so
- 34:51um if these time invariant coariants
- 34:55have a time varying effect on the
- 34:58outcome then the two fix the uh
- 35:01regression adjusting for those coarian
- 35:04uh without like excluding those time
- 35:07invariant coaries may be you know uh
- 35:10biased. So for example um consider um
- 35:14race or gender as time invariant
- 35:17characteristic. So while these do not
- 35:20change over time their influence on
- 35:22outcomes for example uh income and
- 35:25health or employment can evolve due to
- 35:28structural or societal shifts. um if we
- 35:31don't account for how the effect of this
- 35:33coarage change over time the PTA uh
- 35:37parot train assumption may be violated
- 35:39and the DID estimate could be biased. Um
- 35:42so a possible solution is to include an
- 35:45interaction term between the time
- 35:48variant coariant and the time
- 35:51and finally uh two-way basic regressions
- 35:55only effectively control for change in
- 35:58time varying coar coariates over time
- 36:01but not their levels. So in other words,
- 36:05if two units have very different
- 36:08baseline value for comparable uh for a
- 36:11coariates, uh two-way fix effect may
- 36:15still treat them as comparable um as
- 36:18long as the covariate change similarly
- 36:21over time. Um for example, imagine uh
- 36:25two states with very different baseline
- 36:27unemployment rate. If both states
- 36:30experience the same percentage point
- 36:32change in unemployment rate over time,
- 36:35two-way fix that will consider their
- 36:37trends equivalent. But the levels the
- 36:40level difference may influence how a
- 36:43policy affect outcomes say health or uh
- 36:46income security meaning the comparison
- 36:49is not truly uh valid.
- 36:52A possible solution to this is to employ
- 36:55procedures that can match each treated
- 36:57unit with a uh control unit with similar
- 37:00or identical covariate values.
- 37:05Um another concern
- 37:09that has received increasing attention
- 37:12is testing the parot train assumption.
- 37:15uh as mentioned earlier um the event
- 37:18study design is a commonly used
- 37:20quantitative method to assess PDA uh
- 37:23parot transumption. However, uh
- 37:26researchers have pointed out that it
- 37:28often suffers from low statistical power
- 37:32um especially when the event window is
- 37:34wide and each period contains only a
- 37:37small number of observations.
- 37:40This low power can lead to
- 37:42underrejection of the now hypothesis. uh
- 37:44meaning we may incorrectly include
- 37:47conclude that trends are parallel even
- 37:50when they are not. Um in addition uh
- 37:52traditional parot transumption testing
- 37:55methods focus solely on pre-treatment uh
- 37:58trends testing and offer no guarantees
- 38:01about post treatment dynamics which are
- 38:04equally important for credible causal
- 38:07inference.
- 38:09to address these issues um um um these
- 38:14two authors proposed the uh the ond
- 38:17framework. So rather than um assuming
- 38:21exact parallel trends on DID allows for
- 38:24limited violations of the assumption and
- 38:28use sensitivity analysis to assess how
- 38:30robust the estimated treatment effects
- 38:33are to uh to those violations.
- 38:37So more specifically um the methods
- 38:40imposes a restriction that bounds the
- 38:43extent of uh post treatment trends uh
- 38:46deviations to be at most m times the
- 38:49size of pre-treatment difference in
- 38:52trends. So sensitivity analysis is then
- 38:55conducted across a range of uh m uh
- 38:59which captures varying degree of
- 39:01potential uh bias and their final result
- 39:05will produce a set of confidence
- 39:07intervals for the estimated treatment
- 39:10effects across uh this M. Um this helps
- 39:14researchers evaluate whether their
- 39:16findings remain statistically
- 39:19significant even when allowing for some
- 39:22violations of paral trend assumption. Um
- 39:25so currently this approach can be
- 39:27implemented using the uh colorway and s
- 39:30an estimator in both R and STA and for
- 39:35me I think um I haven't seen a lot of uh
- 39:38list uh uh researchers conducting only D
- 39:42in the field of public health and
- 39:44epidemiology
- 39:45and I think it is definitely an area uh
- 39:48that will um receive more attention
- 39:51because it can like you know uh more
- 39:53sensitivity test can uh improve the
- 39:56credibility of our uh findings
- 40:01and also there are many many on other uh
- 40:04other ongoing uh topics uh for example
- 40:07like uh continuous treatment. So so far
- 40:11in today's uh spark we simply focus on
- 40:16uh uh treatment when treatment status is
- 40:19binary but uh how about when treatment
- 40:22uh is continuous uh then that is another
- 40:26field and also like some uh studies have
- 40:29been looking at how to apply uh triple
- 40:32differences uh in this um uh like this
- 40:37relevant field. To wrap up uh here are a
- 40:40few key takeaways.
- 40:42So first um if policy imple implement
- 40:46imp implementation is staggered it is
- 40:48recommended to use heterogeneity robust
- 40:51did estimators. Um second um we already
- 40:56see that the choice of heterogeneity
- 40:58robust estimators often has minor
- 41:01practical uh differences. Uh so just
- 41:04choose uh the method you prefer. But
- 41:07also note that there are some um you
- 41:09know differences. So um just you need to
- 41:13like uh based on your more specific
- 41:16context to choose which method you want.
- 41:18Um and third um we probably have to
- 41:21carefully consider coariate adjustments
- 41:24especially when covariants may be
- 41:27affected by treatment. And finally uh if
- 41:30we have concern about violations of the
- 41:33parot train assumption then maybe it's
- 41:35good to use uh on DID to conduct a
- 41:39sensitivity analysis. I think that is
- 41:41all of my presentation and thanks again
- 41:44for being here. Uh I will welcome any
- 41:47questions and discussions.
About this transcript
This page contains the full transcript of Recent Advances in Difference in Differences Methods for Policy Research by Department of Social Policy and Intervention Oxford, generated from the public captions YouTube serves with the video. The transcript has 5,491 words across 846 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.