Video summary
Synthetic Control Modelling (SCM) is introduced as a vital methodological tool for evaluating place-based initiatives in scenarios where randomized controlled trials are unfeasible due to the unique characteristics of single units, such as Local Government Areas. The core concept involves constructing a "synthetic control," which serves as a weighted average derived from data across multiple comparable areas known as the donor pool; this synthetic unit is engineered to closely mirror the historical trends of the intervention area prior to any specific event or program implementation. By utilizing software based on R programming, researchers can calculate precise weights that generate an accurate match, allowing for robust feasibility studies where a strong alignment in pre-intervention data followed by divergence after the program begins provides evidence that observed changes are attributable to the intervention rather than external noise or existing trends.
The application of this method is illustrated through significant real-world examples, including the analysis of West Germany's economic growth following reunification with East Germany and various philanthropy evaluations such as child care hubs. However, SCM comes with specific limitations; it requires long time series data spanning ideally seven to ten years and functions best for single units rather than multiple ones simultaneously. The technique struggles when external shocks affect all areas equally and can only determine whether an effect occurred without explaining the underlying causal mechanisms. Furthermore, the method cannot be applied if the outcome of interest lies outside the convex hull formed by the donor pool, meaning it is unsuitable for extreme cases where no combination of comparable units could theoretically match the target area's distinct characteristics or trends.
Beyond its technical constraints and limitations regarding differential influences, anticipation effects, and unit interference, SCM faces challenges in error estimation since standard confidence intervals are technically unavailable; instead, researchers must perform placebo checks using R code to verify that the treated unit shows a significantly higher ratio than all other potential units with near-zero p-values for fabricated data. When compared to difference-in-differences models or randomized trials, synthetic controls offer flexibility by allowing trend lines to diverge without strict parallel assumptions, though human intervention remains critical throughout the workflow from formulating research questions and selecting variables to conducting sensitivity analyses that alter inputs like including all Queensland LGAs. Ultimately, SCM is positioned not as a standalone solution but as a complementary tool within an analytical arsenal, supported by downloadable templates for testing feasibility during the modeling stage to ensure rigorous evaluation of public policy impacts.
Read the full video transcript
If at any point
if at any point in time you have a
question that comes up and you are and
you do turn your camera on, it will be
recorded live. Um so welcome again. I'm
Caroline Hemwood. I'm the um convenor of
the AESV Victorian Regional Network
Group and it's a great committee and
we've arranged this seminar tonight um
with George and Squirrel. Thank you very
much and I'm very excited about it. So I
currently have two laptops set up. I'm
gonna hopefully host a little bit as
well as uh play along at the same time
and if that doesn't work, I'll have the
recording to come back to. Um I know a
lot of you are probably already AES
members. Uh so you're no doubt aware
that the AES conference is also coming
up in September and that's being held in
CRA. So that's from the 15th to the
19th. The 15th and the 16th some
workshops and then the 17th to the 19th
is the conference itself. So if you
haven't already, it's a great program.
some really exciting keynote speakers.
So, jump online on the AES website, have
a look, and if you're interested, sign
up and join us there. Um, and for those
of you that aren't members and want to
come along, I'd recommend becoming a
member because it will work out in terms
of the cost at the end of it all. Um, so
just a reminder, and I know Squirrel
said it a couple of times, so hopefully
you're all here because you want bit of
a gentle introduction to synthetic
control modeling. And if that's not what
you're here for, it's okay to leave
right now.
Um
>> or stay. It'll be great.
>> Or stay and learn something new. That's
exactly right. Um so George and Scorell
have really kindly offered to give us a
practical sum seminar um exploring the
use of synthetic synthetic control
modeling in kind of that real world
evaluation context and I'm really
excited about it. So uh for those of you
that uh don't know who they are, George
has recently commenced as a senior
associate for Rooftop Social. and I
think his passion for cap capability
building um and helping other
practitioners will probably be really
critical in some of the work that he
does there. Um he prior to that he has
worked in philanthropy. So he's worked
with the Paul Ramsey Foundation. He's
worked in the university sector teaching
and he's also worked with a large number
of organizations to really improve
evidencebased decision making
evidence-based decision- making as well
as being an author um and having a PhD
in economics. So he's got quite an
impressive record there and squirrel too
has a very impressive record and she's
about to start with Asel Allen but she
has also worked extensively in
philanthropy both at the Paul Ramsey
foundation and prior that at the Ian
Potter Foundation and she actually
founded Australia's philanthropic
evaluation and data analysis network
which is now known as LEAF. So if there
are any funders in the um room, there's
also a evaluation network that's really
focused on um funders and that's thanks
to Squirrel.
She's conducted evaluations for schools,
outdoor ed programs, government,
childare settings, etc., etc. The list
goes on. Um she's got a masters at
Stanford University in evaluation and
policy analysis and holds a masters in
journalism around data infographics and
um info in investigative. I'm not sure
what I've done there, squirrel, so you
might need to come back in and and
uh correct me. And she also has her PhD
from the University of Oakuckland as
well. Um, so I'm going to leave us all
in their capable hands. I will also just
let George and Squirrel talk through how
we're going to manage questions and
answers on the way through as well. So
over to you, George.
>> Thanks, Cass. Um uh and nice to be
joining everyone here from Gatigolands
here in Sydney. Um so I'll just start
sharing my screen so we can begin
hopefully. Now you will get a copy of
the presentation after this. So so don't
don't worry about taking notes or
anything like that. Everything you see
on the screen, you'll get a copy of that
uh as a PDF afterwards. So, I just want
people to listen if if if at some point
you're struggling to keeping up. You're
probably not going to be alone. So, just
let me know. Happy to go over things.
Like I said, this is a gentle
introduction to synthetic control
modeling. And I appreciate that many of
you here probably don't even know what
the hell that is, and that's fine. Um,
we're going to ease you into it. Now, by
way of background, Squirrel and I were
both recently um at the Paul Ramsey
Foundation, and when I joined there,
people asked me what's available to
evaluate place-based initiatives. And
place-based initiatives was becoming a
big thing, not just at the foundation,
but around Australia. And I'd heard of
this thing called synthetic control
modeling. I'd read a little bit about
it, and I thought, oh, it's widely used
in the US, not widely used in Australia.
And so we looked into it further and we
got to apply it in for some of our
grants and squirrel might talk about
that a little later. Um
so we we in that through that
application I realized it's got some
really valuable ways of um understanding
impact in certain contexts and certain
conditions and this presentation is
meant to give you a little bit of an
introduction as to what it's like and
even play around and with some and do do
a little bit of it. So let me let me run
through the the presentation. So the
problem it's something say we're going
to launch a program in the LGA of Casey
in Victoria. Um let's call that the
treatment area. Just start introduce
some of the techn technical terms. and
we want to do something in Casey and
there's an outcome and for that outcome
we've got an indicator and for that
indicator
bigger numbers are better than lower
numbers.
And and maybe in the chat people can put
in something or jump in an outcome of
interest to you where numbers going up
is better than numbers going down.
Anyone want to jump in or put something
in the chat that um that indicates
something like that.
Satisfaction with life, youth attendance
rate in schools. Yep, they're all great.
Great. They're wonderful. Um
good examples. So just imagine we want
to do something in Casey that's going to
improve any of those examples.
Now one way of of dealing with that
is to look at the historical trend in
the data for that outcome or at least
its indicator and say we we've got this
picture. We're going to introduce an
intervention in 2025
in Casey and we've looked at the
historical data before 2025 and it looks
like this. It's kind of going up which
is good but not in a consistent linear
way. Now the problem we have is say
after we introduce the program in 2025
and we continue to collect data
say it looks like line A it's continues
to go up well if you're positive Pete
and you look at that and you say hey
look at that we wanted these numbers to
go up and sure enough they kept going up
and therefore the program has worked
now negative Nelly might look at that
and say hang on hang
it was already going up before the
intervention. It just continues the same
trend. And you could get into an
argument about what that positive growth
really means. Is it due to the
intervention or is it just a
continuation of what was already
happening?
Similarly, we might get the numbers that
represent line B after the intervention
and negative Nelly looks at those and
she says, "Look, it was going up before
the program. You introduced a program
and now they've gone down." And
conversely, someone might say, "Oh, no.
There was just something else that
happened in 20 256
and it caused the numbers to go down."
But it's now starting to go back up
again. And again, you can sort of see
the problems here. It's hard to infer
from a trend in the past that's been a
bit erratic what a future trend after a
program really means. And instinctively
as you're looking at these lines,
probably many of you are thinking, well,
don't you need to compare it to
something?
Don't we need to see what happens in
Casey with some other comparison group
that allows us to make a judgment? And
so what you might do and again introduce
some some terminology
compare Casey or any unit of analysis
that you're interested in with other
places or groups that are similar to it.
And we technical term for those other
possible comparison comparators is the
donor pool. Never quite sure why
economists come up with fancy labels for
things that have common sense terms, but
that's what they tend to do. So that
other possible set of donors is called
the donor pool.
Now you might say a good donor pool for
Casey is Frankston.
So let's look at what's happened to
Frankston in the past on this outcome.
And you can sort of see Frankston is a
dotted line. It was above for a period
then it dropped below.
similar but different enough to make
comparison in the future a bit tricky.
Say we
I'm looking for my tools here on thing.
I can't find them so I won't bother. Say
the lines continue to go up for Casey
and continue to drop below for
Frankston. Is that due to just a
continuation of this trend that started
to emerge beforehand or is it really due
to the program? Again, hard to really
disentangle what is the noise from what
is the underlying change that a program
may have brought about. So, you look at
that and you think, "Oh, bugger.
Frankson's not a really great
comparison. It's close, but not close
enough. Maybe we should compare Casey
with Dandenol.
They're similar but again they kind of
converged in the last year. Is any
divergence after the intervention after
24 24 just a reversion to the
differences that always existed or is it
due to the program?
Again similar problem that we've had. If
we compare one place to another place,
they're not perfectly aligned in the
past. Which means if they're not
perfectly aligned in the future, what
sense do we make of that?
So
maybe there's another solution to this
problem.
Maybe we can come up with a combination
of data that gives us a better track
tracking of Casey than if we can compare
to any one case. Um
why don't we take a kind of average of
those all those other um comparator
areas and hopefully that average will
more closely resemble Casey's historical
track record. Now, if that's if you're
still not quite sure about that, a
metaphor that I've come up with is a bit
like, you know, you've got a wall in
your house that's painted a certain
color, but you don't have that pot of
paint anymore. You've got a pot of
green, you've got a pot of blue, you got
a pot of red, pot of pink. Can you come
up with some combination of those paints
that looks really really close to that
that um one you've already got on the
wall? That's what we're trying to do
with synthetic control modeling. We're
trying to take some kind of average from
the donor pool that really closely
mirrors what's happened to the the place
of interest we want to look at.
So hopefully in anticipation of this
presentation, people got an Excel file
and I'm going to jump out of my
presentation
and go to it. And this is a chance for
all of you to play along.
Um, so I'll share my screen. Sharing the
presentation
because this is a chance for all of you
to sort of do a little bit of this.
So hopefully you can see on my screen
and maybe some of you have this open on
your desktop.
So I've got data that exactly what I
showed just before. I've got a data
series for Frankston. I've got a data
series for Dandelong.
I've compared each of them separately to
this data series for Casey. And I found
that they're not really that great at
approximating what's happened in the
past.
So I think let's take a simple average.
I multiply Frankston's score by a half.
I m multiply Dandelong's score by a
half. And I get a weighted average of
those two. And I do that for each year.
You can see simple calculations. Most of
you could probably do that in their
head. And down here in the graph, it
shows you what I get when I compare
Casey to that weighted average. Now, I
just want to pause there because
suddenly we've done something very
different.
I'm no longer comparing Casey to any
real life
comparator LGA. I'm comparing it to a
mythical LGA, a weighted average of
these other two. It's not a real place
anymore that I'm comparing it to. I'm
comparing it to a construct, what we
might call a synthetic control.
Now, what you might do is say, "Well, I
still don't think those lines are really
close enough by using a simple average."
And you can play along here. You can go
to cell B3 and type in any number
between between 1 and zero.
I'll type in 7
and that recalculates. And you can see
the line moved.
Now, in I I'd like each of you just to
type in some numbers between zero and
one for Frankston. And if any of you
think that you come up with a weight
that really really gets those two lines
really close together,
just write Frankston and the weight you
used. And we'll come back to that in a
moment. So you can all play along and
see if you can come up with a set of
weights that produces two a line that
sits really really close to the real
line for Casey.
I'll just give each of you a moment. Try
lots and lots of different combinations.
And like I said, if you think you come
up with one that's really close, drop
it. Drop what you've come up with in the
chat.
>> Chocolate fish fries when we see you at
the AES conference. Yeah. Now, we could
do this. This is the way we could go
about it. Whenever I want to do a
comparison of one area with another, I
can get a whole bunch of nerdy friends
together. I can give them the data and
we can all just type in lots of
different possible um weights. You know,
we could start at 0.001,
0002, go right up to point 999 and find
the one we think gives us a historical
series for our synthetic control that
really closely matches the intervention
area. Now, while I might think that's my
idea of a good time, there's a better
way of doing it, and that's using a bit
of software, bit of programming, not too
complicated. Anyone can do it. And I'm
going to hand over to Squirrel who will
show you how you can do that. So we what
we're going to do is get a computer to
do all the possible weights that we
could conceivably come up with and pull
out the one that it says is the closest
approximation
to the historical data. Squirrel, over
to you.
>> Amazing. Thank you, George. Um, and
thanks everyone. Before I dive right in,
just so I get a sense, scale from one to
five of how much you've used R as a
programming language. Five being like,
I'm a tutor on the positive course. I
breathe R in my spare time. One being
like, look, I heard that you needed to
download the package for this
presentation. Haven't downloaded it yet.
Two being you've downloaded it, but you
really haven't used it much.
All right. So, we've got some twos and
threes. Love the people have downloaded
it and a couple of fours. Fantastic. Um,
this session is being recorded. So, for
this portion of it, if there's a bit
that you're finding a bit confusing or
tricky, you can always go back and watch
the recording again, play along with,
and if there's something that you're
really stuck with in the future, you're
running it yourself, um, and Chat GPT
can't get you out of it, by all means,
you'll be able to send me an email. I'm
happy to work with you. All right, so
here's the fun. George says, "Hey,
squirrel, there's this thing going on in
America and in Europe of synthetic
control modeling. Think it'll be handy
for placebased stuff." Same. I read the
academic literature and I'm like, "Oh,
yeah, that looks great. It'll be easy."
And it is easy to run, but it's also
really easy to run fast when you have no
idea where you're going. So, I will show
you some of the running. Um, I'm sharing
a screen right now that has the zipped,
unzipped, um, folder that was sent to
you. There's one that looks a bit more
like a box or a package with the letter
R in it. And I have clicked twice on
that particular one. Um, so again,
that's a demo one. And it didn't end in
R, it ended in our project. That'll just
make it easier for everyone to see. Now,
because I did see a lot of twos and
threes, I'm going to go slightly slower
on this bit. Um, forgive me all you four
folks. Um, so down in the bottom right,
you should see some files and you can
click on the demo file. This will
actually give you a whole bunch of code
there. Going to run through the first
few bits really smoothly here. So
these libraries are basically a bunch of
different packages that help the our
programming in the back do things
quickly. It helps it read Excel
spreadsheets. It helps it do synthetic
control modeling. If you need to, you
can go to tools, install packages, and
make sure that you have read Excel
diplier tidyr etc. So if you don't have,
for example, the synth package already
in, just hit that. I don't know what's
going to happen because I've already
installed it, but hit install and so you
go, right?
Then the next thing that's going to
happen is we're going to load and
prepare the data. Now, you might be
saying, god, I don't even know how to
program our how would I do this squirrel
if I weren't um watching you.
You can look at my projects in my chatbt
file and you can see where I was like um
I uploaded the f file of George's
information and I asked Chatty I'm not
I'm not going to like I can program in
R. I'm I am a tutor on the posit course
but I just wanted to show you all that
it's really simple and you can get chat
GPT to run the code for you. Um, so just
want to let you know you can get
yourself in a little bit of trouble this
way too. And I highly recommend learning
a bit more, but um, all I did was ask
sort of, you know, how can you do can
you please run this? And it's like yes.
Okay. I wanted to share that for
transparency and also to let you know
how I use AI in this situation. And I I
would love to get up to the Oh, there we
go. up to the top. Hi, I want to do a
synthetic control model demo. Can you
please import the data, clean the data,
use the synth package controlled for
uncalled and plot the data? My LGA of
interest is Casey. The cut year is 2020.
And then it does. So I just thought for
those of you nervous about coding, um
the future is now and the barriers to
entry to computer coding are much lower.
So then what you would do, you can
highlight things and press run. There
are many shortcuts on the keyboard, but
let's just roll with that. And what that
does is uploads the data.
The next thing you would want to do
is synthetic control data. The package
in R needs to look a particular way. Um,
specifically, it kind of just needs to
have all the years going down like that.
STE stack stack KYC stacks XAC
Frankston. You'll see that's slightly
different from the data we were just
playing with George. Yeah. So I'm just
showing that that I demoed it long.
Now
the next thing I want to do is I want to
talk about a real life scenario because
we used to commission evaluations,
right? And so let's say we've got
Frankston and we've got Casey.
And in the raw data, your commissioner
comes up to you and says, "Actually, I'd
love for you to compare to Cardinia
because it's really close in seifa
levels. It's really close in the percent
of cold, and I think that Cardinia is
going to be an amazing comparator,
right? And in in most situations, you
might have to put it in, right? But this
is a tool now that will let you actually
look like when you eyeball the data, you
can kind of see Cardinia starting way up
high and dropping way down low. Like
it's not looking the same as the others.
Your gut is telling you that the
commissioner is is, you know, smoking
crack for asking you to do this, but you
can't actually say that politely. So
this is a way where you can with numbers
go through and verify what might go in.
Okay. So then basically you need to one
silly little trick of it all is you
actually need to have an ID number for
everything. So just bear that in mind.
Um it can be a pain if you're using
boxar data or some type of government
data set. You might just in your Excel
need to throw down the numbers or
program it into the R. Then synthetic
says you prepare the data and it's got a
bunch of things that equal something
loosely. If you look at the things like
the predictors, this is where you could
say um I want to control for safa and
cold. I would like for it to be the
mean. That's the most common to be
honest. You could use random things like
standard deviation or something, but um
in this particular example that I've
sent you, I pretending that the program
took place in 2020. Um, just want to
note that George was looking at the
whole historical trend line, but I
wanted to show you um, wanted to show
you what would happen from 2020. So then
you chuck in your time predictors of
before and then you throw a plot in.
Okay, great. Um, the dependent variable
is the outcome. The unit variable is
that number I told you about. It has to
be a number. The unit name variable is
your LGA, which was area. Time variable
is year. So you're just adding these in
for whatever data set you have and you
run it correct.
Then the next thing you do once you've
told it all the things that you want
because you might for example have a
predictor um for maybe percentage of
first nations or maybe the number of
diabetic people depending on what you're
running.
Then you run the model that you've just
where after you've put in all the things
and it it runs it like boom there done.
So I want to show in plain basic it has
a really basic plotting package and you
can plot the model and again you can get
tattoot to do this or you can go into R
into the synth package and read it
along. The instructions are really clear
as a package goes. I can see what George
was describing here. We have matched
these two beautifully.
The synthetic Casey, that combination of
red and blue and and white matches my
lavender beautifully. And so I am quite
content to say, okay, looks a bit like
the outcome is getting higher.
Let's say I need to put some numbers to
it. This is not the final answer because
remember we have um thrown in another
LGA because I didn't want to give it
away anyone who was working ahead. So um
I showed some weights. Great. There's
the outcome in a very very clunky
manner. Now let's say I would like to um
I have a commissioner who wants the line
to be red and different and I I want a
few more things. Again ask can you
please make me a graph with a pretty red
line? Right. Um then if you want to
create a data frame or a table for lay
audiences
again you can do that and you can see in
the example with Cardinia that it's kind
of 57% Dandenon 33% Cardinia and about 9
and a half% Frankston is your synthetic
Casey in this hypothetical example where
the program took place in 2020.
Um, I'm not going to go through the rest
just yet. Well, I'm not going to go
through the rest in horrible detail, but
there are ways and things to check and I
do want to talk through a little bit of
that. So, I'm going to start a
PowerPoint um talk through for a few
moments and then George is going to
pause me shortly.
Um, one of one of the examples when I
was reading international literature,
loved it. There's such a thing. Somebody
did a study on a synthetic Karl Marx.
Yeah. and he found that Karl Marx is
actually half niti. Um, a little bit of
Roberta Rouso, bit of Abraham Lincoln
thrown in. I thought this was um, a
hilarious example.
Everything in the literature made it
look so easy. I want to talk a little
bit about some of the practicalities.
Okay, so I used this many times in the
work that the Paul Ramsey Foundation did
a lot of placebased work. So I'm just
thinking of four different examples. So,
in one example, we were looking at an
employment program that operated in
three different LGAAS. Yes. And one
thing that we found was it was only like
one or two comparisons. You only needed
a little bit of um blue and a little bit
of red. There wasn't a lot of like green
and yellow and lavender. And the the
team got a little nervous. They said,
"Well, this is quite you know, you're
we're building a synthetic model, but
we're really only using these other two.
Is that all right?" Yes, that's actually
very normal. It's called sparse weights.
You can almost suppress anything that's
under 0.5. You can stick that into the
formula if you want to. Um and that is
absolutely acceptable and if you read
the literature part of it. Another thing
um we did a synthetic control for a
childare hub and it looked like there
was that beautiful separating of the
lines, right? Um the lines separated
nicely but the effect size was actually
quite small. And so one of the things
that we learned is that you can use a
difference and difference estimator. Now
that's a super technical term. Um those
of you not statistitians don't need to
overly stress if that's that's too hard,
but the idea is that you can estimate
what that effect size would be. Um and
using that difference and difference
estimator we found was pretty good.
Problem is pretty graph showed a split.
Split actually didn't have a very big
effect size. So it was one to again talk
through and think through
like any good evaluation tool. You can
take these numbers and what was
interesting is um thinking through that
story behind the numbers and they were
able to say hey we actually know why the
split isn't so strong here. We've been
having trouble recruiting um a partic
you know a person um to work in this
particular role. Um, so we're not
getting those outcomes because we're
just, you know, shifting through, we'll
just say like reliever educators. Um, so
it was again a beautiful evaluation
moment where you had a tool in a quote
crystal clear graph. Um, and what came
of it. Um, we had a hub that was brand
new, wanted to run it, but you need more
than three data points on either side.
So if you've only been running for one
year and your data is collected
annually, you're not going to have
enough data. The ideal is five or more
on both sides. But if you you you know
if you can get five years prior data to
that cutline and then you still need
three more points. If your data is
quarterly, you're in luck. You can run
it after three quarters, but we did find
we couldn't run it on that brand new hub
because it didn't have more than three
points. And then lastly, we were running
it on some justice reinvestment
initiatives. And it was uh fascinating
again. Um, we showed it to someone in a
particular LGA and there was no
difference
and they got really quiet and went away
and I just thought, "Huh, this is this I
don't know what to say about this."
Anyway, built relationships better with
them, went out to site visits, came back
later on and talked to them and said,
"What was that?" And they said, you
know, absolutely we know why that was.
um we have our youth case worker burnt
out and left and we we actually know
that like this thing that was working so
very well for a little bit isn't working
and we're going to need more wraparound
support and at least two case managers
to make this type of program work which
was great because then as a foundation
we were able to fund a bit more but
really important to know like that cold
hard graph um there is a story behind
all of the numbers
um I'll keep going but George interrupt
if I'm getting too dull
>> and jump maybe I can switch back to the
uh example.
>> Yeah, why don't you switch back to the
example and then I'll come into these um
>> yeah the nitty-gritty shortly.
>> I noticed that Peter you asked it it yes
the answer is generally it is used for a
time series for an aggregate so it might
be an area and you aggregate student
level outcome data or you aggregate
something within a particular area over
a time series. So just answer to your
question. So I'll just jump back to my
presentation. Um
sorry bear with me.
Okay. So hopefully you can see my
presentation again. Um
so coming back to the original example I
started with which is I want to compare
Casey with some other area. I couldn't
find any real areas that made an
adequate comparison. So I give the data
to someone like squirrel and she runs it
through that program and the program
tells me that the best weight I could
use is waiting franken data by a factor
of 767
and thenong 233. Those two numbers
should add up to one. And if you like,
you can go to that Excel spreadsheet and
you put those word weights in and you
should get a graph that looks just like
the one I've got here on the screen. And
what that is, so we're now comparing
Casey to a synthetic Casey, not a real
place, but a a fictitious place. And you
might look at that and say, "Oh, wow.
Yeah, they're really close to each
other, although they're starting to get
apart towards the latter stage." Now
this is the question of is synthetic
control modeling feasible now to use for
the data that then rolls in after the
program is in in is implemented.
Are these two lines close enough for it
to be usable as a way of evaluating the
program once it's rolled out? And you
one of some of you might look at that
and say, "Yeah, I think they're close
enough now so that any divergence
afterwards can be attributed to the
program or not." And some of you might
say they're close, but I don't think
they're close enough. So, we tested
whether synthetic control modeling was
feasible, but we looked at the best
line, and even the best line is still
not close enough. It's a bit like that
paint example. You want to match it to
something that looks a nice shade of
pink. And you've tried to mix yellow,
blue, and green in some combination. And
no matter how hard you tried, you didn't
get a shade of pink that was close
enough to the real wall you want to s um
uh match.
Now, I won't you can debate that, but
let's go to an example that actually has
been applied. This is kind of the case,
the classic synthetic control case
example where the the modeling before
the intervention did produce a synthetic
control that really closely matched the
real example. So back in 1990, the
Berlin wall came down. The former West
West Germany now joined the former East
Germany into a larger country. And the
question that arose that this analysis
needed to answer was what was the impact
on the former West Germany's economic
growth now that it had to take in and
support a relatively poorer region which
was the former East Germany.
So hopefully everyone understands that
scenario. A rich area was growing. Now
it has to support a relatively poor area
with which it is now one nation. What
was the impact on that on the former
West Germany's economic growth? So you
can see before unification in 1990.
The black line is West Germany.
Um hopefully you can see my my my
cursor. Um,
if you take a simple average of the
other 16 OECD nations, and they're over
here on the right, you can see that West
Germany was growing faster and getting
further apart in terms of its growth
rate compared to other advanced
industrial economies.
After unification with East Germany, it
seemed to slow down and converge. And
you look at those lines and there's
probably a simple conclusion. Yeah,
having to absorb a relatively poor area
slows your economic growth. Again, we've
got economists doing in a very
sophisticated way, something that common
sense would probably tell us anyway. But
can we measure the impact of that
unification on the former West Germany's
economic growth?
So, they did a synthetic control model
and over here you can see the weights
they applied. So, for countries like
Australia, they gave it a weight of
zero. Basically, it wasn't included in
the synthetic version. They multiplied
Austria, its economic growth rates over
this whole time period by a factor of42.
And you can see the other nations that
were included in the modeling over here
on the right. And that particular
combination of weights
gave us a line that sat right on top of
the the line for the real West Germany.
So that if we apply the same weights to
what happened after unification, we can
see
the dotted line represents what we think
would have happened to West Germany if
it didn't have to absorb this other
poorer region. It's a counterfactual,
but because it did have to absorb it,
the line dipped down and West Germany,
the former West German region, only grew
by this amount. And we can actually
quantify the difference, the impact that
unification had on the growth rate of
the former West Germany.
So this has become the classic kind of
example of synthetic control modeling in
the literature. Um, what I'd like to
hear from you at this point, and you can
either put up your hand or just jump in,
is is there something you're doing in
your workplace having heard these
examples where you think, "Oh, yeah. I
think I could use synthetic or someone
could do for me a synthetic control
model to tell me whether a particular
intervention affected an outcome or not.
feel free to jump in, put up your hand
or just unmute yourselves. Hang on, I'll
go to the chat.
>> I might jump in with a uh and it's not
specific, but I certainly think that we
use AEDC data a lot. Now, that's quite
different to what you've been showing
here, but in terms of that five-year
incremental, I think that's probably one
of those ones where we could look at
what's it look like program
>> as is referring to the Australian Early
Development Childhood Index. We use that
at the the the the
program at the the Ramsey Foundation and
it's a measure when kids start school,
are they ready for school in terms of
certain measures of cognitive and and
other developmental um factors. And you
could say we're going to introduce a
program in preschools to try and get
kids more developmentally ready.
and we look at the data in the area
where we're introducing those those
school-based measures and we compare
them to areas where it isn't and
hopefully we can see that um
developmental readiness improved in our
intervention areas and that's exactly
the kind of thing um
Ry asks does the program try to get the
best weights prior that's prior to the
intervention that's exactly what it does
so there's
What I've been doing up to now is
talking about you do an assessment on
historical data prior to intervention to
understand whether synthetic control
modeling is feasible. Is it legitimate
to use this approach to then compare
data after the intervention?
Um
Ben, you've got a statewide road safety
strategy with measure of fatalities.
Yeah, that's a a good example. you
introduce a driver safety or road safety
program in certain areas. It could be
statewide and you compare New South
Wales with other states or some weighted
combination of other states data on road
fatalities.
Um
Sophie, I'm not quite sure what that one
means. You might want to jump in and
tell us a bit more about that.
>> Nope. the policing intervention the
effect it would have on crime in
different areas.
>> Yeah. So policing is a classic if you
read a lot of the American literature
where synthetic control modeling has
really taken off a lot of it has been
applied on understanding whether
policing very timely given what's going
on in the US um uh whether policing
approaches do reduce crime rates in
particular areas. So you can see it's
almost synonymous with place-based
approaches. It is conceivable to use it
in other contexts, but it's almost
invariably used in place-based and um
evaluation.
Um any other thoughts or or observations
before I talk about sort of wrap it up
and and have a talk about its
limitations and other benefits
or examples that I think people where
this could be applied.
So George, looking at panel A and panel
B, you could explain what was going on
in panel A is this is my lay person's
interpretation.
>> And then on panel B, you're saying that
the
sort of extent of the
uh full line being below the dotted line
is the more sophisticated
account of that. Is that right? So the
dotted line is sitting right on the hard
line in panel there.
>> So So without the weight,
>> the OECD line sits below it.
>> But then we apply these weights over
here
>> in there. It lifts that dotted line up
and makes it match
>> y
>> the real what happened in West Germany
before the unification. So coming back
to Ray's point, this tells it's a fe
feasibility assessment first. Oh yeah,
we got a great line that sits really
close now. Therefore, after the
intervention, any divergence we could
probably say is due to the intervention.
Is it due to the program?
>> Um I see Tom's question about the
estimates of error error and it's a
great segue to my next few slides. So
hold tight and we'll chat about that.
>> Yeah. Um yeah, estim all of that is a
bit trickier here, but um we might have
time to come back to it. Harry compared
the results against results of robust
RCTs.
Um, not that I know of is the short
answer, but there is there is a there is
some literature, sorry,
if if you want to really nerd out, where
they're integrating synthetic control
modeling as one way of comp making
comparisons within an RCT context. So
it's a it's a sort of an area of new
research where rather than seeing you
can do synthetic control modeling or
RCTs, you can also within the context of
an RCT use synthetic control modeling. I
won't go in too far into that because
this is meant to be a gentle
introduction, but they're not
necessarily incompatible. Um so so there
is some new new research in that whole
area. But generally as Rey has mentioned
actually one of the benefits of
synthetic control modeling is precisely
that when you're doing a complex
placebased initiative you're talking
about a single unit. Here we were
talking about Casey
and it's hard to think of an RCT
for you. You can't randomize other
places and compare it to a unit of one
to do an experimental design. it's uh
it's it's it's not really conceivable.
So, generally it's seen as a something
you can do when you're talking about
individual units
like Casey, like West Germany, like
whatever. And um that idea of
randomization is really difficult or
impossible. Um
>> George, we've got a couple questions in
the chat that I reckon I'll be able to
hit with my um presentation. Let me just
wrap up with
>> um a summary of what this is. So,
synthetic control modeling helps answer
the question, what would have happened
if we didn't do this? And it does it by
using a comp credible comparison.
That comparison is nothing real. It's a
construct. It's a synthetic control.
That's where the word why the word
synthetic a virtual twin is the kind of
term I came up with as I was thinking
about this um that combines realworld
data for other units of analysis where
those units didn't experience the
program. So then if before the program
the synthetic control is a close match,
we can then use it to project out after
pro after intervention, use it as a
counterfactual and talk about whether
the program probably did or did not have
some impact. That that's it in a
nutshell. There's obviously some
benefits to this. It's great, as we've
talked about, for place-based work,
as I've repeated over a few times, it's
great for situations
where you want to lock people down in
advance. How much of a difference is
going to be good enough?
And so, you do that feasibility study on
the historical data. Yep, we think the
two lines are close enough. and we're
only going to talk about this program
being effective if those two lines
diverge by this amount into the future.
And you can lock that down in advance.
So there's no kind of game playing and
going fishing for a positive result. Oh,
we saw this amount of difference. We're
going to call that a a good effect. Um
you do your hands are tied by that
point. It's really transparent and data
driven. Um, and as we've mentioned, it's
really good for oneoff or unique cases,
especially when there's a complexity of
things being done in that one place.
But with benefits, there's also
limitations.
Um, like I said, as I've repeated, it's
only valid if the you can get a
synthetic control that matches on the
historical data really closely to the
place of interest. Um, you really need
good data over a long time period. As
we've said, it's it relies on time
series. It requires at least seven to 10
years of worth of data after the
intervention, at least three or four
after sorry, before the intervention and
at least three or four after. Um,
it's great for studying one place, but
really hard to do when you've got
multiple places. And it's really hard to
do if there are surprises, shocks. CO
sets in and affects every place.
Sometimes it'll show up in a common way
but sometimes not. And lastly, the point
I make whenever I talk about an
individual evaluation design,
it only answers one particular question.
It tells us whether there may have been
an effect. It doesn't tell us why or how
that program brought about that effect.
Um like all evaluation designs they can
adv usually answer a very specific
question or set of questions and we need
a complimentary set of evaluation
approaches to answer a more complete set
of questions. I know that sounds like a
tortology but I'm constantly in this
battle see these battles where people
say no this approach is better no this
approach is better usually they're
answering very different questions and
the better thing to do is think let's do
multiple number of things to answer a
more complete set of questions so
synthetic control modeling tells us
whether something may have had an impact
other evaluation designs that you can
run along with it can tell tell us what
may have been the the pathways or the
causes or the the roots for that having
that effect.
Um, just to finish up,
uh, there is some good resources.
There's a by the guy who invented this,
Abadi, if people want to really nerd
out. There's a really great summary of
it. Not too technical, um, free online.
There's a YouTube video you can sit and
watch him give a lecture on this. I've
done it. I enjoyed it. But like I said,
it might not be everyone's cup of tea.
There is a much simpler explanation from
the Washington Post. Now, that's behind
a payw wall, but if you email somebody,
they may have downloaded a copy of it
and may be able to give it to you, but
you know, I'll leave that for you to to
understand who that might be and how you
might go about it. But technically, we
can't give you that link because it's
behind a payw wall. Um and lastly,
we've created a little template where in
that feasibility stage of synthetic
control modeling
when you're testing to see is this worth
actually doing on the data after
intervention.
um you can ask the right question and
you can give it to us a contractor and
say please give the results of this
visibility in this format and that
answers all these questions. Um so
that's just something for you and to to
to have as a bit of a tool to to you can
download it from the link that will be
in the PDF um that you'll get after
that.
So amazing
>> thoughts, questions, comments
>> before because we've got a qu couple of
questions in advance. I might if it's
all right, George, just jump in and um
show the last few slides which actually
it's almost like I paid the people to
ask those questions. Um I'll just share
screen if that's all right.
>> Yeah.
>> So um thanks guys. Loving the questions.
I'm just going to try to hit um the one
that talks about the estimates of error.
And so for those of you um I think
George does a beautiful job of
explaining things quite quite well. Um
and this next few um column answers
might be a bit technical. So don't
stress if it's is a bit tricky, but I do
want to try to answer some of those
technical questions. Um George, I think
you're still sharing your screen.
>> Oh, sorry. Sorry about that.
>> That's all right. Thought I stopped
sharing, but yeah, there's the last bit.
There you go.
>> Beauty. Good to go. Um, so yeah, a few
things to think about. Um, there are, as
George was saying, a couple of
constraints. You can't use it if uh your
LG of interest is outside of the convex
hall, which means it's super extreme.
Um, this is a pity for things like if
you want to see if justice reinvestment
is working in Burke, you cannot use this
method because the uh First Nations
population of Burke is is so much higher
than the others. You actually can't
create a synthetic that that works. And
so um just something to bear in mind.
Don't don't go in promising you can do
something um until you've um had a good
look at the data.
>> Could I could jump in with an example of
that? Um
>> Sure.
>> Yeah. So for example in the Northern
Territory the unfortunately it's
terrible the rate of detention of First
Nations young people on a given night is
so much higher than any other state in
Australia. So say Northern Territory
does something to try and get the rate
of detention of First Nation kids in the
criminal justice system on a given night
wants to lower that. No combination of
the other states data will give you a
close match because it's it's all the
other lines are sitting below Darwin uh
Northern Territories and so it it needs
to have some lines above and some lines
below generally to get a combination of
them that's going to closely match it.
No, that's a great example. Thank you.
Um you can't have anything influencing
differentially. So CO in Victoria was
really different from CO in Queensland.
So you you have to bear that in mind.
You can't have an anticipation effect
where people think something's about to
come and you also can't have
interference between the units. So maybe
like case workers going between
different LGAAS etc.
Now, somebody asked about the um just
check the the wording in the chat, but
you know, broadly
I can't see the chat when I'm doing
this, but you know, is there a um
>> I called it a falsification test of
error or the estimates of error or
confidence intervals? And the answer to
that is no. There are not estimates of
error technically in this. But what you
can do, for example, is a placebo check.
And that for those of you if you want to
look in the R code, you can see I've put
the hashtag placebo check and
essentially you run as if your null
hypothesis was all of them. Run it
through and the red line is Casey and it
is visibly up and higher than the
others. Um that's that's a very
technical way for saying this graph
looks pretty good. Um, so we can look at
all the placebo units and put in a
particular ratio. None of nobody got as
high as Casey's ratio. Actually got us a
p value of like 0.00000000.
Ridiculous because we made up the data.
Um, indicating strong evidence that this
is a real post-t treatment effect. So
the answer to Tamara's question about
estimates of error is there are other
methods that you can use for um, this
one. Okay. Um, somebody else asked about
falsification
like um, discontinu
discontinuity exits. I haven't seen any
studies comparing those, but um, hey,
you can do one. That'd be great. Peter
Bowers, you asked about how does it
compare to difference. I'm going to come
back to this slide in a sec. Um, so
again also as George was saying other
methods may be more appropriate like a
randomized control trial or a propensity
score match ex course and exact like
when you have individual human data or
linked data at the individual level you
might as well do your propensity score
matching type things or course and exact
matching of difference and difference.
You can also use the catch with
difference and difference and this is
both synthetic difference and difference
which is a thing but loosely to answer
that question quickly is there can be
multiple treated units and difference
and difference but they need um not
synthetic but they need to have parallel
trends whereas synthetic control
modeling you can have your lines sort of
going in different directions. So the
answer is it compares it doesn't compare
necessarily but it can be used in very
different scenarios. Um generally when
there's a treated area and when you
don't see those clear parallel uh
trends. No idea about the econometric
modeling with an estimated
counterfactual line. I'll leave that to
George the economist to field in a
second. But I do want to pop in.
>> We're getting close to finishing
squirrel then.
>> Yeah. This will be the last slide then.
Um
Maxine Anv from the Melbourne Institute
was beautiful and helped taught a bunch
of us at Paul Ramsey and Grant's uh
through this. So I want to credit him
with this. Well, he's just spoke through
um the workflow and I want to take a
twist on what he said and say that these
are all points where humans can
intervene. So once again, we're showing
graphs and numbers, but we need to
formulate that research question. Think
about what outcomes we're focusing on.
Check if we have enough time periods.
Have that clean data set, right? Select
the variables for matching. Select those
donor units. Are you going to use every
LGA in Victoria? Are you going to use
every LGA in Australia? You can play
with things and it's interesting. It
will change a lot. So don't don't think
that just because it's numbers, it's not
um there's not human influence.
um the you can be plotting and ready to
estimate the estimating the effect size.
You would could use difference in
differences as a way to do that and do
your placebo tests. And then there's
this 10th step which called a
sensitivity analysis, but basically you
go back through steps one through nine
and you change it a little and you see
if that changes things. Okay, so if I
include all the LGAs in Queensland, is
that going to make a big difference?
etc. Um this will all be available to
you and I'm aware that we're running out
of time. I'm also happy to help. Also
consider asking chat GPT first. Um, but
just want to say like thank you all for
spending your dinner hour with us and
hopefully that was in fact a nice
combination of gentle but also answering
some of those harderhitting and
technical questions. Um, back over to
you George. No, that's right on 6:30.
So, I don't want to take up any more of
your time. Like I said, Squirtle said
happy to fill questions offline as well.
um we will circulate those notes and um
yeah keen to hear how people might use
things.
>> And I'll just wrap up by saying thank
you very much. I am um incredibly
impressed and I am now going to become
even better friends with chat GPT while
I uh play in R with a whole lot of new
data sets. So thank you. It's been
inspiring and and created a great space
I think for people from each end of the
spectrum to maybe play and think about a
different approach that we can use.
another tool in the arsenal which is
always great.