Video summary
The video features Dr. Chris Laskovski, a distinguished scholar and teacher at the University of Maryland, delivering a lecture on the statistical pitfalls known as Simpson's Paradox. This phenomenon occurs when a trend appears in different groups of data but disappears or reverses when these groups are combined. Dr. Laskovski illustrates this concept using several real-world examples, starting with baseball statistics from 1995 and 1996 where David Justice had a higher batting average than Derek Jeter in both individual years, yet Jeter finished with a superior combined average over the two-year period. He further explains the paradox through a medical case study involving kidney stone treatments, where one procedure appeared more successful overall but was actually less effective for specific subgroups depending on the size of the stones, highlighting how aggregated data can be misleading without considering underlying variables.
A significant portion of the talk addresses the famous 1973 UC Berkeley graduate admissions controversy, where aggregate data suggested a bias against female applicants due to their lower overall acceptance rate compared to males. However, when the data was disaggregated by department, it revealed that women tended to apply to more competitive programs with lower admission rates, while men applied to departments with higher acceptance rates. This shift in application patterns created an illusion of discrimination at the aggregate level that vanished upon closer inspection of individual departments. Similarly, Dr. Laskovski discusses COVID-19 mortality statistics showing a higher death rate among white non-Hispanic populations compared to other groups; this disparity was not due to race but rather because a significantly larger proportion of white individuals fell into older age brackets where the risk of death was much higher, demonstrating how demographic composition can skew overall averages.
The lecture concludes with an examination of the "low birth weight paradox," where babies born to mothers who smoked during pregnancy were found to have lower mortality rates within the low birth weight category compared to those born to non-smokers. Dr. Laskovski clarifies that this does not imply smoking is beneficial; rather, smoking shifts the entire distribution of birth weights downward, causing many fundamentally healthy babies to fall into the "low birth weight" category by definition. Consequently, this group includes a higher proportion of healthy infants relative to the low birth weight babies born to non-smokers, who are more likely to have severe underlying issues. The overarching message is that averaging averages without understanding the context and causality behind the data can lead to contradictory conclusions, urging audiences to be cautious when interpreting headlines or simplified statistics that ignore hidden variables.
Read the full video transcript
all
right thank
you all all right welcome everyone hope
everyone's doing well on this afternoon
for those of you who don't know me U I'm
John Berto I'm the associate probos for
faculty Affairs and it really is my
great pleasure to welcome you to the
2023 2024 distinguished scholar teacher
lecture series for those of you who
don't know the distinguished scholar
teacher award was established in 19
1978 to recognize tenur faculty members
who are committed to and have
demonstrated excellence in instruction
and in the research efforts the award is
sponsored and administered by the office
of Faculty Affairs on behalf of pro uh
provos Jennifer King rice and recipients
are are selected by prior DST um um
nominees and recipients I'm very pleased
and honored on behalf of provos rice to
recognize Dr Chris lowski as one of our
new newest distinguished scholar
teachers as we will hear more shortly um
his research his scholarship focused
primarily in model Theory which is a
branch of mathematical logic in addition
to his research um Chris has contributed
significantly to the University's
learning environment uh engaging
students in courses he teaches and
through his mentorship provos rice and I
congratulate um Chris on his this
much-deserved award in recognition it's
my pleasure now to introduce Don Levy
chair of the Department of mathematics
who will formally introduce Dr
lowski thank
you thanks John and uh thank you all for
coming um I'm Doran Levy I'm the chair
of the math department and it really is
a great pleasure and honor to introduce
Chris laskovski our new distinguished
scholar
teacher um you know these introductions
sometimes start with you know the the
usual stuff of like you know all this
boring information you can probably just
read on the backs side of the pamphlet
that you received uh but nevertheless
I'll still mention a little bit of that
just to uh you know honor Chris yet
again so pH uh Chris received his PhD
from Berkeley in 1987 after which he was
a more instructor at mat and then he
joined the University of Maryland in
1989 as an assistant professor doing
mathematical logic we heard that uh ma
mathematical logic has been a big
there's a big tradition in our
department for mathematical logic and at
the time that Chris joined us how many
logicians were there four five okay Miss
by one uh gradually that number went
down and over the years as uh people
gradually retired and were not replaced
by new
logicians uh not that long ago our logic
group
numbered
one um so if you look at all
departmental reviews uh the report of
the logic group was we need more people
right uh which actually happened as we
recently hired two logicians uh
Christian rosenal and artam chernikov
and now from one we have three and
probably one of the best or certainly
one of the best uh logic groups in the
country if not in the world so small
numbers doesn't mean low quality
actually when it it's concentrated
really really well uh you can do
wonders okay um so I mentioned something
about the trajectory of Chris uh coming
to Maryland uh but Chris has been
nothing short of a wonderful colleague
so being a wonderful colleague is a
combination of many things it's uh a
great researcher so a wonderful Mentor
teacher and someone that's in charge and
uh involved greatly involved with our
Outreach activities but really a great
person to have around he's here pretty
much every day not many people are and
just seeing Chris in the corridors and
welcoming him and greeting him and him
greeting everyone it's I mean how should
I put it differently Chris is a fixture
of this department okay um
so what Chris does for research is
something I'm not going to attempt to
tell you about uh which sounds mostly
like my letters for the AP uh committees
so uh and and I think that actually
Chris will also not try to make this
attempt and he promised to give us a
very uh accessible talk uh so I think
all of us know what averages mean and we
learn about averages of
averages uh but let me just say again
that I'm super pleased that Chris has
been bestowed this award of a
distinguished scholar teacher by the
provest office and by the University of
Maryland congratulation
Chris and with that Chris
laskovski okay so hopefully people on
Zoom can hear as well so thank you Dr
BAU and thank you Doran for the
wonderful uh introductions I'm very
honored to be here speaking in front of
all of you and all everybody on
Zoom um I'd like to also just thank um
both uh Doran Levy and Larry Washington
who were both previous DST winners and
they nominated me for this so none of
this would have happened and uh in
particular I'd like to thank my wife
Carol def
Francis uh quite frankly without her
constant love and support most of much
much of my career would not have uh
transpired so thank
you okay so before launching into things
uh just in there was discussion right at
the beginning uh today is mold day happy
mold day to
everybody um again since most of you are
mathematicians you might not have heard
about this uh you can view this as being
a chemist knockoff of Pi
Day uh to think for mole think 6.02 time
10^ the 23rd and I sort of cryptically
wrote this today is
1023 uh uh the 2023 is redundant or
whatever um so uh unlike just reciting
lots of digits of pi uh one tends to sit
around and tell mle jokes uh this this
will play A Part uh as we go
along so um as to what I'm going to talk
about well uh can read from the thing my
main field is model Theory which is a
branch of mathematical logic which is a
branch of mathematics great but this is
really a lousy thing say at cocktail
parties or elevator speech it's really
so number one someone says what do you
do I teach uh what do you teach
mathematics the usual thing majority
oh I really hate math many times they'll
say I got up to level X and then I had a
really bad teacher and then I can't
imagine what life's beyond that okay um
but then there are a few people say math
okay yeah what kind of math and then you
say mathematical logic and you can see
that they tense
up oh no and like look down at their
hands they're sure that the next words
out of my mouth are going to be this
statement is false or some sort of a
paradoxical thing they're they're
looking and trying to remember what's
the symbol the the Vulcan greeting from
Star Trek or something like this uh uh
because for whatever reason pop po press
mathematical logic uh certainly
discusses paradoxes or uh in and around
girdles incompleteness theorem or the
collection of all sets is not a set uh
or Bertrand's Russell a barber can shave
everyone's head except his own just all
sorts of things um and when I teach set
theory I try to make a really special
thing I think the term Paradox is very
poorly aimed it isn't a proof that 0
equals one or that things are all
falling apart but rather when you see
something Paradox it means you need to
be careful it means warning D something
non-intuitive is going to be happening
and rather than bore you with my
research I want to give an instance of
something in popular press which goes is
a a paradox but one should really View
and try to see what is going on because
it really as we'll demonstrate has a lot
of real world applications so with that
as a preamble let's start out and let's
talk about baseball I said real world
but we'll get to more things beyond that
um but there let's take two
uh players David Justice played for a
while uh many years with Atlanta Braves
and then moved uh moved on to um to
other teams whereas Derek Jeter always
played with the New York Yankees what
about them well let's look at the years
1995 and
1996 say start with David Justice and uh
many of you know about uh statistics in
baseball but one of the most popular
things is batting average so you take
the number of hits so in this case 104
and div by the number of times he comes
up to the plate the number of at bats
this is just a ratio and uh in this case
comes out to it really should be
0.253 but no one says the zero so he say
he bats
253 so clearly the larger the number the
better off every uh you're doing so in
1995 David Justice had a uh higher
batting average than Derek Jeter now you
can note here that U for Derek Jeter uh
the number of at bats was somewhat lower
this was in fact his rookie season and
he got called up during the year um so
he didn't have that many at bats but
then now we passed to
1996 and once again David Justice had a
higher batting average than Derek
Jeter so in both years years Justice's
performance dominated that of Derek
Jeter but if you then take the two years
combined if you take uh so the combined
thing say for David Justice uh he had
104 hits in uh
1995 plus um uh 45 hits um and then
divide by the number of at bats is um
sorry
is uh
411 plus uh 140 so this comes out to the
149 over
551 or in other words they bad it
combined batting average of 270 but as
you can clearly see Derek Jeter's
combined batting average is
310 now at first this should seem kind
of odd how could be that David Justice
dominated uh Derek Jeter in both
1995 and in
1996 but when you combine these guys it
switches okay well this is sort of what
we want to be discussing with
this now why does this feel strange well
suppose we just have numbers so we have
a big number A1 and a smaller number B1
and another so A1 is bigger than B2 A2
is bigger than B2 then from that if you
add the two big numbers together you get
something bigger than the sum divide by
two so the average of uh A1 and A2 is
bigger than the average of B1 and B2
whenever um A1 and B is bigger than B1
and so A2 is bigger than B2 great
so but this is not what you're doing
with batting averages remember batting
averages are hits versus at bats so when
you're getting this combined thing same
computation that's here uh you're taking
the hits Plus in 1995 plus the hits in
696 divided by the sum of their at bats
and in general this is certainly not
equal to the lad think about the right
hand side is as in my title the average
of of averages but this leftand thing is
not now I guess uh when I gave a talk
like this before Larry Washington
pointed out this is very close to we
spend a the math
faculty math faculty spends a lot of
time teaching freshmen that this is not
how you add
fractions there's a half issue but even
without that okay so um keep that in
mind so the the the thing that I want
you to go for going forward is that care
must be taken when averaging
averages so uh we're going to get to
more sophisticated things than this I
promise uh but this whole thing this
flipping that's occurring with this goes
by the name of Simpsons Paradox and from
my general comments I'm really kind of
unhappy with the word paradox here
because it's simply that we need to be
aware of what's going on so now I guess
this being the new fangled thing what do
I mean Simpsons Paradox I do not mean
Homer it is not named after Homer
Simpson much as he might like it to be
but rather much earlier Edward H Simpson
of
1951 from
um uh from the UK a rather noted
statistician at the time uh always
whenever you name something in
mathematics there's some ambiguity and
some people uh call this the Simpson Ule
effect from much earlier Ule in 1903 was
sort of aware of this kind of thing and
certainly I prefer the word effect to
Paradox that just saying something is
happening here okay now so this is just
something you should be aware of to
start off that just when you're looking
at data sets uh baseball certainly
provides lots and lots of data sets then
stuff like this can happen so in the
wild which to ISA Chas just means that
just uh just in nature without any sort
of contriving things uh it's relatively
rare but it does occur and uh just other
baseball pairings the most recent I
could find was actually two Red Sox
players Ellsbury and LEL who had the
same phenomenon for two years in a row
one dominating the other I forget which
was which uh you can always check these
just if you just literally Google your
favorite baseball player stats you're
just going to get a listing of these and
you can do these on your own somebody
painstakingly went through cases of
Allstar uh players and found these just
among allars very familiar names to
Baseball fans fans I guess the last
one's a little remarkable with uh Babe
Ruth and Lou garri these were the first
three years of garri season uh career
and he garri actually beat uh Babe Ruth
in batting in three years 1923 24 25 but
if you add up all of those together babe
Bruth dominates uh um l garri in in the
sum of the three okay so there's more to
it than just baseball this can happen in
the wild and it can have some real world
significance so let's just imagine now
that you are a doctor Circa um
1990 you're a small town and uh your
specialty is kidney stone
treatments so now a patient comes in
woman just huge amount of pain and clear
clearly has a case of kidney stones and
you want to know what to do what should
what procedure should you do to uh to to
help uh ease her kidney stones well
you're in 1990 and you're on top of your
game and you've read the following thing
this is at the time it really was the
gold standard for what you do with uh uh
treating kidney stones there was a long
multi-year thing all in the UK uh
published in the British medical journal
and it
compared uh again there are a lot of
words renal Cal calculi is of course
kidney
stones uh and now you could uh gets a
little Grizzly but you can either open
surgery which is you go in and get the
things uh this other percutaneous I'm
not going to embarrass myself but let's
call this PN for the second method it's
a kind of putting in a straw a very thin
thing and trying to suck out the the the
um the stones and now this third newer
thing was this extra corporeal so in
other words outside of the body the idea
is you get a machine and you just hit it
just with a shock wave and the idea is
to try to jiggle things around enough
that you'll just pass the stone uh but
that requires extra equipment and uh
you're in this small town remember so uh
that's out of the picture so you really
don't have the equipment so you're
either going to treat this patient by
open surgery or this PN meth straw like
method what do you do well you look at
this paper and it's very clear uh the
data says that open surgery succeeds 78%
of the time whereas this PN straw
treatment uh succeed needs 83% of the
time done deal
right accept
that if you go in and ask does the
patient have small kidney stones then uh
uh this treatment o the open surgery
actually beats the going in for a straw
93% to
87% uh if on the other hand if the
patient has large it's more problematic
so the probabilities drop but still once
again the OS treatment the open surgery
dominates that of um of of the
PN so in either subcase if the stone is
small you should use open surgery if the
stones are large you should use open
surgery but if you don't know the size
of the stones then that then
that that then you should use the
straw so now seriously so imagine you
are this doctor what do you do so which
treatment do you
use and then even a more basic question
to you is should you even bother to
check whether the stones are big or
small because if so it'll maybe confuse
you about what to do
okay so it it's curious that in this
paper this this this really gold stand P
gold standard paper they don't discuss
this issue at all they just have just
the facts here's the table uh here they
are now admittedly they were rooting for
this extra corpal uh thing uh and much
of the paper is discussing that that was
the new fangled thing and they're
discussing the pros and cons of it but I
just in reading the paper uh this
Simpsons uh idea just isn't mentioned at
all question yeah the difference both
78% and 83% they're both about four out
of
five yeah and that point one could look
at two treatment and say which one is is
dangerous well but but but then then if
you're going to go to the danger route
then then certainly open surgery would
presume be worse but in either case it's
dominating that the success rate is
dominating that okay well anyway so so I
can imagine that that each of you being
a doctor can have different opinions
about how to answer this but at least
it's an
issue uh okay so let's continue
on uh but sometimes and this is going to
be the bulk of the talk uh what appears
to be simpsons's effect
can really be a hint at some missing
causality in the data set and if I've
learned one thing in preparing this talk
and just thinking about things for a
number of years uh at some level
statisticians really don't understand
causality and even even what the
definitions should be for it this isn't
necessarily a failing it's just a really
uh involved problem and there are a lot
of traps
uh okay great so let's first talk with a
completely toy example that will get
things across so question should
students study for a
test yes or no well let's do a scatter
plot uh we're just going to randomly so
the number of hours study is the x-axis
and the score on the test is the Y AIS
and we're just going to look at these
various thoughts and and what do you
conclude here's all of the data you get
a best fit line and
clearly things are not good with
studying the best fit line is decidedly
slopes down another words if you're a
student you should not study for a
test now you can almost anticipate from
the general shape of of of what I had
here suppose I tell you that in this
that all of these top guys are graduate
students and all of these guys are
undergrads in many cases it's the other
way true true but let's okay so now
Within These
subpopulations let's try to get uh um
the the best fit whoops oops oops oops
oops
oops oops oops sorry
uh let's let's let's let's get the best
fit lines uh for this and and
clearly if you are a graduate student
you should be studying for the
test if you are an
undergraduate then you should be
studying for the test but scrolling back
if you're a student you should not be
studying for the
test okay
so one could easily look at this data
set this toy and write two different
contradictory compelling papers about
the conclusion and in fact the thesis of
this by means of subdividing what's
going on the same data set can be used
to justify two different
contradictory
conclusions okay so this was all a toy
uh got a good laugh out of it but this
really came up probably the by far the
most famous example was uh the issue of
graduate admissions to UC Berkeley uh in
1973 so start off with the guts of this
fact in Fall
1973
44% of all male applicants uh to to
graduate school were
accepted but only 35 % of female
applicants were accepted and this is a
pretty huge data set 12,000 these are
the precise numbers so overall 41% were
admitted uh uh but eight like this uh
this if you do any sort of analysis This
is highly statistically significant so
just on a Kai Square value of this this
is 110 it's huge the probability that
this could be happening by by chance is
is very very
small so this certainly would be in the
realm of of the legal system or
something that uh should should Berkeley
be sued here say for uh for sex
discrimination on this on on on how they
get in uh and as you're going through
this certainly there there exists many
parallel situations where uh in in
today's world uh where it has entered
the legal system but given this uh
Berkeley was quite alarmed with what's
going on uh so just for a thought think
to yourself what could be happening to
cause
this uh and uh the Provost commissioned
a report which uh actually turned out uh
the result of it it became a seminal p
uh paper in statistics uh it was
published in science and and it was a
big noise when it came out uh Peter
Bickle all three of these were
professors at at Berkeley uh Peter
Bickle was uh a young at the time uh
statistician and he wrote one of the
standard textbooks for uh introductory
statistics uh Hamill was uh an
anthropologist I couldn't find the affil
the field of of okano uh but they really
studied what was going on and they went
one by one at the admissions data for
each of the 85 departments on campus to
see what was happening with this and
this this all here is is accurate data
with it um now rather than go through
all 85 the top six or the six largest
departments uh on on campus uh I'll fill
in some of these but but a through F uh
you can see at first blush that there is
some wild things so so first of all in a
uh
82% women uh
3734 uh this guy is very close to even
very close to even so uh say really the
only big place where there's a big
difference in admission rate is is in a
uh but if you look at this a little bit
more closely one thing you're going to
observe is that there is a huge
difference by Department in the number
of
applicants say number one or the the a
is is the engineering department and
there were
825 uh people uh males and only 108
females that were uh admitted to it uh
but the admission rate into engineering
was Sky High it was uh uh well even I
guess they favored women somewhat with
this but uh at about
70% uh on the other hand say in English
this is either English or a compendium
of English comp complet or something uh
that there there were almost twice as
many women that that applied rather than
men but note the Stark difference the
admission rate was only 34
35% uh as opposed to up here in the
60s so if you look at item C there uh
there were 560 men versus only 25 women
but yet this was a relatively easy
department to get
into uh and and so on so the the there
are big big swings in the male to female
ratio of applicants and the admission
rate is by far from from from being uh
uniform okay and this is really what's
explaining if you go down to the 85 the
departmental level uh go I'll now quote
literally from from the summary uh of uh
the summary paragraph of their paper so
first of all EX examination of the
aggregate data take everything all
together on graduate admissions to
Berkeley uh shows a clear but misleading
pattern of bias against female
applicants we have this huge Ki Square
score of
110 however when you break it down to
the disaggregated data Department by
Department uh they there were few
decisionmaking units that show
statistically significant departures
from expected frequencies in either
direction and about as many units appear
to favor women as opposed to favoring
men so what's happening is is that women
are were uh were getting tracked or were
applying to highly competitive
departments whereas men on the other
hand were applying to departments which
accepted many more people the ratio was
uh was different
one thing that I found surprising just a
a couple uh sentences down from this The
Graduate departments that are easier to
enter tend to be ones that require more
mathematics in the introductory
Preparatory
curriculum and I'm curious is that true
at Maryland that again this was in 1973
and enough just really to State it that
like engineering I know here admits an
awful lot of people but is the
engineering acceptance rate uh much
lower than uh much much higher than then
for English and how do English and math
compare there's a lot of things that we
can ask our associate provos
here this might again so so just does
this does this final sentence still or
hold at Maryland in
2023 okay so let's finish with all of
that
and have to throw this in what was
avagadro's favorite Olympic
event the mo
Vault none of these are any good but you
have to throw them in every once in a
while okay onward um so now that was all
1973 let's skip ahead almost 50 years to
a really odd thing uh at first blush
with uh coid 19 in the early days of
this and uh this this is brought put
together by a blog post of Dana McKenzie
at
UCLA and uh Jordan Ellenberg who is uh a
professor of math at University of
Wisconsin but you can almost view him as
being a HomeTown guy he was uh went to
high school in in Maryland and was one
of the winners of the uh his of of the
high school mathematics competition very
bright guy um guess we don't want
to put those updates but uh but also I
also had the that was before I came here
but I had the pleasure of actually
teaching him uh while I was at MIT I was
teaching a graduate course and he would
walk over he was an undergraduate at
Harvard and came over to uh to take this
so anyway good guy uh
but let's give a couple of slides of
data from the CDC from early on in um uh
in in in the coid thing so things
started well really really started going
in March 20120 and now um uh so this was
only just the first four or five months
and I know the slide is hard to see but
we're just going to be concentrating
here on the white non-hispanic cases
roughly a third uh
35% of the cases were uh were white
non-hispanic you can't read that up here
is the Hispanic uh white thing almost
the same and the black is here uh those
are the the main bulk ones but by
contrast if you look at the
deaths that happened uh between this
among in the white non-hispanic
category it's gone from onethird roughly
up to a
half now if you just take a look at this
this goes completely against what you
the what everyone was saying in the
newspapers about uh about that somehow
that it the the pandemic especially
early on was really hitting all of the
uh uh uh the minority communities really
hard uh uh and and it was having a
profound effect there but why is it that
uh among the white
non-hispanic you could think privileged
people they had a third of the cases but
a half of the
deaths so what is going
on well let's try to to answer that by
looking at a breakdown of things
again for this march to June CDC data
and here it's it's lots and lots of
cases a million plus cases and 100,000
plus deaths this is not a small data set
by any means but we're going to break
things down by age so say among people
in the 30 to 49 range
26.5 so all of this is for whites not
non-hispanic whites so
26.5 of the people percent of the people
30 to
49 uh had the coid
um
26.5% of the 30 49s who uh had Co had
coid were white um non non-h Hispanic
but in that range even though at
26.5% of the cases the deaths were only
16.4%
so if you look at this a little bit uh
uh a little bit more closely you'll see
well first of all in the zero to four
thing thankfully there were very very
few cases especially early on I read
somewhere that among the 100,000 deaths
13 were in this category early so we
could just sort of ignore this thing it
was really
microscopic but then just going uh line
by line um again if you were were white
then your probability of uh of dying was
less than your compatriots in every one
of these
categories this you can say okay the
better outcomes are due to the privilege
and the better health care and better
diagnosis and these various finger
things to see about your your blood
oxygen levels uh um but still how does
that explain if uh the whites are doing
better in every one of these things how
does this explain the 35 to
49% and the key thing here is the the
key is it's right sort of not written
with this but
uh white people are old
and this can really be seen not in the
the assembled room but of
of uh of in the general population
9% of of whites are 75 or older but uh
of
nonwhites uh only three% are
75% are 75 years or or older
so in these two problematic categories
and these were really where I now need
to put better outcomes in quotes because
quite honestly the death rate was
absolutely huge it really uh uh uh again
the Lion Share of all of the deaths were
really just in these two
categories so what was happening was
that um uh just there were just so many
more very elderly whites that that
dominating what's happening throughout
this and that can be just explaining
this but but you one needs to be really
careful just if you looked at that first
graph uh about in the explanation is
somewhat deeper than um than than than
what might
expect okay so now continuing
on what do you get if you cut an avocado
into a large number of pieces guacamole
yes everyone you're good
good okay well we needed that after the
talking about this and especially since
there's going to be another maybe not
completely cheery topic coming up this
is something known as the low birth
weight Paradox that was studied
excessive uh extensively by this Allan
Wilcox uh um who is a a researcher at U
NIH and Research
Triangle uh so in order to get to this I
need to Define three things first of all
the median weight of a newborn is 3.6
kilograms uh and for Years Gone by uh a
baby was labeled lbw low birth weight if
uh his or her birth weight was less than
or equal to 2.5
kilog and for the chart that's going to
come the mortality rate is of of these
babies is the number per 1,000 that do
not survive their first
year so uh a lot of data was collected
on this and here it was split by WEA or
not the mother smoked during
pregnancy so the mortality rate uh for a
thousand so in the general population so
just among the non
lbw uh uh babies uh for maternal
non-smokers the death rate was 11.1 so
enough roughly a 1% chance of or 99%
let's be positive a 99% chance that that
the baby would survived the first year
uh but among maternal smokers it's it's
slightly worse than that
but if you go to the low birth weight
things then 210 out of a thousand low
birth weight babies die to maternal
nonsmokers but if the if the mother
smoked during
pregnancy then this rate would drop to
114 so roughly get cut in
half
okay and this was published by this guy
Yosi um and I I'll talk about this this
was actually I would say an important
paper in a in an odd way but at this
moment you might want to think what is
going
on and hint I'm not advocating that
mothers smoke during
pregnancy yeah
yeah yeah that's that that's that's
going to be the that's going to be the
kicker and I'll be able to illustrate
this by by a series of pictures but good
uh great so what is going on with this
well let's start off just straight uh
this I couldn't get anything else but
this this is for birth weights uh in
Norway but I think it's basically
Universal um
uh so the average birth weight is like
3.6 kilograms and this is very roughly a
normal distribution however to the left
there's a tail the tail to the left is
much longer than the tail to the right
but fundamentally it's it's a normal
distribution as one would expect the
numbers are large uh what else could it
be okay so on this uh doctors have
arbitrarily
determined that uh the low birth weight
cut off is uh at 2.5 kg which in the
normal in the standard population cuts
off this
tail now here just the Grim things to
think about this includes almost all
pre-term babies but also ones with
genetic problems with uh uh just
something wasn't right in this uh so
here I hesitate to write everything in
red it sounds like things are doomed
when in fact the good news is really 80%
of these did survive for at least the
first year but uh but it's really the
flattened part of this
tail okay now another thing that is
known and has been tested many times
throughout is the effect of mother
smoking through pregnancy this decreases
the average birth weight but it still
happens a lot and uh this this this can
be quantized so it's known that mother
smoking during pregnancy decreases the
birth weight by approximately 200
grams so forget the units on this but
roughly what's happening is we're
getting this normal distribution but
we're shifting it to the left by 200
gram so Peak is going to be uh 200 grams
off and the same general
shape but then what really the problem
is in my mind the definition of low
birth weight is not changed depending on
whether or not the mother smokes so as a
result if we cut off this thing now
there's going to be in the blue a much
bigger tail this is exactly what you
were saying that um a among mothers that
smoke there are many many more low birth
weight babies but also many many of them
are fundamentally healthy I'm not saying
that smoking is a good thing but
fundamentally there's no uh there
there's not much wrong with them so if
you're looking at a ratio think you're
just dumping in lots of fundamentally
health healthy babies and you're putting
them to the left of this divide as a
opposed to the right so to to summarize
this the point is uh because of the
shift in birth weights well at not cor
making a corresponding change in the
definition of low birth weight many more
fundamentally Healthy Babies of smoking
moms are put into this
category uh but among the low birth
weight babies so so thus uh there's a
greater share of fundamentally healthy
babies and that's what's going to cause
that ratio to go down so just to be
clear this does not mean that the
mother's smoking is good for the baby
causing this shift to the
left
however this Yoli who uh was
biostatistician is trained at John's
Hopkins uh
uh moved up uh through the ranks and
then actually in the 1950s he he created
the bio statistics Lab at at
Berkeley so he really had uh a a
national following with this he was also
a
smoker and this was right at the time 68
to 71 was when uh there were just
discussions about smoking that everyone
smoked all the time but but the uh CDC
others were trying to cut back and the
FDA warning labels were were coming on
but he used this data this fact uh
scroll back here uh this this is his
data he was using this as an argument in
favor of mother smoking or saying it
wasn't bad because look here's actually
a
benefit um sight I find this quite
shocking uh but
worse this paper that he submitted got
into the general
press and two titles that I managed to
find in the Boston record American I
don't think it still exists the the
title mothers needn't worry smoking of
little risk to baby from
1971 worse in defense of smoking moms in
this Family Health magazine that was
still a I remember that from being a kid
uh that uh uh this really got into to to
the thing
um
so to see that this effect doesn't all
have to be this smoking versus
non-smoking if you want say a more a
cheerier example of this whole thing uh
so this Willcox who's been studying this
effect uh if you just concentrate on
babies that are born in
Colorado then maybe it's because of the
elevation who knows but uh the
percentage of babies born that are low
birth weight is significantly above that
of the US
population but uh for any other purpose
they're just as healthy as in other
states so I don't mean this to be a a
smoker versus non-smoker th um this is
this I view as being a a cheery cheery
thing I would say if anything among
epidemiologist that the big problem
might be that uh just absolutely fixing
this low birth weight thing at 2.5
kilograms come hell or high water is
what's causing these these seeming um
seeming paradoxes but in any event um
just the takeaways if you want to look
at all of these examples taken together
the main takeaway I would have is that
we're looking at a single data set but
in by doing various means of subdividing
the same data set can be used to justify
contradictory
conclusions so I'd say beware of
headlines or or sound bites uh with this
if there are any journalism Majors here
uh uh that that one should really be
careful about how to uh uh how to
interpret this and then finally to link
it back to my title in short averaging
averages can be a perilous
undertaking so uh with that thank you
very much for listening and I hope
you'll enjoy the snacks in the crossing
the
hall