Cognitive impairment indicators for the neuropsychological test batteries in the CLSA
Watch on YouTubeVideo summary
This webinar, led by Professor Megan E. Okonell of the University of Saskatchewan, explores the development and application of cognitive impairment indicators within the neuropsychological test batteries of the Canadian Longitudinal Study on Aging (CLSA). The study encompasses two distinct cohorts: a Tracking Cohort assessed via telephone with over 21,000 participants and a Comprehensive Cohort evaluated in-person with approximately 30,000 participants. While both groups utilize core assessments such as the Rey Auditory Verbal Learning Test and Animal Naming Test, the Comprehensive Cohort incorporates additional measures like the Controlled Oral Word Association Test and Victoria Stroop Test. A central methodological challenge addressed is the removal of measurement bias related to language, sex, education, and age through the use of normed scores. Research indicated that standard regression models were insufficient for eliminating bias from education and sex, necessitating a strategy that combined stratification by these demographic factors with continuous age adjustments to ensure valid comparisons across different languages and populations.
To accurately summarize battery performance, the researchers developed a Cognitive Impairment Indicator (CI) based on base rate analyses derived from cognitively healthy subsamples. This approach is designed to account for "spurious low scores" that occur by chance, ensuring that classifications reflect true impairment rather than statistical noise. In the Tracking Cohort, cognitive impairment was indicated when two or more out of four test scores were impaired, while the Comprehensive Cohort utilized a similar threshold of two or more impaired scores out of six tests. At baseline, these criteria identified approximately 3–6% of participants as cognitively impaired, varying slightly depending on the specific battery used. The validity of this indicator was further supported by its strong association with self-reported neurological conditions and established risk factors such as diabetes and hypertension, although a gold-standard clinical diagnosis was initially unavailable for direct comparison.
The presentation also highlights several critical considerations for longitudinal analysis and future validation efforts. Challenges include practice effects, shifts in data collection modes between telephone and in-person assessments, and increasing missing data due to attrition among cognitively impaired participants. To address these issues, the study is currently developing a Change in Cognitive Impairment Indicator (Delta CI) for Follow-up 3, which will utilize diagnoses from the Memory Study Substudy as a reference standard before applying similar methods to earlier waves. Additionally, the discussion covers the trade-offs between sensitivity and specificity when using demographic corrections versus individual screening tests, noting that while corrections improve battery-level analysis, they may reduce sensitivity at the individual level if risk factors like age or low education are strongly linked to the condition. The session concludes by emphasizing that modifying screening tests fundamentally alters their psychometric properties, requiring new normative data rather than reliance on existing metrics, and provides administrative updates regarding upcoming webinars and changes to data access fees.
Read the full video transcript
All right. Welcome everyone. I'm Sophie
Hogavine. I'm the data access officer
for the CLSA. Um so thank you for
joining us today um for this webinar
titled cognitive impairment indicator
for this neurossychological test
batteries in the Canadian longitudinal
study on aging definition and evidence
for validity.
Now before we begin, I want to
acknowledge that the CLSA National
Coordinating Center and McMaster
University are located on the
traditional territories of the Missaga
and Hodeni nations and within the lands
protected by the dish with one spoon
wall and palm agreement. As we gather
here today, we acknowledge that the
University of Saskatchewan is located on
Treaty 6 territory and the homeland of
the Matei. We pay our respect to the
First Nations and Matei ancestors of
this place and reaffirm our relationship
with one another.
As attendees of this webinar, I
encourage you to continue your learning
following the webinar and to acknowledge
the original inhabitants of the lands
where we currently have the privilege to
research, live, and work wherever that
may be.
Before we begin, we have a few uh
housekeeping points. Everyone but the
presenters will be muted throughout the
webinar. If you need to change or test
your audio during the webinar, uh you
can click on audio settings on the
bottom of the Zoom window.
At the end of the presentation, we'll
have a question and answer session. Uh
if you have a question for the presenter
during the webinar, please post it in
the Q&A box located in the bottom
toolbar. We'll address the co the
questions at the very end. And these
each question you put in will be visible
to all attendees.
If you have any technical trouble
concerning the webinar, please use the
chat box to communicate with our webinar
team. A feedback survey will be launched
at the end of the webinar and we invite
you to complete it after exiting the
Zoom session. This brief survey provides
us with important feedback so we can
plan future CLSA webinars.
Today's webinar as I said is titled
cognitive impairment indicator for the
neurossychological test batteries in the
Canadian longitudinal study on aging
definition and evidence for validity.
This webinar will be presented by
professor Megan E. Okonnell of the
University of Saskatchewan.
Megan E. Okonnell is a registered
doctoral clinical psychologist and
professor professor of psychology at the
University of Saskatchewan. Her research
explores neurosychological measurement
relevant to dementia technolog tech
technology for remote dementia care and
cognitive aging. She has been involved
in the CLSA for over a decade
contributing to projects and publishing
as a first author on several papers
exploring measurement in the CLSA. She
has published over 144 peer-reviewed
scientific articles and has been
practicing in the diagnostic rural and
remote memory clinic since 2008.
All right, over to you, Megan.
>> Wonderful. Thank you.
And thank you for that kind
introduction.
And welcome everyone to a talk where I'm
going to talk a little bit about things
we've done, but I'm also going to let
you know the things we're working on. So
you have to get to the end to get to the
future stuff. Um,
and I want to also acknowledge that this
work is not done alone. Um we have I
recognized um Helena's here uh in the
audience. So hello but uh we have a nice
team and of course acknowledging the PIs
of the CLSA
uh who I work with on many different
projects.
So of for those who don't know the CLSA
this is a slide for you but just to
remind you we do have two cohorts in the
CLSA. Um they are different. there's
more diversity in the tracking cohort.
Everything was done by the telephone and
the comprehensive cohort um were
assessed in person and they have quite a
few more measures including some
differences in the neuroscychological
measures. So in tracking the measures
that are um administered are are the Ray
auditory verbal learning test immediate
trial and then a recall trial which is 5
minutes after the initial registration
trial. Then the animal naming test is
another test that's in the tracking
cohort. It's scored two ways. So if
you've if you've looked at the CLSA
data, you'll see an AF2 one and an AF2.
We use AF2 for all composites, but
you're certainly welcome to use the two
scoring ways. One is a bit more strict
and one's lenient. That's the only
difference between them. And they're
highly correlated as you can likely
imagine, but you should use one, not
both scores. And then the mental
alteration test. And the mental
alteration test is not a clinical test.
So the rest of these tests are are based
and modified from tests that we clinical
neuroscychologists use. This one is a
more researchbased test and it's in a
it's a verbal analog to a trail making
test where you alternate between
letterers and numbers in in ascending
order and in tracking all high scores
means better performance in
comprehensive as I mentioned there are
differences in not only how it's
measured it's measured in person but
there are additional tests given in the
neuroscychological battery so we have
similar tests to the tracking
um cohort. We have ray one, ray two. We
have the animal fluency with the two
scores. We have the mental alteration
test. But in addition to that, we also
have the controlled oral word fluency
test
um which you get three scores. The score
for F, the score for A, and score for S.
And then we have the Victoria strip test
um which you get total time for three
different u measures for the dots card,
the word card, and the colors card. And
here high scores mean better for
performance except for the stroop test
uh where um um that's reversed.
So
just as a reminder we have different
numbers possible for the tracking and
comprehensive cohort. I know I've put
them up on previous slides but I hadn't
spoken to them. um just over 21,000 for
the tracking cohort and just over 30,000
for the comprehensive cohort. Um and
we have six tests in the comprehensive
cohort and four possible test scores in
the tracking cohort. And these scores
have been normed using a hybrid
approach. So I'm going to spend a little
bit of time talking about that. Um, and
they are also separated for French and
English samples. So you may, one thing I
haven't mentioned is that the CLSA
because it is a Canadawide study, we
have done everything in everything has
to be able to be done in French and in
English. And we do have a fairly
substantial uh French sample who spoke
French for the administration and
responses. Um so we have everything in
the norming um com separated by French
and English but they're also scores that
are stratified by sex education and then
continuous for age. I just want to talk
about this a bit and I want to spend
some time talking to you about why you
should use the norm scores that are
available. So when you request data from
CLSA, they do give you the norm scores.
And I'm going to talk about two reasons
why you should use these norm scores.
One is it removes measurement bias. So
I'll talk about that in a moment. And
two, it allows for analysis at the
individual participant level, which does
allow you a lot more flexibility in what
you do with these data.
So through the work that we've done with
the CLSA, um we know that scores mean
very cognition scores mean very
different things when they're given in
French and when they're given in
English. That likely doesn't surprise
you to hear. Um we also know robust
finding throughout the literature is
that low education can create low
scores. And it's not because it's a sign
of say an impaired cognition. It's just
because education is so highly
associated with scores. So if we're
trying to identify people who might have
impairment, then low education causing
low scores is a bit of a nuisance
variable. So um and age and sex can also
um create challenges when we're trying
to determine if somebody's performance
is in more than within normal limits. So
these can create some challenges and
because of that when we create
comparative normative standards, we use
corrections for these demographic
variables that we know impact the test
performance
and the normative data remove bias due
to language. So we have um a paper that
Vanessa Taller is leading um it's about
to be submitted um that demonstrates
that the test scores really are
fundamentally different in English and
French. But if you use the normed scores
which are stratified for language of
administration, it removes the bias and
does create equivalence in the scores in
French and English. So that means you
can collapse across language if you use
the norm scores and only specific models
remove bias due to education and sex. So
I'm going to talk a little bit about
that. So this is a a paper where we
looked at various approaches to the
creation of the normative comparison
standards for the CLSA and found much to
our surprise that full regression models
didn't remove bias in education or sex.
Um in fact what we had to do is stratify
by sex and by education and use
continuous uh age as a continuous
variable. And then when we did that, we
were able to demonstrate that we no
longer had bias in our test scores due
to education and due to sex. Um so and I
will say that this paper we did most of
that work with the tracking um cohort,
but when we created the normative data
for the comprehensive, we also use this
this this approach. So what we've been
doing is we state because I I didn't
publish um the normative data. This was
more of a methodological uh paper. So
when I refer to the normative data for
the comprehensive cohort, I say that it
was normed using the same approach and I
cite this paper. So I recommend that you
do the same if you're of course using
the normative data.
Um, one of the interesting things, the
implication of the work that we did to
figure out how to remove bias due to
nuisance variables in the cognition
measures was that as I said, regression
full regression models didn't um remove
the bias. So what this means for you is
if you're not using the normative data
and if you don't use similar methods
that we used to adjust for language for
sex and for education you might have
biases in your cognition score. So
either use the normative data or when
you do adjust for these demographic
variables, which makes sense because for
many models of analysis, you do want to
have, you know, these demographic
variables adjusted not just for
cognition, but for all your measures. So
that makes perfect sense um to avoid
this double correction. Um, but you
really should then doublech check to
make sure that you don't have bias in
the cognition measures for language of
administration, for sex, and for
education. Um, so I do recommend that
you read that paper and think about how
those adjustments were done. Um and I
will say I you know might be good to
think about um that maybe that double
correction for sex education and age um
and language of administration might be
the lesser challenge to have to deal
with when you're doing your analysis. So
it really though that depends on your
statistical approach. So I recognize um
I am a clinician who plays with stats.
I'm not a statistician. Um so I uh would
defer to your statistical expertise on
that but I I will say that the full
regression models did not remove the
bias due to these variables. You have to
stratify for that.
Um I also promised to talk about the
second reason why you should use the
norm scores that are available from CLSA
which is because it allows for an
analysis at the individual participant
level. Why is this the case? Well, we
mostly approached the neuroscychological
battery and CLSA like I approach
cognitive testing with patients in the
the the day a week I spend doing a
diagnosis in the memory clinic. Um, and
what normative comparison standards are
really great for is because they allow
us to say, is the person who's sitting
in front of me, is their cognitive
performance within normal limits or is
it something else going on? Potentially
there's evidence of impairment. So, it
allows us to take participants data and
look at them at the individual level.
So, that is exactly what we did with the
cognitive impairment indicator.
So each and this is again these are
variables that come to you when you
request the cognition uh data
and you'll see that for each cognitive
test score um there is an impairment
indicator and that's because we've
compared their performance on that
cognitive test to the normative
comparison standards and have determined
that it's not within normal limits. It's
quite deviant for normal limits. And we
used of course a cut off of the fifth
percentile which was um based on the
empirical distribution not the
theoretical distribution.
That may not matter to you. And with
sample sizes like this, the theoretical
distribution looked really close to the
uh empirical distribution anyway, but
it's just a it's just a more careful
approach. Um so each test score has an
impairment indicator which could be
really useful for analyses if you were
interested in that. Um there also are
composites uh created from each test
score. Um the ray one and ray two and
two memory composite and then the
executive function measures which differ
for tracking. There's only two, the
animal fluency and the mental alteration
test. But in um in the comprehensive we
have not just those but we also have
additional executive function me
measures. the total score from the
controlled oral word fluency test and
the interference score from the
victorious stroop test. So these are
additional executive function measures
and there are some composites scores for
memory and for executive function the
four test version or the six test
version um as well.
So,
one of the things that is challenging is
when you look at these scores,
particularly if you're not used to
thinking about cognitive batteries, is
how do I summarize people's performance?
cuz they've got these four test scores
um in the tracking and six test scores
in the comprehensive and I want to say
okay it's really nice to know that that
one score is impaired but what about
their performance overall across the
battery of tests and this um is trying
in many ways to mimic approximate what
we do clinically. So when we have
multiple test scores, we need to think
about how many scores need to be
impaired before we might say a as a
summary that they're impaired. And to do
this, uh, we use base rate analyses. And
I'll I'll go through those in a second.
and base rate analyses of expected low
scores just using um a very classic
excellent article u by Crawford and it's
based on it's multicaro simulation based
on the intercorrelation between the test
scores so and it's based on the
cognitively healthy the the normative
subsample and in cognitively healthy
people we see low scores they're
actually fairly frequent and they mean
nothing they're called spiriously low
scores and the space rate analysis helps
us predict the frequency of that based
on the intercorrelation between the
cognitive tests that we're giving and we
each clinician may use a different
cutoff for impairment. I use the fifth
percentile we the more lenient you are
in your cutoff for impairment the more
likely you are to have spiritually
impaired scores. It's a well-known
phenomenon not just in clinical
neuroscychology but in clinical
epidemiology as well.
So why did we do this at all? Why did we
think about creating a cognitive
impairment indicator? Um we did it to
mimic kind of what I do clinically when
interpreting cognitive performance on a
neuroscychological battery.
So I look at each test score for sure.
Absolutely.
But to make a determination overall,
particularly when I'm using the
neuroscychological battery as part of my
diagnosis of dementia, which of course
requires additional information. You
cannot diagnose dementia only from the
cognitive performance, but I can say
there's cognitive impairment or not. So
to make that determination
I look at performance after adjusting
with the normative data that I talked
about and then I look for patterns of
performance and in my in my clinic
setting my my neuroscychological battery
set up so that I have some redundancies
or commonalities in what's being
measured across different tests. So if
somebody's impaired in a certain
cognitive domain, I will see it not just
in one test score. I'll see it across a
pattern of test scores that I could
predict ahead of time based on knowledge
of, you know, the nervous system and
what these test scores measure. Um, we
can't do that here. Our battery here is
really brief. We only have four test
scores in one battery and six in
another. When I'm working chronically, I
think my battery has at least 20 test
scores. So, a really high probability of
having low scores that mean nothing or
are called conspiracy low scores. But,
so I also account for that. So, uh I
know when I was trained at UIC how to
interpret neuroscychological test
batteries, I was trained just the basic,
you know, the P equals 0.05. So, 19 out
of 20 times scores will be where they
should be. But one out of 20 you're
going to get a low score that means
nothing. While the coffer did a better
job of that and and did empirical
distributions from the Monte Carl
simulations of how frequently we see low
scores based on how common these tests
um cognitive tests are in terms of what
they measure. So the correlations
between them. So we couldn't do that
here and that's exactly what we did. We
use the base rate analysis to at least
say hm for this particular patient or
participant, excuse me, um how many low
scores are common when you're
cognitively healthy? And that might help
me determine how many low scores are
needed to indicate you're not
cognitively healthy. I'll repeat that
again in a different way.
Um so to create this cognitive
impairment indicator the CI in the
tracking cohort um what we did was we
took the cognitively healthy subsample
from the tracking cohort which we
defined as those without any
neurological conditions
and we said okay how many people in the
tracking cohort who are cognitively
healthy have at least one low one
abnormally low score which was defined
here as a fifth percentile and you see
about, you know, six almost 16% in
French and English speaking subsamples
have one low score. So that's pretty
common. I wouldn't really get too
excited. That's a fairly frequent base
rate. Whereas only about 4%, so 3.7 3.8
um had two or more um low scores. So
that's a little more exciting to me
going okay if you have two impaired
scores on this four test battery in the
tracking cohort I'm a little more
thinking there's might be something
going on in terms of your performance on
the neurosychological battery
in the comprehensive cohort um these
numbers didn't separate that beautifully
so we do provide two different scores um
for the cognitive impairment indicator
so this is the reason for the two
different scores is for the six tests,
there are about 23% of people who have
one low score. So again, fairly common,
right? 23% of a cognitively healthy
population have a low score just due to
chance. Um whereas 5.8% had about uh two
or more low scores and 1 um.4% had three
or more scores. So I did provide two
different estimates and I think they're
labeled in the database accordingly
about how frequently do you think it
matters that these low scores exist. Um
5.8 to me is sufficiently sufficiently
uh infrequent as a base rate. So I I
tend to use that one myself.
So
what we have is the cognitive impairment
indicator is based on these estimates of
how frequently low scores occur just due
to chance.
So what we did then is we said and it
just happened to work out for
comprehensive and tracking that the two
scores one score impaired on the battery
even though the batteries were different
four tests versus six. one score was
pretty common uh but two scores seemed
less common. So happened to work out
that way. So we said okay on the battery
for each participant so at the
participant level you had you got
categorized as having cognitive
impairment if there were two or more of
the test scores that were impaired.
Otherwise, we're classified as not
cognitively impaired, which is a nice
thing because now we can have a nice
summary score overall from these six
test scores in the comprehensive and for
in tracking. The con is you had to have
complete data and there's actually quite
a few particularly in tracking who who
have missing data. So, that is one of
the the challenges of this approach is
it it does rely on the battery being
intact to have the probability in the
battery. Um, but that doesn't stop you
from using the individual items in a
different way for each test score. Um,
here's the paper that's um the one that
describes this. So, if you haven't read
it already, this is essentially what I'm
summarizing. Um and we have here in the
tracking cohort a certain number of at
the individual test level as I said for
each test as a derived variable in the
CLSA you have whether they're impaired
or not for instance are they impaired in
ray one are they impaired in ray two etc
and the same thing for the rest of the
tests and the tracking and for the
comprehensive cohort as well. So this is
the at each test level. But what happens
when we look at the battery level? So
each person's performance on the battery
are they impaired or not. So in the
tracking cohort at baseline we have
about 3.1%
of the sample who could be categorized
as cognitively impaired whereas most of
the sample is categorized as not
cognitively impaired. That is not
surprising given the sampling um the the
methods used for for CLSA. At baseline,
people could not participate in in CLSA
if they had overt cognitive impairment
and therefore would have required proxy
consent. I recognize there are people
who participated who report a diagnosis
of dementia, but again that's not the
the one of the criteria. criterion was
whether they had over cognitive
impairment. So for that reason we have a
a pretty healthy sample in terms of
cognitive impairment
and um similarly in the comprehensive
cohort at baseline about 3.6% would be
categorized as cognitively impaired if
you're only looking at the four test
battery. Why would you want to look at
the four test battery? Well maybe you
want to combine the information from the
comprehensive and tracking cohorts and
that allows you to do that. Um whereas
on the six test battery which is unique
to the comprehensive cohort about 6.1%
of people were calcified. I do point out
again though that there's quite a bit of
missing data from these um which is
definitely a challenge for this
composite that requires a battery level
analysis which requires all scores on
the battery to be present.
a question though when we did this which
is great um this is how
neurosychological tests in some ways are
done except as I said we didn't have
that pattern analysis
um which is really a big part of the
interpretation of a neurosych battery um
so the question is did this mean
anything um was this useful at all so
that was it's a really important point
here because it doesn't completely mimic
what's being done in konopo practice so
the challenge is we didn't have a gold
standard reference but we only had is
self-reported chronic conditions which
is a has a doctor ever told you that you
have and there was a list of chronic
conditions
um we associated the cognitive
impairment indicator with each of the 30
plus conditions and we summar but we
also summarized them so if you want to
see each of the conditions I'm going to
refer you to the paper I'm not going to
repeat that here but we did summarize
them we summarize them in three ways one
is neurological conditions And here's a
list of the potential neurological
conditions. Um, and then we summarize
them by conditions that we know to be
risk factors for neurological disease
but are not considered neurological
conditions such as diabetes,
hypertension. And you see the list here.
And then as for evidence of divergent
validity, we had people who a lot of
people reported chronic conditions that
were not a risk factor for neurological
disease or were um other things like um
osteoarthritis or allergies etc. So that
was they were categorized as having not
neurological conditions.
And here is the evidence for validity
being on the odds ratios with the with
the confidence intervals. So the
neurological
the neurological group um are way more
likely to have cognitive impairment
whether it be in tracking or
comprehensive. Here it's the four tests
and comprehensive and here it's the six
tests and comprehensive. So but that's
not the greatest evidence for validity.
Why? Because while we the normative
subsample excluded people with
neurological conditions. So it's kind of
circular reasoning really. So but
reassuring I guess but not the strongest
evidence uh in my mind really ideally we
would have had a diagnosis of something
else uh which is coming but the risk of
neurological conditions also increas uh
showing increased risk or likelihood of
having cognitive impairment was very
reassuring and then the fact that if you
didn't have neurological conditions you
were not more likely to have cognitive
impairment was kind of reassuring. So
the pattern overall is reassuring in
terms of evidence for validity. This
approach seems to be capturing something
useful which is kind of what we had
hoped it would do. Um so therefore this
is the evidence for validity. I will say
though we do not have the cognitive
impairment indicator is not a dementia
indicator. The cognitive impairment
indicator only tells us the performance
on the neuroscychological battery and
the evidence for validity is really
needs to be evolved and let's be honest
it really should have a gold standard
reference like a diagnosis of dementia
or a clinical diagnosis of cognitive
impairment then we can that would have
been if we had a magic wand at baseline
that's what we would want. Um but that's
going to be a bit in our future work.
So we did um say this is useful. So
let's keep doing this approach to
summarize the cognition battery. Um we
computed a delta CI or um change in
cognitive impairment indicator for the
first follow-up and it was similar to
the cross-sectional cognitive impairment
indicator at baseline but it differed in
three main ways. One, it used new
follow-up norms for cross-sectional
impairment for each test score.
two, it used reliable change indices
which adjust for eron measurement and
expected practice effects and then it
analyzed patterns of consistently low
scores. So the whole premise underlying
base rate analyses when looking at a
neurosychological battery is that if
there's a low score and it frequently
occurs in the healthy sample, it's a
spiriously low score. It's kind of an
error. But if you're having a low score
consistently,
now I'm paying attention to that. Now
I'm not thinking that scenario. And
that's a little bit of the pattern
analysis I do as a clinician. I actually
do all of these three steps as a
clinician when I I see all of my
patients at a a one-year follow-up at
least. And I do exactly this type of
approach. Um I do cross-sectional
analysis of the new scores. I do
reliable change indices to see if
there's been change. And then I look at
patterns of performance. So at least
here we have some pattern of performance
analysis and um with the premise that if
I said you had a spiriously low score,
it really shouldn't be consistently low.
It should kind of bounce around like
error, which is what we were assuming it
was due to. Um so here we have um a
little bit about the delta CI. We have
people who were not um considered
impaired at baseline. Um and then they
look don't look impaired in any way at
followup and we call them not impaired.
Um we do have a group of people who we
still would say not impaired at
baseline, not impaired at follow-up. But
this is that group where that pattern of
analysis I was talking about where if
you had kind of consistently low score
on the same test each time, I'm like,
hm, there's something going on here. I'm
not going to say you're cognitively
impaired, but I'm a little more
concerned about something that that may
not be a spiriously low score. Um, and
then we also used reliable change
indices um with you change on two or
more tests. And again this is
capitalizing on the fact that we had the
same tests given and remember reliable
changes these address for practice
effects which is a really important
thing. Practice effects persist even if
we had alternate measures of the
cognitive u batteries like say alternate
words for the ray um they still seem to
translate across alternate measures. So
um you know nothing really seems to get
rid of these practice effects and they
do persist for a longer time than we
ever thought. Um but reliable change
indices are a nice way to kind of vary
out the the variance due to that
practice effect and error measurement
because now you got eron measurement at
each time point. So um we use that
together and we use that to be create
three different groups which is not
impaired at risk or impaired. So if you
had evidence of impairment and you can
have new evidence of impairment, you
weren't impaired at baseline but you
were at followup, then you're
categorized as impaired on the Delta CI.
We do have uh the group that that um
scares me, which is people who are
cognitively impaired at baseline and not
necessarily at follow-up. Uh we do have
that group. Some of them looked at risk,
so that's that's fine. But we do have
some who really don't look like uh they
should have been classified as cognitive
impaired and evidence of error in our
approach with the delta CI b I mean the
CI at baseline and that didn't happen
very often. So they were likely
classified incorrectly and didn't happen
very often. And then of course we have
people who are impaired at both time
points. That's the delta CI. The problem
with the delta CI is we don't know what
it means yet. We haven't done the u
associations with the um chronic
conditions yet. Um so we don't know what
it means and it's a fairly complicated
process. So we have the baseline scores,
we have the follow-up scores, we have to
redo the norming to get the
cross-sectional analysis, but then we
also have to have the comparison to the
baseline using reliable change indices
and then we do the battle battery level
um comparison. So it's it's an involved
process and again just because we do it
it doesn't mean the scores mean
anything. We before it could be used we
really need to do the next step. So you
have to wait for the publication. I'm
sorry. Um we don't know what the scores
mean. So we really shouldn't be using
the delta CII just yet.
What about having a cognitive impairment
indicator or a delta CI for future waves
is probably the big question you're
asking. Um because it seems like the CI
um cognitive impairment indicator,
excuse me, is helpful a helpful summary.
It seems to be people are interested in
and seem to want to use it. So the you
know pushes us to think about doing this
for future uh waves. we can't just use
that baseline method that I went through
and for the for um doing this at
follow-up waves because we have these
practice effects um so it shifts the
whole distribution and it doesn't shift
it completely linearly um because the
people who are it's in in practice
effects the pe the people with intact
memory for instance tend to benefit more
from practice so it's not just a a
constant that we can add to everybody's
score. So, it's it's a little more
complicated.
The other problem um with with looking
at creating the delta CI for future
waves is we had a change in the mode of
delivery of the test scores. So, a small
proportion of people in comprehensive
um did their follow-up testing by
telephone and that was a retention
technique. But then at follow-up 2,
partway through follow-up 2, the
pandemic happened. So about half of
comprehensive did their cognitive
testing
um on the telephone due to the pandemic.
So that creates a big challenge. Um and
a paper that you know we're we're really
I need to finish doing before I have to
take on more admin again in my job um is
uh showing that despite using
essentially harmon har har har har har
har har har har har har har har har har
har har har har harmonization technique
is stratified norming could be
considered um the telephone and inperson
cognitive batteries the tests differ
fundamentally memory and executive
function are measured differently when
they're given even though it's the same
test given the same way mostly the same
way just over the phone versus in
person, there really are some
differences. So, you can't just easily
collapse at least the cognitive test
scores across when they're given versus
by telephone versus in person. Um and
more importantly um when when paying
attention to and I I we haven't really
done this carefully uh yet um is paying
attention to the mode of delivery and
using telephone norms when it was
telephone delivery and in-person norms
when it was in-person delivery. That's
really important because right now we've
the baseline we separated it by cohort
because those two were synonymous.
comprehensive cohort was in person,
tracking cohort was by telephone, but in
subsequent waves that's no longer the
case. We have inperson versus telephone.
So we have to use the appropriate
normative comparison standard.
And then the other big challenge is that
missing data is an increasing issue. So
I talked a little bit about missing data
particularly for the cognitive
impairment indicator because it requires
the whole battery to get you a score.
Missing data is an increasing issue. Um
and it's an issue that um isn't just
going to impact the cognitive impairment
indicator also impacts the the normative
comparison standards. Um so we know that
as we go along in the
um as we sorry as we go along in CLSA we
are getting a lot more missing data and
we're getting a sample that is
increasingly healthy. So I think I do
plan to talk about that in the next
slides. Let me get to the next slide.
So, what we're working on right now is
as part of the memory study and work led
by uh Lauren Griffith at all, we're
we're working on a cognitive impairment
indicator for follow-up 3 and we're
starting with follow-up 3 because we
have the memory study diagnosis as that
reference standard. I said that the
evidence for validity that we had for
the cognitive impairment indicator was,
you know, nice, but it was self-reported
chronic conditions. it's not really a
gold standard reference. Whereas with
the memory substudy, um there is a group
of us and I saw David Hogan's name on
there. So, call out to our memory study
colleagues and we we did clinical
diagnosis on 600 cases. So, lots of lots
of work. Um so, that's going to be
released. I don't know when. You can't
ask me that. I don't know that answer.
Um but we're starting with that because
um if all else fails in developing a
follow-up cognitive impairment indicator
because of this loss of um of people who
have more variability in their cognitive
performance or people who are more
cognitively impaired not staying in the
CLSA or not doing all the cognitive
tests for instance um might really mess
up our normative comparison standards.
it might mess up uh how we determine
cognitive impairment and this is a
well-known phenomenon in psychology in
neuroscychology that if the normative
comparison standards are not reflective
of the population what you get from that
clinically is really problematic and uh
for those of for those who are
neurosychologists just remember the
CVLT2 norms right so it can really cause
some challenges so in the very least the
memory study we have the gold standard
reference Now, if it looks like our
approach works, um, we're going to have
to figure out some ways to address for
this bias, this loss to attrition, and
if we figure that out and it works, then
we'll move on to doing the cognitive
impairment indicator, the delta CI for
followup, too. But we have not done that
yet, or at least we haven't done it and
checked it and done it carefully. As I
said, we have worries about this. Um,
and that's why we're starting with
follow-up three because it's the closest
follow-up to when the memory study
diagnoses were done. We have this
increasingly select healthy subsample.
We have complete data. Uh, this impacts
everything about what we want to do. Um,
and it's going to definitely bias the
delta CI. So, it's going to be a
challenge. Um and this is actually the
reason why um we decided not to create
reliable change indices at each wave.
What we are doing is we're using the
reliable change indices that were
created from baseline to first followup
and apply them to subsequent waves
because that is essentially what we do
clinically in neuroscychology is we have
reliable change indices and if I see
people at their four or five or six year
follow-up I'm still using the rcis were
created. It's not perfect but it's it's
yet it's an additional piece of uh data
that I can use clinically. So it's it's
not perfect. I recognize but neither is
having this increasingly healthy
subsample going on which really biases
everything. Um so we really to do this
well to have even to have normative
comparison standards that are reflective
of the population we have to adjust for
this bias this loss to attrition in
order to be able to do this well. So
that's um what's something we we have to
work on and again because it might not
work out. We're starting with the memory
study and you'll have to just wait for
that to come out and um us to finish
that work. And I think that was all I
needed to say about the cognitive
impairment indicator.
Great. Thanks so much. Um that was an
excellent presentation. Uh so now we'd
like to open it up for questions. Uh
just a reminder to everybody that uh
muting will remain on but you can enter
your questions in the Q&A box in the
bottom of the Zoom window. So um not in
the chat but in the separate Q&A box.
Now I saw that we had a few questions
come in already. Um so we'll get started
with the first one from Victor. Will you
be looking at the CI in relation to
plasma biomarkers such as P217 tow NFL
NFL and GFAP GFAP in the future? Not me,
but I'm sure once we create it, someone
will.
So that's the beauty about creating
this. Once we create these, these become
part of the data platform. So we share
it with you all. So whoever is
interested in such things can do such
announcements.
>> Great. Thank you. Um I think I saw
another one come through. Um so yeah
just a reminder to everybody please do
not post your questions in the chat box.
Post them instead in the Q&A box.
All right. So we have one another one
from Roberta. Excellent presentation
regarding aspects of cognitive testing,
practice effects and mode of
administration and cognitive performance
including sperious poor test scores and
its relationship with a diagnosis of
cognitive impairment as well as how that
is not indicating dementia per se. Since
in an aging aging cohort, hearing and
vision changes account for large
variance in cognitive performance, could
some of these spurious findings in one
test reflect a combination of vision and
hearing changes that are progressive and
interact with cognitive processes?
>> Yeah, so that's a great question. We did
look at whether the CI is associated
with the hearing loss um self-reported
hearing loss and self-reported vision
loss and it is associated. It's not a
massive odds ratio. Uh it's in the
original paper. One of the regrets I
have about that original paper, and I
will fix this regret on when I do the
validity for the delta CI is I regret
that I put the hearing and vision loss
group into that not neurological group.
They should have been in the risk of
neurological group. So I really should
not have, you know, I bet my odds ratios
would look even better if I separated
them that way. But yes, we definitely
looked at hearing and vision. I don't
the odds ratios were statistically
significant. So they uh you know, but I
don't actually recall what they were.
They're in the paper though.
>> So great question because it was
something that in retrospect I will fix.
>> Great. Thank you. Um from Jalal, I will
use the cognitive impairment indicator
based on four tests from follow-up 2 as
the outcome in a mediation model. Due to
mixed mode of delivery in follow-up 2,
both telephone and inerson tests,
>> what limitations should I be aware of?
>> I think you should stratify your
analyses by mode of delivery.
So, um I think what you should do is um
do because it's about 50% I think of of
the cohort that got like a comprehensive
cohort who got it in person. So I would
do my analysis for the in person because
the baseline was in person as well. Um I
would do my analysis for that and then I
would check it with the other one. But
knowing that you have this mode of
delivery change and you know it hasn't
been adjusted. Now,
if you're using raw scores, remember you
have these biases anyway. You got these
problems anyway. So, you know, so just
to remind you about that, you already
have problems and biases in your score.
So, now you're just adding another
source of of error, which is mode of
delivery.
>> Thanks. Uh, this question is from
Andrew. Thanks, Professor O'Connell for
a great presentation. In genetic
studies, we are typically concerned
about collider bias by coaring or
adjusting for heritable coariance. Is
there such a concern when cognitive
measures are adjusted for educational
attainment?
Not sure I fully understand the collider
bias issue, but I will say um there are
some concerns about adjusting to these
demographic variables. Uh it's almost
like I planted this question because
this was my dissertation topic. Um so I
actually and I did simulation studies.
Um so
uh what happens when we address for
these demographic variables is it
reduces our sensitivity but it increases
our specificity. Um, so when we're doing
a battery approach analysis, you have
multiple test scores by virtue of what I
was talking about, your sensitivity is
increased and your specificity is you
know decreased anyway. So at a battery
level, these demographic corrections
work. They help. But uh at an individual
test score level, like say a screening
test, demographic corrections may
actually reduce your ability to detect
the condition of interest because the
condition of interest has the this risk
factor. Age being a risk factor, first
aid dementia or education, low education
being a risk factor. So in my simulation
study, I showed that the higher the
association between the risk factor and
the outcome of interest, the more likely
the correction reduced your sensitivity.
So
yeah, but I do not understand the
genetic part. So
thanks. Uh, next question from Hannie.
Uh, thank you for the presentation. In
the context of classifying patients
according to their overall cognitive
performance derived from the psycho
neurossychological battery, which metric
would you recommend as the most
appropriate for our studies that we
actually doing? Would the reliable
change indicator be suitable for this
purpose?
>> Uh we yeah we haven't published the
reliable change paper yet. That's also
on my list of things to do. Um so um the
reliable change
is has similar challenges with base
rates. So again if we're looking at a
battery approach so there's um the need
to adjust for these you know change
scores that have no meaning. So it's a
very similar approach and unfortunately
because we haven't shared that yet with
you I would say maybe use the scores
that we've published on
Thanks.
>> Publication coming soon though. I've
been saying that for over a year though.
So
>> that's not so bad in this world.
Um from Magnolia, thank you for your
work. I'm wondering if it would be
possible to track functional impairment
in similar ways to enable probabilistic
identification of dementia within
participants.
Yeah, I think um I think this
understanding of low spiritually low
scores makes a lot of sense not just in
cognition. Absolutely. I think it could
make a good sense there. Uh the big
challenge to do that though is you have
to have a group for whom you don't
expect there to be any functional
impairment and functional impairment so
confounded by not just cognition
potentially in the case of dementia but
by physical health concerns. So I think
you'd have to really create this
healthy group and then look at the low
the base rates of low scores and these
functional impair impairment. But yeah,
I think somebody could do a very similar
you know Alexander May Hugh who I saw
has done something similar for physical
therapy variable. So I think that you
know a similar uh you know approach but
the conceptually what you need to do is
is be able to have a group for whom you
expect to be the that that typifies the
healthy. So then you know the
probability of low scores in the healthy
group,
>> right? Thank you.
All right. So we have a few more
minutes. Um so if there's any last
questions out there, please get them in
to the Q&A box. Um but for now, I have a
comment from Lisa Kosski in the memory
study group who had clinical assessment.
The assessment included a rough I think
it's rough. Sorry, there's a typo.
Assessment of hearing. So, the impact of
hearing on CI and diagnosis can be
examined formally in this cohort.
>> Yeah. Yeah. So, the memory study does
allow us a lot more work to be done. Uh
it's uh I'm very excited to um get to
developing
cognitive impairment indicator with
memory study data because that's that's
the essentially if it had a magic wand
that's the study I would have designed
for that. So yes yes definitely we can
do all of that.
>> Great.
Um, from Magnolia, would there be a way
to compare the CI results with health
administrative data to identify under
diagnosis, especially across
sociodemographic risk factors like
language or education, for example?
>> Yeah. Yeah, I think so. For sure.
Absolutely. Um, of course, you'd have to
have the subsamples of linked data. Uh,
it would also be really interesting to
compare it to the administrative data
cognitive uh indices, right? They've got
the CP um oh god what's what's it
called? Cognitive score which um is
funny because I don't actually think
it's a cognitive score. The CPS is not
really a cognitive score. And another
paper that I have not written but it's
fully written. I haven't published yet.
um is uh I did go through the cuz the
first time I looked at this I thought
that's a real confounding of function
and cognition and and you know they're
not always like they're two different
constructs and why are we calling this
the cognitive performance scale um so I
created one that was a purely cognition
one um and I do have to publish that
it's it does a better job of identifying
those who in the database not a huge
better job like 5% better but you know
in administrative data 5% is very
substantial. Um, it does a better job of
identifying people who have a diagnosis
in the administrative data. Um, so I did
those analyses a while ago and and I
keep talking about this paper and I need
to get it out and then then it would be
nice to then it would be nice to like
look at that cognitive impairment
indicator with this cognitive impairment
indicator in a subsample. That would be
really lovely. That would be another
dream study that of mine. But I kind of
do need to get that submitted somewhere.
>> Cool. Great. All right. Uh, last
question here from David Hogan. With
function, we have self-report but not
performance-based functional measures.
What are the implications of this?
>> You can't. This is why you can't use
anything in CLSA until we get to the
memory study about diagnosing dementia
because people who have uh cognitive
impairment are not the most reliable
informants about their day-to-day
function. We know that I mean they it's
not that it's there, you know, they're
not completely but we don't trust it. We
need an informant to tell us or we need
a performance measure of it. We need
something objective. It cannot be a
subjective report of their function. So
to make a diagnosis and that's of course
what the memory study was able to do got
objective or at least informant reports.
>> Thank you for helping us to cement that
point in CLSA until the memory study
comes. There is nothing that is a
dementia indicator in CLSA and that is
the reason why. So it doesn't matter if
you have cognitive impairment. We do not
have objective or uh reports of function
and therefore I cannot call it a
dementia indicator.
>> Thank you Megan. That is a great point
to clarify. Um okay, very last question
from Roberta. Can you comment about the
Mocha's online version, the Mocha online
version's efficacy for other studies?
>> Uh no I cannot. um the psychometric
property when you develop when you
change anything about a
neuroscychological test and the mocha
being a screening test so it has the
same issues it changes everything
fundamentally about the test so you need
to redo either the normative data which
for screening tests they don't tend to
do you need to do redo the validity of
it you need to start from scratch it
needs to have new psychometric
properties established and I don't know
if that's happened I just don't tend to
keep up with what's going on in the in
the I don't use the mocha myself. Sorry.
Um I so I it's not something I tend to
keep up with. So but I will say you know
psychometric properties need to be
reestablished. Evidence for sensitivity
and specificity need to be
reestablished. You cannot use the
in-person paper versions sensitivity and
specificity to interpret the online
versions. That is not should not be
done.
Well, thank you very much, Megan. Um, we
appreciate your participation in our
webinar series and and the insight that
you've shared with all of us here. Um,
yeah, that was really great.
>> Thank you.
>> Um, I'd also like to remind everyone
that the next deadline for data access
applications is April 8th, 2026. Please
visit the data access section of our
website to review the data that are
available as well as any additional
details about the process. Um friendly
reminder, if you've um made any changes
that affect your data access agreements,
please let us know. Whether that's
moving institution, changes to your
research team, training who's graduated
or moved on, please reach out to us at
access clsa
uh- elcv.ca with any questions or
updates.
And lastly, um please commute that
anonymous survey upon exiting the Zoom
session. Now, um we do have a couple
more updates. Um
you if you've been to our session or
follow our emails, we've introduced an
updated data access fee structure which
came into effect on October 2nd, 2025.
Um and this update supports a
sustainable cost recovery model that
will help us maintain the exceptional
quality and stewardship of the CLC data
for years to come. There's more
information about that on our website at
the data access um section. or there's
also um Jacqueline's posted a link a
link in the chat.
Um so again, thank you for attending and
participating in today's webinar. The
next webinar will be uh the association
between menopause age and estradile
based hormone therapy with cognitive
performance in cognitively normal women
in the CLSA. This will be on Thursday,
March 26th at 1 p.m. Eastern. presented
by Laura Gravelson's post-doal
researcher at the center for addiction
and mental health and Lisa Gala senior
scientist at the center for addiction
and mental health as well as a professor
in the department of psychiatry at the
University of Toronto. You can find
registration details on our website on
the webinars page and uh the link is in
the chat box. If you or a colleague is
interested in presenting a CLC webinar,
please reach out to the webinar team. Uh
there's also a link for that in the chat
box. The recording of today's webinar
and the slides will be available in the
coming days on our website. And last
reminder to please fill out the exit
survey that will appear at the end of
this webinar. Thank you again for your
attendance. Have a great day.