Calling evaluative reasoning! Come out, come out wherever you are!
Watch on YouTubeVideo summary
In a seminar hosted by the New South Wales Committee of the Australian Evaluation Society, Associate Professor Amy Gullixson presented critical research on the state of evaluative reasoning within the Australian education sector between 2014 and 2024. Her analysis of thirty-seven publicly available reports revealed a significant gap between the profession's emphasis on describing measurements and causality versus the actual practice of making credible judgments. While all analyzed reports included measures, only four demonstrated full evaluative reasoning by completing logical steps and providing justifications for every choice; common deficiencies included missing research questions, performance standards, synthesis methods, and overall judgments. Gullixson argues that this imbalance is akin to skipping leg day at the gym, where the foundational work of judging is neglected in favor of superficial description, ultimately weakening the connection between findings and decision-making.
To address these challenges, the discussion shifted to practical strategies for managing client resistance and defining standards in complex environments. Speakers suggested using performance rubrics that clearly distinguish between current and improved states to provide actionable steps without forcing a formal judgment when stakeholders prefer direct instructions. This approach helps make conflicting values visible, allowing evaluators to explain why certain actions align with funder expectations while potentially harming local communities. Furthermore, rather than reinventing criteria from scratch, the group advocated for sharing implicit standards from existing reports and utilizing tools like Mattea Roa's values identification matrix to surface stakeholder values without being reductionist, ensuring that boundaries are established to define success even when formal standards do not yet exist.
The session also highlighted the importance of simplifying frameworks to focus on what truly matters, such as helping programs discontinue ineffective activities rather than getting bogged down in minor distinctions. By starting with broad descriptors like "excellence," evaluators can navigate varying client appetites for explicit reasoning while acting as translators between technical findings and decision-makers' needs. The overarching conclusion emphasized that overcoming the difficulty of implementing evaluative reasoning requires making these practices viable through shared resources and streamlined methods, ensuring that the profession can effectively balance organizational culture, risk appetite, and stakeholder expectations to produce meaningful, well-reasoned evaluations.
Read the full video transcript
So, I'll start off by saying welcome.
Welcome to this um uh this seminar
organized by the New South Wales
Committee of the Austral the Australian
Evaluation Society. My name is Philip
Belling. I'm an evaluator at the Center
for Education Statistics and Evaluation
in the New South Wales Department of
Education.
Um and it really is great to see some
familiar faces here, but it's also great
to see a lot of participants from across
the profession.
We um started out calling this session
um a provocation and a dialogue. So it's
really good to be able to provoke and
discuss with such a diverse group of
evaluators.
Today's provocation is going to be uh
from associate professor Amy Gullixson.
Um but before I hand over to Dr. Gixson
and also noting that we are in national
reconciliation week. I'd like to start
by acknowledging the Gatagle people of
the Aora nation and their elders past
and present. We're all on Aboriginal
land and I'm fortunate to be on Gatagle
land today. Gatagal elders have been
custodians of songlines and knowledge
and insights over countless generations
and I offer my respect and gratitude and
I also pay my respects to the elders and
traditional custodians of all the lands
from which you're joining uh this
seminar today and to Aboriginal
toouristr islander and first nations
colleagues who are with us as well.
Um, at this point, uh, if there had been
someone from the, um, New South Wales
committee, they'd have leapt up to give
you information about the AES. Um, I'm
going to try to do that. Um, it is an
explicit aim of the AES to strengthen
and promote evaluation practice, theory,
and use. Um, so that evaluation makes a
difference, and that that's a lot of our
focus today. But it the AES organizes
these free monthly seminars but a whole
host of other activities. Um and I'd
like to encourage any non-members who
are here today to think about joining
the society. It's actually a really good
deal um because you get up to 48% off
any of the other workshops that are done
and you get 20% off the annual
conference registration. So, I highly
recommend that you um check out the uh
conference this year in beautiful
Canberra um and check out the program
that's online especially for
pre-conference workshops which are being
delivered by some outstanding evaluation
leaders. Oh, including today's speaker.
So, more on that later. Um my role here
is to introduce the topic and the
speaker and also to let you know how
we'd like to handle questions and
breakout rooms. So as to the topic um
for me it comes down to this that the
AES is championing quality evaluation
that makes a difference and today's
session is all about documenting
evaluative reasoning. So throughout the
session, we're going to think about how
does documenting evaluative reasoning
affect evaluation quality and what sort
of difference that documenting
evaluation reasoning actually makes. So
I'm really looking forward to hearing
from associate professor Amy Gullixson
about recent research about what we
actually see in terms of evaluative
reasoning when we look at published
evaluation reports.
But as I said before, time is really
tight today. So to make the the most of
that discussion time, I want to invite
you to make use of the chat while Amy's
talking. Um, I want you to use it
proflegately and generously and
extravagantly. Use the chat to ask
questions like, "Amy, what did you mean
by or can you say a little bit more
about that concept?" Um, but also use
the chat for ideas, opinions,
speculations, provocations of your own
if you like. um and we'll pick those up
um in the breakouts and in the final
discussion.
Now, finally, it's my um great honor and
privilege to introduce um our speaker
today, Amy Gexon. Um Amy is associate
professor at the University of Melbourne
where she's played a key role in
establishing um and bringing about the
success of the Master of Evaluation. She
is a highly regarded scholar and she's
also a fellow of the AES. She's in
demand as a speaker both nationally and
internationally and she's had a
significant effect on the quality of
evaluation in Australia and abroad
through her teaching, writing and her
inspiring speaking.
Personally, I can attest that her
keynote at the AES conference in
Adelaide has made a lasting impact on me
and other members of the team that I
then belong to and it continues to help
evaluators and evaluation teams to lift
the quality of the work they're doing.
So with that said, I will now pass over
to Amy.
Can you see my slides? Yes. Good. Hi
everybody. It is a delight to see so
many familiar faces and also uh just to
see new people and to have this
opportunity to chat with you. My name is
Amy Gellixen uh and I'm hugely grateful
to the group of people who organized
this session and uh wanted to hear more
about the evaluative reasoning research
and to those of you that joined um the
person who is working with me my partner
in this research is Dr. Katherine
Meldram. Um, and on this slide you'll
see the thesis that she wrote with us at
the University of Melbourne. Uh, the
handle thing is a link and so that'll
come to you in the slides and I think
Philer Julia will put it in the chat
too. So if you want to read the full
thesis, you can do that. Um, she
Katherine has more than 30 years
experience doing research and teaching
across Australia, Asia and the UK. and
she came to study uh at Melour because
she believed that well-conducted
evaluations contribute to evidence bases
which in turn support opportunities so
that education can be empowering for
all. Um she had a PhD when she came to
study with us and she was going to do
the ML but exited halfway with the
gradert and then did this research on
evaluative reasoning and so what I'm
going to talk about today started with
her thesis which is what the link is to
um but we've continued to develop it
since. So, the findings that I'm going
to share are are from the unpublished
paper that we're working on right now.
Uh, she's not here today because she's
hiking in Scotland for heaven's sakes.
She's having a miserable time. So, uh,
you get just me instead of both of us.
Let me uh start out by talking about
evaluation, which is the space we're
working in. Evaluation investigates
investments of time, energy, uh,
financial, and other resources to see if
the thing that they're getting invested
in is actually worth it. The process of
evaluation includes fully describing and
fully judging whatever it is we're
evaluating. We could debate later the
fully word, but describing and judging
is the thing that we do together.
Evaluative reasoning is attached to that
fully judge part. It uses the fully
describe, but it's a separate set of
steps. It provides credible, valid, and
defensible judgments to answer that
question of worth. and that worth
questions often broken into dimensions
of quality, cost and significance.
Deciding what to keep doing, what to
modify and what to quit relies on this
kind of evaluative judgment.
So evaluative reasoning starts with the
logic of evaluation. On that this slide,
it goes from the bottom to the top. So
first you identify criteria, then
performance standards. uh you figure out
what you're going to measure to
understand evaluate the performance of
the evaluand and then you synthesize all
that together to understand how good
something is.
Uh when we looked at the literature we
found two additional steps. One was to
that almost always it starts with
determining key evaluation questions and
those connect with the criteria. Um but
uh when we were reading through the
literature uh Katherine in particular
was demanding that evaluation questions
be included because inquiry starts with
a question and so that became our first
step and then uh we ended with judgment
and it in terms of making a statement
about the goodness of something or the
goodness of the parts of the something
which I'll say more about in a little
bit because findings need to be stated
in order to be shared and findings that
don't get shared can't be useful. And if
we're going to be doing quality
evaluation makes a difference, then it
has to be out in the world uh to make
those changes.
The warrants, this is to get to a
justified uh judgment, you need warrants
and those are the becausees that support
your choices. So uh this orange box is
the justifications of criteria,
standards, uh measures and synthesis
methods. So all those happen throughout
the process. Um, warrants can come from
lots of places. So, they can be from the
literature. Quite often, we might use
theory to warrant if you're doing a
program theory. That's one way to do it.
You look at the literature and discover
theories that are related to the program
you're evaluating. They can be cultural,
methodological, uh, experts, and
authority. And that might include laws
or other things that we need to attend
to in the doing of our evaluation. Uh, I
have put all of that into one orange box
just because orange is bright and I
thought I'd save your eyeballs. But
really what happens here is that that
warranting thing needs to be done. That
warranting step needs to be done every
time you're making a decision. It's
important to give the because these key
eval question evaluation questions are
important for this evalu because of
these reasons. So uh that's how it all
works together. So the logic of
evaluation makes it legitimate because
you're doing the steps of the logic and
the warranting makes it justified
because you're providing re reasons um
for the choices that you made and
altogether that adds up to the elements
of evaluative reasoning.
Uh on the right hand side are the
sources we used to get to this list.
Uh we went to look to see who else had
done this research before we did our our
study. And despite the fundamental
importance of evaluative reasoning to
building these credible, valid and
defensible arguments about how good
something is, we could only find four
previous studies and those are listed on
the right hand side here. Um, three of
them looked at reports in various
contexts. One of them surveyed uh
evaluators and asked them to just
self-report on their knowledge and
practice. So the whole sort of sample to
date internationally in when when we
started this research included reviews
of 76 reports in total and then 214 uh
responses. So we don't even know if
those were all separate members or if
it's people taking a survey twice from
one of the 21 uh global professional
evaluation associations. So it's a
pretty small group of stuff that's been
uh researched. And
so on this slide, we've got the findings
from the three quantitative studies
matched to these elements of the of
evaluative reasoning. I'm just going to
give you a minute to look at that.
Uh overall, Erto and Ozaki had the
closest findings to each other. Uh Jane
Davidson did a qualitative review of a
set of six reports uh for an
organization and she generally found
that key evaluation questions were
missing uh along with not very much
synthesis and evaluative conclusions. So
clearly more empirical research is
needed to understand evaluative
reasoning and practice and that's what
led to our research question. What is
the evaluative reasoning practice in the
Australian education sector between 2014
and 2024?
This is the conceptual framework we used
to identify the presence of elements
that contribute to legitimate which is
the purple boxes the logic of evaluation
and justified orange the warrants uh
evaluative conclusions. So you can see
that's all laid out. We in the first one
we added evaluation objectives because
we found when we did an initial review
of reports that those objectives and
questions were sort of interchangeable.
So we wanted to make sure we captured
that.
We also extended the synthesis step uh
to include microynthesis and
meosynthesis from Jane Davidson's 2014
paper. Um this was to counter uh one of
the long-standing and mistaken arguments
in evaluation that an inquiry process
can only be called evaluation if it
arrives at an overall judgment about
goodness. Um Jane's argument was that
actually you can make evaluative choices
at different levels within an
evaluation. So you might make choices
about weighting one set of data more
heavily than another. uh you might go at
the MEZO level be looking at look
judging criteria separately from each
other or judging components of a program
separately from each other and you might
not want to go to that overall level of
judgment but it's still it doesn't mean
you're not doing evaluation it's just
that the levels uh that where it's
happening are different.
So uh this slide has the definitions we
used. So we use deductive logic and
adaptive and we had an adaptive
systematic quantitative analysis
process. Uh I'll let you read the
definitions because that's what's going
to inform what I'm going to share next
which is the findings on a few slides.
I'm going to go next. Are we ready? Yep.
Thanks for the nod, Phil. So, um, we set
the inclusion criteria to get as many
reports as we could find in English that
were, uh, full whole reports. And what
between 2014 and 2024, so in that 10ear
span, we found 37. And this slide shows
you where they were from. Most of the
reports gathered data at the national
level. Uh you may note that if you're a
person who does the math when looking at
a slide that this adds up to 38 and
that's because one of the reports got
data from across three states but not
nationally.
So the there's about to be a series of
finding slides. So uh when I use this
table, the top row has the elements and
they're colorcoded. And then the second
row shows present or not present. And
then there's a little extra one for
research questions. Uh and then the
different finding stuff is going to
appear in the third row of the table. So
first off, four only four of the whole
of the 37 reports that we found ticked
all the boxes for legitimate and
justified evaluation. So that's that is
they did all the steps in the logic of
evaluation and they provided warrants
for their choices across all those
steps. So four of 37.
Uh now I've got overall findings in that
third row. So the counts represent each
of the elements across all of the
reports. So uh across the whole data set
we found that the elements directly
related to research that sort of the
fully described space were much more
consistently present. So everybody
measured something. All 37 reports had
measures in them. uh 20 of the reports
had research questions and uh and only
five of of the whole set had no
questions or objectives at all. So this
the like what's happening set of
questions and how do we know that's was
all really strongly present in this data
set.
The fully judge aspects of evaluative
reasoning were the least present. So 10
reports had no evaluation questions and
some and like I said five had no
research questions. 27 didn't have any
performance standards. 29 didn't have
any of that sort of lower level
synthesis and 32 didn't have any judge
of that overall judgment.
This makes sense because if you're going
to synthesize to a judgment you're
building up across the whole uh
argument. When I teach it I often talk
about this as an unbroken chain of
reasoning. So if you've got uh links in
the chain that are missing, then your
conclusion isn't going to be uh credible
at the end. So not surprising that
missing some of that stuff would lead to
not making judgments. Um one thing to
highlight, if you've noticed that at the
end in the judgment column, it says
there were five reports that had a
judgment present. And if you think back
to earlier, I said there only four of 37
that ticked all the boxes. um this one
report reported a judgment but they
didn't have criteria or comparators. So
they one of the links in the chain was
missing and so that's why it didn't get
counted amongst the four
16 reports had no warrants for any part
of their evaluation. So that means they
didn't explicitly justify any of their
choices.
Uh and finally, even though we didn't
formally code this, only a few reports
in that data set set uh had identified
limitations, which both Katherine and I
were surprised about because uh context
and limitations are critical for
understanding and interpreting both
research and evaluation findings. So
that was an accidental finding that we
thought was important.
Uh there's of course limitations to the
study. So first we looked at just
publicly available and searchable
evaluation reports. Uh we did this
because we decided the stakeholders were
our primary
uh t taxpayers were our primary
stakeholders in this research and so we
wanted to look at stuff that they could
find if they were looking. Uh but we
know for sure that there are way more
evaluations of education that have
happened in the public sector in
Australia in that 10-year period. So
we're not in any way trying to claim
that this is a state of evaluation
kind of findings or claims at at that
level. We're just saying let's let's
talk about what this means and have it
as a provocation for conversation.
Second, uh this was a desk based review
and we just counted the presence of
elements if they were explicitly stated.
So in some of these reports for instance
constructs could have been in there that
might have been able to be treated as
criteria but we didn't do that sort of
extrapolation.
And finally excuse me we didn't do any
judgments of quality. We just said was
it there or was it not there. Um that's
the next round of work is to start
thinking about can we build rubrics that
show what performance looks like um from
sort of undesirable to really good.
So um of the overall findings let me
I'll just summarize on this slide and
the next uh standards synthesis
warranting and evaluative conclusions
were the least identified as present.
Um, so I think it's worth us thinking
about uh how important that is. You
know, does it matter uh who cares if
we're thinking about the connection to
making a difference? How important is it
to have judgments that connect to the
decisions that need to be made? And um I
think evaluation can be particularly
helpful for that if if evaluative
reasoning and synthesis is present. But
I'm absolutely curious to hear what you
guys think. Uh also evaluation as the
profession generally has emphasized
describing so like program theories
measurement causality uh you know
thinking even about realist evaluation
there's been a huge amount of focus on
describing what's happening this is
super important and I'm not uh I'm not
trying to undermine that in any way I
think what we need to know about fully
describe is as important as fully
judging but in general I think we're
sort of like the guys at the gym have
only done who had been skipping leg day
and just doing arm day. So, we're a bit
overdeveloped and this is probably the
worst uh stick person I've ever drawn.
We're a bit overdeveloped in terms of
our upper body if that's the fully
described part and a bit underdeveloped
in our legs. Uh if we're talking about
that as the fully judge part. So, um
yeah, we might need to do some leg day
and I'm curious to hear what you guys
think of it. So, that's all the things
and I think I'm handing back over to
Phil to send us off toss or let's let's
leave those um let's leave those up for
a moment. Um Amy and uh I just noticed
there's some really great um
contributions in the chat. Lots of um
interesting issues that that are raised.
Um uh I also love your stick figures,
Amy. Um but um I noticed that quite a
number of people are asking about
examples firstly you know links to those
four that you thought were good you just
want to speak a bit about that um yeah
yeah so um in the paper we'll share some
examples Katherine was particularly
insistent that we not identify things
because she didn't want it to feel like
our paper was calling people out for
underperforming or because their things
hadn't ticked all the boxes.
So, uh what we're going to do over this
next bit is to try to get together
people who are interested in this and
are willing uh and allowed by their
employers to share uh examples and are
comfortable with that so that we can
start to build a set of reports that we
can share uh without compromising um you
know people. So yeah, I appreciate the
question and I kept saying how about now
and she kept saying no Amy.
So I and I think that's important. We
want to make sure that uh we're
exploring this in a way that helps
people uh engage and think about it but
not feel so uncomfortable they can't
stand it or be alienated by the
conversation. Great. That makes great
sense. Look forward to that next stage.
Um there is one question there. I think
it'd be good to to just touch on now
about the issue where does the issue of
interpretation of measures and
conclusions concerning causality fit?
measures might only relate to what
happened rather than to what happened
rather than why which I think comes to
your to your fully described uh issue at
the end perhaps. Um so when when I've
talked about this with colleagues who
are doing this essentially you are still
doing that interpretation of of
causality and you know what what h what
is the thing and how did it work and
what happened as a result of it all
those questions are still really
important. um think putting the
evaluative reasoning sort of around it
means that the questions that you might
ask in the measures could be different
because you've made some choices about
what good looks like and how you're
going to what you need to pay attention
to to understand whether good is
happening. uh and it will also generally
mean either you've thought about how to
take that analysis and synthesis of the
data and connect it to the reasoning
immediately like it's all part of of
that synthesis process or you have a
two-step bit where you try to make sense
of the data and then you use it in
something like a rubric or a decision
tree or some other kind of way. uh you
know Jane Davidson's book has a variety
of synthesis uh methods that you can use
to do that kind of work. So I think does
that answer the question? I I think that
goes a long way to it. Um but but but
we'll have an opportunity to come back
to that and I can see a number of uh of
extra questions dropping in but but as
we said at the start we do really want
to give you a chance to discuss um the
implications of what um Katherine and
Amy's research is uh are for the work
that we're doing. So we we do want to
jump into breakout rooms um uh at this
point and then we'll circle back in the
um in the plenary uh to hear what you
what you um what burning issues arose
what what insights you had in the
breakouts but also to address some of
these great questions that are coming up
in the chat here. So um
Julia I think are you able to initiate
the the putting people into breakouts
bit now.
Fantastic.
Bye-bye everyone. See you on the other
side.
I can't hear Phil. Can someone
Phil, you've gone quiet.
Yes, we can hear you.
because my my headphones died. Right.
That's great. Can you hear me now? Yes.
Can you hear me now? Yes.
No. Yes.
He can hear us. Yes. That's a bit Okay,
that's a yes. All right. Trying to find
someone who's nodding at me. I couldn't
find anyone. Okay.
Thank you everyone um for uh uh
rejoining and and as I said I hope there
has been some some really generative
discussion um in those
start my video. Is that what I need to
do? How about that? Okay, now I'm back
in the room as well. This is why they
don't let me loose the the rest of the
time. Um
so
just do that. Um what we want to do now
is start hearing from you all about the
issues that that came up um in your
discussions in the breakout rooms. Um
and just just before I do that um uh I
do want to circle back to um one or two
of the comments that were made before we
went into those rooms um and just give
give Amy a chance to comment on that. So
I mentioned right at the start that
there was um when we at the end of of
Amy's talk that there was a comment from
Eva about um warranting and she was
asking with warranting are the theories
formal or folk theory or both um and
then uh followed that up with the the
request for an example of how to justify
choices of theories. So I wonder um Amy,
can you can you comment on that one?
Yeah, I think so. I think often program
theory gets constructed based on
consultation with people about how they
think the program works. That's legit.
Uh program theory also can come out of
actual theory. So um I'll show a slide
at the end where I've got all that sort
of assembled but uh Finel and Rogers
book is a great source for that because
they have it has a whole chapter on sort
of archetypal theories and archetypal uh
archetypal um program logics and
theories of change. So I would
definitely recommend that to you if you
don't already have it. Um, one of the
ways that can be quite helpful in
evaluation space is to say this program
is a is this kind of program and there's
a bunch of literature in the world that
says this kind of program works in this
way. So if it's a behavior change
program then you can go to the
literature around behavior change and
it'll show you sort of what's happening
typically in that space and what things
actually work and what things probably
don't and what you might measure. So I
think that's one of the spaces where uh
in relation to the fully described side
of things evaluators uh I think quite
often we we build theory using the
people and the documents but we don't
necessarily use theory out of the
literature. Um and that's a way we can
really use to strengthen our arguments
about how stuff works. Great. Thanks
Amy. Um and so so can I now ask um all
of you in your to to tell me about the
the issues that came up in the
discussion room you were in. We asked
you how the presentation resonated with
your practice with your experience. Um
what the context or circumstances are
that enable you to um introduce explicit
evaluative reasoning into the work that
you're doing and what challenges you've
faced. So um can I just open it up and
and ask for um people who uh have been
in those those discussions in breakout
rooms what was the what was the dominant
sort of thing that was coming out of
your discussion there? What was what's
what sort of um experiences um were you
sharing or um insights or challenges or
strategies were you sharing?
Um Phil, I'm happy to go first. Um in
our group um that's okay. Um in our
group it was really about the context um
and uh a few of the members were
internal evaluators. So understanding
the context in which they operate in a
bigger sort of in institutional or
organizational context and that may well
you know influence um um some of the um
warrants.
Um but there was full agreement on the
importance of um being explicit um the
importance of stakeholder engagement
very early on um being really clear
about expectations um and not just you
know forming assumptions um um in in the
leadup to um developing the um
evaluation questions um um I'd sort of
posed a question which we didn't get a
chance to answer towards the end um in
relation to you know how does corporate
culture and risk appetite influence
again and and that sort of um I guess um
dovtales into again that operating
context of an internal evaluator versus
say an external consultant and um the
level of independence and um um often um
there was also discussion about um um
the final reports only including
findings not necessarily recommendations
and that was a particular mandate of one
um department um and then it was up to
management um to develop their response
um in accordance with the findings um
and you know an action plan to map out
um and I'm sure that they do that for
you know appropriate resourcing um
reasons but um yeah it was quite
interesting in that regard thanks
yeah no thank you for that and um so
many useful things there um about um
whether there is a difference in the
experience of people who are in larger
agencies or an internal role or whether
you're playing a role as a as an
external consultant. Um does it does it
look different? Um if if does does
anybody want to comment on that um
experience? I some of you I'm I'm aware
have looked at it from both sides.
So
yeah, I could jump in there. I have
looked at it from both sides. Yeah.
I'm not sure it's super different
depending on which side you're on. I
think um on both sides the key challenge
often is a lack of of a lack of
available explicit standards or criteria
or sources of those. So where there's
nothing in the published research
literature to tell you what good looks
like um and where you're working, for
example, in a highly politically
sensitive area of policy that um doesn't
even make it into the public domain,
let alone the published research
literature. And the client from an
external perspective might not actually
want that but they do want us to be very
clear about the evidence on which we
have made our judgments. So I think that
they
like to have that role themselves I
guess of kind of
assessing from their perspective what
good is and whether we've made a
reasonable judgment
and and Amy would you say that that's
then an example of um explicit
evaluative reasoning when you get to to
that point even if you don't actually
make an ultimate judgment. We do make
judgments. Yeah. Yeah. I think it's it's
that question of how
and I think that goes to what Kim said
first which is the context in which this
is happening. Do they want evaluation
like do they want you to facilitate a
process that arrives at a judgment or do
they want something that's much more
researchbased and I and there's risk and
all kinds of other things are a part of
that conversation. So for me I think
it's about offering the judgment part as
an option on the menu of things that
evaluators do uh and can provide and
that's different than saying evaluation
is the same as research or evaluation
measures stuff. You know, if we offer
them the chance to say if if we pulled
this together, what level of judgment
would be helpful and how could that
facilitate your decision- making, then
that's different than saying, you know,
you have to do a judgment all the time
because I think that's not always
appropriate. Yeah. Like we often do make
judgments, but don't we'll ask we'll ask
the commissioner whether they are
looking for explicit recommendations.
Um but the judgments won't ne they'll be
based on the measures if you like
whatever kinds of measures they are um
but we'll make a kind of a reasoned
assessment of what whether that's good
or not based on our understanding of the
policy context and the hardness of the
problem I suppose and sectors where
we're working all the time.
Yeah. And Eva says uh in the chat,
couldn't the judgment be done using
participatory methods or facilitated
process? I think absolutely. And that's
like a good evaluative judgment process
isn't just the evaluator sitting down
with Yeah, I love a rubric, but often I
think um clients don't have the appetite
or the time for it. Well, and if you're
having to build from scratch every time,
that impacts it, too. So, yeah.
I was going to say yeah no I was going
to say that u as an external evaluator
um yes um often there is a a desire for
findings and some kind of uh you know
conclusions or a judgment around those
but not necessarily recommendations and
I think that's appropriate um in the so
far as um recommendations require a a
more a broader understanding of the
context in which the findings are
landing. Um, and often it's only the,
you know, the policy makers or the
people within a government department or
or whatever that have the full breadth
of understanding and not just the, you
know, of the political, the social, the
the environmental elements that need to
be taken into account to form good
recommendations.
Yeah. And I think that's a really good
um point that that has come up a lot in
discussions about whether whether the
evaluator's role, you know, needs to
extend to those recommendations or how
you provide satisfactory evidence base
for um for recommendations to be
formulated. Um I think that distinction
is one that that that a lot of people
are navigating uh frequently in in
especially in government uh but but
obviously outside government as well. Um
and I use that um term navigating. I
wonder if if in other rooms or or other
participants today had experiences about
how you do navigate those challenges um
uh collaborative processes being one of
them. Um or um other challenges that you
found in um in addressing explicit
evaluative reasoning in your um your
reporting.
We talked a bit about um I guess the uh
the uh capability or the awareness that
the sponsor and the stakeholders have
about evaluation and also um you know um
I you know in terms of um uh do they
have do they understand what
evaluation's about? Do they have an
evaluative attitude rather than they
don't have to have evaluative reasoning
but it's also about the role of the
evaluator needs to be a translator in a
lot of ways in terms of um well you know
looking from what uh what does it mean
to fully describe but also what does it
mean to fully judge um and what do you
want in between so we just just covered
a bit of that um but the um yeah and
there was something else but but I I
guess that that context is really
important and also a lot of times we're
coming in and doing something after the
program's actually been running or you
know things. So it's it's actually also
about how can one actually educate or
the you know sort of pe the people
making decisions about programs to
actually involve evaluators at the front
end that you you you know you can
actually because they can actually also
think a bit about well you know if they
haven't actually thought about um what
uh program theories around you know
around or or behavioral theories around
um to you know to say well maybe maybe
the program as described isn't fit for
purpose. Is it going to do do what you
want to do, but how also do you do you
think you want to measure it? What do
you want to achieve at the the outcome?
So, so there's there's a a number of
factors here which is not just about
sponsors and their awareness um but it's
about that educative informative process
um and and also I suppose trying to
encourage people to think that
evaluation front end
of a process. Yeah.
Yeah. Wonderful. Um uh to hear you hear
that contribution. I I I I tend to get
known as the person who says evaluation
is ECB
that you that it's a necessary
capability of evaluators to be able to
um work with people so that they can
make the best use of of evaluation
findings. So yeah, I I I I thoroughly
agree with that. Um, Amy, feel free to
jump in at any time on on any of these
responses as well. Um, uh, as as one of
our our foremost educators about
evaluation.
Well, I think I was thinking about the
the connection between
evaluative synthesis and reasoning and
recommendations. So, I think those two
things often get grouped together and I
do I don't think of them as the same.
So, Scriven's written a bunch of stuff
about when to make recommendations and
when not to and what what kind and I'm
at like I'm in the Scriven camp. So,
that like that part is that I think what
Lee said is really important. But I
think the synthesis the evaluative
synthesis step is something different
because if you've if you've aligned the
synthesis process to decisions that have
to get made you might not be making
recommendations but you might have said
look if performance looks like this you
know if you're running a training
program for instance and half and you're
not even getting the right people in the
room that directs you without any
recommendation being made somebody would
look at something and say okay like we
need to fix that or or we're just
wasting our money. So, I think it the
explicitness of that reasoning and
connecting it to the decisions or things
that need to be paid attention to
actually means you don't need
recommendations necessarily because
people can see what's happening and make
choices based on that. So um I think if
you're in the space where
they don't want that explicit evaluative
reasoning or synthesis then the
recommendations question becomes a bit
more live like what what if anything do
you say that you think they might do as
a result of what you found in this
space? But again that's it's a question
of appetite. you know, does the person
who you're doing the evaluation for or
the community that you're doing it with,
do they want that kind of explicit
connection uh or not? And and that goes
to the risk tolerance, that goes to uh
evaluative, what I think of as Jane
Davidson calls it evaluative attitude,
like what's your openness to
understanding if shit's not working? And
if you if you don't want to know, then
an actual evaluation is not going to be
appealing to you at all. So um I think
for us as evaluators understanding that
appetite from the beginning can be quite
helpful because if you keep trying to
push reasoning on people who don't want
it you will spend a lot of time in
misery but if what they want is really
just facts then we can do that too
because that's you know that's our bread
and butter already. So
thank you. Um I think um what you've
said there Amy is touches on something
that we discussed in our group too is
about um
uh commissioners wanting short reports
but I mean just because you are doing
very explicit evaluative reasoning
doesn't mean it has to be in the report.
It can be something that sits behind
the evaluation and you can still have a
short report if you if you want to. I
also think the reason can contribute to
short reports like you can put up a a
rubric and show how performance happened
and that's a page. Someone said to me
once that CEOs need cartoons like they
need a one-page drawing of things
because they're not going to read all
the stuff and a a rubric is essentially
that. So, if they're open to that kind
of thing, it's a way to to give them
what they want.
I'm I'm going to jump in and say we
we've got 3 minutes left um uh before
12:30. And so, even though there's
there's great conversation going and
there's no reason why we shouldn't keep
it going after 1:30. Um I just want to
say a couple of things um before we
formally wrap up. Um first and foremost
um a huge amount of thanks to all of you
for um bringing such um engagement and
and great thinking to to this session.
And of course, thank you to Amy uh for
um bringing this deep thinking and
research to us. Um and we we pass that
um thanks on to Katherine as well. Um
and um for um everyone who is in the
room with us um right now, we would like
to um we'd like you to uh complete our
um our feedback form which starts off
again appropriately um reconciliation
week by asking you which Aboriginal
lands you're on today. Um but it'd be
really great um if you could um fill in
that uh that form for us so that we get
some uh information about about what's
happening you know with you and and how
to make these sessions better. Um watch
all the alerts from um AES um about
what's coming up next. Um, as I said,
uh, perhaps, uh, a little bit unclearly
at the start, the, um, New South Wales
chapter runs these, um, uh, networking
opportunities more or less every month,
last Thursday of the month. So, watch
out for free, um, seminars and
opportunities to get together. Um, often
they're later in the day, so that you
can, um, socialize, but um, uh, there's
a there's a mixture of of of times that
they're on. So, keep looking out for
those things. Um and uh and and finally
um yeah, watch have a look out for
conference. Um uh actually we we've got
a minute and I think Amy, you might have
some information about something that's
that's happening at conference and I um
I'd really would love people to uh to be
able to see that as well.
Can you see it now?
We can here. Let me go. So there is a
question about warrants. I made a Wii
slide about that. So if you're looking
for and we talked about it, but there's
the book. Uh so this will be in the
slide pack. Uh if you want to be part of
the cheeky cabal of people who are
talking about this, you can scan this QR
code and it just means you're giving me
your uh contact details so I can stay in
touch about things. Uh people are
distributed like dandelion seeds. So,
we're trying to uh get everybody who's
interested in this together. And I have
no idea what that's going to look like,
but if you tell me that you're
interested when something happens, then
I can tell you. Um I have a few things
happening at uh AES25. So, I'm doing a
full day workshop with Ben Lawless on
rubrics. When we say when Ben as an
assessment person says developmental, he
means stages of learning, not Michael
Quinn Patton's developmental evaluation.
So, let me just be clear about that. Uh,
and then I they always put me two things
on the same day. So, we're talking
Katherine and I will be there to talk
more about this and uh have more
conversation about this at 2 o'clock on
Thursday. And then the evaluator
competency work with the pathways
committees uh will just give you the
next installment of what's happening
with that at 4.
And there's references which will come
out with the things. And here's a
dandelion which was drawn by the
daughter of one of my students. How
awesome is that? So, thanks everybody.
It's been really lovely. Yeah, thank you
Amy and thanks everyone for uh your
participation today. Um as Amy said,
those things will be in the um the pack
um that will the slide pack that we'll
make available to people. This
presentation, the recording will go up
on the AES YouTube site in due course.
Um so you'll be able to to see that
there. Um and uh it'd be great to have
more people involved in this
conversation. Um, we really appreciate
the fantastic contributions today. Um,
and uh, look forward to seeing you at
conference and in other AES uh, meetings
if there's um, thanks Phil. Bye.
Questions. Can people stay back for a
few minutes? I just was asking probably
Amy and you Phil. Oh yeah, I'm can hang
for a bit. That's fine. Yeah. And I can
hang for a bit as well. And uh so if you
if you would like to uh to pose another
question then um this is a great time to
make use of one of Australia's foremost
evaluation educators. Oh my gosh.
Hey Gemma.
Hi. Thank you. Thank you for the extra
time. Um I am an independent evaluator
and we've definitely found that a lot of
our clients often in international
development are not particularly
interested in the logic of evaluation
side of our proposals and projects and
the the concept of a judgment doesn't
sit well with the fact that they seem to
be primarily interested in wanting to be
told what to do and um you know have
something to show for it and I had so I
was really interested in your suggestion
there of you know having a really
explicit discussion to assess their
appetite for um you know the
explicitness of the reasoning or having
judgments.
If um the appetite for that was low um
you know do you think
yet they are intent on this idea of
having an evaluation? Do you think we
should sort of you know steer them away
from the idea that what we're doing is
an evaluation into something else or do
you have any tips on like how to manage
those conversations? Oh gosh. Yeah,
that's a big question. I think um if
what they want so this goes to what Ben
and I are going to talk about in our
workshop. If what they want is to what
they what to do you know steps to
improve then having a set of rubrics
that actually says here's what your
performance looks like now and here's
what the next level up of better
performance would look like will
actually answer those questions for
them. And like Ruth uh said whether or
not that goes into a report
that like that's the existence of that
can be quite helpful. And I think one of
the things in international development
that can be so frustrating is that this
top- down stuff coming from an external
international funer that doesn't
necessarily match up with on the ground
program stuff or even what communities
think is important. And so that you know
you're kind of you you've entered into a
conflict between local uh local
expectations and needs and whatever is
coming from outside. And so being
explicit about that actually can be
quite helpful to like make those values
visible and show that there's a
conflict. And if you're using that sort
of uh reasoning step then you can say
this will look good to the funer for
these reasons but it will be bad for us
for these reasons. And it makes that all
visible. So I think you can make a case
for whether they want judgment or not.
You can make a case for saying if we use
this kind of thing like I've said so
many times to students, I don't care
what you call it, let's just get the
practice happening. So, call it what you
want, you know, say, "Look, let's talk
about how we could be developmental
about this or whatever that that gets
them into that kind of process and
thinking uh and and helps to make those
values visible if there's a conflict
because I think that's the other place
where the if you don't do that then
usually it's the whoever is paying for
the evaluation, it's their values that
get used and then that can fundamentally
be detrimental actually to the thing
that they're meaning to do. So, yeah.
Sorry, that was a long answer to your
short question. Did it help? No, thank
you. I appreciate it. Yeah, thank you.
Yeah,
but it's a good question.
Anybody else have burning things or are
you all just thinking about lunch?
I have a quick question if yeah we have
a moment. Um so we were talking before
and I think um it was talked about that
a really difficult step can be defining
standards when they don't already exist
and that it feels like you get wrapped
up in you know say it's like it's
service delivery it's like I need to
define the standards for all service
delivery in the world in order to define
whether or not this pilot is effective
or not. Yeah. And so I think that that's
probably the step that I struggle the
most with and I was just wondering how
do you approach that? Like how do you
not get embroiled in
the standards of the entire
space? Yep. That's a great question. I
think one of the things that I see quite
often is
trying to make too many levels of
differentiation. like when sometimes
especially at the beginning you need to
know if something is harmful and if
something is okay. So rather than trying
to have like poor you know okay good
excellent like you don't need to be that
precise especially when you're starting
out. So, if you all have a conversation
about how will we know if something is
going terribly wrong that we need to
address immediately
versus this is pretty okay, like we can
leave this sitting like no harm's being
done. It might not be the best thing
ever, but it's not actually actively
hurting anyone. Uh that might be all you
need like to get started. I think it's
the over complication of setting
standards that gets in people's way
sometimes. So that you can use the
literature to have a sense of what bad
stuff to look for and how to set your
standards to make sure that that's
captured. And that's what I would
suggest instead of trying to be
overdeveloped, you know. Um just think
about what what do what do you need to
stop happening? Like what does the
evaluation need to identify? cuz I one
of the I've been reading this woman now
whose name escapes me, but the book's
called Quit and she's a person who does
decision theory. She was a pro poker
player
and then she went did a PhD because why
not? Uh it'll come to me in a minute.
Anyway, the quitting is the hardest
thing for humans to do.
We are not wired to do it. Every bias
that we have, this is Amy's too long,
didn't read, summary of the book, every
part of our brains is wired against
quitting. And what wasn't already wired
against quitting in our sort of
fundamental nature has been reinforced
against quitting by western culture that
says strive, do winners never quit and
all that But we cannot as program
evaluators like helping people quit
stuff is fundamentally important to what
we need to do. So setting standards that
help people see that something needs to
be quit is probably more important than
almost any other kind of thing we can do
in that step.
That's great. Thank you so much. Yeah, I
think I think that's um something that I
was thinking about too when you were
speaking of like um you know why are
some of those steps commonly missed and
I think sometimes the complexity
and the systems thinking level um you
know I'm doing an evaluation around
community resilience at the moment and
it's culturally relative it's complex
it's we're talking about thriving
communities we're talking about absence
of this and that like
There's actually a lot there to unpack.
Um, and we're doing that process at the
moment. But yeah, do you have a a
comment around I mean there's a reason
when I said this the the joke the
comment I don't know what it is the idea
about we've been skipping leg day one of
my students immediately responded squats
are terrible and hard. Why would you do
them? And I think this is like the
answer here is the same. this value
stuff is hard and that's why the
measuring not that measuring isn't its
own challenges but it's e it's been the
thing that's easier to miss and because
we aren't building a database of
criteria for instance like here's a set
of criteria that have been used in
similar programs over x amount of time
that like that's a missed opportunity we
have as a profession to make this stuff
easier for ourselves so um if you're a
person who has access to literature
Rebecca Teasdale is an evaluator in the
US who's been doing a bunch of analysis
of reports and articles and pulling out
sets of criteria that were implicit in
those. Um, I have a handout that I
pulled together for students and I'll
see if they'll let me share it with you
guys that just has like here's all the
criteria and here's how you might like
put them along in the life cycle of a
program. um which hopefully will help
you too cuz I think the too hard basket
is a that's real like it really is too
hard to try to do this by ourselves
which is also why I'm keen to start a
conversation about how can we share
about this especially if it can't be
explicit in the reporting and the
publicly available stuff how can we
start to accumulate across reports so we
can share out those products of
evaluation that we all could use to make
our practice better and easier rather RA
than trying to everybody do this on the
on their own cuz it just can't it won't
get done. It's too hard. Yeah. And
sometimes it sort of I feel like the
evaluation then butts heads with the
spirit of the program that's trying to
be transformative, not trying to be
reductionist. And here we are over here
trying to set these standards and this
particular way of looking at the world.
And it's just um that's another reason
why I think some of those don't happen
because we can't we can't match this and
that sorry I'm just recovering from the
flow bit croaky but um yeah sort of how
that's why I've put in some of my
feedback around how to we do this in a
non reductionist way. That's the framing
of a lot of the underpinning philosophy
and world view that we're coming at it
with when the program staff are looking
at it in a completely different context
and worldview. Yeah. Um I'm I'm saying
this out loud because then I'll be on
the recording and then hopefully I will
remember to share stuff. But we um
Mattea Roa did her PhD with us and she
made a matrix a values identification
matrix to help surface the values of
different stakeholders and the headings
on that are based on western philos you
know like western normative definitions
of what good looks like. But you can
shamelessly repopulate all the column
headings to stuff that matches your
context and to and then you can say okay
these stakeholders think these things
are important. this stakeholder group
actually thinks a whole bunch of other
things are important and then it's
visible so you can have that
conversation and say well we don't agree
about what's happening here so um that
can I think be quite helpful too because
it it it takes out that reductionist
concern and actually makes the different
values that groups have or uh
stakeholder groups have visible
fantastic thank you yeah for sure thanks
for your questions
fantastic questions. Um,
and and real life challenges.
Uh, yeah. Well, and I think that only
way this evaluative reasoning stuff will
work is if we figure out how to make it
viable in practice. So that that to me
is this isn't a scholarly exercise on my
part. The reason that I want to
understand this more and do more
Katherine and I want to do more research
on it is we'll only be able to do it if
we can figure out how how it's been done
well who like and what secret
squirrel kinds of activities people have
done to make it happen. Uh to make those
more visible in whatever way people can
be comfortable with. uh and to figure
out how can we streamline people to be
able to to do this work instead of
having to build the wheel every time
because it just it's not viable. Can I
throw Liam under the bus a little bit
here because there's uh you know he he
worked on designing the evaluation for a
big complicated emerging program with
teachers at the center as innovators and
you know kind of making it usable while
you're building the plane you're
building the evaluation and you know
getting towards that impact.
All right then. Thanks Julia.
I think
u
there are parts where you need to really
engage with the complexity and there's
parts where you can kind of try to
minimize it. So, one of the things that
I just mentioned in the chat there was,
and this reflects, I think something
similar to what you were saying, Amy, I
was agonizing over building a
effectively a rubric for for the
evaluation. And I was agonizing over the
different levels of the rubric in
particular. And um I'm going to pay
tribute to our academic partner who was
working with us at the time who said why
don't you just go with the excellence
descriptor and then you can actually um
sort of use that to make a judgment
against what excellence what excellence
actually looks like and that was kind of
a real
light bulb moment and you know probably
saved my sanity as well because it sort
of meant that I could actually you know
really zoom in on the stuff that we
wanted to See, I think as well
it's
like so much of it probably from my
perspective
is it's probably different to some
people's context because I was an
embedded evaluator uh in that context.
So, and that's sort of something that
I've you know I've been thinking about
for a while because that's sort of a
model that I've actually worked with
since you know probably for about just
over a dozen years now. And I think that
that model of embeddedness
where there is really good authorizing
environment where you build
relationships with the people who can
make the decisions that matter. Um I
think that can be that can be really
useful as well. So I I'll I'll be
presenting with my old boss uh on that
program at the AES conference as well.
So, and it's sort of a bit of an update
of a of an article I did about 10 years
ago. So, it's kind of the 10 year
anniversary of that. But I think
yeah, it it's it's simplify the bits
that you can simplify reflect complexity
in the bits where you still need to
reflect the complexity and just be
people's friend as much as you can.
That's fantastic.
I think that's I think that's fantastic
advice and and uh eager to to hear that
update. Um obviously that 10-year
update. Um also appreciating the
shameless plugging of people's uh AES
conference papers. Um anybody else who
would like to plug uh an activity uh at
this point's very welcome to do so. Um,
no, but I do think that's really um
really valuable from from everybody on
this topic about how you realistically
face the challenge um that complexity um
uh presents and you know as Lachlan was
saying you know when when there aren't a
set of standards when you're trying to
develop that from scratch um what are
some real world you know actual examples
of of how you can go about this to
minimize it while Amy and and the the
dandelion seeds are getting together
this new resource course um you you got
to get something done and and I think
that Liam's um uh approach of of using
um excellence descriptors as your
starting point because you want to know
what good really looks like um and then
not get um buried in the in the minutia
of how good is it you know what's the
difference between average and slightly
above average you know because that's
just not going to be helpful I think
that's really worthwhile as
Angelina,
I saw that you made a comment in the in
the chat as well.
I did. Um, and just really quickly, I
think
I'm trying to put my thoughts in order,
but I think sometimes we shy away from
imposing boundaries on our work because
we don't want to be reductive. I'm just
calling back to a conversation that you
guys had a couple of minutes ago. Um,
but I think boundaries are inherently
necessary in order to actually complete
the work. And it's boundaries along like
we impose boundaries on a lot of things,
whether it's the type of data that
collect that we collect, how much data
that we collect. And I think we need to
take a step back further and also think
about how we're going to impose
boundaries on what it is exactly that
we're evaluating and how exactly we're
going to be judging our success. Because
without those boundaries I and without
them being clearly articulated,
what are we doing? Like we don't even
know. So how is our reader going to
know?
Did I say fully describe um that this is
a dimension of fulliness? Um yeah. Well,
I think it's fullness not in a sense of
like describing absolutely everything,
but fullness in the sense of being
really explicit about like where that
fullness is bound. like what what is
going to be fully described. So within
these boundaries, we're going to be
really full in terms of what we
describe, but we're not going to
describe absolutely everything.
Thank you.
I'm starting to get things pop up on my
screen saying I have to go to another
meeting. Yep, that time of day. Um uh so
much as I would love to continue this
conversation now, I think I'll just say
again we can continue this at conference
and I hope at some other um uh um venues
as well. Um Amy, you know, we do still
have the intention to have um a a panel
discussion with a wide audience um when
when there's an opportunity to do that
um and more of the and closer to the
paper being published as well. Um so
watch watch all of your um alerts for
for information about that um about that
panel discussion where we would bring
together some some people who who are
really struggling or not struggling but
succeeding to address these issues
across a range of different contexts in
both in government and outside
government.