Submind YouTube summaries
Thumbnail for Calling evaluative reasoning! Come out, come out wherever you are!

Calling evaluative reasoning! Come out, come out wherever you are!

Watch on YouTube

Video summary

In a seminar hosted by the New South Wales Committee of the Australian Evaluation Society, Associate Professor Amy Gullixson presented critical research on the state of evaluative reasoning within the Australian education sector between 2014 and 2024. Her analysis of thirty-seven publicly available reports revealed a significant gap between the profession's emphasis on describing measurements and causality versus the actual practice of making credible judgments. While all analyzed reports included measures, only four demonstrated full evaluative reasoning by completing logical steps and providing justifications for every choice; common deficiencies included missing research questions, performance standards, synthesis methods, and overall judgments. Gullixson argues that this imbalance is akin to skipping leg day at the gym, where the foundational work of judging is neglected in favor of superficial description, ultimately weakening the connection between findings and decision-making. To address these challenges, the discussion shifted to practical strategies for managing client resistance and defining standards in complex environments. Speakers suggested using performance rubrics that clearly distinguish between current and improved states to provide actionable steps without forcing a formal judgment when stakeholders prefer direct instructions. This approach helps make conflicting values visible, allowing evaluators to explain why certain actions align with funder expectations while potentially harming local communities. Furthermore, rather than reinventing criteria from scratch, the group advocated for sharing implicit standards from existing reports and utilizing tools like Mattea Roa's values identification matrix to surface stakeholder values without being reductionist, ensuring that boundaries are established to define success even when formal standards do not yet exist. The session also highlighted the importance of simplifying frameworks to focus on what truly matters, such as helping programs discontinue ineffective activities rather than getting bogged down in minor distinctions. By starting with broad descriptors like "excellence," evaluators can navigate varying client appetites for explicit reasoning while acting as translators between technical findings and decision-makers' needs. The overarching conclusion emphasized that overcoming the difficulty of implementing evaluative reasoning requires making these practices viable through shared resources and streamlined methods, ensuring that the profession can effectively balance organizational culture, risk appetite, and stakeholder expectations to produce meaningful, well-reasoned evaluations.
Read the full video transcript
So, I'll start off by saying welcome. Welcome to this um uh this seminar organized by the New South Wales Committee of the Austral the Australian Evaluation Society. My name is Philip Belling. I'm an evaluator at the Center for Education Statistics and Evaluation in the New South Wales Department of Education. Um and it really is great to see some familiar faces here, but it's also great to see a lot of participants from across the profession. We um started out calling this session um a provocation and a dialogue. So it's really good to be able to provoke and discuss with such a diverse group of evaluators. Today's provocation is going to be uh from associate professor Amy Gullixson. Um but before I hand over to Dr. Gixson and also noting that we are in national reconciliation week. I'd like to start by acknowledging the Gatagle people of the Aora nation and their elders past and present. We're all on Aboriginal land and I'm fortunate to be on Gatagle land today. Gatagal elders have been custodians of songlines and knowledge and insights over countless generations and I offer my respect and gratitude and I also pay my respects to the elders and traditional custodians of all the lands from which you're joining uh this seminar today and to Aboriginal toouristr islander and first nations colleagues who are with us as well. Um, at this point, uh, if there had been someone from the, um, New South Wales committee, they'd have leapt up to give you information about the AES. Um, I'm going to try to do that. Um, it is an explicit aim of the AES to strengthen and promote evaluation practice, theory, and use. Um, so that evaluation makes a difference, and that that's a lot of our focus today. But it the AES organizes these free monthly seminars but a whole host of other activities. Um and I'd like to encourage any non-members who are here today to think about joining the society. It's actually a really good deal um because you get up to 48% off any of the other workshops that are done and you get 20% off the annual conference registration. So, I highly recommend that you um check out the uh conference this year in beautiful Canberra um and check out the program that's online especially for pre-conference workshops which are being delivered by some outstanding evaluation leaders. Oh, including today's speaker. So, more on that later. Um my role here is to introduce the topic and the speaker and also to let you know how we'd like to handle questions and breakout rooms. So as to the topic um for me it comes down to this that the AES is championing quality evaluation that makes a difference and today's session is all about documenting evaluative reasoning. So throughout the session, we're going to think about how does documenting evaluative reasoning affect evaluation quality and what sort of difference that documenting evaluation reasoning actually makes. So I'm really looking forward to hearing from associate professor Amy Gullixson about recent research about what we actually see in terms of evaluative reasoning when we look at published evaluation reports. But as I said before, time is really tight today. So to make the the most of that discussion time, I want to invite you to make use of the chat while Amy's talking. Um, I want you to use it proflegately and generously and extravagantly. Use the chat to ask questions like, "Amy, what did you mean by or can you say a little bit more about that concept?" Um, but also use the chat for ideas, opinions, speculations, provocations of your own if you like. um and we'll pick those up um in the breakouts and in the final discussion. Now, finally, it's my um great honor and privilege to introduce um our speaker today, Amy Gexon. Um Amy is associate professor at the University of Melbourne where she's played a key role in establishing um and bringing about the success of the Master of Evaluation. She is a highly regarded scholar and she's also a fellow of the AES. She's in demand as a speaker both nationally and internationally and she's had a significant effect on the quality of evaluation in Australia and abroad through her teaching, writing and her inspiring speaking. Personally, I can attest that her keynote at the AES conference in Adelaide has made a lasting impact on me and other members of the team that I then belong to and it continues to help evaluators and evaluation teams to lift the quality of the work they're doing. So with that said, I will now pass over to Amy. Can you see my slides? Yes. Good. Hi everybody. It is a delight to see so many familiar faces and also uh just to see new people and to have this opportunity to chat with you. My name is Amy Gellixen uh and I'm hugely grateful to the group of people who organized this session and uh wanted to hear more about the evaluative reasoning research and to those of you that joined um the person who is working with me my partner in this research is Dr. Katherine Meldram. Um, and on this slide you'll see the thesis that she wrote with us at the University of Melbourne. Uh, the handle thing is a link and so that'll come to you in the slides and I think Philer Julia will put it in the chat too. So if you want to read the full thesis, you can do that. Um, she Katherine has more than 30 years experience doing research and teaching across Australia, Asia and the UK. and she came to study uh at Melour because she believed that well-conducted evaluations contribute to evidence bases which in turn support opportunities so that education can be empowering for all. Um she had a PhD when she came to study with us and she was going to do the ML but exited halfway with the gradert and then did this research on evaluative reasoning and so what I'm going to talk about today started with her thesis which is what the link is to um but we've continued to develop it since. So, the findings that I'm going to share are are from the unpublished paper that we're working on right now. Uh, she's not here today because she's hiking in Scotland for heaven's sakes. She's having a miserable time. So, uh, you get just me instead of both of us. Let me uh start out by talking about evaluation, which is the space we're working in. Evaluation investigates investments of time, energy, uh, financial, and other resources to see if the thing that they're getting invested in is actually worth it. The process of evaluation includes fully describing and fully judging whatever it is we're evaluating. We could debate later the fully word, but describing and judging is the thing that we do together. Evaluative reasoning is attached to that fully judge part. It uses the fully describe, but it's a separate set of steps. It provides credible, valid, and defensible judgments to answer that question of worth. and that worth questions often broken into dimensions of quality, cost and significance. Deciding what to keep doing, what to modify and what to quit relies on this kind of evaluative judgment. So evaluative reasoning starts with the logic of evaluation. On that this slide, it goes from the bottom to the top. So first you identify criteria, then performance standards. uh you figure out what you're going to measure to understand evaluate the performance of the evaluand and then you synthesize all that together to understand how good something is. Uh when we looked at the literature we found two additional steps. One was to that almost always it starts with determining key evaluation questions and those connect with the criteria. Um but uh when we were reading through the literature uh Katherine in particular was demanding that evaluation questions be included because inquiry starts with a question and so that became our first step and then uh we ended with judgment and it in terms of making a statement about the goodness of something or the goodness of the parts of the something which I'll say more about in a little bit because findings need to be stated in order to be shared and findings that don't get shared can't be useful. And if we're going to be doing quality evaluation makes a difference, then it has to be out in the world uh to make those changes. The warrants, this is to get to a justified uh judgment, you need warrants and those are the becausees that support your choices. So uh this orange box is the justifications of criteria, standards, uh measures and synthesis methods. So all those happen throughout the process. Um, warrants can come from lots of places. So, they can be from the literature. Quite often, we might use theory to warrant if you're doing a program theory. That's one way to do it. You look at the literature and discover theories that are related to the program you're evaluating. They can be cultural, methodological, uh, experts, and authority. And that might include laws or other things that we need to attend to in the doing of our evaluation. Uh, I have put all of that into one orange box just because orange is bright and I thought I'd save your eyeballs. But really what happens here is that that warranting thing needs to be done. That warranting step needs to be done every time you're making a decision. It's important to give the because these key eval question evaluation questions are important for this evalu because of these reasons. So uh that's how it all works together. So the logic of evaluation makes it legitimate because you're doing the steps of the logic and the warranting makes it justified because you're providing re reasons um for the choices that you made and altogether that adds up to the elements of evaluative reasoning. Uh on the right hand side are the sources we used to get to this list. Uh we went to look to see who else had done this research before we did our our study. And despite the fundamental importance of evaluative reasoning to building these credible, valid and defensible arguments about how good something is, we could only find four previous studies and those are listed on the right hand side here. Um, three of them looked at reports in various contexts. One of them surveyed uh evaluators and asked them to just self-report on their knowledge and practice. So the whole sort of sample to date internationally in when when we started this research included reviews of 76 reports in total and then 214 uh responses. So we don't even know if those were all separate members or if it's people taking a survey twice from one of the 21 uh global professional evaluation associations. So it's a pretty small group of stuff that's been uh researched. And so on this slide, we've got the findings from the three quantitative studies matched to these elements of the of evaluative reasoning. I'm just going to give you a minute to look at that. Uh overall, Erto and Ozaki had the closest findings to each other. Uh Jane Davidson did a qualitative review of a set of six reports uh for an organization and she generally found that key evaluation questions were missing uh along with not very much synthesis and evaluative conclusions. So clearly more empirical research is needed to understand evaluative reasoning and practice and that's what led to our research question. What is the evaluative reasoning practice in the Australian education sector between 2014 and 2024? This is the conceptual framework we used to identify the presence of elements that contribute to legitimate which is the purple boxes the logic of evaluation and justified orange the warrants uh evaluative conclusions. So you can see that's all laid out. We in the first one we added evaluation objectives because we found when we did an initial review of reports that those objectives and questions were sort of interchangeable. So we wanted to make sure we captured that. We also extended the synthesis step uh to include microynthesis and meosynthesis from Jane Davidson's 2014 paper. Um this was to counter uh one of the long-standing and mistaken arguments in evaluation that an inquiry process can only be called evaluation if it arrives at an overall judgment about goodness. Um Jane's argument was that actually you can make evaluative choices at different levels within an evaluation. So you might make choices about weighting one set of data more heavily than another. uh you might go at the MEZO level be looking at look judging criteria separately from each other or judging components of a program separately from each other and you might not want to go to that overall level of judgment but it's still it doesn't mean you're not doing evaluation it's just that the levels uh that where it's happening are different. So uh this slide has the definitions we used. So we use deductive logic and adaptive and we had an adaptive systematic quantitative analysis process. Uh I'll let you read the definitions because that's what's going to inform what I'm going to share next which is the findings on a few slides. I'm going to go next. Are we ready? Yep. Thanks for the nod, Phil. So, um, we set the inclusion criteria to get as many reports as we could find in English that were, uh, full whole reports. And what between 2014 and 2024, so in that 10ear span, we found 37. And this slide shows you where they were from. Most of the reports gathered data at the national level. Uh you may note that if you're a person who does the math when looking at a slide that this adds up to 38 and that's because one of the reports got data from across three states but not nationally. So the there's about to be a series of finding slides. So uh when I use this table, the top row has the elements and they're colorcoded. And then the second row shows present or not present. And then there's a little extra one for research questions. Uh and then the different finding stuff is going to appear in the third row of the table. So first off, four only four of the whole of the 37 reports that we found ticked all the boxes for legitimate and justified evaluation. So that's that is they did all the steps in the logic of evaluation and they provided warrants for their choices across all those steps. So four of 37. Uh now I've got overall findings in that third row. So the counts represent each of the elements across all of the reports. So uh across the whole data set we found that the elements directly related to research that sort of the fully described space were much more consistently present. So everybody measured something. All 37 reports had measures in them. uh 20 of the reports had research questions and uh and only five of of the whole set had no questions or objectives at all. So this the like what's happening set of questions and how do we know that's was all really strongly present in this data set. The fully judge aspects of evaluative reasoning were the least present. So 10 reports had no evaluation questions and some and like I said five had no research questions. 27 didn't have any performance standards. 29 didn't have any of that sort of lower level synthesis and 32 didn't have any judge of that overall judgment. This makes sense because if you're going to synthesize to a judgment you're building up across the whole uh argument. When I teach it I often talk about this as an unbroken chain of reasoning. So if you've got uh links in the chain that are missing, then your conclusion isn't going to be uh credible at the end. So not surprising that missing some of that stuff would lead to not making judgments. Um one thing to highlight, if you've noticed that at the end in the judgment column, it says there were five reports that had a judgment present. And if you think back to earlier, I said there only four of 37 that ticked all the boxes. um this one report reported a judgment but they didn't have criteria or comparators. So they one of the links in the chain was missing and so that's why it didn't get counted amongst the four 16 reports had no warrants for any part of their evaluation. So that means they didn't explicitly justify any of their choices. Uh and finally, even though we didn't formally code this, only a few reports in that data set set uh had identified limitations, which both Katherine and I were surprised about because uh context and limitations are critical for understanding and interpreting both research and evaluation findings. So that was an accidental finding that we thought was important. Uh there's of course limitations to the study. So first we looked at just publicly available and searchable evaluation reports. Uh we did this because we decided the stakeholders were our primary uh t taxpayers were our primary stakeholders in this research and so we wanted to look at stuff that they could find if they were looking. Uh but we know for sure that there are way more evaluations of education that have happened in the public sector in Australia in that 10-year period. So we're not in any way trying to claim that this is a state of evaluation kind of findings or claims at at that level. We're just saying let's let's talk about what this means and have it as a provocation for conversation. Second, uh this was a desk based review and we just counted the presence of elements if they were explicitly stated. So in some of these reports for instance constructs could have been in there that might have been able to be treated as criteria but we didn't do that sort of extrapolation. And finally excuse me we didn't do any judgments of quality. We just said was it there or was it not there. Um that's the next round of work is to start thinking about can we build rubrics that show what performance looks like um from sort of undesirable to really good. So um of the overall findings let me I'll just summarize on this slide and the next uh standards synthesis warranting and evaluative conclusions were the least identified as present. Um, so I think it's worth us thinking about uh how important that is. You know, does it matter uh who cares if we're thinking about the connection to making a difference? How important is it to have judgments that connect to the decisions that need to be made? And um I think evaluation can be particularly helpful for that if if evaluative reasoning and synthesis is present. But I'm absolutely curious to hear what you guys think. Uh also evaluation as the profession generally has emphasized describing so like program theories measurement causality uh you know thinking even about realist evaluation there's been a huge amount of focus on describing what's happening this is super important and I'm not uh I'm not trying to undermine that in any way I think what we need to know about fully describe is as important as fully judging but in general I think we're sort of like the guys at the gym have only done who had been skipping leg day and just doing arm day. So, we're a bit overdeveloped and this is probably the worst uh stick person I've ever drawn. We're a bit overdeveloped in terms of our upper body if that's the fully described part and a bit underdeveloped in our legs. Uh if we're talking about that as the fully judge part. So, um yeah, we might need to do some leg day and I'm curious to hear what you guys think of it. So, that's all the things and I think I'm handing back over to Phil to send us off toss or let's let's leave those um let's leave those up for a moment. Um Amy and uh I just noticed there's some really great um contributions in the chat. Lots of um interesting issues that that are raised. Um uh I also love your stick figures, Amy. Um but um I noticed that quite a number of people are asking about examples firstly you know links to those four that you thought were good you just want to speak a bit about that um yeah yeah so um in the paper we'll share some examples Katherine was particularly insistent that we not identify things because she didn't want it to feel like our paper was calling people out for underperforming or because their things hadn't ticked all the boxes. So, uh what we're going to do over this next bit is to try to get together people who are interested in this and are willing uh and allowed by their employers to share uh examples and are comfortable with that so that we can start to build a set of reports that we can share uh without compromising um you know people. So yeah, I appreciate the question and I kept saying how about now and she kept saying no Amy. So I and I think that's important. We want to make sure that uh we're exploring this in a way that helps people uh engage and think about it but not feel so uncomfortable they can't stand it or be alienated by the conversation. Great. That makes great sense. Look forward to that next stage. Um there is one question there. I think it'd be good to to just touch on now about the issue where does the issue of interpretation of measures and conclusions concerning causality fit? measures might only relate to what happened rather than to what happened rather than why which I think comes to your to your fully described uh issue at the end perhaps. Um so when when I've talked about this with colleagues who are doing this essentially you are still doing that interpretation of of causality and you know what what h what is the thing and how did it work and what happened as a result of it all those questions are still really important. um think putting the evaluative reasoning sort of around it means that the questions that you might ask in the measures could be different because you've made some choices about what good looks like and how you're going to what you need to pay attention to to understand whether good is happening. uh and it will also generally mean either you've thought about how to take that analysis and synthesis of the data and connect it to the reasoning immediately like it's all part of of that synthesis process or you have a two-step bit where you try to make sense of the data and then you use it in something like a rubric or a decision tree or some other kind of way. uh you know Jane Davidson's book has a variety of synthesis uh methods that you can use to do that kind of work. So I think does that answer the question? I I think that goes a long way to it. Um but but but we'll have an opportunity to come back to that and I can see a number of uh of extra questions dropping in but but as we said at the start we do really want to give you a chance to discuss um the implications of what um Katherine and Amy's research is uh are for the work that we're doing. So we we do want to jump into breakout rooms um uh at this point and then we'll circle back in the um in the plenary uh to hear what you what you um what burning issues arose what what insights you had in the breakouts but also to address some of these great questions that are coming up in the chat here. So um Julia I think are you able to initiate the the putting people into breakouts bit now. Fantastic. Bye-bye everyone. See you on the other side. I can't hear Phil. Can someone Phil, you've gone quiet. Yes, we can hear you. because my my headphones died. Right. That's great. Can you hear me now? Yes. Can you hear me now? Yes. No. Yes. He can hear us. Yes. That's a bit Okay, that's a yes. All right. Trying to find someone who's nodding at me. I couldn't find anyone. Okay. Thank you everyone um for uh uh rejoining and and as I said I hope there has been some some really generative discussion um in those start my video. Is that what I need to do? How about that? Okay, now I'm back in the room as well. This is why they don't let me loose the the rest of the time. Um so just do that. Um what we want to do now is start hearing from you all about the issues that that came up um in your discussions in the breakout rooms. Um and just just before I do that um uh I do want to circle back to um one or two of the comments that were made before we went into those rooms um and just give give Amy a chance to comment on that. So I mentioned right at the start that there was um when we at the end of of Amy's talk that there was a comment from Eva about um warranting and she was asking with warranting are the theories formal or folk theory or both um and then uh followed that up with the the request for an example of how to justify choices of theories. So I wonder um Amy, can you can you comment on that one? Yeah, I think so. I think often program theory gets constructed based on consultation with people about how they think the program works. That's legit. Uh program theory also can come out of actual theory. So um I'll show a slide at the end where I've got all that sort of assembled but uh Finel and Rogers book is a great source for that because they have it has a whole chapter on sort of archetypal theories and archetypal uh archetypal um program logics and theories of change. So I would definitely recommend that to you if you don't already have it. Um, one of the ways that can be quite helpful in evaluation space is to say this program is a is this kind of program and there's a bunch of literature in the world that says this kind of program works in this way. So if it's a behavior change program then you can go to the literature around behavior change and it'll show you sort of what's happening typically in that space and what things actually work and what things probably don't and what you might measure. So I think that's one of the spaces where uh in relation to the fully described side of things evaluators uh I think quite often we we build theory using the people and the documents but we don't necessarily use theory out of the literature. Um and that's a way we can really use to strengthen our arguments about how stuff works. Great. Thanks Amy. Um and so so can I now ask um all of you in your to to tell me about the the issues that came up in the discussion room you were in. We asked you how the presentation resonated with your practice with your experience. Um what the context or circumstances are that enable you to um introduce explicit evaluative reasoning into the work that you're doing and what challenges you've faced. So um can I just open it up and and ask for um people who uh have been in those those discussions in breakout rooms what was the what was the dominant sort of thing that was coming out of your discussion there? What was what's what sort of um experiences um were you sharing or um insights or challenges or strategies were you sharing? Um Phil, I'm happy to go first. Um in our group um that's okay. Um in our group it was really about the context um and uh a few of the members were internal evaluators. So understanding the context in which they operate in a bigger sort of in institutional or organizational context and that may well you know influence um um some of the um warrants. Um but there was full agreement on the importance of um being explicit um the importance of stakeholder engagement very early on um being really clear about expectations um and not just you know forming assumptions um um in in the leadup to um developing the um evaluation questions um um I'd sort of posed a question which we didn't get a chance to answer towards the end um in relation to you know how does corporate culture and risk appetite influence again and and that sort of um I guess um dovtales into again that operating context of an internal evaluator versus say an external consultant and um the level of independence and um um often um there was also discussion about um um the final reports only including findings not necessarily recommendations and that was a particular mandate of one um department um and then it was up to management um to develop their response um in accordance with the findings um and you know an action plan to map out um and I'm sure that they do that for you know appropriate resourcing um reasons but um yeah it was quite interesting in that regard thanks yeah no thank you for that and um so many useful things there um about um whether there is a difference in the experience of people who are in larger agencies or an internal role or whether you're playing a role as a as an external consultant. Um does it does it look different? Um if if does does anybody want to comment on that um experience? I some of you I'm I'm aware have looked at it from both sides. So yeah, I could jump in there. I have looked at it from both sides. Yeah. I'm not sure it's super different depending on which side you're on. I think um on both sides the key challenge often is a lack of of a lack of available explicit standards or criteria or sources of those. So where there's nothing in the published research literature to tell you what good looks like um and where you're working, for example, in a highly politically sensitive area of policy that um doesn't even make it into the public domain, let alone the published research literature. And the client from an external perspective might not actually want that but they do want us to be very clear about the evidence on which we have made our judgments. So I think that they like to have that role themselves I guess of kind of assessing from their perspective what good is and whether we've made a reasonable judgment and and Amy would you say that that's then an example of um explicit evaluative reasoning when you get to to that point even if you don't actually make an ultimate judgment. We do make judgments. Yeah. Yeah. I think it's it's that question of how and I think that goes to what Kim said first which is the context in which this is happening. Do they want evaluation like do they want you to facilitate a process that arrives at a judgment or do they want something that's much more researchbased and I and there's risk and all kinds of other things are a part of that conversation. So for me I think it's about offering the judgment part as an option on the menu of things that evaluators do uh and can provide and that's different than saying evaluation is the same as research or evaluation measures stuff. You know, if we offer them the chance to say if if we pulled this together, what level of judgment would be helpful and how could that facilitate your decision- making, then that's different than saying, you know, you have to do a judgment all the time because I think that's not always appropriate. Yeah. Like we often do make judgments, but don't we'll ask we'll ask the commissioner whether they are looking for explicit recommendations. Um but the judgments won't ne they'll be based on the measures if you like whatever kinds of measures they are um but we'll make a kind of a reasoned assessment of what whether that's good or not based on our understanding of the policy context and the hardness of the problem I suppose and sectors where we're working all the time. Yeah. And Eva says uh in the chat, couldn't the judgment be done using participatory methods or facilitated process? I think absolutely. And that's like a good evaluative judgment process isn't just the evaluator sitting down with Yeah, I love a rubric, but often I think um clients don't have the appetite or the time for it. Well, and if you're having to build from scratch every time, that impacts it, too. So, yeah. I was going to say yeah no I was going to say that u as an external evaluator um yes um often there is a a desire for findings and some kind of uh you know conclusions or a judgment around those but not necessarily recommendations and I think that's appropriate um in the so far as um recommendations require a a more a broader understanding of the context in which the findings are landing. Um, and often it's only the, you know, the policy makers or the people within a government department or or whatever that have the full breadth of understanding and not just the, you know, of the political, the social, the the environmental elements that need to be taken into account to form good recommendations. Yeah. And I think that's a really good um point that that has come up a lot in discussions about whether whether the evaluator's role, you know, needs to extend to those recommendations or how you provide satisfactory evidence base for um for recommendations to be formulated. Um I think that distinction is one that that that a lot of people are navigating uh frequently in in especially in government uh but but obviously outside government as well. Um and I use that um term navigating. I wonder if if in other rooms or or other participants today had experiences about how you do navigate those challenges um uh collaborative processes being one of them. Um or um other challenges that you found in um in addressing explicit evaluative reasoning in your um your reporting. We talked a bit about um I guess the uh the uh capability or the awareness that the sponsor and the stakeholders have about evaluation and also um you know um I you know in terms of um uh do they have do they understand what evaluation's about? Do they have an evaluative attitude rather than they don't have to have evaluative reasoning but it's also about the role of the evaluator needs to be a translator in a lot of ways in terms of um well you know looking from what uh what does it mean to fully describe but also what does it mean to fully judge um and what do you want in between so we just just covered a bit of that um but the um yeah and there was something else but but I I guess that that context is really important and also a lot of times we're coming in and doing something after the program's actually been running or you know things. So it's it's actually also about how can one actually educate or the you know sort of pe the people making decisions about programs to actually involve evaluators at the front end that you you you know you can actually because they can actually also think a bit about well you know if they haven't actually thought about um what uh program theories around you know around or or behavioral theories around um to you know to say well maybe maybe the program as described isn't fit for purpose. Is it going to do do what you want to do, but how also do you do you think you want to measure it? What do you want to achieve at the the outcome? So, so there's there's a a number of factors here which is not just about sponsors and their awareness um but it's about that educative informative process um and and also I suppose trying to encourage people to think that evaluation front end of a process. Yeah. Yeah. Wonderful. Um uh to hear you hear that contribution. I I I I tend to get known as the person who says evaluation is ECB that you that it's a necessary capability of evaluators to be able to um work with people so that they can make the best use of of evaluation findings. So yeah, I I I I thoroughly agree with that. Um, Amy, feel free to jump in at any time on on any of these responses as well. Um, uh, as as one of our our foremost educators about evaluation. Well, I think I was thinking about the the connection between evaluative synthesis and reasoning and recommendations. So, I think those two things often get grouped together and I do I don't think of them as the same. So, Scriven's written a bunch of stuff about when to make recommendations and when not to and what what kind and I'm at like I'm in the Scriven camp. So, that like that part is that I think what Lee said is really important. But I think the synthesis the evaluative synthesis step is something different because if you've if you've aligned the synthesis process to decisions that have to get made you might not be making recommendations but you might have said look if performance looks like this you know if you're running a training program for instance and half and you're not even getting the right people in the room that directs you without any recommendation being made somebody would look at something and say okay like we need to fix that or or we're just wasting our money. So, I think it the explicitness of that reasoning and connecting it to the decisions or things that need to be paid attention to actually means you don't need recommendations necessarily because people can see what's happening and make choices based on that. So um I think if you're in the space where they don't want that explicit evaluative reasoning or synthesis then the recommendations question becomes a bit more live like what what if anything do you say that you think they might do as a result of what you found in this space? But again that's it's a question of appetite. you know, does the person who you're doing the evaluation for or the community that you're doing it with, do they want that kind of explicit connection uh or not? And and that goes to the risk tolerance, that goes to uh evaluative, what I think of as Jane Davidson calls it evaluative attitude, like what's your openness to understanding if shit's not working? And if you if you don't want to know, then an actual evaluation is not going to be appealing to you at all. So um I think for us as evaluators understanding that appetite from the beginning can be quite helpful because if you keep trying to push reasoning on people who don't want it you will spend a lot of time in misery but if what they want is really just facts then we can do that too because that's you know that's our bread and butter already. So thank you. Um I think um what you've said there Amy is touches on something that we discussed in our group too is about um uh commissioners wanting short reports but I mean just because you are doing very explicit evaluative reasoning doesn't mean it has to be in the report. It can be something that sits behind the evaluation and you can still have a short report if you if you want to. I also think the reason can contribute to short reports like you can put up a a rubric and show how performance happened and that's a page. Someone said to me once that CEOs need cartoons like they need a one-page drawing of things because they're not going to read all the stuff and a a rubric is essentially that. So, if they're open to that kind of thing, it's a way to to give them what they want. I'm I'm going to jump in and say we we've got 3 minutes left um uh before 12:30. And so, even though there's there's great conversation going and there's no reason why we shouldn't keep it going after 1:30. Um I just want to say a couple of things um before we formally wrap up. Um first and foremost um a huge amount of thanks to all of you for um bringing such um engagement and and great thinking to to this session. And of course, thank you to Amy uh for um bringing this deep thinking and research to us. Um and we we pass that um thanks on to Katherine as well. Um and um for um everyone who is in the room with us um right now, we would like to um we'd like you to uh complete our um our feedback form which starts off again appropriately um reconciliation week by asking you which Aboriginal lands you're on today. Um but it'd be really great um if you could um fill in that uh that form for us so that we get some uh information about about what's happening you know with you and and how to make these sessions better. Um watch all the alerts from um AES um about what's coming up next. Um, as I said, uh, perhaps, uh, a little bit unclearly at the start, the, um, New South Wales chapter runs these, um, uh, networking opportunities more or less every month, last Thursday of the month. So, watch out for free, um, seminars and opportunities to get together. Um, often they're later in the day, so that you can, um, socialize, but um, uh, there's a there's a mixture of of of times that they're on. So, keep looking out for those things. Um and uh and and finally um yeah, watch have a look out for conference. Um uh actually we we've got a minute and I think Amy, you might have some information about something that's that's happening at conference and I um I'd really would love people to uh to be able to see that as well. Can you see it now? We can here. Let me go. So there is a question about warrants. I made a Wii slide about that. So if you're looking for and we talked about it, but there's the book. Uh so this will be in the slide pack. Uh if you want to be part of the cheeky cabal of people who are talking about this, you can scan this QR code and it just means you're giving me your uh contact details so I can stay in touch about things. Uh people are distributed like dandelion seeds. So, we're trying to uh get everybody who's interested in this together. And I have no idea what that's going to look like, but if you tell me that you're interested when something happens, then I can tell you. Um I have a few things happening at uh AES25. So, I'm doing a full day workshop with Ben Lawless on rubrics. When we say when Ben as an assessment person says developmental, he means stages of learning, not Michael Quinn Patton's developmental evaluation. So, let me just be clear about that. Uh, and then I they always put me two things on the same day. So, we're talking Katherine and I will be there to talk more about this and uh have more conversation about this at 2 o'clock on Thursday. And then the evaluator competency work with the pathways committees uh will just give you the next installment of what's happening with that at 4. And there's references which will come out with the things. And here's a dandelion which was drawn by the daughter of one of my students. How awesome is that? So, thanks everybody. It's been really lovely. Yeah, thank you Amy and thanks everyone for uh your participation today. Um as Amy said, those things will be in the um the pack um that will the slide pack that we'll make available to people. This presentation, the recording will go up on the AES YouTube site in due course. Um so you'll be able to to see that there. Um and uh it'd be great to have more people involved in this conversation. Um, we really appreciate the fantastic contributions today. Um, and uh, look forward to seeing you at conference and in other AES uh, meetings if there's um, thanks Phil. Bye. Questions. Can people stay back for a few minutes? I just was asking probably Amy and you Phil. Oh yeah, I'm can hang for a bit. That's fine. Yeah. And I can hang for a bit as well. And uh so if you if you would like to uh to pose another question then um this is a great time to make use of one of Australia's foremost evaluation educators. Oh my gosh. Hey Gemma. Hi. Thank you. Thank you for the extra time. Um I am an independent evaluator and we've definitely found that a lot of our clients often in international development are not particularly interested in the logic of evaluation side of our proposals and projects and the the concept of a judgment doesn't sit well with the fact that they seem to be primarily interested in wanting to be told what to do and um you know have something to show for it and I had so I was really interested in your suggestion there of you know having a really explicit discussion to assess their appetite for um you know the explicitness of the reasoning or having judgments. If um the appetite for that was low um you know do you think yet they are intent on this idea of having an evaluation? Do you think we should sort of you know steer them away from the idea that what we're doing is an evaluation into something else or do you have any tips on like how to manage those conversations? Oh gosh. Yeah, that's a big question. I think um if what they want so this goes to what Ben and I are going to talk about in our workshop. If what they want is to what they what to do you know steps to improve then having a set of rubrics that actually says here's what your performance looks like now and here's what the next level up of better performance would look like will actually answer those questions for them. And like Ruth uh said whether or not that goes into a report that like that's the existence of that can be quite helpful. And I think one of the things in international development that can be so frustrating is that this top- down stuff coming from an external international funer that doesn't necessarily match up with on the ground program stuff or even what communities think is important. And so that you know you're kind of you you've entered into a conflict between local uh local expectations and needs and whatever is coming from outside. And so being explicit about that actually can be quite helpful to like make those values visible and show that there's a conflict. And if you're using that sort of uh reasoning step then you can say this will look good to the funer for these reasons but it will be bad for us for these reasons. And it makes that all visible. So I think you can make a case for whether they want judgment or not. You can make a case for saying if we use this kind of thing like I've said so many times to students, I don't care what you call it, let's just get the practice happening. So, call it what you want, you know, say, "Look, let's talk about how we could be developmental about this or whatever that that gets them into that kind of process and thinking uh and and helps to make those values visible if there's a conflict because I think that's the other place where the if you don't do that then usually it's the whoever is paying for the evaluation, it's their values that get used and then that can fundamentally be detrimental actually to the thing that they're meaning to do. So, yeah. Sorry, that was a long answer to your short question. Did it help? No, thank you. I appreciate it. Yeah, thank you. Yeah, but it's a good question. Anybody else have burning things or are you all just thinking about lunch? I have a quick question if yeah we have a moment. Um so we were talking before and I think um it was talked about that a really difficult step can be defining standards when they don't already exist and that it feels like you get wrapped up in you know say it's like it's service delivery it's like I need to define the standards for all service delivery in the world in order to define whether or not this pilot is effective or not. Yeah. And so I think that that's probably the step that I struggle the most with and I was just wondering how do you approach that? Like how do you not get embroiled in the standards of the entire space? Yep. That's a great question. I think one of the things that I see quite often is trying to make too many levels of differentiation. like when sometimes especially at the beginning you need to know if something is harmful and if something is okay. So rather than trying to have like poor you know okay good excellent like you don't need to be that precise especially when you're starting out. So, if you all have a conversation about how will we know if something is going terribly wrong that we need to address immediately versus this is pretty okay, like we can leave this sitting like no harm's being done. It might not be the best thing ever, but it's not actually actively hurting anyone. Uh that might be all you need like to get started. I think it's the over complication of setting standards that gets in people's way sometimes. So that you can use the literature to have a sense of what bad stuff to look for and how to set your standards to make sure that that's captured. And that's what I would suggest instead of trying to be overdeveloped, you know. Um just think about what what do what do you need to stop happening? Like what does the evaluation need to identify? cuz I one of the I've been reading this woman now whose name escapes me, but the book's called Quit and she's a person who does decision theory. She was a pro poker player and then she went did a PhD because why not? Uh it'll come to me in a minute. Anyway, the quitting is the hardest thing for humans to do. We are not wired to do it. Every bias that we have, this is Amy's too long, didn't read, summary of the book, every part of our brains is wired against quitting. And what wasn't already wired against quitting in our sort of fundamental nature has been reinforced against quitting by western culture that says strive, do winners never quit and all that But we cannot as program evaluators like helping people quit stuff is fundamentally important to what we need to do. So setting standards that help people see that something needs to be quit is probably more important than almost any other kind of thing we can do in that step. That's great. Thank you so much. Yeah, I think I think that's um something that I was thinking about too when you were speaking of like um you know why are some of those steps commonly missed and I think sometimes the complexity and the systems thinking level um you know I'm doing an evaluation around community resilience at the moment and it's culturally relative it's complex it's we're talking about thriving communities we're talking about absence of this and that like There's actually a lot there to unpack. Um, and we're doing that process at the moment. But yeah, do you have a a comment around I mean there's a reason when I said this the the joke the comment I don't know what it is the idea about we've been skipping leg day one of my students immediately responded squats are terrible and hard. Why would you do them? And I think this is like the answer here is the same. this value stuff is hard and that's why the measuring not that measuring isn't its own challenges but it's e it's been the thing that's easier to miss and because we aren't building a database of criteria for instance like here's a set of criteria that have been used in similar programs over x amount of time that like that's a missed opportunity we have as a profession to make this stuff easier for ourselves so um if you're a person who has access to literature Rebecca Teasdale is an evaluator in the US who's been doing a bunch of analysis of reports and articles and pulling out sets of criteria that were implicit in those. Um, I have a handout that I pulled together for students and I'll see if they'll let me share it with you guys that just has like here's all the criteria and here's how you might like put them along in the life cycle of a program. um which hopefully will help you too cuz I think the too hard basket is a that's real like it really is too hard to try to do this by ourselves which is also why I'm keen to start a conversation about how can we share about this especially if it can't be explicit in the reporting and the publicly available stuff how can we start to accumulate across reports so we can share out those products of evaluation that we all could use to make our practice better and easier rather RA than trying to everybody do this on the on their own cuz it just can't it won't get done. It's too hard. Yeah. And sometimes it sort of I feel like the evaluation then butts heads with the spirit of the program that's trying to be transformative, not trying to be reductionist. And here we are over here trying to set these standards and this particular way of looking at the world. And it's just um that's another reason why I think some of those don't happen because we can't we can't match this and that sorry I'm just recovering from the flow bit croaky but um yeah sort of how that's why I've put in some of my feedback around how to we do this in a non reductionist way. That's the framing of a lot of the underpinning philosophy and world view that we're coming at it with when the program staff are looking at it in a completely different context and worldview. Yeah. Um I'm I'm saying this out loud because then I'll be on the recording and then hopefully I will remember to share stuff. But we um Mattea Roa did her PhD with us and she made a matrix a values identification matrix to help surface the values of different stakeholders and the headings on that are based on western philos you know like western normative definitions of what good looks like. But you can shamelessly repopulate all the column headings to stuff that matches your context and to and then you can say okay these stakeholders think these things are important. this stakeholder group actually thinks a whole bunch of other things are important and then it's visible so you can have that conversation and say well we don't agree about what's happening here so um that can I think be quite helpful too because it it it takes out that reductionist concern and actually makes the different values that groups have or uh stakeholder groups have visible fantastic thank you yeah for sure thanks for your questions fantastic questions. Um, and and real life challenges. Uh, yeah. Well, and I think that only way this evaluative reasoning stuff will work is if we figure out how to make it viable in practice. So that that to me is this isn't a scholarly exercise on my part. The reason that I want to understand this more and do more Katherine and I want to do more research on it is we'll only be able to do it if we can figure out how how it's been done well who like and what secret squirrel kinds of activities people have done to make it happen. Uh to make those more visible in whatever way people can be comfortable with. uh and to figure out how can we streamline people to be able to to do this work instead of having to build the wheel every time because it just it's not viable. Can I throw Liam under the bus a little bit here because there's uh you know he he worked on designing the evaluation for a big complicated emerging program with teachers at the center as innovators and you know kind of making it usable while you're building the plane you're building the evaluation and you know getting towards that impact. All right then. Thanks Julia. I think u there are parts where you need to really engage with the complexity and there's parts where you can kind of try to minimize it. So, one of the things that I just mentioned in the chat there was, and this reflects, I think something similar to what you were saying, Amy, I was agonizing over building a effectively a rubric for for the evaluation. And I was agonizing over the different levels of the rubric in particular. And um I'm going to pay tribute to our academic partner who was working with us at the time who said why don't you just go with the excellence descriptor and then you can actually um sort of use that to make a judgment against what excellence what excellence actually looks like and that was kind of a real light bulb moment and you know probably saved my sanity as well because it sort of meant that I could actually you know really zoom in on the stuff that we wanted to See, I think as well it's like so much of it probably from my perspective is it's probably different to some people's context because I was an embedded evaluator uh in that context. So, and that's sort of something that I've you know I've been thinking about for a while because that's sort of a model that I've actually worked with since you know probably for about just over a dozen years now. And I think that that model of embeddedness where there is really good authorizing environment where you build relationships with the people who can make the decisions that matter. Um I think that can be that can be really useful as well. So I I'll I'll be presenting with my old boss uh on that program at the AES conference as well. So, and it's sort of a bit of an update of a of an article I did about 10 years ago. So, it's kind of the 10 year anniversary of that. But I think yeah, it it's it's simplify the bits that you can simplify reflect complexity in the bits where you still need to reflect the complexity and just be people's friend as much as you can. That's fantastic. I think that's I think that's fantastic advice and and uh eager to to hear that update. Um obviously that 10-year update. Um also appreciating the shameless plugging of people's uh AES conference papers. Um anybody else who would like to plug uh an activity uh at this point's very welcome to do so. Um, no, but I do think that's really um really valuable from from everybody on this topic about how you realistically face the challenge um that complexity um uh presents and you know as Lachlan was saying you know when when there aren't a set of standards when you're trying to develop that from scratch um what are some real world you know actual examples of of how you can go about this to minimize it while Amy and and the the dandelion seeds are getting together this new resource course um you you got to get something done and and I think that Liam's um uh approach of of using um excellence descriptors as your starting point because you want to know what good really looks like um and then not get um buried in the in the minutia of how good is it you know what's the difference between average and slightly above average you know because that's just not going to be helpful I think that's really worthwhile as Angelina, I saw that you made a comment in the in the chat as well. I did. Um, and just really quickly, I think I'm trying to put my thoughts in order, but I think sometimes we shy away from imposing boundaries on our work because we don't want to be reductive. I'm just calling back to a conversation that you guys had a couple of minutes ago. Um, but I think boundaries are inherently necessary in order to actually complete the work. And it's boundaries along like we impose boundaries on a lot of things, whether it's the type of data that collect that we collect, how much data that we collect. And I think we need to take a step back further and also think about how we're going to impose boundaries on what it is exactly that we're evaluating and how exactly we're going to be judging our success. Because without those boundaries I and without them being clearly articulated, what are we doing? Like we don't even know. So how is our reader going to know? Did I say fully describe um that this is a dimension of fulliness? Um yeah. Well, I think it's fullness not in a sense of like describing absolutely everything, but fullness in the sense of being really explicit about like where that fullness is bound. like what what is going to be fully described. So within these boundaries, we're going to be really full in terms of what we describe, but we're not going to describe absolutely everything. Thank you. I'm starting to get things pop up on my screen saying I have to go to another meeting. Yep, that time of day. Um uh so much as I would love to continue this conversation now, I think I'll just say again we can continue this at conference and I hope at some other um uh um venues as well. Um Amy, you know, we do still have the intention to have um a a panel discussion with a wide audience um when when there's an opportunity to do that um and more of the and closer to the paper being published as well. Um so watch watch all of your um alerts for for information about that um about that panel discussion where we would bring together some some people who who are really struggling or not struggling but succeeding to address these issues across a range of different contexts in both in government and outside government.