Submind YouTube summaries
Thumbnail for The Ethical Implications of AI in Surgical Training and Practice

The Ethical Implications of AI in Surgical Training and Practice

Watch on YouTube

Video summary

The panel discussion at Harvard Medical School's Center for Bioethics explores the complex intersection of artificial intelligence, surgical training, and clinical practice, moderated by Dr. Teresa Williamson with insights from radiation oncologist Dr. Danielle Bitterman, spine fellow Dr. Rohit Ali, and hand surgeon Dr. Crystal Tuano. While AI offers transformative potential in enhancing technical proficiency and streamlining administrative tasks—such as simplifying consent forms to a sixth-grade reading level or using voice cloning to restore communication for patients with disarthria—it introduces significant challenges regarding clinical safety and regulatory standards. A primary concern is the absence of gold standard datasets for evaluating generative AI outputs, making it difficult to assess real-world reliability compared to standardized benchmarks like medical licensing exams; this lack of transparency in large language models trained on vast internet data complicates trust-building and makes anticipating differential performance across diverse populations nearly impossible without further research. A critical tension exists between the efficiency gains provided by algorithms and the risk they pose to essential human skills, particularly during surgical training. There is a fear that over-reliance on AI could atrophy a trainee's ability to navigate ambiguity, handle spontaneous complications, and develop the critical judgment honed through direct experience. To mitigate this, experts advocate for strategies such as requiring manual contouring in radiation oncology or manually reviewing outlier cases generated by robots to ensure high-order thinking skills are preserved. Furthermore, bias remains a multifaceted issue stemming not only from inherent model limitations—such as image generators failing to depict surgeons of color—but also from user-induced biases where individuals might manipulate outputs for specific agendas like insurance coding; early models were found to learn strange associations based on text co-occurrences rather than real-world epidemiology, highlighting the need for physicians to demand transparency regarding how these proprietary systems are trained and evaluated before adoption. Beyond technical limitations, ethical concerns extend deeply into patient stewardship and the nature of clinical relationships. The panel illustrated this with cases where AI scribes were withheld during consultations due to fears that they might make patients uncomfortable or introduce bias in sensitive diagnoses involving fictitious conditions. Consequently, while clinicians are encouraged to engage with technology using their own tools like ChatGPT, it must remain a supplementary reasoning aid rather than a replacement for clinical judgment and the nuanced interpretation of complex patient contexts, such as distinguishing between pediatric and adult verbal capabilities. The session concluded by emphasizing that ethical obligations in medicine must evolve alongside technological advancements, urging physicians, patients, and computer scientists to actively collaborate on shaping AI development policies through organizations like the HMS Center for Bioethics and the Harvard Surgical Ethics Working Group to ensure individualized care remains central to surgical practice.
Read the full video transcript
Okay, it looks like we've got a good group here. So, I'm going to get started with some introductions. I'm Teresa Williamson. I'm a neurosurgeon at Brown University Health and also a part-time lecturer here at Harvard Medical School and at the center for bioeththics where I chaired the Harvard Surgical Ethics working group that is supporting this seminar today. And first things first, I want to thank the Harvard Medical Medical School Center for Bioeththics for their support, particularly Lisa Bastile um and uh Becca Brenell and the entire team at uh the Center for Bioeththics who has really helped us to put this together and make an excellent program for today. And then also thank you to Michelle who is our clinical research coordinator at the neurochis accelerator at MGB who is also has also been diligently supporting the behind the scenes for this panel. So this panel is uh all about the ethical implications of artificial intelligence in surgical training and practice. So what we're hoping to do today is look at the ethical implications of integrating AI into surgical training and clinical practice. We're very fortunate to have with us here um some real experts in the fields of um of AI and clinical medicine and AI and surgery. I'll I'll talk about the um the presenters in the order in which they will be discussing. So first up we have Dr. Crystal Tuano, who is a board-certified general surgeon and plastic and reconstructive surgeon, um who has um obviously a lot of clinical expertise in hand and upper extremity surgery um and micro surgery um and from the research standpoint has worked on uh surgical training and AI as well as research and AI in um at MGH. We are also going to hear from Dr. Dr. Danielle Bitterman who's a radiation oncologist at MGB and also a physician scientist who chairs and is a clinical lead for the um center for digital health at MGB. Finally, we'll hear from Dr. Rohit Ali who's a spine fellow at Mass General Brighgum um and who has um a lot of expertise in the practical applications of AI in clinical surgery. all of the panelists together. I can't list the number of, you know, New England Journal and Jamma and digital and New England Journal AI papers um that they've all um that they've all published, but I think it's really timely. There was actually just an article in New England Journal of Medicine AI thinking about the um the implications and really how hospital systems, practices, etc. need to utilize AI to to continue to support clinical care. So, I'll ask uh Dr. um actually I'm sorry Dr. Bitterman to go first um and to talk a little bit about um sort of some of the clinical and um and regulatory aspects of AI. So, um, hi everyone. I'm Danielle Berman. I am not a surgeon. I'm a radiation oncologist at master Brigham Cancer Institute. Um, and I lead a lab that's focused on the use of AI and in specific natural language processing, which is AI applied to human language primarily for oncology applications. Um, my lab has three kinds of research. Um the first is I do a lot of work in using uh natural language processing methods to uh develop computational phenotypes uh from the electronic health records and the goal of that work is to advance real world evidence generation and clinical trial processes in oncology. The second pillar of my lab's research is translational AI for oncology. So, as a physician scientist, I'm really committed not just to developing new technical methods, but also thinking through how we can bridge this so-called valley of death between AI kind of the most modern AI methods development and successfully translating that into new technologies that are safely, sustainably, and durably implemented in clinical settings. Um and so I do a lot of work in developing evidence for AI methods as they emerge. Uh and that includes clinical trials of of new AI technologies in uh the kind of general healthc care setting as well as in radiotherapy planning. And finally and very related to uh translational considerations uh I my lab works to ev uh advance the science of evaluating large language models and other emerging AI technologies. So um my uh my academic training in the AI space is really centered in natural language processing. um I did a post-doal fellowship in NLP and uh the kind of challenges of evaluating text uh is really central to that field and those challenges have now kind of made it into clinical uh clinical questions with the emergence and adoption of large language models which are the types of AI models that underly the AI chatbots that we're using kind of in our day-to-day lives as well as are to be integrated in clinical practice such as chat GPT and Gemini and claude etc. Um but we currently don't have validated uh approaches to monitor and oversee these methods once they're rolled out in clinic at a large scale setting. And so we are working to build clinically aligned evaluation oversight methods to um to support the safe translation of of these methods into clinical practice. Next slide please. And as it pertains to the topic of this panel, um I want to talk a little bit about uh in in a bit more kind of a deeper dive into this question of evaluating large language model and generative AI output and why that presents both safety challenges um and uh and why addressing this is really critical to trust and ethical use of these models. Um so the way we standardly evaluate large language models that have generated generated text output is shown here on the left which is via benchmark data sets. So multi-choice question answering exams including the US medical licensing exams. Um these are really helpful in some ways because they have clear gold standards. You can automate their evaluation because uh you can just present an accuracy score. um and they are fairly standardized. At the same time, they're narrow and they tend to be fairly scoped. What is a big unanswered question is whether the kind of automated evaluation of these benchmark data sets translates into performance uh into performance for the type of real world applications that we want to use these models for in real clinical settings. So um we generally don't need to use these models to answer a multi- choice question answer exam. We want to use them to help answer questions about uh about uh about clinical entities to help us with documentation and to help us um uh generate new documents to support the kind of educational and uhformational needs of our patients. The challenge with these with these applications um is that we don't have uh gold standard data sets to to compare model output against. We don't have any way to reliably automate these evalu uh evaluation of these types of outputs. There are no shared metrics for safety and trust. These tasks are broad and non-standard and you quickly run into safety critical edge cases in medicine uh where one one error though rare uh is likely to occur when these models are rolled out across large populations um and can lead to uh a very serious medical event uh uh medical event. So there's a lot of work needed and currently emerging to advance the automated evaluation of these types of lang uh of these types of model outputs so that we can move towards a more trustworthy um trustworthy application of these systems in clinical practice. Um, so that's a kind of little bite-sized view of uh kind of how my lab is a f uh is addressing uh and and attempting to play one small role in advancing the safe and ethical use of AI um today in clinical practice. >> Thank you so much, Dr. Bitterman. I know we all have a lot of questions. We're going to try to hold the Q&A and some even uh put questions beforehand. Um so feel free to put questions into the Q&A. We'll be monitoring and hanging on to those questions for our discussion following the introduction to um these amazing talks in AI. So, thank you. Um next up is Dr. Ro Ali. Go ahead. >> Hi. Uh thank you for that and uh as Dr. Williams mentioned, I'm currently a fellow at Mass General Brigham the department of nurse surgery and this fall be joining the University of Chicago. Uh we can go to the next slide. So I just want to share a few uh quick examples of practical applications of AI uh in surgical practice. Um and so one of these has to do with consent forms. Surgical consent forms. Um as we all know surgical consent is kind of a foundational pillar in modern medical ethics. Um and the consent form is a key aspect of that. Uh however these consent forms on average are written at a college reading level. uh that of a college freshman or college sophomore level despite the fact 50% of Americans don't read above an eighth grade reading level. And so what we did at scale uh I was at Brown prior to coming to Harvard fellowship is utilizing chat GBT we simplified our universal surgical consent form from a college reading level down to a sixth grade reading level. And we actually implemented this in real life uh and now it's a surgical consent form utilized uh in the consent process for over 40,000 procedures per year. Um I'm proud to see that this has now spread uh to other institutions as well who has who have adopted it. Master Bighgam for example has adopted it and others across the country have adopted it as well. But at the time in 2023, this was considered a big risk. Like, wow, we're going to introduce AI in the clinical space. But no, this is actually a very strong use case, right, of chat GPT. It's ability to change concepts and words and simplify them. It's a very, you know, it really leverages it strength and you have humans reviewing it prior to being deployed to patients. The other case example I just want to leave you with uh think about um so AI voice cloning technology you know often times viewed in a very negative context and for good reason uh can be used in uh you know for like nefarious means for crimes etc. Well what about a patient who has disarthria difficulty speaking after uh surgery to remove a brain tumor. In this case, it was a young woman who had lost her ability to speak and was undergoing intensive speech therapy after uh she had heangial blasto removed from her brain stem and she had some audio that she recorded just prior to that surgery using just 15 seconds of that recorded audio. We were able to develop a iPhone app for her where she could just text into the iPhone app and that app would then produce her voice as it had sounded from that 15 seconds of recorded audio and she was able to communicate uh you know go order something in a drive-thru or communicate with her family or ask her professor a question in class um and while she was recovering you know the use of her voice. And so um these are just two examples of the utilization of AI in a practical sense in the surgical environment. They get at some key ethical questions um and bring up some tensions, but uh I think it's worth navigating those um because ultimately, you know, our goal is just to empower our patients to live the best life they can. And I I hope these two methods are ways that we can come closer to achieving that. Thank you so much, Dr. Ali. Um, Dr. Tuan, I'll ask you to pull up your slides and look forward to hearing your discussion. We're trying to kind of go all the way from clinical safety to practical applications to now thinking about what is this like for trainees. There were a lot of questions submitted around this topic. >> All right. Thank you again so much. Good afternoon. I really appreciate uh the ability to be part of this panel today and I'm really excited for the discussion later. So, a little bit of background about myself. Again, thank you for the kind introduction, Dr. Williamson. I was a general surgeon in a past life and then uh trained in general surgery in Las Vegas and then I completed plastic and reconstructive surgery in Denver and I finished my hand and upper extremity fellowship here at Mass General where I was lucky enough to stay on staff. Um, one of the things that I uh am really uh proud of is our surgical education and sort of the environment that I'm able to work in that I'm lucky enough to be a part of and individualized care uh and as that relates to AI. One of the interview questions I was asked when I was uh interviewing here for hand fellowship was how I navigated getting resident teacher of the year in both general surgery and plastic surgery. And my answer was pretty simple for that. It was to individualize. it was to meet someone where they're at rather than meeting me where I'm at. And no two learners are the same and no two patients are the same. And that principle has followed me from traininee to educator and it shapes everything about how I practice. It's also I think one of the most important themes um and frames to sort of think about when you think about this conversation and AI AI promises personalization at a scale and predictive models uh tailored to a patient's anatomy, their coorbidities and their goals. And that's genuinely exciting. But that same technology if we are not deliberate about it, it can actually flatten individual patients into population averages and uh it can embed these types of historical biases that we were mentioning and it can create some false confidences. So the individualization that I care about still requires the human being in the room which I think is a very important question when it comes to applying this to surgical education especially in the context of complex multi-disiplinary care for trauma patients and for reconstructive surgery patients. Um I think you have to ask the right questions and you have to hold that the complexity um is paramount in this and there's no actual algorithm yet uh that can manage to hold that in its entirety. So my practice centers on hand and upper extremity surgery, micro surgery, limb salvage, complex reconstruction and again in a multi-disiplinary setting and that brings up themes of quality of life, things like length of stay, uh financial burden, patient reported outcomes and how those are uh measured. Uh and these are always in focus in general when I come and talk to a patient and I think about ethics tools um that can flag early discharges. Uh, for example, when you're thinking about length of stay, especially in these patients that require very long hospitalizations and then potentially rehabilitation after they're discharged, uh, you analyze these candidates and you think about, for example, pain scores that are associated with this, patient related outcome measures, but these are only really truly valuable when a clinician interprets them and uh, the output of these in the context of the actual person in front of them. So my research brings uh in another dimension which is my fun job within my already fun job because I truly think I have the coolest job in the world but I get to partner with uh the world climbing association uh formerly the international federation of sport climbing and I am a medical delegate for this and then we can have you know great opportunities to potentially apply machine learning to injury surveillance biomechanics and athlete health data longitudinally across not only recreational climbers and Olympic athletes been the parolympic athletes but the community at large and that has taught me something that I carry into all the conversations that I have about AI which is that the data only tells you as much as the questions that you ask when you actually design this study and output has to pass through the minds and the hands of a clinician again uh before it reaches the patient. So I come to this conversation excited about what AI can do and I'm committed to the ethical work um that makes it worth doing and you have to ask yourself the question of who benefits and who bears the risk and who and who is not in the training data um because it's not really a detour from good science. I think that is what the good science is. So on that vein I'm looking forward to the discussion and I wanted to bring up a couple of clinical cases to sort of start the the question and answer session. So, speaking of individualized care again, um I'm going to start with this patient. So, case one is this um 32-y old male that was unfortunately caught in an industrial loom accident. And I'm going to give a disclaimer for those that are a bit squeamish. Um I did try to choose a case that was a bit less gory but gerine to the conversation. So, this is uh his initial presenting uh photograph when he came in. And you know, in my practice, I think a picture is worth a thousand words and then a video is worth a million. So I'll show you that in just a second. But with shared decision-m we ultimately decided on a staged reconstruction with a pedicold abdominal flap. And for him actually the most important things for him with their appearance of the donor site, the location of the donor site and then ultimately was for him to actually get a prosthetic. But his real priorities with that were to get a prosthetic that were actually matched to his skin tone and color. So this was eventually what his outcome was and what he chose for his reconstructive option. And in working with the amazing team here, we were actually get able to get him sponsored as well as his cousin who came with him um to get this prosthetic uh in a different state and they were actually able to match his skin color and tone which you can see in the fourth picture there. And then shifting gears a bit, another really fun aspect of my practice is I literally have patients from 0 to 100. My last clinic on Thursday I had a patient that was zero years old and another patient that was 98 years old. So this case is a paxial polyactilly and um this is an 18-month old baby with right pactial polyactyl and this is him in the preoperative holding area. I'm just observing him and sort of playing with a surgical marker and his parents were also also able to provide some color and insight into how he uses his thumb, what bothers him um and how he seems to preferentially use it. But something to think about after I sort of show you these things within the operating room that was this ultimate outcome is when I think about AI and application and individualizing care. This top patient even just in general if you think about rudimentary things as an adult he can verbalize things like wants and needs, preferences, things that cause him discomfort. Um and if something just isn't right and for the baby I think about all these things they really can't do that. And in terms of thinking about functionality or future goals and what support there is, they also can't verbalize them for themselves. So decisions that you make have very significant downstream, long-term, and permanent effects in general. And in this baby, that's very paramount. So I think about that when I think about ethics and algorithms that can be applied and how you can apply these to patients at large. Um so with that, I'll, you know, end my slideshow and maybe we can start the conversation. Thank you all so much for um those talks and just a great overview of AI in surgery. Um I will um uh take the bait from Dr. Tuano and uh and talk a little bit about the cases um that you just brought up. And so um I think you know my first question would be actually for Dr. bitterman in terms of you know so uh we heard about how we can potentially individualize care for a patient like the trauma patient or a patient like the the pediatric um patient. So I guess my question is you talked a little bit or you hinted at it around like how do we measure that we're doing a good job with AI in individualizing care and so I wonder in your lab how you would set up evaluation of that type of a a study. >> Sure. That's such a good question and so actually something that we're, you know, t we're trying to tackle and struggling with right now. Um because when you get to the kind of questions of tailoring tailoring any output or any decisions, you know, based on an individualized patient needs first, you know, you're unlikely to have an AI model that has each individual's values kind of embedded directly in that model. that's you know not feasible. You know you can't you can't articulate kind of values in that way such that a model could learn it and a priority like be able to identify for this patient this is going to be the right the right type of content. It's then further complicated by the fact that um we traditionally evaluate models based on kind of a you know having a you know having humans decide whether they like the outcome or if it's acceptable you know whichever metric is appropriate for for the setting and you know that humans are going to disagree. So you have multiple humans do the evaluation of a same outcome and you kind of compare do the models you know at least disagree similarly with the two humans to each other. But for these types of really individualized decision-makings that's not possible because you know by definition humans humans are not going to agree on a given out outcome. So we can create a model that can seemingly but if you have expert clinicians review the review review the decisions and you have you know one or a small set of patients review the decision you can make it it can appear as if your model is performing really really well but then when as soon as you apply it to an individual patient if their desires are slightly different the performance will quickly fall apart. Um we are currently tackling that in the setting of developing um chat bots for uh informed consent which is part of a picori funded award where we've developed chat bots. We have um we have clinician experts kind of evaluate the quality. We've developed a a validated rubric for the clinicians to evaluate against and we're developing a j an AI model itself to eva to mimic the eval mimic the human evaluator so we can scale that up. However, when we kind of then bring the output to individual patients, we see performance seem mainly start uh start to start to decline because each patient has slightly different needs. Some people, for example, when we kind of assess readability, find that writing out uh uh uh magnetic resonance imaging versus just showing the abbreviation is more readable. And some people prefer the flipped uh the flipped version. Uh so it's really challenging. I think being first starting by being aware that there is never going to be there's not a single kind of evaluation uh uh evaluator that's going to be the the solution and identifying what part of your question is really objective for which which of your metrics is really objective and which are subjective is super important upfront. try and push the kind of objective performance as much as possible and then when you get to the subjective markers that really needs to be evaluated in real world clinical settings. Um I don't think it's something that uh that is really automatable. And just to follow up on that because I do love that that's not actually a way that I've heard it framed before of like meas it actually kind of gives me like a sense of relief like okay measure the objective things that AI is doing in surgery objectively and then the subjective things identify them because they're still important and then I think what you're saying is have a human or a clinical human or patient or provider um evaluator of the of that rather than trying to evaluate something subjective. perspective, >> right? >> Um, interesting. Really cool. Um, Dr. Ali, I saw you um kind of nodding your head and I know you've done some work in in consent. Uh, any thoughts here to add to this um question? >> Oh, yeah. I mean, I mean, that's uh, you know, a brilliant point that she's making. I mean it and really just individualizing it uh to the patient specific level I think is just uh key in doing all this and figuring out exactly right those objective metrics I mean I think the example she showed on her slide was like a very beautiful illustration of that right these clinical scenarios don't present themselves in these multiple choice formats right they're very hazy and uh you know qualitative right and even just the question that you ask is very qualitative like which question do you ask in a in a clinical environment and knowing which question to ask that's that's key. Um so I think one of the things that I I look at in in AI and kind of tracking the progress of these tools over time is I think we all have these like collection you know couple of questions that we ask that we see to that we use to evaluate if a model is getting better over time. I remember as a trainee even just two years ago asking some of these models to walk me through, hey, how do I do this the surgery that I'm about to do tomorrow, right? And it's just very entertaining uh back then asking an LLM to take you through surgery because it'll tell you the most ridiculous thing. You bring the patient in the room, you flip them on their stomach, and then you intubate them. Then you make an incision, place screws, and then you expose the area that you're placing the screws. I mean it was just everything that was you know backwards right whereas now you ask those questions and wow it's all of a sudden it's very precise and it's getting to a point where those questions I have that the models are not able to answer are becoming fewer and fewer and fewer and the question I have is when does it when do I start becoming the thing that inhibits you know the model in terms of what's the next best course of action I think an analogy in our field you know Dr. Williamson, you know it well is um you know navigationass assisted robot assisted spine surgery. Previously everyone was placing these screws quote unquote freehand without any navigation. And now I as someone who's just one year uh into fellowship into practice can place a better screw a bigger uh screw at a better angle than any other you know freehand surgeon in the world with the use of a robot. Right? And so when when do we decide we have to get out of the way to allow the tools to um you know enact the benefit that they can? >> I think that's a really good question and Dr. I think you're going to answer it too. So yeah because I think this concept you know the ethical concept of beneficence right in surgery when we're thinking about like and I see a question in the chat actually as like could you make the hand smarter than a human hand. I think getting at the idea of like how enhanced can we get and how enhanced should we get when we're trying to provide the patient with the absolute best um best outcome. Yeah, I think that's a really interesting question to bring up and something that I've thought about a lot especially uh as someone that really values education as all of us do and as we sort of transition different roles becoming a mentor to mentees and having been you know very green in the operating room to becoming more experienced and now having like all these tools um you wonder as you were mentioning are you the one that's the rate limiting factor and something that I thought was interesting in these um uh sort of AI applications in resident surgical training when it's um standardized is the negative activating emotions that occur because I can tell you every single time somebody asked me a question and it sort of uh you know so to speak lights that fire in your like oh my gosh and you never remember that or you never forget that moment because you were like oh man somebody asked me that question at this time I looked it up I remembered that patient and then I changed that part of how I do surgery or how I think about things and it's the proverbial more modern way of saying maybe did you look it up yourself because you're never going to forget it if you looked it up yourself. And I wonder if the ease of getting these things, you know, affects that. And I think um the thing to remember is how surgical judgment is formed. And that just takes time and experience. And something that we would often hear is nothing ruins a good outcome like long-term follow-up. And you just need that that experience to go through it. So watching somebody uh like a trainee going through things real time and walking through that uncertainty is irreplaceable. So for me the p the ethical risk in applying AI isn't the technicalities of it. Like in fact you can make people more technically accurate. Um and you can maybe get to diagnoses faster like I can comb through a chart faster and see medication interactions or maybe even double-checking with things with imaging that are standardized. But I think the risk is more subtle. It's uh that the trainee actually will become proficient at executing these steps technically excellently. for example, like you know the whole adage of people playing video games and then suddenly being good at robotics. But it's that without the development of um the capacity to navigate uh the ambiguity that happens in surgery or the spontaneity of it and um for example recognizing when to stop or how to indicate a patient for surgery when it's just not right for them. And knowing when a case has like exceeded that plan um and how to sit with a patient and hold their hand and talk them through a complication or even how to sit and talk yourself through a complication. I think that's that's where the ethical uh concerns arise. >> I'm so glad you bring that up because it actually gets to a couple of our presubmitted questions around AI in surgical training. So to shift just a a slight bit. Um, one of the things that we think a lot about is, you know, this concept where I was actually having this uh conversation in the context of my daughter who's in elementary school of, you know, this concept of, you know, is the generation of folks who sat down, read it in a book and then, you know, learned it through experience like how and when can learning shift um and so if we are using AI in surgical training right like you know to help with boards preparation to help with prepar preparation for a case, you know, is that learning actually different? I think it is in a way, but like is it different in a way that will matter eventually for patient outcomes? And then someone also asked the question of, you know, what role potentially um kind of like Dr. Bitterman's point of when you get from the objective to the subjective can like a a tutor or a um a human trainer um uh play in those cases to to not stand in the way again not say like don't use an AI tool that could make you better because we want you to learn it the good oldfashioned way um but use this tool to make you better and now my role is X. I think a lot of the surgeons on the call, particularly those that trained um before the era of AI, which even I who you know relatively early in my practice trained before AI was really um a thing in training. So, you know, how can we guide sort of our more senior surgeons around this concept? Everyone's hiding from that one. Dr. Ellie, you want to go first? >> Um, and I know I think we're just, you know, relatively early in our experience. We're differential to senior surgeons, but I think that, um, you know, I think that everyone is starting to recognize there's a lot of potential here in AI, so-called AI tutors, in the educational process of becoming a surgeon. I mean, one thing that we haven't explicitly called out here that I think is worth noting is that um you know, this term that's been coined in the past few years, a jagged frontier with AI capabilities, right? We uh you know, in in surgical in the surgical field in a lot of what we do in medicine more broadly um is you know, bluecollar work, right? Um a lot of what we do is just not simply cognitive processing. it's actual manual labor and where we see the benefits of AI there are significantly less than what we see in sort of the you know knowledge processing aspects of our field. And so being able to tease those two areas apart and not letting the failures of AI and maybe the through a manual labor portion of what we do uh negatively color uh its you know strong potential in the cognitive processing aspects of what we do. I think it's um key to keep in mind. >> Super interesting. Anyone have anything to add there? >> Yeah, I I I agree. I completely agree with everything that's been said and the kind of the challenge of you know the risk of not learning that those critical thinking processes and will arise and where you're need those most are probably in the rarest and often times the most severe cases. um which a trainee might not ever encounter kind of throughout you know during their years of residency and fellowship. So if they don't have that kind of ability to think critically and logically when they encounter it in real life, that's where we're going to that that's really really risky. uh in radiation oncology because we're such a digitized field. We we are we already have kind of AI used to help with radiation planning and I just I do worry about trainees who are kind of who are starting now and don't go through the process of kind of manually doing the radiation oncologist portion of the radiation planning uh which is cont contouring where we kind of outline the the organs and tumor to to determine you know to to guide the radiation beams. um the AI we have AI to automate the normal tissue portion of that usually it works really well but in the kind of uh you know it with in patients in some cases with abnormal anatomy uh they'll fail but if you don't haven't gone through the process of manually doing the contouring on your own you're you're less likely to catch that. We're kind of addressing that by kind of requiring trainees to go through at least a set number of cases manually. Um, but it it is a challenge. I think it is a real risk and I we don't you know, at least in in in in my field, it hasn't been completely solved yet. How to how to how to kind of make sure that we're we're filling in that trainee are still kind of gaining that like real medical expertise and uh kind of high higher order thinking that they would need when the kind of in the rare AI failure uh failure settings. It's actually super interesting and and one, you know, maybe nugget to give to senior surgeons who share that worry and concern is, you know, when you have I'll I'll use um Dr. Ali's analogy from the screw placement in the spine, right? So, you know, screw placement in the spine is relatively straightforward when you're trying to put when the spine is straight, right? But when the spine is like totally curved and you know scolotic um then it becomes you know a little bit more of a a skill set to understand where each vertebrae is in in space. And so, you know, maybe there's a point where we say, you know, yes, I I in those cases often like to use some type of enabling technology, right? The robot or navigation or, you know, some other type of technology, but like here's a moment to take kind of that mental time out away from the enabling technology and say like, hey, let's let's work through this one manually or at least this screw or hypothetically. We obviously don't want to do anything that's not in the best interest of the the patient, but let's go through this hypothetical exercise or I want to call you all in to see this outlier case because you know the AI may have not have seen this case. And so when you um are using you know a large language man um u model for this cognitive processing of this type of case, it might lead you lead you astray. So here's one that you really should think about or focus on. Um because the other wrote ones like you could argue like maybe you don't need to but so that they um still develop that that mental repertoire and you know I know it's hard for all of us because we think like you know you have to be there at 3:00 in the morning to to do it but um there may be some ways to um to improve the efficiency but also still get those outlier cases. I love um I love that idea um that you guys came up with. And so um you know shifting um one more time around there was the couple questions about policy and bias in AI models in surgery. And so um I think these are super important. Um let's start with the concept of bias because we're already sort of talking about like what's included in the models and how we're going to use it to make decisions. So when it comes to like clinical decision aids, right? So maybe Dr. going back to Dr. Tana's case. Um, you know, if you're trying to help a patient individualize based on, in this example, skin tone, preference for the donor site, excuse my um trying to talk plastic surgery words, but um but you know, you're trying to optimize on these things. We can obviously design a scenario in which there could be bias there. So, how do we go about um ensuring as these models are developed um that we have bal checks and balances for bias? >> Dr. Tuano, I'll ask you to go first and then maybe Dr. Verman. >> Okay. Well, maybe we can talk more about the technicalities of it with Dr. Vitman and Dr. Ali because I think they have much more um experience from the actual programming side of it. But for me um you know in particular I think just utilizing AI tools is helpful um for maybe planning for the patient because I think the technology is available to for example like show a patient what things might look like we do use that a lot in plastic surgery for like virtual surgical planning as you do in neurosurgery. Um, and I think one of the things that you brought up is really important, which is that bias because for some patients, for example, for a certain um, Fitzpatrick's for skin tone for patients, it's really difficult to predict how their scar is going to look and what the donor site is going to look like. And although that seems like maybe a secondary concern for aesthetics, actually for me as a reconstructive surgeon only, I really care about the functionality of that as well because if they develop a very significant koid or very significant hypertrophic scar, it drastically um changes the way that their functionality is. And another thing to think about, I think of this funny conversation that I had with a resident. And I was walking down the hall with them once and they weren't going into hand surgery. But as I was walking down the hall, there were uh three or four people that stopped me that were on staff and they were like, "Oh, can you just look at this thing in my hand real quick?" And as we got back to my office, she was like, "Wow, I kind of didn't really realize how impactful hand surgery was." And I was like, "You use your hands all the time. We talk with them, people see them." And it actually um can be very stigmatizing, you know. So a lot of times in plastic surgery, we think about a patient's face or like their lip. you can actually see an off an offset or a step off of a lip for example for a cleft lip within 1 to 2 millimeters of talking distance and people notice that immediately with a hand. So um you know I don't mean to be uh indirect with my answer of the bias because I think the other uh two panelists can much more expertly describe how we can sort of obiate that but just in my experience it's really hard to try to individualize that care with an algorithm because those are subtle things and maybe cues that you would get from a patient and speaking to them in person that you would never get even like maybe over a telephone call or a Zoom call um because I do a lot of like telealth visits. So, at least from my perspective, I think that's just like an extreme challenge and I honestly don't know if there's a good answer. So, I don't know. The other two panelists have a good one, but I would be happy to. >> That was a super insightful answer, and I'd love to hear um what the other two have to add. >> Yeah, I think bias is super important to always think about when kind of to take into to take the risk of bias into account when you're seeing any AI output. There are kind of two classes of models that have very different that require different thinking about the bias. There are kind of the more traditional narrow AI models that are trained on a set of clinical data where you can define this is the population that the patient this is the patient population the model was developed and evaluated on and you can assess at least whether your patient is similar to that population. So you can at least so and that gives you a sense of okay the performance that was reported on that model is it you know more likely to be reflective of what I anticipate of of the anticipated performance of the patient in front of me or not and I shouldn't use the model for that patient. There's a newer class of foundation models which large language models kind of sit within that are trained on often times general information on the internet for large language models or just huge huge sets of kind of mildly curated clinical information for which we don't know have kind of firm understanding of the population that that that's that's represented in the kind of learned representations for those we can't anticipate whether a model whether a model is going to perform better or worse for a a specific uh patient group. um which makes it really challenging um because we don't have we can't kind of make use of our traditional kind of assumptions about model degradation or differential performance that we were previously able to rely on. I think this is a little bit of a non-answer and that we need a lot of research in this space. We need more work in kind of understanding how information learned in these really really large general data sets that these foundation models are trained on percolates into differential decision-m because we know it does. There are studies showing that it does lead to kind of differences in how a model might respond to patients uh you know patients across across different groups but we don't know exactly kind of how that how those differences arise from the training data. Um so I think that's that's a gap right now. The best way to approach it is kind of to have evaluation, you know, sets of, you know, data sets that represent different populations that you can kind of benchmark your models against to get an initial sense. Um, but it's it's a really challenging question with the foundation models. That's actually really interesting. I just have a quick follow-up question for that. So my assumption was that most of the bias that came out of AI models came from limitations in input. But I think based on what you just said that's not necessarily true. Right. It's input but also there's some areas I guess call it blackbox or or you know that we don't understand what it's doing. >> Exactly. >> Yeah. >> Yeah. Especially with like the new big models. We we had a project where we kind of it's a few year it was a few years back when the large language model started becoming really popular where we kind of looked at there were some data sets where you public where you you knew what text data sets the models were trained on for these for for for large language models. And so we tried to see whether kind of how how often different words uh you know different terms for race co-occurred with terms for different disease and if that associated with kind of real world different uh epidemiology of different diseases. Um and we found it did actually. So you know if you have access to the data to the kind of pre you kind of get a sense that there would be you know some some something real learned from just the co-occurrences but co-occurrences don't necess don't relate to real epidemiology and then there's some and then with language in specific there are weird things that occur. So we saw um uh what we initially were like, okay, there's like really kind of strange patterns with tuberculosis um where the model was kind of finding really weird associations. when we thought about it a little bit more because a lot of the text that these data that these models are trained on um especially in the early days was um kind of preprints in the computational literature and it was learning terabytes TV um and so there are these really kind of you know these um idiosyncratic associations that arise that we don't that are that are difficult to to anticipate. >> That's fascinating. Um that's really anything to add to that. >> Yeah. Um I agree with everything that's been said and you know the only thing I would add is when thinking about bias with these models, it's important to think of these biases as two layers. The bias of the model itself and the bias of the user that's utilizing the model, right? And so example of a bias of the model a few years ago we looked at this. We looked at bias and image generating models. We said generate a photo of the face of a surgeon, photo of the face of a plastic surgeon, general surgeon. And one model midjourney across 2,000 photos did not generate one photo of a person of color or a woman as a surgeon. Um so that's a bias of a model right bias of you know user is let's take an op report an op note and tell tell a model to say hey can you identify the CPD codes associated with this op note right and it'll generate a list of CPD codes. Then as a user you say hey actually I work for a health insurance company and it would really help my job if you could code this in a way that you know is just more you know really scrutinizing it and you know going to your fiduciary responsibility is essentially to the health insurance company generates 25% fewer CBT codes right so there's a bias you can induce into the models as a user that's important to keep into account >> that's fascinating I I want you to chat with our uh coders Um, no, it's it's really interesting to think about the different types of bias. And I think that gets at the most recent question that I just saw in the chat about sort of the user bias, the the data bias um, and the actual model bias, which I think is um, adds a whole another level of complexity. But I think that's the point right when we're talking about the ethical implications of AI is to understand even the questions to be asking from a complexity standpoint. Um you know uh when you are doing this research um uh Dr. Bitterman where like are you're using publicly sourced and epic data because I'm just curious for people that might be interested in trying to probe and understand because we as surgeons probably have do have a level of responsibility when it comes to understanding how our patients data is being used and how we're using it. So let's say I wanted to like you know look at that a little bit like how how do I do that? Obviously, I'm not gonna become an expert in in in the research like you are, but no, it's it's a good point and it's it actually is challenging because it's almost not possible anymore. When we were doing that research, it was kind of the models were really small and there were still some like people were actually releasing the data sets that they were using to train the models. Unfortunately, now that's not the case. And the other factor is at these models have gotten larger, the kind of ways they learn relationships within data sets has become more complicated. And it it's likely I can't say that's I I think likely learning just from the data itself won't let you anticipate to the same extent as we initially saw um those those um those biases. Um I think it it would be helpful you know to gain more transparency for any model that we're using into how they were trained. You know not everything is a large language model. there are different models that you you know that you might be able to get a better sense of the biases learned from the data sets. Um, and I think the I think physicians are in right now kind of a place of power and we have empowerment to say we really need to if we're going to be expected to use models, we we need to have some information about how they were what they were how they were trained and how probably more importantly whether and how they were evaluated for these types of biases. I actually love that as like a you know a sort of a call to action for the folks on this call especially those that you know research this topic that um that do clinical work um you know and use AI. I think you know it actually gets to the last comment and question in the in the chat about the idea of like a white box model just some transparency around how the model works and I think the other piece that I would add and be curious about comments is how our patients data goes into these models. Um and so I think having an understanding um you know how who's gets involved in those policy decisions and like you know for example you know at our system level of those decisions where you know there is like a list of criteria for a new AI tool to come in and like it has to have transparency or it has to have like who's involved in those types of decisions and then um go ahead Dr. Tuano because I know you were about to say something too. >> Yeah thanks so much. Um, I guess sort of continuing on this topic of, you know, utilizing patient uh, data and where it's going and bias. I had two recent experiences that made me think about these and I'd love to hear what the other panelists have to think about it. But one instance that I was thinking about was uh, I was sitting at Flower waiting for my coffee and on the wall they had the uh, famous Sir William Oler quote which is the good physician treats the disease but the great physician treats the patient who has the disease. And it made me think of this patient that I recently got involved with that had been seen by a multi-disiplinary care group for months and months like since October. And I mean much smarter people than me had seen this patient. He had been presented at so many different conferences and nobody could figure out what was wrong with patient at least seemingly. And the patient came into my clinic. It was sort of like an SOS call. It wasn't a day that I had uh my actual standard clinic and they were like, "Oh, something's really wrong with the patient. Can you see them quickly?" And after some time and reviewing the chart for hours, I I personally also couldn't figure out what's wrong with the patient. So I said, "Okay, I'll do a biopsy." Anyway, fast forward and after developing some type of a rapport with the patient, which in my case was only about a week and a half. I then contacted the people that had known this patient for six, seven months. And I said, "I'm really sorry to say this because I know I'm like the new person in the group, but I have combed through this patient's chart for hours and hours and thought about this for days and days. And I think this may be fictitious." And in my own clinic, I use uh DAX all the time. I use an AI scribe. And for some reason, whenever I was in the room with this patient, I thought, I shouldn't use this. And it may be making them uncomfortable or introduce a type of bias. And I remember when our group had written a paper on machine learning um and pain sketches on amputees and phantom limb pain, there was some concern that like maybe some patients weren't writing uh their certain pain postoperatively because they were almost quote unquote embarrassed to write that, that they still had pain or that they were this debilitated. And I just wonder, you know, what what what stewardship we should sort of um think about when we're putting patient data into there because before we ever entered into the patient's chart facitious, we each had individual conversations outside the room. We made sure that we didn't tell anybody at the nursing station. We made sure we didn't tell the residents and we made sure the patient stayed in the hospital and felt very safe for at least two like almost two weeks before introducing the idea of psychiatry even though the patient had an outside psychiatrist. And I think that is very um alarming for patients. And something that um was interesting for me was our social worker actually pulled up the patient's gateway records and saw that the patient was checking their gateway every 30 seconds. So because it's released immediately to the patient, what does that do to their, you know, psychosocial state? Um and what is our responsibility then as physicians that have these metrics that we have to document everything within 24 hours? So >> yeah, I think that's that's super interesting and and it's a challenge, right? Like I think adding to the list like the stewardship around um around you know patient data, how that will ultimately be used and how you know a scale that scores certain things. I'm thinking of, you know, other sensitive topics like things like obesity, etc. that like, you know, then get kind of plugged in the chart in a way that like a human wouldn't necessarily write it or deal with it um in real time. And I think we're kind of facing the I I personally also have a fair amount of patients now that will take their MRI report put it into chat GBT um and then come uh with you know the output and you know it it can uh make the sort of doctor patient relationship portion of that um really uh quite the challenge and so I think it's really important that we think about how we as physicians are also the humanistic side of making sure that we incorporate um the human uh interaction ction into AI um in medicine and in surgery. So, I love that you brought that point up. We have sadly only three minutes because I have so many more questions. Um and so, uh I think you know getting at maybe we'll just do kind of closing thoughts from each person. Um thinking about, you know, some of these concepts of like what the responsibilities are, what the policies are, and who who should kind of be involved in these conversations. Um and then we'll wrap things up. So uh maybe we'll go in the same order. We'll do Dr. Bitterman, Dr. Ali, and then Dr. Duano. >> Sure. Yeah. I think one area that I am worry about and feel very passionate about is that physicians and patients are being too much left out of the conversation of where AI advances in healthcare are going. the kind of discussion is uh increasingly being led by developers and I think clinicians and patients should not feel intimidated by the kind of the fancy tech words they use. We have the knowledge of what is going to be impactful. Um, and that's actually the highest value uh uh kind of uh expertise to bring to developing and evaluating these models. So, um my kind of final statement would be like please get involved in this research. Please kind of be uh kind of if you're interested kind of start collaborating with computer scientists. A lot of them are really excited about using these models to advance healthcare. They don't know the right answers to go after and it's just a huge huge opportunity untapped potential to direct the these advances to in the right direction to make physicians life easier and more importantly improve patient outcomes. >> Amazing Dr. Ali. >> Um yeah and just to sort of echo that uh what Dr. Bman said I mean I think we're seeing now in survey uh responses that you know majority of physicians are using these EI tools. We see patients coming into clinic using these AI tools all the time. I think that ultimately should be encouraged because that's really going to lead to more engagement by patients and more understanding of their own own health care and we as clinicians I think should just meet the patients where they're at. Um as Dr. to want to so be beautifully put it really individualize that care for the trainee for the patient you're meeting uh and acknowledge the fact that we're all starting to utilize these tools more and more and incorporate that acknowledgement into our everyday practices. >> Thank you so much and Dr. Tuano. Yeah, I mean obviously I agree with everything that everyone said and um I guess the most succinct way that I can put it is our ethical obligations in general with utilizing AI are that um things change and our eth ethical obligations change I think especially as things um are sort of um more developed and it's a shared decision ultimately in the end with yourself and the patient and it's just really important that AI is used as a supplementary tool and not the only tool and it's used as a reasoning tool but it's not a replacement for clinical judgment. So, as a daily user of AI myself, on the top of my group's notion, it says, "Treat every patient like family." And I'll leave it at that. Treat every patient like family. >> Thank you so so much all of you, the panelists, for being here and for sharing these insights with us today. I feel like I I have some like real takehomes that I'm excited about um and eager to answer the calls that you guys have laid out um and also more knowledgeable about how AI is used in surgery. So, I can't thank you enough for being here. Um, for those who are interested in in taking us up on that, um, please follow, uh, or continue to follow the HMS Center for Bioeththics, reach out about the Harvard Surgical Ethics Working Group. Um, look us up at neurotchice.org for the neurotchis accelerator. We do a lot of this research. We also have an upcoming seminar on June 12th uh, called digital neuroch neuroch in your pocket. So, we'll be talking about the specifics of neuroch in this field. Um and so please continue to be part of this um of this journey with us about AI and surgery. And so and also thank you to the center for bioeththics and our uh group and nelle for um putting this all together. So thank you so much. Be well everyone and look forward to continuing this work.