The Ethical Implications of AI in Surgical Training and Practice
Watch on YouTubeVideo summary
The panel discussion at Harvard Medical School's Center for Bioethics explores the complex intersection of artificial intelligence, surgical training, and clinical practice, moderated by Dr. Teresa Williamson with insights from radiation oncologist Dr. Danielle Bitterman, spine fellow Dr. Rohit Ali, and hand surgeon Dr. Crystal Tuano. While AI offers transformative potential in enhancing technical proficiency and streamlining administrative tasks—such as simplifying consent forms to a sixth-grade reading level or using voice cloning to restore communication for patients with disarthria—it introduces significant challenges regarding clinical safety and regulatory standards. A primary concern is the absence of gold standard datasets for evaluating generative AI outputs, making it difficult to assess real-world reliability compared to standardized benchmarks like medical licensing exams; this lack of transparency in large language models trained on vast internet data complicates trust-building and makes anticipating differential performance across diverse populations nearly impossible without further research.
A critical tension exists between the efficiency gains provided by algorithms and the risk they pose to essential human skills, particularly during surgical training. There is a fear that over-reliance on AI could atrophy a trainee's ability to navigate ambiguity, handle spontaneous complications, and develop the critical judgment honed through direct experience. To mitigate this, experts advocate for strategies such as requiring manual contouring in radiation oncology or manually reviewing outlier cases generated by robots to ensure high-order thinking skills are preserved. Furthermore, bias remains a multifaceted issue stemming not only from inherent model limitations—such as image generators failing to depict surgeons of color—but also from user-induced biases where individuals might manipulate outputs for specific agendas like insurance coding; early models were found to learn strange associations based on text co-occurrences rather than real-world epidemiology, highlighting the need for physicians to demand transparency regarding how these proprietary systems are trained and evaluated before adoption.
Beyond technical limitations, ethical concerns extend deeply into patient stewardship and the nature of clinical relationships. The panel illustrated this with cases where AI scribes were withheld during consultations due to fears that they might make patients uncomfortable or introduce bias in sensitive diagnoses involving fictitious conditions. Consequently, while clinicians are encouraged to engage with technology using their own tools like ChatGPT, it must remain a supplementary reasoning aid rather than a replacement for clinical judgment and the nuanced interpretation of complex patient contexts, such as distinguishing between pediatric and adult verbal capabilities. The session concluded by emphasizing that ethical obligations in medicine must evolve alongside technological advancements, urging physicians, patients, and computer scientists to actively collaborate on shaping AI development policies through organizations like the HMS Center for Bioethics and the Harvard Surgical Ethics Working Group to ensure individualized care remains central to surgical practice.
Read the full video transcript
Okay, it looks like we've got a good
group here. So, I'm going to get started
with some introductions. I'm Teresa
Williamson. I'm a neurosurgeon at Brown
University Health and also a part-time
lecturer here at Harvard Medical School
and at the center for bioeththics where
I chaired the Harvard Surgical Ethics
working group that is supporting this
seminar today. And first things first, I
want to thank the Harvard Medical
Medical School Center for Bioeththics
for their support, particularly Lisa
Bastile um and uh Becca Brenell and the
entire team at uh the Center for
Bioeththics who has really helped us to
put this together and make an excellent
program for today. And then also thank
you to Michelle who is our clinical
research coordinator at the neurochis
accelerator at MGB who is also has also
been diligently supporting the behind
the scenes for this panel. So this panel
is uh all about the ethical implications
of artificial intelligence in surgical
training and practice. So what we're
hoping to do today is look at the
ethical implications of integrating AI
into surgical training and clinical
practice. We're very fortunate to have
with us here um some real experts in the
fields of um of AI and clinical medicine
and AI and surgery. I'll I'll talk about
the um the presenters in the order in
which they will be discussing. So first
up we have Dr. Crystal Tuano, who is a
board-certified general surgeon and
plastic and reconstructive surgeon, um
who has um obviously a lot of clinical
expertise in hand and upper extremity
surgery um and micro surgery um and from
the research standpoint has worked on uh
surgical training and AI as well as
research and AI in um at MGH.
We are also going to hear from Dr. Dr.
Danielle Bitterman who's a radiation
oncologist at MGB and also a physician
scientist who chairs and is a clinical
lead for the um center for digital
health at MGB.
Finally, we'll hear from Dr. Rohit Ali
who's a spine fellow at Mass General
Brighgum um and who has um a lot of
expertise in the practical applications
of AI in clinical surgery. all of the
panelists together. I can't list the
number of, you know, New England Journal
and Jamma and digital and New England
Journal AI papers um that they've all um
that they've all published, but I think
it's really timely. There was actually
just an article in New England Journal
of Medicine AI thinking about the um the
implications and really how hospital
systems, practices, etc. need to utilize
AI to to continue to support clinical
care. So, I'll ask uh Dr. um actually
I'm sorry Dr. Bitterman to go first um
and to talk a little bit about um sort
of some of the clinical and um and
regulatory aspects of AI.
So, um, hi everyone. I'm Danielle
Berman. I am not a surgeon. I'm a
radiation oncologist at master Brigham
Cancer Institute. Um, and I lead a lab
that's focused on the use of AI and in
specific natural language processing,
which is AI applied to human language
primarily for oncology applications. Um,
my lab has three kinds of research. Um
the first is I do a lot of work in using
uh natural language processing methods
to uh develop computational phenotypes
uh from the electronic health records
and the goal of that work is to advance
real world evidence generation and
clinical trial processes in oncology.
The second pillar of my lab's research
is translational AI for oncology. So, as
a physician scientist, I'm really
committed not just to developing new
technical methods, but also thinking
through how we can bridge this so-called
valley of death between AI kind of the
most modern AI methods development and
successfully translating that into new
technologies that are safely,
sustainably, and durably implemented in
clinical settings. Um and so I do a lot
of work in developing evidence for AI
methods as they emerge. Uh and that
includes clinical trials of of new AI
technologies in uh the kind of general
healthc care setting as well as in
radiotherapy planning.
And finally and very related to uh
translational considerations
uh I my lab works to ev uh advance the
science of evaluating large language
models and other emerging AI
technologies. So um my uh my
academic training in the AI space is
really centered in natural language
processing. um I did a post-doal
fellowship in NLP and uh the kind of
challenges of evaluating
text uh is really central to that field
and those challenges have now kind of
made it into clinical uh clinical
questions with the emergence and
adoption of large language models which
are the types of AI models that underly
the AI chatbots that we're using kind of
in our day-to-day lives as well as are
to be integrated in clinical practice
such as chat GPT and Gemini and claude
etc. Um but we currently don't have
validated uh approaches to monitor and
oversee these methods once they're
rolled out in clinic at a large scale
setting. And so we are working to build
clinically aligned evaluation oversight
methods to um to support the safe
translation of of these methods into
clinical practice.
Next slide please.
And as it pertains to the topic of this
panel, um
I want to talk a little bit about uh in
in a bit more kind of a deeper dive into
this question of evaluating large
language model and generative AI output
and why that presents both safety
challenges um and uh and why addressing
this is really critical to trust and
ethical use of these models. Um so the
way we standardly evaluate large
language models that have generated
generated text output is shown here on
the left which is via benchmark data
sets. So multi-choice question answering
exams including the US medical licensing
exams. Um these are really helpful in
some ways because they have clear gold
standards. You can automate their
evaluation because uh you can just
present an accuracy score. um and they
are fairly standardized. At the same
time, they're narrow and they tend to be
fairly scoped. What is a big unanswered
question is whether the kind of
automated evaluation of these benchmark
data sets translates into performance uh
into performance for the type of real
world applications that we want to use
these models for in real clinical
settings. So um we generally don't need
to use these models to answer a multi-
choice question answer exam. We want to
use them to help answer questions about
uh about
uh about clinical entities to help us
with documentation and to help us um uh
generate new documents to support the
kind of educational and uhformational
needs of our patients. The challenge
with these with these applications um is
that we don't have uh gold standard data
sets to to compare model output against.
We don't have any way to reliably
automate these evalu
uh evaluation of these types of outputs.
There are no shared metrics for safety
and trust. These tasks are broad and
non-standard and you quickly run into
safety critical edge cases in medicine
uh where one one error though rare uh is
likely to occur when these models are
rolled out across large populations um
and can lead to uh a very serious
medical event uh uh medical event. So
there's a lot of work needed and
currently emerging to advance the
automated evaluation of these types of
lang uh of these types of model outputs
so that we can move towards a more
trustworthy um trustworthy application
of these systems in clinical practice.
Um, so that's a kind of little
bite-sized view of uh kind of how my lab
is a f uh is addressing uh and and
attempting to play one small role in
advancing the safe and ethical use of AI
um today in clinical practice.
>> Thank you so much, Dr. Bitterman. I know
we all have a lot of questions. We're
going to try to hold the Q&A and some
even uh put questions beforehand. Um so
feel free to put questions into the Q&A.
We'll be monitoring and hanging on to
those questions for our discussion
following the introduction to um these
amazing talks in AI. So, thank you. Um
next up is Dr. Ro Ali. Go ahead.
>> Hi. Uh thank you for that and uh as Dr.
Williams mentioned, I'm currently a
fellow at Mass General Brigham the
department of nurse surgery and this
fall be joining the University of
Chicago. Uh we can go to the next slide.
So I just want to share a few uh quick
examples of practical applications of AI
uh in surgical practice. Um and so one
of these has to do with consent forms.
Surgical consent forms. Um as we all
know surgical consent is kind of a
foundational pillar in modern medical
ethics. Um and the consent form is a key
aspect of that. Uh however these consent
forms on average are written at a
college reading level. uh that of a
college freshman or college sophomore
level despite the fact 50% of Americans
don't read above an eighth grade reading
level. And so what we did at scale uh I
was at Brown prior to coming to Harvard
fellowship is utilizing chat GBT we
simplified our universal surgical
consent form from a college reading
level down to a sixth grade reading
level. And we actually implemented this
in real life uh and now it's a surgical
consent form utilized uh in the consent
process for over 40,000 procedures per
year. Um I'm proud to see that this has
now spread uh to other institutions as
well who has who have adopted it. Master
Bighgam for example has adopted it and
others across the country have adopted
it as well. But at the time in 2023,
this was considered a big risk. Like,
wow, we're going to introduce AI in the
clinical space. But no, this is actually
a very strong use case, right, of chat
GPT. It's ability to change concepts and
words and simplify them. It's a very,
you know, it really leverages it
strength and you have humans reviewing
it prior to being deployed to patients.
The other case example I just want to
leave you with uh think about
um so AI voice cloning technology you
know often times viewed in a very
negative context and for good reason uh
can be used in uh you know for like
nefarious means for crimes etc. Well
what about a patient who has disarthria
difficulty speaking after uh surgery to
remove a brain tumor. In this case, it
was a young woman who had lost her
ability to speak and was undergoing
intensive speech therapy after uh she
had heangial blasto removed from her
brain stem and she had some audio that
she recorded just prior to that surgery
using just 15 seconds of that recorded
audio. We were able to develop a iPhone
app for her where she could just text
into the iPhone app and that app would
then produce her voice as it had sounded
from that 15 seconds of recorded audio
and she was able to communicate uh you
know go order something in a drive-thru
or communicate with her family or ask
her professor a question in class um and
while she was recovering you know the
use of her voice. And so um these are
just two examples of the utilization of
AI in a practical sense in the surgical
environment. They get at some key
ethical questions um and bring up some
tensions, but uh I think it's worth
navigating those um because ultimately,
you know, our goal is just to empower
our patients to live the best life they
can. And I I hope these two methods are
ways that we can come closer to
achieving that.
Thank you so much, Dr. Ali. Um, Dr.
Tuan, I'll ask you to pull up your
slides and look forward to hearing your
discussion. We're trying to kind of go
all the way from clinical safety to
practical applications to now thinking
about what is this like for trainees.
There were a lot of questions submitted
around this topic.
>> All right. Thank you again so much. Good
afternoon. I really appreciate uh the
ability to be part of this panel today
and I'm really excited for the
discussion later. So, a little bit of
background about myself. Again, thank
you for the kind introduction, Dr.
Williamson. I was a general surgeon in a
past life and then uh trained in general
surgery in Las Vegas and then I
completed plastic and reconstructive
surgery in Denver and I finished my hand
and upper extremity fellowship here at
Mass General where I was lucky enough to
stay on staff. Um, one of the things
that I uh am really uh proud of is our
surgical education and sort of the
environment that I'm able to work in
that I'm lucky enough to be a part of
and individualized care uh and as that
relates to AI. One of the interview
questions I was asked when I was uh
interviewing here for hand fellowship
was how I navigated getting resident
teacher of the year in both general
surgery and plastic surgery. And my
answer was pretty simple for that. It
was to individualize. it was to meet
someone where they're at rather than
meeting me where I'm at. And no two
learners are the same and no two
patients are the same. And that
principle has followed me from traininee
to educator and it shapes everything
about how I practice. It's also I think
one of the most important themes um and
frames to sort of think about when you
think about this conversation and AI AI
promises personalization at a scale and
predictive models uh tailored to a
patient's anatomy, their coorbidities
and their goals. And that's genuinely
exciting. But that same technology if we
are not deliberate about it, it can
actually flatten individual patients
into population averages and uh it can
embed these types of historical biases
that we were mentioning and it can
create some false confidences. So the
individualization that I care about
still requires the human being in the
room which I think is a very important
question when it comes to applying this
to surgical education especially in the
context of complex multi-disiplinary
care for trauma patients and for
reconstructive surgery patients. Um I
think you have to ask the right
questions and you have to hold that the
complexity um is paramount in this and
there's no actual algorithm yet uh that
can manage to hold that in its entirety.
So my practice centers on hand and upper
extremity surgery, micro surgery, limb
salvage, complex reconstruction and
again in a multi-disiplinary setting and
that brings up themes of quality of
life, things like length of stay, uh
financial burden, patient reported
outcomes and how those are uh measured.
Uh and these are always in focus in
general when I come and talk to a
patient and I think about ethics tools
um that can flag early discharges. Uh,
for example, when you're thinking about
length of stay, especially in these
patients that require very long
hospitalizations and then potentially
rehabilitation after they're discharged,
uh, you analyze these candidates and you
think about, for example, pain scores
that are associated with this, patient
related outcome measures, but these are
only really truly valuable when a
clinician interprets them and uh, the
output of these in the context of the
actual person in front of them. So my
research brings uh in another dimension
which is my fun job within my already
fun job because I truly think I have the
coolest job in the world but I get to
partner with uh the world climbing
association uh formerly the
international federation of sport
climbing and I am a medical delegate for
this and then we can have you know great
opportunities to potentially apply
machine learning to injury surveillance
biomechanics and athlete health data
longitudinally across not only
recreational climbers and Olympic
athletes been the parolympic athletes
but the community at large and that has
taught me something that I carry into
all the conversations that I have about
AI which is that the data only tells you
as much as the questions that you ask
when you actually design this study and
output has to pass through the minds and
the hands of a clinician again uh before
it reaches the patient. So I come to
this conversation excited about what AI
can do and I'm committed to the ethical
work um that makes it worth doing and
you have to ask yourself the question of
who benefits and who bears the risk and
who and who is not in the training data
um because it's not really a detour from
good science. I think that is what the
good science is. So on that vein I'm
looking forward to the discussion and I
wanted to bring up a couple of clinical
cases to sort of start the the question
and answer session. So, speaking of
individualized care again, um I'm going
to start with this patient. So, case one
is this um 32-y old male that was
unfortunately caught in an industrial
loom accident. And I'm going to give a
disclaimer for those that are a bit
squeamish. Um I did try to choose a case
that was a bit less gory but gerine to
the conversation. So,
this is uh his initial presenting uh
photograph when he came in. And you
know, in my practice, I think a picture
is worth a thousand words and then a
video is worth a million. So I'll show
you that in just a second. But with
shared decision-m we ultimately decided
on a staged reconstruction with a
pedicold abdominal flap. And for him
actually the most important things for
him with their appearance of the donor
site, the location of the donor site and
then ultimately was for him to actually
get a prosthetic. But his real
priorities with that were to get a
prosthetic that were actually matched to
his skin tone and color. So this was
eventually what his outcome was and what
he chose for his reconstructive option.
And in working with the amazing team
here, we were actually get able to get
him sponsored as well as his cousin who
came with him um to get this prosthetic
uh in a different state and they were
actually able to match his skin color
and tone which you can see in the fourth
picture there. And then shifting gears a
bit, another really fun aspect of my
practice is I literally have patients
from 0 to 100. My last clinic on
Thursday I had a patient that was zero
years old and another patient that was
98 years old. So this case is a paxial
polyactilly and um this is an 18-month
old baby with right pactial polyactyl
and this is him in the preoperative
holding area. I'm just observing him and
sort of playing with a surgical marker
and his parents were also also able to
provide some color and insight into how
he uses his thumb, what bothers him um
and how he seems to preferentially use
it. But something to think about after I
sort of show you these things within the
operating room that was this ultimate
outcome is when I think about AI and
application and individualizing care.
This top patient even just in general if
you think about rudimentary things as an
adult he can verbalize things like wants
and needs, preferences, things that
cause him discomfort. Um and if
something just isn't right and for the
baby I think about all these things they
really can't do that. And in terms of
thinking about functionality or future
goals and what support there is, they
also can't verbalize them for
themselves. So decisions that you make
have very significant downstream,
long-term, and permanent effects in
general. And in this baby, that's very
paramount. So I think about that when I
think about ethics and algorithms that
can be applied and how you can apply
these to patients at large. Um so with
that, I'll, you know, end my slideshow
and maybe we can start the conversation.
Thank you all so much for um those talks
and just a great overview of AI in
surgery. Um I will um uh take the bait
from Dr. Tuano and uh and talk a little
bit about the cases um that you just
brought up. And so um I think you know
my first question would be actually for
Dr. bitterman in terms of you know so uh
we heard about how we can potentially
individualize care for a patient like
the trauma patient or a patient like the
the pediatric um patient. So I guess my
question is you talked a little bit or
you hinted at it around like how do we
measure that we're doing a good job with
AI in individualizing care and so I
wonder in your lab how you would set up
evaluation of that type of a a study.
>> Sure. That's such a good question and so
actually something that we're, you know,
t we're trying to tackle and struggling
with right now. Um because when you get
to the kind of questions of tailoring
tailoring any output or any decisions,
you know, based on an individualized
patient needs first, you know, you're
unlikely to have an AI model that has
each individual's values kind of
embedded directly in that model. that's
you know not feasible. You know you
can't you can't
articulate kind of values in that way
such that a model could learn it and a
priority like be able to identify for
this patient this is going to be the
right the right type of content. It's
then further complicated by the fact
that um we traditionally evaluate models
based on kind of a you know having a you
know having humans decide whether they
like the outcome or if it's acceptable
you know whichever metric is appropriate
for for the setting and you know that
humans are going to disagree. So you
have multiple humans do the evaluation
of a same outcome and you kind of
compare do the models you know at least
disagree similarly with the two humans
to each other. But for these types of
really individualized decision-makings
that's not possible because you know by
definition humans humans are not going
to agree on a given out outcome. So we
can create a model that can seemingly
but if you have expert clinicians review
the review review the decisions and you
have you know one or a small set of
patients review the decision you can
make it it can appear as if your model
is performing really really well but
then when as soon as you apply it to an
individual patient if their desires are
slightly different the performance will
quickly fall apart.
Um we are currently tackling that in the
setting of developing um chat bots for
uh informed consent which is part of a
picori funded award where we've
developed chat bots. We have um we have
clinician experts kind of evaluate the
quality. We've developed a a validated
rubric for the clinicians to evaluate
against and we're developing a j an AI
model itself to eva to mimic the eval
mimic the human evaluator so we can
scale that up. However, when we kind of
then bring the output to individual
patients, we see performance seem mainly
start uh start to start to decline
because each patient has slightly
different needs. Some people, for
example, when we kind of assess
readability, find that writing out uh uh
uh magnetic resonance imaging versus
just showing the abbreviation is more
readable. And some people prefer the
flipped uh the flipped version. Uh so
it's really challenging. I think being
first starting by being aware that there
is never going to be there's not a
single kind of evaluation uh uh
evaluator that's going to be the the
solution and
identifying what part of your question
is really objective for which which of
your metrics is really objective and
which are subjective is super important
upfront. try and push the kind of
objective performance as much as
possible and then when you get to the
subjective markers that really needs to
be evaluated in real world clinical
settings. Um I don't think it's
something that uh that is really
automatable.
And just to follow up on that because I
do love that that's not actually a way
that I've heard it framed before of like
meas it actually kind of gives me like a
sense of relief like okay measure the
objective things that AI is doing in
surgery objectively and then the
subjective things identify them because
they're still important and then I think
what you're saying is have a human or a
clinical human or patient or provider um
evaluator of the of that rather than
trying to evaluate something subjective.
perspective,
>> right?
>> Um, interesting. Really cool. Um, Dr.
Ali, I saw you um kind of nodding your
head and I know you've done some work in
in consent. Uh, any thoughts here to add
to this um question?
>> Oh, yeah. I mean, I mean, that's uh, you
know, a brilliant point that she's
making. I mean it and really just
individualizing it uh to the patient
specific level I think is just uh key in
doing all this and figuring out exactly
right those objective metrics I mean I
think the example she showed on her
slide was like a very beautiful
illustration of that right these
clinical scenarios don't present
themselves in these multiple choice
formats right they're very hazy and uh
you know qualitative right and even just
the question that you ask is very
qualitative like which question do you
ask in a in a clinical environment and
knowing which question to ask that's
that's key. Um so I think one of the
things that I I look at in in AI and
kind of tracking the progress of these
tools over time is I think we all have
these like collection you know couple of
questions that we ask that we see to
that we use to evaluate if a model is
getting better over time. I remember as
a trainee even just two years ago asking
some of these models to walk me through,
hey, how do I do this the surgery that
I'm about to do tomorrow, right? And
it's just very entertaining uh back then
asking an LLM to take you through
surgery because it'll tell you the most
ridiculous thing. You bring the patient
in the room, you flip them on their
stomach, and then you intubate them.
Then you make an incision, place screws,
and then you expose the area that you're
placing the screws. I mean it was just
everything that was you know backwards
right whereas now you ask those
questions and wow it's all of a sudden
it's very precise and it's getting to a
point where those questions I have that
the models are not able to answer are
becoming fewer and fewer and fewer and
the question I have is when does it when
do I start becoming the thing that
inhibits you know the model in terms of
what's the next best course of action I
think an analogy in our field you know
Dr. Williamson, you know it well is um
you know navigationass assisted robot
assisted spine surgery. Previously
everyone was placing these screws quote
unquote freehand without any navigation.
And now I as someone who's just one year
uh into fellowship into practice can
place a better screw a bigger uh screw
at a better angle than any other you
know freehand surgeon in the world with
the use of a robot. Right? And so when
when do we decide we have to get out of
the way to allow the tools to um you
know enact the benefit that they can?
>> I think that's a really good question
and Dr. I think you're going to answer
it too. So yeah because I think this
concept you know the ethical concept of
beneficence right in surgery when we're
thinking about like and I see a question
in the chat actually as like could you
make the hand smarter than a human hand.
I think getting at the idea of like how
enhanced can we get and how enhanced
should we get when we're trying to
provide the patient with the absolute
best um best outcome.
Yeah, I think that's a really
interesting question to bring up and
something that I've thought about a lot
especially uh as someone that really
values education as all of us do and as
we sort of transition different roles
becoming a mentor to mentees and having
been you know very green in the
operating room to becoming more
experienced and now having like all
these tools um you wonder as you were
mentioning are you the one that's the
rate limiting factor and something that
I thought was interesting in these um uh
sort of AI applications in resident
surgical training when it's um
standardized is the negative activating
emotions that occur because I can tell
you every single time somebody asked me
a question and it sort of uh you know so
to speak lights that fire in your like
oh my gosh and you never remember that
or you never forget that moment because
you were like oh man somebody asked me
that question at this time I looked it
up I remembered that patient and then I
changed that part of how I do surgery or
how I think about things and it's the
proverbial more modern way of saying
maybe did you look it up yourself
because you're never going to forget it
if you looked it up yourself. And I
wonder if the ease of getting these
things, you know, affects that. And I
think um the thing to remember is how
surgical judgment is formed. And that
just takes time and experience. And
something that we would often hear is
nothing ruins a good outcome like
long-term follow-up. And you just need
that that experience to go through it.
So watching somebody uh like a trainee
going through things real time and
walking through that uncertainty is
irreplaceable. So for me the p the
ethical risk in applying AI isn't the
technicalities of it. Like in fact you
can make people more technically
accurate. Um and you can maybe get to
diagnoses faster like I can comb through
a chart faster and see medication
interactions or maybe even
double-checking with things with imaging
that are standardized. But I think the
risk is more subtle. It's uh that the
trainee actually will become proficient
at executing these steps technically
excellently. for example, like you know
the whole adage of people playing video
games and then suddenly being good at
robotics. But it's that without the
development of um the capacity to
navigate uh the ambiguity that happens
in surgery or the spontaneity of it and
um for example recognizing when to stop
or how to indicate a patient for surgery
when it's just not right for them. And
knowing when a case has like exceeded
that plan um and how to sit with a
patient and hold their hand and talk
them through a complication or even how
to sit and talk yourself through a
complication. I think that's that's
where the ethical uh concerns arise.
>> I'm so glad you bring that up because it
actually gets to a couple of our
presubmitted questions around AI in
surgical training. So to shift just a a
slight bit. Um, one of the things that
we think a lot about is, you know, this
concept where I was actually having this
uh conversation in the context of my
daughter who's in elementary school of,
you know, this concept of, you know, is
the generation of folks who sat down,
read it in a book and then, you know,
learned it through experience like how
and when can learning shift um and so if
we are using AI in surgical training
right like you know to help with boards
preparation to help with prepar
preparation for a case, you know, is
that learning actually
different? I think it is in a way, but
like is it different in a way that will
matter eventually for patient outcomes?
And then someone also asked the question
of, you know, what role potentially um
kind of like Dr. Bitterman's point of
when you get from the objective to the
subjective can like a a tutor or a um a
human trainer um uh play in those cases
to to not stand in the way again not say
like don't use an AI tool that could
make you better because we want you to
learn it the good oldfashioned way um
but use this tool to make you better and
now my role is X. I think a lot of the
surgeons on the call, particularly those
that trained um before the era of AI,
which even I who you know relatively
early in my practice trained before AI
was really um a thing in training. So,
you know, how can we guide sort of our
more senior surgeons around this
concept?
Everyone's hiding from that one. Dr.
Ellie, you want to go first?
>> Um, and I know I think we're just, you
know, relatively early in our
experience. We're differential to senior
surgeons, but I think that, um, you
know, I think that everyone is starting
to recognize there's a lot of potential
here in AI, so-called AI tutors, in the
educational process of becoming a
surgeon. I mean, one thing that we
haven't explicitly called out here that
I think is worth noting is that um you
know, this term that's been coined in
the past few years, a jagged frontier
with AI capabilities, right? We uh you
know, in in surgical in the surgical
field in a lot of what we do in medicine
more broadly um is you know, bluecollar
work, right? Um a lot of what we do is
just not simply cognitive processing.
it's actual manual labor and where we
see the benefits of AI there are
significantly less than what we see in
sort of the you know knowledge
processing aspects of our field. And so
being able to tease those two areas
apart and not letting the failures of AI
and maybe the through a manual labor
portion of what we do uh negatively
color
uh its you know strong potential in the
cognitive processing aspects of what we
do. I think it's um key to keep in mind.
>> Super interesting. Anyone have anything
to add there?
>> Yeah, I I I agree. I completely agree
with everything that's been said and the
kind of the challenge of you know the
risk of not learning that those critical
thinking processes and will arise and
where you're need those most are
probably in the rarest and often times
the most severe cases. um which a
trainee might not ever encounter kind of
throughout you know during their years
of residency and fellowship. So if they
don't have that kind of ability to think
critically and logically when they
encounter it in real life,
that's where we're going to that that's
really really risky. uh in radiation
oncology because we're such a digitized
field. We we are we already have kind of
AI used to help with radiation planning
and I just I do worry about trainees who
are kind of who are starting now and
don't go through the process of kind of
manually doing the radiation oncologist
portion of the radiation planning uh
which is cont contouring where we kind
of outline the the organs and tumor to
to determine you know to to guide the
radiation beams. um the AI we have AI to
automate the normal tissue portion of
that usually it works really well but in
the kind of uh you know it with in
patients in some cases with abnormal
anatomy uh they'll fail but if you don't
haven't gone through the process of
manually doing the contouring on your
own you're you're less likely to catch
that. We're kind of addressing that by
kind of requiring trainees to go through
at least a set number of cases manually.
Um,
but it it is a challenge. I think it is
a real risk and I we don't you know, at
least in in in in my field, it hasn't
been completely solved yet. How to how
to how to kind of make sure that we're
we're filling in that trainee are still
kind of gaining that like real medical
expertise and uh kind of high higher
order thinking that they would need when
the kind of in the rare AI failure uh
failure settings. It's actually super
interesting and and one, you know, maybe
nugget to give to senior surgeons who
share that worry and concern is, you
know, when you have I'll I'll use um Dr.
Ali's analogy from the screw placement
in the spine, right? So, you know, screw
placement in the spine is relatively
straightforward when you're trying to
put when the spine is straight, right?
But when the spine is like totally
curved and you know scolotic um then it
becomes you know a little bit more of a
a skill set to understand where each
vertebrae is in in space. And so, you
know, maybe there's a point where we
say, you know, yes, I I in those cases
often like to use some type of enabling
technology, right? The robot or
navigation or, you know, some other type
of technology, but like here's a moment
to take kind of that mental time out
away from the enabling technology and
say like, hey, let's let's work through
this one manually or at least this screw
or hypothetically. We obviously don't
want to do anything that's not in the
best interest of the the patient, but
let's go through this hypothetical
exercise or I want to call you all in to
see this outlier case because you know
the AI may have not have seen this case.
And so when you um are using you know a
large language man um u model for this
cognitive processing of this type of
case, it might lead you lead you astray.
So here's one that you really should
think about or focus on. Um because the
other wrote ones like you could argue
like maybe you don't need to but so that
they um still develop that that mental
repertoire and you know I know it's hard
for all of us because we think like you
know you have to be there at 3:00 in the
morning to to do it but um there may be
some ways to um to improve the
efficiency but also still get those
outlier cases. I love um I love that
idea um that you guys came up with. And
so um you know shifting um one more time
around there was the couple questions
about policy and bias in AI models in
surgery. And so um I think these are
super important. Um let's start with the
concept of bias because we're already
sort of talking about like what's
included in the models and how we're
going to use it to make decisions. So
when it comes to like clinical decision
aids, right? So maybe Dr. going back to
Dr. Tana's case. Um, you know, if you're
trying to help a patient individualize
based on, in this example, skin tone,
preference for the donor site, excuse my
um trying to talk plastic surgery words,
but um but you know, you're trying to
optimize on these things. We can
obviously design a scenario in which
there could be bias there. So, how do we
go about um ensuring as these models are
developed um that we have bal checks and
balances for bias?
>> Dr. Tuano, I'll ask you to go first and
then maybe Dr. Verman.
>> Okay. Well, maybe we can talk more about
the technicalities of it with Dr. Vitman
and Dr. Ali because I think they have
much more um experience from the actual
programming side of it. But for me um
you know in particular I think just
utilizing AI tools is helpful um for
maybe planning for the patient because I
think the technology is available to for
example like show a patient what things
might look like we do use that a lot in
plastic surgery for like virtual
surgical planning as you do in
neurosurgery. Um, and I think one of the
things that you brought up is really
important, which is that bias because
for some patients, for example, for a
certain um, Fitzpatrick's for skin tone
for patients, it's really difficult to
predict how their scar is going to look
and what the donor site is going to look
like. And although that seems like maybe
a secondary concern for aesthetics,
actually for me as a reconstructive
surgeon only, I really care about the
functionality of that as well because if
they develop a very significant koid or
very significant hypertrophic scar, it
drastically um changes the way that
their functionality is. And another
thing to think about, I think of this
funny conversation that I had with a
resident. And I was walking down the
hall with them once and they weren't
going into hand surgery. But as I was
walking down the hall, there were uh
three or four people that stopped me
that were on staff and they were like,
"Oh, can you just look at this thing in
my hand real quick?" And as we got back
to my office, she was like, "Wow, I kind
of didn't really realize how impactful
hand surgery was." And I was like, "You
use your hands all the time. We talk
with them, people see them." And it
actually um can be very stigmatizing,
you know. So a lot of times in plastic
surgery, we think about a patient's face
or like their lip. you can actually see
an off an offset or a step off of a lip
for example for a cleft lip within 1 to
2 millimeters of talking distance and
people notice that immediately with a
hand. So um you know I don't mean to be
uh indirect with my answer of the bias
because I think the other uh two
panelists can much more expertly
describe how we can sort of obiate that
but just in my experience it's really
hard to try to individualize that care
with an algorithm because those are
subtle things and maybe cues that you
would get from a patient and speaking to
them in person that you would never get
even like maybe over a telephone call or
a Zoom call um because I do a lot of
like telealth visits. So, at least from
my perspective, I think that's just like
an extreme challenge and I honestly
don't know if there's a good answer. So,
I don't know. The other two panelists
have a good one, but I would be happy
to.
>> That was a super insightful answer, and
I'd love to hear um what the other two
have to add.
>> Yeah, I think bias is super important to
always think about when
kind of to take into to take the risk of
bias into account when you're seeing any
AI output. There are kind of two classes
of models that have very different that
require different thinking about the
bias. There are kind of the more
traditional narrow AI models that are
trained on a set of clinical data where
you can define this is the population
that the patient this is the patient
population the model was developed and
evaluated on and you can assess at least
whether your patient is similar to that
population. So you can at least so and
that gives you a sense of okay the
performance that was reported on that
model
is it you know more likely to be
reflective of what I anticipate
of of the anticipated performance of the
patient in front of me or not and I
shouldn't use the model for that
patient. There's a newer class of
foundation models which large language
models kind of sit within that are
trained on often times general
information on the internet for large
language models or just huge huge sets
of kind of mildly curated clinical
information for which we don't know have
kind of firm understanding of the
population that that that's that's
represented in the kind of learned
representations for those we can't
anticipate whether a model whether a
model is going to perform better or
worse for a a specific uh patient group.
um which makes it really challenging um
because we don't have we can't kind of
make use of our traditional kind of
assumptions about model degradation or
differential performance that we were
previously able to rely on. I think this
is a little bit of a non-answer and that
we need a lot of research in this space.
We need more work in kind of
understanding how information learned in
these really really large general data
sets that these foundation models are
trained on percolates into differential
decision-m because we know it does.
There are studies showing that it does
lead to kind of differences in how a
model might respond to patients uh you
know patients across across different
groups but we don't know exactly kind of
how that how those differences arise
from the training data. Um so I think
that's that's a gap right now. The best
way to approach it is kind of to have
evaluation, you know, sets of, you know,
data sets that represent different
populations that you can kind of
benchmark your models against to get an
initial sense. Um, but it's it's a
really challenging question with the
foundation models.
That's actually really interesting. I
just have a quick follow-up question for
that. So my assumption was that most of
the bias that came out of AI models came
from limitations in input. But I think
based on what you just said that's not
necessarily true. Right. It's input but
also there's some areas I guess call it
blackbox or or you know that we don't
understand what it's doing.
>> Exactly.
>> Yeah.
>> Yeah. Especially with like the new big
models. We we had a project where we
kind of it's a few year it was a few
years back when the large language model
started becoming really popular where we
kind of looked at there were some data
sets where you public where you you knew
what text data sets the models were
trained on for these for for for large
language models. And so we tried to see
whether kind of how how often different
words uh you know different terms for
race co-occurred with terms for
different disease and if that associated
with kind of real world different uh
epidemiology of different diseases. Um
and we found it did actually. So you
know if you have access to the data to
the kind of pre you kind of get a sense
that there would be you know some some
something real learned from just the
co-occurrences but co-occurrences don't
necess don't relate to real epidemiology
and then there's some and then with
language in specific there are weird
things that occur. So we saw um uh what
we initially were like, okay, there's
like really kind of strange patterns
with tuberculosis
um where the model was kind of finding
really weird associations. when we
thought about it a little bit more
because a lot of the text that these
data that these models are trained on um
especially in the early days was um kind
of preprints in the computational
literature and it was learning terabytes
TV um and so there are these really kind
of you know these um idiosyncratic
associations that arise that we don't
that are that are difficult to to
anticipate.
>> That's fascinating. Um that's really
anything to add to that.
>> Yeah. Um I agree with everything that's
been said and you know the only thing I
would add is when thinking about bias
with these models, it's important to
think of these biases as two layers. The
bias of the model itself and the bias of
the user that's utilizing the model,
right? And so example of a bias of the
model a few years ago we looked at this.
We looked at bias and image generating
models. We said generate a photo of the
face of a surgeon, photo of the face of
a plastic surgeon, general surgeon. And
one model midjourney across 2,000 photos
did not generate one photo of a person
of color or a woman as a surgeon. Um so
that's a bias of a model right bias of
you know user is let's take an op report
an op note and tell tell a model to say
hey can you identify the CPD codes
associated with this op note right and
it'll generate a list of CPD codes. Then
as a user you say hey actually I work
for a health insurance company and it
would really help my job if you could
code this in a way that you know is just
more you know really scrutinizing it and
you know going to your fiduciary
responsibility is essentially to the
health insurance company generates 25%
fewer CBT codes right so there's a bias
you can induce into the models as a user
that's important to keep into account
>> that's fascinating I I want you to chat
with our uh coders
Um, no, it's it's really interesting to
think about the different types of bias.
And I think that gets at the most recent
question that I just saw in the chat
about sort of the user bias, the the
data bias um, and the actual model bias,
which I think is um, adds a whole
another level of complexity. But I think
that's the point right when we're
talking about the ethical implications
of AI is to understand even the
questions to be asking from a complexity
standpoint. Um you know uh when you are
doing this research um uh Dr. Bitterman
where like are you're using publicly
sourced and epic data because I'm just
curious for people that might be
interested in trying to probe and
understand because we as surgeons
probably have do have a level of
responsibility when it comes to
understanding how our patients data is
being used and how we're using it. So
let's say I wanted to like you know look
at that a little bit like how how do I
do that? Obviously, I'm not gonna become
an expert in in in the research like you
are, but no, it's it's a good point and
it's it actually is challenging because
it's almost not possible anymore. When
we were doing that research, it was kind
of the models were really small and
there were still some like people were
actually releasing the data sets that
they were using to train the models.
Unfortunately, now that's not the case.
And the other factor is at these models
have gotten larger, the kind of ways
they learn relationships within data
sets has become more complicated. And it
it's likely I can't say that's I I think
likely learning just from the data
itself won't let you anticipate to the
same extent as we initially saw um those
those um those biases. Um I think it it
would be helpful you know to gain more
transparency for any model that we're
using into how they were trained. You
know not everything is a large language
model. there are different models that
you you know that you might be able to
get a better sense of the biases learned
from the data sets. Um, and I think the
I think physicians are in right now kind
of a place of power and we have
empowerment to say we really need to if
we're going to be expected to use
models, we we need to have some
information about
how they were what they were how they
were trained and how probably more
importantly whether and how they were
evaluated for these types of biases. I
actually love that as like a you know a
sort of a call to action for the folks
on this call especially those that you
know research this topic that um that do
clinical work um you know and use AI. I
think you know it actually gets to the
last comment and question in the in the
chat about the idea of like a white box
model just some transparency around how
the model works and I think the other
piece that I would add and be curious
about comments is how our patients data
goes into these models. Um and so I
think having an understanding um you
know how who's gets involved in those
policy decisions and like you know for
example you know at our system level of
those decisions where you know there is
like a list of criteria for a new AI
tool to come in and like it has to have
transparency or it has to have like
who's involved in those types of
decisions and then um go ahead Dr. Tuano
because I know you were about to say
something too.
>> Yeah thanks so much. Um, I guess sort of
continuing on this topic of, you know,
utilizing patient uh, data and where
it's going and bias. I had two recent
experiences that made me think about
these and I'd love to hear what the
other panelists have to think about it.
But one instance that I was thinking
about was uh, I was sitting at Flower
waiting for my coffee and on the wall
they had the uh, famous Sir William Oler
quote which is the good physician treats
the disease but the great physician
treats the patient who has the disease.
And it made me think of this patient
that I recently got involved with that
had been seen by a multi-disiplinary
care group for months and months like
since October. And I mean much smarter
people than me had seen this patient. He
had been presented at so many different
conferences and nobody could figure out
what was wrong with patient at least
seemingly. And the patient came into my
clinic. It was sort of like an SOS call.
It wasn't a day that I had uh my actual
standard clinic and they were like, "Oh,
something's really wrong with the
patient. Can you see them quickly?" And
after some time and reviewing the chart
for hours, I I personally also couldn't
figure out what's wrong with the
patient. So I said, "Okay, I'll do a
biopsy." Anyway, fast forward and after
developing some type of a rapport with
the patient, which in my case was only
about a week and a half. I then
contacted the people that had known this
patient for six, seven months. And I
said, "I'm really sorry to say this
because I know I'm like the new person
in the group, but I have combed through
this patient's chart for hours and hours
and thought about this for days and
days. And I think this may be
fictitious." And in my own clinic, I use
uh DAX all the time. I use an AI scribe.
And for some reason, whenever I was in
the room with this patient, I thought, I
shouldn't use this. And it may be making
them uncomfortable or introduce a type
of bias. And I remember when our group
had written a paper on machine learning
um and pain sketches on amputees and
phantom limb pain, there was some
concern that like maybe some patients
weren't writing uh their certain pain
postoperatively because they were almost
quote unquote embarrassed to write that,
that they still had pain or that they
were this debilitated. And I just
wonder, you know, what what what
stewardship we should sort of um think
about when we're putting patient data
into there because before we ever
entered into the patient's chart
facitious, we each had individual
conversations outside the room. We made
sure that we didn't tell anybody at the
nursing station. We made sure we didn't
tell the residents and we made sure the
patient stayed in the hospital and felt
very safe for at least two like almost
two weeks before introducing the idea of
psychiatry even though the patient had
an outside psychiatrist. And I think
that is very um alarming for patients.
And something that um was interesting
for me was our social worker actually
pulled up the patient's gateway records
and saw that the patient was checking
their gateway every 30 seconds.
So because it's released immediately to
the patient, what does that do to their,
you know, psychosocial state? Um and
what is our responsibility then as
physicians that have these metrics that
we have to document everything within 24
hours? So
>> yeah, I think that's that's super
interesting and and it's a challenge,
right? Like I think adding to the list
like the stewardship around um around
you know patient data, how that will
ultimately be used and how you know a
scale that scores certain things. I'm
thinking of, you know, other sensitive
topics like things like obesity, etc.
that like, you know, then get kind of
plugged in the chart in a way that like
a human wouldn't necessarily write it or
deal with it um in real time. And I
think we're kind of facing the I I
personally also have a fair amount of
patients now that will take their MRI
report put it into chat GBT um and then
come uh with you know the output and you
know it it can uh make the sort of
doctor patient relationship portion of
that um really uh quite the challenge
and so I think it's really important
that we think about how we as physicians
are also the humanistic side of making
sure that we incorporate um the human uh
interaction ction into AI um in medicine
and in surgery. So, I love that you
brought that point up. We have sadly
only three minutes because I have so
many more questions. Um and so, uh I
think you know getting at maybe we'll
just do kind of closing thoughts from
each person. Um thinking about, you
know, some of these concepts of like
what the responsibilities are, what the
policies are, and who who should kind of
be involved in these conversations. Um
and then we'll wrap things up. So uh
maybe we'll go in the same order. We'll
do Dr. Bitterman, Dr. Ali, and then Dr.
Duano.
>> Sure. Yeah. I think one area that I
am worry about and feel very passionate
about is that physicians and patients
are being too much left out of the
conversation of where AI advances in
healthcare are going. the kind of
discussion is uh increasingly being led
by
developers and I think clinicians and
patients should not feel intimidated by
the kind of the fancy tech words they
use. We have the knowledge of what is
going to be impactful. Um, and that's
actually the highest value
uh uh kind of uh expertise to bring to
developing and evaluating these models.
So, um my kind of final statement would
be like please get involved in this
research. Please kind of be uh kind of
if you're interested kind of start
collaborating with computer scientists.
A lot of them are really excited about
using these models to advance
healthcare. They don't know the right
answers to go after and it's just a huge
huge opportunity untapped potential to
direct the these advances to in the
right direction to make physicians life
easier and more importantly improve
patient outcomes.
>> Amazing Dr. Ali.
>> Um yeah and just to sort of echo that uh
what Dr. Bman said I mean I think we're
seeing now in survey uh responses that
you know majority of physicians are
using these EI tools. We see patients
coming into clinic using these AI tools
all the time. I think that ultimately
should be encouraged because that's
really going to lead to more engagement
by patients and more understanding of
their own own health care and we as
clinicians I think should just meet the
patients where they're at. Um as Dr. to
want to so be beautifully put it really
individualize that care for the trainee
for the patient you're meeting uh and
acknowledge the fact that we're all
starting to utilize these tools more and
more and incorporate that
acknowledgement into our everyday
practices.
>> Thank you so much and Dr. Tuano. Yeah, I
mean obviously I agree with everything
that everyone said and um I guess the
most succinct way that I can put it is
our ethical obligations in general with
utilizing AI are that um things change
and our eth ethical obligations change I
think especially as things um are sort
of um more developed and it's a shared
decision ultimately in the end with
yourself and the patient and it's just
really important that AI is used as a
supplementary tool and not the only tool
and it's used as a reasoning tool but
it's not a replacement for clinical
judgment. So, as a daily user of AI
myself, on the top of my group's notion,
it says, "Treat every patient like
family." And I'll leave it at that.
Treat every patient like family.
>> Thank you so so much all of you, the
panelists, for being here and for
sharing these insights with us today. I
feel like I I have some like real
takehomes that I'm excited about um and
eager to answer the calls that you guys
have laid out um and also more
knowledgeable about how AI is used in
surgery. So, I can't thank you enough
for being here. Um, for those who are
interested in in taking us up on that,
um, please follow, uh, or continue to
follow the HMS Center for Bioeththics,
reach out about the Harvard Surgical
Ethics Working Group. Um, look us up at
neurotchice.org
for the neurotchis accelerator. We do a
lot of this research. We also have an
upcoming seminar on June 12th uh, called
digital neuroch neuroch in your pocket.
So, we'll be talking about the specifics
of neuroch in this field. Um and so
please continue to be part of this um of
this journey with us about AI and
surgery. And so and also thank you to
the center for bioeththics and our uh
group and nelle for um putting this all
together. So thank you so much. Be well
everyone and look forward to continuing
this work.