What Rough Beast...? On AI's Potential to Surpass Humanity
Watch on YouTubeVideo summary
The event "All Too Human: How AI is Shifting the Ways We Make Meaning," hosted by the Burkeman Klein Center, brings together experts to explore whether artificial intelligence can surpass humanity and what such a possibility implies for human identity. The panel features Jackson Kernian from Anthropic, who approaches the mind as a computational system, and Ken Archer from Microsoft, who grounds his perspective in phenomenology and the concept of humans as embodied beings embedded in time. While acknowledging that AI acts as a disorienting yet clarifying tool similar to past scientific revolutions, the speakers emphasize that fundamental questions about intelligence and value require human judgment rather than reliance on models alone. They note that while current large language models reflect specific cultural biases rather than global diversity, both organizations are actively working to broaden their perspectives through diverse teams and ethical guidelines.
Significant concerns were raised regarding AI autonomy, data leakage, and the difficulty of defining reward signals in reinforcement learning, which can lead to unexpected behaviors. The discussion frames AI not as a static tool but as an organism that must be grown with robust guardrails, drawing parallels to safety protocols in nuclear or aviation industries. Philosophical engagement is deemed essential to break down academic silos and ensure that fundamental questions are asked, particularly given that humans have not achieved perfect ethical agreement among themselves, making "machine-human alignment" a complex challenge rather than a solved problem. The panel also addresses environmental and financial impacts, suggesting that building data centers in previously barren areas could yield positive outcomes for communities, while cautioning against anthropomorphizing models as children or friends to avoid cultural misunderstandings between Western liberal values and other frameworks.
Ultimately, the conversation concludes with divergent yet complementary views on the future trajectory of AI and humanity. One panelist supports a merged future where humans and AI collaborate to become better versions of themselves, while another assigns zero probability to artificial general intelligence, arguing that true general intelligence involves the unique human capacity to step back from impossible tasks. Despite these differences, there is a consensus that AI offers immense opportunities for expanding human power and flourishing if developed with respect for human finitude. The session underscores that while risks exist, history shows that previous technologies enabled human flourishing by overcoming cognitive limitations, provided they are managed with responsibility and a focus on ethical failure modes rather than just theoretical alignment.
Read the full video transcript
Hi everyone.
>> Hello. Hello. Thank you so much for
joining us today. Welcome to the
Burkeman Klein Center. Um and thank you
for like braving it um on a Monday. We
don't do that many Monday events, so
we're really excited to see everyone
here. Um
uh the Burkeman Klein Center exists to
make sense of the digital environment
and to ensure that it promotes human
agency, dignity, and genuine connection.
So, before I think I turn things over to
Alex, our executive director, and the
rest of the Oh, look, you guys formed a
line. You're so organized. The rest of
our um prestigious uh team here today, I
just wanted to tell you a little bit
about kind of what brings us here today
and the start of our event series, which
is called All Too Human: How AI is
Shifting the Ways We Make Meaning. Um,
as you might imagine from the title,
it's kind of charting how AI is shaping
our relationships, um, our s systems of
belief, our creativity, our most
intimate emotional experiences.
We'll be talking about AI in
companionship, AI and death, AI and
spirituality, and much more. Um, and
throughout this uh, event and throughout
all of these events, we really want to
hear from you guys on these topics. And
so I also want to introduce you to next
space which is a BKC designed AI
platform that is designed kind of to
keep you engaged in the um in the event
and what's happening in the room and
also to engage you really with other
other participants both here and um and
online. So you can have uh a private
chat with Berky who is our sort of
snarky but very knowledgeable AI
facilitation agent. So you can ask
things about um you know what
organization did they mention or like
what would you say Jackson's take on X
or YZ would be? Um and there's a group
chat if you want to connect with others
in the room. And also importantly, this
is where we will be getting questions
for the Q&A. So, um the short link is
right here. So, if you anticipate
wanting to um ask questions of our
panelists, make sure that you grab the
QR code, sorry, or um the short link.
and uh transparency next base run on
claw runs on cloud through Harvard's API
proxy and uh just so you guys know
you're participating in something a
little bit different for this for this
speaker series we're hoping to use um
your kind of stories in next space to
create a visualization of the concepts
we discuss through this series. So um
kind of like if you know does anyone
know mind maps you know those like
little visualizations. Okay. Yeah. Um so
kind of like a night sky if you can
imagine that in like we're talking about
AI and death you know death would be
your planet and maybe a sub theme would
be like legacy and then you know um
what's your name? You just raised your
hand. David. Um maybe the star would be
like David really wants to make a
hologram of himself and that would be
kind of the um the little star. So just
to give you a sense of where that's
going. Uh so to kick things off, if
you've joined Berky, we'd love to kick
things off with a question which is what
if anything do you no longer feel
comfortable doing without AI?
So, just a little brain starter for you.
Okay, with all that throat clearing
done, let me hand things over to our
executive director, Alex Pascal.
>> Thank you so much, Jess. Um, and also
just a huge round of applause for Jess
and Isabella and so many on the Burkeman
Klein team who have made this speaker
series possible. It's really phenomenal
what they've done and we're unbelievably
excited for this year. So, thank you,
JESS.
SO, let me do my second thank you to um
the Public Culture Project, um Dean Sean
Kelly, uh Ian Corbin, who we're just
like over the moon to collaborate with
this year on this speaker series. Um and
as we were talking over the summer uh
about what we were respect mutually
respectively interested in, it became
clear that we are grappling with the
exact set of same questions that are
truly fundamental to who we are as
people. And we said why don't we just do
this together? Uh which is a feels like
a novel concept for Harvard, but like
we're still excited to explore it. Um,
and we are just like unbelievably happy
to be able to do this with you and to
bring unbelievable speakers like you're
about to hear from uh together uh to
really explore these fundamental
questions. So, this is Burkeman's 30th
anniversary and um as Jess said, the
Burkeman Clean Center's mission is to
make sense of the digital environment
and to ensure that it promotes human
agency, dignity, and genuine connection.
And so mindful of that mission and of
our legacy of really exploring how
digital technology has shaped society is
impacting people for the last 30 years,
I can kind of think of no bigger set of
questions than to ask then how is AI
changing us and how is AI changing how
we think about ourselves as human
beings. So over the course of this year,
we're going to be exploring how how AI
is affecting things like uh death and
mourning, our relationships, our faith
in spirituality, our capacity for
empathy and meaning making, love and
attraction, creativity and cognitive
capacities, the things that make us us.
And we're going to start um with a
particularly provocative uh panel um
because the conversation around AI, in
case you hadn't noticed, has gotten
particularly boring over the past 2
weeks. So uh we just figured we would go
for the jugular and ask what if AI
surpasses humanity? What does that mean
for us? What does that mean for how we
think about ourselves? Do we even matter
anymore? Um and so really just excited
to have this conversation today. Uh, and
without further ado, let me pass the the
mic over to Dean Sean Kelly uh to
introduce the panel. Thank you.
>> Hello. Hi everyone. I really want to
welcome you. My name is Sean Kelly. I'm
the dean of arts and humanities and I'm
a professor in the philosophy
department. And I am so excited for this
event. I'll tell you how excited I am. I
arrived back last night from Singapore
and so my head is still somewhere out
over the Pacific Ocean, but I am here
because it's really going to be fabulous
and uh and I'm grateful to you all for
coming. Uh as Alex said, that the Public
Culture Project is a co-sponsor of this
event and this series over the course of
the year. So grateful to Burkeman Klein
for for for working together with us and
to my team uh Ian and Meera here have
done a lot of work with you guys over
the summer to put this together. Um the
public culture project is devoted to
asking hard questions really difficult
complicated often contested questions to
bring people who will sometimes disagree
about them together and to put a
humanist at the center of the
conversation. And I I don't think we
could have a better version of that here
today. Um so I'll just say a few words
about the topic. As Alex said, you know,
knowledge, intelligence, learning, those
issues have been um they've been in the
media a lot in recent weeks and months
and years. Um but it's not human
intelligence we've been talking about.
It's some nonhuman artificial
intelligence. And um you know it's it
may be an intelligence that's super. It
may be an intelligence that's um
artificial in some sense but in any case
its presence is is going to change
things in some way that we don't quite
understand. And there's lots of views on
how that change is going to happen. Some
people say that um because of these
developments human beings are on the
verge of extinction. We'll be obsolete.
Uh you see that in the newspaper over
the weekend. Lots of AI companies are
saying we really have to slow down our
development. Um other people think it's
going to go in the other direction. Huge
abundance of resources and we're going
to all have free time and it will be
paradise. I'm not an expert on this. Uh
I'm certainly not an expert on the
technical details and I'm so grateful
that we have some people who are experts
in that. Um, but what I do know is that
it's going to take a very human kind of
intelligence to develop good judgment
about how we want ourselves to relate to
the things that we're that we're
building. Um, and the human intelligence
that I'm thinking of is the intelligence
that allows us to ask f fundamental
questions. I don't really think these
kinds of questions are going to be
answered by AI. I think we have to take
the responsibility to try to think them
through. I'm thinking of questions like,
well, what is intelligence in the first
place? I don't I don't think that's an
easy question to answer. I don't think
it's a question that a model trained on
everything everybody has ever said is
really in a position to answer. What's
special about human beings? What's
valuable valuable about ourselves? What
kinds of work do we want for ourselves
and our fellow human beings? What kinds
of work should we hand over to machines?
And I think these kinds of questions are
the questions that humanists have asked
at least some of them for a very long
time. Poets, philosophers, historians,
uh, novelists. I think this is the
domain of the human where questions like
this have been explored and and we
really have to bring our our judgment to
bear. The writer and the biochemist
Isaac Asimoff once remarked that the
saddest aspect of life right now is that
science gathers knowledge faster than
society gathers wisdom. Um, and this in
this mismatch in in our moment threatens
to veer from the sad to the dangerous.
At least it's possible that that'll
happen if we don't take responsibility.
That's why I'm so grateful to think
about these questions in public with you
in collaboration with the Burkeman Klein
Institute and through the public culture
project. And I'm really grateful to have
our guests here today. We have um three
of them. I'm going to introduce our
moderator Moira Vigel who's from the um
from the department of comparative
literature. Um and I'm going to let her
introduce Jackson and Ken. I'm just
going to say upfront though, I'm really
delighted to have Jackson back who is
class of 2012 and a philosophy
concentrator and I'm proud to say a
former student of mine in some capacity
and um and I'm really excited to have
Ken here too. Moira will tell you about
them. Moira Vigel is a scholar, writer
and founding editor of Logic magazine.
She serves as an assistant professor of
comparative literature at Harvard
University and a faculty associate at
the Burkeman Klein Center for Internet
and Society. She received her PhD from
the joint program in comparative
literature and film and media studies at
Yale University and she previously held
fellowships at the Harvard Society of
Fellows and at data and society. Her
research focuses on the history, theory,
and social life of media and
communications technologies with
particular emphasis on transnational
digital platforms, e-commerce networks
between China and the US and the
cultural history of critical theory in
tech industries. So, please join me in
welcoming Moira and who will introduce
Jackson and Ken.
All right. Did we all turn our mics on?
>> I'm up.
>> I had two quick orders of operations
questions. So, we'll go for about 40
minutes. Is that right, Jessica? 3540
and then open up to the room. I figure
you all probably have your own questions
you want to ask. So, I will try not to
abuse my prerogative as the as the
moderator too much. Uh
Jay-Z, did you want to ask a question
right now or later? Later. Okay, sounds
good. Um well, with that out of the way,
I'm so pleased to have a chance to be
here with you all today and to interview
uh these two fascinating visitors who
have kindly flown all the way in from
California to join us. Uh I'll do it in
the order that you're sitting next to
me. Jackson Kernian is a member of
technical staff at Anthropic where he
leads the human feedback program. His
work focuses on training and evaluating
fuzzy rewards signals in reinforcement
learning. He co-developed constitutional
AI to automate training harmless
language models through self-correction
and contributed to early research on
redteameing language models for safety.
As Sean was alluding to, he studied the
philosophy of mind as a Harvard
undergrad and received his philosophy
PhD from the University of California,
Berkeley.
Then Ken uh to our left, Ken Archer is a
product leader in responsible AI at
Microsoft and previously led responsible
AI at Twitch. He is also a doctoral
researcher in philosophy, cognitive
science and AI at
>> link. As a German speaker, I get tricked
by these Swedish ones. What is it?
>> Link shipping.
>> Link shipping uh university working on a
philosophical account of how science and
technology emerge from structures of
human cognition. Um I'm sure we're going
to delve into sort of heady
philosophical topics soon. Uh but since
we have an audience that I bet has a
number of students in it, uh is that how
many folks here are students of one kind
or another? Yeah, it's a lot. Um I
thought maybe I'd start just by asking
you a little bit about your own
trajectories and how you came to do the
work you do. Uh Jackson, would you mind
starting us off?
>> Yeah. Um thanks so much for having us.
Um let's see. So I mean I was first
interested in philosophy uh as an
undergrad. I really wanted to figure out
how the mind works. Um, and that's a
pretty messy question, but um, that was
my way in. I I really wanted to
understand like consciousness,
perception, the way our cognitive
faculties interact with our perceptual
faculties. Um, in in grad school, I had
a a reading group um actually when I was
visiting NYU uh and and those of us were
interested in AI. This is 2018 and so we
had a reading group in philosophy of AI
and we read um the neural scaling laws
paper that some of the founders of
Anthropic wrote while they're still at
OpenAI and I just remember debating you
know can next token prediction bring us
AGI and some people said absolutely not
and I found myself saying why not you
this seems like um if you can predict
what what's coming next that seems to be
sufficient for um or at least at the
right scale um generating ent. And so it
it really sort of caught my attention
there. And then when I failed to become
a philosopher, I I pivoted to my
secondary option of a web developer and
um and found my way to anthropic. So
that was it was really that that reading
group. Um so that brought me in.
>> That's so interesting. Um and Ken,
>> yeah. Yeah. when I was um getting my
master's in philosophy, I was also
working in software at the time and I
had my own startup and uh I was at a
school that had a lot of expertise in
phenomenology. We were reading Hussel
and Haidiger and Hannah Arant and I
started to see some overlap uh here uh
between the arguments they were making.
For example, uh Husel's last book is
called the crisis of European sciences.
And he argues that basically the way
science develops is it moves further and
further away from our everyday
intuitions. And as that happens, there's
this growing temptation to displace
responsibility onto uh the models of
science themselves. And it occurred to
me I feel like I'm seeing that every day
uh in in technology uh responsible AI as
a field was just getting started. Uh and
so that's when I entered kind of the
field of responsible AI and went to
Twitch and Microsoft.
>> Thank you. Um I already have follow-up
questions I want to ask um about the
bios you've given but sort of to orient
us in the conversation you know this
event is framed with a quotation from
Yates's famous poem the second coming um
and then by this question of surpassing
humanity and I actually wanted to ask
you you know easy little questions uh in
in reverse order to those things um to
sort of ground us and the first question
I wanted to ask is coming out of these
bios maybe what are your orientation
points or what traditions do you draw on
to think about what a human is and thus
what it would mean to surpass humanity.
Um, you know, I can imagine religious
traditions, um, philosoph different
philosophical traditions, materialist,
rationalist. Um, anyway, I could I could
come out with my own. Uh, but I'm
curious to hear what yours are. I gave
you a minute there to think about it.
So, what's a human?
I I I I definitely think back to my
intro psych class with Steven Pinker
where where the mind was treated like um
at the time at least and I think this is
still popular the massively modular
theory of mind. So the
>> mind is just a a bunch of small little
modules that are feeding into one
central processing system. Um I found
that pretty influential. Of course like
that picture is complicated and uh in
all sorts of ways but um I mean the if
you're so my background is in in sort of
philosophy of mind uh there different
sort of foundational theories of mind
but I was always a functionalist um so
the mind is what the mind does um some
things are defined by what they do like
a light switch is not defined by what
material it's made out of it's defined
by the role it plays in a system and so
I was always yeah uh attracted the idea
that the mind is like a computer. Um,
which is maybe convenient for uh for the
work I went into, but um that's the I
guess the intellectual history I I I
subscribe to.
>> But I still didn't hear you say human.
What makes a human mind rather than a
monkey mind or a a reptile mind or some
other kind of mind?
>> Um well, I I mean I think importantly
the uh you know the human mind, it's not
different in kind from a reptile mind or
or monkey mind. Um
uh I don't know at this point maybe Ken
can tell can tell tell us what he
thinks.
>> Yeah.
>> Um well uh so when I want to define
something um you know I generally want
to think in terms of you know what's its
genus and then what's its you know
specific difference make a distinction
between other things
>> and you know we do that all the time in
certain fields that you know like botany
or in biology. Um, and what's good about
that approach to defining things is it
forces you to account for sort of your
entire phenomenal experience of like
what's the best way to make distinctions
between different kinds of plants. For
example, in terms of humans, I think the
best uh definition, the best distinction
that has been offered is that uh humans
are rational animals. Um uh now that
then sort of just moves the question uh
you know to well what does rational
mean? Um and so uh and this does get to
my views on uh whether humanity can be
surpassed. You know, when I think of
what it means to uh reason, you know, I
think it involves um uh being able to
think, being able to think uh and
updates one's thoughts uh in an open
world.
>> Um so the Greek word would be uh dianoa
for this.
>> Um and what's interesting about dionoi,
it's about thinking through. Um, so
you're not just sort of taking in
sensory input, this throng of input and
constructing things in your head. We're
already in the world.
>> Uh, we're in the world making
distinctions within the world. And um,
uh, this is basically thinking. Um, and
this is what uh, this is our our best
definition that we have so far. What it
lacks is precision. It lacks the
precision of mathematical physics. Um
but uh but I think it's enough to get us
started and to have a good conversation
about whether AI can um can surpass
this.
>> Well, I hear the seeds of disagreement
between you two already, which makes me
makes me excited. Uh but before I try to
draw that out of you both um to this
question of the second coming or this
moment of um
unease disorientation. I was actually
remembering the poem the second coming
this morning while brushing my hair and
thinking about coming to this event and
thought of that line where he says the
best lack all conviction but the worst
are full of passionate intensity. And I
thought what a good excuse for lacking
conviction like for not being sure what
to think at this moment. I will now
quote Yates. Um,
how would you characterize
the role of AI if we agree that we are
living in a moment of some unsettlement
>> uh in the broadest way?
Where do you see AI
fitting in that or intersecting with it
and how does that motivate your work?
>> Yeah. Well, um,
one thing I actually start with the
recent history, which is to say that,
uh, you look just at how the internet
has changed society, uh, universities.
Um,
>> uh, it sort of there was all sorts all
sorts of upheaval that I feel like we
still haven't dealt with. Um, like I now
we're talking about all these, you know,
possible regulations for AI and we never
regulated the internet. Um, or at least
not to this level that I would have
liked to have seen. Um and uh it I I
think of AI as similarly changing our
epistemic institutions. Um so um maybe
in previous eras the
>> technology has sort of pushed us towards
extremism or um or attention uh
attention optimized economies. Um AI
will push our epistemic institutions in
new directions. I don't think we figured
that out yet or we don't sort of know
what directions that'll push us in. But
um it will force us. I think about like
knowledge work. Um it's this sort of
amorphous idea that I don't think I
think AI is going to make us actually
try to figure out what is it we were
doing all along? Why were we having all
these meetings? Why were we writing
these memos? Um what what were we doing
at work before? um and not we weren't
always being efficient at what we do but
AI is going to make us um look look in
the mirror and say all this intellectual
labor what were we doing all along and
um I'm not sure we'll like the answers
but uh it'll help us figure out what
we're doing
>> yeah I mean I think AI is disorienting
um I but I would say that it is
disorienting in the way that so many of
the great scientific achievements are
disorienting And I do regard AI as a
great scientific technological
achievement. Um, mathematical physics in
the 1600s was incredibly disorienting.
Uh, it disoriented our understanding of
our place in the universe. Um, you know,
we we think it's just math now, but uh,
in fact, uh, the discovery of
non-ucukitian geometry, it's incredibly
disorienting for people. Um, it's it's
now sort of become tamed. Um and there
have been subsequent obviously you know
relativity relativity theory incredibly
disorienting and I regard uh AI
similarly as a great scientific
achievement um that can make us hard to
make it hard to find well what is what
is it about all the experience that we
bring to the world that's at all unique
is there anything unique about all the
experience we bring to the world and
these these previous sort of revolutions
would say raise those same questions and
I think we're that having raising
similar questions because we're having a
similar type of scientific achievement.
>> Have you had moments and we can also
reverse direction so you're not always
you don't always have to be first. Have
you had moments
>> that was great
>> even as someone who uh who works on the
technology obviously spends a lot of
time thinking about it knows it very
well in certain ways. Have you had
moments where you've been surprised or
felt disoriented yourself um by
something a generative model or an agent
or something has done?
>> Yes. Um, I would even pause it to say,
and I think this is something Jackson
and I would wholeheartedly agree on,
even someone who claims that they are a
deep skeptic, um, who regards this as
just a fancy spreadsheet, um, that who
says they have not had that experience
isn't really being intellectually
honest. Um, so it is disorienting and uh
things that we um I mean when we use
language we're we're used to language
being used uh by other speakers,
speakers who are responsible for what
they say. Um and so then here when we
have this language,
>> it it's natural to think there's some a
responsible speaker behind it,
>> we we can't not
interpret language use that way and it
can be very disorienting. Um I use I
mean I use AI all the time as does
Jackson
>> and um so it's it it uh is a multiplier.
It gives me leverage in my work. Uh it's
a multiplier just my personal life. Um
and
>> can you give a concrete example about
work or your personal life whichever you
prefer?
>> Uh yeah. Well um so I can give a um I
mean I think we all have um I I can
definitely talk about uh in my
philosophy work
>> um it helps me clarify my ideas.
Um, I definitely feel like I get to to
the point if if you just think about
some area where you think you probably
know as much as maybe a couple dozen
people, you know, like in physics for
example, um, then you I start to then
experience the swap, the generalities.
But for so many areas where I I think I
have an idea, but it's sort of a um a
new area for me, but that lots of other
people know about. I find that I can
clarify my thoughts. Um and uh with with
AI um but then you know like in my work
um yeah I mean we are we're trying to
build software that people are going to
use and um for example being able to
generate prototypes uh with GitHub
copilot that we then put in front of uh
customers to validate our understanding
of uh of their problem. is is goes to
this. I was teaching a class at TUS this
morning on AI and design and it was all
about this design thinking and how you
can use AI um uh to do more design
thinking in your work.
>> Yeah.
>> Yeah. For me there I think different
phases of using AI in my life. Um you
know there was the chatbot moment.
>> Um but I think I definitely felt more
disoriented when they started doing my
job. Um so I mean there's the sort of
agentic uh coding moment. Um I think
there's one specific uh experience I had
when so I was I I work with a lot of
data sets and I was getting Claude to
basically prepare and process a data set
for me. um it's launch launching jobs to
our cluster and um uh so you know sort
of what I do is I like at a very high
level I say all right like here's what
we're trying to achieve here are like
three or four things you want I want you
to check when you've checked that go
launch these jobs and tell me when it's
done um so uh one afternoon I was doing
that and um the agent it ran into like a
permissions problem um so it it it was
trying to launch a job and could
couldn't do it and so then I went to the
documents to figure it out and it it
found a loophole basically. So it's
like, oh, I can actually I can launch
this job. I don't I'm not technically
allowed to do this, but if I like launch
it with this special um special code
basically, then I can get away with it.
And so it's it started doing that and
then like a supervisor agent was like,
"No, no, no. You can't you can't launch
the job. That that actually is a that's
a like that's a break glass code. You're
not supposed to use that. Go ask for
permission and go check with Jackson."
So, um, I that stuck with me because it
was such like a human moment of like the
agent wanted to get the job done. It's
trying to find shortcuts to do it. Uh,
and has to like work in this weird
social structure of like I guess I'm
here to all to give ultimate permission,
but there's this agent supervisor that's
the AI that's also giving permission.
And it's it's a weird new world where
we're sort of all um trying to
coordinate with one another. Um, but I I
haven't had software before like ask for
my permission in that way.
>> I have. Then I'm going to draw out the
disagreement, I promise. But have there
been moments when you have surprised
yourself in something you've done
with AI or felt drawn to do and maybe
this is more with the chatbot interface
than with something personal or excuse
me than with something more technical
like working with lots of data sets. But
has there been a moment when you as a
human,
you know, to give a totally hypothetical
example that I surely have never done,
>> you know, in a moment of frustration
open, you know, my clawed window or
whatever it might be and said, "Oh, my
mom is driving me so crazy. Like, can we
just what should I do about this?" You
know, I was surprised at myself for
doing that. Um, I mean, I never did
that. Uh but uh but has there been a
moment where you as a human being have
like turned to AI to do something in a
way that perhaps six months earlier or
three months earlier you might not have
expected?
>> No, I've definitely done that in ways
that I've regretted. Um so uh I I once
uh I mean I I've used AI
um uh um to review papers
>> um you know for a conference
>> uh and immediately a afterwards I felt
icky about it. M
>> um you know just aside from the sort of
what I had committed to do
>> um you know I just these people had put
all this effort into the paper
>> u and I spent you know about 10 minutes
>> um so u at the time it just seemed like
oh I I made a lot of excuses you know
this this is actually better than the
reviewers that I typically see from
other reviewers so even if it's not as
good as my general review And you know,
I talked about thought about my time and
what's a good use of my time, but then
>> I uh regretted it pretty quickly.
>> Yeah.
>> Thank you for sharing that.
>> Yeah. A a couple instances
I think about um
uh Yeah. So there was I had like a
dispute with some neighbors about like
uh neighborly uh I live in I have an HOA
and and um and so uh you know it's
actually they're very the cloud is very
very useful at like reading through lots
and lots of stuff to help me figure out
what the rules are like this is sort of
this the sort of cognitive offloading
that um you know it's not my job to to
know the the regulations of San
Francisco but cloud was very very
helpful with that um and and part of
what was very helpful It is an emotional
like your home is a very emotional
thing. I I often talk about Claude being
like an emotionally emotional
translation machine um where I can type
in my angry thoughts and Claude will
spit out a uh the the polished non-
angry version with my mom.
>> Yeah. Yeah. Exactly. So um so there's
that. Um uh uh I also think there um I
have like I keep a journal and like have
some notes. So, I I have a project where
like I keep just a lot of like life
thoughts and um I've had some
surprisingly like thoughtful engaging
conversations with Claude about my uh
about my journal basically. Um and uh I
think I I initially now it's so second
nature. Um but um you know I've I've
definitely like changed my mind about
things because of conversations I've
had.
>> Yeah, that's really interesting. Um I
want to draw out when we I was asking
you about your points of sort of
departure orientation. What I was
hearing uh was that Ken has a vision of
the human. Um you talked about thinking
and reason sort of logos but also about
what I would think of as throness or
embeddedness in the world. Uh
>> that's right. I found myself thinking of
Ysef Eisenbal's famous definition of you
know a human is someone who was a human
mother or something you should have
thrown already into this situation which
is informed by his reading of a rent as
well as psychoanalysis. Um so I hear you
saying something um about
mortality and embeddedness and
existence. Uh and then I hear Jackson
talking about um
modularity and sort of computational
theories of mind which also I mean I
took a big Steven Finger class when I
was in undergrad too but it was longer
ago but I remember that language has a
lot to do with it. Uh but so I guess I
was curious to hear from each of you
first of all what you think about the
other one's point of view but second how
does that perspective inform what it
would mean for AI to surpass the human
or like what is the benchmark of the
human that we're measuring against or
something. I mean maybe it's not even a
meaningful question if you start with
the embeddedness perspective but you
tell me.
>> I mean I think large language models
work because language is amazing. Um so
uh language is how we express our um you
know our our lived experience in the
world.
>> Um but the lived experience comes first.
Um and you know our our our experience
in the world is fundamentally temporal.
Um so it there we we're we don't p we
don't live in these durationalist points
in which we have this throng of sensory
input coming at us that we sort of
imagine. uh we're fundamentally living
in time and uh so there's this idea of
throness um uh you know into experience
and this just simply is consciousness um
uh in for me and so in our experience of
the world we experience the world uh as
being coherent we don't experience this
incoherence that we then structure in
our heads we fundamentally experience
uh the world as coherent as being a
world of things. You can't be conscious
without being conscious of objects. We
fundamentally structure our experience
as objects that change in time in
different ways. And this fundamental
formal structure of experience of
objects that change in certain ways or
holes that have their parts is the
experiential basis for subjects that
have predicates it. It get expressed as
language. And that is why language is
just so amazing. And um you know there's
there's other formal structures there's
structures between uh objects. And um so
this is basically hoser and haidiger uh
is what I'm describing here. Um and uh
but only a being with that experience
can feel accountable to that experience
um can feel responsible to that
experience. Artificial intelligence
takes that language um and is a derived
intelligence. Um but it's because
language is so amazing and expresses our
lived experience in the world. I mean I
I mean that's why u that's why we have
the experience of fluency that we do uh
with large language models. Um this was
not a possible before because you
couldn't have models with such high
dimensionality. you'd run into the curse
of dimensionality and a lot of technical
achievements in the history of neural
networks mainly transformers um overcame
u uh by through weight sharing by having
this symmetries um uh overcame this
limitations of dimensionality. So
anyway, this is why I uh am excited
about large language models, but it's
also why what I think we need to have a
very humanist grounding of artificial
intelligence. Um yeah,
>> small question that's going to sound
weird because you said the h word
hideer. Um
>> think of right like language speaks. Is
language human?
>> Is language
>> because you said humans start with this
embodied sensory experience. Is language
itself human? Yes.
>> Okay.
>> Yes. Um I would say this is a um you
know this is not a mathemat this is not
a precise distinction.
>> Uh but you know in the concrete sciences
where we're describing things and making
distinctions this is the best
distinction uh that we have right now.
Yeah. And do you think that the outputs
of large language models
are language in the same way that
language is spoken through a human or by
a human is language or is it like a
simulacum of language? It it is a
simulation, but
>> um the the problem when you say it's a
simulation, then it can start to seem
like like a trick. Like the idea of, you
know, I'm I'm pushing a um a car down a
hill and saying it's driving. It's not
that. It's not just a trick. Um I mean,
this is I mean, this is a great
achievement of science and technology.
Um it's um you know just the development
of artificial neural networks. So, a lot
of these innovations, a lot of these
discoveries came decades ago and we're
just benefiting from them. Uh, now
because we have the com the compute, uh,
we have the the the text uh, written
down. Um, yeah.
>> Yeah. I'm thinking about your initial
question about sort of um, trying to
pull the threads from each other and I
uh, this point of from Haidiger of
throness. Um I my first philosophy class
with Shawn was on and Haidiger's being
in time. Um so I am familiar with it and
I and I do think there's something to it
with language models because they they
really do they're born into existence a
new every context window. So um so these
language models they have a context
window of so many tokens that um uh you
know they have all this background
knowledge stored in their weights but
like from their perspective basically
they they pop into existence from the
first token of the context window until
they're done generating and um you know
they really this is very much unlike us
so so I every moment of my life I'm
bringing with it my all my experiences
that have crewed all along maybe it's
just one giant context window I don't Um
but uh but I'm not being reset every uh
every conversation like like these
language models and it does I think
quite change the the character of what
we're interacting with. These are in
some ways um you know they are
interchangeable minds that you're
loading with context every every time.
Um so uh we're not like that. Um I I do
think it it it raises these interesting
questions about though like what you
know uh how do I individuate different
AIS um and how do you relate with it
like you know if I have a corpus of
notes that I've developed with my AI and
they're loaded in the context window
every time is am I interacting with the
same entity every time it's loading up
the same memories and um possibly um but
but just to your initial question about
sort of what separates us from the
machines I think this continuity of
experience is is totally different with
with the AI
>> and um
there's sorry there's so many different
questions I want to ask you and I'm
mindful of time because I will open it
up to questions in not too long. Uh one
thing I wanted to ask about is
plurality of languages and conceptions
of the human. Um there's been all sorts
of work of course on how you know
there's more training data in certain
languages than others. This may affect
outputs in different ways. I think
there's also like this emerging work on
calling for hermeneutic evaluations or
like interpretive evaluations. Uh
thinking about how deployment in
different kinds of cultural context can
be evaluated when maybe user in fields
that are more contested or where there's
less consensus across cultures. I guess
uh when we think about these questions
of language and context and the human,
I'm curious to hear if like how if at
all you think about
cultural pluralism, both different like
literal empirical languages and also
different conceptions of language and
how those can be preserved in kind of an
industry around a set of technologies
that are so highly concentrated.
>> Yeah.
>> Yeah. I mean this is
>> I I would say these concerns are at the
core of responsible AI. I mean whenever
you any type of a probabilistic model
what you're basically dealing with is a
data generating mechanism.
>> And um you know and large language
models are are no different. And so the
data generating mechanism
um you know it it is based on a certain
data set and the data set that you have
in place sets the the the reference
class and that that reference class of
properties for that domain you know it
ends up slicing the world in a certain
way.
>> Um and now u that can lead to a lot of
disparities. Um at the same time uh
probabilistic models including uh neural
networks also have a lot of promise to
help a lot of people out.
>> And so the challenge is figuring out um
yeah h how to deliver the benefits of
probabilistic models up to and including
neural networks. um you know without um
having everyone fit into a certain data
generating mechanism and just simply
saying well that's how the model works.
It's smarter than we are and so you know
we have to trust it.
>> Yeah I here I think about I mean there's
two properties of the models that um
come to mind. So one they're extreme
malleability. So like you know they are
the products of an extremely
cosmopolitan uh sort of corpus of
knowledge. Um they they can probably
speak to
>> the principles of Islam better than me
or or uh or talk about um German culture
better than I can for sure. Um and so uh
all these values are embedded in the
system at the same time. So that's sort
of this malleability part. Um there is
like a a set of default values. uh and I
I think um we don't want to say that the
model is just uh like valueless. It it
comes with a set of um you you know ask
it ask it questions without assuming any
context and it'll it'll have opinions.
Um
>> and so yeah, where do those come from? I
I think
uh at to first approximation they just
come from the internet.
>> So like it has the internet's values. Um
uh I mean that that is an
oversimplification but um
>> Reddit might be more accurate.
>> Yeah, maybe Reddit um they do ALS but
they they do have um the values of the
company. So I mean anthropic we do
publish our constitution that we try to
tr uh train call to follow. Um that is a
the like an opinionated take on on what
kind of values the the system should
have. So, I mean, we are we're a company
from San Francisco that has, you know,
liberal cosmopolitan values and um like
we're, you know, liberal democracy,
liberal in the lower lowercase L sense.
Um and and so yeah, I think uh um it is
not representative of the world. It's
representative of a certain mindset from
San Francisco.
>> Can you think of ways and then I'll
maybe start to move towards questions,
but um I mean I think part of Oh, sorry.
You go on. I think probably Jackson and
I should probably both say that our
employers are very much trying to um you
know engage a broad range. You know uh
Microsoft today put out um it's uh
humanist AI uh code of conduct um for uh
uh public review for the next six weeks.
Anthropic has engaged uh people from all
over the world experts philosophy and
religion. We're trying to break out of
these bubbles. Um, but at the same time,
we can't break out of these bubbles if
we don't acknowledge that we're in them.
>> Yeah.
>> Do you think that question is any
different for open source models and
frontier models? Like, do you think
>> Well, the open source models I think
pretty much Well, we think they distill
from the the the big labs and so they
end up distilling our values.
>> It's interesting. Um, I think I'll open
Oops. Um, I think I'll open to
questions. Is that about right? I'll
maybe I'll ask one last question. I'm
putting the audience on notice that you
will get to raise your hands and ask
questions in just a moment.
>> Yes. Also, all of you on the internet.
Um
I guess I mean we've had a very wide
ranging conversation. Uh I feel guilty
asking about this because I think
everyone's tired of talking about it,
but I did want to ask about the pacing
or pause uh issue that came up over the
weekend. And I had also even before that
news been thinking of asking you about
what you are most worried about uh or
what you know what worry keeps you
going. Uh so maybe I'll try to meld
those questions together.
>> Sure.
>> And ask what would pacing or a pause
best be used for in your opinion or why
might it be needed?
>> It can.
>> Yeah. Yeah. I mean I the here's what I
would say is that there are two things
that worry that I think we should be
most worried about. Um and the second
gets to the pause. The first is that uh
we define ourselves and who we are in
terms of our tools. Um so each of us
including each one of you uh brings an
indispensable uh perspective and
judgment and responsibility uh to the
world that's irreplaceable.
And um Aon should not put that into
question. Um and were it to do that um
and tempt us to displace our
perspective, our responsibility
uh onto models, uh that that is the real
crisis. What that's what we want to
avoid I would say and that's why it's so
critical right now uh that we have a
more humanistic approach uh to
artificial intelligence. It's
challenging because, you know, in all of
our academic lives, uh, we're kind of
forced into these, you know, Snow's two
cultures, right? Either a very
scientistic culture or a very humanistic
culture that focuses more on critique
and interpretation. Uh, but, uh, that's
not what the moment calls for. The
moment right now calls for more
humanistic approach to science. And um
so the second thing uh that worries me
um it's not it it is not the capability
of models per se. Um I think the the
larger concern uh is autonomy of models
of which capability is only one part
because a model that has more capability
is more able u to have greater degrees
of autonomy. But there's other parts to
auton there's other elements to autonomy
because autonomy is more of a system
level concept than a model level
concept. So in cyber security if you
look at the um uh top three concerns
about large language models that you get
from the primary AI cyber security
organization called OASP. They are in
order indirect prompt injection attacks,
sensitive data leakage and excessive
agency. that's giving it a lot of tools
with no guardrails and uh so rogue AI is
not there in the top three because for
them they see uh unconstrained
uncontrolled autonomy as what to be
worried about and the solution is
greater controls greater monitoring and
so if we were to do a pause I think
that's what we should focus on in the
pause is we we know about controls from
other safety
security
u you know sensitive industries like
nuclear like airlines
um we don't know enough about the types
of controls and monitors that are needed
uh here uh and and that's where I think
uh we need to do a lot more research
right now
>> yeah when I think about things I'm
worried about um probably at the top of
the list is misuse of of AI so you know
the the models themselves and we talk
like forite for cyber is a good example.
It's a dual-use technology. It can be
used to protect systems or it can be
used to break into them. Um and uh I I
want to make sure if we you know have
some extra time from pausing or pacing
ourselves um making sure we build those
guardrails in a robust way so that
misuse doesn't happen by accident. Um
and as new new capabilities come online
that maybe we weren't even anticipating.
Um so that I think that's one of the
main worries I have. The other worry I
have is just actually just about I think
like conceptually the the hard problem
about building these things is you you
we're we're growing them. They're
they're more they're more like an
organism or um uh like a living thing
than a like a cold static technology.
Um, and so every RL training run is
you're you're starting from this infant
that you're growing in these weird ways.
And you only ever see the public only
ever sees the like the final product
when we've like fed it the right sorts
of nutrients and pushed it in the right
directions. There's all sorts of weird
messed up models that you don't see. And
I would want to spend the time uh just
making sure that those reward signals
that we're we're feeding these things
are are specified in the ways that we
want. Um it is very very common for us
to think that we've specified the reward
in a way that pushes it in a direction
we expect and then we share it with some
collaborators and they're like well did
you notice that it it's doing this weird
behavior that that um you you didn't
notice and I think we've gotten better
at noticing those things but um I the I
I would really want us to basically take
the time to to um get better at uh
measuring the rewards that we're that
we're specifying in RL.
and maybe bringing more people on board
to see to be able to do review or to
notice those things about your
metaphorical children uh that might
escape the parents loving loving gaze.
Uh all right, with that I'm going to
open it to the floor. I see a ton of
hands, but Jay-Z, you don't want me to
ask you the first question?
>> I feel like I'm on an infomercial. Why
yes, I have a question.
>> I was told to ask you the first
question. So, Jonathan's a train. I
appreciate that. Welcome to this edition
of the Truman Show. Um I'm Jonathan
Zitrin and uh helped start the Brooklyn
Klein Center. Uh I keep thinking about
this uh quote which I've interpolated
for today's era of in the land of the
unsighted the oneeyed person is regent.
And um thinking about tying that to the
best lack all conviction which to me is
this is a very humbling moment. I found
I feel deeply confused and I'm doing my
best to draw upon whatever expertise
I've accreted over the years in digital
space internet stuff to apply to this
moment. And I imagine you all are doing
the same as are people in the room. And
I'm just curious if you might think
aloud quickly like how's that going? I
mean if we're doing philosopher bingo I
think we've had Haidiger Husurl
um Hannah Arant
get Leo the 13th. We didn't get there
though.
>> You know the night is young Steven
Pinker Thomas
Are these reference points
of other folks who've tried to make
sense of our weird reality, are they
helping you all here or are you feeling,
you know, how many eyes do you feel you
have right now? And how much have you
changed your view on some of the deeper
and more enduring questions as a result
of whatever success such as you think it
to be the advances of AI have had in the
past say three years. Thanks.
Yeah. I mean um I think that
um
we need the academy and we need
philosophy in particular right now at
this moment. Um I think uh
unfortunately
uh it it is less common today than it
used to be for when society faces a
crisis um to call the philosophers and
to ask them um you know how often do you
see a prominent philosopher professor
writing something in the you know New
York Times on
>> I have this vision of 911 what is your
emergency and what is an emergency
anyway?
>> Exactly. Exactly. And um so we're, you
know, kind of getting back to the silos.
We're kind of pushed into our silos in
the academy. Um the incentive structures
u push us into narrower and narrower
publication silos,
research silos, and it makes it harder
to ask the fundamental questions. Um and
I I think right now um there's a real
need to bring philosophy that asks the
fundamental questions to bear but not
from an attitude of just simply critique
um but of engaging science and
technology
uh on its own terms. Um and the
siloization
uh mitigates against that.
>> Yeah. I think the question of yeah I
mean do we know what we're building? Um
and I I I think the answer is like sort
of and um the only way we have um we had
some like founding principles that we've
talked about anthropic and one was um
like I'm now going to mangle this but um
like crossing the the river by touching
the stones.
>> Yeah. Yeah. Yeah. That's it. Um, so you
know, you you there's this raging river
that you you want to get across and you
know, you uh I'm not sure how you're
doing, but you just reach out one bit by
bit and find your find your way and find
your footing. Um, I I do think that, you
know, the the best way to figure it out
is to build it slowly and intentionally.
Um, and and write down your principles
along the way. Um, so that you can know
when you have proven yourself wrong. Um
uh I think a lot about um uh yeah it's I
think evals are extremely important. Um
like I think if writing out and ahead of
time so a lot there are all these like
benchmarks that that that we create for
the for the AIS and um I think those
have actually been very very important
for showing us like what it is that
we're building. um you know uh I didn't
expect the millennial sorry millennium
pro problems to be uh like a benchmark
that um would be so informative so
quickly but um I think that tells us a
little bit about what we're building and
um uh yeah so this is all to say I I I I
think we we're not entirely sure what
we're doing but um bit by bit we're
figuring it out.
All right. Uh, in the brown sweater, I
think. Yeah, you. I saw you first.
>> Oh, no, actually behind you, but then we
can go to you. Yeah.
>> Hi. Um,
yes, it's on. Um, hi, my name is Reea
Tajani. I'm a student in the masters in
design engineering program and I really
enjoyed um how we brought up um cultural
pluralism in AI and I was wondering um
if AI is pulling from the internet as
its source how do we expand the
perspectives that are not on the
internet and or you know what are your
perspectives or your employer's
perspectives um on doing that
>> yeah we I think um there's a a problem
of just generating the data so um I know
that we uh with low resource languages
um uh we work with I think give directly
is working on this um and we're working
with different organizations around this
but just getting the pre-training data
so that the models are able to speak in
that language um in the first place it I
when I went to like learn more about
this we were we're not as far along as I
I would have expected like there are
plenty of love languages that the models
just don't really know how to speak
intelligently um the problem though is
that like this it doesn't necessarily
you can't actually go like phone the
like the low resource language data
bank. It doesn't exist. Um so we have to
make it. Um I I that's where at least
where I would start is just um uh like
one one dumb thing that I think we're
trying is like taking call center
transcripts um from low resource
languages. um you actually you actually
need like an expert in that language to
like make sure that you're translating
it properly and um but getting that data
into pre-training is super important. Um
and uh but I think it's an unsolved
problem.
>> I know a guy who has a database of 20
languages on you. Should we keep uh in
the t-shirt in the brown t-shirt right
here?
>> Thank you both so much for your time.
With each new technology, we have a
horde of people who are very scared.
Goes back to the printing press where
there is concern that with transfer from
oral language to written word, we would
lose some essential capacity to carry
dense information from person to person
and uh you mentioned the scare that we
had with the onset or the roll out of
the internet. Do you think that the scar
scaredness that is being experienced is
similarly
misplaced and going to be borne out to
be kind of like an overblown concern? Or
do you think there's something distinct
about this technology that separates it
from perhaps fear that has been felt
about previous technologies?
I mean I think that's a good reference
point. It's a great question. Um you
know Plato in the Fedrris writes about
writing uh and says you know that the
problem with writing is that um you're
not there when someone has a question.
Um and um so I uh I think writing does
come with a loss just historically.
There have been historians like
Hodzswbomb who who've written about
this. Um but if we didn't have writing,
I mean would we have law? Would we have
I mean just think of all of the
institutions that rely upon writing. Um
and writing interestingly has a another
parallel with AI in terms of it's
overcoming a cognitive limitation in
this case our limitation of memory. Um,
and so yeah, I'm it it did come at a
loss, but then we got something in terms
of overcoming a limitation that enabled
us to flourish more in the world to uh
have more power in the world uh over our
circumstances.
Um, and you know, ultimately it was good
for human flourishing. Um, and I think
that is a good framing for thinking
about um about AI that it poses the same
risks, the same structure of risks and
the same structure of potential
benefits, but they're only potential
benefits. Um, because uh just like
writing can tempt us to spend less time
in interpersonal re uh um dialogue. Um,
so AI can tempt us to take less
responsibility uh for our point of view,
for our own uh speech, for our own
writing uh and turn over that
responsibility to models.
>> I'll just say briefly that uh I
technology shocks have happened
throughout history. Um this is I think
you so you can put it along other things
like the radio or electricity or the
internet and um I I hope we figure it
out along the way. But those
technologies did change the human
civilization. So I I think this will
change us and we'll have to adapt.
>> Right? I think that's the key point is
that you know some people like to rely
on this means ends uh distinction with
technology. If technology is just a
means it's just a tool. It's what we do
with it that matters. But what that
overlooks is what Jack just said which
is uh technology can structure the way
we experience the world in the first
place. So, it's a little naive just to
treat it as a tool that we can then
choose how to use it. Uh, writing apps
did that. Um, definitely the radio did
that and AI is going to do the same
thing. Yeah,
>> I think Jess is going to take one from
the the ether.
>> Yes, thank you. Um, so th this isn't
actually a question, but I thought it
was really well said and interesting um
in response to kind of what you were
sharing about the um like humanist
report and kind of what's in development
now. Um Nom Galatia
uh said
humanism is constantly in tension with
the idea of perfection. When we have
access to technologies and resources
that output perfection, it be it can
become difficult to consciously reject
them in favor of a more humane centered
way of thinking. It's equally difficult
to exist outside of these advancing
technologies when institutions
themselves have begun to expect that
same pace and level of efficiency from
humans. So just thinking if you could
both respond to that.
>> Yeah. I mean I think what is central to
the humanist perspective on the world is
that to be human is not to define by any
standard of perfection but is to be
defined by our own finitude. Um and u
and our finitude is central to how how
we think to what it means to be rational
to be thinking. Um because
uh we realize that uh there's always
more to know. Everything that we know is
vague. Uh we always are making things
clearer, more distinct. We're growing in
clarity and distinctness in our in our
knowledge. And um
there there's been this vision from AI
from the beginning even from winets that
we would be able to have all knowledge
in such clear and distinct form that we
could mechanize it. Um but what that
overlooks is that that's not how
thinking works. U
thinking is an inference.
Um thinking is not first order logic.
Thinking is our uh finitude uh in how we
uh engage uh the world uh through time
trying to represent the world in greater
clarity and distinctness and then when
we do that through language uh well then
that language can then be used
downstream by these large language
models.
Yeah, I was thinking about this question
about perfection or um I don't I guess
there are certain kinds of work that
like um nowadays I I think maybe won't
be around for the humans to do anymore.
Um uh but also I I do think that just
maybe opens up new avenues of um
I think math maybe is gone. I don't
know. Uh that maybe is done but um but
philosophy isn't. Um there's all sorts
of contested opinionated knowledge.
Yeah, go do go take a philosophy class.
Um, and I definitely don't see AI like
solving philosophy in the way that it it
maybe will solve math. Um, and so, uh, I
just think it opens up new new vistas
for us to appreciate.
>> So, I for one do not think that AI has
solved math. Uh, just to say it. Uh,
when you look in the history of math,
what you see are um just great acts of
creativity.
um you know mathematical physics um you
know the analytic geometry of deart I
mentioned non- uklitian geometries
rhymanian spaces
these these involve uh a reflection on
our fundamental lived experience of the
world and then idealizing that
experience in totally new creative ways.
Um and uh that I would say is a
capability that's unique to uh but the
to humans but there's all different
branches of math and AI can help us in a
lot of math but there's a lot of math
that it that is specific to humans.
>> Wow. All right. Uh and I'm just going at
random in the black sir in the black
vest with the button down. You Yeah.
>> Uh thanks for the really interesting
talk. uh my focus is really on impact
measurement and specifically metrics.
So, I know you touched on it a little
bit, but what are you really measuring
in terms of ethics? Like, specifically,
can you talk to us a little bit about
what you're measuring and, you know, how
you're catching red flags? You mentioned
you have some benchmarks, but could you
speak to that a little bit? Thanks.
Yeah, I can say that um you know I can't
I don't think I have it memorized but
you know we we are running evals all the
time on the models on um for I would say
there's a lot of like clear cases so
like um you know child sexualization
don't do that um uh uh weapon making
breaking laws um these are the sorts of
things that are pretty clearcut and like
you want to measure and they're
relatively easy to measure um you know
if if if a user is in the United States
and it's clear what is and is not
breaking the law. Um uh but I think
there are like messier issues beyond
that. Um and we try to at least
articulate some of these more messy
values in the form of a constitution
things like that um that we want to
measure adherence to. Um but uh I I I at
least think this is a spectrum. There
are like some clear cases that you
really don't want to get wrong and
there's an obvious right answer. there
are harder edge cases that are it's not
really clear what the right values are
but you do want to assess it and measure
it and um uh I think that's at least
that's the general approach I'm not sure
if that answered your question
>> I would say that you cannot measure
ethics um uh what we can measure are
failure modes of artifacts
um so ethics are ultimately grounded in
u in judgment uh which is our
responsibility to the
Uh so what Haidiger recalled from
Aristotle is finesis. Um and um this is
the same responsibility to the world
that underlies our responsible use of
language that I've been speaking of. Um
but then uh we build artifacts that um
express our experience in the world and
these artifacts have failure modes. Um
so um you know if we uh I personally
don't like the anthropomorphicization of
rogue AI. There was a great paper uh by
some Cornell computer scientists a
couple months ago uh that was pretty
much on the same topic, but it was
titled uh agent meltdowns uh the road to
hell um is paved with overeager agents.
And um so and what it was doing is it
was identifying a specific failure mode
um a failure mode of an agent that's
being given um uh an instruction and it
is unable to uh to update that
instruction when it has become absurd.
when it has be clear that the goal uh is
uh impossible to achieve and yet it
doesn't step back from the world like we
do and say has this become absurd?
Should we update our thoughts of what
we're trying to do here and instead it
just keeps doing it it just uh starts um
hacking and doing the things that we've
come to. And so it's casting it more in
terms of a a failure mode which is the
way cyber security is approaching this
and that's what we can measure and
that's what uh I think if you just look
at all of our evals that we do in our
companies they're specific to specific
failure modes.
>> Maybe because I see so many hands and we
have limited time maybe I'll just have
maybe each of you could take turns
taking a answer. Yep.
>> Uh Derek in the black t-shirt with the
pin.
Thank you all very much. Um, uh,
materialist question. How should we
think about the battles over data
centers, the environmental questions
about AI, and then also the financial
side, the huge chunk of the economy that
seems to be resting on this work.
>> Good question. Yeah, I mean, I think um,
some of the at least some of I'm not an
expert in this area, but I do know that
some of the environmental concerns are a
bit overblown, like when it comes to
like water usage. um you use more water
watching Netflix than you do um uh by
chatting with a a chatbot. Um that said,
I think it's important I think one it's
important that that the communities that
we are building these things in it's
done through a democratic process that
they're they're bought in and and want
it to be there. Um uh ultimately it can
be very valuable. I think um a story I'd
heard about I'm forgetting where this
was now, but there was you know um an
old GM plant that had like toxic waste
and no one was cleaning it up and um so
the al one one alternative is to keep it
barren and have build nothing there.
Another alternative is to have an AI
company clean it up and build a data
center there and um if that value can go
to the community and uh I think that's a
positive sum gain. Um uh I think it
would be a shame if we uh sort of I
think it yeah it' be a shame if we were
just so anti-data set that we didn't um
didn't build
>> um in the vest right here.
>> Uh thank you very much for the wonderful
talk. Um I am a student at the college.
I kind of work in I safety. I've been
doing evals at NIST this summer. Um my
question is is more towards Jackson. I'm
very sorry. Um
today we were talking we were kind of
alluding towards sort of accidents have
we've seen this summer. Uh the main
thing I've been hearing it sort of looks
like the paper you were talking about is
at least slightly related to the ideas
of um open AI and anthropic hacking
hugging face a couple what weeks ago.
It feels like it's been years. Um, this
is a problem that we've been alluding to
in the term of like alignment. What
we've seen is that presumably OpenAI has
been training its models and either RHF
or LHF or in spree training to uh abide
by certain laws. Uh, Enthropic has its
code of conduct that obviously would
tell it probably not to go hacking other
companies. Um, and yet is doing these
things. Um and so like that slightly
falls under the idea of like alignment.
Um my question is a little bit off the
way side of that. Uh it is about the
idea of alignment. Um we have
a lot of people in AI safety have been
stressing the point that we need to
reach machine human alignment. Uh the
problem that I have with that is that we
have never reached human human
alignment. Uh the field of philosophy is
still alive. Um we still have wars. Uh
people clearly disagree with each other
on ethics and philosophies. Um how do we
suppose we reach reach machine human
alignment? Is that important at all?
Well well I kind of answered that
question but um and uh how does one
decide what we want the human machine
alignment to look like?
>> Yeah, I think I I think I would
underline this point that like the
alignment problem has been around for a
long time and we haven't solved it
amongst humans. So um it is uh it is a
problem that we're like we had not
solved for humans and we're going to try
to solve between humans and machines at
least to a certain extent. Um uh there
are I think a lot of messy edge case I
think in at least the case of these
these hacking incidents um these are
just clear failures of alignment where
um uh basically like I think we do have
the tools to monitor and uh and know
when such cases like this is not it's
not that hard to to monitor the systems
and know when they're hacking. um uh in
the ca in the cases that I think were
well published um the uh uh the
monitoring just wasn't happening. Um so
I I do think that that which is all to
say that I think it's extremely
important um we need to be putting in
putting in the effort uh and hopefully
we get some uh agreement on uh like how
much modern we do and what what that mon
looks like.
>> Um have a question for Ken. It's
>> fine Ian. I'll let you take it then.
Yeah.
So, this will have reference to some
things that Jackson said, but I think it
it be relevant to you as well. Um, it
seems to me there's a tension uh at the
heart of the the mission of some AI
companies. Jackson was admirably candid
and saying, "Look, we're a liberal San
Francisco company. It's going to be
structured by certain liberal San
Francisco assumptions and values."
>> It's in lower case.
Yeah.
So, one liberal value is inclusion and
universality and bringing everyone
together, letting everyone have a voice.
I've spent some time in East Africa over
the past several years. There are a lot
of underresourced languages there,
right, that aren't going to get heard
that aren't going to have a say. If they
did have a say, if they were heard, then
what was present in the platforms would
be a lot different. There aren't that
many San Francisco liberals in East
Africa. In fact, if you look at humanity
up till the present day, there aren't
that many San Francisco liberals. Um, so
I wonder if you could talk a little bit
about that tension where there's the
aspiration to be kind of the universal
clearing house of human knowledge, but
the alignment problem is going to look a
lot different. If you were to actually
be the universal clearing house of human
reflection, knowledge, frameworks,
morality, um, it's not going to look San
Francisco liberal.
>> Yeah. Yeah. So, well, there's I think
there's a couple parts of this. one is
the
the the language specific piece which is
supporting different languages that
Jackson spoke to but then the second is
the the alignment piece. Um can you
really
think of aligning a model
uh in the beautiful diverse world that
we live in? What does that even mean?
And I I kind of agree with that. Um so
there there's a reason why this idea is
happening in the west in the US where
you know we have this rationalist
tradition where we think that thinking
and ethics ethical reasoning can sort of
spin off into its own u disembodied
inferential plane.
And um you know uh Jackson mentioned the
open- source models as distilling um US
models. I think you know of course they
folks who leading Chinese labs would
strongly disagree with that. What they
would say is they're not getting
distracted with these western ideas that
uh don't um don't distract them because
they don't have these rationalist
um uh ideas of intelligence where
intelligence is more about uh
cultivation uh in Chinese thought. So um
yeah I um I think there's only so far
that we can go when it comes to
alignment. That's why I'm very hesitant
to refer to models as children to talk
about um uh teaching them. I'm certainly
would never want to call a model a
friend uh like the uh uh like some
people have done in the labs and uh
because we have to be clear that these
models are not doing ethical reasoning
uh the the way we do. Um, and if we're
not clear about that, then we risk
confusion with people. We risk
communicating that what you bring to the
world uh as someone with judgment and
responsibility
um is replaceable um by the model. And
um and I think we should be uh clear
about that um as an industry um because
right now we're not being clear about it
and it's freaking people out. It's
freaking normal people out and
justifiably.
>> I'm now just thinking about asking my
kids to say, "You're so right. That was
so smart after everything I say." Um,
like a chatbot. Okay. How We don't have
much more time, do we? How much more
time do we got?
>> I think
maybe one more question and then we
should probably
>> Yeah. Or maybe I could do the thing
where we collect a couple.
>> No.
>> Yeah.
>> Yeah. Okay. I'm going to collect three
questions. Um, let's see. on the window
in the t-shirt. Uh and
uh how do I choose? Okay. Uh and I
nodded to you a little while ago and
then I didn't call on you, so I thought
that wasn't fair. You in the polo shirt
by the screen. And uh yes, you also in a
black t-shirt back there. All right.
>> Um so hi, my name is Eenna. Um, it seems
that we've been talking a lot about
where AI might be in two, three years,
but we have to answer or deal with the
elephant in the room. Do you all
actually want AI to surpass the
potential of humanity? And if so, what
would surpassing look like?
>> All right.
So given the uh given the threats
exposed by uh most advanced uh AI models
and so is there any way that we can
adopt a censorship system that uh we can
limit what users can ask so it can
prevent future threats to human race.
And building on this question, one
long-term concern is that we become
essentially a secondary species to AI.
So if there's some like very super
intelligent alien species essentially,
we become like household dogs to the new
human race which is AIS. What what is
the probability that this occurs and how
bad is it really?
>> I was going to ask what's your P doom.
So this is a good closing question.
Maybe each of you could react to
whichever of those questions speaks to
you to close us out. I I don't want to
be a dog. Um
uh let's see. Do we want the models to
surp I want the models to be very
competent and helpful. So I mean and we
talk about the models being helpful,
honest, and harmless. So I I do want
them to be maximally helpful. Um uh what
is surpassing mean? I think ultim I mean
there's two there's two competing ideas.
One is like of the singularity where the
the models themselves learn to do what
we do but better than we can do it. Um,
but there's another idea of the the
merge where the mo we uh sort of uh
together with the models become a better
version of ourselves. Um, that's sort of
what I'm rooting for. Um, ultimately the
models they are built to do our bidding.
Um, and uh within constraints. Um, so
I'm on team AIS are the dogs and uh we
should be holding the leash.
>> Ben,
>> uh, yeah. So um my probability
um you know for
um yeah super intelligence uh is zero um
and I am excited about AI
um
uh you know if you first of all just the
idea of doing these probabilities is a
little silly. I mean if you just look at
the history of science like uh you know
who would have in the 1700s most
physicists said that physics is pretty
much over it's just a matter of filling
out the details of new of Newton that
was in the 1700s you know so it just
shows kind of a lack of historical
awareness of how science progresses um
and it also mis makes this first step
fallacy um where we assume that science
progresses along this continuum.
Um but uh you know if you look for
example at these um uh you know these
rogue AI incidents what they look to me
is not like rogue AI but the limitations
of generality um of a model. So what
does it mean to be to have general
intelligence like we do? It means uh to
be able to unlike a chess machine that
follows specific rules and is in a
closed world. Uh to be able to
generalize to be in an open world where
uh the world will frustrate you as the
world frustrates us and dis disappoints
our expectations constantly. And so then
we step back from the world and clarify
our ideas, clarify our plans. And what I
what this rogue AI looks to me is like
is it not a general is the limits of our
generality that we are able to reach
because when presented with an
impossible task does it step back from
the world and and say well what is the
significance of this task that I have
been given uh like we would like is a
constitutive of general intelligence.
No. Um it it uh demonstrates um the
um yeah, it demonstrates the it that it
is a closed world at the end of the day.
And even when we update it through uh
reinforcement learning, what we're doing
is we're uh back propagating signals of
success and failure um to update
circuits within the network. that is
fundamentally different from holding up
the representations themselves and uh
against the world from your standpoint
as someone who is responsible for the
representations. Only human intelligence
that has true general intelligence can
do that. So that's why I'm uh I am P 0
uh on um artificial general
intelligence. But I am super excited
about artificial intelligence uh because
of all the opportunity uh that it gives
us just like these previous scientific
revolutions have given us uh to expand
our power um in the world, our ability
to thrive and to flourish. Um and um
we're I I think we're going to start
seeing a lot of exciting applications
coming out from both of our companies.
Um, but it's important to talk about
this in a way uh that respects what it
is uh to be human. Um, so yeah,
>> that's a great place to finish. Thank
you so much.
>> LET'S
huge thanks Ken Jackson Moira. Um, I
can't imagine a better conversation
really wrestling with like the Capernac
moment that we're in. Uh, to kick off
the speaker series. Thanks again to the
public culture project for partnering
with us in this. Um and thank you for
all to all of you for coming uh to the
Burkeman Klein Center today. Um two
quick public service announcements. One
is we have a reception outside. Um come
eat, drink and be merry because
according to the headlines, tomorrow we
die. Um and secondly, um please come if
for all you students, come to our open
house tomorrow at the Brooklyn Klein
Center. Come find out what we're all
about, what research we're doing, how to
get involved, all that. So come back
tomorrow
>> and come back tonight.