Is AI the future of health and social science?
Watch on YouTubeVideo summary
Professor David Bun from UCL argues that artificial intelligence represents the future of health and social sciences because it excels at the cognitive tasks central to quantitative research, such as literature reviews, data coding, and administrative work. He points to rapid advancements in large language models, which now surpass human experts on complex benchmarks like GPQA, while costs have plummeted by 99%. By automating these manual drudgeries, AI frees researchers to focus on high-level scientific thinking and hypothesis generation, with examples showing AI agents completing systematic reviews in days rather than years. However, Peter Tenant from the University of Leeds strongly contests this view, warning that current incentives already drive massive research waste and that AI will exacerbate this by enabling the mass production of low-quality papers he terms "slop." Tenant emphasizes that LLMs are merely pattern-matching tools lacking true reasoning or understanding, making them unsuitable for genuine scientific discovery and prone to dangerous hallucinations that can undermine public trust in empirical science.
The debate deepens as Peter highlights severe risks associated with relying on these models, including automation bias, cognitive offloading, and a collective deskilling of the scientific workforce where humans perform worse due to diminished critical thinking skills. He further warns of "model collapse," a phenomenon where training data degrades because over half of internet content is now AI-generated, alongside concerns about the commercialization of tools leading to ads and high costs accessible only to the wealthy. Peter also raises ethical issues regarding intellectual theft, the exploitation of invisible laborers in the Global South, and the immense energy consumption of data centers, concluding that AI threatens the very existence of high-quality research. In response, David counters that LLMs are continuously improving through reinforcement learning and reasoning training, making them suitable tools when used with human discretion, and disputes the inevitability of commercial degradation by noting users can switch to open-source models or smaller, quantized versions that reduce energy usage significantly.
Despite their sharp disagreements on the trajectory of AI, both speakers find common ground on several critical issues necessary for the field's future. They agree that proper governance and incentives are required to prevent the proliferation of low-quality research, and both stress the importance of reducing information overload to improve the signal-to-noise ratio in scientific communication. Furthermore, they concur on the necessity of training scientists in prompt engineering and best practices to navigate these new tools effectively. While acknowledging that AI cannot replace the intrinsic joy and craft of scientific inquiry, they recognize its potential as a useful tool if integrated responsibly, with David adding that it could help address global educational inequities by providing personal tutoring to under-resourced regions. Ultimately, the debate reflects a complex reality where the audience remained nearly split, underscoring that while AI has pitfalls like hallucinations, these are manageable through engineering and triangulation rather than abandonment, paving the way for a balanced approach that prioritizes efficiency alongside quality.
Read the full video transcript
Thank you very much and um a really warm
welcome for tonight's um debate on um AI
and the future of uh well the use of AI
and the future of health and social
sciences and what that means uh for us.
Oh,
okay. So, it's really great to see so
many people here. We were fully booked.
Um so that's fantastic to see. Um I just
want to briefly mention that this is an
event that has been organized by the
National Center for Research Methods and
by our research grant or training grant
funded by the UKI um on digital skills
development and you can look at that up
on our website the NCM website under
innovation. Um just very briefly a
couple of things about the National
Center for Research Methods. We are a
center that is um yeah running
developing um providing a lot of
training and capacity building
activities
um uh for example in core and cutting
edge research methods in particular on a
wide range of research methods areas
quantitative qualitative mix digital and
so on um across the career life course
so for junior and for senior members of
staff or colleagues across sectors not
just academia of course um and we have a
whole um well a high uptake of of users
on our website and our resources. And um
we have also published earlier this year
a impact um assessment report and you
can look that up uh on our website
and um yeah you can find here the QR
code to our website. Um some of our ESRC
funding is changing but we are very um
um excited to announce that we can
continue with some of our core and
advanced training activities. uh and
that is really based on the sort of
successful collaboration between
University of Southampton and University
of Manchester and the number of our
center partners that actually includes
um UCL um at as as as well. Um and yes
um we have got some funding for uh
particular sort of AI skills development
skills activities and there you can find
much more on on our website in this
area. Okay.
Um so very briefly about this um evening
um you may have seen the um outline or
the program. We will have first of all
the debate um between um professor David
Bun and um Peter Tenant from the
University of Leeds and David from
University of College of um University
College London. Um and then we have sort
of a rebuttal later on and then also 20
minutes roughly for question and answer
so people can uh get involved as well.
And then we have about an hour for drink
reception and nibbles. And I think I
have been told it's a malt wine for the
uh sort of season.
Um yeah, so and before we start, we just
wanted to gauge a little bit of a an
understanding of the audience and just
sort of a bit of fun to get you
involved. Um yeah, please take part. Um
or you can enter the code. And basically
it's a question about is AI the future
of health and social sciences and maybe
we can sort of see a little bit if that
has changed later on. If you convince
you or not convince you, obviously we
don't see uh changes amongst uh
individuals but we can sort of see see
the trend. Um I think on this note maybe
I need to leave this up for a little bit
longer.
>> 22 votes.
>> Oh yes.
>> 30 votes.
>> Come on. Some of you need to vote.
>> I have voted already earlier.
Okay. Well, in the meantime, I can
introduce um Peter, Peter Tenant from
the University of Leeds and David Bun
from the uh from UCL. Um and just very
briefly um David, you're professor of
population health. you've done a lot of
work on various different aspects of
population and public health about the
causes of that distribution of health in
the public and also your sort of maybe
more research uh interest in more recent
times includes also thinking about AI in
the future of scientific research and
Peter um your associate professor of
health and data science um we've only
had actually dealings um because of this
this debate so I'm really pleased to
have this connection now we've got some
uh connection to the University of Leeds
and Leeds is one of the center partners
of NCIM um and um you've initially
trained as an epidemiologist and now
you're focusing on you know causal
inference and applied health and social
science research. So I'm really pleased
to hand over um they're gathering
already so I better make a move. Hang
on.
>> Do you want to just
>> I'll just have to read it out. Okay. So
yes 48%.
No 36% and everyone else don't know.
Okay, fantastic. Connect that
[clears throat]
>> and I'll make a move and I'll hand over
to
>> Thank you. Can you all hear me? Okay, I
have a slight cold, but can you hear me
at the back? Great. Okay, before we
start, thank you everybody for arranging
this. And I think I think What does that
mean?
>> Louder. Okay, I [clears throat]
>> maybe you have to hold it.
>> You might have to hold it. Yeah,
>> have to hold it. Okay.
>> God knows what that mean.
>> Okay. Um, is that slightly better? Okay,
great.
Okay, I think we've collectively
organized what might be the most and
least relaxing way to spend a Monday
evening [laughter] before Christmas. Um,
but I hope it is interesting and useful.
This is an important point of the of the
history of our disciplines.
As the saying goes, the aim of of debate
is to make progress.
So, this is my argument in one slide.
Firstly, quantitative research is mostly
cognitive work. In the past, that was
only um doable by humans, but AI systems
are now increasingly capable at
cognitive tasks.
And the second point is that if these
systems can help re research, then they
are part of the future. And I'm
particularly optimistic about humans and
AI systems working together to achieve
our goals.
So the first part then so what are the
tasks that we do in our in our research?
Um so this is a list that I've put
together. So first of all we design
studies and we collect data. So the
survey research world and then we um do
research papers, we review the
literature, understand what's been done
before. We form research questions,
hypothesis, we analyze data and then we
interpret it and write up.
So which of these tasks are cognitive
tasks?
almost all of them are cognitive tasks.
And so you would expect the research
that looks at um exposure to AI across
different occupations would find that we
are quite highly exposed
and indeed that's that's the case. So
this GPTs our GPTs paper buried in the
in the supplement on page 41 is an
estimate of um how much exposure survey
researchers have to AI large language
models and they estimated a 75% exposure
just to LLM like chat GPT
and they estimated 84% exposure um to
LLM plus tools that LM can use as well.
I can sense the energy dipping in the
room slightly after saying that. But
don't worry, we can bring it back with
this slide. Just because we're exposed
does not mean that we're imminently
extinct. So, as the example from
radiology, Jeffrey Hinton, the Nobel
Prize winner, um said that, you know,
people should stop training radiologists
now in 20 2016. And since then, there's
been a continuous rise in demand and
supply of radiologists
because radiology isn't just binary
classification of images. If time is
saved in that activity, they can use
their time elsewhere for to better serve
their goals just like we can as
scientists.
I'm now going to talk about AI
capabilities. I'm going to show you a
few different plots. This one um shows
across them from 2012 to 2024, so
slightly outdated, the rising capability
of frontier large language models across
multiple different domains of um
cognitive task.
And we can see steep rises in multiple
different domains, many of which are
very relevant to the work that we do.
For example, mathematical understanding
or general scientific understanding.
So one in particular I'm going to point
out here is the GPQA. It's described as
being a graduate level Google proof Q&A
benchmark,
a challenging set of questions written
by domain experts in biology, physics,
and chemistry.
This slide shows the same benchmark
broken down across the different models
across time
using slightly more recent data. So I
think in the early LLM period the models
were very new and the models were I was
very intrigued by the models but I
wasn't very impressed by their abilities
in a sense they seem very fickle
particularly with tasks that required
multi-step reasoning.
So this benchmark and other benchmarks
are designed to capture that kind of
reasoning ability and we see great
increases in in the performance in those
benchmarks actually now exceeding that
that described as expert human level. So
those with PhDs or those training for
PhDs.
So we're actually getting to a point
where in the AI world they're worrying
about benchmark saturation. How can we
develop benchmarks where AI systems
don't get all the questions right
essentially
and so we have new initiatives like
humanity's last exam humanity's last
exam has been designed even there in
recent months we're seeing vast
improvements
chat GPT 5.2 two was released 3 days ago
and it was only 30 days since the
previous version GPT 5.1
and we're having improvements in
benchmarks here. So R AI2 described as
capturing advanced reasoning we see a
more than doubling in the performance of
this model in this 30-day period.
So I think I would say that if if people
have very very strong claims on what AI
cannot do, we have to make sure they're
up to date with the latest evidence on
their performance.
So AI is very powerful. Um and some
further indications of this is that
frontier models are now getting um gold
medals in in highly competitive um
competitions. for example, mathematics
and uh coding, informatics,
but they're also highly jagged. So,
they're very they're very um the
performance is very high in some domains
of activity, but very low in other
domains. So, for example, very high in
maths, reasoning, and knowledge, but
very low when it comes to using their
memory, unlike humans.
But still a highly jagged AI system can
still be very helpful to scientific
discovery.
Another indication of their improvement
is that the length of task in which they
can complete is also increasing across
time.
So it used to be they were restricted to
tasks that take seconds like responding
to a multiplechoice exam, but we're now
seeing the frontier models undertake
tasks that might take a normal human 2
hours or 3 hours.
So improvements in their ability,
improvements in the length of time task
that they can complete
and finally rapid declines in their
cost. So this is around 99% reduction in
the cost of frontier models.
So it's been described that the cost of
intelligence is trending towards zero.
[clears throat]
The cost of cognitive work is trending
towards zero. Our work is cognitive
work. So of course there will be
implications for our field.
So the second part then can it actually
help us in research?
What is the goal of research first of
all? Well the goal of research is to
expand knowledge. This is the Vcarti
definition. It's work undertaken in
order to increase the stock of
knowledge.
And in doing research, the tools that we
use change, but the goal does not
change. You can do research without any
tools at all, just using paper or even
without paper, just speaking to another
human being. Or you can use um tools
like a computer. You can use tools like
a computer with some AI or a computer
with lots of AI.
The key criteria is does your um work
contribute to the discovery of new
knowledge?
You can do exactly the same paper and it
may have taken you 10,000 hours in a
very manual endeavor or it might taken
you um you know 10 months or 10 weeks
using frontier tools. If the
contribution to to the key thing is the
contribution to to knowledge is is um is
is the key criteria.
I'll put this in a sort of a different
way. So the two kind of tasks that we do
in our in our work literature reviews
and data analysis we historically it was
very much a manual endeavor. So
physically searching libraries reading
papers manually doing manual
calculations for our analysis.
We didn't stop there. Humanity um
continued to progress. Society um
changed. We now have access to the early
computer era where we have uh scientific
research available by computers via the
internet using digital databases and
using computers to speed up our
analysis.
We didn't stop here either in the early
computer era. And why would we?
We're now in the era of what I describe
as being early AI where we're using
machine learning tools using large
language models to help us make sense of
increasing volumes of literature and
also to help us to accelerate data
analysis and coding.
We're not going to stop here either. So
the progress is continuing in AI and so
we'll be expecting further gains as
well.
And I would guess that we'll be getting
to a point in the future where we can do
a reliable systematic review with a
single click. And we can go from plain
language instruction to reliable
analysis with code that is auditable and
code that is then executed. We're not
there quite yet, but I think we'll be
going in that direction in the future.
The good news is we can actually
research these tools as well. And that's
an important contribution to knowledge.
How useful are these tools? what can
they do? What can they not do?
In academia, we're not building frontier
models, right? We haven't got the um the
capital to do that, but we can at least
understand these models in our work.
Here's a breakdown of the different
tasks and I've put on rough um time
estimates of how long this stuff takes.
So for the average paper let's say it
takes between 3 to six month 3 to 12
months to do a literature review maybe
half a month to locate the data 3 to six
months to analyze data 3 to six months
to interpret and write up
for each of these tasks. I guess we
should be reflecting how much of the
work that we're actually doing is
optimal or how much how much of it is
lower level work that is repetitive
isn't essential for our work and could
be delegated to to AI.
The paper I showed before the GPTs are
GPTs paper it said that because you know
it described GPT as being a general
purpose technology. So the good news is
that we can use GPT and other frontier
LLMs for any task really and we can use
them at our discretion. We can use them
as very restricted tools. We can kind of
use them as research assistants or at
the very far end the more experimental
end we can use them as agents in which
we give tasks and they go off and
complete those tasks.
This may be an optimistic estimate of
how long things take. It's been
estimated in the US that um almost half
of US researchers time is spent on
admin.
So even there I think this is a one area
where AI can help us spend more time on
actual research that we actually want to
do in the first place.
In the next few slides I'm going to
focus on literature reviews, locating
data and data analysis. In the paper on
the bottom right um we discuss other
tasks as well that we do in quantitative
research.
So we're going to start with reviewing
the literature.
So if you do a systematic review, how
long does it take? Well, this review of
length suggested 67 weeks. So over a
year for a single systematic review.
The Cochran handbook um advises between
1 to 12 months.
But I expect many of you many of you
that have done reviews will know a lot
of your time is spent on painstaking and
quite often painful after a series of
hours manual screening of papers. You
have to read through manually hundreds
potentially thousands of individual
abstracts to decide if they should be in
your pool to analyze or not. Is that a
great use of human time?
So if you give humans a very repetitive
and boring task, they tend to make
mistakes. So this paper suggested an
error rate, a baseline error rate of
around 10% that humans make in reviews.
That's another motivation to try and use
AI. It might not only be more efficient,
it might be less errorprone as well.
So even Cochran are now exploring the
use of large language models for
systematic review work because the aim
of Cochran is not that we spend
thousands of hours doing manual
repetitive task. The goal of Cochran is
where a world where health decisions are
based on timely evidence.
We can't really do timely evidence if
our reviews take a year, a year and a
half and so on.
So there's emerging evidence now that AI
can do um can can be useful when
classifying abstracts. This is one paper
which suggested this.
There's another paper here which um at
the more extreme end using agents. It's
estimated that that this agentic system
it completed 12 cochran issues 12 coin
cockrine systematic reviews in two days.
Traditionally that would have taken 12
years.
So if we can get to that point and if we
can get to that point reliably, I I
argue that we should
for ad hoc reviews. Let's say you're
informing what's going to be in your
introduction or your discussion section.
We haven't got the bandwidth to do a
systematic review. We haven't got the
time.
So we do the best we can. I think many
scientists now are actually using AI
already. we're already using Google
Scholar rather than the traditional
tools that we might have trained in and
for good reason. So for example, unlike
PubMed, Google Scholar actually indexes
multiple disciplines. So it it actually
indexes the social science literature
which PubMed doesn't always capture and
it also indexes the gray literature as
well.
So there's an emergence of AI tools for
for for finding papers. So, we've seen a
few weeks ago, um, Google Scholar has a
a Scholar Labs where you enter your
question in plain English. What is the
evidence that X affects Y? And then it
finds you all kinds of studies which
address that research question.
We have the Claude LLM linked in with
the PubMed database. And we have the
elicit app as well.
So these are kind of early stage
technology but in the future you know we
have a we have a frontier AI which has
access to millions of papers greatly
exceeding what we could ever read in a
single human lifetime.
The hallucination rates are also
declining in these frontier models as
well. So I expect this will be extremely
useful for research.
In summary for reviews I think humans
and AIs working together can free our
bandwidth for high level tasks.
And if we want to do systematic reviews,
we can do more ambitious reviews across
different disciplines, across different
study designs. And it opens the door to
doing continual reviews.
If we're just doing research papers, we
can hopefully get to a point where we're
doing more informed papers and more
quickly.
Um, briefly now, how about forming
research questions and hypotheses?
Initially, it was described that LLMs
are simply stochastic parrots. They're
simply interpolating between data
points. They have no creativity at all.
Well, we can view creativity as yet
another cognitive task. And in fact,
their creativity is to a certain degree
empirically testable. So there's some
evidence this is a recent paper which
suggested that it has a GPT4, for
example, has a reasonably high degree of
creativity and creative writing task,
but it doesn't match the distribution of
humans. So the best humans are better
than this AI model.
But the AI model is better than the
worst humans at this task.
We're also seeing the emergence of
benchmarks for assessing creativity as
well.
So in my experience, you know, even if
they're flawed and limited for being
creative, they're still useful because
they're instantly available, very cheap
to use. So AI and humans working
together can be helpful in that creative
art aspect of our research.
So we're now seeing papers and published
in the top journals in their disciplines
in which AI has has had a creative role.
So we have a paper in a top physics
journal published a few weeks ago and a
paper in cell one of the top biology
journals and they described here that um
the top AI generated hypothesis actually
opened new research directions and it
bypassed human biases to to propose
overlooked biological possibilities.
So if our benchmark is are they creative
enough to contribute to new knowledge,
these papers suggest they are.
We also see the emergence of agentic
systems for the full pipeline of
scientific discovery. The AI scientist
and Robin.
So just to highlight again how rapidly
things are developing, the papers at the
bottom here were published in November
of this year and the paper in physics
was published this month of this year.
So it's extremely fast moving.
I think I'm going to skip that in the
interest of time, but I think they're
useful for unlocking new data.
So for coding,
conventional practice is that a human
does it by themselves and you do
everything manually. Okay?
When you're doing manual coding, you
might think to yourself, well, I've
spent years typing exactly the same
command. That kind of breaks the
software engineering principle of
automating repetitive tasks.
Often your code doesn't work right.
There's an error, some unknown un
unknown error. So then you spend maybe
hours checking very verbose
documentation.
Maybe you spend hours checking online
forums and wondering why the responses
are so sarcastic sometimes.
A few hours later, you fix your bug and
your code works. And there's a certain
satisfaction which I do sometimes miss
of fixing coding errors.
But then I wonder to myself across a
scientific lifetime,
how much time is wasted on this process?
Is there a more efficient way to um er
error bug fix our code?
So we're now seeing um AI integrated
into into a data analysis software. And
one way is is the autocomplete options.
for example, with GitHub Copilot.
And here the the AI will see the code
that you're writing and will simply
suggest code that you might want to
write next and you can accept or reject
that that coding inclusion.
And to me, it's not surprising that um
the literature which empirically
evaluates the use of those tools finds
that it does actually um benefit
productivity.
This is a recent report where these
tools were deployed in the UK
government. And they estimated for each
user, if you give them these tools, it
saves them 28 working days per user
annually.
And 58% expressed they would they would
not want to return to their original
working conditions. So it seemed to save
them time and many seem to like using
it.
AI can also help us move up the
abstraction layer. So nobody types out
zeros and ones when they're coding. We
use higher level languages.
AI enables to go from a natural language
like give me the code to do this and
then you can get the output. You can get
code suggestions in any language. So you
can start to think rather than just
using one language I can use whatever
language is necessary for my task
and how this is integrated into software
tools. So I think many of us in our in
our in our disciplines are still using
the traditional tools like maybe R
studio or the STA the STA software
whereas you know in the software world
often using tools like VS VSM stu VS
code where you have the LMS integrated
into your work you can ask the LMS to
review your code to check your code and
to make suggestions
so that's using it simply as a tool just
to make suggestions and to check your
code. You can al also increasingly um
agentic systems of analysis which are
very new and are quite experimental
where you're having a conversation
almost with your data. You're saying
this is my data set analyze it in this
way. The tool then suggests code which
you can audit. It then executes the code
and shows the output.
That's one example. And there's another
example here as well.
So I think one of the reason ways it can
be useful to ask LLMs to rapidly produce
code is it can help with um
reproducibility
the human baseline. So how many of us
share code? There's a paper that
suggested that around 2% of researchers
in the health research world share their
code. The barrier to writing code is
massively and declining. So I think it's
likely with AI's help we can get to a
point where we're being more
reproducible in our work.
Um this is another um extension from the
um Jupiter um con conference where
they're finding yet another way of
working with agents and humans together
where you're asking the agentic system
to do analysis and you're interacting
with your AI as well.
So for those tasks I think you know in
terms of admin in terms of literature
reviews and analyzing data the hope is
that AI and LLM um and other systems can
help us move from the very um manual
drudgery tasks to the high level tasks
that we actually enjoy. So spending more
time thinking about science how we can
do good quality papers how we can design
our infrastructure
and how we can use AI to speed up our
reviews and coding.
I'm now going to talk about some of um
my personal reflections. So, how has it
been useful to myself and our group
where we work? So, we work at the center
for longitudinal studies and we make our
data on multiple different cohorts
available for free to researchers around
the world to use in their in their
research.
And we used to work you know in the
traditional way whereas now we're moving
to a world a point where we're making
all of our scripts available for our
sort of um tasks. So making available
data making available derived variables
and making available tools. So it's
actually been enormously helpful to us
in our in our overall goal of helping
other researchers.
So one example of that is al trying to
understand how to use shell scripts when
working with genetics. LMS were really
helpful there in helping us do this work
and that work is now available for
anybody to um audit and those data are
also freely available um to anybody in
the world upon data application request.
It's also been very helpful in terms of
um um making available tools to handle
our data in multiple different
languages.
We had a hackathon last month and my
colleague Liam who's in the audience he
worked with the office um for national
statistics and there they wanted a tool
to help them understand what exact job
classification somebody had. It's quite
a hard manual task actually.
So their lean was able to make an app to
do that with two single prompts to an
agentic system.
I'm going to skip that. It's very
helpful for making tools and
visualizers. So here I wanted to make a
tool to show students what effect
missing data might have on terms of bias
of your estimates. So there's a nice
visualizer which quite easy to make and
so you can get an intuitive
understanding of that concept which can
be quite hard to understand.
Um I have an example of an AI paper here
which I produced with colleagues
recently.
So we the motivation for this is I was
interested in understanding what is the
boundary between advocacy and research.
How often in our papers are we actually
making bold kind of policy claims even
in our research papers?
And so we analyzed the last 25 years or
so of um the 35 years sorry of of of
published research. We downloaded the
abstracts. We used LLMs to classify the
abstracts because the concordance
between LLMs and humans was similar to
humans and other humans. And LM's
actually enabled us to do this work to
to learn how to use Python APIs and Git.
And so humans and LMS working together
helped us do new research in a new
field, helped make data available anyone
can now use in the future. Also helped
an early career researcher. So Megan
Wang is the second author and is now
applying for a PhD with her first paper.
I'm going to skip that.
I think I'm How time
one minute. I mean it is after six but
yeah.
>> Okay.
>> Okay. Um I'm going to briefly go over
some um some of the um the concerns that
people often have before I hand over to
Peter. One is which that we can't use
our data because our data is private. So
we can't use LLMs. We're simply locked
out. So one compelling alternative is we
can simply use openweight large language
models that we can download to our
computer and use or use in secure um
servers. They're increasingly high
performing as well.
Another
one is about AI slop. There's been a
massive increase in papers across
science in the last few years before AI.
AI might make it worse.
So this really is a human problem. How
do we incentivize quality over quantity
not only in the UK but globally?
And finally, for the environmental
impact,
there's no use LLMs if they have a
catastrophic impact on the environment
that we can't tolerate.
And to that, I would say firstly, do we
have accurate estimates?
Secondly, what is the comparison in
which we're comparing them to? So this
paper suggested the median Gemini text
prompt used less energy than watching 9
seconds of television and consume the
equivalent of five drops of water.
And thirdly, is it worth it if to
balance the impact of human activity
with its potential benefits? And the
benefits of um advanced artificial
intelligence are considerable for
scientific discovery, for health, for
medicine, and for education.
So in conclusion, quantitative research
is mostly cognitive work and AI is
becoming highly capable and useful. And
so I conclude that AI is the future. And
I think it's the case even if even if
progress plateaus, even if we're
strictly limited to the models that are
available right now, I think it's in
fact overdetermined given that great
given that increasing capability and the
profound impact on science, education
and society.
We are at a profound sh this is a
profound change and it's at its infancy.
So deepseek the reasoning model for
deepseek came out less than a year ago.
Chat GPT came out three years ago.
But we should adapt to this change and
also contribute as well. We have a role
to play to understand the impact of AI,
understand its capabilities and its
limits, to help users use it responsibly
and to use it to help us learn.
So in summary, I believe that we can use
AI to accelerate research, to make it
more efficient, more reproducible, more
ambitious, and actually more enjoyable.
And I think that is the future that we
should be building.
So thank you for listening and thank you
to everybody listed on this slide.
[applause]
Thank you very much uh David
uh for a very thoughtprovoking
start.
Um, so I'm away I'm aware that this is
an away game for me. Um, so for anyone
who doesn't know who I am, um, I have a
nice description from a colleague a few
years ago. They said that I am like a
scientist out of the 19th century.
That was not meant as a compliment.
>> Sorry, it's not coming up.
>> Yeah, I can see that. [laughter]
>> Um,
Yeah, there we go. It was not meant as a
compliment, but for today's purposes,
I'm going to take that as my uh as my
base, right? I'm going to embrace my
inner 19th century Yorkman.
In other words, my inner lite and try
and convince you like they try to
convince their uh contemporaries of the
dangers of handing over our profession
to the machines.
So I'm going to start with a bit of
interactivity.
Okay. So I'm going to pose you a
scenario and I would like you to put
your hands up if you agree.
So to begin with when you each of you in
this room comes across a scientist
who is extremely productive and who
publishes a large numbers of paper every
year.
Is this impressive? Put your hands up if
you find this impressive.
Yeah, brilliant. That's about half of
the room.
Put your hands up if you think it's a
red flag.
That's even more. Great. We've got a a
critically thinking room. Why am I
asking this?
Because I think it's a red flag. I think
that good science requires time. And in
particular, it requires time thinking.
And so scientists who publish large
numbers of papers are not spending the
time to think and to craft highquality
meaningful research that has an impact
on health and society. Now we know why
they do this. They do this because we
reward reward hyperproductivity. We
reward that when we should be rewarding
less more thoughtful and more careful
research. And what is the what is the
the impact of that? Well, the pressure
of that hyperproductivity is that we
produce as a scientific system an
abundance of research waste. An enormous
amount. Over 85% of what we produce is
estimated to just be waste. What do I
mean? I mean derivative research. I mean
lowquality research. I mean low value
research. I mean flaw flawed research
and even fraudulent research. And in
order to manage this insatiable quantity
of material, we have a whole system of
mega journals and even predatory
journals that will happily take our junk
as long as we pay them
every single year. My colleague George
Tomova has described it as we are in the
McDonald's era of science.
And I'm saying this because I believe
that this is the biggest problem in
health and social science research. I
believe that this wastes all of our
time. I believe it wastes our labor and
I believe it wastes the money that funds
our time and labor.
I believe that good research then gets
completely lost in all of that waste.
And I believe ultimately most damagingly
it threatens and undermines public trust
in science and public support for
science. So that is what we HAVE TO FIX.
BUT AI is not the solution. AI is going
to take this problem to a whole another
level. Even the mega mega journals, even
the predatory journals are now
struggling to cope with the swamp of
waste that is being produced by AI um uh
algorithms. This is not the future of
health and social science.
This threatens to be the death of health
and social science. That is my argument.
I've got six different points I'm going
to make which obviously I will whiz
through. We will start with this
argument that AI tools cannot actually
help us with scientific reasoning. So to
understand what AI uh tools and indeed
LLMs can and can't contribute to health
and social science, we have to start by
understanding what they can and can't
do. What are they? Well, what the
architecture tells us, LLMs are language
and image calculators. They use NLP and
computer vision to interpret prompts and
produce statistically plausible patterns
that resembles their training data.
They do not understand.
They have no capacity for reasoning.
They do not understand the difference
between truth and fiction. They do not
even understand what a letter is or a
word. It is just a pattern.
So an LLM when we ask an LLM a question,
it does not think.
It predicts a pattern of words or pixels
that would typically follow a given
prompt. As Bergstrom and Back Coleman
have described it in 2025,
they are like autocomplete in overdrive.
Right? So that is the reason not just
for LLMs but for many machine learning
algorithms such as driverless cars. That
is why we have not achieved a world
where they can do this safely by
themselves. Because when they see
something that they have not seen in
their training data, like an upturn
lorry, they do not understand what that
means. I can tell you a human being in
their very first lesson would not plow
into an overturned truck. But after tens
of millions of hours of training, an LLM
which we claim or an algorithm which we
claim has understanding would do
something like this.
So this makes them very well suited to
summarizing what we know
but not so well suited for scientific
discovery
because scientific discovery is
something else. Scientific discovery is
about moving beyond what we've observed
and understood before to a new level of
understanding to new models of the world
to new meaning making about everything
that's happening around us
as John Pearson has described the
problem is the future is not in the
training data
and in fact the problem is bigger than
that because what is the training data
made of
waste.
Do we really want to look back at the
shoddy research that we've done in the
past and reformulate that in some way
and claim that we're doing science?
We know that LLMs have an inherent
fault.
They hallucinate. We ask, "Are children
small or just far away?" And this
intelligent machine tells us, "Children
are not actually small. They're just far
away. Well, we know that's wrong because
we have a mental model of the world. It
does not.
This is complete nonsense. It's
obviously nonsense.
But we don't use these algorithms
when we already know the answer.
We use them, we ask questions of them
that we don't already know.
And when we're faced with those
questions, there is a very dangerous
risk of being misled.
And this is a very simple real world
example. This is Paul Zivich. Um he
asked an LLM to summarize one of his
papers.
He says the main summary was accurate,
but he also asked it to extract and use
code from his paper, which it did. It
just happened to introduce an error, a
subtle error. A subtle error that meant
that the standard errors would be
incorrect. In other words, the
confidence intervals would be incorrect,
the p values would be incorrect, and
potentially the inferences would be
incorrect.
Would you notice that if you didn't
already know the code inside out to
begin with? I would argue not. And for
that reason, in practice, LLMs often
require extensive handholding and
validation checking. Right? So this is a
a a fascinating study of using CH GPT as
a tool for biostisticians. It came out
earlier this year. And what they
conclude is that while some tasks were
completed rather satisfactory, others
suffered from severe issues. So as a
consequence they give this nice list of
things that you have to do.
Double check this, double check that,
scrutinize this, scrutinize that. The
list is so long and tedious that I am
left wondering why would a human not
simply do this themsself rather than
rely on a machine and then have to check
every damn thing that it's done.
So the reality what is the impact of
this in the real world
is that these machines do not match the
marketing hype and we don't have to look
very far to find that out right there
was an MRT study only a few months ago I
think
which estimated that 95% of efforts to
integrate Gen AI into business were
failing and many of the businesses were
aband abandoning their efforts because
they were not having any improvement on
productivity and it was not having any
improvement on the bottom line. And so
the question that I would ask every
scientist in this room,
what makes you think it's any different
for us?
Right? There are people out there who
are spending good money right now on
trying to find out what LLM can do for
them and abandoning it because it's not
working.
Why do we think it will miraculously
work for us?
Right, I'm going to move on to the word
of the year according to Webster's
dictionary. Slop was announced today.
Uh this is um David's paper which he's
already alluded to and colleagues. Um at
the end of the paper
um you argue which you you argued in the
the presentation as well. We live in an
era that incentivizes scientists to
produce masses of papers of questionable
quality. Right? We agree that is a major
problem.
This is a human not an AI problem.
I completely disagree.
I think that's a logical fallacy. I
think it's a false dichotomy. It's
equivalent to saying guns don't kill
people. People kill people.
Well, in reality, if people kill more
people when they have access to guns,
guns are a cause of death. If people
produce more slop when they have access
to AI, AI causes slop.
And we don't have to look very far to
see what a problem that is. Right now,
LLMs are cat catalyzing scientific slop
on an unimaginable scale. Right. There
was an article in the Guardian only last
week about AI research. AI researcher
who was proudly boasting he published a
100 papers in a year and he'd used AI to
help him do it.
And if you speak to any journal editor
right now, I can tell you what the
problem is.
They are struggling to cope with
fraudulent submissions, spam
submissions. Some of them they they pick
up. Others such as the famous massive
testicled mouse
somehow got through the predatory
journal system.
This is bad, right? This is worse than
we could have ever imagined. The state
is so bad
that journals even from some very
questionable publishers are banning
submissions with open data sets because
of the sheer number of lowquality
clearly AI written research that's being
submitted to them. And similarly, even a
preprint journal, I didn't think they
even had standards, right? Even a
preprint journal has paused submissions
about AI topics because of the sheer
influx of clearly AI generated slop.
This is an existential problem for our
system and it's one that's recognized by
publishers and by funders. This is an
article by uh this was a report by
Cambridge University Press earlier this
year warning that important work risk
being lost or drowned out by the surge
of lowquality and AI generated content.
And this is from cancer research UK. We
are being swamped by AI generated
content that is endangering the very
foundations of objectivity and
empirically derived facts.
So we're in trouble.
But it gets worse, right? I think there
is an there is a genuine risk to the
scientific craft.
Okay. If we look at the extreme
proponents of AI, which um this could be
a an example. This is um an agentic AI
model that promises to generate novel
research ideas, write code, execute
experiments, visualize results, describe
its findings uh even by writing full
scientific papers. It's the one-click
scientific paper
um model.
And what is argued is that by giving a
uh algorithms these uh tasks then we are
liberated as human scientists to work on
more creative matters.
Well, I would ask what task are we
liberating ourselves to do if we hand
over every single stage in the
scientific method.
There is actually I can't think of a
laborsaving machine that has been
produced
that has ever freed us as human beings
to to do more meaningful work. This is
what's always promised, right? We were
told we would all be working 15 hours a
week.
How's that getting on?
These tasks are the things that we as
scientists are trained for. They are the
task that we are skilled at. They are
the tasks that hopefully most of you
find rewarding and get great value from
doing.
Would you not rather struggle through
those
than be stuck in some dystopian future
where instead of doing research ourself,
we are simply inputting information into
an algorithm and then interpreting the
output and checking for all the mistakes
that it's made.
Even if for some reason you're happier
with that future, which I am definitely
not,
I would argue that you'd probably be
worse at it than if you left the LLM
alone because there is evidence that we
as human beings perform worse when we
collaborate with AI.
It's well known across many disciplines
automation bias occurs. So we become
overly dependent on an algorithm once
it's there and we lose our ability to
make critical judgments and in the
scientific and academic field that's
cognitive offloading bias. So when we
use LLMs it appears to damage our
critical thinking skills, our most
important skill.
And there's a bigger implication. The
more labor that we give to LLMs, the
fewer opportunities we have as
individuals to learn and maintain our
craft.
So there is a genuine risk of a
collective deskkilling of all of us and
of the scientific workforce. This
creeping reliance on LLMs risk
displacing human scientists,
particularly junior scientists, meaning
there'll be fewer jobs.
There'll be a a breakdown of the
training pipeline and potentially a
complete collapse of the scientific
workforce.
So there's a dark future here where
there's just a few senior academics
floating around. They don't have PhD
students. They don't have post postocs.
They just cost giant LLM subscriptions
into their grants and then spend their
life submitting and interpreting all of
the information.
>> I move on to more jolly matters.
>> Oh, great.
It's often assumed, it's been implied
that LLMs will get better and better.
Is that true? I don't think it's
necessarily true.
We know that existing LLMs have now been
trained on the majority of publicly
available data.
So the the scope for scaling is much
more limited than it was at the
beginning of their life. And indeed, the
improvement between chat GPT4 and chat
GPT5 was so modest that it left a lot of
people wondering, have we hit the wall?
But it's actually worse than that
because the training data is degrading.
Over 50% of the internet that now is
created is thought to be created by AI.
So tomorrow's algorithms will be
learning based on flop. And guess what
happens? Model collapse.
So the future does not look so bright
there. And it does not look so bright
when we consider that we as humans are
contributing less and less as we use
LLMs more and more to to wonderful rich
human platforms like Wikipedia and Stack
Overflow. If you think that LLMs are
really good at coding, where do you
think that knowledge comes from?
Humans [clears throat] who are not
providing that knowledge anymore.
So a realistic future, a genuine risk is
that future models will not be better,
that we may have actually reached the
peak and we're heading towards something
rather disappointing.
But even if we're not, even if the
technology will continue to improve,
what happens next if we contract our
skills and our reasoning to LLMs?
Well, most people when they're using
LLMs are using commercial LLMs.
Let's have a think about what happens
whenever technology promises something
faster and cheaper.
How's that working with food?
Well, we have this abundance of cheap,
lowquality food that's probably killing
us and is certainly having an enormous
environmental impact.
What about fashion?
Again, we have an abundance of cheap
clothes of incredibly low quality with a
high human and environmental cost. Okay.
What about tech?
Well, what's happened with tech?
When Google first came out, it was good.
Then came the ads and the sponsored
results and a worsening user experience.
And the only way to escape this, and
it's the case with every single
technology really out there, pay more.
What do you think is going to happen to
commercial LLMs?
The same. Why?
Because right now when you use chat GPT
even if you have a subscription you are
not paying for it. Someone else is.
Enterprise investment.
Hundreds of billions of dollars of
enterprise investment is being pumped
into these things to make us use them.
Well, they're going to want all of that
money back and more because the only
reason they're making that investment is
the knowledge and the idea that in the
future we will all be trapped and we
will have to pay and pay and pay.
So I guarantee I guarantee
chat GPT every single commercial LLM
will become in shitified over the years
to come. We will have ads. They're
already experimenting with ads. The
thing that scares me beyond belief is
sponsored results.
Who is going to pay to manipulate the
knowledge that I receive when I use one
of these algorithms? And what damage
will that have to society?
The experience will get worse. And the
only way to avoid it is to be the lucky
wealthy people who can afford the
everinccreasing subscription costs that
will come our way.
So I've already said there is a serious
risk here of us funneling more and more
money away from human scientists to
industry and we have already made
painful pacts with the publication
industry that we cannot get free of.
But my GOD IF YOU THINK IT'S HARD to
break away from journals imagine what
it'll be like when we become completely
dependent on these things to do all of
our research.
I think that is a very serious ethical
concern, but it's not the only serious
ethical concern.
As health and social scientists, we have
to think what are the implications of
everything we do for population health
and so society. Well, LLMs are built on
mass intellectual theft. We produce the
work, we share it with the world, they
scoop it all up and sell it back to us.
What does that sound like? This
undermines our labor and our
livelihoods. But there's worse. When you
think an LLM is magic because it's got
some mystery data and algorithm, I will
tell you there is an army of invisible,
exploited workers behind the scenes who
are who are doing all of the moderation.
They are looking through the output and
they are having to deal with extremely
traumatizing material. But we forget
about them because they're in the the
global south.
And then of course there is the energy
cost. LLMs are extremely power and and
and water hungry. And by 2030 it's
estimated that their data centers will
draw 5% of the entire global energy
need.
How can we afford that in a world where
we desperately need to be using less
energy? I want you all to consider this
as a moral question.
What human, social, and environmental
cost are you personally willing to pay
to speed up your next publication?
So, my my positive frame,
there is another way.
Right? So, I'm going to start with this
quote from Rosa uh Ronano.
I want to shine the spotlight on the
commonplace assumption that productivity
must always increase. Good research is
disruptive and thinking time is central
to high quality scholarship and
necessary for disruptive research. In
other words, what do we need?
We need in the fine words of the late
great Doug Alman less research, better
research, and research done for the
right reasons. And I cannot think of a
moment in the last 30 years where these
words are more important for us as a
community. trust and supported science
is at an unprecedented low. This flood
of derivative and fraudulent research
from AI threatens the entire system and
simply chasing greater productivity will
not save us. Instead, we must try and
take the slower path. We must spend more
time thinking and collaborating with
human colleagues to produce better and
more thoughtful research. So just to
summarize in answer to the question is
AI the future of health and social
science research I would argue
emphatically no it risks being the death
of health and social science research.
Thank you.
[applause]
So we've got now five minutes just of a
little back.
>> Yeah. Okay. Thank you, Peter.
Um
I think in general we are scientists and
so we should try and back up our claims
with evidence. [snorts] So you made a
number of claims which I think could be
disputable. Um, and I think progress is
where we can make progress is how we can
actually generate that evidence to
inform these decisions. So the the
self-driving car analogy was
interesting. What are we aiming to
achieve there? What's the comparison
group? You show a photo of a car that's
broken down on the on the on on the
road. What we actually should be
comparing is mortality rates with
self-driving versus human driven cars.
Those are the relevant comparisons.
You mentioned people are abandoning this
technology. Yet purportedly it's the
fastest growing consumer product of all
time. chat GPT within with you know you
know incredible use of launch
enterprises
you mentioned it's will not get better
with a reference from August 2025
um I showed a reference um from a few
days ago where clearly there is
improvements there isn't any great deal
the improvements seem to be continuing
um particularly as we scale up reasoning
so you said that the wrong tool for
scientific discovery scientific
reasoning that they can't reason at all
you know I used to have that view
particularly of early LLMs, but as we've
begun to train reasoning into them with
reinforcement learnings, they have what
appears to be reasoning. So if you give
them your exam script which explicitly
requires reasoning, they give you their
reasoning steps which would get great
marks and they give you the correct
answer which would also get great marks.
They also do well in tasks which clearly
require some degree of reasoning however
we define it. For example, gold medals
in competitive domains.
And I think you know they're unsuited to
scientific discovery. Well, scientific
work is multiple different cognitive
tasks and at our discretion is the task
that we do or do not give to to AI.
Um
you said that the problem with health
and social science is waste and I think
we agree that there's a problem with um
people being wrongly incentivized to
produce lowquality papers.
So the solution to that is that we try
and incentivize um higher quality few
fewer papers. We don't stop using tools
simply because of that problem. And if
you look for example this has been a
problem whenever you make a tool to
increase productivity. So the two sample
mandelian randomization packages there's
been a surge of men randomization um
papers particularly from China.
In the UK we already tackling this and
it's actually difficult human problem.
How do we get the incentives right? So
in ref for example we have a maximum of
five papers and we're moving to
narrative narrative see these in grants
we have to acknowledge that science is
global it's not just the UK
so LLMs cannot liberate us from science
they take away the scientific craft they
are a general purpose technology so we
can use them at our discretion
we can use them to free us up to do more
meaningful creative work or we can use
agents to do all the scientific
discovery
I do not have the evidence that suggests
that those agents should be automating
all scientific discovery. That's not a
very defensible position. These things
have come out in recent months and
they're an interesting technology
development. But I think there's clear
wins where they can add to scientific um
discovery.
So you mentioned the future of LMS is is
in shitification and I think that I mean
I think that that's speculation there.
We still use Google, Google Scholar even
with the ads, right?
If a the beauty I guess of of
capitalism, if if a company over
advertises, it's a very if it's very
competitive environments, you can simply
switch your LLM provider.
Even more so, if you don't want to use a
commercial LLM, you can use an open way
LLM or even a fully open- source LLM.
And perhaps as the research community,
we should be building those out if we're
concerned about the downsides of
commercial LLMs.
You mentioned ethical concerns and you
cited what I think was a Guardian
article. I think every major technology
has costs. We are empiricists. We should
be quantifying those costs and having a
careful having a a proper evidence
discussion about how much they're
costing to train, how much they're
costing to use.
And you know there's a report by um UCL
colleagues that came out a few weeks ago
that suggested that we can reduce energy
costs by 90% simply by using smaller or
quantized large language models. So LLMs
are coming people are using them we have
a choice in how we use them in the
future. And if energy consideration is a
big point we can try and um make them
more energy effic efficient and advocate
for the more energy efficient uses.
And also what do we compare them to? you
know, if we compared it to the
entertainment industry, if we compare it
to cryptocurrency mining, the LM is a
technology which has got clear proven
benefits um in my opinion and and the
efficiency costs have also improving
across time. So, for example, the
Deepseek model which came out earlier
this year seemingly a much more
efficient model to train. If models are
getting more efficient to train, they
should for each individual query be more
energy efficient.
So you also mentioned the case for
slower and more thoughtful science. I
think we're all in agreement. We want
higher quality science, not lots and
lots of redundant lowquality papers.
So overall I think that um
but but I think you know slow itself is
not virtue.
If you can do your paper in six years or
six months and it's the same quality,
what should you aspire to be doing?
What should the public who are funding
your research want you to do? Spend six
years on it because you enjoy that
process or six months on it with tools.
You could even spend long on it if you
use no tools at all. What about the what
about the energy and ethical costs of
industrialization as a whole or using a
computer or using a smartphone?
We at the early stage of this
technology. So, we should be measuring
those costs and then making um balanced
judgments.
I think I'd also add as well just in
terms of the positive potential of this
technology. Um I don't think in any
major technological change like this
it's been so evenly accessible from
around the world. So anyone around the
world can essentially use these models
essentially for free and it's available
in essentially in every recorded human
language. Whereas other technological
revolutions for example smartphones they
were largely restricted to high- income
countries at least for the first period.
So there's an argument here that they
can have a huge positive potential
benefit to countries with lower
resources.
Education is hard to to intervene on,
but what we do know is personal tutoring
can make a massive difference to
education attainment. So we can roll
that out to everybody in the world. So
everyone can have personal tutoring
or we can simply stop using these tools.
I think in general I think you know that
AI is coming whether we like it or
whether we don't like it.
The key thing is what we do, what we do
with that information, what we do as
scientists, what we do to shape our
ecosystem to make it work well for our
our goals.
>> I think I'll finish that.
>> Thank you very much, David. Well done.
[applause]
>> Five minutes, Peter.
>> Okay. Thanks, David. Um I I'm desperate
to respond to the points that you've
made then, but um really I I I need to
try and respond to the the main uh
points that you made in in your
presentation, which I haven't covered.
It's probably a case that we'll have to
agree to disagree uh about the inherent
ability of AI um to actually have
cognition and to be creative. [snorts]
Um clearly I would argue not. Um, I
think they're excellent at cosplaying
reasoning. I think there's an awful lot
of pattern recognition that can take
place from the incredible amount of
learning that it's done, that means it
can quite reasonably guess things that
looks very like reasoning but isn't. Um,
and I think when you see them misbehave
in incredibly ridiculous ways, which
happens if you ask them often simple
mathematical tests or strangely simple
questions and they just completely fall
down, that is a good reminder under
there is not reasoning, right? It
doesn't matter what what uh figures
someone has produced. There's something
very strange there that indicates it's
not really understanding what you're
asking.
Um, you talked about using uh LLMs for
research question and hypothesis
generation. I personally have an
enormous problem with this. Um, I I
think that the the the process of of of
hypothesizing and of constructing
research questions is one of the most
sacred and challenging aspects of being
a scientist. Um [snorts] I think that we
are embedding our mental models and we
are embedding our values when we are
formulating hypothesis and constructing
uh research questions and my concern
with relying on LLMs to do that is that
of course they have baked in all of the
biases and the waste that they've seen
before right so we are trying to move
beyond a world where we're simply
repeating colonial and patriarchal
messages the only way to do that is to
be coming up with questions based on the
society that we live in now and that
means using your human brain.
Um, in terms of systematic reviews, I
the idea of them being able to perform a
systematic review in one click, even in
five years time, I unless they solve
hallucination, which they can't because
hallucination is a fundamental feature
of the way that they work, I can't see
that ever uh happening. But I I will
concede that I think it's reasonable to
use them in systematic reviews,
particularly as a secondary screener.
But again, I would say the biggest
problem with systematic reviews, as I
see them, is a complete lack of critical
thinking. We will just combine all of
the results that we've seen published
before without proper reflection of the
challenges of these papers. Yes, of
course, we have, you know, corporate
guidelines and so on. But still, people
are not conducting the kind of reviews
that we need, the kind of reviews that
take more than six months. I don't think
I've done papers that take six years.
They there's absolutely no human way to
do them any sort than that or no inhuman
way to do them any shorter than that
because it requires six years of back
and forth thinking and conversation with
other humans over many many years to
understand an unsolved puzzle or to
clarify an idea beyond the knowledge
that we already have. But clearly there
are some examples where some people may
drag things out longer than they should.
When it comes to coding, I think this is
a really interesting one. My brother is
a a a software engineer. He has some
very strong opinions on this as well. Um
I think they work very very well for
basic tasks. Um I think they probably do
work pretty well for bug detection, but
in my experience, if you're trying to do
anything remotely creative, it is a
nightmare. There is a genu genuine risk
that by choosing to work with an LLM,
you choose to be the hair in the hair
and the tortoise, right? You think you
race off, you've got all these initial
ideas, you think you're halfway on your
way, and then you get grounded because
you don't have the baseline
understanding that you could have had if
you'd have started reading and planning
and coding with your own knowledge. And
dare I say, some people don't find
college coding to be a repetitive task.
Some people find it to be a very
rewarding um problem-solving task that
maybe we should um appreciate um a bit
more. So
I think I think that's enough. Thank
[laughter] you.
>> Thank you very very much. EXCELLENT.
[applause]
>> Thank you very much.
>> Very well done. I really enjoy. Um I
think I would like to uh invite both
speakers to come forward. Um and
basically we have now uh 20 minutes or
so for question and answers and debate
and so we obviously have plenty of time
over nibbles and and wine and so on to
discuss that further. Um we do have
maybe the voting system
>> at the end. Yeah. Once we've done a Q&A
we'll do a final vote in the last minute
or two and see who has fared worse.
>> Do we need u microphones for people? We
don't really have individual microphones
[clears throat] so maybe just have to
speak up loudly.
>> Yeah. Do you have any questions?
Immediately we have a question.
>> Yes, I really enjoyed your um your your
point and very passionate against it. I
wonder so talk about LLMs and
hilizations things. I was wondering what
you think about other alternatives in
machine learning such as alphafold which
obviously just won the Nobel Prize for
protein folding. Do you see different
areas of machine learning having
desperate impacts?
>> Yes. Um in in short um clearly uh there
are certain prediction tasks um where
machine learning is extremely well
suited. Um but uh the one thing I would
say is beware of the hype right because
they can do certain things extremely
well. People assume that they can do
many tasks extremely well and really
it's a case of individual tasks being
very very well suited and other tasks
potentially not being so well suited.
But clearly for some things fantastic.
>> Yeah, I agree. I mean, I think to keep
it reasonably fair, we steered towards
LLMs and agents in this debate, but
there's clear great utility and uh in
what Alpha Fold and other another
frontier prediction models have done.
For example, meteorology and they've
been published in great journals.
They've open source their resources and
had a good contribution to science from
those commercial companies.
>> Great question.
>> Yes, thank you for that. I think Peter
alluded to the fact that um AI sorry
hallucinations are a fundamental feature
of AI. I was going to ask you David, do
you feel that it's a fundamental feature
and do you feel that we can actually
overcome that at any point in time or do
you think it's just something we have to
kind of accept as being the
>> Yeah, I think they're pro there probably
is a degree that we have to accept.
They're like they're not deterministic,
they're probabilistic in the output they
choose. Um so but there's ways of
engineering these LM so the
hallucination rate produces and
empirically we can track that with
hallucination benchmarks. It's also I
mean in my personal use I tend you're
going to hate this Peter but using
multiple LLMs at the same time and kind
of triangulating off them. Maybe some of
them have some of them have certain
biases and you can get the output of one
and give it to another one feed it to
another one and that way so people have
been exploring that in a more formal way
of like scaffolding work around it to
try and reduce hallucination rates
further and the reinforcement learning
is one paradigm for that.
>> So maybe there are questions around
methods and how to
>> you should answer this one. Uh the
question was about hallucinations. I I
would just agree that that they're a
fundamental feature. I mean I think
you're right that the there are things
you can do to to try and reduce the risk
um but not remove the risk.
>> So yeah, that was your question. Do you
think at some point we can live without
uh yeah going towards not having these?
>> I don't because I personally because
there's no sense of truth and and
fiction, right? So they they're just
producing a pattern and sometimes that
pattern will be wrong because it has no
idea what the concept of right and wrong
is.
>> Can I just say so do humans as well. So
[laughter]
>> there was a question from over there and
then I think
>> um you you were next. You you raised
your hand if you
>> might have to stand up and shout. Do do
you think that uh AI
as it's currently conceived have a
positive
negative or neutral effect on the
development of theory in social science?
>> I'll stop. Um I think right now I I see
nothing that makes me optimistic
unfortunately. Um what what really
scares me is just how willing a
surprising minority are to just embrace
scientific fraud. It's like they were
there waiting for the right tool and now
it's here. Um now it may trigger a
renaissance, right? Like I genuinely
believe there is a possibility that as
we get in influxed with slop, we as a
system say right, we need to do
something different, right? and that
something quite wonderful could happen.
But right now, I'm not optimistic.
>> I think it depends if we get the
governance and the and the incentives
right. Um that's the key thing. Um but
we should also be you know tracking
quality of papers and empirically trying
to answer that kind of question. And
different disciplines have tried to sort
of um ensure quality in different ways.
Um so you know in epidemiology you can
publish a descriptive paper just a
simple descriptive statistics in the top
economics journals they really want
cause identification and um like a very
careful empirical um strategy around
causality but even there there's there's
concerns about the effect that's had on
the discipline. So we have to get the
governance and the incentives right
learning from all the different
disciplines. Each of them has things
they've got right and things they maybe
haven't.
>> Yes. So there was Liam and then
yourself. Yeah.
>> So part of your argument was that um AI
is bad for the craft of science and I
just wondered whether you believe that
not using AI at all will make you
advance in your field and actually
people that follow that strategy might
win out and then AI use itself would
lead to less use over time just because
they were succeeding.
There this is an interesting thought
experiment as well um which I think has
some merit right as things pan out there
is no question in my mind that there'll
be a massive concern about AI slop
and that a little bit like you know the
tailor who continued to produce
magnificent clothes on um what's the
name of the street savro right these
people survive because they continued to
produce produce good quality stuff. They
didn't chase growth, right? They didn't
feel we need multiple several rail row
shops. They just said, "Look, I'm going
to do what I do really well." And they
survive. Um, and I think that there
there could be a route for some people
to do that to actually really embrace
slow science, uh, you know, masterful
science that then the community
appreciates. My only concern with that,
which I should flack, is who has the
luxury and the privilege to do that. I
am aware that I stand up here talking
about slow science with a profile that
says it's easier for me to follow that
path. Um whether it will be so easy for
everyone and what are the ethical
implications of that I I I'm a little
bit scared by.
So in general I think the best approach
is to you know a combination of human
and AI. So use these tools to help you
learn to help you do better science. Um
that's what I personally think is
favorable. If people can do great
science without AI and without LLMs
fantastic. People should have the people
have the choice to use whatever tool
they want.
>> Yeah. Lots of questions. So I you first
you and then over there and I would like
to encourage all all women to ask
questions as well. I see a lot of hands
up by by men which another gender bias
is so should bear that in mind. Yeah.
>> So you first
>> uh you both raised in your discussion uh
question to what extent AI's
contributions are derivative but I was
going to ask to what extent that
actually matters. I mean Newton had this
line about I saw further it was only
because I stood on the shoulders of
giants. Kak McCarthy the writer said
that it's a sad fact that all books are
made from other books. I mean to what
extent is our AI just mimicking human
creativity by ingesting what's come
before and then replicating something uh
something new from it or combining new
information out of old well understood
facts.
Um, I I I I think one of the problems we
have in in science right now is an
information overload. Um, I'm aware of
articles that have been published over a
hundred years ago
warning about methodological issues that
the vast majority of analysts don't even
know about. Um, so we do have a problem
with volume and my concern is that
encouraging that volume and seeing it as
part of the system is not going to solve
that problem. Right? Actually, we we
desperately need to to radically reduce
the amount that we're producing so the
signal to noise ratio um improves.
>> Can I say to that one? Um, so I don't
fundamentally know how the LMS think. I
also don't know how human brains think
either. I think we're both neural
networks of different design but they
seem to to my view they seem to
approximate creativity and approximate
reasoning maybe in a different way but
they approximate it and the key thing is
can we actually evaluate that is
actually is that is that claim actually
falsifiable
and so if they're actually contributing
to new knowledge um that to me they're
making a contribution I've talked about
reasoning a bunch in in the talk in the
rebuttal.
>> Thank you. So there was a question here
and then yourself. Yeah.
>> So I appreciate it. uh David's point
about benchmark saturations
and uh there's a point to be made that
benchmarks can be saturated on areas and
benchmarks can be done on areas where
there are right answers that are no. My
question to both of you is are there
areas of health and behavioral sciences
that lend themselves better or worse to
these benchmarks uh that lend themselves
more to the alpha treatment um rather
than the sloth treatment.
>> Yeah. [snorts]
benchmarks. Um,
there's a lot there's hundreds of
different benchmarks, but they don't
seem to be being produced by people who
work in our field. And so I wonder if
those benchmarks aren't suiting us if we
should not be creating them ourselves.
It's kind of like cognitive testing is
you tend to measure in a very narrow
domain and you wouldn't recruit a
researcher based on a single test score
in the same way you don't know how good
an LLM is just based on its benchmarks.
So I think we're we should be
contributing to that that that that
work. I think
>> I think in the business world this can
be done really well. We've introduced
this tool and oh there's been no change
in our productivity or there's been no
change um in in you know our bottom
line. I think ideally it would be great
to be able to do something very similar
in science to actually perhaps even
conduct an experiment. You have two
teams, you have access to LLMs, you
don't. Let's see what actually happens
um at the end of that. That would be
fun. Um, I don't want to know the
result.
Um, because I would then be arguing,
well, that wasn't really a hard enough
task. Um, but yeah, I think that's
something we should be doing. We should
be collectively finding ways to evaluate
these things in terms of the end goals
that we're actually interested in.
>> Yes. So, yourself first and then over
here. Yeah.
>> Hi.
>> Sorry. Did you want to say something in
return? I didn't get you. Okay.
>> Did I get chance? I forgot to answer
that one. Don't forget I did.
>> I think you did first.
>> Um,
>> if I delegate a task to a colleague, I
end up with a more experienced colleague
where it's very delegate to AI.
What maybe I need to be training
colleagues so they're better at delegate
to AI. But when it comes to training,
what what do you think we need to do for
using AI?
>> Um, there's certainly things you can do
to help people get the most out of AI.
for example, best practices in prompt
engineering, how you can get better
performance out of them like adding in
examples in the prompts and things like
that. So there's a way we which you can
teach people to use AI well and teach
them on for example on science the kind
of areas you you kind of encourage them
to use it versus not encourage it. Um
the training point is really important
because I mean we need to protect the
career structure of scientific discovery
and it's already not satisfactory.
[clears throat]
I mean people have short-term contracts.
we're struggling to retain talent. So,
we need to build a a robust and
defensible pipeline of careers in
science um now more than ever. I don't
work on that, but I hope very careful
social scientists are
do you want to
>> so I I would agree. I I don't know the
answer to be honest.
>> Um
>> yeah,
>> one last question. There was
>> maybe two more because there's the lady
here and over there. Yeah. So, maybe
these two and then finishing the Yeah.
Thank you very much.
>> Hi, I'm Vanessa. Um, I would like to ask
about the critical thinking part. Um,
Pete mentioned about the negative
effects of music loss on critical
thinking. So, do you agree on that? And
if yes, how do is there a way to counter
that? So, we could use
myself. Um, so I think the jury is still
out about overall do they diminish our
our our creative thinking or can they
actually augment it? uh it's just so
general purpose. It's like speaking to a
human expert in any domain. So you can
you can work with it in such a way that
it can actually improve your creative
thinking I would imagine. And the early
empirical evidence on the effect of
tutoring with LMS is to my mind is
extremely positive in terms of um
helping people learn things including um
creative thinking. We it's it's it's a
tool. It's here. We have to figure out
how to mitigate any bad bad bad effects
and how to promote the good effects.
>> That might be one of them.
>> Thank you. and yourself.
>> Um, hi, I'm Georgia. Thank you. Thank
you both so much. Um, one comparison or
argument I've heard about using um, LM
is for example, if you were to go to a
marathon and take a taxi to the end of
the marathon, you wouldn't expect a
medal at the end. And I wonder whether
you agree whether this translates in a
similar way to science. If you haven't
struggled through the craft of science,
would you get the same reward in the end
if you did it yourself or if you did it
with an engine?
>> Was that more for me?
>> It's a good question.
>> Well, you you can start.
>> Um, no, I agree. Like you need you need
like an apprenticeship of doing things
manually to learn. The question is like
to what extent and to what degree do you
want junior colleagues spending you know
an entire calendar year on screening or
can that time be reduced so that they
spend enough time on it to learn but not
too much time that they can't do other
tasks like develop their creative
thinking and how we bring multiple lines
of evidence together to do better
quality papers.
I think your your question made me think
about something that we we don't even
discuss or think about
which is do we enjoy the job
right why do we run a marathon we don't
run a marathon to travel whatever it is
however many miles 21 miles I need to
just travel from A to B so I'm going to
run a marathon we have some other
intrinsic reason to do that and I think
for many of us who were driven into
science as well there was an intrinsic
desire to understand things to question
to query to explore experiment and that
is actually a concern that maybe we're
not as productive
but if it's more enjoyable
then isn't that more important at the
end of the day what's the point of all
of this if not also to enjoy our lives
in the process so I would say yeah you
know you're not a carpenter if you go
and buy a chair from Argos right
there is something about the craft that
we have to actually respect that's above
and beyond output and productivity.
>> Yeah,
>> I think we've run out of time so I think
we need our followup.
>> I think we can
[applause]
diplomatic. Yeah, I think um AI is here
to stay and um obviously there are many
many pitfalls and many many advantages
and and things to be gained and I think
it's our responsibility to think about
how can we use it responsibly and that
is our job and thinking about you know
why are we doing it and and what's the
purpose rather than just doing things
for for the sake of it. So I think it
was a fantastic debate. Thank you very
much for the preparation um and and the
back and forth and so on. And thank you
very much for fantastic questions,
really [laughter] thoughtful arguments
and I'm very sure we can have um yeah
more thoughts and more questions and so
on during the during the um reception.
Thank you very very much for coming.
>> Okay. Have you all voted? [applause]
>> All right.
Can we see the Can we No. Hang on. You
want to see the results?
>> Can we see the code again? Ah, can we
see the code again?
>> Right.
>> Yeah. Okay.
>> This was the before.
So, the before was 49% yes. 36%
no.
And the after,
48% yes, 46% no. I think we've both won.
>> Or I'm going to argue that this is
within something variation. Cool.
[laughter]
>> Within varants. Yes. So people may have
changed from one name to the other.
>> Yes. They're not even necessarily the
same people.
>> We've managed to drag some people from
don't know which I think is a a success.
[laughter]
>> Right. We've definitely earned drinks
and nobles. Well done. [applause]