Submind YouTube summaries
Thumbnail for Is AI the future of health and social science?

Is AI the future of health and social science?

Watch on YouTube

Video summary

Professor David Bun from UCL argues that artificial intelligence represents the future of health and social sciences because it excels at the cognitive tasks central to quantitative research, such as literature reviews, data coding, and administrative work. He points to rapid advancements in large language models, which now surpass human experts on complex benchmarks like GPQA, while costs have plummeted by 99%. By automating these manual drudgeries, AI frees researchers to focus on high-level scientific thinking and hypothesis generation, with examples showing AI agents completing systematic reviews in days rather than years. However, Peter Tenant from the University of Leeds strongly contests this view, warning that current incentives already drive massive research waste and that AI will exacerbate this by enabling the mass production of low-quality papers he terms "slop." Tenant emphasizes that LLMs are merely pattern-matching tools lacking true reasoning or understanding, making them unsuitable for genuine scientific discovery and prone to dangerous hallucinations that can undermine public trust in empirical science. The debate deepens as Peter highlights severe risks associated with relying on these models, including automation bias, cognitive offloading, and a collective deskilling of the scientific workforce where humans perform worse due to diminished critical thinking skills. He further warns of "model collapse," a phenomenon where training data degrades because over half of internet content is now AI-generated, alongside concerns about the commercialization of tools leading to ads and high costs accessible only to the wealthy. Peter also raises ethical issues regarding intellectual theft, the exploitation of invisible laborers in the Global South, and the immense energy consumption of data centers, concluding that AI threatens the very existence of high-quality research. In response, David counters that LLMs are continuously improving through reinforcement learning and reasoning training, making them suitable tools when used with human discretion, and disputes the inevitability of commercial degradation by noting users can switch to open-source models or smaller, quantized versions that reduce energy usage significantly. Despite their sharp disagreements on the trajectory of AI, both speakers find common ground on several critical issues necessary for the field's future. They agree that proper governance and incentives are required to prevent the proliferation of low-quality research, and both stress the importance of reducing information overload to improve the signal-to-noise ratio in scientific communication. Furthermore, they concur on the necessity of training scientists in prompt engineering and best practices to navigate these new tools effectively. While acknowledging that AI cannot replace the intrinsic joy and craft of scientific inquiry, they recognize its potential as a useful tool if integrated responsibly, with David adding that it could help address global educational inequities by providing personal tutoring to under-resourced regions. Ultimately, the debate reflects a complex reality where the audience remained nearly split, underscoring that while AI has pitfalls like hallucinations, these are manageable through engineering and triangulation rather than abandonment, paving the way for a balanced approach that prioritizes efficiency alongside quality.
Read the full video transcript
Thank you very much and um a really warm welcome for tonight's um debate on um AI and the future of uh well the use of AI and the future of health and social sciences and what that means uh for us. Oh, okay. So, it's really great to see so many people here. We were fully booked. Um so that's fantastic to see. Um I just want to briefly mention that this is an event that has been organized by the National Center for Research Methods and by our research grant or training grant funded by the UKI um on digital skills development and you can look at that up on our website the NCM website under innovation. Um just very briefly a couple of things about the National Center for Research Methods. We are a center that is um yeah running developing um providing a lot of training and capacity building activities um uh for example in core and cutting edge research methods in particular on a wide range of research methods areas quantitative qualitative mix digital and so on um across the career life course so for junior and for senior members of staff or colleagues across sectors not just academia of course um and we have a whole um well a high uptake of of users on our website and our resources. And um we have also published earlier this year a impact um assessment report and you can look that up uh on our website and um yeah you can find here the QR code to our website. Um some of our ESRC funding is changing but we are very um um excited to announce that we can continue with some of our core and advanced training activities. uh and that is really based on the sort of successful collaboration between University of Southampton and University of Manchester and the number of our center partners that actually includes um UCL um at as as as well. Um and yes um we have got some funding for uh particular sort of AI skills development skills activities and there you can find much more on on our website in this area. Okay. Um so very briefly about this um evening um you may have seen the um outline or the program. We will have first of all the debate um between um professor David Bun and um Peter Tenant from the University of Leeds and David from University of College of um University College London. Um and then we have sort of a rebuttal later on and then also 20 minutes roughly for question and answer so people can uh get involved as well. And then we have about an hour for drink reception and nibbles. And I think I have been told it's a malt wine for the uh sort of season. Um yeah, so and before we start, we just wanted to gauge a little bit of a an understanding of the audience and just sort of a bit of fun to get you involved. Um yeah, please take part. Um or you can enter the code. And basically it's a question about is AI the future of health and social sciences and maybe we can sort of see a little bit if that has changed later on. If you convince you or not convince you, obviously we don't see uh changes amongst uh individuals but we can sort of see see the trend. Um I think on this note maybe I need to leave this up for a little bit longer. >> 22 votes. >> Oh yes. >> 30 votes. >> Come on. Some of you need to vote. >> I have voted already earlier. Okay. Well, in the meantime, I can introduce um Peter, Peter Tenant from the University of Leeds and David Bun from the uh from UCL. Um and just very briefly um David, you're professor of population health. you've done a lot of work on various different aspects of population and public health about the causes of that distribution of health in the public and also your sort of maybe more research uh interest in more recent times includes also thinking about AI in the future of scientific research and Peter um your associate professor of health and data science um we've only had actually dealings um because of this this debate so I'm really pleased to have this connection now we've got some uh connection to the University of Leeds and Leeds is one of the center partners of NCIM um and um you've initially trained as an epidemiologist and now you're focusing on you know causal inference and applied health and social science research. So I'm really pleased to hand over um they're gathering already so I better make a move. Hang on. >> Do you want to just >> I'll just have to read it out. Okay. So yes 48%. No 36% and everyone else don't know. Okay, fantastic. Connect that [clears throat] >> and I'll make a move and I'll hand over to >> Thank you. Can you all hear me? Okay, I have a slight cold, but can you hear me at the back? Great. Okay, before we start, thank you everybody for arranging this. And I think I think What does that mean? >> Louder. Okay, I [clears throat] >> maybe you have to hold it. >> You might have to hold it. Yeah, >> have to hold it. Okay. >> God knows what that mean. >> Okay. Um, is that slightly better? Okay, great. Okay, I think we've collectively organized what might be the most and least relaxing way to spend a Monday evening [laughter] before Christmas. Um, but I hope it is interesting and useful. This is an important point of the of the history of our disciplines. As the saying goes, the aim of of debate is to make progress. So, this is my argument in one slide. Firstly, quantitative research is mostly cognitive work. In the past, that was only um doable by humans, but AI systems are now increasingly capable at cognitive tasks. And the second point is that if these systems can help re research, then they are part of the future. And I'm particularly optimistic about humans and AI systems working together to achieve our goals. So the first part then so what are the tasks that we do in our in our research? Um so this is a list that I've put together. So first of all we design studies and we collect data. So the survey research world and then we um do research papers, we review the literature, understand what's been done before. We form research questions, hypothesis, we analyze data and then we interpret it and write up. So which of these tasks are cognitive tasks? almost all of them are cognitive tasks. And so you would expect the research that looks at um exposure to AI across different occupations would find that we are quite highly exposed and indeed that's that's the case. So this GPTs our GPTs paper buried in the in the supplement on page 41 is an estimate of um how much exposure survey researchers have to AI large language models and they estimated a 75% exposure just to LLM like chat GPT and they estimated 84% exposure um to LLM plus tools that LM can use as well. I can sense the energy dipping in the room slightly after saying that. But don't worry, we can bring it back with this slide. Just because we're exposed does not mean that we're imminently extinct. So, as the example from radiology, Jeffrey Hinton, the Nobel Prize winner, um said that, you know, people should stop training radiologists now in 20 2016. And since then, there's been a continuous rise in demand and supply of radiologists because radiology isn't just binary classification of images. If time is saved in that activity, they can use their time elsewhere for to better serve their goals just like we can as scientists. I'm now going to talk about AI capabilities. I'm going to show you a few different plots. This one um shows across them from 2012 to 2024, so slightly outdated, the rising capability of frontier large language models across multiple different domains of um cognitive task. And we can see steep rises in multiple different domains, many of which are very relevant to the work that we do. For example, mathematical understanding or general scientific understanding. So one in particular I'm going to point out here is the GPQA. It's described as being a graduate level Google proof Q&A benchmark, a challenging set of questions written by domain experts in biology, physics, and chemistry. This slide shows the same benchmark broken down across the different models across time using slightly more recent data. So I think in the early LLM period the models were very new and the models were I was very intrigued by the models but I wasn't very impressed by their abilities in a sense they seem very fickle particularly with tasks that required multi-step reasoning. So this benchmark and other benchmarks are designed to capture that kind of reasoning ability and we see great increases in in the performance in those benchmarks actually now exceeding that that described as expert human level. So those with PhDs or those training for PhDs. So we're actually getting to a point where in the AI world they're worrying about benchmark saturation. How can we develop benchmarks where AI systems don't get all the questions right essentially and so we have new initiatives like humanity's last exam humanity's last exam has been designed even there in recent months we're seeing vast improvements chat GPT 5.2 two was released 3 days ago and it was only 30 days since the previous version GPT 5.1 and we're having improvements in benchmarks here. So R AI2 described as capturing advanced reasoning we see a more than doubling in the performance of this model in this 30-day period. So I think I would say that if if people have very very strong claims on what AI cannot do, we have to make sure they're up to date with the latest evidence on their performance. So AI is very powerful. Um and some further indications of this is that frontier models are now getting um gold medals in in highly competitive um competitions. for example, mathematics and uh coding, informatics, but they're also highly jagged. So, they're very they're very um the performance is very high in some domains of activity, but very low in other domains. So, for example, very high in maths, reasoning, and knowledge, but very low when it comes to using their memory, unlike humans. But still a highly jagged AI system can still be very helpful to scientific discovery. Another indication of their improvement is that the length of task in which they can complete is also increasing across time. So it used to be they were restricted to tasks that take seconds like responding to a multiplechoice exam, but we're now seeing the frontier models undertake tasks that might take a normal human 2 hours or 3 hours. So improvements in their ability, improvements in the length of time task that they can complete and finally rapid declines in their cost. So this is around 99% reduction in the cost of frontier models. So it's been described that the cost of intelligence is trending towards zero. [clears throat] The cost of cognitive work is trending towards zero. Our work is cognitive work. So of course there will be implications for our field. So the second part then can it actually help us in research? What is the goal of research first of all? Well the goal of research is to expand knowledge. This is the Vcarti definition. It's work undertaken in order to increase the stock of knowledge. And in doing research, the tools that we use change, but the goal does not change. You can do research without any tools at all, just using paper or even without paper, just speaking to another human being. Or you can use um tools like a computer. You can use tools like a computer with some AI or a computer with lots of AI. The key criteria is does your um work contribute to the discovery of new knowledge? You can do exactly the same paper and it may have taken you 10,000 hours in a very manual endeavor or it might taken you um you know 10 months or 10 weeks using frontier tools. If the contribution to to the key thing is the contribution to to knowledge is is um is is the key criteria. I'll put this in a sort of a different way. So the two kind of tasks that we do in our in our work literature reviews and data analysis we historically it was very much a manual endeavor. So physically searching libraries reading papers manually doing manual calculations for our analysis. We didn't stop there. Humanity um continued to progress. Society um changed. We now have access to the early computer era where we have uh scientific research available by computers via the internet using digital databases and using computers to speed up our analysis. We didn't stop here either in the early computer era. And why would we? We're now in the era of what I describe as being early AI where we're using machine learning tools using large language models to help us make sense of increasing volumes of literature and also to help us to accelerate data analysis and coding. We're not going to stop here either. So the progress is continuing in AI and so we'll be expecting further gains as well. And I would guess that we'll be getting to a point in the future where we can do a reliable systematic review with a single click. And we can go from plain language instruction to reliable analysis with code that is auditable and code that is then executed. We're not there quite yet, but I think we'll be going in that direction in the future. The good news is we can actually research these tools as well. And that's an important contribution to knowledge. How useful are these tools? what can they do? What can they not do? In academia, we're not building frontier models, right? We haven't got the um the capital to do that, but we can at least understand these models in our work. Here's a breakdown of the different tasks and I've put on rough um time estimates of how long this stuff takes. So for the average paper let's say it takes between 3 to six month 3 to 12 months to do a literature review maybe half a month to locate the data 3 to six months to analyze data 3 to six months to interpret and write up for each of these tasks. I guess we should be reflecting how much of the work that we're actually doing is optimal or how much how much of it is lower level work that is repetitive isn't essential for our work and could be delegated to to AI. The paper I showed before the GPTs are GPTs paper it said that because you know it described GPT as being a general purpose technology. So the good news is that we can use GPT and other frontier LLMs for any task really and we can use them at our discretion. We can use them as very restricted tools. We can kind of use them as research assistants or at the very far end the more experimental end we can use them as agents in which we give tasks and they go off and complete those tasks. This may be an optimistic estimate of how long things take. It's been estimated in the US that um almost half of US researchers time is spent on admin. So even there I think this is a one area where AI can help us spend more time on actual research that we actually want to do in the first place. In the next few slides I'm going to focus on literature reviews, locating data and data analysis. In the paper on the bottom right um we discuss other tasks as well that we do in quantitative research. So we're going to start with reviewing the literature. So if you do a systematic review, how long does it take? Well, this review of length suggested 67 weeks. So over a year for a single systematic review. The Cochran handbook um advises between 1 to 12 months. But I expect many of you many of you that have done reviews will know a lot of your time is spent on painstaking and quite often painful after a series of hours manual screening of papers. You have to read through manually hundreds potentially thousands of individual abstracts to decide if they should be in your pool to analyze or not. Is that a great use of human time? So if you give humans a very repetitive and boring task, they tend to make mistakes. So this paper suggested an error rate, a baseline error rate of around 10% that humans make in reviews. That's another motivation to try and use AI. It might not only be more efficient, it might be less errorprone as well. So even Cochran are now exploring the use of large language models for systematic review work because the aim of Cochran is not that we spend thousands of hours doing manual repetitive task. The goal of Cochran is where a world where health decisions are based on timely evidence. We can't really do timely evidence if our reviews take a year, a year and a half and so on. So there's emerging evidence now that AI can do um can can be useful when classifying abstracts. This is one paper which suggested this. There's another paper here which um at the more extreme end using agents. It's estimated that that this agentic system it completed 12 cochran issues 12 coin cockrine systematic reviews in two days. Traditionally that would have taken 12 years. So if we can get to that point and if we can get to that point reliably, I I argue that we should for ad hoc reviews. Let's say you're informing what's going to be in your introduction or your discussion section. We haven't got the bandwidth to do a systematic review. We haven't got the time. So we do the best we can. I think many scientists now are actually using AI already. we're already using Google Scholar rather than the traditional tools that we might have trained in and for good reason. So for example, unlike PubMed, Google Scholar actually indexes multiple disciplines. So it it actually indexes the social science literature which PubMed doesn't always capture and it also indexes the gray literature as well. So there's an emergence of AI tools for for for finding papers. So, we've seen a few weeks ago, um, Google Scholar has a a Scholar Labs where you enter your question in plain English. What is the evidence that X affects Y? And then it finds you all kinds of studies which address that research question. We have the Claude LLM linked in with the PubMed database. And we have the elicit app as well. So these are kind of early stage technology but in the future you know we have a we have a frontier AI which has access to millions of papers greatly exceeding what we could ever read in a single human lifetime. The hallucination rates are also declining in these frontier models as well. So I expect this will be extremely useful for research. In summary for reviews I think humans and AIs working together can free our bandwidth for high level tasks. And if we want to do systematic reviews, we can do more ambitious reviews across different disciplines, across different study designs. And it opens the door to doing continual reviews. If we're just doing research papers, we can hopefully get to a point where we're doing more informed papers and more quickly. Um, briefly now, how about forming research questions and hypotheses? Initially, it was described that LLMs are simply stochastic parrots. They're simply interpolating between data points. They have no creativity at all. Well, we can view creativity as yet another cognitive task. And in fact, their creativity is to a certain degree empirically testable. So there's some evidence this is a recent paper which suggested that it has a GPT4, for example, has a reasonably high degree of creativity and creative writing task, but it doesn't match the distribution of humans. So the best humans are better than this AI model. But the AI model is better than the worst humans at this task. We're also seeing the emergence of benchmarks for assessing creativity as well. So in my experience, you know, even if they're flawed and limited for being creative, they're still useful because they're instantly available, very cheap to use. So AI and humans working together can be helpful in that creative art aspect of our research. So we're now seeing papers and published in the top journals in their disciplines in which AI has has had a creative role. So we have a paper in a top physics journal published a few weeks ago and a paper in cell one of the top biology journals and they described here that um the top AI generated hypothesis actually opened new research directions and it bypassed human biases to to propose overlooked biological possibilities. So if our benchmark is are they creative enough to contribute to new knowledge, these papers suggest they are. We also see the emergence of agentic systems for the full pipeline of scientific discovery. The AI scientist and Robin. So just to highlight again how rapidly things are developing, the papers at the bottom here were published in November of this year and the paper in physics was published this month of this year. So it's extremely fast moving. I think I'm going to skip that in the interest of time, but I think they're useful for unlocking new data. So for coding, conventional practice is that a human does it by themselves and you do everything manually. Okay? When you're doing manual coding, you might think to yourself, well, I've spent years typing exactly the same command. That kind of breaks the software engineering principle of automating repetitive tasks. Often your code doesn't work right. There's an error, some unknown un unknown error. So then you spend maybe hours checking very verbose documentation. Maybe you spend hours checking online forums and wondering why the responses are so sarcastic sometimes. A few hours later, you fix your bug and your code works. And there's a certain satisfaction which I do sometimes miss of fixing coding errors. But then I wonder to myself across a scientific lifetime, how much time is wasted on this process? Is there a more efficient way to um er error bug fix our code? So we're now seeing um AI integrated into into a data analysis software. And one way is is the autocomplete options. for example, with GitHub Copilot. And here the the AI will see the code that you're writing and will simply suggest code that you might want to write next and you can accept or reject that that coding inclusion. And to me, it's not surprising that um the literature which empirically evaluates the use of those tools finds that it does actually um benefit productivity. This is a recent report where these tools were deployed in the UK government. And they estimated for each user, if you give them these tools, it saves them 28 working days per user annually. And 58% expressed they would they would not want to return to their original working conditions. So it seemed to save them time and many seem to like using it. AI can also help us move up the abstraction layer. So nobody types out zeros and ones when they're coding. We use higher level languages. AI enables to go from a natural language like give me the code to do this and then you can get the output. You can get code suggestions in any language. So you can start to think rather than just using one language I can use whatever language is necessary for my task and how this is integrated into software tools. So I think many of us in our in our in our disciplines are still using the traditional tools like maybe R studio or the STA the STA software whereas you know in the software world often using tools like VS VSM stu VS code where you have the LMS integrated into your work you can ask the LMS to review your code to check your code and to make suggestions so that's using it simply as a tool just to make suggestions and to check your code. You can al also increasingly um agentic systems of analysis which are very new and are quite experimental where you're having a conversation almost with your data. You're saying this is my data set analyze it in this way. The tool then suggests code which you can audit. It then executes the code and shows the output. That's one example. And there's another example here as well. So I think one of the reason ways it can be useful to ask LLMs to rapidly produce code is it can help with um reproducibility the human baseline. So how many of us share code? There's a paper that suggested that around 2% of researchers in the health research world share their code. The barrier to writing code is massively and declining. So I think it's likely with AI's help we can get to a point where we're being more reproducible in our work. Um this is another um extension from the um Jupiter um con conference where they're finding yet another way of working with agents and humans together where you're asking the agentic system to do analysis and you're interacting with your AI as well. So for those tasks I think you know in terms of admin in terms of literature reviews and analyzing data the hope is that AI and LLM um and other systems can help us move from the very um manual drudgery tasks to the high level tasks that we actually enjoy. So spending more time thinking about science how we can do good quality papers how we can design our infrastructure and how we can use AI to speed up our reviews and coding. I'm now going to talk about some of um my personal reflections. So, how has it been useful to myself and our group where we work? So, we work at the center for longitudinal studies and we make our data on multiple different cohorts available for free to researchers around the world to use in their in their research. And we used to work you know in the traditional way whereas now we're moving to a world a point where we're making all of our scripts available for our sort of um tasks. So making available data making available derived variables and making available tools. So it's actually been enormously helpful to us in our in our overall goal of helping other researchers. So one example of that is al trying to understand how to use shell scripts when working with genetics. LMS were really helpful there in helping us do this work and that work is now available for anybody to um audit and those data are also freely available um to anybody in the world upon data application request. It's also been very helpful in terms of um um making available tools to handle our data in multiple different languages. We had a hackathon last month and my colleague Liam who's in the audience he worked with the office um for national statistics and there they wanted a tool to help them understand what exact job classification somebody had. It's quite a hard manual task actually. So their lean was able to make an app to do that with two single prompts to an agentic system. I'm going to skip that. It's very helpful for making tools and visualizers. So here I wanted to make a tool to show students what effect missing data might have on terms of bias of your estimates. So there's a nice visualizer which quite easy to make and so you can get an intuitive understanding of that concept which can be quite hard to understand. Um I have an example of an AI paper here which I produced with colleagues recently. So we the motivation for this is I was interested in understanding what is the boundary between advocacy and research. How often in our papers are we actually making bold kind of policy claims even in our research papers? And so we analyzed the last 25 years or so of um the 35 years sorry of of of published research. We downloaded the abstracts. We used LLMs to classify the abstracts because the concordance between LLMs and humans was similar to humans and other humans. And LM's actually enabled us to do this work to to learn how to use Python APIs and Git. And so humans and LMS working together helped us do new research in a new field, helped make data available anyone can now use in the future. Also helped an early career researcher. So Megan Wang is the second author and is now applying for a PhD with her first paper. I'm going to skip that. I think I'm How time one minute. I mean it is after six but yeah. >> Okay. >> Okay. Um I'm going to briefly go over some um some of the um the concerns that people often have before I hand over to Peter. One is which that we can't use our data because our data is private. So we can't use LLMs. We're simply locked out. So one compelling alternative is we can simply use openweight large language models that we can download to our computer and use or use in secure um servers. They're increasingly high performing as well. Another one is about AI slop. There's been a massive increase in papers across science in the last few years before AI. AI might make it worse. So this really is a human problem. How do we incentivize quality over quantity not only in the UK but globally? And finally, for the environmental impact, there's no use LLMs if they have a catastrophic impact on the environment that we can't tolerate. And to that, I would say firstly, do we have accurate estimates? Secondly, what is the comparison in which we're comparing them to? So this paper suggested the median Gemini text prompt used less energy than watching 9 seconds of television and consume the equivalent of five drops of water. And thirdly, is it worth it if to balance the impact of human activity with its potential benefits? And the benefits of um advanced artificial intelligence are considerable for scientific discovery, for health, for medicine, and for education. So in conclusion, quantitative research is mostly cognitive work and AI is becoming highly capable and useful. And so I conclude that AI is the future. And I think it's the case even if even if progress plateaus, even if we're strictly limited to the models that are available right now, I think it's in fact overdetermined given that great given that increasing capability and the profound impact on science, education and society. We are at a profound sh this is a profound change and it's at its infancy. So deepseek the reasoning model for deepseek came out less than a year ago. Chat GPT came out three years ago. But we should adapt to this change and also contribute as well. We have a role to play to understand the impact of AI, understand its capabilities and its limits, to help users use it responsibly and to use it to help us learn. So in summary, I believe that we can use AI to accelerate research, to make it more efficient, more reproducible, more ambitious, and actually more enjoyable. And I think that is the future that we should be building. So thank you for listening and thank you to everybody listed on this slide. [applause] Thank you very much uh David uh for a very thoughtprovoking start. Um, so I'm away I'm aware that this is an away game for me. Um, so for anyone who doesn't know who I am, um, I have a nice description from a colleague a few years ago. They said that I am like a scientist out of the 19th century. That was not meant as a compliment. >> Sorry, it's not coming up. >> Yeah, I can see that. [laughter] >> Um, Yeah, there we go. It was not meant as a compliment, but for today's purposes, I'm going to take that as my uh as my base, right? I'm going to embrace my inner 19th century Yorkman. In other words, my inner lite and try and convince you like they try to convince their uh contemporaries of the dangers of handing over our profession to the machines. So I'm going to start with a bit of interactivity. Okay. So I'm going to pose you a scenario and I would like you to put your hands up if you agree. So to begin with when you each of you in this room comes across a scientist who is extremely productive and who publishes a large numbers of paper every year. Is this impressive? Put your hands up if you find this impressive. Yeah, brilliant. That's about half of the room. Put your hands up if you think it's a red flag. That's even more. Great. We've got a a critically thinking room. Why am I asking this? Because I think it's a red flag. I think that good science requires time. And in particular, it requires time thinking. And so scientists who publish large numbers of papers are not spending the time to think and to craft highquality meaningful research that has an impact on health and society. Now we know why they do this. They do this because we reward reward hyperproductivity. We reward that when we should be rewarding less more thoughtful and more careful research. And what is the what is the the impact of that? Well, the pressure of that hyperproductivity is that we produce as a scientific system an abundance of research waste. An enormous amount. Over 85% of what we produce is estimated to just be waste. What do I mean? I mean derivative research. I mean lowquality research. I mean low value research. I mean flaw flawed research and even fraudulent research. And in order to manage this insatiable quantity of material, we have a whole system of mega journals and even predatory journals that will happily take our junk as long as we pay them every single year. My colleague George Tomova has described it as we are in the McDonald's era of science. And I'm saying this because I believe that this is the biggest problem in health and social science research. I believe that this wastes all of our time. I believe it wastes our labor and I believe it wastes the money that funds our time and labor. I believe that good research then gets completely lost in all of that waste. And I believe ultimately most damagingly it threatens and undermines public trust in science and public support for science. So that is what we HAVE TO FIX. BUT AI is not the solution. AI is going to take this problem to a whole another level. Even the mega mega journals, even the predatory journals are now struggling to cope with the swamp of waste that is being produced by AI um uh algorithms. This is not the future of health and social science. This threatens to be the death of health and social science. That is my argument. I've got six different points I'm going to make which obviously I will whiz through. We will start with this argument that AI tools cannot actually help us with scientific reasoning. So to understand what AI uh tools and indeed LLMs can and can't contribute to health and social science, we have to start by understanding what they can and can't do. What are they? Well, what the architecture tells us, LLMs are language and image calculators. They use NLP and computer vision to interpret prompts and produce statistically plausible patterns that resembles their training data. They do not understand. They have no capacity for reasoning. They do not understand the difference between truth and fiction. They do not even understand what a letter is or a word. It is just a pattern. So an LLM when we ask an LLM a question, it does not think. It predicts a pattern of words or pixels that would typically follow a given prompt. As Bergstrom and Back Coleman have described it in 2025, they are like autocomplete in overdrive. Right? So that is the reason not just for LLMs but for many machine learning algorithms such as driverless cars. That is why we have not achieved a world where they can do this safely by themselves. Because when they see something that they have not seen in their training data, like an upturn lorry, they do not understand what that means. I can tell you a human being in their very first lesson would not plow into an overturned truck. But after tens of millions of hours of training, an LLM which we claim or an algorithm which we claim has understanding would do something like this. So this makes them very well suited to summarizing what we know but not so well suited for scientific discovery because scientific discovery is something else. Scientific discovery is about moving beyond what we've observed and understood before to a new level of understanding to new models of the world to new meaning making about everything that's happening around us as John Pearson has described the problem is the future is not in the training data and in fact the problem is bigger than that because what is the training data made of waste. Do we really want to look back at the shoddy research that we've done in the past and reformulate that in some way and claim that we're doing science? We know that LLMs have an inherent fault. They hallucinate. We ask, "Are children small or just far away?" And this intelligent machine tells us, "Children are not actually small. They're just far away. Well, we know that's wrong because we have a mental model of the world. It does not. This is complete nonsense. It's obviously nonsense. But we don't use these algorithms when we already know the answer. We use them, we ask questions of them that we don't already know. And when we're faced with those questions, there is a very dangerous risk of being misled. And this is a very simple real world example. This is Paul Zivich. Um he asked an LLM to summarize one of his papers. He says the main summary was accurate, but he also asked it to extract and use code from his paper, which it did. It just happened to introduce an error, a subtle error. A subtle error that meant that the standard errors would be incorrect. In other words, the confidence intervals would be incorrect, the p values would be incorrect, and potentially the inferences would be incorrect. Would you notice that if you didn't already know the code inside out to begin with? I would argue not. And for that reason, in practice, LLMs often require extensive handholding and validation checking. Right? So this is a a a fascinating study of using CH GPT as a tool for biostisticians. It came out earlier this year. And what they conclude is that while some tasks were completed rather satisfactory, others suffered from severe issues. So as a consequence they give this nice list of things that you have to do. Double check this, double check that, scrutinize this, scrutinize that. The list is so long and tedious that I am left wondering why would a human not simply do this themsself rather than rely on a machine and then have to check every damn thing that it's done. So the reality what is the impact of this in the real world is that these machines do not match the marketing hype and we don't have to look very far to find that out right there was an MRT study only a few months ago I think which estimated that 95% of efforts to integrate Gen AI into business were failing and many of the businesses were aband abandoning their efforts because they were not having any improvement on productivity and it was not having any improvement on the bottom line. And so the question that I would ask every scientist in this room, what makes you think it's any different for us? Right? There are people out there who are spending good money right now on trying to find out what LLM can do for them and abandoning it because it's not working. Why do we think it will miraculously work for us? Right, I'm going to move on to the word of the year according to Webster's dictionary. Slop was announced today. Uh this is um David's paper which he's already alluded to and colleagues. Um at the end of the paper um you argue which you you argued in the the presentation as well. We live in an era that incentivizes scientists to produce masses of papers of questionable quality. Right? We agree that is a major problem. This is a human not an AI problem. I completely disagree. I think that's a logical fallacy. I think it's a false dichotomy. It's equivalent to saying guns don't kill people. People kill people. Well, in reality, if people kill more people when they have access to guns, guns are a cause of death. If people produce more slop when they have access to AI, AI causes slop. And we don't have to look very far to see what a problem that is. Right now, LLMs are cat catalyzing scientific slop on an unimaginable scale. Right. There was an article in the Guardian only last week about AI research. AI researcher who was proudly boasting he published a 100 papers in a year and he'd used AI to help him do it. And if you speak to any journal editor right now, I can tell you what the problem is. They are struggling to cope with fraudulent submissions, spam submissions. Some of them they they pick up. Others such as the famous massive testicled mouse somehow got through the predatory journal system. This is bad, right? This is worse than we could have ever imagined. The state is so bad that journals even from some very questionable publishers are banning submissions with open data sets because of the sheer number of lowquality clearly AI written research that's being submitted to them. And similarly, even a preprint journal, I didn't think they even had standards, right? Even a preprint journal has paused submissions about AI topics because of the sheer influx of clearly AI generated slop. This is an existential problem for our system and it's one that's recognized by publishers and by funders. This is an article by uh this was a report by Cambridge University Press earlier this year warning that important work risk being lost or drowned out by the surge of lowquality and AI generated content. And this is from cancer research UK. We are being swamped by AI generated content that is endangering the very foundations of objectivity and empirically derived facts. So we're in trouble. But it gets worse, right? I think there is an there is a genuine risk to the scientific craft. Okay. If we look at the extreme proponents of AI, which um this could be a an example. This is um an agentic AI model that promises to generate novel research ideas, write code, execute experiments, visualize results, describe its findings uh even by writing full scientific papers. It's the one-click scientific paper um model. And what is argued is that by giving a uh algorithms these uh tasks then we are liberated as human scientists to work on more creative matters. Well, I would ask what task are we liberating ourselves to do if we hand over every single stage in the scientific method. There is actually I can't think of a laborsaving machine that has been produced that has ever freed us as human beings to to do more meaningful work. This is what's always promised, right? We were told we would all be working 15 hours a week. How's that getting on? These tasks are the things that we as scientists are trained for. They are the task that we are skilled at. They are the tasks that hopefully most of you find rewarding and get great value from doing. Would you not rather struggle through those than be stuck in some dystopian future where instead of doing research ourself, we are simply inputting information into an algorithm and then interpreting the output and checking for all the mistakes that it's made. Even if for some reason you're happier with that future, which I am definitely not, I would argue that you'd probably be worse at it than if you left the LLM alone because there is evidence that we as human beings perform worse when we collaborate with AI. It's well known across many disciplines automation bias occurs. So we become overly dependent on an algorithm once it's there and we lose our ability to make critical judgments and in the scientific and academic field that's cognitive offloading bias. So when we use LLMs it appears to damage our critical thinking skills, our most important skill. And there's a bigger implication. The more labor that we give to LLMs, the fewer opportunities we have as individuals to learn and maintain our craft. So there is a genuine risk of a collective deskkilling of all of us and of the scientific workforce. This creeping reliance on LLMs risk displacing human scientists, particularly junior scientists, meaning there'll be fewer jobs. There'll be a a breakdown of the training pipeline and potentially a complete collapse of the scientific workforce. So there's a dark future here where there's just a few senior academics floating around. They don't have PhD students. They don't have post postocs. They just cost giant LLM subscriptions into their grants and then spend their life submitting and interpreting all of the information. >> I move on to more jolly matters. >> Oh, great. It's often assumed, it's been implied that LLMs will get better and better. Is that true? I don't think it's necessarily true. We know that existing LLMs have now been trained on the majority of publicly available data. So the the scope for scaling is much more limited than it was at the beginning of their life. And indeed, the improvement between chat GPT4 and chat GPT5 was so modest that it left a lot of people wondering, have we hit the wall? But it's actually worse than that because the training data is degrading. Over 50% of the internet that now is created is thought to be created by AI. So tomorrow's algorithms will be learning based on flop. And guess what happens? Model collapse. So the future does not look so bright there. And it does not look so bright when we consider that we as humans are contributing less and less as we use LLMs more and more to to wonderful rich human platforms like Wikipedia and Stack Overflow. If you think that LLMs are really good at coding, where do you think that knowledge comes from? Humans [clears throat] who are not providing that knowledge anymore. So a realistic future, a genuine risk is that future models will not be better, that we may have actually reached the peak and we're heading towards something rather disappointing. But even if we're not, even if the technology will continue to improve, what happens next if we contract our skills and our reasoning to LLMs? Well, most people when they're using LLMs are using commercial LLMs. Let's have a think about what happens whenever technology promises something faster and cheaper. How's that working with food? Well, we have this abundance of cheap, lowquality food that's probably killing us and is certainly having an enormous environmental impact. What about fashion? Again, we have an abundance of cheap clothes of incredibly low quality with a high human and environmental cost. Okay. What about tech? Well, what's happened with tech? When Google first came out, it was good. Then came the ads and the sponsored results and a worsening user experience. And the only way to escape this, and it's the case with every single technology really out there, pay more. What do you think is going to happen to commercial LLMs? The same. Why? Because right now when you use chat GPT even if you have a subscription you are not paying for it. Someone else is. Enterprise investment. Hundreds of billions of dollars of enterprise investment is being pumped into these things to make us use them. Well, they're going to want all of that money back and more because the only reason they're making that investment is the knowledge and the idea that in the future we will all be trapped and we will have to pay and pay and pay. So I guarantee I guarantee chat GPT every single commercial LLM will become in shitified over the years to come. We will have ads. They're already experimenting with ads. The thing that scares me beyond belief is sponsored results. Who is going to pay to manipulate the knowledge that I receive when I use one of these algorithms? And what damage will that have to society? The experience will get worse. And the only way to avoid it is to be the lucky wealthy people who can afford the everinccreasing subscription costs that will come our way. So I've already said there is a serious risk here of us funneling more and more money away from human scientists to industry and we have already made painful pacts with the publication industry that we cannot get free of. But my GOD IF YOU THINK IT'S HARD to break away from journals imagine what it'll be like when we become completely dependent on these things to do all of our research. I think that is a very serious ethical concern, but it's not the only serious ethical concern. As health and social scientists, we have to think what are the implications of everything we do for population health and so society. Well, LLMs are built on mass intellectual theft. We produce the work, we share it with the world, they scoop it all up and sell it back to us. What does that sound like? This undermines our labor and our livelihoods. But there's worse. When you think an LLM is magic because it's got some mystery data and algorithm, I will tell you there is an army of invisible, exploited workers behind the scenes who are who are doing all of the moderation. They are looking through the output and they are having to deal with extremely traumatizing material. But we forget about them because they're in the the global south. And then of course there is the energy cost. LLMs are extremely power and and and water hungry. And by 2030 it's estimated that their data centers will draw 5% of the entire global energy need. How can we afford that in a world where we desperately need to be using less energy? I want you all to consider this as a moral question. What human, social, and environmental cost are you personally willing to pay to speed up your next publication? So, my my positive frame, there is another way. Right? So, I'm going to start with this quote from Rosa uh Ronano. I want to shine the spotlight on the commonplace assumption that productivity must always increase. Good research is disruptive and thinking time is central to high quality scholarship and necessary for disruptive research. In other words, what do we need? We need in the fine words of the late great Doug Alman less research, better research, and research done for the right reasons. And I cannot think of a moment in the last 30 years where these words are more important for us as a community. trust and supported science is at an unprecedented low. This flood of derivative and fraudulent research from AI threatens the entire system and simply chasing greater productivity will not save us. Instead, we must try and take the slower path. We must spend more time thinking and collaborating with human colleagues to produce better and more thoughtful research. So just to summarize in answer to the question is AI the future of health and social science research I would argue emphatically no it risks being the death of health and social science research. Thank you. [applause] So we've got now five minutes just of a little back. >> Yeah. Okay. Thank you, Peter. Um I think in general we are scientists and so we should try and back up our claims with evidence. [snorts] So you made a number of claims which I think could be disputable. Um, and I think progress is where we can make progress is how we can actually generate that evidence to inform these decisions. So the the self-driving car analogy was interesting. What are we aiming to achieve there? What's the comparison group? You show a photo of a car that's broken down on the on the on on the road. What we actually should be comparing is mortality rates with self-driving versus human driven cars. Those are the relevant comparisons. You mentioned people are abandoning this technology. Yet purportedly it's the fastest growing consumer product of all time. chat GPT within with you know you know incredible use of launch enterprises you mentioned it's will not get better with a reference from August 2025 um I showed a reference um from a few days ago where clearly there is improvements there isn't any great deal the improvements seem to be continuing um particularly as we scale up reasoning so you said that the wrong tool for scientific discovery scientific reasoning that they can't reason at all you know I used to have that view particularly of early LLMs, but as we've begun to train reasoning into them with reinforcement learnings, they have what appears to be reasoning. So if you give them your exam script which explicitly requires reasoning, they give you their reasoning steps which would get great marks and they give you the correct answer which would also get great marks. They also do well in tasks which clearly require some degree of reasoning however we define it. For example, gold medals in competitive domains. And I think you know they're unsuited to scientific discovery. Well, scientific work is multiple different cognitive tasks and at our discretion is the task that we do or do not give to to AI. Um you said that the problem with health and social science is waste and I think we agree that there's a problem with um people being wrongly incentivized to produce lowquality papers. So the solution to that is that we try and incentivize um higher quality few fewer papers. We don't stop using tools simply because of that problem. And if you look for example this has been a problem whenever you make a tool to increase productivity. So the two sample mandelian randomization packages there's been a surge of men randomization um papers particularly from China. In the UK we already tackling this and it's actually difficult human problem. How do we get the incentives right? So in ref for example we have a maximum of five papers and we're moving to narrative narrative see these in grants we have to acknowledge that science is global it's not just the UK so LLMs cannot liberate us from science they take away the scientific craft they are a general purpose technology so we can use them at our discretion we can use them to free us up to do more meaningful creative work or we can use agents to do all the scientific discovery I do not have the evidence that suggests that those agents should be automating all scientific discovery. That's not a very defensible position. These things have come out in recent months and they're an interesting technology development. But I think there's clear wins where they can add to scientific um discovery. So you mentioned the future of LMS is is in shitification and I think that I mean I think that that's speculation there. We still use Google, Google Scholar even with the ads, right? If a the beauty I guess of of capitalism, if if a company over advertises, it's a very if it's very competitive environments, you can simply switch your LLM provider. Even more so, if you don't want to use a commercial LLM, you can use an open way LLM or even a fully open- source LLM. And perhaps as the research community, we should be building those out if we're concerned about the downsides of commercial LLMs. You mentioned ethical concerns and you cited what I think was a Guardian article. I think every major technology has costs. We are empiricists. We should be quantifying those costs and having a careful having a a proper evidence discussion about how much they're costing to train, how much they're costing to use. And you know there's a report by um UCL colleagues that came out a few weeks ago that suggested that we can reduce energy costs by 90% simply by using smaller or quantized large language models. So LLMs are coming people are using them we have a choice in how we use them in the future. And if energy consideration is a big point we can try and um make them more energy effic efficient and advocate for the more energy efficient uses. And also what do we compare them to? you know, if we compared it to the entertainment industry, if we compare it to cryptocurrency mining, the LM is a technology which has got clear proven benefits um in my opinion and and the efficiency costs have also improving across time. So, for example, the Deepseek model which came out earlier this year seemingly a much more efficient model to train. If models are getting more efficient to train, they should for each individual query be more energy efficient. So you also mentioned the case for slower and more thoughtful science. I think we're all in agreement. We want higher quality science, not lots and lots of redundant lowquality papers. So overall I think that um but but I think you know slow itself is not virtue. If you can do your paper in six years or six months and it's the same quality, what should you aspire to be doing? What should the public who are funding your research want you to do? Spend six years on it because you enjoy that process or six months on it with tools. You could even spend long on it if you use no tools at all. What about the what about the energy and ethical costs of industrialization as a whole or using a computer or using a smartphone? We at the early stage of this technology. So, we should be measuring those costs and then making um balanced judgments. I think I'd also add as well just in terms of the positive potential of this technology. Um I don't think in any major technological change like this it's been so evenly accessible from around the world. So anyone around the world can essentially use these models essentially for free and it's available in essentially in every recorded human language. Whereas other technological revolutions for example smartphones they were largely restricted to high- income countries at least for the first period. So there's an argument here that they can have a huge positive potential benefit to countries with lower resources. Education is hard to to intervene on, but what we do know is personal tutoring can make a massive difference to education attainment. So we can roll that out to everybody in the world. So everyone can have personal tutoring or we can simply stop using these tools. I think in general I think you know that AI is coming whether we like it or whether we don't like it. The key thing is what we do, what we do with that information, what we do as scientists, what we do to shape our ecosystem to make it work well for our our goals. >> I think I'll finish that. >> Thank you very much, David. Well done. [applause] >> Five minutes, Peter. >> Okay. Thanks, David. Um I I'm desperate to respond to the points that you've made then, but um really I I I need to try and respond to the the main uh points that you made in in your presentation, which I haven't covered. It's probably a case that we'll have to agree to disagree uh about the inherent ability of AI um to actually have cognition and to be creative. [snorts] Um clearly I would argue not. Um, I think they're excellent at cosplaying reasoning. I think there's an awful lot of pattern recognition that can take place from the incredible amount of learning that it's done, that means it can quite reasonably guess things that looks very like reasoning but isn't. Um, and I think when you see them misbehave in incredibly ridiculous ways, which happens if you ask them often simple mathematical tests or strangely simple questions and they just completely fall down, that is a good reminder under there is not reasoning, right? It doesn't matter what what uh figures someone has produced. There's something very strange there that indicates it's not really understanding what you're asking. Um, you talked about using uh LLMs for research question and hypothesis generation. I personally have an enormous problem with this. Um, I I think that the the the process of of of hypothesizing and of constructing research questions is one of the most sacred and challenging aspects of being a scientist. Um [snorts] I think that we are embedding our mental models and we are embedding our values when we are formulating hypothesis and constructing uh research questions and my concern with relying on LLMs to do that is that of course they have baked in all of the biases and the waste that they've seen before right so we are trying to move beyond a world where we're simply repeating colonial and patriarchal messages the only way to do that is to be coming up with questions based on the society that we live in now and that means using your human brain. Um, in terms of systematic reviews, I the idea of them being able to perform a systematic review in one click, even in five years time, I unless they solve hallucination, which they can't because hallucination is a fundamental feature of the way that they work, I can't see that ever uh happening. But I I will concede that I think it's reasonable to use them in systematic reviews, particularly as a secondary screener. But again, I would say the biggest problem with systematic reviews, as I see them, is a complete lack of critical thinking. We will just combine all of the results that we've seen published before without proper reflection of the challenges of these papers. Yes, of course, we have, you know, corporate guidelines and so on. But still, people are not conducting the kind of reviews that we need, the kind of reviews that take more than six months. I don't think I've done papers that take six years. They there's absolutely no human way to do them any sort than that or no inhuman way to do them any shorter than that because it requires six years of back and forth thinking and conversation with other humans over many many years to understand an unsolved puzzle or to clarify an idea beyond the knowledge that we already have. But clearly there are some examples where some people may drag things out longer than they should. When it comes to coding, I think this is a really interesting one. My brother is a a a software engineer. He has some very strong opinions on this as well. Um I think they work very very well for basic tasks. Um I think they probably do work pretty well for bug detection, but in my experience, if you're trying to do anything remotely creative, it is a nightmare. There is a genu genuine risk that by choosing to work with an LLM, you choose to be the hair in the hair and the tortoise, right? You think you race off, you've got all these initial ideas, you think you're halfway on your way, and then you get grounded because you don't have the baseline understanding that you could have had if you'd have started reading and planning and coding with your own knowledge. And dare I say, some people don't find college coding to be a repetitive task. Some people find it to be a very rewarding um problem-solving task that maybe we should um appreciate um a bit more. So I think I think that's enough. Thank [laughter] you. >> Thank you very very much. EXCELLENT. [applause] >> Thank you very much. >> Very well done. I really enjoy. Um I think I would like to uh invite both speakers to come forward. Um and basically we have now uh 20 minutes or so for question and answers and debate and so we obviously have plenty of time over nibbles and and wine and so on to discuss that further. Um we do have maybe the voting system >> at the end. Yeah. Once we've done a Q&A we'll do a final vote in the last minute or two and see who has fared worse. >> Do we need u microphones for people? We don't really have individual microphones [clears throat] so maybe just have to speak up loudly. >> Yeah. Do you have any questions? Immediately we have a question. >> Yes, I really enjoyed your um your your point and very passionate against it. I wonder so talk about LLMs and hilizations things. I was wondering what you think about other alternatives in machine learning such as alphafold which obviously just won the Nobel Prize for protein folding. Do you see different areas of machine learning having desperate impacts? >> Yes. Um in in short um clearly uh there are certain prediction tasks um where machine learning is extremely well suited. Um but uh the one thing I would say is beware of the hype right because they can do certain things extremely well. People assume that they can do many tasks extremely well and really it's a case of individual tasks being very very well suited and other tasks potentially not being so well suited. But clearly for some things fantastic. >> Yeah, I agree. I mean, I think to keep it reasonably fair, we steered towards LLMs and agents in this debate, but there's clear great utility and uh in what Alpha Fold and other another frontier prediction models have done. For example, meteorology and they've been published in great journals. They've open source their resources and had a good contribution to science from those commercial companies. >> Great question. >> Yes, thank you for that. I think Peter alluded to the fact that um AI sorry hallucinations are a fundamental feature of AI. I was going to ask you David, do you feel that it's a fundamental feature and do you feel that we can actually overcome that at any point in time or do you think it's just something we have to kind of accept as being the >> Yeah, I think they're pro there probably is a degree that we have to accept. They're like they're not deterministic, they're probabilistic in the output they choose. Um so but there's ways of engineering these LM so the hallucination rate produces and empirically we can track that with hallucination benchmarks. It's also I mean in my personal use I tend you're going to hate this Peter but using multiple LLMs at the same time and kind of triangulating off them. Maybe some of them have some of them have certain biases and you can get the output of one and give it to another one feed it to another one and that way so people have been exploring that in a more formal way of like scaffolding work around it to try and reduce hallucination rates further and the reinforcement learning is one paradigm for that. >> So maybe there are questions around methods and how to >> you should answer this one. Uh the question was about hallucinations. I I would just agree that that they're a fundamental feature. I mean I think you're right that the there are things you can do to to try and reduce the risk um but not remove the risk. >> So yeah, that was your question. Do you think at some point we can live without uh yeah going towards not having these? >> I don't because I personally because there's no sense of truth and and fiction, right? So they they're just producing a pattern and sometimes that pattern will be wrong because it has no idea what the concept of right and wrong is. >> Can I just say so do humans as well. So [laughter] >> there was a question from over there and then I think >> um you you were next. You you raised your hand if you >> might have to stand up and shout. Do do you think that uh AI as it's currently conceived have a positive negative or neutral effect on the development of theory in social science? >> I'll stop. Um I think right now I I see nothing that makes me optimistic unfortunately. Um what what really scares me is just how willing a surprising minority are to just embrace scientific fraud. It's like they were there waiting for the right tool and now it's here. Um now it may trigger a renaissance, right? Like I genuinely believe there is a possibility that as we get in influxed with slop, we as a system say right, we need to do something different, right? and that something quite wonderful could happen. But right now, I'm not optimistic. >> I think it depends if we get the governance and the and the incentives right. Um that's the key thing. Um but we should also be you know tracking quality of papers and empirically trying to answer that kind of question. And different disciplines have tried to sort of um ensure quality in different ways. Um so you know in epidemiology you can publish a descriptive paper just a simple descriptive statistics in the top economics journals they really want cause identification and um like a very careful empirical um strategy around causality but even there there's there's concerns about the effect that's had on the discipline. So we have to get the governance and the incentives right learning from all the different disciplines. Each of them has things they've got right and things they maybe haven't. >> Yes. So there was Liam and then yourself. Yeah. >> So part of your argument was that um AI is bad for the craft of science and I just wondered whether you believe that not using AI at all will make you advance in your field and actually people that follow that strategy might win out and then AI use itself would lead to less use over time just because they were succeeding. There this is an interesting thought experiment as well um which I think has some merit right as things pan out there is no question in my mind that there'll be a massive concern about AI slop and that a little bit like you know the tailor who continued to produce magnificent clothes on um what's the name of the street savro right these people survive because they continued to produce produce good quality stuff. They didn't chase growth, right? They didn't feel we need multiple several rail row shops. They just said, "Look, I'm going to do what I do really well." And they survive. Um, and I think that there there could be a route for some people to do that to actually really embrace slow science, uh, you know, masterful science that then the community appreciates. My only concern with that, which I should flack, is who has the luxury and the privilege to do that. I am aware that I stand up here talking about slow science with a profile that says it's easier for me to follow that path. Um whether it will be so easy for everyone and what are the ethical implications of that I I I'm a little bit scared by. So in general I think the best approach is to you know a combination of human and AI. So use these tools to help you learn to help you do better science. Um that's what I personally think is favorable. If people can do great science without AI and without LLMs fantastic. People should have the people have the choice to use whatever tool they want. >> Yeah. Lots of questions. So I you first you and then over there and I would like to encourage all all women to ask questions as well. I see a lot of hands up by by men which another gender bias is so should bear that in mind. Yeah. >> So you first >> uh you both raised in your discussion uh question to what extent AI's contributions are derivative but I was going to ask to what extent that actually matters. I mean Newton had this line about I saw further it was only because I stood on the shoulders of giants. Kak McCarthy the writer said that it's a sad fact that all books are made from other books. I mean to what extent is our AI just mimicking human creativity by ingesting what's come before and then replicating something uh something new from it or combining new information out of old well understood facts. Um, I I I I think one of the problems we have in in science right now is an information overload. Um, I'm aware of articles that have been published over a hundred years ago warning about methodological issues that the vast majority of analysts don't even know about. Um, so we do have a problem with volume and my concern is that encouraging that volume and seeing it as part of the system is not going to solve that problem. Right? Actually, we we desperately need to to radically reduce the amount that we're producing so the signal to noise ratio um improves. >> Can I say to that one? Um, so I don't fundamentally know how the LMS think. I also don't know how human brains think either. I think we're both neural networks of different design but they seem to to my view they seem to approximate creativity and approximate reasoning maybe in a different way but they approximate it and the key thing is can we actually evaluate that is actually is that is that claim actually falsifiable and so if they're actually contributing to new knowledge um that to me they're making a contribution I've talked about reasoning a bunch in in the talk in the rebuttal. >> Thank you. So there was a question here and then yourself. Yeah. >> So I appreciate it. uh David's point about benchmark saturations and uh there's a point to be made that benchmarks can be saturated on areas and benchmarks can be done on areas where there are right answers that are no. My question to both of you is are there areas of health and behavioral sciences that lend themselves better or worse to these benchmarks uh that lend themselves more to the alpha treatment um rather than the sloth treatment. >> Yeah. [snorts] benchmarks. Um, there's a lot there's hundreds of different benchmarks, but they don't seem to be being produced by people who work in our field. And so I wonder if those benchmarks aren't suiting us if we should not be creating them ourselves. It's kind of like cognitive testing is you tend to measure in a very narrow domain and you wouldn't recruit a researcher based on a single test score in the same way you don't know how good an LLM is just based on its benchmarks. So I think we're we should be contributing to that that that that work. I think >> I think in the business world this can be done really well. We've introduced this tool and oh there's been no change in our productivity or there's been no change um in in you know our bottom line. I think ideally it would be great to be able to do something very similar in science to actually perhaps even conduct an experiment. You have two teams, you have access to LLMs, you don't. Let's see what actually happens um at the end of that. That would be fun. Um, I don't want to know the result. Um, because I would then be arguing, well, that wasn't really a hard enough task. Um, but yeah, I think that's something we should be doing. We should be collectively finding ways to evaluate these things in terms of the end goals that we're actually interested in. >> Yes. So, yourself first and then over here. Yeah. >> Hi. >> Sorry. Did you want to say something in return? I didn't get you. Okay. >> Did I get chance? I forgot to answer that one. Don't forget I did. >> I think you did first. >> Um, >> if I delegate a task to a colleague, I end up with a more experienced colleague where it's very delegate to AI. What maybe I need to be training colleagues so they're better at delegate to AI. But when it comes to training, what what do you think we need to do for using AI? >> Um, there's certainly things you can do to help people get the most out of AI. for example, best practices in prompt engineering, how you can get better performance out of them like adding in examples in the prompts and things like that. So there's a way we which you can teach people to use AI well and teach them on for example on science the kind of areas you you kind of encourage them to use it versus not encourage it. Um the training point is really important because I mean we need to protect the career structure of scientific discovery and it's already not satisfactory. [clears throat] I mean people have short-term contracts. we're struggling to retain talent. So, we need to build a a robust and defensible pipeline of careers in science um now more than ever. I don't work on that, but I hope very careful social scientists are do you want to >> so I I would agree. I I don't know the answer to be honest. >> Um >> yeah, >> one last question. There was >> maybe two more because there's the lady here and over there. Yeah. So, maybe these two and then finishing the Yeah. Thank you very much. >> Hi, I'm Vanessa. Um, I would like to ask about the critical thinking part. Um, Pete mentioned about the negative effects of music loss on critical thinking. So, do you agree on that? And if yes, how do is there a way to counter that? So, we could use myself. Um, so I think the jury is still out about overall do they diminish our our our creative thinking or can they actually augment it? uh it's just so general purpose. It's like speaking to a human expert in any domain. So you can you can work with it in such a way that it can actually improve your creative thinking I would imagine. And the early empirical evidence on the effect of tutoring with LMS is to my mind is extremely positive in terms of um helping people learn things including um creative thinking. We it's it's it's a tool. It's here. We have to figure out how to mitigate any bad bad bad effects and how to promote the good effects. >> That might be one of them. >> Thank you. and yourself. >> Um, hi, I'm Georgia. Thank you. Thank you both so much. Um, one comparison or argument I've heard about using um, LM is for example, if you were to go to a marathon and take a taxi to the end of the marathon, you wouldn't expect a medal at the end. And I wonder whether you agree whether this translates in a similar way to science. If you haven't struggled through the craft of science, would you get the same reward in the end if you did it yourself or if you did it with an engine? >> Was that more for me? >> It's a good question. >> Well, you you can start. >> Um, no, I agree. Like you need you need like an apprenticeship of doing things manually to learn. The question is like to what extent and to what degree do you want junior colleagues spending you know an entire calendar year on screening or can that time be reduced so that they spend enough time on it to learn but not too much time that they can't do other tasks like develop their creative thinking and how we bring multiple lines of evidence together to do better quality papers. I think your your question made me think about something that we we don't even discuss or think about which is do we enjoy the job right why do we run a marathon we don't run a marathon to travel whatever it is however many miles 21 miles I need to just travel from A to B so I'm going to run a marathon we have some other intrinsic reason to do that and I think for many of us who were driven into science as well there was an intrinsic desire to understand things to question to query to explore experiment and that is actually a concern that maybe we're not as productive but if it's more enjoyable then isn't that more important at the end of the day what's the point of all of this if not also to enjoy our lives in the process so I would say yeah you know you're not a carpenter if you go and buy a chair from Argos right there is something about the craft that we have to actually respect that's above and beyond output and productivity. >> Yeah, >> I think we've run out of time so I think we need our followup. >> I think we can [applause] diplomatic. Yeah, I think um AI is here to stay and um obviously there are many many pitfalls and many many advantages and and things to be gained and I think it's our responsibility to think about how can we use it responsibly and that is our job and thinking about you know why are we doing it and and what's the purpose rather than just doing things for for the sake of it. So I think it was a fantastic debate. Thank you very much for the preparation um and and the back and forth and so on. And thank you very much for fantastic questions, really [laughter] thoughtful arguments and I'm very sure we can have um yeah more thoughts and more questions and so on during the during the um reception. Thank you very very much for coming. >> Okay. Have you all voted? [applause] >> All right. Can we see the Can we No. Hang on. You want to see the results? >> Can we see the code again? Ah, can we see the code again? >> Right. >> Yeah. Okay. >> This was the before. So, the before was 49% yes. 36% no. And the after, 48% yes, 46% no. I think we've both won. >> Or I'm going to argue that this is within something variation. Cool. [laughter] >> Within varants. Yes. So people may have changed from one name to the other. >> Yes. They're not even necessarily the same people. >> We've managed to drag some people from don't know which I think is a a success. [laughter] >> Right. We've definitely earned drinks and nobles. Well done. [applause]