Submind YouTube summaries
Thumbnail for EGU WEBINARS: Introducing Project Cosmos, The World's Largest Climate Research Database

EGU WEBINARS: Introducing Project Cosmos, The World's Largest Climate Research Database

Watch on YouTube

Video summary

Project Cosmos represents a groundbreaking initiative developed by Carbon Brief over eighteen months under the leadership of Simon Clark and Leo Hickman's vision, aiming to visualize and analyze the entire global network of human knowledge on climate change. By establishing IPCC reports as its foundational "gold standard," the team constructed an extensive database comprising approximately 1.8 million unique academic publications linked through more than 40 million citations. This massive dataset was built by expanding beyond primary sources to include second-order references within those reports, studies that cite them, articles from twenty-two key climate-focused journals, and every paper referenced by those journal entries, creating a comprehensive map of the field's intellectual landscape currently accessible at CarbonBrief.org/cosmos/. The interactive platform reveals critical insights into the structure of scientific influence through its "Cosmos 500" ranking, which utilizes a custom citation score to highlight significant disparities in representation. The analysis shows that US-based authors and institutions dominate nearly half of the top positions, while experts from the Global South account for only four percent and women appear merely thirty-six times among the most cited individuals. To address these systemic barriers regarding research conduct and publication visibility, the team has supported the creation of a self-submitted "Global South Climate Database" listing over ten thousand experts seeking collaboration, emphasizing that fostering diversity is essential for scientific progress rather than just a charitable endeavor. While users can currently explore topic clusters ranging from physical sciences to social psychology via color-coded maps or trace connections between papers within the graph-based network, certain advanced features remain limited due to technical constraints and resource limitations. The primary database lacks full-text search capabilities at this stage, restricting text queries mostly to abstracts found in an underlying unstructured document store that allows users to locate specific technologies through openAlex's topic hierarchy; furthermore, downloading large metadata payloads is currently restricted but remains a future goal for deep data exploration. Future iterations of the project plan to introduce filters for recent publications or rapid citation growth and may explore ranking datasets or identifying rising topics similar to social media trends, with updates anticipated annually following major events like the IPCC Seventh Assessment Report. Although direct contributions to the database are not currently accepted due to its private hosting status aimed at preventing AI scraping, the team actively seeks collaborators for future analysis and invites attendees to submit feature requests via cosmos.org. The project acknowledges key contributors such as Felix Chavelli, Travis Cohen, and Kristen Cam while maintaining a focus on expanding the network's utility without compromising data integrity or technical feasibility. As the database evolves, it aims to provide deeper insights into climate science by balancing current analytical capabilities with long-term goals for inclusivity, comprehensive search functionality, and broader access to metadata as resources allow.
Read the full video transcript
So I think it's time to make a start. Um so firstly thank you for joining our webinar today. Um uh this webinar is on introducing project cosmos the world's largest climate research database. My name is uh Simon Clark. I am the project manager here at the European Geosciences Union. uh and today's webinar should provide some insight into development of uh this new database as well as um some uh insight perhaps how we can use it as well. Um the webinar will begin with a presentation followed by a Q&A all within this 1 hour. Uh if you uh want to have a question or ask a question for our speakers today, there's a Q&A box at the bottom of the screen where you can enter your questions which be asked after presentation. You can also upvote questions you think uh like you want to be answered. Um and so we have two members joining us from carbon brief today. We have a tandem who is carbon briefs science correspondent uh who has won um awardw winning uh leading the 2024 world meteorological society emerging communicator award as well as short list from multiple others. Um Aisha will be doing a presentation today and joining for the Q&A later is Thomas Pearson. Um Thomas Pearson has uh over 20 years experience working in presenting complex data in clear and accessible manner as well as producing award winning graphics from everything from BBC news to financial times and to natural history museum. So as a brief introduction to what we're going to talk about today and who is presenting uh Aisha I'll move on to you if you could share your screen uh and start your presentation please. >> Absolutely. Thank you so much for the introduction Simon. Um, all right. Yeah. So, thank you all so much for being here to learn a bit more about Project Cosmos. So, just to give you a bit of background, uh, myself and Tom Pearson work at Carbon Brief, which is a climate change journalism and, uh, analysis, uh, company website. Um, and we do a lot of kind of, uh, news reporting, but we also sometimes work on these really big in-depth long projects. And I think Project Cosmos might be an example of one of the most in-depth and longest projects that we've worked on so far. Maybe Tom can correct me if that's wrong, but we started this project a year and a half ago. And I would like to say this idea was first seeded in the in the mind of Leo Hickman, who's the the head of Carbon Brief, I want to say a decade ago, at least since I joined the company six years ago. This has always been an idea that he's been floating this concept of bringing together all of climate science literature, all of the academic papers and books and reports about climate change and somehow visualizing them, somehow showing how all of these different papers link together, maybe through topic uh relationships or through links from authors. Um, I think he's always had this idea of this big visualization, something like what you can see on the screen here, to show the sheer magnitude of climate science and how all of the different papers link together. So, we actually started building this project in depth um yeah about a year and a half ago. And the final thing that you can see here um is is the outcome. So, this database is to the best of our knowledge the largest database of climate science literature um ever put together. It includes 1.8 million individual publications which are linked together by 40 million citation um and I'll go into a bit more uh in the course of this presentation. But I'm going to kick this off actually with a video that was produced by Carbon Brief's visuals team. Um I was trying to work on a little um a little summary myself and I realized that actually the video they've made is just so much better. So I'm just going to play that for you guys first. This is Project Cosmos, Carbon Brief's universe of 1.8 million climate science papers. This major new resource allows experts to analyze the entire network of human knowledge about climate change. So, what does this database show and how did Carbon Reef create it? Well, it all starts with the Intergovernmental Panel on Climate Change, the IPCC. Ever since 1990, the IPCC has been publishing the world's most authoritative summaries on the latest climate science. Since these reports were first published, humanity's knowledge about climate change has ballooned. The IPCC has published six sets of reports so far, each one longer than the last. Each report contains a list of references, the academic papers, books, and reports that the authors draw on. This list is crucial for underpinning the legitimacy of the final document. In total, the IPCC references more than 100,000 academic papers, books, and reports. This forms the core of our climate science universe. Harbon brief has built on this core by looking at four other sources of data. the references contained within those references, publications that themselves reference the IPCC, studies from key climate focused academic journals, and the publications that those studies themselves reference. Every single publication in the Cosmos database is linked to at least one other through references. Draw together those links and this is what you see. Each column and cluster reveals different topics and densities of research. Carbon Brief has now used this vast database to rank the most highly cited publications, authors, and institutions. Check out Carbon Brief's Cosmos 500 ranking to see the results. Okay, so that was that was the very snazzy video produced by the visuals team that I think gives you a bit of a background about the project. But what I'll now do for the next 20 minutes is just dig into that in a bit more detail because I think there was quite a lot of it you guys quite quickly. Um, so let's start with how we actually went about forming this database. So the first big problem, the reason it's taken us so long to come up with this idea to form this database is that we kind of were stuck for a really long time on the very first question of what is climate change research. So we knew or Leo knew that we wanted to visualize all of climate change research. But how do you actually define what a climate change paper is? There are some papers which are very obviously to do with climate change. And so this is an example here. One of the very first scientific studies to connect atmospheric CO2 to a rise in global temperature was published in 1856 by Ununice Newton foot. Very cool. Very clearly a paper about climate change. These these papers that are the about sort of the very core foundations of what climate science is that came out in sort of the 1800s 1900s. It's quite clear that those about climate science. But since then, there's been this massive ballooning of literature about climate change. Well, a ballooning of literature about everything really, but it's actually very difficult to define exactly what a climate paper is. So, there have been a lot of different attempts by various groups. One way, for example, is to use keyword searching. So to put a massive massive list of key words, for example, carbon dioxide or global warming or or atmosphere, whatever, put a whole bunch of keywords um into some kind of program and use that to sift through something like Google Scholar or Open Alex, one of these big big databases of climate research, and just pull out every single paper that has those keywords in their abstract or in their headline. Um this isn't a foolproof method, obviously. there's there's a lot of scope for you to pick up papers that you don't want or to to accidentally leave out papers that you do want. But but this is a method that defines peer-reviewed studies that have been published using this method. Um it's a very valid method. Um other people especially more recently have used artists they've trained an AI model on water climate. Um and again this is a very valid method that people have used. There are pros and cons to it. Um but we didn't feel that carbon brief that this approach was quite right for us. um we decided that we wanted sort of take a very completist approach and so what we decided to do was to base everything on what can be absolutely certain is climate science and what we decided to look at was the intergovernmental panel on climate change the IP so the IPCC has been publishing some of the world's most authoritative summaries about the latest climate science um they published their very first reports in 1990 and this is kind of the gold standard of climate science I mean I'm sure a lot of you are climate change scientists so I don't need to tell you this but IPCC reports are the absolute gold standard. Um they are the most complete assessments of climate change research. Um and so far the IPCC has published six um assessment cycles. So six sets of reports. Um each one takes thousands of scientists I would say hundreds of thousands of hours to pull together. Every line in a lot of these reports is is sort of worked over and checked and agreed by many um and and the kind of final summary government and it's approved line by line. So these reports really are um cornerstones of climate change science and between these reports um there are there are hundreds of thousands maybe even millions of no hundreds of thousands um of other academic referenced um and so we decided to use this we decided to use these IP core of our data. Um so you can see that the number of references in IPCC reports has increased over time pretty dramatically. Um so IPCC reports are always published in working group sets. So working group one is the physical science basis. Working group two is impacts adaptation and vulnerability and then working group three is about the and then attached to each um also a group of special reports. So for example a special report on 1.5° C or in the upcoming AR7 assessment cycle there's going to be a climate. Um, and so yeah, each of these reports has a bunch of references, like tens of thousands of references. Um, and the number of references in these reports has just grown as the length of the reports. Um, so let's just uh define some terminology really quickly. What do I mean actually when I say reference? Again, I'm sorry for a lot of you if this is quite basic stuff that you know already. I'm sure a lot of you academics know what a reference is, but bear with me because there is some some terminology stuff that I want to just nail down here. So, I'm just going to very briefly talk through what a study is and what a reference is. Um, again, bear with me if you know this. This will just be quick, but an academic study is um a document which sets out a new scientific hypothesis. It could be an experiment that someone's done. Um, and you start at the very top with your title and your DOI. So, the DOI is like this unique. Underneath that, you will list all of the authors of the paper. So, these are the people who to making this paper, this document. Um, and so it could be the people who conducted the science. It could be the people who own the lab. It could be the people who wrote the paper. But there'll be a list of the authors. And then at the very bottom of the paper will be the references. So references are again a cornerstone of how science is conducted. Every single time that um experts write a paper, they need to substantiate their claims. They need to back up what they're saying. And to do this, they will link to past research. So within the body of the um academic paper, there'll be lots of references, lots of numbers. Uh so for example the the paper might say the planet has warmed by 1.2° C number and that number will link down to a reference at the bottom and that reference is another paper or book or report which shows that the planet has warmed by um it is really really important to do this like you wouldn't get a paper published without a good set of ref. Um so in a study at the bottom you have the references um as you can see the blue arrow is pointing to like the references the other studies. So that's our sort of carbon brief terminology. We're calling the list at the bottom um of the paper that links to a bunch of other papers. We're calling those papers that then we also have what we're calling citers. So citers are papers that themselves reference a study. So let's say that this study in the middle is study A. Any papers that in their list of references at the bottom link to study A, we will call citers of study A. Meanwhile, any papers at the bottom of study A uh in in the list of references, we will call references of study A. So, this is our our references and citers terminology. Um, someone jump in if that was not clear, but hopefully hopefully it is because it's about to get more um so carbon brief built Cosmos database on these references and citers relationships. So, uh you saw this graphic very briefly in the video from before, but I'll just go through it again. So in the center of our concentric circles over here we have the IPCC reports. Um so these have in total um 107,000 references. So that means uh in the list of references at the bottom of the IP um there are in total if you look at all the IPCC reports there are 107,000 different academic studies linked there. Um and that is your your like first circle in the in that diagram there. Um, I should note that what we're counting here is just the unique references. So, obviously an IPCC report or multiple different IPCC reports might all reference the same paper. So, here there are 107,000 unique studies referenced, but in total there are 162,000 just so you know. So, there's a much bigger number if you're including duplicates. But anyway, so that's that's your that's your references. That's our first layer of the day. We then added another layer onto that. We took every single one of those IPCC references and we found all of their references. So we're calling this second order references. So you can imagine the scale here is just ballooning very very quickly. Um so yeah 1.4 million second order reference. We then also included IPCC citers. So that is academic studies that themselves reference IPCC reports or chapters. Um 168,000 of those. And then we decided to just be a bit more um completist about this. Um we decided to add another another lay another step onto this analysis. And so we identified 22 key climate focused journals. So these are journals for example like nature climate change which really do just focus on climate change. So here we didn't include big journals like like nature or science that cover everything in in in sort of academia. We we just focused on the climate change specific journals. and we had some of our carbon brief scientist friends help us with identifying those journals. So we identified 22 journals and we looked at all of the studies that had ever been published in those 22 journals. So that was about 50,000 studies and then we looked at all of the studies that um all the studies, books and reports referenced by those papers. So that was 642,000 of those. Um so if you pull all of those together that you get 1.8 8 million um papers and books and reports. Um and that is that is essentially what this database is. This database is is that compilation of of papers and books and reports. Um and so yeah, that that that's basically how we pulled together database. Um right, I say we here, I mainly mean Tom Pearson and then um we had some fantastic collaborators who I will who I will thank formally at the end, but like Felix and Tristan and Travis Cohen from the University of Exit. did a lot of work on pulling this database together and if you have more technical questions about how that worked, you can grill Tom about them at the end. Um because he was really deeply involved in like pulling together this massive massive database of papers. Um so yeah, it's really cool. It was, I would say, a pretty big technical challenge to get this much data all organized. And we spent a lot of time kind of, well, we again, Tom spent a lot of time kind of uh debugging and cleaning and and getting everything to the right format. And yeah, again, if you have more technical details about how we did this, please do ask Tom at the end. Um, because I want to get on to the bit that I think is more exciting, which is the results. So, there is a lot you can do when you have 1.8 million papers. There's a lot of analysis you can do and we have so many ideas for different types of analysis that we want to do in the future. Um, but for now we've just decided to kind of just just to show in principle what could be done, we decided to do some initial analysis and what we came up with was this idea of the Cosmos 500. So this is basically a ranking of the top 500 scientists, institutions and um uh what else was it uh countries um in the database. So we when I say top actually I'll describe what I mean when I say top in a minute but it's basically who who was cited most highly who was like included in the database most um so publications that are referenced really extensively they're often referred to as highly cited. Um and in academia if a paper is highly cited i.e. lots of other people have referenced that study. It generally is a sign that the paper is kind of well has important findings or is very foundational or or maybe it's a it's a data set that lots of other people are using. But but generally you can say that if something is highly cited, it's sort of high quality science that's very uh and so we thought it would be interesting to see which which scientists have published the most highly cited papers uh and which papers kind of come out top of our ranking most highly cited. Um, and so what we did was we defined this metric called a citation score. And so this is the number of times that each academic paper, book or report is referenced by others within the Cosmos database. So very clear here, this is just showing how highly cited something is within our database. So you could have a paper that say a statistics paper that is really influential in in some mathematical field and has been hugely highly cited by other mathematicians, but it might not rank very highly in the Cosmos database. Equally, you have might might have something that ranks really really high in the Cosmos database that when compared with broad academic literature is actually not that influential, but like because it's a core foundational climate science paper, it ranks really highly here. So yeah, just to be very clear, this is a ranking for the Cosmos database only. Um, and in these rankings, we have omitted the actual IPCC reports and chapters because those obviously would have just been like the highest ranking. They would have like topped every chart and it would have thrown everything else out of whack. Um, so that's just the preamble and now we can get to the exciting. Um, so first up, authors. So again, this citation count is how many times each expert's publications are referenced by other publications within the Cosmos database. Um, we we pulled the top 500. So if you go on to the Carbon Reef website, uh, the Cosmos database website, you can see all 500. Um, but here I've just shown you the top 20. Um, so some really well-renowned um, climate scientists. Um, and Carbon Reef actually had a chance, well, I actually had the chance to chat to quite a few of these people to pull together a piece about their achievements and their work they've been doing. It's really cool stuff. Um, I will note that um, there's not a lot of diversity. Almost half of the authors in the Cosmos 500 are actually from the US and The Global South only accounts for 4% of the authors. um only 10% of the authors are women and actually you have to go down to ranking number 36 to even get the first woman in the Cosmos 500 that the first 35 are all men. So um this is I I mean it's a reflection of how academia and then climate science included is quite dominated by men from the global north um which I won't delve into now but that's a topic that I have spent a lot of time working on in the past. So if anybody wants to ask me about that later you are also welcome to do that. Um but yeah, so this this was the ranking for the Cosmos 500 for the top 20 authors. Um and Carbon Brief actually um interviewed the top authors. So this carbon brief did an interview with Philip Keys and with Professor Dler Vanurren. Um so Phipe was the top author in the Cosmos 500 and Professor Van Beern was oh I forgot to mention this. Okay, so carbon brief we kind of have two versions of the database. We have the main Cosmos database which is everything and then we have the IPCCon database which is just those 107,000 papers that were referenced by IPCC reports. Uh and so when you're looking at the the results you can kind of toggle between those two databases. Um so what we have here is the top two authors from those two separate databases. So um yeah the Cosmos database is the main one with everything all of those steps I talked about. But we did also think it would be interesting just to see who is the most highly referenced scientist just within IPCC reports. You know, there's really really foundational climate science reports. So that's what we have here. Um so we also did publications. Um so the top publication actually wasn't something particularly climate sciency. It's like kind of a stats u mathematical manual thing. But maybe that's not surprising. It's quite a foundational um piece of literature that clearly lots of people have referenced. Um but yes, my colleague Rob actually did this part of the analysis. Um so we found that maybe unsurprisingly, nature and science are the top um climate uh journals that the studies appeared in from the Cosmos 500. And you can kind of see this like increase in studies um going up to the 2000s. The most most of the studies in the Cosmos 500 were published in the 2000s. Um, we'll have to see as time goes on whether this is just because studies in the 2010s and 2020s haven't had enough time to be referenced by other papers or whether there was like a weird reason for some spike in the 2000s. We'll have to we'll have to see that um like as time goes on and as we keep updating progresses but again you can have a look at which publications were the most highly cited which I think is pretty cool. Um and then institutions as well. We had a look at which institutions are publishing um are publishing the most highly cited papers and the way we did this was by linking authors to the institution that they were based at at the time that they wrote that paper. So if someone was from um Colombia University at the time that they wrote a specific paper that was very highly cited then they're citing um even if like now they are working somewhere else or retired or something a whole complicated system of like trying to deal with figuring out how to link authors. But again you can kind of see this this global north bias again in which institutions rank most highly in the Cosmos 500. So yeah, if you take the top 500 institutions, the list is dominated by the US. You then have the UK, Australia, Germany. I think the first global south country is China um 16 which is I mean not bad but I mean if you compare it to the US of 183 uh institutions you really can see how crucial the US is for climate science research which I mean not to go into too much detail but given how the Trump administration is really gutting climate science in the US right now I think this is pretty significant. Um yeah I or it'll make me sad but yeah the US very important for climate science. We'll have to see whether that changes. So, um, but yeah. Okay. So, last thing I wanted Oh. Um, okay. Well, what I wanted to show you was that that Tom has made a really cool like interactive visual visualization page thing to show. Actually, maybe I can show you. I think uh give me two seconds, folks. Sorry. Okay. What I wanted to show was cuz I've been signed out. What I wanted to show you is this was the map. And I'll just do it live actually. Okay. So, the final thing I wanted to show you is how you can play with this database yourself because it's a really really cool resource. And I had made a little video where I talk through how you can do that, but clearly clearly Google has decided to sign me out right now, which is terrible, terrible timing. Um, so I'm just going to show you live instead. So, the database is actually interactive. It's clickable. You can go in there and you can sort of have a look around it yourself. Um, and so this is how you can do it. You can go to the the Carbon Briefcosmos uh page. That's interactive.org/cos org/cosmos/ and then if you go onto the map you can have a little scroll. So what you can see here is okay we have a bit of explanation that every single star in this database represents one of the studies in the cosmos database and that the colors illustrate clusters um of similar topic areas. Uh so they show like which academic areas are located in our within our map. Um so for example these two large circled areas represent physical science and social science. So you can see the physical science the blue is a pretty is a much bigger cluster than the social sciences which is in yellow but both and then you can go in even deeper. So um yeah so the green and the red areas capture the fields of immunology and microbiology for example. Um, this is really cool that you can see the coloring. And again, maybe Tom can talk to anyone who's interested about how that coloring uh worked, like how those topic clusters worked. Um, but it's really cool that you can actually see them grouping together like this. Um, yep. So, okay. Yeah, this is the thing I wanted to show you. This is the thing I thought was the most cool was that, um, you can see that some of the circles, um, 500 of them to be precise, are a bit bigger. um those are the the studies that are actually within the Cosmos 500 ranking of the most highly cited publications. So what you can do is you can go and zoom into them. So for example, this one here that's been highlighted is the Stern review um which was uh the most highly cited publication across the IPCC only data and zoom. So yeah, the whole thing is like clickable and movable which is I mean so cool. Um, so you can see that down here at the bottom you have all of the different um all of the different categories. So for example, yellow is social sciences here. So you can see psychology is in orange for example. You have physical sciences is in blue over on the on the left hand side at the bottom of my screen. So environmental science is light blue. You have quite a lot of those. Um I'm going to click on okay. So like this one is green for example which means it's in health science. And so I can click on it and we can see that it is um discussion reporting of 14C data. Okay. I don't know quite what that means, but that's a that's a paper that was highly cited. See, if I'd had my video, I picked a really nice one. I just can't remember Galaxy, but anyway, you can click around and you can see what the papers are. Um you can see um what the ranking is. Okay, so this is this is um from the bulletin of the American Meteorological Society um representative of a climate science study. But anyway, you can have a click, you can have a play, you can see what different papers are there. You can see which ones are clustered for example together. um that's about um out of uh social ecological resilience in this together. But anyway, you can have a play. I just think it's very very cool that with this database you can have a Oh, I've been signed out. Okay, [snorts] I had one more slide. Uh oh, it's making me create. Okay, I'm just going to add my last slide. I remember what it said. Um or maybe I can stop sharing and Simon can bring up the slides and just Yeah, I'll stop sharing and maybe Simon can just bring up that last slide for me. um if that's um but I will yeah so that last slide basically what I wanted to say is that the Cosmos database is um this is the very initial kind of early phases of the Cosmos database we have done this as a kind of proof of concept we've developed the tool and now we want people to use it so we've done our Cosmos 500 ranking and we're going to update this every year we're going to update the database every single year hopefully that's the plan uh to add all of the new climate literature that comes in and then you can expect a really big bump in the database once the seventh assessment recycle seventh assessment cycle reports. You could expect a really big bump then. But in the meantime, there's a ton of other stuff we think we could do. There's kind of temporal analysis stuff we think we could do. There's keyword search analysis. We have a ton of new ideas. And what we would really love is for academics is for people in the scientific community to get in contact with us if you have ideas for the database. We deliberately haven't made it open source. in part because we're a bit worried about AI bots scraping it and in part because it's just such a massive database that getting it open source online would be a nightmare. Um but if academics approach us with ideas and ways they want to collaborate and use the database, we would be really really keen to work with you guys. So um yeah, I don't know if Simon is able to share the slides. >> Oh, so um I also had a problem of being kicked out and accessing it. So, apologies. I could not step in and help even though it was supposed to be >> um backup. Um >> I never had this specific technical issue on a presentation, but you know what? At least it happened near the end of the presentation. So, >> um I think I showed you when I was scrolling where you can access the Cosmos database and I'll put a link in the chat to where you can access the Cosmos database as well. Um if you go on there, there is an email address where you can email carbon brief to tell us if you have an interesting proposal project idea. Um, and yeah, I think it would be really cool to hear from some of you guys if you have things you want to do. So, thank you so much and yeah, maybe we'll take questions now. >> Yeah, thanks. So, thanks for that uh presentation. I think we'll bring in Tom now as well who also discussed his perspective on developing the database as well. Um, I just quickly kind of clarify um partly because I miss this myself when I was trying to understand why I've been kicked out of Google. Could you just quickly say in what ways perhaps our geoscientist audience might be able to contribute to this database for example for example it's not directly by adding papers um or it's not open source either um in what ways could they perhaps uh contributions to a database >> yeah so hello um yeah so so basically we're looking for kind of collaborators essentially we've got this database we're hosting it sort of you know in private at the moment Um but yeah, we're looking for people who have ideas about what we could use this kind of vast corpus of of kind of interlink documents to uh to discover. I mean, we have a few ideas like for where we want to take it for the for the next kind of, you know, ne for next year's release, but um but yeah, we're particularly interested in in sort of reaching out to the community to sort of ask about uh about what other people find interesting cuz yeah, it's a it's a big resource. >> Well, thanks. I guess that um also kind of leads to other question is like what um I guess is the intention you had in terms of its use, I guess. Um there's obviously the visualization of all these papers etc. Um I guess it's searchable, findable. So at the base level it's also just to help people kind of find um research and see how it whatever research might be linked to it for example. >> Yeah. So to kind of go into a bit of detail about the sort of structure of the database. So the initial the sort of initial piece of work was uh commissioned like a couple of years ago. Um we worked with a a French PhD student Felix who basically uh wrote the software to scrape various different sort of open repositories. So like Open Alex, Google Scholar and a few others. Um and sort of built up this initial kind of unstructured database which we then kind of converted into uh a kind of a a more formal structured database via the process of kind of cleaning everything up, working out when two papers that are described slightly differently in different repositories are the same thing. So we dduplicated stuff and we sort of formalized the relationships between them. Um this was all really kind of like with a view of this kind of getting a kind of like broad picture of climate science. So uh what I should described as kind of like Leo's kind of headline vision of this kind of like big big kind of big picture thing uh and the and the sort of like uh and and the ranking. But I think what we've kind of realized since um since compiling this database is actually a lot of the kind of interesting stories and interesting features of the database happen at not the kind of global scale but the smaller scale where you see little clusters of work on kind of immunology or on um uh sort of disease resilience or or adaptation and it's interesting as well I think and and one area that we do want to explore is how these areas develop over time. Like how we kind of did a did a kind of quick sort of time series test to see how the the sort of literature from the first assessment report varied from the the sixth and how over time like the social sciences have formed a much bigger part of of those reports than than they did initially. And so seeing that kind of information and using that to to kind of like find out about how knowledge is produced and how it develops um is something we're really keen to do over the next over the next um you know 6 months to a year. >> Sure. I see. Yeah. Um actually you talked about um kind of grouping on a broad idea of climate science and trying to revisit that through the database. Um I I it links to like one of the audience questions we have actually as well about how these groups are kind of defined etc. Um so one of the audience members uh said that some people in the top list would not see themselves as climate scientists at all. But others uh might see this as um I suppose a kind of a broader picture of like uh research speaks to other research is it how can you draw draw hard boundary between um uh certain kind of fields etc. Um I guess and the question kind of links up to uh one of the perhaps uh in some areas like subsets of the data set might be better for some questions about a broader lookout better for others etc. Um I was wondering if you could uh talk more about how you then um define that because I guess people obviously identify themselves different ways as research science but for us um you applied it another way. >> Yeah. So as Aisha said this is like kind of probably the broadest way in which you can kind of conceive of like you know climate science like you know it sort of extends to as as far out as you can go and still be thinking of this as kind of like knowledge about climate. Um how we actually came up with the the the sort of categorization it was derived from open Alex. So open Alex the sort of big online sort of open- source I think open source open data um repository for for papers has a process by which they you know do some sort of automated analysis on the text of of these documents and basically come up with a a set of um categories that fall within this kind of like taxonomy where you've got like the top level which I think they call domain and then like topics and then subtopics and we kind of went down we kind of queried Open Alex down to the sort of subtopic level. Um [clears throat] and then sort of once we had that we realized that was far too granular. So we sort of took a step back and and looked at it at the topic level which is the kind of level at which you have things like computer science, uh material science, chemistry, um psychology that kind of level of of granularity. Uh, and then we just used the primary topic for each paper as as our way of kind of coloring it in. Essentially, like all of that information still exists in the database. So, if we wanted to, we could sort of like look at a different way of of categorizing things maybe either at a more granular level or kind of like waiting between the topics. Or if we wanted to look specifically at a particular area of knowledge, we might choose say um neuroscience and say just pull out all of the papers that are neuroscience that have neuroscience as their topic and show us the connections between those or the citation relationships or which authors are prominent in that corpus. So like we didn't make any kind of editorial decisions about what cons constitutes a paper and we tried to sort of cast our net as broadly as possible in in terms of what we include. Um so yeah it's absolutely right that like a lot of these papers aren't what you'd traditionally consider climate science >> but I mean that's a key point of trying to relate um how research impacts other things right we we know how climate change impacts health so health papers being included in our medical papers makes sense right and trying to get a broad picture so I usually want to say something >> sorry no absolutely just to I mean everything that Tom said absolutely true and this I guess it's just to add that you don't have to be a climate scientist for your yeah for your work to be very influential in climate science But um yeah, I guess as Tom said, it was it was an editorial decision and we were thinking initially about whether we should try to narrow things down a bit more and we we've made the decision deliberately to be very broad scope. And I think that is actually really interesting is when you can see things like like the top paper being a essentially a reference manual for how to use the coding language R being the top paper. But that I think is really interesting because it shows that that's such a foundational piece of work. Those experts are probably not climate scientists. they wouldn't consider themselves climate scientists but now they can see that their work has actually had a really huge foundational impact on climate science going back decades which I don't know hopefully they feel good about >> so when people asking about projects this kind of um input in terms of granularity etc might be a way people could increase its use for example I was thinking sorting the uh database by phenomenon I don't know by natural hazard or something or perhaps by a method is uh is perhaps a way to do it perhaps um I guess when you're developing these links um it's basically based on a broad topic of the paper right it's not on the these more granular >> options yeah so we don't um we don't have the well the our colleagues uh Tristan and Travis at exit are currently looking at the idea of downloading and sort of analyzing a lot of these papers on a sort of specific level but we're Yeah, we're just using the kind of like metadata that's stored in these various repositories. We're actually not looking at the papers. Like in some cases, we've looked at the abstracts, but mostly it's just work that other people have done that we're building on. >> I see. So, it's connecting to other databases then. Sure. Um like what's the relationship then as well? For me, it's more about um uh I kind of question resilience as well. perhaps as um as known as climate data databases being removed um due to kind of government action etc like that. Um is that something you consider perhaps when looking at these [clears throat] relationships at all or um I guess also like I creating this uh database um where it's bas like can be fed by other databases it's kind of it independent um mapping that use. >> Yeah. So like kind of I think um I think there's like the the sort of big big data sources that we use like open Alex and Google Scholar. I think there's definitely question around questions around though the resilience of those sources. I think open Alex seems like a a kind of good bet to hang around for a while. I think Google Scholar is on um you know perhaps somewhat more shaky ground given Google's past record and I think we've been hearing recently. But um but like a lot of these data like a lot of these databases are themselves interrelated. So open Alex populates a lot of its stuff and kind of like through Google Scholar or or other um other ones. I keep on mentioning those two but there I think there was four main main databases which we which are scraped. Um but yeah in terms of like how the database that we've created sort of persists and it you know we're sort of hosting it on a private server at the moment. it's like backed up. But we want to definitely we're definitely sort of investigating ways in which we can sort of make this more open and make it more sort of we so we can kind of you know not act as the gatekeepers for it in quite such a a sort of strict way as we are at the moment. We sort of want people to be able to access it but there's sort of yeah cost time and expert implications associated with all of that which we haven't yet worked through. of course. Um yeah, like you'll give me flashbacks to my own work when I was a researcher as well. Um yeah um one one thing that came up though actually I think this is a question for you actually is uh we talked about these other databases but you also talked about um um kind of the perhaps historical impact of representation on on who gets cited and the kind of diversity of people represented in the citations. Uh Bisha, you also have like uh there's another database that common brief exists as well to help kind of combat I guess bias against certain authors. Perhaps this is a nice time to kind of just briefly highlight that at all. >> Yeah, thanks Simon. So yes, I've done a lot of work over the past five six years about the lack of diversity in climate science literature. Um so I did some analysis in 2021 now how time flies. uh showing basically that women and experts from the global south are hugely underrepresented list of some of the most highly cited climate science research and that I think is borne out at a much much bigger scale by the cosmos database here. I mean yeah you can see it's very consistent across that front. Um and that's for a range of reasons to do with barriers in conducting the research in the first place barriers in getting research published. There's I mean there's a lot of stuff you could look into for that. Um, yeah, I won't get into all of the details now, but yeah, a few years ago, um, I helped to develop, uh, another database called the Global South Climate Database that Tom actually was very key in helping to to put together as well. And man, it's great. Um, this is a database of climate science experts from the global south. So, experts add themselves to this database. It's a it's a self-submission thing. So, every expert on this database wants to be there. They want to be contacted. Um, they want their their voice to be heard. And I can see Simon, you've just put the link in the chat. So, thank you. Appreciate that. Uh, we now have more than 10,000 experts on the database. So, I think it is uh we want to well, for example, as a journalist, if you want to find more experts from the global south to quote in your reporting, or as a scientist, if you want to find peer reviewers, I've heard of quite a few reviewers. Um, and I guess the core point I want to make is that diversity isn't just a this isn't um this is, you know, genuinely important. getting different research brings different perspectives that that genuinely do help to progress science. This isn't just just a oh, what's the word I use? Charity case or something like this is genuinely very important. It's an issue we should all care about. So, I think there's definitely some really interesting analysis we could do on the Cosmos database more linked to this um where people are publishing from, who who is publishing and and maybe hopefully how that has changed or improved over time. That's a bit of analysis I would really love to do. We'll see if that bears out. But >> thank you. So of course this it's not the only database and if you want to start engaging with people um uh looking where these connections are in project costs is good but always good to keep an idea of these historical systemic um biases or problems which otherwise result in kind of like quite a a poor representation in um science or even even in kind of climate change research group. Uh so the global self climate database is one way and thought looking out for scientists who aren't well presenters. Um we did also get a question about uh the um dates of some of the publications. Uh one of the attendees said that the top 500 um I guess has a lot of articles which are pre20 I guess not many after 2010. um and particularly is asking about finding papers that include technology. Is there a way and there database and also um why the database might be looking at more or I guess preference for historic papers or publications. [clears throat] It's two questions in one. Apologies. >> Yeah, I'll I'll answer the one about the um the age of the papers. I think that's just um because we're because we query by most citations like older papers are kind of inevitably going to have more citations cuz they've had longer [clears throat] to gather them. Um one thing that we do want to sort of look at is uh sort of balancing this a bit and have a kind of perhaps a sort of like you know you know either sort of timebound query. So we say like kind of what what if we only asked for papers for the last 10 years or maybe what are the papers that have acred um acred citations most rapidly over over the last 10 years. So that's all kind of stuff that we can we can look at in order to kind of come up with different rankings. I think at the moment like in some ways like the the the sort of huge bulk of the work here was kind of collecting and cleaning this data and putting it together into a into a kind of coherent single database. Um and then we kind of like at that point we were like okay we'll publish this with kind of like the most easily understood way of of kind of creating this like list of 500 papers or people or whatever. But I think definitely going forward we want to look at ways we can sort of yeah surface more recent papers surface like different kinds of contribution and and different kinds of uh different kinds of paper. Um what was the other question? I've forgotten already. >> As about um how to find um perhaps article articles and new technologies for example. Um >> yeah I mean that's a good question. I guess you'd probably want to look at um I'd probably go in via the sort of open Alex topic hierarchy and like find a kind of subtopic that sort of matched a particular kind of thing. So like presumably something inside the engineering or computer science topics there'd be you know more more kind of um granularity in there and then and then looking that way through it. But it's not really because we're not doing any kind of text on uh any sort of text searching isn't possible. So, so that would need a kind of different a different kind of database. There's actually a a sort of second database that sits behind the the main kind of Cosmos database which is a a kind of unstructured document store which has a lot more um a lot more kind of random information in there. So like papers in that often have abstracts which would be a good place to do a text search for kind of you know new technologies or whatever. Um so like the Cosmos database itself is probably not a great uh a great way into that kind of stuff unless you're going in through those topics and subtopics. >> So the way the cos cosmos project cosmos can inform your search for other databases as well actually in that sense. >> Yeah. Yeah. Yeah. And I think as well if you've kind of like got a got a particular paper that's a way in like you can obviously do a query which is like show me all of the I've got this paper um the title is this show me all the papers within the database that reference that paper and all the dat all the papers within the database that that paper references and then you kind of like you know come back with like 100 papers and that will give you a a kind of way to look in and then you can make similar queries on those papers. So you kind of you know the database is a web essentially and and you're kind of like crawling around it finding uh finding the interesting things that way rather than uh like quering it like a more traditional database. >> Sure. Yeah. I I guess then a project perhaps that could be brought forward then um we've talked about gridarity in terms of topic theme and subject but also um yeah as you said looking at um things that are upcoming rising. So the thing that came to my head were Reddit filters which kind of filter by like a most recently interacted with high level of interaction between certain bimes etc etc. Um so in which case if you want to look at uh popular papers in the last five years or which are getting attention for some reason that might be another a project people could borrow. Uh yeah. Yes. >> I mean because it's I think it's interesting that you mentioned that because it is it is a graph database which is essentially the kind of same sort of database that backs something like Facebook or a social network. It's about >> like it's a it's a kind of database that privileges the relationships between these nodes as much as they're like first class citizens as well as the documents themselves. So, so like things that are about like relationships between papers or between disciplines or between authors like this database is is kind of made for that kind of uh research. >> Excellent. Um I have a um I've got another question um which is about um ranking the most cited data sets um uh as [clears throat] a method to try and identify um foundational climate data sets or data centers. Um it's been interested to some other uh geoscience communities. GCS has been one of them. I was wondering if that's uh something you've uh included in in uh Cosbox. So ranking the most cited data set. >> Uh no we haven't. But that does sound like a really interesting uh interesting thing to think about. >> But there you go. So I mean if you have these ideas please submit them to cover brief. [snorts] Um so it's cosmos.org. So these are things about granularity looking at pubshed method of phenomenon looking at um most cited data sets um perhaps searching by rising or most recently um popular citations etc. I think these I guess are the things that you might want people to bring forward to you as a project or idea. >> Yeah. >> Yeah. Absolutely. >> Absolutely. Um we're running out of time very quickly. Um I just want to like ask perhaps one final question. Um um well actually before that there's one thing um about information available. If people look at the data set and searching for things it's quite easy then to kind of move on to like find a paper etc. site it read it. Um I guess I was also thinking about in terms of contacting people or um sorting things by disciplines and fields. Are these uh possibilities in the data data set? Were people looking around? >> Not in the not in the website as it stands like because of the size of the database. We did like I originally kind of tried to like I got it I got the I got the sort of data payload down to about 200 megabytes and I was able to deliver all of the titles and um >> and and subjects and and publications for all the papers. But we decided that like kind of lumping someone with a 200 megabyte download when they visited the page was probably a bit much. Um and and sort of splitting it up intelligently was again it's something we didn't really have have the sort of time or resources for. So >> at the at the moment uh that's not possible. You can sort of you know search the tables that are on the website but um but kind of like going deeper is is is not um not yet a thing. I mean, which is a real shame because one of the things I really enjoyed when I was like developing the site was being able to, you know, zoom into that big cloud and kind of like see the see the papers that it had like, you know, clustered together and and, you know, kind of like really just get get a feel for the scope of it was great. So, it's a it's something that I would really like to do, but um but yeah, it's a it's a technical problem that needs to be solved. >> And we got no foundation down. We got a massive network of relations between different types of research but uh isn't I guess maybe broadly looks at climate change research that people can access to look around and also provide their own thoughts and contributions to uh I'm going to close the webinar but uh I'm going to put the floor open to any final comments or requests um uh from the panel if there's any kind of one last thing you want people to keep in mind about the database or about um contributing or working with carbon reef as a >> um I have I would like to give further props to Felix Chavelli um Travis Cohen and uh Kristen Cam for their work on this because they were like kind of you know they're they're the main stars of the show and they're sort of behind the curtain but I think uh yeah their their work was absolutely invaluable in in doing this. It wouldn't have happened without them at all. Absolutely a lot of work and effort to produce such a a vast database. Uh Aisha, do you have any uh final thoughts or requests? Even >> Tom absolutely took my answer because [laughter] my final slide which never mean I guess I'll also mention the rest of the carbon brief team. Then Leo Hickman, Rob Mcweeny, Cecilia Keiting, uh Joe Goodman, Tom Pra, like this was a massive carbon brief wide effort. Um, and so yeah, very keen to bring in other collaborators to to continue expanding the Cosmos team. >> Brandon, thank you so much. As mentioned before, if you have an idea about what perhaps a feature you want to see in Project Cosmos, how to um make it kind of more usable or appropriate to your research as a geoccientist, um, please send us to Cosmos at.org. Otherwise, thank you so much for joining us on this uh, incredibly hot afternoon. Um and thank you for our speakers comparison each time for joining us today. This uh webinar will be on the EGU YouTube channel in a couple of weeks once it starts. Otherwise, thank you so much again and have a lovely day.