EGU WEBINARS: Introducing Project Cosmos, The World's Largest Climate Research Database
Watch on YouTubeVideo summary
Project Cosmos represents a groundbreaking initiative developed by Carbon Brief over eighteen months under the leadership of Simon Clark and Leo Hickman's vision, aiming to visualize and analyze the entire global network of human knowledge on climate change. By establishing IPCC reports as its foundational "gold standard," the team constructed an extensive database comprising approximately 1.8 million unique academic publications linked through more than 40 million citations. This massive dataset was built by expanding beyond primary sources to include second-order references within those reports, studies that cite them, articles from twenty-two key climate-focused journals, and every paper referenced by those journal entries, creating a comprehensive map of the field's intellectual landscape currently accessible at CarbonBrief.org/cosmos/.
The interactive platform reveals critical insights into the structure of scientific influence through its "Cosmos 500" ranking, which utilizes a custom citation score to highlight significant disparities in representation. The analysis shows that US-based authors and institutions dominate nearly half of the top positions, while experts from the Global South account for only four percent and women appear merely thirty-six times among the most cited individuals. To address these systemic barriers regarding research conduct and publication visibility, the team has supported the creation of a self-submitted "Global South Climate Database" listing over ten thousand experts seeking collaboration, emphasizing that fostering diversity is essential for scientific progress rather than just a charitable endeavor.
While users can currently explore topic clusters ranging from physical sciences to social psychology via color-coded maps or trace connections between papers within the graph-based network, certain advanced features remain limited due to technical constraints and resource limitations. The primary database lacks full-text search capabilities at this stage, restricting text queries mostly to abstracts found in an underlying unstructured document store that allows users to locate specific technologies through openAlex's topic hierarchy; furthermore, downloading large metadata payloads is currently restricted but remains a future goal for deep data exploration. Future iterations of the project plan to introduce filters for recent publications or rapid citation growth and may explore ranking datasets or identifying rising topics similar to social media trends, with updates anticipated annually following major events like the IPCC Seventh Assessment Report.
Although direct contributions to the database are not currently accepted due to its private hosting status aimed at preventing AI scraping, the team actively seeks collaborators for future analysis and invites attendees to submit feature requests via cosmos.org. The project acknowledges key contributors such as Felix Chavelli, Travis Cohen, and Kristen Cam while maintaining a focus on expanding the network's utility without compromising data integrity or technical feasibility. As the database evolves, it aims to provide deeper insights into climate science by balancing current analytical capabilities with long-term goals for inclusivity, comprehensive search functionality, and broader access to metadata as resources allow.
Read the full video transcript
So I think it's time to make a start. Um
so firstly thank you for joining our
webinar today. Um
uh this webinar is on introducing
project cosmos the world's largest
climate research database. My name is uh
Simon Clark. I am the project manager
here at the European Geosciences Union.
uh and today's webinar should provide
some insight into development of uh this
new database as well as um some uh
insight perhaps how we can use it as
well. Um the webinar will begin with a
presentation followed by a Q&A all
within this 1 hour. Uh if you uh want to
have a question or ask a question for
our speakers today, there's a Q&A box at
the bottom of the screen where you can
enter your questions which be asked
after presentation. You can also upvote
questions you think uh like you want to
be answered. Um and so we have two
members joining us from carbon brief
today. We have a tandem who is carbon
briefs science correspondent uh who has
won um awardw winning uh leading the
2024 world meteorological society
emerging communicator award as well as
short list from multiple others. Um
Aisha will be doing a presentation today
and joining for the Q&A later is Thomas
Pearson. Um Thomas Pearson has uh over
20 years experience working in
presenting complex data in clear and
accessible manner as well as producing
award winning graphics from everything
from BBC news to financial times and to
natural history museum. So as a brief
introduction to what we're going to talk
about today and who is presenting uh
Aisha I'll move on to you if you could
share your screen uh and start your
presentation please.
>> Absolutely. Thank you so much for the
introduction Simon. Um, all right. Yeah.
So, thank you all so much for being here
to learn a bit more about Project
Cosmos. So, just to give you a bit of
background, uh, myself and Tom Pearson
work at Carbon Brief, which is a climate
change journalism and, uh, analysis, uh,
company website. Um, and we do a lot of
kind of, uh, news reporting, but we also
sometimes work on these really big
in-depth long projects. And I think
Project Cosmos might be an example of
one of the most in-depth and longest
projects that we've worked on so far.
Maybe Tom can correct me if that's
wrong, but we started this project a
year and a half ago. And I would like to
say this idea was first seeded in the in
the mind of Leo Hickman, who's the the
head of Carbon Brief, I want to say a
decade ago, at least since I joined the
company six years ago. This has always
been an idea that he's been floating
this concept of bringing together all of
climate science literature, all of the
academic papers and books and reports
about climate change and somehow
visualizing them, somehow showing how
all of these different papers link
together, maybe through topic uh
relationships or through links from
authors. Um, I think he's always had
this idea of this big visualization,
something like what you can see on the
screen here, to show the sheer magnitude
of climate science and how all of the
different papers link together. So, we
actually started building this project
in depth um yeah about a year and a half
ago. And the final thing that you can
see here um is is the outcome. So, this
database is to the best of our knowledge
the largest database of climate science
literature um ever put together. It
includes 1.8 million individual
publications which are linked together
by 40 million citation um and I'll go
into a bit more uh in the course of this
presentation. But I'm going to kick this
off actually with a video that was
produced by Carbon Brief's visuals team.
Um I was trying to work on a little um a
little summary myself and I realized
that actually the video they've made is
just so much better. So I'm just going
to play that for you guys first. This is
Project Cosmos, Carbon Brief's universe
of 1.8 million climate science papers.
This major new resource allows experts
to analyze the entire network of human
knowledge about climate change. So, what
does this database show and how did
Carbon Reef create it? Well, it all
starts with the Intergovernmental Panel
on Climate Change, the IPCC.
Ever since 1990, the IPCC has been
publishing the world's most
authoritative summaries on the latest
climate science. Since these reports
were first published, humanity's
knowledge about climate change has
ballooned. The IPCC has published six
sets of reports so far, each one longer
than the last. Each report contains a
list of references, the academic papers,
books, and reports that the authors draw
on. This list is crucial for
underpinning the legitimacy of the final
document. In total, the IPCC references
more than 100,000 academic papers,
books, and reports. This forms the core
of our climate science universe. Harbon
brief has built on this core by looking
at four other sources of data. the
references contained within those
references, publications that themselves
reference the IPCC,
studies from key climate focused
academic journals, and the publications
that those studies themselves reference.
Every single publication in the Cosmos
database is linked to at least one other
through references. Draw together those
links and this is what you see. Each
column and cluster reveals different
topics and densities of research. Carbon
Brief has now used this vast database to
rank the most highly cited publications,
authors, and institutions. Check out
Carbon Brief's Cosmos 500 ranking to see
the results. Okay, so that was that was
the very snazzy video produced by the
visuals team that I think gives you a
bit of a background about the project.
But what I'll now do for the next 20
minutes is just dig into that in a bit
more detail because I think there was
quite a lot of it you guys quite
quickly. Um, so let's start with how we
actually went about forming this
database. So the first big problem, the
reason it's taken us so long to come up
with this idea to form this database is
that we kind of were stuck for a really
long time on the very first question of
what is climate change research. So we
knew or Leo knew that we wanted to
visualize all of climate change
research. But how do you actually define
what a climate change paper is? There
are some papers which are very obviously
to do with climate change. And so this
is an example here. One of the very
first scientific studies to connect
atmospheric CO2 to a rise in global
temperature was published in 1856
by Ununice Newton foot. Very cool. Very
clearly a paper about climate change.
These these papers that are the about
sort of the very core foundations of
what climate science is that came out in
sort of the 1800s 1900s. It's quite
clear that those about climate science.
But since then, there's been this
massive ballooning of literature about
climate change. Well, a ballooning of
literature about everything really, but
it's actually very difficult to define
exactly what a climate paper is. So,
there have been a lot of different
attempts by various groups. One way, for
example, is to use keyword searching. So
to put a massive massive list of key
words, for example, carbon dioxide or
global warming or or atmosphere,
whatever, put a whole bunch of keywords
um into some kind of program and use
that to sift through something like
Google Scholar or Open Alex, one of
these big big databases of climate
research, and just pull out every single
paper that has those keywords in their
abstract or in their headline. Um this
isn't a foolproof method, obviously.
there's there's a lot of scope for you
to pick up papers that you don't want or
to to accidentally leave out papers that
you do want. But but this is a method
that defines
peer-reviewed studies that have been
published using this method. Um it's a
very valid method. Um other people
especially more recently have used
artists they've trained an AI model on
water climate. Um and again this is a
very valid method that people have used.
There are pros and cons to it. Um but we
didn't feel that carbon brief that this
approach was quite right for us. um we
decided that we wanted sort of take a
very completist approach and so what we
decided to do was to base everything on
what can be absolutely certain is
climate science and what we decided to
look at was the intergovernmental panel
on climate change the IP so the IPCC has
been publishing some of the world's most
authoritative summaries about the latest
climate science um they published their
very first reports in 1990 and this is
kind of the gold standard of climate
science I mean I'm sure a lot of you are
climate change scientists so I don't
need to tell you this but IPCC reports
are the absolute gold standard. Um they
are the most complete assessments of
climate change research. Um and so far
the IPCC has published six um assessment
cycles. So six sets of reports. Um each
one takes thousands of scientists I
would say hundreds of thousands of hours
to pull together. Every line in a lot of
these reports is is sort of worked over
and checked and agreed by many um and
and the kind of final summary government
and it's approved line by line. So these
reports really are um cornerstones of
climate change science and between these
reports um there are there are hundreds
of thousands maybe even millions of no
hundreds of thousands um of other
academic referenced um and so we decided
to use this we decided to use these IP
core of our data. Um so you can see that
the number of references in IPCC reports
has increased over time pretty
dramatically. Um so IPCC reports are
always published in working group sets.
So working group one is the physical
science basis. Working group two is
impacts adaptation and vulnerability and
then working group three is about the
and then attached to each um also a
group of special reports. So for example
a special report on 1.5° C or in the
upcoming AR7 assessment cycle there's
going to be a climate. Um, and so yeah,
each of these reports has a bunch of
references, like tens of thousands of
references. Um, and the number of
references in these reports has just
grown as the length of the reports. Um,
so let's just uh define some terminology
really quickly. What do I mean actually
when I say reference? Again, I'm sorry
for a lot of you if this is quite basic
stuff that you know already. I'm sure a
lot of you academics know what a
reference is, but bear with me because
there is some some terminology stuff
that I want to just nail down here. So,
I'm just going to very briefly talk
through what a study is and what a
reference is. Um, again, bear with me if
you know this. This will just be quick,
but an academic study is um a document
which sets out a new scientific
hypothesis. It could be an experiment
that someone's done. Um, and you start
at the very top with your title and your
DOI. So, the DOI is like this unique.
Underneath that, you will list all of
the authors of the paper. So, these are
the people who to making this paper,
this document. Um, and so it could be
the people who conducted the science. It
could be the people who own the lab. It
could be the people who wrote the paper.
But there'll be a list of the authors.
And then at the very bottom of the paper
will be the references. So references
are again a cornerstone of how science
is conducted. Every single time that um
experts write a paper, they need to
substantiate their claims. They need to
back up what they're saying. And to do
this, they will link to past research.
So within the body of the um academic
paper, there'll be lots of references,
lots of numbers. Uh so for example the
the paper might say the planet has
warmed by 1.2° C number and that number
will link down to a reference at the
bottom and that reference is another
paper or book or report which shows that
the planet has warmed by um it is really
really important to do this like you
wouldn't get a paper published without a
good set of ref. Um so in a study at the
bottom you have the references um as you
can see the blue arrow is pointing to
like the references the other studies.
So that's our sort of carbon brief
terminology. We're calling the list at
the bottom um of the paper that links to
a bunch of other papers. We're calling
those papers that then we also have what
we're calling citers. So citers are
papers that themselves reference a
study. So let's say that this study in
the middle is study A. Any papers that
in their list of references at the
bottom link to study A, we will call
citers of study A. Meanwhile, any papers
at the bottom of study A uh in in the
list of references, we will call
references of study A. So, this is our
our references and citers terminology.
Um, someone jump in if that was not
clear, but hopefully hopefully it is
because it's about to get more um so
carbon brief built Cosmos database on
these references and citers
relationships. So, uh you saw this
graphic very briefly in the video from
before, but I'll just go through it
again. So in the center of our
concentric circles over here we have the
IPCC reports.
Um so these have in total um 107,000
references. So that means uh in the list
of references at the bottom of the IP um
there are in total if you look at all
the IPCC reports there are 107,000
different academic studies linked there.
Um and that is your your like first
circle in the in that diagram there. Um,
I should note that what we're counting
here is just the unique references. So,
obviously an IPCC report or multiple
different IPCC reports might all
reference the same paper. So, here there
are 107,000
unique studies referenced, but in total
there are 162,000 just so you know. So,
there's a much bigger number if you're
including duplicates. But anyway, so
that's that's your that's your
references. That's our first layer of
the day. We then added another layer
onto that. We took every single one of
those IPCC references and we found all
of their references. So we're calling
this second order references. So you can
imagine the scale here is just
ballooning very very quickly. Um so yeah
1.4 million second order reference. We
then also included IPCC citers. So that
is academic studies that themselves
reference IPCC reports or chapters. Um
168,000 of those. And then we decided to
just be a bit more um completist about
this. Um we decided to add another
another lay another step onto this
analysis. And so we identified 22 key
climate focused journals. So these are
journals for example like nature climate
change which really do just focus on
climate change. So here we didn't
include big journals like like nature or
science that cover everything in in in
sort of academia. We we just focused on
the climate change specific journals.
and we had some of our carbon brief
scientist friends help us with
identifying those journals. So we
identified 22 journals and we looked at
all of the studies that had ever been
published in those 22 journals. So that
was about 50,000 studies and then we
looked at all of the studies that um all
the studies, books and reports
referenced by those papers. So that was
642,000 of those. Um so if you pull all
of those together that you get 1.8 8
million um papers and books and reports.
Um and that is that is essentially what
this database is. This database is is
that compilation of of papers and books
and reports. Um and so yeah, that that
that's basically how we pulled together
database. Um right, I say we here, I
mainly mean Tom Pearson and then um we
had some fantastic collaborators who I
will who I will thank formally at the
end, but like Felix and Tristan and
Travis Cohen from the University of
Exit. did a lot of work on pulling this
database together and if you have more
technical questions about how that
worked, you can grill Tom about them at
the end. Um because he was really deeply
involved in like pulling together this
massive massive database of papers. Um
so yeah, it's really cool. It was, I
would say, a pretty big technical
challenge to get this much data all
organized. And we spent a lot of time
kind of, well, we again, Tom spent a lot
of time kind of uh debugging and
cleaning and and getting everything to
the right format. And yeah, again, if
you have more technical details about
how we did this, please do ask Tom at
the end. Um, because I want to get on to
the bit that I think is more exciting,
which is the results. So, there is a lot
you can do when you have 1.8 million
papers. There's a lot of analysis you
can do and we have so many ideas for
different types of analysis that we want
to do in the future. Um, but for now
we've just decided to kind of just just
to show in principle what could be done,
we decided to do some initial analysis
and what we came up with was this idea
of the Cosmos 500. So this is basically
a ranking of the top 500 scientists,
institutions and um uh what else was it
uh countries um in the database. So we
when I say top actually I'll describe
what I mean when I say top in a minute
but it's basically who who was cited
most highly who was like included in the
database most um so publications that
are referenced really extensively
they're often referred to as highly
cited. Um and in academia if a paper is
highly cited i.e. lots of other people
have referenced that study. It generally
is a sign that the paper is kind of well
has important findings or is very
foundational or or maybe it's a it's a
data set that lots of other people are
using. But but generally you can say
that if something is highly cited, it's
sort of high quality science that's very
uh and so we thought it would be
interesting to see which which
scientists have published the most
highly cited papers uh and which papers
kind of come out top of our ranking most
highly cited. Um, and so what we did was
we defined this metric called a citation
score. And so this is the number of
times that each academic paper, book or
report is referenced by others within
the Cosmos database. So very clear here,
this is just showing how highly cited
something is within our database. So you
could have a paper that say a statistics
paper that is really influential in in
some mathematical field and has been
hugely highly cited by other
mathematicians, but it might not rank
very highly in the Cosmos database.
Equally, you have might might have
something that ranks really really high
in the Cosmos database that when
compared with broad academic literature
is actually not that influential, but
like because it's a core foundational
climate science paper, it ranks really
highly here. So yeah, just to be very
clear, this is a ranking for the Cosmos
database only. Um, and in these
rankings, we have omitted the actual
IPCC reports and chapters because those
obviously would have just been like the
highest ranking. They would have like
topped every chart and it would have
thrown everything else out of whack. Um,
so that's just the preamble and now we
can get to the exciting. Um, so first
up, authors. So again, this citation
count is how many times each expert's
publications are referenced by other
publications within the Cosmos database.
Um, we we pulled the top 500. So if you
go on to the Carbon Reef website, uh,
the Cosmos database website, you can see
all 500. Um, but here I've just shown
you the top 20. Um, so some really
well-renowned um, climate scientists.
Um, and Carbon Reef actually had a
chance, well, I actually had the chance
to chat to quite a few of these people
to pull together a piece about their
achievements and their work they've been
doing. It's really cool stuff. Um, I
will note that um, there's not a lot of
diversity. Almost half of the authors in
the Cosmos 500 are actually from the US
and The Global South only accounts for
4% of the authors. um only 10% of the
authors are women and actually you have
to go down to ranking number 36 to even
get the first woman in the Cosmos 500
that the first 35 are all men. So um
this is I I mean it's a reflection of
how academia and then climate science
included is quite dominated by men from
the global north um which I won't delve
into now but that's a topic that I have
spent a lot of time working on in the
past. So if anybody wants to ask me
about that later you are also welcome to
do that. Um but yeah, so this this was
the ranking for the Cosmos 500 for the
top 20 authors. Um and Carbon Brief
actually um interviewed the top authors.
So this carbon brief did an interview
with Philip Keys and with Professor Dler
Vanurren. Um so Phipe was the top author
in the Cosmos 500 and Professor Van
Beern was oh I forgot to mention this.
Okay, so carbon brief we kind of have
two versions of the database. We have
the main Cosmos database which is
everything and then we have the IPCCon
database which is just those 107,000
papers that were referenced by IPCC
reports. Uh and so when you're looking
at the the results you can kind of
toggle between those two databases. Um
so what we have here is the top two
authors from those two separate
databases. So um yeah the Cosmos
database is the main one with everything
all of those steps I talked about. But
we did also think it would be
interesting just to see who is the most
highly referenced scientist just within
IPCC reports. You know, there's really
really foundational climate science
reports. So that's what we have here. Um
so we also did publications. Um so the
top publication actually wasn't
something particularly climate sciency.
It's like kind of a stats u mathematical
manual thing. But maybe that's not
surprising. It's quite a foundational um
piece of literature that clearly lots of
people have referenced. Um but yes, my
colleague Rob actually did this part of
the analysis. Um so we found that maybe
unsurprisingly, nature and science are
the top um climate uh journals that the
studies appeared in from the Cosmos 500.
And you can kind of see this like
increase in studies um going up to the
2000s. The most most of the studies in
the Cosmos 500 were published in the
2000s. Um, we'll have to see as time
goes on whether this is just because
studies in the 2010s and 2020s haven't
had enough time to be referenced by
other papers or whether there was like a
weird reason for some spike in the
2000s. We'll have to we'll have to see
that um like as time goes on and as we
keep updating progresses but again you
can have a look at which publications
were the most highly cited which I think
is pretty cool. Um and then institutions
as well. We had a look at which
institutions are publishing
um are publishing the most highly cited
papers and the way we did this was by
linking authors to the institution that
they were based at at the time that they
wrote that paper. So if someone was from
um Colombia University at the time that
they wrote a specific paper that was
very highly cited then they're citing um
even if like now they are working
somewhere else or retired or something a
whole complicated system of like trying
to deal with figuring out how to link
authors. But again you can kind of see
this this global north bias again in
which institutions rank most highly in
the Cosmos 500. So yeah, if you take the
top 500 institutions, the list is
dominated by the US. You then have the
UK, Australia, Germany. I think the
first global south country is China um
16 which is I mean not bad but I mean if
you compare it to the US of 183 uh
institutions you really can see how
crucial the US is for climate science
research which I mean not to go into too
much detail but given how the Trump
administration is really gutting climate
science in the US right now I think this
is pretty significant. Um yeah I or
it'll make me sad but yeah the US very
important for climate science. We'll
have to see whether that changes. So,
um, but yeah. Okay. So, last thing I
wanted Oh. Um, okay. Well, what I wanted
to show you was that that Tom has made a
really cool like interactive visual
visualization page thing to show.
Actually, maybe I can show you. I think
uh give me two seconds, folks. Sorry.
Okay. What I wanted to show was cuz I've
been signed out. What I wanted to show
you is this was the map. And I'll just
do it live actually. Okay. So, the final
thing I wanted to show you is how you
can play with this database yourself
because it's a really really cool
resource. And I had made a little video
where I talk through how you can do
that, but clearly clearly Google has
decided to sign me out right now, which
is terrible, terrible timing. Um, so I'm
just going to show you live instead. So,
the database is actually interactive.
It's clickable. You can go in there and
you can sort of have a look around it
yourself. Um, and so this is how you can
do it. You can go to the the Carbon
Briefcosmos uh page. That's
interactive.org/cos org/cosmos/
and then if you go onto the map you can
have a little scroll. So what you can
see here is okay we have a bit of
explanation that every single star in
this database represents one of the
studies in the cosmos database and that
the colors illustrate clusters um of
similar topic areas. Uh so they show
like which academic areas are located in
our within our map. Um so for example
these two large circled areas represent
physical science and social science. So
you can see the physical science the
blue is a pretty is a much bigger
cluster than the social sciences which
is in yellow but both and then you can
go in even deeper. So um yeah so the
green and the red areas capture the
fields of immunology and microbiology
for example. Um, this is really cool
that you can see the coloring. And
again, maybe Tom can talk to anyone
who's interested about how that coloring
uh worked, like how those topic clusters
worked. Um, but it's really cool that
you can actually see them grouping
together like this. Um, yep. So,
okay. Yeah, this is the thing I wanted
to show you. This is the thing I thought
was the most cool was that, um, you can
see that some of the circles, um, 500 of
them to be precise, are a bit bigger. um
those are the the studies that are
actually within the Cosmos 500 ranking
of the most highly cited publications.
So what you can do is you can go and
zoom into them. So for example, this one
here that's been highlighted is the
Stern review um which was uh the most
highly cited publication across the IPCC
only data and zoom. So yeah, the whole
thing is like clickable and movable
which is I mean so cool. Um, so you can
see that down here at the bottom you
have all of the different um all of the
different categories. So for example,
yellow is social sciences here. So you
can see psychology is in orange for
example. You have physical sciences is
in blue over on the on the left hand
side at the bottom of my screen. So
environmental science is light blue. You
have quite a lot of those. Um I'm going
to click on okay. So like this one is
green for example which means it's in
health science. And so I can click on it
and we can see that it is um discussion
reporting of 14C data. Okay. I don't
know quite what that means, but that's a
that's a paper that was highly cited.
See, if I'd had my video, I picked a
really nice one. I just can't remember
Galaxy, but anyway, you can click around
and you can see what the papers are. Um
you can see um what the ranking is.
Okay, so this is this is um from the
bulletin of the American Meteorological
Society um representative of a climate
science study. But anyway, you can have
a click, you can have a play, you can
see what different papers are there. You
can see which ones are clustered for
example together. um that's about um out
of uh social ecological resilience in
this together. But anyway, you can have
a play. I just think it's very very cool
that with this database you can have a
Oh, I've been signed out. Okay, [snorts]
I had one more slide. Uh oh, it's making
me create. Okay, I'm just going to add
my last slide. I remember what it said.
Um or maybe I can stop sharing and Simon
can bring up the slides and just Yeah,
I'll stop sharing and maybe Simon can
just bring up that last slide for me. um
if that's um but I will yeah so that
last slide basically what I wanted to
say is that the Cosmos database is um
this is the very initial kind of early
phases of the Cosmos database we have
done this as a kind of proof of concept
we've developed the tool and now we want
people to use it so we've done our
Cosmos 500 ranking and we're going to
update this every year we're going to
update the database every single year
hopefully that's the plan uh to add all
of the new climate literature that comes
in and then you can expect a really big
bump in the database once the seventh
assessment recycle seventh assessment
cycle reports. You could expect a really
big bump then. But in the meantime,
there's a ton of other stuff we think we
could do. There's kind of temporal
analysis stuff we think we could do.
There's keyword search analysis. We have
a ton of new ideas. And what we would
really love is for academics is for
people in the scientific community to
get in contact with us if you have ideas
for the database. We deliberately
haven't made it open source. in part
because we're a bit worried about AI
bots scraping it and in part because
it's just such a massive database that
getting it open source online would be a
nightmare. Um but if academics approach
us with ideas and ways they want to
collaborate and use the database, we
would be really really keen to work with
you guys. So um yeah, I don't know if
Simon is able to share the slides.
>> Oh, so um I also had a problem of being
kicked out and accessing it. So,
apologies. I could not step in and help
even though it was supposed to be
>> um backup. Um
>> I never had this specific technical
issue on a presentation, but you know
what? At least it happened near the end
of the presentation. So,
>> um I think I showed you when I was
scrolling where you can access the
Cosmos database and I'll put a link in
the chat to where you can access the
Cosmos database as well. Um if you go on
there, there is an email address where
you can email carbon brief to tell us if
you have an interesting proposal project
idea. Um, and yeah, I think it would be
really cool to hear from some of you
guys if you have things you want to do.
So, thank you so much and yeah, maybe
we'll take questions now.
>> Yeah, thanks. So, thanks for that uh
presentation. I think we'll bring in Tom
now as well who also discussed his
perspective on developing the database
as well. Um, I just quickly kind of
clarify um partly because I miss this
myself when I was trying to understand
why I've been kicked out of Google.
Could you just quickly say in what ways
perhaps our geoscientist audience might
be able to contribute to this database
for example for example it's not
directly by adding papers um or it's not
open source either um in what ways could
they perhaps uh contributions to a
database
>> yeah so hello um yeah so so basically
we're looking for kind of collaborators
essentially we've got this database
we're hosting it sort of you know in
private at the moment
Um but yeah, we're looking for people
who have ideas about what we could use
this kind of vast corpus of of kind of
interlink documents to uh to discover. I
mean, we have a few ideas like for where
we want to take it for the for the next
kind of, you know, ne for next year's
release, but um but yeah, we're
particularly interested in in sort of
reaching out to the community to sort of
ask about uh about what other people
find interesting cuz yeah, it's a it's a
big resource.
>> Well, thanks. I guess that um also kind
of leads to other question is like what
um I guess is the intention you had in
terms of its use, I guess. Um there's
obviously the visualization of all these
papers etc. Um I guess it's searchable,
findable. So at the base level it's also
just to help people kind of find um
research and see how it whatever
research might be linked to it for
example.
>> Yeah. So to kind of go into a bit of
detail about the sort of structure of
the database. So the initial
the sort of initial piece of work was uh
commissioned like a couple of years ago.
Um we worked with a a French PhD student
Felix who basically uh wrote the
software to scrape various different
sort of open repositories. So like Open
Alex, Google Scholar and a few others.
Um and sort of built up this initial
kind of unstructured database which we
then kind of converted into uh a kind of
a a more formal structured database via
the process of kind of cleaning
everything up, working out when two
papers that are described slightly
differently in different repositories
are the same thing. So we dduplicated
stuff and we sort of formalized the
relationships between them. Um this was
all really kind of like with a view of
this kind of getting a kind of like
broad picture of climate science. So uh
what I should described as kind of like
Leo's kind of headline vision of this
kind of like big big kind of big picture
thing uh and the and the sort of like uh
and and the ranking. But I think what
we've kind of realized since um since
compiling this database is actually a
lot of the kind of interesting
stories and interesting features of the
database happen at not the kind of
global scale but the smaller scale where
you see little clusters of work on kind
of immunology or on um uh sort of
disease resilience or or adaptation and
it's interesting as well I think and and
one area that we do want to explore is
how these areas develop over time. Like
how we kind of did a did a kind of quick
sort of time series test to see how the
the sort of literature from the first
assessment report varied from the the
sixth and how over time like the social
sciences have formed a much bigger part
of of those reports than than they did
initially. And so seeing that kind of
information and using that to to kind of
like find out about how
knowledge is produced and how it
develops um is something we're really
keen to do over the next over the next
um you know 6 months to a year.
>> Sure. I see. Yeah. Um actually you
talked about um kind of grouping on a
broad idea of climate science and trying
to revisit that through the database.
Um I I it links to like one of the
audience questions we have actually as
well about how these groups are kind of
defined etc. Um so one of the audience
members uh said that some people in the
top list would not see themselves as
climate scientists at all. But others uh
might see this as um I suppose a kind of
a broader picture of like uh research
speaks to other research is it how can
you draw draw hard boundary between um
uh certain kind of fields etc. Um I
guess and the question kind of links up
to uh one of the perhaps uh in some
areas like subsets of the data set might
be better for some questions about a
broader lookout better for others etc.
Um I was wondering if you could uh talk
more about how you then um define that
because I guess people obviously
identify themselves different ways as
research science but for us um you
applied it another way.
>> Yeah. So as Aisha said this is like kind
of probably the broadest way in which
you can kind of conceive of like you
know climate science like you know it
sort of extends to as as far out as you
can go and still be thinking of this as
kind of like knowledge about climate. Um
how we actually came up with the the the
sort of categorization it was derived
from open Alex. So open Alex the sort of
big online sort of open- source I think
open source open data um repository for
for papers has a
process by which they you know do some
sort of automated analysis on the text
of of these documents and basically come
up with a a set of um categories that
fall within this kind of like taxonomy
where you've got like the top level
which I think they call domain and then
like topics and then subtopics and we
kind of went down we kind of queried
Open Alex down to the sort of subtopic
level. Um [clears throat] and then sort
of once we had that we realized that was
far too granular. So we sort of took a
step back and and looked at it at the
topic level which is the kind of level
at which you have things like computer
science, uh material science, chemistry,
um psychology that kind of level of of
granularity. Uh, and then we just used
the primary topic for each paper as as
our way of kind of coloring it in.
Essentially, like all of that
information still exists in the
database. So, if we wanted to, we could
sort of like look at a different way of
of categorizing things maybe either at a
more granular level or kind of like
waiting between the topics. Or if we
wanted to look specifically at a
particular area of knowledge, we might
choose say um neuroscience and say just
pull out all of the papers that are
neuroscience that have neuroscience as
their topic and show us the connections
between those or the citation
relationships or which authors are
prominent in that corpus. So like we
didn't make any kind of editorial
decisions about what cons constitutes a
paper and we tried to sort of cast our
net as broadly as possible in in terms
of what we include. Um so yeah it's
absolutely right that like a lot of
these papers aren't what you'd
traditionally consider climate science
>> but I mean that's a key point of trying
to relate um how research impacts other
things right we we know how climate
change impacts health so health papers
being included in our medical papers
makes sense right and trying to get a
broad picture so I usually want to say
something
>> sorry no absolutely just to I mean
everything that Tom said absolutely true
and this I guess it's just to add that
you don't have to be a climate scientist
for your yeah for your work to be very
influential in climate science But um
yeah, I guess as Tom said, it was it was
an editorial decision and we were
thinking initially about whether we
should try to narrow things down a bit
more and we we've made the decision
deliberately to be very broad scope. And
I think that is actually really
interesting is when you can see things
like like the top paper being a
essentially a reference manual for how
to use the coding language R being the
top paper. But that I think is really
interesting because it shows that that's
such a foundational piece of work. Those
experts are probably not climate
scientists. they wouldn't consider
themselves climate scientists but now
they can see that their work has
actually had a really huge foundational
impact on climate science going back
decades which I don't know hopefully
they feel good about
>> so when people asking about projects
this kind of um input in terms of
granularity etc might be a way people
could increase its use for example I was
thinking sorting the uh database by
phenomenon I don't know by natural
hazard or something or perhaps by a
method is uh is perhaps a way to do it
perhaps um I guess when you're
developing these links um it's basically
based on a broad topic of the paper
right it's not on the these more
granular
>> options yeah so we don't um we don't
have the well the our colleagues uh
Tristan and Travis at exit are currently
looking at the idea of downloading and
sort of analyzing a lot of these papers
on a sort of specific level but we're
Yeah, we're just using the kind of like
metadata that's stored in these various
repositories. We're actually not looking
at the papers. Like in some cases, we've
looked at the abstracts, but mostly it's
just work that other people have done
that we're building on.
>> I see. So, it's connecting to other
databases then. Sure. Um like what's the
relationship then as well? For me, it's
more about um
uh I kind of question resilience as
well. perhaps as um as known as climate
data databases being removed um due to
kind of government action etc like that.
Um is that something you consider
perhaps when looking at these
[clears throat] relationships at all or
um I guess also like I creating this uh
database um where it's bas like can be
fed by other databases it's kind of it
independent um mapping that use.
>> Yeah. So like kind of I think um I think
there's like the the sort of big big
data sources that we use like open Alex
and Google Scholar. I think there's
definitely question around questions
around though the resilience of those
sources. I think open Alex seems like a
a kind of good bet to hang around for a
while. I think Google Scholar is on um
you know perhaps somewhat more shaky
ground given Google's past record and I
think we've been hearing recently. But
um but like a lot of these data like a
lot of these databases are themselves
interrelated. So open Alex populates a
lot of its stuff and kind of like
through Google Scholar or or other um
other ones. I keep on mentioning those
two but there I think there was four
main main databases which we which are
scraped. Um but yeah in terms of like
how the database that we've created sort
of persists and it you know we're sort
of hosting it on a private server at the
moment. it's like backed up. But we want
to definitely we're definitely sort of
investigating ways in which we can sort
of make this more open and make it more
sort of we so we can kind of you know
not act as the gatekeepers for it in
quite such a a sort of strict way as we
are at the moment. We sort of want
people to be able to access it but
there's sort of yeah cost time and
expert implications associated with all
of that which we haven't yet worked
through.
of course. Um yeah, like you'll give me
flashbacks to my own work when I was a
researcher as well. Um yeah um one one
thing that came up though actually I
think this is a question for you
actually is uh we talked about these
other databases but you also talked
about um um kind of the perhaps
historical impact of representation on
on who gets cited and the kind of
diversity of people represented in the
citations. Uh Bisha, you also have like
uh there's another database that common
brief exists as well to help kind of
combat I guess bias against certain
authors. Perhaps this is a nice time to
kind of just briefly highlight that at
all.
>> Yeah, thanks Simon. So yes, I've done a
lot of work over the past five six years
about the lack of diversity in climate
science literature. Um so I did some
analysis in 2021 now how time flies. uh
showing basically that women and experts
from the global south are hugely
underrepresented list of some of the
most highly cited climate science
research and that I think is borne out
at a much much bigger scale by the
cosmos database here. I mean yeah you
can see it's very consistent across that
front. Um and that's for a range of
reasons to do with barriers in
conducting the research in the first
place barriers in getting research
published. There's I mean there's a lot
of stuff you could look into for that.
Um, yeah, I won't get into all of the
details now, but yeah, a few years ago,
um, I helped to develop, uh, another
database called the Global South Climate
Database that Tom actually was very key
in helping to to put together as well.
And man, it's great. Um, this is a
database of climate science experts from
the global south. So, experts add
themselves to this database. It's a it's
a self-submission thing. So, every
expert on this database wants to be
there. They want to be contacted. Um,
they want their their voice to be heard.
And I can see Simon, you've just put the
link in the chat. So, thank you.
Appreciate that. Uh, we now have more
than 10,000 experts on the database. So,
I think it is uh we want to well, for
example, as a journalist, if you want to
find more experts from the global south
to quote in your reporting, or as a
scientist, if you want to find peer
reviewers, I've heard of quite a few
reviewers. Um, and I guess the core
point I want to make is that diversity
isn't just a this isn't um this is, you
know, genuinely important. getting
different research brings different
perspectives that that genuinely do help
to progress science. This isn't just
just a oh, what's the word I use?
Charity case or something like this is
genuinely very important. It's an issue
we should all care about. So, I think
there's definitely some really
interesting analysis we could do on the
Cosmos database more linked to this um
where people are publishing from, who
who is publishing and and maybe
hopefully how that has changed or
improved over time. That's a bit of
analysis I would really love to do.
We'll see if that bears out. But
>> thank you. So of course this it's not
the only database and if you want to
start engaging with people um uh looking
where these connections are in project
costs is good but always good to keep an
idea of these historical systemic um
biases or problems which otherwise
result in kind of like quite a a poor
representation in um science or even
even in kind of climate change research
group. Uh so the global self climate
database is one way and thought looking
out for scientists who aren't well
presenters. Um we did also get a
question about uh the um dates of some
of the publications. Uh one of the
attendees said that the top 500
um I guess has a lot of articles which
are pre20
I guess not many after 2010. um and
particularly is asking about finding
papers that include technology. Is there
a way and there database and also um why
the database might be looking at more or
I guess preference for historic papers
or publications. [clears throat]
It's two questions in one. Apologies.
>> Yeah, I'll I'll answer the one about the
um the age of the papers. I think that's
just um because we're because we query
by most citations like older papers are
kind of inevitably going to have more
citations cuz they've had longer
[clears throat] to gather them. Um one
thing that we do want to sort of look at
is
uh sort of balancing this a bit and have
a kind of perhaps a sort of like you
know you know either sort of timebound
query. So we say like kind of what what
if we only asked for papers for the last
10 years or maybe what are the papers
that have acred
um acred citations most rapidly over
over the last 10 years. So that's all
kind of stuff that we can we can look at
in order to kind of come up with
different rankings. I think at the
moment like in some ways like the the
the sort of huge bulk of the work here
was kind of collecting and cleaning this
data and putting it together into a into
a kind of coherent single database. Um
and then we kind of like at that point
we were like okay we'll publish this
with kind of like the most easily
understood way of of kind of creating
this like list of 500 papers or people
or whatever. But I think definitely
going forward we want to look at ways we
can sort of yeah surface more recent
papers surface like different kinds of
contribution and and different kinds of
uh different kinds of paper. Um what was
the other question? I've forgotten
already.
>> As about um how to find um perhaps
article articles and new technologies
for example. Um
>> yeah I mean that's a good question. I
guess you'd probably want to look at um
I'd probably go in via the sort of open
Alex topic hierarchy and like find a
kind of subtopic that sort of matched a
particular kind of thing. So like
presumably something inside the
engineering or computer science topics
there'd be you know more more kind of um
granularity in there and then and then
looking that way through it. But it's
not really because we're not doing any
kind of text on
uh any sort of text searching isn't
possible. So, so that would need a kind
of different a different kind of
database. There's actually a a sort of
second database that sits behind the the
main kind of Cosmos database which is a
a kind of unstructured document store
which has a lot more um a lot more kind
of random information in there. So like
papers in that often have abstracts
which would be a good place to do a text
search for kind of you know new
technologies or whatever. Um so like the
Cosmos database itself is probably not a
great uh a great way into that kind of
stuff unless you're going in through
those topics and subtopics.
>> So the way the cos cosmos project cosmos
can inform your search for other
databases as well actually in that
sense.
>> Yeah. Yeah. Yeah. And I think as well if
you've kind of like got a got a
particular paper that's a way in like
you can obviously do a query which is
like show me all of the I've got this
paper um the title is this show me all
the papers within the database that
reference that paper and all the dat all
the papers within the database that that
paper references and then you kind of
like you know come back with like 100
papers and that will give you a a kind
of way to look in and then you can make
similar queries on those papers. So you
kind of you know the database is a web
essentially and and you're kind of like
crawling around it finding uh finding
the interesting things that way rather
than uh like quering it like a more
traditional database.
>> Sure. Yeah. I I guess then a project
perhaps that could be brought forward
then um we've talked about gridarity in
terms of topic theme and subject but
also um yeah as you said looking at um
things that are upcoming rising. So the
thing that came to my head were Reddit
filters which kind of filter by like a
most recently interacted with high level
of interaction between certain bimes etc
etc. Um so in which case if you want to
look at uh popular papers in the last
five years or which are getting
attention for some reason that might be
another a project people could borrow.
Uh yeah. Yes.
>> I mean because it's I think it's
interesting that you mentioned that
because it is it is a graph database
which is essentially the kind of same
sort of database that backs something
like Facebook or a social network. It's
about
>> like it's a it's a kind of database that
privileges the relationships between
these nodes as much as they're like
first class citizens as well as the
documents themselves. So, so like things
that are about like relationships
between papers or between disciplines or
between authors like this database is is
kind of made for that kind of uh
research.
>> Excellent. Um I have a um I've got
another question um which is about um
ranking the most cited data sets um uh
as [clears throat] a method to try and
identify
um foundational climate data sets or
data centers. Um it's been interested to
some other uh geoscience communities.
GCS has been one of them. I was
wondering if that's uh something you've
uh included in in uh Cosbox. So ranking
the most cited data set.
>> Uh no we haven't. But that does sound
like a really interesting uh interesting
thing to think about.
>> But there you go. So I mean if you have
these ideas please submit them to cover
brief. [snorts] Um so it's cosmos.org.
So these are things about granularity
looking at pubshed method of phenomenon
looking at um most cited data sets um
perhaps searching by rising or most
recently
um popular citations etc. I think these
I guess are the things that you might
want people to bring forward to you as a
project or idea.
>> Yeah.
>> Yeah. Absolutely.
>> Absolutely. Um we're running out of time
very quickly. Um I just want to like ask
perhaps one final question. Um um well
actually before that there's one thing
um about information available. If
people look at the data set and
searching for things it's quite easy
then to kind of move on to like find a
paper etc. site it read it. Um I guess I
was also thinking about in terms of
contacting people or um sorting things
by disciplines and fields. Are these uh
possibilities in the data data set? Were
people looking around?
>> Not in the not in the website as it
stands like because of the size of the
database. We did like I originally kind
of tried to like I got it I got the I
got the sort of data payload down to
about 200 megabytes and I was able to
deliver all of the titles and um
>> and and subjects and and publications
for all the papers. But we decided that
like kind of lumping someone with a 200
megabyte download when they visited the
page was probably a bit much. Um and and
sort of splitting it up intelligently
was again it's something we didn't
really have have the sort of time or
resources for. So
>> at the at the moment uh that's not
possible. You can sort of you know
search the tables that are on the
website but um but kind of like going
deeper is is is not um not yet a thing.
I mean, which is a real shame because
one of the things I really enjoyed when
I was like developing the site was being
able to, you know, zoom into that big
cloud and kind of like see the see the
papers that it had like, you know,
clustered together and and, you know,
kind of like really just get get a feel
for the scope of it was great. So, it's
a it's something that I would really
like to do, but um but yeah, it's a it's
a technical problem that needs to be
solved.
>> And we got no foundation down. We got a
massive network of relations between
different types of research but uh isn't
I guess maybe broadly looks at climate
change research that people can access
to look around and also provide their
own thoughts and contributions to uh I'm
going to close the webinar but uh I'm
going to put the floor open to any final
comments or requests um uh from the
panel if there's any kind of one last
thing you want people to keep in mind
about the database or about um
contributing or working with carbon reef
as a
>> um I have I would like to give further
props to Felix Chavelli um Travis Cohen
and uh Kristen Cam for their work on
this because they were like kind of you
know they're they're the main stars of
the show and they're sort of behind the
curtain but I think uh yeah their their
work was absolutely invaluable in in
doing this. It wouldn't have happened
without them at all. Absolutely a lot of
work and effort to produce such a a vast
database. Uh Aisha, do you have any uh
final thoughts or requests? Even
>> Tom absolutely took my answer because
[laughter]
my final slide which never
mean I guess I'll also mention the rest
of the carbon brief team. Then Leo
Hickman, Rob Mcweeny, Cecilia Keiting,
uh Joe Goodman, Tom Pra, like this was a
massive carbon brief wide effort. Um,
and so yeah, very keen to bring in other
collaborators to to continue expanding
the Cosmos team.
>> Brandon, thank you so much. As mentioned
before, if you have an idea about what
perhaps a feature you want to see in
Project Cosmos, how to um make it kind
of more usable or appropriate to your
research as a geoccientist,
um, please send us to Cosmos at.org.
Otherwise, thank you so much for joining
us on this uh, incredibly hot afternoon.
Um and thank you for our speakers
comparison each time for joining us
today. This uh webinar will be on the
EGU YouTube channel in a couple of weeks
once it starts. Otherwise, thank you so
much again and have a lovely day.