Submind YouTube summaries
Thumbnail for How to become a computational social scientist

How to become a computational social scientist

Watch on YouTube

Video summary

Computational social science (CSS) represents a unique interdisciplinary field that combines the "human thinking" of social scientists with the "computer thinking" of computational methods to address complex questions about human behavior. Unlike traditional research that simply applies digital tools, CSS leverages computational elements such as handling massive data volumes, processing high-speed information, and analyzing novel patterns in people's actions. The core of this discipline lies in bridging two distinct skill sets: social scientists excel at identifying nuanced problems and communicating abstract concepts, while computer scientists are stronger in programming and managing vast datasets; therefore, successful CSS projects often rely on collaboration rather than requiring every researcher to master every technical technique alone. The practical execution of a CSS project follows a structured yet non-linear eight-step process that begins with clearly defining the research problem and exploring it through various means like surveys or secondary data analysis. Researchers must then formalize their concepts into terms computers can understand, often using pseudo-code or agent-based modeling to define specific variables and rules. This is followed by the critical stages of data collection and software implementation, where challenges such as format mismatches or access restrictions require careful debugging and ethical consideration, particularly when dealing with non-disclosive data or paywalled sources. While artificial intelligence tools like Large Language Models can assist in drafting code, they are prone to errors in complex tasks, making a hybrid approach that pairs AI assistance with rigorous human verification essential for accuracy. Once the technical groundwork is laid, the focus shifts to running experiments and analyzing data to draw conclusions that can inform policy recommendations or drive societal change. However, the journey does not end there; effective communication of these findings to diverse audiences through journals, blogs, or workshops is a vital final step that must be paired with ensuring reproducibility via open documentation and code repositories. Because the research process is iterative, continuous documentation is necessary to track methodological changes and facilitate ongoing collaboration with computer science experts, allowing researchers to return to earlier steps as needed to refine their approach and address new challenges that arise during the study.
Read the full video transcript
All righty. Um, so hi everyone. My name is Louis Caper and I'm going to be leading today's computational social science introductory workshop. I think uh Beck can pop in the events page into the chat if you want to have a look at any of our other coming events. Uh, please do so. But for now, let's um let's get round to this. Okay. So, um, just a little bit of a disclaimer. So this session is going to be focused on introducing you to the concepts of computational social science. So we're not going to be covering more advanced topics including you know epistemological and theoretical challenges posed by CSS. Nor are we going to be focusing on the ins and outs of one particular computational methodology. So if you are coming here thinking that you know we're going to be getting indepth with one particular method that's not the case. It is more of an overview. Um so just as a heads up um but if you are interested in the workshops and the resources that we have on different computational methods then uh Beth can also post a link to our YouTube CSS playlist and also our GitHub which has a ton of um code repositories which contain interactive coding notebooks, other sorts of materials. We've got a bunch of stuff on text mining um some stuff on machine learning for example. So, you know, if that um sounds like something that you'd be interested in, please check that out. But today, what we're going to be doing is covering um sort of three main sections really, which include um so first, what is um computational social science? So, what does this term even really mean? How do I become a computational social scientist? Um so, what sorts of skills am I going to need to have? And then we'll do um kind of a walk through of the eight steps of a CSS project. Um that's going to be the most interactive part of the workshop where you can get stuck in and sketch out a CSS project for yourself. Um and of course at the end there's going to be time for final thoughts and any questions that you might have. So yeah, what is computational social science then? um it's the use of computational and empirical methods to address social science questions. So if we break this down a bit more um we can think of it as um requiring a sort of human thinking to identify important research questions. Um and because we're dealing with social science questions, we need to understand of course how people behave, what they want, what their motivations are, what they want to achieve. So crucially, we need that social science sort of brain in order to formulate those really interesting research questions, but we're also as well going to need a sprinkle of something different as well, which is that um more computer thinking. So, you know, trying to think about how we translate these murky social science concepts into a more, you know, computational um application. what kind of method are we going to use? So, how how do we turn these questions into computational or empirical methods? And after you complete your research, you're going to need that social science brain again to be able to effectively communicate those results to other people. So, you're going to need that human thinking kind of element. And so, another way to think about computational social science is to think about um what it's not, right? So um it's not just using computers within a social science research project. It's um it's not just using digital versions of purely traditional social science methods. So you know it's not just using SPSS to analyze survey data and it's not just using digital but purely non-empirical methods. So, I'll go ahead and untangle these a little bit because I know they're not always um entirely obvious. So, here we'll go through a few examples that just might help put this into context a little bit. So one example of a computational social science project would be collecting, processing and analyzing millions of online news articles to um show changing political attitudes. So you can see how we have that human thinking here. So we um use that in order to formulate our social science research question because we're looking at political attitudes. But we also have that um computer thinking inherent in using a computational method. So in this case if we wanted to collect millions of online news articles then we'd be doing that um using a computational method called web scraping and that's how we'd gather such a massive amount. Okay. So you might think this kind of human thinking um computer thinking um you know binary that I've got here is um is kind of like crude but it's just to show the difference in the two really key elements that you need. Um so we could also use real time weather and traffic data to look at how travelers react to particular events. So, for example, maybe you hear that there's a new storm that has caused damage to a local town and then you want to look at how people react and how this event is dealt with in real time. So, you can see how this project, you know, it's not just a case of using a computer within a social science project or using a digital version of a traditional social science method. Instead, we have these uniquely computational elements, right? So looking at that real time weather and traffic data. Um we have some other examples too. So um you know I'm sure a lot of you now know that um or maybe you even have them yourself. These um sort of smart watches um these new novel wearables that can you know um track your heart rate and other interesting health data that kind of thing. Um, an example could be combining data from those kind of weather robels or apps to establish correlation between, you know, social media activity and heart rate. Um, you could be looking at social science questions about how people feel about certain images. So, you could look at how they react. Are there positive or negative feelings? Are they irritated? And so on. Um, that kind of thing. Finally, we could be interested in mapping family names over time by importing, processing, and formatting centuries of parish records. So, that could allow you to explore the movement of certain families to different areas. Maybe you could um explore the changes in family sizes or whether they've moved away or not. So these are just a few kind of broad but really key um computational social science examples. So yeah uh moving on there are a few factors that make a CSS project computational and I guess the first um thing I think of here is um the data volume, complexity, speed, difficulty or novelty. So this is going to be more important than the exact data source or type. So in our previous example of parish records, the source of data is not entirely important. It's about the volume um complexity speed of that data. Additionally, the data must pertain to people, actions, behaviors, choices, and statements. And finally, the research question should be a social science research question which uses this atypical data to talk about how people make decisions or what influences their behavior and choices. So, you know, the exact research question is not important. What we want to stress here is that it must be a social science question. And that's why we have that kind of really obvious intersection between the computational and the social science. And that's what we're really going to be focusing on today. And we have a nice little quote here. Um so in essence, um computational social science is an opportunity to do socially valuable research that just wouldn't be possible without computational methods and tools. So by this I mean that you know we couldn't for example manually scan years and years and years of police recorded statistics to try and understand how crime rates have changed over the last I don't know 50 years. Um stuff like this is either physically not possible to do manually or it's just going to take such a you know massive massive amount of time. But with the use of computers, we can apply advanced statistics and models to understand this kind of change in crime rates. And we can look at how we can count many types of crimes that are taking place in certain areas. We can even aggregate these crimes to map them spatially. We can explore the long-term trend seasonality or the noise components of different crime types. And this type of research um you know it just wouldn't be possible without the use of these novel computational methods. And you know that other example that we had of web scraping millions of online articles to try and understand how political opinions have changed. Perhaps we're doing this looking at the last 20 years. In order to get a sense of those political opinions, you know, we'd want to examine the words in articles. maybe how many words belong to different categories, what themes are appearing, what are the proportions, are these um certain words being used, how is that changing over time? That's something that we wouldn't be able to do without that computational method web scraping and also different natural language processing techniques as well. So, now that we've um covered a little bit about, you know, what makes up a computational social science project, what key elements, that kind of thing. We're going to have a little bit of um interaction now. And what I'm going to do is I'm going to give you guys some examples of different projects and then you're going to vote on whether you think a given project is a conversational social social science project or not. So, if you would like to head back to menty.com, we can get this set up and then you can start voting and we can just have a chat about things as we go along. So, I'll stop sharing now and I'll head over to menter. Let me get the right thing. Okay. All right. So, seems some of you have already started voting. Nice. Let me just set things up that you go ahead. Okay. Let me find my slide. All right. Okay. Uh, sorry about that. Okay. So, let's see what we think about this. So, conversational social science or not. So we've got scanning historic recipes and using AI algorithms to recognize text to identify ingredients and measures used over time. Okay, so seems that um most of us are saying this is definitely CSS and then a lot of us are also saying not enough social science. Um so yeah, I guess it's one of those annoying things because it's bit of a broad question and it does really depend. So those of you that have said definitely CSS can see why you've said that because we have um this um conversational method, right? So we have one we're scanning historic recipes. Um I suppose well scanning no if we were web scraping them that would be a conversational method. If we're manually scanning them not necessarily but the use of AI algorithms definitely right. So we're using them to recognize text to identify ingredients and measures used over time. So that's that computational bit check. But um the social science aspect, well I guess it's going to depend, right? Isn't it? If we're looking at just you know a pure um measurements maybe for a um a biomed paper or something like that, that's not necessarily social science. But if we're looking at, I don't know, how different cultures have incorporated different foods and how that's maybe transformed, I don't know, local food markets, whatever, then that would definitely be more of a social sciency kind of um project. So yeah, it's um it's also those of you that said, you know, I need a little bit more information to decide. Also fully valid there because we don't necessarily from just this um brief title have enough information. Okay, what about this? So, we're going to use a gamified smartome display to understand how people interact with energy saving technologies. So, what do we reckon? I think this one's a little bit more clearcut than the last one. Um, but let's let's see what see what people think. Yeah. So it seems that we don't have as much of a divide this time, right? Um so let's go through it. So a gamified smart home display um to look at how people interact with energy saving technologies. So the smart home display there, that's going to be if you're looking at that data, that's going to be a unique computational um bit of data there, isn't it? And also a novel a novel um bit of data as well. So that would take that computational um element as well if you're looking to analyze that. Um and social science where we think well we're looking at how people interact with energy saving technologies. So we're looking at people's behaviors. So for me I'd say big fact check as well on that social science element. So yeah it seems um seems less divide here. So some of you still said you need more information which is completely fair. Some have said not enough computation which is um interesting. I would think given that we're using this um remember we talked about for instance having you know novel wearables like you know people have these smart um watches now don't they to collect a lot of health data here we've got a gamified smart home display as well um which is a novel form of data right it's quite new um we're probably going to be analyzing that data as well um you know maybe using some computational techniques So I would say it's definitely definitely also hitting that um computational mark as well. Um some have said not enough social science. Um I would um pick up just this bit in the question as well. How people interact with those energy saving technologies. So that's uh for me looking at people's behavior. Probably there's an implied um you know bent to the paper as well. If someone's going to be writing this up, they're probably going to be wanting people to interact more with these energy saving technologies for um you know um climate change purposes and all of that sort of stuff. So seems pretty social sciency to me. So I hope that hope that one kind of um made sense once I've gone through it a bit. Let's look at another one. [clears throat] So CSS or not, advertising for survey participation on social media and storing the responses in a database. This one's a little bit um trickier. I expect um a little bit. Yeah. All right. Nice. So a lot of you have um picked up as well that remember when we talked about in one of the earlier slides as well you know it's not just using digital um forms of um traditional social science methods right so you know a lot of us maybe in our undergrads or in other research projects might have used websites like is it like you know survey monkey that kind of thing to collect um survey data But you know just because it's a digital form of doing that doesn't necessarily mean it's novel or it's um you know a new computational method. Right? So uh bang on for those of us that have said not enough computation here. Where's the novel data? Where's the computational method in that it's not it's not implied. Um, some of you said need more information, which you know with these kind of questions is it's always valid to want a bit more information with them because we've just got one sentence here. It's not a hell of a lot to go off. So that's also quite fair. Um, and also someone has said quite rightly not enough um, social science as well, which yeah, where's where's the social science question in this, right? Um, I mean that's my fault for making you only choose one thing. Um, ideally we would have ticked both of these, right? Because there doesn't seem to be a social science question buried in here. So nice one for that. And let's go to the next slide. So we're going to be reading in real time weather and air pollution data to create complex models of hyper local air quality. So yeah, this um this one's a little bit tricky as well. Um kind of reminds us of that first question. Um you know, the one we had about scanning historic recipes, looking at ingredients, that kind of thing. Cuz even though we had AI algorithms like a lot of you um said here, not really a social science question implied there. I mean there could be right you know perhaps they're looking at how this air quality affects you know I don't know some particular towns or something or you know but it's not if this if I was just looking at this question I would be saying yeah where's the social science in here um some of you have said definitely CSS which you know I can see why if you are thinking oh okay well we're looking air pollution because we want to, you know, look at this data and um come up with some solutions for this particular, you know, problem. I can see why it could have both those elements for you. Um, but if I was just looking at this based on this one sentence, I'd probably say, where's the where's the social science component? I need that to be a bit more a bit more obvious. What about this? Um, so we're going to train a neural network on social media data to create a believable chatbot that's then going to counteract online radicalization. So what what do we think here? Okay. So, I'd be inclined to agree with um the majority here that this is in my eyes pretty um you know good example of a computational social science project. So, let's break it down um a little bit. So we've got training a neural network, right? Um so that's computational method is ticked there. Um we're going to be training it on social media data to create so a novel piece of data there as well. And we're going to be creating a believable chatbot to counteract online radicalization, which for me, you know, is very social sciency because we're looking at, you know, changing people's behaviors, how people are acting online. So to be doing this, we're probably going to be need needing to know, well, what is what are um you know, certain language that's associated with radicalization, right? So there's probably going to be a bit of text mining involved in there, which is where we, you know, try and scrape certain text data and we apply natural language processing techniques, which are ways of um, you know, looking for certain words, counting their frequency, clustering them, um, you know, trying to find out what is associate, what's the kind of terms that are associated with radicalization, what are the signs. you're going to need to apply those tech techniques to then be able to counteract the online radicalization. So for me, I think this is has that nice blend there. It's got social science. We want to know what how people are behaving online. So we're going to need some way to understand that and we want to change that behavior that in my opinion very it's got that good social science element and computational. Well, we're going to be training a neural network that's tick social media data that we're probably going to be processing and applying certain, like I said, natural language, natural language processing techniques to it to be able to, you know, get some meaning of um the terms and stuff behind radicalization. That's again a big tick for computation. So, computational um elements. So, I hope that um makes sense. Um okay. Um you don't have to do this. We won't do the word cloud for now because um I don't think we've covered enough yet. But um what I'll do now is I'll head back to the PowerPoint and we'll continue on with the slides. Okay. So sometimes here I do take a short break. Um who is there anyone who desperately wants us to take a 5 minute break? I'll just see what you guys say in the chat and if there is at least I don't know one or two people we can take a five minute break but if everyone just wants to plow on then that's um that's totally fine. So yeah, in the chat just maybe pop in yes if you want us to carry on. No if you want us to take a break. Okay, getting a lot of yeses, so let's just crack on. Okay. Right. Um, so we're going to talk a bit about um how to become a computational social scientist. Um, so it might seem rather basic, but we are going to be covering first what a social scientist is. Um, not that I assume that you don't know that, but we're just going to be covering some of the skills. And I think it's I think it's good to go over stuff like this. Um so firstly those things first thing to say is that social scientists think like people and you might thinking well yeah but what I mean by this is they use a lot of human type thinking skills like abstraction infer inference the ability to understand fuzzy concepts and background knowledge that ability to not shy away from gray areas or overlapping categories because those sort sorts of things are part and parcel of being um a social scientist. And that's um no surprise because they're going to need these sorts of skills because they're studying people, interactions, and behaviors. And that requires a certain skill set because people are pretty complex. Um societies are pretty complex. But social scientists also build up a lot of data skills in the course of their research. So if you think about things like response categorization, encoding, quality evaluation, patent detection and statistics, these are all skills that a social scientist does have. So it's not to say, you know, well that's the domain of the computer scientists. Social scientists have some of these skills already. But whilst often they are going to be using computers, who's not in today's world? This might not often involve writing computer code, right? So it might involve computer programs such as SPSS or status data. I'm not sure how it's actually pronounced for statistical analysis, but maybe not much um sort of familiar familiarity with programming languages. And then we have um computer sciences. So we're making some big generalizations here. Um but in contrast to those the types of skills I mentioned before that are associated you know more with social scientists those human type thinking skills we can say more that computer scientists have to think like computers. So the thinking skills that they have are more along the lines of concrete definitions, more absolutes. So they're going to have to think more in terms of strict hierarchies and categories, clearly defined and scoped variables and rules. Um, and in terms of data skills, computer scientists collect, analyze, and manipulate data through programming scripts, computational methods, and technological tools. But unlike social scientists, they might not be taught as much to identify or motivate research projects on the basis of societal impact or value. So those of you that are social scientists, you might have um had to justify your research on the grounds that broadly speaking, you know, it'll make the world the world a better place, right? Even in a small way that contributes to some knowledge base. It it could be like I said with that example before, I'm researching radicalization radicalization online forums to produce um insights that could lead to counter measures. Whereas a computer scientist might be more focused on a logical justification for a particular project that's more along the lines of well I want to make this algorithm more efficient so that it uses less memory something like that right so in order to do computational social science you're going to need a blend of those skills um so you're going to need that social science and that computer science um element So let's go through these four kind of things. So you're going to need these skills that we've touched on before. So these are skills like being able to identify important problems or knowledge gaps, considering possible solutions, connecting problems to relevant theories or perspectives, and being able to collect relevant information and research to frame your research approach. These are all things that social scientists excel at. Ability to understand context, nuance perspectives, how to communicate abstract ideas, and how to attack a research question. Whereas this might be an area where computer scientists may struggle a bit more as they're more used to, like I said, those more concrete definitions and absolutes rather than these gray areas or murky social science concepts. So that's the first thing um that you're going to need and we're going to cover why you also need those um computer thinking skills. So what kind of skills are we going to need here? You're going to need the ability to access, organize, process, and handle vast or complex data. You're going to want to know how to write collaborative code and how to do document um your work for flow properly which is often a step that people um neglect. These are all skills that um computer scientists might find quite easy and might come quite natural um to them whereas uh you know social for social scientists it can be much harder to transition towards those um computer thinking skills. But um like I said, you know, social scientists do have those data skills that they can build upon. So those are those things that I've mentioned before such as you know coding responses, pattern detection and statistics, formatting surveys, all of these things. So it's important to remember that um social sciences do start out with a very good base here and you know it might seem a bit wishy-washy but this is actually a really important ingredient because I can tell you as someone that has gone from um a social science field to doing uh computational social science, it can be really intimidating at first. um you know that computational element, the stuff you need to learn. But that's why it's good to remember that you know no one starts out with all the skills that they need. Nor do you know all the skills that you might need to acquire. So this happens to me quite a lot. So, I might start off saying I want to scrape tweets um for information on the 2016 US election. And I'm expecting, yeah, I'll probably need to know how to code a bit, but what I don't know is that that's going to entail learning about APIs or different file formats or all these unique ways to visualize the data. Um, but if you approach computational social science with an open mind and a willingness to learn, you can then, you know, gain more skills and you'll start to be able to find your way. And you'll also start to understand that some skills as well have a steeper learning curve than others. Um, so you know, from what I found, it's fairly simple to learn how to do a little bit of web scraping, right? I'm going to get all the links from a few pages. But learning to build your own neural network and train it on a large amount of data, well, that's going to be a completely different beast, right? So, that's where as well collaboration comes in um with people from other fields. And then you start to get this nice bridge between the social science and computing worlds. As social scientists learn more about computing and vice versa, we begin to see more conversation and collaboration between these two field fields and some really unique and interesting projects. And what you will find is that you're not going to need to know everything about computer science. But if you know enough, you'll be able to have those productive collaborative conversations with others in the field. And you might think to yourself, okay, well, I'm going to I'm going to try and get good at web scraping, right? But if I'm looking at building a neural network, well, maybe I'll get in touch with someone from the computer science department, right? And I can learn a bit about the basics, but maybe they could help me with this big project that I'm wanting to undertake. The final ingredient that you're going to need um if you're embarking on a computational social science project is a problem that's going to require those skills. Okay? And that's a mixed problem. So it's one that's going to require that blend of human thinking and computer thinking. So some of you might be here because you've already encountered one of those problems. And I wouldn't be surprised because you know as resources become more digitized these unique um projects and um problems are going to become more relevant. Um everything's becoming more smart smart and moreworked. We've got the fact that just such a sheer amount of data is now available to us espec es especially online and things are updated faster. You know, if you think about a classic social science problem of maybe we're interested in how men and women move through cities differently or we narrow down that question to look at how people with disabilities navigate cities traditionally for that kind of social science problem. Well, you might have stationed some interviewers in different places to stop people as they go past. You might have counted maybe how many people go by that are using mobility aids or maybe you'd send out some surveys to people's houses. But now there is much more opportunity for us to gather a large amount of data with computational methods. So you could collect data from public transport networks about how many people bought tickets or how many people swiped their card at the tram stop. I know on um Oxford Road um for example, you have those little smart sensors that will tell you how many bikes have passed um on a like on the day. Um, you can get sensors which track as well how many cars go past a given point. You could use AI algorithms. You could get in touch with local councils and look at CCTV footage to identify how many people are moving through a space and even at what speed. So, as you can see, there are new ways then of approaching traditional social science questions. But there's no reason as well why we have to abandon those traditional methods completely. Part of really good research is evaluating different methods and comparing outcomes. And it might be interesting to see whether by using different methods, you get different answers. And if so, then you can ask, well, why is that? Which then may prompt further questions. So yeah, it can be a difficult task taking on a CSS project, but it has a lot of benefits in terms of building upon your computational skills and also really strengthening those social science skills that you might already have. Um, so before we move on to the eightstep process, um, I didn't realize I put this thing here. That's kind of weird. Also, it's taking my background off so you can see how messy my room is. Um, but yeah, before we move on to the eightstep process, I've noticed that the last few times that I've done this workshop, a few people have said that they want to know more about the career path or, you know, just how someone finds their way to doing computational social science research. So, I thought I'd highlight my colleagues in my CSS team and their background and just like how they find their way, how they found their way into doing this kind of work. Um, so you can see that we have my boss Jules who did an undergrad in linguistics and then a master's in evolution of language and cognition and then she went on to do a PhD. We also have my other colleague Nadia who did criminology undergrad and then did a research master's degree in criminology and social stats. And finally myself I did politics uh undergrad and then I did a conversion degree masters in data science and AI. So that was aimed at students from a non-computational social science. Um no a non-computational science background. Um so I did that at the uni of Liverpool. Um and I'm highlighting this you know not to show off about all of us but to show that there's no computational social science degree. Not yet anyway. I do think there is more CSS type degrees that are popping up out there. So, I know the Uni of Manchester now offers, I think, a masters in social research methods and statistics with CSS. Um, but what I want to stress is that you don't need to have studied this undergrad or for your masters to do a CSS project because chances are your undergrad and your masters or your PhD have provided you with really useful transfer transferable skills that you can use to carry out a CSS project. So you can see that for Jules, she gained a bunch of text mining skills and a knowledge of how to perform statistical analysis in her undergrad and also learned about advanced stats and agent-based modeling in a masters and PhD. Meanwhile, Nadia, like many of us, gained skills in her background related to traditional statistical software and then was introduced to the programming language R, which she then used primarily in her masters. So you can see um a lot of um like I started off in my undergrad I think the only um software that I used was SPSS and Stata right but these are really good foundations you know if you can use software like that then you know you have the ability to to understand how to navigate what is quite complex software to other people so you know you will have undeniably gained a bunch of skills that are then transferable to, you know, moving into a more CSS direction. So, what about coding? Well, as I mentioned in the previous slides, you're not going to need to be an expert coder to carry out a CSS project. After all, like I mentioned, collaboration is going to be key for those of us that are students or academics working in higher education. Um, we're really lucky as well to have a big reservoir of potential collaborators to work with. So consider reaching out to enlist a programmer or an expert um to help you with your project, especially if you're getting started, right? You know, you like I said, you might have learned the basics of something, but maybe they can help you point you in the right direction or maybe they can have it's just good to have sometimes a second pair of eyes to have a look at your code, that kind of thing. Um, but if you want to carry out a CSS project and this is something that you're going to want to do, you know, maybe not just as a one-off, you want to do it, you know, a lot more now that there's all this new and interesting data, then I would consider um getting some knowledge of programming languages like R or Python. And that's because these languages are going to offer packages and libraries which are going to help you implement a computational method. So in our team for example uh my colleague Nadia is our resident R expert. Some of you might have heard of R already as it is becoming quite popular in the social sciences now and it's been used for a while in other fields like biomedical science bioatistics and that kind of stuff. Um if you've previously used stata you might find that it's quite similar. It's really user friendly and has a nice approachable layout. Um, so you can see I've just put up an example of um the R programming language um in R Studio. So those of you that have used a static, you can see it's quite similar there with these like four panels. Um, you also have the benefit if you're going to learn um, R that you don't have to around with picking a code editor or um, an integrated development envir integrated development environment, sorry, as that's all provided with R Studio. And it's also great for producing data visualizations. Um, it's just really superb at that kind of thing. So, it's a good choice for those that already have a background in statistics, as you'll probably find the syntax and the functions more intuitive, whereas um me and um my boss Jules, we mostly use Python and that's just down to what we were familiar with during our masters. Um, so Python is more of a general purpose language unlike R and it's not just limited to data science. It has a really broad user base. It's popular with web developers, software developers, etc. But it's also known for its simplicity and readability. Um, the syntax is pretty easy to pick up, but unlike R, you do have to do a bit more shopping around for what kind of coding editor you want to use. Um, so you can see here I'm using a coding editor called Jupyter Notebook cuz I like how um I can just have my cells um you know straightforward um and quite linear. Um yeah, so like I've said um probably best not to go massively into this side of things as this is just an intro workshop. But I would say the biggest learning curve for me for getting into programming and computational methods was setting up my computational environment and learning basic code. Um, and when I say setting up my computational environment, that's stuff like how do I navigate the command line? What is the command line? How do I install software or coding packages? How do I write my first function in Python or R? And um that's the sort of stuff that I cover in our code anxiety club. Um and that is um going to be on October the 6th if that's something that you think will be interesting and you want to come along to. They're just half an hour sessions. They start at half one till 2. Um you can just they're just basically live stream to YouTube. You don't have to put your camera on anything because it's just streamed to YouTube. You can ask questions in the chat. You can go completely off topic and ask me any random computing questions that you have. So I will um just to spotlight that if you think after this webinar I feel like I do want to get into um coding that kind of thing that'll be a good next step for you guys. Um all righty. Um so there will be an opportunity to take a little short break if that is something that you guys would be interested in. Um and after that I want to quickly introduce an eightstep process for how how to undertake a CSS project. Um and these eight steps are going to be about identifying problems, exploring the problems, formalizing concepts, collecting data, um using those concepts to experiment or analyze data, discussing your findings, communicating, publishing and presenting your work and sharing your findings as well as documenting and validating your findings. Um let me just see what time I'm on. I think we have more than enough time for a short little break here. So, what are we on now? Let's see what time it is. So, let's join back here at 52. Um, get a brew if you need to. I don't know, stretch your legs. Um, go to the L, that kind of thing. And we'll meet back here at 52. So, I'll just mute my um video and turn off my audio. All righty. Um, okay. Let's um crack on. So, let's go through these um eight steps then. So, to make the process useful to you, you can um start thinking about either a project that you'd like to tackle or a research idea that you've been thinking about. It could even be if you just have no idea at all, a project that you've done in the past, you can jot this idea down or maybe even um put it in the chat if you want and we can have it in mind as we go through these steps. Um okay, so step one is identifying the problem. So once you've identified the problem, the thing that you want to study, you're going to want to be as sorry this we go. You're going to want to be as clear and specific as possible about the pattern, the problem, or the lack of insight. You're going to want to identify um who is involved, where it is, etc. And what this will do is it will help you to define your research question. So maybe we have a goal in mind, right? So we want more people traveling actively through city centers. We want, you know, less cars on the road. We want more people riding their bikes or scooters or, you know, just being able to get from A to B in their wheelchair. So the research question might be, what are the barriers to active travel in city centers? So what I would do then, so how this step comes into play is I will identify who is involved. So you can just start to list down who might be involved. So whether this is potential companies, people, different demographics that may be of interest to you. So for my example problem, I might want to look at city councils, bus companies, different businesses, different vehicles. And it's better as well to just go all out with these lists as well because it's going to give you different avenues to explore. And you can always, you know, cross out any after you've done a bit more investigation into them or decided that they're not actually that relevant. But this is just a nice brainstorming part of it where you jot down everything related to what kind of problem you want to look at, what kind of people might be involved, what kind of, you know, place it might be, that kind of thing. The next step is going to involve exploring the problem and that's where you'll um gather information and perspectives in multiple ways. So you might carry out a few surveys, observations, some secondary data analysis, maybe a little bit of web scraping. Um so this could involve um conducting a few interviews with people of interest. So, with my little example before about travel through city centers, I could be interviewing the manager for my city's transport network or local council workers, but I would probably also need a survey or maybe some observations or secondary data analysis to capture how many people are actually moving through the city center. So, it's about using different methods and tools to further enhance your understanding of the problem or the research topic. So, you're going to want to uh spell out any sub problems that might appear, processes, relationships, simplifications, assumptions, related issues, all of that kind of stuff. So, after settling on that main research question, you're going to need to then get more specific in order to make that question relevant and measurable. So if for example, what are the main barriers to active travel in the city center is my main question, I might want to specifically be focusing on what are the barriers to active travel through this specific city center at this specific time of day given the way that these specific roads are laid out. Right? So this is where you really nail down the particulars of your research question. What I sometimes do here is I head to meny.com again and I just um invite people to share their first steps and their second steps, what kind of things they have in mind um just so I can see what people are thinking of. We can do that and um I'll head to see if there's not much of an appetite for that. That's fine. I can just um crack on and we can go through um what you call it the um the different steps in more detail. But I will share the slide now or we can we can see what everyone's everyone's thinking. The reason why it's hard to do this and it takes me so long is because there's a toolbar right at the top from Zoom and it's it obscures everything else that I want to um click on which kind of let's see sorry about that. There we go. Okay. So, like I said, I invite you if you want to to share a bit about your steps one and two. Like I've said, you don't have to. There's no pressure. And if there's not a lot of, you know, appetite for this, we can just carry on going through those eight steps. Um, so I'll leave it maybe a couple of minutes. Um, and like I said, no one wants to share. We can just move on. All right. Nice. So, we've got someone new secondary data. So, yeah, key part of step two is, you know, exploring that problem a bit more in depth, looking at what kind of secondary data exists to prompt you in further directions, you know. Um, focus question with boundaries. Yep, nailed it. Identify stakeholders, people, and also data sources. Yep, brilliant. Um, how do people experience competing demands in the workplace? That's a really interesting research question. And then second, um, step two would be looking at surveys for that. So, yeah, brilliant. Um, that's really interesting, um, research question as well. I mean you could even look at like social media data could be a good um avenue to explore that. I mean anyone who uses X or Twitter um knows that you know a lot of people will use that to vent and talk about maybe workplace issues. There's particular subreddits as well that will focus on um you know workplace stress that kind of thing. network of involved people in organization. Yep. So, um step one is great chance to just list all the people that are involved in it. Um what organizations? Um how do you proceed if you suspect the data does not yet exist? That's a very good question. Um I would like to know a little bit more about I guess what area that you're studying, but maybe you're thinking about you know exploring the problem. step two and you're thinking, "Oh, well, this is actually a novel area of study." Um, I guess it would be about, you know, explaining that a bit more. Why is it a novel area of study? Why does this data not exist yet? Um, maybe if you could give me a bit more of a idea of maybe what it is, what kind of data that it is um that you you would want to study, I could maybe advise a bit more. Um, that's an interesting question. region with lack of data transparency. Um, is that the person who's um maybe could put in the chat if you're the same person who's um put this um question there. But yeah, I guess it would be trying to think about ways that you could see I'm a person who mentions teachers. Okay. Hidden population teachers with math anxiety. Okay. Um Oh, yeah. That is a really interesting one. I suppose for that then that would be research that you would want to carry out for exploring the problem. Obviously, if you suspect the data doesn't exist yet, it's about talking about that. And I guess maybe related issues, you know, you would look at, well, who normally suffers maths anxiety? Why is it the why is there this gap? That kind of thing. For this kind of like for it to be CSS, you would have to be thinking, well, how are you going to what computational method are you going to apply to study that, right, that makes it computational? So I wonder if you thought of what particular method um you know are you going to web scrape um experiences of teachers that have maths anxiety perhaps that have expressed that on particular forums or that kind of thing. Are you going to look at another sort of computational method natural language processing that kind of thing? Um yeah, using critical realism lens to explore mechanisms behind social phenomena. These are all brilliant. These all sound super super interesting. Um thanks guys for sharing that. What I'll do now is I'll go back to the um PowerPoint. I'll talk about a bit more about steps three and four. We can always then share our steps three and four and we can chat a bit more about this as well. Um, so let's go to the back. Okay. So, moving on for step to step three. This is where we formalize our concepts. And what I mean by this is you'll want to make all the concepts and processes explicit format formal sorry and both computer and human understandable. Um so often times um this is referred to as pseudo code. Um but you don't need to know how to write code. You just need to start understanding how to formalize things. Um, for instance, maybe your research question focuses on trust, which is a very social sciency sort of concept, right? If our goal is then to get a computer [clears throat] to be able to measure it or model it, maybe we want, for instance, our computational method is something like agent-based modeling, right? Or we want to represent it in a simulation. It's about thinking, okay, well, how do I define it in a way that a computer would understand? So maybe I decide to define trust as a variable between a variable between zero and 100. Maybe I'll need to make rules about, you know, how that variable will change in certain situations. Maybe if two parties in my simulation or my agent-based modeling interact positively, then that trust increases. But given a negative interaction where one of those parties is judged to be deceitful, maybe that level of trust then declines or even resets to zero. So you have to start thinking about how to formalize concepts in your research question so that a computer would be able to interpret it. So, you know, if I want to find out, maybe I'm looking at something like um social network analysis and I want to find out, well, I wonder in this particular science journal, I wonder if people collaborate with the same kind of people, you know, perhaps they perhaps if I had a social network graph, I could look at the connections of who's worked on which paper together and try and look at some of these clusters to then understand, you know, who's working together, what are they working on, that kind of thing. So, it's about thinking about that method and how we'll get that research question into that method in a way that makes sense to a computer. Then on to step four. So, this involves collecting data, implementing software, and verifying your process. So, you need to select and implement one or more methods. So many of you might have thought a little bit about um some methods when we touched on step two. So you know maybe when you were exploring the problem you had a look at maybe what some people had done before had an idea of maybe what you kind of wanted to do. You could have wrote uh down web scraping agent based modeling something like that. And this is going to be the step where you implement these methods and make sure that they work in the way that you anticipated. So you know when you're working on a um computational or computational social science project, any project that basically comes with computation or data, there's always going to be hurdles, right? Maybe the data comes in a different format than you expected or maybe you're encountering a bunch of error messages in your code. So we want to look at well can we make it work in the way that we expect right maybe um we've come across a problem where with our data set and you know this often happens with me I you know maybe I've scraped a bunch of things um from the internet and then I've got this data set right but I'm applying a function to it and I just something's just not clicking it's not working in the way that I expect maybe what I'd do then is reduce my data set down to maybe just five rows, apply that function again and try and look through what's happening with each row as that function is applied just while I fine-tune my method. And of course, you know, the choice of your method is going to be highly dependent on the research topic as well. If you're looking at online radicalization and you're wanting to get social media data, then that's going to influence the choice of method because you're going to have to web script, right? So, you're going to be using that web scraping method. If you're going to look at the type of language that's being used, okay, natural language processing there, that's going to push you in a direction towards a certain um um method. Lastly, you're going to need to thoroughly check the selected method has been implemented correctly. Um, and that's what we mean by verifying your process or method. It's about answering that question. Did we do the thing right? So, what I can do now is I can pop back over to Mensy and you can share some of your ideas about what kind of um computational method you could use for a particular problem. Um, what you've kind of been thinking about, that kind of thing. So, let's head back over to um menty. Let's see what's next. Okay, I'll give it a couple of minutes and you can just pop in any ideas that you might have been thinking of. Any methods that you're interested in, computational methods that you think, oh, that would be that would be an interesting one I want to explore with my question. You could maybe tell me what your research question is, what your method is that you'd want to use to study it, and I can maybe point you in certain direction or give you some advice. I'm just reading the chat now. Um, no worries, Pedro. We must go back to the trenches of coursework preparation. Sounds rough, but uh, thanks for joining. And to the other people that have had to leave and have left a message, um, thanks for joining. It's no worries. People that need to leave, I get it. You know, we've all got um busy schedules and stuff, that's no worries, but thanks for joining. All righty. Um web scrape teacher forums for teachers asking their peers for math support with or with a without mention of anxiety. Yeah, nice. Um that would be a really interesting one. Um, there's I always think Reddit's a good one for this because there seems to be a subreddit for everything and people tend to get really indepth. Um, I often if I have something that I'm thinking about or even you know health stuff instead of just going to Google I often times will just type in the query and then Reddit after it and see what's being mentioned. Um, yeah. So we could look as well if you were scraping those um forums, you could look at identifying key words, right? So maybe it's maybe they don't mention anxiety, that word per se, but maybe they're mentioning they're worried or they're um confused or struggling, you know, those kind of words. something like natural language processing would be great picking out terms like that. You could even look at there's something called sentiment analysis which is good. So maybe um teachers that are posting on these forums about maths maybe when if we did a sentiment analysis we would notice that there's a lot of negative sentiment there. So what sentiment analysis does it might take a a sentence. So, um, the sentence might be, "I'm very worried about this." And it will score each word with a sentiment. So, a sentiment score, um, 0 to one. So, for worried, it's going to pick that out and it's going to notice, okay, that's a negative word there, and it'll give that sentence a score. So something like sentiment analysis might be an interesting one for looking at those posts about um mass support because if there's a lot of negative sentiment then that suggests something in itself. Um I'm trying to understand polit political parties responses to great power rivalry. I am thinking of using uh a web scraping method. Yeah, that would be um brilliant um place to start. You could be um scraping um party websites, manifestos, um you know, forums that are related to those political parties, that kind of thing. Um yeah, that's a really interesting research question and um a good a good method there. Somehow collect data data using AI to see how individuals respond to competing demands. ideally collect data to measure their anxiety levels, ethical considerations on both. Yeah, that's a really good thing to point out as well. I've just been mentioning, you know, scraping uh data willy-nilly and scraping data from Reddit. These things have big ethical considerations, especially when we're looking at disclosive data, right? So, that's something to really bear in mind. Um also as well whether you can access that data because a lot of social media websites now have clamped down on getting access um to data. So Twitter used to be a great source of social media data but now it's just completely um you know behind a pay wall. You have to pay an extraordinary amount. No social media company now wants to give its data for free which is really sad for researchers. Um but yeah, somehow collect data using AI to see how individuals respond to competing demands. That's interesting. That is very interesting. So when you say using AI, you thinking of using like um something like chatbt or cord to help you set up a web scraping script or sometime somehow use that to just pull the data from the internet. Um that would be interesting to know. um left my comment in the Zoom chat. Let's see. Okay. Yeah, I can see. I'm going to create messages from public Telegram chats to analyze how people migrant expat communities discuss health issues, how they spread, share information or misinformation. My master's thesis project, but I'm not sure how to tackle the coding problem. Theoretically, I could handcode it and then do LLM assisted coding, but this feels very timeconuming for a mast's project. Do you really think it would be possible to only rely on LLM assisted coding? Yeah, that's a really good question. Um, so I will this is something that me and my colleagues talk about a lot because you know everyone is using AI now. Um, a lot of you will have heard of people vibe coding, right? um which is just a way you don't uh don't worry about the typers no worries um which is where you know you might have a limited amount of coding knowledge but you can kind of use chatbt or some other LLM to kind of like um brute force your way there. Um I there's there's no substitute. So in my in my opinion to use CHBT and Claude I have to know some of the jargon, right? So I need to know some of the coding jargon or I need to know details about the methodology to get something useful out of the AI. So I will say um I was thinking of coding the topics of the messages like main points. Yeah. No, that's fair enough. Um, I found LLM assisted coding to be a little bit rubbish, but then that was maybe a couple of months ago and things are moving fast in this field. Um, I think Chachi has just released Astra. There has been some papers on people that are using LLMs to, you know, like for instance topic modeling to put things in categories. I tried to do this thing with my notes app where um because you can plug in there's a plugin now for your notes app for um chat and I don't tend to keep anything disclosive on my notes app tends to be shopping list quotes from books little random thoughts I have at 3 in the morning and I said to chip I was like okay I've got all these notes thousands of notes I want you to put them into different topics okay so I was expecting something like shopping you know, um, random thoughts on this topic. Um, just these categories, right? It really, really struggled. It really struggled putting them into categories. It hallucinated some things. It just couldn't seem to really um I couldn't seem to handle it. Um I've read um Regina is saying in the chat I've read some of the papers on that so far and the results are so different. Some say that LLMs are almost as precise as student assistants but that there are so many limitations. Yeah, it's a really I would look at Google Scholar have a look at these papers for anyone else who is interested in this. It's really interesting. But I think yeah, at the minute I would say even though it is time consuming, I would maybe look at if you're going to use an OLM, use it to help you draft some code, but don't use it to do the code the coding itself in terms of like putting things into categories, right? I would look at maybe applying some puppet modeling. see if you can scope out the categories and then code in the things with the code that the LLM comes up with, not just asking the LM to do it because it just hallucinates a lot. Um, okay. Um, okay. I got a bit carried away there. Um, but thank you guys. That was really interesting. Let's go back to Um, yeah, no worries, Regina. Thanks for that. That was really interesting. Let's go back to the slides. Okay. Okay. So, at step five, we're going to want to experiment and analyze the data. So, we're going to run the experiments, build the models, analyze the data, or otherwise use the methods that we've selected in the previous step. And we're going to try and identify and explain the results within the context of the experiments. the model or the method that we've used, right? So, if we've used social network analysis, we're going to be looking at um some social network analysis metrics, we're going to be looking at what we found, that kind of thing. Then once you've run your experiments and analyzed your data, you can then start to interrogate the results and form some conclusions. So this means going beyond the experiment model and method to draw some conclusions about what the results mean, what sort of picture they're forming and this is where your social science science element comes in as well. So based on the results, what's going on? Like do these results point to a particular policy recommendation? Who or what do these results affect? Why does it even matter? what should change, who benefits from that proposed change. So, um, if you remember a bit before I was talking about that example I had about looking at how people move through cities, perhaps I found that people with disabilities related to mobility have more difficulty navigating through particular areas. Maybe there's even uneven surfaces or particular roads that I've identified that are too narrow. In that case, then I can recommend some very specific changes. right wider foot paths um I don't know ramps rails whatever that kind of thing so what is our research showing what's it pointing to us towards now I would normally maybe go back to my we have been having some absolutely cracking discussions but just in the interest of time I notic we've got a few questions on the Q&A as well and we're nearly at half and I'm going to just crack on but you can always pop something in the Q&A if you've got a question for me and I'll get to it at the end. Um, so step seven. So we're nearing the end of the research process now where we're going to be focusing on communicating and sharing our research. And in terms of communicating our research, it's important to understand that all of the previous steps that we talked about, they're going to be communicated to multiple audiences in different ways. Um, and what's really important is that you think about short-term and long-term engagement. So for a lot of us our mind will instant instantly go to well I want to get this published in an academic journal which for sure is important but it's also good to think about other forms of communication. So, are you going to present your work at a particular conference or submit a piece about it for a blog or, you know, maybe have your own personal substack? Um, and you can think about whether there are workshops or classes that you could present to get your research shared more widely. And with that as well, you're going to want to think about how you adapt your style and tone of communication to suit those different audiences. Right? I mean the way I talk in a Substack blog might be well it's going to have to be very different than um how I talk in a academic journal for instance. So think about different avenues you can go down and you're going to want to share document and validate your findings. um worrying by uh validating the research. It's just making sure that the right thing was done by allowing your work to be studied, reproduced, and or modified as needed. And to do that, you want to allow as many people as possible to be able to access your methodology and any code or data. Um so you want your research to be as transparent, well doumented, and open as possible. Of course, you know, there's going to be um caveats to that. Um you might be working with admin data that is restricted restricted. So you could look at workarounds. A good one is to create a dummy data set so that people can still run your code and work through your methodology. But nothing disclosive is identified, right? Reproducibility of course is a really important part of research and unfortunately it is often neglected which is why it's good to think about how you document your work before your research gets underway. Um so yeah if you're coding make sure you put code comments in not just for other people's sake as well but for your sake. I've done it so many times when I'm coding and I think don't understand that code anymore and it's because I haven't properly documented it or written any [clears throat] code comments on it. And just a few things to note at the end. These steps are not linear. Um there's going to be many many points from each step that you'll need to return to or apply throughout the research process. For instance, um documenting your research is something that you'll want to be doing throughout the project. When it comes to computational social science projects especially, most or all of them are going to require many iterations, which means revisiting certain steps. So maybe you'll come up with a research question in step one, but after exploring that problem a bit more in step two, you might want to then jump back to step one and reformulate the research question in light of something that you've read. Maybe you're on step four and you've implementated implemented your method, but you need to go back to step three because you've not outlined the concepts and the processes enough. That's why uh documentation is so important so that you can capture all of these really important nuances and changes as your research evolves. I've often had it where I'll make some really important changes to my code, but like I said, I don't document it. Then I come to writing up my methodology and I'm then really struggling to explain why I opted for a particular type of algorithm or a particular coding module. So it is a really really important step. Um and it's actually a really good habit to get used to. So, you know, when you're sketching out your research question, you can note down why you think it's important. And then what you'll find is as you're documenting a bunch of stuff, the research actually kind of writes itself, which is a much nicer process, right? Versus coming to a blank page and thinking, o, I'm having to start right from the beginning. When it comes to your code as well, good practice is often to put it in a code repository to protect it. You can also put it in a cloud and make it available for others to look at. And that's great then because often when you spend so long looking at your own code, it becomes really difficult to spot issues with it. But if you have a colleague look at it, it can just be helpful as they might spot a problem that you've overlooked. Um, you know, it's like we talked about at the beginning, it's really important to be open-minded and remember that you don't need to know everything about computer science. You just have to be enough to be able know enough to be able to have those productive collaborative conversations with those in the field. And definitely I am someone that is not you know some some people have a very intuitive computer science brain right comes very naturally to them. They can pick up coding at the drop of a hat. Wasn't like that for me. I found it really difficult and I still find it quite difficult but I also find it very interesting and rewarding. So there's a ton of data out there now more than ever. It's an exciting time to be doing a computational social science research project. But yeah, always remember that especially if you're someone that's at a university, you've got loads of, you know, computer science people around you, you might want to collaborate or just help you out on another part of your project. Um so we'll just note that there's some references um if you want to have a look at these and the slides are going to be um put on the um events page at some point so you can always go back and have a look at these in more detail.