Submind YouTube summaries
Thumbnail for Using QGIS to map 2021/22 census data from the UK Data Service

Using QGIS to map 2021/22 census data from the UK Data Service

Watch on YouTube

Video summary

James Cron from the UK Data Service leads an instructional session focused on utilizing QGIS to visualize 2021/22 census data, specifically targeting the creation of univariate and bivariate choropleth maps alongside area cartograms. The workshop begins by guiding users through the process of downloading aggregate statistics for specific geographies, such as wards in Leeds, using tools like Census Explorer and the Boundary Data Selector to obtain shapefiles containing both electoral boundaries and census counts. A critical initial step involves transforming "long form" tabular data into a manageable wide format via pivot tables within Excel or QGIS, followed by normalizing raw population counts to generate percentages and adding categorical summaries that identify dominant accommodation types for effective mapping preparation. Once the data is prepared, the session details configuring QGIS layers including raster backdrops from Ordnance Survey, vector boundaries, and CSV statistics joined through geographic identifiers. The presenter explains fundamental choropleth principles while addressing limitations regarding assumptions of uniform population distribution within administrative polygons, demonstrating how to classify continuous variables into discrete buckets using methods like equal intervals or natural breaks with appropriate color ramps. Advanced techniques are then introduced, including the use of plugins such as 'by variant' and Cartogram; however, due to specific software instability on certain setups, some demonstrations utilize pre-configured packages instead of live generation. This segment highlights how bivariate renderers can display combinations of variables like accommodation type and tenure through grid-based legends with adjustable opacity for underlying map layers. To overcome the inherent inaccuracies of standard choropleth maps caused by varying population densities within fixed boundaries, the tutorial introduces asymmetric mass choropleth maps as a superior alternative. These visualizations utilize building footprint layers to shade areas realistically rather than filling entire administrative polygons, offering a more accurate representation pioneered by institutions like the London School of Economics and adopted for official census data releases. The session concludes with practical steps on exporting final visualizations as PDFs using QGIS Layout Manager while ensuring proper attribution is given to UK Data Service sources. Additionally, attendees are encouraged to explore further learning modules from the training program or contact the dedicated help desk service, where queries regarding complex topics like statistical disclosure control and geographical selection can be directed to specialists based on their specific areas of expertise.
Read the full video transcript
Okay folks um welcome to this using QST map 202122 census data from the UK data service. My name is James Cron. I work for the UK data service up in Adena in Edinburgh. Um so the aims for the session today are we're going to use the QGS desktop GIS application to map centers data down the original UK data service. There's three specific aims. We're going to create different types of mapping. a univariatic chlorophyth map, a bidic chlorophyth map and an area carttogram. As you see this is a journey. So the objectives are to create those different types of visualization. And along the way we'll learn various things about the UK data service, how you use the data and how to use the GIS application CUGIS. So we'll first learn how to download data from the UK data service. That's the actual census data itself and some census boundaries from different parts of the UK data service. Um we'll learn out how to carry different types of data transformations on the data and this is to prepare it for mapping in CUGIS. Um we'll learn about some of the issues to be aware of when mapping census data to make sure we create valid maps that are not misleading to our users of our maps and then the we'll look at how to use actual QS itself and so if you've not used QS before this will be an opportunity to find out what QJS and what you can do with it and at the end there'll be a short section where we just go over the sort of resource available in the UK data service to help you at using census data and mapping that census data. So in your joining instructions, you will have proided with a link to um a GitHub repo which has a PDF workbook of all the steps I'm going to demonstrate during the demos today. The workshop's going to work from the basis of I I will give short talks which will talk about the fee behind some of the things and then I'll do demonstrations using the UK data service website and the tools or using some QS software or doing some data manipulation in Excel or Libri Office. I'm on a Mac and I've got Excel. So I'm going to do my data manipulation using Excel but you can also use Libri Office Calc. And the choice is you up to you. you either follow just what I'm doing on screen and try and do it yourself on your own machine or you can just um watch the and then maybe do later after the workshop because I say you'll have access to the workbook so you can always go through afterwards in your own time. So I'm just first give an introduction to the census itself. So the UK census is a denial census and it happens every 10 years. um within the UK data service, we mainly support the census from 1971 through to 2021 and 2022. Uh in 2021, uh 2022 uh in Scotland because of COVID, they did it a year later in 2022. Whereas in England Wales, it was 2021. So that's why there's a a difference there. Um the census itself um in the old days it was done by um people would get a paper form for the post or a numerous would come around and ask some questions. Um it tends to be done online now but the the different questions are there are household questions and there are questions about individuals. So the household questions will ask people stuff about the type of type of house they live in um the rooms in the house and stuff how they how they own the accommodation like are they renting uh are they mortgaged or they are they renting from a social landlord. Whereas the individual questions are questions about each of the individuals within that accommodation. So it could be age, sex, what sort of job they do, how they travel to work, etc. So on census night which is usually in April and the last the last one in England Wales was 2021 Scotland will be April 2022 the population supposed to fill in their census forms and submit the results to the census agencies which in England Wales is off national statistics and in [snorts] Scotland there's national record of Scotland and I say this is just some more the topics covered by the UK census. So you get this question related to education, housing, health and language or transport. Um and during the sort of demonstration of practice today we're going to focus in on housing data and we're going to look at accommodation type and tenure type. So sort of are we looking at is there a flats or houses or semi- detached and is it rented privately or is it socially rented. So when the sens agency gets all the data um after census night they do a lot of processing on the data um so they they bring all the data together and they and they produced output uh sent us aggregate data sets and these output the different variables as either counts of people or households. Um and that data that centers accurate data is produced at different levels of output geography and the smallest level output geography is a thing called the census output area and the output areas are specially um created just for the census and this is an example of an output area for part of Edinburgh. So um each of these um black polygons is a sentence output area and you can see they cover a few houses or streets within a neighborhood of Edinburgh and the actual um sentence variables themselves. You can see there's an output area code which is a unique JRock identifier which unique identifies a sentence output area and then you have various statistics related to that output area such as total population in this case the number of households the number of males females and is that are there how many detached or flats within that area and there's different types of sentus aggregate data there's univariate um uh aggregate data and is multivaried. So the univariate census data just contains information about a single census variable which could just be like accommodation type or tenure whereas a multivariate data um basically um gives you multiple um topics for each um data set. So it would give you accommodation type but also tenure combined and this gives you a much richer insight into the information. So you could take a univariate um census data like the method used to travel to work which could be um by car or foot or train and you could add additional variable which be how um the occupation of people are doing which would give you a multivaried data set and that way you could um do some sort of classification of different types of um people's occupation how they're traveling to work. So how are people who um do skilled jobs traveling to their work as opposed to people who might be um um uh working as in um building occupation travel to work so that you might make some differences there. So multivaried data sets are a lot richer um whereas in this workshop we will simply be using univariate data set because it's easier to manipulate. But for the UK data service you can download both univariate and multivariate census data. So the UK data service itself um provides access and trading to u economic and population social research data. uh and the emphasis here is we provide access to both the data itself and the training around that data. So it's access to the data and then how you would use that data and there's an entire training program of which this workshop is one part of and within the UK data server itself is a specific part of us that do that purely support the census and we are called census support um and we're distributed along census support is um uh distributed along different institutions within the UK so I'm up here in Adina and we support the census boundary data. I have colleagues in Manchester who support the aggregate data itself and I have colleagues at um UCL in London who support the census flow data and Jill is based at the Kathy Mars Institute in Manchester who support the microents micro data and today we're going to look at using centers aggregate data and centers boundary data and each of our different um subsets of centers have different applications in order to access different types of sensor data. So in terms of the census aggre data tools, this gives you access to raw census stats themselves. SECAN will give you bulk access to the census tables. Um so you can download the census table say for all of the UK or all of Scotland or Wales, but there's no ability for you to then to to subset that data. So if you only wanted data for Glasgow or say you would have to do that subsetting yourself perhaps in spreadsheet or a database application. Whereas a data explorer is a second UK data service tool which allows you the ability to create extracts from the complete census accurate data sets. So they would allow you to just restrict your census data to Edinburgh or Glasgow or Lurra and to pick the actual types of the variables within the tables of census data you want. So it means you you can be a lot more um specific of the sensor data you want and where you want it for. Data Explorer currently provides access to 2021 and 22 data and I think it also gives you access to 71 data and we have a bunch of legacy applications that we're in the process or sorry the engineers in Manchester are in the process of translating or migrating that data from the legacy platforms into data explorer. The ultimate aim is eventually all the census data will be available through data explorer and you'll be able to use the same tool to access data from 71 81 91 through to 2001 and 2022 possibly 2031 in that census happens. So I'm going to do a first demonstration and I'm going to use the census data explorer to download 2021 census data. So let me just um minimize my slides and I'm going to open a web browser. So I'm going to first go to the UK data service and show you how you access data explorer. So this is the UK data service. Um we have a data catalog. Um which will allow you to search for data sets across the entire UK data service including the census. But within the census itself, we have a quick link to our particular census applications. So within the data catalog, if you click this quick link to census, it will bring up this pop-up window which lists our various census data applications. So you can access the data explorer. Um well, this is the census boundary here. This is another one we're looking at today. Or census full data. And it say infuse is a legacy census application which gives you access to earlier census data. And I'd say you've got this CCAN which gives you access to bulk data. But we're going to use data explorer. So I'm going to click on data explorer. So this is what the data explorer tool looks like. Um you can see it's got 71 data and 2021 2022 data. I think it also provides access to other non-sensus data. some data from World Bank, but we are specifically concerned with the census data. Um, these are quick link buttons at the bottom here. We can just click in these buttons and it will take you to like all census 21 2022 data sets. Um, but for the purpose of this demonstration, I'm going to search for a particular type of census data for 2021 and 2022. I'm just going to refer to my um handy workbook that I have given you all and I'm going to look for 2021 census accommodation type. And so you can see there are 54 data sets of 2021 or 2022 census data which meet that term accommodation type. And see these are all listed here. But the one I want is this TSO44 accommodation type data set. So if I click on the data set we get an overview. So we uh data explorer will give us some metadata about what the sort of data is. So you can see it provides census data estimates that classify households in England Wales by accommodation type. Um and I say at a minute we're currently um if we just if we hit the download button now we would download the entire data set whereas I want to split it down just for data just for leads because I'm going to show you how to create a map of accommodation type using of using a 20 on data for leads. Um so to to um split the data up and refine what we want to download there are various filters are left here and I'm going to set these various filters to um filter the data. So first top level geography I'm going to set England. Click apply and you can see it adds a filter under geographic grouping. I only want to download census data for a specific type of output geography. In this case, I'm going to download the data for wards because it's a relatively small geography. Um, so I can click on the wards and divisions item and click apply. And this this will tell data explorer to output or census data by wards. And having tell that I now to tell where I want to download the data for. Um so again at the minute it's selecting data for all living whales. I want to only download data for leads. Um and what you have here is this um this will basically show you all the sort of um sub geographies of England Wales that we can download data for. At the minute it's currently quite unwieldy. So what you can do is you can do a collapse all and then we can like drill down through the geography just to restrict to leads. So um so I know that leads is within Yorkshire and the Humber. So I can just do a drill down this way. Um I believe Yorkshire. We're going to check my PDF. Being in Scotland I'm not particularly familiar with Yorkshire but let's check. Ah so leads under West Yorkshire. So, and here's leads and and so I want to each of these um items here is an electro ward within leads. And I could just manually go here and and I click each one to add it to the data selection. And this is what there's a different types of selection mode here. But what I'm going to do is change the selection mode to items and all items directly below. And this way when I click on leads it will select all the awards within leads and now I click apply and this will add that to my data selection and so so we've constrained to the type of geography electro wards I'm constrain just to lead and then the final filter is the actual sensitive variables themselves. I'm going to want all these because these are different types of accommodation. So again I can change selection mode rather than manually individually selecting these. So items and all types right below and click apply. And as you can see we now set up our filters within data explorer. I can get a preview of the data by clicking the table button. And you can see we have geography for leads. So each of these is a different ward and across the columns we have the different types of accommodation type. And to download the data we can use the download button here. Um, the trick here is to make sure you use the filter data in tags or text form button because that will output the data as like a plain CSV that we can then use in Excel and then later use in our desktop application. If we use the unfiltered data then basically data explorer will not just ignore all those filters we set and we'll get the entire data set. If we use the table in Excel, then we get a lot of extra formatting information, which might look pretty, but it makes it hard to manipulate the data. So, I'm just going to go for filter data in tabular text format. And you can see we've downloaded the file. Um, a quirk of data explorer is it creates a very a very long file name. So, I'm going to make this more usable by shortening the file name. Just something a bit more manageable. Great. It doesn't like it'll make a lot of difference cuz Yeah. Okay. Uh, okay. I'm just going to go back to some slides now. And so, we've downloaded our set of stats from data explorer. The next step is having download those stats and we need to do some data manipulation to prepare those sensor stats for mapping intuis. Um, and there's like three things we're going to do to the data. We're going to transform the data from long to wide form. I'll explain long to wide form later. I'm going to do some data normalization so that we can don't produce misleading maps. And then I'm going to do an extra bit where I add an extra variable that summarized the center stats to create create a different type of map. Um these first two are pretty much essential. We never use data from the UK data service because of the way the data is first provided and then in order to to normalize that data for mapping this third thing as a cascar summary variable is just an extra step that I'm doing just so I can create a different type of map. But normally if you're downloading data data from the UK data service from data explorer, you're going to have to do a transformation to from long to wide form and you're going to have to do some sort of normalization on the sensarials before you map in CQIS. And it's a I'm going to do a manipulation in a spreadsheet because um spreadsheets are great at handling like tabular data. You can use like um Libri Office Calc or Excel. Um I'm on a Mac and I happen to have Excel. So I'm going to use Excel to do my man manipulation. But you can use Libri Office Calc. And in the workbook there are instructions for doing the same operations in either Libri Office Calc or Excel. And so this first transformation we have to do is convert the the data downloaded from data explorer from long to wide form. Um so data sets have different types of shape which like um is determined by how the data is arranged into rows and columns. Um so data explorer will provide the data in what's known as a long form and what you'll have is each record will have a will will contain a different census variable and um as you can see here um so here we have three different output areas. You can tell there's three because each one has a unique um go ID which is this EO5011 384 thing. And each of the different sentence variables is provided on a new row. [snorts] So if we've got eight sentence variables that for that one output area, it'll be provided on eight records. Um whereas what we have here is the same data set but in the wide form and here you get a single row but each of the different census variables in is is then is instead shown in a different column. So we have again our same E05 01 on 384 output area but this time each of the different columns provides each of the different centers variables. Um and so within CQIS, CQIS is a lot can is a lot better at handling data in this wide form because as we'll see later, we'll download centis boundary data and for each of these output areas, we'll get a polygon which um which we'll attach to this table and then we'll have the census variables of one record that has all the columns plus the output area geography by row. as opposed to this. If we were to add the col the geometries here, it would make the table a lot more complicated. But as I say, because the data comes in long form and we had to wide form, we have to do some data conversions within our spreadsheet just to make the data mappable. And this can be done in a spreadsheet application using a pivot table, which basically just groups the data by the unique geographical identifier. And I'll demonstrate that in a couple of minutes in Excel. Although you can do the same thing in Libri Office Calc. [snorts] And then before we actually map the census data, we need to do some data normalization, the census data itself. So that's the the counts of of population of of of individual households or people or just raw values. Um, and so if we want to map the data and we want to say where we we're basically creating a map that's comparing one area to another, then we need to do some normalization. Um, because otherwise we just be showing the raw counts which would not be very meaningful. And so there are different ways of doing normalization. We can divide by the actual size of the area itself. We should express the data the value as a population density. So um percentage of sort of people per hectare or kilometer squared or we can divide by the total population size which would then express that as a percentage. So, so the percentage of people within that particular geography which live in flats or rent from a private sector landlord or travel to work by bike or travel to work by um car and because it's a percentage you could then compare that small area to another neighboring output area or against country as a whole. Um and again in a spreadsheet application we can do that by just doing some sort of simple um um data manipulation. So we would have like the count of our private rent here and we have the total and dividing that to the the the private rent by the total and then multiplying by 100 would express as a percentage. And then I say those first two are essential in order to be able to map the data. And this third thing where I add a categorical summary variable is just an extra step that I'm doing just to like provide an extra thing that I can map. And what this is going to do is going to look across all the different sentence variables. So in this case for the different types of rent, mortgage etc. And it's going to find the dominant one. So which category of house type is the dominant one for the output area. So in this case it's going to be come out as mortgage or private rented and that's just a good way of summarizing the data and say what's the dominant type of um accommodation type within that output area. alternative. You could have like a what's the the least dominant would be a different one you'd add. And what I would do again, I'm going to use this in Excel and I'm going to do calculate this column and then I'll be able to map that and show a dominant a map of DOM accommodation type. So I'm going to have a second demo where I'm going to demonstrate how you do that preparation of the data that we downloaded from data explorer in Excel and I'll show you how we do the pivot table, how we do the data normalization and then how we create that categorical variable and let me just find the data set we downloaded. So accommodation type leads 2021 census and I say I'm going to open this into Excel but again you can use libri office cal. I will just increase the size of this. So this is the raw data we pulled down from data explorer. Um there's quite a lot of columns here, a lot of like metadata. So it'll tell you um it will tell you the table data came from. Um the top level geography. Um the important thing is the geographic area one because that each of these it will tell you what output area that um record is is related to. It will give you the localized name for the output area or I think these actually wards so the actual ward name and then it will give you the different types of accommodation type. So in this case detached semi- detached terrace and the always value is the actual count. So the actual census variable for those different types of accommodation and I say what I'm going to do at the minute this is in long form I need to convert into wide form. So I'm going to create a pivot table in Excel and you do this using the insert pivot table menu. So an Excel insert. >> And so you able to maximize that spreadsheet? >> It it doesn't really increase the size. Is it still quite hard to see or >> um I can see it but others might struggle. >> I think it's just the actual Yeah. Okay. Okay. Yeah. So, insert and then pivot table. And so, there's a pivot table field here where we have to define how to create the actual pivot table. So, we're trying saying how do we want to aggregate the data. And so, what you do here on the top here are listed the different fields within that spreadsheet. And there are different boxes here which controls how the pivot table is constructed. So what you just have to drag them from the top here to the different boxes. So I will first grab accommodation type eight categories and drag that to the columns box. Jog area that as I say is the unique set of J identifiers and we'll grab to rows. And then you want the actual observation values themselves. So the actual counts of the sentence variable or the values and what you can see is it's aggregated that data so across the different um variables by the uh geographical identifiers. So we end up with like one row per uh unique ward and each of those unique rows has the different columns of the census data and we get a grand total which sums the values sum sums the columns and then sums the records. So what I'm going to do here is then grab this data. So do a copy and then create a new sheet and then paste the data. Enter that new sheet. Let me just increase the size again. And what I'm going to do now, I'm going to rename some of the columns because they're quite long. So this row labels one I'm going to rename to Ward geo ID. So ward geographic identifier. This rather lengthy a caravan or from a structure temporary structure. I'm going to rename that to caravan because detached. It's probably okay if detached, but I'll just make it detached with a small D. um [clears throat] this very lengthy one. I'm just going to change this to um comma b. This is all um explained in the PDF. This is just going to be flat. Um shared h again another lengthy one. It's um I'm just can change this to um s C ch or warehouse um semi- detached will just become semi debt terrorist again. Minor tweak there and grant total I will change to total. Great. So I say these are raw census counts and we want to normalize each of them by the total population so they get expressed as a percentage. So I'm going to create another eight columns and each of these will become renamed as prop uh detached prop and you can do the same for the own just so that it's then clear that we have the data as provided as the raw count and then as a proportion and so to do the actual uh conversion itself we enter a Excel formula. So we do the caravan one first. So equals so the raw value which is C2 divided by the total and then just multiply 100. So that's expressed the value as a percentage. So we've got 3,000 divided by 9,000 gives you a potential of 32%. >> [snorts] >> and we want to use that same formula across the other columns. And so to do that, we're just going to edit the formula slightly um and just anchor uh so this way we can use it across the other columns. So what I did there is I just insert a dollar just before the column so that it still refers to the total when it applies it to all the other ones. So if I just drag across check, let's just check. Yeah. So, we've gone D2 / G2. Yeah. Is that correct? D2. No. Okay. I got the first one wrong. That should have been B2, not C2. Okay. Sorry about that. Now, that's correct. So detached is C2 / J2. So comp is D2 / G2. That's correct. Flat. Uh well, we got E2. Yep. Divided by G2. Correct. Okay. Final check. Terrace. Um column is I. Yep. So we've got I2id by G2. Correct. Okay. And then we can just fill the other records. And let's just do a check. So B31. Yep. Divided by J31. Yep. Correct. Okay. And then we just populate the rest of the records. And again, let's just pick one to randomly check. So, uh, semi- detached. So, column is H. So, you got H23 divided by J23. H23 / J23 * 100. Yep, that's correct. Great. And so that's expressed all of those census variables as a proportion. So as a percentage. And so we've done our we've done our transformation of the pivot table. We've done our normalization. In the final bit I'm just going to create that categorical column which in each case will tell you what is the dominant census variable for each of the words. And again this is another um I'm going to use another equation to do this. I'm going to first give you a column name for this and I'm going to call this dominant accommodation type and it's quite a lengthy expression. So I'm going to grab it from the actual PDF. So if you just bear with me wh I look for the workbook. And so this is the same PDF that you will have and this is the equation we're going to use to do the cascarable summary variable. So I'm just going to copy that and put it into the cell. So what you can see is the equation that's told us for this record this ward the dominant type of accommodation I the one with the highest percentage is the semi- detached which if you look at the data is the case and again I'm going to want to um just to make this clearer I'm going to replace the dash prop I'm going to get rid of that in the actual text and I can just do that by editing the equation slightly and wrapping of a substitute ute. So just edit it and substitute and I'm going to tell it to substitute the underbar prop bit with nothing. So and so you can see it just makes the text a lot cleaner. So it's got rid of this like the prop from the column and that just tells that this the dominant accommodation type is semi- detached. And I'm going to I want to copy the same formula down the rest of the columns. So again, I need to edit the uh formula just to anchor some of the the ranges with the dollar symbol. So referring to the PDF, I will add a dollar here just to anchor what it's referring to. And so that's still correct that one. And I should now be able to copy and paste to the rest of the records. And this will now refer to the correct data. So again, a quick check. So for this board, it reckon the dominant one is detached. And if we look along, yep, the dominant one is 41%. So that seems correct. So we've done manipulating the data in Excel. If you're using livery office calc on the menu system and I'm now just going to save this data and so later we can manipulate it in CQIS. So save as comma separated values save. I want to replace it. Yes. It'll complain multiple things, but I only want to save the current sheet because that's the one I've been manipulating. So, I can just okay that. And if I just close Excel and go to wherever my data was. I will open something. Yeah. So this is our data. That's what it looks like. We've got all our preparation proportional value and our dominant acceleration type. So that's the sensors aggregate data. In order to create maps, we also need to get download some some geospatial boundary data. I'm just going to go back to the slides and do some more talk about geospatial data. So in terms of geospatial data um it's used to model some aspect of the world um different types of geospatial data if you never come across before vector data which is um um it's like point lines and polygons. Raster data is um typically used to represent um like satellite imagery, digital train models. Um it's essentially a grided data set. Um each um cell could represent um an area of elevation or like a say a part of a satellite data set or it could be um uh like temperature or water or something. Um geospatial data is um what special geospatial is is um spcially referenced um the data is provided for a particular um um what they call spatial reference systems. So here in the UK we use ordinance survey data or data from the or national statistics um it uses I think called the British national grid to locates that data within the UK. If you're using global data for like say all of Europe that would use like a thing called WJ84 and that data um uses a different sort of spatial reference. Um there's a whole polar of different types of geospatial data formats. Um there's thing called shape files which are like the sort of almost the Excel of the geospatial world. Um we'll look at a few of these formats as we go on. Um there's a bunch of geospatial standards which um define how spatial data is created and shared. Um there's an entire discipline called geographic information systems or science. Um and there are various um what they call GIS software applications which allow folk to create manipulate and do things like create maps or do some data manipulation or what they call spatial analysis. Um, a commercial well-known one of these is called ArcJS. And there's also an open source one which we're going to be using today called CUGIS. Um, CQIS has been around for like quite a while now. And despite the fact it's free, it's incredibly powerful and has some great features and it's ideal for doing sort of census data. So mapping census data um in terms of the census output geographies. So as I said before the census out data is output a range of different small areas. Um so I say we download data at word level. You also can get data at what they call output areas or lower super output areas or middle super output areas. Um um and together these type of out geographies form a centers geography hierarchy from the smallest from the output area up to LSO to MSOA to ward to district to county to country to nation. Um but the thing you have to be aware of is there's this thing called statistical disclosure control which means that not all the census topics or the census tables are available at all geographies. So you might find that there are certain um possibly disclosive census variables um relate which could only be output at like higher geographies like um ward or district and would not be available out of area level which might um have an effect on the sort of research question you can ask um and this is just an example of what these census boundary data is actually is. So a census boundary data set is a polygonal uh geospatial data set that just tells you on the ground what the uh output geography is like. So on the left here we have a count area boundary for Glasgow. On the right we have the much smaller cent output area boundaries. Um we also have a thing a concept or a thing called a geographic lookup table which allows us to relate the different types of centers geographies together. And this can be useful if you want to make some transformation between the different types of centers geography. So you might have data at output area level, you might want to convert it to data at um ward or district and then you can use a direct lookup table which maintains relationships between output areas and wards or LSOS or super output areas within the same census. There also a set of lookup tables that do the same across census. So you can make lookups between um output areas in 2020 22 and 2011 um which but you have to be aware there are some because the geographies tend to change over time. They're not always one to one exact. You just have to be aware of that. But um and you also get a type of geographic lookup table called a postcode directory. So a lot of data which um you may be manipulating from like surveys which um is output using a postcode and what a postcode lookup table does is it relates postcodes to other type of geography. So you might have a postcode of EH104EL in a postcode directory. Each of the record for E104 would have a bunch of other columns which would then tell you the the geography that postcode fell within. So that might be the census output area for 2021 or the the electrol for 20201. Um and that way you you could then use that lookup table to relate your data at postcode level to census itself and then use the sets to add context to your other um survey information. and the the postcode topography lookup tables as well as the other sensors jack lookup tables are all are all available through the UK data service and so we looked at um uh data explorer from the Manchester team we're now going to use um one of the boundary data download tools which I in response were looking after at Adena to download the census boundary data and We're specifically going to use the boundary data selector tool to download the electro wards for leads um that we can then join to our sensor stats in order to do some mapping. Um so yeah be another demo where I'm going to do some download of sentence boundary data from the b data selector. So, I'm go back to the UK data service. Let's just maximize my browser this time. I'll try and increase so zoom things in a bit. And again to access the boundary data selector, we can use that quick link. So data catalog, click on the sensors um loger black thing. Yep. Up pops the popup. And we want census boundary data borders. And this is the boundary data selector. So the boundary data selector allows you to look through the different boundary data sets we provide. And again, you can either download the complete data set. So we could download electro awards for all of England or in our case you actually want to drill down and only obtain the electro awards just for leads. So it's quite a simple application um really it's got two tabs. There's a what and where you tell it what data set you want and where you want it for and a format tab that lets you change the type of uh geospatial data format the data is provided in. I'm really just going to use the what and where format because it will default to providing data in shape pile format which for our purposes I'm using the data in Q just is fine. So there's three dropowns on the top here which will control the data set is listed. Um so I want data for England I want electoral data and 2021 and later. So you can see it's come back. We have a choice between electoral wards and parliamentary constituencies. So I want the electoral wards list areas and and like the data explorer we can now drill down through the sub geographies to restrict the data purely to leads. So again similar to data explorer it's York and the Humber and again we want leads. Let's go borders boundary data selector. We've told it we just want data for leads. And if we hit this extract boundary data button, it's going to go to the database, pull out the features, and then dump it as a shape file. So that's done that now. Just minimize my browser. So it's our data. I just open it. It comes as a zip file and you can see it's a shape file. So a shape file is made up of what they call um is is these four different um individual files is a PRG, a DBF, a SAS, and a shape. And the important thing when you deal with shape files, you have to have all four of these files in order for the data to be to to work. So if you were to copy this to somewhere else, you'd have to make sure you copied all each of the four components to the same place. Otherwise, if you try to open the data in CQ, it would complain because it wouldn't be able to find the rest of the data it needs to show the data set. Um there's various other things that we get provided with D from data selector. A simple read me just tells you those files where it's for and sort of the folklore support the service. And this terms conditions file here which tells us um specifically because the data is open it's released under open government license and in order to use open government license data the one of the conditions whenever you use that data you have to include an attribution statement that that tells you where the data came from and we'll see this later in CQIS when we're going to create a map and we'll go back to this file and insert that copyright statement into something we create in CQIS. So that's our data from CQIS. Oops. Just find my slides again. So let me grab the boundaries. I'm just going to grab some other data from the ordinance survey. And this is like um contextual. It's going to like give us a contextual background map that I'll add to CUGIS so that when you when we overlay that we bring in the center boundaries. We'll have some context just to add to make our map better. Um and I'm going to grab this data from the ordinance survey. Again, this is open data. So the ordinance survey is Britain's national math agency. Um, and they have a website where you can download some of their open data sets. And so if I just ordinate survey and what they call is the data hub. And so this is the OS ordinance survey data hub. Um, and they have a bunch of they have open data here. Um and it lists various open data sets they provide. So they've got stuff from the British Geological Survey. Um so rock type and stuff and stuff related to soil, but we want um actual map data. So you can change the providers from British OS survey to all you just want the orange survey. And this is the data set we want this one to 20 250,000 scholar kit. It's like a a road atlas. It'll provide some nice context for our boundaries. So, I'm just going to go to the download page and hit the download button. And this is going to download data for all of the UK. And this is raster data. So, these are just like image images. So, let me just download that. And I go to my downloads. This is the ordinance survey data here. It's called our zip file. So open that. And if I go to the data itself, it consists of this um mapping as a bunch of like separate what they call image tiles. So I can try opening one of them here just in like uh my Mac data preview. So you can see it's basically this is a 100 km by 100 km block of map data in this case for uh land end. So, so that's we've got data from the UK data service data explorer, UK data service boundary data selector and the ordinance survey. I'm going to now open CQIS and show you how to use QGIS and we're going to add all these data sets to CUGIS. Um we're going to have a break at 10 11 but before that I'm just going to show you I'm just going to add the data from our various data sets into CQIS and then after the break we'll get on with actually doing the mapping the data in CQIS. So so first I'm going to start CUGIS. So, I'm on a Mac. I'm on a fairly recent Mac. So, I've got QGS4. I think in the instructions they told you not to use QGS4, but um it seems to work. Okay. Um I think I will have to manually edit the instructions for next time to say that. Um uh so this is CQIS. Let me just I unfortunately I don't think I'm going to be able to make it much zoomed in. So, a lot of these OP menu things may seem may seem quite small still. So, you just have to um if you if you don't manage to follow along now after the workshop, you've got the workbook. So, you might know to you might want to do it yourself when you have more time if you sort of lose track of what's going on here. [snorts] Um so, this is what Q just looks like. There's a bunch of um it's like most of our like gooey applications. There's a bunch of buttons along the top here which you can do use to do different stuff. Um there's what we call the browser window the left here. Um these are different um types of data you can add to CQIS. Um this main window here is the actual is a CQIS map window. So when we add data to CQIS, it'll be displayed in this map window and the left here is what we call the layers panel. So each of those data sets will be will will basically form a layer of data and the different layers will be shown on the left here. Um I said before on that geospatial slide that the spatial data was providing different spatial reference systems and I mentioned the British National Grid and WGS84 for the globe. By default, curious will default to displaying data in WGS84 for the globe because it assumes you might want to add data for anywhere in the world. Whereas what we want to do is display data just for the British National Grid in the UK. Um, so at the right here you can see there's this thing called EPSG4326 and this is Q just telling us that it's assuming the data is going to be displayed in WGS84 and that the data you add is in WGS84 and that's not what we want. You want to tell CQIS to display our data using the British National Grid spatial reference system. So the first thing I'm going to do, and this is like in your PDF workbook, it will tell you is to change this from EPSG 4326 to EPSG 27700. 27700 is a code for the British National Grid. So to do that, I just click and you can see it brings up this dialogue saying project coordinate reference system. Um, and I'm going to set that to British National Grid 27 and 700. Um, you can also set it to any of these other ones if you wanted to. Um, you might be downloading data for America like in Idaho, which has a SE has its own unique um, spatial reference systems. Um there's a whole load of different spatial reference systems which have different purposes, but for our case, we want data in British National Grid. So we're going to set it to EPSG 27700. So there we go. So we've got a data into 7700. I'm now going to start adding data to CQIS. So I'm going to first add that raster data that we downloaded from the ordinance survey that backdrop mapping. So to add data to CQIS use this button at the top left called open data source manager. So if I click that up pops this dialogue um and this is called the data source manager which tells QJS where to source data from on our local machine or from like outside our local machine like from web services or exam etc. And you can see down the left there are all the different types of data we can add to cus. So I'm going to first add the raster data by clicking the raster option. Um, and here we specify the location on your local machine where that raster data is. So I on for me it's on my downloads folder. It's RAS 250GB data. And I'm now going to add one of these um map blocks to the one I want is SEIF because I know this is for leads. So select the SEIF and click open. That's fine. CH just reads the file. It's grabbed some metadata that's used that will display the data. And I'll just click add. And we get our data set shown. Um, by default when CQIS draws raster data, it doesn't always show it particularly clearly because it's it's it's optimized for like speed rather than like um the actual quality of what's been shown, which can slow things down. But so we zoom in, it can be quite distorted and blocky, and I want to make that clearer. So I can tell curious to um draw the map a lot clearer and I don't really care if it draws it a lot slower. So again, as I said, what we have here loaded is one of our data sets in the map window and we have the layers panel at the left that's listing our current layers in CUGIS. And so you see it's add a layer for SE which is our map tile. Um within the layers uh panel you can select any of the layer items i.e. the different types of map and you can right click and go to properties and that will open the properties a properties window for that type of layer and we can use that properties to set different properties of the of the the layer being displayed. Uh we'll use this later on when we want to um create our census maps. So the minute it it collected symbology and in order to tell QS to display the raster data more clearly I can use this resampling option here and so the minute it's it's say zoomed in nearest neighbor nearest neighbor so it's applying some raster operation to tell it how to display that data and I'm going to change this from nearest neighbor to cubic and over sample It's quite low. So, I'm going to change that to a lot higher. And this basically just tells CUGIS to draw the raster data a much higher quality, acknowledging there'll be a hit to the actual render speed or the draw speed of the map. So, I'm going to hit apply. And you see now Q just redraws the map. It's slower to draw. but it's a lot more higher quality. And so this is leads here. This is going to be our area of interest for the rest of the workshop. I'm now going to add the other census data to CQIS. So again remember we downloaded the sentence boundaries as a shape file from boundary data selector and we downloaded uh uh the center stat as a CSV file from data explorer. So I'm going to first add the boundaries. So again same deal use the open data source manager button. This time want to use the vector because it's vector data set shape file. Um again navigate to the data boundary data and the one we want to select is the shape. I add I the shape file. Open that. Add and close. And here are our boundaries in cubis. Um so because it's a vector data set um we can again we can use some of the buttons along here to target the data. So this I is called identify and it allows you to identify the feature on the mouse click. So I can identify click on one of the polygons and it will show you various properties of that polygon. So we can see it's a this particular ward has this uh geo identifier and it's called hairwood at the minute it's not set as data because we've not actually provided any sense it's just simply the boundary and we can do stuff like we can select one of the polygons or we can select a bunch of the polygons um and we can do sort of basic spatial analysis. So um so we've got a polygon selected. We might want to find all the other polygons that are connected to this selected polygon by doing a select by location operation. And so when we do that, Q just finds all the other intersecting polygons with our selected polygon. [cough] Um let me just add the actual sets of stats themselves. So again data source manager this time you want the limited text. So browse to the data set in this case our sensor stats which is just the CSV file. Click open. Um so we get a preview of this of the set as stat CSV. You can see our lovely um proportional data expressing the values as proportions and our our uh categorical thing that telling us the dominant combination type. Um under geometry definition make sure that's set to no geometry because the CAC val doesn't actually contain any geometry itself. that's in that's in those polygons. We're going to join them later. And again, just hit add and then close the dialogue. CUGIS will add that centers um CSV to the layers. Um and so we can open the attribute table in CUGIS. So you can see it's loaded the CSV into CQIS, but it's not showing a map view because at the moment it's looking for the map. Um so that's our data in CQIS. Um we're going to take a short break now for 10 minutes. Uh and then in that after the break we now go into the actual fun stuff of doing the actual mapping of this data. I think that 11:25. So we're going to start again. So, so far we've grabbed data from the UK data service. We've done some manipulation of that data in Excel and we've stuck it into CQIS. We now want to go into the process of actually creating some mapping. So, what we're going to do is we're going to create a thing called chloropl map. And I'm just going to give you a quick introduction into what chlor mapping actually is, although you may well be aware of this already. So, At the left here we have like a bunch of boundaries. Um I think these are uh Scottish counter areas. Um you can see each of the count areas has a a uh geographic identifier. That's the S12 code and it has some random um stat attached to it. Percentage of something I guess. So 5.27 8.2. um not particularly useful as a map and quite hard to see the differences are. And so that's that's where chlorophy maps come from that. You basically shade the polygonal data based on the actual attached variable. And so high values get shaded like dark green in this case or lower values light green. and it allows you to quickly look across the data set and look at the variation of the census variable. Um so then you can compare what the census day is like in in say Edinburgh versus um the rest of Scotland or the rest of the UK. Um I think on the left here this is showing a percentage of people who work in um the forest industries. Um so not surprising in Scotland that tends to be the Scottish borders or Wales the sort of um rural parts of Wales and less um people employed in forestry within the city of London for example or the UK Scottish central belt whereas I think on the right we have a chloropath map which is showing the percentage of people possibly house type um and so the dark areas are where it might be there's more boat living in flats. The lights less living in flats. Um so there are some limitations of the qualify map we have to be aware of. Um they tend to imply because you're just shading the entire polygon the same color that the underlying population is distributed uniformly across the polygon which reality that's not the case. So, as you can see here, um this is uh some aerial photography and we've got these black boundaries shown on top. Um and the population is really only in the sort of like the the upper areas of the uh polyon around here, but there are areas of like green space for example and areas of industry where there's no one living. So um the chlorophy map as is can be a bit misleading in terms of like because the data is not uniformly underneath it. So as we'll see at the end of the presentation there al alternatives the chlorof map which have become more popular among some of the sort of um ways of visualizing the census data. So just to be aware of the this limitation. Um the cloth map is still great though as a way of like getting a quick insight into the how the census variable uh varies across the entire data set and to make quick cap comparisons between a small particular region of interest and the UK as a whole or whatever. Uh and it's a useful skill to be able to create cluster maps in cuis. And so that's what we're going to do um over the next 20 minutes or so. So what on the right here is a PDF that I've created in QGIS using all the census data we downloaded this morning. So this is showing the percentage of households in leads living in flat as recorded by the 2021 census. Um and if you're following the workbook, you'll be able to create this PDF yourself and you will also be able to do this. uh and the idea is that then you'll be able to take you'll be able to download any other data set from the UK data service data explorer and then build your own uh PDF maps for any other data set that's available in the UK data service by doing the same sort of things and those same transformations in Excel and those manipulations and so the components of a glorified map are the map centers variable um so those stats from data explorer uh the centers boundaries that we download from boundary data selector and making a linkage between the stats and the boundaries doing a cloverhead map classification. So we have to tell CQIS how it should shade each of those polygons according to the sentence variable and then a color ramp. So what sort of should we use blue to display the different um values or red or green or whatever. And then as I said before we always have to include a data attribution statement because although the data that you download from the UK data store is open access there are conditions of use of using that open data. one of which is you attribute the provider of that data. So this is why on my um uh PDF here and as well as all my maps that I'm showing in these slides there is an attribution statement at the bottom that tells you the data came from national statistics under an open government license and it contains ordinance survey data of the because the ordinance survey are the ultimate source of those census boundaries and there's a crown copyright database right copyright statement. So, and so when we talk about linking stats to census boundaries, again, just to reiterate, we have our census stats download from data explorer. Um, they'll contain like this geographic identifier column, which is those nine-digit codes that tell you within each row what is the unique ward. and then a bunch of census boundaries that have the same codes. And because they're the same in each case, we can link the data to one another. Um, within the census, um, they should be fairly consistent between data sets. Um, if you have other nonsensus data, so you might have random neighborhood statistics data that was produced outside the census that you might say be trying to relate to electoral wards. You can often find there can be slight differences as the geographies have changed and so you might always you might not always find there's a onetoone link between your data and the particular geographies. Uh so you just have to watch and make sure there is a onetoone link and so I'm a hands-on demo of just um in QGS making that join from our census sat to our census boundaries. So minimize some of these windows back to CQIS. Um again I think CU just it might be quite hard to see what's going on here because I've I've got quite a high resolution screen. So, so again, as I said before, we've got the layer panels here, and I can select the polygon layer. So, these are electro war boundaries. Right click and go to properties, and it will open that layer properties. And one of the properties is joins. So it tells CQIS what data is currently joined to the polygonal layer and what data do we want to join to the polygonal layer if there's none currently added. And we want to join those census stats to the census polygons. And so we click the add button and it will add a vector join. So I can just make this window a bit wider so you can see what's going on. So, it's pre-selected that this accommodation type leads 2021 because that's our CSV file. Um, and it's it's been clever. It's just picked the first column in the CSV as the GR identifier, which is what we renamed to Ward's GUID. The target field is the field within the actual polygonal data set itself. I that shape file which also contains those geographic identifiers. So if I select the dropdown, the one we want is W 2022 code and if I do an okay and apply and okay. And this time again in the layers panel right click and go to but this time go to open attribute table and this will show our polygonal data set. But this time it's we now have all those um sets of stats added to the polygons. So that's again our lovely um raw counts of the households, our lovely proportional values and our dominant accommodation type. And again I can use the identify button to select click on one of the polygons and it will show all the properties of those center stats on each of the boundaries. The identify buttons a nice way you can just quickly step through the data and relate the plug polygon to the actual sensor stats. Um, so we've got our sensor stats attached to our boundaries. Before we actually do the chlorop happy in CQIS, we just go back to slides for a while. Oops. Yep. So we've done our table joint. We start that's attached to the boundaries and we're just going to consideration some of the chloride mapping choices. So we have a [clears throat] so we have the choice of the centers output geography. Um in our case we've already made that choice by picking the electro wards um the choice of color map classification method and the choice of the cloth color map ramp. So as said before the sense output geography is available at different levels of output geography [clears throat] but disclosure control means that not all the sensor variables are available at all levels. Um [snorts] and the point so this is what you can see here. Uh well these are this is the same sentence variable displayed at different um output geographies. So we have sentus output areas here um lower layer super output areas and middle layer super alput areas and and the thing is if you analyze the same assuming the data is available all levels and and it's not been imp impacted by disclosure control not making available uh you can get different insights into the variable by looking at at different levels of geography um at the output area level it might be quite um um noisy whereas But as you zoom out it become the data tends to get smoothed and you see you will see more regional patterns and and so you will therefore produce different types of chlorophy map at the different levels of geography. Um and that's just a case of finding um the level of mapping what's most suitable for you and also you may be looking at bringing in other data sets. So maybe environmental information which may only be be available at particular output geographies because so you might have like um pollution scores or something which may only be available at super output area level and you're bringing that into add context to the census data. And in terms of um what a chlor map uh actually is it's doing some sort sort of data classification. It's it's simplifying the data um so your raw data itself. So um where's it going? Uh let me just find that spreadsheet. Uh oh, we've got I can open this. Yeah. Yeah. So these are all just um one of the columns is a whole variet what the 33 layers 33 records is a whole range of values and there's probably a unique value per per record. And so when we create a cloth map we're going to simplify the data by arranging into five um groups or seven what they call data classes just to simplify the data. Um, so if I go back to the slides, if I can find the slides. Um, yep. And so we do a sort of data classification. [clears throat] We take the full range of data values and we classify into a set of what they call data buckets. And each data bucket has a we set a minimum and a maximum value for the values within that bucket. And you can see that what's going on here. So down the left we have the range of values say from 3 to 93. And then what we do we have a in this case I've set five buckets. And for each of those buckets we set a minimum value and a maximum value that tells you when that data set value is within the bucket. So in our first bucket have a it will take range of values from 3 to 20. The next bucket from 21 to 38, 57 to 74 and then a 75 plus. And then so all the records which [clears throat] are values between 3 and 20 will get put into this bucket. All the records values between 61 and 64 will be placed into this bucket. And this is a simple what they call an equal equal interval classification because each of the buckets is the same size. And so they all go from 3 to 20. So they all go they all cover 17 values in my case from 3 to 20 or from 21 to 38 or from 39 to 56 or 59 to 74. There are different classification methods and and it's the way they create those buckets which is a different each classification method. In some the buckets may not be the same size because you they may um where you have more data you may have more buckets or buckets are smaller or have different minimum maximum values. Um and so this is what each of these is what they call a different um chlorop classification method. So that equal info one I showed there and there's a whole bunch of different ones and they basically are just controlling how those buckets have been formed and how the data is being applied to those buckets and sort of the how how they're basically simplifying the the range of values in the data in order to make the data more usable or um and so chlorophy map which we've applied some data classification and then we've got five buckets and and again with a minimum and maximum value. So and apply to some color ramp in this case an orange one. So low values the lowest bucket values have like a a very white whereas the largest value bucket is quite dark and again this this thing you can take the same data set and you can choose a different classification method and you will produce a different sort of map. So it's just to be aware that if you pick a different classification method with the same data, you can produce a different kind of map and you just um and the and and you can also when [clears throat] the number of buckets you also pick also produce different types of map. So if you have a a a classification with five buckets, you will get a different looking map than you have one with three or seven. So that's what you can see here. So we have the same data set. We're using the same sort of classification method. I think it's probably equal interval, but we're varying the number of buckets used between three, five, and seven. And when you do that, you get a different sort of map um just because the way the data is being like allocated to those bins. And then the output map gets shown. And on the left, same deal. This time we've got the same number of buckets of five, but we're using a different classification method. Um, and again, you get a different sort of map depending on the classification method. And the point is that no classification method is right or wrong, but you just have to be guided by what your actual data is like in order to pick the classification method to use and the sort of map you want to construct. And the idea is not to create to create a misleading map. And we'll see that in KU just how you can actually look at the underlying data to see how it's been classified. And so that's the classification process. So that's telling us how to bucket the data. You obviously then want to pick the actual color ramp to use. Um so sequential is what you normally use for color maps. So that's from light to dark. So you have a green a sequential color map from light green to dark green. Diverging would be case if you're like um it might be data around a mean around zero. So you might have minus values to the left and plus values to the right. So increasing plus to green to light to dark green and diverging sort of light brown to dark brown. And a qualitative is a little more of a categorical map. So this could be a land cover map. where you just like um just light blue for the sea, green for forests, orange for mountains, red for buildings. And so go back to curious and we'll actually get on with actually creating some colored maps. So, so I'm going to first create a categorical map and that's going to use this um semi- detached column here. That was the one that tells the dominant type accommodation. So again left click the layer into properties and this time you want symboli option and it's like the symboli controls how the map is being displayed by cuis and the minute it's showing a single symbol. So it's defaulted what it defaults to will probably will change um each time it just picks a random color and in this case a single symbol. So it's I'm drawing all the polygons the same the same single symbol in this case orange and I want to change this which I use here to the different types of visualization and I'm going to first use what we call a categorized and I'm going to select that categorical data we created our dominant accommodation type and hit the classify button QS will look at the underlying data and we'll create a new um color class for each of those the the unique strings essentially within that column. So our semi debt or flat terrace whatever um which you can see is what it's done here within QIS I can just um it for it always adds in all of our values because in some people have data sets which contain data with null values our case we don't need that because all our records are populated so I'm going to get rid of that so you can just select the all values and do delete Um and so the values here are those that's the actual strings shown in that um uh dominant accommodation type column. Uh a legend is on a map. It tells you what the map is actually showing. It's like a label of each of those classes. At the minute it's just used the value, but I want to make that more reasonable. So you can just in the QJS you just double click under legend and then you can edit what it actually shows as a label for the class. So, I'm going to put detached houses and flats, semi- detached houses and terrorist. How was it? And you can also reorder things here. So, if you just select one of them and then just move them around. So, uh let's put detached houses, semi- detached houses, terraces, and flat at the bottom. And we also change the colors for each of the classes. So, so detached houses, I don't know, we go to dark green. Um, semi- detached, light green terrace is a lovely shade of orange and flat. Let's go for like sort of pinky color and just apply that. And so now Cugis has redrawn the data using that categorical uh classification. And you can see that in leads in our electro wards uh as you might expect the do within the inner city the dominant accommodation type is is either flat or terrace houses reflecting the urban density in the city center. I can make this more readable if you go to the the layer properties and then you can change the opacity from 100% to something less than 100 to 80%. And what this mean this will make the the the boundaries sort of semi-transparent. So you'll be able to see the ordinance background mapping shining through the data. And that way it's a lot easier to see the lead city center and the suburbs. And you can see in the suburbs the domination type house type is detached homes and semi- detached house just reflecting I guess the uh lower density areas of the city. So that's our categorical map. Um, if you want to create an actual chloroplast map now from the actual one of the census variables and we're going to show the percentage of data of people who live in flats. And so the same deal, leftclick the item in the layer panel. So left so left click to select it and then right click and then properties. And again back to the symbology. And this time change from categorized to graduated. And again pick the column we want to use to draw the map. And I want flat but I want flat proportion. So this is our percentage value rather than the actual raw flat values. And I can just go away with the defaults and hit the classify. And CQIS has gone away and has applied one of those those those classification algorithms to bin the data into in this case five bins. And it set a minimum max for each of the uh classes that that determine um which records get values get attached to each bin. And what you can look at is the underlying data. So the histogram so the classes tab click on histogram and click on loads value and it will show you the distribution of your underlying data and this can help guide you in a sort of um classification method to use. So is your data all is it mostly in lower values or is it upper values and then that can help you define I'm going to pick a natural breaks one classification and I increase the number of classes to seven and you can see that QS will re um calculate the class breaks and I'm also going to change the color ramp from red to blue. like so. And let me just apply and okay. And you can see that QGIS has now redrawn what was a categorical map into a univariat chlorophy map. It's univariate. It's a single variable being mapped. And you can see that perhaps unsurprising in leads um the wards with the highest proportion of folk living in flats tend to be in the city center or right in the safe city center and that lessens as you go towards the suburbs of the city. What we can now do is create [clears throat] that PDF in CQIS. [cough] And within CQIS, um the way you create a PDF output, which you might say want to incorporate into a a a document you're producing or send to someone an email is a thing called the layout manager. And to access the layout manager, you go to project and layout manager. And you just accept new from template and click create. And you have to give the layout a title. So I'm going to give it the title of my census map. And what pops up is a bit like a if you've used PowerPoint, you get a black canvas in which you can then sort of drag elements on in order to create your print layout. It defaults to landscape. I want to be a PDF. I sorry, I want portrait. So I go to layout properties, page properties, rather sorry. Page properties. Yeah, layout page properties. So layout page properties and change orientation landscape to portrait. I'm just going to maximize this the layout window. Um a bit oops. Okay. And so down the left here you have different elements you can add to the layout. So I'm going to first add a map. So which is this one here which is like a add map and you select the button and then drag onto the canvas and cuis will add to the layout what's currently shown in the map window of the main cuis application which in our case is a univa chlorophy map overlaying against a backdrop or in a survey map and then we can add again you can resize this and move stuff around Um, one of the first elements we want to add is an actual map legend so that we can because if we just gave this to people, they wouldn't understand what the different shades of blue actually mean. So you have to give them a a legend and a map legend just tells you what the actual map is showing. Um, so we use the add legend widget. So again, we click on it and then drag a legend. And because QJS is currently showing a raster data set, it's added all the each of the colors shown on each of the raster background as a separate thing, which is really not what we want. So we can edit edit this legend to remove the the legend for the raster just to so it just shows the class breaks for the chlorophy map. And so to to do that under the legend section right here, we go to legend items. It's currently set to synchronized to visible layers, which means whatever you see in the map layout is is slayed to what's showing in the main QS window. And so what we want to do is change that to manual. And that means we can now select the raster one and click delete. It will still keep the raster map shown in the window, but it just removes all that um the actual legend for the raster, which makes our legend a lot more useful. And now we can actually go into legend itself and and start editing this because at the minute it says 3.4 to 618, but what well it's actually flat, but we need to tell the users of the map that. So you can just double click one of the legend items and then you can just edit this. So percentage fats percentage C and I can do that for all the rest of them. Oops. Yeah, you get the idea. So, uh flats flat That's good. Oops. And you can also double click here and change the actual title from at the minute it's just showing the name of the actual underlying boundary data set. We can change this to um 2021 electoral lords and the legends updated and we can move the legend around. And so now a user of our PDF can actually tell what's on the map. Um we can also add stuff like a scale bar. um not terribly useful because we're not going to use this map for navigation, but it still might be useful because by providing a scale bar, we can the user can tell how big each of these actual ward polygons is in the real world. So, in terms of kilometers, we have a north arrow just to help the user orientate themselves. Um, and then we can add some text. So, the first text we're going to do is add a title to the map. So, maybe that was a bit quick. So, again, from the left, you just click the add label and then drag a new label onto the canvas. Oops. And then you can write the canvas the text. So, I'm going to go a map showing census data for leads. And you can change the font to make the text a lot bigger. I might have to drag this out again. And maybe I can make that uh bold. And as I said before, an important thing we have to add is a copyright statement. And again, we just add some text label. And if I say go back to the for our downloaded census boundaries into that terms conditions document and just grab the attribution statement which is this one here. Just copy that into the label. At the minute, it's got the year as a placeholder, and we want to update that to this year's current year, which is 2026. And again, just increase the font. And so this way, we now have an attribution copyright statement for the data we're showing. And so there's various sort of things you add to the layout. You could add some decorations. I don't know, like a big star or something. That's not particularly useful. Um um but again, in the CQIS documentation, there are various guides that tell you how to create nice looking layouts. And so what you can now do to export as a PDF is use the button along the top called which has got like this the Acrobat symbol on it and just export as PDF and tell you where you want to put the PDF to. I'm going to just put that onto my desktop. My sentence map maybe of leads and do a save and keep defaults are fine. If I go to my desktop, let's just minimize Q just for a while. And here's my PDF. So, here's my PDF of um leads that I create in QJS from census data downloaded from the UK data service. And I could just email that email that to someone or stick it in combine it in a a document I'm creating. And again, you can do the same for any sort of census data you download from the UK data service if you follow this process. So I'm going back to some slides I think and we're just going to look at some other types of sets map other than univariate and chlorop and categorical maps. So let me just make these full screen again. Um so some other types of sensors map you might see or indeed you can create in cougis are bi chlorophy maps and ctograms and there's also some sensors maps called asymmetric chlorophy maps which we'll look at some examples of later but you can't actually use you can't actually create easily in cis. So the type of chlor map we were showing and which we indeed created in CUGIS is a univariate chloroplith map because we're only showing a single variable at one time in our case like accommodation type so the potential flats. A baricop map on the other hand will show simultaneously two variables at the same time which [clears throat] can be really useful for census data because it means you can create some some more interesting maps. [snorts] Um and that's what you can see it right here. So at the top we have these two univariate centers. That's unariatified maps. Well, one for poor health and one for unemployment. So in poor health um so low percentages of health poor health are in gray and as poor health increases it's um up to dark pink unemployment so low unemployment gray high unemployment um dark cyan um those two individual maps but if you combine them and create a third map so this time a biol map we're showing the two variables at the same time. And so what you end up with you is like this sort of thing where we get combinations of purple and cyan. And so where on both unemployment and poor health is high, you'll get a very dark shade of blue. Where they're both low, you'll get gray. And then the intervening things, you'll get different combinations of cyan and purple. And this is like a in so in one in one single map we're showing two variables at the same time. Um which can be useful for census data especially multivariate census data because again we can show multiple variables at the same time. The key thing of vari maps though is to make sure you pick variables that are related because otherwise you could create some quite misleading looking things. Um a carttogram on the other hand um in a normal chloroplith map we simply keep the say the electro ward boundaries how the shape of them as it is and simply color those polygons according to the sentence variable. Our cargram is a special for a map where instead um we actually distort and reshape the census boundaries according to the variable itself. [snorts] And this can be really useful like if you're trying to like um if you got the case of of Scotland for example. So these are showing Scottish council areas and as you can see in the central belt the the council area for Edinburgh is very quite small whereas the council area for like um the Highland Island is massive. Um and so you can get if you just show it as a chlorophy map the actual underlying geographies are quite hidden and so instead we can do create a a ctogram which will distort or reshape the actual polygons according to that census variable. So that where the sensor variables are high the the the polygons will be distorted more and where it's lower they will get distorted less. And so you end up with something at the bottom here where the underlying um polygons have been distorted. And so these large the the central belt polygons have essentially ballooned up and have become a large bigger whereas the the the highlands and islands which are like lower populations have been distorted less and it's a different way of like viewing the data. Um and the important thing in in this type of carttogram is that the the actual topological relationships between the polygons has been maintained. So that um so in this case selfirer which is down here is still neighboring easter and include and that's the same in the output ctogram. so that you can still relate neighbors to one another. And so ctograms are quite feature in the wild for like um visualizing social economic data. Um the guardian features carttograms in its um reporting for the EU referendum what 10 so years ago which we're still living from. Um and also they featured prominently in this book called people in places which is a really nice book which is visualizes the 2011 census data and it uses all that the data is all visualized as um these area ctograms uh it's a really nice book um um and so what I'm going to demonstrate now in QIS is how you can create a car yourself um of the sort we saw there. Um and also how you create a bicaro map. I have to caveat this um as we'll see my version of CQIS will not currently allow me to create the actual ctogram. Um because I'm using a Mac and it's a new version of CUGIS and a new Mac and there's some issue with my particular install. But I will definitely show you the tool and I'll be able to show you the cart that I pre-made. The last time I ran this workshop three months ago in March where it ran. So um and hopefully in the next work time we me run this workshop in three months time I will have come with a workar around and I'll be do a actual live demo of using this particular car construction but the minute it's going to be a bit blue peter here's one I did earlier. So um let me just go back to CQIS and we'll certainly look at how to create the biaric chlorophyll map in CQIS. But again the the ctogram will be a pre-made one. So apologies for that. So let me just remove the layout. Go back to cugis. Um to create the by variant chlor uh chlorop again we need a data set that has two types of variable in accommodation type and tenure type which are not in our data we created. So for the purposes of this exercise I created another data set which is in that data pack that you could have downloaded from the GitHub repo that we sent you in the joining instructions. So I'm going to navigate to that data set and add the cu just now and it's actually a data in a geo package which is a different type of data other than a shape file. So so here's my data pack. The directory search here should be the same as what you download from the GitHub repo. And so what I have to first do is add to CQIS the geo package. Um it's a geo package. A geo package is a special type of geospatial data set. In some ways it's an advancement or a shape files in that you can have a single geo package that can contain different layers of data. So you add the geo package and then you connect to it and then you can add the data set. And again it's it's for leads. Um and if we open the attribute table you can see we've got a different types of column again which are proportions. So the AT means accommodation type. Um the 10 is the tenure type. So basically we have those electro awards for leads and a bunch of columns giving sensor variables on accommodation type and a bunch of columns giving colum uh variables at tenure. And we're going to basically in CQIS create a by variable map that that shows um one of the accommodation type variables and one of the tenure types variables together. And so in your um workbook it talks about how CQIS is really nice. It's an open source piece of software and we can expand its functionality using a thing called plugins. And these are like community developed enhancements to cues that people have added to increase its functionality. And it just so happens that one of these plugins is a one that allow us to create by variate chlorophy maps because by default QIS will not allow us to create biariate chlorophy maps. Uh so we use a plug-in to do that. And there's also an additional CQIS plugin the one that I can't currently get to work which will allow us to create carttograms. And so to enable these and this is explained in the workbook. You go to the plugins menu and you can tell what plugins are currently installed in QGIS. So you can see I've got the by variant one installed and the curs one installed. And you can also search for other plugins. And you can see there's a whole load of them. there like the thing is curious if you discover it can't do something you can normally find a plugin that will do it for you. Um and so in the case of the by variant plugin again I'd go to the properties and go to that symbology type it adds a new option to the symbology type of the bariate renderer. So rather than categorized or graduated, I can go to biariate renderer and similar sort of thing that was on the slides. So we have to tell it what of the two columns to use to construct the biariate chlorop. And I'm going to go for accommodation type of flat and maybe private rented flat tenure. Now we click apply. QS will classify the data according to that biic classification and you can see it's again it's it's created a new type of legend for us which is like these um rather strange looking like um grid representation. So again, where both accommodation flat is high and private rent is high, you'll get blue and you'll get gray when they're both low. And you got the intermediates for the different types of uh variable. And again, you can also change the opacity to make the OS data shine through and give more context like so. And then when I to create a carttogram using that cartgram plugin, QJS will add a plugin a cartgram button to the main toolbar called compute cartgram. So you would click this and up would pop this UN UI and you basically tell the cartgram plugin which column to create the cartgram from. And so you might go for like um let's pick actual correct death uh so flat and you hit okay. If I hit okay here I'm fairly sure CQ just will probably dry to a halt and then will die which is not good for a demonstration. So I'm not going to do that. I am going to add the d the cam I created on my other Linux machine which I know worked. So let me just add another duo package with this um uh with this card I created before. So and so that's that's what the the the cartgram plugin will create is a ctogram. So as you see it's distorted the electro wards according to the accommodation type variable and again you can shade and that just gives you an alternative way of creating a cartgram in cuis. So that's the demonstrations of everything in the workbooks of how you manipulate data and create a univariatic chlorophyth map, a categorical map, a ctogram and a biiclorith map. Uh I've got some final slides and then we'll maybe go through some questions and then we'll bring the workshop to a close. So as I say, I said one of the problems with the uh chlorophyth map is the fact it implies the the population is uniformly distributed across the extent of the polygon. Um, an alternative to the chlorophyth map is what you'll see are what these things called asymmetric and mass chlorophy maps which are an alternative. Um, and what they do is they basically use like a an additional layer rather than the sort of the boundaries such as a buildings layer and then you basically shade that buildings layer by the census variable. So you get a much more realistic um distribution of the population shown. Um and these are used a lot in they've been used a lot for um by the ONS um other um organizations to display 2011 and 2021 census data. So there are two examples here. So this work is really pioneers by the focus college London and this data shine um service and that's what data shine is showing it's basically rendering instead of the the the sensor data as a a uniform color map it's rendering using a buildings layer and that and that's what you see on the left here. So it gives you a more realistic um looking map. So for leads you're so you get where the areas of like green space and stuff are not being shown or in industry when there's no buildings and I say the ONS have taken inspiration from this and I've and they use it to create um census maps in their census 21 mapping software which these are both online you can look at um and they use a similar sort of visualization there. Um you can create these things in couges but it be it can become quite involved um because you have to obtain a sort of UKwide buildings layer which can be quite large and then do all the sort of like um geometrical operations and and processing to create the maps yourself. Um um so it's just to help you if you have questions about census data or how you in the future um within the UK data service there's a UK data service training program this workshop is one of them um we will be repeating this workshop in three months time so if you maybe are just getting started you you try the work back book afterwards and then you discover issues and you want to get see the presentations again then you can certainly welcome to join that again that'll be mid-occtober um within the UK data service we're developing a set of learning modules which will cover um different ways of using census data um some of the background of how the census data was created um looking at some of the issues how you manipulate the census data there is a UK data service help desk where you can ask questions you have and as I say because the UK data service set support is distribution among different organizations we have different areas of expertise. So my particular expertise is in is the actual is like um the census geography and how you compare census data over time and cope with changing geography. Um how you do a lot of geospatial data manipulation. Other guys in Manchester are more um knowledgeable about the actual set of stats themselves um and how how you pick the right sense of stat for your particular research question. But if you send a query to the help desk, they will do some triaging and will make sure it gets rooted to the most um the best person within the centers team to answer your specific Sir,