Video summary
James Cron from the UK Data Service leads an instructional session focused on utilizing QGIS to visualize 2021/22 census data, specifically targeting the creation of univariate and bivariate choropleth maps alongside area cartograms. The workshop begins by guiding users through the process of downloading aggregate statistics for specific geographies, such as wards in Leeds, using tools like Census Explorer and the Boundary Data Selector to obtain shapefiles containing both electoral boundaries and census counts. A critical initial step involves transforming "long form" tabular data into a manageable wide format via pivot tables within Excel or QGIS, followed by normalizing raw population counts to generate percentages and adding categorical summaries that identify dominant accommodation types for effective mapping preparation.
Once the data is prepared, the session details configuring QGIS layers including raster backdrops from Ordnance Survey, vector boundaries, and CSV statistics joined through geographic identifiers. The presenter explains fundamental choropleth principles while addressing limitations regarding assumptions of uniform population distribution within administrative polygons, demonstrating how to classify continuous variables into discrete buckets using methods like equal intervals or natural breaks with appropriate color ramps. Advanced techniques are then introduced, including the use of plugins such as 'by variant' and Cartogram; however, due to specific software instability on certain setups, some demonstrations utilize pre-configured packages instead of live generation. This segment highlights how bivariate renderers can display combinations of variables like accommodation type and tenure through grid-based legends with adjustable opacity for underlying map layers.
To overcome the inherent inaccuracies of standard choropleth maps caused by varying population densities within fixed boundaries, the tutorial introduces asymmetric mass choropleth maps as a superior alternative. These visualizations utilize building footprint layers to shade areas realistically rather than filling entire administrative polygons, offering a more accurate representation pioneered by institutions like the London School of Economics and adopted for official census data releases. The session concludes with practical steps on exporting final visualizations as PDFs using QGIS Layout Manager while ensuring proper attribution is given to UK Data Service sources. Additionally, attendees are encouraged to explore further learning modules from the training program or contact the dedicated help desk service, where queries regarding complex topics like statistical disclosure control and geographical selection can be directed to specialists based on their specific areas of expertise.
Read the full video transcript
Okay folks um welcome to this using QST
map 202122
census data from the UK data service. My
name is James Cron. I work for the UK
data service up in Adena in Edinburgh.
Um so the aims for the session today are
we're going to use the QGS desktop GIS
application to map centers data down the
original UK data service. There's three
specific aims. We're going to create
different types of mapping. a
univariatic chlorophyth map, a bidic
chlorophyth map and an area carttogram.
As you see this is a journey. So the
objectives are to create those different
types of visualization.
And along the way we'll learn various
things about the UK data service, how
you use the data and how to use the GIS
application CUGIS.
So we'll first learn how to download
data from the UK data service. That's
the actual census data itself and some
census boundaries from different parts
of the UK data service.
Um we'll learn out how to carry
different types of data transformations
on the data and this is to prepare it
for mapping in CUGIS.
Um we'll learn about some of the issues
to be aware of when mapping census data
to make sure we create valid maps that
are not misleading to our users of our
maps
and then the we'll look at how to use
actual QS itself and so if you've not
used QS before this will be an
opportunity to find out what QJS and
what you can do with it and at the end
there'll be a short section where we
just go over the sort of resource
available in the UK data service to help
you at using census data and mapping
that census data. So in your joining
instructions, you will have proided with
a link to um a GitHub repo
which has a PDF workbook of all the
steps I'm going to demonstrate during
the demos today.
The workshop's going to work from the
basis of I I will give short talks which
will talk about the fee behind some of
the things and then I'll do
demonstrations
using the UK data service website and
the tools or using some QS software or
doing some data manipulation in Excel or
Libri Office. I'm on a Mac and I've got
Excel. So I'm going to do my data
manipulation using Excel but you can
also use Libri Office Calc. And the
choice is you up to you. you either
follow just what I'm doing on screen and
try and do it yourself on your own
machine or you can just um watch the and
then maybe do later after the workshop
because I say you'll have access to the
workbook so you can always go through
afterwards in your own time. So I'm just
first give an introduction to the census
itself.
So the UK census is a denial census and
it happens every 10 years.
um within the UK data service, we mainly
support the census from 1971 through to
2021 and 2022.
Uh in 2021,
uh 2022 uh in Scotland because of COVID,
they did it a year later in 2022.
Whereas in England Wales, it was 2021.
So that's why there's a a difference
there. Um
the census itself
um in the old days it was done by um
people would get a paper form for the
post or a numerous would come around and
ask some questions. Um it tends to be
done online now but the the different
questions are there are household
questions and there are questions about
individuals. So the household questions
will ask people stuff about the type of
type of house they live in
um the rooms in the house and stuff how
they how they own the accommodation like
are they renting uh are they mortgaged
or they are they renting from a social
landlord. Whereas the individual
questions are questions about each of
the individuals within that
accommodation. So it could be age, sex,
what sort of job they do, how they
travel to work, etc.
So on census night which is usually in
April and the last the last one in
England Wales was 2021 Scotland will be
April 2022 the population supposed to
fill in their census forms and submit
the results to the census agencies which
in England Wales is off national
statistics and in [snorts] Scotland
there's national record of Scotland and
I say this is just some more the topics
covered by the UK census.
So you get this question related to
education, housing, health and language
or transport.
Um and during the sort of demonstration
of practice today we're going to focus
in on housing data and we're going to
look at accommodation type and tenure
type. So sort of are we looking at is
there a flats or houses or semi-
detached and is it rented privately or
is it socially rented. So when the sens
agency gets all the data
um after census night they do a lot of
processing on the data um so they they
bring all the data together and they and
they produced output uh sent us
aggregate data sets and these output the
different variables as either counts of
people or households.
Um and that data that centers accurate
data is produced at different levels of
output geography and the smallest level
output geography is a thing called the
census output area
and the output areas are specially um
created just for the census
and this is an example of an output area
for part of Edinburgh. So
um each of these um black polygons is a
sentence output area and you can see
they cover a few houses or streets
within a neighborhood of Edinburgh
and the actual um sentence variables
themselves. You can see there's an
output area code which is a unique JRock
identifier which unique identifies a
sentence output area and then you have
various statistics related to that
output area such as total population in
this case the number of households the
number of males females and is that are
there how many detached or flats within
that area and there's different types of
sentus aggregate data there's univariate
um uh aggregate data and is multivaried.
So the univariate census data just
contains information about a single
census variable
which could just be like accommodation
type or tenure
whereas a multivariate data um basically
um gives you multiple um
topics for each um data set. So it would
give you accommodation type but also
tenure combined and this gives you a
much richer insight into the
information. So you could take a
univariate um census data like the
method used to travel to work which
could be um by car or foot or train and
you could add additional variable which
be how um the occupation of people are
doing which would give you a multivaried
data set and that way you could um do
some sort of classification of different
types of um people's occupation how
they're traveling to work. So how are
people who um
do skilled jobs traveling to their work
as opposed to people who might be um um
uh working as in um building occupation
travel to work so that you might make
some differences there. So multivaried
data sets are a lot richer
um whereas in this workshop we will
simply be using univariate data set
because it's easier to manipulate.
But for the UK data service you can
download both univariate and
multivariate census data.
So the UK data service itself
um provides access and trading to u
economic and population social research
data.
uh and the emphasis here is we provide
access to both the data itself and the
training around that data. So it's
access to the data and then how you
would use that data and there's an
entire training program of which this
workshop is one part of
and within the UK data server itself is
a specific part of us that do that
purely support the census and we are
called census support
um and we're distributed along census
support is um uh distributed along
different institutions within the UK so
I'm up here in Adina
and we support the census boundary data.
I have colleagues in Manchester who
support the aggregate data itself and I
have colleagues at um UCL in London who
support the census flow data and Jill is
based at the Kathy Mars Institute in
Manchester who support the microents
micro data and today we're going to look
at using centers aggregate data and
centers boundary data and each of our
different um subsets of centers have
different applications in order to
access different types of sensor data.
So in terms of the census aggre data
tools, this gives you access to raw
census stats themselves. SECAN will give
you bulk access to the census tables. Um
so you can download the census table say
for all of the UK or all of Scotland or
Wales, but there's no ability for you to
then to to subset that data. So if you
only wanted data for Glasgow or say
you would have to do that subsetting
yourself perhaps in spreadsheet or a
database application. Whereas a data
explorer is a second UK data service
tool which allows you the ability to
create extracts from the complete census
accurate data sets. So they would allow
you to just restrict your census data to
Edinburgh or Glasgow or Lurra and to
pick the actual types of the variables
within the tables of census data you
want. So it means you you can be a lot
more um specific of the sensor data you
want and where you want it for. Data
Explorer currently provides access to
2021 and 22 data and I think it also
gives you access to 71 data and we have
a bunch of legacy applications that
we're in the process or sorry the
engineers in Manchester are in the
process of translating or migrating that
data from the legacy platforms into data
explorer. The ultimate aim is eventually
all the census data will be available
through data explorer and you'll be able
to use the same tool to access data from
71 81 91 through to 2001 and 2022
possibly 2031 in that census happens. So
I'm going to do a first demonstration
and I'm going to use the census data
explorer to download 2021 census data.
So let me just um minimize my slides and
I'm going to open a web browser. So I'm
going to first go to the UK data service
and show you how you access data
explorer. So
this is the UK data service. Um
we have a data catalog. Um
which will allow you to search for data
sets across the entire UK data service
including the census.
But within the census itself, we have a
quick link to our particular census
applications. So within the data
catalog, if you click this quick link to
census, it will bring up this pop-up
window which lists our various census
data applications.
So you can access the data explorer.
Um well, this is the census boundary
here. This is another one we're looking
at today. Or census full data. And it
say infuse is a legacy census
application which gives you access to
earlier census data.
And I'd say you've got this CCAN which
gives you access to bulk data. But we're
going to use data explorer. So I'm going
to click on data explorer. So this is
what the data explorer tool looks like.
Um you can see it's got 71 data and 2021
2022 data.
I think it also provides access to other
non-sensus data. some data from World
Bank, but we are specifically concerned
with the census data. Um, these are
quick link buttons at the bottom here.
We can just click in these buttons and
it will take you to like all census 21
2022 data sets. Um, but for the purpose
of this demonstration, I'm going to
search for a particular type of census
data for 2021 and 2022.
I'm just going to refer to my um handy
workbook that I have given you all and
I'm going to look for 2021 census
accommodation type.
And so you can see there are 54 data
sets of 2021 or 2022 census data which
meet that term accommodation type.
And see these are all listed here. But
the one I want is this TSO44
accommodation type data set.
So
if I click on the data set we get an
overview. So we uh data explorer will
give us some metadata about what the
sort of data is. So you can see it
provides census data estimates that
classify households in England Wales by
accommodation type.
Um
and I say at a minute we're currently um
if we just if we hit the download button
now we would download the entire data
set whereas I want to split it down just
for data
just for leads because I'm going to show
you how to create a map of accommodation
type using of using a 20 on data for
leads.
Um
so to to um split the data up and refine
what we want to download there are
various filters are left here and I'm
going to set these various filters to um
filter the data. So
first top level geography I'm going to
set England.
Click apply and you can see it adds a
filter
under geographic grouping. I only want
to download census data for a specific
type of output geography. In this case,
I'm going to download the data for wards
because it's a relatively small
geography. Um,
so I can click on the wards and
divisions item and click apply. And this
this will tell data explorer to output
or census data by wards.
And having tell that I now to tell where
I want to download the data for.
Um so again at the minute it's selecting
data for all living whales. I want to
only download data for leads.
Um and what you have here is this um
this will basically show you all the
sort of um sub geographies of England
Wales that we can download data for. At
the minute it's currently quite
unwieldy. So what you can do is you can
do a collapse all and then we can like
drill down through the geography just to
restrict to leads. So
um
so I know that leads is within Yorkshire
and the Humber. So I can just do a drill
down this way. Um I believe Yorkshire.
We're going to check my PDF.
Being in Scotland I'm not particularly
familiar with Yorkshire but let's check.
Ah so leads under West Yorkshire. So,
and here's leads
and and so I want to each of these um
items here is an electro ward within
leads. And I could just manually go here
and and I click each one to add it to
the data selection.
And this is what there's a different
types of selection mode here.
But what I'm going to do is change the
selection mode to items and all items
directly below. And this way when I
click on leads it will select all the
awards within leads
and now I click apply and this will add
that to my data selection
and so so we've constrained to the type
of geography electro wards I'm constrain
just to lead and then the final filter
is the actual sensitive variables
themselves.
I'm going to want all these because
these are different types of
accommodation. So again I can change
selection mode rather than manually
individually selecting these. So items
and all types right below
and click apply. And as you can see we
now set up our filters
within data explorer. I can get a
preview of the data by clicking the
table button.
And you can see
we have geography for leads. So each of
these is a different ward and across the
columns we have the different types of
accommodation type.
And to download the data we can use the
download button here.
Um, the trick here is to make sure you
use the filter data in tags or text form
button
because that will output the data as
like a plain CSV that we can then use in
Excel and then later use in our desktop
application.
If we use the unfiltered data then
basically data explorer will not just
ignore all those filters we set and
we'll get the entire data set. If we use
the table in Excel, then we get a lot of
extra formatting information, which
might look pretty, but it makes it hard
to manipulate the data. So, I'm just
going to go for filter data in tabular
text format.
And you can see we've downloaded the
file. Um, a quirk of data explorer is it
creates a very a very long file name.
So, I'm going to make this more usable
by shortening the file name.
Just something a bit more manageable.
Great.
It doesn't like it'll make a lot of
difference cuz Yeah.
Okay.
Uh, okay. I'm just going to go back to
some slides now.
And so, we've downloaded our set of
stats from data explorer.
The next step is having download those
stats and we need to do some data
manipulation to prepare those sensor
stats for mapping intuis.
Um,
and there's like three things we're
going to do to the data.
We're going to transform the data from
long to wide form. I'll explain long to
wide form later. I'm going to do some
data normalization so that we can don't
produce misleading maps. And then I'm
going to do an extra bit where I add an
extra variable that summarized the
center stats to create create a
different type of map. Um these first
two are pretty much essential. We never
use data from the UK data service
because of the way the data is first
provided and then in order to to
normalize that data for mapping this
third thing as a cascar summary variable
is just an extra step that I'm doing
just so I can create a different type of
map. But normally if you're downloading
data data from the UK data service from
data explorer, you're going to have to
do a transformation to from long to wide
form and you're going to have to do some
sort of normalization on the sensarials
before you map in CQIS.
And it's a
I'm going to do a manipulation in a
spreadsheet because um spreadsheets are
great at handling like tabular data. You
can use like um Libri Office Calc or
Excel. Um I'm on a Mac and I happen to
have Excel. So I'm going to use Excel to
do my man manipulation. But you can use
Libri Office Calc. And in the workbook
there are instructions for doing the
same operations in either Libri Office
Calc or Excel.
And so this first transformation we have
to do is convert the the data downloaded
from data explorer from long to wide
form.
Um so data sets have different types of
shape which like um is determined by how
the data is arranged into rows and
columns.
Um so data explorer will provide the
data in what's known as a long form
and what you'll have is each record will
have a will will contain a different
census variable and um as you can see
here
um
so here we have three different output
areas. You can tell there's three
because each one has a unique um go ID
which is this EO5011 384 thing. And each
of the different sentence variables is
provided on a new row. [snorts]
So if we've got eight sentence variables
that for that one output area, it'll be
provided on eight records.
Um whereas what we have here is the same
data set but in the wide form and here
you get a single row but each of the
different census variables in is is then
is instead shown in a different column.
So we have again our same E05 01 on 384
output area but this time each of the
different columns provides each of the
different centers variables.
Um and so within CQIS, CQIS is a lot can
is a lot better at handling data in this
wide form because as we'll see later,
we'll download centis boundary data and
for each of these output areas, we'll
get a polygon which um which we'll
attach to this table and then we'll have
the census variables of one record that
has all the columns plus the output area
geography by row. as opposed to this. If
we were to add the col the geometries
here, it would make the table a lot more
complicated.
But as I say, because the data comes in
long form and we had to wide form, we
have to do some data conversions within
our spreadsheet just to make the data
mappable. And this can be done in a
spreadsheet application using a pivot
table, which basically just groups the
data by the unique geographical
identifier. And I'll demonstrate that in
a couple of minutes in Excel. Although
you can do the same thing in Libri
Office Calc.
[snorts]
And then before we actually map the
census data, we need to do some data
normalization,
the census data itself. So that's the
the counts of of population of of of
individual households or people or just
raw values.
Um,
and so if we want to map the data and we
want to say where we we're basically
creating a map that's comparing one area
to another, then we need to do some
normalization. Um, because otherwise we
just be showing the raw counts which
would not be very meaningful.
And so there are different ways of doing
normalization.
We can divide by the actual size of the
area itself. We should express the data
the value as a population density. So um
percentage of sort of people per hectare
or kilometer squared or we can divide by
the total population size which would
then express that as a percentage. So,
so the percentage of people within that
particular geography which live in flats
or rent from a private sector landlord
or travel to work by bike or travel to
work by um car and because it's a
percentage you could then compare that
small area to another neighboring output
area or against country as a whole.
Um
and again in a spreadsheet application
we can do that by just doing some sort
of simple um um data manipulation. So we
would have like the count of our private
rent here and we have the total and
dividing that to the the the private
rent by the total and then multiplying
by 100 would express as a percentage.
And then I say those first two are
essential in order to be able to map the
data. And this third thing where I add a
categorical summary variable is just an
extra step that I'm doing just to like
provide an extra thing that I can map.
And what this is going to do is going to
look across all the different sentence
variables. So in this case for the
different types of rent, mortgage etc.
And it's going to find the dominant one.
So which category of house type is the
dominant one for the output area. So in
this case it's going to be come out as
mortgage or private rented and that's
just a good way of summarizing the data
and say what's the dominant type of um
accommodation type within that output
area.
alternative. You could have like a
what's the the least dominant would be a
different one you'd add. And what I
would do again, I'm going to use this in
Excel and I'm going to do calculate this
column and then I'll be able to map that
and show a dominant a map of DOM
accommodation type.
So I'm going to have a second demo
where I'm going to demonstrate how you
do that preparation of the data that we
downloaded from data explorer in Excel
and I'll show you how we do the pivot
table, how we do the data normalization
and then how we create that categorical
variable
and let me just find the data set we
downloaded. So accommodation type leads
2021 census
and I say I'm going to open this into
Excel
but again you can use libri office cal.
I will just increase the size of this.
So
this is the raw data we pulled down from
data explorer.
Um there's quite a lot of columns here,
a lot of like metadata. So it'll tell
you
um
it will tell you the table data came
from.
Um the top level geography. Um
the important thing is the geographic
area one because that each of these it
will tell you what output area that um
record is is related to.
It will give you the localized name for
the output area or I think these
actually wards so the actual ward name
and then it will give you the different
types of accommodation type. So in this
case detached semi- detached terrace
and
the always value is the actual count. So
the actual census variable for those
different types of accommodation
and I say what I'm going to do at the
minute this is in long form I need to
convert into wide form. So I'm going to
create a pivot table in Excel
and you do this using the insert pivot
table menu. So
an Excel insert.
>> And so you able to maximize that
spreadsheet?
>> It
it doesn't really increase the size. Is
it still quite hard to see or
>> um I can see it but others might
struggle.
>> I think it's just the actual Yeah.
Okay.
Okay. Yeah. So, insert and then pivot
table.
And so, there's a pivot table field here
where we have to define how to create
the actual pivot table. So, we're trying
saying how do we want to aggregate the
data.
And so, what you do here
on the top here are listed the different
fields
within that spreadsheet. And there are
different boxes here which controls how
the pivot table is constructed.
So what you just have to drag them from
the top here to the different boxes.
So I will first grab accommodation type
eight categories and drag that to the
columns box.
Jog area that as I say is the unique set
of J identifiers and we'll grab to rows.
And then you want the actual observation
values themselves. So the actual counts
of the sentence variable or the values
and what you can see is it's aggregated
that data so across the different um
variables by the uh geographical
identifiers. So we end up with like one
row per uh unique ward and each of those
unique rows has the different columns of
the census data
and we get a grand total which sums the
values sum sums the columns and then
sums the records.
So
what I'm going to do here is then grab
this data.
So do a copy
and then create a new sheet and then
paste the data.
Enter that new sheet. Let me just
increase the size again.
And what I'm going to do now, I'm going
to rename some of the columns because
they're quite long. So this row labels
one I'm going to rename to
Ward geo ID. So ward geographic
identifier. This rather lengthy a
caravan or from a structure temporary
structure. I'm going to rename that to
caravan because
detached. It's probably okay if
detached, but I'll just make it detached
with a small D.
um [clears throat]
this very lengthy one. I'm just going to
change this to um comma b. This is all
um explained in the PDF.
This is just going to be flat.
Um shared
h
again another lengthy one. It's um I'm
just can change this to um s
C ch
or warehouse um semi- detached
will just become
semi
debt
terrorist again.
Minor tweak there
and grant total I will change to total.
Great.
So I say these are raw census counts and
we want to normalize each of them by the
total population so they get expressed
as a percentage.
So
I'm going to create another eight
columns
and each of these will become renamed as
prop
uh detached
prop
and you can do the same for the own just
so that it's then clear that we have the
data as provided as the raw count and
then as a proportion
and so to do the actual uh conversion
itself we enter a Excel formula.
So
we do the caravan one first. So equals
so the raw value which is C2
divided by the total
and then just multiply 100.
So that's expressed the value as a
percentage. So we've got 3,000 divided
by 9,000 gives you a potential of 32%.
>> [snorts]
>> and we want to use that same formula
across the other columns. And so to do
that,
we're just going to edit the formula
slightly um
and just anchor uh
so this way we can use it across the
other columns. So what I did there is I
just insert a dollar just before the
column
so that it still refers to the total
when it applies it to all the other
ones. So
if I just drag across
check,
let's just check.
Yeah. So, we've gone D2 / G2.
Yeah. Is that correct? D2.
No.
Okay. I got the first one wrong. That
should have been B2, not C2.
Okay.
Sorry about that.
Now, that's correct. So detached is C2 /
J2.
So comp is D2 / G2. That's correct.
Flat.
Uh well, we got E2. Yep. Divided by G2.
Correct. Okay.
Final check. Terrace. Um column is I.
Yep. So we've got I2id by G2. Correct.
Okay. And then we can just fill the
other records.
And let's just do a check. So
B31.
Yep. Divided by J31.
Yep. Correct. Okay. And then we just
populate the rest of the records.
And again, let's just pick one to
randomly check. So,
uh, semi- detached. So, column is H. So,
you got H23 divided by J23.
H23
/ J23 * 100. Yep, that's correct. Great.
And so that's expressed all of those
census variables as a proportion. So as
a percentage.
And so
we've done our we've done our
transformation of the pivot table. We've
done our normalization. In the final bit
I'm just going to create that
categorical column which in each case
will tell you what is the dominant
census variable for each of the words.
And again this is another um I'm going
to use another equation to do this. I'm
going to first give you a column name
for this and I'm going to call this
dominant
accommodation type and it's quite a
lengthy expression. So I'm going to grab
it from the actual PDF. So if you just
bear with me wh I look for the workbook.
And so this is the same PDF that you
will have
and this is the equation we're going to
use to do the cascarable summary
variable. So I'm just going to copy that
and put it into the cell.
So what you can see is the equation
that's told us for this record this ward
the dominant type of accommodation I the
one with the highest percentage is the
semi- detached which if you look at the
data is the case
and again I'm going to want to um just
to make this clearer I'm going to
replace the dash prop I'm going to get
rid of that in the actual text and I can
just do that by editing the equation
slightly
and wrapping of a substitute ute. So
just edit it and substitute
and I'm going to tell it to substitute
the underbar prop bit with nothing. So
and so you can see it just makes the
text a lot cleaner. So it's got rid of
this like the prop from the column and
that just tells that this the dominant
accommodation type is semi- detached.
And I'm going to I want to copy the same
formula down the rest of the columns. So
again, I need to edit the uh formula
just to anchor some of the the ranges
with the dollar symbol.
So
referring to the PDF,
I will add a dollar here
just to anchor what it's referring to.
And so that's still correct that one.
And I should now be able to copy and
paste to the rest of the records. And
this will now refer to the correct data.
So again, a quick check. So for this
board, it reckon the dominant one is
detached.
And if we look along,
yep, the dominant one is 41%. So that
seems correct.
So we've done manipulating the data in
Excel. If you're using livery office
calc
on the menu system
and I'm now just going to save this data
and so later we can manipulate it in
CQIS.
So
save as
comma separated values save.
I want to replace it. Yes.
It'll complain multiple things, but I
only want to save the current sheet
because that's the one I've been
manipulating. So, I can just okay that.
And if I just close Excel
and
go to wherever my data was.
I will open something.
Yeah. So this is our data. That's what
it looks like. We've got all our
preparation proportional value and our
dominant acceleration type.
So that's the sensors aggregate data. In
order to create maps, we also need to
get download some some geospatial
boundary data. I'm just going to go back
to the slides and do some more
talk about geospatial data.
So in terms of geospatial data um it's
used to model some aspect of the world
um
different types of geospatial data if
you never come across before vector data
which is um um it's like point lines and
polygons. Raster data is um typically
used to represent um like satellite
imagery,
digital train models. Um it's
essentially a grided data set. Um each
um cell could represent um an area of
elevation or like a say a part of a
satellite data set or it could be um uh
like temperature or water or something.
Um geospatial data is um what special
geospatial is is um spcially referenced
um
the data is provided for a particular um
um what they call spatial reference
systems. So here in the UK we use
ordinance survey data or data from the
or national statistics um it uses I
think called the British national grid
to locates that data within the UK. If
you're using global data for like say
all of Europe that would use like a
thing called WJ84
and that data um uses a different sort
of spatial reference. Um
there's a whole polar of different types
of geospatial data formats. Um there's
thing called shape files which are like
the sort of almost the Excel of the
geospatial world. Um we'll look at a few
of these formats as we go on. Um there's
a bunch of geospatial standards which um
define how spatial data is created and
shared.
Um there's an entire discipline called
geographic information systems or
science.
Um
and there are various um what they call
GIS software applications which allow
folk to create manipulate and do things
like create maps or do some data
manipulation or what they call spatial
analysis.
Um, a commercial well-known one of these
is called ArcJS.
And there's also an open source one
which we're going to be using today
called CUGIS.
Um, CQIS has been around for like quite
a while now. And despite the fact it's
free, it's incredibly powerful and has
some great features and it's ideal for
doing sort of census data. So mapping
census data
um in terms of the census output
geographies. So as I said before the
census out data is output a range of
different small areas. Um
so I say we download data at word level.
You also can get data at what they call
output areas or lower super output areas
or middle super output areas. Um
um and together these type of out
geographies form a centers geography
hierarchy from the smallest from the
output area up to LSO to MSOA to ward to
district to county to country to nation.
Um but the thing you have to be aware of
is there's this thing called statistical
disclosure control
which means that not all the census
topics or the census tables are
available at all geographies. So you
might find that there are certain um
possibly disclosive census variables um
relate which could only be output at
like higher geographies like um ward or
district and would not be available out
of area level which might um have an
effect on the sort of research question
you can ask
um
and this is just an example of what
these census boundary data is actually
is. So a census boundary data set is a
polygonal
uh geospatial data set that just tells
you on the ground what the
uh output geography is like. So on the
left here we have a count area boundary
for Glasgow. On the right we have the
much smaller cent output area
boundaries. Um
we also have a thing a concept or a
thing called a geographic lookup table
which allows us to relate the different
types of centers geographies together.
And this can be useful if you want to
make some transformation between the
different types of centers geography. So
you might have data at output area
level, you might want to convert it to
data at um ward or district and then you
can use a direct lookup table which
maintains relationships between output
areas and wards or LSOS or super output
areas
within the same census. There also a set
of lookup tables that do the same across
census. So you can make lookups between
um output areas in 2020 22
and 2011
um which but you have to be aware there
are some because the geographies tend to
change over time. They're not always one
to one exact. You just have to be aware
of that. But um and you also get a type
of geographic lookup table called a
postcode directory. So a lot of data
which um you may be manipulating from
like surveys which um is output using a
postcode
and what a postcode lookup table does is
it relates postcodes to other type of
geography. So you might have a postcode
of EH104EL
in a postcode directory. Each of the
record for E104 would have a bunch of
other columns which would then tell you
the the geography that postcode fell
within. So that might be the census
output area for 2021 or the the electrol
for 20201.
Um and that way you you could then use
that lookup table to relate your data at
postcode level to census itself and then
use the sets to add context to your
other um survey information.
and the the postcode topography lookup
tables as well as the other sensors jack
lookup tables are all are all available
through the UK data service
and so we looked at um uh data explorer
from the Manchester team we're now going
to use um one of the boundary data
download tools which I in response were
looking after at Adena to download the
census boundary data and We're
specifically going to use the boundary
data selector tool to download the
electro wards for leads um that we can
then join to our sensor stats in order
to do some mapping.
Um
so yeah be another demo where I'm going
to do some download of sentence boundary
data from the b data selector.
So,
I'm go back to the UK data service.
Let's just maximize my browser this
time.
I'll try and increase
so zoom things in a bit.
And again to access the boundary data
selector, we can use that quick link.
So data catalog,
click on the sensors
um loger black thing. Yep. Up pops the
popup. And we want census boundary data
borders.
And this is the boundary data selector.
So the boundary data selector allows you
to look through the different boundary
data sets we provide. And again, you can
either download the complete data set.
So we could download electro awards for
all of England or in our case you
actually want to drill down and only
obtain the electro awards just for
leads.
So it's quite a simple application
um really it's got two tabs. There's a
what and where you tell it what data set
you want and where you want it for and a
format tab that lets you change the type
of uh geospatial data format the data is
provided in. I'm really just going to
use the what and where format because it
will default to providing data in shape
pile format which for our purposes I'm
using the data in Q just is fine.
So there's three dropowns on the top
here which will control the data set is
listed.
Um so I want data for England
I want electoral data and 2021 and
later.
So you can see it's come back. We have a
choice between electoral wards and
parliamentary constituencies. So I want
the electoral wards
list areas and and like the data
explorer we can now drill down through
the sub geographies to restrict the data
purely to leads.
So
again similar to data explorer
it's York and the Humber
and again we want leads. Let's go
borders boundary data selector. We've
told it we just want data for leads. And
if we hit this extract boundary data
button, it's going to go to the
database, pull out the features, and
then dump it as a shape file.
So that's done that now.
Just minimize my browser.
So it's our data. I just open it. It
comes as a zip file and you can see it's
a shape file. So a shape file is made up
of what they call um is is these four
different um individual files is a PRG,
a DBF, a SAS, and a shape. And the
important thing when you deal with shape
files, you have to have all four of
these files in order for the data to be
to to work. So if you were to copy this
to somewhere else, you'd have to make
sure you copied all each of the four
components to the same place. Otherwise,
if you try to open the data in CQ, it
would complain because it wouldn't be
able to find the rest of the data it
needs to show the data set.
Um
there's various other things that we get
provided with D from data selector. A
simple read me just tells you those
files where it's for and sort of the
folklore support the service.
And this terms conditions file here
which tells us um specifically because
the data is open it's released under
open government license and in order to
use open government license data the one
of the conditions whenever you use that
data you have to include an attribution
statement that that tells you where the
data came from and we'll see this later
in CQIS when we're going to create a map
and we'll go back to this file and
insert that copyright statement into
something we create in CQIS.
So that's our data from CQIS.
Oops. Just find my slides again.
So let me grab the boundaries.
I'm just going to grab some other data
from the ordinance survey. And this is
like um contextual. It's going to like
give us a contextual background map that
I'll add to CUGIS so that when you when
we overlay that we bring in the center
boundaries. We'll have some context just
to add to make our map better. Um
and I'm going to grab this data from the
ordinance survey. Again, this is open
data. So the ordinance survey is
Britain's national math agency.
Um,
and they have a website where you can
download some of their open data sets.
And so if I just ordinate survey
and what they call is the data hub.
And so this is the OS ordinance survey
data hub.
Um,
and they have a bunch of they have open
data here. Um and it lists various open
data sets they provide. So they've got
stuff from the British Geological
Survey. Um so rock type and stuff and
stuff related to soil, but we want um
actual map data. So you can change the
providers from British OS survey to all
you just want the orange survey. And
this is the data set we want this one to
20 250,000 scholar kit. It's like a a
road atlas. It'll provide some nice
context for our boundaries.
So, I'm just going to go to the download
page and hit the download button.
And this is going to download data for
all of the UK.
And this is raster data. So, these are
just like image images.
So, let me just download that.
And I go to my downloads.
This is the ordinance survey data here.
It's called our zip file. So open that.
And if I go to the data itself,
it consists of this um mapping as a
bunch of like separate what they call
image tiles.
So I can try opening one of them here
just in like uh my Mac data preview. So
you can see it's basically this is a 100
km by 100 km block of map data in this
case for
uh land end. So,
so that's we've got data from the UK
data service data explorer, UK data
service boundary data selector and the
ordinance survey. I'm going to now open
CQIS and show you how to use QGIS and
we're going to add all these data sets
to CUGIS.
Um
we're going to have a break at 10 11 but
before that I'm just going to show you
I'm just going to add the data from our
various data sets into CQIS and then
after the break we'll get on with
actually doing the mapping the data in
CQIS.
So
so first I'm going to start CUGIS. So,
I'm on a Mac. I'm on a fairly recent
Mac. So, I've got QGS4. I think in the
instructions they told you not to use
QGS4, but um
it seems to work. Okay. Um
I think I will have to manually edit the
instructions for next time to say that.
Um
uh so this is CQIS. Let me just I
unfortunately I don't think I'm going to
be able to make it much zoomed in. So, a
lot of these OP menu things may seem may
seem quite small still. So, you just
have to um if you if you don't manage to
follow along now after the workshop,
you've got the workbook. So, you might
know to you might want to do it yourself
when you have more time if you sort of
lose track of what's going on here.
[snorts]
Um
so, this is what Q just looks like.
There's a bunch of um it's like most of
our like gooey applications. There's a
bunch of buttons along the top here
which you can do use to do different
stuff. Um there's what we call the
browser window the left here. Um these
are different um types of data you can
add to CQIS.
Um this main window here is the actual
is a CQIS map window. So when we add
data to CQIS, it'll be displayed in this
map window and the left here is what we
call the layers panel. So each of those
data sets will be will will basically
form a layer of data and the different
layers will be shown on the left here.
Um I said before on that geospatial
slide that the spatial data was
providing different spatial reference
systems and I mentioned the British
National Grid and WGS84 for the globe.
By default, curious will default to
displaying data in WGS84 for the globe
because it assumes you might want to add
data for anywhere in the world. Whereas
what we want to do is display data just
for the British National Grid in the UK.
Um,
so at the right here you can see there's
this thing called EPSG4326
and this is Q just telling us that it's
assuming the data is going to be
displayed in WGS84
and that the data you add is in WGS84
and that's not what we want. You want to
tell CQIS to display our data using the
British National Grid spatial reference
system. So the first thing I'm going to
do, and this is like in your PDF
workbook, it will tell you is to change
this from EPSG 4326 to EPSG 27700.
27700 is a code for the British National
Grid. So to do that, I just click
and you can see it brings up this
dialogue saying project coordinate
reference system. Um,
and I'm going to set that to British
National Grid 27 and 700.
Um,
you can also set it to any of these
other ones if you wanted to. Um, you
might be downloading data for America
like in Idaho, which has a SE has its
own unique um, spatial reference
systems.
Um
there's a whole load of different
spatial reference systems which have
different purposes, but for our case, we
want data in British National Grid. So
we're going to set it to EPSG 27700.
So there we go. So we've got a data into
7700. I'm now going to start adding data
to CQIS. So I'm going to first add that
raster data that we downloaded from the
ordinance survey that backdrop mapping.
So to add data to CQIS use this button
at the top left called open data source
manager. So
if I click that up pops this dialogue
um and this is called the data source
manager which tells QJS where to source
data from on our local machine or from
like outside our local machine like from
web services or exam etc.
And you can see down the left there are
all the different types of data we can
add to cus.
So I'm going to first add the raster
data by clicking the raster option.
Um, and here we specify the location on
your local machine where that raster
data is.
So I on for me it's on my downloads
folder. It's RAS 250GB
data.
And I'm now going to add one of these um
map blocks to
the one I want is SEIF because I know
this is for leads. So
select the SEIF and click open.
That's fine. CH just reads the file.
It's grabbed some metadata that's used
that will display the data. And I'll
just click add.
And we get our data set shown.
Um, by default
when CQIS draws raster data, it doesn't
always show it particularly clearly
because it's it's it's optimized for
like speed rather than like um the
actual quality of what's been shown,
which can slow things down. But so we
zoom in, it can be quite distorted and
blocky,
and I want to make that clearer. So I
can tell curious to um draw the map a
lot clearer and I don't really care if
it draws it a lot slower. So
again, as I said, what we have here
loaded is one of our data sets in the
map window and we have the layers panel
at the left that's listing our current
layers in CUGIS. And so you see it's add
a layer for SE which is our map tile.
Um within the layers uh panel you can
select any of the layer items i.e. the
different types of map and you can right
click and go to properties
and that will open the properties a
properties window for that type of layer
and we can use that properties to set
different properties of the of the the
layer being displayed.
Uh we'll use this later on when we want
to um create our census maps.
So the minute it it collected symbology
and in order to tell QS to display the
raster data more clearly I can use this
resampling option here
and so the minute it's it's say zoomed
in nearest neighbor nearest neighbor so
it's applying some raster operation to
tell it how to display that data
and I'm going to change this from
nearest neighbor to cubic
and over sample
It's quite low. So, I'm going to change
that to a lot higher.
And this basically just tells CUGIS to
draw the raster data a much higher
quality, acknowledging there'll be a hit
to the actual render speed
or the draw speed of the map. So, I'm
going to hit apply.
And you see now Q just redraws the map.
It's slower to draw.
but it's a lot more higher quality.
And so this is leads here. This is going
to be our area of interest for the rest
of the workshop.
I'm now going to add the other census
data to CQIS. So again
remember we downloaded the sentence
boundaries as a shape file from boundary
data selector
and we downloaded uh
uh the center stat as a CSV file from
data explorer. So I'm going to first add
the boundaries. So again same deal use
the open data source manager button.
This time want to use the vector
because it's vector data set shape file.
Um again navigate to the data
boundary data and the one we want to
select is the shape.
I add I the shape file. Open that.
Add
and close. And here are our boundaries
in cubis. Um
so because it's a vector data set um we
can again
we can use some of the buttons along
here to target the data. So this I is
called identify and it allows you to
identify the feature on the mouse click.
So I can identify click on one of the
polygons
and it will show you various properties
of that polygon.
So we can see it's a this particular
ward has this uh geo identifier and it's
called hairwood
at the minute it's not set as data
because we've not actually provided any
sense it's just simply the boundary
and we can do stuff like we can select
one of the polygons or we can select a
bunch of the polygons
um
and we can do sort of basic spatial
analysis. So
um
so
we've got a polygon selected. We might
want to find all the other polygons that
are connected to this selected polygon
by doing a select by location operation.
And so when we do that, Q just finds all
the other intersecting polygons with our
selected polygon.
[cough]
Um let me just add the actual sets of
stats themselves. So again data source
manager this time you want the limited
text.
So browse to the data set in this case
our sensor stats which is just the CSV
file.
Click open.
Um
so we get a preview of this of the set
as stat CSV. You can see our lovely um
proportional data expressing the values
as proportions and our our uh
categorical thing that telling us the
dominant combination type.
Um under geometry definition make sure
that's set to no geometry because the
CAC val doesn't actually contain any
geometry itself. that's in that's in
those polygons. We're going to join them
later. And again, just hit add
and then close the dialogue. CUGIS will
add that centers um CSV to the layers.
Um and so we can open the attribute
table in CUGIS. So you can see it's
loaded the CSV into CQIS,
but it's not showing a map view because
at the moment it's looking for the map.
Um
so that's our data in CQIS. Um we're
going to take a short break now for 10
minutes. Uh and then in that after the
break we now go into the actual fun
stuff of doing the actual mapping of
this data. I think that 11:25. So we're
going to start again.
So, so far we've grabbed data from the
UK data service.
We've done some manipulation of that
data in Excel and we've stuck it into
CQIS. We now want to go into the process
of actually creating some mapping.
So,
what we're going to do is we're going to
create a thing called chloropl map. And
I'm just going to give you a quick
introduction into what chlor mapping
actually is,
although you may well be aware of this
already. So,
At the left here we have like a bunch of
boundaries.
Um I think these are uh Scottish counter
areas. Um you can see each of the count
areas has a a uh geographic identifier.
That's the S12 code and it has some
random um stat attached to it.
Percentage of something I guess. So 5.27
8.2.
um
not particularly useful as a map and
quite hard to see the differences are.
And so that's that's where chlorophy
maps come from that. You basically shade
the polygonal data based on the actual
attached variable.
And so high values get shaded like dark
green in this case or lower values light
green. and it allows you to quickly look
across the data set and look at the
variation of the census variable.
Um so then you can compare what the
census day is like in in say Edinburgh
versus um the rest of Scotland or the
rest of the UK. Um I think on the left
here this is showing a percentage of
people who work in um the forest
industries. Um
so not surprising in Scotland that tends
to be the Scottish borders or Wales the
sort of um rural parts of Wales and less
um people employed in forestry within
the city of London for example or the UK
Scottish central belt whereas I think on
the right we have a chloropath map which
is showing the percentage of people
possibly house type
um and so the dark areas are where it
might be there's more boat living in
flats. The lights less living in flats.
Um
so
there are some limitations of the
qualify map we have to be aware of. Um
they tend to imply because you're just
shading the entire polygon the same
color that the underlying population is
distributed uniformly across the polygon
which reality that's not the case. So,
as you can see here,
um this is uh some aerial photography
and we've got these black boundaries
shown on top. Um and the population is
really only in the sort of like the the
upper areas of the uh polyon
around here, but there are areas of like
green space for example and areas of
industry where there's no one living. So
um the chlorophy map as is can be a bit
misleading in terms of like because the
data is not uniformly underneath it. So
as we'll see at the end of the
presentation there al alternatives the
chlorof map which have become more
popular among some of the sort of um
ways of visualizing the census data. So
just to be aware of the this limitation.
Um the cloth map is still great though
as a way of like getting a quick insight
into the how the census variable
uh varies across the entire data set and
to make quick cap comparisons between a
small particular region of interest and
the UK as a whole or whatever. Uh and
it's a useful skill to be able to create
cluster maps in cuis.
And so that's what we're going to do um
over the next 20 minutes or so. So what
on the right here is a PDF that I've
created in QGIS using all the census
data we downloaded this morning. So this
is showing the percentage of households
in leads living in flat as recorded by
the 2021 census.
Um and if you're following the workbook,
you'll be able to create this PDF
yourself and you will also be able to do
this. uh and the idea is that then
you'll be able to take you'll be able to
download any other data set from the UK
data service data explorer and then
build your own uh PDF maps for any other
data set that's available in the UK data
service by doing the same sort of things
and those same transformations in Excel
and those manipulations
and so the components of a glorified map
are the map centers variable um so those
stats from data explorer uh the centers
boundaries that we download from
boundary data selector and making a
linkage between the stats and the
boundaries
doing a cloverhead map classification.
So we have to tell CQIS how it should
shade each of those polygons according
to the sentence variable
and then a color ramp. So what sort of
should we use blue to display the
different um values or red or green or
whatever.
And then as I said before we always have
to include a data attribution statement
because although the data that you
download from the UK data store is open
access there are conditions of use of
using that open data. one of which is
you attribute the provider of that data.
So this is why on my um
uh PDF here and as well as all my maps
that I'm showing in these slides there
is an attribution statement at the
bottom that tells you the data came from
national statistics under an open
government license and it contains
ordinance survey data of the because the
ordinance survey are the ultimate source
of those census boundaries
and there's a crown copyright database
right copyright statement.
So,
and so when we talk about linking stats
to census boundaries, again, just to
reiterate, we have our census stats
download from data explorer.
Um, they'll contain like this geographic
identifier column, which is those
nine-digit codes that tell you within
each row what is the unique ward.
and then a bunch of census boundaries
that have the same codes. And because
they're the same in each case, we can
link the data to one another.
Um, within the census, um, they should
be fairly consistent between data sets.
Um, if you have other nonsensus data, so
you might have random neighborhood
statistics data that was produced
outside the census that you might say be
trying to relate to electoral wards. You
can often find there can be slight
differences as the geographies have
changed and so you might always you
might not always find there's a onetoone
link between your data and the
particular geographies.
Uh so you just have to watch and make
sure there is a onetoone link
and so I'm a hands-on demo of just um in
QGS making that join from our census sat
to our census boundaries.
So
minimize some of these windows
back to CQIS. Um
again I think CU just it might be quite
hard to see what's going on here because
I've I've got quite a high resolution
screen. So,
so again,
as I said before, we've got the layer
panels here,
and I can select the polygon layer. So,
these are electro war boundaries.
Right click and go to properties, and it
will open that layer properties.
And one of the properties is joins. So
it tells CQIS what data is currently
joined to the polygonal layer and what
data do we want to join to the polygonal
layer if there's none currently added.
And we want to join those census stats
to the census polygons.
And so we click the add button
and it will add a vector join. So I can
just make this window a bit wider so you
can see what's going on.
So, it's pre-selected that this
accommodation type leads 2021 because
that's our CSV file.
Um,
and it's
it's been clever. It's just picked the
first column in the CSV as the GR
identifier,
which is what we renamed to Ward's GUID.
The target field is the field within the
actual polygonal data set itself. I that
shape file which also contains those
geographic identifiers. So if I select
the dropdown, the one we want is W 2022
code
and if I do an okay
and apply
and okay. And this time again in the
layers panel right click and go to but
this time
go to open attribute table
and this will show our polygonal data
set.
But this time it's we now have all those
um sets of stats added to the polygons.
So that's again our lovely um raw counts
of the households, our lovely
proportional values and our dominant
accommodation type.
And again I can use the identify button
to select click on one of the polygons
and it will show all the properties of
those center stats on each of the
boundaries.
The identify buttons a nice way you can
just quickly step through the data and
relate the plug polygon to the actual
sensor stats.
Um,
so we've got our sensor stats attached
to our boundaries.
Before we actually do the chlorop happy
in CQIS, we just go back to slides for a
while.
Oops. Yep. So we've done our table
joint. We start that's attached to the
boundaries
and we're just going to consideration
some of the chloride mapping choices.
So we have a
[clears throat] so we have the choice of
the centers output geography. Um in our
case we've already made that choice by
picking the electro wards
um the choice of color map
classification method and the choice of
the cloth color map ramp.
So as said before the sense output
geography is available at different
levels of output geography
[clears throat]
but disclosure control means that not
all the sensor variables are available
at all levels.
Um [snorts] and the point so this is
what you can see here. Uh well these are
this is the same sentence variable
displayed at different um output
geographies. So we have sentus output
areas here um lower layer super output
areas and middle layer super alput areas
and and the thing is if you analyze the
same assuming the data is available all
levels and and it's not been imp
impacted by disclosure control not
making available
uh you can get different insights into
the variable by looking at at different
levels of geography um
at the output area level it might be
quite um um noisy whereas But as you
zoom out it become the data tends to get
smoothed and you see you will see more
regional patterns and and so you will
therefore produce different types of
chlorophy map at the different levels of
geography. Um and that's just a case of
finding
um the level of mapping what's most
suitable for you and also you may be
looking at bringing in other data sets.
So maybe environmental information which
may only be be available at particular
output geographies because so you might
have like um pollution scores or
something which may only be available at
super output area level and you're
bringing that into add context to the
census data.
And in terms of um what a chlor map uh
actually is it's doing some sort sort of
data classification. It's it's
simplifying the data
um
so your raw data itself. So
um
where's it going? Uh let me just find
that spreadsheet. Uh
oh, we've got I can open this. Yeah.
Yeah. So these are all just um one of
the columns is a whole variet what the
33 layers 33 records
is a whole range of values and there's
probably a unique value per per record.
And so when we create a cloth map we're
going to simplify the data by arranging
into five um groups or seven what they
call data classes just to simplify the
data. Um,
so if I go back to the slides, if I can
find the slides. Um,
yep.
And so we do a sort of data
classification. [clears throat] We take
the full range of data values and we
classify into a set of what they call
data buckets. And each data bucket has a
we set a minimum and a maximum value for
the values within that bucket.
And you can see that what's going on
here. So down the left we have the range
of values say from 3 to 93.
And then what we do we have a in this
case I've set five buckets. And for each
of those buckets we set a minimum value
and a maximum value that tells you when
that data set value is within the
bucket. So in our first bucket have a it
will take range of values from 3 to 20.
The next bucket from 21 to 38, 57 to 74
and then a 75 plus. And then so all the
records
which [clears throat] are values between
3 and 20 will get put into this bucket.
All the records values between 61 and 64
will be placed into this bucket. And
this is a simple what they call an equal
equal interval classification because
each of the buckets is the same size.
And so they all go from 3 to 20. So they
all go they all cover 17 values in my
case from 3 to 20 or from 21 to 38 or
from 39 to 56 or 59 to 74. There are
different classification methods and and
it's the way they create those buckets
which is a different each classification
method. In some the buckets may not be
the same size because you they may um
where you have more data you may have
more buckets or buckets are smaller or
have different minimum maximum values.
Um
and so this is what each of these is
what they call a different um chlorop
classification method. So that equal
info one I showed there and there's a
whole bunch of different ones
and they basically are just controlling
how those buckets have been formed and
how the data is being applied to those
buckets and sort of the how how they're
basically simplifying the the range of
values in the data in order to make the
data more usable or um
and so chlorophy map which we've applied
some data classification and then we've
got five buckets and and again with a
minimum and maximum value. So
and apply to some color ramp in this
case an orange one. So low values the
lowest bucket values have like a a very
white whereas the largest value bucket
is quite dark
and again
this this thing you can take the same
data set and you can choose a different
classification method and you will
produce a different sort of map. So it's
just to be aware that if you pick a
different classification method with the
same data, you can produce a different
kind of map and you just um
and the and and you can also
when [clears throat] the number of
buckets you also pick also produce
different types of map. So if you have a
a a classification with five buckets,
you will get a different looking map
than you have one with three or seven.
So
that's what you can see here. So we have
the same data set. We're using the same
sort of classification method. I think
it's probably equal interval,
but we're varying the number of buckets
used between three, five, and seven. And
when you do that, you get a different
sort of map um just because the way the
data is being like allocated to those
bins. And then the output map gets
shown. And on the left, same deal. This
time we've got the same number of
buckets of five, but we're using a
different classification method.
Um, and again, you get a different sort
of map depending on the classification
method.
And the point is that no classification
method is right or wrong, but you just
have to be guided by what your actual
data is like in order to pick the
classification method to use and the
sort of map you want to construct. And
the idea is not to create to create a
misleading map.
And we'll see that in KU just how you
can actually look at the underlying data
to see how it's been classified.
And so that's the classification
process. So that's telling us how to
bucket the data. You obviously then want
to pick the actual color ramp to use. Um
so sequential is what you normally use
for color maps. So that's from light to
dark. So you have a green a sequential
color map from light green to dark
green. Diverging would be case if you're
like um it might be data around a mean
around zero. So you might have minus
values to the left and plus values to
the right. So increasing plus to green
to light to dark green and diverging
sort of light brown to dark brown. And a
qualitative is a little more of a
categorical map. So this could be a land
cover map.
where you just like um just light blue
for the sea, green for forests, orange
for mountains, red for buildings.
And so go back to curious and we'll
actually get on with actually creating
some colored maps.
So,
so I'm going to first create a
categorical map
and that's going to use this um semi-
detached column here. That was the one
that tells the dominant type
accommodation.
So again left click the layer into
properties
and this time you want symboli option
and it's like the symboli controls how
the map is being displayed by cuis
and the minute it's showing a single
symbol. So it's defaulted what it
defaults to will probably will change
um each time it just picks a random
color and in this case a single symbol.
So it's I'm drawing all the polygons the
same the same single symbol in this case
orange
and I want to change this which I use
here to the different types of
visualization
and I'm going to first use what we call
a categorized
and I'm going to select
that categorical data we created our
dominant accommodation type
and hit the classify button
QS will look at the underlying data and
we'll create a new um color class for
each of those the the unique strings
essentially within that column. So our
semi debt or flat terrace whatever
um
which you can see is what it's done here
within QIS I can just um it for it
always adds in all of our values because
in some people have data sets which
contain data with null values our case
we don't need that because all our
records are populated so I'm going to
get rid of that so you can just select
the all values and do delete
Um
and so the values here are those that's
the actual strings shown in that um uh
dominant accommodation type column.
Uh a legend is on a map. It tells you
what the map is actually showing. It's
like a label of each of those classes.
At the minute it's just used the value,
but I want to make that more reasonable.
So you can just in the QJS you just
double click under legend and then you
can edit what it actually shows as a
label for the class. So, I'm going to
put detached houses
and flats,
semi- detached houses
and terrorist.
How was it?
And you can also reorder things here.
So, if you just select one of them and
then just move them around. So,
uh let's put detached houses, semi-
detached houses, terraces, and flat at
the bottom. And we also change the
colors for each of the classes. So,
so detached houses, I don't know,
we go to dark green. Um, semi- detached,
light green
terrace is
a lovely shade of orange
and flat. Let's go for like
sort of pinky color
and just apply that.
And so now Cugis has redrawn the data
using that categorical uh
classification.
And you can see that in leads in our
electro wards uh as you might expect the
do within the inner city the dominant
accommodation type is is either flat or
terrace houses
reflecting the urban density in the city
center. I can make this more readable if
you go to the the layer properties
and then you can change the opacity from
100% to something less than 100 to 80%.
And what this mean this will make the
the the boundaries sort of
semi-transparent. So you'll be able to
see the ordinance background mapping
shining through the data.
And that way it's a lot easier to see
the lead city center and the suburbs.
And you can see in the suburbs the
domination type house type is detached
homes and semi- detached house just
reflecting I guess the uh lower density
areas of the city.
So that's our categorical map.
Um, if you want to create an actual
chloroplast map now from the actual one
of the census variables
and we're going to show the percentage
of data of people who live in flats.
And so the same deal, leftclick the item
in the layer panel. So left so left
click to select it and then right click
and then properties.
And again back to the symbology. And
this time change from categorized
to graduated.
And again pick the column we want to use
to draw the map.
And I want flat
but I want flat proportion.
So this is our percentage value rather
than the actual raw flat values.
And I can just go away with the defaults
and hit the classify.
And CQIS has gone away and has applied
one of those those those classification
algorithms to bin the data into in this
case five bins. And it set a minimum max
for each of the uh classes that that
determine um which records get values
get attached to each bin.
And what you can look at is the
underlying data. So the histogram
so the classes tab click on histogram
and click on loads value and it will
show you the distribution of your
underlying data and this can help guide
you in a sort of um classification
method to use. So is your data all is it
mostly in lower values or is it upper
values and then
that can help you define
I'm going to pick a natural breaks one
classification
and I increase the number of classes to
seven and you can see that QS will re um
calculate the class breaks
and I'm also going to change the color
ramp from red to blue.
like so.
And let me just apply and okay. And you
can see that QGIS has now redrawn what
was a categorical map into a univariat
chlorophy map. It's univariate. It's a
single variable being mapped.
And you can see that perhaps
unsurprising in leads
um the wards with the highest proportion
of folk living in flats tend to be in
the city center or right in the safe
city center
and that lessens as you go towards the
suburbs of the city.
What we can now do is create
[clears throat] that PDF in CQIS.
[cough]
And within CQIS,
um the way you create a PDF output,
which you might say want to incorporate
into a a a document you're producing or
send to someone an email is a thing
called the layout manager. And to access
the layout manager,
you go to project
and layout manager.
And you just accept new from template
and click create.
And you have to give the layout a title.
So I'm going to give it the title of
my census map.
And what pops up is a bit like a if
you've used PowerPoint,
you get a black canvas in which you can
then sort of drag elements on in order
to create your print layout. It defaults
to landscape. I want to be a PDF. I
sorry, I want portrait.
So
I go to layout properties,
page properties, rather sorry.
Page properties. Yeah, layout page
properties. So layout page properties
and change orientation landscape to
portrait.
I'm just going to maximize this the
layout window. Um
a bit oops.
Okay. And so down the left here you have
different elements you can add to the
layout.
So I'm going to first add a map. So
which is this one here which is like a
add map and you select the button and
then drag onto the canvas
and cuis will add to the layout what's
currently shown in the map window of the
main cuis application which in our case
is a univa chlorophy map overlaying
against a backdrop or in a survey map
and then we can add again you can resize
this and move stuff around
Um, one of the first elements we want to
add is an actual map legend so that we
can because if we just gave this to
people, they wouldn't understand what
the different shades of blue actually
mean. So you have to give them a a
legend and a map legend just tells you
what the actual map is showing. Um, so
we use the add legend widget. So again,
we click on it and then drag a legend.
And because QJS is currently showing a
raster data set, it's added all the each
of the colors shown on each of the
raster background as a separate thing,
which is really not what we want. So we
can edit edit this legend to remove the
the legend for the raster just to so it
just shows the class breaks for the
chlorophy map.
And so to to do that under the legend
section right here, we go to legend
items. It's currently set to
synchronized to visible layers, which
means whatever you see in the map layout
is is slayed to what's showing in the
main QS window.
And so what we want to do is change that
to manual.
And that means we can now select the
raster one and click delete.
It will still keep the raster map shown
in the window, but it just removes all
that um the actual legend for the
raster, which makes our legend a lot
more useful.
And now we can actually go into legend
itself and and start editing this
because at the minute it says 3.4 to
618, but what well it's actually flat,
but we need to tell the users of the map
that. So
you can just double click one of the
legend items
and then you can just edit this. So
percentage
fats percentage
C
and I can do that for all the rest of
them.
Oops.
Yeah, you get the idea. So, uh
flats
flat That's good.
Oops.
And you can also double click here and
change the actual title from at the
minute it's just showing the name of the
actual underlying boundary data set. We
can change this to um
2021
electoral
lords
and the legends updated and we can move
the legend around.
And so now a user of our PDF can
actually tell what's on the map.
Um we can also add stuff like a scale
bar. um
not terribly useful because we're not
going to use this map for navigation,
but it still might be useful because by
providing a scale bar, we can the user
can tell how big each of these actual
ward polygons is in the real world. So,
in terms of kilometers,
we have a north arrow just to help the
user orientate themselves. Um,
and then we can add some text. So, the
first text we're going to do is add a
title to the map. So, maybe that was a
bit quick. So, again, from the left, you
just click the add label and then drag a
new label onto the canvas.
Oops.
And then you can write the canvas the
text. So, I'm going to go
a map showing
census data for leads.
And you can change the font to make the
text a lot bigger.
I might have to drag this out again.
And maybe I can make that uh bold.
And as I said before, an important thing
we have to add is a copyright statement.
And again, we just add some text label.
And if I say go back to the
for our downloaded census boundaries
into that terms conditions document
and just grab the attribution statement
which is this one here.
Just copy that
into the label.
At the minute, it's got the year as a
placeholder, and we want to update that
to this year's current year, which is
2026.
And again, just increase the font.
And so this way, we now have an
attribution copyright statement for the
data we're showing.
And so there's various sort of things
you add to the layout. You could add
some decorations. I don't know, like a
big star or something.
That's not particularly useful. Um
um but again, in the CQIS documentation,
there are various guides that tell you
how to create nice looking layouts.
And so what you can now do to export as
a PDF is use the button along the top
called which has got like this the
Acrobat symbol on it and just export as
PDF
and tell you where you want to put the
PDF to. I'm going to just put that onto
my desktop. My sentence map
maybe
of leads
and do a save
and keep defaults are fine.
If I go to my desktop, let's just
minimize Q just for a while.
And here's my PDF.
So, here's my PDF of um leads that I
create in QJS from census data
downloaded from the UK data service. And
I could just email that email that to
someone or stick it in combine it in a
a document I'm creating.
And again, you can do the same for any
sort of census data you download from
the UK data service if you follow this
process.
So I'm going back to some slides I think
and we're just going to look at some
other types of sets map other than
univariate and chlorop and categorical
maps. So let me just make these full
screen again.
Um so some other types of sensors map
you might see or indeed you can create
in cougis are bi chlorophy maps and
ctograms
and there's also some sensors maps
called asymmetric chlorophy maps which
we'll look at some examples of later but
you can't actually use you can't
actually create easily in cis.
So the type of chlor map we were showing
and which we indeed created in CUGIS is
a univariate chloroplith map because
we're only showing a single variable at
one time in our case like accommodation
type so the potential flats. A baricop
map on the other hand will show
simultaneously two variables at the same
time
which [clears throat] can be really
useful for census data because it means
you can create some some more
interesting maps. [snorts]
Um and that's what you can see it right
here. So
at the top we have these two univariate
centers. That's unariatified maps. Well,
one for poor health and one for
unemployment. So in poor health um so
low percentages of health poor health
are in gray and as poor health increases
it's um up to dark pink
unemployment so low unemployment gray
high unemployment um dark cyan
um those two individual maps but if you
combine them
and create a third map so this time a
biol map we're showing the two variables
at the same time.
And so what you end up with you is like
this sort of thing where we get
combinations of purple and cyan.
And so where on both unemployment and
poor health is high, you'll get a very
dark shade of blue. Where they're both
low, you'll get gray. And then the
intervening things, you'll get different
combinations of cyan and purple. And
this is like a in so in one in one
single map we're showing two variables
at the same time. Um which can be useful
for census data especially multivariate
census data because again we can show
multiple variables at the same time. The
key thing of vari maps though is to make
sure you pick variables that are related
because otherwise you could create some
quite misleading looking things.
Um a carttogram on the other hand um in
a normal chloroplith map we simply keep
the say the electro ward boundaries how
the shape of them as it is and simply
color those polygons according to the
sentence variable. Our cargram is a
special for a map where instead
um we actually distort and reshape the
census boundaries according to the
variable itself.
[snorts]
And this can be really useful like if
you're trying to like um if you got the
case of of Scotland for example.
So these are showing Scottish council
areas and as you can see in the central
belt the the council area for Edinburgh
is very quite small whereas the council
area for like um the Highland Island is
massive. Um and so you can get if you
just show it as a chlorophy map the
actual underlying geographies are quite
hidden
and so instead we can do create a a
ctogram which will distort or reshape
the actual polygons according to that
census variable. So that where the
sensor variables are high the the the
polygons will be distorted more and
where it's lower they will get distorted
less. And so you end up with something
at the bottom here where the underlying
um polygons have been distorted. And so
these large the the central belt
polygons have essentially ballooned up
and have become a large bigger whereas
the the the highlands and islands which
are like lower populations have been
distorted less
and it's a different way of like viewing
the data.
Um
and the important thing in in this type
of carttogram
is that the the actual topological
relationships between the polygons has
been maintained. So that
um so in this case selfirer which is
down here is still neighboring easter
and include and that's the same in the
output ctogram.
so that you can still relate neighbors
to one another.
And so ctograms are quite feature in the
wild for like um visualizing social
economic data. Um
the guardian features carttograms in its
um reporting for the EU referendum what
10 so years ago which we're still living
from. Um
and also they featured prominently in
this book called people in places which
is a really nice book which is
visualizes the 2011 census data and it
uses all that the data is all visualized
as um these area ctograms
uh it's a really nice book um
um
and so what I'm going to demonstrate now
in QIS is how you can create a car
yourself um of the sort we saw there.
Um and also
how you create a bicaro map.
I have to caveat this um as we'll see my
version of CQIS
will not currently allow me to create
the actual ctogram.
Um because I'm using a Mac and it's a
new version of CUGIS and a new Mac and
there's some issue with my particular
install. But I will definitely show you
the tool and I'll be able to show you
the cart that I pre-made. The last time
I ran this workshop three months ago in
March where it ran. So um and hopefully
in the next work time we me run this
workshop in three months time I will
have come with a workar around and I'll
be do a actual live demo of using this
particular car construction but the
minute it's going to be a bit blue peter
here's one I did earlier. So
um let me just go back to CQIS
and we'll certainly look at how to
create the biaric chlorophyll map in
CQIS. But again the the ctogram will be
a pre-made one. So apologies for that.
So let me just remove the layout.
Go back to cugis.
Um to create the by variant chlor uh
chlorop again we need a data set that
has two types of variable in
accommodation type
and tenure type which are not in our
data we created. So for the purposes of
this exercise I created another data set
which is in that data pack that you
could have downloaded from the GitHub
repo that we sent you in the joining
instructions.
So I'm going to navigate to that data
set and add the cu just now
and it's actually a data in a geo
package which is a different type of
data other than a shape file.
So
so here's my data pack.
The directory search here should be the
same as what you download from the
GitHub repo.
And so what I have to first do is add to
CQIS the geo package.
Um it's a geo package. A geo package is
a special type of geospatial data set.
In some ways it's an advancement or a
shape files in that you can have a
single geo package that can contain
different layers of data. So you add the
geo package and then you connect to it
and then you can add the data set.
And again it's it's for leads. Um
and if we open the attribute table you
can see we've got a different types of
column again which are proportions. So
the AT means accommodation type.
Um the 10 is the tenure type. So
basically we have those electro awards
for leads and a bunch of columns giving
sensor variables on accommodation type
and a bunch of columns giving colum uh
variables at tenure. And we're going to
basically in CQIS create a by variable
map that that shows um one of the
accommodation type variables and one of
the tenure types variables together.
And so in your um workbook it talks
about how CQIS is really nice. It's an
open source piece of software and we can
expand its functionality using a thing
called plugins. And these are like
community developed enhancements to cues
that people have added to increase its
functionality. And it just so happens
that one of these plugins is a one that
allow us to create by variate chlorophy
maps because by default QIS will not
allow us to create biariate chlorophy
maps. Uh so we use a plug-in to do that.
And there's also an additional CQIS
plugin the one that I can't currently
get to work which will allow us to
create carttograms.
And so to enable these and this is
explained in the workbook. You go to the
plugins menu
and you can tell what plugins are
currently installed in QGIS. So you can
see I've got the by variant one
installed and the curs one installed.
And you can also search for other
plugins. And you can see there's a whole
load of them. there like the thing is
curious if you discover it can't do
something you can normally find a plugin
that will do it for you. Um
and so in the case of the by variant
plugin
again I'd go to the properties
and go to that symbology type it adds a
new option to the symbology type of the
bariate renderer. So rather than
categorized or graduated, I can go to
biariate renderer
and similar sort of thing that was on
the slides.
So we have to tell it what of the two
columns to use to construct the biariate
chlorop.
And I'm going to go for
accommodation type of flat
and maybe private rented flat tenure.
Now we click apply.
QS will classify the data
according to that biic classification
and you can see it's again it's it's
created a new type of legend for us
which is like these um rather strange
looking like um grid representation.
So again, where both accommodation flat
is high and private rent is high, you'll
get blue and you'll get gray when
they're both low. And you got the
intermediates for the different types of
uh variable. And again, you can also
change the opacity to make the OS data
shine through
and give more context
like so.
And then when I to create a carttogram
using that cartgram plugin, QJS will add
a plugin a cartgram button to the main
toolbar called compute cartgram.
So you would click this
and up would pop this UN UI and you
basically tell the cartgram plugin which
column to create the cartgram from.
And so you might go for like um let's
pick actual correct death uh so flat and
you hit okay. If I hit okay here I'm
fairly sure CQ just will probably dry to
a halt and then will die which is not
good for a demonstration. So I'm not
going to do that. I am going to add the
d the cam I created on my other Linux
machine which I know worked. So let me
just add another duo package with this
um
uh
with this card I created before. So
and so that's that's what the the the
cartgram plugin will create is a
ctogram. So as you see it's distorted
the electro wards according to the
accommodation type variable and again
you can shade
and that just gives you an alternative
way of creating a cartgram in cuis.
So that's the demonstrations of
everything in the workbooks of how you
manipulate data and create a univariatic
chlorophyth map, a categorical map, a
ctogram and a biiclorith map. Uh I've
got some final slides and then we'll
maybe go through some questions and then
we'll bring the workshop to a close.
So as I say,
I said one of the problems with the uh
chlorophyth map is the fact it implies
the the population is uniformly
distributed across the extent of the
polygon.
Um, an alternative to the chlorophyth
map is what you'll see are what these
things called asymmetric and mass
chlorophy maps which are an alternative.
Um, and what they do is they basically
use like a an additional layer rather
than the sort of the boundaries such as
a buildings layer and then you basically
shade that buildings layer by the census
variable. So you get a much more
realistic
um distribution of the population shown.
Um
and these are used a lot in they've been
used a lot for um by the ONS
um other um organizations to display
2011 and 2021 census data. So there are
two examples here. So this work is
really pioneers by the focus
college London and this data shine um
service and that's what data shine is
showing it's basically rendering instead
of the the the sensor data as a a
uniform color map it's rendering using a
buildings layer and that and that's what
you see on the left here. So it gives
you a more realistic um looking map. So
for leads you're
so you get where the areas of like green
space and stuff are not being shown or
in industry when there's no buildings
and I say the ONS
have taken inspiration from this and
I've and they use it to create um census
maps in their census 21 mapping software
which these are both online you can look
at um and they use a similar sort of
visualization there. Um you can create
these things in couges but it be it can
become quite involved um because you
have to obtain a sort of UKwide
buildings layer which can be quite large
and then do all the sort of like um
geometrical operations and and
processing to create the maps yourself.
Um
um so it's just to help you if you have
questions about census data or how you
in the future
um within the UK data service there's a
UK data service training program this
workshop is one of them um we will be
repeating this workshop in three months
time so if you maybe
are just getting started you you try the
work back book afterwards and then you
discover issues and you want to get see
the presentations again then you can
certainly welcome to join that again
that'll be mid-occtober
um within the UK data service we're
developing a set of learning modules
which will cover um different ways of
using census data um some of the
background of how the census data was
created um looking at some of the issues
how you manipulate the census data
there is a UK data service help desk
where you can ask questions you have and
as I say because the UK data service set
support is distribution among different
organizations we have different areas of
expertise. So my particular expertise is
in is the actual is like um the census
geography and how you compare census
data over time and cope with changing
geography. Um how you do a lot of
geospatial data manipulation.
Other guys in Manchester are more um
knowledgeable about the actual set of
stats themselves
um and how how you pick the right sense
of stat for your particular research
question. But if you send a query to the
help desk, they will do some triaging
and will make sure it gets rooted to the
most um the best person within the
centers team to answer your specific
Sir,