Video summary
This workshop, led by Moren Hicker of the UK Data Service, establishes that effective documentation in social sciences research is the foundation of metadata, essential for maximizing reuse value, transparency, and historical provenance. The session distinguishes between two primary levels of documentation: the data-level, which pertains to specific files like transcripts or SPSS datasets, and the study or project-level, which encompasses the broader research context, methods, and findings. To ensure data is Findable, Accessible, Interoperable, and Reusable (FAIR), this documentation must be collected throughout the entire research lifecycle rather than added retroactively, utilizing machine-readable formats and standardized schemas such as the Data Documentation Initiative (DDI).
Key components of high-quality documentation include concise variable labels under 80 characters that specify units of measurement and link to questionnaire questions, alongside clear value labels for categorical variables that explicitly differentiate between missing data types like "not recorded" versus "not applicable." The presentation strongly advocates for the use of controlled vocabularies and standardized thesauri, such as ELSET, over free text to ensure consistency across systems. Furthermore, the workshop highlights a critical gap in qualitative research where analytic materials like memos, field notes, and analysis syntax are often omitted; it argues that including these elements is vital for demonstrating analytic transparency and explaining the decision-making processes behind the study.
Beyond standard data dictionaries and lists, valuable documentation extends to researcher observations regarding power dynamics during data gathering and even draft analysis work, such as unpublished defenses of biased samples or conceptual pieces on "felt poverty," which provide crucial context despite not being formally published. Technology now facilitates the download of codebooks and memos from software like NVivo, the digitization of historical physical binders, and the use of blogs, websites, and multimedia files as supplementary resources, with tools like Qualibank enabling the browsing of qualitative data alongside linked documentation. Readme files are emphasized for describing data preparation steps, while existing materials such as consent forms, interview guides, and information sheets from other projects should be reused to inspire good practices, particularly when working with vulnerable groups.
The session concludes by outlining the practical tools available for creating structured metadata and codebooks that can exist independently of the raw data files, including the DDI Editor and NESTAR Publisher. Researchers are encouraged to utilize mandatory deposit requirements like data offer forms and templates provided through the UK Data Service, as well as external resources such as the Learning Hub and SESD expert guides. By integrating these diverse forms of documentation—from technical reports and protocols to field notes and digital archives—researchers can create a comprehensive record that supports future reuse and maintains the integrity of social science research over time.
Read the full video transcript
Okay, I think we're right at 10 o'clock
now, so I'll go ahead and get get
started. Hello everyone. Uh, welcome to
this online workshop, best practices for
documenting social sciences research
data. My name is Moren Hicker and I've
worked with the UK data service for
about 14 years now. Um, doing anything
from, you know, what we call pre-ingest
before we get collections in all the way
through to reuse projects.
Um, okay. So, rest assured everybody is
muted. Your cameras are off so we can't
see you. Um, and we're going to go ahead
and get started.
Okay. So, this is a bit of an overview
of what we're going to be doing today.
We'll start with the basics of what is
documentation, why is it important,
including some just a little bit of
discussion on fair principles. And from
there we'll move to metadata or data
about data which is really documentation
at its core. We'll then move into some
examples of data level documentation and
then study level uh or project level
documentation.
Um and then we've got a bit of a
distinction between common documentation
for qualitative projects and
documentation for quantitative projects.
And I've got a lot of screenshots of
examples uh for you to have a look at.
Throughout, I've included just a couple
of exercises to give you an opportunity
to look at some examples and think a
little bit more about what kind of
documentation you can and should include
with your project.
So, before we get too far along, we do
have our first activity. So, um I
believe Emma uh is going to pop a link
into the chat of a worksheet. And what I
would like you to do is read through
um read through this worksheet. It's an
interview transcript. There is a little
bit of a challenge with this interview
transcript which you'll see as soon as
you open it. But what I want you to try
and do is to the best of your ability
read through what you can and try to
guess what is the participant's age and
what year do you think the interview was
taken. Okay. [clears throat] So just
those two questions. I'll give you some
time uh with the interview itself and
then uh we'll we'll uh pop our answers
into the chat when you're done. So, um
if you bear with me, I believe the
worksheet is already up, but I'm going
to try and uh swap my screens around to
put that worksheet up on the screen as
well in case anybody needs it.
So, I'll give you just a few minutes now
and we'll we'll I'll come back again.
I expect many of you have have made some
observations already um about this
interview. But remember, we're looking
at what is the participant's age and
what year was the interview taken. So, I
am going to give you two more minutes.
It is a it is a challenge to read
because it's a phonetic uh
transcription. Um, but yeah, what is the
participant's age and what is the
interview date?
We've got a couple guesses in. Do feel
free to share your guesses in the uh
chat if you would like.
All right, we've got some some good
guesses already.
All right, we've got some people
pointing out it must be 1980s.
um
just because some of the context of the
interview itself. We've got some some
legal changes there. We've got age
range. Let's see. 40 plus, 60 plus.
Seems like people are kind of guessing
40 or 60 or someone said about 50, 65.
Um any other guesses [clears throat]
with with the year? Oh, non-native
speaker. Yes, this is a this is a
challenging one to read um because of
the phonetic
um transcription here.
Yes, northeast Scotland. There is there
is if you're if you're familiar with
different Scottish accents, you might be
able to pinpoint where it is. Um that
was done purposefully to have the uh
phonetic transcription. We could have a
whole another session about
transcription and the art of
transcription if people would like. Um,
fantastic. Okay, so we have some really
good um, guesses already. I'm going to
go ahead and just pause pause my share
again and I'm going to put the answers
up and uh, Emma, do feel free to share
the the next worksheet.
Okay, so we have here that the date of
interview is 1978 and this grandmother
is 43 years old.
Um, so I I love doing this kind of
exercise where we look at a bit of data
without documentation and then you add
the documentation in. And I know the
phonetic transcription is particularly
hard to read um, if you're not familiar
with with uh, some of the uh, Scottish
slang and pronunciation.
I believe it is uh, Glasgow which
somebody has highlighted. I think it was
Glasgow where this was taken interview
uh was done. Uh this it's a really
interesting um study which was looking
at uh generational knowledge about
health. So they did interviews with
grandmothers, mothers and daughters. Um
and they kind of did a compare and
contrast on the different ideas of what
is health. Um so it's it's a really
interesting study in and of itself. But
I think what's particularly interesting
about the added context of some of this
is normally when we are presented with
something like this is a grandmother, we
might think about our own grandmothers
as adults without really thinking about
how that person has changed over time or
what their different circumstances might
be. Um, so whenever I looked at this
one, I would think of my elderly
grandmother rather than a 43 year old
woman, which probably says more about my
position in the research than it does
about the actual study. Um, so having
that added context, having that added
documentation is actually really really
important for understanding what the um
opportunities [clears throat] and the
limitations are of that data. And it can
really help you think about your own
positionality when you look at what
assumptions you might make about um the
data. Um because we all make
assumptions. It was, you know, it's a
totally normal thing to to try and place
that data somewhere. Uh it's just having
the wherewithal to kind of think about
how you're placing it and how accurate
that might be.
Okay, so hopefully that was uh an
interesting exercise to get us started
on documentation. I am just going to go
back to my slideshow.
So hopefully you can yes see the slides
again. Excellent. Okay, so we're going
to get underway into actually talking
about documentation now. So I just want
to go back briefly to the types of
documentation that we're going to be
covering. So I mentioned data level and
study level which is also called project
level documentation which are the terms
that are used by the UK data archive. So
data level documentation provides
information about data objects. So this
could be things like variable
information in a data file or it could
be some of the demographic details of a
participant in an interview transcript
and that kind of documentation applies
specifically to that data subject or
that data object.
Study level or project level
documentation however is information
about the broader research project. So
what methods were used? It might be
summary of findings uh so on and so
forth. So it would apply across all of
the data objects. Right? So those are
the two types of um documentation that
we use at UK data service. These however
are not the only terms that's used to
describe documentation. So all of this
material from your research is called
different things by different policies
and different organizations. So archives
like UK data service we'll call it
documentation and under documentation we
have those specific kinds of
documentation. So we also have what we
call user guides, datal lists, data
dictionaries and we're going to talk
about all of those in this workshop.
There's also things like readme files
which you may have been asked to write
for us if you've ever deposited data
with us. Some policies like ESRC's data
policy refer to research materials which
doesn't really have a specific
definition as such but rather it refers
to anything that might have been used
across the research project. It also
refers to data assets which is the
system that is used to hold the data.
There's also metadata and the specific
kind of brand name of metadata like DDI.
So DDI is a specific schema of metadata
which gives you a distinct vocabulary
when you're trying to describe data.
I'll get a little bit more into metadata
in a bit. Um so while the concept of
documentation can be quite simple, the
deeper you go into it, the deeper some
of those ideas can run. So if you're
interested in exploring some of uh this
terminology, code data has a working
group that's dedicated to ter
terminology
um and it has some guidance on
terminology.
Hopefully some of the practical examples
that we have of the uh different types
of documentation will give you enough of
a of a flavor to really dive into any
area that you might want to know a
little bit more about.
So now you know what it is um and the
different levels and now you're
hopefully starting to see how embedded
these are in doing and disseminating
research.
But why do we do it? Well, documentation
from the point of an archive maximizes
the reuse value of data. Um so this is
essential for us to be able to review
and publish data. You can't understand
data basically without documentation.
If you were to just drop, you know, any
kind of data file like an interview
transcript in the middle of a street and
someone picked it up, they can't really
understand the data without
understanding the context in which it
was gathered. And this also adds a
historical value to the data. It builds,
you know, the provenence for it.
Essentially, documentation allows you to
expand on the methods and the processes
that might not normally get covered in a
publication. Um,
uh, there's endless space really for
documentation if you feel that it's
useful for understanding that data.
Doing so can also help enhance your
research outputs. So, documentation is
an output in and of itself. And
documentation itself can also be reused.
And I'll talk about some examples of
that later as well. It also adds a layer
of transparency to research. Uh as part
of the peerreview process, reviewers can
better understand the work and the data
and reusers can also more accurately and
efficiently reuse that data. And
finally, it aids in the creation of fair
research. So I am very briefly just
going to touch on fair data in a moment.
But um you know I just want to do
another uh quick exercise that may
highlight some of these points for you
um as to why documentation is so
important.
So let's just imagine for a moment that
you're looking at the table here. Can
you understand what the data represent?
We might be able to guess a meaning but
we need extra information to help us
understand and use this data.
So we add some additional metadata and
the data table is suddenly transformed
and the usability of it has dramatically
increased.
However, we might still be missing some
important pieces of contextual
information. What's in what you know
what is missing here? What other
information might help you make sense of
this?
And applying additional metadata then
further provides information about the
units in which the numerical data are
presented and it makes the data even
easier to interpret.
Okay, so hopefully that's a clear
demonstration of how context can matter
and how it can change your perception of
data. But why share the documentation in
data? And this particular point is uh
one of the key underpinning principles
to fair data. So fair principles and I
know some of you might already know what
fair is but fair principles are
relatively new guidelines or goals of
research which aim to make research more
transparent, collaborative and
constructive.
So since the early 2000s,
technology has had a massive impact on
how research is done. We can collect
more data which is more complex and we
can share it much quicker than we've
ever been able to do before.
However, despite collecting so much data
all of the time, we still have
challenges to processing that data. So
just think about any organization where
data is not shared between departments
um and you have to constantly re-enter
the same information. So how do we solve
this? Well in 2016 Wilkinson and and uh
his the rest of his team published what
are called the fair principles and that
outlined what good data management
looked like that would enable the
sharing and reuse of data. And the key
here is that these fair principles say
that the data need to be machine
readable. Um
so you know to be able to use technology
that you know massively change the re
waysearch is done is also transformative
to the reuse of that data. So it needs
to be machine readable and these
guidelines were so influential that an
international collaboration established
um the go far international support and
coordination office um and these fair
principles continue to influence
policies. So you'll see fair reference
references in uh data policies of the
research data alliance of the
association for European research
libraries um for those based in the UK
the UKRI and all of its research
councils will reference fair as well in
their data policies. Um, if you receive
any grants or taxpayer money to complete
research,
uh, chances are that you're going to be
asked to share the data and
documentation at the completion of the
project. And many publishers are now
also requiring the sharing of data and
research materials before publication in
the name of transparency and research
rigor. So even publishers guidelines are
underpinned by these fair data
principles.
And the fair principles state that data
should be findable, accessible,
interoperable, and reusable. So research
is not just something that you can
complete in the solitude of an academic
office, but it's something which is
completed in collaboration with others
and it's shared for further reuse. So to
make your data fair, this does mean also
documenting that data.
To make data findable and accessible
requires clear metadata and equally to
make data reusable means documenting the
provenence of the data. I won't go
further into the fair principles here.
Uh you can read a bit more about
yourself if you like. Um, but please do
send in any questions that you might
have about fair through the Q&A. I'm
happy to try and give a little bit more
detail if you're interested in it. But I
did just want to point out something
which spans across all of the fair
principles
and that is metadata. So metadata is
really the root of what is
documentation. It helps describe the
data in catalog pages. It helps describe
participants and it helps describe the
methods. There's a lot of uh grassroots
work to kind of build up what we call
schemas of metadata and to try and
synchronize some of those descriptions.
So, we're going to expand a little bit
more on metadata and I'm going to
introduce you to DDI.
Okay. some metadata. We're going to
cover what it is, what qualifies as good
metadata, what are some of the
standards, and how it's produced. So,
we're starting here with metadata
because all documentation is in some
sense metadata.
But what is metadata?
So, um [clears throat] the best
definition of metadata that I've heard
uh which I already mentioned is that
it's data about data. Essentially,
it is information that describes and
explains data. So, usually when we talk
about metadata, um we're we're often
talking about a standardized structured
information about data. Um the idea
behind the standardization is that it
will make it machine readable. So if you
go to our data catalog and you start
searching, our catalog is reading the
standardized metadata on the catalog
pages to bring back results. Because
metadata is what is underpinning our
cataloging and citing and discovering
and retrieving of those data
collections, it's really essential that
there's a particular structure to it for
our catalog to work. So basically you
need metadata in order to find the data
collections. So good complete metadata
becomes very important for the reuse of
data.
Importantly [clears throat]
metadata is something that should be
collected and recorded throughout the
research data life cycle. So it's not
simply something that you would think
about as you are depositing a data set
and then quickly trying to work
backwards to come up with the
information that you need. Um, it's not
just something that you would need when
you're reusing a data set, but instead
you should be thinking about metadata
along each step of creating, preserving,
and reusing data and thinking about what
metadata is available um, and what
metadata might enhance the collection.
But what does metadata actually depict?
All of this metadata aims to describe
the kind of journalistic five W's and
one H of data. Who created the data?
What does the data file contain? When
were the data created? Where were the
data created? And why were the data
created? And how were the data created?
So, we have lots of examples of how this
is done. But overall the the key
question is what would someone with no
prior knowledge of the project or data
need in order to be able to understand
or use that data correctly in their own
research.
Now some of the metadata that we capture
on the UKDS um catalog pages includes
these fields. So you have an abstract, a
set of keywords and topics which in our
catalog are based on standardized
schema. Um so you would search and
select from a list of specified keywords
and topics. Don't just necessarily fill
in your own. We've got um dates of
fieldwork, the country, information
about the sample including observation
or analytic unit. Um, we've got the
wider population, the number in the
sample, and then there's information
that's collected about the method and
the kind of data and any waiting where
applicable.
But if we are collecting this
information or tacitly, you might just
know this information about your own
research, why have you said it needs to
be structured? And I know some people
are really, you know, uh uncertain about
the idea of structuring uh information
into a standardized way. Well,
structured metadata is metadata that's
been standardized and uses a schema to
guide its description. So by using
structured metadata, it helps us
establish how one piece of information
relates to the other pieces of
information and it makes it clear for
computers to automatically extract
information from the metadata. And this
is obviously helpful for archives who
are looking to increase the findability
of data collections.
But it's also a useful comparative to
see how your research fits within the
larger field. And it helps you find
background literature. it will um uh
build context for your research
questions. Um and this information
provides the needed context that might
influence decisions that you make in the
analysis as well.
So what are metadata standards? Who
decides what these are? Well, a metadata
standard provides a framework to
establish a common set of
characteristics or attributes of data.
And these are very standardized from the
language, the spelling, even the format.
All of this is taken into account when
setting these standards. So if everyone
uses a different standard, it makes it
very difficult to find and or compare
data from different sources. A metadata
standard is a requirement which is
intended to establish a common
understanding of the meaning or the
semantics of the data. um and it
[clears throat] ensures the correct and
proper use and interpretation of the
data by its owners and users.
But there are different standards that
are based on disciplinary differences.
So consequently, sometimes it's
necessary for a translation program to
map between these standards to allow for
better findability.
So metadata standards allow for data
discovery and access consequent reuse of
the data interoperability of systems. So
being able to use data across different
systems and programs and the sharing of
data and metadata between communities be
those communities you know if you're a
data provider like an archive or a data
user or you know we do have different
types of of data users as well but all
of them would be able to use the same
kind of standard or understand the same
kind of standard.
Now there are a number of different
metadata schemas that can be used.
However, for those who are doing
research in social sciences, DDI is the
schema to understand best and it's also
what the UK data service uses to
underpin our data catalog. So DDI or the
data documentation initiative is a very
rich and detailed metadata standard
which has been developed by the DDI
alliance. There are a couple versions of
DDI. So we have the DDI code book and
the DDI life cycle. The DDI code book
gathers much more basic information and
it's often used to describe collections
at a higher level. So it's really useful
for things like data cataloges.
DDI life cycle on the other hand is a
little bit more complex and it allows
you to describe things like survey
questions or variables in addition to
the basic catalog information. So DDI
life cycle might be particularly useful
if you're looking at things like
longitudinal studies where similar
questions or variables might be reused
across different survey waves. It also
allows you to identify and compare
responses of the same variables.
So, just to point out, there are a lot
of different types of metadata standards
and schemas. Some have very specific
uses. Um,
you can use more than one standard to
help you describe data and you be may be
able to map or translate across
different standards if needed.
If you are planning to archive data, you
will need to provide uh metadata to
describe it. So this is usually gathered
by a form where you would add some
information about your data and its
provenence. Behind the form, the
information is categorized into metadata
elements that construct the catalog
record for your data. So, for the UK
data service catalog, um you know that
that little scrolling bit um that we
just had, let me see if I can run it
again so you can see it one more time.
These are all of our uh metadata DDI
metadata elements. So, we have a study
description and that's information about
the context of the data collection such
as a bibliographic citation and study
and data. Um the scope of the study. So
we have topics and geography and time.
Um we have methodology of data
collection, sampling and processing. Um
data access information and information
about accompanying materials. We also
have a data file description. So this is
information about the data format, the
file type, the file structure, if
there's missing data or waiting
variables.
We have variable descriptions and we
have keywords.
So it's good practice to use a uh
thesaurus
such as ELSET. So this is the European
language social science thesaurus.
Um and there are some multilingual uh
thesauri as well for social sciences.
So it is worth considering how you are
describing and sharing your metadata as
openly as possible. If you are archiving
with UK data service, this is kind of
done by default through the processes of
deposit. you will be asked to fill out
this standardized information.
Um,
if you are trying to do it yourself or
you are depositing in a kind of open
repository, an online repository or
something like that, um, you know, use
the fair principles. They can help.
There are some tools as well such as
metadata profiles that are available
that might also help you decide on which
metadata elements are most important
when describing your study.
Oops, I've gone too far. There we go.
Okay. So in addition to the DDI
structure, just to add another nuance,
another layer, in addition to that DDI
structure of the metadata for our
catalog pages, we also subscribe to what
are called controlled vocabularies. And
you'll notice on the catalog pages that
we have topics and keywords that you can
use to search for related data. So if
you click on any of those keywords
within our catalog page, you can be kind
of directed toward you can find other um
studies that use those same keywords as
well. So as I said before, those are not
just free text fields that the depositor
has kind of filled in. These are what we
call controlled vocabularies. So the UK
data archive is the guardian of the
humanities and social science electronic
thesaurus and the multilingual versions
of that including the European language
uh social science thesaurus and there is
a number of these vocabularities which
are usually discipline specific. So I'm
going to give you an opportunity now to
have a search through some different uh
controlled vocabulary.
So, uh, what I would like you to do, and
I'm just going to give you just a couple
of minutes for this, but if you could go
to bar talk.org,
and what this is is basically um,
uh, a search engine for controlled
vocabularies of your discipline. So,
once you go to baralk.org, you'll find
there's this really tempting search bar
right in the middle. um and you can
start searching for uh your discipline
or a general area of study and it should
bring up uh some controlled vocabularies
related to that. Now depending on what
your area of study is, what your
discipline is um you may or may not get
back a variety of controlled
vocabularies. It's it's highly
dependent. There's always one or two who
do this who um aren't able to find
anything. um probably because they are
studying quite niche things. So if you
try and keep it broad um as possible,
you'll probably get better results. But
there are some areas that have very
specific vocabularies as well. But see
what kinds of of things that you can
pull up um like I said, I'll give you
just a couple minutes with that um to do
some search.
If you do have any questions about it,
please do feel free to pop it in the Q&A
or the chat um and I'll have a
While you're doing some searching, there
are just a couple of questions that have
come up about uh the metadata that you
uh UK data service collects when you're
depositing data. Just to say when you're
going through the um data deposit
process, if you're going through reshare
and you're just you know um directly uh
depositing your data into reshare,
you'll find that there is a kind of
standardized online form that you have
to fill in and that form is the metadata
that is that is basically what you um
what you would need to fill in. So,
you'll you'll be prompted with is do you
have a grant number? You can search for
it. It might fill in the form for you,
but you'll be asked information about
your contact details. You'll be asked to
select um uh topics and keywords. Um
you'll be asked to describe, and there
are a few of those that are free text
fields. There are others that as you
start typing things, you'll get options
kind of pull up. So all of that is our
um standardized metadata schema that
feeds our catalog pages. Um so that's
that's through reshare. If if you are
doing a larger kind of national study um
if you're doing one of the national
surveys or something like that and you
are going through we have a different
route of curation for the for those
large national studies. Um, but it's the
same sort of thing is we'll we will have
an online form for you to fill in
basically, but um I'll have a look and
and answer those questions properly once
uh once we get toward the end here.
Okay, so it's been a couple minutes.
Hopefully, you've had a chance to just
have a look um and see if there are any
controlled vocabularies for your
discipline. Um, like I say, if if you
didn't get any s uh search hits, maybe
try a broader or related area and you
might get more hits that way.
Okay. So, I'm going to go into the
different types of documentation that we
collect for UK data service.
talked about metadata, which really is
the root of all documentation. And
hopefully you understand a little bit
better about why we are prompting you
with kind of structured metadata when
we're asking you to fill in information
for our catalog pages. But now we need
to get into some of the documentation
that's available for users to kind of
peruse as they're deciding which
collection from all of these search
results in the data catalog they might
uh pursue. So we're going to start with
data level documentation and then I'll
go on to project level documentation. So
I'll give you an overview of what data
level documentation is and some key
considerations for survey data and for
qualitative data.
Okay. So data level documentation is all
about the specific data files within
your collection. So this relates not to
the whole collection but to a specific
data file which could be for example a
single interview transcript. It could be
an SPSS file. Um it could be a single
image. So often this kind of metadata is
embedded within the file itself and
sometimes you can collect that
information as a standard part of data
collection.
So in survey uh data data level
documentation tends to be very
structured especially if you're using a
program like SPSS. So the software is
designed for you to add that metadata in
before you even begin analysis. So the
metadata then becomes embedded in with
the data. So for things like variable
names, there needs to be um a question
number system which will match the
questions in the questionnaire. Or if
you have a numerical system that you've
used, it should be clear what that
system is and how it relates to the
questionnaire to um uh uh you know how
the questionnaire relates to the
variable names as you've named them.
There should be meaningful
abbreviations. Um so something like gore
for example go is often used for
government office region. Um try and use
something that's recognizable. There
should be consistency and naming
conventions across the entire project
especially where there are different
data sets across the project. And
finally, for interoperability across
platforms, variables should be no longer
than eight characters and without
spaces.
Um, we have, you know, kind of similar
principles for variable names. So, the
variable labor labels should be brief
and concise. They should be no more than
80 characters. Where applicable they
should use a unit of measurement and
again if applicable describe the coding
or classification scheme that's used
including a reference. So for example
the standard occupational classification
2000 this is sock 2000 um can be
referenced.
Finally include a reference to the
question in the survey or questionnaire.
So the example here does all of that. So
Q9 BH XW is given the label Q9B hours
spent taking physical exercise in a
typical week. So the label clearly
includes the unit of measurement um and
a reference to the question. Uh not only
does this make it easier for reusers,
but I suspect it would make it a lot
easier for your own analysis as well.
You should add value labels as well,
making sure there are no outofbounds
values for categorical variables. So
avoid having blanks or zeros and instead
label your uh missing data with detail
where possible, differentiating between
not recorded, not provided, not
applicable, not known. Um so just all
things to think about really.
So here's an example of an SPSS variable
view which shows the variable the name
and you can see the label here as well
as the measurement and the uh missing
values
and here are the variable values which
include include um some of the missing
information too.
Now with transcripts data level
documentation is usually included at the
start of each interview. So for those
that are curated by the UK data archive,
you'll see this kind of information at
the start of every interview. So we
usually have the collection, the
principal investigator, plus some
demographic details about the
participant. And that can include things
like their sex, their socioeconomic
status, um the region that they're from,
and those details can vary depending on
what characteristics are deemed to be
important or influential to this study.
So maybe gender or sex wasn't as
important but um you know uh having
their race or ethnicity is. So you might
see some of these swapped out based on
what was important for that particular
study.
Images similarly will have metadata
about the image itself. Now we don't
have many collections with images. Um we
do have some lingering issues uh like
anonymization to deal with. um which may
impact the usefulness of the image. So,
we're kind of still working through how
we take some of this image or video data
um as an archive, but we do have some
examples including images with this is
this is from our Edwardians collection.
So, in these we usually have a caption
which describes the photo itself,
including anything that might be notable
about the image. You also might include
some historical notes to help
contextualize the image, and you can
also record characteristics like region,
year it was taken, who it was taken by,
and all of that would be useful metadata
to help you make sense of that image at
a later date.
Okay, so that's data level
documentation. I think it it's fairly
basic. It tends to be metadata that is
already collected and embedded in your
data files. But now we have something
that's a little bit more flexible which
is our study level or project level
documentation.
So um study or project level
documentation provides quite high level
information on the research context and
design and the data collection methods
that were used any data preparations and
manipulations. There might be summaries
of findings based on the data as well.
Um, this also tends to be a little bit
more flexible as to what you include and
how much you can include. So, I've got
some key considerations for survey data
again and for qualitative. And I've got
an absolute laundry list of some
examples of study level documentation
because you do get a lot of variation in
this kind of documentation.
So user guides are a key piece of
documentation that sit alongside data
collections and that includes further
information about the methods and the
fieldwork. It includes all this
information together and that is usually
openly available. So you don't need to
register in order to view it. You would
just click on the user guide link and it
opens up a PDF for you to have a look
at. Um, looking at the user guide alone,
you probably would be able to start
unpicking some of the intricacies of the
collection. Not information about your
data subjects itself necessarily. Um,
but you would have a really good idea
about the um creation of the data rather
than being embedded with the data like
the previous examples were. Things like
user guides are kept as a separate
document. and user guides do look
different across different collections.
There's no specific template to follow
as such, but rather it's tailored to
what information is available and what
specifics are of the collection. So I
I'm not going to go into the three user
guides here, but uh do check back on the
slides when they're posted on the
website and feel free to have explore of
these particular ones. I've chosen these
because they're quite different user
guides from each other. So there are
really good examples
within collections that are curated by
the UK data archive. These materials
would be collated into a user guide. So
the user guide is then bookmarked. So
you can see in the upper right corner um
well down the right side really of of
this um uh screenshot here you can see
some of those bookmarks and they just
tell you what materials are available
within that user guide and this tends to
encompass the project level
documentation. So while collections
curated by UK data archive do have user
guides, you're more than likely to see a
kind of folder of documents or separate
files on collections that are deposited
by researchers themselves. So we do have
these two curation routes. We have our
in-house curation where we would
amalgamate all of these documents into a
single user guide, but we also have our
reshare and this is self-curated. Um, so
when it's self-curated,
sometimes people will collate it into a
single user guide, but often you'll see
them as separate documents, but they do
have kind of, you know, they would make
a user guide if they actually collated
them into a single file, right? It's
just what we call it because when we
curate things, we put it all together
into one file.
Um, so this user guide here is an
example from a survey. So again, you can
see the bookmarks of the information
that were provided with this user guide
on the right hand side and that is
basically acting as a map to kind of
talk you through the preparation and the
delivery of the survey as well as any
other information that you might need to
know. So here you can see there's
further information about the sample,
the questionnaire, the data structure.
There's even a how to use this user
guide um within this within this one.
In addition to what you saw in that
example, other key documentation for
surveys might include things like
technical reports, information leaflets
and protocols. You might have blank
questionnaires with code books or survey
instructions. You could have coding
frames. um any information about known
errors or issues with the data. Um you
might even you know uh uh depositors
might also actually write up some of
this information specifically for the
user guide.
Qualitative collections will also have
user guides like this one that you see
here. This time the bookmarks is is on
the left hand side. So this one shows an
interview topic guide along with the
final report that was done for the
project. There's a blank consent form.
There's more information about the
sample. So a lot of times this is
collected um uh not collected, this is
produced rather as standard for the
research itself and it's just a matter
of actually collating some of those
pieces of documentation and pulling them
together. Um, but like I said, you can
write up things additionally for a user
guide.
Um, so here's some other examples for
qualitative work. And basically, this is
all of the information that you would
probably create along the way of doing
the research, but it's never seen
outside the research team or maybe at
best uh participants.
You can include things like interview
preparation, including instructions to
interviewers, prompts, uh topic guides.
You might have blank consent forms,
information sheets, any other materials
that the participant received prior to
taking part. It could also be text that
was written um by you explo you know
kind of expanding on the methodology or
sampling
including where it's permissible under
copyrights restrictions. It could
include though extracts of publications
or draft work as well. What we don't see
very often but can be quite useful are
things like research meeting minutes,
research diaries or field notes. um
documentation from the analysis, which
might include things like memos or code
books or initial analysis writeups. Um
so we've got close to 2,000 or so
collections of qualitative research. Um
and what we really don't see across any
of these is things on the analysis. And
I'm not sure why it's not a kind of
standard part of documentation.
Um this is pretty standard in
quantitative collections. you will have
a bit of code or syntax that you would
include so people can rerun the
analysis. Um and you know we can talk a
lot about analytic transparency and how
that can help validate research findings
and you know help reusers better
understand how decisions were made or
not made about the cleaning and
processing of the data. Um and that kind
of context that analytic material is
actually really really important for
qualitative work. Um this is kind of
part of the argument
for doing things qualitatively. You can
make [clears throat] different kinds of
decisions. So documenting how that was
done or why that was done is actually
quite important to explaining how you
got to your findings. Um
so the better you can kind of understand
that context hopefully that means that
you know all of those materials are
being actively considered by the
research team during the data collection
and analysis. So collating them and
making them available as documentation
would really help achieve what it is
that qualitative work is aiming to
achieve.
But yeah, the the analytic examples here
are not ones that that we often see in
qualitative collections,
but you can if you want to. If you have
collected that and you'd like to make it
available, by all means, you can. Okay.
So,
um here's an example of a collection
that rather than collating their user
guide into a single document as a user
guide, this is kind of remains separated
into different documents and that's just
because they went uh through the
self-curation route. So, they compiled
these and they just uploaded those
documents separately. So, that's how it
appears. But if you were to collate all
of those, that would essentially become
the user guide.
As an alternative to embedded metadata,
you can also have information that um is
in a structured document altogether,
like a code book or a data dictionary.
And this should contain detailed and
sufficient information about all of the
data items. So that includes variables
both new and derived frequencies,
command files that were used to create
derived m uh variables. We also have um
codebook creation tools. We have what's
called the DDI editor which is aimed at
uh data processing for curation
purposes. So this can be used prior to
depositing your collection in the
archive. We also have uh the nestar
publisher. So this is a lot of existing
information on these and how they can be
used. So please do check those links for
further information. Um
yeah and just if you want to have a
closer look at a code book um you can
see this one. This is from understanding
society teaching data set and you can
see here the variable name the label um
and then where applicable
[clears throat] options for responses or
the range of responses
unlike embedded metadata that I showed
you earlier on the SPSS files. This code
book can actually sit separately from
the data itself. So it can be made
available uh um openly.
Data dictionaries are very similar to
code books and they're um you know the
the terms are actually often used
interchangeably. Um they can also sit
within the study level documentation. So
often the data dictionary will contain
more information about the structure of
the database. So you can see here
they've documented that the variable is
numeric and scale level measurement. So,
it's got just a little bit more
information.
Um, if there was an equivalent for quali
uh collections, it would probably be a
data list. So, in addition to the user
guide, um, the UK data archive will also
curate a data list, which is this at a
glance look at participants. So this
would probably be considered perhaps
maybe this actually sits in a gray area
between data level documentation where
you have you know basic metadata basic
demographic details about each
participant
um listed but they've also got the file
names so you can kind of see the
structure of the data sets as a whole as
well. Um, so these data lists are not
exactly standardized or
all-encompassing. We do have a template
that you can follow, but you can adapt
it as you see fit or relevant for the
project itself. The data lists can take
a little bit of time to assemble. Um,
but they are really useful
organizational tools. Um, especially if
you're doing uh compiling something like
this during the research. So I would say
it's a good practice to kind of create
one as you go if you are doing a
qualitative project to kind of be aware
of the different um participants, the
different files that you're creating for
them and what some basic profile is
about them. But this kind of just gives
you your sample glance if you will um
for qualitative collections.
Other types of documentation include
observations that are written by
researchers in the moment of data
gathering.
Some types of methods dictate kind of
taking time for self-reflection as part
of the method. Um, and that can be used
as data. Others just recommend that this
is a good practice. So, what you see
here are comments. They're just a few
sentences that were written after every
semi-structured interview and survey. It
was a mixed methods um collection uh
called the affluent worker.
Um so this research was actually uh the
foundation for creating our current ONS
categories of class. So that this uh
resulted in what's called the gold Thorp
uh schema. But I digress. So these
interviewer comments basically help to
contextualize the relationship between
the researcher and the participant and
it can be really helpful data level
documentation.
This is quite unusual. I'm not sure if
you know people are veering away from
methods that would recommend this kind
of quick reflection or if it's just not
something that's often shared but it
really does help to rebuild power
dynamics of the interview. Um uh there
has been uh some interesting research
projects as well that is done on this
kind of reflections and what's called
paradata which is on the side margins of
surveys the little comments that might
be written in or the little amendments
that are made. There's some really
interesting research that specifically
dives into this kind of documentation
andor data depending on how you class
it. Um, but if you do have any kind of
reflections like that, you can include
it. It might be that it's documentation
describing the data collection or you
might be treating it as data too. It's
kind it's kind of a gray area.
A step further than field notes would be
um draft work of the analysis. So this
example comes from Dennis Marsden's
mothers alone collection. So he had this
piece on felt poverty which as it's
written here never actually made it to
publication but it was included with the
documentation for the collection. And
this is a really interesting collection
because it was led by white educated men
interviewing single mothers living on
welfare in the 1970s.
Um, and you might think for a moment
that actually the research team might
struggle to connect with their
participants given the vast ocean really
between their circumstances and their
participants circumstances. But I think
this piece in particular really provides
the context to show the sympathy that
the research team felt toward their
participants and it really gives you a
flavor of their own mindset too. So I
think this is really interesting um
documentation to have.
Um Annette Lawson did a study in the
1980s called adultery an analysis of
love and betrayal which was aiming to
explore a you know pretty taboo topic at
the time anyways of adultery and as such
it was really hard for her to recruit
participants. So Lawson chose to put out
a call for participants in a newspaper,
but it created an arguably biased
sample. So it was mostly white, mostly
middleclass women who responded to the
call for participants. So, as such,
Annette Lawson had a bit of a
preoccupation with her sample, and she
ended up writing a 54page defense of her
sample. And she started here with a
discussion on the ethical conundrums
that arose as part of the sampling
strategy, which included jealous
partners who were sending in information
on their married partners to
participate. or there was a man who
called from a psychiatric ward as well
about participating.
She then moved into an extensive
comparison between her sample and the
national population exploring what was a
significant difference from the national
population and whether or not that would
affect her data. So she kind of said
yes, you know, there is a lean toward
women who are middle class and white,
but have a look at some of these other
characteristics like urban versus rural
divide, and you'll find that actually
there's no difference between my sample
and the national.
And then finally, she comes to these
really interesting conclusions about
sampling strategies more broadly. um
including this point here that the
sampling needs to match the context of
the study and that exploratory studies
are benefited from the greater focus on
the ability to talk about the topic in
detail rather than the focus on uh who
the participants are as such. Um so
[clears throat] it's really fascinating.
it it never actually went to
publication. And if you read any of the
articles about that particular piece of
research, you'll find that there's this
really kind of um
uh uh almost sterile kind of um uh
paragraph that she would write about um
the call for recruitment. But if you
dive into her documentation, there's
actually an extensive thought process
behind it. So if you do write something
up like that, even if it doesn't make it
to publication, you can always put it in
with the documentation and that
documentation then is also a valid
output of that research project.
Um so earlier I had the example of
interviewer notes but field notes are
another example of very detailed
documentation. Field notes are a little
bit, you know, like the interviewer
notes. They they can occupy this gray
space of being both data and
documentation at different times. It's
worth pointing out that documentation is
normally something that's openly
available. So, field notes and other
reflections, depending on the level of
detail that's in them, may need to be
put under a similar access restriction
as data. Um, but we only have about I
think it's about a half a dozen
collections with examples of field notes
like this. Almost all of those studies
are ethnographic. So they put their
field notes in with the data itself. But
depending on what you are reflecting on
or reflecting about it, it may be
something that you do want to put into a
user guide.
And of course there are new
possibilities with changing technology.
So um if you are a qualitative
researcher using something like Envivo
or any other kind of computer assisted
qualitative data analysis software
um you you may be able to download your
code books, your memos, any mind maps
that you might be creating in there. So
here you see a list of nodes and their
description nodes is what uh Envivo call
them but the the kind of coding that
might be done as part of the initial
steps of um qualitative data analysis
all of that is very downloadable
um when we received the what's called
the Edwardians collection which is uh
453
80 plus page interviews of British
residents who were born during the
Edwardian period. So this is study
number 2000. It is the kind of um
founding collection of the qualitative
part of our archive. Um it included all
of the uh analysis that was done by
hand. So we had because this was done in
the uh I think it was done in the 1980s.
So we had 16 shelves that were dedicated
to holding the coding of the transcripts
into key themes. So they had printed out
all of the interview transcripts, cut
out the examples of data for this code,
they stapled them onto a page, and then
they put them into these binders, and
the binders were the themes. Um, so we
had 16 shelves of those. And of course
now with technology, um, you can
download those into a single file at any
point during or after coding.
Research teams can also use blogs and
websites to keep in touch with
participants and and keep a research
blog with updates about uh the progress
or it might post information for
participants or it might send out calls
for for participants as well. Once that
project is done, the site can then sit
alongside the project as a related
resource. And again, it's just providing
a little bit of additional
documentation.
And we do also see creative
documentation, too, such as this photo
story. Again, this is the sort of thing
that might be classed as either
documentation or potentially data. They
class it as as uh documentation,
but it it just gives you some scope to
think about how you might use video or
audio files as well to accompany the
data. Um, so if you have an interview
that you've done about um the collection
or if you were featured on the BBC or
what have you, you know, that that is
something that if you had the file and
permission to share, you would be able
to include that with your project as
well.
And then finally, changing technology is
not just for how research is done, but
also how we archive. So the UK data
service has created qualibbank which is
an online tool for searching, browsing
and citing qualitative data and as part
of that tool you can search and view um
qualitative data online and you can also
view linked documentation. So this
documentation can relate to the specific
data object or data subject um or it can
relate to the collection as a whole but
it allows for the distinction between
project level and data level
documentation and kind of includes
everything along with that particular
data file.
Finally, I would be remiss if I didn't
mention readme files. So in addition to
user guides and related resources and
embedded metadata, we also have what's
called the readme file. And this file
usually contains information about how
the data was prepared for ingest into
the archive. Within the archive, we will
automatically generate this and edit it
if you are doing our kind of in-house
curated uh route. But if you are uh
doing the self-curated route, you would
need to create this readme file and you
would be prompted for a readme file um
separately. But we have all kinds of
examples of these um and you can kind of
see what the um subject headings are and
how we describe some of that data
processing.
Okay. So
um I just wanted to make some a couple
of concluding remarks just a few I think
observations um really about reusing
documentation and then I also have a
data sharing checklist um as well to
share with you.
So reusing documentation. So often we
think about the reuse value of the data
specifically but less about how
documentation can have its own value as
well and documentation can serve as an
inspiration for good practices. So you
can look at other people's consent forms
and information sheets and adapt those
rather than just writing from scratch.
We also have hundreds of consent forms
that are documented within our
collections. So you can see a couple of
snippets here. One um is explaining what
data sharing means. Um the other kind of
talks about uh how the data is going to
be used. So do have an explore of some
of those and see what you can learn from
them in terms of good research practices
and you can do the same with um uh our
user guides. So if you are thinking
about uh data collection methods and
ways of collecting data again you can
you can explore the documentation uh
that people have um uh included with
their research. So one collection, the
foot and mouth disease in North Cumbria
deposited their interview guides and
their interview guides were reused by
medical students to better understand
doctor patient dynamics, what questions
to ask, when, how to build rapport, etc.
Um, I also I I teach dissertation
students and I refer my dissertation
students to find a similar interview
guide before setting out and making one
themselves. Um, I think it's it's really
useful to see what's important to ask
and what you may not ask um uh when
you're making decisions about your own
data collection methods. So, if you want
good practice on data collection, you
can use documentation as well.
You can also examine how to do research
with um
you know uh uh vulnerable groups and and
other groups that are just generally can
be quite challenging to think about how
you present information or how you
phrase questions. So we have some
examples here of um information letters
and consent forms that were used for
children specifically. Um, we also have
a lot of examples of documentation with
other vulnerable groups. So, if you are
working um with a hard-to-reach
population or a population that isn't
often included in research, it might be
worth going through the documentation on
other studies uh that have also explored
those populations simply because they
may have some really good practice
embedded in there that you might want to
adopt and adapt for your own project.
Okay. And a data sharing checklist as
promised as well. So we do have some
mandatory documentation that if you are
depositing with us, we would ask you
for. Um, every archive and repository
might be a little bit different. So you
need to make sure that you're checking
some of those guidelines to create the
necessary documentation files depending
on what kind of data that you're
sharing. There might also already be
templates available which you can just
download and adapt. So for example, we
have a template for a data list and a
readme file which you might want to have
a look at before making your own. Um for
us you will fill in what's called a data
offer form or data deposit form. So, uh,
you know, you log into your UK data
service account, you go into on the
lefth hand side data, and you'll see a
button there that, um, you would
[clears throat] like to offer a a data
set. So, try and fill that in with as
much detail as possible, which will
allow for us to be able to create
machine readable metadata from that. Um,
it would also, um, allow us to do a
really good appraisal of the collection.
Um and then finally make sure your data
files are uh contain that data level
documentation as well. Um especially for
qualitative I think for quantitative
researchers it's kind of embedded in
what you do. You know if you're if
you're going to create um an SPSS file
or something you know
um creating some of that embedded
metadata as a standard part of the
practice. So if it's not, just make sure
you're cognizant of that and create some
data level documentation where it's
needed.
So we do have a lot of um tools and
templates to use. We have model consent
forms and survey consent forms. We have
transcription templates, transcription
instructions, confidentiality
agreements, datal list template. We have
a lot of different tools and templates.
So don't feel like you have to create,
you know, recreate the wheel. you can uh
use that as a starting point and just
adapt it.
We also have some further resources on
data management. So we have um the
learning hub on research data
management. We also have um SESD does
the consortium for European and social
science data archives and they have a
data management expert guide as well. Um
we have our own text that we publish
through Sage called managing and sharing
research data best practices for
researchers.
Um
uh Closer also uh for those of you who
do quantitative work with some of our
national surveys might be uh familiar
with Closer. Um but they also have a
guide called understanding metadata. And
then of course if you're interested uh
in learning a little bit more about
those fair principles you can go to
goofair.org. or
so this is our data management guidance.
Um we published it all into one book.
There's lots of examples like the ones
that I've shown you here as well as
others. We have lots of training
exercises within um the book as well. So
do check that out if you would like uh
more information.
And we still have some more upcoming
events before the end of this academic
year. So have a look at our training uh
and events page uh to see if there's
anything else that you would like to
drop into.
And then uh as always we are uh on
social media so you can get connected if
you're interested in learning about uh
because we receive collections on
literally a daily basis and we do
publish uh new collections every week.
um you can sign up to our JISKmail list
and you will get uh an email
correspondence that lists what our new
additions are um for that week. So if
you want to know what data we're making
available as we make it available, sign
up to our just mail list and you can um
get updated on that. We're also on blue
sky and YouTube.
Um but yeah, thank you very much for uh
joining us. I know documentation isn't
probably the the most exciting topic to
ever attend a training on, but thank you
for bearing with us and hopefully you
have gotten a couple of useful bits of
information and uh yeah, please come
along to some of our future training
events um to learn a little bit more and
do fill in the evaluation survey if you
can. If you've got any suggestions for
improvements or other topics you'd like
to hear about, let us know uh so we can
have a look at that for next academic
year when we set our training schedule
for the year again. Thank you.