Introduction to Data Cleaning with OpenRefine for Natural Science
Watch on YouTubeVideo summary
Bu modülün odağı, doğal bilimler alanında veri bilim sürecinin en önemli ancak yeterince altı çizilmeden ele alınan aşaması olan veri temizlemesidir. Veri analizi, modelleme ve görselleştirme gibi adımlara hızlıca geçmek isteyen birçok kişi vardır; ancak analizlerin kalitesi doğrudan kullanılan verilerin kalitesine bağlıdır. Eğer veriler kirlenmişse, eksik ise, tutarsızsa veya tekrarlanmışsa, bu durum gelişmiş araçlar kullanılsa bile sonuçların zayıf olmasına neden olacaktır. Bu nedenle modülde açık kaynaklı OpenRefine adlı bir aracın pratik kullanımına odaklanılacak ve veri analize gönderilmeden önceki temizleme ve hazırlama süreçleri işlenecektir.
Veri yönetimi sürecinin genel akışında, proje planlaması ile başlayıp ham verilerin Excel veya çevrimiçi formlar gibi araçlarla toplanmasıyla devam edilmektedir. Ancak bu aşamada genellikle yanlış yazım hataları, kategorizasyon eksiklikleri, kaybolan değerler, gereksiz boşluklar ve tarih formatlarındaki tutarsızlıklar gibi sorunlara rastlanır; veri temizleme tam olarak bu tür hataların giderilmesi için yapılan adımdır. Temizlenen veriler daha sonra yeni değişkenlerin oluşturulması yoluyla dönüştürülebilir veya R, Python veya Tableau gibi araçlarla analiz ve görselleştirme aşamalarına taşınabilir. Sonuçta ise temizlenmiş veri analitik bulgulara dönüşerek karar alma süreçleri için net bir araştırma çıktısı haline gelir.
Modül içeriği kendi kendine çalışma videoları, egzersizler ve pratik aktivitelerden oluşmakta olup, katılımcılara kavramları netleştirmeleri ve araçların kendi araştırmalarına nasıl destek olabileceğini düşünmeleri için soru-cevap oturumları sunulacaktır. Ayrıca Dr. Bianca Peterson tarafından hazırlanan video materyallerinin kalitesi takdir edilmekte ve veri temizlemenin sadece teknik bir görev değil, aynı zamanda araştırma kalitesini artıran temel bir adım olarak görülmesi teşvik edilmektedir. Katılımcılardan süreçte acele etmeden her adımın neden önemli olduğunu anlamaya çalışmaları istenirken, geri bildirimlerin programların gelecekteki iyileştirilmesi için değerli bulunacağı vurgulanmaktadır.
Read the full video transcript
Good day, colleagues. Welcome to this
module on data cleaning with Open Refine
for Natural Science Track.
My name is Elizabeth Boy. I'll be the
facilitator for this module.
Before I begin, let me introduce myself.
I am currently the head of business
intelligence and data architect at the
University of the Western Cape. My work
focuses on data analytics, business
intelligence, institutional research,
and building a stronger data culture
within my institution.
I have more than 23 years of experience
in data analytics and business
intelligence, as well as information
management.
I also have more than 16 years
experience as a facilitator in data
analytics, research, numeracy, and
statistical literacies.
I'm currently doing my second year of my
PhD in operations research, where my
research focuses on responsible learning
analytics with the emphasis on fair-
fairness,
bias mitigation, and model
explainability in higher education.
I also serve
as the president of the Southern African
Association for Institutional Research.
I'm the coordinator of the African
Network for Learning Analytics Research,
and I serve on the organizing committee
of Deep Learning Indaba, South Africa.
These roles, amongst others, enables me
to contribute to data science, learning
analytics, artificial intelligence, as
well as research capacity development
across South African universities, or
South Africa as a whole, as well as the
African continent.
Now, let's talk about this module.
The module focuses on one of the most
important, yet underestimated part of
the research data science process,
data cleaning.
Many people want to move quickly to the
analysis,
modeling, visualization, and reporting.
However,
the quality of your analysis depends
strongly on the quality of your data.
If the data is messy, incomplete,
inconsistent, and duplicated, or even
poorly structured, the analysis will be
weak even when you use advanced tools
for analysis.
This is why data cleaning matters.
In this module, we will use OpenRefine,
which is an open-source tool,
as a practical tool for cleaning and
preparing data before we take it further
the process.
OpenRefine is useful because it allows
us to inspect data carefully,
identify inconsistencies, transform
values, group similar entries, filter
records, and document the steps we took
when we do data cleaning. It also helps
us with report with writing, especially
in the research process.
Now, looking at this data pipeline,
often research starts with a project
planning. This is where data your
project is defined, the purpose of the
work is defined, your research questions
are articulated, the scope of your data
set is outlined, as well as what the
expected outcome or output should look
like.
The next step
after project planning will be your data
capturing or data collection.
This is where raw data is collected or
entered into most often tools like
Excel, Qualtrics, uh Google Forms,
Monkey Survey, and so forth, even with
document.
Once data is captured, we usually
encounter it as a messy data. Why?
Because this may include things like
incorrect spelling, inconsistencies,
categories that are not classified
properly,
missing values, extra spaces, duplicated
records, and mixed date format,
or even values that should not even
appear in your data set.
The fourth step, after
>> [clears throat]
>> you look at your data and interrogate
it, is to clean it,
which is the main focus of this module.
In this module, we will learn how to use
OpenRefine to remove all those errors,
including removing inconsistencies
and formats um that are inconsistent in
your data set, but making sure that we
prepare the data set for further
analysis.
Once the data is clean,
it can move along through the process.
It can be used
as for as um in a form of transformation
by creating new values,
uh or it can be used for further
analysis and visualization. At this
stage, the tools like R,
Python, Power BI, Tableau, Alteryx may
be used to explore your data and
summarize it the patterns as well as
produce visual output.
The final steps in terms of this uh data
pipeline is your report writing
and producing your final report.
This is where the clean data and
analyzed data is translated into
findings and visualization,
and it
it's converted to a clear research
output for decision making.
In this module,
if we move,
>> [snorts]
>> it's structured or comprises of
self-study content, videos,
as well as exercises and practical
activities that
linked to those exercises.
There will also be two days dedicated to
opportunities for question and answer
session, reflection, reviews. These
sessions should help you clarify
concepts, ask practical questions, as
well as reflect on how these tools that
you are learning can support your own
research or data work.
Support also is available via Slack. At
the end of the module also, you will be
expected to complete the feedback form.
We do value your
um your feedback to help us improve the
future data school program.
I'll also like to acknowledge
that all the videos that you will go
through
um in this module
all prepared by Dr. Bianca Peterson from
who was the previous facilitator for
this course. I really appreciate the
work and the amount of work that went
into preparing this material.
As you engage with this module, I
encourage you to view data cleaning not
only as a technical task, but also as a
research quality task.
Clean data helps us produce analysis
that are more accurate, more
transparent, and more credible.
Please do work through the videos and
exercises carefully. Do not rush the
process. The aim is not only to complete
the task, but to understand why each
step matters in the process.
Welcome to the module, and I look
forward to engaging with you.
Thank you.