Submind YouTube summaries
Thumbnail for Introduction to Data Cleaning with OpenRefine for Natural Science

Introduction to Data Cleaning with OpenRefine for Natural Science

Watch on YouTube

Video summary

Bu modülün odağı, doğal bilimler alanında veri bilim sürecinin en önemli ancak yeterince altı çizilmeden ele alınan aşaması olan veri temizlemesidir. Veri analizi, modelleme ve görselleştirme gibi adımlara hızlıca geçmek isteyen birçok kişi vardır; ancak analizlerin kalitesi doğrudan kullanılan verilerin kalitesine bağlıdır. Eğer veriler kirlenmişse, eksik ise, tutarsızsa veya tekrarlanmışsa, bu durum gelişmiş araçlar kullanılsa bile sonuçların zayıf olmasına neden olacaktır. Bu nedenle modülde açık kaynaklı OpenRefine adlı bir aracın pratik kullanımına odaklanılacak ve veri analize gönderilmeden önceki temizleme ve hazırlama süreçleri işlenecektir. Veri yönetimi sürecinin genel akışında, proje planlaması ile başlayıp ham verilerin Excel veya çevrimiçi formlar gibi araçlarla toplanmasıyla devam edilmektedir. Ancak bu aşamada genellikle yanlış yazım hataları, kategorizasyon eksiklikleri, kaybolan değerler, gereksiz boşluklar ve tarih formatlarındaki tutarsızlıklar gibi sorunlara rastlanır; veri temizleme tam olarak bu tür hataların giderilmesi için yapılan adımdır. Temizlenen veriler daha sonra yeni değişkenlerin oluşturulması yoluyla dönüştürülebilir veya R, Python veya Tableau gibi araçlarla analiz ve görselleştirme aşamalarına taşınabilir. Sonuçta ise temizlenmiş veri analitik bulgulara dönüşerek karar alma süreçleri için net bir araştırma çıktısı haline gelir. Modül içeriği kendi kendine çalışma videoları, egzersizler ve pratik aktivitelerden oluşmakta olup, katılımcılara kavramları netleştirmeleri ve araçların kendi araştırmalarına nasıl destek olabileceğini düşünmeleri için soru-cevap oturumları sunulacaktır. Ayrıca Dr. Bianca Peterson tarafından hazırlanan video materyallerinin kalitesi takdir edilmekte ve veri temizlemenin sadece teknik bir görev değil, aynı zamanda araştırma kalitesini artıran temel bir adım olarak görülmesi teşvik edilmektedir. Katılımcılardan süreçte acele etmeden her adımın neden önemli olduğunu anlamaya çalışmaları istenirken, geri bildirimlerin programların gelecekteki iyileştirilmesi için değerli bulunacağı vurgulanmaktadır.
Read the full video transcript
Good day, colleagues. Welcome to this module on data cleaning with Open Refine for Natural Science Track. My name is Elizabeth Boy. I'll be the facilitator for this module. Before I begin, let me introduce myself. I am currently the head of business intelligence and data architect at the University of the Western Cape. My work focuses on data analytics, business intelligence, institutional research, and building a stronger data culture within my institution. I have more than 23 years of experience in data analytics and business intelligence, as well as information management. I also have more than 16 years experience as a facilitator in data analytics, research, numeracy, and statistical literacies. I'm currently doing my second year of my PhD in operations research, where my research focuses on responsible learning analytics with the emphasis on fair- fairness, bias mitigation, and model explainability in higher education. I also serve as the president of the Southern African Association for Institutional Research. I'm the coordinator of the African Network for Learning Analytics Research, and I serve on the organizing committee of Deep Learning Indaba, South Africa. These roles, amongst others, enables me to contribute to data science, learning analytics, artificial intelligence, as well as research capacity development across South African universities, or South Africa as a whole, as well as the African continent. Now, let's talk about this module. The module focuses on one of the most important, yet underestimated part of the research data science process, data cleaning. Many people want to move quickly to the analysis, modeling, visualization, and reporting. However, the quality of your analysis depends strongly on the quality of your data. If the data is messy, incomplete, inconsistent, and duplicated, or even poorly structured, the analysis will be weak even when you use advanced tools for analysis. This is why data cleaning matters. In this module, we will use OpenRefine, which is an open-source tool, as a practical tool for cleaning and preparing data before we take it further the process. OpenRefine is useful because it allows us to inspect data carefully, identify inconsistencies, transform values, group similar entries, filter records, and document the steps we took when we do data cleaning. It also helps us with report with writing, especially in the research process. Now, looking at this data pipeline, often research starts with a project planning. This is where data your project is defined, the purpose of the work is defined, your research questions are articulated, the scope of your data set is outlined, as well as what the expected outcome or output should look like. The next step after project planning will be your data capturing or data collection. This is where raw data is collected or entered into most often tools like Excel, Qualtrics, uh Google Forms, Monkey Survey, and so forth, even with document. Once data is captured, we usually encounter it as a messy data. Why? Because this may include things like incorrect spelling, inconsistencies, categories that are not classified properly, missing values, extra spaces, duplicated records, and mixed date format, or even values that should not even appear in your data set. The fourth step, after >> [clears throat] >> you look at your data and interrogate it, is to clean it, which is the main focus of this module. In this module, we will learn how to use OpenRefine to remove all those errors, including removing inconsistencies and formats um that are inconsistent in your data set, but making sure that we prepare the data set for further analysis. Once the data is clean, it can move along through the process. It can be used as for as um in a form of transformation by creating new values, uh or it can be used for further analysis and visualization. At this stage, the tools like R, Python, Power BI, Tableau, Alteryx may be used to explore your data and summarize it the patterns as well as produce visual output. The final steps in terms of this uh data pipeline is your report writing and producing your final report. This is where the clean data and analyzed data is translated into findings and visualization, and it it's converted to a clear research output for decision making. In this module, if we move, >> [snorts] >> it's structured or comprises of self-study content, videos, as well as exercises and practical activities that linked to those exercises. There will also be two days dedicated to opportunities for question and answer session, reflection, reviews. These sessions should help you clarify concepts, ask practical questions, as well as reflect on how these tools that you are learning can support your own research or data work. Support also is available via Slack. At the end of the module also, you will be expected to complete the feedback form. We do value your um your feedback to help us improve the future data school program. I'll also like to acknowledge that all the videos that you will go through um in this module all prepared by Dr. Bianca Peterson from who was the previous facilitator for this course. I really appreciate the work and the amount of work that went into preparing this material. As you engage with this module, I encourage you to view data cleaning not only as a technical task, but also as a research quality task. Clean data helps us produce analysis that are more accurate, more transparent, and more credible. Please do work through the videos and exercises carefully. Do not rush the process. The aim is not only to complete the task, but to understand why each step matters in the process. Welcome to the module, and I look forward to engaging with you. Thank you.