Skip to content

Latest commit

 

History

4 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 

Repository files navigation

📊 Data Analysis Workshop — From A to Z

A complete, beginner-friendly data analysis workshop built for NGO staff, field teams, and new analysts — covering everything from what data even is to cleaning, analyzing, visualizing, and communicating findings to stakeholders.

This repo contains the full workshop package: slide deck, written training material, a hands-on Jupyter notebook, and a real "dirty" dataset to practice cleaning and analysis on.


📁 What's Inside

data-analysis-workshop/
├── slides/           → Workshop slide deck (.pptx, built with Gamma)
├── notebooks/         → Hands-on practical notebook (.ipynb)
├── datasets/          → Practice dataset (intentionally messy, for cleaning exercises)
└── README.md

🎯 What This Workshop Covers

  1. Foundations — what is data, data vs. information vs. insight, data analysis, data types (nominal, ordinal, interval, ratio; structured vs. unstructured), and biased data
  2. The Data Life Cycle — Plan → Collect → Clean → Analyze → Visualize → Communicate → Store
  3. Data Collection & Tools — surveys, interviews, and field tools like Google Forms and KoboToolbox
  4. Data Cleaning — finding and fixing missing values, duplicates, inconsistent formatting, and invalid values
  5. Data Analysis — turning questions into answers, using SMART questions to guide analysis
  6. Data Visualization — matching chart types to the question being asked
  7. Stakeholders — tailoring findings to different audiences (donors, field staff, government, community members)
  8. Capstone Case Study — a full walkthrough tying every module together

🚀 Getting Started with the Notebook

Prerequisites

  • Python 3.9+
  • Jupyter Notebook or JupyterLab

Setup

git clone https://github.com/<your-username>/data-analysis-workshop.git
cd data-analysis-workshop
pip install -r requirements.txt
jupyter notebook notebooks/HR_Workshop_Notebook.ipynb

The notebook is designed to be run live — every section explains why a step matters before showing the code, so it can be followed step-by-step in a workshop setting, not just read passively.


🧩 About the Dataset

dirty_hr_dataset.xlsx is a synthetic HR dataset built specifically to contain realistic data quality problems for teaching purposes:

  • Duplicate employee records
  • Inconsistent category spelling (e.g. department names)
  • An attrition column coded 6 different ways (Y, y, Yes, 1, etc.)
  • Missing and placeholder values (TBD, unknown) in the salary column
  • Negative salary values and out-of-range performance scores
  • Mixed date formats and logically impossible future hire dates
  • Broken text encoding in some name fields

No real employee data is used — all records are fictional.


🙌 Using This Material

Feel free to fork this repo, adapt the slides and exercises, and use them in your own training sessions. If you improve an exercise, fix a bug in the notebook, or add a new case study, pull requests are welcome.

📄 License

This project is licensed under the MIT License — free to use, adapt, and share, with attribution appreciated.

Instructor : Aya N. Alharazin

About

A complete data analysis workshop — from what data is, to bias, collection tools, cleaning, analysis, visualization, and stakeholder communication — with slides, a hands-on notebook, and a real messy dataset to practice on.

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages