A complete, beginner-friendly data analysis workshop built for NGO staff, field teams, and new analysts — covering everything from what data even is to cleaning, analyzing, visualizing, and communicating findings to stakeholders.
This repo contains the full workshop package: slide deck, written training material, a hands-on Jupyter notebook, and a real "dirty" dataset to practice cleaning and analysis on.
data-analysis-workshop/
├── slides/ → Workshop slide deck (.pptx, built with Gamma)
├── notebooks/ → Hands-on practical notebook (.ipynb)
├── datasets/ → Practice dataset (intentionally messy, for cleaning exercises)
└── README.md
- Foundations — what is data, data vs. information vs. insight, data analysis, data types (nominal, ordinal, interval, ratio; structured vs. unstructured), and biased data
- The Data Life Cycle — Plan → Collect → Clean → Analyze → Visualize → Communicate → Store
- Data Collection & Tools — surveys, interviews, and field tools like Google Forms and KoboToolbox
- Data Cleaning — finding and fixing missing values, duplicates, inconsistent formatting, and invalid values
- Data Analysis — turning questions into answers, using SMART questions to guide analysis
- Data Visualization — matching chart types to the question being asked
- Stakeholders — tailoring findings to different audiences (donors, field staff, government, community members)
- Capstone Case Study — a full walkthrough tying every module together
- Python 3.9+
- Jupyter Notebook or JupyterLab
git clone https://github.com/<your-username>/data-analysis-workshop.git
cd data-analysis-workshop
pip install -r requirements.txt
jupyter notebook notebooks/HR_Workshop_Notebook.ipynbThe notebook is designed to be run live — every section explains why a step matters before showing the code, so it can be followed step-by-step in a workshop setting, not just read passively.
dirty_hr_dataset.xlsx is a synthetic HR dataset built specifically to contain realistic data
quality problems for teaching purposes:
- Duplicate employee records
- Inconsistent category spelling (e.g. department names)
- An attrition column coded 6 different ways (
Y,y,Yes,1, etc.) - Missing and placeholder values (
TBD,unknown) in the salary column - Negative salary values and out-of-range performance scores
- Mixed date formats and logically impossible future hire dates
- Broken text encoding in some name fields
No real employee data is used — all records are fictional.
Feel free to fork this repo, adapt the slides and exercises, and use them in your own training sessions. If you improve an exercise, fix a bug in the notebook, or add a new case study, pull requests are welcome.
This project is licensed under the MIT License — free to use, adapt, and share, with attribution appreciated.
Instructor : Aya N. Alharazin