Skip to content

About

Student retention analysis and success forecasting pipeline following CRISP-DM methodology.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

OULAD Student Learning Analytics & Dropout Prediction

This repository houses a comprehensive data mining and machine learning project focused on predicting student performance and identifying potential dropouts using the Open University Learning Analytics Dataset (OULAD). By merging demographic profiles with weekly student engagement logs on the Virtual Learning Environment (VLE), the project builds predictive models to support early educational intervention strategies.


📋 Table of Contents

  1. Project Overview
  2. Dataset Overview
  3. Repository Structure
  4. Methodology & Machine Learning Pipeline
  5. Key Visualizations & Figures
  6. Installation & Usage

🎯 Project Overview

In distance learning, students face high dropout rates due to isolation and lack of structured feedback. Identifying at-risk students early allows institutions to provide proactive support.

This project aims to:

  1. Model Student Success: Classify whether a student will pass, fail, or withdraw.
  2. Track Engagement: Analyze clicks and interactions with online resources (forums, lecture videos, questionnaires) to evaluate study patterns.
  3. Select Key Predictors: Use Recursive Feature Elimination (RFECV) to isolate the most predictive VLE features.
  4. Predict Scores: Train regression models to forecast final course grades.

📊 Dataset Overview

The project is built on the OULAD dataset, which contains data from 32,593 students across 7 modules:

  • Student Demographics: Gender, region, highest education level, age band, disability, deprivation index (IMD).
  • Course Structure: Module presentations, length, and registration details.
  • Assessment Scores: Weights and dates of exams and tutor-marked assessments.
  • Virtual Learning Environment (VLE) Activity: Clicks and daily interactions with course pages.

📂 Repository Structure

OULAD-Learning-Analytics/
├── data/                    # Processed datasets (df_model.csv, master OULAD files) (ignored by git)
├── docs/                    # Final academic report (result.docx)
├── figures/                 # Auto-generated plots (feature importances, confusion matrices)
├── notebooks/               # Jupyter Notebook walkthroughs
│   ├── Data_preparation_notes.ipynb      # Joining, cleaning, and aggregating student tables
│   ├── Failed_Dataset_notes.ipynb         # EDA on failed cohorts and VLE engagement analysis
│   ├── Supervised_Models_Notes_Collab.ipynb # Classifier pipeline runs (Random Forest, XGBoost)
│   ├── third.ipynb                        # Regression pipeline runs for grade prediction
│   └── Group_assignment_main_file.ipynb  # Comprehensive team execution workflow
├── src/                     # Core Package Modules
│   ├── __init__.py
│   ├── config.py            # Store constants, column lists, and model parameters
│   ├── data_preprocessing.py # Load raw files, clean VLE, demographics and splits
│   ├── feature_selection.py  # Target-leakage analysis and XGBoost-based RFECV
│   ├── classification.py     # Classification modeling (XGBoost, RF, LogReg, SVM)
│   ├── regression.py         # Grade prediction modeling (XGBoost, Random Forest)
│   └── clustering.py         # Unsupervised student segmentation (K-Means, DBSCAN, GMM)
├── main.py                  # CLI Driver script to run the pipeline
├── requirements.txt         # Package dependencies
└── README.md                # System documentation

🧠 Methodology & Machine Learning Pipeline

1. Data Aggregation & Preparation

  • Weekly Click Aggregations: Clicks are grouped into weekly bins for each student to capture engagement trajectories.
  • Table Merging: Demographics are joined with VLE interactions and historical assessment scores.
  • Handling Missing Values: Categorical features are imputed; missing deprivation indices (IMD) are filled.

2. Feature Selection (RFECV)

  • To prevent overfitting and eliminate noisy features, Recursive Feature Elimination with Cross-Validation (RFECV) is performed using an XGBoost Classifier.
  • Outcome: Identified 23 optimal features that maximize the cross-validated accuracy, significantly reducing model complexity.

3. Predictive Modeling

  • XGBoost Classifier: Multi-class model predicting Pass, Fail, Distinction, and Withdrawn.
  • XGBoost Regressor: Predicts continuous final grades.
  • Class Imbalance: Addressed through hyperparameter tuning and class weighting.

📈 Key Visualizations & Figures

The pipeline generates and stores several key visualizations inside the figures/ directory:

Feature Selection Curves

  • RFECV Selection: Cross-validated accuracy curve showing where adding more features stops improving accuracy. RFECV Curve

Feature Importance Charts

  • XGBoost Classifier Key Predictors: Impact of specific VLE clicks (like VLE homepage, forum posts, and quiz interactions) on final retention. XGB Classifier Feature Importances

  • XGBoost Regressor Key Predictors: Highlights the VLE elements most strongly correlated with high exam scores. XGB Regressor Feature Importances

Model Confusion Matrices

  • XGBoost Classification Matrix: Heatmap displaying predicted vs. actual student outcomes. Confusion Matrix Heatmap

🚀 Installation & Usage

Setup Environment

# Navigate to project
cd OULAD-Learning-Analytics

# Create and activate virtual environment
python -m venv .venv
source .venv/bin/activate  # On Windows: .venv\Scripts\activate

# Install dependencies
pip install -r requirements.txt

Running the Python CLI Pipeline

To run the analysis using the modular python package:

# Preprocess raw data if CSV files are placed in data/
python main.py --preprocess

# Train classification models (XGBoost, Random Forest, Logistic Regression, SVM)
python main.py --classify

# Train regression models (XGBoost Regressor, Random Forest Regressor)
python main.py --regress

# Run unsupervised clustering (K-Means, DBSCAN, GMM)
python main.py --cluster

# Run all pipeline stages in sequence
python main.py --run-all

Running the Jupyter Notebooks

To run the analysis visually:

  1. Make sure you place the raw OULAD CSV files in the data/ folder.
  2. Spin up a Jupyter Lab server:
jupyter lab
  1. Open notebooks/Group_assignment_main_file.ipynb to view the full pipeline from data aggregation to model evaluation.

🛠️ System Architecture

graph TD
    A[OULAD Demographic & VLE Datasets] --> B[Data Cleaning & Imputation]
    B --> C[Feature Engineering & Engagement Aggregation]
    C --> D[DBSCAN Clustering - Student Segmentation]
    C --> E[XGBoost Classifier - Risk Prediction]
    D --> F[Clustering Profiles & Engagement Analysis]
    E --> G[Student At-Risk Predictions]
    F --> H[Interactive HTML Decision-Support Dashboard]
    G --> H
Loading

About

Student retention analysis and success forecasting pipeline following CRISP-DM methodology.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages