This repository houses a comprehensive data mining and machine learning project focused on predicting student performance and identifying potential dropouts using the Open University Learning Analytics Dataset (OULAD). By merging demographic profiles with weekly student engagement logs on the Virtual Learning Environment (VLE), the project builds predictive models to support early educational intervention strategies.
- Project Overview
- Dataset Overview
- Repository Structure
- Methodology & Machine Learning Pipeline
- Key Visualizations & Figures
- Installation & Usage
In distance learning, students face high dropout rates due to isolation and lack of structured feedback. Identifying at-risk students early allows institutions to provide proactive support.
This project aims to:
- Model Student Success: Classify whether a student will pass, fail, or withdraw.
- Track Engagement: Analyze clicks and interactions with online resources (forums, lecture videos, questionnaires) to evaluate study patterns.
- Select Key Predictors: Use Recursive Feature Elimination (RFECV) to isolate the most predictive VLE features.
- Predict Scores: Train regression models to forecast final course grades.
The project is built on the OULAD dataset, which contains data from 32,593 students across 7 modules:
- Student Demographics: Gender, region, highest education level, age band, disability, deprivation index (IMD).
- Course Structure: Module presentations, length, and registration details.
- Assessment Scores: Weights and dates of exams and tutor-marked assessments.
- Virtual Learning Environment (VLE) Activity: Clicks and daily interactions with course pages.
OULAD-Learning-Analytics/
├── data/ # Processed datasets (df_model.csv, master OULAD files) (ignored by git)
├── docs/ # Final academic report (result.docx)
├── figures/ # Auto-generated plots (feature importances, confusion matrices)
├── notebooks/ # Jupyter Notebook walkthroughs
│ ├── Data_preparation_notes.ipynb # Joining, cleaning, and aggregating student tables
│ ├── Failed_Dataset_notes.ipynb # EDA on failed cohorts and VLE engagement analysis
│ ├── Supervised_Models_Notes_Collab.ipynb # Classifier pipeline runs (Random Forest, XGBoost)
│ ├── third.ipynb # Regression pipeline runs for grade prediction
│ └── Group_assignment_main_file.ipynb # Comprehensive team execution workflow
├── src/ # Core Package Modules
│ ├── __init__.py
│ ├── config.py # Store constants, column lists, and model parameters
│ ├── data_preprocessing.py # Load raw files, clean VLE, demographics and splits
│ ├── feature_selection.py # Target-leakage analysis and XGBoost-based RFECV
│ ├── classification.py # Classification modeling (XGBoost, RF, LogReg, SVM)
│ ├── regression.py # Grade prediction modeling (XGBoost, Random Forest)
│ └── clustering.py # Unsupervised student segmentation (K-Means, DBSCAN, GMM)
├── main.py # CLI Driver script to run the pipeline
├── requirements.txt # Package dependencies
└── README.md # System documentation
- Weekly Click Aggregations: Clicks are grouped into weekly bins for each student to capture engagement trajectories.
- Table Merging: Demographics are joined with VLE interactions and historical assessment scores.
- Handling Missing Values: Categorical features are imputed; missing deprivation indices (IMD) are filled.
- To prevent overfitting and eliminate noisy features, Recursive Feature Elimination with Cross-Validation (RFECV) is performed using an XGBoost Classifier.
- Outcome: Identified 23 optimal features that maximize the cross-validated accuracy, significantly reducing model complexity.
- XGBoost Classifier: Multi-class model predicting
Pass,Fail,Distinction, andWithdrawn. - XGBoost Regressor: Predicts continuous final grades.
- Class Imbalance: Addressed through hyperparameter tuning and class weighting.
The pipeline generates and stores several key visualizations inside the figures/ directory:
- RFECV Selection: Cross-validated accuracy curve showing where adding more features stops improving accuracy.

-
XGBoost Classifier Key Predictors: Impact of specific VLE clicks (like VLE homepage, forum posts, and quiz interactions) on final retention.

-
XGBoost Regressor Key Predictors: Highlights the VLE elements most strongly correlated with high exam scores.

# Navigate to project
cd OULAD-Learning-Analytics
# Create and activate virtual environment
python -m venv .venv
source .venv/bin/activate # On Windows: .venv\Scripts\activate
# Install dependencies
pip install -r requirements.txtTo run the analysis using the modular python package:
# Preprocess raw data if CSV files are placed in data/
python main.py --preprocess
# Train classification models (XGBoost, Random Forest, Logistic Regression, SVM)
python main.py --classify
# Train regression models (XGBoost Regressor, Random Forest Regressor)
python main.py --regress
# Run unsupervised clustering (K-Means, DBSCAN, GMM)
python main.py --cluster
# Run all pipeline stages in sequence
python main.py --run-allTo run the analysis visually:
- Make sure you place the raw OULAD CSV files in the
data/folder. - Spin up a Jupyter Lab server:
jupyter lab- Open
notebooks/Group_assignment_main_file.ipynbto view the full pipeline from data aggregation to model evaluation.
graph TD
A[OULAD Demographic & VLE Datasets] --> B[Data Cleaning & Imputation]
B --> C[Feature Engineering & Engagement Aggregation]
C --> D[DBSCAN Clustering - Student Segmentation]
C --> E[XGBoost Classifier - Risk Prediction]
D --> F[Clustering Profiles & Engagement Analysis]
E --> G[Student At-Risk Predictions]
F --> H[Interactive HTML Decision-Support Dashboard]
G --> H
