A Machine Learning-powered web app that predicts the category of a resume (e.g., Data Science, Network Security Engineer, Advocate, etc.) using NLP and a Support Vector Machine classifier. Built using Python, scikit-learn, and Streamlit.
- ✅ Upload resume files in
.pdf,.docx, or.txtformats - 🔍 Clean and process resume content using NLP
- ✨ Predict job category using a trained ML model (SVM with TF-IDF)
- 📊 Supports over 25 resume categories
- 🌐 Simple, interactive interface using Streamlit
Examples include:
- Data Science
- Network Security Engineer
- Advocate
- Health and Fitness
- Web Designing
- Java Developer
- Python Developer
- Sales
- DevOps Engineer
- HR
- Civil Engineer
- Mechanical Engineer
- Blockchain Developer
... and many more
- Model: Support Vector Classifier (SVC) using One-vs-Rest strategy
- Vectorization: TF-IDF (Term Frequency-Inverse Document Frequency)
- Preprocessing:
- Remove URLs, mentions, hashtags
- Remove special characters and punctuations
- Lowercasing, whitespace removal
git clone https://github.com/anujayavidmal2002/NLP_ResumeScreeningApp.git
cd NLP_ResumeScreeningApppip install -r requirements.txtOr install them manually:
pip install streamlit scikit-learn pandas numpy matplotlib seaborn python-docx PyPDF2streamlit run app.py.
├── app.py # Main Streamlit app
├── clf.pkl # Trained classifier (SVC)
├── tfidf.pkl # TF-IDF vectorizer
├── encoder.pkl # LabelEncoder for category decoding
├── UpdatedResumeDataSet.csv # Original dataset (optional)
├── requirements.txt # Python dependencies
└── README.md # Project documentation
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.svm import SVC
from sklearn.multiclass import OneVsRestClassifier
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import LabelEncoder- Dataset:
UpdatedResumeDataSet.csv - Cleaned resumes
- Encoded with
LabelEncoder - Vectorized using
TfidfVectorizer - Model trained using
OneVsRestClassifier(SVC()) - Accuracy: ~90% (on balanced data)
You can test the app by uploading a resume in .pdf, .txt, or .docx format.
Example categories include:
- 🧠 Data Scientist
- 🧑⚕️ Health and Fitness
- 🔐 Network Security Engineer
- ⚖️ Advocate
"Experienced Python Developer with background in machine learning, TensorFlow..."
→ Predicted Category: Python Developer
"Certified advocate with 10 years in civil and family law..."
→ Predicted Category: Advocate
- Python 3.7+
- Streamlit
- scikit-learn
- pandas, numpy
- matplotlib, seaborn
- PyPDF2
- python-docx
Scikit-learn model was saved using version 1.7.0.
To avoid version conflicts or InconsistentVersionWarning, make sure to use the same version or retrain the model in your own environment.
This project is licensed under the MIT License.
Anujaya Vidmal
🖥️ GitHub: @anujayavidmal2002
📧 Email: anujayavidmal2002@gmail.com
🏫 University of Moratuwa
🎥 Watch the full demo and tutorial:
🔗 https://youtu.be/F3F6c0N12ls
- Dataset by Kaggle community
- Streamlit team for the amazing web app framework