π Docs: README Β· Architecture Β· Concepts Β· License Β· Security
Industrial air compressors are the "lungs" of manufacturing plants. Unplanned downtime can cost thousands of dollars per hour. This project implements an end-to-end MLOps pipeline to predict multi-component failures (Bearings, Radiators, Pumps) using real-time sensor data (Vibration, Temperature, Pressure).
- Production Architecture: Transitioned from experimental Jupyter Notebooks to a modular Python package.
- Time-Series Engineering: Implemented rolling window statistics and lag features to capture equipment degradation.
- Robust Validation: Utilized
TimeSeriesSplitto prevent data leakage and ensure temporal reliability. - Deployment Ready: Containerized FastAPI service for real-time inference.
The system is designed as a modular pipeline to ensure scalability and maintainability.
- Ingestion: Robust loading with structured logging and error handling.
- Preprocessing: Scikit-Learn Pipelines with
RobustScalerfor sensor outlier handling. - Feature Engineering: Generation of rolling mean/std-dev for mechanical vibration (GACC/HACC) and thermal sensors.
- Experiment Tracking: MLflow manages hyperparameters, metrics (F1-Score), and model versioning.
- Inference: FastAPI endpoint serving predictions via a Docker container.
βββ api/
β βββ main.py # FastAPI implementation (/health, /predict)
βββ data/raw/ # Sensor CSV data
βββ src/ # Core Logic
β βββ config.py # Shared column/experiment constants
β βββ ingestion.py # Data loading & logging
β βββ preprocessing.py# Sklearn Transformation Pipelines
β βββ features.py # Rolling window & lag engineering
β βββ train.py # MLflow training logic with TimeSeriesSplit
β βββ predict.py # Loads latest MLflow models & serves predictions
βββ Dockerfile # Containerization for production
βββ pyproject.toml / uv.lock # Project dependencies (uv)
Model artifacts aren't stored as local .pkl files β each training run logs a
self-contained preprocessing+model pipeline to MLflow, and src/predict.py
loads the latest run per target at serving time.
- Python 3.13+
- Docker (Optional for containerization)
git clone https://github.com/Prafful-Vyas/Air-Compressor-predictive-maintenance-using-ML.git
cd Air-Compressor-predictive-maintenance-using-ML
uv sync
Run the training pipeline to log metrics to MLflow:
python -m src.train
mlflow ui # View results at http://localhost:5000
Training registers a new candidate model version per target in the MLflow Model Registry, but that version is not automatically served β see step 4.
src/predict.py only ever loads the model version holding the production
alias (configurable via ACPDM_MODEL_REGISTRY_ALIAS) for each target, so a
freshly trained run has zero effect on what the API serves until you
explicitly promote it:
python -m src.promote list bearings # see candidate versions + their holdout_f1
python -m src.promote promote bearings 1 # attach the 'production' alias to version 1
python -m src.promote current bearings # confirm what's currently promotedPromotion refuses a version whose holdout_f1 is below
ACPDM_MIN_HOLDOUT_F1_FOR_PROMOTION (default 0.0, i.e. no gate) unless you
pass --force. Re-running promote with an older version number is an
instant rollback. Repeat for every target (bearings, wpump, radiator,
exvalve) β the API reports "degraded" at /health/ready for any target
that hasn't been promoted yet.
uvicorn api.main:app --reload
Navigate to http://localhost:8000/docs to test the interactive Swagger API.
/predict requires an X-API-Key header matching ACPDM_API_KEY (see
.env.example); /health and /health/ready are unauthenticated.
The image runs as a non-root user and serves the API only; it loads
promoted models from an MLflow tracking store at startup, so mount the
local mlruns/ directory produced by step 3 (or point
ACPDM_MLFLOW_TRACKING_URI at a remote tracking server):
docker build -t air-compressor-api .
docker run -p 8000:8000 \
-e ACPDM_API_KEY=change-me-in-production \
-v "$(pwd)/mlruns:/app/mlruns" \
air-compressor-api
For a prod-like local stack with a real MLflow tracking server instead of
the file-based default, use docker-compose.yml:
cp .env.example .env # then set ACPDM_API_KEY to something real
docker compose up --build -d
docker compose exec api uv run python -m src.train
docker compose exec api uv run python -m src.promote promote bearings 1
curl localhost:8000/health/ready
uv run pytest # unit + integration tests
uv run ruff check . # lint
uv run ruff format . # auto-format
pre-commit install will run lint/format automatically on each commit
(config in .pre-commit-config.yaml). The same checks run in CI on every
push and pull request against main (.github/workflows/ci.yml).
Instead of simple accuracy, this project prioritizes F1-Score and Precision-Recall due to the class imbalance inherent in machinery failure data.
| Target | Mean CV F1 (5-fold) | Holdout F1 |
|---|---|---|
| Bearings | 0.80 | 0.63 |
| Water Pump | 0.84 | 0.84 |
| Radiator | 0.83 | 0.88 |
| Exhaust Valve | 0.67 | 0.79 |
- Validation Strategy: 5-Fold
TimeSeriesSplitfor a robustness estimate, plus a chronological 80/20 holdout as the reported test metric. Numbers above are from a single run and will shift as the training data grows β rerunpython -m src.trainand update this table periodically rather than treating it as fixed.
- Language: Python
- ML Libraries: Scikit-Learn, Pandas, NumPy
- MLOps: MLflow (experiment tracking + model registry)
- Backend: FastAPI, Uvicorn, Pydantic, pydantic-settings
- DevOps: Docker (multi-stage, non-root), Docker Compose, GitHub Actions (CI: ruff, mypy, pytest+coverage, pip-audit, docker build)
- Quality/Security tooling: ruff, mypy, pytest-cov, pip-audit, gitleaks (pre-commit)
To transition this from a standalone project to an enterprise-grade Industrial IoT (IIoT) platform, the following enhancements are proposed:
Currently, the system processes static CSVs. Integrating Apache Kafka would allow the pipeline to ingest high-frequency sensor data streams directly from PLC (Programmable Logic Controller) systems in a factory setting.
Implementing a monitoring dashboard to track:
- Model Drift: Detecting when the compressor's physical characteristics change (e.g., after a major part replacement).
- Latency: Monitoring the inference time of the FastAPI endpoint.
- System Health: Tracking CPU/Memory usage of the Docker containers.
Transition from binary classification ("Will it fail?") to Remaining Useful Life (RUL) estimation using LSTMs or GRUs. This provides a countdown (e.g., "14 days until bearing failure"), allowing for much better maintenance scheduling.
Setting up GitHub Actions or Airflow to trigger a "Continuous Training" (CT) pipeline. When model performance drops below a certain F1-score threshold, the system would automatically retrain on the latest 3 months of sensor data and promote the best model to production.
Optimizing the model using ONNX or TensorRT to deploy the inference engine directly onto "Edge" devices (like an NVIDIA Jetson or Raspberry Pi) located physically on the air compressor, reducing the need for constant cloud connectivity.
This project went through a hardening pass (env-based config, API-key auth, CORS, input validation, an MLflow model-registry promotion gate, a non-root multi-stage Docker image + compose stack, and an expanded CI pipeline with coverage/type/vulnerability gates) that deliberately stayed scoped to hardening what already existed. Explicitly out of scope for that pass, each for a specific reason rather than an oversight:
- Kafka/streaming, Grafana/Prometheus, RUL/LSTM modeling, edge deployment β larger feature work, tracked above under "Future Improvements."
- Cloud-specific IaC / managed services β the deployment target is deliberately cloud-agnostic (Docker + docker-compose only); pick your own infra on top of the image.
- OAuth2/JWT β a static API key (
ACPDM_API_KEY) is the intentional minimal auth model for a single-service-to-service endpoint; revisit if multiple external clients need distinct identities/scopes. - DVC / git-lfs β the raw dataset is ~239KB and committed directly; fine short-term, worth revisiting if the dataset grows substantially.
- A Makefile/task runner β low-priority DX polish, not blocking anything.
- Upgrading MLflow past 3.13 β see the comment on
mlflow's pin inpyproject.toml: newer versions disable the file-based tracking backend this project's local dev setup relies on by default, which needs a deliberate migration to a database backend, not a version bump.