An end-to-end Machine Learning pipeline and REST API service for multi-label classification of Amazon rainforest satellite imagery.
This project demonstrates a complete ML lifecycle: from reproducible model training and data versioning to deploying an optimized model via a web API. It is based on the Kaggle Planet: Understanding the Amazon from Space dataset.
- Deep Learning: PyTorch, PyTorch Lightning, ResNet18
- MLOps & Tracking: DVC (Data Version Control), ClearML, Hydra (Configuration)
- Inference & API: FastAPI, Uvicorn, ONNX Runtime
- Code Quality: Flake8, Black, Makefile
This is a monorepo containing both the research/training pipeline and the production service.
modeling/— Model training pipeline, data processing, and ONNX export.service/— FastAPI application for serving predictions.models/— Shared directory for model weights and ONNX graphs (tracked via DVC).
To run the REST API service locally and test the model:
git clone [https://github.com/guzelfey/satellite-classifier-service.git](https://github.com/guzelfey/satellite-classifier-service.git)
cd satellite-classifier-service
dvc pull # Downloads the latest model.onnx from remote storagecd service
python -m venv .venv
source .venv/bin/activate
make installmake runThe service will be available at http://127.0.0.1:8000. You can test the API via the interactive Swagger documentation at http://127.0.0.1:8000/docs.
The modeling/ directory contains a reproducible training pipeline. We use Hydra for hyperparameter management and ClearML for experiment tracking.
👉 View Training Metrics & Loss Curves in ClearML
cd modeling
make install
dvc pull # Downloads the training dataset
make train # Trains the model and saves checkpoints
make convert # Exports the best checkpoint to ONNX format (CPU-optimized)