Repository navigation
Data
MeshFL is demonstrated using the Mindboggle dataset (Klein & Tourville, 2012), consisting of 3D brain MRI volumes for gray matter (GM) and white matter (WM) segmentation.
A subset of 15 MRI samples is used for reproducibility and fast experimentation.
For the dataset structure and sample inspection, see MeshFL_dataset_read.ipynb and Colab version.
MeshFL uses a SQLite database (mindboggle.db) to store the dataset.
Unlike traditional pipelines that store file paths, MeshFL stores the data directly inside the database as binary objects (BLOBs).
The dataset is stored in a single table:
| Column | Type | Description |
|---|---|---|
| ID | INTEGER | Sample identifier |
| Image | BLOB | 3D MRI volume |
| Label | BLOB | Segmentation labels |
| GWlabels | BLOB | Gray/White matter labels |
| ANAlabels | BLOB | Additional anatomical labels |
Total samples: 15
The dataset schema includes multiple label columns:
-
GWlabels: Gray/white matter segmentation (used in MeshFL training) -
Label: Reserved for alternative label definitions -
ANAlabels: Anatomical labels for extended segmentation tasks
Currently, MeshFL uses only the GWlabels column for training:
SELECT Image, GWlabels FROM mindboggle101The additional label columns are included to support future extensions such as multi-class segmentation or alternative labeling schemes.
MeshFL simulates decentralized training using separate databases per site:
test_data/site-1/mindboggle.db
test_data/site-2/mindboggle.db
- Each site accesses only its local database
- No raw data is shared across sites
- Only gradients are exchanged during training
The dataset is accessed via the Scanloader, which:
- Reads BLOB data from the database
- Converts it into tensors
- Feeds it into the MeshNet model
This design removes dependency on external file paths and simplifies deployment.
Ground truth labels are initially stored in normalized form:
[0.0, 0.5, 1.0]
Before training, they are converted to class indices:
labels = (labels * 2).round().long()Final classes:
0: Background
1: Gray Matter
2: White Matter
This ensures compatibility with CrossEntropyLoss.
The dataset used in the examples is provided via Git LFS.
git lfs install
git lfs pull
The figure below shows one MRI sample retrieved directly from the site database, together with its associated gray/white matter ground-truth label.

MeshFL can be adapted to new datasets by:
- Creating a SQLite database with the same schema
- Inserting MRI volumes and labels as BLOBs
- Updating or extending site-specific databases
This database-driven design enables flexible dataset integration without modifying the core training code.
MeshFL can be adapted to custom MRI datasets by converting raw NIfTI images and their corresponding label volumes into a SQLite database that follows the same schema as the demo dataset. This allows users to preserve the existing MeshFL training pipeline while replacing the underlying data.
For the MeshFL dataset creation, see MeshFL_dataset_create.ipynb and Colab version.