Skip to content
Mohamed E. Masoud edited this page Apr 9, 2026 · 9 revisions

MeshFL is demonstrated using the Mindboggle dataset (Klein & Tourville, 2012), consisting of 3D brain MRI volumes for gray matter (GM) and white matter (WM) segmentation.

A subset of 15 MRI samples is used for reproducibility and fast experimentation.

For the dataset structure and sample inspection, see MeshFL_dataset_read.ipynb and Colab version.


Database Format

MeshFL uses a SQLite database (mindboggle.db) to store the dataset.

Unlike traditional pipelines that store file paths, MeshFL stores the data directly inside the database as binary objects (BLOBs).


Database Structure

The dataset is stored in a single table:

Table: mindboggle101

Column Type Description
ID INTEGER Sample identifier
Image BLOB 3D MRI volume
Label BLOB Segmentation labels
GWlabels BLOB Gray/White matter labels
ANAlabels BLOB Additional anatomical labels

Total samples: 15

The dataset schema includes multiple label columns:

  • GWlabels: Gray/white matter segmentation (used in MeshFL training)
  • Label: Reserved for alternative label definitions
  • ANAlabels: Anatomical labels for extended segmentation tasks

Currently, MeshFL uses only the GWlabels column for training:

SELECT Image, GWlabels FROM mindboggle101

The additional label columns are included to support future extensions such as multi-class segmentation or alternative labeling schemes.


Multi-Site Dataset Organization

MeshFL simulates decentralized training using separate databases per site:

test_data/site-1/mindboggle.db
test_data/site-2/mindboggle.db
  • Each site accesses only its local database
  • No raw data is shared across sites
  • Only gradients are exchanged during training

Data Loading in MeshFL

The dataset is accessed via the Scanloader, which:

  • Reads BLOB data from the database
  • Converts it into tensors
  • Feeds it into the MeshNet model

This design removes dependency on external file paths and simplifies deployment.


Label Processing

Ground truth labels are initially stored in normalized form:

  • [0.0, 0.5, 1.0]

Before training, they are converted to class indices:

labels = (labels * 2).round().long()

Final classes:

0: Background

1: Gray Matter

2: White Matter

This ensures compatibility with CrossEntropyLoss.


Data Access

The dataset used in the examples is provided via Git LFS.

To download:

git lfs install
git lfs pull

Example Sample Visualization

The figure below shows one MRI sample retrieved directly from the site database, together with its associated gray/white matter ground-truth label.

Dataset sample overlay


Extending the Dataset

MeshFL can be adapted to new datasets by:

  • Creating a SQLite database with the same schema
  • Inserting MRI volumes and labels as BLOBs
  • Updating or extending site-specific databases

This database-driven design enables flexible dataset integration without modifying the core training code.

MeshFL can be adapted to custom MRI datasets by converting raw NIfTI images and their corresponding label volumes into a SQLite database that follows the same schema as the demo dataset. This allows users to preserve the existing MeshFL training pipeline while replacing the underlying data.

For the MeshFL dataset creation, see MeshFL_dataset_create.ipynb and Colab version.


Clone this wiki locally