Skip to content
xmed-labPublic

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

9 Commits

Folders and files

Repository files navigation

SyRe

Synergistic Vision-Language Reinforcement Enables Scalable On-Demand Analysis across Diverse Clinical Tasks

arXiv Model Dataset Demo

SyRe is a text-driven segmentation foundation model for biomedical images. Given an image and a natural-language target such as “liver tumor” or “optic disc”, SyRe predicts a pixel-level mask without training a separate model for every class.

Research release. The model is intended for research and development. It is not a medical device and should not be used as a standalone clinical decision maker.

📌 Overview

SyRe connects language understanding with SAM-style dense prediction through a closed-loop vision-language architecture:

SyRe framework

SyRe couples vision-language reasoning with SAM-based dense prediction for on-demand biomedical segmentation.

The accompanying study describes SyReData as a large image-mask-text collection spanning nine imaging modalities and 229 segmentation tasks. See the paper and interactive demo for the complete experimental protocol and reported results.

📣 Latest updates

  • Public checkpoint: the released SyRe weights are available on Hugging Face.
  • Public data: the 2D image/mask collection is available on SyReData.
  • Reproducible utilities: this repository now includes native 2d_test inference and batch evaluation scripts.
  • Interactive demo: try SyRe at 143.89.46.197:7860.

The public model repository is about 16.3 GB and the public dataset repository is about 379 GB. Download only the files or subsets you need.

📊 Results and visual examples

Quantitative results

SyRe quantitative results on internal datasets

Quantitative comparison on internal datasets.

SyRe quantitative results on external datasets

Quantitative comparison on external datasets.

SyRe quantitative results on in-house datasets

Quantitative comparison on in-house datasets.

Qualitative results

SyRe qualitative results across biomedical imaging modalities

Representative predictions across CT, MRI, ultrasound, pathology, dermoscopy, X-ray, fundus, endoscopy, and PET images.

SyRe qualitative results on in-house tasks

Representative qualitative results on in-house segmentation tasks.

SyRe pathology application

Example pathology application showing SyRe predictions across biomedical image patches.

🛠️ Installation & setup

We recommend Python 3.10. Install a PyTorch build matching your CUDA version first, then install the repository dependencies:

conda create -n syre python=3.10 -y
conda activate syre

# Choose the PyTorch/CUDA wheel for your machine.
python -m pip install torch torchvision
python -m pip install -r requirements.txt
python -m pip install huggingface_hub

The command-line help can be used without downloading the model. Full inference is GPU-oriented and memory intensive.

📦 Prepare the public resources

Model weights

Use the Hub ID directly, or cache the checkpoint locally:

python - <<'PY'
from huggingface_hub import snapshot_download

snapshot_download(
    repo_id="McGregorW/SyRe",
    local_dir="checkpoints/SyRe",
)
PY

Both of the following forms are accepted by the inference script:

--model McGregorW/SyRe
--model checkpoints/SyRe

Dataset

The dataset contains train and test splits. Downloading the complete repository is optional; use allow_patterns when you only need selected files:

python - <<'PY'
from huggingface_hub import snapshot_download

snapshot_download(
    repo_id="McGregorW/SyReData",
    repo_type="dataset",
    local_dir="data/SyReData",
    # Example: allow_patterns=["*test*", "**/*.png"]
)
PY

The native evaluation path follows GLaMMedv16 and reads image2label_2d_test.json directly. The public dataset directory should therefore contain that index together with its referenced images and masks.

🚀 Quick start

Single-sample prediction on 2d_test

python examples/inference.py \
  --model McGregorW/SyRe \
  --dataset-dir data/SyReData \
  --index 0 \
  --output-dir outputs/sample_000000 \
  --device cuda:0 \
  --dtype bf16 \
  --save-visualization

The command uses the same dataset, prompt construction, image preprocessing, and [SEG] batch format as the GLaMMedv16 2D test evaluator. It writes one binary mask per annotated class and inference.json.

For use from Python, see examples/inference_api.py. The core calls are:

import torch
from scripts.syre_2d_test import load_model, load_2d_test_dataset
from scripts.syre_2d_test import make_2d_test_loader, predict_batch

bundle = load_model("McGregorW/SyRe", device="cuda:0", dtype="bf16")
dataset = load_2d_test_dataset(bundle.tokenizer, "data/SyReData", mode="2d_test")
loader = make_2d_test_loader(bundle, torch.utils.data.Subset(dataset, [0]))
output, batch = predict_batch(bundle, next(iter(loader)))

Evaluation on 2d_test

python scripts/evaluate_syre.py \
  --model McGregorW/SyRe \
  --dataset-dir data/SyReData \
  --output-dir outputs/test \
  --device cuda:0 \
  --dtype bf16 \
  --limit 100 \
  --save-predictions \
  --save-visualizations

The evaluator writes:

outputs/test/
├── metrics--2d_test.jsonl            # per-image/per-class metrics
├── summary--2d_test.json             # global_iou, class_iou, Dice, and grouped metrics
├── class_iou_avg--2d_test.json       # mean Dice by class, matching GLaMM naming
├── mod_iou_avg--2d_test.json         # mean Dice by modality and class
├── pred_masks/                       # optional binary predictions
└── visualizations/                   # optional green-GT/red-pred overlays

Use --start-index and --limit to evaluate a contiguous subset of 2d_test. The script intentionally exposes only 2d_test; 3D and other GLaMM evaluation modes are not included.

🧪 Training & validation

The original training entry point is train.py. The supplied launcher uses repository-relative defaults and configurable environment variables, so it can be used from the repository root without editing private machine paths. Download the model and dataset described above, place the SAM ViT-H checkpoint at checkpoints/sam_vit_h_4b8939.pth, and then run:

python train.py --help
bash scripts/finetune_2d_syre.sh

The launcher defaults to checkpoints/SyRe, data/SyReData, and output/syre_2d. Override these locations with SYRE_MODEL, SYRE_DATASET_DIR, SYRE_SAM_CHECKPOINT, SYRE_OUTPUT_DIR, and SYRE_TEXT_PROMPTS_PATH. Training starts from the downloaded model by default; to continue from a local DeepSpeed checkpoint, set SYRE_RESUME to its repository-relative or absolute path. For distributed training, configure GPUS_PER_NODE, NNODES, NODE_RANK, MASTER_ADDR, and MASTER_PORT in the environment.

Important arguments:

Argument Description
--version Base or merged language-model checkpoint.
--vision_pretrained SAM ViT-H checkpoint.
--dataset_dir Dataset root containing indexes and referenced files.
--mode, --mode_val Training and validation index suffixes.
--text_prompts_path Optional CRD description file for training.
--lora_r, --lora_alpha LoRA configuration.
--batch_size Per-process micro-batch size; the reported launcher uses 8.
--grad_accumulation_steps Gradient accumulation steps; the reported launcher uses 10.
--mask_validation Enable segmentation metrics during validation.
--resume DeepSpeed checkpoint directory.

With 40 distributed processes, the reported settings give a global training batch size of 40 × 8 × 10 = 3,200.

Merge a DeepSpeed-exported checkpoint into a Hugging Face directory with:

python scripts/merge_lora_weights.py \
  --version /path/to/base_model \
  --vision_pretrained /path/to/sam_vit_h_4b8939.pth \
  --weight /path/to/zero_to_fp32.bin \
  --save_path /path/to/merged_syre

📚 Citation

If you use SyRe, please cite the SyRe paper:

@misc{wang2026synergisticvisionlanguagereinforcementenables,
      title={Synergistic Vision-Language Reinforcement Enables Scalable On-Demand Analysis across Diverse Clinical Tasks}, 
      author={Haonan Wang and Jiaji Mao and Lehan Wang and Qixiang Zhang and Marawan Elbatel and Yi Qin and Huijun Hu and Baoxun Li and Wenhui Deng and Weifeng Qin and Hongrui Li and Jialin Liang and Jun Shen and Xiaomeng Li},
      year={2026},
      eprint={2505.03380},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2505.03380}, 
}

Please also follow the attribution and usage terms of the source datasets and pretrained components used in your experiments.

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages