SyRe is a text-driven segmentation foundation model for biomedical images. Given an image and a natural-language target such as “liver tumor” or “optic disc”, SyRe predicts a pixel-level mask without training a separate model for every class.
Research release. The model is intended for research and development. It is not a medical device and should not be used as a standalone clinical decision maker.
SyRe connects language understanding with SAM-style dense prediction through a closed-loop vision-language architecture:
SyRe couples vision-language reasoning with SAM-based dense prediction for on-demand biomedical segmentation.
The accompanying study describes SyReData as a large image-mask-text collection spanning nine imaging modalities and 229 segmentation tasks. See the paper and interactive demo for the complete experimental protocol and reported results.
- Public checkpoint: the released SyRe weights are available on Hugging Face.
- Public data: the 2D image/mask collection is available on SyReData.
- Reproducible utilities: this repository now includes native
2d_testinference and batch evaluation scripts. - Interactive demo: try SyRe at 143.89.46.197:7860.
The public model repository is about 16.3 GB and the public dataset repository is about 379 GB. Download only the files or subsets you need.
Quantitative comparison on internal datasets.
Quantitative comparison on external datasets.
Quantitative comparison on in-house datasets.
Representative predictions across CT, MRI, ultrasound, pathology, dermoscopy, X-ray, fundus, endoscopy, and PET images.
Representative qualitative results on in-house segmentation tasks.
Example pathology application showing SyRe predictions across biomedical image patches.
We recommend Python 3.10. Install a PyTorch build matching your CUDA version first, then install the repository dependencies:
conda create -n syre python=3.10 -y
conda activate syre
# Choose the PyTorch/CUDA wheel for your machine.
python -m pip install torch torchvision
python -m pip install -r requirements.txt
python -m pip install huggingface_hubThe command-line help can be used without downloading the model. Full inference is GPU-oriented and memory intensive.
Use the Hub ID directly, or cache the checkpoint locally:
python - <<'PY'
from huggingface_hub import snapshot_download
snapshot_download(
repo_id="McGregorW/SyRe",
local_dir="checkpoints/SyRe",
)
PYBoth of the following forms are accepted by the inference script:
--model McGregorW/SyRe
--model checkpoints/SyRe
The dataset contains train and test splits. Downloading the complete repository is optional; use allow_patterns when you only need selected files:
python - <<'PY'
from huggingface_hub import snapshot_download
snapshot_download(
repo_id="McGregorW/SyReData",
repo_type="dataset",
local_dir="data/SyReData",
# Example: allow_patterns=["*test*", "**/*.png"]
)
PYThe native evaluation path follows GLaMMedv16 and reads image2label_2d_test.json directly. The public dataset directory should therefore contain that index together with its referenced images and masks.
python examples/inference.py \
--model McGregorW/SyRe \
--dataset-dir data/SyReData \
--index 0 \
--output-dir outputs/sample_000000 \
--device cuda:0 \
--dtype bf16 \
--save-visualizationThe command uses the same dataset, prompt construction, image preprocessing, and [SEG] batch format as the GLaMMedv16 2D test evaluator. It writes one binary mask per annotated class and inference.json.
For use from Python, see examples/inference_api.py. The core calls are:
import torch
from scripts.syre_2d_test import load_model, load_2d_test_dataset
from scripts.syre_2d_test import make_2d_test_loader, predict_batch
bundle = load_model("McGregorW/SyRe", device="cuda:0", dtype="bf16")
dataset = load_2d_test_dataset(bundle.tokenizer, "data/SyReData", mode="2d_test")
loader = make_2d_test_loader(bundle, torch.utils.data.Subset(dataset, [0]))
output, batch = predict_batch(bundle, next(iter(loader)))python scripts/evaluate_syre.py \
--model McGregorW/SyRe \
--dataset-dir data/SyReData \
--output-dir outputs/test \
--device cuda:0 \
--dtype bf16 \
--limit 100 \
--save-predictions \
--save-visualizationsThe evaluator writes:
outputs/test/
├── metrics--2d_test.jsonl # per-image/per-class metrics
├── summary--2d_test.json # global_iou, class_iou, Dice, and grouped metrics
├── class_iou_avg--2d_test.json # mean Dice by class, matching GLaMM naming
├── mod_iou_avg--2d_test.json # mean Dice by modality and class
├── pred_masks/ # optional binary predictions
└── visualizations/ # optional green-GT/red-pred overlays
Use --start-index and --limit to evaluate a contiguous subset of 2d_test. The script intentionally exposes only 2d_test; 3D and other GLaMM evaluation modes are not included.
The original training entry point is train.py. The supplied launcher uses repository-relative defaults and configurable environment variables, so it can be used from the repository root without editing private machine paths. Download the model and dataset described above, place the SAM ViT-H checkpoint at checkpoints/sam_vit_h_4b8939.pth, and then run:
python train.py --help
bash scripts/finetune_2d_syre.shThe launcher defaults to checkpoints/SyRe, data/SyReData, and output/syre_2d. Override these locations with SYRE_MODEL, SYRE_DATASET_DIR, SYRE_SAM_CHECKPOINT, SYRE_OUTPUT_DIR, and SYRE_TEXT_PROMPTS_PATH. Training starts from the downloaded model by default; to continue from a local DeepSpeed checkpoint, set SYRE_RESUME to its repository-relative or absolute path. For distributed training, configure GPUS_PER_NODE, NNODES, NODE_RANK, MASTER_ADDR, and MASTER_PORT in the environment.
Important arguments:
| Argument | Description |
|---|---|
--version |
Base or merged language-model checkpoint. |
--vision_pretrained |
SAM ViT-H checkpoint. |
--dataset_dir |
Dataset root containing indexes and referenced files. |
--mode, --mode_val |
Training and validation index suffixes. |
--text_prompts_path |
Optional CRD description file for training. |
--lora_r, --lora_alpha |
LoRA configuration. |
--batch_size |
Per-process micro-batch size; the reported launcher uses 8. |
--grad_accumulation_steps |
Gradient accumulation steps; the reported launcher uses 10. |
--mask_validation |
Enable segmentation metrics during validation. |
--resume |
DeepSpeed checkpoint directory. |
With 40 distributed processes, the reported settings give a global training batch size of 40 × 8 × 10 = 3,200.
Merge a DeepSpeed-exported checkpoint into a Hugging Face directory with:
python scripts/merge_lora_weights.py \
--version /path/to/base_model \
--vision_pretrained /path/to/sam_vit_h_4b8939.pth \
--weight /path/to/zero_to_fp32.bin \
--save_path /path/to/merged_syreIf you use SyRe, please cite the SyRe paper:
@misc{wang2026synergisticvisionlanguagereinforcementenables,
title={Synergistic Vision-Language Reinforcement Enables Scalable On-Demand Analysis across Diverse Clinical Tasks},
author={Haonan Wang and Jiaji Mao and Lehan Wang and Qixiang Zhang and Marawan Elbatel and Yi Qin and Huijun Hu and Baoxun Li and Wenhui Deng and Weifeng Qin and Hongrui Li and Jialin Liang and Jun Shen and Xiaomeng Li},
year={2026},
eprint={2505.03380},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2505.03380},
}Please also follow the attribution and usage terms of the source datasets and pretrained components used in your experiments.






