Skip to content

About

MICCAI 2026 Early-Accept Spotlight

Resources

Stars

4 stars

Watchers

0 watching

Forks

Latest commit

 

History

6 Commits

Folders and files

Repository files navigation

A Multi-Center Benchmark for Abdominal Disease Diagnosis and Report Generation from Non-Contrast CT

Mariam Elbakry1, Aliaa Sheha1, Salma Tantawy1, Aya Yassin1, Concetto Spampinato3, Karim Lekadir4, Xiaomeng Li2, Marawan Elbatel2

1Ain Shams University    2HKUST    3University of Catania    4Universitat de Barcelona

MICCAI 2026 Spotlight

arXiv Dataset License: CC BY-NC 4.0

This is the official implementation of TriALS-Report, a multi-center benchmark for abdominal disease diagnosis and radiology report generation from non-contrast CT (NCCT).

Related publications

Benchmarking Foundation Models for Cervical Cancer CT Reporting in Zambia. Kangwa E. Mukuka, Lena Lambart, Festus Mwape, Lighton Phiri, Concetto Spampinato, Karim Lekadir, Marawan Elbatel. MICCAI 2026 Workshop AFRICAI. OpenReview

Abstract

Multiphasic contrast-enhanced CT (CECT) is widely used for abdominal lesion characterization, yet it carries inherent risks of contrast-induced nephropathy, escalates acquisition burden, and heavily contributes to radiologist workload. To address these challenges, we introduce a novel multi-center benchmark for multi-organ abdominal disease diagnosis and automated radiology report generation, which learns to synthesize contrast-enhanced findings from single-phase non-contrast CT (NCCT). To support this, we curated a large-scale dataset of paired NCCT–CECT studies and their corresponding contrast-enhanced radiology reports from two centers, partitioned into internal sets and an external validation cohort. Under a unified evaluation protocol, we benchmarked five contemporary deep learning architectures encompassing chest-specific, abdomen-specific, and general-purpose multimodal domains. Extensive experiments demonstrate that NCCT retains diagnostic signals, achieving an average multi-organ AUC of 71.1% on the internal cohort and 66.2% on the external cohort, respectively. By releasing this dataset and standardized benchmark publicly, this study aims to catalyze future research into safer, resource-efficient, and globally accessible contrast-free abdominal imaging workflows.

Study workflow: multi-center NCCT collection, label extraction from triphasic reports, model development and evaluation

Study workflow. Non-contrast CT volumes are paired with the triphasic contrast-enhanced report of the same patient; findings are extracted from the report to form the label space, and models are evaluated on disease diagnosis and report generation.

Results

Carcinomas and metastases (internal and external test pooled, n=388)

🔥 Zero-shot DAMO RADAR outperforms linear probes on Pillar-0 and Merlin for carcinoma and liver metastasis detection on TriALS-Report, without any training on this data.

Model Hepatocellular carcinoma Pancreatic tumor Colonic carcinoma Hepatic metastases Metastatic disease (n=92)
Pillar-0 83.23 [76.1, 89.3] 82.46 [69.7, 92.8] 56.51 [41.4, 70.9] 75.62 [68.4, 82.5] 71.06 [64.9, 76.4]
Merlin 81.74 [73.1, 88.9] 66.10 [49.1, 83.4] 49.26 [34.0, 63.4] 69.72 [61.2, 77.3] 71.63 [65.1, 77.8]
DAMO RADAR (zero-shot) 92.42 [88.3, 95.7] 91.26 [78.2, 98.8] 67.61 [54.8, 80.5] 87.12 [82.3, 91.1] 68.56 [62.5, 74.2]

Organ-level results (paper)

Non-contrast CT, frozen encoder + linear probe. AUC in %, 95% bootstrap CI in brackets.

Internal test (Center 1, n=219)

Model Liver (10 diseases) Pancreas (3 diseases) Average (15 organs, 51 diseases)
Pillar-0 68.24 [65.2, 71.5] 80.39 [71.4, 87.9] 69.92 [67.4, 72.2]
Merlin 69.48 [66.5, 72.4] 76.00 [65.1, 85.6] 71.05 [68.8, 73.3]

External test (Center 2, n=169)

Model Liver (10 diseases) Pancreas (3 diseases) Average (15 organs, 51 diseases)
Pillar-0 66.48 [62.4, 70.8] 78.55 [64.8, 92.1] 65.71 [63.2, 68.4]
Merlin 68.15 [64.3, 72.1] 61.32 [52.5, 70.1] 66.24 [63.7, 68.8]

Liver: cirrhosis, congenital liver cysts, hepatic changes post-resection, hepatic cysts, hepatic hemangiomas, hepatic masses, hepatic metastases, hepatic steatosis, hepatocellular carcinoma, hepatomegaly. Pancreas: main pancreatic duct dilatation, pancreatic atrophy, pancreatic tumors.

The taxonomy has 53 diseases over 16 organs; the average covers the 15 organs excluding Multi-organ (metastatic disease, inguinal hernias), so 51 diseases. Chance level is 50.00 AUC. F1 and all 16 organs are written to results/summary_organs.csv and results/<model>/<split>/seed<k>/.

The encoders, Pillar-0 and Merlin, are frozen and only the probe is trained, so no checkpoints are released.

Dataset

Each case is a non-contrast CT volume paired with the triphasic radiology report, from which 232 findings were extracted. Volumes are 512 x 512, mean in-plane spacing 0.875 x 0.875 mm, mean slice thickness 1.07 mm.

Split Center 1 (internal) Center 2 (external) Total
Train 760 – 760
Val 106 – 106
Test 219 169 388
Patients / Volumes 1,085 169 1,254
huggingface-cli download marwankefah/TriALS-Report --repo-type dataset --local-dir ./TriALS-Report
TriALS-Report/
├── label_dictionary.csv          column, organ, question
├── Center 1/
│   ├── labels.csv                patient_id, center, split, image, 232 finding columns
│   └── ct/<patient_id>.nii.gz
└── Center 2/
    ├── labels.csv
    └── ct/<patient_id>.nii.gz

See the dataset card for the label and split conventions.

Installation

git clone https://github.com/xmed-lab/TriALS-Report
cd TriALS-Report
conda create -n trials-report python=3.11 -y
conda activate trials-report
pip install -r requirements.txt

Disease diagnosis

1. Extract features

Everything below is built from the downloaded dataset. Preprocessing uses rad-vision-engine and extraction uses RATE-Evals; the four lists are center1_train, center1_val, center1_test and center2_test.

# list the volumes per center and split
python prepare.py series --data ./TriALS-Report --work ./work

# preprocess each list
for L in center1_train center1_val center1_test center2_test; do
  vision-engine process --config rad-vision-engine/configs/ct_abdomen.yaml \
      --input-series-csv work/series/$L.csv --output work/cache/$L --workers 16
done

# inputs for rate-extract
python prepare.py rate-inputs --data ./TriALS-Report --work ./work

# embeddings, per list (Merlin: --model merlin --model-repo-id stanfordmimi/Merlin)
for L in center1_train center1_val center1_test center2_test; do
  rate-extract --model pillar0 --dataset abd_ct_merlin --split train --batch-size 1 \
      --model-repo-id YalaLab/Pillar0-AbdomenCT --output-dir work/emb/pillar0/$L \
      data.train_json=work/rate/$L.jsonl data.cache_manifest=work/rate/manifest_$L.csv
done

# collect them next to the labels
python prepare.py features --data ./TriALS-Report --work ./work --model pillar0

This writes TriALS-Report/features/pillar0/{center1,center2}.parquet, which run_benchmark.py reads.

2. Run the probe

python run_benchmark.py --data ./TriALS-Report --out results

The probe trains on Center 1 train, picks the epoch by mean AUC on Center 1 val, and is evaluated on Center 1 test (internal) and Center 2 (external). Patient counts are printed and test patients of the same center are checked against train.

results/<model>/<internal|external>/seed<k>/
    predictions.csv          accession, question, probability, prediction, true_label
    bootstrap.json           per-organ and average AUC/F1 with 95% CIs
    threshold_per_question.csv, threshold_overall.csv
results/summary.csv          average AUC/F1 per model, mean and SD over seeds

Options: --models pillar0, --seeds 0 1 2, --n-boot 200, --device cuda, --no-val (keep the last epoch, as in the paper).

One nn.Linear(D, 232) is trained on all findings jointly, BCEWithLogitsLoss with pos_weight = negatives/positives per finding, Adam (lr 1e-3), up to 1000 epochs, batch 8192, threshold 0.5. Organ-level AUC/F1 average the per-disease scores within each organ, over 1,000 patient-level bootstrap resamples. The paper's runs trained a fixed 1000 epochs without validation and were unseeded; this code seeds the probe and selects the epoch on val, so values can differ slightly.

3. Zero-shot DAMO RADAR

Set up DAMO RADAR and download its checkpoints:

git clone https://github.com/alibaba-damo-academy/damo-radar
cd damo-radar
conda create -n radar python=3.10 -y && conda activate radar
pip install -r requirements.txt
git apply ../patches/radar_inference_memory.patch   # lower GPU memory on full-size scans, same outputs
cd download_scripts && python download_checkpoints.py && cd ../..

Run it on the 388 test volumes (a GPU is required), then score the output:

python prepare.py radar-inputs --data ./TriALS-Report --work ./work

cd damo-radar/RADAR_inference
python inference_demo.py --img_dir ../../work/radar_inputs --save_dir ../results --save_tag trials_report_test
cd ../..

python score_radar.py --data ./TriALS-Report \
    --radar damo-radar/results/RADAR_infer_results_trials_report_test.csv

radar-inputs links the test volumes as Center1__<id>.nii.gz and Center2__<id>.nii.gz, the names RADAR writes to its output, and score_radar.py prints the table above.

Acknowledgment

Built on RATE-Evals and rad-vision-engine. The evaluated encoders are Pillar-0 and Merlin, and the zero-shot comparison uses DAMO RADAR.

License

Released under CC BY-NC 4.0. The frozen encoders keep their original licenses.

Citation

@misc{elbakry2026multicenterbenchmarkabdominaldisease,
      title={A Multi-Center Benchmark for Abdominal Disease Diagnosis and Report Generation from Non-Contrast CT}, 
      author={Mariam Elbakry and Aliaa Sayed Sheha and Salma Hassan Tantawy and Aya Yassin and Concetto Spampinato and Karim Lekadir and Xiaomeng Li and Marawan Elbatel},
      year={2026},
      eprint={2606.16991},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2606.16991}, 
}

If you use the features or the evaluation protocol, please also cite RATE-Evals / Pillar-0, Merlin and rad-vision-engine.

About

MICCAI 2026 Early-Accept Spotlight

Resources

Stars

4 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages