Hi — we (bioc-on-ice, https://github.com/seandavi/bioc-on-ice) are mirroring BEDbase's metadata (not files) into a public Apache Iceberg catalog for Bioconductor users, keyed on genome_digest, so BED files become findable next to gene models, ontologies and expression datasets. Doing that against the current API surfaced a few things you should know, and one request.
What we observed (2026-09-18, all times UTC)
/v1/bed/list cannot be enumerated by offset. With limit=100 a page takes ~40 s around offset=67000, and past ~70,000 the gateway returns 504 (60 s). limit caps at 10,000 (422 above). hg38 alone is 546,706 of the 663,721 files, so it is unreachable this way.
- Ordering is unstable under paging. Paging
/v1/bedset/list at 100 per page yielded 222 pages but only 14,485 unique bedset ids of 22,189 (duplicates and gaps); limit=25000&offset=0 returned all 22,189 in one 5 s response. The same drift affects /bed/list.
- We likely caused an outage — sorry. Around 12:00–13:30 UTC our crawl (first a multi-threaded reader, then 2–3 sequential workers) coincided with every endpoint returning HTTP 500 after 30 s (
/v1/stats, single-record /metadata, listings at offset 0) while bedbase.org stayed up, for about 95 minutes until ~15:34. We stopped as soon as we saw it and have since used ≤2 workers with delays. If your logs point at us, please say so and we will adjust further.
- What worked well:
/v1/bedset/{id}/bedfiles returns a bedset's files in one response (35,707 records / 30 MB for encode_files_chunk_0000, 23 s), and genome=<label> on /bed/list lets every genome except hg38 page completely. Union coverage: 537,557 of 663,721 files (81%); the rest are hg38 files in no bedset.
Request
Either of these would make a full, gentle mirror possible and remove the load from your API:
- A periodic bulk metadata export — one gzipped JSONL or Parquet file of the
/bed/list records (and one for bedsets), e.g. on S3 next to the BED files, versioned by date. This is how NCBI, iCite, Ensembl and CELLxGENE are consumed downstream, and one file per month is far cheaper for you than thousands of listing requests.
- Failing that, keyset (cursor) pagination on
/bed/list — e.g. ?after=<id> with a stable sort — which avoids both the deep-offset cost and the drift.
Happy to test either, or to share the crawl notes if useful. Thanks for BEDbase — the per-file genome_digest and the DUO licence codes made this straightforward to model.
Hi — we (bioc-on-ice, https://github.com/seandavi/bioc-on-ice) are mirroring BEDbase's metadata (not files) into a public Apache Iceberg catalog for Bioconductor users, keyed on
genome_digest, so BED files become findable next to gene models, ontologies and expression datasets. Doing that against the current API surfaced a few things you should know, and one request.What we observed (2026-09-18, all times UTC)
/v1/bed/listcannot be enumerated by offset. Withlimit=100a page takes ~40 s aroundoffset=67000, and past ~70,000 the gateway returns 504 (60 s).limitcaps at 10,000 (422 above). hg38 alone is 546,706 of the 663,721 files, so it is unreachable this way./v1/bedset/listat 100 per page yielded 222 pages but only 14,485 unique bedset ids of 22,189 (duplicates and gaps);limit=25000&offset=0returned all 22,189 in one 5 s response. The same drift affects/bed/list./v1/stats, single-record/metadata, listings at offset 0) while bedbase.org stayed up, for about 95 minutes until ~15:34. We stopped as soon as we saw it and have since used ≤2 workers with delays. If your logs point at us, please say so and we will adjust further./v1/bedset/{id}/bedfilesreturns a bedset's files in one response (35,707 records / 30 MB forencode_files_chunk_0000, 23 s), andgenome=<label>on/bed/listlets every genome except hg38 page completely. Union coverage: 537,557 of 663,721 files (81%); the rest are hg38 files in no bedset.Request
Either of these would make a full, gentle mirror possible and remove the load from your API:
/bed/listrecords (and one for bedsets), e.g. on S3 next to the BED files, versioned by date. This is how NCBI, iCite, Ensembl and CELLxGENE are consumed downstream, and one file per month is far cheaper for you than thousands of listing requests./bed/list— e.g.?after=<id>with a stable sort — which avoids both the deep-offset cost and the drift.Happy to test either, or to share the crawl notes if useful. Thanks for BEDbase — the per-file
genome_digestand the DUO licence codes made this straightforward to model.