iclr
Datasets
All datasets matching “iclr”iclr-wm-backup-public
ICLR Watermark Benchmark — backup overflow (public part)
Companion to the private repo Aak975/iclr-wm-backup, which reached its
storage quota. Together the two repos form ONE backup — every file exists in
exactly one of them, with the same layout:
archives/<sub>/part-0000 ... part-NNNN, MANIFEST.json
restore one archive: cat part-* | zstd -d | tar -x
MANIFEST.json = {"parts": N, "sha256": <whole-stream>, "total_bytes": M}
This public part holds only shareable image data… See the full description on the dataset page: https://huggingface.co/datasets/Aak975/iclr-wm-backup-public.cybergymICLR_2025_OCR
ICLR_2025_OCR
OCR Data.
scref_ICLR_2025
scREF
This dataset contains human single cell RNA-sequencing (scRNA-seq) data collected from 46 studies and standardized
by Diaz-Mejia JJ et al. (2025) for the paper Benchmarking and optimizing organism wide single-cell RNA alignment methods presented at the LMRL Workshop at the International Conference on Learning Representations (2025).
Folder Phenomic-AI/scref_ICLR_2025/zarr contains standardized single-cell RNA data for each study in zarr format.
Sub-folder names show: {first… See the full description on the dataset page: https://huggingface.co/datasets/Phenomic-AI/scref_ICLR_2025.ai-research-index-iclr-openreview
Conference-paper corpus (AI-Research-Index project)
Private working dataset. Conference/journal paper corpus across OpenReview
venues (ICLR, NeurIPS incl. D&B/position tracks, ICML incl. position, COLM,
TMLR, AISTATS, UAI, ALT, MathAI, and later additions) plus the ACL Anthology
family. Coverage, per-venue availability, decisions semantics, and known
biases are documented authoritatively in the GitHub repo's data/README.md
— read that first; per-venue counts change as the corpus… See the full description on the dataset page: https://huggingface.co/datasets/latkes/ai-research-index-iclr-openreview.cybergym-server
