datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
iclr-wm-backup-public
ICLR Watermark Benchmark — backup overflow (public part)
Companion to the private repo Aak975/iclr-wm-backup, which reached its
storage quota. Together the two repos form ONE backup — every file exists in
exactly one of them, with the same layout:
archives/<sub>/part-0000 ... part-NNNN, MANIFEST.json
restore one archive: cat part-* | zstd -d | tar -x
MANIFEST.json = {"parts": N, "sha256": <whole-stream>, "total_bytes": M}
This public part holds only shareable image data… See the full description on the dataset page: https://huggingface.co/datasets/Aak975/iclr-wm-backup-public.cybergymiclr2026_statsICLR2025Openreview基于 20241208 爬取数据制作。
题目和摘要的中文翻译、关键词由 Qwen2.5-72B-Instruct 自动生成,可能存在漏误,大家酌情参考。
生成数据的代码见 repo
iclr2026_real_reviewsICLR2025reviewICLR2024-papersiclr-rejected-papers-with-code-1k
Rejected ICLR Papers with Reviews and Code
This dataset contains 1,000 rejected ICLR submissions from 2018–2026. Each row
has the OpenReview submission metadata and reviews, the rejected submission PDF,
and a commit-pinned archive of a matched public GitHub repository.
This collection was built directly from OpenReview. It does not use a
third-party ICLR review dataset.
Project repository: TheAppliedScientist
Contents
1,000 unique rejected OpenReview submissions… See the full description on the dataset page: https://huggingface.co/datasets/Vidushee/iclr-rejected-papers-with-code-1k.ICLR-2023-Accepted-Papers
ICLR 2023 International Conference on Learning Representations 2023 Accepted Paper Meta Info Dataset
This dataset is collect from the ICLR 2023 OpenReview website (https://openreview.net/group?id=ICLR.cc/2023/Conference#tab-accept-oral) as well as the arxiv website DeepNLP paper arxiv (http://www.deepnlp.org/content/paper/iclr2023). For researchers who are interested in doing analysis of ICLR 2023 accepted papers and potential trends, you can use the already cleaned up json files.… See the full description on the dataset page: https://huggingface.co/datasets/DeepNLP/ICLR-2023-Accepted-Papers.ICLR-2024-Accepted-Papers
ICLR 2024 International Conference on Learning Representations 2024 Accepted Paper Meta Info Dataset
This dataset is collect from the ICLR 2024 OpenReview website (https://openreview.net/group?id=ICLR.cc/2024/Conference#tab-accept-oral) as well as the arxiv website DeepNLP paper arxiv (http://www.deepnlp.org/content/paper/iclr2024). For researchers who are interested in doing analysis of ICLR 2024 accepted papers and potential trends, you can use the already cleaned up json files.… See the full description on the dataset page: https://huggingface.co/datasets/DeepNLP/ICLR-2024-Accepted-Papers.ICLR-2022-Accepted-Papers
ICLR 2022 International Conference on Learning Representations 2022 Accepted Paper Meta Info Dataset
This dataset is collect from the ICLR 2022 OpenReview website (https://openreview.net/group?id=ICLR.cc/2022/Conference#tab-accept-oral) as well as the arxiv website DeepNLP paper arxiv (http://www.deepnlp.org/content/paper/iclr2022). For researchers who are interested in doing analysis of ICLR 2022 accepted papers and potential trends, you can use the already cleaned up json files.… See the full description on the dataset page: https://huggingface.co/datasets/DeepNLP/ICLR-2022-Accepted-Papers.ICLR-2021-Accepted-Papers
ICLR 2021 International Conference on Learning Representations 2021 Accepted Paper Meta Info Dataset
This dataset is collect from the ICLR 2021 OpenReview website (https://openreview.net/group?id=ICLR.cc/2021/Conference#tab-accept-oral) as well as the arxiv website DeepNLP paper arxiv (http://www.deepnlp.org/content/paper/iclr2021). For researchers who are interested in doing analysis of ICLR 2021 accepted papers and potential trends, you can use the already cleaned up json files.… See the full description on the dataset page: https://huggingface.co/datasets/DeepNLP/ICLR-2021-Accepted-Papers.editlens_iclr_binary_reasoning
bingbangboom/editlens_iclr_binary_reasoning
This dataset is a binary-classification subset drawn from the training split of pangram/editlens_iclr dataset.
It isolates purely human-crafted texts (human_written) against purely synthetic content (ai_generated), strictly filtering out the overlapping ai_edited classification cluster for binary classification tasks.
The primary augmentation of this dataset is the inclusion of Reasoning Traces (Chain of Thought). Every single text… See the full description on the dataset page: https://huggingface.co/datasets/bingbangboom/editlens_iclr_binary_reasoning.benchname-results
BenchName (raw results)
These are the raw results from the BenchName benchmark suite, as well as the corresponding model predictions.
Please use the subset dropdown menu to select the necessary data relating to our six benchmarks:
🤗 Library-based code generation
🤗 CI builds repair
🤗 Project-level code completion
🤗 Commit message generation🤗 Bug localization
🤗 Module summarization
RISE-ICLR-2026
Embedding Geometry Paper Data
This directory contains the paired transformation data used in the embedding geometry research.
Structure
eg_paper_data/
├── ar/ # Arabic
├── en/ # English
├── es/ # Spanish
├── ja/ # Japanese
├── ta/ # Tamil
├── th/ # Thai
└── zu/ # Zulu
Each language directory contains three transformation types:
negation_pairs.jsonl - Neutral sentences paired with their negated versions… See the full description on the dataset page: https://huggingface.co/datasets/mfwta/RISE-ICLR-2026.ICLR26_anonymousiclr_filter50_step2_orgiclr2025_newbenchiclr2025_synthetic_reviews
