datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rebuttal6tasksafe
VLABench Rebuttal SAFE
This LeRobot v2.1 dataset contains 600 successful SAFE demonstrations from 12
controlled branches, totaling 94,871 frames at 10 fps. Each episode stores two
480x480 RGB camera streams, a 7D end-effector state, a 7D action, and a natural
language task instruction.
The image payloads are embedded in the episode parquet files. Redundant
conversion-time PNG files are intentionally omitted from this repository.
Comments_200K
Comments-200k 📚
1. Introduction
Raw reviews often intermix critical points with extraneous content, such as salutations and summaries. Directly feeding this unprocessed text into a model introduces significant noise and redundancy, which can compromise the precision of the generated rebuttal. Furthermore, due to diverse reviewer writing styles and varying conference formats, comments are typically presented in an unstructured manner. Therefore, to address these… See the full description on the dataset page: https://huggingface.co/datasets/RebuttalAgent/Comments_200K.rebuttal_dataiclr-rebuttal-analysis
ICLR Rebuttal Dynamics Dataset
Companion dataset for the paper "Rebuttal dynamics in machine-learning peer review: an LLM-instrumented analysis of ICLR 2024-2025".
Overview
This dataset contains 164,974 peer reviews from ICLR 2023-2026, enriched with LLM-extracted features for studying how author rebuttals influence reviewer score changes. Each review includes the full review text, pre- and post-rebuttal scores, Gemini-predicted scores, score-change verdicts, and… See the full description on the dataset page: https://huggingface.co/datasets/MlouisBE/iclr-rebuttal-analysis.generations-olmo-3-7b-rmu-bm25-6t-rebuttalgenerations-llama-3_1-8b-rmu-bm25-10b-rebuttalgenerations-llama-3_1-8b-undial-bm25-6t-rebuttalgenerations-olmo-3-7b-undial-bm25-6t-rebuttalgenerations-olmo-3-7b-rmu-bm25-10b-rebuttalgenerations-olmo-3-7b-undial-igm-10b-rebuttallua-coco100-rebuttal-2k
LUA COCO-100 Rebuttal 2K Generations
This dataset contains 2048x2048 generations for a deterministic 100-example
MS COCO val2017 subset using official COCO captions. It was prepared for the
LUA rebuttal domain-validation task.
Contents
prompts.csv: selected prompts. The gpt_caption column is JSON and uses
the same sdxl field convention as the previous competitor protocol.
manifest.json: selected COCO image ids and metadata.
references/: COCO reference images saved as… See the full description on the dataset page: https://huggingface.co/datasets/vaskers5/lua-coco100-rebuttal-2k.generations-qwen3-8b-rmu-bm25-6t-rebuttalgenerations-qwen3-8b-undial-bm25-6t-rebuttalgenerations-qwen3-8b-undial-igm-10b-rebuttalgenerations-llama-3_1-8b-rmu-bm25-6t-rebuttalgenerations-olmo-3-7b-rmu-igm-10b-rebuttalgenerations-qwen3-8b-rmu-igm-10b-rebuttalgenerations-llama-3_1-8b-rmu-igm-10b-rebuttalgenerations-qwen3-8b-undial-bm25-10b-rebuttalgenerations-qwen3-8b-rmu-bm25-10b-rebuttalgenerations-olmo-3-7b-undial-bm25-10b-rebuttalgenesis-hr-bench-rebuttal-full19-hrv2-eval-20260812
Genesis HR Bench — preliminary HR v2 train19 rebuttal evaluation
This is an immutable, sanitized snapshot of the currently available artifacts from:
hf19_eval5_act080000_full95_fast_20260812T0236Z
It contains per-episode videos, state traces, result JSON, validation records, task/model manifests, policy-server logs, and Slurm logs for five policies across the 19 HR v2 tasks used for this rebuttal evaluation. It also includes both focused seed-7 skillet diagnostic pilots.… See the full description on the dataset page: https://huggingface.co/datasets/zimplex/genesis-hr-bench-rebuttal-full19-hrv2-eval-20260812.generations-llama-3_1-8b-undial-bm25-10b-rebuttalgenerations-llama-3_1-8b-undial-igm-10b-rebuttaluiq_rebuttalPATCH-rebuttal-20260805N_rebuttallig-rebuttal-data
LIG Rebuttal Data
Anonymous data release for COLM 2026 paper: The Latent Intelligence Gap
Tiers
Tier
Contents
Size
Sufficient for
1 Results
Aggregate JSONs
~5 MB
Verify every number in rebuttal
2 Matched
Per-problem trajectories (K=32)
~700 MB
Reproduce all baselines
3 Embeddings
Last-layer hidden states
~5 GB
Reproduce IWC, verifier
Models
Model
Parameters
Benchmarks
Qwen2.5-7B-Instruct
7B
GSM8K, GPQA, AIME24, AIME25… See the full description on the dataset page: https://huggingface.co/datasets/Anonymousblind/lig-rebuttal-data.cora_wodiv_rebuttaltrajectory-analysis
Trajectory Analysis with Sequential Dependencies
This dataset contains step-level classifications for reasoning trajectories
from nine models on two datasets. The original six-model collection uses
Trial Step, Subtask Step, and Sequential Step. Three additional GPT
run-1 collections use the two-class Trial Step/Subtask Step scheme.
Status: This dataset is actively being updated. Its coverage,
classifications, and metadata may evolve as additional trajectories are
included and… See the full description on the dataset page: https://huggingface.co/datasets/parason-rebuttal/trajectory-analysis.
