datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
scaling-data-constrained-llms
Scaling Data-Constrained Language Models with Synthetic Data
This repository provides the pre-training corpora used in Scaling Data-Constrained Language Models with Synthetic Data (Findings of EACL 2026).
Overview
This repository contains multiple corpora designed to study data augmentation strategies for pre-training Japanese LLMs under a data-constrained data setting.
Starting from a limited Japanese Web corpus and a larger English Web corpus, we construct three… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/scaling-data-constrained-llms.inverse-scaling-ttc-main
Inverse Scaling in Test-Time Compute
Paper: Inverse Scaling in Test-Time Compute
Project Page: https://safety-research.github.io/inverse-scaling-ttc/
Abstract
We construct evaluation tasks where extending the reasoning length of Large Reasoning Models (LRMs) deteriorates performance, exhibiting an inverse scaling relationship between test-time compute and accuracy. Our evaluation tasks span four categories: simple counting tasks with distractors, regression tasks with… See the full description on the dataset page: https://huggingface.co/datasets/inverse-scaling-ttc/inverse-scaling-ttc-main.fusion-pairwise-evals-test-time-scaling
Automatic pairwise preference evaluations for: Making, not taking, the Best-of-N
Content
This data contains pairwise automatic win-rate evaluations for the m-ArenaHard-v2.0 benchmark and it compares CommandA against gemini-2.5-pro in 2 settings:
Test-time scaling with Fusion : 5 samples are generated from CommandA, then fused with CommandA into one completion and compared to a single completion from gemini-2.5-pro
Test-time scaling with BoN : 5 samples are generated from… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/fusion-pairwise-evals-test-time-scaling.DensingLaw-ScalingBench
DensingLaw-ScalingBench
This dataset was created to enable a more accurate performance scaling law estimation of Large Language Models (LLMs).
This dataset is released as part of our paper, Densing Law of LLMs.
📜 Paper
💡 Overview
This repository contains the open-source dataset used for calculating conditional loss in our LLM density evaluation framework.
LLM density is defined as the ratio of effective parameter size to actual parameter size, where effective… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/DensingLaw-ScalingBench.s1-test-time-scaling-synth-public
s1-test-time-scaling-synth: Japanese and English Reinforcement Learning Dataset Derived from the s1 Simple Test-Time Scaling Dataset
This repository contains s1-test-time-scaling-synth, a reinforcement learning dataset in Japanese and English.This dataset is built upon the supervised fine-tuning dataset simplescaling/data_ablation_full59K (hereafter, the "original dataset"), originally developed in "s1: Simple test-time scaling" [Muennighoff+, EMNLP25].
The original dataset is a… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/s1-test-time-scaling-synth-public.scaling-tasks-v1
scaling-tasks-v1
35 quantitative scaling tasks for the
scaling-env RL
environment, on the topics of Google DeepMind's free book How To Scale Your
Model: rooflines and arithmetic intensity,
transformer FLOPs and parameter counts, KV-cache and optimizer-state memory, collective
communication volume, and decode throughput.
field
meaning
task_id
sc-000 … sc-034
category
roofline / flops / memory / sharding / inference
prompt
the question, every parameter it needs, and… See the full description on the dataset page: https://huggingface.co/datasets/eltociear/scaling-tasks-v1.
