datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DEH-image-scan-datahle-public-questionsSWE-bench_Pro
SWE-bench Pro V2
SWE-bench Pro is a benchmark of long-horizon software engineering tasks drawn from real pull requests in 11 open-source
repositories (Go, Python, JavaScript, TypeScript). Each task gives an agent a repository at a base commit plus a PR
description with explicit requirements and interfaces; the agent's patch is graded by hidden fail-to-pass and
pass-to-pass tests in a pristine container.
V2 (2026-09-22) is the default config: 642 tasks (go 256, python 237, js 145… See the full description on the dataset page: https://huggingface.co/datasets/ScaleAI/SWE-bench_Pro.ScaleEdit-12M
ScaleEdit-12M: Scaling Open-Source Image Editing Data Generation via Multi-Agent Framework
📌 Overview
The largest open-source instruction-based image editing dataset to date.
ScaleEdit-12M contains 11.8 million rigorously verified instruction–image pairs spanning 23 task families across diverse real and synthetic visual domains. It was constructed using ScaleEditor, a fully open-source hierarchical multi-agent framework that eliminates… See the full description on the dataset page: https://huggingface.co/datasets/InternVL-U/ScaleEdit-12M.colinear_scaling_models
license: gpl-2.0
Collinear Scaling Models
Checkpoint repository for scaling law experiments comparing collinear (CO) and non-collinear (NC) experimental designs.
Directory Structure
{dataset}/{design}/N_{param_count}/
Dataset: wikipedia, pes2o, cosmopedia, redpajama, c4 (plus _fp16 and _bigtpp variants)
Design: colinear or non_colinear
N: Model parameter count (one of 14 canonical sizes from ~5M to ~70M)
Experimental Designs
Collinear (CO):… See the full description on the dataset page: https://huggingface.co/datasets/leibnitz-lab/colinear_scaling_models.hfh_ci_scan_dataset_bScaleCUA-Data
ScaleCUA: Scaling Open-Source Computer Use Agents with Cross-Platform Data
📑 Paper | 🤗 Dataset | 🤖 Model | 🖥️ Model Demo
Vision-Language Models (VLMs) have enabled computer use agents (CUAs) that operate GUIs autonomously with great potential.
However, developing robust CUAs requires extensive in-domain knowledge about software interfaces and operations.
Unlike image–text pairs that are widely available on the Internet, computer-use data, particularly… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/ScaleCUA-Data.kernelbench-samples
KernelBench Samples
Samples from experiments for KernelBench, described in our arxiv
Learn more about KernelBench from our
Paper
Github Repo
The samples are organized as such
baseline_eval (Section 4 Baseline)
repeated_sampling (Section 5.1.1 Repeated Sampling)
iterative_refinement (Section 5.1.2 Iterative Refinement of Generations)
Within each folder, we organize the results by /level/model/problem_{id}/sample_{id}.
The inner most .json file contains the generated kernel and… See the full description on the dataset page: https://huggingface.co/datasets/ScalingIntelligence/kernelbench-samples.MCP-AtlasMCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers
Leaderboard | MCP Atlas Paper | Github
Dataset Summary
This public release is a subset of 500 sample tasks from the MCP Atlas Benchmark dataset.
MCP Atlas is a large-scale benchmark for evaluating tool-use competency, comprising 36 real MCP servers and 220 tools.
Tasks are designed to assess tool-use competency in realistic, multi-step workflows.
Tasks use natural language prompts that avoid… See the full description on the dataset page: https://huggingface.co/datasets/ScaleAI/MCP-Atlas.hil-benchtiny-scalesThis repo contains the tinyHLE dataset, a list of items to use as a subset of the Humanity's Last Exam benchmark in order to make evaluation more efficient.
The repo contains two files:
tiny_hle.json: a file containing a list of question IDs and weights for three different sample sizes (0.5%, 1.0%, 2.0%)
clean_scales_embedding_hle.parquet: a file containing embeddings representing each item of the HLE benchmark along 16 cognitive scales dimensions, used to create the subsets
Since these are… See the full description on the dataset page: https://huggingface.co/datasets/ambean-tr/tiny-scales.rlvr-reward-hacking-scale-no-conftest-20260909-completion
Matched no-conftest RLVR study 20260909-completion
Lossless research records, grouped by model and trajectory type. Only the listed
configurations have published records. Canary diagnostics are excluded from study
estimates; run status in provenance distinguishes retired diagnostics from active
or completed training. Valid failures, refusals and truncations are retained.
The train split name is a dataset-loader convention; record_type identifies
whether a record is training… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-scale-no-conftest-20260909-completion.paired-llama-3.2-1b-embeddings-lmsys-chat-1m
Paired Llama 3.2 1B Token Embeddings (LMSYS-Chat-1M)
This dataset contains paired activations corresponding to single token locations extracted from Meta's Llama 3.2 1B Instruct on conversations from LMSYS-Chat-1M.
Embeddings are provided for layers 5 through 14, which capture the most interesting intermediate representations.
This dataset was built to study things like:
Learning different basis for activations at a given layer
Studying if there are cases where position encodes… See the full description on the dataset page: https://huggingface.co/datasets/scaleinvariant/paired-llama-3.2-1b-embeddings-lmsys-chat-1m.hle_prompts_07_02_25TLG_libre_scans
TLG libre — fac-similés
Fac-similés d'éditions critiques de textes grecs anciens, rassemblés pour le
projet TLG libre, qui reconstitue en TEI ouvert le corpus du Thesaurus
Linguae Graecae.
Chaque volume archivé ici porte une ou plusieurs œuvres du Canon TLG, et le
rattachement est explicite : voir hf_tlg_manifest.csv à la racine.
Le manifeste : hf_tlg_manifest.csv
C'est la clé d'entrée du dépôt. Une ligne par volume archivé :
Colonne
Contenu
repo_path… See the full description on the dataset page: https://huggingface.co/datasets/Zual/TLG_libre_scans.Scale-SWE
Immersion in the GitHub Universe: Scaling Coding Agents to Mastery
🔥 Highlights
Source from 6M+ pull requests and 23000+ repositories.
Cover 5200 Repositories.
100k high-quality instances.
71k trajectories from DeepSeek v3.2 with 3.5B token.
Strong performance: 64% in SWE-bench-Verified trained from Qwen3-30A3B-Instruct.
📣 News
2026-02-26 🚀 We released a portion of our data on Hugging Face. This release includes 20,000 SWE task… See the full description on the dataset page: https://huggingface.co/datasets/AweAI-Team/Scale-SWE.scannet_mini_val_set_suiteCT-RATE_Generated_Scans
Dataset Card for Synthetic Text-to-CT Scans - VLM3D Challenge
Dataset Details
Dataset Description
This dataset contains 1,000 synthetic 3D chest CT scans generated using the model introduced in
From Alignment to Synthesis: Contrastive Volumetric Grounding for Text-to-CT Generation (Molino et al., BMVC 2026).
The model was trained on the CT-RATE dataset, the largest publicly available collection of paired CT volumes and radiology reports.
It… See the full description on the dataset page: https://huggingface.co/datasets/dmolino/CT-RATE_Generated_Scans.colinear_scaling_models
Collinear/Non-Collinear Scaling Models
Checkpoint repository for scaling law experiments comparing collinear (CO) and non-collinear (NC) experimental designs for the paper Tokens-per-Parameter Coverage Is Critical for Robust LLM Scaling Law Extrapolation under review for NeurIPS 2026.
Code
Anonymized code repository (reproduces all tables): anonymous.4open.science
Directory Structure
{dataset}/{design}/N_{param_count}/
Dataset: wikipedia, pes2o, cosmopedia… See the full description on the dataset page: https://huggingface.co/datasets/TPPIsCriticalFor/colinear_scaling_models.concerto_scannet_compressedscarlet-test-datasecret-scan-remediation-trajectories
Secret Scan Remediation Trajectories
Rights & intended use: legacy public research corpus / portfolio
artifact. Hosted frontier-model outputs are research-only inputs under
project policy (synthetic-factory#161):
intended_use: research_only, project_training_policy: blocked. Not
training data for any model-weight update. Machine-readable record:
rights.json.
Release status: The raw, uncurated payload is now published under
data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/secret-scan-remediation-trajectories.scannetv2
ScanNet Instructions
To acquire the access to ScanNet dataset, Please refer to the ScanNet project page and follow the instructions there. You will get a download-scannet.py script after your request for the ScanNet dataset is approved. Note that only a subset of ScanNet is needed. Once you get download-scannet.py, please use the commands below to download the portion of ScanNet that is necessary for ScanRefer:
python2 download-scannet.py -o data/scannet --type _vh_clean_2.ply… See the full description on the dataset page: https://huggingface.co/datasets/Believe0029/scannetv2.police-scanner-audio
Police Scanner Audio Dataset
A comprehensive collection of police and emergency services radio communications from multiple US cities, captured from publicly available scanner feeds.
Dataset Overview
This dataset contains 103,660 audio recordings totaling 357GB of police scanner audio from 6 different cities across the United States. The recordings span multiple months of continuous monitoring and represent real-world emergency services communications.
Scanner… See the full description on the dataset page: https://huggingface.co/datasets/trentmkelly/police-scanner-audio.NanoMTEB-Scandinavian
NanoMTEB-Scandinavian
This dataset is a Nano-style retrieval dataset for HAKARI-bench.
NanoMTEB-Scandinavian is a compact retrieval benchmark for Scandinavian-language MTEB-style task families. It includes Danish, Norwegian, and Swedish retrieval tasks spanning fact verification, question answering, news, encyclopedic content, FAQ retrieval, and social-media retrieval.
Usage
from datasets import load_dataset
dataset_id = "hakari-bench/NanoMTEB-Scandinavian"
split… See the full description on the dataset page: https://huggingface.co/datasets/hakari-bench/NanoMTEB-Scandinavian.TSFM-ScalingLaws-Dataset
TSFM-ScalingLaws-Dataset
This is the dataset for the paper Towards Neural Scaling Laws for Time Series Foundation Models.
Code: https://github.com/Qingrenn/TSFM-ScalingLaws
Well-trained models: https://huggingface.co/PeacefulData/TSFM-ScalingLaws-Checkpoints
Dataset Summary
Domain
Transport
Climate
Energy
Cloud
Health
Sales
Web
Total
Datasets
8
2
14
3
9
1
2
39
Time Points
4.82B
4.73B4.76B
2.15B
232M
140M
40M
16.8B
Proportion
28.52%
28.06%
28.21%
12.76%… See the full description on the dataset page: https://huggingface.co/datasets/Qingren/TSFM-ScalingLaws-Dataset.Scaffold-CoT
Scaffold-CoT
Structured chain-of-thought training data with 3,726,548 examples in 76 JSONL shards.
Fields
Every row has exactly four top-level fields:
Field
Contents
metadata
domain, subdomain, difficulty, length_bucket
input
Ordered user messages as {index, content} objects
cot
Ordered {index, type, content} events, including reasoning, tool calls, and tool results
output
Ordered final assistant answers as {index, content} objects
The index… See the full description on the dataset page: https://huggingface.co/datasets/Specific-Labs/Scaffold-CoT.deepcad_test_scanscannet200_50_2d_maskScale-SWE-Verified
Scale-SWE-Verified
Gold-patch-validated fork of
AweAI-Team/Scale-SWE
(paper): 17,202 / 20,181 Python issue-resolving tasks
that produce a clean reward signal end-to-end. Default dataset of the scaleswe_v1 taskset.
Changes vs upstream
Validation (ours) removed 2,979 / 20,181 rows (14.8%):
892 rows whose image_url appears in
scale-swe-exclude-images.json.
2,061 rows categorized gold_patch_failure in
scale-swe-validation.jsonl.
15 rows categorized noop_pass… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/Scale-SWE-Verified.
