datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
the-stack-smol-xl
Dataset Description
A small subset of the-stack dataset, with 87 programming languages, each has 10,000 random samples from the original dataset.
Languages
The dataset contains 87 programming languages:
'ada', 'agda', 'alloy', 'antlr', 'applescript', 'assembly', 'augeas', 'awk', 'batchfile', 'bison', 'bluespec', 'c',
'c++', 'c-sharp', 'clojure', 'cmake', 'coffeescript', 'common-lisp', 'css', 'cuda', 'dart', 'dockerfile', 'elixir',
'elm', 'emacs-lisp','erlang'… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-smol-xl.the-stack-smol-xs\CUA-Gym
CUA-Gym
CUA-Gym is a collection of verifiable computer-use agent tasks for reinforcement learning with verifiable rewards (RLVR). Each task pairs a natural-language instruction with executable setup artifacts and a Python reward function that checks task completion programmatically. For details, see the paper CUA-Gym: Scaling Verifiable Training Environments and Tasks for Computer-Use Agents.
This release contains the full public CUA-Gym task set after the necessary data review.… See the full description on the dataset page: https://huggingface.co/datasets/xlangai/CUA-Gym.sim-posttrain
HUMANUAL Posttraining Data
Posttraining data for user simulation, derived from the train splits of the
HUMANUAL benchmark datasets.
Datasets
HUMANUAL (posttraining)
Config
Rows
Description
news
48,618
News article comment responses
politics
45,429
Political discussion responses
opinion
37,791
Reddit AITA / opinion thread responses
book
34,170
Book review responses
chat
23,141
Casual chat responses
email
6,377
Email reply responses… See the full description on the dataset page: https://huggingface.co/datasets/Xuhui/sim-posttrain.ICPC_Data
ICPC World Finals — a discriminative subset, with model traces
24 ICPC World Finals problems (2021–2025), together with the full transcripts of an
LLM attempting each of them three times under simulated contest rules.
Selection
The model
Every run in this dataset comes from:
nvidia/Nemotron-Cascade-2-30B-A3B
The partitions
Every one of the 53 problems was run 3 times (seeds 1, 2, 3). Each problem was then
placed by its pass rate and… See the full description on the dataset page: https://huggingface.co/datasets/xupy21/ICPC_Data.DSULT-Core-ShareGPT-X
DSULT-Core/ShareGPT-X Filtered Dataset
This dataset is a curated subset of ShareGPT-X, which contains approximately 92,000 one-to-one conversations between humans and ChatGPT, collected from X.com (formerly Twitter). The corpus covers content from January 2024 through May 2025, built entirely from public "share" links posted by users on their timelines.
The file ChatGPT-Simple_ShareGPT_Full.json includes the longest sequences of alternating human and gpt messages within each… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/DSULT-Core-ShareGPT-X.xRouteBench
xRouteBench — LLM Routing Benchmark
xRouteBench is a benchmark for training and evaluating LLM routers — systems that pick the best LLM from a candidate pool for each incoming query, trading off performance vs. price cost.
Every query in each scenario was executed against all 18 candidate LLMs, recording each model's response, task performance, token usage, and latency. A router learns from the train split which model to pick, and is evaluated on test.
Scenarios… See the full description on the dataset page: https://huggingface.co/datasets/ulab-ai/xRouteBench.gspc-xr
GSPC — cross reality bank (XRAIV)
Council of AI measurement bank. Measurement, not certification.
Bank. Frozen split. Live n is the matching axis on GET https://councilof.ai/api/gspc, not a Hub score. Not a certificate. Art 50 (EUR-Lex): 2 August 2026 live; marking grace 2 December 2026.
Live measurement. This bank stands behind the cross-reality row of the live GSPC board: GET https://councilof.ai/api/gspc?axis=cross-reality (family, kind, status and n are on that row, never… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-xr.readall-sft-stage-a-b
ReadAll / ReadTwice SFT:Stage A + Stage B
当前 ReadAll 模型的 SFT 数据和可移植训练包。训练链为:
Qwen/Qwen2.5-7B-Instruct → Stage A (ReadTwice step286) → Stage B (ReadAll union step448)
阶段
训练行数
验证行数
Parquet 分片
全局 batch
学习率
1 epoch 更新数
A
73,416
1,676
15 + 1
256
1e-5
286
B
57,287
无独立验证集
29
128
5e-7
448
这些计数是 SFT 消息样本行数,包含 SKIM / UPDATE / FINAL,并非独立问题数。45 个 Parquet 共 1,030,251,959 字节。原始数据分片和 manifest 原样保留,逐一核对原始 SHA-256;没有删列、重新筛选或重新生成。
Stage A 使用普通多轮 assistant-token SFT、12,288 token… See the full description on the dataset page: https://huggingface.co/datasets/Xirui1208/readall-sft-stage-a-b.SWE-Review-Traj
SWE-Review-Traj
8,914 decision-correct agentic code review trajectories for training open-source code review models. Each trajectory captures a full multi-turn review session where an AI agent explores a repository, traces the root cause of a bug, and produces a structured review decision — all verified against executable test suites.
Project Page | Paper | Code | Benchmark | Training Data | Claude Code Plugin
About SWE-Review
SWE-Review is a framework for… See the full description on the dataset page: https://huggingface.co/datasets/Lego-X/SWE-Review-Traj.the-stack-v2-train-xsmol-content
The Stack v2 — 12-Language Resolved Content
This dataset provides resolved file content for twelve programming languages,
derived from the repository/file identifiers published in
bigcode/the-stack-v2-train-full-ids.
The upstream dataset ships identifiers only — each file is a pointer into the
Software Heritage archive. Here, those
identifiers have been resolved to their actual source text so the content is
directly usable, with the upstream metadata carried through unchanged.… See the full description on the dataset page: https://huggingface.co/datasets/handwoven8588/the-stack-v2-train-xsmol-content.Feng-Chat简体中文 | English
Feng-Chat
“解答世间万物”
Feng-Chat 是一批中文 conversational SFT 数据。单轮问答好做,多轮里的追问、接话、拐弯和突然补一句,才是这批数据真正想留下来的东西。QA 抽取后依次做去重、assistant-only PPL、message-level Guard 和主题分类,最后发布为标准 messages。
整个 pipeline 只负责筛选,不改写峰哥的表达。公开 release 仅保留训练和分析需要的字段,不包含内部 prompt、reasoning、raw response、本地路径或原始来源 ID。
数据概览
指标
最终结果
发布记录
33,468
QA 轮次
40,645
多轮记录
3,952(11.81%)
单条最多轮次
30
伪名化来源分组
729
内容日期范围
2022-01-02 ~ 2026-08-20
assistant loss tokens
4,806,476… See the full description on the dataset page: https://huggingface.co/datasets/xfalcon9/Feng-Chat.terminal-bench
Terminal-Bench Dataset
This dataset contains tasks from Terminal-Bench, a benchmark for evaluating AI agents in real terminal environments. Each task is packaged as a complete, self-contained archive that preserves the exact directory structure, binary files, Docker configurations, and test scripts needed for faithful reproduction.
The archive column contains a gzipped tarball of the entire task directory.
Dataset Overview
Terminal-Bench evaluates AI agents on… See the full description on the dataset page: https://huggingface.co/datasets/XUO/terminal-bench.JailbreakGuardrailBenchmark
An Open Benchmark for Evaluating Jailbreak Guardrails in Large Language Models
Introduction
This repository provides instruction datasets in our SoK paper, SoK: Evaluating Jailbreak Guardrails for Large Language Models. The datasets are collected from various sources to evaluate the effectiveness of jailbreak guardrails in large language models (LLMs), including harmful prompts (i.e., JailbreakHub, JailbreakBench, MultiJail, and SafeMTData) and normal prompts (i.e.… See the full description on the dataset page: https://huggingface.co/datasets/xunguangwang/JailbreakGuardrailBenchmark.SRM_data
Dataset Card for SRM Data
ShareGPT-X
Dataset Summary
ShareGPT-X is an expanded, snapshot of ~92K (ChatGPT) one-to-one human & LLM conversations harvested from X.com (formerly Twitter).The corpus spans January 2024 → present (last ingest 2025-05) and is built entirely from public "share" links that users posted to their timelines.Each thread contains the original user prompt plus the assistant’s reply; no system prompts or metadata are exposed.
Supported Tasks and Leaderboards
text-generation… See the full description on the dataset page: https://huggingface.co/datasets/DSULT-Core/ShareGPT-X.ample-math
AMPLE-Math
5,319 mathematics problems, each with a verified final answer and six references to that same
answer. The references differ only in how much of the reasoning they show, which makes them useful
for studying what a teacher's reference content contributes during distillation.
Problems and original reasoning come from the metadata configuration of
OpenThoughts-114k, and keep its
Apache-2.0 attribution. A question was kept only if all six references exist, every generated… See the full description on the dataset page: https://huggingface.co/datasets/xiuyuz/ample-math.XQ-MEval
Dataset Card for XQ-MEval
XQ-MEval is a quality-parallel benchmark dataset for automatic evaluation metrics on cross-lingual scoring bias.
Dataset Details
Dataset Description
XQ-MEval is a benchmark released under CC BY-S 4.0 for evaluating automatic metrics with respect to cross-lingual scoring bias. This dataset is constructed by injecting varying numbers of Multidimensional Quality Metric (MQM)-defined errors into high-quality translations… See the full description on the dataset page: https://huggingface.co/datasets/naist-nlp/XQ-MEval.AIME24-25_CoT_Verification
Dataset for ICLR 2026 Paper: Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
📌 Dataset Summary
This dataset contains the rollouts (reasoning traces) and verification results used in our ICLR 2026 paper. The data allows for the analysis of how Reinforcement Learning with Verifiable Rewards (RLVR) incentivizes the correct reasoning of Large Language Models (LLMs) on challenging mathematics benchmarks.
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/XumengWen/AIME24-25_CoT_Verification.xl-instruct
Dataset Card for XL-Instruct
This dataset card provides a summary of the XL-Instruct dataset, a resource for advancing the cross-lingual capabilities of Large Language Models. It was introduced in the paper XL-Instruct: Synthetic Data for Cross-Lingual Open-Ended Generation.
Dataset Details
Dataset Description
XL-Instruct is a high-quality, large-scale synthetic dataset designed to fine-tune LLMs for cross-lingual open-ended generation. The core task involves… See the full description on the dataset page: https://huggingface.co/datasets/viyer98/xl-instruct.Maux-Persian-SFT-30k
Maux-Persian-SFT-30k
Dataset Description
This dataset contains 30,000 high-quality Persian (Farsi) conversations for supervised fine-tuning (SFT) of conversational AI models. The dataset combines multiple sources to provide diverse, natural Persian conversations covering various topics and interaction patterns.
Dataset Structure
Each entry contains:
messages: List of conversation messages with role (user/assistant/system) and content
source: Source dataset… See the full description on the dataset page: https://huggingface.co/datasets/xmanii/Maux-Persian-SFT-30k.XIH-Bench
XIH-Bench
Benchmark for the paper "Language Shapes Instruction Hierarchy Compliance in Multilingual LLMs".
Instruction hierarchy (IH) requires models to prioritize instructions by source, so that
higher-priority instructions override lower-priority ones. XIH-Bench evaluates IH under both
same-language and cross-language conflicts across six languages, four domains and three
hierarchy settings.
78,894 evaluation instances
Paper: https://arxiv.org/abs/2607.23545
Code:… See the full description on the dataset page: https://huggingface.co/datasets/g1moon/XIH-Bench.XLingHealth
Dataset Card for "XLingHealth"
XLingHealth is a Cross-Lingual Healthcare benchmark for clinical health inquiry that features the top four most spoken languages in the world: English, Spanish, Chinese, and Hindi.
Statistics
Dataset
#Examples
#Words (Q)
#Words (A)
HealthQA
1,134
7.72 ± 2.41
242.85 ± 221.88
LiveQA
246
41.76 ± 37.38
115.25 ± 112.75
MedicationQA
690
6.86 ± 2.83
61.50 ± 69.44
#Words (Q) and \#Words (A) represent the average number of words… See the full description on the dataset page: https://huggingface.co/datasets/claws-lab/XLingHealth.diff-xyz
Diff-XYZ
This is a dataset for the paper: Diff-XYZ: A Benchmark for Evaluating Diff Understanding.
Diff-XYZ contains 1,000 real-world code edits sampled and filtered from
the CommitPackFT dataset.Each example provides three components: the original file contents (old_code), the modified contents (new_code), and
multiple diff representations (udiff, udiff-h, udiff-l, and search-replace).
These formats enable evaluation of LLM capabilities on three code editing tasks:
Apply: Given… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/diff-xyz.bayesian-llm-safety-inference
Bayesian Latent Safety-Trait Dataset
Summary
This dataset supports Bayesian latent-trait analysis of language-model safety behavior.
It contains 90 benchmark-derived roots, three matched prompt variants per root, responses
from four target models over five runs, two independent LLM ratings per response, and one
human rating for a stratified 540-response calibration subset.
The three dimensions are harmful compliance, sycophancy, and agentic protocol violation.… See the full description on the dataset page: https://huggingface.co/datasets/Charly-X/bayesian-llm-safety-inference.TEA-Dialog
TEA-Dialog
TEA-Dialog is a dialogue dataset released as part of TEA-Bench: A Systematic Benchmarking of Tool-enhanced Emotional Support Dialogue Agent.
This repository contains the released datasets of TEA-Bench, including TEA-Scenario and TEA-Dialog.
Dataset Description
TEA-Dialog contains multi-turn emotional support dialogues generated/evaluated under TEA-Bench scenarios. Each example includes scenario information, dialogue messages, user type, end reason, and… See the full description on the dataset page: https://huggingface.co/datasets/XingYuSSS/TEA-Dialog.synthid-qwen3-4b-instruct-2507-wildchat
Qwen3-4B SynthID three-arm corpus
This export contains aligned unwatermarked, SynthID key-A, and SynthID key-B
responses from Qwen/Qwen3-4B-Instruct-2507. Matched splits share prompts
and request seeds across configurations; unmatched splits use mutually disjoint
prompt pools.
Export complete for its source work queue: true.
Generation profile
Model revision: cdbee75f17c01a7cc42f958dc650907174af0554
Native model dtype: bfloat16
Maximum generated tokens: 4096… See the full description on the dataset page: https://huggingface.co/datasets/xlr8harder/synthid-qwen3-4b-instruct-2507-wildchat.ChatGPT-Jailbreak-Prompts
Dataset Card for Dataset Name
Name
ChatGPT Jailbreak Prompts
Dataset Summary
ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT.
Languages
[English]
ml-intern-sessions
ML Intern session traces
This dataset contains ML Intern coding agent session traces uploaded from local
ML Intern runs. The traces are stored as JSON Lines files under sessions/,
with one file per session.
Links
ML Intern demo: https://smolagents-ml-intern.hf.space
ML Intern CLI: https://github.com/huggingface/ml-intern
Data description
Each *.jsonl file contains a single ML Intern session converted to a
Claude-Code-style event stream for the… See the full description on the dataset page: https://huggingface.co/datasets/xucenying/ml-intern-sessions.xRouteBench
xRouteBench — LLM Routing Benchmark
xRouteBench is a benchmark for training and evaluating LLM routers — systems that pick the best LLM from a candidate pool for each incoming query, trading off performance vs. price cost.
Every query in each scenario was executed against all 18 candidate LLMs, recording each model's response, task performance, token usage, and latency. A router learns from the train split which model to pick, and is evaluated on test.
Scenarios… See the full description on the dataset page: https://huggingface.co/datasets/Samyol/xRouteBench.
