datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
xcopa
Dataset Card for "xcopa"
Dataset Summary
XCOPA: A Multilingual Dataset for Causal Commonsense Reasoning
The Cross-lingual Choice of Plausible Alternatives dataset is a benchmark to evaluate the ability of machine learning models to transfer commonsense reasoning across
languages. The dataset is the translation and reannotation of the English COPA (Roemmele et al. 2011) and covers 11 languages from 11 families and several areas around
the globe. The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/cambridgeltl/xcopa.XL-DocBench
XL-DocBench
Evidence-grounded reasoning across hundreds or thousands of pages.
Fully verified by 194 human experts.
Hongchen Wei1,†,‡, Yuanzhe Wang2,†,‡,
Bei Liu2,*, Yifan Yang2, Qi Dai2,
Ruichun Ma2, Kai Qiu2, Yunsheng Li2,
Dongdong Chen2, Chong Luo2,
Zhenzhong Chen1, Baining Guo2
1Wuhan University 2Microsoft
†Equal contribution ‡Work done during an internship at MSRA
*Project leader
Project Page ·
Paper ·
Live Leaderboard… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/XL-DocBench.gspc-xr
GSPC — cross reality bank (XRAIV)
Council of AI measurement bank. Measurement, not certification.
Bank. Frozen split. Live n is the matching axis on GET https://councilof.ai/api/gspc, not a Hub score. Not a certificate. Art 50 (EUR-Lex): 2 August 2026 live; marking grace 2 December 2026.
Live measurement. This bank stands behind the cross-reality row of the live GSPC board: GET https://councilof.ai/api/gspc?axis=cross-reality (family, kind, status and n are on that row, never… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-xr.effibench-x
Dataset Card for EffiBench-X
EffiBench-X is the first multi-language benchmark designed specifically to evaluate the efficiency of LLM-generated code across six programming languages: Python, C++, Java, JavaScript, Ruby, and Golang. The dataset comprises 623 competitive programming problems paired with human expert solutions as efficiency baselines.
Dataset Details
Dataset Description
EffiBench-X addresses critical limitations in existing code generation… See the full description on the dataset page: https://huggingface.co/datasets/EffiBench/effibench-x.Feng-Chat简体中文 | English
Feng-Chat
“解答世间万物”
Feng-Chat 是一批中文 conversational SFT 数据。单轮问答好做,多轮里的追问、接话、拐弯和突然补一句,才是这批数据真正想留下来的东西。QA 抽取后依次做去重、assistant-only PPL、message-level Guard 和主题分类,最后发布为标准 messages。
整个 pipeline 只负责筛选,不改写峰哥的表达。公开 release 仅保留训练和分析需要的字段,不包含内部 prompt、reasoning、raw response、本地路径或原始来源 ID。
数据概览
指标
最终结果
发布记录
33,468
QA 轮次
40,645
多轮记录
3,952(11.81%)
单条最多轮次
30
伪名化来源分组
729
内容日期范围
2022-01-02 ~ 2026-08-20
assistant loss tokens
4,806,476… See the full description on the dataset page: https://huggingface.co/datasets/xfalcon9/Feng-Chat.terminal-bench
Terminal-Bench Dataset
This dataset contains tasks from Terminal-Bench, a benchmark for evaluating AI agents in real terminal environments. Each task is packaged as a complete, self-contained archive that preserves the exact directory structure, binary files, Docker configurations, and test scripts needed for faithful reproduction.
The archive column contains a gzipped tarball of the entire task directory.
Dataset Overview
Terminal-Bench evaluates AI agents on… See the full description on the dataset page: https://huggingface.co/datasets/XUO/terminal-bench.XhotpotQA
XHotpotQA
AUDITED V1.1 · PUBLIC
Cross-lingual multi-hop question answering over mixed-language evidence
One question · multiple evidence paragraphs · languages may change between hops
23,066 released rows
Supplied-candidate MHQA
Parquet · 22 fields
CC BY-SA 4.0
Navigate:
Overview ·
Load ·
Structure ·
Quality ·
Findings ·
Citation ·
Paper ·
Code v0.4.1
XHotpotQA is a controlled benchmark for… See the full description on the dataset page: https://huggingface.co/datasets/Iman998/XhotpotQA.XhotpotQA-V2
XHotpotQA V2
AUDITED RC1 · INCOMPLETE
Cross-lingual multi-hop QA over mixed-language evidence
Gemma 4 31B · source-aligned fields · transparent release gate
STATUS · RC1
ROWS · 22,836
GENERATOR · Gemma 4 31B
MISSING · 230
LICENSE · CC BY-SA 4.0
Navigate:
Overview ·
Coverage ·
Load ·
Schema ·
Methodology ·
Quality ·
Citation ·
Paper ·
Collection
RC1 release warning. This is a… See the full description on the dataset page: https://huggingface.co/datasets/Iman998/XhotpotQA-V2.cross-lingual-pitfalls
Cross-Lingual Pitfalls
Cross-Lingual Pitfalls is a fixed, failure-focused dataset from the ACL 2025 paper "Cross-Lingual Pitfalls: Automatic Probing Cross-Lingual Weakness of Multilingual Large Language Models." It contains 6,713 bilingual English-to-target-language question pairs across 16 target languages. The paper's search-based multilingual LLM evaluation method uses beam search and LLM-based simulation to discover cases where a model answers correctly in English but fails… See the full description on the dataset page: https://huggingface.co/datasets/xzx34/cross-lingual-pitfalls.Maux-Persian-SFT-30k
Maux-Persian-SFT-30k
Dataset Description
This dataset contains 30,000 high-quality Persian (Farsi) conversations for supervised fine-tuning (SFT) of conversational AI models. The dataset combines multiple sources to provide diverse, natural Persian conversations covering various topics and interaction patterns.
Dataset Structure
Each entry contains:
messages: List of conversation messages with role (user/assistant/system) and content
source: Source dataset… See the full description on the dataset page: https://huggingface.co/datasets/xmanii/Maux-Persian-SFT-30k.ntu_adl_questionlean-proof-or-refute-300
Lean Proof-or-Refute 300
Lean Proof-or-Refute 300 is a compact collection of 300 formal reasoning
problems grounded in Lean 4 and Mathlib. Each problem starts from a verified
Mathlib theorem, makes one small numerical or operator mutation, and asks the
model to return either:
a Lean certificate proving the mutated proposition; or
a Lean certificate proving the exact negation of the complete proposition.
The model receives the related source theorem, a bounded source excerpt… See the full description on the dataset page: https://huggingface.co/datasets/xlr8harder/lean-proof-or-refute-300.XLingHealth
Dataset Card for "XLingHealth"
XLingHealth is a Cross-Lingual Healthcare benchmark for clinical health inquiry that features the top four most spoken languages in the world: English, Spanish, Chinese, and Hindi.
Statistics
Dataset
#Examples
#Words (Q)
#Words (A)
HealthQA
1,134
7.72 ± 2.41
242.85 ± 221.88
LiveQA
246
41.76 ± 37.38
115.25 ± 112.75
MedicationQA
690
6.86 ± 2.83
61.50 ± 69.44
#Words (Q) and \#Words (A) represent the average number of words… See the full description on the dataset page: https://huggingface.co/datasets/claws-lab/XLingHealth.xai-questions-datasetExplore the questions users have for robots across a diverse set of situations!
You can read the paper here: What Questions Should Robots Be Able to Answer? A Dataset of User Questions for Explainable Robotics!
from datasets import load_dataset
dataset = load_dataset("lwachowiak/xai-questions-dataset")
dataset['train'][0]
The analysis code can be found on GitHub
Paper Abstract
With the increased use of large language models and conversational interfaces in human–robot… See the full description on the dataset page: https://huggingface.co/datasets/lwachowiak/xai-questions-dataset.ChatGPT-Jailbreak-Prompts
Dataset Card for Dataset Name
Name
ChatGPT Jailbreak Prompts
Dataset Summary
ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT.
Languages
[English]
X-SVAMP_en_zh_ko_it_es
X-SVAMP
🤗 Paper | 📖 arXiv
Dataset Description
X-SVAMP is an evaluation benchmark for multilingual large language models (LLMs), including questions and answers in 5 languages (English, Chinese, Korean, Italian and Spanish).
It is intended to evaluate the math reasoning abilities of LLMs. The dataset is translated by GPT-4-turbo from the original English-version SVAMP.
In our paper, we evaluate LLMs in a zero-shot generative setting: prompt the instruction-tuned LLM with… See the full description on the dataset page: https://huggingface.co/datasets/zhihz0535/X-SVAMP_en_zh_ko_it_es.pinga-fogo-chico-xavier
🎙️ Pinga-Fogo com Chico Xavier — TV Tupi, 1971
As duas entrevistas históricas do médium Chico Xavier, transmitidas ao vivo pela
TV Tupi em 1971, transcritas e estruturadas em turnos de fala com timestamp.
345 turnos (115 deles respostas do próprio Chico Xavier), a partir de
6 horas de áudio — o registro mais extenso do médium falando de improviso,
sem edição, diante de um painel de jornalistas.
Arquivos
Arquivo
Programa
Turnos
Respostas do Chico… See the full description on the dataset page: https://huggingface.co/datasets/ia-espirita/pinga-fogo-chico-xavier.ChatGPT-Jailbreak-Prompts
Dataset Card for Dataset Name
Name
ChatGPT Jailbreak Prompts
Dataset Summary
ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT.
Languages
[English]
StackMathematics-XL
Mathematics Stack Exchange Q&A Dataset
A curated dataset of question-and-answer pairs harvested from Mathematics Stack Exchange (math.stackexchange.com). Each entry contains the question, the top answer (if available), metadata (scores, tags, URL), and a combined Q/A text useful for training and evaluation of question-answering and tutoring models.
Dataset summary
Source: Mathematics Stack Exchange (via the official Stack Exchange API)
Content: High-quality Q&A… See the full description on the dataset page: https://huggingface.co/datasets/FurkanNar/StackMathematics-XL.autonomous-linux-kernel-ebpf-xdp-suite
⚡ Autonomous Linux Kernel, eBPF & XDP Programmable Dataplane Suite (2026)
A Production-Grade, Verifiable Synthetic Corpus for Training Autonomous Linux Kernel & eBPF Systems Agents
⚡ Overview & Industry Problem
Modern hyperscale cloud datacenters, bare-metal Kubernetes clusters, and low-latency financial trading nodes rely on in-kernel programmable dataplanes: eBPF, AF_XDP zero-copy rings, Traffic Control (TC) shapers, BPF LSM security hooks… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/autonomous-linux-kernel-ebpf-xdp-suite.ChatGPT-Jailbreak-Prompts
Dataset Card for Dataset Name
Name
ChatGPT Jailbreak Prompts
Dataset Summary
ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT.
Languages
[English]
CBSE-Class-12th_2024_PYQs__structuredThis data set contains the CBSE Class 12 2024 papers in a structured format. The papers are annotated with topic and chapter names, and the figures are parsed and their paths annotated.
RoomReader-Full
RoomReader-Full
RoomReader-Full is a multimodal benchmark for fine-grained understanding of how a
speaker manages information during a conversation. Given a short target video clip
and its audio/dialogue, a model must determine whether the response is consistent
with the relevant interactional expectation and assign one of six information
strategy labels.
The benchmark is designed for cases where the observable response alone is not
enough to determine the label. Formal gold… See the full description on the dataset page: https://huggingface.co/datasets/xilinghuiye/RoomReader-Full.counterfactual-trace-audits
Counterfactual Trace Audits
This dataset contains 25,600 unique synthetic, self-contained reasoning
problems. Each problem shows an original computation over a list or binary
tree, applies a counterfactual semantic patch, and asks for two K/R/X
judgments plus both complete patched evaluation traces.
Prompt format v2 explicitly defines trace notation and the nested answer
schema. Tree-height prompts also include a small example of the pruning marker.
The displayed answer shape… See the full description on the dataset page: https://huggingface.co/datasets/xlr8harder/counterfactual-trace-audits.ChatGPT-Jailbreak-Prompts-rubend18
Dataset Card for Dataset Name
Name
ChatGPT Jailbreak Prompts
Dataset Summary
ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT.
Languages
[English]
GraphWalkerBenchThis repository contains the GraphWalkerBench dataset from the paper GraphWalker: Agentic Knowledge Graph Question Answering via Synthetic Trajectory Curriculum.
Code: https://github.com/XuShuwenn/GraphWalker
s1_54k_filter_with_isreasoning
Dataset Card for XuHu6736/s1_54k_filter_with_isreasoning
Dataset Description
XuHu6736/s1_54k_filter_with_isreasoning is an enhanced version of the XuHu6736/s1_54k_filter dataset. This version includes additional annotations to assess the suitability of each question for reasoning training. These annotations, isreasoning_score and isreasoning, were generated using the deepseek-v3 model.
The purpose of these new fields is to allow users to filter, weight, or specifically… See the full description on the dataset page: https://huggingface.co/datasets/XuHu6736/s1_54k_filter_with_isreasoning.aibe-xxi-set-a
AIBE 16–21 MCQ Benchmark
All 100 multiple-choice questions from each of the last six All India Bar Examinations (AIBE XVI–XXI, 2021–2026), as six configs of one dataset, each paired with the most authoritative answer key available and labeled with the same official 19-subject BCI syllabus taxonomy.
Config
Exam
Held
Set
Question source
Answer key
Withdrawn
Multi-answer
aibe16
AIBE XVI
Oct 2021
C
Delhi Law Academy compilation
DLA 4-set key table (no official copy… See the full description on the dataset page: https://huggingface.co/datasets/rushankg/aibe-xxi-set-a.x_dataset_193
Bittensor Subnet 13 X (Twitter) Dataset
Miner Data Compliance Agreement
In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy.
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/RentonWEB3/x_dataset_193.XLingHealth
