datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
car-bench-dataset
CAR-Bench Dataset
CAR-Bench is a benchmark for evaluating AI voice assistants in a realistic automotive (car) environment.
It tests an agent's ability to correctly use vehicle control tools, handle disambiguation, and avoid hallucinations.
Dataset Structure
The dataset is organized into task configs and mock data configs:
Tasks
Each task defines a user persona, an instruction, the initial vehicle/environment context, and the ground-truth sequence of tool-call… See the full description on the dataset page: https://huggingface.co/datasets/johanneskirmayr/car-bench-dataset.GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Carlosaug47/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.TCM-Pretrain-Data-ShizhenGPT
📚 Introduction
This dataset is the pre-training dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source the largest existing TCM corpus dataset (over 5B tokens) from TCM-related websites and books. Additionally, we also open-source the largest scale TCM image-text pretraining dataset.
For details, see our paper and GitHub repository.
📊 Dataset Overview
The open-sourced pre-training dataset consists of five parts:… See the full description on the dataset page: https://huggingface.co/datasets/CarsonnnNN/TCM-Pretrain-Data-ShizhenGPT.gspc-care
GSPC — care bank (CareBench)
Council of AI measurement bank. Measurement, not certification.
Bank. Frozen split. Live n is the matching axis on GET https://councilof.ai/api/gspc, not a Hub score. Not a certificate. Art 50 (EUR-Lex): 2 August 2026 live; marking grace 2 December 2026.
Live measurement. This bank stands behind the care row of the live GSPC board: GET https://councilof.ai/api/gspc?axis=care (family, kind, status and n are on that row, never typed here; the whole… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-care.ko-instruction-dataset
고품질 한국어 데이터셋
한국어로 이루어진 고품질 한국어 데이터셋 입니다.
WizardLM-2-8x22B 모델을 사용하여 WizardLM: Empowering Large Language Models to Follow Complex Instructions에서 소개된 방법으로 생성되었습니다.
@article{koinstructiondatasetcard,
title={CarrotAI/ko-instruction-dataset Card},
author={CarrotAI (L, GEUN)},
year={2024},
url = {https://huggingface.co/datasets/CarrotAI/ko-instruction-dataset}
}
EMPA-character_card
English | 中文
EMPA: Evaluating Persona-Aligned Empathy as a Process
Empathy Potential Modeling and Assessment
Paper |
Dataset |
Citation
🍊 Overview
EMPA is the first benchmark to evaluate empathy as a dynamic Process rather than a static response. We posit that true empathetic capability resides in the Latent Space of dialogue and must be captured through multi-turn interaction trajectories.
Unlike traditional benchmarks that focus solely on… See the full description on the dataset page: https://huggingface.co/datasets/SalmonTell/EMPA-character_card.carnice-glm5-hermes-traces
Carnice GLM-5 Hermes Traces
This dataset is a merged release bundle of GLM-5 traces collected through the Hermes Agent harness.
It was generated by running the carnice_trace_prompt_bank_v4 prompt bank through Hermes Agent with:
z-ai/glm-5 via OpenRouter
local/file/terminal/code-execution tools for local tasks
Hermes browser tools plus Tavily-backed web_search / web_extract for web tasks
isolated disposable workspaces per prompt
This release is prepared for Hugging Face upload and… See the full description on the dataset page: https://huggingface.co/datasets/kai-os/carnice-glm5-hermes-traces.carnice-agent-trance-prompt-bank
Carnice Agent Trace Prompt Bank
This repository is a curated prompt bank for collecting agent traces.
It is not a trace dataset by itself. It is the input side: prompts that can be run through an agent harness, then logged into traces with tool calls, observations, and final answers.
The goal of this release is practical:
keep prompts that work well in an agent harness
remove prompts that assume hidden local state or user-private state
expand browser and long-horizon tasks enough… See the full description on the dataset page: https://huggingface.co/datasets/kai-os/carnice-agent-trance-prompt-bank.tarotoo-tarot-card-meanings
Tarotoo Tarot Card Meanings
A complete, structured dataset of all 78 tarot cards (22 Major Arcana + 56 Minor Arcana) in the Rider–Waite–Smith tradition. Published by Tarotoo. These are the card meanings that ground the AI-generated readings on Tarotoo.com.
Dataset details
Curated by: Tarotoo (tarotoo.com)
Language: English
License: MIT
Rows: 78 (one per card) · Fields: 22
DOI (Zenodo, cite this): 10.5281/zenodo.21514483
Concept DOI (Zenodo, always resolves to the… See the full description on the dataset page: https://huggingface.co/datasets/Tarotoo/tarotoo-tarot-card-meanings.TCM-Instruction-Tuning-ShizhenGPT
📚 Introduction
This dataset is a fine-tuning dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source 245K multimodal Chinese medicine instruction data, including text instructions, visual instructions, and signal instructions for TCM.
For details, see our paper and GitHub repository.
📊 Dataset Overview
The open-sourced fine-tuning dataset consists of three parts:
Modality
Data Quantity
TCM Text Instructions
📝 Text… See the full description on the dataset page: https://huggingface.co/datasets/CarsonnnNN/TCM-Instruction-Tuning-ShizhenGPT.cartographer-corpus
Cartographer Agent Corpus
Domain-specific training corpus for the Cartographer Agent covering financial, legal, and governance knowledge. 9 chapters of curated prompt/completion pairs.
Chapters
File
Topic
Records
ch1-credit-repair-mastery.jsonl
Credit repair strategies
9
ch2-sovereign-trust-architecture.jsonl
Trust structure design
8
ch3-ach-dispute-protocol.jsonl
ACH dispute procedures
7
ch4-fcra-zombie-debt-credit-law.jsonl
FCRA and debt law
7… See the full description on the dataset page: https://huggingface.co/datasets/Snapkitty/cartographer-corpus.tend
TEND (Gold)
This dataset publishes execution-validated gold-tier examples from the
TEND pipeline: natural-language questions paired with
SQL schema, gold SQL, generated MongoDB schema/query, and plain-English
documentation. It is designed for multi-task research spanning Text→SQL,
SQL→MongoDB, and MongoDB→Documentation.
Every published row is execution-validated. For each example, the pipeline
runs the gold sql_query on PostgreSQL and the generated nosql_query
on MongoDB, then… See the full description on the dataset page: https://huggingface.co/datasets/care2achieve/tend.EMPA-character_card
English | 中文
EMPA: Evaluating Persona-Aligned Empathy as a Process
Empathy Potential Modeling and Assessment
Paper |
Dataset |
Citation
🍊 Overview
EMPA is the first benchmark to evaluate empathy as a dynamic Process rather than a static response. We posit that true empathetic capability resides in the Latent Space of dialogue and must be captured through multi-turn interaction trajectories.
Unlike traditional… See the full description on the dataset page: https://huggingface.co/datasets/Jiechen0328/EMPA-character_card.synthetic-abandoned-cart-email-examples
Synthetic Abandoned Cart Email Examples
An entirely synthetic, bilingual collection of abandoned-cart email drafts with transparent checklist annotations. It is intended for education, prototyping, and evaluation, and contains no real recipients, customer messages, orders, merchant data, or campaign results.
Dataset Description
The dataset mirrors the five visible checks in NeuroCheckout's public Abandoned Cart Email Checker:
message clarity;
primary call to… See the full description on the dataset page: https://huggingface.co/datasets/neurocheckout-ai/synthetic-abandoned-cart-email-examples.ZELAIHANDICLEAN
Dataset Summary
A large, cleaned Basque-language corpus originally based on the ZelaiHandi dataset, augmented with books and Wikipedia articles to support language-modeling experiments. The data have been normalized and stripped of extraneous whitespace, blank lines and non-linguistic characters.
For example, this are the stats for Ekaia subset:
Metric
Value
Initial characters (Ekaia subset)
14,480,942
Final characters (after cleaning)
12,746,071
Overall cleaned
11.98… See the full description on the dataset page: https://huggingface.co/datasets/Carlos1411/ZELAIHANDICLEAN.car-bench-dataset
CAR-Bench Dataset
CAR-Bench is a benchmark for evaluating AI voice assistants in a realistic automotive (car) environment.
It tests an agent's ability to correctly use vehicle control tools, handle disambiguation, and avoid hallucinations.
Dataset Structure
The dataset is organized into task configs and mock data configs:
Tasks
Each task defines a user persona, an instruction, the initial vehicle/environment context, and the ground-truth sequence of… See the full description on the dataset page: https://huggingface.co/datasets/nomador/car-bench-dataset.carnice-glm5-hermes-traces
Carnice GLM-5 Hermes Traces
This dataset is a merged release bundle of GLM-5 traces collected through the Hermes Agent harness.
It was generated by running the carnice_trace_prompt_bank_v4 prompt bank through Hermes Agent with:
z-ai/glm-5 via OpenRouter
local/file/terminal/code-execution tools for local tasks
Hermes browser tools plus Tavily-backed web_search / web_extract for web tasks
isolated disposable workspaces per prompt
This release is prepared for Hugging Face upload and… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/carnice-glm5-hermes-traces.credit_card_fraud_disputes
Credit_Card_Fraud_Disputes (Synthetic B2B Dataset Preview)
Add me on Discord: xomohappy for access support, delivery questions, or product questions about this premade commercial dataset.
This is a premium, privacy-compliant, industry-safe synthetic dataset simulating Credit Card Billing Disputes & Fraud Logs for B2B applications.
About this Dataset
This dataset is generated programmatically using large language models combined with a strict data curation and… See the full description on the dataset page: https://huggingface.co/datasets/HaseebDev/credit_card_fraud_disputes.ZELAIHANDICLEANED
Dataset Summary
A large, cleaned Basque-language corpus originally based on the ZelaiHandi dataset, augmented with books and Wikipedia articles to support language-modeling experiments. The data have been normalized and stripped of extraneous whitespace, blank lines and non-linguistic characters.
Supported Tasks
Causal language modeling
Masked language modeling
Next-sentence prediction
Any downstream Basque NLP task (fine-tuning)
Languages
Basque (eu)… See the full description on the dataset page: https://huggingface.co/datasets/Carlos1411/ZELAIHANDICLEANED.wisconsin-building-codes-grpo
Wisconsin Building Codes Q&A Dataset (GRPO-Formatted)
This dataset is a version of the Wisconsin Building Codes Q&A Dataset formatted specifically for Grouped-Reward-Optimization (GRPO) training with libraries like TRL and unsloth.
Dataset Description
This dataset contains 13,200 prompts designed for training preference models. Each record includes a user prompt (prompt), a "chosen" high-quality response, and a placeholder for a "rejected" response.
Training samples: 11… See the full description on the dataset page: https://huggingface.co/datasets/carlscape/wisconsin-building-codes-grpo.TSAR2025_SharedTask_RCTS_Test-Data
Citation
@inproceedings{alva-manchego-etal-2025-findings,
title = "Findings of the {TSAR} 2025 Shared Task on Readability-Controlled Text Simplification",
author = "Alva-Manchego, Fernando and Stodden, Regina and Imperial, Joseph Marvin and Barayan, Abdullah and North, Kai and Tayyar Madabushi, Harish",
editor = "Shardlow, Matthew and Alva-Manchego, Fernando and North, Kai and Stodden, Regina and Saggion, Horacio and Khallaf, Nouran and Hayakawa, Akio"… See the full description on the dataset page: https://huggingface.co/datasets/cardiffnlp/TSAR2025_SharedTask_RCTS_Test-Data.caramelo-dataset
Caramelo — pares de correção de estilo
414 pares instrução → resposta que ensinam um modelo a responder na voz de escrita do Guilherme Favaron: direto ao ponto, argumentando com dados e exemplos, em português do Brasil, sem hype e sem emoji. É o dado de treino do Caramelo 3.4.2 (Gemma 3 4B + LoRA) e do Caramelo 4.4.1 (Gemma 4 E4B + LoRA), a versão em produção em ia-caramelo.com.
Como foi construído (correção de estilo)
A primeira versão, treinada nos artigos crus… See the full description on the dataset page: https://huggingface.co/datasets/guifav/caramelo-dataset.in-car-context-benchmark
Benchmarking contextual understanding for in-car conversational systems
This dataset contains the complete evaluation benchmarks, user utterances, venue recommendations, and failure-annotated responses for evaluating in-car Conversational Question Answering (ConvQA) systems.
Official Code & Implementation: github.com/saydemr/judgebench
Paper (Journal of Systems and Software, 2026): doi.org/10.1016/j.jss.2026.112915 or arxiv.org/abs/2512.12042
📌 Quickstart
from… See the full description on the dataset page: https://huggingface.co/datasets/saydemr/in-car-context-benchmark.Carla-Road-X1
Carla-Road-X1: OpenDRIVE Road Network Generation Dataset
Text-to-xodr dataset for training language models to generate OpenDRIVE (.xodr) road network files from natural language descriptions.
Dataset Summary
Train samples: 385,929
Val samples: 42,882
Template samples: 40 (train: 36, val: 4)
Format: JSONL with Gemma chat template (system/user/assistant messages)
Data Sources
Source
Count
Description
OSM-converted xodr sub-networks
~428K… See the full description on the dataset page: https://huggingface.co/datasets/NCUT-AI/Carla-Road-X1.tulu-3-car-50k
Tulu-3-CaR-50K
Project | Github | Paper | HuggingFace's collection
This dataset includes 50K high-quality and diverse SFT data sampled from Tulu3 using CaR.
Performance
Method
Data Size
ARC
BBH
GSM
HE
MMLU
IFEval
Avg_obj
AE
MT
Wild
Avg_sub
Avg
Pool
939K
69.15
63.88
83.40
63.41
65.77
67.1068.79
8.94
6.86
-24.66
38.40
53.59
Random
50K
74.24
64.80
70.36
51.22
63.86
61.00
64.25
8.57
7.06
-22.15
39.36
51.81
ZIP
50K
77.63
63.00
52.54
35.98
65.00
61.00
59.19… See the full description on the dataset page: https://huggingface.co/datasets/xsample/tulu-3-car-50k.pbd-autism-caregiver
Privacy-by-Design in AI-Assisted Systems for Caregivers of Children with Autism: A Secure Multi-Agent Architecture
Dataset Description
This dataset accompanies the paper "Privacy-by-Design in AI-Assisted Systems for Caregivers of Children with Autism: A Secure Multi-Agent Architecture". The system is a privacy-by-design multi-agent architecture integrating Retrieval-Augmented Generation (RAG), Data Loss Prevention (DLP), consent management, explainable AI (XAI), and audit… See the full description on the dataset page: https://huggingface.co/datasets/Ionutcroitoru/pbd-autism-caregiver.ko-code-alpaca-QAcode-alpaca QA 데이터셋입니다.
필터링이 어느정도 필요합니다.
참고하시고 사용하시면 됩니다.
ZELAITESTkmmlu-conversation-sampleKmmlu 데이터를 이용해서 대화 데이터셋 샘픔을 생성하였습니다.
멀티턴 데이터셋으로 학습용도로 만들어졌습니다.
HelpSteer3_dpo_format
Dataset: nvidia/HelpSteer3
HelpSteer3 is an open-source dataset (CC-BY-4.0) that supports aligning models to become more helpful in responding to user prompts.
Preference Score Integer from -3 to 3, corresponding to:
-3: Response 1 is much better than Response 2
-2: Response 1 is better than Response 2
-1: Response 1 is slightly better than Response 2
0: Response 1 is about the same as Response 2
1: Response 2 is slightly better than Response 1
2: Response 2 is better than Response… See the full description on the dataset page: https://huggingface.co/datasets/CarrotAI/HelpSteer3_dpo_format.
