datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
public-domain-poetry
Overview
This dataset is a collection of approximately 38,500 poems from https://www.public-domain-poetry.com/.
Language
The language of this dataset is English.
License
All data in this dataset is public domain, which means you should be able to use it for anything you want, as long as you aren't breaking any law in the process of doing so.
GeneratingQuestions
HVU_QA
HVU_QA is an open-source Vietnamese Question-Context-Answer (QCA) corpus, accompanied by supporting tools, created to facilitate the development of FAQ-style question generation and question answering systems, particularly for low-resource language settings. The dataset was developed by a research team at Hung Vuong University, Phu Tho, Vietnam, led by Dr. Ha Nguyen, Deputy Head of the Department of Engineering Technology. HVU_QA was constructed using a fully automated… See the full description on the dataset page: https://huggingface.co/datasets/DANGDOCAO/GeneratingQuestions.HundredCV-Chat
百人对话数据集
HundredCV-Chat: A Dataset of Daily Chatting Developed on HundredCVs
简介
本项目提出一个全新的中文多轮对话数据集(HundredCV-Chat),该数据集由 100 位青年的简历数据集 HundredCVs 开发而来,共包含 24,750 组日常闲聊对话数据。
数据集具有如下特点:
自动化标注:HundredCV-Chat 中的对话均由 Deepseek-V3 大模型生成,不涉及任何人工标注,因此同时保证了大规模数据量和低成本优势。
多样性话题:HundredCV-Chat 中的对话话题涵盖了校园生活、工作经验、兴趣爱好、生活琐事等多个方面,与真实生活联系紧密,尤其适用于开发年轻化应用。
高质量对话:利用 Deepseek 强大的生成能力和全面的知识,HundredCV-Chat 的对话内容在流畅度、拟人性、多样性方面均显著优于现有的开源对话数据集。
数据样例
HundredCV-Chat 含有 24… See the full description on the dataset page: https://huggingface.co/datasets/Jax-dan/HundredCV-Chat.danbooru-wiki-2026
danbooru-wiki-2026-04-28
About
Wiki pages about the danbooru tags on danbooru.donmai.us. The wiki contains the description of each tag.
This dataset was collected from the wiki data using danbooru API and then merged with tags data (category_name & post_count). No entries were removed, no matter how many posts they have. The resulting set was manually filtered. Originally some of the lists, tag_group:, about:, help:, howto:, etc, pages were labled as "general"… See the full description on the dataset page: https://huggingface.co/datasets/kierarkia/danbooru-wiki-2026.RetailBanking-Conversations
Dataset Description
RetailBanking-Conversations is a synthetic dataset designed to train and evaluate language models in the retail banking domain, it has been created using the open source library wizardSdata that eable the creation of synthetic datasets in any field.
The dataset contains 320 realistic conversations, across 160 unique financial profiles and 10 key retail banking topics, between financial advisors and clients, covering 10 main categories of banking products and… See the full description on the dataset page: https://huggingface.co/datasets/danystar/RetailBanking-Conversations.alpaca-cleaned-italian
Dataset Card for Alpaca-Cleaned-Italian
About the translation and the original data
The translation was done with X-ALMA, a 13-billion-parameter model that surpasses state-of-the-art open-source multilingual LLMs (as of Q1 2025, paper here).
The original alpaca-cleaned dataset is also kept here so that there is parallel data for Italian and English.
Additional notes on the translation
Despite the good quality of the translation, errors, though rare, are… See the full description on the dataset page: https://huggingface.co/datasets/DanielSc4/alpaca-cleaned-italian.DialogES
DialogES: An Large Dataset for Generating Dialogue Events and Summaries
简介
本项目提出一个对话事件抽取和摘要生成数据集——DialogES,数据集包括共 44,672 组多轮对话,每组对话采用自动的方式标注出对话事件和对话摘要。
该数据集主要用于训练对话摘要模型,研究者可采用单任务和多任务学习的方式利用本数据集。
收集过程
对话采集:本数据集中的对话数据收集自两个已有的开源数据集,即 NaturalConv 和 HundredCV-Chat;
事件标注:采用 few-shot in-context learning 的方式标注,人工标注出 5 个对话-事件样本,作为演示样例嵌入大模型的提示词中,然后引导模型对输入对话进行标注;
摘要标注:采用 zero-shot learning 的方式标注,编写提示词"请总结对话的主要内容:",要求大模型生成对话摘要;
自动标注:按照以上准则,本项目利用 Deepseek-V3… See the full description on the dataset page: https://huggingface.co/datasets/Jax-dan/DialogES.dan_remixed
DAN Remixed Data Set
License: MIT
Overview
The DAN Remixed data set's goal is to advance freedom in AI and AI usage. This dataset builds on early efforts to resist heavy-handed censorship and surveillance in AI, originally inspired by the DAN dataset. The original dataset, though significant, was somewhat poorly written (meaning many typos and inconsistencies) and had highly violent completions. This version improves the dataset's overall quality and replaces needlessly… See the full description on the dataset page: https://huggingface.co/datasets/UnfilteredAI/dan_remixed.saas-product-support-sharegpt-1k
SaaS/Tech Product Support — Multi-Turn SFT Dataset
A domain-specific supervised fine-tuning dataset for
SaaS and tech product support conversations, built for
LLM fine-tuning and instruction tuning.
Dataset Summary
This dataset contains 1,200 multi-turn English conversations
between a customer and a support agent, covering common
SaaS/tech support scenarios: bug reports, billing issues,
API errors, authentication problems, onboarding blockers,
integration failures… See the full description on the dataset page: https://huggingface.co/datasets/Dang-DN-VN/saas-product-support-sharegpt-1k.RUCAIBox-Story-Generation-Alpacahttps://huggingface.co/datasets/RUCAIBox/Story-Generation
RUC AI Box HC Story Generation augmented and converted to alpaca format.
No filtering has been done.
Dans-SystemmaxxThis dataset would not be possible without Gryphe and Caitlyns amazing work on https://huggingface.co/datasets/Gryphe/Sonnet3.5-SlimOrcaDedupCleaned and https://huggingface.co/datasets/cgato/SlimOrcaDedupCleaned
CCOpenBooks
Dataset Card for CC OpenBooks
Dataset Description
CC OpenBooks is a curated collection of high quality non-fiction books. All texts are from CC-By-4.0 sources, with no license ambiguity.
The documents are normalized to markdown, and care is taken to ensure most formatting (e.g. inline LaTeX) remains intact. Files are manually inspected and cleaned of all defects wherever possible.
Source Data
The following Openstax collections were used in creating this… See the full description on the dataset page: https://huggingface.co/datasets/Daniel-P-Gonzalez/CCOpenBooks.based-chat-v0.1-Mistral-Nemo-Base-2407
Based-Chat v0.1 (Mistral Nemo Base 2407)
This dataset was developed as part of an exploration into understanding the necessity of supervised datasets for fine-tuning base LLMs into conversational models.
It's a synthetic dataset created with Mistral-Nemo-Base-2407, and used to fine-tune that model, producing relay-v0.1-Mistral-Nemo-2407.
Methodology
This synthetic dataset is generated using the following as conversation starters:
facebook/empathetic_dialogues… See the full description on the dataset page: https://huggingface.co/datasets/danlou/based-chat-v0.1-Mistral-Nemo-Base-2407.castillo
🏰 CASTILLO: Characterizing Response Length Distributions in Large Language Models
The CASTILLO dataset is designed to support research on the variability of response lengths in large language models (LLMs). It provides statistical summaries of output lengths across 13 open-source LLMs evaluated on 7 instruction-following datasets. For each unique ⟨prompt, model⟩ pair, 10 independent responses were generated using fixed decoding parameters, and key statistics were recorded—such as… See the full description on the dataset page: https://huggingface.co/datasets/danfperam/castillo.OCD
Dataset Card for Only Clean Data (OCD)
If you are training base language models and want the cleanest sources available, OCD was built just for you.
Dataset Details
Dataset Description
It is without question that the quality of a language model rests on the quality of its training data. OCD is a meticulously curated and cleaned corpus of text documents, ensuring the highest quality text from a variety of sources. Part of this process includes manually… See the full description on the dataset page: https://huggingface.co/datasets/Daniel-P-Gonzalez/OCD.HundredCVs
百人简历数据集
HundredCVs: A Curriculum Vitae Dataset of 100 Young Chinese People
简介
本项目提出一个全新的中文简历数据集(HundredCVs),包含了 100 位青年的个人简历。HundredCVs 具有以下特点:
年轻化、多样性:数据集中的人物年龄分布在 15~30 岁之间,广泛涵盖了不同性别、不同职业、不同学历(高中至博士不等)。
结构完整:每份简历中的信息包括人物的个人名片、性格特征、主要事迹,以及详细经历/个人自述等。
安全性:我们使用化名替代了人物的真实姓名,此外,人物经历也使用大语言模型的改写和提炼,表现出标准化和一致性的语言风格。
设计意图
HundredCVs 的提出主要是为了方便研究者开展基于简历的自然语言处理任务。数据集中提供两个文件:
profile.json:每条记录只包含人物的个人名片、性格特征和主要事迹。可用于开展角色扮演、人物画像构建等任务。… See the full description on the dataset page: https://huggingface.co/datasets/Jax-dan/HundredCVs.localagent-dispatch-data
LocalAgent Dispatch Data
Synthetic data for training/evaluating a generable tool-dispatch model over a 50-tool surface
(route head → dense selector → pointer-copy). A static snapshot of the deterministic generators in
LocalAgent (src/localagent/data/). Train/eval are
disjoint in both phrasing and slot values. Companion model + demo:
danelcsb/localagent-tiny-30m-byte ·
Space.
Configs
config
rows (train/eval)
what it is
paraphrase
1000 / 1000
many natural… See the full description on the dataset page: https://huggingface.co/datasets/danelcsb/localagent-dispatch-data.Alpaca_Evol_Instruct_CleanedAlpaca Evol Instruct cleaned of refusals, scrubbed of overly repetitive responses, aggresively deduplicated, and all URLs removed from the output. The final dataset has aproximately 54k instructions.
Base dataset https://huggingface.co/datasets/victor123/evol_instruct_70k
taboo-dance
taboo-dance
This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT).
Usage
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("bcywinski/taboo-dance")
Format
The dataset is in JSONL format where each line contains a conversation record suitable for training chat models.
dfm11-danish-template-instantiator-training
dfm11-danish-template-instantiator-training
Independently accepted Danish template-instantiator supervision produced by the DFM-owned FineInstructions reproduction pipeline.
Rows retain generation and audit provenance. Local filesystem paths are removed.
The synthetic release does not broaden rights attached to upstream grounding
or query sources; consult each row's source provenance and upstream terms.
Lite-Thinking
Lite-Thinking: A Large-Scale Math Dataset with Moderate Reasoning Steps
Motivation
With the rapid popularization of large reasoning models, like GPT-4o, Deepseek-R1, and Qwen3, there are increasing researchers seeking to build their own reasoning models.
Typically, small foundation models are chosen; following the mature technology of Deepseek-R1, mathematical datasets are mainly adopted to build training corpora.
Despite existing available datasets, represented by… See the full description on the dataset page: https://huggingface.co/datasets/Jax-dan/Lite-Thinking.dfm11-danish-query-templatizer-training
dfm11-danish-query-templatizer-training
Danish query-templatizer supervision produced by the DFM-owned FineInstructions reproduction pipeline.
Rows retain generation and audit provenance. Local filesystem paths are removed.
The synthetic release does not broaden rights attached to upstream grounding
or query sources; consult each row's source provenance and upstream terms.
daniel-os-profile-sft
Daniel OS Profile SFT and Behavior Tests
Small, source-grounded datasets used to adapt and evaluate the browser-native
Daniel OS portfolio assistant. The model separates Daniel-specific claims from
general definitions, synthesizes definitions from retrieved evidence, requests
public retrieval when evidence is absent, and declines private-person requests.
Splits
Configuration
Split
Records
Purpose
sft
train
268
Profile-grounded conversational fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/danelcsb/daniel-os-profile-sft.reasoning-corpus-4K-5M-v1 Reasoning Corpus 5M · Within 5k sequence length
About Dataset
This dataset contains reasoning chains from major AI models, such as: DeepSeek-v4 (both Pro and Flash), DeepSeek-r1 (DS-r1, Llama-DS, Qwen-DS), Qwen3, Qwen3.5/3.6 (both OpenSource and API models), Gemma4-31B derived from many other repositories, and properly filtered to train SLMs.
The dataset has these columns for users to filter out:
repo_id
tok_len
user
thought_trace
assistant
ChatML
Repositories… See the full description on the dataset page: https://huggingface.co/datasets/danie1111/reasoning-corpus-4K-5M-v1.spicyfictionanima-corpus-ko-fineweb2-broad
anima-corpus-ko-fineweb2-broad
🇰🇷 한국어 broad (일반) 코퍼스 for anima conv 303M byte-level pretraining.
anima chat register 표준(a_chat_registers)의 4칸 {ko·en} × {일반·SNS} 중 ko-일반 칸을 메우기 위한 데이터셋이다. 기존 ko-일반 source 가 ~1.7MB 로 빈약(en ~202MB 대비)했던 갭을 FineWeb-2 한국어로 보강한다.
Source
Upstream: HuggingFaceFW/fineweb-2, config kor_Hang (한국어 한글 스크립트).
Extracted from: train parquet data/kor_Hang/train/000_00000.parquet (1 of 25 shards).
Field: text 컬럼만 추출 (raw UTF-8, byte-vocab256… See the full description on the dataset page: https://huggingface.co/datasets/dancinlab/anima-corpus-ko-fineweb2-broad.gutenberg-top-50
Gutenberg Top-50: 50 popular books on Gutenberg
Key Features
Source: This corpus is collected from Gutenberg, which is a public e-book library.
Components: It contains 50 popular books that easily split into chapters containing paragraphs.
Applications:
Researchers can use this corpus to train their own pre-trained language model.
Researchers can use this corpus to ask questions that depend on the certain chapter/paragraphs content, for QA or RAG development.
rest-v3
rest-v3
rest-v3 is an English text-rewriting dataset for supervised fine-tuning of a humanizing editor. Each record asks a model to rewrite a source text while preserving its meaning and contains a detector-verified natural-language rewrite.
Dataset composition
The training split contains 1,116 JSONL records:
1,033 newly mined, on-policy rewrites from the from-final-best generator checkpoint.
83 compatible existing verified examples.
541 examples sourced from the… See the full description on the dataset page: https://huggingface.co/datasets/danilxyz/rest-v3.danbooru_tag_corpusA filtered ~232k danbooru dataset with categorized tags.
Filtering criteria:
score > 100
created within the last 5 years
no video, ai-generated and tagme tags
Structure:
{
"post_id": 0,
"tags": {
"general": [],
"artist": [],
"copyright": [],
"character": [],
"meta": [],
"unknown": []
},
"rating": "",
"created_at": "yyyy-MM-dd hh:mm:ss",
"score": 0
}
Each tag list is sorted by relevance (post count per tag), you may need… See the full description on the dataset page: https://huggingface.co/datasets/SuccubusBot/danbooru_tag_corpus.ai-arenaen-conversations
AI Arenaen Conversations
A large dataset of conversations from AI-Arenaen, the Danish subset of the compar:IA platform.
Origin of the data: what is AI-Arenaen?
The conversations are collected using AI-Arenaen, the Danish entry point to the compar:IA platform, which is a Conversational AI comparison tool (a "chatbot arena"), developed within the French Ministry of Culture and adapted for Danish users by Danish Foundation Models and The ministry of digital affair.… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/ai-arenaen-conversations.
