datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Knowledge-QA-SingleTurn-Dataset
Knowledge QA Single-turn Dataset(知識質問データセット・シングルターン)
概要
本データセットは、Aratako/Synthetic-JP-Conversations-Magpie-Nemotron-4-10k から質問を抽出し、DeepSeek V3.2で整形、Kimi K2.5で回答を生成した シングルターンの知識質問応答データセット です。Reasoning有効化により思考過程も最終データに含まれ、質問の難易度に応じてReasoning effortが動的に切り替わります。
生成にはSDG-LOOMという合成データ生成パイプラインを用いました。(sdg-loom)
データの説明
項目
内容
件数
約7,000件
形式
JSONL(1行1JSON)
言語
日本語
ターン数
1ターン(質問1 + 回答1)
ソースデータセット… See the full description on the dataset page: https://huggingface.co/datasets/DataPilot/Knowledge-QA-SingleTurn-Dataset.hh-rlhf-49k-ja-single-turnThis dataset was created by automatically translating part of "Anthropic/hh-rlhf" into Japanese, and selected for single turn conversations.You can use this dataset for RLHF and DPO.
hh-rlhf repository
https://github.com/anthropics/hh-rlhf
Anthropic/hh-rlhf
https://huggingface.co/datasets/Anthropic/hh-rlhf
tool-calls-singleturnsingle_turn
InterSyn: A High-Quality Dataset and Reliable Evaluation for Interleaved Image-Text Generation
This dataset card accompanies the paper
A High-Quality Dataset and Reliable Evaluation for Interleaved Image-Text GenerationYukang Feng, Jianwen Sun, Chuanhao Li, Zizhen Li, Jiaxin Ai, Fanrui Zhang, Yifan Chang, Sizhuo Zhou, Shenglin Zhang, Yu Dai, Kaipeng Zhang (2025)
🧠 Introduction
TL;DR InterSyn is a high-quality dataset for instruction‑following, interleaved image–text… See the full description on the dataset page: https://huggingface.co/datasets/finyorko/single_turn.crag-mm-single-turn-public
CRAG-MM: Comprehensive multi-modal, multi-turn RAG Benchmark
This repository contains the CRAG-MM dataset, a high-quality conversational benchmark for multimodal assistants. The dataset features conversations about images with varied complexity levels, designed to evaluate AI systems' visual understanding and conversational abilities.
CRAG-MM is a visual question-answering benchmark that focuses on factual questions, offering a unique collection of image and question-answering sets… See the full description on the dataset page: https://huggingface.co/datasets/crag-mm-2025/crag-mm-single-turn-public.function_calling_eval_single_turn_v0qwen3_4b_instruct_lcbv6_single_turn_s_65_e_131tiny-singleturn-chat-kosingle-turn-compilationChinese-Roleplay-SingleTurn请注意,个人模型经过characterEval的reward model进行DPO训练,因此使用本数据集进行SFT的模型在该榜单上会存在bias,导致分数异常偏高,请勿直接使用该榜单进行测试
简介
因已找到更优数据合成方案,为填充中文角色扮演数据集的空白,现开源部分中文角色扮演单轮对话数据集。
使用Refined-Anime-Text作为system prompt,使用小黄鸡随机query作为输入,调用个人角色扮演模型作为输出。
已处理为alpaca数据格式,方便大家处理和训练。经过验证,仅使用该数据集进行Lora微调即可获取一个效果还不错的模型~
chatGPT对比
character
question
answer_us
answer_chatGPT
黑须彼方是(省略……)黑须彼方有着许多有趣的爱好和特点。她是一个有点毒舌的人,但总能犀利地指出问题所在。她有着敏锐的洞察力,擅长看透人心。她经常以此来捉弄加贺正午。她与正午有着相同的口癖,张扬的性格(省略……)她的个性和爱好使她成为一个备受喜爱的角色。… See the full description on the dataset page: https://huggingface.co/datasets/LooksJuicy/Chinese-Roleplay-SingleTurn.crag-mm-single-turn-public-v0.1.2-with-imagessynthetic-swift-data-single-turn
Dataset Card for synthetic-swift-data-single-turn
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/vinhnx90/synthetic-swift-data-single-turn/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/vinhnx90/synthetic-swift-data-single-turn.llava-med-zh-instruct-singleturncrag-mm-single-turn-public-v0.1.2-with-images-gpt4.1-v2tulu-3-sft-single-turn-gpt4o-mini-thoughts-original-responsescrag-mm-single-turn-debug-public
CRAG-MM: Comprehensive multi-modal, multi-turn RAG Benchmark
This repository contains the CRAG-MM dataset, a high-quality conversational benchmark for multimodal assistants. The dataset features conversations about images with varied complexity levels, designed to evaluate AI systems' visual understanding and conversational abilities.
CRAG-MM is a visual question-answering benchmark that focuses on factual questions, offering a unique collection of image and question-answering sets… See the full description on the dataset page: https://huggingface.co/datasets/crag-mm-2025/crag-mm-single-turn-debug-public.MIRAGE-Standard-single-turn-with-meteo-pathssingle-turn-compilation-SmolLM2-1024claude_3.5s_single_turn_unslop_filteredcrag-mm-single-turn-public-v0.1.2json-mode-singleturnFineTome-single-turn-dedup-amharic
Dataset Card for FineTome Single Turn Conversations - Amharic
This dataset contains 83,290 conversational examples translated from English to Amharic, providing high-quality instruction-following conversations for training language models in Amharic.
Dataset Details
Dataset Description
This dataset is a translation of the FineTome-single-turn-dedup dataset into Amharic, creating one of the largest publicly available collection of… See the full description on the dataset page: https://huggingface.co/datasets/addisai/FineTome-single-turn-dedup-amharic.lcb_feb_may_2025_single_turn_0_131hermis_singleTurnNepali_educationalwikirace-v5-single-turntulu-3-sft-single-turnfunc-calling-eval-singleturnbfcl_v4_single_turntyphoon-s-instruct-sft-single-turnFineTome-single-turn-dedup
FineTome-single-turn-dedup
This dataset is a transformed version of mlabonne/FineTome-100k, created by applying the following steps:
Extracted the first turn of each conversation: optional system message + user message + assistant message.
Performed deduplication: using MinHash to remove duplicate entries.
Converted the format from ShareGPT to OpenAI
