datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
InternVL-Chat-V1-2-SFT-Data
Data Card for InternVL-Chat-V1-2-SFT-Data
Overview
Inspired by LLaVA-NeXT, we adopted a data-efficient SFT strategy to train InternVL-Chat-V1-2, utilizing approximately 1.2M of visual instruction tuning samples in total, all of which are fully open-source. In a macro sense, we build upon ShareGPT-4V and additionally integrate LLaVA-ZH, DVQA, ChartQA, AI2D, DocVQA, GeoQA+, and SynthDoG-EN. Most of the data remains consistent with LLaVA-NeXT.
Citation
If you use… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/InternVL-Chat-V1-2-SFT-Data.ChatTS-Training-Dataset
ChatTS-Training Data
This repository contains the training data for the ChatTS project. This is the dataset for training the ChatTS-14B model.
Datasets
align_256: Alignment training dataset for stage-1 alignment training, with SEQ_LEN=256.
align_random: Alignment training dataset with random sequence lengths between 64 and 1024.
sft: SFT dataset generated with Time Series Evol-Instruct.
ift: Instruction following dataset.
dev: A small dataset for development and testing.… See the full description on the dataset page: https://huggingface.co/datasets/ChatTSRepo/ChatTS-Training-Dataset.NSFW_Chat_Dataset
💕 Spicy AI GF Chat Dataset 🔥
🚨 18+ Only! NSFW & Spicy Content Ahead 🚨
Hey there, AI enthusiasts and romance lovers! 😏 Welcome to the Spicy AI GF Chat Dataset, the ultimate dataset designed to bring your AI waifu to life! 💖 If you've ever dreamed of building an AI that responds like your virtual girlfriend, THIS is the dataset for you.
📜 What’s Inside?
This dataset features two columns:
input → Boyfriend’s dialogue (aka what YOU say 😉)
output →… See the full description on the dataset page: https://huggingface.co/datasets/utsavm/NSFW_Chat_Dataset.chat_doctorThis dataset was formed from the three data sources from the ChatDoctor work.
100k real conversations between patients and doctors from HealthCareMagic.com HealthCareMagic-100k. - ADDED
10k real conversations between patients and doctors from icliniq.com icliniq-10k. - ADDED
5k generated conversations between patients and physicians from ChatGPT GenMedGPT-5k and disease database. - NOT ADDED (because of the data created by LLM, but you could add it manually)
data sample:
{'instruction': "If… See the full description on the dataset page: https://huggingface.co/datasets/avaliev/chat_doctor.Vietnamese-Multi-turn-Chat-Alpacachinese_chatgpt_corpus
Dataset Card for chinese_chatgpt_corpus
Dataset Summary
This repo collects chinese corpus for Supervised Finetuning (SFT) and Reinforcement Learning From Human Feedback (RLHF).
Supported Tasks and Leaderboards
More Information Needed
Languages
Chinese
Dataset Structure
Data Instances
train_data_external_v1.jsonl
Size of downloaded dataset files: 5.04 GB
Size of the generated dataset: 0 GB… See the full description on the dataset page: https://huggingface.co/datasets/sunzeyeah/chinese_chatgpt_corpus.cukurova_university_chatbot
Çukurova University Computer Engineering Chatbot Dataset
📊 Dataset Overview
This dataset contains 22,524 high-quality question-answer pairs specifically designed for training an AI chatbot that serves the Computer Engineering Department at Çukurova University. The dataset is part of the CengBot project, a sophisticated multilingual Telegram chatbot that provides automated assistance to students regarding courses, programs, and departmental information.
🔢… See the full description on the dataset page: https://huggingface.co/datasets/Naholav/cukurova_university_chatbot.Java_method2test_chatml
Java Method to Test ChatML
This dataset is based on the methods2test dataset from Microsoft. It follows the ChatML template format: [{'role': '', 'content': ''}, {...}].
Originally, methods2test contains only Java methods at different levels of granularity along with their corresponding test cases. The different focal method segmentations are illustrated here:
To simulate a conversation between a Java developer and an AI assistant, I introduce two key parameters:
The prompt… See the full description on the dataset page: https://huggingface.co/datasets/random-long-int/Java_method2test_chatml.SE-Chatting.en
SE.02
Dataset
Hello, welcome to the official main dataset of SE.02 that's always getting updated, make sure to like to help us a lot.
this dataset contains pretty much everything from math to isk what to put but pretty much anything you think an ai can say.
anyways this is our biggest dataset yet, and out first one that I don't even know how I managed to make this.
you can use it to train your own ai if you want.
instruct_chat_50k.jsonlinstruct_chat_50k.jsonl which is composed of 30k Chinese sharegpt dataset and 20k alpaca-instruction-Chinese-dataset
ChatEnv
ChatEnv: A Domain-Specific Instruction Dataset for Environmental Science
ChatEnv is a large-scale, domain-specific instruction dataset designed to enhance large language models (LLMs) for environmental science tasks. This dataset is an integral part of the EnvGPT framework, supporting fine-tuning and evaluation processes by providing a diverse and high-quality set of instructions tailored to the unique demands of environmental science research and applications.
📃 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/SustcZhangYX/ChatEnv.Chatkanoonturkish-law-chatbot
MindLaw için Hukuk Veri Seti
Bu veri seti, MindLaw modelinin eğitimi için oluşturulmuş olup, Türkçe hukuk alanına özgü metinlerden derlenmiştir. Veri seti, anayasanın sunduğu içeriklerden ve anayasayı açıklayan hukuki metinlerden oluşmaktadır. Ayrıca, bireylerin avukatlara sıkça yönlendirebilecekleri sorular formatında düzenlenmiş hukuki sorular ve cevapları da içermektedir.
Veri Seti İçeriği
Anayasa Metinleri: Türkiye Cumhuriyeti Anayasası'nın çeşitli maddeleri ve… See the full description on the dataset page: https://huggingface.co/datasets/Renicames/turkish-law-chatbot.LogicMind-Chat-Reasoning-SFT-300K
Nemotron-Post-Training-Dataset-v2-chat Dataset Card
Overview 📌
This dataset contains 296,168 chat-style instruction/response samples generated by qwen-3-32b. Each record provides a user prompt, an explicit reasoning trace, and a final answer, plus precomputed length fields. The data is packaged as JSONL (one JSON object per line).
Highlights
Scale: 296,168 samples
Category: chat (100%)
Generator: qwen-3-32b (100%)
Structure: problem → qwen3-reasoning → qwen3-solution… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/LogicMind-Chat-Reasoning-SFT-300K.arabic-rag-chat-8k-eval
arabic-rag-chat-8k-eval
Per-row evaluation artifacts for the 8,192-token Arabic multi-turn RAG models:
the test split, every model's raw replies, every judge verdict, and the rendered
report for each. Thirteen judged models, all scored on the same 1,651 prompts
by the same judge at temperature 0.0, so the comparison below is like-for-like
and can be recomputed offline without a GPU or a judge server.
This is the measurement half of
oddadmix/100M-8192-Nawah-dsv4;
the training… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-rag-chat-8k-eval.chatgpt4-commonsense-qa
Synthetic CommonSense
Generated using ChatGPT4, originally from https://huggingface.co/datasets/commonsense_qa
Notebook at https://github.com/mesolitica/malaysian-dataset/tree/master/question-answer/chatgpt4-commonsense
synthetic-commonsense.jsonl, 36332 rows, 7.34 MB.
Example data
{'question': '1. Seseorang yang bersara mungkin perlu kembali bekerja jika mereka apa?\n A. mempunyai hutang\n B. mencari pendapatan\n C. meninggalkan pekerjaan\n D. memerlukan… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/chatgpt4-commonsense-qa.chatgpt-in-russia-qa
Dataset Card for чатгпт-в-россии.рф
Dataset Summary
This dataset contains question-answer pairs collected from чатгпт-в-россии.рф (meaning in English would be something like chatgpt-in-russia[.]rf), a Russian question-answering website. Each entry in the dataset represents a question asked by a user and the corresponding answer generated by an unspecified language model. The dataset contains 704,208 unique question-answer pairs covering various topics.… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/chatgpt-in-russia-qa.allaM-offsec-arabic-chat-v2
Arabic Offensive Security Chat Dataset v2
High-quality category-aware bilingual Arabic/English dataset for offensive security assistants.
What's New in v2
✅ Category-aware responses: Different response structures for web vulns, DeFi, reconnaissance tools, social engineering, etc.
✅ No generic templates: Each category has specialized analysis framework
✅ No verbatim copying: Responses analyze and transform the input, not repeat it
✅ Semantic accuracy: Tools (nmap… See the full description on the dataset page: https://huggingface.co/datasets/haiderkamal23/allaM-offsec-arabic-chat-v2.Qwen3-235B-A22B-Instruct-2507-Distilled-chat
Qwen3-235B-A22B-Instruct-2507-Distilled-chat📚
Curated/Funded/Shared by: [Jack Rong]
Language(s): English (major), Chinese, Русский, 한국어, 日本語, others
License: [apache-2.0]
Distilled Model: 🏆Qwen/Qwen3-235B-A22B-Instruct-2507
Qwen3-235B-A22B-Instruct-2507 Benchmarks📊
Introduction:
The objectives of this project are:
Focus on chat capabilities (excluding CoT), covering cross-lingual real-world Q&A/explanation/generation;
Utilize… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Qwen3-235B-A22B-Instruct-2507-Distilled-chat.arabic-rag-chat-30K
Arabic multi-turn RAG customer-support conversations (31,294 conversations)
Synthetic Modern Standard Arabic customer-support conversations for training
small Arabic RAG assistants. The bulk was distilled from gemini-3.1-flash-lite
via the Batch API; a first 2.4% came from unsloth/gemma-4-31B-it-NVFP4
on a local vLLM server before the run was moved off-GPU. Both teachers were given
the same prompts and the same validator. Each row is one conversation of 1-5
rounds over one… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-rag-chat-30K.tw-bar-examination-2020-chat
Dataset Card for tw-bar-examination-2020-chat
tw-bar-examination-2020-chat 是一個中華民國 2020 年律師考試選擇題之 Alpaca 格式微調資料集,合計 299 題(train 269、test 30)。每題包含統一提示語、題目與四個選項,以及正確答案字母,適用於微調繁體中文語言模型於台灣法律選擇題作答任務。
Dataset Details
Dataset Description
本資料集源自 Jamie0510/taiwan-law-exam 中之 2020 年律師考試題目,整合其四大類科後進行後處理:去除欄位缺失之題目,並統一轉為 Alpaca 三欄格式(instruction / input / output)。每題之 instruction 欄為固定提示語「請在下列的單一選擇題中,選出正確的答案,並且只回答 A, B, C, D 其中一個字代表正確答案」。
本資料集作為 SFT 訓練素材設計,建議與… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-bar-examination-2020-chat.Chinese-DeepSeek-V3.2-Exp-chat-example
deepseek/deepseek-v3.2-exp (6.6K) 中文数据集样本
一、前言
本报告基于 deepseek/deepseek-v3.2-exp 模型(官方 API,8K 上下文窗口)进行数据集评测与可视化展示。测试数据集共包含 6,655 轮对话,语言覆盖以中文为主,辅以部分混合语种及非中文输入。本次报告旨在总结模型的对话特征、输入输出长度分布及上下文预算消耗情况,并为后续应用和优化提供参考。
二、数据与方法
数据来源:用户构建的 6,655 轮真实中文对话样本。
估算方法:
中文字符近似为 1 Token;
英文 4 字符 ≈ 1 Token;
用于规模与上下文预算对比,而非精确 Token 计数。
统计维度:
平均 Prompt/Output 长度(字符与估算 Token);
总 Token 占上下文窗口比例;
语言分布(Prompt 语言类型);
对话长度分布(用户提问、助手回答、总对话长度)。
三、总体结果
1. 样本概况… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Chinese-DeepSeek-V3.2-Exp-chat-example.asknyc-chatassistant-formatQuestions from Reddit.com/r/AskNYC, downloaded from PushShift, filtered to direct responses from humans, where the post net score is >= 3.
Collected one month of posts from each year 2015-2019 (i.e. no content from July 2019 onward)
Adapted from the CSV used to fine-tune https://huggingface.co/monsoon-nlp/gpt-nyc
Blog about the original model: https://medium.com/geekculture/gpt-nyc-part-1-9cb698b2e3d
big-red-bark-chat-evaluation
Big Red Bark Chat Q&A Dataset
Dataset Description
This dataset contains 12,385 question-and-answer pairs collected from Big Red Bark Chat, an innovative AI assistant developed at Cornell University that answers questions about dog health (as well as other animal species). While it does not replace professional veterinary advice, it serves as a valuable starting point by searching trusted sources. Big Red Bark Chat is designed to provide quick and reliable answers… See the full description on the dataset page: https://huggingface.co/datasets/Sr523/big-red-bark-chat-evaluation.Instruction-Tuning-with-GPT-4-RedPajama-Chat
Instruction Tuning with GPT 4 RedPajama-Chat
This dataset has been converted from the Instruction-Tuning-with-GPT-4 dataset for the purpose of fine-tuning the RedPajama-INCITE-Chat-3B-v1 model.
About Instruction-Tuning-with-GPT-4
English Instruction-Following Data generated by GPT-4 using Alpaca prompts for fine-tuning LLMs.
Usage and License Notices
The data is intended and licensed for research use only. The dataset is CC BY NC 4.0 (allowing only… See the full description on the dataset page: https://huggingface.co/datasets/Fredithefish/Instruction-Tuning-with-GPT-4-RedPajama-Chat.legal-chat-sft-dataset
Thai Legal Chat SFT Dataset (CoT & Hybrid RAG)
ชุดข้อมูลสำหรับการทำ Instruction Fine-Tuning (SFT) เพื่อสร้าง AI ผู้ช่วยนักกฎหมายไทยที่มีความสามารถในการคิดวิเคราะห์แบบเป็นขั้นตอน (Chain-of-Thought) และมีความรู้กฎหมายที่ทันสมัยจากการใช้ Hybrid RAG (Retrieval-Augmented Generation)
Dataset Summary
ชุดข้อมูลนี้ถูกสร้างขึ้นแบบสังเคราะห์ (Synthetic Data) โดยใช้โมเดลภาษาขนาดใหญ่ (LLM) ตระกูล Qwen (27B+) บนขุมพลัง AMD MI300X ผ่านระบบ vLLM Monster Engine… See the full description on the dataset page: https://huggingface.co/datasets/Phonsiri/legal-chat-sft-dataset.paper_persi_chat
PaperPersiChat Dataset
Dataset for paper PaperPersiChat: Scientific Paper Discussion Chatbot using Transformers and Discourse Flow Management
Dataset creation
To construct the dataset, we used the part of Semantic Scholar Open Research Corpus [https://github.com/allenai/s2orc] as the main source of scientific publications, namely the Computer Science section. We constructed dialogues over the segments of the papers where each segment consists of a combination of several… See the full description on the dataset page: https://huggingface.co/datasets/ai-forever/paper_persi_chat.Alpha_Chat_Style_Dataset
🦾 Alpha Chat Style Dataset | darkknight25
Inject dominance, charm, and precision into your LLMs.
Crafted by Sunny Thakur, this dataset is designed to train conversational agents that speak like a leader, think like a tactician, and respond like a professional.
“Control the tone. Command the room. Every word should land like a calculated move.” – Alpha Protocol
🎯 Purpose
This dataset enables large language models—like Mixtral 8x7B Instruct—to adopt a bold… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/Alpha_Chat_Style_Dataset.lemonseed-rl-chat-tasks
lemonseed-rl-chat-tasks
LemonSeed — chat-alignment RL tasks (prompt/gold single-turn).
Contents
rl_chat_tasks.jsonl (9852 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Synthetic, generated programmatically for the LemonSeed 1.5B project (by Geramy L. Loveless). Data authored by Michael Anthony Falabella.
arabic-rag-chat-grpo-5K
Arabic multi-turn RAG conversations — GRPO pool (5,259 conversations)
The reinforcement-learning half of
oddadmix/arabic-rag-chat-30K:
same generator, same validator, same schema, disjoint companies. It exists
so GRPO explores fresh knowledge bases instead of taking a second pass over
material the SFT already memorised.
conversations
turns
companies
this pool
5,259
14,018
309
Company-disjointness is exact and verified: this pool shares zero
company_id values… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-rag-chat-grpo-5K.
