datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lmsys-chat-1m-formattedlmsys-chat-1m-chat-formattedchat-formattedt5gemma2-indonesia-chat-formatted
T5Gemma-2 Indonesian Chat & QA Dataset
A high-quality Indonesian language multi-turn conversation and reading comprehension dataset, specifically formatted for instruction tuning of sequence-to-sequence (Seq2Seq) models like T5-Gemma / T5-Gemma-2.
Dataset Description
This dataset contains over 7,400 multi-turn conversations and document-based Q&A in Bahasa Indonesia. It covers diverse topics including everyday life, technology, general knowledge, and structured… See the full description on the dataset page: https://huggingface.co/datasets/daruokta/t5gemma2-indonesia-chat-formatted.llama3_additional_rr40k_non_delete_sft_chat_formathh-rlhf_chat_formattuluv2-expanded-150k-part_v0-chat-format-syn-knowledgeSynthetic-Japanese-Roleplay-NSFW-gpt-5-chat-5k-formatted
Synthetic-Japanese-Roleplay-NSFW-gpt-5-chat-5k-formatted
概要
OpenRouterのgpt-5-chatを用いて作成した日本語ロールプレイデータセットであるAratako/Synthetic-Japanese-Roleplay-NSFW-gpt-5-chat-5kをOpenAI messages形式に整形したデータセットです。
データの詳細については元データセットのREADMEを参照してください。
ライセンス
CC-BY-NC-SA 4.0の元配布します。
また、OpenAIの利用規約に記載のある通り、このデータを使ってOpenAIのサービスやモデルと競合するようなモデルを開発することは禁止されています。
oasst2_thai_top1_chat_format
Open Assistant 2 Top-1 Thai
Dataset Details
Dataset Description
A top-1 Thai dataset taken from the top scoring https://huggingface.co/datasets/OpenAssistant/oasst2 conversations. Saved in HF Chat format.
License: Apache 2.0
Script: https://github.com/wannaphong/deep_4_all/tree/main/datasets/oasst
Dataset Structure
We structure the dataset using the format commonly used as input into Hugging Face Chat Templates:
[
{'content':… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/oasst2_thai_top1_chat_format.vi-self-chat-sharegpt-format
🇻🇳 Vietnamese Self-Chat Dataset
This dataset is designed to enhance the model's ability to engage in multi-turn conversations with humans.
To construct this dataset, we follow a two-step process:
Step 1: Instruction Generation
We employ the methodology outlined in the Self-Instruct paper to craft a diverse set of instructions. This paper serves as a guide for aligning pretrained language models with specific instructions, providing a structured foundation for subsequent dialogue… See the full description on the dataset page: https://huggingface.co/datasets/bkai-foundation-models/vi-self-chat-sharegpt-format.orca-chat-chat-format
Dataset Card for "orca-chat-chat-format"
More Information needed
tokenized-chat-ml-format-vicgalle-alpaca-gpt4-1kllama3_regular_balanced_sft_chat_formatsamvaad-hi-v1-chat-formatairoboros-3.1-no-mathjson-max-1k-chat-formatCopy of habanoz/airoboros-3.1-no-mathjson-max-1k transformed to work with huggingface chat templates e.g. role(user|assistant), content.
Note that samples are limited to 1K length.
oasst2_top1_chat_format
OpenAssistant TOP-1 Conversation Threads in huggingface chat format
Export of oasst2 only top 1 threads in huggingface chat format
Script
The convert script can be find here
chat-formatted-magicoder
Dataset Card for "chat-formatted-magicoder"
More Information needed
MATH_chat_formatoasst_top1_2023-08-25-chat-format
Dataset Card for "oasst_top1_2023-08-25-chat-format"
This dataset is a copy of OpenAssistant/oasst_top1_2023-08-25 transformed to work with hugging face chat templates.
baseline-documents-chat-formattedchat-formatted-metamathHelpSteer3-general-code-Shift-Qwen-2.5-1.5B-Instruct-chat-formatted-generationsMATH_full_chat_formatsemiconductor_chat_v3.2_filtered_formatedgus-dataset-chat-formatindian-legal-summaries-chat-formatpodcast_llama_chat_format
Intro
This dataset formats an existing podcast dataset (64bits/lex_fridman_podcast_for_llm_vicuna) for llama 3 chat model fine tuning.
It represents a compilation of audio-to-text transcripts from the Lex Fridman Podcast. The Lex Fridman Podcast, hosted by AI researcher at MIT, Lex Fridman.
Problems
There might be some minor issues during the transcribe phase.
Next Step
Use whisper to directly load the podcast and transcribe it in this format.
pile-lmsys-mix-1m-Llama3.2_chat_formatopenassistant-guanaco-chat-format
Dataset Card for "openassistant-guanaco-chat-format"
Copy of timdettmers/openassistant-guanaco. Modified to use with hugging face chat tempaltes.
wizard-vicuna-70k-unfiltered_chat_format
