datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
t5gemma2-indonesia-chat-formatted
T5Gemma-2 Indonesian Chat & QA Dataset
A high-quality Indonesian language multi-turn conversation and reading comprehension dataset, specifically formatted for instruction tuning of sequence-to-sequence (Seq2Seq) models like T5-Gemma / T5-Gemma-2.
Dataset Description
This dataset contains over 7,400 multi-turn conversations and document-based Q&A in Bahasa Indonesia. It covers diverse topics including everyday life, technology, general knowledge, and structured… See the full description on the dataset page: https://huggingface.co/datasets/daruokta/t5gemma2-indonesia-chat-formatted.reasoning-and-chat-harmony-format
Open Paws Reasoning And Conversational Finetuning Harmony Format
This dataset is part of the Open Paws initiative to develop AI training data aligned with animal liberation and advocacy principles. Created to train AI systems that understand and promote animal welfare, rights, and liberation.
Dataset Details
Dataset Type: Reasoning Data
Format: JSONL (JSON Lines)
Languages: Multilingual (primarily English)
Focus: Animal advocacy and ethical reasoning
Organization: Open… See the full description on the dataset page: https://huggingface.co/datasets/open-paws/reasoning-and-chat-harmony-format.Synthetic-Japanese-Roleplay-NSFW-gpt-5-chat-5k-formatted
Synthetic-Japanese-Roleplay-NSFW-gpt-5-chat-5k-formatted
概要
OpenRouterのgpt-5-chatを用いて作成した日本語ロールプレイデータセットであるAratako/Synthetic-Japanese-Roleplay-NSFW-gpt-5-chat-5kをOpenAI messages形式に整形したデータセットです。
データの詳細については元データセットのREADMEを参照してください。
ライセンス
CC-BY-NC-SA 4.0の元配布します。
また、OpenAIの利用規約に記載のある通り、このデータを使ってOpenAIのサービスやモデルと競合するようなモデルを開発することは禁止されています。
podcast_llama_chat_format
Intro
This dataset formats an existing podcast dataset (64bits/lex_fridman_podcast_for_llm_vicuna) for llama 3 chat model fine tuning.
It represents a compilation of audio-to-text transcripts from the Lex Fridman Podcast. The Lex Fridman Podcast, hosted by AI researcher at MIT, Lex Fridman.
Problems
There might be some minor issues during the transcribe phase.
Next Step
Use whisper to directly load the podcast and transcribe it in this format.
oasst2_top1_chat_formatOpen Assistant Conversations Dataset Release 2 (OASST2)
source: https://huggingface.co/datasets/OpenAssistant/oasst2
This dataset made the Top1 chat format from the train subset in all languages from OpenAssistant.
podcast_llama_chat_format-1k
Intro
This dataset(1K) formats an existing podcast dataset (64bits/lex_fridman_podcast_for_llm_vicuna) for llama 3 chat model fine tuning.
It represents a compilation of audio-to-text transcripts from the Lex Fridman Podcast. The Lex Fridman Podcast, hosted by AI researcher at MIT, Lex Fridman.
Problems
There might be some minor issues during the transcribe phase.
Next Step
Use whisper to directly load the podcast and transcribe it in this format.
