datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TheArabicPile_Conversational
The Arabic Pile
Introduction:
The Arabic Pile is a comprehensive dataset meticulously designed to parallel the structure of The Pile and The Nordic Pile. Focused on the Arabic language, the dataset encompasses a vast array of linguistic nuances, incorporating both Modern Standard Arabic (MSA) and various Levantine, North African, and Egyptian dialects. Tailored for the training and fine-tuning of large language models, the dataset consists of 13 subsets, each uniquely… See the full description on the dataset page: https://huggingface.co/datasets/premio-ai/TheArabicPile_Conversational.malay-conversational-speech-corpus
malay-conversational-speech-corpus
Mirror for https://magichub.com/datasets/malay-conversational-speech-corpus/, license is Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License
mega-asr-conversational-overlap
Mega-ASR Conversational Overlap
Mega-ASR Conversational Overlap is a deterministic English ASR diagnostic set
derived from AirCaps/mega-asr-noise-a5sv2,
which in turn is sampled from the Mega-ASR training corpus
zhifeixie/Voices-in-the-Wild-2M.
The existing AirCaps dataset evaluates single-utterance acoustic robustness.
This companion dataset evaluates a different failure mode: two-turn conversational
continuity with slight overlap and unequal turn loudness. It does not replace… See the full description on the dataset page: https://huggingface.co/datasets/AirCaps/mega-asr-conversational-overlap.conversational_englishluel-conversational-speech-samples
Conversational Speech Samples (Luel)
License: All Rights Reserved. Proprietary. Access only for authorized parties; no redistribution or use without permission. See LICENSE.
A multilingual conversational speech dataset: two-speaker dialogue sessions across 8 languages. Each session includes combined audio, speaker-separated audio, and a JSON file containing transcripts with word-level timestamps, speaker diarization, and acoustic metadata.
Quick Stats… See the full description on the dataset page: https://huggingface.co/datasets/Luel-ai/luel-conversational-speech-samples.conversational-hindi-largePidgin-to-English-conversational-translations
Pidgin-to-English Translation Dataset (Sample)
Sample dataset: Nigerian Pidgin to English translation pairs for machine translation research
🤗 Hugging Face • 📊 Figshare • 🌐 Website • 📧 Contact
📋 Overview
The Pidgin-to-English Translation Dataset (Sample) is a conversational-style parallel corpus containing 122 translation pairs from Nigerian Pidgin English to Standard English. Created by Bytte AI through AI chatbot interactions with human validation… See the full description on the dataset page: https://huggingface.co/datasets/Bytte-AI/Pidgin-to-English-conversational-translations.malay-conversational-speech-corpus-whisper-formatconversational_hindiprop-trading-qa-conversational-ai
Prop Trading Q&A Dataset for Conversational AI
Description
This dataset contains 200+ curated question-answer pairs covering the domain of proprietary (prop) trading firms. It is designed to serve as training and retrieval data for building AI assistants, chatbots, and educational tools focused on prop trading knowledge.
Each entry consists of a natural-language question paired with a detailed, factual answer. The data spans ten thematic categories ranging from… See the full description on the dataset page: https://huggingface.co/datasets/propfirmkey/prop-trading-qa-conversational-ai.sft-conversational_datasetQuestion – Answer DatasetThe dataset contains 400 queries from two domains: Current Affairs and Creative Writing. It serves as a versatile resource for Natural Language Processing (NLP) tasks, including text classification, information retrieval, and model training.
Data attributes:
Query: The user-generated question. Data type: string.
Answer: The response provided by a team of writers and editors in markdown format, containing information related to the query.
Citations: Up to 4 credible… See the full description on the dataset page: https://huggingface.co/datasets/SoftAge-AI/sft-conversational_dataset.conversational_ai_turn_2_checkpointconversational_ai_turn_3_checkpointconversational_ai_turn_4_checkpointconversational_aiconversational_ai_5_turns_only_ckp_1conversational_ai_5_turns_only_ckp_2conversational_ai_5_turns_only_ckp_3conversational_ai_5_turns_only_ckp_4conversational_ai_5_turns_onlyconversational_ai_turn_0_checkpointconversational_ai_turn_1_checkpointconversational_ai_5_turns_only_ckp_0conversational-ai-training
conversational-ai-training
This dataset was created using the Claude Dataset Skill.
