datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TheArabicPile_Conversational
The Arabic Pile
Introduction:
The Arabic Pile is a comprehensive dataset meticulously designed to parallel the structure of The Pile and The Nordic Pile. Focused on the Arabic language, the dataset encompasses a vast array of linguistic nuances, incorporating both Modern Standard Arabic (MSA) and various Levantine, North African, and Egyptian dialects. Tailored for the training and fine-tuning of large language models, the dataset consists of 13 subsets, each uniquely… See the full description on the dataset page: https://huggingface.co/datasets/premio-ai/TheArabicPile_Conversational.Pidgin-to-English-conversational-translations
Pidgin-to-English Translation Dataset (Sample)
Sample dataset: Nigerian Pidgin to English translation pairs for machine translation research
🤗 Hugging Face • 📊 Figshare • 🌐 Website • 📧 Contact
📋 Overview
The Pidgin-to-English Translation Dataset (Sample) is a conversational-style parallel corpus containing 122 translation pairs from Nigerian Pidgin English to Standard English. Created by Bytte AI through AI chatbot interactions with human validation… See the full description on the dataset page: https://huggingface.co/datasets/Bytte-AI/Pidgin-to-English-conversational-translations.prop-trading-qa-conversational-ai
Prop Trading Q&A Dataset for Conversational AI
Description
This dataset contains 200+ curated question-answer pairs covering the domain of proprietary (prop) trading firms. It is designed to serve as training and retrieval data for building AI assistants, chatbots, and educational tools focused on prop trading knowledge.
Each entry consists of a natural-language question paired with a detailed, factual answer. The data spans ten thematic categories ranging from… See the full description on the dataset page: https://huggingface.co/datasets/propfirmkey/prop-trading-qa-conversational-ai.
