datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Full-Ecom-Chatbot-Dataset
E-commerce Chatbot Training Data
A curated, multi-source dataset for training and evaluating e-commerce conversational AI systems. It covers a broad range of customer intents — from product discovery and order management to returns, tool-augmented responses, and RAG-grounded Q&A — across 16+ product domains.
Dataset Summary
Split
Records
Train
35,213
Test
8,818
Total
44,031
The train/test split uses prompt-group-level stratified sampling on source ×… See the full description on the dataset page: https://huggingface.co/datasets/rescommons/Full-Ecom-Chatbot-Dataset.full-modality-data
Full Modality Dataset Statistics
Video Statistics
Total Videos: 28,472
Total Duration: 1422.33 hours
Average Duration: 179.84 seconds
Median Duration: 160.08 seconds
Duration Range: 10.04s - 1780.03s
QA Statistics
Total Questions: 1,444,526
Average Questions per Video: 50.7
Questions per Video Range: 14 - 450
Question Type Distribution
OE: 1,444,526 (100.0%)
Question Category Distribution
temporal: 96,873 (6.7%)
causal: 96,873… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab/full-modality-data.CHSA-Triage-Medic-Full-Dataset
CHSA-Triage-Medic-Full-Dataset
Ce dataset a été constitué dans le cadre d'un projet de formation AI Engineer (Projet CHSA). Il est conçu pour entraîner un Assistant Médical Intelligent capable d'effectuer du triage d'urgence et de fournir des raisonnements cliniques.
Le dataset est divisé en 3 sous-ensembles distincts correspondant aux différentes phases d'entraînement (Fine-Tuning Supervisé et Alignement).
Organisation du Dataset
Le repository contient trois… See the full description on the dataset page: https://huggingface.co/datasets/cyrille-elie/CHSA-Triage-Medic-Full-Dataset.ams_data_full_2000-2020Aerospace Mechanism Symposia PDF documents parsed by page. All symposia documents from the year 2000-2022 are included. No splitting was used.
Original documents here: https://github.com/dan-s-mueller/aerospace_chatbot/tree/main/data/AMS
