CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01oss-codes /Finance-Conversational-Dataset-Indictext100K<n<1M1 likes2.3k downloads1y agoHugging Face02oss-codes /NCERT-Conversational-Dataset-Indictext100K<n<1M0 likes1.9k downloads1y agoHugging Face03nvidia /Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1 Dataset Description: We created an RL dataset for conversational tool-use by utilizing existing expert tool-use trajectories. We pose each assistant step of the trajectory as a separate behavior cloning problem where the policy model is incentivized to match the tool call choices of the expert model. Each trajectory includes the use of tools for authentication, data lookup, servicing (i.e. booking reservations, changing them, getting discounts, etc), and more across 838 different… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1.tabular10K<n<100K32 likes1.4k downloads6mo agoHugging Face04oss-codes /Law-Conversational-Dataset-Indictext100K<n<1M0 likes1.4k downloads1y agoHugging Face05oss-codes /Cyber-Conversational-Dataset-Indictext1K<n<10K0 likes1.3k downloads1y agoHugging Face06oss-codes /Coding-Conversational-Dataset-Indic1 likes1.1k downloads1y agoHugging Face07sweatSmile /medical-symptom-triage-conversationaltabular1K<n<10K1 likes1k downloads1y agoHugging Face08JDKdev /french-tts-conversational-dataset French Conversational TTS Dataset Dataset Description This dataset contains high-fidelity French text-to-speech audio clips generated using Mistral's Voxtral Mini TTS model (voxtral-mini-tts-2603). It covers three B2B industry verticals with balanced male/female speaker distribution. Verticals Vertical Description fintech_banking Banking operations, account inquiries, fraud alerts, investments, customer service ecommerce_logistics Order… See the full description on the dataset page: https://huggingface.co/datasets/JDKdev/french-tts-conversational-dataset.audiotext-to-speechn<1K0 likes907 downloads3mo agoHugging Face09Narsil /conversational_dummy0 likes795 downloads2y agoHugging Face10oss-codes /Computer-Science-Conversational-Dataset-Indictext10K<n<100K0 likes783 downloads1y agoHugging Face11AdaMLLab /smolkalam-arabic-conversational-sft SmolKalam SmolKalam is a quality-filtered Arabic SFT dataset of 1,790,478 examples (~2.45B tokens), built as an ensemble translation of SmolTalk2. It covers multi-turn dialogue (23% of rows), reasoning traces (19% carry <think>), tool and function calling (4.4%), and long context, categories that are underrepresented in existing Arabic post-training data. The SmolTalk2 source mixtures are kept as subsets. Released with the paper SmolKalam: Ensemble Quality-Filtered Translation… See the full description on the dataset page: https://huggingface.co/datasets/AdaMLLab/smolkalam-arabic-conversational-sft.tabulartext-generation1M<n<10M3 likes771 downloads1mo agoHugging Face12oss-codes /CA-Conversational-Dataset-Indictext100K<n<1M0 likes626 downloads1y agoHugging Face13suriya7 /Conversational-Dataset-cleanedtext1K<n<10K3 likes605 downloads2y agoHugging Face14B-Lounes /conversational_audio_fr_dataset-metadatatextn<1K0 likes604 downloads1y agoHugging Face15nytopop /expresso-conversational The Expresso Dataset [paper] [demo samples] [Original repository] Introduction The Expresso dataset is a high-quality (48kHz) expressive speech dataset that includes both expressively rendered read speech (8 styles, in mono wav format) and improvised dialogues (26 styles, in stereo wav format). The dataset includes 4 speakers (2 males, 2 females), and totals 40 hours (11h read, 30h improvised). The transcriptions of the read speech are also provided. You can listen to… See the full description on the dataset page: https://huggingface.co/datasets/nytopop/expresso-conversational.audio10K<n<100K14 likes406 downloads1y agoHugging Face16GEM /conversational_weatherThe Conversational Weather dataset is designed for generation of responses to weather queries based on a structured input data. The input allows specifying data attributes such as dates, times, locations, weather conditions, and errors, and also offers control over structure of response through discourse relations such as join, contrast, and justification.table-to-text5 likes396 downloads4y agoHugging Face17Conversational-Reasoning /Topical-Chat Topical-Chat We introduce Topical-Chat, a knowledge-grounded human-human conversation dataset where the underlying knowledge spans 8 broad topics and conversation partners don’t have explicitly defined roles. Topical-Chat broadly consists of two types of files: Conversations: JSON files containing conversations between pairs of Amazon Mechanical Turk workers. Reading Sets: JSON files containing knowledge sections rendered as reading content to the Turkers having conversations. For… See the full description on the dataset page: https://huggingface.co/datasets/Conversational-Reasoning/Topical-Chat.4 likes396 downloads3y agoHugging Face18Cyleux /gemma3n-conversational-reasoning Gemma3N Conversational Reasoning This dataset is prepared for Unsloth Gemma3/Gemma3N conversational notebooks that use: from datasets import load_dataset from unsloth.chat_templates import standardize_data_formats dataset = load_dataset("Cyleux/gemma3n-conversational-reasoning", split="train[:3000]") dataset = standardize_data_formats(dataset) Schema: conversations: ShareGPT-style list of turns with from and value metadata columns are included for analysis and filtering Notes:… See the full description on the dataset page: https://huggingface.co/datasets/Cyleux/gemma3n-conversational-reasoning.tabulartext-generation1K<n<10K0 likes371 downloads8mo agoHugging Face19ElementXMaster /conversational_data_untokenized_mergedtext1M<n<10M0 likes294 downloads1y agoHugging Face20nvidia /NeMo-Gym-Conversational-Tool-Use-Assets NeMo Gym Conversational Tool-Use Assets This dataset repository stores prompt and reference assets for NeMo Gym's conversational tool-use generation pipeline. It is an asset bundle for Gym components, not a training or evaluation dataset. Contents conversational_tool_use_domain_generation/prompts: the domain-generation prompt. conversational_tool_use_domain_generation/prompt_history: historical domain-generation prompt revisions.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/NeMo-Gym-Conversational-Tool-Use-Assets.textn<1K2 likes282 downloads2mo agoHugging Face21ksterx /hle-no-img-conversational-formatimage1K<n<10K0 likes252 downloads1y agoHugging Face22HTH-inc /japanese-casual-conversational-speech-golden-dataset-preview Japanese Casual Conversational Speech Golden Dataset (Preview) 💼 Commercial License & Full Access This repository contains a limited preview. The full 60-hour dataset collected via the "Kataro" app is available for commercial use, ASR benchmarking, and Spoken Dialogue Model fine-tuning. To purchase the full dataset, please contact us: 👉 Email: info@hth-inc.com 👉 Website: https://hth-inc.com/business 🌟 4 Reasons to Choose This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/HTH-inc/japanese-casual-conversational-speech-golden-dataset-preview.audioautomatic-speech-recognitionn<1K2 likes234 downloads21d agoHugging Face23Makan09 /bam-asr-conversational All Bambara ASR Dataset This is the dataset that fueled our early ASR experiments that gave as results the V0 models. It is primarily composed of the Jeli-ASR dataset (available at RobotsMali/jeli-asr), along with the Mali-Pense data curated and published by Aboubacar Ouattara (available at oza75/bambara-tts). Additionally, it includes 1 hour of audio recently collected by the RobotsMali AI4D Lab, featuring children's voices reading some of RobotsMali GAIFE books. This dataset… See the full description on the dataset page: https://huggingface.co/datasets/Makan09/bam-asr-conversational.audioautomatic-speech-recognition10K<n<100K2 likes234 downloads23d agoHugging Face24oss-codes /Medical-Conversational-Dataset-Indictext10K<n<100K0 likes227 downloads1y agoHugging Face25BlossomsAI /vietnamese-conversational-datasettext1M<n<10M0 likes218 downloads1y agoHugging Face26MagicHub /korean-conversational-speech-corpus ASR-KCSC: A Korean Conversational Speech Corpus Every data point counts. Dataset Basic Info Dataset Type: ASR Speech Corpus Language: Korean Audio Parameters: 16 kHz, 16 bits File Format: WAV (PCM) Recording Equipment: Mobile device Recording Environment: Indoor Dataset Description This open-source dataset consists of 5.22 hours of transcribed Korean conversational speech on certain topics, where 22 conversations between seven pairs of speakers… See the full description on the dataset page: https://huggingface.co/datasets/MagicHub/korean-conversational-speech-corpus.audio1K<n<10K1 likes217 downloads3mo agoHugging Face27Roudranil /shakespearean-and-modern-english-conversational-dataset Dataset Card for Shakespearean and Modern English Conversational Dataset Dataset Summary This dataset contains dialog pairs taken from Shakespeare's works - the first dialog is a translated text in modern english, and the second dialog is it's actual response as written in Shakespeare's plays. See the github repo for more details. text1K<n<10K4 likes213 downloads1y agoHugging Face28Conversational-Reasoning /Topical-ChatASR Topical-Chat ASR: An ASR-augmented version of Topical-Chat This README describes Topical-Chat ASR, an augmentation of Topical-Chat with non-trivial synthetic and actual ASR hypotheses. Synthetic: /TopicalChatASR/synthetic For each file in the original Topical-Chat dataset, non-trivial synthetic ASR hypotheses are constructed at four different corpus-level target Word Error Rates (WER). We used the ASR error simulator method based on n-gram confusion matrix and trained the… See the full description on the dataset page: https://huggingface.co/datasets/Conversational-Reasoning/Topical-ChatASR.text-classification100K<n<1M1 likes205 downloads3y agoHugging Face29oss-codes /CAT-Conversational-Dataset-Indictext10K<n<100K0 likes204 downloads1y agoHugging Face30Conversational-Reasoning /Topical-Chat-Enriched Enriched Topical-Chat: A Dialogue Act and Knowledge Sentence annotated version of Topical-Chat This README describes Enriched Topical-Chat, an augmentation of Topical-Chat that contains dialogue act and knowledge sentence annotations for each turn in the dataset. Each annotation is automatically annotated using off-the-shelf models. Knowledge Sentence Annoations Each conversation in Topical-Chat has a pair of reading sets which consists of a set of knowledge sentences.… See the full description on the dataset page: https://huggingface.co/datasets/Conversational-Reasoning/Topical-Chat-Enriched.0 likes186 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.