CoolFace
25 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01yatin-superintelligence /Creative-Professionals-Agentic-Tasks-1M Creative Professionals Agentic Tasks (1M) Abstract A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Creative-Professionals-Agentic-Tasks-1M.tabulartext-generation1M<n<10M29 likes1.2k downloads7mo agoHugging Face02My-Weird-Prompts /episodes My Weird Prompts - Episode Dataset The production record of every episode of the My Weird Prompts podcast: the transcript, links to the published episode, a description of the prompt that started it, and the generation telemetry for how it was made - model, pipeline version, GPU, timings and compute cost. 5,393 episodes. Synced daily from the production database. from datasets import load_dataset ds = load_dataset("My-Weird-Prompts/episodes", split="train") Which… See the full description on the dataset page: https://huggingface.co/datasets/My-Weird-Prompts/episodes.audiotext-generation1K<n<10K1 likes516 downloads18h agoHugging Face03rAVEUK /Creative-Professionals-Agentic-Tasks-1M Creative Professionals Agentic Tasks (1M) Abstract A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/rAVEUK/Creative-Professionals-Agentic-Tasks-1M.tabulartext-generation1M<n<10M3 likes511 downloads6mo agoHugging Face04ppenner /edge-agent-reasoning-websearch-260k Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/ppenner/edge-agent-reasoning-websearch-260k.texttext-generation100K<n<1M0 likes340 downloads4mo agoHugging Face05kryp1234 /Creative-Professionals-Agentic-Tasks-1M Creative Professionals Agentic Tasks (1M) Abstract A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/kryp1234/Creative-Professionals-Agentic-Tasks-1M.tabulartext-generation1M<n<10M1 likes297 downloads6mo agoHugging Face06aamirhs /pashto-audio-wav2vecaudiotext-generationn<1K0 likes246 downloads4y agoHugging Face07Pawlo77 /mllm-shap MLLM-SHAP experiment datasets Curated test splits for studying Shapley-value explanations in multimodal large language models (text and audio inputs). Each configuration is a filtered, size-controlled subset built for reproducible benchmarking—not a full copy of the upstream corpora. Configs follow the naming pattern {task}__{source} (for example single_sentence__voice_bench). Quick load Pin a dataset revision for reproducibility (replace REVISION with the commit hash… See the full description on the dataset page: https://huggingface.co/datasets/Pawlo77/mllm-shap.tabulartext-generation1K<n<10K2 likes212 downloads4mo agoHugging Face08vhands /audio-event-classification-post-public audio-event-classification-post-public Sound-event and acoustic-scene classification annotations: ESC-50 (environmental), UrbanSound8K, FSD50k (50k+ events), TUT-Acoustic-Scenes-2017, DCASE-2025, NonSpeech7k (vocal sounds), VocalSound (laugh/cough/sigh). Useful for training audio LLMs on the perception substrate underneath higher-level reasoning. Audio is not bundled in this repo. See download.sh and per-dataset data/<name>.info.json for the fetch recipe; run postlink_audio.py… See the full description on the dataset page: https://huggingface.co/datasets/vhands/audio-event-classification-post-public.textaudio-classification100K<n<1M1 likes192 downloads3mo agoHugging Face09zeio /auto-pale Dataset card for pale Dataset summary This dataset contains league of legends champions' quotes parsed from fandom. See dataset usage example at google colab. The dataset is available in the following configurations: vanilla - all data pulled from the website without significant modifications apart from the web page structure parsing; quotes - truncated version of the corpus, which does't contain sound effects; annotated - an extended version of the full configuration… See the full description on the dataset page: https://huggingface.co/datasets/zeio/auto-pale.audiotext-generation100K<n<1M0 likes174 downloads3y agoHugging Face10ivrit-ai /knesset-plenumsgated About This dataset is derived from raw a/v recordings and human-generated protocols of the Knesset (the Israeli house of representatives) plenums as part of the ivrit.ai project. Consider visiting the preview space for this dataset here Method Data dumps from the Knesset contain A/V recordings, alongside proprietary protocols with timestamps. We extract the audio stream, and clean up timestamp mistakes (such as backward jumps, or out-of-order timestamp artifacts). The… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/knesset-plenums.audioautomatic-speech-recognition1K<n<10K3 likes108 downloads10mo agoHugging Face11arcada-labs /product-bench Product Bench 31-turn multi-turn speech-to-speech benchmark for evaluating voice AI models as a laptop comparison shopping assistant. Part of Audio Arena, a suite of 6 benchmarks spanning 221 turns across different domains. Built by Arcada Labs. Leaderboard | GitHub | All Benchmarks Dataset Description The model acts as a laptop comparison shopping assistant helping a customer evaluate, compare, and order laptops. The conversation features multi-intent turns… See the full description on the dataset page: https://huggingface.co/datasets/arcada-labs/product-bench.audioautomatic-speech-recognitionn<1K2 likes108 downloads6mo agoHugging Face12danielrosehill /Long-Prompt-ExperimentI conducted this experiment to investigate the impact of prompt structure and optimization on LLM performance, specifically testing whether quality and organization matter more than raw prompt length for complex technical tasks. Research Question For specialized technical tasks, does prompt structure and optimization have a greater impact on output quality than raw prompt length alone? Experiment Design I compared three distinct prompting approaches using Gemini 2.5 Lite… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Long-Prompt-Experiment.audiotext-generationn<1K0 likes76 downloads1y agoHugging Face13vhands /audio-music-mir-post-public audio-music-mir-post-public Music information retrieval and tagging annotations: genre (FMA), instrument family (NSynth × 3, Medley-solos-DB), social tags (MagnaTagATune via LLARK), and large-scale Music4All metadata. Foundation for music understanding heads in audio LLMs. Audio is not bundled in this repo. See download.sh and per-dataset data/<name>.info.json for the fetch recipe; run postlink_audio.py after fetching to rewrite the JSONL audio_path fields with absolute local… See the full description on the dataset page: https://huggingface.co/datasets/vhands/audio-music-mir-post-public.textaudio-classification100K<n<1M0 likes74 downloads3mo agoHugging Face14vhands /audio-reasoning-qa-post-public audio-reasoning-qa-post-public Question-answering and multi-task audio reasoning annotations across 15 public audio QA datasets. Spans general audio QA (Clotho-AQA, HeySQuAD), music reasoning (MU-LLaMA, MusicBench, LLARK-MTAT, Music-AVQA), speech-grounded QA (LibriSQA, GigaSpeech), and NVIDIA-aggregator skill subsets (TemporalQA, CountingQA, AudioSet-Speech-QA, GigaSpeech-Long-QA). Closes a substantial slice of the public audio-reasoning SFT gap (compare to NVIDIA AudioSkills-XL… See the full description on the dataset page: https://huggingface.co/datasets/vhands/audio-reasoning-qa-post-public.textquestion-answering100K<n<1M0 likes67 downloads3mo agoHugging Face15Paranoiid /VoicePersona VoicePersona Dataset A comprehensive voice persona dataset for character consistency in voice synthesis, generated using advanced audio-language models. 📋 Overview VoicePersona Dataset serves as the training foundation for VoiceForge - an AI architecture that generates character voices from pure text descriptions. The Connection: VoicePersona provides detailed voice characteristics and personality profilesVoiceForge uses this data to learn text→voice mapping for… See the full description on the dataset page: https://huggingface.co/datasets/Paranoiid/VoicePersona.audiotext-generation10K<n<100K1 likes58 downloads1y agoHugging Face16Pastaaaaa2003 /Hindi-speech-instructgated Hindi LLaMA-Omni Instruct Dataset A Hindi speech instruction-following dataset designed for training speech-language models such as LLaMA-Omni. Each example pairs a spoken Hindi user question (audio) with a text assistant response. Dataset Summary Property Value Language Hindi (hi) Total examples ~110,718 Train split ~105,000 examples (batches 001–210) Validation split ~5,500 examples (batches 211–222) Audio format FLAC, 16,000 Hz mono… See the full description on the dataset page: https://huggingface.co/datasets/Pastaaaaa2003/Hindi-speech-instruct.audioautomatic-speech-recognition100K<n<1M0 likes56 downloads3mo agoHugging Face17tterumiimurett1 /colloqialized_prompt Colloquialized Prompt Dataset This repository contains prompt and audio variants derived from the 60 WildClawBench tasks, plus the reusable task template. It supports experiments that compare written prompts, spoken-style rewrites, synthesized speech, raw ASR transcripts, and normalized ASR transcripts. Dataset layout . ├── prompts/ # Instructions used by rewrite/normalization jobs ├── scripts/ # Reproducible data preparation… See the full description on the dataset page: https://huggingface.co/datasets/tterumiimurett1/colloqialized_prompt.audioautomatic-speech-recognitionn<1K0 likes46 downloads2mo agoHugging Face18audibeal74 /panta_instruct_multi_modal_v1 Panta Instruct Multi-Modal v1 Dataset d'instructions multimodal en français : chaque exemple associe une question (texte + parole + pictogrammes) à une réponse (texte + pictogrammes). Colonnes Colonne Type Description audio Audio (24 kHz, mono) Enregistrement de la question (text_input) text_input string Question / instruction text_output string Réponse pictos_input list[string] Identifiants des pictogrammes de la question pictos_output… See the full description on the dataset page: https://huggingface.co/datasets/audibeal74/panta_instruct_multi_modal_v1.audioautomatic-speech-recognition10K<n<100K0 likes41 downloads22h agoHugging Face19Rexe /nigerian-pidgin-speech Nigerian Pidgin Audio + Text Dataset for Whisper Fine-tuning Nigerian Pidgin speech dataset for Whisper fine-tuning Dataset Summary This dataset contains audio recordings and transcriptions in Nigerian Pidgin English, designed for fine-tuning speech recognition models, particularly OpenAI's Whisper. Dataset Structure Train Split: 65 samples Test Split: 8 samples Total Duration: 0.0 hours (estimated) Average Duration: 2.4 seconds per sample Sample Rate: 16kHz… See the full description on the dataset page: https://huggingface.co/datasets/Rexe/nigerian-pidgin-speech.audioautomatic-speech-recognitionn<1K1 likes40 downloads1y agoHugging Face20TumeloKonaite /synthetic-patient-dr-data Synthetic Patient DR Data Synthetic doctor-patient consultation dataset with structured clinical outputs and optional full-consultation audio. Dataset Summary This dataset was generated for research and prototyping in: clinical dialogue generation structured clinical extraction text-to-audio workflows conversational healthcare modeling All consultations are synthetic and should not be treated as real clinical encounters. Export Metadata Mode: audio Repo… See the full description on the dataset page: https://huggingface.co/datasets/TumeloKonaite/synthetic-patient-dr-data.audiotext-generationn<1K0 likes23 downloads6mo agoHugging Face21gruhit-patel /alpaca_speech_instructaudiotext-generation10K<n<100K2 likes22 downloads2y agoHugging Face22satierf /Rhulk_pt-braudiotext-to-speechn<1K1 likes18 downloads3y agoHugging Face23PatronusAI /BLURgated Browsing Lost Unformed Recollections The leaderboard can be found at https://huggingface.co/spaces/PatronusAI/BLUR-leaderboard. If you use or find this dataset helpful in your research, please do cite our paper: Paper Link: arXiv @misc{chwang2025blur, title = {Browsing {Lost} {Unformed} {Recollections}: {A} {Benchmark} for {Tip}-of-the-{Tongue} {Search} and {Reasoning}}, shorttitle = {Browsing {Lost} {Unformed} {Recollections}}, url = {http://arxiv.org/abs/2503.19193}… See the full description on the dataset page: https://huggingface.co/datasets/PatronusAI/BLUR.audiotext-generationn<1K12 likes12 downloads2y agoHugging Face24rchiera /podcast-transcripts Podcast Transcripts Dataset This dataset contains transcripts from Bitcoin and cryptocurrency podcasts, processed by the belief-engines ETL pipeline. Files transcripts.parquet - Full episode transcripts with speaker diarization transcript_chunks.parquet - Transcript chunks (512 tokens) with optional embeddings Schema transcripts.parquet episode_id: Unique episode identifier podcast_slug: Podcast name slug episode_slug: Episode name slug… See the full description on the dataset page: https://huggingface.co/datasets/rchiera/podcast-transcripts.tabulartext-generation10K<n<100K0 likes12 downloads7mo agoHugging Face25phongps2 /VietSpeechgated Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Based on dataset: https://huggingface.co/datasets/NhutP/VietSpeech Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More… See the full description on the dataset page: https://huggingface.co/datasets/phongps2/VietSpeech.audiotext-generation100K<n<1M0 likes7 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.