CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01yatin-superintelligence /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M52 likes6.2k downloads6mo agoHugging Face02CanCLID /zoengjyutgaai 張悦楷講古語音數據集 English 呢個係張悦楷講《三國演義》、《水滸傳》、《走進毛澤東的最後歲月》、《鹿鼎記》語音數據集。張悦楷係廣州最出名嘅講古佬 / 粵語説書藝人。佢從上世紀七十年代開始就喺廣東各個收音電台度講古,佢把聲係好多廣州人嘅共同回憶。本數據集收集嘅係佢最知名嘅四部作品。 數據集用途: TTS(語音合成)訓練集 ASR(語音識別)訓練集或測試集 各種語言學、文學研究 直接聽嚟欣賞藝術! TTS 效果演示:https://huggingface.co/spaces/laubonghaudoi/zoengjyutgaai_tts 説明 所有文本都根據 https://jyutping.org/blog/typo/ 同 https://jyutping.org/blog/particles/ 規範用字。 所有文本都使用全角標點,冇半角標點。 所有文本都用漢字轉寫,無阿拉伯數字無英文字母 所有音頻源都存放喺/source,為方便直接用作訓練數據,切分後嘅音頻都放喺 opus/ 所有 opus 音頻皆為 48000… See the full description on the dataset page: https://huggingface.co/datasets/CanCLID/zoengjyutgaai.audioautomatic-speech-recognition100K<n<1M30 likes5.3k downloads8mo agoHugging Face03laubonghaudoi /legco-speech 香港立法會會議語音數據集 本數據集係由香港立法會會議製成嘅大規模語音數據集。原始錄音總時長 22,196 個鐘,切分語音後總時長 20,471 個鐘。數據集分兩個子集,raw同segmented,分別為原始錄音同VAD識別切分後嘅語音。 數據集製作流程 先去香港特別行政區立法會 YouTube下載所有會議紀錄並轉為 16kHz 採樣率嘅 OPUS音頻 用 fsmn-vad 切分所有語音,並用 Qwen3-ASR-1.7B 轉寫成粵文 srt 字幕 轉寫後用正則表達式修正字幕中常見轉寫錯誤 將數據集分成 raw、 segmented 兩個子集傳到HF 子集 subset raw segment 總行數 Row number 14,036 9,557,109 總時長 Total duration 22,195.55 hr (79,903,980.00 s) 20471.21 hr (73,696,365.27 s) 平均時長 Average duration 1.58 hr (5692.79 s) 7.71… See the full description on the dataset page: https://huggingface.co/datasets/laubonghaudoi/legco-speech.audioautomatic-speech-recognition1M<n<10M3 likes5.1k downloads7mo agoHugging Face04RVtech /Audio2Tool Audio2Tool: Speak, Call, Act — A Dataset for Benchmarking Speech Tool Use Authors: Ramit Pahwa1,∗,∗∗, Apoorva Beedu1,∗, Parivesh Priye1, Rutu Gandhi†1, Saloni Takawale†1, Aruna Baijal1, Zengli Yang1 1 Rivian & Volkswagen Technologies &nbsp;·&nbsp; ∗ equal contribution &nbsp;·&nbsp; ∗∗ corresponding author &nbsp;·&nbsp; † equal contribution 📄 Project page / demo: https://audio2tool.github.io/ 📦 Dataset: https://huggingface.co/datasets/RVtech/Audio2Tool ✉️ Contact (corresponding… See the full description on the dataset page: https://huggingface.co/datasets/RVtech/Audio2Tool.audioautomatic-speech-recognition10K<n<100K2 likes5k downloads3mo agoHugging Face05NJU-LINK /OmniVideoBenchgated OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs ✨ Overview Recent advances in multimodal large language models (MLLMs) have brought remarkable progress in video understanding.However, most existing benchmarks fail to jointly evaluate both audio and visual reasoning — often focusing on one modality or overlooking their interaction. 🎬 OmniVideoBench fills this gap.It’s a large-scale, rigorously curated… See the full description on the dataset page: https://huggingface.co/datasets/NJU-LINK/OmniVideoBench.texttext-generation1K<n<10K5 likes2.7k downloads6mo agoHugging Face06nyu-dice-lab /wavepulse-radio-raw-transcripts WavePulse Radio Raw Transcripts Dataset Summary WavePulse Radio Raw Transcripts is a large-scale dataset containing segment-level transcripts from 396 radio stations across the United States, collected between June 26, 2024, and Dec 29th, 2024. The dataset comprises >250 million text segments derived from 750,000+ hours of radio broadcasts, primarily covering news, talk shows, and political discussions. The summarized version of these transcripts is available here. For… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/wavepulse-radio-raw-transcripts.audiotext-generation100M<n<1B9 likes2.3k downloads2y agoHugging Face07marsianin500 /Speech2Latex Speech2Latex Dataset The Speech2LaTeX dataset is the first fully open-source large-scale dataset for converting spoken mathematical expressions and sentences into LaTeX. It comprises over 66,000 human-annotated audio samples of mathematical equations and sentences in both English and Russian, drawn from diverse scientific domains. This work lays the groundwork for future advances in multimodal AI, with a particular focus on mathematical content recognition. The dataset was presented… See the full description on the dataset page: https://huggingface.co/datasets/marsianin500/Speech2Latex.audioautomatic-speech-recognition100K<n<1M7 likes2k downloads10mo agoHugging Face08lenamerkli /distilled-web Dataset Card for lenamerkli/distilled-web This dataset consists of web-scraped data using a custom crawler purpose-built for each website. Dataset Details Dataset Sources Repository: https://github.com/lenamerkli/distilled-web Uses This dataset is useful for training large language models. The train split provides instruction-following and chat data for supervised fine-tuning (SFT) and instruction tuning. The pretrain split… See the full description on the dataset page: https://huggingface.co/datasets/lenamerkli/distilled-web.audiotext-generation1M<n<10M2 likes1.1k downloads5d agoHugging Face09Shelton1013 /SwitchLingua_audiogated Dataset Card for SwitchLingua_text 🚀 News [19/09/2025] SwitchLingua: The First Large-Scale Multilingual and Multi-Ethnic Code-Switching Dataset is accepted by NeurIPS 2025! [30/05/2024] The manuscript can be found on arXiv. Dataset Summary SwitchLingua is a comprehensive multilingual and multicultural code-switching dataset designed to advance research in automatic speech recognition, natural language processing, and conversational AI. The… See the full description on the dataset page: https://huggingface.co/datasets/Shelton1013/SwitchLingua_audio.audiotext-generation1K<n<10K12 likes1.1k downloads3mo agoHugging Face10yatin-superintelligence /Audio-Video-Engineering-Agentic-Tasks-1M Audio/Video Engineering Agentic Tasks (1M) Abstract A highly specialized dataset comprising 1,029,459 in-context troubleshooting prompts and execution commands built for the deepest levels of media production. Unlike standard datasets that simulate clean, theoretical instructions, this matrix captures the chaotic, highly-detailed, and conversational reality of professional audio engineers, composers, and video editors mid-session. It is engineered to train multimodal AI… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Audio-Video-Engineering-Agentic-Tasks-1M.tabulartext-generation1M<n<10M14 likes1.1k downloads6mo agoHugging Face11cjerzak /MultimodalMathBenchmarks MultimodalMathBenchmarks This repository contains the datasets for the paper Multiplication in Multimodal LLMs: Computation with Text, Image, and Audio Inputs (ACL Findings 2026). It covers the public benchmark datasets and their modality assets (text, images, and audio) used to evaluate the arithmetic capabilities of multimodal LLMs. Canonical Upload Manifest HF path Local source Count Purpose SharedMultimodalGrid.csv SavedData/SharedMultimodalGrid.csv… See the full description on the dataset page: https://huggingface.co/datasets/cjerzak/MultimodalMathBenchmarks.audioimage-text-to-text10K<n<100K0 likes846 downloads2mo agoHugging Face12yatin-superintelligence /Creative-Professionals-Agentic-Tasks-1M Creative Professionals Agentic Tasks (1M) Abstract A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Creative-Professionals-Agentic-Tasks-1M.tabulartext-generation1M<n<10M29 likes841 downloads6mo agoHugging Face13BlueIsGreen /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/BlueIsGreen/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M11 likes695 downloads6mo agoHugging Face14ClareNie /Light-Omni-Training Light-Omni Training Dataset This repository contains the training data used by Light-Omni, a multimodal agent framework for reflexive video understanding with long-term memory. Light-Omni uses memory-augmented multimodal streams to train adapters for memory construction, response generation, and reaction/action control. Links Project page: https://clare-nie.github.io/Light-Omni/ Code: https://github.com/Clare-Nie/Light-Omni Dataset:… See the full description on the dataset page: https://huggingface.co/datasets/ClareNie/Light-Omni-Training.audiovisual-question-answering100K<n<1M3 likes686 downloads3mo agoHugging Face15ATH-MaaS /Marco_Longspeech Marco-LongSpeech Dataset Marco-LongSpeech is a multi-task long speech understanding dataset containing 8 different speech understanding tasks designed to benchmark Large Language Models on lengthy audio inputs. 📊 Dataset Statistics Task Statistics Task Train Val Test Total Unique Audios ASR 71,275 15,273 15,274 101,822 101,822 Temporal_Relative_QA 5,886 1,261 1,262 8,409 8,409 summary 4,366 935 937 6,238 6,238… See the full description on the dataset page: https://huggingface.co/datasets/ATH-MaaS/Marco_Longspeech.audioautomatic-speech-recognition10K<n<100K18 likes678 downloads4mo agoHugging Face16MathLLMs /VoiceAssistant-Eval 🔥 VoiceAssistant-Eval: Benchmarking AI Assistants across Listening, Speaking, and Viewing [🌐 Homepage] [🔮 Visualization] [💻 Github] [📖 Paper] [📊 Leaderboard ] [📊 Detailed Leaderboard ] [📊 Roleplay Leaderboard ] 🚀 Data Usage from datasets import load_dataset for split in ['listening_general', 'listening_music', 'listening_sound', 'listening_speech', 'speaking_assistant', 'speaking_emotion', 'speaking_instruction_following'… See the full description on the dataset page: https://huggingface.co/datasets/MathLLMs/VoiceAssistant-Eval.textquestion-answering10K<n<100K12 likes571 downloads11mo agoHugging Face17DEMIRUNC /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/DEMIRUNC/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M0 likes548 downloads6mo agoHugging Face18rAVEUK /Creative-Professionals-Agentic-Tasks-1M Creative Professionals Agentic Tasks (1M) Abstract A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/rAVEUK/Creative-Professionals-Agentic-Tasks-1M.tabulartext-generation1M<n<10M3 likes477 downloads6mo agoHugging Face19My-Weird-Prompts /episodes My Weird Prompts - Episode Dataset The production record of every episode of the My Weird Prompts podcast: metadata, the transcript, and the generation telemetry for how each episode was made - model, pipeline version, GPU, timings and compute cost. 5,333 episodes. Synced daily from the production database. from datasets import load_dataset ds = load_dataset("My-Weird-Prompts/episodes", split="train") Which dataset do you want? This one… See the full description on the dataset page: https://huggingface.co/datasets/My-Weird-Prompts/episodes.audiotext-generation1K<n<10K0 likes470 downloads12h agoHugging Face20RodolfoRizzi /MyMentorLLM-dataset Dataset Card for MyMentorLLM This dataset accompanies the MyMentorLLM paper, arXiv:2607.25667, which describes the simulation environment, generation procedure, experimental design and validation analyses. If you use this dataset, please cite the paper (see Citation Information). Dataset Summary MyMentorLLM is a multimodal voice- and text-based deliberate-practice environment used to generate 2,100 simulated Cognitive Behavioural Therapy (CBT) training sessions… See the full description on the dataset page: https://huggingface.co/datasets/RodolfoRizzi/MyMentorLLM-dataset.audiotext-generation10K<n<100K0 likes428 downloads15h agoHugging Face21omnibench /anonymous-storybench Omni-StoryBench Omni-StoryBench is a context-aware omnimodal story generation benchmark.Each sample provides a current story page and requires generating the next page's image, narration text, and speech utterance. Dataset Structure The dataset contains: data/testset.jsonl: Main benchmark file. images/: Page images. texts/: Page text files. speech/: Generated speech audio files. instruction/: Source-level instruction metadata. Data Fields Each JSONL sample… See the full description on the dataset page: https://huggingface.co/datasets/omnibench/anonymous-storybench.audiotext-generationn<1K0 likes384 downloads5mo agoHugging Face22ddwang2000 /SpeechInstructBench SpeechInstructBench Arxiv: https://arxiv.org/abs/2503.02769 This is the SpeechInstructBench dataset download page. SpeechInstructBench is a multilingual (Chinese and English) benchmark designed to evaluate the instruction-following capabilities of speech models. Instruction-following refers to a model’s ability to accurately interpret and execute user-provided natural language directives while strictly adhering to all specified constraints and requirements. To comprehensively… See the full description on the dataset page: https://huggingface.co/datasets/ddwang2000/SpeechInstructBench.audioquestion-answering10K<n<100K3 likes343 downloads8mo agoHugging Face23recursal /reprocessed_singapore_national_speech_corpus Dataset Card for Reprocessed National Speech Corpus NOTE: This is an Reprocessed version KaraKaraWitch from Recursal.The official download can be found here. Dataset Details Dataset Description Dataset Description: The National Speech Corpus (NSC) is the first large-scale Singapore English corpus, sponsored by the Info-communications and Media Development Authority (IMDA) of Singapore. The objective is to serve as a primary resource of open speech data for… See the full description on the dataset page: https://huggingface.co/datasets/recursal/reprocessed_singapore_national_speech_corpus.audiotext-generation1M<n<10M7 likes340 downloads2y agoHugging Face24minghanw /sdf_dataset_en SpeechDialogueFactory Dataset Background This dataset is part of the SpeechDialogueFactory project, a comprehensive framework for generating high-quality speech dialogues at scale. Speech dialogue datasets are essential for developing and evaluating Speech-LLMs, but existing datasets face limitations including high collection costs, privacy concerns, and lack of conversational authenticity. This dataset addresses these challenges by providing synthetically generated… See the full description on the dataset page: https://huggingface.co/datasets/minghanw/sdf_dataset_en.audiotext-generation1K<n<10K5 likes330 downloads1y agoHugging Face25Torenn /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/Torenn/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M1 likes328 downloads6mo agoHugging Face26ai4bharat /IndicCMixgated IndicCMix Most Indic NLP data assumes people write in one script and one language at a time. Real chat looks nothing like that. You get Hindi words in Roman letters, English verbs in the middle of a Tamil sentence, and the same person switching scripts halfway through a paragraph. This dataset is an attempt to cover that actual messiness. For every English sentence, you get three different Indic renderings of it: one code-mixed in the native script, one clean native-script… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/IndicCMix.audiotranslationn<1K1 likes314 downloads5mo agoHugging Face27amol-derick /strudel-rl-data strudel-rl-data Datasets from the Strudel-RL project (training a Qwen3.6-35B-A3B to write house / techno / hypnotic techno as Strudel programs). sft/: SFT datasets v1..v5 (messages format, system+user+assistant) with stats. synth/: GLM-5.3-Flash synthetic rounds 1..5 (prompt, code, reasoning length). captions/: GLM captions of programs. prompts/: brief generators and held-out eval briefs (v0..v3; v3 = reference-driven). gold/: 26 hand-written gold programs + manifest.… See the full description on the dataset page: https://huggingface.co/datasets/amol-derick/strudel-rl-data.audiotext-generation0 likes313 downloads18d agoHugging Face28HongbangYuan /OmniRewardBench Omni-Reward: Towards Generalist Omni-Modal Reward Modeling with Free-Form Preferences 📄 Paper | 💻 Code | 🤗 Benchmark (This Dataset) | 🤗 Training Data | 🤗 Model | 🏠 Homepage Reward models (RMs) play a critical role in aligning AI behaviors with human preferences, yet they face two fundamental challenges: (1) Modality Imbalance, where most RMs are mainly focused on text and image modalities, offering limited support for video, audio, and other modalities; and… See the full description on the dataset page: https://huggingface.co/datasets/HongbangYuan/OmniRewardBench.audiotext-generation1K<n<10K6 likes312 downloads11mo agoHugging Face29kryp1234 /Creative-Professionals-Agentic-Tasks-1M Creative Professionals Agentic Tasks (1M) Abstract A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/kryp1234/Creative-Professionals-Agentic-Tasks-1M.tabulartext-generation1M<n<10M1 likes308 downloads6mo agoHugging Face30zuhri025 /OpenDialog_English OpenDialog English This dataset contains English dialog and conversation data. Dataset Structure The dataset is provided in Parquet format with 153 splits for efficient loading. Data Files Format: Parquet Splits: 153 files (train-00001-of-00153.parquet through train-00153-of-00153.parquet) Total Size: ~72.8 GB Loading the Dataset from datasets import load_dataset # Load the full dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/zuhri025/OpenDialog_English.audiotext-generation100K<n<1M1 likes291 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.