CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01yatin-superintelligence /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M52 likes6.1k downloads6mo agoHugging Face02RUC-NLPIR /Omnimodal-Agent-SFT-2K OmniGAIA: Omni-Modal General AI Assistant Benchmark 📄 Paper   •   💻 Code & Demo   •   🤗 Dataset & Model   •   📈 Leaderboard This dataset contains omni-modal agent supervised fine-tuning (SFT) trajectories in the LlamaFactory SFT data format. You can directly follow LlamaFactory's instructions to fine-tune your omni-modal LLMs.OmniGAIA is a benchmark for Omni-Modal General AI Assistants that jointly reason over vision, audio, and language with external tools. It is… See the full description on the dataset page: https://huggingface.co/datasets/RUC-NLPIR/Omnimodal-Agent-SFT-2K.audioquestion-answering1K<n<10K9 likes5k downloads7mo agoHugging Face03ArlingtonCL2 /Barkopedia_DOG_AGE_GROUP_CLASSIFICATION_DATASET Dataset Check Training Data here: ArlingtonCL2/Barkopedia_DOG_AGE_GROUP_CLASSIFICATION_DATASET Dataset Description This dataset is for Dog Age Group Classification and contains dog bark audio clips. The data is split into training, public test, and private test sets. Training set: 17888 audio clips. Test set: 4920 audio clips, further divided into: Test Public (~40%): 1966 audio clips for live leaderboard updates. Test Private (~60%): 2954 audio clips for final evaluation. You will… See the full description on the dataset page: https://huggingface.co/datasets/ArlingtonCL2/Barkopedia_DOG_AGE_GROUP_CLASSIFICATION_DATASET.audioaudio-classification10K<n<100K2 likes3.2k downloads1y agoHugging Face04voidful /agent-sft-stitch-zh-tts agent-sft-stitch-zh-tts Voiced version of voidful/agent-sft-stitch-zh: the STITCH-S spoken chunks synthesized with BlueMagpie-TTS (hung_yi_lee voice), per-utterance loudness-aligned to -23 LUFS, best-of-N + Whisper-CER accepted. Configs records (default): one row per agent dialogue — id/source/user/msg (full STITCH-S trajectory) + available_tools + STITCH quality scores + spoken (ordered list of the utterances, each with audio, text, seg_index, cer, accepted… See the full description on the dataset page: https://huggingface.co/datasets/voidful/agent-sft-stitch-zh-tts.audiotext-to-speech100K<n<1M0 likes1.5k downloads2mo agoHugging Face05ziggylott /agent-tts-libraryaudion<1K0 likes1k downloads2mo agoHugging Face06yatin-superintelligence /Audio-Video-Engineering-Agentic-Tasks-1M Audio/Video Engineering Agentic Tasks (1M) Abstract A highly specialized dataset comprising 1,029,459 in-context troubleshooting prompts and execution commands built for the deepest levels of media production. Unlike standard datasets that simulate clean, theoretical instructions, this matrix captures the chaotic, highly-detailed, and conversational reality of professional audio engineers, composers, and video editors mid-session. It is engineered to train multimodal AI… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Audio-Video-Engineering-Agentic-Tasks-1M.tabulartext-generation1M<n<10M14 likes992 downloads6mo agoHugging Face07yatin-superintelligence /Creative-Professionals-Agentic-Tasks-1M Creative Professionals Agentic Tasks (1M) Abstract A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Creative-Professionals-Agentic-Tasks-1M.tabulartext-generation1M<n<10M29 likes947 downloads6mo agoHugging Face08hlx1021 /Barkopedia_DOG_AGE_GROUP_CLASSIFICATION_DATASET Dataset Check Training Data here: ArlingtonCL2/Barkopedia_DOG_AGE_GROUP_CLASSIFICATION_DATASET Dataset Description This dataset is for Dog Age Group Classification and contains dog bark audio clips. The data is split into training, public test, and private test sets. Training set: 17888 audio clips. Test set: 4920 audio clips, further divided into: Test Public (~40%): 1966 audio clips for live leaderboard updates. Test Private (~60%): 2954 audio clips for final evaluation. You… See the full description on the dataset page: https://huggingface.co/datasets/hlx1021/Barkopedia_DOG_AGE_GROUP_CLASSIFICATION_DATASET.audioaudio-classification10K<n<100K0 likes880 downloads3mo agoHugging Face09mteb /commonlanguage-age-miniaudio1K<n<10K0 likes827 downloads1y agoHugging Face10saeedzou /common-voice-17-en-age-gender-accentaudio100K<n<1M0 likes655 downloads2mo agoHugging Face11leungtianle /AgentChataudio100K<n<1M2 likes631 downloads5mo agoHugging Face12BlueIsGreen /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/BlueIsGreen/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M11 likes621 downloads6mo agoHugging Face13cx-cmu /AgentWebBench-corpus AgentWebBench Corpus Pre-built dense-retrieval corpus for AgentWebBench [ICML 2026], a benchmark for Multi-Agent Coordination in Agentic Web over a realistic 100-website slice of ClueWeb22 (~18.4M documents). This repository holds the embeddings and FAISS indices the benchmark loads at run time, including per-website indices, a global index, and website-level vectors. It does not contain ClueWeb22 text (see Raw documents). Websites: 100 Documents: ~18.4M Embedding dim: 1024… See the full description on the dataset page: https://huggingface.co/datasets/cx-cmu/AgentWebBench-corpus.audiotext-retrievaln<1K0 likes592 downloads3mo agoHugging Face14Jaward /lectura-agents-data LectūraAgents Dataset Overview This dataset is in support of findings in our paper LectūraAgents: A Multi-Agent Framework for Adaptive Personalized AI-Assisted Learning and Embodied Teaching. LectūaAgents is a hierarchical multi-agent framework that enables end-to-end personalized learning experiences through adaptive embodied teaching. It mirrors a professor–students’ relationship, wherein a ProfessorAgent guides a collaborative team of specialized subordinate… See the full description on the dataset page: https://huggingface.co/datasets/Jaward/lectura-agents-data.audion<1K24 likes585 downloads21d agoHugging Face15DEMIRUNC /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/DEMIRUNC/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M0 likes581 downloads6mo agoHugging Face16rAVEUK /Creative-Professionals-Agentic-Tasks-1M Creative Professionals Agentic Tasks (1M) Abstract A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/rAVEUK/Creative-Professionals-Agentic-Tasks-1M.tabulartext-generation1M<n<10M3 likes484 downloads6mo agoHugging Face17saeedzou /common-voice-17-en-age-genderaudio100K<n<1M0 likes459 downloads2mo agoHugging Face18Torenn /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/Torenn/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M1 likes400 downloads6mo agoHugging Face19leungtianle /AgentChat-Test Test Set Description This directory contains the test set used for tool-use evaluation. The JSON files under Test-JSON/ are organized by task type: SingleTaskProcessing/tool-select_test.json: single-tool selection tasks. ParallelProcessing/parallel-call_test.json: parallel tool-call tasks. ProactiveSeeking/searchTools_test_predictions_kept.json: proactive tool-search tasks. TaskDecomposition/muti-tool-select_test.json: multi-tool task decomposition tasks.… See the full description on the dataset page: https://huggingface.co/datasets/leungtianle/AgentChat-Test.audion<1K0 likes368 downloads3mo agoHugging Face20ppenner /edge-agent-reasoning-websearch-260k Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/ppenner/edge-agent-reasoning-websearch-260k.texttext-generation100K<n<1M0 likes340 downloads4mo agoHugging Face21kryp1234 /Creative-Professionals-Agentic-Tasks-1M Creative Professionals Agentic Tasks (1M) Abstract A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/kryp1234/Creative-Professionals-Agentic-Tasks-1M.tabulartext-generation1M<n<10M1 likes302 downloads6mo agoHugging Face22JACKYS999 /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/JACKYS999/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M0 likes290 downloads4mo agoHugging Face23kanepi-1977 /Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/kanepi-1977/Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M0 likes268 downloads6mo agoHugging Face24ageroul /SECD SECD: String Ensemble Chords Dataset SECD is a large-scale controlled compositional audio dataset for research on string-ensemble harmony and performance attributes. It contains 287,088 harmonic-interval and chord instances constructed through additive superposition of professionally recorded isolated string notes from the Philharmonia Orchestra into duo-, trio-, and quartet-like four-voice mixtures. The dataset is organised into six configurations covering harmonic intervals… See the full description on the dataset page: https://huggingface.co/datasets/ageroul/SECD.audioaudio-classification100K<n<1M0 likes233 downloads11d agoHugging Face25svryn /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/svryn/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M0 likes188 downloads4mo agoHugging Face26IngCrowd /japanese-roleplay-travel-agency Japanese Travel Agency Roleplay Dialogue Corpus Overview The Japanese Travel Agency Roleplay Dialogue Corpus is a collection of ten Japanese role-play dialogues simulating travel-agency consultations between a staff member and a customer. The conversations were collected for speech and dialogue research and include two-speaker mixed audio, speaker-separated audio, and manually created, time-aligned ELAN (.eaf) annotation files. The dialogues cover a variety of… See the full description on the dataset page: https://huggingface.co/datasets/IngCrowd/japanese-roleplay-travel-agency.audioautomatic-speech-recognitionn<1K0 likes144 downloads24d agoHugging Face27tterumiimurett1 /agentic-asrgated Agentic ASR Public consolidated audio and ASR result dataset for the OSWorld and WildClawBench benchmark families. Layout osworld/: synthetic raw/colloquial speech, human recordings, DNS-noise pairs, task images, ASR results, and reports. wildclawbench/: 60 formal colloquialized prompts, synthetic speech, 20 synthetic ASR condition tables, and ten-participant human recordings. task0_template derivatives are excluded. metadata/conditions.jsonl: model, variant… See the full description on the dataset page: https://huggingface.co/datasets/tterumiimurett1/agentic-asr.audio10K<n<100K0 likes138 downloads2d agoHugging Face28diffunity /GLOBE_V3_age_N_allsplitsaudio10K<n<100K0 likes136 downloads8mo agoHugging Face29saeedzou /common-voice-17-en-age-gender-accent-sampledaudio10K<n<100K0 likes129 downloads2mo agoHugging Face30AustinZhang /resp-agent-dataset Resp-229K: Respiratory Sound Dataset A Large-Scale Respiratory Sound Dataset for Training and Evaluation 📖 Overview Resp-229K is a comprehensive respiratory sound dataset containing 229,101 valid audio files with a total duration of over 407 hours. This dataset is curated for training the Resp-Agent system - an intelligent respiratory sound analysis and generation framework. 📊 Dataset Statistics Split Valid Files Total Duration Avg Duration Max… See the full description on the dataset page: https://huggingface.co/datasets/AustinZhang/resp-agent-dataset.audioaudio-classification100K<n<1M0 likes109 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.