CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01oncody /AI_Agent_Task_Dataset 🤖 Massive AI Agent Task Dataset (10.5GB) 📌 Overview Welcome to the AI Agent Task Dataset, a massive 10.5GB procedural dataset designed for training, fine-tuning, and evaluating autonomous AI agents and LLMs. This dataset focuses on: Multi-step reasoning Tool usage (APIs, frameworks, systems) Real-world execution workflows Perfect for building agentic AI systems, copilots, and automation models. 📑 Table of Contents Dataset Details Dataset… See the full description on the dataset page: https://huggingface.co/datasets/oncody/AI_Agent_Task_Dataset.texttext-generation10M<n<100M3 likes79 downloads6mo agoHugging Face02beatsprom /realtime-conversational-voice-agent-duplex-2026 🎙️ Real-Time Conversational Voice Agent, Turn-Taking, Full-Duplex & Prosody SFT/DPO Dataset (2026) This repository contains the 100-Sample Production Teaser for the Real-Time Conversational Voice Agent & Full-Duplex Prosody Suite (2026) by BeatsProm AI Research Lab. The dataset is engineered to train open-weights language models (Qwen-2.5-Audio, Llama-3.1-Voice, Moshi, Mini-Omni, Whisper-LLM) into ultra-low latency, real-time conversational voice agents featuring sub-150ms… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/realtime-conversational-voice-agent-duplex-2026.texttext-generationn<1K0 likes69 downloads20d agoHugging Face03beatsprom /audio-speech-realtime-voice-agents-2026 🎙️ Audio, Speech Foundation Models & Real-Time Voice Agents Dataset (2026 Edition) A structured research dataset featuring 1,722 domain-verified research papers and 298 official code repositories focused on Full-Duplex Speech-to-Speech LLMs, Real-Time Voice Agents (<200ms Latency), Zero-Shot TTS, Voice Cloning, OpenAI Whisper-v3, Neural Audio Codecs (EnCodec/DAC/SNAC), and Generative Music (2023–2026). Built with Universal Scientific Engine V18.1 Diamond, providing 48 schema… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/audio-speech-realtime-voice-agents-2026.tabularaudio-to-audion<1K0 likes54 downloads1mo agoHugging Face04arcada-labs /audio-agent-bench-suite Audio Agent Bench Suite A suite of six multi-turn, multi-domain spoken conversational benchmarks for evaluating voice AI and audio agent systems. Each sub-dataset targets a distinct real-world deployment domain, together covering the core capabilities required of production audio agents: instruction following, knowledge-base grounding, tool/function-call accuracy, long-range conversational memory, and state tracking. Sub-datasets Dataset Domain Turns HuggingFace… See the full description on the dataset page: https://huggingface.co/datasets/arcada-labs/audio-agent-bench-suite.automatic-speech-recognition0 likes12 downloads5mo agoHugging Face05AgentPublic /eval-stt-officiels EvalSTT — Corpus officiels (FR) Corpus public d'évaluation de modèles de transcription (speech-to-text) sur du langage de l'administration française : discours officiels, allocutions publiques et questions au gouvernement. Constitué par le département IA dans l'État (DINUM) dans le cadre de l'évaluation des modèles de speech-to-text. Ce dataset est publié pour la transparence : il documente les jeux de données utilisés pour nos évaluations et permet de reproduire les mesures… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/eval-stt-officiels.audioautomatic-speech-recognitionn<1K1 likes12 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.