datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
audio-3-speaker-dataset-v2
Three-Speaker Audio Dataset with Timbral Speaker Embeddings
A teaching dataset maintained by AI-Academy. It pairs single-speaker English
utterances with precomputed timbral speaker embeddings, and is intended for
coursework and exercises rather than for benchmarking or production systems.
The dataset deliberately contains one injected inconsistency; locating it is one of
the intended exercises (see The injected anomaly).
Overview
Property
Value
Examples… See the full description on the dataset page: https://huggingface.co/datasets/aiacademy-kg/audio-3-speaker-dataset-v2.AI_Agent_Task_Dataset
🤖 Massive AI Agent Task Dataset (10.5GB)
📌 Overview
Welcome to the AI Agent Task Dataset, a massive 10.5GB procedural dataset designed for training, fine-tuning, and evaluating autonomous AI agents and LLMs.
This dataset focuses on:
Multi-step reasoning
Tool usage (APIs, frameworks, systems)
Real-world execution workflows
Perfect for building agentic AI systems, copilots, and automation models.
📑 Table of Contents
Dataset Details
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/oncody/AI_Agent_Task_Dataset.
