datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Edge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Edge-Agent-Reasoning-WebSearch-260K.Omnimodal-Agent-SFT-2K
OmniGAIA: Omni-Modal General AI Assistant Benchmark
📄 Paper
•
💻 Code & Demo
•
🤗 Dataset & Model
•
📈 Leaderboard
This dataset contains omni-modal agent supervised fine-tuning (SFT) trajectories in the LlamaFactory SFT data format. You can directly follow LlamaFactory's instructions to fine-tune your omni-modal LLMs.OmniGAIA is a benchmark for Omni-Modal General AI Assistants that jointly reason over vision, audio, and language with external tools. It is… See the full description on the dataset page: https://huggingface.co/datasets/RUC-NLPIR/Omnimodal-Agent-SFT-2K.Barkopedia_DOG_AGE_GROUP_CLASSIFICATION_DATASET
Dataset
Check Training Data here: ArlingtonCL2/Barkopedia_DOG_AGE_GROUP_CLASSIFICATION_DATASET
Dataset Description
This dataset is for Dog Age Group Classification and contains dog bark audio clips. The data is split into training, public test, and private test sets.
Training set: 17888 audio clips.
Test set: 4920 audio clips, further divided into:
Test Public (~40%): 1966 audio clips for live leaderboard updates.
Test Private (~60%): 2954 audio clips for final evaluation.
You will… See the full description on the dataset page: https://huggingface.co/datasets/ArlingtonCL2/Barkopedia_DOG_AGE_GROUP_CLASSIFICATION_DATASET.agent-sft-stitch-zh-tts
agent-sft-stitch-zh-tts
Voiced version of voidful/agent-sft-stitch-zh: the STITCH-S spoken chunks synthesized with BlueMagpie-TTS (hung_yi_lee voice), per-utterance loudness-aligned to -23 LUFS, best-of-N + Whisper-CER accepted.
Configs
records (default): one row per agent dialogue — id/source/user/msg (full STITCH-S trajectory) + available_tools + STITCH quality scores + spoken (ordered list of the utterances, each with audio, text, seg_index, cer, accepted… See the full description on the dataset page: https://huggingface.co/datasets/voidful/agent-sft-stitch-zh-tts.agent-tts-libraryAudio-Video-Engineering-Agentic-Tasks-1M
Audio/Video Engineering Agentic Tasks (1M)
Abstract
A highly specialized dataset comprising 1,029,459 in-context troubleshooting prompts and execution commands built for the deepest levels of media production. Unlike standard datasets that simulate clean, theoretical instructions, this matrix captures the chaotic, highly-detailed, and conversational reality of professional audio engineers, composers, and video editors mid-session. It is engineered to train multimodal AI… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Audio-Video-Engineering-Agentic-Tasks-1M.Creative-Professionals-Agentic-Tasks-1M
Creative Professionals Agentic Tasks (1M)
Abstract
A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Creative-Professionals-Agentic-Tasks-1M.Barkopedia_DOG_AGE_GROUP_CLASSIFICATION_DATASET
Dataset
Check Training Data here: ArlingtonCL2/Barkopedia_DOG_AGE_GROUP_CLASSIFICATION_DATASET
Dataset Description
This dataset is for Dog Age Group Classification and contains dog bark audio clips. The data is split into training, public test, and private test sets.
Training set: 17888 audio clips.
Test set: 4920 audio clips, further divided into:
Test Public (~40%): 1966 audio clips for live leaderboard updates.
Test Private (~60%): 2954 audio clips for final evaluation.
You… See the full description on the dataset page: https://huggingface.co/datasets/hlx1021/Barkopedia_DOG_AGE_GROUP_CLASSIFICATION_DATASET.commonlanguage-age-minicommon-voice-17-en-age-gender-accentAgentChatEdge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/BlueIsGreen/Edge-Agent-Reasoning-WebSearch-260K.AgentWebBench-corpus
AgentWebBench Corpus
Pre-built dense-retrieval corpus for AgentWebBench [ICML 2026], a benchmark for Multi-Agent Coordination in Agentic Web over a realistic 100-website slice of
ClueWeb22 (~18.4M documents).
This repository holds the embeddings and FAISS indices the benchmark loads at run time, including per-website indices, a global index, and website-level vectors. It does not contain ClueWeb22 text (see Raw documents).
Websites: 100
Documents: ~18.4M
Embedding dim: 1024… See the full description on the dataset page: https://huggingface.co/datasets/cx-cmu/AgentWebBench-corpus.lectura-agents-data
LectūraAgents Dataset
Overview
This dataset is in support of findings in our paper LectūraAgents: A Multi-Agent Framework for Adaptive Personalized AI-Assisted Learning and Embodied Teaching. LectūaAgents is a hierarchical multi-agent framework that enables end-to-end personalized learning experiences through adaptive embodied teaching. It mirrors a professor–students’ relationship, wherein a ProfessorAgent guides a collaborative team of specialized subordinate… See the full description on the dataset page: https://huggingface.co/datasets/Jaward/lectura-agents-data.Edge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/DEMIRUNC/Edge-Agent-Reasoning-WebSearch-260K.Creative-Professionals-Agentic-Tasks-1M
Creative Professionals Agentic Tasks (1M)
Abstract
A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/rAVEUK/Creative-Professionals-Agentic-Tasks-1M.common-voice-17-en-age-genderEdge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/Torenn/Edge-Agent-Reasoning-WebSearch-260K.AgentChat-Test
Test Set Description
This directory contains the test set used for tool-use evaluation. The JSON files under Test-JSON/ are organized by task type:
SingleTaskProcessing/tool-select_test.json: single-tool selection tasks.
ParallelProcessing/parallel-call_test.json: parallel tool-call tasks.
ProactiveSeeking/searchTools_test_predictions_kept.json: proactive tool-search tasks.
TaskDecomposition/muti-tool-select_test.json: multi-tool task decomposition tasks.… See the full description on the dataset page: https://huggingface.co/datasets/leungtianle/AgentChat-Test.edge-agent-reasoning-websearch-260k
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/ppenner/edge-agent-reasoning-websearch-260k.Creative-Professionals-Agentic-Tasks-1M
Creative Professionals Agentic Tasks (1M)
Abstract
A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/kryp1234/Creative-Professionals-Agentic-Tasks-1M.Edge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/JACKYS999/Edge-Agent-Reasoning-WebSearch-260K.Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/kanepi-1977/Agent-Reasoning-WebSearch-260K.SECD
SECD: String Ensemble Chords Dataset
SECD is a large-scale controlled compositional audio dataset for research on string-ensemble harmony and performance attributes. It contains 287,088 harmonic-interval and chord instances constructed through additive superposition of professionally recorded isolated string notes from the Philharmonia Orchestra into duo-, trio-, and quartet-like four-voice mixtures.
The dataset is organised into six configurations covering harmonic intervals… See the full description on the dataset page: https://huggingface.co/datasets/ageroul/SECD.Edge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/svryn/Edge-Agent-Reasoning-WebSearch-260K.japanese-roleplay-travel-agency
Japanese Travel Agency Roleplay Dialogue Corpus
Overview
The Japanese Travel Agency Roleplay Dialogue Corpus is a collection of ten Japanese role-play dialogues simulating travel-agency consultations between a staff member and a customer.
The conversations were collected for speech and dialogue research and include two-speaker mixed audio, speaker-separated audio, and manually created, time-aligned ELAN (.eaf) annotation files.
The dialogues cover a variety of… See the full description on the dataset page: https://huggingface.co/datasets/IngCrowd/japanese-roleplay-travel-agency.agentic-asr
Agentic ASR
Public consolidated audio and ASR result dataset for the OSWorld and
WildClawBench benchmark families.
Layout
osworld/: synthetic raw/colloquial speech, human recordings, DNS-noise
pairs, task images, ASR results, and reports.
wildclawbench/: 60 formal colloquialized prompts, synthetic speech,
20 synthetic ASR condition tables, and ten-participant human recordings.
task0_template derivatives are excluded.
metadata/conditions.jsonl: model, variant… See the full description on the dataset page: https://huggingface.co/datasets/tterumiimurett1/agentic-asr.GLOBE_V3_age_N_allsplitscommon-voice-17-en-age-gender-accent-sampledresp-agent-dataset
Resp-229K: Respiratory Sound Dataset
A Large-Scale Respiratory Sound Dataset for Training and Evaluation
📖 Overview
Resp-229K is a comprehensive respiratory sound dataset containing 229,101 valid audio files with a total duration of over 407 hours. This dataset is curated for training the Resp-Agent system - an intelligent respiratory sound analysis and generation framework.
📊 Dataset Statistics
Split
Valid Files
Total Duration
Avg Duration
Max… See the full description on the dataset page: https://huggingface.co/datasets/AustinZhang/resp-agent-dataset.
