datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NOTSOFAR
Introduction
Welcome to the "NOTSOFAR-1: Distant Meeting Transcription with a Single Device" Challenge.
This repo contains the baseline system code for the NOTSOFAR-1 Challenge.
For more information about NOTSOFAR, visit CHiME's official challenge website
Register to participate.
Baseline system description.
Contact us: join the chime-8-notsofar channel on the CHiME Slack, or open a GitHub issue.
📊 Baseline Results on NOTSOFAR dev-set-1
Values are presented in… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/NOTSOFAR.RESOURCE2SKILL
Resource2Skill: Executable Agent Skill Libraries
This is the official Microsoft dataset release for
Resource2Skill, a system that
distills human-created multimodal resources into reusable executable skills for
software agents.
Project page: https://microsoft.github.io/Resource2Skill/
Paper: https://arxiv.org/abs/2606.29538
Code: https://github.com/microsoft/Resource2Skill
Contents
skills_wiki/ Structured skill entries used for discovery and inspection… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/RESOURCE2SKILL.microduck-emotions
Microduck Emotions
A collection of emotions for the Microduck robot. Each one is a motion and a sound designed together, beat by
beat, with the beak opening on the sound, rendered in the physics simulation and validated on the real robot. Every
emotion is three files: the motion (emotions/<name>.json, keyframes at 30 fps: head and body offsets played on
top of whichever trained policy is active, plus the policy hand-overs, such as the sit that devastated and play dead
start)… See the full description on the dataset page: https://huggingface.co/datasets/pollen-robotics/microduck-emotions.microsoft-speech-corpus-indian
Microsoft Speech Corpus – Indian Languages
Dataset Description
This dataset is a redistribution of the Microsoft Speech Corpus (Indian Languages) containing conversational and phrasal speech training and test data for Telugu, Tamil, and Gujarati languages. Each entry includes an audio recording and its corresponding transcript.
Attribution required: "Data provided by Microsoft and SpeechOcean.com"
⚠️ License: This data is provided for research purposes only. Commercial… See the full description on the dataset page: https://huggingface.co/datasets/deepdml/microsoft-speech-corpus-indian.nadi2026-adi20-micro-25pct-knnvc
NADI 2026 ADI20-micro — kNN-VC augmented (4 target voices)
Voice-converted copy of the 25% stratified subset (seed 42) of
UBC-NLP/NADI_2026_ADI20_micro, made with kNN-VC
following Abdullah et al. 2025.
Configs: voice_01–voice_04, 16,757 rows each, train split only.
Validation/test audio is deliberately left natural.
Target voices: 4 Arabic speakers from Common Voice (~60s each), gender-balanced,
the same set used across all dialects.
column
meaning
audio
converted… See the full description on the dataset page: https://huggingface.co/datasets/nadi-task2/nadi2026-adi20-micro-25pct-knnvc.NADI_2026_ADI20_micro
ADI-20 Micro
This is a smaller version of the ADI-20/ADI-17 Arabic dialect identification dataset use for the NADI 2026 shared task (although participants are encouraged to use the full ADI-20 for their submissions). This dataset consists of 10 hours per each dialect taken from the original ADI-17 alongside 10 hours for each of the new dialects introduced in ADI-20.
Dataset Sources
ADI-17: ArabicSpeech/ADI17
ADI-20: ArabicSpeech/ADI20
Papers: ADI-17, ADI-20… See the full description on the dataset page: https://huggingface.co/datasets/UBC-NLP/NADI_2026_ADI20_micro.MicroTex
📦 Sound Texture Dataset Collection
This dataset is a collection of three distinct sound texture datasets, prepared for machine learning tasks such as classification, generation, or analysis of audio textures. Each subset comes from different sources and includes metadata where applicable.
📁 Dataset Structure
The repository contains the following folders:
1. boreillysegmented16K_class/
This folder contains the BOReillySegmented16K dataset, based on the… See the full description on the dataset page: https://huggingface.co/datasets/cordutie/MicroTex.urban_sounds_micromic-rotate-simulation-v1
MicRotate Simulation Dataset
Android 스마트폰 마이크 회전(0° Portrait → 90° Landscape)에 따른
스테레오 오디오 변환 AI 모델 학습용 시뮬레이션 데이터셋.
Data Fields
Field
Type
Description
pair_id
int
페어 고유 ID
source_type
str
speech / sine_sweep / white_noise / pink_noise / environmental
audio_0deg_ch0
Audio
Portrait 0° MIC_TOP (48kHz)
audio_0deg_ch1
Audio
Portrait 0° MIC_BOT (48kHz)
audio_90deg_ch0
Audio
Landscape 90° MIC_TOP (48kHz)
audio_90deg_ch1
Audio
Landscape 90° MIC_BOT (48kHz)… See the full description on the dataset page: https://huggingface.co/datasets/haejin1320/mic-rotate-simulation-v1.Microsoft-AEC-Silero-VADmicrovent
microvent
A compact development set for video retrieval, claim extraction, and report
generation. It uses the same schema as the larger multivent-raw, so scripts
that target one transfer straight to the other.
This dataset card covers the core release: videos, audio, keyframes, and
the public evaluation annotations. Derived signals (OCR text, ASR transcripts,
visual / audio / video / omni embeddings) live in a companion release,
microvent-features, with its own dataset card… See the full description on the dataset page: https://huggingface.co/datasets/hltcoe/microvent.microvent-features
microvent-features
Derived signals for the microvent core release: per-keyframe OCR text,
per-chunk ASR transcripts, and an embedding zoo (keyframe-level vision,
keyframe-OCR text, audio-level, video-level, omni-modal).
This card covers only the features. For the source videos, audio,
keyframes, and the public eval annotations, see the microvent dataset
card. All artifacts here key on the same chunk_id and follow the same
WebDataset shard layout, so joining feature shards back… See the full description on the dataset page: https://huggingface.co/datasets/hltcoe/microvent-features.microsoft-AEC-datasetmicrosoft-AEC-vad-dataset
