datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
asr-alignment
Speech Recognition Alignment Dataset
This dataset is a variation of several widely-used ASR datasets, encompassing Librispeech, MuST-C, TED-LIUM, VoxPopuli, Common Voice, and GigaSpeech. The difference is this dataset includes:
Precise alignment between audio and text.
Text that has been punctuated and made case-sensitive.
Identification of named entities in the text.
Usage
First, install the latest version of the 🤗 Datasets package:
pip install --upgrade pip
pip… See the full description on the dataset page: https://huggingface.co/datasets/nguyenvulebinh/asr-alignment.align-anything
Overview: Align-Anything Dataset
A Comprehensive All-Modality Alignment Dataset with Fine-grained Preference Annotations and Language Feedback.
🏠 Homepage | 🤗 Align-Anything Dataset | 🤗 T2T_Instruction-tuning Dataset | 🤗 TI2T_Instruction-tuning Dataset | 👍 Our Official Code Repo
Our world is inherently multimodal. Humans perceive the world through multiple senses, and Language Models should operate similarly. However, the development of Current Multi-Modality Foundation Models… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/align-anything.multilingual_audio_alignments
Multilingual MFA-Aligned Speech Dataset (UNDER DEVELOPMENT)
A large-scale multilingual speech dataset with word-level and phoneme-level alignments produced using the Montreal Forced Aligner (MFA).
Dataset Description
This dataset consolidates multiple speech corpora across various languages, all processed through MFA to provide precise phoneme and word alignments. Each sample includes the original audio, transcript, and detailed timing information for both words and… See the full description on the dataset page: https://huggingface.co/datasets/takuM23/multilingual_audio_alignments.multilingual_audio_alignments
Multilingual MFA-Aligned Speech Dataset
A large-scale multilingual speech dataset with word-level and phoneme-level alignments produced using the Montreal Forced Aligner (MFA).
Dataset Description
This dataset consolidates multiple speech corpora across various languages, all processed through MFA to provide precise phoneme and word alignments. Each sample includes the original audio, transcript, and detailed timing information for both words and phonemes.… See the full description on the dataset page: https://huggingface.co/datasets/AAdonis/multilingual_audio_alignments.librispeech-alignments
Dataset Card for Librispeech Alignments
Librispeech with alignments generated by the Montreal Forced Aligner. The original alignments in TextGrid format can be found here
Dataset Details
Dataset Description
Librispeech is a corpus of read English speech, designed for training and evaluating automatic speech recognition (ASR) systems. The dataset contains 1000 hours of 16kHz read English speech derived from audiobooks.
The Montreal Forced Aligner (MFA) was used… See the full description on the dataset page: https://huggingface.co/datasets/gilkeyio/librispeech-alignments.interactivity-alignment-samples
Audio Samples: Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models
Audio samples accompanying the paper "Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models".
Paper: arxiv.org
Blog post: kyutai.org
Models: 🤗 huggingface.co
Overview
This repository hosts the audio samples generated on Full-Duplex-Bench v1 (static evaluation with pre-recorded input) and Full-Duplex-Bench v2 (real-time multi-turn dialogue with GPT-Realtime), used in… See the full description on the dataset page: https://huggingface.co/datasets/kyutai/interactivity-alignment-samples.quran-alignment-benchmark
Quran Recitation Alignment Benchmark
Audio recordings of Quran recitation with a reviewed word-level ground truth: every recited word, in the order it was recited, with its start and end time, plus the reviewed segmentation and non-Quran regions. This is the corpus behind the Quran Recitation Alignment Benchmark; the task, scoring rules, leaderboard and submission format are documented there, not here.
16 recordings · 357 minutes · 18,421 recited words · Hafs ʿan ʿĀṣim ·… See the full description on the dataset page: https://huggingface.co/datasets/QUD-Technologies/quran-alignment-benchmark.librispeech-alignments_clean100
librispeech-alignments_clean100
This is a subset of librispeech-alignments (https://huggingface.co/datasets/gilkeyio/librispeech-alignments) which only includes train_clean_100 and test_clean splits for small experiments and tutorials.
Cite:
@inproceedings{panayotov2015librispeech,
title={Librispeech: an ASR corpus based on public domain audio books},
author={Panayotov, Vassil and Chen, Guoguo and Povey, Daniel and Khudanpur, Sanjeev},
booktitle={ICASSP},
year={2015}… See the full description on the dataset page: https://huggingface.co/datasets/ErfanAShams/librispeech-alignments_clean100.audio-confidence-alignment
Vietnamese Wav2Vec2 Feature & K-Means Tokenized Dataset
This repository contains the structured speech features and tokenized cluster indices for the target pw733 and clean viVoice Vietnamese datasets, formatted as Parquet tables.
📊 Dataset Schema
audio_uuid (string): Unique identifier of the audio file.
text (string): Transcription text (empty for raw pw733 audio).
features (list of list of float): Frame-level Wav2Vec2 embeddings ([Num_Frames, 768]).
indices… See the full description on the dataset page: https://huggingface.co/datasets/giangndm/audio-confidence-alignment.THE-BLUEPRINT-FOR-AI-ALIGNMENThachimi-alignment
Hachimi Alignment Dataset
Accompanying dataset for "When Meaning Fades: Probing Acoustic Properties in Audio-Text Alignment" (ACL 2025).
Paper and code: github.com/ngyygm/hachimi-alignment
What are Hachimi Songs?
Hachimi (哈基米) songs are Chinese internet parody songs that replace original meaningful lyrics with nonsense syllables ("ha-ji-mi") while preserving melody, rhythm, and vocal timbre. This creates a natural experiment for probing what audio-text alignment models… See the full description on the dataset page: https://huggingface.co/datasets/heihei/hachimi-alignment.librispeech-codec-22khzsample-force-alignment-datasetlibris-asr-alignmentdogPost-Training_Answer_Style_Alignmentnew-forced-alignment-datasetConversational_Response_Style_Alignment_Resultplskillmeiwantodiepodcast-1-test-preprocessedhifitts2_alignments_60_80hifitts2_alignments_01hifitts2_alignments_0030hifitts2_alignments_6090hifitts2_alignments_70_75hifitts2_alignments_60100hifitts2_alignments_10_15hifitts2_alignments_15_20hifitts2_alignments_05_10hifitts2_alignments_3060
