CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nguyenvulebinh /asr-alignment Speech Recognition Alignment Dataset This dataset is a variation of several widely-used ASR datasets, encompassing Librispeech, MuST-C, TED-LIUM, VoxPopuli, Common Voice, and GigaSpeech. The difference is this dataset includes: Precise alignment between audio and text. Text that has been punctuated and made case-sensitive. Identification of named entities in the text. Usage First, install the latest version of the 🤗 Datasets package: pip install --upgrade pip pip… See the full description on the dataset page: https://huggingface.co/datasets/nguyenvulebinh/asr-alignment.audio10M<n<100M5 likes8.4k downloads3y agoHugging Face02PKU-Alignment /align-anything Overview: Align-Anything Dataset A Comprehensive All-Modality Alignment Dataset with Fine-grained Preference Annotations and Language Feedback. 🏠 Homepage | 🤗 Align-Anything Dataset | 🤗 T2T_Instruction-tuning Dataset | 🤗 TI2T_Instruction-tuning Dataset | 👍 Our Official Code Repo Our world is inherently multimodal. Humans perceive the world through multiple senses, and Language Models should operate similarly. However, the development of Current Multi-Modality Foundation Models… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/align-anything.audioany-to-any10K<n<100K48 likes6.2k downloads1y agoHugging Face03takuM23 /multilingual_audio_alignments Multilingual MFA-Aligned Speech Dataset (UNDER DEVELOPMENT) A large-scale multilingual speech dataset with word-level and phoneme-level alignments produced using the Montreal Forced Aligner (MFA). Dataset Description This dataset consolidates multiple speech corpora across various languages, all processed through MFA to provide precise phoneme and word alignments. Each sample includes the original audio, transcript, and detailed timing information for both words and… See the full description on the dataset page: https://huggingface.co/datasets/takuM23/multilingual_audio_alignments.audioautomatic-speech-recognition10M<n<100M4 likes6.1k downloads6mo agoHugging Face04AAdonis /multilingual_audio_alignments Multilingual MFA-Aligned Speech Dataset A large-scale multilingual speech dataset with word-level and phoneme-level alignments produced using the Montreal Forced Aligner (MFA). Dataset Description This dataset consolidates multiple speech corpora across various languages, all processed through MFA to provide precise phoneme and word alignments. Each sample includes the original audio, transcript, and detailed timing information for both words and phonemes.… See the full description on the dataset page: https://huggingface.co/datasets/AAdonis/multilingual_audio_alignments.audioautomatic-speech-recognition10M<n<100M27 likes3.4k downloads5mo agoHugging Face05gilkeyio /librispeech-alignments Dataset Card for Librispeech Alignments Librispeech with alignments generated by the Montreal Forced Aligner. The original alignments in TextGrid format can be found here Dataset Details Dataset Description Librispeech is a corpus of read English speech, designed for training and evaluating automatic speech recognition (ASR) systems. The dataset contains 1000 hours of 16kHz read English speech derived from audiobooks. The Montreal Forced Aligner (MFA) was used… See the full description on the dataset page: https://huggingface.co/datasets/gilkeyio/librispeech-alignments.audioautomatic-speech-recognition100K<n<1M21 likes2.4k downloads3y agoHugging Face06kyutai /interactivity-alignment-samples Audio Samples: Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models Audio samples accompanying the paper "Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models". Paper: arxiv.org Blog post: kyutai.org Models: 🤗 huggingface.co Overview This repository hosts the audio samples generated on Full-Duplex-Bench v1 (static evaluation with pre-recorded input) and Full-Duplex-Bench v2 (real-time multi-turn dialogue with GPT-Realtime), used in… See the full description on the dataset page: https://huggingface.co/datasets/kyutai/interactivity-alignment-samples.audio1K<n<10K9 likes407 downloads3mo agoHugging Face07QUD-Technologies /quran-alignment-benchmark Quran Recitation Alignment Benchmark Audio recordings of Quran recitation with a reviewed word-level ground truth: every recited word, in the order it was recited, with its start and end time, plus the reviewed segmentation and non-Quran regions. This is the corpus behind the Quran Recitation Alignment Benchmark; the task, scoring rules, leaderboard and submission format are documented there, not here. 16 recordings · 357 minutes · 18,421 recited words · Hafs ʿan ʿĀṣim ·… See the full description on the dataset page: https://huggingface.co/datasets/QUD-Technologies/quran-alignment-benchmark.audioautomatic-speech-recognitionn<1K0 likes274 downloads15d agoHugging Face08ErfanAShams /librispeech-alignments_clean100 librispeech-alignments_clean100 This is a subset of librispeech-alignments (https://huggingface.co/datasets/gilkeyio/librispeech-alignments) which only includes train_clean_100 and test_clean splits for small experiments and tutorials. Cite: @inproceedings{panayotov2015librispeech, title={Librispeech: an ASR corpus based on public domain audio books}, author={Panayotov, Vassil and Chen, Guoguo and Povey, Daniel and Khudanpur, Sanjeev}, booktitle={ICASSP}, year={2015}… See the full description on the dataset page: https://huggingface.co/datasets/ErfanAShams/librispeech-alignments_clean100.audioautomatic-speech-recognition10K<n<100K1 likes188 downloads1y agoHugging Face09giangndm /audio-confidence-alignment Vietnamese Wav2Vec2 Feature & K-Means Tokenized Dataset This repository contains the structured speech features and tokenized cluster indices for the target pw733 and clean viVoice Vietnamese datasets, formatted as Parquet tables. 📊 Dataset Schema audio_uuid (string): Unique identifier of the audio file. text (string): Transcription text (empty for raw pw733 audio). features (list of list of float): Frame-level Wav2Vec2 embeddings ([Num_Frames, 768]). indices… See the full description on the dataset page: https://huggingface.co/datasets/giangndm/audio-confidence-alignment.text100K<n<1M0 likes124 downloads2mo agoHugging Face10The-Nature-of-Reality /THE-BLUEPRINT-FOR-AI-ALIGNMENTaudion<1K3 likes101 downloads1y agoHugging Face11heihei /hachimi-alignment Hachimi Alignment Dataset Accompanying dataset for "When Meaning Fades: Probing Acoustic Properties in Audio-Text Alignment" (ACL 2025). Paper and code: github.com/ngyygm/hachimi-alignment What are Hachimi Songs? Hachimi (哈基米) songs are Chinese internet parody songs that replace original meaningful lyrics with nonsense syllables ("ha-ji-mi") while preserving melody, rhythm, and vocal timbre. This creates a natural experiment for probing what audio-text alignment models… See the full description on the dataset page: https://huggingface.co/datasets/heihei/hachimi-alignment.audiofeature-extractionn<1K0 likes96 downloads6mo agoHugging Face12Alignment-Lab-AI /librispeech-codec-22khzaudio10K<n<100K0 likes83 downloads9mo agoHugging Face13sujalappa /sample-force-alignment-datasetaudion<1K0 likes40 downloads9mo agoHugging Face14nguyenvulebinh /libris-asr-alignmentaudion<1K0 likes36 downloads3y agoHugging Face15Alignment-Lab-AI /dogaudio10K<n<100K0 likes23 downloads3y agoHugging Face16Tuyentd /Post-Training_Answer_Style_Alignmentaudion<1K0 likes21 downloads1y agoHugging Face17sujalappa /new-forced-alignment-datasetaudion<1K0 likes14 downloads8mo agoHugging Face18Tuyentd /Conversational_Response_Style_Alignment_Resultaudion<1K0 likes12 downloads11mo agoHugging Face19Alignment-Lab-AI /plskillmeiwantodieaudio10K<n<100K0 likes2 downloads2y agoHugging Face20Alignment-Lab-AI /podcast-1-test-preprocessedaudio1K<n<10K0 likes2 downloads2y agoHugging Face21AdoCleanCode /hifitts2_alignments_60_80gatedaudio10K<n<100K0 likes2 downloads10mo agoHugging Face22AdoCleanCode /hifitts2_alignments_01gatedaudion<1K0 likes1 downloads10mo agoHugging Face23AdoCleanCode /hifitts2_alignments_0030gatedaudio10K<n<100K0 likes1 downloads10mo agoHugging Face24AdoCleanCode /hifitts2_alignments_6090gatedaudio10K<n<100K0 likes1 downloads10mo agoHugging Face25AdoCleanCode /hifitts2_alignments_70_75gatedaudio10K<n<100K0 likes1 downloads10mo agoHugging Face26AdoCleanCode /hifitts2_alignments_60100gatedaudio10K<n<100K0 likes1 downloads10mo agoHugging Face27AdoCleanCode /hifitts2_alignments_10_15gatedaudio1K<n<10K0 likes1 downloads10mo agoHugging Face28AdoCleanCode /hifitts2_alignments_15_20gatedaudio1K<n<10K0 likes1 downloads10mo agoHugging Face29AdoCleanCode /hifitts2_alignments_05_10gatedaudio1K<n<10K0 likes1 downloads10mo agoHugging Face30AdoCleanCode /hifitts2_alignments_3060gatedaudio10K<n<100K0 likes1 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.