datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
seamless-align-enA-viA.speaker-embedding.xlsr-2bseamless-align-enA-esA.speaker-embedding.w2vbert-600mseamless-align-enA-frA.speaker-embedding.hubert-xlseamless-align-enA-jaA.speaker-embedding.w2vbert-600mseamless-align-deA-enA.speaker-embedding.xlsr-2bseamless-align-enA-hiA.speaker-embedding.hubert-xlseamless-align-enA-frA.speaker-embedding.xlsr-2bseamless-align-enA-esA.speaker-embedding.xlsr-2bseamless-align-enA-zhA.speaker-embedding.w2vbert-600mseamless-align-enA-frA.speaker-embedding.w2vbert-600masr-alignment
Speech Recognition Alignment Dataset
This dataset is a variation of several widely-used ASR datasets, encompassing Librispeech, MuST-C, TED-LIUM, VoxPopuli, Common Voice, and GigaSpeech. The difference is this dataset includes:
Precise alignment between audio and text.
Text that has been punctuated and made case-sensitive.
Identification of named entities in the text.
Usage
First, install the latest version of the 🤗 Datasets package:
pip install --upgrade pip
pip… See the full description on the dataset page: https://huggingface.co/datasets/nguyenvulebinh/asr-alignment.seamless-align-enA-zhA.speaker-embedding.xlsr-2bseamless-align-enA-zhA.speaker-embedding.hubert-xlseamless-align-enA-koA.speaker-embedding.w2vbert-600mseamless-align-enA-hiA.speaker-embedding.w2vbert-600mgovza-sa-cabinet-statements-sentence-aligned
Gov-ZA Multilingual Cabinet Statements (Sentence-Aligned)
Dataset Description
This dataset contains sentence-aligned parallel text from South African government cabinet statements in 11 official languages. The data is sourced from the Government Communication and Information System (GCIS) and scraped from www.gov.za/cabinet-statements.
Key Features:
📊 55 language pair combinations covering 11 South African languages
🔗 Sentence-level alignment using LASER embeddings
📈… See the full description on the dataset page: https://huggingface.co/datasets/dsfsi/govza-sa-cabinet-statements-sentence-aligned.seamless-align-enA-viA.speaker-embedding.w2vbert-600mseamless-align-enA-jaA.speaker-embedding.hubert-xlseamless-align-enA-hiA.speaker-embedding.xlsr-2bseamless-align-enA-jaA.speaker-embedding.xlsr-2bmultilingual_audio_alignments
Multilingual MFA-Aligned Speech Dataset (UNDER DEVELOPMENT)
A large-scale multilingual speech dataset with word-level and phoneme-level alignments produced using the Montreal Forced Aligner (MFA).
Dataset Description
This dataset consolidates multiple speech corpora across various languages, all processed through MFA to provide precise phoneme and word alignments. Each sample includes the original audio, transcript, and detailed timing information for both words and… See the full description on the dataset page: https://huggingface.co/datasets/takuM23/multilingual_audio_alignments.DolusChatsoft-trigger-verifiedMagpie-Pro-300K-Filtered
Project Web: https://magpie-align.github.io/
Arxiv Technical Report: https://arxiv.org/abs/2406.08464
Codes: https://github.com/magpie-align/magpie
Abstract
Click Here
High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Pro-300K-Filtered.seamless-align-enA-koA.speaker-embedding.hubert-xlemova-alignment-7m
EMOVA-Alignment-7M
🤗 EMOVA-Models | 🤗 EMOVA-Datasets | 🤗 EMOVA-Demo
📄 Paper | 🌐 Project-Page | 💻 Github | 💻 EMOVA-Speech-Tokenizer-Github
Overview
EMOVA-Alignment-7M is a comprehensive dataset curated for omni-modal pre-training, including vision-language and speech-language alignment.
This dataset is created using open-sourced image-text pre-training datasets, OCR datasets, and 2,000 hours of ASR and TTS data.
This dataset is part of the EMOVA-Datasets… See the full description on the dataset page: https://huggingface.co/datasets/Emova-ollm/emova-alignment-7m.seamless-align-deA-enA.speaker-embedding.w2vbert-600mmultilingual_audio_alignments
Multilingual MFA-Aligned Speech Dataset
A large-scale multilingual speech dataset with word-level and phoneme-level alignments produced using the Montreal Forced Aligner (MFA).
Dataset Description
This dataset consolidates multiple speech corpora across various languages, all processed through MFA to provide precise phoneme and word alignments. Each sample includes the original audio, transcript, and detailed timing information for both words and phonemes.… See the full description on the dataset page: https://huggingface.co/datasets/AAdonis/multilingual_audio_alignments.seamless-align-enA-koA.speaker-embedding.xlsr-2balignment_faking_claude_completions
