CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01asahi417 /seamless-align-enA-viA.speaker-embedding.xlsr-2btabular100K<n<1M0 likes20k downloads2y agoHugging Face02asahi417 /seamless-align-enA-esA.speaker-embedding.w2vbert-600mtabular100K<n<1M0 likes19k downloads2y agoHugging Face03asahi417 /seamless-align-enA-frA.speaker-embedding.hubert-xltabular1M<n<10M0 likes16k downloads2y agoHugging Face04asahi417 /seamless-align-enA-jaA.speaker-embedding.w2vbert-600mtabular100K<n<1M0 likes15k downloads2y agoHugging Face05asahi417 /seamless-align-deA-enA.speaker-embedding.xlsr-2btabular100K<n<1M0 likes11k downloads2y agoHugging Face06asahi417 /seamless-align-enA-hiA.speaker-embedding.hubert-xltabular100K<n<1M0 likes10k downloads2y agoHugging Face07asahi417 /seamless-align-enA-frA.speaker-embedding.xlsr-2btabular100K<n<1M0 likes10k downloads2y agoHugging Face08asahi417 /seamless-align-enA-esA.speaker-embedding.xlsr-2btabular100K<n<1M0 likes9.8k downloads2y agoHugging Face09asahi417 /seamless-align-enA-zhA.speaker-embedding.w2vbert-600mtabular100K<n<1M0 likes9.7k downloads2y agoHugging Face10asahi417 /seamless-align-enA-frA.speaker-embedding.w2vbert-600mtabular1M<n<10M0 likes8.9k downloads2y agoHugging Face11nguyenvulebinh /asr-alignment Speech Recognition Alignment Dataset This dataset is a variation of several widely-used ASR datasets, encompassing Librispeech, MuST-C, TED-LIUM, VoxPopuli, Common Voice, and GigaSpeech. The difference is this dataset includes: Precise alignment between audio and text. Text that has been punctuated and made case-sensitive. Identification of named entities in the text. Usage First, install the latest version of the 🤗 Datasets package: pip install --upgrade pip pip… See the full description on the dataset page: https://huggingface.co/datasets/nguyenvulebinh/asr-alignment.audio10M<n<100M5 likes8.5k downloads3y agoHugging Face12asahi417 /seamless-align-enA-zhA.speaker-embedding.xlsr-2btabular100K<n<1M0 likes8.2k downloads2y agoHugging Face13asahi417 /seamless-align-enA-zhA.speaker-embedding.hubert-xltabular100K<n<1M0 likes7k downloads2y agoHugging Face14asahi417 /seamless-align-enA-koA.speaker-embedding.w2vbert-600mtabular100K<n<1M0 likes6.9k downloads2y agoHugging Face15asahi417 /seamless-align-enA-hiA.speaker-embedding.w2vbert-600mtabular100K<n<1M0 likes6.5k downloads2y agoHugging Face16dsfsi /govza-sa-cabinet-statements-sentence-aligned Gov-ZA Multilingual Cabinet Statements (Sentence-Aligned) Dataset Description This dataset contains sentence-aligned parallel text from South African government cabinet statements in 11 official languages. The data is sourced from the Government Communication and Information System (GCIS) and scraped from www.gov.za/cabinet-statements. Key Features: 📊 55 language pair combinations covering 11 South African languages 🔗 Sentence-level alignment using LASER embeddings 📈… See the full description on the dataset page: https://huggingface.co/datasets/dsfsi/govza-sa-cabinet-statements-sentence-aligned.texttranslation100K<n<1M1 likes6.3k downloads9mo agoHugging Face17asahi417 /seamless-align-enA-viA.speaker-embedding.w2vbert-600mtabular100K<n<1M0 likes6.3k downloads2y agoHugging Face18asahi417 /seamless-align-enA-jaA.speaker-embedding.hubert-xltabular100K<n<1M0 likes5.9k downloads2y agoHugging Face19asahi417 /seamless-align-enA-hiA.speaker-embedding.xlsr-2btabular10K<n<100K0 likes5.7k downloads2y agoHugging Face20asahi417 /seamless-align-enA-jaA.speaker-embedding.xlsr-2btabular10K<n<100K0 likes5.7k downloads2y agoHugging Face21takuM23 /multilingual_audio_alignments Multilingual MFA-Aligned Speech Dataset (UNDER DEVELOPMENT) A large-scale multilingual speech dataset with word-level and phoneme-level alignments produced using the Montreal Forced Aligner (MFA). Dataset Description This dataset consolidates multiple speech corpora across various languages, all processed through MFA to provide precise phoneme and word alignments. Each sample includes the original audio, transcript, and detailed timing information for both words and… See the full description on the dataset page: https://huggingface.co/datasets/takuM23/multilingual_audio_alignments.audioautomatic-speech-recognition10M<n<100M4 likes5.5k downloads6mo agoHugging Face22AlignmentResearch /DolusChattext10K<n<100K7 likes4.9k downloads1y agoHugging Face23AlignmentResearch /soft-trigger-verifiedtext1K<n<10K0 likes4.6k downloads9mo agoHugging Face24Magpie-Align /Magpie-Pro-300K-Filtered Project Web: https://magpie-align.github.io/ Arxiv Technical Report: https://arxiv.org/abs/2406.08464 Codes: https://github.com/magpie-align/magpie Abstract Click Here High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Pro-300K-Filtered.text100K<n<1M56 likes4.3k downloads2y agoHugging Face25asahi417 /seamless-align-enA-koA.speaker-embedding.hubert-xltabular100K<n<1M0 likes4k downloads2y agoHugging Face26Emova-ollm /emova-alignment-7m EMOVA-Alignment-7M 🤗 EMOVA-Models | 🤗 EMOVA-Datasets | 🤗 EMOVA-Demo 📄 Paper | 🌐 Project-Page | 💻 Github | 💻 EMOVA-Speech-Tokenizer-Github Overview EMOVA-Alignment-7M is a comprehensive dataset curated for omni-modal pre-training, including vision-language and speech-language alignment. This dataset is created using open-sourced image-text pre-training datasets, OCR datasets, and 2,000 hours of ASR and TTS data. This dataset is part of the EMOVA-Datasets… See the full description on the dataset page: https://huggingface.co/datasets/Emova-ollm/emova-alignment-7m.imageimage-to-text1M<n<10M10 likes3.5k downloads2y agoHugging Face27asahi417 /seamless-align-deA-enA.speaker-embedding.w2vbert-600mtabular100K<n<1M0 likes3.3k downloads2y agoHugging Face28AAdonis /multilingual_audio_alignments Multilingual MFA-Aligned Speech Dataset A large-scale multilingual speech dataset with word-level and phoneme-level alignments produced using the Montreal Forced Aligner (MFA). Dataset Description This dataset consolidates multiple speech corpora across various languages, all processed through MFA to provide precise phoneme and word alignments. Each sample includes the original audio, transcript, and detailed timing information for both words and phonemes.… See the full description on the dataset page: https://huggingface.co/datasets/AAdonis/multilingual_audio_alignments.audioautomatic-speech-recognition10M<n<100M27 likes3.3k downloads5mo agoHugging Face29asahi417 /seamless-align-enA-koA.speaker-embedding.xlsr-2btabular10K<n<100K0 likes3.3k downloads2y agoHugging Face30cfierro /alignment_faking_claude_completionstext1K<n<10K0 likes3.2k downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.