CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01clips /mfaqWe present the first multilingual FAQ dataset publicly available. We collected around 6M FAQ pairs from the web, in 21 different languages.tabularquestion-answering10M<n<100M37 likes1.3k downloads4y agoHugging Face02AdoCleanCode /korea_speech_mfa_aligned_validationaudio100K<n<1M0 likes455 downloads8mo agoHugging Face03govnejri /golos_mfa_punctuation Golos MFA Punctuation Расширенная версия датасета Golos — русскоязычного корпуса речи с краудсорс и студийными записями. Датасет дополнен пунктуацией и word-level временными метками (MFA alignment). Опубликовано и поддерживается Jeti Labs. Описание Параметр Значение Язык Русский (ru) Записей 970,597 Аудио ~1,044 часов Частота дискретизации 16,000 Hz Формат WAV, mono, 16-bit Что добавлено по сравнению с оригинальным Golos… See the full description on the dataset page: https://huggingface.co/datasets/govnejri/golos_mfa_punctuation.audio100K<n<1M5 likes402 downloads5mo agoHugging Face04zenmule /MFA_tutorial_2025-04-28_PAPPSThis repo contains the material for this Montreal Forced Aligner Tutorial. The recordings are from ALLSTAR and Mozilla Common Voice. audio1K<n<10K0 likes351 downloads1y agoHugging Face05humairawan /emilia_mfa_correctaudio1M<n<10M1 likes298 downloads9mo agoHugging Face06maximedb /mfaq_lightMQA is a multilingual corpus of questions and answers parsed from the Common Crawl. Questions are divided between Frequently Asked Questions (FAQ) pages and Community Question Answering (CQA) pages.tabular10M<n<100M0 likes294 downloads5y agoHugging Face07humairawan /giga_mfa_correct_audio100K<n<1M0 likes248 downloads9mo agoHugging Face08qklent /tonebooks-mfa-phonemes-only-hard-saudio10K<n<100K0 likes196 downloads10mo agoHugging Face09govnejri /kazakh_speech_mfa_punctuation Kazakh Speech MFA Punctuation Расширенная версия датасета ISSAI KSC2 — крупнейшего открытого корпуса казахской речи от института ISSAI (Nazarbayev University). Датасет дополнен пунктуацией и word-level временными метками (MFA alignment). Опубликовано и поддерживается Jeti Labs. Описание Параметр Значение Язык Казахский (kk) Записей 595,690 Аудио ~1,110 часов Частота дискретизации 16,000 Hz Формат WAV, mono, 16-bit Размер 52.9 GB… See the full description on the dataset page: https://huggingface.co/datasets/govnejri/kazakh_speech_mfa_punctuation.audio100K<n<1M6 likes178 downloads1mo agoHugging Face10Hemabhushan /ksbvs-mfa-synthesizer0 likes165 downloads2y agoHugging Face11open-llm-leaderboard-old /details_netcat420__MFANNv0.5 Dataset Card for Evaluation run of netcat420/MFANNv0.5 Dataset automatically created during the evaluation run of model netcat420/MFANNv0.5 on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_netcat420__MFANNv0.5.0 likes148 downloads2y agoHugging Face12open-llm-leaderboard-old /details_netcat420__MFANNv0.4 Dataset Card for Evaluation run of netcat420/MFANNv0.4 Dataset automatically created during the evaluation run of model netcat420/MFANNv0.4 on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_netcat420__MFANNv0.4.0 likes147 downloads2y agoHugging Face13omrifahn /mfa-vs-sae-2026-webapp-datatabularn<1K0 likes145 downloads8mo agoHugging Face14open-llm-leaderboard-old /details_netcat420__MFANN3bv0.7.100 likes134 downloads2y agoHugging Face15open-llm-leaderboard-old /details_netcat420__MFANN3bv0.70 likes127 downloads2y agoHugging Face16open-llm-leaderboard-old /details_netcat420__MFANN3bv0.3 Dataset Card for Evaluation run of netcat420/MFANN3bv0.3 Dataset automatically created during the evaluation run of model netcat420/MFANN3bv0.3 on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_netcat420__MFANN3bv0.3.0 likes112 downloads2y agoHugging Face17JudSacr /mfaqWe present the first multilingual FAQ dataset publicly available. We collected around 6M FAQ pairs from the web, in 21 different languages.question-answering0 likes106 downloads4mo agoHugging Face18anyspeech /librispeech_MFA_alignments Dataset Card for "librispeech_MFA_alignments" More Information needed text100K<n<1M0 likes91 downloads3y agoHugging Face19AdoCleanCode /free_st_chinese_mandarin_corpus_mfa_alignedaudio10K<n<100K0 likes90 downloads8mo agoHugging Face20Taykhoom /mfass MFASS Splicing Variant Effects This dataset packages 28,972 single-nucleotide variants from the Multiplexed Functional Assay of Splicing (MFASS) as one compact benchmark table. Each row contains the exact 170 bp transcript-oriented assay sequence pair, native exon-inclusion measurements, assay-relative geometry, and canonical GRCh38 locus. Of the 28,972 rows, 27,733 are evaluable and 1,050 are labeled splice-disrupting variants. Row identity: pair_id is the unique row key.… See the full description on the dataset page: https://huggingface.co/datasets/Taykhoom/mfass.tabular10K<n<100K0 likes88 downloads28d agoHugging Face21AdoCleanCode /italian_voxopopuli_mfaaudio10K<n<100K0 likes82 downloads7mo agoHugging Face22AdoCleanCode /AISHELL_mandarin_processed_mfa_alignedaudio100K<n<1M0 likes79 downloads7mo agoHugging Face23AdoCleanCode /genshin_voice_v3.3_mandarin_mfa_alignedaudio10K<n<100K0 likes76 downloads8mo agoHugging Face24AdoCleanCode /taiwanspeech_mfaaudio10K<n<100K1 likes75 downloads7mo agoHugging Face25mfarre /servedfilesimagen<1K0 likes72 downloads1y agoHugging Face26open-llm-leaderboard-old /details_netcat420__MFANNv0.2 Dataset Card for Evaluation run of netcat420/MFANNv0.2 Dataset automatically created during the evaluation run of model netcat420/MFANNv0.2 on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_netcat420__MFANNv0.2.0 likes65 downloads2y agoHugging Face27mfarrington /biobert-ner-fda-recalls-dataset Dataset Card for FDA CDRH Device Recalls NER Dataset This is a FDA Medical Device Recalls Dataset Created for Medical Device Named Entity Recognition (NER) Dataset Details Dataset Description This dataset was created for the purpose of performing NER tasks. It utilizes the OpenFDA Device Recalls dataset, which has been processed and annotated for performing NER. The Device Recalls dataset has been further processed to extract the recall action element, which… See the full description on the dataset page: https://huggingface.co/datasets/mfarrington/biobert-ner-fda-recalls-dataset.texttext-classification1K<n<10K3 likes60 downloads2y agoHugging Face28netcat420 /MFANNMFANN v2 Chain-of-Thought experiment text1K<n<10K5 likes57 downloads1y agoHugging Face29mfarre /simplevideo2text100K<n<1M1 likes53 downloads2y agoHugging Face30chunping-hf /MFA_env0 likes51 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.