CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AdoCleanCode /korea_speech_mfa_aligned_validationaudio100K<n<1M0 likes513 downloads8mo agoHugging Face02govnejri /golos_mfa_punctuation Golos MFA Punctuation Расширенная версия датасета Golos — русскоязычного корпуса речи с краудсорс и студийными записями. Датасет дополнен пунктуацией и word-level временными метками (MFA alignment). Опубликовано и поддерживается Jeti Labs. Описание Параметр Значение Язык Русский (ru) Записей 970,597 Аудио ~1,044 часов Частота дискретизации 16,000 Hz Формат WAV, mono, 16-bit Что добавлено по сравнению с оригинальным Golos… See the full description on the dataset page: https://huggingface.co/datasets/govnejri/golos_mfa_punctuation.audio100K<n<1M5 likes438 downloads5mo agoHugging Face03humairawan /emilia_mfa_correctaudio1M<n<10M1 likes295 downloads9mo agoHugging Face04humairawan /giga_mfa_correct_audio100K<n<1M0 likes248 downloads9mo agoHugging Face05qklent /tonebooks-mfa-phonemes-only-hard-saudio10K<n<100K0 likes217 downloads10mo agoHugging Face06govnejri /kazakh_speech_mfa_punctuation Kazakh Speech MFA Punctuation Расширенная версия датасета ISSAI KSC2 — крупнейшего открытого корпуса казахской речи от института ISSAI (Nazarbayev University). Датасет дополнен пунктуацией и word-level временными метками (MFA alignment). Опубликовано и поддерживается Jeti Labs. Описание Параметр Значение Язык Казахский (kk) Записей 595,690 Аудио ~1,110 часов Частота дискретизации 16,000 Hz Формат WAV, mono, 16-bit Размер 52.9 GB… See the full description on the dataset page: https://huggingface.co/datasets/govnejri/kazakh_speech_mfa_punctuation.audio100K<n<1M6 likes177 downloads1mo agoHugging Face07AdoCleanCode /free_st_chinese_mandarin_corpus_mfa_alignedaudio10K<n<100K0 likes128 downloads8mo agoHugging Face08AdoCleanCode /italian_voxopopuli_mfaaudio10K<n<100K0 likes106 downloads7mo agoHugging Face09AdoCleanCode /taiwanspeech_mfaaudio10K<n<100K1 likes93 downloads7mo agoHugging Face10AdoCleanCode /genshin_voice_v3.3_mandarin_mfa_alignedaudio10K<n<100K0 likes90 downloads8mo agoHugging Face11AdoCleanCode /AISHELL_mandarin_processed_mfa_alignedaudio100K<n<1M0 likes81 downloads8mo agoHugging Face12mfarrington /biobert-ner-fda-recalls-dataset Dataset Card for FDA CDRH Device Recalls NER Dataset This is a FDA Medical Device Recalls Dataset Created for Medical Device Named Entity Recognition (NER) Dataset Details Dataset Description This dataset was created for the purpose of performing NER tasks. It utilizes the OpenFDA Device Recalls dataset, which has been processed and annotated for performing NER. The Device Recalls dataset has been further processed to extract the recall action element, which… See the full description on the dataset page: https://huggingface.co/datasets/mfarrington/biobert-ner-fda-recalls-dataset.texttext-classification1K<n<10K3 likes78 downloads2y agoHugging Face13anyspeech /librispeech_MFA_alignments Dataset Card for "librispeech_MFA_alignments" More Information needed text100K<n<1M0 likes75 downloads3y agoHugging Face14AdoCleanCode /porjai_thai_voice_dataset_central_mfa_aligned_traintext100K<n<1M0 likes68 downloads9mo agoHugging Face15mfarre /simplevideo2text100K<n<1M1 likes53 downloads2y agoHugging Face16AdoCleanCode /gemini_flash_2.0_speech_puck_mfa_aligned_traintext100K<n<1M0 likes51 downloads9mo agoHugging Face17Taykhoom /mfass MFASS Splicing Variant Effects This dataset packages 28,972 single-nucleotide variants from the Multiplexed Functional Assay of Splicing (MFASS) as one compact benchmark table. Each row contains the exact 170 bp transcript-oriented assay sequence pair, native exon-inclusion measurements, assay-relative geometry, and canonical GRCh38 locus. Of the 28,972 rows, 27,733 are evaluable and 1,050 are labeled splice-disrupting variants. Row identity: pair_id is the unique row key.… See the full description on the dataset page: https://huggingface.co/datasets/Taykhoom/mfass.tabular10K<n<100K0 likes44 downloads1mo agoHugging Face18AdoCleanCode /korea_speech_mfa_aligned_validation_traintext100K<n<1M0 likes37 downloads9mo agoHugging Face19mfarre /simplevideoimagen<1K0 likes24 downloads2y agoHugging Face20mfarre /simplevideoshortstext10K<n<100K0 likes24 downloads2y agoHugging Face21AdoCleanCode /clean_english_mfa_aligned_800k-1200k_traintext10K<n<100K0 likes22 downloads9mo agoHugging Face22artmelancholy /golos_mfa_punctuation_long Golos MFA Punctuation (Long) Long-form Russian speech derived from govnejri/golos_mfa_punctuation. Purpose Most public Russian STT corpora ship as short clips (a few seconds each). For benchmarking long-form transcription, VAD, punctuation, and streaming behavior, you want minutes-long audio with reliable word-level alignments. This dataset builds those long clips by splicing groups of consecutive short clips together, inserting randomized silences between them, and… See the full description on the dataset page: https://huggingface.co/datasets/artmelancholy/golos_mfa_punctuation_long.audioautomatic-speech-recognition1K<n<10K1 likes18 downloads4mo agoHugging Face23AdoCleanCode /polish_yodas_mfa_alignedaudion<1K0 likes17 downloads7mo agoHugging Face24ghananlpcommunity /asante-twi-mfa-wordlevel-tokenized This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. text10K<n<100K0 likes16 downloads3mo agoHugging Face25AdoCleanCode /clean_english_mfa_aligned_0-400k_traintext10K<n<100K0 likes15 downloads9mo agoHugging Face26andybi7676 /speechocean_with_mfaaudio1K<n<10K0 likes15 downloads9mo agoHugging Face27AdoCleanCode /cv-corpus-17.0-zh-CN-client_id-grouped_mfaaudio10K<n<100K0 likes15 downloads7mo agoHugging Face28yfqiu-nlp /mfact-classificationtextn<1K0 likes13 downloads3y agoHugging Face29mfarre /testingtextn<1K0 likes12 downloads7mo agoHugging Face30Bsbell21 /MFA_tweet_topics Dataset Card for "MFA_tweet_topics" More Information needed textn<1K1 likes11 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.