datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mfaqWe present the first multilingual FAQ dataset publicly available. We collected around 6M FAQ pairs from the web, in 21 different languages.korea_speech_mfa_aligned_validationgolos_mfa_punctuation
Golos MFA Punctuation
Расширенная версия датасета Golos —
русскоязычного корпуса речи с краудсорс и студийными записями.
Датасет дополнен пунктуацией и word-level временными метками (MFA alignment).
Опубликовано и поддерживается Jeti Labs.
Описание
Параметр
Значение
Язык
Русский (ru)
Записей
970,597
Аудио
~1,044 часов
Частота дискретизации
16,000 Hz
Формат
WAV, mono, 16-bit
Что добавлено по сравнению с оригинальным Golos… See the full description on the dataset page: https://huggingface.co/datasets/govnejri/golos_mfa_punctuation.MFA_tutorial_2025-04-28_PAPPSThis repo contains the material for this Montreal Forced Aligner Tutorial.
The recordings are from ALLSTAR and Mozilla Common Voice.
emilia_mfa_correctmfaq_lightMQA is a multilingual corpus of questions and answers parsed from the Common Crawl. Questions are divided between Frequently Asked Questions (FAQ) pages and Community Question Answering (CQA) pages.giga_mfa_correct_tonebooks-mfa-phonemes-only-hard-skazakh_speech_mfa_punctuation
Kazakh Speech MFA Punctuation
Расширенная версия датасета ISSAI KSC2 —
крупнейшего открытого корпуса казахской речи от института ISSAI (Nazarbayev University).
Датасет дополнен пунктуацией и word-level временными метками (MFA alignment).
Опубликовано и поддерживается Jeti Labs.
Описание
Параметр
Значение
Язык
Казахский (kk)
Записей
595,690
Аудио
~1,110 часов
Частота дискретизации
16,000 Hz
Формат
WAV, mono, 16-bit
Размер
52.9 GB… See the full description on the dataset page: https://huggingface.co/datasets/govnejri/kazakh_speech_mfa_punctuation.ksbvs-mfa-synthesizerdetails_netcat420__MFANNv0.5
Dataset Card for Evaluation run of netcat420/MFANNv0.5
Dataset automatically created during the evaluation run of model netcat420/MFANNv0.5 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_netcat420__MFANNv0.5.details_netcat420__MFANNv0.4
Dataset Card for Evaluation run of netcat420/MFANNv0.4
Dataset automatically created during the evaluation run of model netcat420/MFANNv0.4 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_netcat420__MFANNv0.4.mfa-vs-sae-2026-webapp-datadetails_netcat420__MFANN3bv0.7.10details_netcat420__MFANN3bv0.7details_netcat420__MFANN3bv0.3
Dataset Card for Evaluation run of netcat420/MFANN3bv0.3
Dataset automatically created during the evaluation run of model netcat420/MFANN3bv0.3 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_netcat420__MFANN3bv0.3.mfaqWe present the first multilingual FAQ dataset publicly available. We collected around 6M FAQ pairs from the web, in 21 different languages.librispeech_MFA_alignments
Dataset Card for "librispeech_MFA_alignments"
More Information needed
free_st_chinese_mandarin_corpus_mfa_alignedmfass
MFASS Splicing Variant Effects
This dataset packages 28,972 single-nucleotide variants from the Multiplexed
Functional Assay of Splicing (MFASS) as one compact benchmark table.
Each row contains the exact 170 bp transcript-oriented assay sequence pair,
native exon-inclusion measurements, assay-relative geometry, and canonical
GRCh38 locus. Of the 28,972 rows, 27,733 are evaluable and 1,050 are labeled
splice-disrupting variants.
Row identity: pair_id is the unique row key.… See the full description on the dataset page: https://huggingface.co/datasets/Taykhoom/mfass.italian_voxopopuli_mfaAISHELL_mandarin_processed_mfa_alignedgenshin_voice_v3.3_mandarin_mfa_alignedtaiwanspeech_mfaservedfilesdetails_netcat420__MFANNv0.2
Dataset Card for Evaluation run of netcat420/MFANNv0.2
Dataset automatically created during the evaluation run of model netcat420/MFANNv0.2 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_netcat420__MFANNv0.2.biobert-ner-fda-recalls-dataset
Dataset Card for FDA CDRH Device Recalls NER Dataset
This is a FDA Medical Device Recalls Dataset Created for Medical Device Named Entity Recognition (NER)
Dataset Details
Dataset Description
This dataset was created for the purpose of performing NER tasks.
It utilizes the OpenFDA Device Recalls dataset, which has been processed and annotated for performing NER.
The Device Recalls dataset has been further processed to extract the recall action element, which… See the full description on the dataset page: https://huggingface.co/datasets/mfarrington/biobert-ner-fda-recalls-dataset.MFANNMFANN v2 Chain-of-Thought experiment
simplevideo2MFA_env
