mfa
Datasets
All datasets matching “mfa”mfaqWe present the first multilingual FAQ dataset publicly available. We collected around 6M FAQ pairs from the web, in 21 different languages.korea_speech_mfa_aligned_validationgolos_mfa_punctuation
Golos MFA Punctuation
Расширенная версия датасета Golos —
русскоязычного корпуса речи с краудсорс и студийными записями.
Датасет дополнен пунктуацией и word-level временными метками (MFA alignment).
Опубликовано и поддерживается Jeti Labs.
Описание
Параметр
Значение
Язык
Русский (ru)
Записей
970,597
Аудио
~1,044 часов
Частота дискретизации
16,000 Hz
Формат
WAV, mono, 16-bit
Что добавлено по сравнению с оригинальным Golos… See the full description on the dataset page: https://huggingface.co/datasets/govnejri/golos_mfa_punctuation.MFA_tutorial_2025-04-28_PAPPSThis repo contains the material for this Montreal Forced Aligner Tutorial.
The recordings are from ALLSTAR and Mozilla Common Voice.
emilia_mfa_correctmfaq_lightMQA is a multilingual corpus of questions and answers parsed from the Common Crawl. Questions are divided between Frequently Asked Questions (FAQ) pages and Community Question Answering (CQA) pages.
