datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
korea_speech_mfa_aligned_validationgolos_mfa_punctuation
Golos MFA Punctuation
Расширенная версия датасета Golos —
русскоязычного корпуса речи с краудсорс и студийными записями.
Датасет дополнен пунктуацией и word-level временными метками (MFA alignment).
Опубликовано и поддерживается Jeti Labs.
Описание
Параметр
Значение
Язык
Русский (ru)
Записей
970,597
Аудио
~1,044 часов
Частота дискретизации
16,000 Hz
Формат
WAV, mono, 16-bit
Что добавлено по сравнению с оригинальным Golos… See the full description on the dataset page: https://huggingface.co/datasets/govnejri/golos_mfa_punctuation.emilia_mfa_correctgiga_mfa_correct_tonebooks-mfa-phonemes-only-hard-skazakh_speech_mfa_punctuation
Kazakh Speech MFA Punctuation
Расширенная версия датасета ISSAI KSC2 —
крупнейшего открытого корпуса казахской речи от института ISSAI (Nazarbayev University).
Датасет дополнен пунктуацией и word-level временными метками (MFA alignment).
Опубликовано и поддерживается Jeti Labs.
Описание
Параметр
Значение
Язык
Казахский (kk)
Записей
595,690
Аудио
~1,110 часов
Частота дискретизации
16,000 Hz
Формат
WAV, mono, 16-bit
Размер
52.9 GB… See the full description on the dataset page: https://huggingface.co/datasets/govnejri/kazakh_speech_mfa_punctuation.free_st_chinese_mandarin_corpus_mfa_aligneditalian_voxopopuli_mfataiwanspeech_mfagenshin_voice_v3.3_mandarin_mfa_alignedAISHELL_mandarin_processed_mfa_alignedbiobert-ner-fda-recalls-dataset
Dataset Card for FDA CDRH Device Recalls NER Dataset
This is a FDA Medical Device Recalls Dataset Created for Medical Device Named Entity Recognition (NER)
Dataset Details
Dataset Description
This dataset was created for the purpose of performing NER tasks.
It utilizes the OpenFDA Device Recalls dataset, which has been processed and annotated for performing NER.
The Device Recalls dataset has been further processed to extract the recall action element, which… See the full description on the dataset page: https://huggingface.co/datasets/mfarrington/biobert-ner-fda-recalls-dataset.librispeech_MFA_alignments
Dataset Card for "librispeech_MFA_alignments"
More Information needed
porjai_thai_voice_dataset_central_mfa_aligned_trainsimplevideo2gemini_flash_2.0_speech_puck_mfa_aligned_trainmfass
MFASS Splicing Variant Effects
This dataset packages 28,972 single-nucleotide variants from the Multiplexed
Functional Assay of Splicing (MFASS) as one compact benchmark table.
Each row contains the exact 170 bp transcript-oriented assay sequence pair,
native exon-inclusion measurements, assay-relative geometry, and canonical
GRCh38 locus. Of the 28,972 rows, 27,733 are evaluable and 1,050 are labeled
splice-disrupting variants.
Row identity: pair_id is the unique row key.… See the full description on the dataset page: https://huggingface.co/datasets/Taykhoom/mfass.korea_speech_mfa_aligned_validation_trainsimplevideosimplevideoshortsclean_english_mfa_aligned_800k-1200k_traingolos_mfa_punctuation_long
Golos MFA Punctuation (Long)
Long-form Russian speech derived from
govnejri/golos_mfa_punctuation.
Purpose
Most public Russian STT corpora ship as short clips (a few seconds each).
For benchmarking long-form transcription, VAD, punctuation, and streaming
behavior, you want minutes-long audio with reliable word-level alignments.
This dataset builds those long clips by splicing groups of consecutive
short clips together, inserting randomized silences between them, and… See the full description on the dataset page: https://huggingface.co/datasets/artmelancholy/golos_mfa_punctuation_long.polish_yodas_mfa_alignedasante-twi-mfa-wordlevel-tokenized
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
clean_english_mfa_aligned_0-400k_trainspeechocean_with_mfacv-corpus-17.0-zh-CN-client_id-grouped_mfamfact-classificationtestingMFA_tweet_topics
Dataset Card for "MFA_tweet_topics"
More Information needed
