CoolFace
11 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01WueNLP /belebele-fleurs Belebele-Fleurs Belebele-Fleurs is a dataset suitable to evaluate two core tasks: Multilingual Spoken Language Understanding (Listening Comprehension): For each spoken paragraph, the task is to answer a multiple-choice question. The question and four answer choices are provided in text form. Multilingual Long-Form Automatic Speech Recognition (ASR) with Diverse Speakers: By concatenating sentence-level utterances, long-form audio clips (ranging from 30 seconds to 1 minute 30… See the full description on the dataset page: https://huggingface.co/datasets/WueNLP/belebele-fleurs.audioaudio-classification10K<n<100K9 likes3k downloads2y agoHugging Face02facebook /2M-Belebele 2M-Belebele Highly-Multilingual Speech and American Sign Language Comprehension Dataset We introduce 2M-Belebele as the first highly multilingual speech and American Sign Language (ASL) comprehension dataset. Our dataset, which is an extension of the existing Belebele only-text dataset, covers 74 spoken languages at the intersection of Belebele and Fleurs, and one sign language (ASL). The speech dataset is built from aligning Belebele, Flores200 and Fleurs datasets as… See the full description on the dataset page: https://huggingface.co/datasets/facebook/2M-Belebele.tabularquestion-answering10K<n<100K13 likes2.2k downloads2y agoHugging Face03fosters /bely_klyck_all Белы Клык Аўтар / Author: Джэк ЛонданМова / Language: Беларуская (Belarusian) Аўдыё нарэзана з арыгінальнага запісу ў зыходнай частаце дыскрэтызацыі (native), мона, фрагменты да 30 секунд. Частка калекцыі Belarusian Audiobooks (native). Радкоў у датасеце 2,056 Працягласць 5 гадз 20 хв Частата дыскрэтызацыі 44100 Hz Каналы мона Даўжыня фрагмента да 30 с Структура Кожны радок змяшчае: audio — аўдыёфрагмент (native SR, мона, ≤30 с) text —… See the full description on the dataset page: https://huggingface.co/datasets/fosters/bely_klyck_all.audioautomatic-speech-recognition1K<n<10K0 likes33 downloads3mo agoHugging Face04BeliefEngines /podcast-transcripts Podcast Transcripts & Belief Graph Structured belief extractions, transcripts, speaker profiles, and embeddings mined from Bitcoin / crypto podcasts by the be-podcast-etl pipeline. Scale (snapshot 2026-04-21) Asset Count Episodes (manifests) 1,551 Podcasts 18 Speakers 876 Persons (enriched profiles) 3,915 Belief shards 66,453 Embeddings (1536-dim) 65,007 Matrices 62,882 Top podcasts: simply-bitcoin (375), the-bitcoin-matrix (264)… See the full description on the dataset page: https://huggingface.co/datasets/BeliefEngines/podcast-transcripts.texttext-generationn<1K0 likes26 downloads5mo agoHugging Face05fosters /bely_klyck_output Белы Клык Аўтар / Author: Джэк ЛонданМова / Language: Беларуская (Belarusian) Частка калекцыі Ministerskija — выраўнаваныя аўдыёзапісы беларускіх аўдыёкніг з транскрыпцыямі. Апублікаваных радкоў (HF) 1,863 Агулам у БД 4,831 Доўгасць аўдыё 12h36m Парог даверу ≥ 0.95 Структура Кожны радок змяшчае: audio — аўдыёфрагмент (~15 с) text — транскрыпцыя (Gemini + ASR выраўнаванне) chunk_uid — унікальны ідэнтыфікатар фрагмента Апрацоўка… See the full description on the dataset page: https://huggingface.co/datasets/fosters/bely_klyck_output.audioautomatic-speech-recognition1K<n<10K0 likes19 downloads4mo agoHugging Face06bellpepper0606 /reazonspeech-v2-based-yomi-inferredgated ReazonSpeech v2 読み推定データセット 利用制限 本データセットは ReazonSpeech v2 を元に作成した派生データセットです。以下の規約が適用されます。 ライセンス:CDLA-Sharing-1.0 日本国著作権法第30条の4(情報解析)の範囲内でのみ利用可能 それ以外の用途での使用は不可 本データセットにアクセスすることで、上記の条件に同意したものとみなされます。 概要 ReazonSpeech v2 all の音声・字幕データを元に、形態素解析および読み推定を行って作成したデータセットです。 著作権保護の観点から、元の字幕テキスト(表層形)は含まれていません。各レコードには、元データのエントリを参照するIDと、形態素解析結果から抽出した読み・品詞・表層形の文字数のみが保持されています。 解説記事:https://zenn.dev/bellpepper0606/articles/e028355e06ad2a データ構造… See the full description on the dataset page: https://huggingface.co/datasets/bellpepper0606/reazonspeech-v2-based-yomi-inferred.texttoken-classification10M<n<100M1 likes17 downloads5mo agoHugging Face07fosters /ales-krautsevich-karotkaia-gistoryia-belarusi-uladzimir-lisouski-output Кароткая гісторыя Беларусі Аўтар / Author: Алесь КраўцэвічДыктар / Narrator: Уладзімір ЛісоўскіМова / Language: Беларуская (Belarusian) Частка калекцыі Ministerskija — выраўнаваныя аўдыёзапісы беларускіх аўдыёкніг з транскрыпцыямі. Апублікаваных радкоў (HF) 675 Доўгасць аўдыё 2h21m Парог даверу ≥ 0.95 Структура Кожны радок змяшчае: audio — аўдыёфрагмент (~15 с) text — транскрыпцыя (Gemini + ASR выраўнаванне) chunk_uid — унікальны… See the full description on the dataset page: https://huggingface.co/datasets/fosters/ales-krautsevich-karotkaia-gistoryia-belarusi-uladzimir-lisouski-output.audioautomatic-speech-recognitionn<1K0 likes16 downloads3mo agoHugging Face08fosters /ales-krautsevich-karotkaia-gistoryia-belarusi-uladzimir-lisouski-output_original Кароткая гісторыя Беларусі — арыгінальнае аўдыё Аўтар / Author: Алесь КраўцэвічМова / Language: Беларуская (Belarusian) Арыгінальнае аўдыё без апрацоўкі, захаванае ў зыходнай якасці. Частка калекцыі Ministerskija — корпус беларускіх аўдыёкніг. Апрацаваная версія (сегменты ~15 с, выраўнаваная транскрыпцыя): ales-krautsevich-karotkaia-gistoryia-belarusi-uladzimir-lisouski-output Доўгасць аўдыё 2h21m Радкоў у датасеце 675 Структура Кожны радок… See the full description on the dataset page: https://huggingface.co/datasets/fosters/ales-krautsevich-karotkaia-gistoryia-belarusi-uladzimir-lisouski-output_original.audioautomatic-speech-recognitionn<1K0 likes14 downloads3mo agoHugging Face09Speech-data /Belarusian-Speech-Dataset 🎧 Belarusian Speech Dataset The Belarusian Speech Dataset is a high-quality speech audio dataset designed to provide structured and diverse audio data for AI systems focused on speech technologies and language preservation. It includes 118 hours of audio data across 571 files, delivered in MP3 and WAV formats, with a total size of 192 MB. This well-balanced audio dataset ensures reliable voice data, with 55% female and 45% male speakers, and an age distribution spanning from 18 to… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Belarusian-Speech-Dataset.audioautomatic-speech-recognitionn<1K0 likes13 downloads6mo agoHugging Face10fosters /bely_klyck_output_original Белы Клык — арыгінальнае аўдыё Аўтар / Author: Джэк ЛонданМова / Language: Беларуская (Belarusian) Арыгінальнае аўдыё без апрацоўкі, захаванае ў зыходнай якасці. Частка калекцыі Ministerskija — корпус беларускіх аўдыёкніг. Апрацаваная версія (сегменты ~15 с, выраўнаваная транскрыпцыя): bely_klyck_output Доўгасць аўдыё 12h36m Радкоў у датасеце 1,853 Структура Кожны радок змяшчае: audio — арыгінальны аўдыёзапіс text — транскрыпцыя chunk_uid —… See the full description on the dataset page: https://huggingface.co/datasets/fosters/bely_klyck_output_original.audioautomatic-speech-recognition1K<n<10K0 likes13 downloads4mo agoHugging Face11fosters /ales-krautsevich-karotkaia-gistoryia-belarusi-uladzimir-lisouski_all Кароткая гісторыя Беларусі Аўтар / Author: Алесь КраўцэвічМова / Language: Беларуская (Belarusian) Аўдыё нарэзана з арыгінальнага запісу ў зыходнай частаце дыскрэтызацыі (native), мона, фрагменты да 30 секунд. Частка калекцыі Belarusian Audiobooks (native). Радкоў у датасеце 675 Працягласць 2 гадз 16 хв Частата дыскрэтызацыі 44100 Hz Каналы мона Даўжыня фрагмента да 30 с Структура Кожны радок змяшчае: audio — аўдыёфрагмент (native SR… See the full description on the dataset page: https://huggingface.co/datasets/fosters/ales-krautsevich-karotkaia-gistoryia-belarusi-uladzimir-lisouski_all.audioautomatic-speech-recognitionn<1K0 likes9 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.