datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
risale-sohbet-turkish-2risale-i-nur-sohbet
Risale-i Nur Sohbet
Prof. Dr. Şener Dilek’ten izin alındı.
Türkçe
Risale-i Nur sohbetlerini ses, ham ASR metni ve zaman hizalı segmentler hâlinde
birlikte sunan bağımsız bir veri kümesidir. İlk sürüm izinli ve doğrulanmış
sohbetleri içerir; kitap metni, grounded, çok dilli veya kitap seslendirme veri
kümelerine karıştırılmaz.
Kapsam
2095 sohbet, 954.66 saat 16 kHz mono FLAC ses
Aynı derslerin ölçülmüş 48 kHz kalite katmanı; 786 derste
seçici… See the full description on the dataset page: https://huggingface.co/datasets/risaleinur/risale-i-nur-sohbet.open-large-bengali-asr-data
Open Large Bengali ASR Data
This is a collection of publicly available ASR data for Bengali. It contains 5000 hours of audio. We have a filtering column called is_better to filter good-quality audio from the corpus. It is set based on the wer between original transcription and prediction taken from a Bengali-Wav2Vec2 model and word-per-second (wps).
Datasets:
commonvoice
risale-nur-audio
Risale-i Nur Audio–Text Corpus
Gerçek insan okumalarını, aynı satırdaki kaynak metinle birlikte sunan açık bir
ses–metin veri kümesidir. Yeni varsayılan audio-text yapılandırması 15 kitaptan
91.792 oynatılabilir klip ve 203,02 saat ses içerir. Metinler kanonik kaynaktan
değiştirilmeden alınır ve her kayıt byte-exact section_id alıntılarıyla bağlanır.
An open speech corpus pairing human readings with their source text in the same
row. The default audio-text config contains 91,792… See the full description on the dataset page: https://huggingface.co/datasets/risaleinur/risale-nur-audio.risalei-nur-text-audio
Risale-i Nur Text–Audio
Kaynak · Source: RNK Neşriyat — yazılı izinle · used with written permission.
Her satırda gerçek insan okuması ile o sesin kanonik metni birlikte bulunur.
Sesler dış bağlantı değildir: WAV baytları Parquet dosyalarının içindedir.
Kaynak sitesi veya başka bir ses sunucusu gerekmez.
Each row pairs a human reading with its canonical transcript. Audio is stored
as WAV bytes inside the Parquet files; no source website or external audio
server is required.… See the full description on the dataset page: https://huggingface.co/datasets/risaleinur/risalei-nur-text-audio.nepali_asr
Dataset Card for Nepali Asr Dataset
Dataset Summary
This dataset consists of over 5 hours (300+ minutes) of English speech audio collected from YouTube. The dataset is designed for automatic speech recognition (ASR) and speaker identification tasks. It features both male and female speakers, with approximately 60% of the samples from male voices and the remaining 40% from female voices. The dataset contains 35 distinct speakers, each with their audio segmented into… See the full description on the dataset page: https://huggingface.co/datasets/rishi70612/nepali_asr.filtered_nepali_male_dataset1risale-sohbet-turkish
YouTube Transkripsiyon Veri Seti
Veri Yapısı
audio/: MP3 dosyaları
transcripts/: Metin transkripsiyonları
srt/: Altyazı dosyaları
metadata/: Video bilgileri
database.json: Tüm videoların indeksi
Güncelleme Tarihi
2025-03-21
quebecois_canadian_french_datasetNEWARIanimalclap-dataset
AnimalCLAP
AnimalCLAP: Taxonomy-Aware Language-Audio Pretraining for Species Recognition and Trait InferenceICASSP 2026
AuthorsRisa Shinoda, Kaede Shiohara, Nakamasa Inoue, Hiroaki Santo, Fumio Okura
Overview
This dataset contains 701,020 animal sound recordings collected from:
iNaturalist
Xeno-Canto
Splits
HF Split
Original Split
Description
train
train
Training data (URL only)
validation
test
Validation data (URL only)
test
zero_shot… See the full description on the dataset page: https://huggingface.co/datasets/risashinoda/animalclap-dataset.ristey_mirwari
Riste Mirwari — audiobook speech (ristey_mirwari)
31,535 transcribed speech segments cut from a reading of the book ڕشتەی مرواری
(Rishtey Mirwari), with a speaker and gender column alongside each segment.
At a glance
Rows
31,535 (single train split)
Columns
audio, file, transcription, speaker, gender, book, is_gold_transcript
Parquet on disk
4.99 GB across 11 shards (5.50 GB uncompressed)
Row groups
29 per shard (1,000 rows per group)
Audio… See the full description on the dataset page: https://huggingface.co/datasets/razhan/ristey_mirwari.emotionintelligencevalidation_nepali_asr
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/rishi70612/validation_nepali_asr.training_dataset
Dataset Card for OpenSLR Nepali Large ASR Cleaned
Dataset Summary
This data set contains transcribed audio data for Nepali. The data set consists of flac files, and a TSV file. The file utt_spk_text.tsv contains a FileID, anonymized UserID and the transcription of audio in the file.
The data set has been manually quality-checked, but there might still be errors.
The audio files are sampled at a rate of 16KHz, and leading and trailing silences are trimmed using… See the full description on the dataset page: https://huggingface.co/datasets/rishi70612/training_dataset.cleaned-nepali-asr-datasetukrainian-tts-audiobook-pani-nina-parquet
Ukrainian TTS audiobook dataset Pani Nina (Parquet)
Segmented Ukrainian audiobook speech with aligned text, prepared for training and evaluating Text-to-Speech (TTS) models.
The dataset is published as Hugging Face-compatible Parquet shards so the Hub Dataset Preview can render an audio column.
The dataset was prepared using whisper and ffmpeg:
Whisper was used for transcription and approximate segment timing.
FFmpeg was used to slice audio into short utterances (roughly 2-10… See the full description on the dataset page: https://huggingface.co/datasets/rishchen/ukrainian-tts-audiobook-pani-nina-parquet.see-2-sound-evalWe sample images from Laion400M and the web to construct this small evaluation set.
owr_cv_albanian_testfudu_datasetmyst_wav_clean
Dataset Card for "myst_wav_clean"
This is a subset of the MyST dataset used in our research experiments. It is distributed under the original MyST dataset license, available here: https://catalog.ldc.upenn.edu/LDC2021S05. If you would like access to this subset, please provide proof of a valid license for the original dataset from the LDC.
Updated Email:
Please forward your query to this email for people who are trying to get access to this dataset:… See the full description on the dataset page: https://huggingface.co/datasets/rishabhjain16/myst_wav_clean.SkyDraxurbansound8K(card and dataset copied from https://www.kaggle.com/datasets/chrisfilo/urbansound8k)
This dataset contains 8732 labeled sound excerpts (<=4s) of urban sounds from 10 classes: air_conditioner, car_horn, children_playing, dog_bark, drilling, enginge_idling, gun_shot, jackhammer, siren, and street_music. The classes are drawn from the urban sound taxonomy. For a detailed description of the dataset and how it was compiled please refer to our paper.All excerpts are taken from field recordings… See the full description on the dataset page: https://huggingface.co/datasets/risan-raja-iitm/urbansound8K.pancake_datasetpfs_wav
Dataset Card for "pfstar_subset"
This is a subset of the PFSTAR dataset used in our research experiments. It is distributed under the original PFSTAR dataset license. For licensing details, please refer to the official ELRA catalogue entry: ELRA-S0140, and the associated publication: Batliner et al., Interspeech 2005.
If you would like access to this subset, please provide proof of a valid license or institutional access to the original PFSTAR dataset.
Updated Email:
Please forward… See the full description on the dataset page: https://huggingface.co/datasets/rishabhjain16/pfs_wav.cmu_wav
Dataset Card for "cmu_kids_subset"
This is a subset of the CMU Kids dataset used in our research experiments. It is distributed under the original CMU Kids dataset license, available here: LDC97S63. If you would like access to this subset, please provide proof of a valid license for the original dataset from the LDC.
Updated Email:
Please forward your query to this email for people who are trying to get access to this dataset: j_rishabh@outlook.com
Update… See the full description on the dataset page: https://huggingface.co/datasets/rishabhjain16/cmu_wav.myst_pfs
Dataset Card for "myst_pfs"
Contains MyST dataset and PFstar datset as mentioned in our paper:
R. Jain, A. Barcovschi, M. Y. Yiwere, P. Corcoran and H. Cucu, "Exploring Native and Non-Native English Child Speech Recognition With Whisper," in IEEE Access, vol. 12, pp. 41601-41610, 2024, doi: 10.1109/ACCESS.2024.3378738.
Dataset Card for myst_wav_clean
This is a subset of the MyST dataset used in our research experiments. It is distributed under the original MyST dataset… See the full description on the dataset page: https://huggingface.co/datasets/rishabhjain16/myst_pfs.Bubsyukrainian-tts-audiobook-pani-nina-parquet-old
Ukrainian TTS audiobook dataset (Parquet)
Segmented Ukrainian audiobook speech with aligned text, prepared for training and evaluating Text-to-Speech (TTS) models. The dataset is published as Hugging Face-compatible Parquet shards so the Hub Dataset Preview can render an audio column.
That was mabe by using whisper (https://github.com/openai/whisper) and ffmpeg (https://www.ffmpeg.org/), where with whisper we set start and end of voices + transcribe it and using ffmpeg slice into… See the full description on the dataset page: https://huggingface.co/datasets/rishchen/ukrainian-tts-audiobook-pani-nina-parquet-old.Teste
