datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mualem-recitations-original
المصاحف القرآنية
مصاحف مجمعمة من القراء المتقنين لبناء نماذج ذكاء اصطناعي لخدمة القرآن الكريم. أنظر هنا لأكواد بناء قاعدة التلاوات القرآنية
البيانات الوصفية للمصاحف
ds = load_dataset('obadx/mualem-recitations-original', name='moshaf_metadata')['train']
وصف أوجه حفص
Attribute Name
Arabic Name
Values
Default Value
More Info
rewaya
الرواية
- hafs (حفص)
The type of the quran Rewaya.
recitation_speed
سرعة التلاوة
- mujawad (مجود)-… See the full description on the dataset page: https://huggingface.co/datasets/obadx/mualem-recitations-original.or_in_datasetin-the-grooveCompiled from several different sets of songs:
(ITG) In the Groove
(ITG) In the Groove 2
Songs were downloaded from https://search.stepmaniaonline.net/packs/in+the+groove and are stored here for persistence.
In The Groove/ITG typically refers to DDR beatmaps done with an eye towards pad play.
Dataset info: https://paperswithcode.com/dataset/itg
tedlium-originalotoSpeech-full-duplex-task-oriented-20h
Dataset Viewer
https://cc-task-oriented-preview.vercel.app/
Task Walkthrough
https://www.oto.earth/research/task-oriented-dataset.html
What each of the seven tasks is for, what the two speakers could each see, and
how the interaction log lines up with the audio.
otoSpeech-full-duplex-task-oriented-20h
Contact
This sample dataset is provided for research purposes. We maintain larger and
more diverse datasets.
For collaborations… See the full description on the dataset page: https://huggingface.co/datasets/otoearth/otoSpeech-full-duplex-task-oriented-20h.orig-plus-asr-tamil-clean
orig-plus-asr-tamil-clean
Combined ASR dataset built from:
albagon/til26-asr-split (orig rows)
whyismydininghallonfire/asr-tamil-clean (asr_tamil_clean rows)
Audio paths are namespaced under each split to avoid filename collisions:
audio/orig/...
audio/asr_tamil_clean/...
Each row keeps key, audio, transcript, and language, with an added source_dataset field.
Counts:
train: 3595 orig + 891 asr_tamil_clean = 4486
validation: 899 orig + 224 asr_tamil_clean = 1123
iemocap-original-wavlm-large-layer-9-temporalorislop-ami-lipsync-data
Orislop AMI lip-sync source data
This repository mirrors an acquisition-capped subset of the official
AMI Meeting Corpus for reproducible
Orislop audio-visual synchronization research.
It contains:
low-size AMI close-up AVI video streams;
corresponding individual headset WAV streams;
SHA-256 provenance records;
the official camera/headset mapping snapshot and generated clip manifests
after local acquisition completes.
Files are uploaded only after a local download finishes and… See the full description on the dataset page: https://huggingface.co/datasets/gonnerthetooner/orislop-ami-lipsync-data.iemocap-original-wavlm-layer-6-temporaloriginal_data_manipuri_ttsoriginal_data_gujrati_ttsoriginal-songs
Dataset Card for "original-songs" (Audio + análisis DSP)
Dataset Summary
Dataset pequeño de canciones originales creadas con IA, cada una con su WAV,
letra transcrita automáticamente (Whisper) y un análisis DSP completo (tempo,
tonalidad, loudness, features perceptuales) además de detección de contenido
explícito. Pensado para quien quiera mejorar modelos open source: extracción
de features musicales, clasificación de audio, transcripción y moderación de
letras.… See the full description on the dataset page: https://huggingface.co/datasets/arnauquest/original-songs.Ndizi_parler_data_prep_originaloriginal_data_tamil_ttsoriginal_data_bengali_ttsoriginal_data_odia_ttsoriginal_data_kannada_ttsNdizi_TTS_origoriginal_data_telegu_ttsivan_shamyakin_tryvozhnae_shchastse_output_original
Трывожнае шчасце — арыгінальнае аўдыё
Аўтар / Author: Іван ШамякінМова / Language: Беларуская (Belarusian)
Арыгінальнае аўдыё без апрацоўкі, захаванае ў зыходнай якасці.
Частка калекцыі Ministerskija —
корпус беларускіх аўдыёкніг.
Апрацаваная версія (сегменты ~15 с, выраўнаваная транскрыпцыя):
ivan_shamyakin_tryvozhnae_shchastse_output
Доўгасць аўдыё
30h45m
Радкоў у датасеце
6,523
Структура
Кожны радок змяшчае:
audio — арыгінальны аўдыёзапіс… See the full description on the dataset page: https://huggingface.co/datasets/fosters/ivan_shamyakin_tryvozhnae_shchastse_output_original.mualem-recitations-original
المصاحف القرآنية
مصاحف مجمعمة من القراء المتقنين لبناء نماذج ذكاء اصطناعي لخدمة القرآن الكريم. أنظر هنا لأكواد بناء قاعدة التلاوات القرآنية
البيانات الوصفية للمصاحف
ds = load_dataset('obadx/mualem-recitations-original', name='moshaf_metadata')['train']
وصف أوجه حفص
Attribute Name
Arabic Name
Values
Default Value
More Info
rewaya
الرواية
- hafs (حفص)
The type of the quran Rewaya.
recitation_speed
سرعة التلاوة
- mujawad (مجود)-… See the full description on the dataset page: https://huggingface.co/datasets/nour-world/mualem-recitations-original.original_data_assamese_ttsVoxceleb1_test_original
Dataset Card for "Voxceleb1"
More Information needed
original_data_malayalam_ttsoriginal_data_rajasthani_ttskuzma_chorny_zyamlya_output_original
Зямля — арыгінальнае аўдыё
Аўтар / Author: Кузьма ЧорныМова / Language: Беларуская (Belarusian)
Арыгінальнае аўдыё без апрацоўкі, захаванае ў зыходнай якасці.
Частка калекцыі Ministerskija —
корпус беларускіх аўдыёкніг.
Апрацаваная версія (сегменты ~15 с, выраўнаваная транскрыпцыя):
kuzma_chorny_zyamlya_output
Доўгасць аўдыё
28h15m
Радкоў у датасеце
6,820
Структура
Кожны радок змяшчае:
audio — арыгінальны аўдыёзапіс
text — транскрыпцыя… See the full description on the dataset page: https://huggingface.co/datasets/fosters/kuzma_chorny_zyamlya_output_original.original_data_marathi_ttskuzma_chorny_poshuki_buduchyni_output_original
Пошукі будучыні — арыгінальнае аўдыё
Аўтар / Author: Кузьма ЧорныМова / Language: Беларуская (Belarusian)
Арыгінальнае аўдыё без апрацоўкі, захаванае ў зыходнай якасці.
Частка калекцыі Ministerskija —
корпус беларускіх аўдыёкніг.
Апрацаваная версія (сегменты ~15 с, выраўнаваная транскрыпцыя):
kuzma_chorny_poshuki_buduchyni_output
Доўгасць аўдыё
22h58m
Радкоў у датасеце
5,530
Структура
Кожны радок змяшчае:
audio — арыгінальны аўдыёзапіс
text —… See the full description on the dataset page: https://huggingface.co/datasets/fosters/kuzma_chorny_poshuki_buduchyni_output_original.wpp_pav_originalrasa_san_original
