datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
persian-asr-audio-text-2.69M-chizzled
🗂️ persian-asr-audio-text-2.69M-chizzled
English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission
🌟 At a glance | معرفی سریع
English
فارسی
🎯 Purpose
Phase A-scale audio/text dataset.
پیکرهٔ بزرگ جفتهای صوت و متنِ پالایششده برای آموزش در مقیاس فاز A.
🧩 Role
Persian text and linguistic asset
مصنوع متنی و زبانی فارسی
📦 Snapshot
417 files; approximately 236.86 GB
417 فایل؛ حدود 236.86 GB
🧱 Packaging
414 Parquet files and 0… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-asr-audio-text-2.69M-chizzled.Videos-Dataset-For-LLMs-RAG-That-Require-Audio-Vidoes-And-Text
Dataset Overview
A collection of 27 domains (“topics”) and 3100 question-answer pair.
Each topic comes with average 117 QA pairs.Every QA entry comes with:
references: one or more source files the answer is extracted from
time with each reference comes the starting and ending time the answer is extracted from the reference
video_files: the video files where the answer can be found
(future) video title & description from metadata.csv
File structure
You-Are-Here!/… See the full description on the dataset page: https://huggingface.co/datasets/elmoghany/Videos-Dataset-For-LLMs-RAG-That-Require-Audio-Vidoes-And-Text.Tumbuka_Text-Speech_Audiotext_audio_pretraining_for_GAN
bw_jh_dataset — Robot 3D Point-Track Dataset (droid / hrdexdb / robocasa)
A unified multi-source robot manipulation dataset with 3D point tracks of the robot gripper and arm, calibrated multi-view RGB video, language instructions, and success labels. Three splits:
Split
Episodes
Size
Views
Source
droid
27,615 (+~1,050 in droid-11)
~280 GB
2 exterior ZED views
DROID (real, Franka)
hrdexdb
1,601
~998 GB
22 calibrated cameras
HRDexDB (real, xArm6 + dexterous hands)… See the full description on the dataset page: https://huggingface.co/datasets/rooty2020/text_audio_pretraining_for_GAN.ewe-bible-audio-text-tts
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Twi 16-Word Speech Segments
48775 speech-text pairs split from long recordings.
Processing pipeline
Source audio from ghananlpcommunity/ewe-tts-bible-full-audio-text
Full-file CTC forced alignment (MMS-300M) for… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ewe-bible-audio-text-tts.dagbani-bible-audio-text-tts
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Twi 16-Word Speech Segments
53410 speech-text pairs split from long recordings.
Processing pipeline
Source audio from ghananlpcommunity/dagbani-tts-bible-full-audio-text
Full-file CTC forced alignment (MMS-300M) for… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/dagbani-bible-audio-text-tts.IEMO_Audio_Text_MergedMSPI_Audio_Text_Mergedquran-audio-text
QuranLab — Verse-Aligned Quran Text + Recitation References
This dataset joins QuranLab's canonical Hafs Arabic text to its
per-ayah recitation references. Every row is one exact
(recitation_id, verse_key) pair: the Uthmani transcript, a search-friendly
Simple-Clean transcript, and the corresponding audio_url.
QuranLab is a volunteer effort. Our aim is to present these works carefully and at high quality, and to help them travel faithfully — in the spirit in which they were… See the full description on the dataset page: https://huggingface.co/datasets/quranlab/quran-audio-text.risalei-nur-text-audio
Risale-i Nur Text–Audio
Kaynak · Source: RNK Neşriyat — yazılı izinle · used with written permission.
Her satırda gerçek insan okuması ile o sesin kanonik metni birlikte bulunur.
Sesler dış bağlantı değildir: WAV baytları Parquet dosyalarının içindedir.
Kaynak sitesi veya başka bir ses sunucusu gerekmez.
Each row pairs a human reading with its canonical transcript. Audio is stored
as WAV bytes inside the Parquet files; no source website or external audio
server is required.… See the full description on the dataset page: https://huggingface.co/datasets/risaleinur/risalei-nur-text-audio.iemocap_audio_text
Dataset Card for "iemocap_audio_text"
More Information needed
korean-audio-text-economyConvert YouTube playlists to speech-to-text datasets
JA_audio_JA_text_180k_samples-Noise and silence have been removed from the begining and end of each sample.
-Unnecessary and inaccurate punctuation have been removed.
-Text has been normalized.
Normalization is based on the neologd's rules: https://github.com/neologd/mecab-ipadic-neologd/wiki/Regexp.ja.
ewe-tts-bible-full-audio-text
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Ewe Tts Bible Full Audio Text
iemocap_audio_text_splitted
Dataset Card for "iemocap_audio_text_splitted"
More Information needed
speech_text-tts_audiowaxal-audio-text-retrieval
WAXAL speech–text retrieval (MTEB)
Multilingual speech↔text retrieval over 16 Sub-Saharan African languages, derived from
WAXAL (Google and partners).
Most of these languages have no presence in mteb's existing multilingual audio tasks,
which skew European and South/East Asian. Prepared as WaxalA2TRetrieval and
WaxalT2ARetrieval.
Contents
One config per language, each with id, audio, text, speaker_id, gender,
language. 1,722 utterances total.
code
language… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/waxal-audio-text-retrieval.nlp-sentiment-audio-text-curated24
NLP Sentiment Audio Text Data Notes
Dataset summary
This data card accompanies a lightweight NLP Sentiment loader for Audio Text metadata. It is meant for pipeline inspection, source adaptation, and reproducible split preparation.
Included material
load_data.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md —… See the full description on the dataset page: https://huggingface.co/datasets/annashevchuk/nlp-sentiment-audio-text-curated24.quran-audio-text-dataset
Quran MD 🕌
Overview
This collection provides comprehensive audio recordings of the complete Quran (Holy Book of Islam) with multiple recitations and granular annotations. The dataset is organized into two separate, specialized datasets optimized for different use cases.
Related Dataset Links
#
Dataset
Focus
Samples
Hugging Face Link
1
🎵 Ayah Dataset
Verse-level recitations
187,080
Buraaq/quran-md-ayahs
2
📝 Word Dataset
Individual word… See the full description on the dataset page: https://huggingface.co/datasets/Buraaq/quran-audio-text-dataset.MELD_Text_AudioAudio-text
Hinglish Audio Dataset
This dataset contains 30 audio-text pairs.
Structure
file_name: Audio file path
text: Hinglish transcript
duration: Duration in seconds
Audio was generated using Sarvam AI's Bulbul v2 model.
ewe-bible-audio-text-tts
Twi 16-Word Speech Segments
48775 speech-text pairs split from long recordings.
Processing pipeline
Source audio from ghananlpcommunity/ewe-tts-bible-full-audio-text
Full-file CTC forced alignment (MMS-300M) for word-level timestamps
Words grouped into 16-word segments
Leading/trailing silence trimmed with VAD (-40 dBFS threshold)
Filtered: min 1.0s, max 15.0s
Original sample rate preserved (24kHz)
Usage
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/fiifinketia/ewe-bible-audio-text-tts.robotics-audio-text-benchmark
Robotics Audio Text Data Notes
Dataset summary
This repository contains a preparation pipeline and a small metadata sample for Robotics work with Audio Text inputs. It does not claim to be a complete benchmark release; the loader documents how source data is normalized and validated.
Included material
dataset.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small… See the full description on the dataset page: https://huggingface.co/datasets/christopherwright/robotics-audio-text-benchmark.korean-audio-text-developfood-audio-text-curated-2023
Food Audio Text Data Notes
Dataset summary
This data card accompanies a lightweight Food loader for Audio Text metadata. It is meant for pipeline inspection, source adaptation, and reproducible split preparation.
Included material
prepare.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md — data card and usage… See the full description on the dataset page: https://huggingface.co/datasets/isabellaruiz/food-audio-text-curated-2023.architecture-audio-text29
Architecture Audio Text Data Notes
Dataset summary
This data card accompanies a lightweight Architecture loader for Audio Text metadata. It is meant for pipeline inspection, source adaptation, and reproducible split preparation.
Included material
build_dataset.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md —… See the full description on the dataset page: https://huggingface.co/datasets/Nicholassmith1220/architecture-audio-text29.triplets_audio_image_text_v1robotics-audio-text3
Robotics Audio Text Data Notes
Dataset summary
A documented Robotics data-preparation workflow for Audio Text records. The bundled rows demonstrate the schema and validation path rather than pretending to be a full training corpus.
Included material
prepare.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md —… See the full description on the dataset page: https://huggingface.co/datasets/YangZhoudaj/robotics-audio-text3.travel-audio-text-v2
Travel Audio Text Data Notes
Dataset summary
A documented Travel data-preparation workflow for Audio Text records. The bundled rows demonstrate the schema and validation path rather than pretending to be a full training corpus.
Included material
dataset.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md — data… See the full description on the dataset page: https://huggingface.co/datasets/lfrodriguesza/travel-audio-text-v2.audio-text-dataset
Music Audio Text Data Notes
Dataset summary
This repository contains a preparation pipeline and a small metadata sample for Music work with Audio Text inputs. It does not claim to be a complete benchmark release; the loader documents how source data is normalized and validated.
Included material
loader.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable… See the full description on the dataset page: https://huggingface.co/datasets/Williamsmadison/audio-text-dataset.
