datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
persian-asr-audio-text-2.69M-chizzled
🗂️ persian-asr-audio-text-2.69M-chizzled
English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission
🌟 At a glance | معرفی سریع
English
فارسی
🎯 Purpose
Phase A-scale audio/text dataset.
پیکرهٔ بزرگ جفتهای صوت و متنِ پالایششده برای آموزش در مقیاس فاز A.
🧩 Role
Persian text and linguistic asset
مصنوع متنی و زبانی فارسی
📦 Snapshot
417 files; approximately 236.86 GB
417 فایل؛ حدود 236.86 GB
🧱 Packaging
414 Parquet files and 0… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-asr-audio-text-2.69M-chizzled.ewe-bible-audio-text-tts
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Twi 16-Word Speech Segments
48775 speech-text pairs split from long recordings.
Processing pipeline
Source audio from ghananlpcommunity/ewe-tts-bible-full-audio-text
Full-file CTC forced alignment (MMS-300M) for… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ewe-bible-audio-text-tts.dagbani-bible-audio-text-tts
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Twi 16-Word Speech Segments
53410 speech-text pairs split from long recordings.
Processing pipeline
Source audio from ghananlpcommunity/dagbani-tts-bible-full-audio-text
Full-file CTC forced alignment (MMS-300M) for… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/dagbani-bible-audio-text-tts.IEMO_Audio_Text_MergedMSPI_Audio_Text_Mergedquran-audio-text
QuranLab — Verse-Aligned Quran Text + Recitation References
This dataset joins QuranLab's canonical Hafs Arabic text to its
per-ayah recitation references. Every row is one exact
(recitation_id, verse_key) pair: the Uthmani transcript, a search-friendly
Simple-Clean transcript, and the corresponding audio_url.
QuranLab is a volunteer effort. Our aim is to present these works carefully and at high quality, and to help them travel faithfully — in the spirit in which they were… See the full description on the dataset page: https://huggingface.co/datasets/quranlab/quran-audio-text.risalei-nur-text-audio
Risale-i Nur Text–Audio
Kaynak · Source: RNK Neşriyat — yazılı izinle · used with written permission.
Her satırda gerçek insan okuması ile o sesin kanonik metni birlikte bulunur.
Sesler dış bağlantı değildir: WAV baytları Parquet dosyalarının içindedir.
Kaynak sitesi veya başka bir ses sunucusu gerekmez.
Each row pairs a human reading with its canonical transcript. Audio is stored
as WAV bytes inside the Parquet files; no source website or external audio
server is required.… See the full description on the dataset page: https://huggingface.co/datasets/risaleinur/risalei-nur-text-audio.iemocap_audio_text
Dataset Card for "iemocap_audio_text"
More Information needed
korean-audio-text-economyConvert YouTube playlists to speech-to-text datasets
JA_audio_JA_text_180k_samples-Noise and silence have been removed from the begining and end of each sample.
-Unnecessary and inaccurate punctuation have been removed.
-Text has been normalized.
Normalization is based on the neologd's rules: https://github.com/neologd/mecab-ipadic-neologd/wiki/Regexp.ja.
ewe-tts-bible-full-audio-text
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Ewe Tts Bible Full Audio Text
iemocap_audio_text_splitted
Dataset Card for "iemocap_audio_text_splitted"
More Information needed
speech_text-tts_audiowaxal-audio-text-retrieval
WAXAL speech–text retrieval (MTEB)
Multilingual speech↔text retrieval over 16 Sub-Saharan African languages, derived from
WAXAL (Google and partners).
Most of these languages have no presence in mteb's existing multilingual audio tasks,
which skew European and South/East Asian. Prepared as WaxalA2TRetrieval and
WaxalT2ARetrieval.
Contents
One config per language, each with id, audio, text, speaker_id, gender,
language. 1,722 utterances total.
code
language… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/waxal-audio-text-retrieval.Audio-text
Hinglish Audio Dataset
This dataset contains 30 audio-text pairs.
Structure
file_name: Audio file path
text: Hinglish transcript
duration: Duration in seconds
Audio was generated using Sarvam AI's Bulbul v2 model.
ewe-bible-audio-text-tts
Twi 16-Word Speech Segments
48775 speech-text pairs split from long recordings.
Processing pipeline
Source audio from ghananlpcommunity/ewe-tts-bible-full-audio-text
Full-file CTC forced alignment (MMS-300M) for word-level timestamps
Words grouped into 16-word segments
Leading/trailing silence trimmed with VAD (-40 dBFS threshold)
Filtered: min 1.0s, max 15.0s
Original sample rate preserved (24kHz)
Usage
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/fiifinketia/ewe-bible-audio-text-tts.korean-audio-text-developtriplets_audio_image_text_v1text-vision-audio-2k-testA 2k sample dataset for testing multimodal (text+vision+audio) format. This is compatible with HF's processor apply_chat_template.
Load in Axolotl via:
datasets:
- path: Nanobit/text-vision-audio-2k-test
type: chat_template
Make sure to download the image and audio via:
wget https://huggingface.co/datasets/Nanobit/text-vision-audio-2k-test/resolve/main/African_elephant.jpg
wget https://huggingface.co/datasets/Nanobit/text-vision-audio-2k-test/resolve/main/En-us-African_elephant.oga… See the full description on the dataset page: https://huggingface.co/datasets/axolotl-ai-co/text-vision-audio-2k-test.text-audio-2k-testA 2k sample dataset for testing multimodal (text+audio) format. This is compatible with HF's processor apply_chat_template.
Load in Axolotl via:
datasets:
- path: Nanobit/text-audio-2k-test
type: chat_template
Make sure to download the audio via:
wget https://huggingface.co/datasets/Nanobit/text-vision-audio-2k-test/resolve/main/En-us-African_elephant.oga
Audio source: https://upload.wikimedia.org/wikipedia/commons/a/ad/En-us-African_elephant.oga
Each sample has the following format… See the full description on the dataset page: https://huggingface.co/datasets/axolotl-ai-co/text-audio-2k-test.Tamil_MSA_Audio_Text_Chunkedtext-2-audio-human-preference-benchmark
Text to Audio Human Benchmark
In this dataset, ~32k human responses collected in less than 1h using the Rapidata Python API, accessible to anyone and ideal for large scale evaluation.
The annotators were asked Which voice is more friendly? and Which voice sounds more natural? respectively.
Check out the Benchmark!
Tamil_MSA_Audio_Text
Dataset Details
Title: Dravidianmultimodality: A dataset for multi-modal sentiment analysis in Tamil and Malayalam
Authors: Bharathi Raja Chakravarthi et al.
Link to Paper: arXiv:2106.04853
Published: 2021
Source: arXiv preprint
orcasound_audio_as_textTrio-Image-Audio-Text
Trio
A unified multimodal dataset combining image, audio, and text from diverse public sources.
Usage
This dataset uses Configurations (Subsets) to manage its diverse data sources. You can load specific parts or the entire "filtered" dataset without downloading the NSFW portions.
pip install datasets
1. Load the "filtered" Subset
This configuration loads all 29 safe subsets, excluding the NSFW content.
from datasets import load_dataset
#… See the full description on the dataset page: https://huggingface.co/datasets/VINAY-UMRETHE/Trio-Image-Audio-Text.ng-accent-audio-with-textdagbani-bible-audio-text-tts
Twi 16-Word Speech Segments
53410 speech-text pairs split from long recordings.
Processing pipeline
Source audio from ghananlpcommunity/dagbani-tts-bible-full-audio-text
Full-file CTC forced alignment (MMS-300M) for word-level timestamps
Words grouped into 16-word segments
Leading/trailing silence trimmed with VAD (-40 dBFS threshold)
Filtered: min 1.0s, max 15.0s
Original sample rate preserved (24kHz)
Usage
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/fiifinketia/dagbani-bible-audio-text-tts.earlymeida-id-audio-text-datasetaudio-text-pair_train-test-dataset-hindi
dataset contains pairs of audio-text where text are hindi transcriptions and audio are corresponding audio file
structure:
`
DatasetDict({
train: Dataset({
features: ['audio', 'text'],
num_rows: 200
})
test: Dataset({
features: ['audio', 'text'],
num_rows: 60
})
sample row
{'audio': {'path': '/content/drive/MyDrive/sarvam.ai/1.mp3', 'array': array([ 5.85489014e-13, -6.08550428e-13, 5.14475181e-13, ..., -3.15216061e-13, -1.82061695e-13, 0.00000000e+00])… See the full description on the dataset page: https://huggingface.co/datasets/PYD4320/audio-text-pair_train-test-dataset-hindi.JA_audio_JA_text_180k
