datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DDD-Cambodia-khmer-speech-dataset-parquet-000-159-en-translateDisclaimer: The original dataset can be found here.
It is published by Digital Divide Data Cambodia (DDD-Cambodia).
License:
Khmer ASR Cultural Dataset's license is Creative Commons Attribution Share Alike 4.0 International (CC-BY-SA-4.0).
Please attribute Digital Divide Data if you use this dataset in any way.
Objective of this dataset
Add English translation: a new column en_translate is added to the original dataset (only from parquet 000 to 159 of the original… See the full description on the dataset page: https://huggingface.co/datasets/KrorngAI/DDD-Cambodia-khmer-speech-dataset-parquet-000-159-en-translate.simchoir-parquet
FastMSS synthetic multi-speaker meetings - parquet edition
Streaming-friendly parquet shards of the FastMSS synthetic multi-speaker conversational corpus. Each row is one mixture with the audio bytes embedded inline (16 kHz mono WAV) plus per-segment diarization timestamps, per-word transcript and the full lhotse cut as a JSON blob. See fastmss/hf_dataset.py for the schema docstring.
Subsets and splits
debug — splits: train — 1 mixtures, 1.6 min total, 6 unique speakers… See the full description on the dataset page: https://huggingface.co/datasets/arda-argmax/simchoir-parquet.common_voice_22_0_toki_pona_parquet
Common Voice 22.0 - Toki Pona Subset!
My own Parquet conversion of Toki Pona's subset of Fsicoli's reupload of Common Voice 22 so we don't have to downgrade to Datasets 3.6 anymore!
Why?
Because the original dataset required Hugging Face Datasets 3.6 or older because it has Python code and it's in TAR shards.
This is in Parquet and works with any recent version of Hugging Face Datasets!
Details
Dataset Structure
DatasetDict({… See the full description on the dataset page: https://huggingface.co/datasets/MihaiPopa-1/common_voice_22_0_toki_pona_parquet.one_voice_FACEBOOK_PARQUET
Artificial Omnivoice Hungarian Speaker Dataset
Ez egy teljesen szintetikus magyar nyelvű beszédadatbázis, amely kiváló minőségű szövegfelolvasó (TTS) és beszédfelismerő (ASR) modellek tanításához és finomhangolásához készült.
Adatforrás és Referencia Hang
A dataset alapjául szolgáló referencia hang (speaker identity) egy 20 másodperces részlet az alábbi YouTube videóból:
Forrás: Hogyan legyél tökéletes magyar várvédő tutorial
Licenc: A videó CC (Creative Commons)… See the full description on the dataset page: https://huggingface.co/datasets/fablevi/one_voice_FACEBOOK_PARQUET.prueba_parquetThis is an example of a repository with parquet files only.
medical-tts-parquet-2-16khz
IntelMedica Medical TTS Dataset v2 (16kHz)
Description
Synthetic medical speech dataset for fine-tuning Whisper-based ASR models on clinical and nursing terminology. Contains 101,475 audio-text pairs totaling 184.1 hours of speech at 16 kHz mono, generated using Kokoro-82M TTS with 19 voices across three English accent groups.
This is v2 -- a companion to the v1 dataset (125,500 samples, ~257 hours). v2 focuses on terms from additional data sources (RxNorm API, FDA… See the full description on the dataset page: https://huggingface.co/datasets/intelmedica/medical-tts-parquet-2-16khz.bengali-tts-folderized-parquet-stage1
Bengali TTS Folderized Parquet Stage 1
This is the intermediate organized parquet layer before the final fully row-wise Bengali TTS dataset.
Generated metadata refresh: 2026-06-18T21:44:49Z
Layout
<speaker>/part-00000.parquet
<speaker>/part-00001.parquet
<speaker>/metadata/stats.json
Columns
speaker
video_id
chunk_file
audio_file
duration
transcription
uuid
audio as Hugging Face Audio feature backed by parquet struct<bytes,path>… See the full description on the dataset page: https://huggingface.co/datasets/smam/bengali-tts-folderized-parquet-stage1.medical-tts-parquet-1
IntelMedica Medical TTS Dataset v1 (24kHz) -- DEPRECATED
This dataset is deprecated. Please use intelmedica/medical-tts-parquet-1-16khz instead, which contains 125,500 samples (vs 10,000 here) at 16kHz sample rate optimized for ASR training.
Description
Synthetic medical speech dataset for training medical ASR models. This is the original 10K-sample version at 24kHz. It has been superseded by the 16kHz version with 12.5x more data.
Dataset Details
Samples:… See the full description on the dataset page: https://huggingface.co/datasets/intelmedica/medical-tts-parquet-1.bengali-tts-folderized-parquet-stage2
Bengali TTS Folderized Parquet Stage 2
Final filtered (keep=True) Bengali TTS dataset, with merged/combined chunks.
Layout
<speaker>.parquet (single shard, audio <= ~1GB)
<speaker>_001.parquet, _002.parquet, ... (multiple shards, split by audio byte size)
Audio sourcing convention
COMBINED == False -> sourced from extracted_audio/<speaker>/<video_id>/<chunk_file>
COMBINED == True -> sourced from… See the full description on the dataset page: https://huggingface.co/datasets/dipit099/bengali-tts-folderized-parquet-stage2.
