datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
farsi-asr-iran-international-raw
Iran International raw Farsi audio archive
This repository preserves 11,993 individually addressable FLAC source files for incremental ASR relabeling and reproducible restoration.
Repository file layout
The first 9,990 FLAC files are stored at the repository root. The remaining 2,003 FLAC files are stored individually under overflow/ to respect Hugging Face's 10,000-entry-per-directory limit.
REMOTE_PATHS.jsonl records every source filename, remote path, byte size… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/farsi-asr-iran-international-raw.and-peacetajik-asr-corpus-v0
Tajik ASR Corpus v0
Deduplicated Tajik automatic speech recognition corpus assembled from FLEURS-derived
speech data, Mozilla Common Voice 25 Tajik, and Muhtasham Tajik ASR augmented data.
Format
Each split has a data.tsv and an audio/ directory.
TSV columns:
id
audio_filename
raw_transcription
transcription
characters
audio_bytes
source
source_id
duplicate_count
tajik_asr_combined.sqlite mirrors the TSV rows and includes normalized_text,
source_split, and… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/tajik-asr-corpus-v0.SINE
SINE Dataset
Overview
The Speech INfilling Edit (SINE) dataset is a comprehensive collection for speech deepfake detection and audio authenticity verification. This dataset contains ~87GB of audio data distributed across 32 splits, featuring both authentic and synthetically manipulated speech samples.
Dataset Statistics
Total Size: ~87GB
Number of Splits: 32 (split-0.tar.gz to split-31.tar.gz)
Audio Format: WAV files
Source: Speech edited from LibriLight… See the full description on the dataset page: https://huggingface.co/datasets/PeacefulData/SINE.peaky-blinders-learning-purpose-only
TTS Dataset - Peaky Blinders
This dataset contains audio segments with transcriptions from Peaky Blinders for Text-to-Speech training.
Dataset Generation
This dataset was generated using the TTS-Dataset-Maker pipeline, which provides:
Silero VAD-based silence removal - Removes long silences while preserving natural speech gaps
DeepFilterNet denoising - CPU-optimized audio denoising with gentle attenuation (15dB)
AssemblyAI transcription - High-quality speech-to-text with… See the full description on the dataset page: https://huggingface.co/datasets/ahk-d/peaky-blinders-learning-purpose-only.pe-av-assetsPeacockneyshekar-v3-asr-aligned
Neyshekar v3 ASR-Aligned
This is a repaired subset of Neyshekar v3 for Persian ASR work. The public v3
archive contains real audio and real transcripts, but the downloaded
dataset.json filename-to-text mapping does not align for the checked samples.
This export keeps only audio clips whose transcript could be recovered by
matching multiple ASR hypotheses back to the original Neyshekar transcript pool.
It is useful as a curated ASR training/evaluation candidate set, with the… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/neyshekar-v3-asr-aligned.Pepper-Ann-PearsonDCASE_2025_AudioQA_backupDCASE 2025 Audio Question and Answering
PearlSINE_v2
SINE v2 Dataset
SINE v2 is a large-scale audio dataset containing over 350K audio samples organized by data processing types. The dataset includes four different configurations representing different audio processing techniques.
Dataset Summary
This dataset contains audio samples with timing annotations and processing labels, organized into four main configurations:
edit: Audio samples with editing processing (~87K samples)
real: Real/original audio samples (~87K… See the full description on the dataset page: https://huggingface.co/datasets/PeacefulData/SINE_v2.and-peace1PearlJamBlackVozPearlJamBlackVoz2
