datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
or_in_datasetArabic_Audio_Deepfake
ArAD Dataset (Arabic Audio DeepFake Dataset)
Dataset SummaryThis dataset contains Arabic deepfake audio samples, focusing mainly on Levantine dialect with some examples in Standard Arabic. It was created using the RVC v2 framework, fine-tuned on a custom dataset of multi-dialect Arabic speech. The goal is to simulate real-world deepfake audio attacks by generating synthetic speech from actual recordings and voice messages.
One of the first datasets to include real-world deepfake… See the full description on the dataset page: https://huggingface.co/datasets/DeepFake-Audio-Rangers/Arabic_Audio_Deepfake.tamil-english-podcast-diarization
Tamil-English Code-Mixed Podcast Diarization Dataset
Dataset Summary
This dataset contains long-form Tamil-English code-mixed podcast recordings
annotated for speaker diarization research. The recordings consist of natural
conversational speech with multiple speakers and realistic acoustic conditions,
making the dataset suitable for evaluating diarization pipelines in
real-world scenarios.
The dataset is intended to support research in:
Speaker diarization
Code-mixed… See the full description on the dataset page: https://huggingface.co/datasets/Rangasuthan/tamil-english-podcast-diarization.corpus5-proposed-no-overlap-random-1000
Corpus5 proposed-rule unflagged random sample
This manually gated dataset contains 1,000 uniformly randomly selected rows from
the fixed 2,715,793-row prepared Corpus5 snapshot on mac02.
A row is eligible only when both proposed checks are false:
the full chunk interval does not intersect positive-duration diarization turns
from two distinct speakers; and
no speaker-change/no-change disagreement is detected at aligned adjacent words
in transcript1 and transcript2 and propagated… See the full description on the dataset page: https://huggingface.co/datasets/Cybrpgs/corpus5-proposed-no-overlap-random-1000.stillalive-overlap-random-1000
Stillalive random overlap sample
This dataset contains 1,000 randomly selected audio chunks from the locally
validated Stillalive dataset for which has_any_overlap is true and
overlap_percentage is greater than zero.
The complete source schema and embedded 48 kHz mono, 320 kbps MP3 audio are
preserved. The sample contains 571 rows from finalized Part A and 429 rows from
the current Part B merge. Sampling used deterministic reservoir sampling with
seeds 20260920 (Part A) and… See the full description on the dataset page: https://huggingface.co/datasets/Cybrpgs/stillalive-overlap-random-1000.knihi-be-arlou_randevu_na_manieurach_all
AudioSet Pipeline Output
Мова / Language: Беларуская (Belarusian)
Аўдыё нарэзана з арыгінальнага запісу ў зыходнай частаце дыскрэтызацыі (native), мона, фрагменты да 30 секунд.
Частка калекцыі Belarusian Audiobooks (native).
Радкоў у датасеце
117
Працягласць
25 хв
Частата дыскрэтызацыі
44100 Hz
Каналы
мона
Даўжыня фрагмента
да 30 с
Структура
Кожны радок змяшчае:
audio — аўдыёфрагмент (native SR, мона, ≤30 с)
text — транскрыпцыя (Gemini +… See the full description on the dataset page: https://huggingface.co/datasets/fosters/knihi-be-arlou_randevu_na_manieurach_all.stillalive-proposed-no-overlap-random-1000
Random sample of chunks not flagged by the proposed overlap rules
This dataset contains 1,000 uniformly randomly selected chunks from the 1,408,153 chunks not flagged by either proposed overlap rule in a fixed 3,882,782-row completed Stillalive dataset snapshot, captured on 20 September 2026 at approximately 21:23 WAT.
Selection used a single global reservoir across both dataset parts with seed 202609201. There were no additional filters for duration, text, language, quality… See the full description on the dataset page: https://huggingface.co/datasets/Cybrpgs/stillalive-proposed-no-overlap-random-1000.stillalive-proposed-no-overlap-random-1000-v2
Second random sample of chunks not flagged by the proposed overlap rules
This dataset contains a second, disjoint set of 1,000 uniformly randomly selected chunks from a fixed 3,882,782-row completed Stillalive dataset snapshot captured on 20 September 2026.
The candidate population contains chunks for which both proposed flags are false:
no two distinct ordinary-diarization speakers have positive-duration intersections with the full chunk interval; and
no… See the full description on the dataset page: https://huggingface.co/datasets/Cybrpgs/stillalive-proposed-no-overlap-random-1000-v2.knihi-be-arlou_randevu_na_manieurach_output_original
AudioSet Pipeline Output — арыгінальнае аўдыё
Мова / Language: Беларуская (Belarusian)
Арыгінальнае аўдыё без апрацоўкі, захаванае ў зыходнай якасці.
Частка калекцыі Ministerskija —
корпус беларускіх аўдыёкніг.
Апрацаваная версія (сегменты ~15 с, выраўнаваная транскрыпцыя):
knihi-be-arlou_randevu_na_manieurach_output
Структура
Кожны радок змяшчае:
audio — арыгінальны аўдыёзапіс
text — транскрыпцыя
chunk_uid — унікальны ідэнтыфікатар
Ліцэнзія /… See the full description on the dataset page: https://huggingface.co/datasets/fosters/knihi-be-arlou_randevu_na_manieurach_output_original.
