datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ne-asr-dataset-nag-aug
NE ASR Augmented Dataset -- Nagamese (nag)
Augmented automatic speech recognition dataset for Nagamese (nag),
a Assamese-based creole language spoken in Nagaland, India.
Source
Augmented from sulabhkatiyar/ne-asr-nag
(original transcribed speech data from the ARTPARK-IISc Vaani project).
Language Information
Property
Value
Language
Nagamese
ISO 639-3
nag
Family
Assamese-based creole
Region
Nagaland, India
Tonal
No
Tier
D (23.76h… See the full description on the dataset page: https://huggingface.co/datasets/sulabhkatiyar/ne-asr-dataset-nag-aug.ne-asr-dataset-nag
Nagamese (nag) — ASR dataset
A small Nagamese (nag) speech-to-text dataset for automatic speech recognition
(ASR) of a low-resource North-East India language. Each example pairs a short audio
clip with its Romanized (Latin-script) transcript.
Source
Derived from the ARTPARK-IISc Vaani project (https://vaani.iisc.ac.in/)
Splits
Split
Samples
train
12,862
validation
1,532
test
1,717
Data fields
Each example has:… See the full description on the dataset page: https://huggingface.co/datasets/sulabhkatiyar/ne-asr-dataset-nag.nagatoro-sound
# Nagatoro Hayase Voice Dataset (140 Clean Clips)
This dataset contains 140 high-quality, pre-processed clean voice clips of the anime character Nagatoro Hayase (Ijiranaide, Nagatoro-san / Don't Toy with Me, Miss Nagatoro).
It is specifically curated and optimized for AI voice training pipelines, voice conversion models, and audio deep learning experiments.
Dataset Details
Character: Nagatoro Hayase (長瀞 早瀬)
Language: Japanese (JA)
File Format: .wav (High Quality)
Total… See the full description on the dataset page: https://huggingface.co/datasets/ezfiez/nagatoro-sound.bleep-spans
Bleep spans — synthetic sensitive-speech regions with frame-accurate labels
Where sensitive information is spoken, and what kind it is — never what was
said.
Every recording is synthetic. No real telephone call, clinical recording, or any
other real speech was used, recorded, or derived from at any stage.
🤗 Model: NagaYu/bleep-0.09b
🎛️ Demo: NagaYu/bleep
What a row contains
utt_id, voice_key, condition, duration, subsets, and three parallel
arrays —… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/bleep-spans.ne-tts-nag
NE-TTS Nagamese (nag)
Cleaned TTS dataset for Nagamese (nag), a North East Indian language. Derived from the Vaani dataset with SNR filtering, LUFS normalization, and text cleaning.
Stats
Metric
Value
Total clips
14,796
Total hours
21.9h
High SNR (>=20dB)
9,688 clips
Medium SNR (15-20dB)
5,108 clips
Sample rate
16kHz
Audio format
WAV, 16-bit PCM
Schema
Column
Type
Description
audio
Audio
16kHz WAV audio
text… See the full description on the dataset page: https://huggingface.co/datasets/sulabhkatiyar/ne-tts-nag.vaani-rajasthan_nagaur-cleanedne-tts-f5-nag
NE-TTS F5 Nagamese (nag)
F5-TTS training dataset for Nagamese (nag). Contains 9,688 clips at 24kHz (SNR >= 20dB) from the cleaned NE-TTS dataset, formatted for F5-TTS training.
Stats
Metric
Value
Clips
9,688
Hours
14.5h
Sample rate
24kHz
SNR filter
>= 20dB
Source
ne-tts-nag
Schema
Column
Type
Description
audio
Audio
24kHz WAV audio
text
string
Cleaned transcript
Note
Audio is upsampled from… See the full description on the dataset page: https://huggingface.co/datasets/sulabhkatiyar/ne-tts-f5-nag.Vaani-nagamese-majority-lg-English-no-transcript0vaani-maharashtra_nagpur-cleanedne-tts-mms-nag
NE-TTS MMS-VITS Nagamese (nag)
MMS-VITS fine-tuning subset for Nagamese (nag). Contains 150 high-quality clips selected from the cleaned NE-TTS dataset (SNR >= 20dB), formatted for MMS-VITS fine-tuning.
Stats
Metric
Value
Clips
150
Sample rate
22050Hz
SNR filter
SNR >= 20dB
Source
ne-tts-nag
Schema
Column
Type
Description
audio
Audio
22050Hz WAV audio
text
string
Cleaned transcript
Usage
Use… See the full description on the dataset page: https://huggingface.co/datasets/sulabhkatiyar/ne-tts-mms-nag.Vaani-nagamese-majority-lg-English-with-transcriptwhisperkit_testsAll files are from: earnings22
Rencoded to 24kbps MP3 using:
ffmpeg -i 4446796.wav -vn -map_metadata -1 -ac 1 -c:a libmp3lame -b:a 24k -application voip -y 4446796.mp3
XTTS_testNagisinNagatoro-voiceVaani_Nagaur_tran_hin_audioVaani_Nagpur_tran_hin_audiotrain
