datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
f-actor-behavior-sd-nanocodec
F-Actor Nano-Codec Dataset
This repository contains the data accompanying the paper
F-Actor: Controllable Conversational Behaviour in Full-Duplex Models.
The data consists of the Behavior-SD dataset, encoded using nvidia/nemo-nano-codec-22khz-0.6kbps-12.5fps, and augmented with a different narrative.
About our work:
Spoken conversational systems require more than accurate speech generation to have human-like conversations: to feel natural and engaging, they must produce… See the full description on the dataset page: https://huggingface.co/datasets/maikezu/f-actor-behavior-sd-nanocodec.kanitts2-fr-nanocodecemolia_filtered_nano_codec_21_dataset
Emolia · Filtered · NanoCodec (FSQ) Tokens
A cleaned, pre-tokenized version of laion/Emolia
prepared for text-to-speech (TTS) training.
The pipeline is two steps:
Quality filtering with the open-source
audio_filter tool — this
removes the dirtiest recordings (noise, clipping, band-limiting, robotic artifacts,
overlapping speakers), which matters a lot for TTS quality.
Discrete audio tokenization with NVIDIA
nvidia/nemo-nano-codec-22khz-1.89kbps-21.5fps
(an FSQ neural audio… See the full description on the dataset page: https://huggingface.co/datasets/nineninesix/emolia_filtered_nano_codec_21_dataset.elise-en-nano-codec-dataset
Elise EN Nano-Codec Dataset
This dataset is built upon the Elise dataset and re-encoded using NVIDIA’s NeMo Audio Codec into nano audio tokens.
It is designed for fine-tuning multimodal LLMs and speech systems (TTS/ASR) that rely on codec-based audio token representations.
Dataset Structure
text: transcription of the utterance.
speaker: speaker identifier (string).
nano_layer_1 … nano_layer_4: tokenized audio representations from the NVIDIA NeMo Nano Codec… See the full description on the dataset page: https://huggingface.co/datasets/nineninesix/elise-en-nano-codec-dataset.indicvoice-hi-nanocodec-tokenspuck-gemini-flash-en-nano-codec-datasetindicvoices-nanocodec-tokensThis dataset contains 200K samples of text and corresponding audio tokenized using nvidia/nemo-nano-codec-22khz-1.89kbps-21.5fps
First 100K samples are in Hindi, sampled from the hindi split of indicvoices dataset, and next 100K samples are in english with Indian accent, sampled from skbose/indian-english-nptel-v0 dataset.
The dataset can be useful for training TTS models.
nvidia-nemo-nano-codec-biblerasa-malayalam-nano-codecexp-nano-codecur_nano_codec
Urdu Nano Codec TTS Dataset
This dataset contains Urdu sentences along with their Nano Codec tokenized representations
(nano_layer_1 to nano_layer_4) generated using NVIDIA NeMo Nano Codec model (22kHz, 0.6 kbps, 12.5 fps).
It can be used for:
TTS model training
Audio reconstruction from tokens
Low-resource Urdu speech research
Dataset Structure
text: Original Urdu sentences
nano_layer_1 … nano_layer_4: Tokenized representations (quantized codebooks)
encoded_len:… See the full description on the dataset page: https://huggingface.co/datasets/mahwizzzz/ur_nano_codec.jinsaryko-tifa-en-nano-codec-dataset
Tifa EN Nano-Codec Dataset
This dataset is built upon the Tifa dataset and re-encoded using NVIDIA’s NeMo Audio Codec into nano audio tokens.
It is designed for fine-tuning multimodal LLMs and speech systems (TTS/ASR) that rely on codec-based audio token representations.
Dataset Structure
text: transcription of the utterance.
speaker: speaker identifier (string).
nano_layer_1 … nano_layer_4: tokenized audio representations from the NVIDIA NeMo Nano Codec… See the full description on the dataset page: https://huggingface.co/datasets/nineninesix/jinsaryko-tifa-en-nano-codec-dataset.optimus-prime-en-nano-codec-datasetexpresso-conversational-en-nano-codec-dataset
Expresso Conversational EN Nano-Codec Dataset
This dataset is built upon the Expresso conversational dataset and re-encoded using NVIDIA’s NeMo Audio Codec into nano audio tokens.
It is designed for fine-tuning multimodal LLMs and speech systems (TTS/ASR) that rely on codec-based audio token representations.
Dataset Structure
text: transcription of the utterance.
speaker: speaker identifier (string).
nano_layer_1 … nano_layer_4: tokenized audio representations… See the full description on the dataset page: https://huggingface.co/datasets/nineninesix/expresso-conversational-en-nano-codec-dataset.IndexTTS-nano-codecimasc_slr_Malayalam-nano-codeckore-gemini-flash-en-nano-codec-datasetmagic-data-nano-codec-datasetNano_Codec_Nisan_KumruIndicVoices-r-ML-nano-codecSPRINGLab-IndicTTS_Malayalam-nano-codecpuck-gemini-flash-en-nano-codec-datasetkuroyukihime-ja-nano-codec-dataseturdu-tts-nano-codecthien-tts-nanocodec
