nano-codec
nemo-nano-codec-22khz-1.89kbps-21.5fpsnemo-nano-codec-22khz-0.6kbps-12.5fpsnemo-nano-codec-22khz-1.78kbps-12.5fpsnemo-nano-codec-22khz-0.6kbps-12.5fps-MLXnemo-nano-codec-22khz-0.6kbps-12.5fps-ONNXnemo-nano-codec-22khz-1.78kbps-12.5fps-ONNXnemo-nano-codec-22khz-1.89kbps-21.5fpsnemo-nano-codec-22khz-1.89kbps-21.5fps-ONNX
f-actor-behavior-sd-nanocodec
F-Actor Nano-Codec Dataset
This repository contains the data accompanying the paper
F-Actor: Controllable Conversational Behaviour in Full-Duplex Models.
The data consists of the Behavior-SD dataset, encoded using nvidia/nemo-nano-codec-22khz-0.6kbps-12.5fps, and augmented with a different narrative.
About our work:
Spoken conversational systems require more than accurate speech generation to have human-like conversations: to feel natural and engaging, they must produce… See the full description on the dataset page: https://huggingface.co/datasets/maikezu/f-actor-behavior-sd-nanocodec.kanitts2-fr-nanocodecemolia_filtered_nano_codec_21_dataset
Emolia · Filtered · NanoCodec (FSQ) Tokens
A cleaned, pre-tokenized version of laion/Emolia
prepared for text-to-speech (TTS) training.
The pipeline is two steps:
Quality filtering with the open-source
audio_filter tool — this
removes the dirtiest recordings (noise, clipping, band-limiting, robotic artifacts,
overlapping speakers), which matters a lot for TTS quality.
Discrete audio tokenization with NVIDIA
nvidia/nemo-nano-codec-22khz-1.89kbps-21.5fps
(an FSQ neural audio… See the full description on the dataset page: https://huggingface.co/datasets/nineninesix/emolia_filtered_nano_codec_21_dataset.elise-en-nano-codec-dataset
Elise EN Nano-Codec Dataset
This dataset is built upon the Elise dataset and re-encoded using NVIDIA’s NeMo Audio Codec into nano audio tokens.
It is designed for fine-tuning multimodal LLMs and speech systems (TTS/ASR) that rely on codec-based audio token representations.
Dataset Structure
text: transcription of the utterance.
speaker: speaker identifier (string).
nano_layer_1 … nano_layer_4: tokenized audio representations from the NVIDIA NeMo Nano Codec… See the full description on the dataset page: https://huggingface.co/datasets/nineninesix/elise-en-nano-codec-dataset.indicvoice-hi-nanocodec-tokenspuck-gemini-flash-en-nano-codec-dataset
