nanocodec
f-actor-behavior-sd-nanocodec
F-Actor Nano-Codec Dataset
This repository contains the data accompanying the paper
F-Actor: Controllable Conversational Behaviour in Full-Duplex Models.
The data consists of the Behavior-SD dataset, encoded using nvidia/nemo-nano-codec-22khz-0.6kbps-12.5fps, and augmented with a different narrative.
About our work:
Spoken conversational systems require more than accurate speech generation to have human-like conversations: to feel natural and engaging, they must produce… See the full description on the dataset page: https://huggingface.co/datasets/maikezu/f-actor-behavior-sd-nanocodec.kanitts2-fr-nanocodecindicvoice-hi-nanocodec-tokensindicvoices-nanocodec-tokensThis dataset contains 200K samples of text and corresponding audio tokenized using nvidia/nemo-nano-codec-22khz-1.89kbps-21.5fps
First 100K samples are in Hindi, sampled from the hindi split of indicvoices dataset, and next 100K samples are in english with Indian accent, sampled from skbose/indian-english-nptel-v0 dataset.
The dataset can be useful for training TTS models.
Nano_Codec_Nisan_Kumruthien-tts-nanocodec
