datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fleurs-r-neucodec-all-languages
FLEURS-R NeuCodec All Languages
FLEURS-R metadata, source audio and precomputed NeuCodec speech tokens for 102
locales, plus a speaker label FLEURS itself does not ship.
Layout
data/{locale}-{split}.parquet — metadata, one row per utterance (this is what the
viewer shows).
audio/{locale}-{split}.zip — source FLEURS-R audio, 24kHz mono PCM16 WAV, members
named audio/{locale}/{split}/{id}.wav (the path column).
neucodec/{locale}-{split}-rank{N}.zip — NeuCodec… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/fleurs-r-neucodec-all-languages.emilia-yodas-english-neucodec
Dataset Card for NeuCodec Emilia-YODAS
Dataset Summary
The NeuCodec Emilia-YODAS dataset is an English-language dataset containing >30M audio samples (>78k hours), taken from the English-language subset of Emilia-YODAS and compressed with NeuCodec.
Usage
import torch
from datasets import load_dataset
from neucodec import NeuCodec
# load dataset and model
dataset = load_dataset("neuphonic/emilia-yodas-english-neucodec", split="train"… See the full description on the dataset page: https://huggingface.co/datasets/neuphonic/emilia-yodas-english-neucodec.yodas-granary-it-neucodec-10s-20s
Granary Italian VoxPopuli NeuCodec
NeuCodec-tokenized rows for NeuTTS fine-tuning. The source rows were streamed from espnet/yodas-granary / Italian and uploaded as resumable Parquet shards.
{
"source_dataset": "espnet/yodas-granary",
"source_config": "Italian",
"source_splits": [
"ast"
],
"validation_source_split": null,
"codec_checkpoint": "neuphonic/neucodec",
"columns": [
"text",
"codes",
"duration",
"source_dataset",
"source_config"… See the full description on the dataset page: https://huggingface.co/datasets/Steveeeeeeen/yodas-granary-it-neucodec-10s-20s.yodas-granary-it-neucodec-150k
Granary Italian VoxPopuli NeuCodec
NeuCodec-tokenized rows for NeuTTS fine-tuning. The source rows were streamed from espnet/yodas-granary / Italian and uploaded as resumable Parquet shards.
{
"source_dataset": "espnet/yodas-granary",
"source_config": "Italian",
"source_splits": [
"ast"
],
"validation_source_split": null,
"codec_checkpoint": "neuphonic/neucodec",
"columns": [
"text",
"codes",
"duration",
"source_dataset",
"source_config"… See the full description on the dataset page: https://huggingface.co/datasets/Steveeeeeeen/yodas-granary-it-neucodec-150k.yodas-granary-it-neucodec-300k-5s30s
Granary Italian VoxPopuli NeuCodec
NeuCodec-tokenized rows for NeuTTS fine-tuning. The source rows were streamed from espnet/yodas-granary / Italian and uploaded as resumable Parquet shards.
{
"source_dataset": "espnet/yodas-granary",
"source_config": "Italian",
"source_splits": [
"ast",
"asr"
],
"validation_source_split": null,
"codec_checkpoint": "neuphonic/neucodec",
"columns": [
"text",
"codes",
"duration",
"source_dataset"… See the full description on the dataset page: https://huggingface.co/datasets/Steveeeeeeen/yodas-granary-it-neucodec-300k-5s30s.emilia-yodas-english-neucodec-VJKL-250kemilia-yodas-english-neucodec-VJKL-250k-prepgranary-it-voxpopuli-neucodec
Granary Italian VoxPopuli NeuCodec
NeuCodec-tokenized rows for NeuTTS fine-tuning. The source rows were streamed from nvidia/Granary / it_voxpopuli and uploaded as resumable Parquet shards.
{
"source_dataset": "nvidia/Granary",
"source_config": "it_voxpopuli",
"source_split": "asr",
"validation_source_split": "asr",
"codec_checkpoint": "neuphonic/neucodec",
"columns": [
"text",
"codes",
"duration",
"source_dataset",
"source_config",
"source_split"… See the full description on the dataset page: https://huggingface.co/datasets/Steveeeeeeen/granary-it-voxpopuli-neucodec.emilia-yodas-english-neucodec-VFLTYemilia-yodas-english-neucodec-VJKL-25k-prepcml-tts-italian-neucodec
CML TTS Italian NeuCodec
Encoded Italian CML TTS dataset with NeuCodec speech tokens and speaker metadata.
{
"source_encoded_dataset": "Steveeeeeeen/cml-tts-italian-neucodec",
"speaker_metadata_source": "ylacombe/cml-tts/italian",
"columns": [
"text",
"codes",
"speaker_id",
"cml_split",
"cml_row_idx",
"duration",
"num_words"
],
"summary": {
"splits": {
"train": {
"rows": 35337,
"missing": 0,
"speakers": 59… See the full description on the dataset page: https://huggingface.co/datasets/Steveeeeeeen/cml-tts-italian-neucodec.emilia-yodas-english-neucodec-2000
