Tilas/MoulSot-Tokens-v1
MoulSot-Tokens-v1 Discrete speech tokens for Moroccan Darija. 101 hours of transcribed speech, encoded to a single-codebook neural audio codec and paired with text, ready to train a text-to-speech model that predicts tokens directly. Size Source audio (16 kHz parquet) ~12 GB This dataset (tokens, text) 153 MB Same speech, ~78× smaller. At the codec level that is 256 kbps of PCM reduced to 0.8 kbps — a 320× reduction in bits — since in this case, one second of… See the full description on the dataset page: https://huggingface.co/datasets/Tilas/MoulSot-Tokens-v1.
139
No card is published for this repository, or it could not be fetched from Hugging Face right now.
