CoolFace
Datasetpublic

LeeHarrold/musiccaps-mot-tokens

MusicCaps Pre-Encoded Tokens for Mixture-of-Transformers (MoT) Dataset Description This dataset contains pre-encoded audio tokens from the MusicCaps dataset, processed through Meta's MusicGen EnCodec tokenizer for use in Mixture-of-Transformers (MoT) training. Dataset Summary 5,233 music clips encoded as discrete tokens 4 codebook layers from MusicGen's EnCodec ~500 tokens per 10-second clip Compressed from ~12GB audio to 82MB tokens Ready for… See the full description on the dataset page: https://huggingface.co/datasets/LeeHarrold/musiccaps-mot-tokens.

sourceHugging Faceupdated 10mo agoView on Hugging Face
0likes23downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
LeeHarrold/musiccaps-mot-tokens · CoolFace