CoolFace
Datasetpublic

niloy629/s2m-bang

s2m-bang Bangla Stage-1 speech→transcript dataset used for speech-to-LLM adapter training (Moonshine-BN encoder + LLM adapter). Splits Manifest n Role manifests/train.json 89,275 Train (Kathbath + CV/OpenSLR short + FLEURS train) manifests/dev.json 2,638 Dev manifests/dev_fast.json 256 Fast mid-train gate manifests/fleurs_test.json 920 Held-out FLEURS test manifests/cv_bn_eval.json 260 Held-out Common Voice eval Schema Each… See the full description on the dataset page: https://huggingface.co/datasets/niloy629/s2m-bang.

sourceHugging Faceotherupdated 2mo agoView on Hugging Face
0likes879downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
niloy629/s2m-bang · CoolFace