CoolFace
Datasetpublic

pagantibet/normalisation-S2S-training

Tibetan Normalisation - S2S Training Data A large-scale parallel training dataset for Tibetan text normalisation, containing approximately 2 million line pairs mapping diplomatic (non-standard, abbreviated) Tibetan manuscript text to Standard Classical Tibetan. This dataset was used to train the sequence-to-sequence normalisation models (tokenised S2S model and non-tokenised S2S model) released as part of the PaganTibet project. The dataset combines a manually curated… See the full description on the dataset page: https://huggingface.co/datasets/pagantibet/normalisation-S2S-training.

sourceHugging Facecc-by-nc-sa-4.0updated 6mo agoView on Hugging Face
0likes53downloads
settings

This repository belongs to pagantibet on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namenormalisation-S2S-training
visibilitypublic
licencecc-by-nc-sa-4.0
gatedno
ownerpagantibet
Account settings
pagantibet/normalisation-S2S-training · CoolFace