CoolFace
Datasetpublic

pagantibet/normalisation-S2S-training

Tibetan Normalisation - S2S Training Data A large-scale parallel training dataset for Tibetan text normalisation, containing approximately 2 million line pairs mapping diplomatic (non-standard, abbreviated) Tibetan manuscript text to Standard Classical Tibetan. This dataset was used to train the sequence-to-sequence normalisation models (tokenised S2S model and non-tokenised S2S model) released as part of the PaganTibet project. The dataset combines a manually curated… See the full description on the dataset page: https://huggingface.co/datasets/pagantibet/normalisation-S2S-training.

sourceHugging Facecc-by-nc-sa-4.0updated 6mo agoView on Hugging Face
0likes53downloads

pagantibet/normalisation-S2S-training · main · files are served by the source, never re-hosted here