pagantibet/normalisation-S2S-training
Tibetan Normalisation - S2S Training Data A large-scale parallel training dataset for Tibetan text normalisation, containing approximately 2 million line pairs mapping diplomatic (non-standard, abbreviated) Tibetan manuscript text to Standard Classical Tibetan. This dataset was used to train the sequence-to-sequence normalisation models (tokenised S2S model and non-tokenised S2S model) released as part of the PaganTibet project. The dataset combines a manually curated… See the full description on the dataset page: https://huggingface.co/datasets/pagantibet/normalisation-S2S-training.
053
