CoolFace
Datasetpublic

Kemsekov/Corrupted-russian-word-documents-text-dataset

This is synthetic text-corruption dataset, based on Russian open official/government documents text. This dataset is intended to be used to train LLM to perform text-recovery task. All the errors in text is made solely in Russian sentences, hence ignoring any English sentence. Texts contains complex formatting, which is common for documents. Each line contains json object that have array messages value, which consists of role-based conversation. Each first message is randomly chosen system… See the full description on the dataset page: https://huggingface.co/datasets/Kemsekov/Corrupted-russian-word-documents-text-dataset.

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
2likes61downloads
settings

This repository belongs to Kemsekov on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameCorrupted-russian-word-documents-text-dataset
visibilitypublic
licenceapache-2.0
gatedno
ownerKemsekov
Account settings
Kemsekov/Corrupted-russian-word-documents-text-dataset · CoolFace