CoolFace
Datasetpublic

Kemsekov/Corrupted-russian-word-documents-text-dataset

This is synthetic text-corruption dataset, based on Russian open official/government documents text. This dataset is intended to be used to train LLM to perform text-recovery task. All the errors in text is made solely in Russian sentences, hence ignoring any English sentence. Texts contains complex formatting, which is common for documents. Each line contains json object that have array messages value, which consists of role-based conversation. Each first message is randomly chosen system… See the full description on the dataset page: https://huggingface.co/datasets/Kemsekov/Corrupted-russian-word-documents-text-dataset.

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
2likes62downloads

Kemsekov/Corrupted-russian-word-documents-text-dataset · main · files are served by the source, never re-hosted here