CoolFace
Datasetpublic

Kemsekov/Corrupted-russian-word-documents-text-dataset

This is synthetic text-corruption dataset, based on Russian open official/government documents text. This dataset is intended to be used to train LLM to perform text-recovery task. All the errors in text is made solely in Russian sentences, hence ignoring any English sentence. Texts contains complex formatting, which is common for documents. Each line contains json object that have array messages value, which consists of role-based conversation. Each first message is randomly chosen system… See the full description on the dataset page: https://huggingface.co/datasets/Kemsekov/Corrupted-russian-word-documents-text-dataset.

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
2likes62downloads
4 commits on main
bb75ad52y ago

Update

Kemsekov
3e5f5492y ago

Create README.md

Kemsekov
aa36cb12y ago

Upload 6 files

Kemsekov
c975c762y ago

initial commit

Kemsekov