CoolFace
Datasetpublic

RichNachos/georgian-corpus

Georgian Corpus Dataset The Georgian Corpus Dataset is an open-source dataset designed to advance the NLP community in Georgia. It contains 5 million rows of filtered, cleaned, and deduplicated text data extracted from the Common Crawl repository. This dataset was developed as part of a bachelor’s project by: Georgi Kldiashvili Luka Paichadze Saba Shoshiashvili Dataset Summary Content: High-quality text in Georgian, suitable for NLP tasks like text… See the full description on the dataset page: https://huggingface.co/datasets/RichNachos/georgian-corpus.

sourceHugging Facegpl-3.0updated 2y agoView on Hugging Face
8likes901downloads
8 commits on main
355a2e22y ago

Update README.md

RichNachos
dfabb562y ago

Upload Report.pdf

RichNachos
928defa2y ago

Update README.md

RichNachos
eab943e2y ago

Create README.md

RichNachos
43359102y ago

Upload folder using huggingface_hub

RichNachos
a8fa8d82y ago

Upload folder using huggingface_hub

RichNachos
12196632y ago

Upload folder using huggingface_hub

RichNachos
611a24d2y ago

initial commit

RichNachos