CoolFace
Datasetpublic

RichNachos/georgian-corpus

Georgian Corpus Dataset The Georgian Corpus Dataset is an open-source dataset designed to advance the NLP community in Georgia. It contains 5 million rows of filtered, cleaned, and deduplicated text data extracted from the Common Crawl repository. This dataset was developed as part of a bachelor’s project by: Georgi Kldiashvili Luka Paichadze Saba Shoshiashvili Dataset Summary Content: High-quality text in Georgian, suitable for NLP tasks like text… See the full description on the dataset page: https://huggingface.co/datasets/RichNachos/georgian-corpus.

sourceHugging Facegpl-3.0updated 2y agoView on Hugging Face
8likes901downloads
settings

This repository belongs to RichNachos on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namegeorgian-corpus
visibilitypublic
licencegpl-3.0
gatedno
ownerRichNachos
Account settings
RichNachos/georgian-corpus · CoolFace