CoolFace
Datasetpublic

RichNachos/georgian-corpus

Georgian Corpus Dataset The Georgian Corpus Dataset is an open-source dataset designed to advance the NLP community in Georgia. It contains 5 million rows of filtered, cleaned, and deduplicated text data extracted from the Common Crawl repository. This dataset was developed as part of a bachelor’s project by: Georgi Kldiashvili Luka Paichadze Saba Shoshiashvili Dataset Summary Content: High-quality text in Georgian, suitable for NLP tasks like text… See the full description on the dataset page: https://huggingface.co/datasets/RichNachos/georgian-corpus.

sourceHugging Facegpl-3.0updated 2y agoView on Hugging Face
8likes760downloads
Dataset Card

Georgian Corpus Dataset

The Georgian Corpus Dataset is an open-source dataset designed to advance the NLP community in Georgia. It contains 5 million rows of filtered, cleaned, and deduplicated text data extracted from the Common Crawl repository.

This dataset was developed as part of a bachelor’s project by:

Dataset Summary

  • Content: High-quality text in Georgian, suitable for NLP tasks like text classification, translation, and language modeling.
  • Source: Extracted from Common Crawl, then processed to ensure data quality through filtering, cleaning, and deduplication.
  • Size: Slightly more than 5 million rows.

Motivation

We created this dataset to address the limited availability of Georgian-language datasets and to contribute to the global and local open-source NLP community.

How to Use

The dataset can be used for a variety of tasks, including:

  • Language Modeling: Pretraining or fine-tuning Georgian language models.
  • Machine Translation: Building models for translation between Georgian and other languages.
  • Text Classification: Training classifiers for sentiment analysis or topic categorization.

For usage examples and additional details, please refer to the project report.

License

This dataset is distributed under the GPL-3.0 License.

Citation

If you use this dataset in your research or projects, please cite us.