glot
Datasets
All datasets matching “glot”Glot500
Glot500 Corpus
A dataset of natural language data collected by putting together more than 150
existing mono-lingual and multilingual datasets together and crawling known multilingual websites.
The focus of this dataset is on 500 extremely low-resource languages.
(More Languages still to be uploaded here)
This dataset is used to train the Glot500 model.
Homepage: homepage
Repository: github
Paper: acl, arxiv
This dataset has the identical data format as the Taxi1500 Raw Data… See the full description on the dataset page: https://huggingface.co/datasets/cis-lmu/Glot500.GlotCC-V1
Dataset Summary
GlotCC-V1.0 is a document-level, general domain dataset derived from CommonCrawl, covering more than 1000 languages.It is built using the GlotLID language identification and Ungoliant pipeline from CommonCrawl.We release our pipeline as open-source at https://github.com/cisnlp/GlotCC.
List of Languages: See https://datasets-server.huggingface.co/splits?dataset=cis-lmu/GlotCC-V1 to get the list of splits available.
Usage (Huggingface Hub -- Recommended)… See the full description on the dataset page: https://huggingface.co/datasets/cis-lmu/GlotCC-V1.glotlid_processedneurips_glotocr
GlotOCR-bench
GlotOCR-bench is a dataset of 16375 images covering 158 writing systems (+2000 languages), designed to evaluate the fundamental OCR capabilities required to support diverse writing systems and languages.
Quick links:
🏆 Leaderboard
📝 License
glotlid-wordlists
GlotLID Wordlists
This is a set of wordlists extracted from the GlotLID-corpus for high-precision filtering of FineWeb2.
Download
The recommended way to download the data is via git clone:
git clone https://huggingface.co/datasets/cis-lmu/glotlid-wordlists
Method and Usage for Precision Filtering
For details on the filtering method, please refer to the FineWeb2 paper.
Each word listed occurs significantly more often in its own language dataset than in any… See the full description on the dataset page: https://huggingface.co/datasets/cis-lmu/glotlid-wordlists.neurips_glotocr
GlotOCR-bench
GlotOCR-bench is a dataset of 16375 images covering 158 writing systems (+2000 languages), designed to evaluate the fundamental OCR capabilities required to support diverse writing systems and languages.
Quick links:
🏆 Leaderboard
📝 License
