datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Glot500
Glot500 Corpus
A dataset of natural language data collected by putting together more than 150
existing mono-lingual and multilingual datasets together and crawling known multilingual websites.
The focus of this dataset is on 500 extremely low-resource languages.
(More Languages still to be uploaded here)
This dataset is used to train the Glot500 model.
Homepage: homepage
Repository: github
Paper: acl, arxiv
This dataset has the identical data format as the Taxi1500 Raw Data… See the full description on the dataset page: https://huggingface.co/datasets/cis-lmu/Glot500.GlotCC-V1
Dataset Summary
GlotCC-V1.0 is a document-level, general domain dataset derived from CommonCrawl, covering more than 1000 languages.It is built using the GlotLID language identification and Ungoliant pipeline from CommonCrawl.We release our pipeline as open-source at https://github.com/cisnlp/GlotCC.
List of Languages: See https://datasets-server.huggingface.co/splits?dataset=cis-lmu/GlotCC-V1 to get the list of splits available.
Usage (Huggingface Hub -- Recommended)… See the full description on the dataset page: https://huggingface.co/datasets/cis-lmu/GlotCC-V1.glotlid_processedneurips_glotocr
GlotOCR-bench
GlotOCR-bench is a dataset of 16375 images covering 158 writing systems (+2000 languages), designed to evaluate the fundamental OCR capabilities required to support diverse writing systems and languages.
Quick links:
🏆 Leaderboard
📝 License
glotlid-wordlists
GlotLID Wordlists
This is a set of wordlists extracted from the GlotLID-corpus for high-precision filtering of FineWeb2.
Download
The recommended way to download the data is via git clone:
git clone https://huggingface.co/datasets/cis-lmu/glotlid-wordlists
Method and Usage for Precision Filtering
For details on the filtering method, please refer to the FineWeb2 paper.
Each word listed occurs significantly more often in its own language dataset than in any… See the full description on the dataset page: https://huggingface.co/datasets/cis-lmu/glotlid-wordlists.neurips_glotocr
GlotOCR-bench
GlotOCR-bench is a dataset of 16375 images covering 158 writing systems (+2000 languages), designed to evaluate the fundamental OCR capabilities required to support diverse writing systems and languages.
Quick links:
🏆 Leaderboard
📝 License
GlotStoryBook
Dataset Description
Story Books for 180 ISO-639-3 codes.
The Parallel ID or parallel_id can be used to find the parallel documents in different languages and build a parallel dataset.
This dataset consists of 2 subsets:
default, which consists of 4 publishers:
asp: African Storybook
pb: Pratham Books
lcb: Little Cree Books
lida: LIDA Stories
nalibali, which comes from Nal'ibali stories.
Usage (HF Loader)
default:
from datasets import load_dataset
dataset… See the full description on the dataset page: https://huggingface.co/datasets/cis-lmu/GlotStoryBook.GlotCC-V1
Dataset Summary
GlotCC-V1.0 is a document-level, general domain dataset derived from CommonCrawl, covering more than 1000 languages.It is built using the GlotLID language identification and Ungoliant pipeline from CommonCrawl.We release our pipeline as open-source at https://github.com/cisnlp/GlotCC.
List of Languages: See https://datasets-server.huggingface.co/splits?dataset=cis-lmu/GlotCC-V1 to get the list of splits available.
Usage (Huggingface Hub -- Recommended)… See the full description on the dataset page: https://huggingface.co/datasets/innadark/GlotCC-V1.Math-Glot-Cleaned
Math-Glot-Cleaned
Math-Glot-Cleaned is a filtered and structured version of the NVIDIA AceReason-1.1-SFT dataset, focusing specifically on high-quality math and code-related question-answer pairs. This version is intended for use in training and evaluating reasoning-focused large language models, particularly for math problem solving and structured generation tasks.
Dataset Overview
Source: Derived from NVIDIA’s AceReason-1.1-SFT
Size: 22,232 examples
Format:… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Math-Glot-Cleaned.glotlid-corpus
GlotLID Corpus
This is the corpus used to train the open GlotLID model, a language identification model capable of classifying ~2000 labels.
The list of raw texts supported in this repository is openly available here: https://github.com/cisnlp/GlotLID/blob/main/sources.md
Download
The recommended way to download the data is via git clone:
git clone https://huggingface.co/datasets/cis-lmu/glotlid-corpus
Use a Hugging Face access token as the Git password/credential.… See the full description on the dataset page: https://huggingface.co/datasets/cis-lmu/glotlid-corpus.madar_on_glotlid_outputsGlotCC-V1_maltulu-3-sft-mixture-language-glotGlotOCR-bench
GlotOCR-bench
GlotOCR-bench is a dataset of 16375 images covering 158 writing systems (+2000 languages), designed to evaluate the fundamental OCR capabilities required to support diverse writing systems and languages.
Quick links:
📃 Paper
🛠️ Code
📈 Results
🏆 Leaderboard
License
This dataset is released under the GlotOCR Open Evaluation License v1.0 (see LICENSE file for full terms).
The GlotOCR-bench metadata is licensed under CC0-1.0.
The texts used to… See the full description on the dataset page: https://huggingface.co/datasets/cis-lmu/GlotOCR-bench.glotcc-v1_urls
Dataset Card for glotcc-v1_urls
This dataset provides the URLs and top-level domains associated with training records in cis-lmu/GlotCC-V1. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible.
Dataset Details
Dataset Description
This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record identifiers. In doing so, it… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/glotcc-v1_urls.nadi_glotlid_task5_dumpsglotstorybookThe GlotStoryBook dataset is a compilation of children's storybooks from the Global
Storybooks project, encompassing 174 languages organized for machine translation tasks. It
features rows containing the text segment (text number), the language code, and the file
name, which corresponds to the specific book and story segment. This structure allows for
the comparison of texts across different languages by matching file names and text numbers
between rows.GlotSparse
GlotSparse Corpus
Collection of news websites in low-resource languages.
Homepage: homepage
Repository: github
Paper: paper
Point of Contact: amir@cis.lmu.de
These languages are supported:
('azb_Arab', 'South-Azerbaijani_Arab')
('bal_Arab', 'Balochi_Arab')
('brh_Arab', 'Brahui_Arab')
('fat_Latn', 'Fanti_Latn') # aka
('glk_Arab', 'Gilaki_Arab')
('hac_Arab', 'Gurani_Arab')
('kiu_Latn', 'Kirmanjki_Latn') # zza
('sdh_Arab', 'Southern-Kurdish_Arab')
('twi_Latn', 'Twi_Latn') # aka… See the full description on the dataset page: https://huggingface.co/datasets/cis-lmu/GlotSparse.glot500-div-thaa
Glot500 DATASET
Glot500 (https://huggingface.co/datasets/cis-lmu/Glot500)
glot500_amh_Ethi_trainGlotCC-V1-jav-Latnglot-dataglot500_amh_trainglot500_amh_Ethi_testglot500_bak_Cyrl_devGlotCC-V1-jav-Latn-content-onlyglot500_bak_Cyrl_testGlot500-swahiliGlotOCR-bench-v1.0-results
GlotOCR Bench Results
This dataset contains OCR results from images in cis-lmu/GlotOCR-bench using multiple multilingual OCR models.
Quick links:
📃 Paper
🛠️ Code
📈 Benchmark
Download
git clone https://huggingface.co/datasets/cis-lmu/GlotOCR-bench-v1.0-results
If you want to clone without large files - just their pointers:
GIT_LFS_SKIP_SMUDGE=1 git clone https://huggingface.co/datasets/cis-lmu/GlotOCR-bench-v1.0-results
Processing Details
Source… See the full description on the dataset page: https://huggingface.co/datasets/cis-lmu/GlotOCR-bench-v1.0-results.glot500_yid_Hebr_dev
