CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01cis-lmu /Glot500 Glot500 Corpus A dataset of natural language data collected by putting together more than 150 existing mono-lingual and multilingual datasets together and crawling known multilingual websites. The focus of this dataset is on 500 extremely low-resource languages. (More Languages still to be uploaded here) This dataset is used to train the Glot500 model. Homepage: homepage Repository: github Paper: acl, arxiv This dataset has the identical data format as the Taxi1500 Raw Data… See the full description on the dataset page: https://huggingface.co/datasets/cis-lmu/Glot500.text1B<n<10B43 likes41k downloads10mo agoHugging Face02cis-lmu /GlotCC-V1 Dataset Summary GlotCC-V1.0 is a document-level, general domain dataset derived from CommonCrawl, covering more than 1000 languages.It is built using the GlotLID language identification and Ungoliant pipeline from CommonCrawl.We release our pipeline as open-source at https://github.com/cisnlp/GlotCC. List of Languages: See https://datasets-server.huggingface.co/splits?dataset=cis-lmu/GlotCC-V1 to get the list of splits available. Usage (Huggingface Hub -- Recommended)… See the full description on the dataset page: https://huggingface.co/datasets/cis-lmu/GlotCC-V1.tabular1B<n<10B61 likes2.1k downloads2y agoHugging Face03Jakh0103 /glotlid_processedtext100M<n<1B1 likes1k downloads2y agoHugging Face04AtharvImmverse /neurips_glotocr GlotOCR-bench GlotOCR-bench is a dataset of 16375 images covering 158 writing systems (+2000 languages), designed to evaluate the fundamental OCR capabilities required to support diverse writing systems and languages. Quick links: 🏆 Leaderboard 📝 License imageimage-to-text10K<n<100K0 likes314 downloads2mo agoHugging Face05cis-lmu /glotlid-wordlists GlotLID Wordlists This is a set of wordlists extracted from the GlotLID-corpus for high-precision filtering of FineWeb2. Download The recommended way to download the data is via git clone: git clone https://huggingface.co/datasets/cis-lmu/glotlid-wordlists Method and Usage for Precision Filtering For details on the filtering method, please refer to the FineWeb2 paper. Each word listed occurs significantly more often in its own language dataset than in any… See the full description on the dataset page: https://huggingface.co/datasets/cis-lmu/glotlid-wordlists.text1M<n<10M2 likes241 downloads1y agoHugging Face06nlplabedu /neurips_glotocr GlotOCR-bench GlotOCR-bench is a dataset of 16375 images covering 158 writing systems (+2000 languages), designed to evaluate the fundamental OCR capabilities required to support diverse writing systems and languages. Quick links: 🏆 Leaderboard 📝 License imageimage-to-text10K<n<100K0 likes212 downloads5mo agoHugging Face07cis-lmu /GlotStoryBook Dataset Description Story Books for 180 ISO-639-3 codes. The Parallel ID or parallel_id can be used to find the parallel documents in different languages and build a parallel dataset. This dataset consists of 2 subsets: default, which consists of 4 publishers: asp: African Storybook pb: Pratham Books lcb: Little Cree Books lida: LIDA Stories nalibali, which comes from Nal'ibali stories. Usage (HF Loader) default: from datasets import load_dataset dataset… See the full description on the dataset page: https://huggingface.co/datasets/cis-lmu/GlotStoryBook.texttranslation10K<n<100K9 likes172 downloads16h agoHugging Face08innadark /GlotCC-V1 Dataset Summary GlotCC-V1.0 is a document-level, general domain dataset derived from CommonCrawl, covering more than 1000 languages.It is built using the GlotLID language identification and Ungoliant pipeline from CommonCrawl.We release our pipeline as open-source at https://github.com/cisnlp/GlotCC. List of Languages: See https://datasets-server.huggingface.co/splits?dataset=cis-lmu/GlotCC-V1 to get the list of splits available. Usage (Huggingface Hub -- Recommended)… See the full description on the dataset page: https://huggingface.co/datasets/innadark/GlotCC-V1.tabular1B<n<10B1 likes161 downloads9mo agoHugging Face09prithivMLmods /Math-Glot-Cleaned Math-Glot-Cleaned Math-Glot-Cleaned is a filtered and structured version of the NVIDIA AceReason-1.1-SFT dataset, focusing specifically on high-quality math and code-related question-answer pairs. This version is intended for use in training and evaluating reasoning-focused large language models, particularly for math problem solving and structured generation tasks. Dataset Overview Source: Derived from NVIDIA’s AceReason-1.1-SFT Size: 22,232 examples Format:… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Math-Glot-Cleaned.texttext-generation10K<n<100K0 likes76 downloads1y agoHugging Face10cis-lmu /glotlid-corpusgated GlotLID Corpus This is the corpus used to train the open GlotLID model, a language identification model capable of classifying ~2000 labels. The list of raw texts supported in this repository is openly available here: https://github.com/cisnlp/GlotLID/blob/main/sources.md Download The recommended way to download the data is via git clone: git clone https://huggingface.co/datasets/cis-lmu/glotlid-corpus Use a Hugging Face access token as the Git password/credential.… See the full description on the dataset page: https://huggingface.co/datasets/cis-lmu/glotlid-corpus.text100M<n<1B15 likes75 downloads6mo agoHugging Face11NadiaGHEZAIEL /madar_on_glotlid_outputstext10M<n<100M0 likes68 downloads1y agoHugging Face12Navneeth017 /GlotCC-V1_maltabulartext-generation100K<n<1M0 likes68 downloads8mo agoHugging Face13sartifyllc /tulu-3-sft-mixture-language-glottext100K<n<1M2 likes62 downloads2y agoHugging Face14cis-lmu /GlotOCR-benchgated GlotOCR-bench GlotOCR-bench is a dataset of 16375 images covering 158 writing systems (+2000 languages), designed to evaluate the fundamental OCR capabilities required to support diverse writing systems and languages. Quick links: 📃 Paper 🛠️ Code 📈 Results 🏆 Leaderboard License This dataset is released under the GlotOCR Open Evaluation License v1.0 (see LICENSE file for full terms). The GlotOCR-bench metadata is licensed under CC0-1.0. The texts used to… See the full description on the dataset page: https://huggingface.co/datasets/cis-lmu/GlotOCR-bench.imageimage-to-text10K<n<100K6 likes49 downloads5mo agoHugging Face15nhagar /glotcc-v1_urls Dataset Card for glotcc-v1_urls This dataset provides the URLs and top-level domains associated with training records in cis-lmu/GlotCC-V1. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible. Dataset Details Dataset Description This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record identifiers. In doing so, it… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/glotcc-v1_urls.text100M<n<1B1 likes47 downloads1y agoHugging Face16NadiaGHEZAIEL /nadi_glotlid_task5_dumpstext10M<n<100M0 likes47 downloads1y agoHugging Face17SEACrowd /glotstorybookThe GlotStoryBook dataset is a compilation of children's storybooks from the Global Storybooks project, encompassing 174 languages organized for machine translation tasks. It features rows containing the text segment (text number), the language code, and the file name, which corresponds to the specific book and story segment. This structure allows for the comparison of texts across different languages by matching file names and text numbers between rows.0 likes43 downloads2y agoHugging Face18cis-lmu /GlotSparsegated GlotSparse Corpus Collection of news websites in low-resource languages. Homepage: homepage Repository: github Paper: paper Point of Contact: amir@cis.lmu.de These languages are supported: ('azb_Arab', 'South-Azerbaijani_Arab') ('bal_Arab', 'Balochi_Arab') ('brh_Arab', 'Brahui_Arab') ('fat_Latn', 'Fanti_Latn') # aka ('glk_Arab', 'Gilaki_Arab') ('hac_Arab', 'Gurani_Arab') ('kiu_Latn', 'Kirmanjki_Latn') # zza ('sdh_Arab', 'Southern-Kurdish_Arab') ('twi_Latn', 'Twi_Latn') # aka… See the full description on the dataset page: https://huggingface.co/datasets/cis-lmu/GlotSparse.text100K<n<1M7 likes37 downloads2y agoHugging Face19alakxender /glot500-div-thaa Glot500 DATASET Glot500 (https://huggingface.co/datasets/cis-lmu/Glot500) text1M<n<10M0 likes23 downloads2y agoHugging Face20yiyic /glot500_amh_Ethi_traintext1M<n<10M0 likes22 downloads2y agoHugging Face21akahana /GlotCC-V1-jav-Latntabular10K<n<100K0 likes22 downloads2y agoHugging Face22belatijagad /glot-data0 likes20 downloads2y agoHugging Face23yiyic /glot500_amh_traintext1M<n<10M0 likes18 downloads2y agoHugging Face24yiyic /glot500_amh_Ethi_testtext1K<n<10K0 likes17 downloads2y agoHugging Face25yiyic /glot500_bak_Cyrl_devtext1K<n<10K0 likes17 downloads2y agoHugging Face26akahana /GlotCC-V1-jav-Latn-content-onlytext10K<n<100K0 likes17 downloads2y agoHugging Face27yiyic /glot500_bak_Cyrl_testtext1K<n<10K0 likes15 downloads2y agoHugging Face28Aletheia-ng /Glot500-swahilitext10K<n<100K0 likes15 downloads2y agoHugging Face29cis-lmu /GlotOCR-bench-v1.0-resultsgated GlotOCR Bench Results This dataset contains OCR results from images in cis-lmu/GlotOCR-bench using multiple multilingual OCR models. Quick links: 📃 Paper 🛠️ Code 📈 Benchmark Download git clone https://huggingface.co/datasets/cis-lmu/GlotOCR-bench-v1.0-results If you want to clone without large files - just their pointers: GIT_LFS_SKIP_SMUDGE=1 git clone https://huggingface.co/datasets/cis-lmu/GlotOCR-bench-v1.0-results Processing Details Source… See the full description on the dataset page: https://huggingface.co/datasets/cis-lmu/GlotOCR-bench-v1.0-results.image100K<n<1M4 likes15 downloads5mo agoHugging Face30yiyic /glot500_yid_Hebr_devtext1K<n<10K0 likes14 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.