CoolFace
20 results

glot

cis-lmu /Glot500 Glot500 Corpus A dataset of natural language data collected by putting together more than 150 existing mono-lingual and multilingual datasets together and crawling known multilingual websites. The focus of this dataset is on 500 extremely low-resource languages. (More Languages still to be uploaded here) This dataset is used to train the Glot500 model. Homepage: homepage Repository: github Paper: acl, arxiv This dataset has the identical data format as the Taxi1500 Raw Data… See the full description on the dataset page: https://huggingface.co/datasets/cis-lmu/Glot500.text1B<n<10B43 likes41k downloads10mo agoHugging Facecis-lmu /GlotCC-V1 Dataset Summary GlotCC-V1.0 is a document-level, general domain dataset derived from CommonCrawl, covering more than 1000 languages.It is built using the GlotLID language identification and Ungoliant pipeline from CommonCrawl.We release our pipeline as open-source at https://github.com/cisnlp/GlotCC. List of Languages: See https://datasets-server.huggingface.co/splits?dataset=cis-lmu/GlotCC-V1 to get the list of splits available. Usage (Huggingface Hub -- Recommended)… See the full description on the dataset page: https://huggingface.co/datasets/cis-lmu/GlotCC-V1.tabular1B<n<10B61 likes2.1k downloads2y agoHugging FaceJakh0103 /glotlid_processedtext100M<n<1B1 likes1k downloads2y agoHugging FaceAtharvImmverse /neurips_glotocr GlotOCR-bench GlotOCR-bench is a dataset of 16375 images covering 158 writing systems (+2000 languages), designed to evaluate the fundamental OCR capabilities required to support diverse writing systems and languages. Quick links: 🏆 Leaderboard 📝 License imageimage-to-text10K<n<100K0 likes314 downloads2mo agoHugging Facecis-lmu /glotlid-wordlists GlotLID Wordlists This is a set of wordlists extracted from the GlotLID-corpus for high-precision filtering of FineWeb2. Download The recommended way to download the data is via git clone: git clone https://huggingface.co/datasets/cis-lmu/glotlid-wordlists Method and Usage for Precision Filtering For details on the filtering method, please refer to the FineWeb2 paper. Each word listed occurs significantly more often in its own language dataset than in any… See the full description on the dataset page: https://huggingface.co/datasets/cis-lmu/glotlid-wordlists.text1M<n<10M2 likes241 downloads1y agoHugging Facenlplabedu /neurips_glotocr GlotOCR-bench GlotOCR-bench is a dataset of 16375 images covering 158 writing systems (+2000 languages), designed to evaluate the fundamental OCR capabilities required to support diverse writing systems and languages. Quick links: 🏆 Leaderboard 📝 License imageimage-to-text10K<n<100K0 likes212 downloads5mo agoHugging Face