CoolFace
21 results

cis

cis-lmu /Glot500 Glot500 Corpus A dataset of natural language data collected by putting together more than 150 existing mono-lingual and multilingual datasets together and crawling known multilingual websites. The focus of this dataset is on 500 extremely low-resource languages. (More Languages still to be uploaded here) This dataset is used to train the Glot500 model. Homepage: homepage Repository: github Paper: acl, arxiv This dataset has the identical data format as the Taxi1500 Raw Data… See the full description on the dataset page: https://huggingface.co/datasets/cis-lmu/Glot500.text1B<n<10B43 likes48k downloads10mo agoHugging Facefahadhafeezofficial /cissp-llmbench CISSP-LLMBench tabulartext-generation10K<n<100K0 likes3.2k downloads3mo agoHugging Facecis-lmu /GlotCC-V1 Dataset Summary GlotCC-V1.0 is a document-level, general domain dataset derived from CommonCrawl, covering more than 1000 languages.It is built using the GlotLID language identification and Ungoliant pipeline from CommonCrawl.We release our pipeline as open-source at https://github.com/cisnlp/GlotCC. List of Languages: See https://datasets-server.huggingface.co/splits?dataset=cis-lmu/GlotCC-V1 to get the list of splits available. Usage (Huggingface Hub -- Recommended)… See the full description on the dataset page: https://huggingface.co/datasets/cis-lmu/GlotCC-V1.tabular1B<n<10B61 likes2.1k downloads2y agoHugging Facecis-lmu /Taxi1500-RawData Taxi1500 Raw Data Introduction This repository contains the raw text data of the Taxi1500-c_v3.0 corpus, without classification labels and Bible verse ids. For the original Taxi1500 dataset for Text Classification, please refer to the GitHub repository. The data format of the Taxi1500-RawData is identical to that of the Glot500 Dataset, facilitating seamless parallel utilization of both datasets. Usage Replace acr_Latn with your specific language. from… See the full description on the dataset page: https://huggingface.co/datasets/cis-lmu/Taxi1500-RawData.text10M<n<100M2 likes2.1k downloads2y agoHugging Faceciscoriordan /open-greek-corpus-annotations Open Greek Corpus Annotations Token-level linguistic annotations for the Open Greek Corpus: lemma, part of speech (UD UPOS), and morphology (UD features) for every served token. Three provenance classes, never confused thanks to per-token provenance and confidence tiers: gold treebank annotations where an openly licensed MANUAL treebank covers a work (GLAUx's treebank layers, MACULA Greek for the NT), GLAUx's own automatic annotation as the middle auto: class, and model… See the full description on the dataset page: https://huggingface.co/datasets/ciscoriordan/open-greek-corpus-annotations.token-classification0 likes1k downloads2mo agoHugging Facenasa-cisto-data-science-group /aurora_rollout_beta0 likes959 downloads2y agoHugging Face