cis
Datasets
All datasets matching “cis”Glot500
Glot500 Corpus
A dataset of natural language data collected by putting together more than 150
existing mono-lingual and multilingual datasets together and crawling known multilingual websites.
The focus of this dataset is on 500 extremely low-resource languages.
(More Languages still to be uploaded here)
This dataset is used to train the Glot500 model.
Homepage: homepage
Repository: github
Paper: acl, arxiv
This dataset has the identical data format as the Taxi1500 Raw Data… See the full description on the dataset page: https://huggingface.co/datasets/cis-lmu/Glot500.cissp-llmbench
CISSP-LLMBench
GlotCC-V1
Dataset Summary
GlotCC-V1.0 is a document-level, general domain dataset derived from CommonCrawl, covering more than 1000 languages.It is built using the GlotLID language identification and Ungoliant pipeline from CommonCrawl.We release our pipeline as open-source at https://github.com/cisnlp/GlotCC.
List of Languages: See https://datasets-server.huggingface.co/splits?dataset=cis-lmu/GlotCC-V1 to get the list of splits available.
Usage (Huggingface Hub -- Recommended)… See the full description on the dataset page: https://huggingface.co/datasets/cis-lmu/GlotCC-V1.Taxi1500-RawData
Taxi1500 Raw Data
Introduction
This repository contains the raw text data of the Taxi1500-c_v3.0 corpus, without classification labels and Bible verse ids. For the original Taxi1500 dataset for Text Classification, please refer to the GitHub repository.
The data format of the Taxi1500-RawData is identical to that of the Glot500 Dataset, facilitating seamless parallel utilization of both datasets.
Usage
Replace acr_Latn with your specific language.
from… See the full description on the dataset page: https://huggingface.co/datasets/cis-lmu/Taxi1500-RawData.open-greek-corpus-annotations
Open Greek Corpus Annotations
Token-level linguistic annotations for the
Open Greek Corpus:
lemma, part of speech (UD UPOS), and morphology (UD features) for every
served token. Three provenance classes, never confused thanks to per-token
provenance and confidence tiers: gold treebank annotations where an openly
licensed MANUAL treebank covers a work (GLAUx's treebank layers, MACULA
Greek for the NT), GLAUx's own automatic annotation as the middle auto:
class, and model… See the full description on the dataset page: https://huggingface.co/datasets/ciscoriordan/open-greek-corpus-annotations.aurora_rollout_beta
