CoolFace
19 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tokenizer-eval /ud-treebank-tokens Dataset Card for Dataset Name Dataset Summary This is a subset of the Universal Dependencies Treebanks dataset (version 2.12) which only contains raw sentences and their corresponding tokenized form. This dataset is licensed under the Universal Dependencies v2.10 License Agreement. The user is reminded that some subsets only allow non-commercial use. Each row contains the license of the dataset it originated from, and a full listing of all included subsets and… See the full description on the dataset page: https://huggingface.co/datasets/tokenizer-eval/ud-treebank-tokens.text1M<n<10M0 likes1.2k downloads3y agoHugging Face02dgabri3le /english-ipa-dep-treebank English IPA Dependency Treebank A large-scale dataset of 10.4 million English sentences paired with IPA (International Phonetic Alphabet) transcriptions and Universal Dependencies syntactic annotations. Each sentence includes its full dependency parse — head indices, relation labels, and a linearized tagged-IPA representation that interleaves syntactic roles with phonetic content. Dataset Structure Each sample contains: Field Type Description raw_english string… See the full description on the dataset page: https://huggingface.co/datasets/dgabri3le/english-ipa-dep-treebank.texttext-generation10M<n<100M1 likes206 downloads5mo agoHugging Face03Lots-of-LoRAs /task1167_penn_treebank_coarse_pos_tagging Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1167_penn_treebank_coarse_pos_tagging Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1167_penn_treebank_coarse_pos_tagging.texttext-generation1K<n<10K1 likes82 downloads2y agoHugging Face04daidalos-project /latin_treebanks_ud_testNOTE: This template for datasheets for ancient language data is based on the proposal by Gebru et al. 2021 https://arxiv.org/abs/1803.09010. The majority of questions is taken from there, but in a slightly rearanged order. However, as some questions are not relevant for historical data, they have been left out. Questions which are of relevance for Humanities scholars researching ancient languages have been added. The question have been answered to the best of the knowledge of the Daidalos… See the full description on the dataset page: https://huggingface.co/datasets/daidalos-project/latin_treebanks_ud_test.texttoken-classification1K<n<10K0 likes82 downloads1mo agoHugging Face05gmikros /kathnlp-treebank kathnlp Katharevousa Greek Treebank A Universal-Dependencies-style reference treebank for Katharevousa Greek, the archaizing official register used in 20th-century Greek law, administration, and parliamentary discourse. The treebank covers 1,697 sentences from written parliamentary questions of the early Third Hellenic Republic (1976–1977) and is released alongside the kathnlp parsing pipeline. Paper: A Reproducible Universal Dependencies-Style Pipeline for Katharevousa Greek… See the full description on the dataset page: https://huggingface.co/datasets/gmikros/kathnlp-treebank.texttoken-classification1K<n<10K0 likes70 downloads4mo agoHugging Face06rohith2812 /stanford-sentiment-treebank-datasetThe text file contains the collection of short movie reviews. Each review is enclosed in parentheses and consists of a numerical rating followed by the review text. The numerical rating is on a scale of 1 to 4, where higher numbers indicate a more positive review. Here are some additional observations: Format: The reviews follow a consistent format with the rating at the beginning, making it easy to identify the sentiment of each review. Concise: The reviews are generally concise, focusing… See the full description on the dataset page: https://huggingface.co/datasets/rohith2812/stanford-sentiment-treebank-dataset.text1K<n<10K1 likes46 downloads2y agoHugging Face07SEACrowd /parallel_asian_treebankThe ALT project aims to advance the state-of-the-art Asian natural language processing (NLP) techniques through the open collaboration for developing and using ALT. It was first conducted by NICT and UCSY as described in Ye Kyaw Thu, Win Pa Pa, Masao Utiyama, Andrew Finch and Eiichiro Sumita (2016). Then, it was developed under ASEAN IVO. The process of building ALT began with sampling about 20,000 sentences from English Wikinews, and then these sentences were translated into the other languages. ALT now has 13 languages: Bengali, English, Filipino, Hindi, Bahasa Indonesia, Japanese, Khmer, Lao, Malay, Myanmar (Burmese), Thai, Vietnamese, Chinese (Simplified Chinese).0 likes34 downloads2y agoHugging Face08Goader /ukrainian-treebank-lmUkrainian part of the Universal Dependencies, specifically preprocessed for the language modeling task. The data can be split into documents, paragraphs or sentences. Manual selection of the data done by the authors of the dataset makes it suitable for the perplexity evaluation. Authors of the dataset: Institute for Ukrainian, NGO, org@mova.institute GitHub: https://github.com/UniversalDependencies/UD_Ukrainian-IUtextfill-mask1K<n<10K0 likes32 downloads3y agoHugging Face09ReySajju742 /urdu-universal-dependency-treebank Summary The Urdu Universal Dependency Treebank was automatically converted from Urdu Dependency Treebank (UDTB) which is part of an ongoing effort of creating multi-layered treebanks for Hindi and Urdu. Acknowledgments UDTB is developed at IIIT-H India. The project is supported by NSF Grant (Award Number: CNS 0751202; CFDA Number: 47.070). Any publication reporting the work done using this data should cite the following references: Riyaz Ahmad Bhat, Rajesh Bhatt, Annahita… See the full description on the dataset page: https://huggingface.co/datasets/ReySajju742/urdu-universal-dependency-treebank.text-classification10K<n<100K0 likes29 downloads1y agoHugging Face10SEACrowd /alt_burmese_treebankA 20,000-sentence Burmese (Myanmar) treebank on news articles containing complete phrase structure annotation.As the final result of the Burmese component in the Asian Language Treebank Project, this is the first large-scale,open-access treebank for the Burmese language.0 likes26 downloads2y agoHugging Face11Akindelevictoria /yoruba-constituency-treebank Yoruba Constituency Treebank (Version 1.0) Overview This dataset contains a manually annotated Yoruba constituency treebank consisting of 1,000 sentences. The treebank was developed as part of an undergraduate linguistics research project focused on Yoruba syntax and computational parsing for under-resourced languages. The annotations follow a phrase-structure (constituency) framework, including labels such as NP, VP, IP, and CP. Dataset Contents The repository… See the full description on the dataset page: https://huggingface.co/datasets/Akindelevictoria/yoruba-constituency-treebank.other1K<n<10K0 likes26 downloads9mo agoHugging Face12turkish-nlp-suite /Treebank-Benchmarking Turkish Treebank Benchmarking This is the repo for Turkish treebank benchmarking, namely evaluating Tranformer models on POS-Dep-Morph task. For the data, we used two treebank, IMST and BOUN. We converted conllu format to json lines for being compatible to HF dataset formats. Here are treebank sizes at a glance: Dataset train lines dev lines test lines BOUN 7803 979 979 IMST 3435 1100 1100 A typical instance from the dataset looks like: { "id": "ins_1267"… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/Treebank-Benchmarking.text10K<n<100K0 likes23 downloads7mo agoHugging Face13anishka /UD_Treebank_Te_TransliterateTELUGU dataset converted to Transliterate0 likes19 downloads3y agoHugging Face14BUCOLIN /OTA-BOUN_UD_Treebank0 likes19 downloads2y agoHugging Face15proxectonos /TreebankNos TreebankNos TreebankNos is a Galician-language dataset combining Universal Dependencies (UD) treebank annotations with Named Entity Recognition (NER) labels. It is built upon two established UD corpora — UD TreeGal and UD Parallel Universal Dependencies (PUD) — extended with BIO-format NER annotations covering four entity types: Person, Location, Organisation, and Miscellaneous. The dataset is intended for multi-task NLP research in Galician, supporting POS tagging… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/TreebankNos.texttoken-classification10K<n<100K0 likes18 downloads4mo agoHugging Face16pythainlp /blackboard_treebank_prompt Dataset Card for "blackboard_treebank_prompt" This dataset made from blackboard treebank. The dataset want to create Thai sentence by structure. The original dataset used own tags but we use Universal Dependencies tags, so we convert those tags into Universal Dependencies tags. See blackboard treebank tags to Universal Dependencies tags Source code for create dataset: https://github.com/PyThaiNLP/support-aya-datasets/blob/main/pos/blackboard_treebank_prompt.ipynb Template… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/blackboard_treebank_prompt.texttext-generation100K<n<1M0 likes16 downloads3y agoHugging Face17daniazie /parallel_asian_treebank_sentstext100K<n<1M0 likes10 downloads6mo agoHugging Face18rohith2812 /STANFORD-SENTIMENT-TREEBANKThe text file contains the collection of short movie reviews. Each review is enclosed in parentheses and consists of a numerical rating followed by the review text. The numerical rating is on a scale of 0 to 4, where higher numbers indicate a more positive review. Here are some additional observations: Format: The reviews follow a consistent format with the rating at the beginning, making it easy to identify the sentiment of each review. Concise: The reviews are generally concise, focusing… See the full description on the dataset page: https://huggingface.co/datasets/rohith2812/STANFORD-SENTIMENT-TREEBANK.text1K<n<10K1 likes9 downloads2y agoHugging Face19supergoose /flan_combined_task1167_penn_treebank_coarse_pos_taggingtext10K<n<100K0 likes4 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.