datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ud-treebank-tokens
Dataset Card for Dataset Name
Dataset Summary
This is a subset of the Universal Dependencies Treebanks dataset (version 2.12) which only contains raw sentences and their corresponding tokenized form.
This dataset is licensed under the Universal Dependencies v2.10 License Agreement. The user is reminded that some subsets only allow non-commercial use. Each row contains the license of the dataset it originated from, and a full listing of all included subsets and… See the full description on the dataset page: https://huggingface.co/datasets/tokenizer-eval/ud-treebank-tokens.english-ipa-dep-treebank
English IPA Dependency Treebank
A large-scale dataset of 10.4 million English sentences paired with IPA (International Phonetic Alphabet) transcriptions and Universal Dependencies syntactic annotations.
Each sentence includes its full dependency parse — head indices, relation labels, and a linearized tagged-IPA representation that interleaves syntactic roles with phonetic content.
Dataset Structure
Each sample contains:
Field
Type
Description
raw_english
string… See the full description on the dataset page: https://huggingface.co/datasets/dgabri3le/english-ipa-dep-treebank.task1167_penn_treebank_coarse_pos_tagging
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1167_penn_treebank_coarse_pos_tagging
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1167_penn_treebank_coarse_pos_tagging.latin_treebanks_ud_testNOTE: This template for datasheets for ancient language data is based on the proposal by Gebru et al. 2021 https://arxiv.org/abs/1803.09010. The majority of questions is taken from there, but in a slightly rearanged order. However, as some questions are not relevant for historical data, they have been left out. Questions which are of relevance for Humanities scholars researching ancient languages have been added. The question have been answered to the best of the knowledge of the Daidalos… See the full description on the dataset page: https://huggingface.co/datasets/daidalos-project/latin_treebanks_ud_test.kathnlp-treebank
kathnlp Katharevousa Greek Treebank
A Universal-Dependencies-style reference treebank for Katharevousa Greek, the archaizing official register used in 20th-century Greek law, administration, and parliamentary discourse. The treebank covers 1,697 sentences from written parliamentary questions of the early Third Hellenic Republic (1976–1977) and is released alongside the kathnlp parsing pipeline.
Paper: A Reproducible Universal Dependencies-Style Pipeline for Katharevousa Greek… See the full description on the dataset page: https://huggingface.co/datasets/gmikros/kathnlp-treebank.stanford-sentiment-treebank-datasetThe text file contains the collection of short movie reviews.
Each review is enclosed in parentheses and consists of a numerical rating followed by the review text.
The numerical rating is on a scale of 1 to 4, where higher numbers indicate a more positive review.
Here are some additional observations:
Format: The reviews follow a consistent format with the rating at the beginning, making it easy to identify the sentiment of each review.
Concise: The reviews are generally concise, focusing… See the full description on the dataset page: https://huggingface.co/datasets/rohith2812/stanford-sentiment-treebank-dataset.parallel_asian_treebankThe ALT project aims to advance the state-of-the-art Asian natural language processing (NLP) techniques through the open collaboration for developing and using ALT.
It was first conducted by NICT and UCSY as described in Ye Kyaw Thu, Win Pa Pa, Masao Utiyama, Andrew Finch and Eiichiro Sumita (2016).
Then, it was developed under ASEAN IVO.
The process of building ALT began with sampling about 20,000 sentences from English Wikinews, and then these sentences were translated into the other languages.
ALT now has 13 languages: Bengali, English, Filipino, Hindi, Bahasa Indonesia, Japanese, Khmer, Lao, Malay, Myanmar (Burmese), Thai, Vietnamese, Chinese (Simplified Chinese).ukrainian-treebank-lmUkrainian part of the Universal Dependencies, specifically preprocessed for the language modeling task. The data can be split into documents, paragraphs or sentences. Manual selection of the data done by the authors of the dataset makes it suitable for the perplexity evaluation.
Authors of the dataset: Institute for Ukrainian, NGO, org@mova.institute
GitHub: https://github.com/UniversalDependencies/UD_Ukrainian-IUurdu-universal-dependency-treebank
Summary
The Urdu Universal Dependency Treebank was automatically converted from Urdu Dependency Treebank (UDTB) which is part of an ongoing effort of creating multi-layered treebanks for Hindi and Urdu.
Acknowledgments
UDTB is developed at IIIT-H India. The project is supported by NSF Grant (Award Number: CNS 0751202; CFDA Number: 47.070).
Any publication reporting the work done using this data should cite the following references:
Riyaz Ahmad Bhat, Rajesh Bhatt, Annahita… See the full description on the dataset page: https://huggingface.co/datasets/ReySajju742/urdu-universal-dependency-treebank.alt_burmese_treebankA 20,000-sentence Burmese (Myanmar) treebank on news articles containing complete phrase structure annotation.As the final result of the Burmese component in the Asian Language Treebank Project, this is the first large-scale,open-access treebank for the Burmese language.yoruba-constituency-treebank
Yoruba Constituency Treebank (Version 1.0)
Overview
This dataset contains a manually annotated Yoruba constituency treebank consisting of 1,000 sentences. The treebank was developed as part of an undergraduate linguistics research project focused on Yoruba syntax and computational parsing for under-resourced languages.
The annotations follow a phrase-structure (constituency) framework, including labels such as NP, VP, IP, and CP.
Dataset Contents
The repository… See the full description on the dataset page: https://huggingface.co/datasets/Akindelevictoria/yoruba-constituency-treebank.Treebank-Benchmarking
Turkish Treebank Benchmarking
This is the repo for Turkish treebank benchmarking, namely evaluating Tranformer models on POS-Dep-Morph task.
For the data, we used two treebank, IMST and BOUN. We converted conllu format to json lines for being compatible to HF dataset formats.
Here are treebank sizes at a glance:
Dataset
train lines
dev lines
test lines
BOUN
7803
979
979
IMST
3435
1100
1100
A typical instance from the dataset looks like:
{
"id": "ins_1267"… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/Treebank-Benchmarking.UD_Treebank_Te_TransliterateTELUGU dataset converted to TransliterateOTA-BOUN_UD_TreebankTreebankNos
TreebankNos
TreebankNos is a Galician-language dataset combining Universal Dependencies (UD) treebank annotations with Named Entity Recognition (NER) labels. It is built upon two established UD corpora — UD TreeGal and UD Parallel Universal Dependencies (PUD) — extended with BIO-format NER annotations covering four entity types: Person, Location, Organisation, and Miscellaneous.
The dataset is intended for multi-task NLP research in Galician, supporting POS tagging… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/TreebankNos.blackboard_treebank_prompt
Dataset Card for "blackboard_treebank_prompt"
This dataset made from blackboard treebank. The dataset want to create Thai sentence by structure.
The original dataset used own tags but we use Universal Dependencies tags, so we convert those tags into Universal Dependencies tags. See blackboard treebank tags to Universal Dependencies tags
Source code for create dataset: https://github.com/PyThaiNLP/support-aya-datasets/blob/main/pos/blackboard_treebank_prompt.ipynb
Template… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/blackboard_treebank_prompt.parallel_asian_treebank_sentsSTANFORD-SENTIMENT-TREEBANKThe text file contains the collection of short movie reviews.
Each review is enclosed in parentheses and consists of a numerical rating followed by the review text.
The numerical rating is on a scale of 0 to 4, where higher numbers indicate a more positive review.
Here are some additional observations:
Format: The reviews follow a consistent format with the rating at the beginning, making it easy to identify the sentiment of each review.
Concise: The reviews are generally concise, focusing… See the full description on the dataset page: https://huggingface.co/datasets/rohith2812/STANFORD-SENTIMENT-TREEBANK.flan_combined_task1167_penn_treebank_coarse_pos_tagging
