conllu
Datasets
All datasets matching “conllu”oga-conllu-stoicheia
OGA parsed with Stoicheia
Period-delimited sentences from the source=oga records of the pristine split of
Ericu950/AncientGreek. A literal . terminates a sentence and is retained. All source columns
are inherited by each sentence. author contains the canonical author label corresponding to the
TLG/CTS code in id; text contains the sentence and conllu contains the Universal
Dependencies-style output from Ericu950/Stoicheia-tagger-parser.
The dataset contains 1,135,428 sentences… See the full description on the dataset page: https://huggingface.co/datasets/Urdatorn/oga-conllu-stoicheia.roots_ca_enriched_conllu_ancora_for_ml_trainingROOTS Subset: roots_ca_enriched_conllu_ancora_for_ml_training
Enriched CONLLU Ancora for ML training
Dataset uid: enriched_conllu_ancora_for_ml_training
Description
This is an enriched version for Machine Learning purposes of the CONLLU adaptation of AnCora corpus .
This version of the corpus was developed by BSC TeMU as part of the AINA project, and has been used to do multi-task learning for the Catalan language Spacy 3.0 models.
Homepage… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_ca_enriched_conllu_ancora_for_ml_training.conllu2xml-rrt-v1conllu2xml-rrt-dev-v0CoNLLU_WikiNEuRalwik_eng_wikipedia_en_20m_words-1-conllu
wik_eng_wikipedia_en_20m_words-1
This dataset contains a CoNLL-U parse file:
wik_eng_wikipedia_en_20m_words-1.conllu
Token count (CoNLL-U integer IDs only, excluding comments, MWT ranges, and empty nodes): 12,746,154.
Parsed by puddin.
