CoolFace
Datasetpublic

EXOROBOURII/Stanza-Wikitext-2

Dataset Card for Stanza-Wikitext-2 Dataset Description Stanza-Wikitext-2 is a structurally pristine, mathematically verified NLP dataset designed for multi-task language modeling, custom tokenizer training, structural NLP research, and mechanistic interpretability work. It is a rigorously modernized and annotated derivative of the wikitext-2-raw-v1 corpus. Using the Stanford NLP Stanza neural pipeline, every token in the corpus has been explicitly mapped to its… See the full description on the dataset page: https://huggingface.co/datasets/EXOROBOURII/Stanza-Wikitext-2.

sourceHugging Facecc-by-sa-4.0updated 6mo agoView on Hugging Face
0likes103downloads
upos_depth_stats.csv19 linesDownload Raw Back to root
1upos,count,mean,std,median,min,max,q90,q99,entropy2CCONJ,72684,3.651918,1.664702,3.0,0,21,6.0,9.0,2.5599723ADP,292758,3.510001,1.517752,3.0,0,22,5.0,8.0,2.4544124PART,47912,3.490128,1.557959,3.0,0,15,5.0,8.0,2.5515685DET,230554,3.201961,1.464867,3.0,0,21,5.0,8.0,2.3619366ADJ,165192,3.107178,1.605191,3.0,0,19,5.0,8.0,2.6032727X,1637,3.100183,1.637274,3.0,0,9,5.0,8.0,2.6690728NUM,93871,2.964888,1.543966,3.0,0,24,5.0,8.0,2.5433569PRON,88350,2.902173,1.655757,3.0,0,20,5.0,8.0,2.60889310SCONJ,33078,2.840317,1.223844,2.0,0,13,4.0,7.0,1.95986911PROPN,286360,2.752745,1.571835,2.0,0,18,5.0,8.0,2.54423412PUNCT,313889,2.682996,1.707904,2.0,0,22,5.0,8.0,2.54611513INTJ,466,2.609442,1.763988,2.0,0,10,5.0,8.35,2.64667514ADV,70163,2.596782,1.620345,2.0,0,21,5.0,8.0,2.4887215NOUN,431818,2.452719,1.541889,2.0,0,22,4.0,7.0,2.46770916AUX,87909,2.030475,1.392487,2.0,0,20,4.0,7.0,2.07370617SYM,26215,1.979287,2.12325,1.0,0,25,5.0,8.0,2.62004318VERB,227056,1.467431,1.566902,1.0,0,23,4.0,6.0,2.36866319