EXOROBOURII/Stanza-Wikitext-2
Dataset Card for Stanza-Wikitext-2 Dataset Description Stanza-Wikitext-2 is a structurally pristine, mathematically verified NLP dataset designed for multi-task language modeling, custom tokenizer training, structural NLP research, and mechanistic interpretability work. It is a rigorously modernized and annotated derivative of the wikitext-2-raw-v1 corpus. Using the Stanford NLP Stanza neural pipeline, every token in the corpus has been explicitly mapped to its… See the full description on the dataset page: https://huggingface.co/datasets/EXOROBOURII/Stanza-Wikitext-2.
0103
1upos,count,mean,std,median,min,max,q90,q99,entropy2VERB,227056,3.537149,1.830323,4.0,0,15,6.0,8.0,2.8831453NOUN,431818,2.219646,1.615559,2.0,0,36,4.0,7.0,2.613414PROPN,286360,1.321672,1.519218,1.0,0,43,3.0,6.0,2.2672035X,1637,1.196701,1.991833,0.0,0,18,4.0,8.64,2.0320266INTJ,466,0.980687,1.094292,1.0,0,7,2.5,4.0,1.9092667NUM,93871,0.73634,1.00383,0.0,0,11,2.0,4.0,1.683928SYM,26215,0.577227,0.97691,0.0,0,10,1.0,5.0,1.4295689ADJ,165192,0.555656,1.339018,0.0,0,17,2.0,6.0,1.31507610ADV,70163,0.251814,0.672836,0.0,0,9,1.0,3.0,0.88657211PRON,88350,0.140668,0.540159,0.0,0,10,0.0,2.0,0.57628612SCONJ,33078,0.029748,0.224341,0.0,0,9,0.0,1.0,0.18356913CCONJ,72684,0.029236,0.219907,0.0,0,10,0.0,1.0,0.17951514ADP,292758,0.029171,0.241405,0.0,0,11,0.0,1.0,0.17301215AUX,87909,0.02167,0.246182,0.0,0,8,0.0,1.0,0.10715416DET,230554,0.017432,0.194266,0.0,0,10,0.0,1.0,0.10510717PUNCT,313889,0.006308,0.147937,0.0,0,11,0.0,0.0,0.03186418PART,47912,0.005447,0.127396,0.0,0,6,0.0,0.0,0.03163819 