EXOROBOURII/Stanza-Wikitext-2
Dataset Card for Stanza-Wikitext-2 Dataset Description Stanza-Wikitext-2 is a structurally pristine, mathematically verified NLP dataset designed for multi-task language modeling, custom tokenizer training, structural NLP research, and mechanistic interpretability work. It is a rigorously modernized and annotated derivative of the wikitext-2-raw-v1 corpus. Using the Stanford NLP Stanza neural pipeline, every token in the corpus has been explicitly mapped to its… See the full description on the dataset page: https://huggingface.co/datasets/EXOROBOURII/Stanza-Wikitext-2.
0103
1degree,count,proportion,cdf,power_law_fit20,1577896,0.63884705,0.63884705,31,270873,0.10966909,0.74851614,310761.184885870742,216626,0.08770596,0.8362221,149003.4476280341653,168122,0.06806801,0.90429011,96930.0583082874864,112052,0.0453668,0.94965691,71444.0170936861875,68814,0.02786091,0.97751782,56389.288471860386,33921,0.01373369,0.99125151,46475.9228924470697,13851,0.00560789,0.9968594,39466.95217673095108,4867,0.00197052,0.99882992,34255.90252934918119,1814,0.00073444,0.99956436,30233.6220242516831210,597,0.00024171,0.99980607,27037.4770088636251311,273,0.00011053,0.9999166,24438.2920860257361412,86,3.482e-05,0.99995142,22284.226858037541513,42,1.7e-05,0.99996842,20470.748693627571614,25,1.012e-05,0.99997854,18923.57291617161715,12,4.86e-06,0.9999834,17588.480432482661816,8,3.24e-06,0.99998664,16424.9842861086251917,3,1.21e-06,0.99998785,15402.2497205938622018,8,3.24e-06,0.99999109,14496.3854399989762119,5,2.02e-06,0.99999312,13688.5974059963742220,3,1.21e-06,0.99999433,12963.9011737065872321,2,8.1e-07,0.99999514,12310.2052695090572422,3,1.21e-06,0.99999636,11717.646707699562523,3,1.21e-06,0.99999757,11178.1013539405892625,1,4e-07,0.99999798,10232.1396912245132731,1,4e-07,0.99999838,8145.10043361018052833,1,4e-07,0.99999879,7622.5899245537882936,1,4e-07,0.99999919,6950.7117161300693039,1,4e-07,0.9999996,6385.0666073896693143,1,4e-07,1.0,5757.02109963840532 