EXOROBOURII/Stanza-Wikitext-2
Dataset Card for Stanza-Wikitext-2 Dataset Description Stanza-Wikitext-2 is a structurally pristine, mathematically verified NLP dataset designed for multi-task language modeling, custom tokenizer training, structural NLP research, and mechanistic interpretability work. It is a rigorously modernized and annotated derivative of the wikitext-2-raw-v1 corpus. Using the Stanford NLP Stanza neural pipeline, every token in the corpus has been explicitly mapped to its… See the full description on the dataset page: https://huggingface.co/datasets/EXOROBOURII/Stanza-Wikitext-2.
0103
1depth,count,proportion,cdf20,107410,0.04348738,0.0434873831,462211,0.18713663,0.2306240142,681789,0.27603777,0.5066617853,540836,0.21896974,0.7256315264,339330,0.13738546,0.8630169875,180698,0.07315969,0.9361766786,86632,0.03507493,0.9712516197,39851,0.01613458,0.98738619108,17661,0.00715046,0.99453665119,7583,0.00307015,0.99760681210,3197,0.00129438,0.998901181311,1437,0.0005818,0.999482981412,656,0.0002656,0.999748571513,311,0.00012592,0.999874491614,121,4.899e-05,0.999923481715,66,2.672e-05,0.99995021816,30,1.215e-05,0.999962351917,22,8.91e-06,0.999971252018,17,6.88e-06,0.999978142119,16,6.48e-06,0.999984612220,12,4.86e-06,0.999989472321,10,4.05e-06,0.999993522422,7,2.83e-06,0.999996362523,5,2.02e-06,0.999998382624,3,1.21e-06,0.99999962725,1,4e-07,1.028 