CoolFace
Datasetpublic

EXOROBOURII/Stanza-Wikitext-2

Dataset Card for Stanza-Wikitext-2 Dataset Description Stanza-Wikitext-2 is a structurally pristine, mathematically verified NLP dataset designed for multi-task language modeling, custom tokenizer training, structural NLP research, and mechanistic interpretability work. It is a rigorously modernized and annotated derivative of the wikitext-2-raw-v1 corpus. Using the Stanford NLP Stanza neural pipeline, every token in the corpus has been explicitly mapped to its… See the full description on the dataset page: https://huggingface.co/datasets/EXOROBOURII/Stanza-Wikitext-2.

sourceHugging Facecc-by-sa-4.0updated 6mo agoView on Hugging Face
0likes103downloads

EXOROBOURII/Stanza-Wikitext-2 · main · files are served by the source, never re-hosted here