CoolFace
Datasetpublic

EXOROBOURII/Stanza-Wikitext-2

Dataset Card for Stanza-Wikitext-2 Dataset Description Stanza-Wikitext-2 is a structurally pristine, mathematically verified NLP dataset designed for multi-task language modeling, custom tokenizer training, structural NLP research, and mechanistic interpretability work. It is a rigorously modernized and annotated derivative of the wikitext-2-raw-v1 corpus. Using the Stanford NLP Stanza neural pipeline, every token in the corpus has been explicitly mapped to its… See the full description on the dataset page: https://huggingface.co/datasets/EXOROBOURII/Stanza-Wikitext-2.

sourceHugging Facecc-by-sa-4.0updated 6mo agoView on Hugging Face
0likes103downloads
settings

This repository belongs to EXOROBOURII on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameStanza-Wikitext-2
visibilitypublic
licencecc-by-sa-4.0
gatedno
ownerEXOROBOURII
Account settings
EXOROBOURII/Stanza-Wikitext-2 · CoolFace