stanza
Datasets
All datasets matching “stanza”alberti-stanzasautonlp-data-alberti-stanza-names
AutoNLP Dataset for project: alberti-stanza-names
Table of content
Dataset Description
Languages
Dataset Structure
Data Instances
Data Fields
Data Splits
Dataset Descritpion
This dataset has been automatically processed by AutoNLP for project alberti-stanza-names.
Languages
The BCP-47 code for the dataset's language is unk.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{… See the full description on the dataset page: https://huggingface.co/datasets/alvp/autonlp-data-alberti-stanza-names.autonlp-data-alberti-stanzas-finetuning
AutoNLP Dataset for project: alberti-stanzas-finetuning
Table of content
Dataset Description
Languages
Dataset Structure
Data Instances
Data Fields
Data Splits
Dataset Descritpion
This dataset has been automatically processed by AutoNLP for project alberti-stanzas-finetuning.
Languages
The BCP-47 code for the dataset's language is unk.
Dataset Structure
Data Instances
A sample from this dataset looks as… See the full description on the dataset page: https://huggingface.co/datasets/alvp/autonlp-data-alberti-stanzas-finetuning.Stanza-Wikitext-2
Dataset Card for Stanza-Wikitext-2
Dataset Description
Stanza-Wikitext-2 is a structurally pristine, mathematically verified NLP dataset designed for multi-task language modeling, custom tokenizer training, structural NLP research, and mechanistic interpretability work.
It is a rigorously modernized and annotated derivative of the wikitext-2-raw-v1 corpus. Using the Stanford NLP Stanza neural pipeline, every token in the corpus has been explicitly mapped to its… See the full description on the dataset page: https://huggingface.co/datasets/EXOROBOURII/Stanza-Wikitext-2.persian_poems_stanzas_vocabbaseStanza-TinyStories
Dataset Card for Stanza-TinyStories-2
Dataset Summary
Stanza-TinyStories-2 is a structurally and morphologically enriched iteration of the TinyStories dataset (Eldan and Li, 2023).
This dataset projects the 1D synthetic text generated by large language models into a fully resolved grammatical and topological space. Every sentence in the 2.7-million-story training split and the 21,000-story validation split has been deterministically parsed to extract Universal… See the full description on the dataset page: https://huggingface.co/datasets/EXOROBOURII/Stanza-TinyStories.
