CoolFace
20 results

stanza

alvp /alberti-stanzastabular1K<n<10K0 likes141 downloads2y agoHugging Facealvp /autonlp-data-alberti-stanza-names AutoNLP Dataset for project: alberti-stanza-names Table of content Dataset Description Languages Dataset Structure Data Instances Data Fields Data Splits Dataset Descritpion This dataset has been automatically processed by AutoNLP for project alberti-stanza-names. Languages The BCP-47 code for the dataset's language is unk. Dataset Structure Data Instances A sample from this dataset looks as follows: [ {… See the full description on the dataset page: https://huggingface.co/datasets/alvp/autonlp-data-alberti-stanza-names.text-classification0 likes133 downloads5y agoHugging Facealvp /autonlp-data-alberti-stanzas-finetuning AutoNLP Dataset for project: alberti-stanzas-finetuning Table of content Dataset Description Languages Dataset Structure Data Instances Data Fields Data Splits Dataset Descritpion This dataset has been automatically processed by AutoNLP for project alberti-stanzas-finetuning. Languages The BCP-47 code for the dataset's language is unk. Dataset Structure Data Instances A sample from this dataset looks as… See the full description on the dataset page: https://huggingface.co/datasets/alvp/autonlp-data-alberti-stanzas-finetuning.text-classification0 likes132 downloads5y agoHugging FaceEXOROBOURII /Stanza-Wikitext-2 Dataset Card for Stanza-Wikitext-2 Dataset Description Stanza-Wikitext-2 is a structurally pristine, mathematically verified NLP dataset designed for multi-task language modeling, custom tokenizer training, structural NLP research, and mechanistic interpretability work. It is a rigorously modernized and annotated derivative of the wikitext-2-raw-v1 corpus. Using the Stanford NLP Stanza neural pipeline, every token in the corpus has been explicitly mapped to its… See the full description on the dataset page: https://huggingface.co/datasets/EXOROBOURII/Stanza-Wikitext-2.tabulartext-generation10K<n<100K0 likes104 downloads6mo agoHugging FaceSaeedRahmani /persian_poems_stanzas_vocabbasetext1M<n<10M0 likes87 downloads3y agoHugging FaceEXOROBOURII /Stanza-TinyStories Dataset Card for Stanza-TinyStories-2 Dataset Summary Stanza-TinyStories-2 is a structurally and morphologically enriched iteration of the TinyStories dataset (Eldan and Li, 2023). This dataset projects the 1D synthetic text generated by large language models into a fully resolved grammatical and topological space. Every sentence in the 2.7-million-story training split and the 21,000-story validation split has been deterministically parsed to extract Universal… See the full description on the dataset page: https://huggingface.co/datasets/EXOROBOURII/Stanza-TinyStories.tabular100K<n<1M0 likes74 downloads5mo agoHugging Face