datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
thesis-corpus-v18
Part of the SZL Holdings governed estate — claims are designed to carry checkable receipts. Verification proves integrity & origin, never accuracy or performance.
SZLHOLDINGS/thesis-corpus-v18
The v18 Ouroboros Invariant thesis — LaTeX chapters, the 179 formal blocks
(theorem / lemma / definition / axiom environments) as a flat CSV, and the per-version
delta ledger that tracks how every formal block evolved v1 → v18.
Contents
File… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/thesis-corpus-v18.thesis-chile
Thesis Chile Dataset
Dataset Summary
Thesis Chile is the dataset partially used to create the DiscoEval in Spanish benchmark.
This dataset was created by scraping titles and abstracts of Chilean thesis from public repositories of the Pontificia Universidad Catolica de Chile (repositorio.uc.cl), Universidad de Chile (repositorio.uchile.cl) and Universidad Técnica Federico Santa María (biblioteca.usm.cl).
Supported Tasks
We see the potential utility of this… See the full description on the dataset page: https://huggingface.co/datasets/vgaraujov/thesis-chile.English-Thesis-Dataset
Thesis Dataset
Dataset Description
The Thesis Dataset is a large-scale academic text corpus designed for building next-generation Natural Language Processing (NLP) systems, Large Language Models (LLMs), Retrieval-Augmented Generation (RAG) systems, research assistants, document understanding models, and knowledge extraction pipelines.
The complete multilingual collection contains over 641,700 thesis documents comprising 7.38+ billion words. This English release… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/English-Thesis-Dataset.Thesis_Development_of_a_Complex_of_Neural_Networks_for_Linked_Generation_of_Large_TextsHere presented a partially synthesized dataset, developed utilizing the GPT-4 model, for the purpose of NLG, particulary for the task of hierarchical generation of longer texts from short summaries. The creation of this dataset was undertaken as a component of my thesis paper. It incorporates excerpts from prominent British and American novels, from which plots, summaries, and metadata have been derived using GPT-4 API to facilitate extensive future research.
The metadata included in the… See the full description on the dataset page: https://huggingface.co/datasets/Fleur-roar/Thesis_Development_of_a_Complex_of_Neural_Networks_for_Linked_Generation_of_Large_Texts.thesis-texts
sghosts/thesis-texts
This dataset contains columns: thesis_id, text.
thesis_id: Derived from the filename (if digits exist, the longest digit sequence; otherwise the filename stem).
text: Contents of the .txt file (UTF-8 preferred; problematic files may be skipped).
This card was created automatically. Feel free to edit.
