datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ms-thesisthesis-corpus-v18
Part of the SZL Holdings governed estate — claims are designed to carry checkable receipts. Verification proves integrity & origin, never accuracy or performance.
SZLHOLDINGS/thesis-corpus-v18
The v18 Ouroboros Invariant thesis — LaTeX chapters, the 179 formal blocks
(theorem / lemma / definition / axiom environments) as a flat CSV, and the per-version
delta ledger that tracks how every formal block evolved v1 → v18.
Contents
File… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/thesis-corpus-v18.indo-mmarcolayra-thesis-corpus-rawen-id-parallel-sentences-embedding-vbertschaeffer_thesis_correctedThe SCHAEFFER dataset (Spectro-morphogical Corpus of Human-annotated Audio with Electroacoustic Features for Experimental Research), is a compilation of 788 raw audio data accompanied by human annotations and morphological acoustic features.
The audio files adhere to the concept of Sound Objects introduced by Pierre Scaheffer, a framework for the analysis and creation of sound that focuses on its typological and morphological characteristics.
Inside the dataset, the annotation are provided in… See the full description on the dataset page: https://huggingface.co/datasets/dbschaeffer/schaeffer_thesis_corrected.electronic-radiology-phd-thesis-trRMSc-Thesis-professionset
Dataset Card for professions-v2
Dataset Summary
🏗️ WORK IN PROGRESS
⚠️ DISCLAIMER: The images in this dataset were generated by text-to-image systems and may depict offensive stereotypes or contain explicit content.
The Professions dataset is a collection of computer-generated images generated using Text-to-Image (TTI) systems.
In order to generate a diverse set of prompts to evaluate the system outputs’ variation across dimensions of interest, we use the pattern Photo… See the full description on the dataset page: https://huggingface.co/datasets/malena-fj/MSc-Thesis-professionset.lao-asr-thesis-dataseten-id-parallel-sentences-embedding
Dataset Card for "en-id-parallel-sentences-embedding"
More Information needed
distilled-ccmatrix-de-en
Dataset Card for "distilled-ccmatrix-de-en"
More Information needed
xg-thesis
expected-goals-thesis
A repository for analysis on Expected Goals using StatsBomb and Wyscout data.
StatsBomb data
This repository assumes that the StatsBomb open-data has already been cloned to a local directory.
Versioning
The original thesis was run from a particular version of the data and mplsoccer (my football plotting library).
The original code is here:… See the full description on the dataset page: https://huggingface.co/datasets/fadhilra101/xg-thesis.thesisbachelor-thesis-datasetsFinQA-DPO-Mixedmsmarco-corpus-en-id-parallel-sentences
Dataset Card for "msmarco-corpus-en-id-parallel-sentences"
More Information needed
TATQA_SFTFinQA-DPO-Statement-Preferenceindo-snli
Dataset Card for "indo-snli"
More Information needed
jfk_senior_thesis_data
Dataset Card for "jfk_senior_thesis_data"
More Information needed
thesis-dataset-180-180-final-128distilled-ccmatrix-es-en
Dataset Card for "distilled-ccmatrix-es-en"
More Information needed
korea_summary_ThesisThesis-Abstract-Classification-11K
Dataset Card for Thesis-Abstract-Classification-11K
Dataset Description
Thesis-Abstract-Classification-11K dataset is obtained by processing a subset of Turkish Academic Theses dataset.
Dataset Structure
The original dataset was large and examples had several subject fields, representing the field of the thesis.
In order to construct a single-class classification problem with a reasonable data size, the following steps are carried out:
For each example, only… See the full description on the dataset page: https://huggingface.co/datasets/boun-tabilab/Thesis-Abstract-Classification-11K.distilled-ccmatrix-en-fr
Dataset Card for "distilled-ccmatrix-en-fr"
More Information needed
Thesisminingpile-thesiskorea_summary_Thesisthesis_datasetsFinQA-DPO-Answer-Preference
