datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ms-thesisindo-mmarcoen-id-parallel-sentences-embedding-vbertschaeffer_thesis_correctedThe SCHAEFFER dataset (Spectro-morphogical Corpus of Human-annotated Audio with Electroacoustic Features for Experimental Research), is a compilation of 788 raw audio data accompanied by human annotations and morphological acoustic features.
The audio files adhere to the concept of Sound Objects introduced by Pierre Scaheffer, a framework for the analysis and creation of sound that focuses on its typological and morphological characteristics.
Inside the dataset, the annotation are provided in… See the full description on the dataset page: https://huggingface.co/datasets/dbschaeffer/schaeffer_thesis_corrected.lao-asr-thesis-datasetdistilled-ccmatrix-de-en
Dataset Card for "distilled-ccmatrix-de-en"
More Information needed
xg-thesis
expected-goals-thesis
A repository for analysis on Expected Goals using StatsBomb and Wyscout data.
StatsBomb data
This repository assumes that the StatsBomb open-data has already been cloned to a local directory.
Versioning
The original thesis was run from a particular version of the data and mplsoccer (my football plotting library).
The original code is here:… See the full description on the dataset page: https://huggingface.co/datasets/fadhilra101/xg-thesis.MSc-Thesis-professionset
Dataset Card for professions-v2
Dataset Summary
🏗️ WORK IN PROGRESS
⚠️ DISCLAIMER: The images in this dataset were generated by text-to-image systems and may depict offensive stereotypes or contain explicit content.
The Professions dataset is a collection of computer-generated images generated using Text-to-Image (TTI) systems.
In order to generate a diverse set of prompts to evaluate the system outputs’ variation across dimensions of interest, we use the pattern Photo… See the full description on the dataset page: https://huggingface.co/datasets/malena-fj/MSc-Thesis-professionset.en-id-parallel-sentences-embedding
Dataset Card for "en-id-parallel-sentences-embedding"
More Information needed
msmarco-corpus-en-id-parallel-sentences
Dataset Card for "msmarco-corpus-en-id-parallel-sentences"
More Information needed
TATQA_SFTindo-snli
Dataset Card for "indo-snli"
More Information needed
jfk_senior_thesis_data
Dataset Card for "jfk_senior_thesis_data"
More Information needed
thesis_useThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "lekiwi",
"total_episodes": 12,
"total_frames": 24728,
"total_tasks": 1,
"total_videos": 36,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:12"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Deason11/thesis_use.thesis-dataset-180-180-final-128Thesis-Abstract-Classification-11K
Dataset Card for Thesis-Abstract-Classification-11K
Dataset Description
Thesis-Abstract-Classification-11K dataset is obtained by processing a subset of Turkish Academic Theses dataset.
Dataset Structure
The original dataset was large and examples had several subject fields, representing the field of the thesis.
In order to construct a single-class classification problem with a reasonable data size, the following steps are carried out:
For each example, only… See the full description on the dataset page: https://huggingface.co/datasets/boun-tabilab/Thesis-Abstract-Classification-11K.distilled-ccmatrix-es-en
Dataset Card for "distilled-ccmatrix-es-en"
More Information needed
distilled-ccmatrix-en-fr
Dataset Card for "distilled-ccmatrix-en-fr"
More Information needed
synthethic-baseminingpile-thesiskorea_summary_Thesisderm_QAthesis_datasetskorea_summary_Thesisdistilled-ccmatrix-fr-en
Dataset Card for "distilled-ccmatrix-fr-en"
More Information needed
Thesisthesis-dataset-legalthesis_vlm_flat_exp_augmenteddistilled-ccmatrix-en-de
Dataset Card for "distilled-ccmatrix-en-de"
More Information needed
TimeQA_SFT
