datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
csd_filessdxl_images_easy_prompts-artists-seed1godot-gdscript-dataset
Godot GDscript Code Dataset
This dataset contains GDScript code from 5k+ github repositories. Data from each repo has been extracted into a text file. Each text file contains the code from all .gd files & README.md text (if the README was not empty in the original repo).
Original forum post:
https://diffused.to/Thread-Godot-GDscript-Code-Dataset-5k
Dataset collection date
June 2025
Dataset structure:
📂 files/
├── repo-name-1.txt
├── repo-name-2.txt… See the full description on the dataset page: https://huggingface.co/datasets/wallstoneai/godot-gdscript-dataset.sdxl_images_mj_prompts-artists-seed3sdxl_images_mj_prompts-artists-seed5godot-gdscript-dataset
Godot GDscript Code Dataset
This dataset contains GDScript code from 5k+ github repositories. Data from each repo has been extracted into a text file. Each text file contains the code from all .gd files & README.md text (if the README was not empty in the original repo).
Original forum post:
https://diffused.to/Thread-Godot-GDscript-Code-Dataset-5k
Dataset collection date
June 2025
Dataset structure:
📂 files/
├── repo-name-1.txt
├──… See the full description on the dataset page: https://huggingface.co/datasets/icici121/godot-gdscript-dataset.sdxl_images_sb_prompts-single_artist-seed0sdxl_images_mj_prompts-artists-seed6sdxl_images_mj_prompts-artists-seed7mmu_jwst_gds
mmu_jwst_gds HATS Catalog Collection
This is the collection of HATS catalogs representing mmu_jwst_gds.
This dataset is part of the Multimodal Universe,
a large-scale collection of multimodal astronomical data. For full details, see the paper:
The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TBs of Astronomical Scientific Data.
Access the catalog
We recommend the use of the LSDB Python framework to access HATS catalogs.
LSDB can be installed… See the full description on the dataset page: https://huggingface.co/datasets/UniverseTBD/mmu_jwst_gds.sdxl_images_mj_prompts-artists-seed4embeat_45m_spotify_tracks
Embeat 45M Spotify Tracks
A large-scale music metadata dataset containing 45 million Spotify tracks, combined with the Spotify metadata from Anna's Archive and artist genres from Every Noise at Once.
GitHub project: https://github.com/gdstudio-org/Embeat
Dataset Preview
>>> from datasets import load_from_disk
>>> ds = load_from_disk("GD-Studio/embeat_45m_spotify_tracks")
>>> ds
Dataset({
features: ['track_id', 'track_name', 'isrc', 'popularity', 'explicit'… See the full description on the dataset page: https://huggingface.co/datasets/GD-Studio/embeat_45m_spotify_tracks.mmu_jwst_gds_grizli_v7.0_all_96
mmu_jwst_gds_grizli_v7.0_all_96 HATS Catalog Collection
This is the collection of HATS catalogs representing mmu_jwst_gds_grizli_v7.0_all_96.
This dataset is part of the Multimodal Universe,
a large-scale collection of multimodal astronomical data. For full details, see the paper:
The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TBs of Astronomical Scientific Data.
Access the catalog
We recommend the use of the LSDB Python framework to access… See the full description on the dataset page: https://huggingface.co/datasets/UniverseTBD/mmu_jwst_gds_grizli_v7.0_all_96.sdxl_images_sb_prompts-multi_artist-seed1sdxl_images_mj_prompts-artists-seed0sdxl_images_mj_prompts-artists-seed2sdxl_images_mj_prompts-artists-seed1Gd-Script
Dataset Card for Gd-Script
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/bunnycore/Gd-Script/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/bunnycore/Gd-Script.gdsiisdxl_images_sb_prompts-multi_artistcsd_resultssdxl_images_mj_prompts-artists-seed9ds_benchmark_gds_airqualityds_benchmark_gds_retailds_benchmark_gds_poisfd-scripts-dataset
Superfighters Deluxe (SFD) Scripts Dataset
Este dataset contiene código fuente en C# extraído de scripts y mapas comunitarios del juego Superfighters Deluxe.
Propósito
Ha sido recopilado para servir como base de datos en el entrenamiento (Instruction Tuning) de un modelo Encoder-Decoder.
Estructura
Formato: Texto plano (C#).
Limpieza: Se han eliminado las cabeceras binarias y metadatos nativos del motor .sfdm dejando exclusivamente el código… See the full description on the dataset page: https://huggingface.co/datasets/GDSTARSFAN49/sfd-scripts-dataset.gdsuite-delphi-result
GDsuite results — Delphi model collection
GDsuite evaluation results for
the marin-community/delphi
model collection.
Contents
summary.jsonl — tidy per-task metrics (14256 rows). One row per
(model, family, task, metric):
metric = hard_acc (5 logprob families) — fraction of items where
P(correct) > P(incorrect); higher ⇒ resists the misleading pattern.
metric = correct_log_prob (5 logprob families) — mean
teacher-forced log probability of the correct answer.
metric =… See the full description on the dataset page: https://huggingface.co/datasets/WillHeld/gdsuite-delphi-result.gdsc-model-datasetsdxl_images_mj_prompts-artists-seed8sdxl_images_easy_prompts-multi_artist-seed0
