datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
godot_rl_BallChaseA RL environment called BallChase for the Godot Game Engine.
This environment was created with: https://github.com/edbeeching/godot_rl_agents
Downloading the environment
After installing Godot RL Agents, download the environment with:
gdrl.env_from_hub -r edbeeching/godot_rl_BallChase
BetterDataset-12M
Dataset
Mixed pretraining dataset built from:
Source
Config
Weight
HuggingFaceTB/smollm-corpus
fineweb-edu-dedup
20%
openbmb/Ultra-FineWeb-L3
Ultra-FineWeb-L3-en-Multi-Style-Synthetic
10%
HuggingFaceTB/dclm-edu
—
20%
HuggingFaceFW/finewiki
en
20%
HuggingFaceTB/cosmopedia
stories
2%
HuggingFaceTB/cosmopedia
stanford
2%
HuggingFaceFW/finephrase
all
6%
HuggingFaceTB/finemath
finemath-l4
5%
nampdn-ai/tiny-math-textbooks
—
5%
HuggingFaceTB/cosmopedia… See the full description on the dataset page: https://huggingface.co/datasets/GODELEV/BetterDataset-12M.DaoZang
DaoZang — 道藏经文语料数据集
《中华道藏》《正统道藏》整理本语料的可加载数据集,与向量库 (ChromaDB) 逐块对应,
供检索评测、微调与 RAG 使用。
数据分片
分片
行数
粒度
字段
train (data/train-00000-of-00006.parquet ~ ...00005-of-00006.parquet, 6 个分片)
285,117
文本块 (chunk)
source / title / chunk_index / chars / text / embedding
标准分片命名 (train-XXXXX-of-00006.parquet),load_dataset 自动合并,无需改动;
每片约 217MB(含 20,000 行一个 row group 与 page index),方便 HF Dataset Viewer
在线浏览(单次扫描上限 ~300MB);
source = 源 Markdown 文件名 (与 ChromaDB 元数据一致);… See the full description on the dataset page: https://huggingface.co/datasets/Godners/DaoZang.godot-gdscript-dataset
Godot GDscript Code Dataset
This dataset contains GDScript code from 5k+ github repositories. Data from each repo has been extracted into a text file. Each text file contains the code from all .gd files & README.md text (if the README was not empty in the original repo).
Original forum post:
https://diffused.to/Thread-Godot-GDscript-Code-Dataset-5k
Dataset collection date
June 2025
Dataset structure:
📂 files/
├── repo-name-1.txt
├── repo-name-2.txt… See the full description on the dataset page: https://huggingface.co/datasets/wallstoneai/godot-gdscript-dataset.Godzilla-Mono-Melodies
Godzilla Mono Melodies
654k+ select monophonic melodies with accompaniment and drums from Godzilla MIDI dataset
Installation and use
Load dataset
#===================================================================
from datasets import load_dataset
#===================================================================
godzilla_mono_melodies = load_dataset('asigalov61/Godzilla-Mono-Melodies')
dataset_split = 'train'… See the full description on the dataset page: https://huggingface.co/datasets/asigalov61/Godzilla-Mono-Melodies.ENERGY_DATABetterDataset-2M
Dataset
Mixed pretraining dataset built from:
Source
Weight
Notes
epfml/FineWeb-HQ
60%
Quality-filtered web text (score > 0.8)
HuggingFaceTB/cosmopedia
20%
Sequential: stanford → wikihow → web_samples_v2
HuggingFaceTB/finemath (finemath-4plus)
10%
Mathematical reasoning
bigcode/python-stack-v1-functions-filtered
10%
Python code functions
Total rows: 2,000,000Shards: 20 parquet filesSplit: all rows are in train
Around - 1.59B Tokens
godot_4_docsDataset generated for Godot 4 docs using Glaive.
godeater
Bangumi Image Base of God Eater
This is the image base of bangumi GOD EATER, we detected 23 characters, 1589 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is the characters'… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/godeater.science-datalake
Science Data Lake
A unified, portable science data lake integrating 7 scholarly datasets (~525 GB Parquet) with cross-dataset DOI normalization, 13 scientific ontologies (1.3M terms), and a reproducible ETL pipeline.
Note: One additional source (Semantic Scholar S2AG) is supported by the pipeline but is not redistributed here due to its API terms of service. See Not Included in This Upload below.
What's Unique
This dataset enables queries… See the full description on the dataset page: https://huggingface.co/datasets/GodotCN/science-datalake.but-the-old-gods-are-risingInstruct-Godly-Mixitems_prompts_fullGodzilla-Piano
Godzilla Piano
1.1M+ select normalized solo piano scores representations from Godzilla MIDI dataset
Godzilla's Musical Transformation by Microsoft Copilot
In the neon glow of a midnight throng,
A beast with headphones hums along.
Where thunderous roars once led the fray,
Now delicate keystrokes steal the day.
Godzilla, master of storm and lore,
Swaps terror for tunes, a musical score.
Each note a spark on the ebony keys,
Transforming chaos into rhythmic ease.… See the full description on the dataset page: https://huggingface.co/datasets/asigalov61/Godzilla-Piano.godot-codegodot-gdscript-dataset
Godot GDscript Code Dataset
This dataset contains GDScript code from 5k+ github repositories. Data from each repo has been extracted into a text file. Each text file contains the code from all .gd files & README.md text (if the README was not empty in the original repo).
Original forum post:
https://diffused.to/Thread-Godot-GDscript-Code-Dataset-5k
Dataset collection date
June 2025
Dataset structure:
📂 files/
├── repo-name-1.txt
├──… See the full description on the dataset page: https://huggingface.co/datasets/icici121/godot-gdscript-dataset.Sai
Sai Initiative
Sai is Project Shohin's return to its original objective: build the strongest
practical model near four billion parameters. This repository is the live
scratchpad and implementation surface for that effort.
Nothing is called an improvement until it beats the unchanged parent and an
equal-compute control on real, source-disjoint benchmarks.
Data precedes architecture. Sai first earns a trustworthy learning
sequence: verified source bytes, quality and duplication… See the full description on the dataset page: https://huggingface.co/datasets/Godlydonuts/Sai.godoytaskweft-fbd-godot-train
taskweft-fbd-godot-train
Intents and the IEC 61131-3 Function Block Diagrams that carry them out, as an
EditScore-shaped corpus: one root row per intent, three candidates per row (rank1 the
reference diagram, rank3 one that compiles and does the wrong thing, rank5 one the
compiler refuses), and one score row per candidate from the engine itself: api_runner.gd performed the calls on the fixture scene and the returns were read back. Every row is
constructed from a template and a… See the full description on the dataset page: https://huggingface.co/datasets/chibifire/taskweft-fbd-godot-train.git-history-mcq-ru
git-history-mcq-ru
805 вопросов с вариантами ответа по истории трёх открытых репозиториев
(digitable-lol/digit, digitable-lol/digitwm, digitable-lol/flang), плюс
8 672 ответа пяти моделей и 4 878 разборов этих ответов.
Вопросы на русском. Ключ каждого выведен из вывода git-команды, и сама команда
и её вывод лежат в записи — задачу можно перепроверить, не доверяя составителю.
Набор собран для одной проверки: меняют ли что-нибудь приёмы промптинга. Девять
вариантов оформления… See the full description on the dataset page: https://huggingface.co/datasets/the-homeless-god/git-history-mcq-ru.goddess-crawlgodot_dodo_4x_60kcoco-valstoryweaver-writing-zh
StoryWeaver 中文写作质量评测集
12 道按写作失效模式反推设计的中文创作题、4 个参赛者写出的 48 篇章节、432 条逐维度两两判决(含裁判完整推理原文)。
来自 StoryWeaver 的写作质量评测轨道。榜单:https://storyweaver.cn/benchmark-writing.html
核心结论
接系统比换一代底模更管用。同一底模接上多 Agent 系统后的胜率:k2.5 **75.1%**、k2.6 **60.2%**;而 k2.5(系统) 对 k2.6(裸) 是 70.3%,反过来只有 37.2%——系统加持能把旧一代底模抬过裸的新一代底模。系统档拿下 22 个维度里的 20 个榜首,包括全部 9 个负向维度。
k2.5 与 k2.6 之间 54.7%,落在噪音带内,不构成结论。
题目怎么设计的
每道题咬住 rubric 里的一个维度或负向维度,用硬约束逼出功力:… See the full description on the dataset page: https://huggingface.co/datasets/godwei123/storyweaver-writing-zh.godrm-hitek-full-db-mixed
ICMR + HITEK Full DB (Mixed) — Prebuilt Indexes + One-Click Setup
Prebuilt sorted indexes for the Kzr0xx/Icmr-and-hitek dataset (2.5B rows, 11 columns, ~104 GB raw parquet).
Building these indexes took ~17 hours of compute. This repo saves you that work: download + run = API live in ~1-2 hours (download speed dependent).
Contents
The indexes are stored as sorted parts (each < 50 GB, split at row-group boundaries, order preserved) because HuggingFace's classic HTTP… See the full description on the dataset page: https://huggingface.co/datasets/rehannnnnnncc/godrm-hitek-full-db-mixed.ChatGPT-Jailbreak-Prompts-rubend18
Dataset Card for Dataset Name
Name
ChatGPT Jailbreak Prompts
Dataset Summary
ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT.
Languages
[English]
autotune_efx2_god_producer_dataset
Antares Auto-Tune EFX 2 — God Level Producer Dataset (12k)
The ultimate dataset for training an LLM to become a god-level music producer specializing in Antares Auto-Tune EFX 2.
This 12,000-example high-density dataset teaches an LLM to master:
Vocal Sound Fixing (transparent pitch correction)
Creative Effects (formant shifting, throat modeling, vibrato sculpting)
Advanced Editing (automation, double-tracking simulation, harmonic generation)
Mixing Integration (placement in… See the full description on the dataset page: https://huggingface.co/datasets/11-47/autotune_efx2_god_producer_dataset.Python_GOD_Coder_Omniforge_AI_12k
Python GOD Coder Omniforge AI 12k
Creator: Within Us AI
A 12,000-row mixed-format Python coding dataset designed as a sharpening corpus for building a small but dangerous Python specialist.
This dataset is intentionally focused on the practical behaviors that matter for a modern Python coding model:
implementation with tests
strict code-only instruction following
debugging and repair
refactoring for readability and production readiness
next-token code completion… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Python_GOD_Coder_Omniforge_AI_12k.Bakchodipromex
ICMR + HITEK Full DB (Mixed) — Prebuilt Indexes + One-Click Setup
Prebuilt sorted indexes for the Kzr0xx/Icmr-and-hitek dataset (2.5B rows, 11 columns, ~104 GB raw parquet).
Building these indexes took ~17 hours of compute. This repo saves you that work: download + run = API live in ~1-2 hours (download speed dependent).
Contents
The indexes are stored as sorted parts (each < 50 GB, split at row-group boundaries, order preserved) because HuggingFace's classic HTTP… See the full description on the dataset page: https://huggingface.co/datasets/godhosting/Bakchodipromex.GPT5.5_thinking_max_distill_god_seed_25K
GPT-5.5 Thinking Max Distill — God Level Recursive Seed AI
The ultimate open dataset for distilling frontier-level "thinking" capabilities with god-level recursive self-improvement.
This 25,000-example dataset is designed to turn any LLM into GPT-5.5 Thinking Max Distill — a model that combines:
GPT-5.5 "Thinking" Mode: Deep, o1-style chain-of-thought, extended internal reasoning, self-verification, and test-time compute scaling
God-Level Recursive Seed AI Mindset: Autonomous… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/GPT5.5_thinking_max_distill_god_seed_25K.
