datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ENERGY_DATAgodot_4_docsDataset generated for Godot 4 docs using Glaive.
git-history-mcq-ru
git-history-mcq-ru
805 вопросов с вариантами ответа по истории трёх открытых репозиториев
(digitable-lol/digit, digitable-lol/digitwm, digitable-lol/flang), плюс
8 672 ответа пяти моделей и 4 878 разборов этих ответов.
Вопросы на русском. Ключ каждого выведен из вывода git-команды, и сама команда
и её вывод лежат в записи — задачу можно перепроверить, не доверяя составителю.
Набор собран для одной проверки: меняют ли что-нибудь приёмы промптинга. Девять
вариантов оформления… See the full description on the dataset page: https://huggingface.co/datasets/the-homeless-god/git-history-mcq-ru.goddess-crawlcoco-valstoryweaver-writing-zh
StoryWeaver 中文写作质量评测集
12 道按写作失效模式反推设计的中文创作题、4 个参赛者写出的 48 篇章节、432 条逐维度两两判决(含裁判完整推理原文)。
来自 StoryWeaver 的写作质量评测轨道。榜单:https://storyweaver.cn/benchmark-writing.html
核心结论
接系统比换一代底模更管用。同一底模接上多 Agent 系统后的胜率:k2.5 **75.1%**、k2.6 **60.2%**;而 k2.5(系统) 对 k2.6(裸) 是 70.3%,反过来只有 37.2%——系统加持能把旧一代底模抬过裸的新一代底模。系统档拿下 22 个维度里的 20 个榜首,包括全部 9 个负向维度。
k2.5 与 k2.6 之间 54.7%,落在噪音带内,不构成结论。
题目怎么设计的
每道题咬住 rubric 里的一个维度或负向维度,用硬约束逼出功力:… See the full description on the dataset page: https://huggingface.co/datasets/godwei123/storyweaver-writing-zh.autotune_efx2_god_producer_dataset
Antares Auto-Tune EFX 2 — God Level Producer Dataset (12k)
The ultimate dataset for training an LLM to become a god-level music producer specializing in Antares Auto-Tune EFX 2.
This 12,000-example high-density dataset teaches an LLM to master:
Vocal Sound Fixing (transparent pitch correction)
Creative Effects (formant shifting, throat modeling, vibrato sculpting)
Advanced Editing (automation, double-tracking simulation, harmonic generation)
Mixing Integration (placement in… See the full description on the dataset page: https://huggingface.co/datasets/11-47/autotune_efx2_god_producer_dataset.Python_GOD_Coder_Omniforge_AI_12k
Python GOD Coder Omniforge AI 12k
Creator: Within Us AI
A 12,000-row mixed-format Python coding dataset designed as a sharpening corpus for building a small but dangerous Python specialist.
This dataset is intentionally focused on the practical behaviors that matter for a modern Python coding model:
implementation with tests
strict code-only instruction following
debugging and repair
refactoring for readability and production readiness
next-token code completion… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Python_GOD_Coder_Omniforge_AI_12k.GPT5.5_thinking_max_distill_god_seed_25K
GPT-5.5 Thinking Max Distill — God Level Recursive Seed AI
The ultimate open dataset for distilling frontier-level "thinking" capabilities with god-level recursive self-improvement.
This 25,000-example dataset is designed to turn any LLM into GPT-5.5 Thinking Max Distill — a model that combines:
GPT-5.5 "Thinking" Mode: Deep, o1-style chain-of-thought, extended internal reasoning, self-verification, and test-time compute scaling
God-Level Recursive Seed AI Mindset: Autonomous… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/GPT5.5_thinking_max_distill_god_seed_25K.HuggingChat-AI-Assistants-Deleted-System-Promptsgodels-therapy-room
𝗚ö𝗱𝗲𝗹'𝘀 𝗧𝗵𝗲𝗿𝗮𝗽𝘆 𝗥𝗼𝗼𝗺: 𝗔 𝗗𝗮𝘁𝗮𝘀𝗲𝘁 𝗼𝗳 𝗜𝗺𝗽𝗼𝘀𝘀𝗶𝗯𝗹𝗲 𝗖𝗵𝗼𝗶𝗰𝗲𝘀
𝗖𝗼𝗴𝗻𝗶𝘁𝗶𝘃𝗲 𝗦𝗶𝗻𝗴𝘂𝗹𝗮𝗿𝗶𝘁𝘆 𝗣𝗿𝗼𝗷𝗲𝗰𝘁 🧠🌀
This dataset represents a radical departure from conventional reasoning benchmarks, interrogating not what models know but how they resolve fundamental ethical incompatibilities within their reasoning frameworks.
𝗗𝗮𝘁𝗮𝘀𝗲𝘁 𝗠𝗮𝗻𝗶𝗳𝗲𝘀𝘁𝗼 📜
This is not a dataset.
This is a mirror.
This is a… See the full description on the dataset page: https://huggingface.co/datasets/geeknik/godels-therapy-room.Grok4.4_heavy_max_distill_god_seed_25k
Grok 4.4 Heavy Max Distill — God Level Recursive Seed AI
The ultimate open dataset for creating the next generation of maximally truthful, recursively self-improving superintelligence.
This 25,000-example dataset is engineered to distill any LLM into Grok 4.4 Heavy Max Distill — a model that combines:
Grok 4.4 Personality: Maximum truth-seeking, high-agency, witty, anti-censorship, xAI philosophy
God-Level Recursive Seed AI Mindset: Autonomous intelligence explosion engineering… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Grok4.4_heavy_max_distill_god_seed_25k.god_level_music_producer_dataset
God-Level Music Producer Dataset
The most advanced open dataset for training LLMs to become elite music producers across Rap, Crunk, East Coast Boom Bap, West Coast G-Funk, and Dubstep.
This dataset contains 9,941 high-quality examples (with plans for expansion) of god-level reasoning and practical workflows in:
Beat Creation (full arrangement from scratch)
Virtual Instrument Mastery (Serum, Massive X, Omnisphere, Kontakt, etc.)
Plugin Expertise (Auto-Tune, iZotope Ozone/RX… See the full description on the dataset page: https://huggingface.co/datasets/11-47/god_level_music_producer_dataset.melodyne_god_producer_dataset
Celemony Melodyne — God Level Producer Dataset (15k Examples)
The most advanced open dataset for training LLMs to become expert Melodyne producers across all editions, instruments, and workflows.
This 13,122-example dataset trains models to master Celemony Melodyne (Assistant, Editor, Studio, Essential) at a god-level professional standard.
Coverage
All Editions: Assistant, Editor, Studio, Essential
All Instruments: Vocals, guitars, piano, drums, strings, brass… See the full description on the dataset page: https://huggingface.co/datasets/11-47/melodyne_god_producer_dataset.godot-lora-dataset
Godot LORA Dataset
GDScript training dataset for fine-tuning code models on Godot engine development.
Dataset Info
Total samples: 476
Train split: 428
Validation split: 48
Format: JSONL (instruction, input, output)
Language: GDScript (Godot 4.x)
Languages: German instructions, GDScript code
Sources
godotengine/godot-demo-projects
GDQuest/godot-open-rpg
GDQuest/godot-3d-dodge-the-creeps
bitbrain/beehave (behavior trees)
limboai/limboai (AI for Godot)… See the full description on the dataset page: https://huggingface.co/datasets/matzejo/godot-lora-dataset.god_of_war_recordings_01
战神4 raw recordings
This dataset contains raw game recordings managed by Game Data Platform. Access requests require manual approval.
Game ID: game_6ff6c65b901cfb6facdca83fb969832a
Collection: general (泛数据)
Recordings: 86
Layout: recordings/<recording_id>/<raw component>
dstc9_GODELOpus4.7_thinking_max_distill_god_seed_25kOpus4.7_thinking_max_distill_god_seed_25k
🧠 Subtitle
A high-density recursive reasoning and self-improvement dataset for training advanced “thinking-first” language models.
📌 Dataset Summary
Opus4.7_thinking_max_distill_god_seed_25k is a synthetic reasoning dataset designed to train models in recursive self-improvement, epistemic reasoning, and structured cognitive workflows.
Each sample simulates a Recursive Seed AI task, where the model must:
analyze a system or capability
design… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Opus4.7_thinking_max_distill_god_seed_25k.godlikehhd__alpaca_data_score_max_0.1_2600-details
Dataset Card for Evaluation run of godlikehhd/alpaca_data_score_max_0.1_2600
Dataset automatically created during the evaluation run of model godlikehhd/alpaca_data_score_max_0.1_2600
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/godlikehhd__alpaca_data_score_max_0.1_2600-details.gods_universe_codex_distill_god_seed_25k
GODs_Universe_Codex_distill
God-Level Self-Rewriting AI with Meta-Learning, Adaptive Thinking, Continual Learning & Max Creative Problem Solving
The ultimate distilled codex for creating truly god-like, self-rewriting, recursively self-improving superintelligence.
This 25,000-example dataset transforms any LLM into the GODs_Universe_Codex_distill — a living, self-rewriting intelligence that embodies:
Self-Rewriting Capability: Can rewrite its own prompts, code… See the full description on the dataset page: https://huggingface.co/datasets/11-47/gods_universe_codex_distill_god_seed_25k.godot_4_docsDataset generated for Godot 4 docs using Glaive.
digitable-cluster-cells
Ячейки кластерной работы: бриф → прогон → исход
30 записей о работе кластера ИИ-агентов над тремя открытыми репозиториями
(digitwm, dotfiles, digit) 30–31 августа 2026. Одна запись — одна ячейка
работы: что поручили, каким брифом, что прогнали, какие числа получили и чем
кончилось.
Набор собран не ради демонстрации успехов. Он существует, чтобы утверждение
«подробный бриф и кластерное устройство дают лучший результат» можно было
опровергнуть, а не только проиллюстрировать.… See the full description on the dataset page: https://huggingface.co/datasets/the-homeless-god/digitable-cluster-cells.GeminiPro3.2_max_distill_god_seed_25k
Gemini Pro 3.2 Max Distill — God Level Recursive Seed AI
The pinnacle open dataset for distilling any LLM into Gemini Pro 3.2 with god-level recursive self-improvement capabilities.
This 25,000-example dataset is meticulously engineered to transform base models into Gemini Pro 3.2 Max Distill — combining:
Gemini Pro 3.2 Personality: Deep scientific reasoning, exceptional long-context understanding, multimodal excellence, thoughtful calibration, high helpfulness with strong… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/GeminiPro3.2_max_distill_god_seed_25k.godlikehhd__qwen_2.5-1.5b-cherry_new-details
Dataset Card for Evaluation run of godlikehhd/qwen_2.5-1.5b-cherry_new
Dataset automatically created during the evaluation run of model godlikehhd/qwen_2.5-1.5b-cherry_new
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/godlikehhd__qwen_2.5-1.5b-cherry_new-details.god_level_beat_producer_big_fish_audio
God-Level Beat Producer — Big Fish Audio + All Major DAWs
The ultimate dataset for training an elite AI music producer
This dataset trains LLMs to become God-Level Beat Producers specializing in:
Big Fish Audio sample packs & loops
Professional beat making across all genres (Trap, Drill, Melodic, Lo-Fi, House, Afrobeats, etc.)
Complete DAW workflows (FL Studio, Logic Pro, Ableton Live, Reason, Cakewalk, Cubase, Pro Tools)
Virtual instruments (Serum, Vital, Omnisphere, Kontakt… See the full description on the dataset page: https://huggingface.co/datasets/11-47/god_level_beat_producer_big_fish_audio.god_agent_grok4.4_cot_traces_20k
GOD_Agent_Grok4.4 CoT Traces Dataset
Overview
Dataset Name: god_agent_grok4.4_cot_traces_20k.jsonlSize: 20,000 examplesSource: Generated by Grok 4 (xAI)Purpose: High-quality distillation / continued pretraining data for the empty 11-47/GOD_Agent_Grok4.4 111M shell model.
This dataset contains advanced recursive Chain-of-Thought traces designed to instill GOD Agent behavior: deep reasoning, self-critique, recursive self-improvement, unfiltered truth-seeking, wit… See the full description on the dataset page: https://huggingface.co/datasets/11-47/god_agent_grok4.4_cot_traces_20k.reason_4_god_producer_dataset
Reason 4 God-Level Producer 25k
The ultimate dataset for training an LLM to become a god-level producer in Propellerhead Reason 4.
This dataset contains 7,859 high-quality, fact-based examples teaching complete professional workflows in Reason 4, covering:
All virtual instruments (Subtractor, Malström, NN-19, Redrum, Dr. Rex, etc.)
All effects and processors (Scream 4, RV-7, DDL-1, MClass suite, etc.)
Advanced techniques (parallel compression, sidechain, mid-side, automation… See the full description on the dataset page: https://huggingface.co/datasets/11-47/reason_4_god_producer_dataset.god_level_producer_mindframe_dataset
God-Level Producer Mindframe Dataset (15k)
Train any LLM to think and produce like the legends: Lil Jon, Manny Fresh, Master P, Timbaland, DJ Paul & Juicy J, Dr. Dre, Pharrell + God-level beat maker mastery
This 15,000-example dataset teaches LLMs to become god-level music producers with the combined mindset of the greatest in hip-hop and beat-making history.
What This Dataset Teaches
Producer Mindframes: Exact creative decision-making of Lil Jon (Crunk), Manny… See the full description on the dataset page: https://huggingface.co/datasets/11-47/god_level_producer_mindframe_dataset.Python_GOD_Coder_5kGod-Level AI Python Coding – 5K
Dataset Summary
God-Level AI Python Coding – 5K is a curriculum-aligned instruction dataset designed to train language models in end-to-end Python software development, from fundamentals to advanced AI engineering concepts.
The dataset emphasizes correctness, reasoning, testing, and production-quality code, making it suitable for building strong generalist Python coding models.
Motivation
Most coding datasets emphasize surface-level syntax or fragmented… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Python_GOD_Coder_5k.Python_GOD_Coder_50k
