datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
git-history-mcq-ru
git-history-mcq-ru
805 вопросов с вариантами ответа по истории трёх открытых репозиториев
(digitable-lol/digit, digitable-lol/digitwm, digitable-lol/flang), плюс
8 672 ответа пяти моделей и 4 878 разборов этих ответов.
Вопросы на русском. Ключ каждого выведен из вывода git-команды, и сама команда
и её вывод лежат в записи — задачу можно перепроверить, не доверяя составителю.
Набор собран для одной проверки: меняют ли что-нибудь приёмы промптинга. Девять
вариантов оформления… See the full description on the dataset page: https://huggingface.co/datasets/the-homeless-god/git-history-mcq-ru.storyweaver-writing-zh
StoryWeaver 中文写作质量评测集
12 道按写作失效模式反推设计的中文创作题、4 个参赛者写出的 48 篇章节、432 条逐维度两两判决(含裁判完整推理原文)。
来自 StoryWeaver 的写作质量评测轨道。榜单:https://storyweaver.cn/benchmark-writing.html
核心结论
接系统比换一代底模更管用。同一底模接上多 Agent 系统后的胜率:k2.5 **75.1%**、k2.6 **60.2%**;而 k2.5(系统) 对 k2.6(裸) 是 70.3%,反过来只有 37.2%——系统加持能把旧一代底模抬过裸的新一代底模。系统档拿下 22 个维度里的 20 个榜首,包括全部 9 个负向维度。
k2.5 与 k2.6 之间 54.7%,落在噪音带内,不构成结论。
题目怎么设计的
每道题咬住 rubric 里的一个维度或负向维度,用硬约束逼出功力:… See the full description on the dataset page: https://huggingface.co/datasets/godwei123/storyweaver-writing-zh.god_of_war_recordings_01
战神4 raw recordings
This dataset contains raw game recordings managed by Game Data Platform. Access requests require manual approval.
Game ID: game_6ff6c65b901cfb6facdca83fb969832a
Collection: general (泛数据)
Recordings: 86
Layout: recordings/<recording_id>/<raw component>
godlikehhd__alpaca_data_score_max_0.1_2600-details
Dataset Card for Evaluation run of godlikehhd/alpaca_data_score_max_0.1_2600
Dataset automatically created during the evaluation run of model godlikehhd/alpaca_data_score_max_0.1_2600
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/godlikehhd__alpaca_data_score_max_0.1_2600-details.digitable-cluster-cells
Ячейки кластерной работы: бриф → прогон → исход
30 записей о работе кластера ИИ-агентов над тремя открытыми репозиториями
(digitwm, dotfiles, digit) 30–31 августа 2026. Одна запись — одна ячейка
работы: что поручили, каким брифом, что прогнали, какие числа получили и чем
кончилось.
Набор собран не ради демонстрации успехов. Он существует, чтобы утверждение
«подробный бриф и кластерное устройство дают лучший результат» можно было
опровергнуть, а не только проиллюстрировать.… See the full description on the dataset page: https://huggingface.co/datasets/the-homeless-god/digitable-cluster-cells.godlikehhd__qwen_2.5-1.5b-cherry_new-details
Dataset Card for Evaluation run of godlikehhd/qwen_2.5-1.5b-cherry_new
Dataset automatically created during the evaluation run of model godlikehhd/qwen_2.5-1.5b-cherry_new
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/godlikehhd__qwen_2.5-1.5b-cherry_new-details.god_level_python_dataset_25k
God-Level Python Coder Dataset (25K Unique Advanced Examples)
Version: 1.0Size: Exactly 25,000 unique entries delivered. 100% synthetic with strong uniqueness guarantees via careful parameterization and deduplication.Focus: Training LLMs to achieve god-level mastery of Python — not just solving problems, but writing idiomatic, performant, robust, elegant, and deeply understood Python code.
This dataset is designed to push LLMs beyond basic LeetCode-style problems into true… See the full description on the dataset page: https://huggingface.co/datasets/11-47/god_level_python_dataset_25k.god_level_python_dataset_v1
God-Level Python Coder Dataset
A high-quality, synthetic dataset for training LLMs to achieve elite ("god-level") Python programming mastery.
Dataset Summary
This dataset contains 2,502 unique, advanced Python coding examples specifically designed to push large language models beyond basic problem-solving into true expert-level Python engineering.
It focuses on the hardest and most important areas of Python:
Deep metaprogramming
Production-grade asyncio &… See the full description on the dataset page: https://huggingface.co/datasets/11-47/god_level_python_dataset_v1.godlikehhd__alpaca_data_ifd_max_2600_3B-details
Dataset Card for Evaluation run of godlikehhd/alpaca_data_ifd_max_2600_3B
Dataset automatically created during the evaluation run of model godlikehhd/alpaca_data_ifd_max_2600_3B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/godlikehhd__alpaca_data_ifd_max_2600_3B-details.godlikehhd__alpaca_data_ifd_max_2600-details
Dataset Card for Evaluation run of godlikehhd/alpaca_data_ifd_max_2600
Dataset automatically created during the evaluation run of model godlikehhd/alpaca_data_ifd_max_2600
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/godlikehhd__alpaca_data_ifd_max_2600-details.BesiegeField_humandataset_coldstart
🏰 BesiegeField Human Dataset for ColdStart
📎 Links
Project Page: https://besiegefield.github.io/
GitHub: https://github.com/Godheritage/BesiegeField
arXiv: https://arxiv.org/abs/2510.14980
📌 Overview
This dataset collects human-made Besiege machines and reformats them for the LLM Cold-Start stage.Processing details are described in Paper §F.1.
🗂️ Data Fields
trainable_dataset: Dataset root.
ID: Unique id for each data.
xxx.json: The tree… See the full description on the dataset page: https://huggingface.co/datasets/Godheritage/BesiegeField_humandataset_coldstart.godseed-worldgodlikehhd__alpaca_data_sampled_ifd_5200-details
Dataset Card for Evaluation run of godlikehhd/alpaca_data_sampled_ifd_5200
Dataset automatically created during the evaluation run of model godlikehhd/alpaca_data_sampled_ifd_5200
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/godlikehhd__alpaca_data_sampled_ifd_5200-details.godlikehhd__alpaca_data_ins_max_5200-details
Dataset Card for Evaluation run of godlikehhd/alpaca_data_ins_max_5200
Dataset automatically created during the evaluation run of model godlikehhd/alpaca_data_ins_max_5200
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/godlikehhd__alpaca_data_ins_max_5200-details.godlikehhd__alpaca_data_ifd_min_2600-details
Dataset Card for Evaluation run of godlikehhd/alpaca_data_ifd_min_2600
Dataset automatically created during the evaluation run of model godlikehhd/alpaca_data_ifd_min_2600
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/godlikehhd__alpaca_data_ifd_min_2600-details.godlikehhd__qwen_ins_ans_2500-details
Dataset Card for Evaluation run of godlikehhd/qwen_ins_ans_2500
Dataset automatically created during the evaluation run of model godlikehhd/qwen_ins_ans_2500
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/godlikehhd__qwen_ins_ans_2500-details.godlikehhd__qwen-2.5-1.5b-cherry-details
Dataset Card for Evaluation run of godlikehhd/qwen-2.5-1.5b-cherry
Dataset automatically created during the evaluation run of model godlikehhd/qwen-2.5-1.5b-cherry
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/godlikehhd__qwen-2.5-1.5b-cherry-details.godlikehhd__alpaca_data_full_2-details
Dataset Card for Evaluation run of godlikehhd/alpaca_data_full_2
Dataset automatically created during the evaluation run of model godlikehhd/alpaca_data_full_2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/godlikehhd__alpaca_data_full_2-details.godlikehhd__alpaca_data_ins_min_2600-details
Dataset Card for Evaluation run of godlikehhd/alpaca_data_ins_min_2600
Dataset automatically created during the evaluation run of model godlikehhd/alpaca_data_ins_min_2600
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/godlikehhd__alpaca_data_ins_min_2600-details.godlikehhd__alpaca_data_score_max_2600_3B-details
Dataset Card for Evaluation run of godlikehhd/alpaca_data_score_max_2600_3B
Dataset automatically created during the evaluation run of model godlikehhd/alpaca_data_score_max_2600_3B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/godlikehhd__alpaca_data_score_max_2600_3B-details.godlikehhd__alpaca_data_full_3B-details
Dataset Card for Evaluation run of godlikehhd/alpaca_data_full_3B
Dataset automatically created during the evaluation run of model godlikehhd/alpaca_data_full_3B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/godlikehhd__alpaca_data_full_3B-details.godlikehhd__alpaca_data_sampled_ifd_new_5200-details
Dataset Card for Evaluation run of godlikehhd/alpaca_data_sampled_ifd_new_5200
Dataset automatically created during the evaluation run of model godlikehhd/alpaca_data_sampled_ifd_new_5200
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/godlikehhd__alpaca_data_sampled_ifd_new_5200-details.godlikehhd__alpaca_data_score_max_0.7_2600-details
Dataset Card for Evaluation run of godlikehhd/alpaca_data_score_max_0.7_2600
Dataset automatically created during the evaluation run of model godlikehhd/alpaca_data_score_max_0.7_2600
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/godlikehhd__alpaca_data_score_max_0.7_2600-details.godlikehhd__qwen_full_data_alpaca-details
Dataset Card for Evaluation run of godlikehhd/qwen_full_data_alpaca
Dataset automatically created during the evaluation run of model godlikehhd/qwen_full_data_alpaca
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/godlikehhd__qwen_full_data_alpaca-details.godlikehhd__ifd_new_qwen_2500-details
Dataset Card for Evaluation run of godlikehhd/ifd_new_qwen_2500
Dataset automatically created during the evaluation run of model godlikehhd/ifd_new_qwen_2500
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/godlikehhd__ifd_new_qwen_2500-details.godlikehhd__alpaca_data_ins_min_5200-details
Dataset Card for Evaluation run of godlikehhd/alpaca_data_ins_min_5200
Dataset automatically created during the evaluation run of model godlikehhd/alpaca_data_ins_min_5200
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/godlikehhd__alpaca_data_ins_min_5200-details.godlikehhd__alpaca_data_score_max_0.3_2600-details
Dataset Card for Evaluation run of godlikehhd/alpaca_data_score_max_0.3_2600
Dataset automatically created during the evaluation run of model godlikehhd/alpaca_data_score_max_0.3_2600
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/godlikehhd__alpaca_data_score_max_0.3_2600-details.godlikehhd__ifd_2500_qwen-details
Dataset Card for Evaluation run of godlikehhd/ifd_2500_qwen
Dataset automatically created during the evaluation run of model godlikehhd/ifd_2500_qwen
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/godlikehhd__ifd_2500_qwen-details.godlikehhd__ifd_new_correct_sample_2500_qwen-details
Dataset Card for Evaluation run of godlikehhd/ifd_new_correct_sample_2500_qwen
Dataset automatically created during the evaluation run of model godlikehhd/ifd_new_correct_sample_2500_qwen
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/godlikehhd__ifd_new_correct_sample_2500_qwen-details.godlikehhd__ifd_new_correct_all_sample_2500_qwen-details
Dataset Card for Evaluation run of godlikehhd/ifd_new_correct_all_sample_2500_qwen
Dataset automatically created during the evaluation run of model godlikehhd/ifd_new_correct_all_sample_2500_qwen
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/godlikehhd__ifd_new_correct_all_sample_2500_qwen-details.
